Back to Module 1: LLM FoundationIn Progress
AI Syllabus Module
Module 1.9: Self-Attention
Deconstruct dot-product attention steps, QKV matrices, and context calculations mathematically.
Lessons & Submodules
Why Self-Attention Was Needed
Compare recurrent constraints against parallel attention calculations.
Query, Key and Value
Deconstruct the role of projection matrices in generating Q, K, and V.
Scaled Dot Product Attention
Derive the mathematical formula of scaled dot-product attention.
Attention Weights
Visualize token attention relationships using heatmaps.
Causal Masking
Implement causal mask filters to enforce autoregressive generation constraints.
Multi-head Attention
Split query, key, and value vectors across multiple parallel heads.
Attention Interview Guide
Prepare for mathematical questions focused on attention layers.
Key Skills
- •Compute Query, Key, and Value vectors from raw inputs
- •Map scaled dot-product attention matrices mathematically
Interview Value
- Explain how causal masking prevents models from looking at future token values during autoregressive generation.