Module 1: Mathematical Foundations of Attention & Sequences
1.1 The Vanishing Gradient Problem & The Need for Attention
An exploration of sequence modeling bottlenecks, memory bottlenecks in LSTMs, and the intuition behind associative memory tables.
1.2 Scaled Dot-Product Attention & Multi-Head Projections
Rigorous derivation of queries, keys, and values matrices, softmax normalization, and head concatenation projections.