Course Module

Module 1: Mathematical Foundations of Attention & Sequences

1.1 The Vanishing Gradient Problem & The Need for Attention

An exploration of sequence modeling bottlenecks, memory bottlenecks in LSTMs, and the intuition behind associative memory tables.

45 mins →

1.2 Scaled Dot-Product Attention & Multi-Head Projections

Rigorous derivation of queries, keys, and values matrices, softmax normalization, and head concatenation projections.

55 mins →