Mathematical Foundations
The Mathematics of Transformers: Attention, Projections & Softmax Geometry
A rigorous mathematical deconstruction of Scaled Dot-Product Attention, Query-Key-Value projection spaces, and temperature scaling proofs.
The Transformer architecture eliminated recurrence and convolutions in favor of pure self-attention mechanisms. While intuitive conceptually, its mathematical formulation reveals deep geometric properties.
1. Scaled Dot-Product Attention
Given an input sequence represented as a matrix , we project into three distinct vector spaces using learned weight matrices: where and .
The attention formula is expressed as:
2. Why Scale by ?
Consider two random vectors whose components are independent random variables with mean and variance :
Calculating the expectation and variance of their dot product:
As grows large (e.g., ), the variance of the dot product scales linearly with . Extremely large positive or negative values push the function into regions with near-zero gradients (vanishing gradients). Dividing by normalizes the variance back to .