Mathematical Foundations

The Mathematics of Transformers: Attention, Projections & Softmax Geometry

A rigorous mathematical deconstruction of Scaled Dot-Product Attention, Query-Key-Value projection spaces, and temperature scaling proofs.

Mar 20249 min read · Mathematical Derivation

The Transformer architecture eliminated recurrence and convolutions in favor of pure self-attention mechanisms. While intuitive conceptually, its mathematical formulation reveals deep geometric properties.


1. Scaled Dot-Product Attention

Given an input sequence represented as a matrix X∈Rn×dX \in \mathbb{R}^{n \times d}, we project XX into three distinct vector spaces using learned weight matrices: Q=XWQ,K=XWK,V=XWVQ = X W_Q, \quad K = X W_K, \quad V = X W_V where WQ,WK∈Rd×dkW_Q, W_K \in \mathbb{R}^{d \times d_k} and WV∈Rd×dvW_V \in \mathbb{R}^{d \times d_v}.

The attention formula is expressed as: Attention(Q,K,V)=softmax(QKTdk)V\text{Attention}(Q, K, V) = \text{softmax}\left( \frac{Q K^T}{\sqrt{d_k}} \right) V


2. Why Scale by 1dk\frac{1}{\sqrt{d_k}}?

Consider two random vectors q,k∈Rdkq, k \in \mathbb{R}^{d_k} whose components are independent random variables with mean 00 and variance 11: qi,ki∼N(0,1)q_i, k_i \sim \mathcal{N}(0, 1)

Calculating the expectation and variance of their dot product: E[S]=0,Var(S)=dk\mathbb{E}[S] = 0, \quad \text{Var}(S) = d_k

As dkd_k grows large (e.g., dk=128d_k = 128), the variance of the dot product scales linearly with dkd_k. Extremely large positive or negative values push the softmax\text{softmax} function into regions with near-zero gradients (vanishing gradients). Dividing by dk\sqrt{d_k} normalizes the variance back to 11.