Chapter 02

Attention, per decode step

What a new token needs from the past.

tokensThe␣capital␣of␣France␣is

Same five tokens, one forward pass. This time we watch what one attention layer does with them.

01 / 14
The math, once

For one head, with Q=XWQQ = XW_Q, K=XWKK = XW_K, V=XWVV = XW_V:

Attention(Q,K,V)=softmax ⁣(QKdk)V\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V

Row ii of the result only uses row ii of QQ, but every row i\le i of KK and VV. That asymmetry is the whole next chapter.