Chapter 02
Attention, per decode step
What a new token needs from the past.
Same five tokens, one forward pass. This time we watch what one attention layer does with them.
01 / 14
The math, once
For one head, with , , :
Row of the result only uses row of , but every row of and . That asymmetry is the whole next chapter.