Attention
Each word decides which other words matter
When you read a sentence, every word quietly asks which other words it should pay attention to. Attention is that decision done with arithmetic. The sentence here is The cat chased the mouse. Each word has three small arrows: a query (the question this word is asking), a key (the label it wears so other words can find it) and a value (the information it hands over when it is read). Pick a word to let it ask its question. Every word’s key is scored by how well it matches that question, the scores become percentages that add up to 100%, and the answer is a blend of the words’ values using those percentages. Drag a key to see the percentages change. The arrows are placed by hand so that chased is asking “who did it, and to what?” and so favours cat and mouse. Try the experiments in the right-hand panel.Attention Explorer: One Query Row
- q·khow well they match
- ÷ √dkeep numbers modest
- ÷ Thow decisive
- maskhide later words
- softmaxpercentages, total 1
- Σ aiviblend the values
The gold arrow is the question the selected word is asking. Each dot is a word’s key (the asking word’s own key is marked “own key”). The more a key points the same way as the arrow, the better the match. The dotted line shows how far each key reaches along the arrow, and that reach is its score. Drag a dot to change what a word advertises; the arrow stays put. Hover any word, bar or cell to follow it across every panel. Keyboard: click the plane, press 1–6 to pick a key, then use the arrow keys to move it (Shift for bigger steps). Exact numbers are in the table on the right.
Each dot is the information a word would hand over. The rose diamond is the blend: a percentage of each value, added together. Because the percentages are never negative and add up to 100%, the blend always lands inside the cloud of dots, leaning toward the words that matched the question best.
Show the arithmetic for the best match
Temperature sets how decisive the choice is. T = 1 is the standard formula. Lower it and the best match takes nearly all the weight. Raise it and the weights even out until every word counts about the same.
Switch this on to hide every word that comes after the asking word, the way a model writing a sentence cannot see words it has not written yet. Their dots stay on the plane, but their weight is exactly zero.
Key coordinates (type exact values)
This tool computes one row of scaled dot-product attention on a six-token sequence whose vectors are two-dimensional, so the same numbers that enter the formula can be drawn as arrows. The sections below state the formula exactly, say what temperature and the causal mask change, and name what this picture is not.
1. The formula
For one query vector $q$ and a matrix $K$ whose rows are the key vectors, the scores are the scaled dot products $qK^{\mathsf T}/\sqrt{d}$. Softmax turns that row into weights $a$, and the output is the same weights applied to the value matrix $V$:
$$\mathrm{Attention}(q,K,V)=\mathrm{softmax}\!\left(\frac{qK^{\mathsf T}}{\sqrt{d}}\right)V$$Here $d=2$, so $\sqrt{d}\approx 1.41$. The lab always divides by it. With the temperature slider at $T=1$, the number on each heatmap cell is exactly one entry of $qK^{\mathsf T}/\sqrt{d}$, and the bars are exactly the softmax of that row.
2. What $q$, $k$, and $v$ are in this demo
Each token is handed three 2-D vectors: a query, a key, and a value. They are placed by hand so that chased starts aligned with cat and mouse, and the two determiners sit near each other. Dragging a key changes only that token’s $k$. The gold query arrow does not follow the dot.
In a real attention layer those three vectors are not stored independently. They are three projections of the same input, $q=xW_Q$, $k=xW_K$, $v=xW_V$. Different matrices are why a token’s query can point somewhere its key does not. This lab skips $W_Q$, $W_K$, and $W_V$ so the dot product stays a picture. Nothing here is learned.
3. Why divide by $\sqrt{d}$
A dot product of two $d$-dimensional vectors grows with $d$ if the coordinates stay order-1. Softmax then saturates: one weight goes to 1 and the rest go to 0, and the gradient through softmax vanishes. Dividing by $\sqrt{d}$ is the fix from Attention Is All You Need, written for $d$ of 64 or 128. At $d=2$ the factor is mild. It is kept so the formula on the canvas is the real one, not a cousin that drops the scale because the picture is small.
4. Softmax is a mixture
Softmax exponentiates the scores and divides by their sum. Every weight is positive, and the weights sum to 1, so the output $\sum_i a_i v_i$ is a convex combination of the value vectors. On the value plane the rose marker therefore sits inside the cloud of dots, closer to the values whose keys point along $q$. It is not “the” value of the winning token unless one weight is nearly 1.
5. Temperature
Temperature is not part of the 2017 formula. The slider divides the scaled scores by $T$ before softmax. At $T=1$ nothing extra happens. Lower $T$ and the largest score pulls away from the others, toward a hard lookup. Raise $T$ and the row flattens toward $1/6$ each (or $1/n$ for however many tokens the mask still allows). The heatmap stays on $q\cdot k/\sqrt{d}$ either way — temperature changes the bars and the mixture, not the geometry.
6. The causal mask
With the mask on, every key after the query has its score replaced by $-\infty$ before softmax, which forces that weight to exactly 0. The dot stays on the key plane: the mask does not move vectors, it refuses to read them. Turn it on with chased selected and mouse drops out, even though its key still points along the query. That is the decoder rule. A position may mix the past and itself, not the future. Encoder self-attention leaves the mask off and lets every token see every other token.
7. What this is not
One head, six tokens, vectors in $\mathbb{R}^2$. There is no multi-head split, no residual add, no layer norm, no feed-forward block, and no positional encoding. The Matrices page writes $\text{Attention Scores}=QK^{\mathsf T}$ and $\text{Attention Output}=AV$ for every query at once. This lab is a single row of that product — the row for the token you clicked. The dot product itself is the same operation the Latent Space lab uses for nearness. Attention is that nearness, turned into a weighted read of the values.
Check yourself
Predict the answer, try it in the Explorer if you can, then open the answer.
If every key were identical, what would the weights and the answer be?
Every word would match the question equally well, so each gets the same share: $1/6$ (or $1/n$ if some words are hidden). The answer is then just the plain average of the values, whatever the question is.
Why can the rose diamond never leave the cloud of value dots?
The blend uses percentages that are never negative and add up to 100%. A mix like that can only land somewhere between the values, never outside them. It sits exactly on one value only if that word gets 100%.
Doubling every key vector at $T=1$ has the same effect as what temperature setting?
$T=0.5$. Doubling the keys doubles every match score, and dividing by $0.5$ also doubles them. So a more decisive choice can come from longer keys or from a lower temperature.
With the causal mask on and the first token selected, what is the output?
The first word can see only itself, so it gets 100% of the weight and the answer is just its own value.
Why is a weight never exactly 0 unless the mask is on?
The softmax step turns every score into a positive number, so every word keeps at least a sliver of weight. A tiny weight can look like 0.00 on screen, but only hiding the word with the mask makes it truly zero.
Why does the short query from period give a nearly flat row?
A short query is a faint question: it hardly prefers any key, so every score is close to 0 and every word gets about the same share. The length of the query matters as well as its direction. A long query is a firm question and acts like a lower temperature.
