The word that needed its neighbours
You have done this today
You ask a chatbot a question with a long sentence, and it understands which word refers to which.
What happens behind the screen
Read this: 'The driver parked the truck because it was heavy.' What does 'it' mean? You looked back at 'truck' without thinking. Older models read words in order and often lost that thread by the time they reached 'it'.
Step through how attention lets a word look at the others.
Take the word 'it'. On its own it means nothing.
The idea in plain words
Attention gives every word a way to look at every other word and decide how much each matters. The word's meaning is rebuilt as a weighted blend of the words it attended to.
Each word has three vectors: a query (what am I looking for), a key (what do I offer) and a value (what I pass on). Compare the query with every key to get scores, turn the scores into weights that add up to 1, and blend the values. The cost is that every word looks at every other, so the work grows with the square of the text length, which is why long documents are expensive.
Scaled dot-product attention
import numpy as np
def attention(Q, K, V):
d = Q.shape[-1]
scores = Q @ K.T / np.sqrt(d) # how well each query matches each key
scores -= scores.max(axis=-1, keepdims=True) # numerical stability
weights = np.exp(scores)
weights /= weights.sum(axis=-1, keepdims=True) # each row adds up to 1
return weights @ V # blend values by those weightsIf an interviewer asks
"Explain self-attention in simple terms. Why is it expensive for long inputs?"
You could say
Self-attention lets each token compare itself with every other token and build its representation as a weighted mix of them, using queries, keys and values. Because every token attends to every other, the cost grows quadratically with sequence length, which is why long contexts are costly.
Check yourself
In attention, what do the weights add up to?
Why is attention costly for very long text?
Was this clear?