Kaizoi
0
All stories
Self-attentionAI and ML8 min

The word that needed its neighbours

You have done this today

You ask a chatbot a question with a long sentence, and it understands which word refers to which.

What happens behind the screen

Read this: 'The driver parked the truck because it was heavy.' What does 'it' mean? You looked back at 'truck' without thinking. Older models read words in order and often lost that thread by the time they reached 'it'.

Step through how attention lets a word look at the others.

How the word 'it' finds its meaning

Take the word 'it'. On its own it means nothing.

Step 1 of 4

The idea in plain words

Attention gives every word a way to look at every other word and decide how much each matters. The word's meaning is rebuilt as a weighted blend of the words it attended to.

Each word has three vectors: a query (what am I looking for), a key (what do I offer) and a value (what I pass on). Compare the query with every key to get scores, turn the scores into weights that add up to 1, and blend the values. The cost is that every word looks at every other, so the work grows with the square of the text length, which is why long documents are expensive.

Scaled dot-product attention

import numpy as np

def attention(Q, K, V):
    d = Q.shape[-1]
    scores = Q @ K.T / np.sqrt(d)                 # how well each query matches each key
    scores -= scores.max(axis=-1, keepdims=True)  # numerical stability
    weights = np.exp(scores)
    weights /= weights.sum(axis=-1, keepdims=True)  # each row adds up to 1
    return weights @ V                            # blend values by those weights

If an interviewer asks

"Explain self-attention in simple terms. Why is it expensive for long inputs?"

You could say

Self-attention lets each token compare itself with every other token and build its representation as a weighted mix of them, using queries, keys and values. Because every token attends to every other, the cost grows quadratically with sequence length, which is why long contexts are costly.

Check yourself

In attention, what do the weights add up to?

Why is attention costly for very long text?

Was this clear?

Send on WhatsApp