How an app knows what you will like next
You have done this today
You finish a film and the app offers three more that you actually want to watch.
What happens behind the screen
Imagine describing every film with a few numbers: how funny, how tense, how romantic. Each film becomes a point on a map, and films that feel alike sit close together.
Step through how a recommendation is found.
Every film gets a list of numbers, such as [funny 0.9, tense 0.1, romantic 0.3]. That list is a vector.
The idea in plain words
A vector is a list of numbers that describes something. 'Close' is usually measured by the angle between two arrows, called cosine similarity, because direction captures taste better than length. Same direction scores 1, perpendicular scores 0.
Real systems use hundreds of dimensions learned by a model and fast indexes to find neighbours. The same idea powers retrieval for language models: turn the question and your documents into vectors, and fetch the nearest documents.
Cosine similarity from scratch
import math
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
len_a = math.sqrt(sum(x * x for x in a))
len_b = math.sqrt(sum(y * y for y in b))
return dot / (len_a * len_b)
# funny, tense, romantic
film_a = [0.9, 0.1, 0.3]
film_b = [0.8, 0.2, 0.4]
print(cosine(film_a, film_b)) # close to 1 means similar tasteIf an interviewer asks
"What is an embedding, and how would you use it to build a recommendation feature?"
You could say
An embedding turns an item into a vector so that similar items sit close together. For recommendations I would embed the items, take the one the user just engaged with, and return its nearest neighbours by cosine similarity, using an approximate nearest-neighbour index to keep it fast at scale.
Check yourself
Two vectors point in exactly the same direction. What is their cosine similarity?
Why use direction instead of length to compare tastes?
Was this clear?
Up next
The word that needed its neighbours
Attention, the idea inside every large language model.