Kaizoi
0
All stories
Chunking and shardingData engineering7 min

Why a film is cut into small pieces

You have done this today

You press play on a long film on your phone and it starts almost immediately.

What happens behind the screen

The service did not send you the whole file. It had already cut the film into short pieces, a few seconds each, and it sends the first piece while you watch.

Cutting also lets the work be shared across many machines. Step through the journey of one film.

From one huge file to instant playback

A two-hour film arrives as one huge file.

Step 1 of 5

The idea in plain words

Chunking means splitting big data into small pieces. Sharding means spreading those pieces over machines using a key. The art is choosing the key: sharding by film alone sends a blockbuster to one machine, while film plus chunk number spreads it out.

Each step is safe to retry, which is what makes the pipeline reliable.

Choosing a shard for a chunk

import hashlib

SHARDS = 8

def shard_for(film_id: str, chunk: int) -> int:
    # Hash film + chunk so one popular film spreads across all shards.
    key = f"{film_id}:{chunk}".encode()
    digest = hashlib.md5(key).hexdigest()
    return int(digest, 16) % SHARDS

If an interviewer asks

"Design the storage for a video platform. How do you avoid one machine becoming a bottleneck?"

You could say

I would split each video into small chunks, encode them in parallel, and store them across shards. I would pick a shard key that spreads a popular video, like video id plus chunk number, and make every step idempotent so a failed chunk can simply be retried.

Check yourself

Why shard by film plus chunk number instead of film alone?

If one chunk fails to encode, what is retried?

Was this clear?

Send on WhatsApp

Up next

The report that was wrong every Monday

Why a data job must be safe to run twice.