The Interview Edge Blog
← Back to all guides
ML · Data pipelines

The Training-Data Pipeline: Where an LLM’s Knowledge Actually Comes From

Nobody hands a model its knowledge. Engineers manufacture it: crawl the web at petabyte scale, throw away ninety percent, kill the duplicates, scrub the dangerous bits, chop text into tokens, and stream it to the GPUs in exactly the right mix. Here is the pipeline interviewers now expect you to design.

Explain it like I’m five

Imagine you’re building a library for a reader who reads a million books a day — and believes everything she reads.

You can’t just hand her the internet. So you hire a ruthless librarian. First she photocopies the entire web — every page, every forum rant, every ad. That’s the easy part. Then the real work starts: she throws ninety percent of it in the bin. Junk mail, spam, the same article copied a thousand times, pages that are mostly punctuation. What survives gets a second pass — she blacks out every phone number and email address with a marker, and pulls anything hateful or explicit off the shelf. Then she takes scissors to the remaining pages and cuts them into small word-chunks, because your reader doesn’t read words, she reads chunks. Finally she shuffles the whole library, so the reader never gets ten cookbooks in a row and starts believing the world is only recipes. Only then does the reading begin.

That librarian is the training-data pipeline. Crawling is the photocopying; filtering, deduplication, scrubbing, tokenization, and shuffling are everything after. And here’s the uncomfortable truth the interview is really about: the model’s intelligence is capped by the librarian’s taste. Garbage in, garbage out — at a trillion tokens, the garbage is just better hidden.

Intuition: the interview question wearing an ML costume

“How would you build the data pipeline for training an LLM?” is the new “design a load balancer” — a systems question in disguise. The candidates who pass talk about data for twenty minutes before they mention the model.

Look at the numbers. GPT-3 (2020) trained on about 300 billion tokens. Llama 3 (2024) trained on over 15 trillion — fifty times more data in four years. The architecture barely changed; the data did. The Chinchilla results taught the field that models were systematically undertrained: data, not parameters, is the scarce input. Whoever manufactures the best trillion tokens wins.

And the pipeline is where the money burns. Every junk document that survives filtering wastes GPU time — and GPU time is the most expensive line item in AI. A bad pipeline trains on spam; a good one is a moat. Notice what frontier labs publish (model weights, sometimes) versus what they never publish (their filtering stack). That silence tells you where the value is.

One sentence worth memorizing

The pipeline turns the raw web into training tokens — acquire at petabyte scale, keep the best tenth, kill duplicates, scrub PII and toxicity, tokenize once, and stream a shuffled, weighted mix to the GPUs.

How it works: six stages, one ruthless librarian

1. Acquire: photocopy the internet

It starts with a crawl. Common Crawl is the open one — a nonprofit that has crawled the web every month since 2007, petabytes of pages, free for anyone to download. Frontier labs run their own crawlers too, ingesting terabytes a day. Acquisition is the easy, solved part: bandwidth and storage.

On top of the crawl sit curated corpora — Wikipedia, books, code, papers. Small in volume, enormous in signal. The crawl gives you breadth; the curated sources give you quality. The mixture between them is a design decision, not an accident.

2. Filter: keep the best tenth

Raw web text is mostly junk — ads, SEO spam, boilerplate, machine-generated sludge. Filtering runs in layers. First language identification: keep the languages you want, drop the rest. Then heuristic filters: documents that are too short, too repetitive, or heavy on symbols fail fast — the cheap tests catch the obvious garbage.

Then the serious layer: model-based quality classifiers. Train a classifier on known-good text (Wikipedia, books, edited articles) versus raw crawl, and keep only documents that score like the good stuff. Typical yield: on the order of ten percent of the crawl survives. Nine-tenths of the internet is, by the pipeline’s standards, not worth training on.

Filtering, in one line

Heuristics are cheap and catch the obvious junk; a classifier trained on good text catches the subtle junk. Together they decide what the model is allowed to read.

3. Deduplicate: kill the copies

The web is the same article copied a thousand times. Exact duplicates die by hashing — normalize the text, SHA-256 it, drop repeats. The hard part is near-duplicates: the same article slightly reworded across a thousand sites.

That’s MinHash (Broder, 1997): fingerprint each document’s word shingles, compare fingerprints instead of full documents, and bucket similar ones with locality-sensitive hashing. Billions of comparisons become feasible because you never compare the documents themselves.

Why it matters beyond hygiene: duplicates don’t just waste compute — they teach memorization. A model that has seen the same paragraph fifty times will recite it verbatim. Deduplication is a safety feature wearing a janitor’s uniform.

4. Scrub: remove the dangerous bits

Two things must not reach the GPUs. PII — emails, phone numbers, ID numbers — found by patterns and classifiers, then redacted or dropped with the whole document. Miss this and the model leaks real people’s phone numbers; it has happened. And toxicity — hate speech, explicit content, graphic violence — filtered hard by classifiers. The scrub is never perfect, which is why it’s defense in depth: filter at the pipeline, then again with safety training later.

5. Tokenize: cut text into numbers

Models don’t read words or characters — they read tokens, chunks learned by byte-pair encoding: frequent character pairs merged until the vocabulary covers the language. Modern models use roughly 100,000 tokens. The discipline that matters: tokenize once. Convert the whole corpus to token IDs a single time and store them — re-tokenizing petabytes is a bill nobody pays twice. Tokens, not words, are the currency: context windows, training budgets, and API prices are all measured in them.

6. Mix and stream: the feeding schedule

The last stage decides what the model eats, in what order. Domain mixture sets the weights — more code for a code model, more papers for a research model. Global shuffle destroys ordering bias: the model must never read ten cookbooks in a row and conclude the world is recipes. Documents are packed into fixed-length sequences, sharded across storage, and streamed during training so thousand-GPU clusters never sit idle waiting for data.

Six pipeline stages from raw web crawl to GPU training, with ninety percent discarded at the filter stepcrawlpetabytesfilterkeep ~10%dedupeMinHashscrubPII + toxicitytokenizeBPE · onceGPUsshuffled mix90% discarded

The pipeline, drawn: the crawl is the widest stage and the filter the narrowest — nine-tenths of the web never makes it past step two.

Three pipelines, one playbook

Every training-data pipeline you will ever meet is the same six stages wearing different uniforms. Here are the uniforms that matter in interviews.

01 · The open crawl

Common Crawl is the raw material the open world builds on: a nonprofit crawl of the web, petabytes deep, refreshed monthly since 2007, free to download. Almost every open dataset — and quietly, many closed ones — starts here. It solves acquisition; everything after it is your pipeline’s job.

02 · The open dataset

The Pile (EleutherAI) — 825 GiB across 22 sources, the dataset behind GPT-J and GPT-NeoX — and Dolma (AI2) — 3 trillion tokens with the pipeline code itself open-sourced — show what a serious filtering and dedupe stack looks like in public. They’re the closest thing to reading a frontier lab’s playbook.

03 · The frontier lab

Llama 3 trained on over 15 trillion tokens — a data scale that only works because the filtering stack is industrial. Meta publishes the token count, not the classifier. That asymmetry is the tell: the crawl is commoditized, the filter is the moat.

Source: Common Crawl — petabyte-scale open web crawl, monthly since 2007 ↗
Source: The Pile (EleutherAI) — 825 GiB open pretraining dataset ↗
Source: Dolma (AI2) — 3T-token open corpus with open pipeline tooling ↗

The uniforms change — open crawl, open dataset, frontier lab — but the playbook doesn’t: acquire wide, filter ruthlessly, dedupe, scrub, tokenize once, stream a shuffled mix.

The scale, worked by hand

Four numbers that reframe the whole topic — the kind of arithmetic worth doing out loud in the interview.

300B → 15TGPT-3 (2020) trained on ~300 billion tokens; Llama 3 (2024) on over 15 trillion. Fifty times more data in four years — the pipeline, not the architecture, absorbed that growth.
~10%The share of a raw web crawl that typically survives quality filtering. Nine-tenths of the internet is, by the pipeline’s standards, not worth training on.
825 GiBThe Pile’s size — the largest open dataset of its era fit on a single hard drive. Everything since came from better pipelines, not just bigger disks.
~100kTokens in a modern BPE vocabulary. Every token the model ever reads or writes is one of these hundred thousand chunks — tokens, not words, are the currency.

Code: the filter and the fingerprint

The two pieces interviewers ask you to whiteboard: a heuristic quality filter and a MinHash near-duplicate check. Simplified, but the shape is real.

Python · pipeline sketch
import re, hashlib

def quality_score(text: str) -> float:
    """Heuristic junk filter: 1.0 = keep, 0.0 = discard."""
    words = text.split()
    if len(words) < 50:            # too short to learn from
        return 0.0
    if len(set(words)) / len(words) < 0.3:  # repetitive boilerplate
        return 0.0
    punct = sum(1 for c in text if not c.isalnum() and not c.isspace())
    if punct / max(len(text), 1) > 0.3:  # punctuation soup
        return 0.0
    return 1.0

def shingles(text: str, k: int = 5):
    words = re.sub(r"\W+", " ", text.lower()).split()
    return {" ".join(words[i:i+k]) for i in range(len(words) - k + 1)}

def minhash_signature(shingle_set, num_hashes: int = 64):
    """Simplified MinHash: minimum hash value per hash function."""
    sig = []
    for seed in range(num_hashes):
        h = min(
            int(hashlib.md5(f"{seed}:{s}".encode()).hexdigest(), 16)
            for s in shingle_set) if shingle_set else 0
        sig.append(h)
    return sig

def jaccard_estimate(sig_a, sig_b) -> float:
    """Fraction of matching signature slots ≈ Jaccard similarity."""
    return sum(1 for a, b in zip(sig_a, sig_b) if a == b) / len(sig_a)

# pipeline rule: keep a doc only if quality_score == 1.0
# and its signature is dissimilar (< 0.8) to everything kept so far.
# Production systems add LSH banding so billions of docs stay feasible.

Interview prompts

How the question actually sounds in the room — and what the interviewer is listening for.

Meta
“Design the data pipeline for training an LLM.”

Walk the six stages in order, but spend most of your time on filtering and deduplication — that’s where judgment lives. Name the yield (~10% survives) and the failure mode of skipping each stage.

Google
“How do you find near-duplicate documents across billions of pages?”

MinHash plus locality-sensitive hashing: fingerprint shingle sets, compare fingerprints — never the documents. Say why exact hashing alone isn’t enough.

OpenAI
“Why shuffle the training data?”

Ordering bias. The model must not learn that cookbooks arrive in clumps; global shuffle plus a deliberate domain mixture is part of the design, not housekeeping.

Anthropic
“How do you keep PII out of the training set?”

Detection (patterns plus classifiers), then redact or drop the document — and admit it’s never perfect, so it’s defense in depth with safety training downstream.

Meta
“When does synthetic data help, and when does it hurt?”

Gold for code and math, where correctness is verifiable; risky at scale — too much and quality drifts, a copy of a copy. Used sparingly, with provenance logged.

Takeaways

  1. Crawl wide, filter ruthlessly. Acquisition is the solved part; the pipeline’s value is what it throws away.

  2. Deduplication is a safety feature, not just hygiene. Duplicates waste compute and teach verbatim memorization.

  3. Tokenize once. BPE turns text into the model’s currency — never re-pay the cost of tokenizing petabytes.

  4. Shuffle and mix deliberately. Ordering bias is real; the feeding schedule is part of the design.

  5. The filter is the moat. Everyone can download Common Crawl. The quality classifier is what frontier labs don’t publish.

Sources

Back toAll guides →