The Training-Data Pipeline: Where an LLM’s Knowledge Actually Comes From
Nobody hands a model its knowledge. Engineers manufacture it: crawl the web at petabyte scale, throw away ninety percent, kill the duplicates, scrub the dangerous bits, chop text into tokens, and stream it to the GPUs in exactly the right mix. Here is the pipeline interviewers now expect you to design.
Explain it like I’m five
Imagine you’re building a library for a reader who reads a million books a day — and believes everything she reads.
You can’t just hand her the internet. So you hire a ruthless librarian. First she photocopies the entire web — every page, every forum rant, every ad. That’s the easy part. Then the real work starts: she throws ninety percent of it in the bin. Junk mail, spam, the same article copied a thousand times, pages that are mostly punctuation. What survives gets a second pass — she blacks out every phone number and email address with a marker, and pulls anything hateful or explicit off the shelf. Then she takes scissors to the remaining pages and cuts them into small word-chunks, because your reader doesn’t read words, she reads chunks. Finally she shuffles the whole library, so the reader never gets ten cookbooks in a row and starts believing the world is only recipes. Only then does the reading begin.
That librarian is the training-data pipeline. Crawling is the photocopying; filtering, deduplication, scrubbing, tokenization, and shuffling are everything after. And here’s the uncomfortable truth the interview is really about: the model’s intelligence is capped by the librarian’s taste. Garbage in, garbage out — at a trillion tokens, the garbage is just better hidden.
Intuition: the interview question wearing an ML costume
“How would you build the data pipeline for training an LLM?” is the new “design a load balancer” — a systems question in disguise. The candidates who pass talk about data for twenty minutes before they mention the model.
Look at the numbers. GPT-3 (2020) trained on about 300 billion tokens. Llama 3 (2024) trained on over 15 trillion — fifty times more data in four years. The architecture barely changed; the data did. The Chinchilla results taught the field that models were systematically undertrained: data, not parameters, is the scarce input. Whoever manufactures the best trillion tokens wins.
And the pipeline is where the money burns. Every junk document that survives filtering wastes GPU time — and GPU time is the most expensive line item in AI. A bad pipeline trains on spam; a good one is a moat. Notice what frontier labs publish (model weights, sometimes) versus what they never publish (their filtering stack). That silence tells you where the value is.
The pipeline turns the raw web into training tokens — acquire at petabyte scale, keep the best tenth, kill duplicates, scrub PII and toxicity, tokenize once, and stream a shuffled, weighted mix to the GPUs.
How it works: six stages, one ruthless librarian
1. Acquire: photocopy the internet
It starts with a crawl. Common Crawl is the open one — a nonprofit that has crawled the web every month since 2007, petabytes of pages, free for anyone to download. Frontier labs run their own crawlers too, ingesting terabytes a day. Acquisition is the easy, solved part: bandwidth and storage.
On top of the crawl sit curated corpora — Wikipedia, books, code, papers. Small in volume, enormous in signal. The crawl gives you breadth; the curated sources give you quality. The mixture between them is a design decision, not an accident.
2. Filter: keep the best tenth
Raw web text is mostly junk — ads, SEO spam, boilerplate, machine-generated sludge. Filtering runs in layers. First language identification: keep the languages you want, drop the rest. Then heuristic filters: documents that are too short, too repetitive, or heavy on symbols fail fast — the cheap tests catch the obvious garbage.
Then the serious layer: model-based quality classifiers. Train a classifier on known-good text (Wikipedia, books, edited articles) versus raw crawl, and keep only documents that score like the good stuff. Typical yield: on the order of ten percent of the crawl survives. Nine-tenths of the internet is, by the pipeline’s standards, not worth training on.
Heuristics are cheap and catch the obvious junk; a classifier trained on good text catches the subtle junk. Together they decide what the model is allowed to read.
3. Deduplicate: kill the copies
The web is the same article copied a thousand times. Exact duplicates die by hashing — normalize the text, SHA-256 it, drop repeats. The hard part is near-duplicates: the same article slightly reworded across a thousand sites.
That’s MinHash (Broder, 1997): fingerprint each document’s word shingles, compare fingerprints instead of full documents, and bucket similar ones with locality-sensitive hashing. Billions of comparisons become feasible because you never compare the documents themselves.
Why it matters beyond hygiene: duplicates don’t just waste compute — they teach memorization. A model that has seen the same paragraph fifty times will recite it verbatim. Deduplication is a safety feature wearing a janitor’s uniform.
4. Scrub: remove the dangerous bits
Two things must not reach the GPUs. PII — emails, phone numbers, ID numbers — found by patterns and classifiers, then redacted or dropped with the whole document. Miss this and the model leaks real people’s phone numbers; it has happened. And toxicity — hate speech, explicit content, graphic violence — filtered hard by classifiers. The scrub is never perfect, which is why it’s defense in depth: filter at the pipeline, then again with safety training later.
5. Tokenize: cut text into numbers
Models don’t read words or characters — they read tokens, chunks learned by byte-pair encoding: frequent character pairs merged until the vocabulary covers the language. Modern models use roughly 100,000 tokens. The discipline that matters: tokenize once. Convert the whole corpus to token IDs a single time and store them — re-tokenizing petabytes is a bill nobody pays twice. Tokens, not words, are the currency: context windows, training budgets, and API prices are all measured in them.
6. Mix and stream: the feeding schedule
The last stage decides what the model eats, in what order. Domain mixture sets the weights — more code for a code model, more papers for a research model. Global shuffle destroys ordering bias: the model must never read ten cookbooks in a row and conclude the world is recipes. Documents are packed into fixed-length sequences, sharded across storage, and streamed during training so thousand-GPU clusters never sit idle waiting for data.
The pipeline, drawn: the crawl is the widest stage and the filter the narrowest — nine-tenths of the web never makes it past step two.
Three pipelines, one playbook
Every training-data pipeline you will ever meet is the same six stages wearing different uniforms. Here are the uniforms that matter in interviews.
Common Crawl is the raw material the open world builds on: a nonprofit crawl of the web, petabytes deep, refreshed monthly since 2007, free to download. Almost every open dataset — and quietly, many closed ones — starts here. It solves acquisition; everything after it is your pipeline’s job.
The Pile (EleutherAI) — 825 GiB across 22 sources, the dataset behind GPT-J and GPT-NeoX — and Dolma (AI2) — 3 trillion tokens with the pipeline code itself open-sourced — show what a serious filtering and dedupe stack looks like in public. They’re the closest thing to reading a frontier lab’s playbook.
Llama 3 trained on over 15 trillion tokens — a data scale that only works because the filtering stack is industrial. Meta publishes the token count, not the classifier. That asymmetry is the tell: the crawl is commoditized, the filter is the moat.
Source: The Pile (EleutherAI) — 825 GiB open pretraining dataset ↗
Source: Dolma (AI2) — 3T-token open corpus with open pipeline tooling ↗
The uniforms change — open crawl, open dataset, frontier lab — but the playbook doesn’t: acquire wide, filter ruthlessly, dedupe, scrub, tokenize once, stream a shuffled mix.
The scale, worked by hand
Four numbers that reframe the whole topic — the kind of arithmetic worth doing out loud in the interview.
Code: the filter and the fingerprint
The two pieces interviewers ask you to whiteboard: a heuristic quality filter and a MinHash near-duplicate check. Simplified, but the shape is real.
import re, hashlib
def quality_score(text: str) -> float:
"""Heuristic junk filter: 1.0 = keep, 0.0 = discard."""
words = text.split()
if len(words) < 50: # too short to learn from
return 0.0
if len(set(words)) / len(words) < 0.3: # repetitive boilerplate
return 0.0
punct = sum(1 for c in text if not c.isalnum() and not c.isspace())
if punct / max(len(text), 1) > 0.3: # punctuation soup
return 0.0
return 1.0
def shingles(text: str, k: int = 5):
words = re.sub(r"\W+", " ", text.lower()).split()
return {" ".join(words[i:i+k]) for i in range(len(words) - k + 1)}
def minhash_signature(shingle_set, num_hashes: int = 64):
"""Simplified MinHash: minimum hash value per hash function."""
sig = []
for seed in range(num_hashes):
h = min(
int(hashlib.md5(f"{seed}:{s}".encode()).hexdigest(), 16)
for s in shingle_set) if shingle_set else 0
sig.append(h)
return sig
def jaccard_estimate(sig_a, sig_b) -> float:
"""Fraction of matching signature slots ≈ Jaccard similarity."""
return sum(1 for a, b in zip(sig_a, sig_b) if a == b) / len(sig_a)
# pipeline rule: keep a doc only if quality_score == 1.0
# and its signature is dissimilar (< 0.8) to everything kept so far.
# Production systems add LSH banding so billions of docs stay feasible.Interview prompts
How the question actually sounds in the room — and what the interviewer is listening for.
Walk the six stages in order, but spend most of your time on filtering and deduplication — that’s where judgment lives. Name the yield (~10% survives) and the failure mode of skipping each stage.
MinHash plus locality-sensitive hashing: fingerprint shingle sets, compare fingerprints — never the documents. Say why exact hashing alone isn’t enough.
Ordering bias. The model must not learn that cookbooks arrive in clumps; global shuffle plus a deliberate domain mixture is part of the design, not housekeeping.
Detection (patterns plus classifiers), then redact or drop the document — and admit it’s never perfect, so it’s defense in depth with safety training downstream.
Gold for code and math, where correctness is verifiable; risky at scale — too much and quality drifts, a copy of a copy. Used sparingly, with provenance logged.
Takeaways
Crawl wide, filter ruthlessly. Acquisition is the solved part; the pipeline’s value is what it throws away.
Deduplication is a safety feature, not just hygiene. Duplicates waste compute and teach verbatim memorization.
Tokenize once. BPE turns text into the model’s currency — never re-pay the cost of tokenizing petabytes.
Shuffle and mix deliberately. Ordering bias is real; the feeding schedule is part of the design.
The filter is the moat. Everyone can download Common Crawl. The quality classifier is what frontier labs don’t publish.
Sources
- Common Crawl (petabyte-scale open web crawl, monthly since 2007)
- Broder, 1997 — “On the resemblance and containment of documents” (MinHash)
- The Pile (EleutherAI) (825 GiB open pretraining dataset)
- Dolma (AI2) (3T-token corpus with open pipeline tooling)
- Brown et al., 2020 — GPT-3 (~300B training tokens)
- Meta — Llama 3 announcement (15T+ training tokens)