Design YouTube: Async Uploads, Transcode Farms & the Edge
The interview’s boss fight: one logo, two completely different systems. Uploads go async to blob storage and through a transcode farm; playback serves skewed demand from the edge in tiny segments. Get the split right and the rest of the interview runs itself.
Explain it like I’m five
Imagine a giant bakery that also owns delivery trucks. When you bring in your birthday cake recipe, the bakery doesn’t make you wait while it bakes — it gives you a ticket with a number and says “come back later.” Behind the counter, ovens bake your cake in every size, from cupcake to wedding-cake. The ticket goes into a card catalog so anyone can find it. And when a million kids want your cake, the bakery doesn’t ship from one kitchen — it stocks little corner stores near every kid, and each kid gets their cake one slice at a time, small slices on a slow day, big slices when the truck is fast.
YouTube is that bakery. Uploading is handing over the recipe and getting a ticket (an upload ID) while the work happens in the background. Transcoding is the ovens baking every size — 144p to 4K. The metadata database is the card catalog: titles and view counts, not the cakes themselves. The edge cache is the corner stores near the viewers. And segmented streaming is serving the cake one slice at a time, sized to how fast your truck is moving. Everything in this guide is just the bakery and the trucks, engineered for billions of cakes.
Intuition: two systems wearing one logo
“Design YouTube” is the interview’s boss fight because it punishes the candidate who designs one system. Uploads and playback have opposite shapes — one is write-heavy, bursty, and latency-tolerant; the other is read-heavy, skewed, and latency-obsessed. The candidates who pass say the split out loud in the first minute; the candidates who struggle try to serve both with the same boxes.
The split gives you two clean stories to tell. Story one — the factory: a creator uploads a file, and everything after the HTTP request is asynchronous: chunks land in blob storage, a queue feeds a farm of transcoders, and the video becomes a family of files, one per resolution. The creator gets an upload ID in milliseconds and walks away; the factory works for minutes. Story two — the delivery network: a viewer presses play, and the system must start the first frame in about a second, at a quality their connection can sustain, from a copy that is physically close to them. The database never touches the video bytes — it points at them. The hot videos live on the edge; the cold ones are fetched from origin on demand.
Async upload to blob storage, a transcode farm for every resolution, metadata in a database sharded by video ID, hot videos pushed to the edge, cold ones served from origin, and everything streamed in tiny adaptive segments.
How the pieces fit — the factory, then the trucks
Upload: chunks, an upload ID, and blob storage
The creator’s file arrives over HTTP, but never as one giant blob and never synchronously. The client chops it into chunks — a few megabytes each — and uploads them in parallel, each tagged with an upload ID minted at the start. The server reassembles the chunks into blob storage: dumb, cheap object storage that holds bytes and asks no questions. Hand the creator the upload ID the moment the first chunk lands, and do everything else in the background — transcoding, thumbnails, metadata extraction. The upload request returns in milliseconds; the video is watchable minutes later.
Chunks buy you resumability for free. A 4 GB upload that dies at 90% doesn’t restart — the client asks which chunks are missing and re-sends only those. In the interview, say “resumable chunked upload” before the interviewer has to ask; it’s the detail that separates a whiteboard answer from a shipped one.
The transcode farm: one upload becomes a dozen files
Raw uploads are unplayable as-is: wrong codecs, no adaptive ladder, one giant file. So every finished upload drops a job onto a queue, and a fleet of transcode workers pulls jobs off it. Each worker re-encodes the video into every resolution in the ladder — 144p, 240p, 360p, 480p, 720p, 1080p, 1440p, 4K — plus thumbnails and a preview sprite. One upload becomes a family of a dozen-plus files, all back in blob storage.
The queue is load-bearing, not decorative. Uploads spike — think a breaking news event — and the transcode fleet scales with the queue depth, not with the upload rate. Slow creators never block fast ones, and a poisoned file that crashes a worker only costs one job, retried with a backoff. Prioritize the ladder smartly: mint the low resolutions first so the video is watchable quickly, then fill in 4K behind it. The interview line: decide the upload is done the moment bytes land; everything after is a background job.
Metadata: a database that never touches video bytes
Titles, descriptions, uploaders, view counts, likes — that lives in a regular relational database, sharded by video ID. The database row never holds video bytes; it points at them. This matters because the two access patterns want different things: the metadata is small, hot, and read constantly (every page load, every recommendation), while the bytes are huge, cold, and streamed sequentially.
Two consequences follow. First, hot metadata gets cached aggressively — everybody reads the same trending page, so a cache in front of the database absorbs the skew. Second, view counts are a write problem disguised as a read problem: millions of increments a day. You don’t UPDATE a row per view; you buffer counters (in-memory aggregation, periodic flush) and accept that the displayed count is eventually consistent. Say that out loud — “the view count you see is approximate” — and the interviewer knows you’ve thought about write amplification.
Playback: the skew is the architecture
Views on video platforms are brutally skewed: a tiny slice of videos collects the overwhelming majority of views. So the playback architecture is just skew management. Hot videos are pushed to edge caches — CDN points of presence physically close to viewers — where the first byte travels tens of milliseconds instead of crossing an ocean. Cold videos are served from origin on demand: the first unlucky viewer waits for the origin fetch, the edge caches it, and the next viewer gets it fast. Never pay to cache what nobody watches; never serve the hits from origin.
The origin shield sits between the edge and the core: one caching layer that collapses duplicate origin fetches, so a viral video doesn’t thundering-herd your storage when a thousand edge nodes all miss at once. This is the piece candidates forget and interviewers probe — “a video goes viral at midnight, what breaks?” The answer: nothing, if the shield absorbs the miss storm and the hot set was pre-warmed; otherwise your origin melts.
Segments: why nothing ships as one giant file
Each rendition is chopped into segments of a few seconds, and the player fetches them one at a time. The server also serves a manifest — a playlist listing the segments and the available bitrates. Two open standards define this dance: MPEG-DASH, the industry-standard adaptive streaming format maintained by the DASH Industry Forum, and Apple’s HLS (HTTP Live Streaming), Apple’s protocol for the same job. Both are plain HTTP: no special streaming servers, just files and playlists, which is exactly why they scale — every CDN on earth already knows how to cache files.
Segments are what make adaptive bitrate possible. The player measures its download speed and picks the next segment’s quality to match: strong wifi, 4K segments; a tunnel, 144p. That’s why video goes blurry on bad wifi instead of freezing — the player is trading quality for continuity, segment by segment. And segments make seeking cheap: jumping to minute 40 means fetching segment 400, not downloading the first 39 minutes. In the interview, “segmented adaptive streaming” is the phrase that ends the playback discussion.
Uploads finish the moment bytes land; a background farm makes every resolution. Playback serves the hot head from the edge and the cold tail from origin — everything in segments.
Deletes: a million copies don’t vanish at once
When a creator deletes a video, the metadata row flips to deleted instantly — but the bytes live on in blob storage, transcode outputs, and a million edge caches. Purge propagation is not instant, and the design should admit it. Two honest approaches: versioned URLs — every rendition URL embeds a version hash, so a re-upload never collides with a stale cached copy and deletes simply stop issuing new URLs — or TTL-based staleness: edge entries carry short time-to-live values, so a deleted video may remain viewable for minutes while caches expire, which you state as an accepted tradeoff. The interviewer’s real question is whether you know the cache is a copy, not the truth. Say it: “the database is the truth; the edge is a rumor with a TTL.”
Upload is a write problem solved with asynchrony — chunks, an upload ID, and a transcode queue — so the creator never waits. Playback is a read problem solved with geography — edge caches for the hot head, origin for the cold tail, adaptive segments for the last mile. Two systems, one logo.
Who runs this at scale — and what they learned
Nobody designs video delivery from first principles anymore — the patterns are settled, and each one has a canonical source. Interviewers don’t expect you to have built YouTube; they expect you to know where the settled answers live.
The YouTube team’s own retrospective — “7 Years of YouTube Scalability Lessons in 30 Minutes,” published on the High Scalability blog — is the closest thing to a primary source on how the architecture grew under explosive demand: separate the serving concerns, keep the critical path dumb, and let the system degrade gracefully. When the interviewer asks “how do you know this works,” the answer is that the team that runs it wrote the postmortem in public. Read the lessons →
MPEG-DASH is the open standard for adaptive streaming over HTTP, maintained by the DASH Industry Forum. The model is exactly this guide’s playback path: the server publishes a manifest describing segments at multiple bitrates, and the client — not the server — decides which segment to fetch next. That client-driven choice is the whole trick: the server stays stateless and cacheable, and every CDN on earth already knows how to serve files. dashif.org →
Apple’s HTTP Live Streaming (HLS) is Apple’s protocol for the same job — playlists (.m3u8 manifests) plus short media segments over plain HTTP. Apple’s developer documentation is the reference implementation of the segment-and-manifest pattern, and on Apple devices it’s the native path. Knowing both names and that they’re the same idea in different packaging is the interview-ready answer. Apple’s HLS docs →
Cloudflare’s CDN explainer is the clearest public account of what the “edge” actually does: cache static content close to users so the origin rarely sees a request. The hot-head/cold-tail split in this guide is that explainer turned into an architecture — popular videos live on the edge because the math says they must, and the origin survives because the edge absorbs the skew. What a CDN does →
The through-line: nothing here is invented. The factory pattern is YouTube’s retrospective, the playback pattern is two public standards, and the edge pattern is the CDN industry’s daily bread — the interview is testing whether you can assemble settled parts, not invent new ones.
Work the storage math by hand
State your assumptions out loud, then derive. Take ~500 hours of video uploaded per minute (the figure YouTube has publicly cited), an average stored bitrate of ~1 GB per hour of 1080p, and 6 renditions per upload. None of these are exact — the interview grades the derivation, not the constants.
Five numbers, one story: uploads are a petabytes-per-day write problem solved with async queues; playback is an exabyte-per-day read problem solved with geography; metadata is a rounding error. Derive them in this order and the interview runs itself.
Chunks, manifests, and hot-or-cold — in Python
The three ideas that survive contact with the interview: reassembling a chunked upload, writing an HLS-style manifest, and routing a playback request to edge or origin. Under sixty lines.
import hashlib
# --- 1. Chunked upload: the server only reassembles, never waits ---
uploads = {} # upload_id -> {total, chunks: {idx: bytes}}
def start_upload(total_chunks: int) -> str:
uid = hashlib.sha256(str(total_chunks).encode()).hexdigest()[:12]
uploads[uid] = {"total": total_chunks, "chunks": {}}
return uid # handed to the creator in ms; work happens later
def put_chunk(uid: str, idx: int, data: bytes):
u = uploads[uid]
u["chunks"][idx] = data
if len(u["chunks"]) == u["total"]:
blob = b"".join(u["chunks"][i] for i in range(u["total"]))
del uploads[uid]
return blob # -> blob storage, then enqueue transcode job
return len(u["chunks"]) # resume: client re-sends only missing idx
# --- 2. HLS-style manifest: the player, not the server, adapts ---
def manifest(video_id: str, renditions):
"""renditions: [(bandwidth_bps, segment_urls...), ...]"""
lines = ["#EXTM3U"]
for bw, segs in renditions:
lines.append(f"#EXT-X-STREAM-INF:BANDWIDTH={bw}")
lines.extend(f"#EXTINF:6.0,{u}" for u in segs)
return "\n".join(lines)
# --- 3. Edge routing: hot head served close, cold tail from origin ---
edge, HOT = {}, set()
def serve(video_id: str, segment: int):
key = (video_id, segment)
if video_id in HOT and key in edge:
return "edge", edge[key] # ms away, not an ocean away
data = fetch_from_origin(video_id, segment) # cold: pay once
if video_id in HOT:
edge[key] = data # now the next million views are free
return "origin", dataThree ideas, one file: the upload ID decouples the creator from the factory; the manifest hands adaptation to the player so the server stays stateless; and the edge cache turns the hot head of demand into free reads — the origin only ever pays for the cold tail once.
Six questions that test the real understanding
What to say (≈90 sec): “First the split: upload and playback are different systems. Uploads: the client chunks the file, gets an upload ID in milliseconds, and everything after is async — chunks reassemble in blob storage, a queue feeds a transcode farm that mints every resolution from 144p to 4K. Metadata — titles, view counts — goes in a database sharded by video ID; the bytes never touch it. Playback: views are brutally skewed, so hot videos live on edge caches near viewers and cold ones are fetched from origin on demand, behind an origin shield so a miss storm doesn’t melt storage. Everything streams as few-second segments with a manifest — DASH or HLS — so the player adapts quality to bandwidth. Rough math: ~500 upload-hours a minute is petabytes a day of new bytes, and ~1B watched hours a day is an exabyte served — which is why the edge exists.”
Likely follow-up: “Why not one system for both?” → Opposite shapes: uploads are bursty writes that tolerate minutes of latency; playback is skewed reads that need the first frame in a second. One box compromises both.
The answer that sinks you: “Store the video in the database and stream it from the app server.” Why it fails: it fuses the bytes to the metadata path and puts the exabyte read load on boxes sized for kilobytes — the interview is over at that sentence.
What to say (≈90 sec): “A 4K upload can take minutes; holding an HTTP request open that long ties up a connection, a thread, and the creator’s patience — and any blip restarts the whole thing. Async means the request ends at ‘bytes received’: the client gets an upload ID in milliseconds and the transcode farm works the queue at its own pace. Chunking adds resumability — a failure at 90% re-sends only the missing chunks. If it were synchronous, one slow creator would hold a worker for the entire transcode, throughput would collapse to the slowest upload, and a single network hiccup would cost the whole file.”
Likely follow-up: “How does the creator know when it’s watchable?” → Poll the upload ID, or push a notification on job completion — the ID is the handle for the whole lifecycle.
The answer that sinks you: “Make the request timeout longer.” Why it fails: it treats a design problem as a configuration problem — the queue and the workers are the design, not the timeout.
What to say (≈90 sec): “If the hot set was pre-warmed, nothing — the edge absorbs it. What breaks without preparation is the origin: the first wave of edge misses all fetch the same segments simultaneously, a miss storm. The fix is the origin shield — one caching layer that collapses a thousand duplicate fetches into one origin read. After that, the metadata database feels the view-count writes, which is why counts are buffered and eventually consistent, not updated per view. The answer the interviewer wants: name the miss storm, name the shield, and say pre-warming is a product decision — trending predictions push bytes to the edge before the spike.”
Likely follow-up: “The shield itself gets hot?” → Shields are regional and anycast — the load spreads across shield PoPs, and the origin still only sees one fetch per region per segment.
The answer that sinks you: “Auto-scale the origin.” Why it fails: the origin holds exabytes it can’t replicate in minutes — you don’t scale the source of truth, you shield it.
What to say (≈90 sec): “Three reasons. Adaptive bitrate: the player picks each segment’s quality to match current bandwidth, which is why video goes blurry instead of freezing — impossible with one file. Seeking: jumping to minute 40 fetches segment 400, not the first 39 minutes. Caching: segments are small, immutable, cacheable files — every CDN already knows how to serve them, which is why DASH and HLS won. The manifest is just the playlist tying them together. One big file gives you none of this: no adaptation, expensive seeks, and a cache story that fights the CDN instead of using it.”
Likely follow-up: “Who picks the quality — client or server?” → The client, always. The server stays stateless and cacheable; adaptation logic lives in the player. That’s the deep reason the standards look the way they do.
The answer that sinks you: “Segments load faster.” Why it fails: true but shallow — the real answers are adaptation, seeking, and CDN-cacheability, and “faster” names none of them.
What to say (≈90 sec): “The metadata row flips to deleted instantly — that’s the truth. The edge is a rumor with a TTL: cached copies expire on their own schedule, so for minutes the video may still play. Two honest designs: versioned URLs, where every rendition embeds a content hash so a re-upload never collides with stale copies and deletes just stop issuing URLs; or short TTLs, accepting brief staleness as a tradeoff you state out loud. What you must never do is claim the purge is instant — the interviewer is testing whether you know the cache is a copy, not the source of truth.”
Likely follow-up: “Legal takedown needs it gone in seconds?” → Active purge APIs exist — the CDN walks its edge nodes invalidating the key — but it’s an explicit, expensive operation, not the default path.
The answer that sinks you: “Delete it from the database; the caches will figure it out.” Why it fails: caches don’t figure anything out — without versioning, TTLs, or an explicit purge, the deleted video plays until every copy independently expires.
What to say (≈90 sec): “State assumptions: ~500 upload-hours per minute, ~1 GB per hour of 1080p, ~6 renditions per upload. That’s 720,000 hours a day, ~4.3 petabytes of new bytes daily — over an exabyte a year. Metadata is the contrast: ~30M uploads a day at ~1 KB per row is ~30 GB a day, a rounding error, which is exactly why the database is sharded by video ID without breaking a sweat while the bytes live in blob storage. And the playback number dwarfs both: ~1B watched hours a day is ~1 exabyte served — no origin fleet does that, so the edge exists. The interview grades the derivation: assumptions stated, then upload → storage → playback, in that order.”
Likely follow-up: “Cut storage cost 30% — where?” → Lifecycle policies: drop 4K renditions for videos with no views in 90 days, re-transcode on demand. The cold tail pays for the hot head.
The answer that sinks you: “A few petabytes, it’s fine.” Why it fails: no assumptions, no derivation, no contrast between the byte problem and the metadata problem — the math is the interview.
Key takeaways
- Say the split first: uploads are an async write problem, playback is a skewed read problem — two systems, one logo.
- Chunked upload with an upload ID: the creator gets an answer in milliseconds, the transcode farm works the queue for minutes, and failures resume at the missing chunk.
- Metadata is sharded by video ID and never touches bytes; view counts are buffered and eventually consistent — the displayed number is approximate.
- The edge exists because of the skew: hot videos served in milliseconds from nearby caches, cold ones fetched once from origin behind a shield.
- Segments make everything else possible: adaptive bitrate, cheap seeks, and CDN-cacheable files — the player adapts, the server stays stateless.
- The cache is a rumor with a TTL: deletes flip the metadata row instantly, but bytes expire on their own schedule — version URLs or accept brief staleness.
Sources & further reading
Every claim in this guide traces to one of these — the wording is ours, the ideas are credited.
- 7 Years of YouTube Scalability Lessons in 30 Minutes — the YouTube team’s own retrospective on High Scalability: how the architecture grew under explosive demand.
- DASH Industry Forum (dashif.org) — the home of MPEG-DASH, the open standard for adaptive streaming over HTTP: manifests, segments, client-driven bitrate choice.
- Apple HTTP Live Streaming — Apple’s developer documentation for HLS: .m3u8 playlists plus short media segments over plain HTTP.
- What is a CDN? — Cloudflare — what edge caching actually does: serve content close to users so the origin rarely sees a request.