The Interview Edge Blog
← Back to all guides
Systems · System design

Design a Twitter Timeline: Push vs Pull Fan-Out

The interview question that separates seniors from everyone else: your timeline is read ~100x more than it’s written, so the whole design is one bet — optimize for reads. Push fan-out, pull fan-out, and the hybrid real systems actually run.

Explain it like I’m five

Imagine you run a newspaper for a whole city, and every reader wants their own personal front page.

When someone writes a story, you could photocopy it and slip it into every subscriber’s personal stack right away — that’s fan-out on write (push). It’s lovely for readers, because their stack is always ready. But when a famous columnist writes one story, you’re suddenly photocopying it a hundred thousand times.

Or you could never pre-copy anything. When a reader opens their paper, you sprint around collecting the latest stories from everyone they follow — that’s fan-out on read (pull). No photocopying marathons, but every reader waits while you run.

The grown-up answer: photocopy for ordinary writers, sprint for the famous ones. That hybrid is the whole interview.

Intuition: the timeline is read ~100x more than it’s written

Say this first in the interview and the rest of the design falls out of it.

Twitter’s own engineers have described the service as primarily a consumption mechanism, not a production mechanism — on the order of 300K queries per second spent reading timelines against only a few thousand per second spent on writes. Millions of people refresh constantly; each person posts rarely.

That asymmetry is the entire strategy. Reads must be nearly free; writes are allowed to be expensive. Every decision below is just this sentence wearing different clothes.

One sentence worth memorizing

Precompute timelines on write so reads are a single cache lookup — except for celebrities, whose fan-out is too expensive to precompute, so merge them at read time.

How the pieces fit — write path, then the two fan-outs

The write path: save the tweet first, fan out second

Someone posts. Step one has nothing to do with timelines: persist the tweet itself to a tweet store, sharded by tweet ID. That’s the easy part — a key-value write.

Step two is the actual interview question: getting that tweet into every follower’s timeline. Two classic moves, opposite tradeoffs.

Fan-out on write (push): precompute everyone’s timeline

On every post, look up the author’s followers and append the tweet ID to each follower’s timeline cache. A normal user with 200 followers costs 200 cache writes. Reads become trivial: open the app, read your precomputed list, done.

The failure mode has a name: write amplification. A celebrity with 100K followers turns one tweet into 100K writes. When celebrities tweet at each other, the write path catches fire — and slow fan-out is why replies can visibly arrive before the original tweet.

Fan-out on read (pull): compute nothing, fetch live

The mirror image: store nothing per user. When someone opens the app, fetch the latest tweets from everyone they follow, right now, and merge them. Write amplification drops to zero.

But now the read path does all the work — every timeline load fans out to hundreds of stores and merges on the fly. It’s slow, and it melts under load. You traded a write problem for a read problem, and reads are the thing you were supposed to protect.

The hybrid: push for the many, pull for the few

Real systems split the difference. Normal users get push — their follower counts are small, so precomputing is cheap and reads stay fast. Celebrities get pull — their tweets are fetched at read time and merged into the precomputed timeline. You only pay for what you actually use.

This is the answer that gets the nod. Pure push dies on celebrities; pure pull dies on reads; the hybrid is what Twitter converged on — doing more of the work on reads specifically for high-fan-out users.

Storage: two layers, both cached

Each user’s timeline cache holds a list of recent tweet IDs, newest first, capped at a few hundred — not the tweet contents. Contents live in a separate store keyed by tweet ID. Rendering a timeline is two hops: read the ID list, then batch-fetch the contents. Both hops hit cache in the common case.

Splitting IDs from contents is what makes deletes and edits survivable: the timeline list never embeds anything that can go stale.

Who runs this at scale — and what they learned

This isn’t a thought experiment; it’s a retelling of how Twitter actually evolved.

Twitter’s VP of Engineering at the time, Raffi Krikorian, described the shift in a talk on timelines at scale: as high-follower accounts became the common case rather than the exception, Twitter moved from doing all the work on writes toward doing more work on reads for high-value users — the hybrid, stated plainly. The system was pushing roughly 300K timeline-read queries per second against ~6K writes per second, serving 150M active users with a target of getting a celebrity’s tweet to tens of millions of followers in under 5 seconds.

Two operational lessons from that era are worth stealing for your interview. First, replies arriving before the original tweet is the visible symptom of slow fan-out — ordering across a distributed fan-out is genuinely hard, which is why you order by timestamp at merge time. Second, the storage layer behind the timelines was a fleet of Redis clusters holding the precomputed ID lists, because a timeline is fundamentally a list data structure with range reads.

The general principle travels: Instagram, LinkedIn, and every feed product since has faced the same push/pull decision and landed on the same hybrid.

Work the write-amplification math by hand

Interviewers love this because the arithmetic kills one design in thirty seconds.

Assume 400M tweets per day (Twitter’s stated volume in the early 2010s) and an average of 200 followers per user. Pure push fan-out costs 400M × 200 = 80B timeline writes per day — about a million writes per second, sustained, just to keep timelines fresh.

Now add one celebrity with 30M followers tweeting 5 times a day: 5 × 30M = 150M writes per day from a single user. That one account costs nearly twice the entire average-user population’s fan-out. This is the moment in the interview where you say “so we exempt celebrities from push” — and the hybrid designs itself.

On the read side, cap timeline caches at ~800 tweet IDs of ~8 bytes each: ~6.4 KB per user. For 150M active users that’s roughly a terabyte of timeline cache — entirely reasonable for a Redis fleet, and the reason the ID-list design wins.

The hybrid, in Python

Sketch-quality code that shows the decision, not the infrastructure.

CELEBRITY_THRESHOLD = 10_000  # followers

def post_tweet(author, tweet_id):
    tweet_store.save(tweet_id, author.content)          # sharded by tweet_id
    if author.follower_count < CELEBRITY_THRESHOLD:
        for follower in author.followers:               # PUSH
            timeline_cache.prepend(follower, tweet_id, cap=800)
    # celebrities: nothing precomputed; merged at read time

def read_timeline(user):
    ids = timeline_cache.get(user.id, limit=800)        # precomputed push part
    for celeb in user.followed_celebrities:             # PULL part
        ids += tweet_store.latest(celeb, limit=50)
    ids = merge_by_timestamp(ids)                       # ordering happens here
    return tweet_store.batch_get([i for i in ids if not tombstoned(i)])

Notice where each interview talking point lives in the code: the threshold is the hybrid, merge_by_timestamp is the ordering answer, and the tombstoned check is the delete answer.

Six questions that test the real understanding

Meta
“Design a Twitter timeline.” Walk me through it.

What to say (≈90 sec): Lead with the read/write asymmetry — reads ~100x writes, so optimize reads. Write path: persist tweet sharded by ID. Fan-out: push for normal users (precompute timeline caches), pull for celebrities (merge at read time). Storage: timeline caches hold ID lists, contents fetched by ID. Close with ordering and deletes.

Likely follow-up: “Why not just push for everyone?”

The answer that sinks you: Picking pure push or pure pull without naming the celebrity failure mode.

Google
A celebrity with 30M followers tweets. Walk me through exactly what happens.

What to say (≈90 sec): Under push, that’s 30M cache writes from one tweet — write amplification. So celebrities are exempted: their tweet is only persisted, and each follower’s read path fetches it live and merges it into their precomputed timeline by timestamp.

Likely follow-up: “What does the follower’s read latency look like now?”

The answer that sinks you: Saying the celebrity’s followers get the tweet “instantly” — the merge adds read-time cost, that’s the tradeoff.

Amazon
How do you keep timelines ordered newest-first when writes race?

What to say (≈90 sec): Don’t rely on write order — fan-out to 100K caches can’t be atomic. Stamp every tweet with a timestamp (or better, a hybrid logical clock) at persist time, and let the read-time merge sort by it. Out-of-order arrival is expected; ordering is a read concern.

Likely follow-up: “What about clock skew across machines?”

The answer that sinks you: Assuming a single global clock exists.

Apple
A user deletes a tweet. It’s already in a million caches — now what?

What to say (≈90 sec): You can’t unwrite a million caches cheaply. Write a tombstone for the tweet ID; the content fetch checks it and skips the tweet. The ID may linger in timeline lists until it ages out of the 800-item cap — acceptable, since it never renders.

Likely follow-up: “How long until it’s fully gone?”

The answer that sinks you: Proposing to synchronously delete from every follower’s cache.

Meta
Why store tweet IDs in the timeline instead of the full tweets?

What to say (≈90 sec): Three reasons: size (an 800-ID list is ~6KB; 800 full tweets is megabytes), consistency (one copy of the content means edits/deletes propagate via the content store, not via rewriting every timeline), and cache efficiency (hot tweet contents are shared across millions of timelines).

Likely follow-up: “So what’s the extra cost?”

The answer that sinks you: Not naming the second fetch — it’s a deliberate two-hop read.

Google
How much write traffic does pure push fan-out generate? Show me the math.

What to say (≈90 sec): 400M tweets/day × 200 avg followers = 80B timeline writes/day, ~1M writes/sec sustained. One 30M-follower celebrity tweeting 5x/day adds 150M writes/day alone. That arithmetic is the entire justification for the hybrid.

Likely follow-up: “Where would you put the celebrity threshold?”

The answer that sinks you: Guessing a number without tying it to write capacity.

Key takeaways

  1. Reads outnumber writes ~100:1 — the whole design optimizes reads.
  2. Fan-out on write (push) makes reads trivial but explodes on celebrities.
  3. Fan-out on read (pull) kills write amplification but melts the read path.
  4. The hybrid — push for normal users, pull + merge for celebrities — is what real systems run.
  5. Timeline caches store ID lists, not contents; rendering is two cached hops.
  6. Ordering is a read-time merge concern; deletes are tombstones, not rewrites.

Sources & further reading

Every claim in this guide traces to one of these — the wording is ours, the ideas are credited.

Back toAll guides →