The Interview Edge Blog
← Back to all guides
Systems · System design

Design a Chat System: the Switchboard Behind Every DM

HTTP hangs up after every request — chat can’t work that way. Here’s the persistent-connection architecture behind WhatsApp and iMessage: how a million open sockets get managed, how every message is stored and delivered exactly as promised, and why group chats are the real boss fight.

Explain it like I’m five

Imagine you and your best friend each have a walkie-talkie tuned to the same channel. You press the button, talk, and she hears you instantly. She presses hers and answers. Simple. Now imagine a million pairs of friends, all with walkie-talkies, and one giant switchboard in the middle. When you press your button, the switchboard has to know which channel your friend is listening on, patch you straight through — and because walkie-talkies cut out in tunnels, it writes down everything you said, so nothing ever gets lost.

A chat system is that switchboard. It keeps a million lines open at once, remembers who is listening on which line, delivers every word, and keeps a copy of everything said. Everything in this guide is the switchboard’s playbook, written out: how the lines stay open, how a message finds your friend, and the arithmetic of a thousand-person group chat.

Intuition: a phone call that never hangs up

HTTP is a doorbell: ring, answer, door closes. Chat is a phone call that never hangs up — a persistent, two-way pipe between your phone and a server. That pipe is a WebSocket: it starts life as an ordinary HTTP request, then both sides agree to upgrade it into a lasting connection. After that, either side can send at any time. No polling, no refresh button, no “new messages” spinner.

The whole design has three jobs. Hold the pipes: millions of open connections, each cheap to keep alive. Route the words: know which pipe your friend is on and hand the message across. Remember everything: store every message so history, search, and offline catch-up all work. Nail those three and the rest is detail.

How it works: pipes, maps, and promises

Connections. One server can’t hold a million sockets, so you run a fleet of connection servers — machines whose only job is babysitting open WebSockets. When you open the app, a load balancer (yes, the thing from the last guide) hands you to one of them. Then the system writes down the mapping: user → server 7. Lookups go in something blazing fast, like Redis, keyed by user ID. To find your friend, a server just asks: where is she connected?

Message flow. You send “hey”. It lands on your server first. Your server looks up your friend’s server, forwards the message server-to-server, and her server pushes it down her open socket. Three hops, milliseconds. If she’s offline, the message waits in a queue instead — and when her phone reconnects, it pulls everything it missed.

Storage. Every message lands in a messages table: message ID, conversation ID, sender, timestamp, text — partitioned by conversation, so one chat’s history lives together and pages fast. Photos and videos never touch the database: they go to blob storage, and the message stores only the link.

Delivery promises. Networks drop things, so the system acknowledges everything: the server ACKs the sender, the receiver ACKs the server. No ACK? Resend. That’s at-least-once delivery — the message always arrives, possibly twice — so the client dedupes by message ID before showing anything. A per-conversation sequence number keeps ordering sane even when retries arrive out of order.

Presence and receipts. “Online” is just a heartbeat: your app pings every ~30 seconds, and a few missed pings means offline. “Last seen” is the timestamp of your last ping. Read receipts are tiny messages themselves — they ride the exact same pipeline, flagged as receipts instead of text.

Switchboards of the real internet

WhatsApp. Built on Erlang, famous for a 2012 engineering post reporting over 2 million concurrent TCP connections on a single server — with CPU and memory to spare. Erlang’s featherweight processes made each idle connection nearly free, which is exactly the property a chat switchboard needs. Later they traded peak density for headroom as messages got richer, but the lesson stands: pick the runtime that matches the shape of the problem.

Discord. Runs its real-time layer on Elixir (Erlang’s modern cousin) and has written publicly about scaling to millions of concurrent users, sharding chat by guild so one server’s explosion never takes down the rest. Same switchboard, different wiring.

Telegram. Uses its own MTProto protocol and draws the storage line in two places: cloud chats live on Telegram’s servers (history everywhere, instantly), while secret chats are end-to-end encrypted and live only on the two devices. One architecture, two trust models.

iMessage / push notifications. When your socket is dead — phone asleep, app killed — the message still reaches you as a push notification via Apple’s or Google’s push service. Push is the offline path every chat system quietly depends on: the queue drains through a different pipe.

The fan-out math, by hand

Ten million users, each sending 40 messages a day: 400M messages/day, about 4,600/sec on average. Peaks run ~10× average, so provision for ~46,000/sec. Storage at ~500 bytes per message with metadata: 200 GB/day, ~73 TB/year before replication. That’s the one-to-one math — now the group math.

Fan-outOne message into a 1,000-member group is 1,000 deliveries. A thousand messages a second into big groups is a million deliveries a second — two hundred times the one-to-one rate for the same message count. Groups don’t add load, they multiply it. This is why WhatsApp caps group sizes: the cap is a load-shedding device wearing a product feature’s clothes.
Push vs pullPush to online members immediately; queue for offline members to pull on reconnect. If 30% of a big group is offline at any moment, 300 of every 1,000 deliveries wait in the queue — the queue absorbs the burst so the hot path never sees it. Sizing the offline queue is sizing for your worst commute-hour reconnect storm.
ConnectionsTwo million concurrent sockets at ~50 KB of state each is ~100 GB of connection state fleet-wide. WhatsApp’s 2M-connections-per-box result shows a single tuned machine can hold an astonishing share of that — the fleet exists for failure domains and geography, not because one box can’t.
Dedup windowAt-least-once means duplicates; the client dedupes by message ID. Keep a rolling set of recent IDs per conversation — at 46,000 msgs/sec, a 60-second window is ~2.8M IDs, a few tens of MB in a hash set. Cheap insurance against the double-send your users will absolutely notice.

Three patterns in Python

A consistent-hash ring assigning users to connection servers, an idempotent message store that dedupes retries, and an ACK-and-retry sender — the three patterns behind almost every “how does the message get there safely” follow-up.

python · consistent hash ring, idempotent store, ack-and-retry sender
import hashlib
from bisect import bisect


class ConnectionRing:
    # user -> connection server: the "this user lives on server 7" map
    def __init__(self, servers, replicas=200):
        self.ring = {}
        for s in servers:
            for i in range(replicas):
                h = int(hashlib.md5(f"{s}#{i}".encode()).hexdigest(), 16)
                self.ring[h] = s
        self.keys = sorted(self.ring)

    def server_for(self, user_id):
        h = int(hashlib.md5(str(user_id).encode()).hexdigest(), 16)
        return self.ring[self.keys[bisect(self.keys, h) % len(self.keys)]]


class MessageStore:
    # at-least-once in, exactly-once out: dedupe by message id
    def __init__(self):
        self.by_conversation = {}
        self.seen_ids = set()

    def store(self, msg):
        if msg["id"] in self.seen_ids:
            return False  # duplicate retry: drop it
        self.seen_ids.add(msg["id"])
        conv = self.by_conversation.setdefault(msg["conv"], [])
        conv.append(msg)
        return True

    def history(self, conv, before_seq, limit=50):
        msgs = [m for m in self.by_conversation.get(conv, [])
                if m["seq"] < before_seq]
        return msgs[-limit:]


def send_with_ack(send_fn, msg, max_retries=3):
    # no ACK? resend. duplicates are harmless: the store dedupes
    for attempt in range(max_retries):
        ack = send_fn(msg)
        if ack:
            return True
    raise TimeoutError("receiver unreachable: queued for pull on reconnect")

Six questions that test the real understanding

Meta
“Design WhatsApp.” Walk me through it.

What to say (≈90 sec): “Start with the pipe: every client holds a persistent WebSocket to a connection server — no polling. A million users means a fleet of connection servers, and a fast lookup maps each user ID to their server. Sending is three hops: my server, a lookup, your server, down your socket. Everything is stored in a messages table partitioned by conversation, media in blob storage with the message holding the link. Delivery is at-least-once with ACKs at every hop and client-side dedup by message ID, plus per-conversation sequence numbers for ordering. Offline users get queued messages pulled on reconnect, with push notifications as the backup pipe. Presence is heartbeats; read receipts ride the same pipeline. And the scaling trap is group chats — one message fans out to every member, so you push to online members and queue for offline ones.”

Likely follow-up: “How do you handle a user with two devices?” — each device holds its own connection and gets its own entry in the user-to-server map; the fan-out treats a user’s devices like tiny group members.

Discord
How do read receipts work without melting the servers?

What to say: “Receipts are just tiny messages — same pipeline, flagged as receipts. The trick is aggregation: don’t fan out a receipt per message. When I open a chat with 40 unread, the client sends one receipt for the highest sequence number read, and the server treats everything below it as read. In a group, receipts fan out like messages but they’re small and coalescible — batch them per second per conversation and the receipt traffic stays a rounding error next to the messages.”

Google
A message arrives twice. How do you stop the user seeing it twice?

What to say: “Idempotency at the client: every message carries a unique ID assigned by the sender. The client keeps a set of recently seen IDs per conversation and drops duplicates before rendering. The server dedupes too, so retries don’t double-store. This is the standard answer to at-least-once delivery — you don’t prevent the duplicate on the wire, you make it harmless at both ends.”

Amazon
How would you implement “user is typing…”?

What to say: “It’s an ephemeral event, not a message — never stored, never ACKed, never retried. Keystrokes are noisy, so the client throttles: send ‘typing started’ at most every few seconds, and ‘typing stopped’ on send or after ~5 idle seconds. The server fans it out like a receipt but with a TTL — if the stop event is lost, the indicator just expires. Cheap, lossy, and nobody notices the loss.”

Startup
Design storage for ten years of chat history.

What to say: “Partition the messages table by conversation, ordered by sequence number, so a chat’s history is one contiguous range scan — pagination is just ‘give me 50 before seq N’. Hot conversations stay on fast storage; cold ones tier down to cheap object storage after, say, 90 days of silence. Ten years at 200 GB/day is ~730 TB raw, ~2 PB with triple replication — the tiering is what makes the bill survivable. Media was never in the database to begin with, so history stays text-cheap.”

WhatsApp
Why not just use HTTP polling every second?

What to say: “Do the math: a million users polling every second is a million requests a second of pure overhead — each carrying headers, TLS setup, and server work — for a chat where most polls return nothing. A WebSocket is one handshake and then silence until there’s actually a message. Polling burns battery, bandwidth, and servers to ask a question whose answer is usually ‘no’. The persistent pipe inverts the cost: you pay per message, not per second.”

Key takeaways

  1. A chat system has three jobs — hold millions of persistent connections, route each message to the right socket, and store everything. The interview is really about the second and third: the user-to-server map and the delivery promises.
  2. WebSockets beat polling because you pay per message, not per second — a million users polling every second is a million empty requests a second.
  3. At-least-once delivery plus idempotent dedup is the industry answer to “never lose a message” — don’t prevent duplicates on the wire, make them harmless at both ends.
  4. Partition message storage by conversation so history is one range scan; media lives in blob storage, never in the database.
  5. Group chats multiply load instead of adding it — one message, N deliveries. Push to online members, queue for offline ones, and respect the fan-out math when sizing the system.

Sources & further reading

Every claim in this guide traces to one of these — the wording is ours, the ideas are credited.

Back toAll guides →