A chat system lets people send short messages to one person (1:1 chat) or to a group, and see them arrive almost instantly. WhatsApp, Telegram, Signal and Messenger are the familiar examples. Behind the simple interface sit several hard problems: keeping hundreds of millions of devices connected, delivering every message exactly once in the right order, reaching people who are offline, and showing who is online without flooding the servers.
It is one of the most commonly asked system design problems. Interviewers use it to check whether you understand long-lived connections, at-least-once delivery with deduplication, ordering, and the difference between data that must be durable (messages) and data that can be lossy (presence, typing indicators).
What interviewers typically probe:
- Why WebSockets rather than polling, and how a connection gateway layer works.
- How a message travels from sender to receiver, and how sent, delivered and read ticks are produced.
- How you order messages and generate message IDs.
- Storage choice (usually a wide-column store) and how you deliver to offline users.
- Group chat fan-out, presence, media, end-to-end encryption and multi-device sync.
A shorter version of this design appears in Real-world designs. This lesson goes much deeper. Useful background: Networking and edge delivery, Message queues and Databases at scale.
Problem and scope
Design a mobile-first messaging service supporting:
- 1:1 chat with text messages.
- Group chat with up to 1,024 members.
- Delivery states: sent (server received it), delivered (recipient device received it), read.
- Offline delivery: messages wait for the recipient and trigger a push notification.
- Online/last-seen presence and typing indicators.
- Images, videos and documents.
- Several devices per user (phone plus laptop) kept in sync.
- End-to-end encryption, at a conceptual level.
Out of scope: voice and video calls, payments, channels with millions of subscribers, and message search on the server (with end-to-end encryption the server cannot read messages anyway).
Clarifying questions
| Question | Why it matters | Assumption here |
|---|---|---|
| How many users? | Connections and throughput | 500 million daily active users (DAU) |
| Messages per user per day? | Write rate | 40 sent |
| Maximum group size? | Fan-out cost | 1,024 members |
| Does the server keep history forever, or only until delivered? | Storage size and design | Keep for 30 days for multi-device sync; clients keep long-term history |
| Is end-to-end encryption required? | Server cannot read content | Yes |
| How many devices per user? | Fan-out per user | Up to 5 (1 phone + 4 linked) |
| Must ordering be global or per chat? | Coordination cost | Per conversation only |
Interview tip
Ask whether the server stores messages permanently. A "store until delivered" model (keep messages only until every device has them) needs far less storage than a "cloud history" model. Stating the choice and its consequence early makes your storage estimates meaningful.
Functional and non-functional requirements
Functional:
- Send a message to a user or group; receive it in real time when online.
- Persist undelivered messages and deliver them when the device reconnects.
- Report sent, delivered and read states back to the sender.
- Show presence (online or last seen) and typing indicators.
- Send media files.
- Sync conversations across a user's devices.
Non-functional:
- Low latency: under about 200–500 ms end to end for online recipients in the same region.
- Reliability: no message lost after the sender sees "sent"; no duplicates shown.
- Ordering: messages in one conversation appear in the same order for everyone.
- High availability: chat must keep working through server and data-centre failures.
- Efficiency on mobile: minimise battery and data use.
- Privacy: end-to-end encryption; minimal metadata retained.
Back-of-the-envelope estimates
Message rate. 500M DAU × 40 messages = 20 billion messages per day. 20 × 10⁹ ÷ 86,400 ≈ 231,000 messages per second on average, about 463,000 per second at a 2× peak (festival greetings at midnight can be much higher, so plan headroom).
Ingress bandwidth for text. At about 200 bytes per message including metadata: 231,000 × 200 B ≈ 46 MB/s. Text bandwidth is small; the challenge is the number of operations and connections.
Storage.
- 20B × 200 B = 4 TB per day of message data, 12 TB/day with 3 replicas.
- With a 30-day retention for sync: 4 TB × 30 = 120 TB before replication.
- If you instead kept everything forever: 4 TB × 365 ≈ 1.46 PB per year before replication. That difference is why the retention question matters.
Concurrent connections. Suppose 40% of DAU are connected at peak: 200 million open connections. If one gateway server holds 100,000 connections comfortably, you need about 2,000 gateway servers, plus headroom for failures and deploys. The per-server number depends heavily on hardware, language runtime and message rate, so treat it as an assumption to benchmark.
Heartbeats. If each connection sends a heartbeat (a tiny "I am alive" packet) every 30 seconds: 200M ÷ 30 ≈ 6.7 million heartbeats per second. Writing each one to a database would be absurd; gateways must handle them in memory. This number drives the presence design.
Media. If 5% of messages carry media averaging 100 KB: 20B × 0.05 × 100 KB = 100 TB per day. Media dwarfs text, so it goes to object storage and a CDN (content delivery network), never through the chat path.
What the numbers tell you
Text throughput is high in operations but low in bytes. Connections (hundreds of millions) and heartbeats (millions per second) are the unusual scale factors. Media is the storage and bandwidth giant. Design three paths: a connection layer, a message path and a media path.
API design
Chat uses two kinds of interface:
A persistent WebSocket for real-time traffic. A WebSocket is a long-lived, two-way connection that starts as an HTTP request and then upgrades, so either side can send messages at any time without a new request.
Client -> server frames
{"type":"send", "client_msg_id":"c-7f3a", "conv_id":"g-55",
"ciphertext":"...", "media_ref":null}
{"type":"ack", "conv_id":"g-55", "up_to_seq":1043} # delivered
{"type":"read", "conv_id":"g-55", "up_to_seq":1043}
{"type":"typing", "conv_id":"g-55"}
{"type":"ping"}
Server -> client frames
{"type":"sent", "client_msg_id":"c-7f3a", "msg_id":"...", "seq":1044}
{"type":"message", "conv_id":"g-55", "seq":1044, "msg_id":"...",
"sender":"u-9", "ciphertext":"..."}
{"type":"receipt", "conv_id":"g-55", "user":"u-2",
"state":"read", "up_to_seq":1044}
{"type":"presence", "user":"u-2", "online":true}
{"type":"pong"}
Plain HTTPS APIs for everything that is not real time:
POST /v1/media/upload-url -> {"upload_url": "...", "media_id": "m-1"}
GET /v1/conversations?cursor=... -> list with last seq per conversation
GET /v1/conversations/g-55/messages?after_seq=1000&limit=100
POST /v1/groups -> create group
POST /v1/devices -> register device and push token
Notes:
client_msg_idis generated by the sender's device. If the device resends after a timeout, the server recognises the ID and returns the original result instead of storing a duplicate. This is idempotency for sends.- Acknowledgements are cumulative (
up_to_seq), which is cheaper than acknowledging every message individually. - The
after_seqsync API lets a reconnecting device fetch exactly what it missed.
Data model and storage choice
Message traffic is write-heavy, append-only and almost always read as "the latest N messages of one conversation". That fits a wide-column store such as Cassandra, ScyllaDB or HBase: data is grouped into partitions by a key, and rows inside a partition are kept sorted by a clustering key, so reading a range of a conversation is one sequential read.
-- CQL-style (Cassandra) schema
CREATE TABLE messages (
conv_id text,
bucket int, -- e.g. days since epoch, bounds partition size
seq bigint, -- per-conversation sequence number
msg_id bigint, -- globally unique, time-sortable
sender_id text,
ciphertext blob,
media_ref text,
created_at timestamp,
PRIMARY KEY ((conv_id, bucket), seq)
) WITH CLUSTERING ORDER BY (seq DESC)
AND default_time_to_live = 2592000; -- 30 days
CREATE TABLE inbox ( -- per-device pending deliveries
device_id text,
conv_id text,
seq bigint,
PRIMARY KEY (device_id, conv_id, seq)
);
Other data:
| Data | Store | Why |
|---|---|---|
| Messages | Wide-column, partitioned by (conv_id, bucket) | Huge append-only writes, range reads per chat |
| Conversation metadata, group membership | Relational DB or wide-column, cached | Small, read on every send |
| Per-conversation sequence counters | Strongly consistent store (or the conversation's owner shard) | Must not hand out the same number twice |
| Session registry (user/device → gateway) | In-memory store such as Redis, with TTL | Changes constantly, small, rebuilt on reconnect |
| Presence | In-memory, TTL keys | Lossy is fine |
| Media | Object storage + CDN | Large immutable blobs |
| Push tokens | KV store per device | Lookup by device |
The bucket column stops one very old, very busy group from growing a single partition forever: each day (or week) starts a new partition.
High-level design
Phone A Phone B
| WebSocket WebSocket |
+----v-----------+ +----------v----+
| Gateway 17 | | Gateway 902 |
| (holds sockets,| | |
| heartbeats) | | |
+----+-----------+ +----------^----+
| |
| send deliver to B |
+----v-------------------------------------------+----+
| Chat service |
| dedupe -> assign seq -> persist -> fan-out |
+--+------------+---------------+----------------+----+
| | | |
+--v-----+ +----v-------+ +-----v------+ +------v------+
|Message | |Seq counter | |Session | |Group service|
|store | |per conv | |registry | |(members) |
|(wide- | +------------+ |device -> | +-------------+
| column)| |gateway |
+--------+ +------------+
|
| offline devices
+--v--------------+ +---------------+
| Push service +----->| APNs / FCM |
+-----------------+ +---------------+
Media: client -> object storage (pre-signed URL) -> CDN -> client
Presence: gateways -> presence service (in memory, TTL)
- Gateways terminate WebSockets, authenticate devices, handle heartbeats and forward frames. They hold no durable state. A load balancer spreads new connections across them.
- Chat service is stateless business logic: deduplicate, assign a sequence number, persist, then route.
- Session registry maps each online device to the gateway holding its socket.
- Push service wakes offline devices through Apple Push Notification service (APNs) and Firebase Cloud Messaging (FCM).
Request flows
Flow 1: 1:1 message to an online user
- Alice's app encrypts the message for Bob's devices and sends a
sendframe withclient_msg_id = c-7f3aover her WebSocket to gateway 17. - Gateway 17 forwards it to the chat service.
- The chat service checks a short-lived dedup cache for
(alice, c-7f3a). New, so continue. - It obtains the next sequence number for the conversation, say 1044, and a unique
msg_id. - It writes the message to the message store and records pending deliveries for Bob's devices. Only after the write succeeds does it reply
sentto Alice. Alice sees one tick. - It looks up Bob's devices in the session registry: Bob's phone is on gateway 902.
- It forwards the message to gateway 902, which pushes it down Bob's socket.
- Bob's phone stores and displays it, then sends
ack up_to_seq 1044. - The chat service clears the pending delivery and sends Alice a
deliveredreceipt. Two ticks. - When Bob opens the chat, his phone sends
read up_to_seq 1044; Alice gets areadreceipt. Blue ticks.
Alice GW17 Chat svc Store GW902 Bob
|--send--->| | | | |
| |--send---->| | | |
| | |--write---->| | |
| | |<---ok------| | |
|<-sent----|<--sent----| | | |
| | |--deliver------------>|--msg--->|
| | |<-------------ack-----|<--ack---|
|<-deliv---|<-receipt--| | | |
| | |<-------------read----|<-read---|
|<-read----|<-receipt--| | | |
Flow 2: recipient is offline
- Steps 1–5 as before; the message is safely stored.
- The registry has no live session for Bob's phone.
- The chat service asks the push service to notify Bob's devices. With end-to-end encryption, the push payload is either the encrypted message itself or just "new message" and the app fetches the content when it wakes.
- When Bob's phone reconnects, it sends its last seen
seqper conversation (or the server reads Bob's inbox table). - The server streams everything after that point, in order. Bob acknowledges, and Alice receives delivered receipts, possibly hours after sending.
Flow 3: a group message
- Alice sends to group g-55 (200 members).
- The chat service assigns the group's next
seqand stores the message once in the group's partition. - It fetches the member list (cached) and, for each member's devices, either forwards to the online gateway or records a pending delivery and triggers push.
- Delivered and read receipts are aggregated: the sender sees "delivered" when all members have received it, and can open "message info" to see per-member states.
Deep dive 1: the connection layer
Why WebSockets
| Technique | How it works | Problem for chat |
|---|---|---|
| Short polling | Client asks "anything new?" every few seconds | Wasted requests, battery drain, delay up to the interval |
| Long polling | Server holds the request until there is news or a timeout | Better, but a new request per message and awkward for sending |
| Server-sent events | Server streams to client over HTTP | One-way only; client sends via separate requests |
| WebSocket | One persistent two-way connection | Best fit; needs stateful gateways |
Some apps use other persistent protocols (for example MQTT or a custom protocol over TCP), but the architecture is the same: a long-lived connection per device to a gateway.
Connection gateways
Gateways do one job: hold connections. Keeping them separate from business logic means you can deploy chat logic without disconnecting 200 million devices.
- Connecting: the device resolves the gateway hostname, a layer-4 load balancer picks a gateway, the device authenticates with a token, and the gateway writes
device → gateway 17into the session registry with a TTL. - Heartbeats: the device sends a ping every 30 seconds or so (mobile networks and NAT devices drop idle connections, so a ping keeps the path open). If the gateway hears nothing for, say, 2–3 intervals, it closes the connection and removes the registry entry. Heartbeats are handled in gateway memory; only state changes (connect, disconnect) go to the registry and presence service.
- Routing to a device: the chat service looks up the gateway in the registry and sends the frame there over an internal RPC, or publishes to that gateway's dedicated channel in a pub/sub system.
- Draining: during deploys, a gateway stops accepting new connections and asks existing clients to reconnect gradually, so you avoid a stampede of millions of reconnects at once.
Common mistake
Do not let the gateway write to the database on every heartbeat. At about 6.7 million heartbeats per second, that would be the biggest write load in the whole system for data nobody needs. Keep liveness in memory and publish only transitions.
Deep dive 2: message IDs, ordering and delivery guarantees
Two kinds of identifier
msg_id: globally unique and roughly time-sortable, used for deduplication and references (replies, reactions). A common design is a Snowflake-style 64-bit ID: 41 bits of milliseconds since a custom epoch, 10 bits of machine ID and 12 bits of per-millisecond sequence. 2⁴¹ milliseconds is about 69.7 years of IDs, and 12 bits allow 4,096 IDs per millisecond per machine (about 4 million per second) with no coordination.seq: a per-conversation counter (1, 2, 3, …) that defines the official order of messages in that conversation.
Why not just order by msg_id or by timestamp? Two senders' messages can be created on different machines with slightly different clocks, so timestamps can disagree with the order the server actually accepted them. A per-conversation counter assigned at one place gives a single, gap-detectable order that every member's device agrees on.
Who assigns seq?
All messages of one conversation are routed to the same owner (for example, by hashing conv_id to a chat-service partition or by using an atomic counter in a strongly consistent store). The owner increments the counter and writes the message. Ordering across different conversations is not needed, so this scales by spreading conversations across many owners.
Delivery guarantees
Networks fail mid-flight, so the system uses at-least-once delivery plus deduplication, which gives effectively-exactly-once display:
- The sender retries until it receives
sent; the server deduplicates onclient_msg_id. - The server redelivers until the device acknowledges; the device deduplicates on
seq(ormsg_id).
Worked example: gaps and duplicates on the client
Bob's phone has shown conversation g-55 up to seq 40. Messages then arrive out of order and with a duplicate after a reconnect:
| Arrives | Buffer before | Action | Shown |
|---|---|---|---|
| seq 41 "hi" | empty | 41 is next, show it | 41 |
| seq 43 "free?" | empty | 42 missing, hold 43 | nothing |
| seq 41 "hi" (resent) | 43 | 41 already shown, drop | nothing |
| seq 42 "are you" | 43 | show 42, then 43 | 42, 43 |
If 42 had not arrived within a few seconds, the phone would call GET .../messages?after_seq=41 to fill the gap. The runnable sketch:
class ConversationView:
"""Client-side buffer: shows messages in seq order, no gaps, no duplicates."""
def __init__(self, last_seq=0):
self.last_shown = last_seq # highest seq displayed contiguously
self.pending = {} # seq -> message, arrived early
def receive(self, seq, text):
if seq <= self.last_shown or seq in self.pending:
return [] # duplicate: already shown or buffered
self.pending[seq] = text
shown = []
while self.last_shown + 1 in self.pending:
self.last_shown += 1
shown.append((self.last_shown, self.pending.pop(self.last_shown)))
return shown
def missing(self):
"""Seqs to fetch from the server if the gap does not fill soon."""
if not self.pending:
return []
return [s for s in range(self.last_shown + 1, max(self.pending))
if s not in self.pending]
v = ConversationView(last_seq=40)
print(v.receive(41, "hi")) # [(41, 'hi')]
print(v.receive(43, "free?")) # [] -- 42 is missing, buffer 43
print(v.missing()) # [42]
print(v.receive(41, "hi")) # [] -- duplicate after a retry
print(v.receive(42, "are you")) # [(42, 'are you'), (43, 'free?')]
The sender's own messages
Alice's app shows her message immediately with a clock icon (optimistic UI), then moves it to its official position when sent returns with the seq. If two people type at the same moment, the server's seq decides the final order, and both screens converge on it.
Deep dive 3: group fan-out
When Alice posts in a group of n members, n − 1 other members (and all their devices) must receive it.
Fan-out on write (push): the server stores the message once and sends or queues a copy reference for each member device. For a group of 256 members, one message means 255 member deliveries; a busy group sending 100 messages an hour generates 25,500 deliveries per hour. Because groups are capped at 1,024 members, fan-out is bounded, and push is the right model.
Store once, deliver references. The message body (ciphertext) is written once to the group's partition. Per-device inbox entries store only (conv_id, seq), which are tiny. Offline members catch up by reading the group partition after their last seq.
Large broadcast channels (millions of subscribers) would flip to fan-out on read: subscribers pull new posts when they open the channel, exactly like the celebrity problem in Design a Twitter timeline. That is why such products usually treat channels as a separate feature.
Encryption and groups: with end-to-end encryption, the sender cannot encrypt one copy per member for every message cheaply at scale, so group protocols use a shared group key (or a sender key distributed to members once) so each message is encrypted once and fanned out as the same ciphertext.
Receipts in groups are aggregated to avoid a storm: devices send cumulative read up_to_seq, and the server sends the sender a summary rather than one event per member per message.
Deep dive 4: presence, push, media and multi-device
Presence
Presence is the "online" or "last seen at 10:42" label. It is useful but not critical, so it is designed to be cheap and lossy.
- Gateways report connect and disconnect transitions to a presence service, which keeps
presence:<user> = onlinewith a TTL a little longer than the heartbeat timeout. - Users do not get every contact's presence pushed all the time. Presence is sent when you open a chat or the contact list, and updates are pushed only to users currently viewing that person (subscribe on open, unsubscribe on close).
- Debounce: a phone moving between Wi-Fi and mobile data may disconnect and reconnect within seconds. Delay the "offline" transition by a few seconds so contacts do not see flicker.
- Typing indicators are sent directly through gateways to the other participants and never stored.
Push notifications
Mobile operating systems suspend background apps, so an offline or sleeping app cannot keep a socket open. The server sends a push notification through APNs or FCM, which the OS delivers. Practical points:
- Store push tokens per device and remove tokens the provider reports as invalid.
- Collapse multiple notifications for the same chat to avoid buzzing the phone 20 times.
- With end-to-end encryption, send either the ciphertext for the app to decrypt in a notification extension, or a content-free "new message" signal.
- Push providers are best-effort; the message store remains the source of truth, so a lost push only delays delivery until the app next connects.
Media
Large files never travel over the chat socket.
- The sender encrypts the file with a random key, asks for a pre-signed upload URL (a short-lived URL that grants permission to upload one object), and uploads directly to object storage.
- The chat message contains only
media_id, the decryption key (inside the end-to-end-encrypted payload) and a small thumbnail. - Recipients download through a CDN, which caches the encrypted blob near them; only holders of the key can decrypt it.
- Unreferenced media expires after the retention period.
See Design a file sync and storage service for chunked uploads and resumable transfers.
End-to-end encryption, conceptually
End-to-end encryption (E2EE) means only the communicating devices hold the keys; the server relays ciphertext it cannot read. Conceptually:
- Each device has a long-term identity key pair and uploads public "pre-keys" to a key server.
- To start a chat, the sender fetches the recipient device's public keys and performs a key agreement (Diffie-Hellman style) to derive a shared secret without ever sending it.
- Protocols such as the Signal protocol then change keys continuously (a "ratchet"), so compromising one key does not expose past messages (forward secrecy).
- Every recipient device is a separate encryption target. With 5 devices per user, a 1:1 message is encrypted for each of the recipient's devices and the sender's other devices.
Consequences for the design: the server cannot search, moderate or index message content; push payloads must be encrypted or empty; and multi-device sync must share keys or re-encrypt per device.
Multi-device sync
Each device is a first-class recipient with its own device_id, session, inbox entries and keys.
- A message from Alice's phone is delivered to Bob's devices and Alice's own laptop, so her laptop shows what she sent.
- Each device keeps its own
last_seqper conversation and syncs withafter_seqon reconnect. - Read state syncs too: when Bob reads on his laptop, a
read up_to_seqevent clears unread badges on his phone. - A newly linked device can receive recent history from the 30-day server store (or, in stricter E2EE designs, from the user's primary phone, which encrypts and transfers history to the new device).
Scaling and bottlenecks
| Bottleneck | Symptom | Fix |
|---|---|---|
| Connections | Gateways exhaust memory or file descriptors | More gateways, tune kernels, keep gateways stateless and lightweight |
| Reconnect storm after a gateway or region failure | Millions reconnect at once | Client exponential backoff with jitter; gradual draining |
| Hot group | One partition and one sequence owner overloaded | Cap group size, bucket partitions by time, batch receipts |
| Session registry load | Lookups on every delivery | Shard the registry, cache locally on chat servers with short TTL |
| Message store writes | ~460k writes/s at peak | Wide-column store scaled horizontally; partition by (conv_id, bucket) |
| Media bandwidth | 100 TB/day | Object storage + CDN; compress and resize on the client |
| Cross-region chats | Higher latency | Users connect to the nearest region; messages routed between regions over backbone links |
For geo-distribution, a common pattern is to home each conversation in one region (where its sequence counter lives) and let gateways in every region forward to it. See Multi-region design.
Failure handling
- Gateway crashes. Its connections drop; devices reconnect (with jittered backoff) to another gateway and sync from their
last_seq. Nothing is lost because messages were persisted beforesentwas returned. - Chat service crashes after storing but before replying
sent. The sender times out and retries with the sameclient_msg_id; the dedup check returns the existing message. No duplicate. - Delivery to a device fails. The pending inbox entry remains; the message is delivered on reconnect or via push.
- Session registry entry is stale (device moved gateways). Forwarding fails; the chat service treats the device as offline, keeps the inbox entry, and sends push. The device's reconnect fixes the registry.
- Message store replica fails. The wide-column store keeps serving with quorum reads and writes from remaining replicas.
- Push provider outage. Messages wait in the store; users get them when they next open the app.
- Region outage. Clients fail over to another region; conversations homed in the failed region either fail over (with their counters) or are briefly unavailable, depending on the chosen consistency.
See Reliability and recovery for failover patterns.
Trade-offs and alternatives
| Decision | Chosen | Alternative | Why |
|---|---|---|---|
| Transport | WebSocket via gateways | Long polling | Lower latency, less battery and overhead |
| Ordering | Per-conversation seq | Timestamps or global order | Agreed order without clock trust; global order is unnecessary and costly |
| Delivery | At-least-once + dedup | Exactly-once protocol | Simpler and robust to any failure |
| Storage | Wide-column, 30-day TTL | Relational DB; keep forever | Write-heavy append workload; retention limits cost and data exposure |
| Group fan-out | Write (push), store once | Read (pull) | Groups are bounded; pull suits huge channels |
| Presence | In-memory, lossy, subscribe-on-view | Persisted, broadcast to all contacts | Presence is high-volume and low-value per event |
| Media | Object storage + CDN, pre-signed URLs | Through chat servers | Keeps the chat path fast and cheap |
| Security | End-to-end encryption | Server-side encryption only | Server cannot read content; costs server-side search |
What interviewers probe
"How do you guarantee a message is not lost?" Persist before acknowledging sent; keep a pending delivery per device until it acknowledges; retry and use push for offline devices. Clients retry sends with a stable client_msg_id so duplicates are removed.
"How do you know which server Bob is connected to?" A session registry maps each device to its gateway, written on connect and expired by TTL. The chat service looks it up for each delivery.
"Two people send at the same instant; what order do they see?" Whatever order the conversation's sequence owner assigns. Both devices reorder by seq, so everyone converges on the same order.
"How would you implement last seen for 500 million users?" Track connect and disconnect transitions in memory with TTLs, store a last-seen timestamp only on disconnect, push updates only to users viewing that contact, and debounce flapping connections.
"How does the system change for groups of 100,000?" Switch to fan-out on read, rate-limit senders, aggregate receipts or drop them, and treat it as a broadcast channel feature.
"How do you sync a new laptop?" Register it as a new device with its own keys; deliver new messages to it; backfill recent history from the server's retention window or from the user's phone.
Interview questions
Q1. Why use WebSockets for chat instead of HTTP polling?
Polling sends requests even when there is nothing new, wasting battery and bandwidth, and adds delay up to the polling interval. A WebSocket keeps one two-way connection open so the server can push messages instantly and the client can send without a new handshake.
Q2. What does a connection gateway do, and why separate it from the chat service?
It holds device connections, authenticates them, handles heartbeats and forwards frames. Separating it lets you deploy and scale business logic without dropping hundreds of millions of connections, and keeps the stateful part as simple as possible.
Q3. How are sent, delivered and read states produced?
"Sent" is returned after the server has durably stored the message. "Delivered" comes when the recipient device acknowledges receipt. "Read" comes when the recipient opens the conversation and sends a cumulative read marker. In groups, these are aggregated per member.
Q4. How do you order messages within a conversation?
A single owner per conversation assigns a monotonically increasing sequence number when it accepts each message. Clients display messages by sequence number, buffer early arrivals and fetch missing numbers. Timestamps are not used for ordering because device and server clocks differ.
Q5. How do you prevent duplicate messages?
The sender attaches a client-generated ID and the server deduplicates retries on it. The server redelivers until acknowledged, and devices drop any seq they have already shown. At-least-once delivery plus deduplication gives an exactly-once experience.
Q6. Why a wide-column store for messages?
The workload is a very high rate of append-only writes, and reads are almost always "the latest messages of one conversation". Partitioning by conversation (and time bucket) with rows sorted by sequence makes those reads sequential and lets the store scale horizontally across many nodes.
Q7. How are messages delivered to offline users?
They are stored with per-device pending entries, and a push notification is sent through APNs or FCM. When the device reconnects, it reports its last sequence numbers and the server streams everything newer, in order.
Q8. How do you design presence at scale?
Keep it in memory with TTLs, update only on connect and disconnect (not on every heartbeat), deliver updates only to users currently viewing that contact, and debounce short disconnects. Losing a presence update is acceptable.
Q9. How does group messaging work?
Store each group message once in the group's partition with the group's next sequence number, then fan out deliveries to member devices: push to online ones and queue plus notify offline ones. Bounded group sizes keep fan-out on write affordable; huge channels use fan-out on read.
Q10. How is media handled?
The client encrypts and uploads the file directly to object storage using a pre-signed URL, then sends a small message containing the media reference and key. Recipients download the encrypted blob through a CDN. This keeps large data off the chat servers.
Q11. What does end-to-end encryption change about the server design?
The server only sees ciphertext, so it cannot search, filter or generate previews of content. It must store per-device public keys, fan out per-device ciphertexts or use group sender keys, and push notifications must carry encrypted or empty payloads.
Q12. How do you support multiple devices per user?
Treat each device as its own recipient with its own session, keys, inbox entries and last-seen sequence numbers. Deliver messages to all of the recipient's devices and to the sender's other devices, and sync read markers between them.
Q13. What happens when a gateway with 100,000 connections crashes?
All those devices reconnect, ideally with jittered exponential backoff to avoid a storm, land on other gateways, update the session registry and sync from their last sequence numbers. No message is lost because delivery state lives in the store, not the gateway.
Q14. How would you estimate the number of gateway servers?
Estimate peak concurrent connections (for example 40% of 500 million DAU, so 200 million) and divide by the benchmarked connections per server (for example 100,000), giving about 2,000, then add headroom for failures and deploys.
Key takeaways
- Separate three paths: a connection layer (WebSocket gateways), a durable message path, and a media path through object storage and a CDN.
- Persist before acknowledging "sent"; deliver at least once and deduplicate with client message IDs and sequence numbers.
- Order messages with a per-conversation sequence number assigned by one owner, not with timestamps.
- Store messages in a wide-column store partitioned by conversation and time bucket, sorted by sequence, with a retention TTL.
- Offline devices get push notifications and catch up from their last sequence number on reconnect.
- Presence and typing are lossy, in-memory and subscription-based; never write heartbeats to a database.
- Bounded groups use fan-out on write with one stored copy; huge channels need fan-out on read.
- End-to-end encryption and multi-device support make every device a separate recipient and keep content opaque to the server.
Next lesson
Continue with Design a video streaming service.

