Match the transport to the interaction
A stock dashboard receives updates; a multiplayer game sends and receives frequent events; a payment provider notifies a server. These are different workloads.
| Method | Direction | Strength | Cost or limitation |
|---|---|---|---|
| Short polling | Client asks repeatedly | Simple deployment | Empty requests and delayed discovery |
| Long polling | Server holds each request until update/timeout | Works over ordinary HTTP | Reconnect churn and held resources |
| SSE | Server to browser stream | Event IDs and reconnect support | Primarily one-way text events |
| WebSocket | Bidirectional connection | Interactive updates | Application-level reconnect and flow control |
| Webhook | Server to registered server | Decoupled external notification | Duplicate delivery, validation and retries |
SSE is often sufficient for generation progress or a notification feed. Its event format supports IDs and reconnect behavior; retention and replay remain application responsibilities.
Follow the transport lifecycle
With long polling, a client sends a request, the server waits for an update or timeout, and the client issues another request after the response. Every cycle repeats HTTP request handling. Bound the wait and use a cursor so events arriving between requests are not silently lost.
For the HTTP/1.1 WebSocket handshake, the client requests an upgrade and a successful server response uses status 101. Subsequent communication uses WebSocket frames over that connection. Other HTTP versions have different establishment mechanisms; do not hardcode the HTTP/1.1 explanation into every deployment. Use wss for TLS-protected communication.
A webhook instead starts with an agreed destination and event subscription. An event producer calls that destination; the receiver validates it and durably accepts work before responding. A socket's transport success and a webhook's accepted HTTP response are both different from confirmation that the business effect completed.
Connection establishment is not delivery correctness
For WebSockets, authenticate the handshake and validate origin in the relevant browser context. Recheck permissions as needed for long-lived sessions. Heartbeats detect a broken path, not certainty that the remote process is dead. Proxies need compatible idle timeouts and graceful connection draining.
Define a message envelope with event ID, channel, sequence number, schema version and payload. Ordering is scoped to a connection or channel; different connections can race. Reconnect using the last acknowledged sequence, replay retained events, then move to live delivery. If the cursor is too old, return a snapshot plus a new cursor.
Worked chat architecture
Client -> connection gateway -> conversation service -> durable messages
-> fan-out bus -> recipient gateway
Offline notification workers read a separate durable task queue
Persist a message before reporting the durability guarantee promised to the sender. Route live delivery using a connection directory. Store message history separately from presence; presence can expire and tolerate approximation. A device-delivered acknowledgment and a human-read receipt are distinct states.
Give a conversation a sequencing authority or another explicit ordering rule. Deduplicate resends using a stable client message ID scoped to the sender/conversation. A socket disconnect after persistence but before acknowledgment otherwise invites duplicate messages.
Backpressure and scale
At 100,000 connections and a hypothetical 20 KB of per-connection state, the connection layer alone needs about 2 GB before process overhead and buffers. Measure actual implementation costs rather than treating this as a universal constant.
Bound outbound buffers. A slow client should not consume unlimited memory; coalesce replaceable updates, drop explicitly disposable telemetry or disconnect and resume from a cursor. Reliable business events must remain recoverable from durable history. Partition channels and avoid broadcasting every event to every gateway.
Webhook receiver: acknowledge durable acceptance
- Read the raw body within a size limit and verify the provider's signature and timestamp policy.
- Apply an atomic unique event-ID insert into an inbox with the intended work payload.
- Acknowledge after durable acceptance; process asynchronously.
- Retry failed work with bounded backoff; quarantine poison events and expose a reconciliation path.
Do not return success before durable acceptance if the sender will stop retrying. Avoid parsing and reserializing JSON before signature verification when the signature covers raw bytes. Providers can deliver duplicates and out-of-order events, so handlers should not assume a perfect sequence.
Sender design and outbound safety
Persist the business change and notification intent in one transaction through an outbox. Each delivery attempt records destination, event ID, attempt count, status and next retry time. Sign the body with a rotating secret and expose delivery history and manual replay.
User-supplied destinations create SSRF risk. Validate the target, restrict internal/reserved network destinations, account for DNS resolution changes and redirects, bound response sizes/timeouts and isolate outbound workers. Rate-limit by destination so one failing endpoint does not stall other subscribers.
Interview exercise
Design notifications for job deadlines. Explain SSE vs WebSocket, authorization on channels, duplicate protection, lost-connection recovery, email fallback, quiet hours and slow-client behavior. For a webhook payment callback, explain why event arrival order cannot be your only source of truth; reconcile with authoritative payment state.
Next: message queues and transactions and CDC.