Problem and scope
A notification system is the shared service that every other team in a company uses to tell users that something happened. Your food order is out for delivery, an OTP has arrived by SMS, someone replied to your comment, your bank statement is ready by email, or a sale starts tonight. The order service, the payments service and the marketing team should not each build their own way to send messages. They call one notification service, which decides whether, when, how and through which channel each message reaches the user.
A channel is a delivery route. The four common ones are:
- Push notifications: the banner on a phone's lock screen. They go through Apple Push Notification service (APNs) for iPhones and Firebase Cloud Messaging (FCM) for Android. You never talk to the phone directly; you hand the message to Apple or Google, who deliver it.
- SMS: text messages, sent through an SMS gateway provider that connects to mobile operators. Each message costs money.
- Email: sent through an email provider (an email service provider, or ESP) such as Amazon SES, SendGrid or Resend.
- In-app: the bell icon with a red badge inside the app or website. This one you store and serve yourself.
It is a popular interview problem because the core idea is simple ("put a message on a queue, a worker sends it") but real systems fail in the details: duplicate OTPs, a marketing blast delaying a password reset, a 3 a.m. promotional buzz, or a provider outage that silently drops messages. Interviewers usually probe how you prioritise urgent messages over bulk ones, how you retry without creating duplicates, how you respect user preferences and rate limits, and how you track whether anything was actually delivered.
Clarifying questions
| Question | Assumed answer |
|---|---|
| Which channels? | Push (iOS and Android), SMS, email and in-app |
| Who sends? | Internal services only (orders, payments, social, marketing), not external customers |
| Kinds of messages? | Transactional (OTP, order updates, security alerts) and promotional (offers, digests) |
| Scale? | About 1 billion notifications per day, with large marketing campaigns |
| Real-time requirement? | OTPs and security alerts within seconds; promotions may wait minutes or hours |
| Do users control what they receive? | Yes: per-category, per-channel opt-outs, plus quiet hours |
| Must we guarantee delivery? | We guarantee we hand it to the provider at least once, and track the outcome; the final hop (Apple, Google, the operator) is outside our control |
| Do we need analytics? | Yes: sent, delivered, opened and clicked rates per campaign and template |
| Scheduled sends? | Yes, for example "send at 9 a.m. in each user's local time" |
Two answers shape the whole design. First, transactional and promotional traffic must never share a single queue, because a 50-million-user campaign would bury an OTP. Second, we cannot promise exactly-once delivery to a phone; the realistic promise is "at least once to the provider, deduplicated as well as we can, and tracked".
Functional and non-functional requirements
Functional:
- Internal services send a notification by naming a user (or a list of users), a template and some data, for example
order_shippedwith the order id and expected date. - The system picks channels based on the notification type and the user's preferences.
- Templates are rendered per channel and per language: a short push text, a longer email with HTML, a 160-character-friendly SMS.
- Users can opt out by category and channel, and set quiet hours (a nightly window in their local time when non-urgent notifications are held back).
- Scheduled and bulk sends: "send this to everyone in segment X at 9 a.m. local time".
- In-app inbox: list notifications, mark as read, show an unread count.
- Delivery tracking: each notification's status per channel, and aggregate analytics.
Non-functional:
- Low latency for urgent messages: an OTP should leave our system within a second or two.
- High throughput for bulk: tens of thousands of sends per second during campaigns.
- Reliability: no silent loss. A message we accepted must either be delivered to the provider or end up in a place where someone can see it failed.
- No duplicates (as far as possible): a user should not get the same OTP or order update twice.
- Respect limits: per-user caps (do not spam) and per-provider caps (Apple, Google, SMS and email vendors throttle you).
- Availability: sending services should be able to submit even if a provider is down; the backlog drains later.
Transactional versus promotional
A transactional notification is triggered by something the user did or needs to know now: an OTP, a payment receipt, a login from a new device. A promotional (or marketing) notification is something the company wants to tell many users: an offer or a weekly digest. They differ in urgency, legal rules (many countries require an unsubscribe option and consent for marketing), and how much you can delay or drop them. Most design decisions in this lesson split along this line.
Back-of-the-envelope estimates
Assumptions:
- 1 billion notifications per day across all channels.
- Channel mix: 600 million push, 300 million in-app, 80 million email, 20 million SMS.
- Peak traffic is about 5 times the daily average.
- A big campaign targets 50 million users and the marketing team wants it out within 30 minutes.
- A stored notification record (ids, template, rendered text, status fields) is about 500 bytes.
- Each notification produces about 3 tracking events (sent, delivered, opened), about 100 bytes each.
Average and peak rate. 1,000,000,000 ÷ 86,400 seconds ≈ 11,600 notifications per second on average. At 5 times that, peak is about 58,000 per second.
Campaign rate. 50,000,000 users ÷ 1,800 seconds ≈ 27,800 sends per second for that one campaign. So a single campaign can roughly match half our peak load. That alone justifies separate queues and per-provider throttling.
Storage for notifications. 1,000,000,000 × 500 bytes = 500 GB per day. If we keep 90 days for the in-app inbox and support lookups, that is 500 GB × 90 = 45 TB. That is a job for a horizontally partitioned store, not a single database server.
Tracking events. 1,000,000,000 × 3 × 100 bytes = 300 GB per day of events, append-only, perfect for a log (Kafka) feeding a columnar analytics store.
Provider limits change the picture. Suppose our email provider account allows 1,000 sends per second (an illustrative figure; real limits depend on the vendor and your contract). A 10-million-user email campaign then takes 10,000,000 ÷ 1,000 = 10,000 seconds, about 2.8 hours, no matter how many workers we run. If password-reset emails sit behind that campaign in the same queue, users wait hours to log in. Provider limits, not our servers, are often the real bottleneck.
Deduplication memory. If we remember an idempotency key (a short hash, about 64 bytes with overhead) for every notification for 24 hours: 1,000,000,000 × 64 bytes ≈ 64 GB. That fits in a modest Redis cluster.
Interview tip
State the conclusion, not just the numbers: "Our own throughput, about 58,000 per second at peak, is manageable with partitioned queues and stateless workers. The hard limits are external: provider rate limits and per-user caps. So the design is mostly about queues, priorities and throttling in front of the providers."
API design
Internal services call the notification service; users call it only for their inbox and preferences.
# Internal: send to one user (transactional)
POST /v1/notifications
headers: Idempotency-Key: order-881-shipped
body: {
"userId": "u_42",
"type": "order_shipped", # maps to category + template
"priority": "high", # critical | high | normal | low
"data": { "orderId": "881", "eta": "2026-10-13" },
"channels": ["push", "in_app"], # optional override
"sendAt": null # or an ISO time for scheduling
}
202 -> { "notificationId": "n_9f3", "status": "ACCEPTED" }
# Internal: bulk / campaign
POST /v1/campaigns
body: { "segmentId": "seg_diwali_buyers", "type": "promo_sale",
"data": {...}, "sendAtLocal": "09:00", "rateLimitPerSec": 20000 }
202 -> { "campaignId": "c_12", "status": "SCHEDULED" }
# Status
GET /v1/notifications/n_9f3
200 -> { "status": { "push": "DELIVERED", "in_app": "READ" } }
# User-facing
GET /v1/me/inbox?cursor=...&limit=20
POST /v1/me/inbox/n_9f3/read
GET /v1/me/preferences
PUT /v1/me/preferences
body: { "promo_sale": { "push": false, "email": true },
"quietHours": { "start": "22:00", "end": "07:00",
"tz": "Asia/Kolkata" } }
# Devices register for push
POST /v1/me/devices
body: { "platform": "ios", "token": "a1b2...", "appVersion": "8.4.0" }
Points worth saying out loud:
202 Accepted, not200 OK. The request means "we stored it and will send it", not "it has arrived". Sending is asynchronous.Idempotency-Keyis mandatory for internal callers. If the order service retries after a timeout, the same key returns the samenotificationIdand nothing is sent twice. The key should be derived from the business event (order-881-shipped), not random, so that even two different instances of the caller produce the same key.typemaps to a category and template on the server. Callers should not send raw text; they send data and a type. That keeps wording, translations and legal footers in one place and lets preferences work by category.- Device tokens. A device token is an opaque string that APNs or FCM gives the app; it identifies one app install on one device. The app sends it to us on login and whenever it changes.
Data model and storage choice
| Data | Store | Why |
|---|---|---|
| Templates and notification types | Relational database, cached in memory | Small, changes rarely, needs versioning and review |
| User preferences and quiet hours | Key-value or relational, cached | Read on every send; tiny per user |
| Device tokens | Key-value store keyed by user id | A user has a few devices; read on every push |
| Notifications and per-channel status | Wide-column or partitioned store (Cassandra, DynamoDB) keyed by user id | 45 TB, write-heavy, read by user for the inbox |
| Idempotency keys | Redis with a time-to-live | Fast check-and-set, expires automatically |
| Rate-limit counters | Redis | Atomic increments, short-lived |
| Queues | Kafka or a managed queue (SQS, RabbitMQ) | Buffering, priorities, retries |
| Tracking events | Kafka, then a columnar store (ClickHouse, BigQuery) | Append-only, aggregated for dashboards |
A sketch of the main tables:
CREATE TABLE notification_types (
type VARCHAR(64) PRIMARY KEY, -- 'order_shipped'
category VARCHAR(32) NOT NULL, -- 'orders', 'security', 'promo'
kind VARCHAR(16) NOT NULL, -- 'transactional' | 'promotional'
default_prio VARCHAR(8) NOT NULL,
channels VARCHAR(64) NOT NULL -- 'push,in_app,email'
);
CREATE TABLE templates (
type VARCHAR(64) NOT NULL,
channel VARCHAR(8) NOT NULL, -- 'push','sms','email','in_app'
locale VARCHAR(8) NOT NULL, -- 'en-IN', 'hi-IN'
version INT NOT NULL,
subject TEXT,
body TEXT NOT NULL, -- 'Your order {{orderId}} ...'
PRIMARY KEY (type, channel, locale, version)
);
CREATE TABLE preferences (
user_id BIGINT NOT NULL,
category VARCHAR(32) NOT NULL,
channel VARCHAR(8) NOT NULL,
enabled BOOLEAN NOT NULL,
PRIMARY KEY (user_id, category, channel)
);
And the notification record in a wide-column store, partitioned by user so the inbox is one partition read:
partition key: user_id
clustering key: created_at DESC, notification_id
{ user_id, created_at, notification_id, type, priority,
rendered_title, rendered_body, deep_link,
status_push, status_sms, status_email, status_in_app,
read_at, expires_at }
Templates use placeholders such as {{orderId}}. They are versioned: editing a template creates a new version, so you can tell which wording a user actually received, and roll back a bad edit.
High-level design
+-----------+ +-----------+ +-----------+
| Orders | | Payments | | Marketing |
| service | | service | | campaigns |
+-----+-----+ +-----+-----+ +-----+-----+
| | |
v v v
+--------------------------------------------+
| Notification API (validate, idempotency, |
| store record, enqueue) |
+----------------------+---------------------+
v
+--------------------------------------------+
| Intake queues by priority |
| [critical] [high] [normal] [low/bulk] |
+----------------------+---------------------+
v
+--------------------------------------------+
| Processor workers |
| - load prefs, quiet hours, devices |
| - per-user rate limit, dedup |
| - render template per channel/locale |
+----+---------+---------+---------+---------+
v v v v
[push q] [sms q] [email q] [in-app q] <- channel queues
v v v v
+--------+ +--------+ +--------+ +--------+
| Push | | SMS | | Email | | In-app |
| sender | | sender | | sender | | writer |
+---+----+ +---+----+ +---+----+ +---+----+
v v v v
APNs/FCM SMS gw ESP inbox store
| | | + websocket
+----------+----------+
v delivery callbacks
+--------------------------------------------+
| Tracking: status updates + Kafka events |
| --> analytics store, dashboards |
+--------------------------------------------+
Side paths: retry queues per channel, dead-letter
queues, scheduler for delayed sends, Redis for
dedup keys and rate-limit counters.
Main components:
- Notification API: checks the caller's permission and the idempotency key, writes the record and enqueues it by priority. It does no slow work.
- Intake queues by priority: separate queues (or Kafka topics) so urgent traffic never waits behind bulk.
- Processor workers: the "brain". For each notification they decide the final channel list after preferences, apply quiet hours and per-user limits, render the templates and fan out one message per channel.
- Channel queues and senders: one queue and one pool of sender workers per channel (often per provider). Senders know the provider's API, its rate limit and its error codes.
- Scheduler: holds notifications whose send time is in the future (scheduled sends and quiet-hour delays) and releases them when due.
- Tracking service: receives provider callbacks (called webhooks: the provider makes an HTTP request to us when something happens, such as "email bounced") and client events ("opened"), updates the status and emits analytics events.
- In-app writer and real-time gateway: stores the inbox entry and, if the user is online, pushes it over a WebSocket so the bell updates instantly.
Request flows step by step
Flow 1: an OTP by SMS (critical)
- The auth service calls
POST /v1/notificationswith typelogin_otp, prioritycritical, and idempotency keyotp-u_42-attempt-3. - The API runs
SET key NX EX 600in Redis (set only if it does not exist, expire in 10 minutes). If the key already existed, it returns the original notification id and stops. - It writes the record and enqueues it on the critical queue.
- A processor reads it. Security and OTP types ignore quiet hours and promotional caps, but still check a per-user OTP limit (for example, at most 5 per 10 minutes, to stop abuse and SMS cost attacks).
- It renders the SMS template in the user's language and puts it on the SMS queue with high priority.
- The SMS sender calls the gateway. On success it records
SENTwith the provider's message id. - Later, the gateway's webhook reports
DELIVERED(orFAILED), and the tracking service updates the status.
The whole path from step 1 to step 6 should take well under a second when nothing is backed up.
Flow 2: an order update (push and in-app)
- The order service sends
order_shippedfor useru_42with priorityhigh. - The processor loads
u_42's preferences: categoryordersis enabled for push and in-app. It loads two device tokens: an iPhone and an Android tablet. - The in-app writer stores the inbox entry and, because the user has a live WebSocket, pushes it immediately so the badge updates.
- The push sender sends one message to APNs and one to FCM.
- APNs answers that the iPhone token is no longer valid (the app was uninstalled). The sender deletes that token so we stop wasting calls on it.
Flow 3: a campaign to 50 million users
- Marketing creates a campaign for a segment at 9 a.m. local time.
- A campaign service resolves the segment in batches (for example 10,000 user ids at a time, using a cursor) instead of loading 50 million ids into memory.
- For each batch it computes each user's send time in their time zone and hands the notifications to the scheduler.
- When each time zone reaches 9 a.m., the scheduler releases that slice onto the low priority queue, at the campaign's configured rate.
- Processors drop users who opted out of promotions or already hit their promotional cap for the day, then render and fan out.
- Senders throttle to provider limits. The campaign drains over minutes or hours while critical traffic flows past it on its own queue.
Deep dive 1: priority queues and isolation
The single most important design choice: do not let bulk traffic delay urgent traffic. If everything shares one FIFO (first in, first out) queue, a campaign that enqueues 50 million messages at 9:00 means an OTP enqueued at 9:01 waits behind all of them.
There are three common ways to give priority:
- Separate queues per priority, with dedicated consumers.
critical,high,normalandloweach have their own queue and their own worker pool. Critical workers never touch bulk work, so a flood of bulk cannot starve them. This is the simplest and most robust option. - Weighted consumption. One worker pool reads from all queues but takes, say, 8 messages from critical for every 1 from low when both have work. Better utilisation, but a bug in weighting can still hurt urgent traffic.
- A true priority queue (RabbitMQ supports per-message priorities). Convenient, but Kafka has no per-message priority at all, so with Kafka you use separate topics.
+-------------------+ +-------------------+
critical -->| queue (small) |---->| workers: 20 |
+-------------------+ +-------------------+
high -->| queue |---->| workers: 40 |
+-------------------+ +-------------------+
normal -->| queue |---->| workers: 60 |
+-------------------+ +-------------------+
low/bulk -->| queue (huge) |---->| workers: 80, |
+-------------------+ | rate-capped |
+-------------------+
Isolation must continue past the processor. The SMS queue and the email queue also need priority lanes, and the per-provider rate limit (deep dive 2) should reserve part of the provider's capacity for critical traffic. For example, if the SMS gateway allows 500 messages per second, reserve 100 per second for OTPs and let promotions use at most 400. Many teams also use a separate provider account for transactional email, so spam complaints about marketing cannot hurt the sender reputation of password-reset mail.
Common mistake
Giving priority at the first queue but sending everything through one shared channel queue. The OTP skips the line at intake and then waits behind a million promotional SMS at the sender. Priority has to exist at every queue on the path, including the provider's rate budget.
Deep dive 2: preferences, quiet hours and rate limits
Before sending anything, the processor runs a short pipeline of checks. Order matters: cheap checks first, and checks that drop a message before checks that delay it.
notification
|
v
1. type allowed for caller? no -> reject
2. user exists, not blocked? no -> drop
3. channel enabled in preferences? no -> drop that channel
4. device token / email / phone known? no -> drop that channel
5. per-user rate limit ok? no -> drop (promo) or delay
6. inside quiet hours (non-urgent)? yes -> schedule for later
7. dedup key already used? yes -> drop
|
v
render and send
Preferences
Preferences are stored per category and channel, not per template: "orders by push: on; promotions by SMS: off". Users understand categories, and a new template in an existing category automatically follows the user's choice. Some categories are mandatory: security alerts and OTPs cannot be switched off, though the user may choose the channel.
Preferences are read for every notification, roughly 58,000 times per second at peak, so they are cached with a short time-to-live or invalidated when the user changes them. A few seconds of staleness is tolerable; days of it is a legal problem for marketing consent in many countries.
Quiet hours
Quiet hours are a window, such as 22:00 to 07:00 in the user's local time, during which non-urgent notifications are held back and delivered when the window ends. Two details trip people up:
- The window often crosses midnight, so "is 23:30 inside 22:00–07:00" cannot be a simple
start <= t < endcheck. - It is in the user's time zone, not the server's.
Here is a small, runnable function that returns when a notification may be sent:
from datetime import datetime, time, timedelta, timezone
from zoneinfo import ZoneInfo
def deliver_at(now_utc, tz_name, quiet_start, quiet_end):
"""Return the UTC time at which a non-urgent notification may be sent."""
local = now_utc.astimezone(ZoneInfo(tz_name))
t = local.time()
if quiet_start > quiet_end: # window crosses midnight
in_quiet = t >= quiet_start or t < quiet_end
else:
in_quiet = quiet_start <= t < quiet_end
if not in_quiet:
return now_utc
end = local.replace(hour=quiet_end.hour, minute=quiet_end.minute,
second=0, microsecond=0)
if end <= local:
end += timedelta(days=1)
return end.astimezone(timezone.utc)
now = datetime(2026, 10, 11, 17, 0, tzinfo=timezone.utc)
start, end = time(22, 0), time(7, 0)
for tz in ["Asia/Kolkata", "Europe/London", "America/New_York"]:
local = now.astimezone(ZoneInfo(tz)).strftime("%H:%M")
print(tz, local, "->", deliver_at(now, tz, start, end).isoformat())
Output:
Asia/Kolkata 22:30 -> 2026-10-12T01:30:00+00:00
Europe/London 18:00 -> 2026-10-11T17:00:00+00:00
America/New_York 13:00 -> 2026-10-11T17:00:00+00:00
Walking through it. At 17:00 UTC it is 22:30 in India (UTC+5:30), which is inside the quiet window. The window ends at 07:00 the next morning, India time, which is 01:30 UTC on 12 October, so the notification is scheduled for then. In London (on British Summer Time, UTC+1, on that date) it is 18:00, and in New York (UTC−4) it is 13:00, both outside the window, so they go now.
Notifications held overnight should not all fire at 07:00: collapse them into one summary ("You have 6 new updates") and spread the release over a few minutes.
Per-user rate limits
A per-user rate limit caps how many notifications one user receives in a period, for example "at most 3 promotional pushes per day" or "at most 5 OTP SMS per 10 minutes". It protects users from spam and protects you from SMS cost attacks, where an attacker repeatedly triggers OTPs to someone's number.
A simple, effective implementation is a fixed-window counter in Redis:
key = "rl:u_42:promo_push:2026-10-11"
count = INCR key
if count == 1: EXPIRE key 86400
if count > 3: drop (promotional) or delay (non-urgent)
INCR is atomic, so two processors handling the same user at once cannot both see "2" and both send. For smoother limits, a sliding window or token bucket works too (see building blocks explained for the general algorithms). What to do when the limit is hit depends on the type: promotions are dropped, transactional updates are usually never limited except for abuse-prone ones like OTPs, and the limit check for those returns an error to the caller ("too many OTP requests, try later").
Per-provider rate limits
Providers throttle you. APNs and FCM accept very high volumes but will reject or slow you if you misbehave; SMS gateways and email providers typically give each account a fixed number of sends per second. If you exceed it, you get errors (HTTP 429 Too Many Requests is common), and repeatedly ignoring them can get the account suspended.
So each sender pool enforces a global rate per provider account. Because senders run on many machines, the limit must be shared: a token bucket in Redis that all senders draw from, or a fixed share of the limit assigned to each sender instance (simpler, slightly less efficient). A token bucket holds up to B tokens and refills at R tokens per second; each send takes one token; with no tokens left you wait. It allows short bursts up to B while keeping the average at R.
Interview tip
Name both limits explicitly: "Per-user limits protect the user and are enforced in the processor with Redis counters. Per-provider limits protect our provider accounts and are enforced in the senders with a shared token bucket, with a reserved slice for critical traffic."
Deep dive 3: retries, backoff and dead-letter queues
Provider calls fail all the time: timeouts, 5xx errors, rate limiting, network blips. A sender must classify every failure:
| Failure | Example | Action |
|---|---|---|
| Transient | Timeout, HTTP 500/503, connection reset | Retry with backoff |
| Throttled | HTTP 429 | Retry after the time the provider asks for, slow the bucket |
| Permanent for this recipient | Invalid device token, unregistered app, hard email bounce, invalid phone number | Do not retry; mark the address or token invalid |
| Permanent for this message | Template rendering error, payload too large | Do not retry; send to dead-letter queue and alert |
Exponential backoff with jitter
Exponential backoff means waiting longer after each failed attempt: 1 second, then 2, then 4, then 8, doubling up to a cap. It stops you from hammering a struggling provider. Jitter means picking a random delay between zero and that ceiling, so thousands of failed messages do not all retry at the same instant and cause a second spike (a "thundering herd").
import random
def backoff_delays(attempts, base=1.0, factor=2.0, cap=300.0):
"""Delay before each retry: exponential, capped, with full jitter."""
for n in range(attempts):
ceiling = min(cap, base * factor ** n)
yield ceiling, random.uniform(0, ceiling)
random.seed(7)
total = 0
for n, (ceiling, delay) in enumerate(backoff_delays(8), start=1):
total += ceiling
print(f"retry {n}: wait up to {ceiling:>5.0f}s, chosen {delay:6.1f}s")
print("worst-case total wait:", total, "s")
Worked example. With base 1 second, factor 2, cap 300 seconds and 8 retries, the ceilings are 1, 2, 4, 8, 16, 32, 64 and 128 seconds. The cap is not reached. The worst-case total wait is 1 + 2 + 4 + 8 + 16 + 32 + 64 + 128 = 255 seconds, a little over 4 minutes. With full jitter the actual waits are random values below each ceiling, so the real total is usually much lower.
Is 4 minutes right? For promotions, you could retry for hours. For an OTP, it is pointless after a minute or two: the code expires in 5 or 10 minutes and the user has probably asked for another. So retry policy belongs to the notification type: OTPs get a few fast retries and possibly a fallback channel (if SMS fails, try a voice call or a messaging app), while digests get long, patient retries.
Where do retries wait?
A sender should not sleep for 2 minutes holding a worker. Instead it publishes the message to a retry queue with a "not before" time and moves on. Common shapes:
- A few delay queues (for example 10 seconds, 1 minute, 10 minutes). After a failure, the message goes to the next one; a small mover re-publishes due messages to the main channel queue.
- A managed queue with per-message delay or visibility timeout (SQS allows a delivery delay of up to 15 minutes).
- The scheduler service itself, which already handles "send later".
Each message carries an attempt counter. After the last allowed attempt, it goes to a dead-letter queue (DLQ): a separate queue that holds messages the system gave up on. Nothing in a DLQ is lost or forgotten; it is visible on a dashboard, alerts fire when it grows, and an engineer can inspect, fix the cause (a broken template, an expired provider key) and replay the messages.
channel queue --> sender --ok--> provider
|
| transient failure
v
retry queue (delay = backoff(attempt))
|
| due
+--> back to channel queue
|
| attempt > max, or permanent error
v
dead-letter queue --> alert, inspect, replay
Common mistake
Retrying permanent errors. If APNs says a device token is unregistered, retrying eight times wastes capacity and can look abusive to the provider. Classify errors, retry only transient ones, and clean up invalid tokens, bounced email addresses and invalid numbers so you stop sending to them at all.
Provider failover
If a whole SMS or email provider is down, retries alone will just fill the retry queue. A circuit breaker watches the error rate per provider; when it crosses a threshold, it "opens" and stops sending to that provider for a short time, then lets a few test requests through. Many companies keep a secondary provider for SMS and email; when the breaker opens, senders route to the secondary. Push has no alternative provider (only Apple can reach iPhones), so for push you wait and retry.
Deep dive 4: deduplication and idempotency
Duplicates come from three places:
- The caller retries. The order service times out and calls us again.
- Our queue redelivers. Most queues give at-least-once delivery: if a worker crashes after sending but before acknowledging the message, the queue hands it to another worker, which sends again.
- The caller emits the same event twice. A bug or a replayed event stream fires
order_shippedtwice.
Idempotent means "doing it twice has the same effect as doing it once". We make each stage idempotent:
- At the API: the
Idempotency-Keyfrom the caller, stored in Redis withSET key NX EX ttl.NXmeans "only if not already present", so exactly one request wins; the others look up and return the same notification id. Deriving the key from the business event (order-881-shipped) catches case 3 as well. - Per channel send: before calling the provider, the sender checks and sets a key such as
sent:n_9f3:push:device_7. If a redelivered message finds the key already set, it skips the provider call. - At the provider, where supported: some providers accept an idempotency key on their API, and APNs and FCM support a collapse key (APNs calls it
apns-collapse-id): a later notification with the same key replaces an earlier one on the device instead of adding a second banner. That turns a duplicate into a harmless overwrite.
There is an unavoidable gap. If the sender calls the provider, the provider sends the message, and then the sender crashes before writing sent:..., the redelivered message will send again. No design can close this window completely without the provider's cooperation, because the "send" and the "remember that I sent" happen in two different systems. This is why you say "at-least-once, with deduplication that makes duplicates rare", and use collapse keys where the channel supports them.
Interview tip
If asked for exactly-once delivery, explain the gap honestly: "We can make our side idempotent at every hop, but the phone is reached through Apple, Google or an operator. A crash between their acceptance and our bookkeeping can cause a duplicate. We reduce it with per-hop dedup keys and collapse keys, and design messages so a rare duplicate is harmless."
How long to keep dedup keys? Long enough to cover realistic retries: a caller retry happens within minutes, a queue redelivery within the visibility timeout. 24 hours is a common, safe choice, and our estimate showed about 64 GB of keys at that window, which a small Redis cluster handles.
Delivery tracking and analytics
Each notification moves through statuses per channel:
ACCEPTED -> QUEUED -> SENT -> DELIVERED -> OPENED -> CLICKED
\ \
\ +-> FAILED (bounced, invalid token)
+-> SUPPRESSED (opted out, rate-limited,
or held and expired)
Where each status comes from:
- SENT: our sender got a success response from the provider.
- DELIVERED: SMS gateways and email providers report delivery through webhooks. For push, APNs and FCM confirm they accepted the message, not that it was displayed; to know about display you send a small event from your app when the notification is received (possible on Android, and on iOS with a notification service extension).
- OPENED / CLICKED: the app reports a tap on a push; email opens are measured with a tiny tracking image (unreliable, since many mail clients block or pre-load images) and clicks through redirect links.
- SUPPRESSED: we deliberately did not send. Recording why (opted out, cap reached) is essential for debugging "why didn't I get my notification?" tickets.
Every status change becomes an event on Kafka. One consumer updates the notification record (so GET /v1/notifications/n_9f3 and the support tools show the latest state). Another loads events into a columnar analytics store, where dashboards compute delivery rate, open rate and click-through rate by template, campaign, channel and day.
Webhook safety. Verify provider webhook signatures, and make the handler idempotent: providers retry webhooks and may deliver them out of order. A DELIVERED arriving after OPENED must not move the status backwards.
In-app notifications
The in-app channel is the only one you fully own:
- The in-app writer inserts a row into the user's inbox partition.
- It increments the user's unread counter (a small counter in Redis or a column updated alongside).
- If the user is connected to the real-time gateway over a WebSocket, it publishes the notification to that connection so the bell updates instantly. If not, the app fetches the inbox the next time it opens.
The inbox read is GET /v1/me/inbox?cursor=..., served from the partition keyed by user_id and ordered by created_at DESC, using cursor pagination (the cursor encodes the last item's timestamp and id), so it stays fast however long the history. Old entries expire with a time-to-live (here 90 days) instead of a bulk delete job.
Scaling and bottlenecks
| Component | Pressure | Approach |
|---|---|---|
| Notification API | 58,000 requests/s peak | Stateless, behind a load balancer; does only validation, dedup check, write and enqueue |
| Intake queues | Bursty campaigns | Kafka topics per priority, partitioned by user id so one user's notifications stay ordered |
| Processors | Preference and device lookups per message | Cache preferences and tokens; batch lookups; scale horizontally |
| Senders | Provider limits, slow provider APIs | Many concurrent requests per sender (HTTP/2 for APNs), shared token bucket, reserved critical capacity |
| Notification store | 500 GB/day, 45 TB retained | Partition by user id; TTL-based expiry |
| Redis (dedup, counters) | Hundreds of thousands of ops/s | Cluster mode, keys spread by user id |
| Campaign fan-out | 50 million recipients | Batch segment reads, schedule by time zone, rate-capped release |
| Tracking | 300 GB/day of events | Kafka, batched loads into a columnar store |
Two details that come up:
- Ordering. Partitioning queues by
user_idkeeps one user's notifications in order within a priority, but not across priorities. If order matters, include a sequence number so the app can ignore stale updates. - Hot users. A celebrity account that triggers "X posted" to 10 million followers is a campaign in disguise. Treat large fan-outs as bulk jobs at
lowornormalpriority, never as 10 millionhighmessages. The Twitter timeline lesson covers fan-out trade-offs in depth.
Failure handling
- Provider outage: circuit breaker opens; switch to a secondary provider for SMS and email; push messages wait in retry queues and drain when APNs or FCM recovers. Critical messages past their usefulness (an expired OTP) are dropped rather than sent late.
- Processor crash mid-message: the queue redelivers to another worker; dedup keys prevent double sends.
- Redis down: decide per check. For dedup, the safe choice for OTPs is to send anyway (a rare duplicate beats a missing code); for promotions, pause rather than risk spamming. For rate limits, fall back to a coarse local limit per worker.
- Bad template deployed: rendering fails, messages go to the DLQ, alerts fire. Because templates are versioned, roll back and replay the DLQ.
- Runaway caller: a buggy service sends a million notifications in a loop. Per-caller quotas at the API (requests per second per service and per type) stop it before users notice; per-user caps are the second line of defence.
Trade-offs and alternatives
| Decision | Option A | Option B | Choice and reason |
|---|---|---|---|
| Priority | One queue, priority field | Separate queues and worker pools | Separate: bulk can never starve critical traffic |
| Queue technology | Kafka topics | RabbitMQ / SQS | Kafka for volume and replayable tracking events; a broker with per-message delay is convenient for retries. Either works |
| Delivery guarantee | At-most-once (never duplicate, may lose) | At-least-once with dedup | At-least-once: a lost OTP or payment alert is worse than a rare duplicate |
| Rendering | At API time | In processors | Processors: locale and channel are decided there, and the API stays fast |
| Preference checks | Caller filters | Notification service filters | Service: one place for consent rules and opt-outs |
| Rate-limit algorithm | Fixed window counters | Token bucket / sliding window | Fixed windows for per-user daily caps; token bucket for provider throughput |
| Retries | Sleep in the worker | Delay queues and DLQ | Delay queues: workers stay free; DLQ keeps failures visible |
| Provider | Single | Primary plus fallback | Fallback for SMS and email; push has no alternative |
What interviewers probe
"How do you make sure an OTP is not delayed by a marketing campaign?" Separate queues and worker pools per priority from intake to sender, plus a reserved slice of each provider's rate limit for critical traffic, and ideally a separate provider account for transactional email and SMS.
"A user complains they got the same notification three times. Where would you look?" The tracking data for that user: three notification ids means the caller sent three times without a stable idempotency key; one id with three sends means queue redelivery bypassed the per-channel dedup key, or retries after an ambiguous timeout. Then fix the corresponding stage.
"How do you send at 9 a.m. in every user's time zone?" Group recipients by time zone, compute the UTC send time for each group, and let the scheduler release each group when due. Spread each release over a few minutes so that one time zone does not create a single spike.
"How do you handle invalid device tokens?" When APNs or FCM reports a token as unregistered or invalid, delete it. Also expire tokens that have not been refreshed in a long time, since the app re-registers its token on launch.
"How would you batch notifications?" For noisy events (likes, comments), hold them for a short window per user and send one summary: "Asha and 12 others liked your post". This is a digest: a buffer per user plus a timer, implemented with the scheduler.
Interview questions
Q1. Why is a notification service usually asynchronous?
Sending depends on slow and unreliable external providers, and one request may fan out to several channels and devices. The API stores the request and enqueues it, returning 202 Accepted in milliseconds. Workers then do the slow work and retry without the caller waiting.
Q2. How do you prioritise urgent notifications?
Classify each type as critical, high, normal or low, and give each priority its own queue and worker pool so bulk traffic cannot block urgent traffic. Keep the separation at the channel senders too, and reserve part of each provider's rate limit for critical messages.
Q3. What is the difference between per-user and per-provider rate limiting?
Per-user limits cap how many notifications one person receives in a period, protecting users from spam and you from OTP cost attacks; they are checked in the processor with atomic Redis counters. Per-provider limits keep the total send rate under each vendor's allowance, enforced by senders with a shared token bucket. They protect different things and both are needed.
Q4. How do you implement quiet hours?
Store a window and time zone per user. For non-urgent notifications, convert the current time to the user's zone; if it falls in the window (handling windows that cross midnight), schedule the notification for the end of the window instead of sending now. Critical notifications such as OTPs and security alerts bypass quiet hours.
Q5. Explain exponential backoff with jitter.
After each failed attempt, the maximum wait doubles (1, 2, 4, 8 seconds and so on) up to a cap, so a struggling provider is not hammered. Jitter picks a random wait below that maximum, so many failing messages do not retry at the same instant. With 8 retries starting at 1 second, the worst-case total wait is 255 seconds.
Q6. What is a dead-letter queue and why use one?
It is a separate queue for messages that failed all retries or hit a permanent error. It keeps failures visible instead of silently dropping them, lets you alert on its size, and lets engineers fix the root cause and replay the messages.
Q7. How do you avoid sending duplicate notifications?
Require a business-derived idempotency key from callers and store it with set-if-absent in Redis. Add a dedup key per notification, channel and device before each provider call, so queue redeliveries skip already-sent messages. Use collapse keys on push so any remaining duplicate replaces the earlier banner instead of adding one.
Q8. Can you guarantee exactly-once delivery?
Not end to end. The final hop is through Apple, Google, SMS operators or email servers, and a crash between the provider accepting the message and our recording that fact can cause a resend. We guarantee at-least-once hand-off and make duplicates rare and harmless.
Q9. How do you know whether a push notification was delivered?
APNs and FCM confirm acceptance, not display. To track delivery and opens, the app reports when a notification is received (where the platform allows) and when it is tapped. SMS and email providers report delivery and bounces through webhooks.
Q10. How would you send a campaign to 50 million users?
Read the segment in batches, compute per-time-zone send times, and schedule them. Release each batch onto the low-priority queue at a capped rate, apply opt-outs and promotional caps in processors, and let senders throttle to provider limits. Critical traffic is unaffected because it uses separate queues.
Q11. What happens when an SMS provider goes down?
A circuit breaker notices the error rate and stops sending to that provider. Senders switch to a secondary provider if one is configured; otherwise messages wait in retry queues. Time-sensitive messages such as OTPs that expire are dropped rather than delivered late.
Key takeaways
- A notification service accepts requests synchronously and sends asynchronously through queues, returning
202 Accepted. - Keep transactional and promotional traffic apart at every stage: queues, workers and provider rate budgets.
- Check preferences, quiet hours (in the user's time zone) and per-user caps before rendering and sending.
- Enforce provider rate limits globally with a shared token bucket; provider limits are often the true bottleneck.
- Classify failures; retry transient ones with exponential backoff and jitter; park the rest in a dead-letter queue.
- Make each hop idempotent with business-derived keys and per-channel dedup keys; exactly-once to the device is not achievable.
- Track statuses from provider webhooks and app events, and stream them to analytics.
- Treat large fan-outs as scheduled, rate-capped bulk jobs.
Next lesson
Continue with Design typeahead autocomplete.

