Why this problem matters
"Design Twitter" (or "design a news feed" for Facebook, Instagram or LinkedIn) is one of the most commonly asked system design questions. The heart of it is one deceptively simple feature: when you open the app, show the most recent posts from everyone you follow, in a fraction of a second.
That is hard because of the shape of the data. One person may follow 500 accounts. One celebrity may be followed by 50 million people. Every time someone posts, the post has to appear in many other people's timelines. Every time someone opens the app, their timeline has to be built from many other people's posts. You must decide when to do that work: at write time, at read time, or a mix.
Interviewers typically probe:
- whether you can estimate the fan-out cost (how many timeline writes one post causes),
- the trade-off between fan-out on write and fan-out on read, and the celebrity problem,
- how you store timelines in a cache and what happens when the cache is missing,
- how you generate unique, time-sortable tweet ids across many machines,
- how search works (an inverted index), and how media and notifications fit in.
A summary of the feed design appears in Real-world designs. This lesson is the full version. It builds on Caching and Message queues.
Step 1: Problem and scope
One sentence:
Users post short messages (tweets); each user sees a home timeline of recent tweets from accounts they follow, can view any user's own tweets, and can search tweets by keyword.
Some vocabulary first:
- Tweet: a post of up to 280 characters, optionally with images or video.
- Follow: a one-way relationship. If you follow a celebrity, they do not have to follow you back.
- User timeline: all tweets written by one user, newest first. This is what you see on someone's profile.
- Home timeline: tweets from everyone you follow, merged, newest first (or ranked). This is the main screen.
- Fan-out: copying or delivering one item to many recipients. A tweet "fans out" to its author's followers.
In scope: post a tweet, view home timeline, view user timeline, follow and unfollow, keyword search, notifications for mentions, media attachments.
Out of scope unless asked: ads, trends, direct messages, a machine-learning ranking model in detail (we will describe where it plugs in).
Step 2: Clarifying questions
| Question | Assumed answer | Why it matters |
|---|---|---|
| How many daily active users (DAU)? | 200 million | Drives read rate |
| Tweets per day? | 400 million | Drives write and fan-out rate |
| Average followers per user? | 200 | Multiplies writes in fan-out on write |
| Largest accounts? | Tens of millions of followers | The celebrity problem |
| Timeline order: time or ranked? | Reverse chronological first; ranking as an extension | Ranking adds a scoring step |
| How fresh must timelines be? | A few seconds of delay is fine | Allows asynchronous fan-out |
| How far back can users scroll? | About 800 tweets in the fast path | Bounds cache size |
| How often is the home timeline read? | 10 times per DAU per day | Read rate |
| Searches per day? | 400 million | Search cluster size |
| Fraction of tweets with media? | 10%, average 500 KB | Media storage and CDN |
The key insight from these answers: timeline reads (2 billion per day) outnumber tweets (400 million per day) by 5 to 1, and each tweet goes to about 200 followers. This is the trade-off you will spend most of the interview on.
Step 3: Requirements
Functional requirements
- Post a tweet, with optional media.
- Delete your own tweet.
- Follow and unfollow a user.
- View your home timeline, newest first, paginated.
- View any user's timeline.
- Search tweets by keywords.
- Receive notifications when mentioned, followed or liked.
Non-functional requirements
- Fast timeline reads: p99 under 200 ms for the first page. This is the screen people see every time they open the app.
- High availability: reading timelines should keep working even if posting is degraded. Showing a slightly stale timeline is far better than showing an error.
- Eventual consistency: a new tweet may take a few seconds to appear in followers' timelines. The author should see their own tweet immediately (read-your-writes for the author).
- Durability: a posted tweet is never lost.
- Scalability: handle celebrity accounts and traffic spikes during big events.
Step 4: Back-of-the-envelope estimates
A day has 86,400 seconds. Round freely; the point is the order of magnitude.
Write traffic
- Tweets: 400 million per day ÷ 86,400 ≈ 4,630 tweets per second. At a 2x peak, about 9,300 per second.
Fan-out traffic
If every tweet is pushed into every follower's home timeline:
- 400 million tweets × 200 followers = 80 billion timeline inserts per day.
- 80 billion ÷ 86,400 ≈ 926,000 inserts per second on average.
That is almost a million writes per second just to keep timelines up to date. This number is why fan-out is the central design decision.
Read traffic
- Home timeline reads: 200 million DAU × 10 = 2 billion per day ≈ 23,000 reads per second.
- Search: 400 million per day ≈ 4,630 searches per second.
Storage
- Tweet text and metadata: about 300 bytes per tweet (text up to 280 characters, which can take more than 280 bytes in UTF-8 for non-Latin scripts, plus ids, timestamps and counters).
- 400 million × 300 bytes = 120 GB per day, about 44 TB per year.
- Media: 10% of 400 million = 40 million items × 500 KB = 20 TB per day. Media is by far the largest storage cost and goes to object storage with a CDN in front.
Timeline cache
If we keep the newest 800 tweet ids for each active user, and each entry is 16 bytes (an 8-byte tweet id and an 8-byte author id):
- 800 × 16 bytes = 12,800 bytes ≈ 12.8 KB per user.
- 200 million active users × 12.8 KB ≈ 2.56 TB of raw data.
Redis adds per-entry overhead, so plan for more than this, spread across a cluster of many memory nodes. It is large but affordable, and it only holds ids, not tweet text.
What the numbers told us
Tweets are written about 4,600 times a second, but pushing them into timelines costs almost a million writes a second. Timelines are read about 23,000 times a second. Storing only ids in timelines keeps the cache at a few terabytes. Media dominates storage.
Step 5: API design
POST /api/v1/tweets
{ "text": "...", "mediaIds": ["m_123"], "replyTo": null }
201 { "id": "1852...", "createdAt": "..." }
DELETE /api/v1/tweets/{id} 204
POST /api/v1/users/{id}/follow 204
DELETE /api/v1/users/{id}/follow 204
GET /api/v1/timeline/home?limit=20&cursor=<lastSeenId>
200 { "tweets": [ ... ], "nextCursor": "1851..." }
GET /api/v1/users/{id}/tweets?limit=20&cursor=<lastSeenId>
GET /api/v1/search?q=world+cup&limit=20&cursor=...
POST /api/v1/media (returns an upload URL; client uploads
the file directly to object storage)
Cursor pagination. Use the id of the last tweet the client saw as the cursor, not a page number. With offsets ("page 3"), new tweets arriving at the top shift everything down and users see duplicates. A cursor says "give me tweets older than this id", which stays correct as new tweets arrive. This only works because tweet ids are sortable by time, which is the job of the id generator described below.
Media upload separately. The client asks for an upload URL, uploads the file straight to object storage, and then posts the tweet with the media id. Large files never pass through the tweet service.
Step 6: Data model and storage choice
users user_id (pk), handle (unique), name, created_at,
follower_count, is_celebrity
tweets tweet_id (pk, Snowflake), author_id, text,
media_ids, reply_to, created_at, deleted
-- shard by tweet_id (or by author_id, see below)
user_tweets author_id, tweet_id desc -- user timeline
-- shard by author_id
follows follower_id, followee_id, created_at
-- two copies: by follower ("who do I follow")
-- and by followee ("who follows me")
home_timeline Redis list per user: key "home:<user_id>"
values: tweet ids, newest first, capped at 800
media media_id, owner_id, object_key, type, size
Choosing stores
- Tweets: a large, append-mostly dataset read by id. A wide-column store such as Cassandra, or sharded MySQL, both work. Twitter historically used sharded MySQL behind its own layers; many teams would pick Cassandra for easy horizontal scaling. The access pattern matters more than the brand: point reads by id, and "latest tweets by author".
- User timeline (
user_tweets): partition byauthor_idwith rows sorted bytweet_iddescending. "Latest 20 tweets by author X" is then one partition, one sorted read. This is a natural fit for a wide-column store. - Follow graph: store each edge twice, once keyed by follower and once by followee, so both "who do I follow" and "who follows me" are single-partition reads. The graph has its own lesson: Design a social graph.
- Home timelines: an in-memory store (Redis lists) because they are read constantly and can be rebuilt if lost.
- Media: object storage plus a CDN.
- Search: a separate search index (Elasticsearch or a custom inverted index), fed asynchronously.
Interview tip
Say explicitly that the home timeline is a cache, not the source of truth. The source of truth is tweets plus the follow graph. If a timeline is lost, it can be rebuilt. This answers half of the failure questions before they are asked.
Step 7: High-level design
clients
|
+-v-------------+ +-----------+
| Load balancer | | CDN |<---- media reads
+--+-----+---+--+ +-----^-----+
| | | |
| | +----------+ +----+--------+
| | | | Object store|
| | v +-------------+
| | +-----------+
| | | Search |--> [ search index cluster ]
| | | service |
| | +-----------+
| v
| +----------------+ +-----------------+
| | Timeline svc |---->| Timeline cache |
| | (read) | | Redis lists |
| +--+---------+---+ +--------^--------+
| | | |
| | hydrate | celebrity | push ids
| v v pull |
| +---------------+ +-----+--------+
| | Tweet cache + | | Fan-out |
| | Tweet store | | workers |
| +-------^-------+ +-----^--------+
| | |
v | |
+--------------+ event "tweet +--+-------+
| Tweet svc |---created"------->| Queue |
| (write) | +--+-------+
+------+-------+ |
| +--> search indexer
v +--> notification svc
+--------------+
| ID generator | (Snowflake, inside each tweet server)
+--------------+
Separate services handle writing tweets, building timelines, reading timelines, search and notifications. They communicate through a queue, so a slow consumer (for example the search indexer) never slows down posting.
Step 8: Request flows
Write path: posting a tweet
- The client sends
POST /tweets. Any media was uploaded earlier and is referenced by id. - The tweet service authenticates the user, checks the rate limit, and validates the text length.
- It generates a tweet id with the Snowflake-style generator (no network call needed).
- It writes the tweet to the tweet store and appends the id to the author's
user_tweetspartition. Once both writes succeed, the tweet is durable. - It returns 201 to the client. The client shows the tweet at the top of the author's own home timeline immediately (read-your-writes for the author, done on the client or by inserting into the author's own timeline synchronously).
- It publishes a
tweet_createdevent to a queue. - Fan-out workers consume the event. If the author is a normal user, they read the author's followers in pages of, say, 5,000, and for each follower whose timeline is cached,
LPUSHthe tweet id ontohome:<follower>andLTRIMthe list to 800 entries. If the author is a celebrity, they skip this step (see the hybrid deep dive). - The search indexer and the notification service consume the same event independently.
Read path: opening the home timeline
- The client sends
GET /timeline/home?limit=20. - The timeline service reads the first 20 ids from
home:<user>in Redis (LRANGE home:<user> 0 19). - It finds which celebrities this user follows (a small list, cached) and fetches each one's latest tweet ids from their
user_tweetspartition (also heavily cached, because millions of people read the same celebrity's latest tweets). - It merges the lists by id (ids are time-ordered, so larger means newer), removes duplicates, and keeps the top 20.
- It hydrates the ids: fetches tweet bodies, author names and avatars, and like counts, mostly from a tweet cache with a multi-get (one request for many keys).
- It filters out deleted tweets, tweets from blocked or muted accounts, and anything the user is not allowed to see.
- It returns the page with a
nextCursorequal to the last tweet id.
The merge step is short to write. Since every list is already sorted newest first, a k-way merge works:
import heapq
def home_timeline(precomputed, celebrity_lists, limit=5):
# Each input list holds tweet ids, newest first. Snowflake ids sort by
# time, so "newest first" is simply "largest id first".
merged = heapq.merge(precomputed, *celebrity_lists, reverse=True)
out, seen = [], set()
for tweet_id in merged:
if tweet_id not in seen:
seen.add(tweet_id)
out.append(tweet_id)
if len(out) == limit:
break
return out
cached = [990, 870, 640, 500] # pushed by fan-out on write
star_a = [950, 700] # pulled at read time
star_b = [980, 870, 300] # 870 also arrived via a retweet
print(home_timeline(cached, [star_a, star_b]))
# [990, 980, 950, 870, 700]
Read path: a user timeline
GET /users/{id}/tweets reads one user_tweets partition sorted by id, then hydrates. It is simple and cache-friendly. Popular profiles are cached as a whole first page.
Step 9: Deep dives
Deep dive 1: Fan-out on write vs fan-out on read vs hybrid
This is the question the whole interview revolves around.
Fan-out on write (push model)
When a user tweets, immediately push the tweet id into every follower's home timeline cache. Reading is then a single list read.
Asha tweets (200 followers)
|
v
fan-out worker --> home:f1 [new, ...]
--> home:f2 [new, ...]
--> ...
--> home:f200
Later, f1 opens app: LRANGE home:f1 0 19 -> done
- Reads are cheap and fast: one Redis read per page.
- Writes are expensive: about 926,000 inserts per second under our numbers.
- Wasted work: many followers never open the app that day, yet we updated their timeline. Fix: only push to users active in the last N days. Inactive users get their timeline rebuilt when they return.
- The celebrity problem: one tweet by an account with 50 million followers means 50 million inserts. Even if the cluster can do one million inserts per second, that single tweet takes about 50 seconds of the entire cluster's write capacity, and followers near the end of the queue see it very late. A few celebrities tweeting during a live event can swamp the system.
Fan-out on read (pull model)
Store nothing per follower. When a user opens the app, look up everyone they follow, fetch each one's recent tweets, and merge.
f1 opens app (follows 500 accounts)
|
v
timeline svc --> user_tweets[a1] latest 20
--> user_tweets[a2] latest 20
--> ... 500 reads ...
--> merge, keep top 20
- Writes are cheap: posting a tweet is one write.
- Reads are expensive: a user who follows 500 accounts needs 500 reads (or a big scatter-gather across shards) every time they open the app, at 23,000 reads per second. That is about 11.5 million partition reads per second, and the slowest of 500 calls decides the latency.
- No wasted work for inactive users, and no celebrity problem on write.
Hybrid (what real systems do)
Use push for ordinary accounts and pull for celebrities.
- Mark accounts above a threshold (say, 1 million followers, or a threshold based on measured fan-out cost) as celebrities.
- Tweets from normal accounts are pushed to followers' timelines.
- Tweets from celebrities are not pushed. At read time, the timeline service pulls the latest tweets of the few celebrities the user follows and merges them in, as in the read path above.
Why this works: a typical user follows only a handful of celebrities, so the read-time pull is a few extra reads, not 500. And a celebrity's latest tweets are read by millions of people, so they sit in cache all the time. The expensive cases of both models are avoided.
| Push (write) | Pull (read) | Hybrid | |
|---|---|---|---|
| Cost to post | O(followers) | O(1) | O(followers), normal users only |
| Cost to read | O(1) list read | O(followees) reads | O(1) + O(celebrities followed) |
| Celebrity tweet | Millions of writes | Cheap | Cheap (pulled) |
| Inactive users | Wasted writes | No cost | Push only to active users |
| Freshness | Seconds of queue delay | Always fresh | Mixed, both fine |
| Complexity | Medium | Low | Highest |
Interview tip
Do the multiplication on the whiteboard: "400 million tweets a day times 200 followers is 80 billion inserts, about a million per second. A 50-million-follower celebrity would need 50 million inserts per tweet. So I'll push for normal users and pull for celebrities." Numbers turn an opinion into a decision.
Common mistake
Do not answer "fan-out on write is better" or "fan-out on read is better" as a universal rule. Each is right for a different kind of account, which is exactly why the hybrid exists.
Deep dive 2: The timeline cache
Structure. Each active user has a Redis list home:<user_id> containing tweet ids (and optionally the author id, so you can filter blocked authors without a lookup). New ids are pushed to the front with LPUSH, and LTRIM home:<user_id> 0 799 keeps only the newest 800.
Why store ids, not tweets? A tweet that fans out to 200 followers would otherwise be copied 200 times. Edits, deletes and like counts would need 200 updates. With ids, the tweet exists once and is hydrated at read time from a tweet cache.
Sharding. Spread timeline keys across Redis nodes by hashing the user id (consistent hashing; see Design a search query cache). Each node has replicas for availability.
Only active users. Keep timelines only for users active in, say, the last 30 days. This cuts memory and fan-out work. When an inactive user returns, there is a cache miss: rebuild their timeline using the pull model once (read recent tweets of everyone they follow, merge, store the result), then switch back to push. The first load is slower; every later one is fast.
Scrolling past 800. Very few users scroll that far. Beyond the cached list, fall back to a pull-based query over followees for older tweets. It is slower, but rare.
Deletes. When a tweet is deleted, mark it deleted in the tweet store. You do not need to remove the id from millions of lists: hydration sees the deleted flag and drops it. Optionally, a background job removes it from lists.
Unfollows and blocks. When you unfollow someone, their tweets may still be in your cached list. Filter at read time by checking the author against your current follow list, and lazily clean the list. A block must take effect immediately, so the read-time filter is required, not optional.
Deep dive 3: Generating tweet ids (Snowflake-style)
We need ids that are:
- unique across hundreds of tweet servers,
- roughly sorted by time, so "newest first" is just "largest id first", and cursor pagination works,
- 64 bits, to fit a standard integer column,
- generated without a network call, so a central counter does not become a bottleneck.
The Snowflake layout, introduced by Twitter and widely copied, splits a 64-bit integer into fields:
| 1 bit | 41 bits | 10 bits | 12 bits |
| 0 | ms since custom epoch | machine id | sequence |
sign ~69.7 years of ms 1,024 4,096 ids
machines per ms each
- 41 bits of milliseconds: 2 to the power 41 milliseconds is about 69.7 years from a custom epoch (a start date you pick, such as 2024-01-01).
- 10 bits of machine id: up to 1,024 generators. Each server gets a unique id at startup, for example from a coordination service.
- 12 bits of sequence: up to 4,096 ids per millisecond per machine, which is about 4 million per second per machine.
Because the timestamp is in the high bits, sorting ids numerically sorts them by creation time (to the millisecond; ids from different machines in the same millisecond are ordered by machine id, which is fine).
import threading
import time
EPOCH_MS = 1_704_067_200_000 # 2024-01-01 00:00:00 UTC, our custom epoch
class Snowflake:
def __init__(self, machine_id: int):
assert 0 <= machine_id < 1024 # 10 bits
self.machine_id = machine_id
self.last_ms = -1
self.sequence = 0
self.lock = threading.Lock()
def next_id(self) -> int:
with self.lock:
now = int(time.time() * 1000)
if now < self.last_ms:
raise RuntimeError("clock moved backwards")
if now == self.last_ms:
self.sequence = (self.sequence + 1) & 0xFFF # 12 bits
if self.sequence == 0: # 4096 used
while now <= self.last_ms: # wait 1 ms
now = int(time.time() * 1000)
else:
self.sequence = 0
self.last_ms = now
return ((now - EPOCH_MS) << 22) | (self.machine_id << 12) | self.sequence
gen = Snowflake(machine_id=7)
ids = [gen.next_id() for _ in range(10_000)]
assert ids == sorted(ids) and len(set(ids)) == len(ids)
Clock problems. If a server's clock jumps backwards (for example after an NTP correction, the protocol that syncs clocks over the network), it could reissue ids it already used. The code above refuses to generate ids until time catches up. Production systems either wait, or fail and let the load balancer route to another server.
Alternatives. UUIDs (128-bit random ids) need no coordination but are twice the size and not time-sorted (UUID version 7 adds a timestamp prefix and fixes the ordering). A database auto-increment is sorted but becomes a bottleneck. Snowflake is the standard interview answer.
Deep dive 4: Search with an inverted index
Searching "world cup" across billions of tweets cannot scan every tweet. Instead, build an inverted index: a map from each word to the list of tweet ids containing it. It is called "inverted" because a normal document maps a tweet to its words; this maps a word to its tweets.
Tweets:
101: "World Cup final tonight"
102: "cup of chai before the final"
103: "World news today"
Tokenize (lowercase, split, drop very common words like "the"):
world -> [101, 103]
cup -> [101, 102]
final -> [101, 102]
tonight -> [101]
chai -> [102]
news -> [103]
today -> [103]
Query "world cup" = world AND cup
[101, 103] intersect [101, 102] = [101]
The lists of tweet ids are called posting lists. Keeping each list sorted lets you intersect two lists by walking them together in linear time.
Building it. The search indexer consumes tweet_created events, tokenizes the text (lowercasing, splitting on spaces and punctuation, handling hashtags and mentions as special tokens, and possibly stemming, which reduces "running" to "run"), and appends the tweet id to each word's posting list. New tweets are searchable within seconds.
Scaling it. The index is too large for one machine, so it is split across many shards. Two ways to split:
- By document (tweet): each shard indexes a subset of tweets, for example by time range or by tweet id hash. A query goes to every shard (scatter), each returns its best matches, and an aggregator merges them (gather). This is the common choice because writes are spread evenly.
- By term (word): each shard holds all posting lists for some words. A query touches only the shards for its words, but popular words create hot shards and multi-word queries must combine results across shards.
Recent tweets are searched far more often than old ones, so the newest days are kept in a fast in-memory tier and older tweets in a larger, slower tier. Results are ranked by recency plus signals like engagement and author reputation. Popular queries are cached; that design is the subject of Design a search query cache. For index internals, see Storage, search and geo.
Deep dive 5: Notifications and media (brief)
Notifications. The notification service consumes events (tweet_created with mentions, followed, liked) from the queue. For each, it decides who to notify, checks their preferences, deduplicates and batches ("Ravi and 23 others liked your tweet"), stores the notification in a per-user inbox, and sends a push message through the phone platforms' push services. It runs entirely off the hot path: a slow push provider never delays posting. The full design is in Design a notification system.
Media. The client uploads the image or video directly to object storage using a pre-signed URL (a temporary URL that grants upload permission for one object). A media worker creates thumbnails and several resolutions and, for video, transcodes into streaming formats. The tweet stores only the media id. Reads go through a CDN, so popular images are served from servers near users. See Design a video streaming service for the video pipeline.
Step 10: Scaling and bottlenecks
- Fan-out workers scale horizontally by adding consumers to the queue. Partition the queue so that one huge account does not block others, and give celebrity-sized jobs their own lower-priority lane (or skip them, with the hybrid).
- Timeline cache scales by adding Redis shards. Because it holds only active users and only ids, growth is bounded.
- Tweet store scales by sharding. Sharding by
tweet_idspreads writes evenly; sharding byauthor_idkeeps a user's tweets together but makes celebrity partitions hot. Common practice is tweets by id, plus a separateuser_tweetsindex by author. - Hot tweets (a viral tweet read by millions): a replicated tweet cache with an in-process cache layer on each timeline server.
- Big events (finals, elections) produce sudden spikes of both tweets and reads. Queues absorb write spikes; fan-out falls behind for a few seconds, which is acceptable because freshness is eventual.
- Ranking (if required) runs after the merge: take a few hundred candidate tweets, score each with a model using features like author affinity and engagement, and return the best 20. Cache the ranked page for a short time.
Step 11: Failure handling
| Failure | Effect | Mitigation |
|---|---|---|
| Redis timeline shard lost | Some users' timelines missing | Replicas; on miss, rebuild by pulling from followees |
| Fan-out workers lag | New tweets appear late | Autoscale workers; alert on queue lag; timelines still readable |
| Duplicate event delivery | Same tweet pushed twice | Idempotent push: check the list head or deduplicate at read time |
| Tweet store shard down | Hydration fails for some tweets | Replicas; skip unavailable tweets rather than failing the page |
| Search cluster down | Search errors | Search is isolated; timelines unaffected |
| Celebrity tweet during spike | Hot keys on their user timeline | Cache their latest tweets with replication and local caches |
| Clock skew on id generator | Duplicate or out-of-order ids | Refuse ids until clock catches up; monitor NTP |
The design degrades gracefully: if fan-out is slow, timelines are a bit stale; if hydration of one tweet fails, that tweet is skipped; if search is down, the rest works. See Reliability and recovery.
Step 12: Trade-offs and alternatives
| Decision | Choice here | Alternative | When the alternative wins |
|---|---|---|---|
| Timeline building | Hybrid push/pull | Pure pull | Small scale, few users, simpler code |
| Timeline storage | Redis list of ids, capped at 800 | Persistent store per user | Users must scroll years back quickly |
| Tweet ids | Snowflake 64-bit | UUIDv7, DB sequence | No custom infrastructure wanted (UUIDv7) |
| Tweet store sharding | By tweet id | By author id | User timeline reads dominate and no celebrities |
| Order | Reverse chronological | Ranked by model | Engagement is the goal; more compute |
| Search index split | By document | By term | Very selective queries with rare terms |
| Fan-out targets | Active users only | All followers | Small graph, cheap memory |
What interviewers probe
"What threshold makes someone a celebrity?" It is a tunable number based on fan-out cost, not a fixed rule. Start with something like a million followers, measure fan-out queue time, and adjust. Some designs also treat accounts that tweet extremely often as "pull" accounts.
"How does a user see their own tweet immediately?" Insert it into the author's own home timeline synchronously during the write, or have the client add it locally. Everyone else can wait for asynchronous fan-out.
"What if a user follows 5,000 celebrities?" That user is rare. Cap the pull to the most recently active celebrities, or precompute a merged list of celebrity tweets for that user periodically.
"How do you handle unfollow?" Filter at read time against the current follow list, and lazily remove ids. Never block the unfollow on cleaning the cache.
"How do you paginate without duplicates?" Cursor by tweet id: "tweets with id less than the last one I saw". Because Snowflake ids are time-ordered, new tweets never shift older pages.
"How do you count likes on a viral tweet?" Do not update one row per like. Buffer increments in memory or in Redis counters and flush periodically; or shard the counter into several sub-counters and add them on read.
"How fresh is search?" New tweets are indexed from the event stream within seconds, into an in-memory tier for recent tweets.
Interview questions
Q1. What is the difference between a home timeline and a user timeline?
A user timeline is one author's own tweets, newest first; it maps to a single partition keyed by author. A home timeline merges tweets from everyone a user follows; it is either precomputed (push) or assembled at read time (pull). The home timeline is the hard one.
Q2. Explain fan-out on write.
When a user posts, a worker inserts the tweet id into the cached home timeline of each follower. Reads become one cheap list read. The cost is moved to write time and grows with follower count, which breaks down for celebrities.
Q3. Explain fan-out on read.
Nothing is precomputed. When a user opens the app, the system fetches recent tweets from every account they follow and merges them. Posting is cheap, but each read touches many partitions, so latency and read load grow with how many accounts a user follows.
Q4. Why does the hybrid model work?
Most accounts have modest follower counts, so pushing their tweets is affordable. The few celebrity accounts would cause millions of writes per tweet, so their tweets are pulled at read time instead. A typical user follows only a few celebrities, and those celebrities' latest tweets are always cached, so the extra pull is cheap.
Q5. Why store tweet ids instead of tweet bodies in the timeline cache?
Ids are small (8 bytes), so the cache fits more users. A tweet is stored once, so edits, deletes and counter changes need one update, not one per follower. Hydration at read time fetches the bodies from a tweet cache in a single multi-get.
Q6. How do you size the timeline cache?
Multiply active users by entries per user by bytes per entry. With 200 million active users, 800 entries and 16 bytes per entry, that is about 2.56 TB of raw data, before Redis overhead, spread across a sharded cluster.
Q7. What happens when a user's timeline is not in the cache?
Rebuild it with a pull: fetch recent tweets from their followees, merge, store the top 800 ids, and continue with push from then on. This is how inactive users who return are handled, and how the system recovers from losing a cache shard.
Q8. Describe the Snowflake id format.
A 64-bit integer with 1 unused sign bit, 41 bits of milliseconds since a custom epoch (about 69.7 years), 10 bits of machine id (1,024 machines) and 12 bits of per-millisecond sequence (4,096 ids per millisecond per machine). Ids are unique without coordination and sort by time.
Q9. What goes wrong if a server's clock moves backwards?
It could generate an id with a timestamp it already used, risking duplicates and misordering. The generator should detect the backward step and wait or refuse to issue ids until the clock passes the last used timestamp.
Q10. How does keyword search work?
An inverted index maps each word to a sorted list of tweet ids containing it. A query looks up each word's list and intersects them, then ranks the results. The index is sharded by document, queries scatter to all shards and results are merged.
Q11. How is a deleted tweet removed from millions of timelines?
It is marked deleted in the tweet store, and the read path drops it during hydration. Removing it from every cached list is optional background cleanup. This makes delete a single write.
Q12. How do you make a block take effect immediately?
Always filter at read time against the viewer's block and mute lists, because cached timelines may still contain the blocked author's tweets. Storing the author id next to each tweet id in the list makes this filter cheap.
Q13. How do notifications fit in without slowing tweets?
The notification service consumes the same tweet and interaction events from the queue asynchronously. It deduplicates, batches, checks preferences and sends pushes. Posting never waits on it.
Q14. Where would a ranking model plug in?
After candidates are gathered (cached timeline plus pulled celebrity tweets), score a few hundred candidates with a model and return the top results. Keep a reverse-chronological fallback in case the ranking service is slow or down.
Key takeaways
- The timeline problem is about when to do the fan-out work: at write time, at read time, or a mix.
- Under our assumptions, push costs about 80 billion inserts per day (about 926,000 per second); a single 50-million-follower tweet would cost 50 million inserts.
- The hybrid pushes for normal accounts and pulls for celebrities, avoiding the worst case of each model.
- Home timelines are a cache of tweet ids, capped (800), only for active users, and rebuildable from the source of truth.
- Snowflake ids give unique, time-sortable 64-bit ids without coordination, which also enables cursor pagination.
- Search is an inverted index fed asynchronously from the tweet event stream, sharded by document.
- Media, search and notifications are separate pipelines off the posting path, connected by a queue.
- Read-time filtering handles deletes, unfollows and blocks correctly even when caches are stale.
Next lesson
Continue with Design a Social Graph.

