Why estimate at all
A back-of-the-envelope estimate is a quick calculation, rough enough to do on scrap paper, that tells you the size of a problem. In a system design interview it answers questions like: Does this fit on one database? Do we need a cache? Is bandwidth the real problem, or storage? Without numbers, every design choice is a guess. With numbers, you can say "about 4,000 reads per second at peak, so one primary with two replicas and a cache is enough", and the interviewer can follow your reasoning.
The goal is the order of magnitude: is it 100, 10,000 or 1,000,000 requests per second? Being off by 20% does not matter. Being off by 1,000 times does, because it changes the architecture. Interviewers commonly probe three things: whether you state your assumptions, whether your arithmetic is sound, and whether you use the result to make a design decision.
This lesson builds on How systems grow and Building blocks explained. The fundamentals lesson has a short estimation section; here we go slowly, with tables to memorise, a repeatable method, six fully worked examples and the mistakes beginners make. Every number below was checked with Python.
The reference numbers
You only need a small set of numbers. Learn these and you can estimate almost anything.
Powers of two and data sizes
Computers count in powers of two, but for estimation you can treat each step of 2^10 as roughly 1,000.
| Power | Exact value | Approximately | Size name |
|---|---|---|---|
| 2^10 | 1,024 | 1 thousand | 1 KB (kilobyte) |
| 2^20 | 1,048,576 | 1 million | 1 MB (megabyte) |
| 2^30 | 1,073,741,824 | 1 billion | 1 GB (gigabyte) |
| 2^40 | about 1.1 trillion | 1 trillion | 1 TB (terabyte) |
| 2^50 | about 1.1 quadrillion | 10^15 | 1 PB (petabyte) |
| 2^32 | 4,294,967,296 | about 4.3 billion | range of a 32-bit integer |
| 2^64 | about 1.8 x 10^19 | effectively unlimited for IDs | range of a 64-bit integer |
Two quick consequences: a 32-bit ID runs out after about 4 billion rows, which a popular app can hit, so use 64-bit IDs. And a byte is 8 bits, which matters for bandwidth (see the mistakes section).
Typical sizes of things
These are assumptions you state out loud; adjust them to the problem.
| Item | Rough size |
|---|---|
| An ID (64-bit integer) | 8 bytes |
| A timestamp | 8 bytes |
| A short text post or tweet with metadata | 300 to 500 bytes |
| A chat message with metadata | about 100 bytes |
| A URL record (short code, long URL, owner, dates) | about 500 bytes |
| A compressed photo | 200 KB to 3 MB |
| One minute of standard-definition video | a few MB to tens of MB, depending on quality |
Time conversions
| Period | Seconds | Easy approximation |
|---|---|---|
| 1 day | 86,400 | about 10^5 (100,000) |
| 1 month (30 days) | 2,592,000 | about 2.5 million |
| 1 year | 31,536,000 | about 3 x 10^7 |
The single most useful shortcut: 1 million requests per day is about 12 per second (1,000,000 / 86,400 = 11.6). So 100 million per day is about 1,200 per second, and 1 billion per day is about 12,000 per second.
If you round a day to 100,000 seconds, your answers come out about 14% low (1 billion / 100,000 = 10,000 instead of 11,574). That is fine for estimation, but say which you used.
Latency numbers (orders of magnitude)
These are approximate orders of magnitude to build intuition. Exact values vary by hardware, generation and distance, so never quote them as precise measurements.
| Operation | Rough time | Order of magnitude |
|---|---|---|
| Read from CPU L1 cache | about 1 ns | nanoseconds |
| Read from main memory (RAM) | about 100 ns | hundreds of nanoseconds |
| Read 1 MB sequentially from RAM | a few microseconds | microseconds |
| Random read from an SSD | tens of microseconds | tens of microseconds |
| Round trip inside one data centre | about 0.5 ms | sub-millisecond |
| Read 1 MB sequentially from SSD | tens of microseconds to about 1 ms | sub-millisecond |
| Hard disk seek | several milliseconds | milliseconds |
| Round trip between continents | about 100 to 200 ms | hundreds of milliseconds |
The ratios are what matter: memory is roughly a thousand times faster than a random disk seek, and a cross-continent round trip costs as much as many thousands of in-data-centre calls. That is why caches live in memory and why multi-region designs keep each user's traffic local. The fundamentals lesson has a similar latency table.
The 5-step method
Use the same five steps every time. Writing them down in order keeps you calm and makes your reasoning easy to follow.
Step 1: Write your assumptions. Daily active users (DAU), actions per user per day, read-to-write ratio, size of each item, how long you keep data, peak factor. Say them out loud and ask the interviewer if they look reasonable.
Step 2: Compute average QPS. QPS means queries (requests) per second.
average QPS = (DAU x actions per user per day) / 86,400
Step 3: Compute peak QPS. Traffic is not flat; evenings, launches and events spike it. Multiply by a peak factor, commonly 2 to 3 for everyday traffic and more for planned events.
peak QPS = average QPS x peak factor
Step 4: Compute storage and bandwidth.
storage per day = new items per day x size per item
storage total = storage per day x days retained x replication factor
bandwidth = QPS x size per response (bytes/second; x8 for bits)
Step 5: Turn numbers into decisions. Cache memory, number of servers, whether one database is enough, whether you need a CDN. An estimate that does not change a decision was not worth doing.
Interview tip
Narrate in round numbers: "100 million DAU, 10 reads each, that is 1 billion reads a day, divided by roughly 100,000 seconds is about 10,000 per second, call it 12,000 using 86,400, so around 35,000 at peak." Then write the result in a corner of the board. You will refer to it all interview.
A small Python helper is handy for practising. You will not run code in an interview, but checking your practice answers this way builds trust in your mental arithmetic.
SECONDS_PER_DAY = 86_400
def qps(dau, actions_per_user, peak_factor=3):
avg = dau * actions_per_user / SECONDS_PER_DAY
return round(avg), round(avg * peak_factor)
def storage_tb(items_per_day, bytes_per_item, years, replicas=1):
total = items_per_day * bytes_per_item * 365 * years * replicas
return total / 1e12
print(qps(100_000_000, 20)) # (23148, 69444)
print(storage_tb(3_333_333, 500, 5)) # about 3.04
Worked example 1: QPS and peak QPS for a Twitter-like timeline
Question. A Twitter-like service has 100 million DAU. How many requests per second must the timeline (home feed) and posting systems handle?
Step 1: assumptions.
- 100 million DAU.
- Each user opens their timeline 20 times per day.
- Each user posts 2 times per day on average (many post nothing, a few post a lot).
- Each user has 200 followers on average.
- Peak factor 3.
Step 2: average QPS.
Timeline reads: 100,000,000 x 20 = 2,000,000,000 per day
2,000,000,000 / 86,400 = 23,148 per second
Posts (writes): 100,000,000 x 2 = 200,000,000 per day
200,000,000 / 86,400 = 2,315 per second
Step 3: peak QPS.
Peak reads = 23,148 x 3 = 69,444 (about 70,000 per second)
Peak writes = 2,315 x 3 = 6,944 (about 7,000 per second)
Step 4: a hidden number, fan-out. If we precompute each user's timeline by copying every new post into each follower's timeline (called fan-out on write), every post creates 200 timeline inserts:
200,000,000 posts/day x 200 followers = 40,000,000,000 inserts/day
40,000,000,000 / 86,400 = 462,963 inserts per second (average)
Step 5: what it means.
- Reads outnumber writes 10 to 1, so this is a read-heavy system: caches and precomputed timelines pay off.
- 70,000 reads per second at peak is well beyond one database, so timelines should be served from an in-memory cache.
- Fan-out on write costs about 460,000 inserts per second on average. That is affordable for normal users but explodes for celebrities with tens of millions of followers. This number is the reason real designs use a hybrid: fan-out on write for ordinary users and fan-out on read (merge at read time) for celebrities. The Twitter timeline case study builds this out.
Worked example 2: storage for 5 years for a URL shortener
Question. A URL shortener creates 100 million new short links per month. How much storage do we need for 5 years, and how long must the short code be?
Step 1: assumptions.
- 100 million new URLs per month.
- Read-to-write ratio 100 to 1 (links are created once and clicked many times).
- Each record: short code (7 bytes), long URL (up to a few hundred bytes), owner ID (8), created and expiry times (16), plus overhead. Call it 500 bytes.
- Keep everything for 5 years.
- 3 copies for durability.
Step 2: QPS (needed later).
Writes: 100,000,000 / (30 x 86,400) = 100,000,000 / 2,592,000
= 38.6 per second
Reads: 38.6 x 100 = 3,858 per second
Step 3: number of records.
100,000,000 per month x 12 months x 5 years = 6,000,000,000 URLs
Step 4: storage.
6,000,000,000 x 500 bytes = 3,000,000,000,000 bytes = 3 TB
With 3 replicas: 3 TB x 3 = 9 TB
Step 5: key length. The short code uses 62 characters (a to z, A to Z, 0 to 9), called base62. How many codes does each length give?
62^6 = 56,800,235,584 (about 57 billion)
62^7 = 3,521,614,606,208 (about 3.5 trillion)
We need 6 billion codes in 5 years. Six characters (57 billion) is enough with roughly a tenfold margin; seven characters (3.5 trillion) gives vast headroom and makes guessing random codes harder. Most designs choose 7.
What it means. 3 TB of data (9 TB with replicas) fits comfortably on a few database machines; this is not a "must shard on day one" problem by size. The hard part is the read rate and low latency, which points to caching (example 4). The URL shortener case study uses these numbers.
Common mistake
Forgetting replication and indexes. Raw data is the floor, not the bill. Real usage adds replicas (often 3 copies), indexes (can be a large fraction of table size), backups and free-space headroom. Say "3 TB raw, roughly 9 TB with three replicas, more with indexes and backups".
Worked example 3: bandwidth for a video upload and streaming service
Question. A video platform receives 500,000 uploads per day and serves 500 million views per day. What upload (ingress) and viewing (egress) bandwidth does it need, and how much storage over 5 years?
Step 1: assumptions.
- 500,000 uploads per day, average upload 200 MB.
- After processing, we store the original plus several resolutions (renditions) totalling about 2 times the original size.
- 500 million views per day; an average view streams about 50 MB. (Check: a 5-minute view at about 1.3 Mbps is 1.3 million bits x 300 seconds / 8 = about 49 MB.)
- Bandwidth is quoted in bits per second, so multiply bytes by 8.
Step 2: ingress (uploads).
500,000 x 200 MB = 100,000,000 MB = 100 TB per day
100 TB / 86,400 s = 1.16 GB per second
1.16 GB/s x 8 = about 9.3 Gbps (average)
Step 3: egress (viewing).
500,000,000 x 50 MB = 25,000,000,000 MB = 25 PB per day
25 PB / 86,400 s = 289 GB per second
289 GB/s x 8 = about 2.3 Tbps (average)
Peak at 2x = about 4.6 Tbps
Step 4: storage over 5 years.
Stored per day = 100 TB x 2 (renditions) = 200 TB
5 years = 200 TB x 365 x 5 = 365,000 TB = 365 PB
Step 5: what it means.
- Egress (about 2.3 Tbps average) is roughly 250 times ingress. No single data centre should serve this directly; a CDN carrying the popular videos is mandatory.
- Storage of hundreds of petabytes means object storage, and probably tiering: popular videos on fast storage, rarely watched old videos on cheaper cold storage.
- Uploads (about 9 Gbps) are large but manageable; uploads should go directly to object storage via signed URLs, with transcoding done by workers from a queue.
See the video streaming case study for the full design.
Worked example 4: cache memory size
Question A. How much memory does the URL shortener need to cache its hot links?
The usual starting rule is the 80/20 rule: a small share of items (around 20%) gets most of the traffic. A common estimate is to cache 20% of a day's read requests' worth of records.
Step 1: assumptions. 3,858 reads per second (from example 2), 500 bytes per record, cache 20% of daily reads.
Step 2: daily reads.
3,858 x 86,400 = 333,333,333 reads per day (about 333 million)
Step 3: memory.
20% of 333,333,333 = 66,666,667 entries
66,666,667 x 500 bytes = 33,333,333,333 bytes = about 33 GB
This is a generous upper bound: many of those reads are repeats of the same popular URLs, so the number of distinct entries is smaller. 33 GB fits on one large cache server; use two or three for redundancy and headroom.
Question B. How much memory to keep the latest timeline for every daily user in the Twitter-like service?
Step 1: assumptions. 100 million DAU; store the most recent 800 post IDs per timeline; 8 bytes per ID (we store IDs, not full posts, and fetch post bodies from a separate post cache).
Step 2: memory.
100,000,000 users x 800 IDs x 8 bytes = 640,000,000,000 bytes = 640 GB
Step 3: what it means. 640 GB does not fit in one machine's memory comfortably, so the cache is partitioned across nodes, for example 10 nodes of 64 GB each (640 / 64 = 10) before replication and overhead. Partition by user ID with consistent hashing. In practice you would also add overhead for the data structure and keep a replica of each partition, so plan for roughly double.
Interview tip
Storing IDs instead of full objects in list caches is a classic trick worth saying aloud. 800 IDs cost 6.4 KB per user; 800 full posts at 300 bytes each would cost 240 KB, 37.5 times more.
Worked example 5: number of servers
The formula.
servers = peak QPS / (QPS one server handles x target utilisation)
Target utilisation is how busy you let servers get at peak, often 60% to 70%, leaving headroom for spikes, a failed server and slow requests. The QPS one server handles depends entirely on the work per request; in an interview you state an assumption (or measure it in real life with a load test).
Question A. How many app servers for the URL shortener's redirect path?
Peak reads = 3,858 x 3 = 11,574 per second
Assume one server handles 1,000 simple redirects per second
Target utilisation 70%
servers = 11,574 / (1,000 x 0.7) = 11,574 / 700 = 16.5 -> 17
Then add spare capacity so that losing a server, or a whole availability zone, does not overload the rest. With servers spread across 3 zones, losing one zone removes a third of them; one simple approach is to round up to a multiple of 3 and add margin, for example 18 to 24 servers. The precise answer matters less than showing headroom and failure thinking.
Question B. How many servers to assemble timelines?
Peak timeline reads = 69,444 per second (example 1)
Assume one server handles 500 timeline requests per second
(each request reads a cached ID list and fetches ~20 posts)
Target utilisation 70%
servers = 69,444 / (500 x 0.7) = 69,444 / 350 = 198.4 -> about 200
What it means. Two hundred stateless servers behind load balancers is normal. If the answer had been 20,000, that would tell you the per-request work is too heavy and you need more caching or precomputation.
Common mistake
Sizing from average QPS. A fleet sized for 23,000 requests per second falls over at the 70,000 evening peak. Always compute peak, then add headroom for failures.
Worked example 6: a full estimate for a chat app
This example puts every step together, as you would in an interview for a one-to-one and group chat service. See the chat messaging case study (coming to this track as design-chat-messaging) for the full design.
Step 1: assumptions.
- 50 million DAU.
- Each user sends 40 messages per day.
- Each message, with metadata (IDs, timestamps, status), is about 100 bytes.
- Keep messages for 5 years, 3 replicas.
- At peak, 20% of DAU are online at once, each holding one long-lived connection (such as a WebSocket).
- One connection server can hold about 50,000 open connections (an assumption; real numbers depend on memory per connection and message rate).
- Peak factor 3.
Step 2: message QPS.
Messages per day = 50,000,000 x 40 = 2,000,000,000 (2 billion)
Average = 2,000,000,000 / 86,400 = 23,148 per second
Peak = 23,148 x 3 = 69,444 per second
Step 3: storage.
Per day = 2,000,000,000 x 100 bytes = 200 GB
5 years = 200 GB x 365 x 5 = 365,000 GB = 365 TB
3 replicas: 365 TB x 3 = 1,095 TB = about 1.1 PB
Step 4: bandwidth for message ingress at peak.
69,444 messages/s x 100 bytes = 6,944,400 bytes/s = about 6.9 MB/s
x 8 = about 56 Mbps
Small. Text chat is not a bandwidth problem; media attachments would be, and they go to object storage and a CDN.
Step 5: connection servers.
Concurrent connections = 50,000,000 x 20% = 10,000,000
Servers at full capacity = 10,000,000 / 50,000 = 200
At 70% target utilisation = 10,000,000 / 35,000 = 285.7 -> about 286
What it means.
- Connections, not message rate or bandwidth, drive the size of the real-time tier: about 300 stateful servers, and you need a way to find which server holds a given user's connection (a session registry).
- 69,444 writes per second and 1.1 PB over 5 years call for a horizontally scalable, write-optimised store, partitioned by conversation ID, such as a wide-column database.
- Messages to offline users must be stored and delivered later, so delivery goes through a queue.
A small practice set
Try these before checking the answers (all verified with Python).
| Question | Answer |
|---|---|
| 10 million requests per day: average QPS? | 10,000,000 / 86,400 = about 116 |
| A paste-bin gets 1 million new pastes a day, 10 KB each. Storage for 5 years (no replicas)? | 10 GB/day x 365 x 5 = 18.25 TB |
| Same paste-bin, 10 reads per paste. Peak read QPS at factor 3? | 10,000,000 / 86,400 x 3 = about 347 |
| Recipebox: 10 million DAU, 30 requests each. Average and peak (x3)? | about 3,472 and about 10,417 |
| 100 Mbps link: megabytes per second? | 100 / 8 = 12.5 MB/s |
Common estimation mistakes
- Mixing bits and bytes. Network speeds are in bits per second (Mbps, Gbps); file sizes are in bytes (MB, GB). Divide bits by 8 to get bytes. Mistaking them is an 8-times error.
- Mixing 10^3 and 2^10 inconsistently. Pick decimal (1 KB = 1,000 bytes) for estimation and stay with it. The 2.4% difference per step does not matter; inconsistency confuses you.
- Forgetting peak. Average QPS is a planning input, not a capacity target.
- Forgetting replication, indexes and growth. Multiply raw storage by replicas and note overhead. Also ask whether usage grows each year.
- Losing a factor of 1,000 in units. Write units on every line ("per day", "per second", "bytes", "GB"). Most errors are unit errors.
- False precision. "We need 16.53 servers" sounds precise but each input is a guess. Round, and explain the margin.
- Estimating without using it. Every number should lead to a decision: cache or not, shard or not, CDN or not.
- Spending ten minutes on it. In a 45-minute interview, estimation should take about 5 minutes. Estimate what matters for the design; skip the rest.
- Silent assumptions. If you assume 200 followers per user, say it. Interviewers cannot give credit for reasoning they cannot see, and they may want a different assumption.
Interview tip
If the interviewer gives you no numbers, propose them: "Shall I assume 10 million daily users and a 10 to 1 read-to-write ratio?" Proposing reasonable numbers shows confidence, and they will correct you if they had something else in mind.
Interview questions
Q1. How do you convert daily active users into requests per second?
Multiply DAU by actions per user per day to get daily requests, then divide by 86,400 seconds (about 10^5). Multiply by a peak factor, commonly 2 to 3, to get peak QPS. For example, 50 million users making 40 requests a day is 2 billion per day, about 23,000 per second on average and about 70,000 at peak.
Q2. Why do we estimate peak QPS and not just average?
Systems fail at peak, not on average. Traffic follows daily cycles and spikes during events, so a system sized for the average will overload every evening. Capacity plans size for peak plus headroom for failures and unexpected bursts.
Q3. How would you estimate storage for a photo-sharing app over 5 years?
Assume uploads per day and average photo size, multiply to get bytes per day, then multiply by 365 x 5 and by the number of stored versions (original, thumbnails) and replicas. For example, 10 million photos a day at 2 MB is 20 TB per day, about 36.5 PB in 5 years before thumbnails and replication. Then conclude: object storage plus CDN, with cold tiers for old photos.
Q4. What is the 80/20 rule in cache sizing?
It is the observation that a small fraction of items, around 20%, often receives most requests. A common starting estimate is to cache 20% of a day's worth of reads or records. It is a heuristic; you then refine it with the measured hit ratio.
Q5. How do you estimate the number of servers needed?
Divide peak QPS by the QPS one server handles at a safe utilisation, such as 70%. The per-server QPS is an assumption you state, or a load-test result. Then add redundancy so the system survives losing a server or a zone.
Q6. Why does the base62 code length matter for a URL shortener?
The number of unique codes is 62 to the power of the length. Six characters give about 57 billion codes and seven give about 3.5 trillion; you compare that with the number of URLs you expect to create over the system's life. A larger space also makes random codes harder to guess.
Q7. Bits or bytes: how do you compute bandwidth?
Bandwidth is QPS times response size in bytes, multiplied by 8 to get bits per second, which is how network capacity is quoted. For example, 1,000 requests per second of 100 KB responses is 100 MB/s, or 800 Mbps.
Q8. Your estimate says 640 GB of cache. What do you do?
That is too much for one server comfortably, so partition the cache across nodes, for example ten 64 GB nodes, using consistent hashing on the key, and add replicas for availability. I would also look for ways to shrink it: store IDs rather than full objects, cache only active users, or shorten the list length.
Q9. How precise should estimates be?
Order-of-magnitude accurate, with round numbers and stated assumptions. The decision usually changes only when a number crosses a boundary, such as "fits on one machine" vs "must be partitioned". False precision wastes time and suggests you do not understand what the estimate is for.
Q10. Why are latency numbers useful in design?
They show where time goes: memory is far faster than disk, and a cross-continent round trip is very slow. That justifies caching in memory, batching network calls, avoiding chatty service-to-service calls and keeping users in nearby regions. You quote them as orders of magnitude, not exact figures.
Q11. What does fan-out mean for estimation in a social feed?
Fan-out is how many downstream writes or reads one action triggers. If each post is copied to every follower's timeline, write load is posts per second times average followers, which can be hundreds of thousands per second. That number decides between fan-out on write, fan-out on read or a hybrid.
Q12. What dominates the cost of a chat system: messages, storage or connections?
For text chat, usually concurrent connections and storage. Message bandwidth is small (tens of Mbps in our example), but tens of millions of persistent connections need hundreds of stateful servers, and years of messages reach petabytes with replication.
Key takeaways
- Estimates aim for the right order of magnitude, not precision.
- Memorise: 1 day is about 86,400 (about 10^5) seconds; 1 million per day is about 12 per second.
- Follow five steps: assumptions, average QPS, peak QPS, storage and bandwidth, then decisions.
- Always size for peak, and add replicas, indexes and headroom to raw numbers.
- Bandwidth is in bits; storage is in bytes. Multiply or divide by 8.
- Cache estimates start from the 80/20 heuristic; store IDs rather than objects in list caches.
- Servers = peak QPS divided by per-server QPS times target utilisation, plus failure headroom.
- Every number should change a design decision, or it was not worth computing.
Next lesson
Continue with Interview playbook.

