What this interview actually tests
A system design interview is a 45 to 60 minute conversation in which you design a software system, such as a URL shortener, a chat app or a news feed, on a whiteboard or a shared drawing tool. There is no single correct answer and no code to compile. Instead, the interviewer watches how you think:
- Do you clarify a vague problem before solving it?
- Do you make reasonable estimates and use them?
- Can you draw a working end-to-end design, then improve it where it matters?
- Do you see trade-offs and failure cases, and explain your choices clearly?
- Do you drive the conversation while listening to hints?
For freshers and junior engineers, the bar is a sound, simple design with clear reasoning. For experienced engineers, the bar rises to depth, operational judgment and leading the discussion. Either way, structure is what separates a calm, convincing interview from a scattered one.
This lesson gives you that structure. It assumes you have read the earlier beginner lessons: How systems grow for why components exist, Building blocks explained for what each one does, and Estimation practice for the numbers. Here you will learn the framework, a time budget, phrases to use at each step, a full worked mini-interview, how you are scored, the common mistakes, and a study plan through this track.
How the session runs
Most system design interviews follow a similar arc, whether in person or online:
0 min Introductions, problem statement ("Design a paste-bin")
|
v
~2 min You lead: requirements, estimates, API, data model
|
v
~20 min High-level design on the board
|
v
~30 min Deep dives: interviewer picks areas, or you propose them
|
v
~40 min Bottlenecks, failures, trade-offs, extensions
|
v
~45 min Wrap-up; your questions for the interviewer
The problem statement is deliberately vague. "Design a paste-bin" does not say how many users, whether pastes expire, or whether users log in. That vagueness is part of the test: the interviewer wants to see you turn it into a concrete, scoped problem.
The interviewer usually plays two roles. Early on, they are the product owner who answers questions about requirements. Later, they are a senior colleague who pushes on weaknesses: "What happens if this cache dies?" Treat both as collaboration, not attack.
The framework: eight steps
Here is the framework we will use. Memorise the order; the content of each step follows.
| Step | What you produce | 45-min budget | 60-min budget |
|---|---|---|---|
| 1. Requirements | Functional and non-functional lists, scope | 5 min | 6 min |
| 2. Estimates | QPS, peak, storage, bandwidth | 4 min | 5 min |
| 3. API | Main endpoints with inputs and outputs | 4 min | 5 min |
| 4. Data model | Tables or collections, keys, storage choice | 4 min | 5 min |
| 5. High-level design | Boxes and arrows for the main flows | 10 min | 12 min |
| 6. Deep dives | Two or three components in detail | 12 min | 18 min |
| 7. Bottlenecks and trade-offs | Failures, scaling limits, alternatives | 4 min | 6 min |
| 8. Wrap-up | Summary, extensions, questions | 2 min | 3 min |
| Total | 45 min | 60 min |
The budget is a guide, not a stopwatch. If the interviewer wants to jump to a deep dive early, follow them. But if you notice you have spent 15 minutes on requirements, move on: running out of time before drawing a high-level design is one of the most common ways to fail.
Interview tip
Say the plan at the start: "I will clarify requirements, do quick estimates, sketch the API and data model, draw a high-level design, then go deep on the parts you find most interesting." It takes ten seconds and tells the interviewer you have a process.
Step 1: Requirements
Split requirements into two kinds.
Functional requirements are what the system does: the features. For a paste-bin: create a paste, view a paste by its link, optionally set an expiry.
Non-functional requirements are how well it does them: scale, latency, availability, consistency, durability, security. For a paste-bin: reads should be fast, pastes must not be lost, the service should stay up, and links should not be guessable.
Then scope: say explicitly what you will not design. Analytics dashboards, user accounts and abuse detection might be out of scope. Scoping protects your time and shows judgment.
Questions worth asking for almost any problem:
- Who are the users and how many? Daily active users?
- What are the two or three core features? Which are most important?
- Read-heavy or write-heavy?
- How fresh must data be? Is a few seconds of delay acceptable?
- Any latency target? Any availability target?
- How long do we keep data?
- Global or one country? Mobile, web or both?
Sample phrases:
- "Before designing, I want to make sure I build the right thing. Can I ask a few questions?"
- "I will focus on creating and reading pastes; I will treat user accounts and analytics as out of scope unless you want them."
- "Is it acceptable if a new paste takes a second or two to be readable worldwide?"
Step 2: Estimates
Use the 5-step method from Estimation practice: assumptions, average QPS, peak QPS, storage and bandwidth, then decisions. Keep it to a few minutes and write the results in a corner of the board.
Sample phrases:
- "Let me assume 1 million new pastes a day and 10 reads per paste. Does that sound reasonable?"
- "That gives about 12 writes and 116 reads per second on average, roughly 350 reads per second at peak."
- "So this is read-heavy and modest in QPS; storage over 5 years is the bigger number."
If the interviewer says "skip the numbers", do so gracefully but mention scale qualitatively: "I will assume read-heavy with a few hundred requests per second."
Step 3: API
Define the main operations as endpoints. This turns features into a contract and often reveals missing requirements. Keep it to the core: 2 to 4 endpoints.
Sample phrases:
- "I will use a REST-style API since this is client-facing."
- "Create returns the short key so the client can build the link."
- "I will include an idempotency key on create so a retried request does not make two pastes." (This is a strong signal, but only if you can explain it.)
For API conventions, pagination and versioning, see API design.
Step 4: Data model
List the main entities and their fields, pick the primary key, name the indexes your queries need, and choose a storage type with a reason.
Sample phrases:
- "The main query is lookup by key, so the key is the primary key and no other index is needed for reads."
- "Paste content can be up to 1 MB, so I will keep content in object storage and metadata in a database."
- "I would start with a relational database because the metadata is small and structured; a key-value store would also fit since all reads are by key."
Step 5: High-level design
Draw boxes and arrows for the main flows, end to end: client, load balancer, services, data stores. Walk through each core use case on the drawing ("a create request goes here, then here"). Start simple, as in stage 3 or 4 of How systems grow, and add components only when a requirement or number calls for them.
Sample phrases:
- "Let me first draw a design that works, then we can scale it."
- "Here is the write path... and here is the read path."
- "Given 350 peak reads per second, a cache in front of the database will absorb most of them."
Step 6: Deep dives
Pick two or three components and go deep: algorithms, data flow, failure handling, scaling. Often the interviewer chooses; if not, propose the most interesting or risky parts. For a paste-bin: key generation, the read path with caching, and expiry clean-up.
Sample phrases:
- "Which area would you like me to go deeper on? I think key generation and caching are the most interesting here."
- "There are two options for generating keys. Let me compare them."
Step 7: Bottlenecks and trade-offs
Walk through the design looking for weak points: single points of failure, hot keys, overloaded components, data loss windows. State trade-offs you made and alternatives you rejected.
Sample phrases:
- "The single database is a single point of failure; I would add a replica with automatic failover."
- "I chose eventual consistency for the cache; the trade-off is that a deleted paste may be visible for up to the TTL, so on delete I also remove the cache entry."
- "If traffic grew 100 times, the first thing to break would be..."
Step 8: Wrap-up
Summarise the design in 30 seconds, mention what you would add with more time (monitoring, rate limiting, analytics), and ask a thoughtful question if time allows.
Sample phrase:
- "To summarise: stateless API servers behind a load balancer, metadata in a replicated database, content in object storage behind a CDN, a cache for hot pastes and a background job for expiry. With more time I would add rate limiting on create and abuse detection."
A worked mini-interview: design a paste-bin
Below is a condensed transcript-style walkthrough of a 45-minute interview. "I" is the interviewer and "C" is the candidate. Notes in square brackets explain what the candidate is doing well. Read it once for the flow, then again for the phrasing.
Minutes 0 to 5: requirements
I: Let us design a paste-bin: users paste text and get a link they can share.
C: Sure. I will start with requirements, do quick estimates, then API, data model and a high-level design, and then go deeper where you like. First, functional requirements. Users paste text and get a unique link. Anyone with the link can view the paste. Should pastes expire?
I: Yes, users can choose an expiry; default never.
C: Do users need accounts, or is it anonymous?
I: Anonymous is fine for now.
C: Is there a size limit per paste?
I: Say 1 MB maximum.
C: Do we need editing, or are pastes immutable?
I: Immutable.
C: Great, immutability makes caching much easier. Out of scope unless you want them: accounts, syntax highlighting, analytics and abuse detection. Non-functional: reads should be fast, say under 200 ms at p99 for most users; pastes must not be lost once created; the service should be highly available; and links should not be guessable, since pastes may be private-by-obscurity. Is that right?
I: That is right.
[The candidate stated a plan, asked targeted questions, noticed that immutability matters, scoped explicitly, and turned "fast" and "reliable" into concrete targets.]
Minutes 5 to 9: estimates
C: Let me assume 1 million new pastes per day, 10 reads per paste, average paste 10 KB, keep data for 5 years.
Writes: 1,000,000 / 86,400 = about 12 per second
Reads: 10,000,000 / 86,400 = about 116 per second
Peak reads (x3) = about 350 per second
Storage: 1,000,000 x 10 KB = 10 GB per day
5 years: 10 GB x 365 x 5 = 18.25 TB (about 55 TB with 3 copies)
Pastes in 5 years: 1M x 365 x 5 = 1.825 billion
Peak read bandwidth: 350 x 10 KB = about 3.5 MB/s
C: So the QPS is modest; one well-provisioned database could almost handle it, but 18 TB of content over 5 years is large for a database, and content is up to 1 MB. That suggests storing content in object storage and only metadata in the database. Reads dominate 10 to 1, so caching will help.
[Numbers were quick, rounded and immediately turned into two design decisions.]
Minutes 9 to 13: API
C: Two core endpoints:
POST /api/v1/pastes
body: { "content": "...", "expires_in_seconds": 86400 }
returns: 201 { "key": "aZ3kQ9x", "url": "https://paste.example/aZ3kQ9x",
"expires_at": "2026-10-11T10:00:00Z" }
GET /api/v1/pastes/{key}
returns: 200 { "content": "...", "created_at": "...", "expires_at": "..." }
404 if missing or expired
C: I would also accept an optional Idempotency-Key header on create, so if the client retries after a timeout, we return the same paste instead of creating a duplicate. And I would rate-limit create per IP to limit abuse.
Minutes 13 to 17: data model
C: One metadata table:
CREATE TABLE pastes (
paste_key CHAR(7) PRIMARY KEY, -- base62 key
object_key VARCHAR(64) NOT NULL, -- where content lives
size_bytes INTEGER NOT NULL,
created_at TIMESTAMP NOT NULL,
expires_at TIMESTAMP -- NULL = never
);
CREATE INDEX idx_pastes_expires_at ON pastes (expires_at);
C: Reads are always by paste_key, which is the primary key. The expires_at index serves the clean-up job that finds expired pastes. Content lives in object storage under object_key. Each row is under 100 bytes, so 1.8 billion rows is roughly 180 GB of metadata, which fits on one relational database with replicas for a long time. A key-value store would also work since every read is by key; I am choosing relational for simplicity and the expiry query.
[A schema with a reason for every index, a quick size check and an honest alternative.]
Minutes 17 to 27: high-level design
C: Here is the design:
+---------+
client --------------->| CDN |---- (miss) ----+
| GET /aZ3kQ9x +---------+ |
| v
| POST, GET (API) +----+ +------------------+
+------------------>| LB |---->| API servers (xN) |
+----+ +------------------+
| | |
+---------+ | +------------+
v v v
+-----------+ +-------------+ +----------------+
| Cache | | Metadata DB | | Object storage |
| (hot keys)| | primary + | | (content) |
+-----------+ | replica | +----------------+
+-------------+
+---------------------------+
| Expiry job (scheduled) |
+---------------------------+
C: Write path: the API server validates size, generates a unique key, writes content to object storage, then inserts the metadata row, and returns the URL. Writing content first means a crash leaves at worst an orphaned object, never a row pointing to nothing; the expiry job can clean orphans.
C: Read path: the API server checks the cache for the key. On a hit, return. On a miss, read the metadata row; if it is missing or expired, return 404; otherwise fetch the content from object storage, put it in the cache with a TTL and return it. Because pastes are immutable, caching is safe; the only invalidation is on expiry or deletion.
C: API servers are stateless, so we scale them behind the load balancer. For public pastes, we could even let the CDN cache the GET response with a TTL no longer than the paste's remaining life.
I: Good. How do you generate the key?
Minutes 27 to 39: deep dives
C: Key generation has three common options:
| Option | How | Pros | Cons |
|---|---|---|---|
| Hash the content | First 7 chars of a hash of content + salt | No coordination | Collisions must be checked; identical pastes need a salt |
| Counter + base62 | A global counter, encoded in base62 | No collisions, short | Sequential keys are guessable; counter is a coordination point |
| Random + check | Generate 7 random base62 chars; insert; retry on conflict | Unguessable, simple | Rare retries |
C: Since links should not be guessable, I prefer random keys. With 62^7, about 3.5 trillion possible keys, and 1.8 billion used after 5 years, the chance a new random key collides is about 1.8 billion divided by 3.5 trillion, roughly 0.05%, so retries are rare. The primary-key constraint detects collisions safely: insert, and if it fails with a duplicate key, generate another. An alternative is a separate key-generation service that pre-generates unused keys in batches, which avoids retries but adds a component.
I: What about the cache: how big?
C: Using the 80/20 heuristic, cache 20% of a day's reads: 20% of 10 million reads times 10 KB is about 20 GB, which fits on one cache node; I would run two for redundancy. Eviction is LRU. If the cache dies, reads fall through to the database and object storage, which at 350 per second they can absorb, so the cache is a performance layer, not a correctness one.
I: And expiry?
C: Two parts. On read, we always check expires_at, so an expired paste returns 404 immediately even if it has not been deleted yet. Separately, a scheduled job runs every few minutes, finds rows with expires_at in the past using the index, in batches of, say, 1,000, deletes the object and then the row, and removes the cache key. Batching keeps the job from hammering the database. If the job falls behind, correctness is unaffected because reads check expiry.
[The candidate compared options in a table, chose one with a reason tied to a requirement, checked collision math, and explained how failures degrade gracefully.]
Minutes 39 to 43: bottlenecks and trade-offs
C: Weak points and fixes:
- The metadata database is a single point of failure. Add a replica with automatic failover. Reads can also go to replicas; a just-created paste might briefly miss on a lagging replica, so for a few seconds after create, the creator's read could go to the primary.
- A viral paste is a hot key. The cache and CDN absorb it; since content is immutable, a long CDN TTL is safe.
- If writes grew 100 times, metadata would still fit for a while; eventually I would shard by
paste_key, which is random, so load spreads evenly. - Abuse: rate limits on create per IP and a size cap; scanning for malware or illegal content would be a separate asynchronous pipeline.
Minutes 43 to 45: wrap-up
C: To summarise: stateless API servers behind a load balancer; random 7-character base62 keys; metadata in a replicated relational database; content in object storage; a cache and CDN for hot reads; and a scheduled expiry job, with expiry also enforced on read. With more time I would add monitoring on p99 latency, error rates and cache hit ratio, and discuss user accounts. Could I ask how your team decides between building on managed cloud services and running your own storage?
How you are scored
Interviewers usually take notes against a few signals (observable behaviours) and then write a recommendation. Exact rubrics vary between companies, but the signals below are widely used.
| Signal | What it looks like when strong |
|---|---|
| Problem exploration | Asks focused questions, states assumptions, scopes clearly |
| Estimation and data | Quick, sane numbers that drive decisions |
| High-level design | Complete, working end-to-end flow; nothing magic |
| Component knowledge | Picks the right building blocks and explains why |
| Trade-offs | Compares options, names costs, ties choices to requirements |
| Depth | Can go several levels down on at least one or two components |
| Failure thinking | Considers outages, retries, data loss, hot spots |
| Communication | Structured, clear, checks in, uses the board well, responds to hints |
Junior vs senior expectations
| Area | Fresher or junior | Senior |
|---|---|---|
| Driving | Follows the framework; accepts guidance | Leads the whole session; proposes the deep dives |
| Design | Correct, simple design; knows standard blocks | Design fitted to numbers; knows when not to add complexity |
| Depth | Reasonable detail on one component | Deep detail on several, including internals and operations |
| Trade-offs | Names main pros and cons | Quantifies them; discusses cost, team and migration impact |
| Failures | Mentions replicas, retries | Discusses failure modes, recovery, degradation, monitoring and alerting |
| Hints | Uses hints well | Rarely needs them; anticipates follow-ups |
For freshers, a clean simple design with clear reasoning, sensible estimates and willingness to discuss trade-offs is usually a pass. You are not expected to know the internals of every database. You are expected not to hand-wave: if you add a component, you should be able to say what it does and why.
What a 'no' usually means
Most rejections are not about missing an obscure technology. They come from a missing end-to-end design, no requirements gathering, buzzwords without reasons, or freezing instead of thinking aloud. All of these are fixable with practice.
Common mistakes
- Jumping into boxes immediately. Drawing a design before requirements means solving the wrong problem. Always spend the first few minutes clarifying.
- Over-engineering. Adding Kafka, microservices, sharding and three regions to a system with 100 requests per second. Fit the design to the numbers.
- Buzzword listing. "We use Redis, Kafka and Cassandra" without saying what each does. Every box needs a reason.
- Going silent. The interviewer cannot score thoughts they cannot hear. Think aloud, even when unsure: "I am deciding between two options..."
- Never finishing the high-level design. Spending 30 minutes on one component and never connecting the full flow. Breadth first, then depth.
- Ignoring hints. When the interviewer asks "what if this server dies?", it is a hint that something is missing. Engage with it rather than defending.
- Refusing to commit. "It depends" is a start, not an answer. Say what it depends on and then choose.
- Ignoring failure. A design that only works when everything is healthy is incomplete.
- Messy board. Arrows going everywhere, unlabelled boxes. Keep the flow left to right or top to bottom, label every arrow's purpose for the core flows.
- Arguing instead of discussing. If the interviewer suggests a different approach, compare it honestly. You can disagree, with reasons, and still be collaborative.
Common mistake
Memorising a "standard" design for each common question and reciting it. Interviewers change one requirement ("pastes must be editable", "now support 1 billion users") precisely to see whether you can reason. Learn the reasons behind designs, not the diagrams.
Study plan through this track
Here is a practical order through the System Design track. Adjust the pace to your background; a fresher with some time can do one phase per week.
Phase 1: Foundations (beginner path).
- How systems grow: why each component exists.
- Building blocks explained: what each block does and when.
- Estimation practice: numbers you can use.
- This playbook: the interview process.
- Fundamentals: requirements, CAP, consistency, availability in depth.
Phase 2: Core building blocks in depth.
- Networking, proxies and edge delivery
- API design and identity and security
- Load balancing
- Databases at scale and caching
- Scalability
- Message queues
Phase 3: Practise case studies with the framework. For each one, first try it yourself for 45 minutes on paper using the eight steps, then read the case study and note what you missed.
- Design a URL shortener: the closest relative of the paste-bin; start here.
- Design a Twitter-like timeline: fan-out and caching.
- Design a social graph: relationships and sharding.
- Design a sales ranking system: aggregation and batch vs stream.
- Design video streaming: object storage, CDN, transcoding.
- Design file sync and storage: chunking and sync.
- Further case studies being added to this track: web crawler (design-web-crawler), search query cache (design-search-query-cache), personal finance app (design-personal-finance-app), chat and messaging (design-chat-messaging), ride-hailing (design-ride-hailing), notification system (design-notification-system) and typeahead autocomplete (design-typeahead-autocomplete).
- Real-world designs and design workshops for more practice.
Phase 4: Advanced topics (especially for experienced candidates).
- Microservices, Kafka and event streaming, transactions and CDC
- Consensus and coordination, reliability and recovery, observability and performance
- Multi-region design, storage, search and geo, booking and payments design, capacity, SLOs and cost
Phase 5: Mock interviews. Do at least three timed mocks with a friend playing interviewer, using problems you have not studied. Record yourself if you can, and check: did you clarify, estimate, finish a high-level design, go deep, and discuss failures? The course guide has a concept index for quick review, and system design videos help between sessions.
Interview tip
Practise drawing on the same medium you will use: a physical whiteboard, a drawing tool or a shared document. Many candidates lose minutes fighting the tool. A few practice sessions make your diagrams quick and tidy.
Interview questions
These are questions about the interview process that candidates often wonder about, plus meta-questions interviewers sometimes ask.
Q1. What should you do in the first five minutes?
Restate the problem, state your plan, and clarify functional and non-functional requirements with focused questions. Agree on scope, including what you will leave out. Do not draw architecture yet; a few minutes of clarification prevents solving the wrong problem.
Q2. How much time should estimation take?
About 4 to 5 minutes in a 45-minute interview. Estimate only what affects the design, usually QPS, peak QPS, storage and sometimes bandwidth, and say how each number changes your choices. If the interviewer wants to skip it, state the scale qualitatively and move on.
Q3. What if you do not know a technology the interviewer mentions?
Say so honestly and reason from first principles: "I have not used that, but if it is a distributed log, I would expect it to give ordering within a partition." Interviewers value honest reasoning far more than bluffing, which is quickly exposed by follow-up questions.
Q4. How do you decide what to deep dive on?
Choose the components that are most specific to the problem or carry the most risk, such as key generation for a URL shortener, fan-out for a feed, or connection handling for chat. Offer two or three options and let the interviewer pick. Generic components like the load balancer are rarely the best deep dive unless asked.
Q5. Should you mention specific products like Redis or Cassandra?
Yes, as examples, but always lead with the role: "a distributed in-memory cache, such as Redis". This shows you understand the category and makes your design portable. Be ready to explain why that kind of product fits.
Q6. What do you do if the interviewer disagrees with your choice?
Listen, ask what concern they have, and compare both options against the requirements. If their option is better, say so and adapt. If you still prefer yours, explain the trade-off clearly and respectfully. Interviewers often push back to test reasoning, not because you are wrong.
Q7. How is a junior candidate evaluated differently from a senior one?
Juniors are expected to produce a correct, simple end-to-end design, use standard building blocks sensibly and discuss main trade-offs, with some guidance. Seniors must drive the session, go deep on several components, quantify trade-offs and discuss failure modes, operations and evolution. The framework is the same; the expected depth differs.
Q8. Is it bad to start with a single server design?
No, if you say why and then scale it based on numbers. Starting simple shows you understand that complexity has a cost. The mistake is either staying simple when the numbers demand more, or starting with a giant architecture that the numbers do not justify.
Q9. How should you handle non-functional requirements like availability?
Turn them into concrete targets and design choices: "99.9% availability means under about 9 hours of downtime a year, so every tier needs redundancy and the database needs automatic failover." Then check your design against them during the bottleneck step.
Q10. What makes a high-level design "complete"?
Every core functional requirement has a traceable path from the client through the components to storage and back, with no magic boxes. You can walk the read and write paths aloud. Components that matter for the non-functional requirements, such as replication or caching, are present with reasons.
Q11. How do you show trade-off thinking?
For each important choice, name at least one alternative, compare them against the requirements, choose, and state the cost of your choice and how you mitigate it. A compact comparison table on the board is an efficient way to do this.
Q12. What should you do if you run out of time?
Prioritise a complete high-level design over perfect detail, then summarise what you would deep dive on next and why. A clear summary of remaining risks and next steps shows judgment even when time runs short.
Key takeaways
- A system design interview tests reasoning, structure and communication, not one correct answer.
- Use the eight steps in order: requirements, estimates, API, data model, high-level design, deep dives, bottlenecks, wrap-up.
- Budget roughly 5, 4, 4, 4, 10, 12, 4 and 2 minutes in a 45-minute interview, and always finish the high-level design.
- Clarify and scope first; turn vague goals like "fast" into concrete targets.
- Every component needs a reason; every important choice needs an alternative and a trade-off.
- Think aloud, use hints, and treat pushback as a discussion.
- Juniors pass with a clean, reasoned, simple design; seniors must lead and go deep.
- Practise case studies with the framework and timed mock interviews.
Next lesson
Continue with System design fundamentals.

