Learn the building blocks, then design a system
This course combines fundamentals, practical distributed-systems reasoning and end-to-end design workshops. Use the sidebar as your syllabus. Start with requirements and estimates, or find a concept in the index below.
The curriculum develops 40 core concepts alongside networking, security, consensus, observability and production recovery. System design is an open-ended field; this is a substantial interview and engineering foundation, not a claim that every possible specialization is covered.
Study in this order
| Phase | Lessons | Produce this evidence |
|---|---|---|
| Scope and requests | Fundamentals; networking; API design; identity | Requirements, capacity estimate and an authorized request trace |
| Data and scale | Database scaling; caching; scalability; storage/search/geo | Schema, indexes, cache policy and partition key |
| Events and coordination | Queues; Kafka; real-time/webhooks; transactions/CDC; consensus | Delivery contract and crash/recovery timeline |
| Operating systems at scale | Microservices; reliability; observability | SLO, overload policy and tested recovery plan |
| Apply and review | Worked designs; design workshops | Ten designs with alternatives and failure cases |
A four-week example: spend week one on request/identity/data basics, week two on cache/partition/storage, week three on events/correctness/coordination, and week four on operations and timed workshops. Choose pace by your prior knowledge rather than treating four weeks as a guarantee.
Interactive experiments
- Kafka partition and consumer lab: publish, commit, crash, replay and stop brokers.
- Consensus experiment: remove voters and observe the majority boundary.
- Bloom filter experiment: observe shared bits and probable membership.
The simulations are explicitly simplified browser models, not production benchmarks or live clusters.
Find a concept
Use this index to find the lesson that develops each concept. Related topics are linked within the lessons.
| Concept | Read |
|---|---|
| APIs | APIs |
| API Gateways | API Gateways |
| REST vs GraphQL | REST vs GraphQL |
| JWTs | JWTs |
| Webhooks | Webhooks |
| Long Polling vs WebSockets | Long Polling vs WebSockets |
| Load Balancing | Load Balancing |
| Proxy vs Reverse Proxy | Proxy vs Reverse Proxy |
| CDN | CDN |
| Rate Limiting | Rate Limiting |
| Caching | Caching |
| Distributed Caching | Distributed Caching |
| Caching Strategies | Caching Strategies |
| Cache Eviction Policies | Cache Eviction Policies |
| Consistent Hashing | Consistent Hashing |
| SQL vs NoSQL | SQL vs NoSQL |
| ACID Transactions | ACID Transactions |
| Database Index | Database Index |
| Database Sharding | Database Sharding |
| Database Scaling | Database Scaling |
| Data Replication | Data Replication |
| Data Redundancy | Data Redundancy |
| Change Data Capture | Change Data Capture |
| CAP Theorem | CAP Theorem |
| Strong vs Eventual Consistency | Strong vs Eventual Consistency |
| Scalability | Scalability |
| Availability | Availability |
| Single Point of Failure | Single Point of Failure |
| Latency vs Throughput | Latency vs Throughput |
| Stateful vs Stateless | Stateful vs Stateless |
| Message Queues | Message Queues |
| Circuit Breaker | Circuit Breaker |
| Idempotency | Idempotency |
| Fault Tolerance | Fault Tolerance |
| Disaster Recovery | Disaster Recovery |
| Distributed Locking | Distributed Locking |
| Bloom Filters | Bloom Filters |
| Concurrency vs Parallelism | Concurrency vs Parallelism |
| Batch vs Stream Processing | Batch vs Stream Processing |
| Geohashing | Geohashing |
Advanced topics
| Area | Additional coverage |
|---|---|
| Communication | DNS/TCP/TLS/HTTP, connection reuse, SSE, gRPC, reconnect cursors, backpressure |
| Correctness | Isolation anomalies, optimistic concurrency, outbox/inbox, CDC, sagas, 2PC, CQRS/event sourcing |
| Coordination | Quorums, Raft, election epochs, logical clocks, leases/fencing, split brain and CRDT boundaries |
| Storage | B-trees/LSM/WAL, inverted indexes, object storage, compaction, approximate membership, spatial candidates |
| Security | OAuth/OIDC/PKCE, token rotation, revocation, CSRF/XSS, tenant isolation and service identity |
| Operations | Deadline budgets, retry amplification, bulkheads, admission control, SLOs, traces, profiles and recovery drills |
| Architecture | Monolith vs services, discovery, service mesh, control/data planes, rollout safety and cost |
For specialized work, continue into database implementation, formal protocol verification, GPU/ML infrastructure, real-time media, regulatory security or domain-specific systems. The appropriate depth depends on the role.
Review before an interview
Take one requirement and explain its effect on API, data, latency and failure behavior. Calculate one capacity number with assumptions. Walk through a crash between each pair of writes. Explain who owns retries, where duplicate prevention is atomic and how the system recovers. Finish by naming the simplest viable design and what would justify the next layer of complexity.