Design fundamentals · 2 / 4
Design a distributed rate limiter
Your challenge
Limit an account to a burst of 20 requests and an average of 5 requests/second across multiple API servers.
Try it first. Write down your assumptions and explain your reasoning.
1.Choose the policy before the algorithm
The requested burst and sustained rate fit a token bucket: capacity 20, refill 5 tokens/second, one token consumed per request. Identify the authenticated account on the server; an arbitrary account id or untrusted forwarding header must not choose the bucket.
2.Make shared updates atomic
Store the token count and last refill time in a shared store. Calculate refill and consume inside one atomic operation, such as a Redis Lua script. Use a consistent time source and expiry for inactive buckets. Per-instance buckets would multiply the allowed traffic across servers.
3.Explain failure and boundary decisions
Return 429 with a useful retry indication when empty. Decide whether an unavailable limiter fails open or closed based on the protected operation: a public read and an expensive payment action may need different policies. A single bucket can become hot; partition by account and discuss multi-region consistency versus latency.
Take it one step further
- 1.How is a fixed-window counter different near a boundary?
- 2.What changes for a global shared quota?
Self-review
Can you explain each point without looking at the solution?
- Correct capacity and refill semantics
- Atomic shared state
- Trust boundaries and failure policy
