Learn/System Design/Capacity, SLOs, Autoscaling & Cost
Advanced~18 min read

Capacity, SLOs, Autoscaling & Cost

Calculate failure headroom, tail latency, burn rates, backlog recovery and cost per useful operation.

System DesignDistributed SystemsProduction

Capacity belongs to a workload

Requests/second is incomplete without payload sizes, read/write mix, cache behavior, fan-out and latency objectives. Measure sustainable throughput while meeting the objective, not the highest burst rate before failure. Keep peak, average and failure-mode load separate.

For every dependency record service rate, saturation indicator, concurrency limit and queue boundary. An application instance can look idle while every request waits for the same database pool.

Estimate from demand

Assume 100,000 daily users, six sessions/user/day and eight API requests/session. That is 4.8 million requests/day, about 56 requests/second on average. A stated 12x peak factor gives about 667 requests/second. This factor is an assumption to validate, not a universal traffic law.

At 30 KB average response payload, peak application egress is about 20 MB/second before protocol overhead and static assets. Separate CDN delivery from application egress; adding them together can suggest the wrong bottleneck. Include retries and internal fan-out when sizing downstream load.

Queue stability and recovery

Let arrival rate be 800 tasks/second and service rate 1,000. A 600,000-task backlog drains at a net 200 tasks/second while arrivals continue: about 50 minutes. Dividing backlog by 1,000 incorrectly assumes no new work.

When arrivals exceed service rate, the queue cannot recover without throttling, shedding, pausing producers or increasing effective capacity. Oldest-item age often describes user harm better than depth. Define expiry and priority so stale disposable work does not consume the entire recovery window.

Little's Law with boundaries

Average concurrency equals throughput times average residence time for a stable boundary. At 667 requests/second and 150 ms average residence, about 100 requests are in flight. Include queue waiting if it lies inside the measured boundary.

This is an average relation, not an instruction to set every pool to 100. Pools serve different portions of a request and have different bottlenecks. Load tests reveal the latency curve and safe headroom.

Tail latency and fan-out

If a request waits for 100 independent branches and each has a 1% probability of exceeding a threshold, the probability that at least one exceeds it is 1 - 0.99^100, about 63.4%. Real dependencies are often correlated, so this is an illustrative model rather than a prediction.

Reduce unnecessary fan-out, bound concurrency and consider partial results for optional components. Hedged requests can reduce tails for safe reads but consume extra capacity; launch only under a measured policy and cancel redundant work. Do not hedge a non-idempotent payment mutation.

SLO and burn-rate calculation

Suppose eligible requests have a 99.9% success SLO over 30 days. The allowed error fraction is 0.1%. Observed 1% errors consume the budget at 10x the sustainable rate. For steady traffic, one hour at 10x uses about 1.39% of a 30-day budget: 10 * 1 / 720.

Use several windows to distinguish urgent degradation from noise, and make each alert actionable. Define eligible requests and whether rejected excess traffic counts. A time-based uptime objective and a request-success objective can give different results during a high-traffic outage.

Autoscaling is delayed control

A scaler observes a signal, chooses capacity, provisions resources and waits for readiness. Traffic can change throughout that delay. CPU-only scaling misses queue waits, I/O-bound work and blocked connection pools. Queue-based scaling should consider age and service rate, not just task count.

Use minimum warm capacity, maximum limits, stabilization and cooldown behavior appropriate to the platform. Predictable events can justify scheduled pre-scaling. Adding workers without enough database capacity increases contention instead of throughput.

Survive a zone loss

Three zones each carry 2,000 requests/second. Losing one requires the other two to carry 3,000 each. If their sustainable limit is 2,500 each, the service lacks failure headroom despite looking healthy normally.

Specify whether the objective allows load shedding, delayed batch work or reduced optional features. Cache cold starts and reconnect storms can add demand at the worst time. Test those conditions, not only the steady two-zone state.

Cost per useful result

Calculate compute, storage, replicas, requests, egress, queue operations, logging and external API/model usage. Divide by successful useful operations rather than raw attempts. Retrying a failing dependency can increase cost while completed work decreases.

For illustration, 2 TB/month public downloads served from an origin have a different cost shape from 2 TB delivered through a caching edge; exact prices depend on deployment. Evaluate locality, hit ratio and contractual pricing before optimizing. Compress and batch where latency permits, and eliminate needless payloads before adding machines.

Performance tests that answer decisions

Run stepped load for the saturation curve, spikes for admission behavior, soak tests for leaks, and failure tests for headroom. Include hot keys, cold caches, realistic payloads, slow clients and retries. Record offered load as well as achieved throughput so the generator cannot hide demand while blocked.

Use representative environments and report assumptions. A small local benchmark can validate a calculation or algorithm, but it is not evidence of production fleet capacity. Compare before/after at the same useful workload and SLO.

Exercise

Size a college portal for a placement-test announcement. Estimate registrations, burst traffic, email queue drain and DB connections. Add one-zone failure, a cold-cache deploy and an email-provider outage. Produce an admission policy, alert plan, scale trigger and monthly cost model with assumptions.

Section navigation

Keep it in your account.

Your progress is saved securely and available when you return.

Sign in Create a free account