Quick reference for scaling axes, replication topologies, partitioning strategies, consistent hashing, and hotspot fixes.
ShardingReplicationConsistent HashingAKF Cube
Scaling Axes
Vertical vs Horizontal
| Vertical (up) | Horizontal (out) |
| Bigger box, no code change | More boxes + LB |
| Hard ceiling, SPOF | Near-unbounded, fault tolerant |
| Simple | Needs stateless services |
AKF Scale Cube
X clone identical instances (LB across copies)
Y split by function/service (microservices)
Z split by data attribute (sharding, by region)
Replication
| Topology | Notes |
| Leader-follower | 1 writer, N readers; leader SPOF, lag |
| Multi-leader | Multi-region writes; conflict resolution |
| Leaderless | Quorum R+W>N (Dynamo/Cassandra) |
sync = durable, slower async = fast, may lose on failover
read replicas: send reads to followers (read-heavy wins)
replica lag -> read-your-writes: read from leader briefly
web read:write ratio ~ 10:1 (varies)
Sharding / Partitioning
| Strategy | Pro / Con |
| Range | range scans; sequential-key hotspots |
| Hash % N | even spread; resharding remaps everything |
| Consistent hash | only K/N keys move on resize |
Consistent Hashing Ring
0 / 2^32
. A
k4 . . k1
D . . B -> key walks clockwise to next node
k3 . . k2
. C
add/remove node -> only adjacent arc's keys move (~K/N)
use 100-200 VIRTUAL NODES per physical node to even out load
Caching Patterns
| Pattern | Behavior |
| Cache-aside | App reads cache, falls back to DB, populates |
| Write-through | Write cache + DB together (consistent, slower) |
| Write-back | Write cache, async flush to DB (fast, risky) |
CDN -> static + edge-cached content, ~20 ms vs 150 ms origin
TTL -> always set one; plan invalidation
async -> queue slow work (email, resize, fan-out) to workers
denorm -> duplicate data for one-hop reads; sync on write
Hotspots & Capacity
Celebrity / Hotspot Fixes
| Fix | Use |
| Key salting | Spread one hot key across shards |
| Dedicated cache | Serve hot items from replicated cache |
| Hybrid fan-out | Push for normal, pull for celebrities |
| Request coalescing | Collapse duplicate concurrent misses |
Capacity Rule of Thumb
servers = peak_QPS / (cores / per_req_seconds)
e.g. 30k / (16 / 0.005) = 30k / 3200 ~= 10
add 30-50% headroom (spikes, deploys, node loss)
always provision for PEAK, never average