contentintech

Scalability Cheatsheet

Quick reference for scaling axes, replication topologies, partitioning strategies, consistent hashing, and hotspot fixes.

ShardingReplicationConsistent HashingAKF Cube
NotesCheatsheet

Scaling Axes

Vertical vs Horizontal

Vertical (up)Horizontal (out)
Bigger box, no code changeMore boxes + LB
Hard ceiling, SPOFNear-unbounded, fault tolerant
SimpleNeeds stateless services

AKF Scale Cube

X  clone identical instances   (LB across copies)
Y  split by function/service   (microservices)
Z  split by data attribute     (sharding, by region)

Replication

TopologyNotes
Leader-follower1 writer, N readers; leader SPOF, lag
Multi-leaderMulti-region writes; conflict resolution
LeaderlessQuorum R+W>N (Dynamo/Cassandra)
sync  = durable, slower   async = fast, may lose on failover
read replicas: send reads to followers (read-heavy wins)
replica lag -> read-your-writes: read from leader briefly
web read:write ratio ~ 10:1 (varies)

Sharding / Partitioning

StrategyPro / Con
Rangerange scans; sequential-key hotspots
Hash % Neven spread; resharding remaps everything
Consistent hashonly K/N keys move on resize

Consistent Hashing Ring

         0 / 2^32
           . A
      k4 .      . k1
   D .              . B   -> key walks clockwise to next node
      k3 .      . k2
           . C

add/remove node -> only adjacent arc's keys move (~K/N)
use 100-200 VIRTUAL NODES per physical node to even out load

Caching Patterns

PatternBehavior
Cache-asideApp reads cache, falls back to DB, populates
Write-throughWrite cache + DB together (consistent, slower)
Write-backWrite cache, async flush to DB (fast, risky)
CDN     -> static + edge-cached content, ~20 ms vs 150 ms origin
TTL     -> always set one; plan invalidation
async   -> queue slow work (email, resize, fan-out) to workers
denorm  -> duplicate data for one-hop reads; sync on write

Hotspots & Capacity

Celebrity / Hotspot Fixes

FixUse
Key saltingSpread one hot key across shards
Dedicated cacheServe hot items from replicated cache
Hybrid fan-outPush for normal, pull for celebrities
Request coalescingCollapse duplicate concurrent misses

Capacity Rule of Thumb

servers = peak_QPS / (cores / per_req_seconds)
  e.g. 30k / (16 / 0.005) = 30k / 3200 ~= 10
add 30-50% headroom (spikes, deploys, node loss)
always provision for PEAK, never average

Section navigation