Learn/system design/Load Balancing
Beginner~25 min read

Load Balancing

L4 vs L7 balancing, algorithms, health checks, session persistence, SSL termination, global load balancing, and LB high availability.

L4 vs L7AlgorithmsHealth ChecksGSLB

What a Load Balancer Does

A load balancer (LB) sits between clients and a pool of backend servers and distributes incoming traffic across them. It is the linchpin of horizontal scaling: it lets many identical servers act as one endpoint, spreads load so no single node is overwhelmed, and routes around failed instances.

text
                 +---------------------+
   clients  ----> |    Load Balancer    | ----> server 1
                  | (VIP: 203.0.113.10) | ----> server 2
                  +---------------------+ ----> server 3
                          |
              health checks remove dead nodes,
              add capacity is transparent to clients

Beyond distribution, LBs provide health checking, SSL termination, session persistence, and a stable virtual IP (VIP) so clients never need to know individual server addresses.

Layer 4 vs Layer 7

Load balancers operate at different layers of the network stack, which determines how much they understand about traffic.

AspectL4 (transport)L7 (application)
SeesIP + TCP/UDP port onlyFull HTTP: URL, headers, cookies
RoutingBy connection / 5-tupleBy path, host, header, cookie
SpeedVery fast, low overheadSlower — parses payload
FeaturesMinimalSSL term, content routing, WAF, rewrites
AWS exampleNLBALB

When to use which

L4 for raw throughput, non-HTTP protocols, or ultra-low latency (millions of connections). L7 when you need path/host routing, TLS termination, or per-request logic — which is most web traffic.

Load Balancing Algorithms

AlgorithmHow it picks a backend
Round robinCycle through servers in order — simple, assumes equal capacity
Weighted round robinRound robin biased by server capacity weights
Least connectionsSend to server with fewest active connections — good for long-lived
Least response timeFewest connections + lowest latency; adapts to slow nodes
IP hashhash(client IP) → same client always to same server (sticky)
Consistent hashingRing-based; minimizes remap on server changes (cache affinity)
text
Round robin:      r1 r2 r3 r1 r2 r3 ...   (equal, stateless)
Least conn:       pick min(active_conns)   (uneven request cost)
IP hash:          server = hash(ip) % N    (stickiness, no shared state)
Consistent hash:  ring position            (cache/shard affinity)

Rule of thumb: round robin for uniform, short requests; least connections/response time when request cost varies; hashing when you need a client to keep landing on the same backend (cache locality or session affinity).

Health Checks

An LB must know which backends are alive. It periodically probes each server and stops routing to any that fail. Health checks are what make an LB fault-tolerant rather than just a distributor.

TypeChecks
PassiveWatches live traffic; ejects a node after N real failures
Active (TCP)Can we open a connection to the port?
Active (HTTP)GET /healthz → expect 200 within a timeout

Design the health endpoint carefully

A shallow check only proves the process is up. A deep check (DB reachable, dependencies OK) is more accurate but can cause cascading ejections if a shared dependency blips. Use readiness vs liveness distinctions.

Session Persistence (Sticky Sessions)

If a backend stores per-user state in local memory, requests from that user must return to the same server. LBs achieve this stickiness via a cookie the LB sets, or by IP hashing.

text
Cookie-based: LB sets a cookie -> routes that cookie to server 2
IP-based:     hash(client IP)   -> same server (breaks behind NAT/proxy)

Prefer stateless over sticky

Sticky sessions cause uneven load, break autoscaling, and lose state when a node dies. The better fix is to externalize session state (Redis / JWT) so any server can serve any request.

SSL / TLS Termination

TLS handshakes are CPU-intensive. Terminating TLS at the LB decrypts once, hands plaintext (or re-encrypted traffic) to backends, and centralizes certificate management.

text
Termination:   client --TLS--> [LB decrypts] --HTTP--> backend
               + offloads crypto, central certs, L7 inspection
               - plaintext on internal network

Passthrough:   client --TLS----------------------> backend (LB just forwards)
               + true end-to-end encryption
               - LB can't inspect L7

Re-encryption: LB terminates then opens a new TLS to backend (best of both)

Reverse Proxy vs Load Balancer

The two overlap and are often the same box, but the intent differs. A reverse proxy is a front door that forwards requests to backends (adding caching, TLS, compression, routing) even to a single server. A load balancer specifically spreads traffic across many backends. Nginx/Envoy do both.

Global Server Load Balancing (GSLB)

A single-region LB can't help users on another continent — physics adds ~150 ms. GSLB directs users to the nearest healthy region.

TechniqueHow it routes
DNS-based LBReturn a region-specific IP based on client location; TTL controls agility
AnycastOne IP announced from many sites; BGP routes to nearest PoP
GeoDNS + latencyRoute by measured latency / geography to the best datacenter
text
EU user --DNS query--> GSLB --> returns EU LB IP  --> EU region
US user --DNS query--> GSLB --> returns US LB IP  --> US region
region down? GSLB drops it from DNS / anycast withdraws the route

DNS LB caveat: resolvers cache records for the TTL, so failover is not instant. Anycast reacts faster (route-level) and is how CDNs and DNS providers absorb DDoS and steer to the nearest edge.

Load Balancer High Availability

The LB itself must not be a single point of failure. It is deployed redundantly.

TopologyBehavior
Active-passiveStandby takes the VIP if primary fails (failover); half the capacity idle
Active-activeAll LBs serve traffic; more capacity, no idle standby
text
VIP + keepalived (VRRP):
  LB-A (MASTER) holds VIP 203.0.113.10  <-- clients hit this
  LB-B (BACKUP) heartbeats with LB-A
  LB-A dies -> LB-B claims the VIP (gratuitous ARP) in ~seconds

Cloud reality

Managed LBs (AWS ELB/ALB/NLB, GCP, Azure) are themselves horizontally scaled and multi-AZ, so you rarely run keepalived yourself — but the same active-active principle is running under the hood.

Hardware vs Software & Common Tools

Dedicated hardware LBs (F5, Citrix) offer maximum throughput at high cost and low flexibility. Modern stacks overwhelmingly use software LBs and managed cloud offerings, which are cheaper, programmable, and scale elastically.

ToolStrength
HAProxyBattle-tested L4/L7; extremely high performance
NginxReverse proxy + LB + web server + TLS; ubiquitous
EnvoyCloud-native L7, xDS dynamic config, service-mesh data plane
AWS ELB familyALB (L7), NLB (L4, static IP), GLB (appliances); fully managed

Section navigation