What a Load Balancer Does
A load balancer (LB) sits between clients and a pool of backend servers and distributes incoming traffic across them. It is the linchpin of horizontal scaling: it lets many identical servers act as one endpoint, spreads load so no single node is overwhelmed, and routes around failed instances.
+---------------------+
clients ----> | Load Balancer | ----> server 1
| (VIP: 203.0.113.10) | ----> server 2
+---------------------+ ----> server 3
|
health checks remove dead nodes,
add capacity is transparent to clients
Beyond distribution, LBs provide health checking, SSL termination, session persistence, and a stable virtual IP (VIP) so clients never need to know individual server addresses.
Layer 4 vs Layer 7
Load balancers operate at different layers of the network stack, which determines how much they understand about traffic.
| Aspect | L4 (transport) | L7 (application) |
|---|---|---|
| Sees | IP + TCP/UDP port only | Full HTTP: URL, headers, cookies |
| Routing | By connection / 5-tuple | By path, host, header, cookie |
| Speed | Very fast, low overhead | Slower — parses payload |
| Features | Minimal | SSL term, content routing, WAF, rewrites |
| AWS example | NLB | ALB |
When to use which
L4 for raw throughput, non-HTTP protocols, or ultra-low latency (millions of connections). L7 when you need path/host routing, TLS termination, or per-request logic — which is most web traffic.
Load Balancing Algorithms
| Algorithm | How it picks a backend |
|---|---|
| Round robin | Cycle through servers in order — simple, assumes equal capacity |
| Weighted round robin | Round robin biased by server capacity weights |
| Least connections | Send to server with fewest active connections — good for long-lived |
| Least response time | Fewest connections + lowest latency; adapts to slow nodes |
| IP hash | hash(client IP) → same client always to same server (sticky) |
| Consistent hashing | Ring-based; minimizes remap on server changes (cache affinity) |
Round robin: r1 r2 r3 r1 r2 r3 ... (equal, stateless)
Least conn: pick min(active_conns) (uneven request cost)
IP hash: server = hash(ip) % N (stickiness, no shared state)
Consistent hash: ring position (cache/shard affinity)
Rule of thumb: round robin for uniform, short requests; least connections/response time when request cost varies; hashing when you need a client to keep landing on the same backend (cache locality or session affinity).
Health Checks
An LB must know which backends are alive. It periodically probes each server and stops routing to any that fail. Health checks are what make an LB fault-tolerant rather than just a distributor.
| Type | Checks |
|---|---|
| Passive | Watches live traffic; ejects a node after N real failures |
| Active (TCP) | Can we open a connection to the port? |
| Active (HTTP) | GET /healthz → expect 200 within a timeout |
Design the health endpoint carefully
A shallow check only proves the process is up. A deep check (DB reachable, dependencies OK) is more accurate but can cause cascading ejections if a shared dependency blips. Use readiness vs liveness distinctions.
Session Persistence (Sticky Sessions)
If a backend stores per-user state in local memory, requests from that user must return to the same server. LBs achieve this stickiness via a cookie the LB sets, or by IP hashing.
Cookie-based: LB sets a cookie -> routes that cookie to server 2
IP-based: hash(client IP) -> same server (breaks behind NAT/proxy)
Prefer stateless over sticky
Sticky sessions cause uneven load, break autoscaling, and lose state when a node dies. The better fix is to externalize session state (Redis / JWT) so any server can serve any request.
SSL / TLS Termination
TLS handshakes are CPU-intensive. Terminating TLS at the LB decrypts once, hands plaintext (or re-encrypted traffic) to backends, and centralizes certificate management.
Termination: client --TLS--> [LB decrypts] --HTTP--> backend
+ offloads crypto, central certs, L7 inspection
- plaintext on internal network
Passthrough: client --TLS----------------------> backend (LB just forwards)
+ true end-to-end encryption
- LB can't inspect L7
Re-encryption: LB terminates then opens a new TLS to backend (best of both)
Reverse Proxy vs Load Balancer
The two overlap and are often the same box, but the intent differs. A reverse proxy is a front door that forwards requests to backends (adding caching, TLS, compression, routing) even to a single server. A load balancer specifically spreads traffic across many backends. Nginx/Envoy do both.
Global Server Load Balancing (GSLB)
A single-region LB can't help users on another continent — physics adds ~150 ms. GSLB directs users to the nearest healthy region.
| Technique | How it routes |
|---|---|
| DNS-based LB | Return a region-specific IP based on client location; TTL controls agility |
| Anycast | One IP announced from many sites; BGP routes to nearest PoP |
| GeoDNS + latency | Route by measured latency / geography to the best datacenter |
EU user --DNS query--> GSLB --> returns EU LB IP --> EU region
US user --DNS query--> GSLB --> returns US LB IP --> US region
region down? GSLB drops it from DNS / anycast withdraws the route
DNS LB caveat: resolvers cache records for the TTL, so failover is not instant. Anycast reacts faster (route-level) and is how CDNs and DNS providers absorb DDoS and steer to the nearest edge.
Load Balancer High Availability
The LB itself must not be a single point of failure. It is deployed redundantly.
| Topology | Behavior |
|---|---|
| Active-passive | Standby takes the VIP if primary fails (failover); half the capacity idle |
| Active-active | All LBs serve traffic; more capacity, no idle standby |
VIP + keepalived (VRRP):
LB-A (MASTER) holds VIP 203.0.113.10 <-- clients hit this
LB-B (BACKUP) heartbeats with LB-A
LB-A dies -> LB-B claims the VIP (gratuitous ARP) in ~seconds
Cloud reality
Managed LBs (AWS ELB/ALB/NLB, GCP, Azure) are themselves horizontally scaled and multi-AZ, so you rarely run keepalived yourself — but the same active-active principle is running under the hood.
Hardware vs Software & Common Tools
Dedicated hardware LBs (F5, Citrix) offer maximum throughput at high cost and low flexibility. Modern stacks overwhelmingly use software LBs and managed cloud offerings, which are cheaper, programmable, and scale elastically.
| Tool | Strength |
|---|---|
| HAProxy | Battle-tested L4/L7; extremely high performance |
| Nginx | Reverse proxy + LB + web server + TLS; ubiquitous |
| Envoy | Cloud-native L7, xDS dynamic config, service-mesh data plane |
| AWS ELB family | ALB (L7), NLB (L4, static IP), GLB (appliances); fully managed |