The earlier lessons in this track explained the protocols. This lesson is about how those protocols are used in real production systems, and how engineers debug them when something breaks. The first half covers the infrastructure every modern web service sits behind: CDNs, load balancers, proxies, API gateways, NAT traversal, software-defined networking and data-centre network design. The second half is a practical toolkit: what ping, traceroute, ss, dig, curl -v, tcpdump and Wireshark show you, what to look for in their output, and a worked "the site is slow" investigation.
Interviewers ask about this material in two ways. Conceptual questions: "L4 versus L7 load balancer", "forward versus reverse proxy", "how does a CDN decide what to cache", "what load-balancing algorithm would you choose". And practical ones, especially in SRE, DevOps and backend roles: "a user says the site is slow, how do you debug it", "how would you check if a port is open", "how do you find which process is using port 8080". Being able to name the right tool and read its output signals real experience.
CDNs: content delivery networks
A CDN (content delivery network) is a large set of servers, called edge servers or points of presence (PoPs), spread across many cities. Users connect to a nearby edge instead of your origin (your own servers). The edge serves cached content directly and forwards the rest to the origin.
Users in Mumbai Users in London Users in Sao Paulo
| | |
[Edge: Mumbai] [Edge: London] [Edge: Sao Paulo]
\ | /
\______ cache miss: fetch from origin ________/
|
[Origin: us-east servers]
Why CDNs make sites faster
- Shorter round trips. The TCP and TLS handshakes happen with an edge maybe 5 ms away instead of an origin 200 ms away. Since a new HTTPS connection needs two or more round trips, this alone saves hundreds of milliseconds (see the round-trip table in what happens when you type a URL).
- Edge caching. Static files (images, CSS, JavaScript, video segments) and even some HTML are served from the edge's cache without touching the origin.
- Origin offload. If 95 percent of requests are cache hits, the origin handles only 5 percent of the traffic.
- Warm, optimised backbone. Edges keep persistent connections to the origin and route over the CDN's own network.
- Protection. The CDN absorbs DDoS traffic across its huge capacity and often runs a WAF in front of your site.
How users reach the nearest edge
Two common techniques:
- DNS-based: your domain is a CNAME to the CDN's host name. The CDN's DNS returns the IP of an edge close to the user's resolver (see the DNS lesson).
- Anycast: every edge announces the same IP address through BGP. The Internet's routing naturally delivers each user's packets to the topologically nearest edge (see the routing lesson).
Pull versus push CDNs
| Pull CDN | Push CDN | |
|---|---|---|
| How content arrives | Edge fetches from origin on the first request (cache miss), then caches it | You upload content to the CDN's storage ahead of time |
| First request | Slower (miss goes to origin) | Fast, already there |
| Effort | Low: just point DNS at the CDN | You manage uploads and updates |
| Storage | Only what users actually request | Everything you push, whether requested or not |
| Best for | Websites, APIs, frequently changing content | Large, rarely changing files: video libraries, game patches, software downloads |
Most web traffic uses pull CDNs. A push model is common for large media where you want everything pre-positioned before a launch.
How edge caching is controlled
The CDN decides what and how long to cache mainly from HTTP headers sent by your origin (explained in the HTTP lesson):
Cache-Control: public, max-age=31536000, immutablefor fingerprinted static files (app.3f9a1c.js).Cache-Control: public, s-maxage=60to cache HTML at the edge for 60 seconds while browsers use their ownmax-age.Cache-Control: privateorno-storefor personalised pages, so the CDN never serves one user's page to another.Vary: Accept-Encodingto store separate compressed versions.
The cache key is what the CDN uses to decide whether two requests are "the same": by default, host plus path plus query string. Including irrelevant query parameters (such as tracking tags like utm_source) in the key causes needless misses, while leaving out a parameter that changes the response causes users to see the wrong content.
Invalidation (purge): when content changes before its TTL expires, you ask the CDN to purge it by URL, by prefix or by tag. Purges take seconds to propagate across the network. The cleaner strategy for static assets is to change the file name on every deploy, so no purge is needed.
Useful metrics: cache hit ratio (hits divided by total requests) and origin offload (bytes served from cache). Response headers such as X-Cache: HIT, cf-cache-status: HIT or Age: 342 tell you whether a particular response came from cache.
Common mistake
Caching a page that contains personal data at the CDN because the origin forgot Cache-Control: private. One user's account page then gets served to everyone. Always mark personalised responses private or no-store, and make sure Set-Cookie responses are never cached publicly.
Load balancers
A load balancer (LB) sits in front of a group of servers (a pool or target group) and distributes incoming requests among them. It provides:
- Scalability: add servers behind it, and clients still use one address.
- Availability: health checks (periodic probes such as
GET /healthzevery 10 seconds) detect failed servers, and the LB stops sending them traffic. - Zero-downtime deploys: drain one server at a time (finish its existing requests, send it no new ones), update it, return it to the pool.
Layer 4 versus Layer 7
The "layer" refers to the OSI model (see the OSI and TCP/IP models lesson).
A Layer 4 (L4) load balancer works at the transport layer. It sees IP addresses and TCP/UDP ports, not the content. It picks a backend when a connection starts and forwards all packets of that connection there.
A Layer 7 (L7) load balancer works at the application layer. It terminates the client's TCP and TLS connection, reads the HTTP request, and makes decisions based on its content: path, host name, headers, cookies. It then opens (or reuses) its own connection to a backend.
L4: client ==TCP==> [LB: forward packets by IP:port] ==same TCP flow==> backend
L7: client ==TLS/HTTP==> [LB: terminate, read request] ==new HTTP==> backend
route /api -> api pool
route /img -> image pool
| L4 | L7 | |
|---|---|---|
| Sees | IPs, ports, TCP/UDP | Full HTTP: URL, headers, cookies, body |
| Routing decisions | Per connection | Per request |
| TLS | Usually passed through to backends | Usually terminated at the LB |
| Speed and cost | Very fast, little CPU, millions of connections | More CPU per request |
| Features | Simple balancing, any TCP/UDP protocol (databases, games, MQTT) | Path and host routing, header rewrites, retries, auth, rate limiting, caching, compression, WebSocket and gRPC awareness |
| Examples | AWS Network Load Balancer, Linux IPVS, HAProxy in TCP mode | AWS Application Load Balancer, Nginx, Envoy, HAProxy in HTTP mode |
A common architecture uses both: an L4 layer (often anycast) spreads traffic across a fleet of L7 proxies, which route to services.
Interview tip
Explain L7's key advantage with HTTP/2: since many requests share one long-lived connection, an L4 balancer sends all of them to one backend. An L7 balancer can spread individual requests across backends. This matters a lot for gRPC, which runs on HTTP/2.
Load-balancing algorithms
| Algorithm | How it picks | Good for | Watch out for |
|---|---|---|---|
| Round robin | Next server in order | Identical servers, similar requests | Ignores current load |
| Weighted round robin | In proportion to weights (a 2x server gets 2x traffic) | Mixed server sizes, canary releases | Weights are static |
| Least connections | Server with fewest active connections | Long or variable-length requests, WebSockets | Needs connection tracking |
| Least response time | Fastest recent responses (and fewest connections) | Heterogeneous backends | Can herd onto a server that just got fast |
| Random with two choices | Pick two at random, send to the less loaded | Large fleets, many independent balancers | Slightly more complex |
| IP hash | Hash of client IP picks server | Simple stickiness without cookies | Uneven when many users share one NAT IP; reshuffles when servers change |
| Consistent hashing | Hash of a key (user id, URL) on a ring | Caches and shards, where the same key should hit the same server | Needs virtual nodes for even spread |
Worked example: round robin versus least connections
Three servers, A, B and C. Requests arrive in this order with these durations: R1 (10 s), R2 (1 s), R3 (1 s), R4 (10 s), R5 (1 s), R6 (1 s). Assume all six arrive within the first fraction of a second, so nothing finishes before the last assignment.
Round robin: R1 to A, R2 to B, R3 to C, R4 to A, R5 to B, R6 to C. Server A now has two 10-second requests (20 s of work), while B and C each have 2 s of work. A is overloaded while B and C sit idle after 2 seconds.
Least connections (ties broken in order A, B, C): R1 to A (A=1). R2 to B (B=1). R3 to C (C=1). R4: all have 1, tie goes to A (A=2). R5: B and C have 1, goes to B (B=2). R6 to C (C=2). Same result, because everything arrived at once.
Now change one assumption: requests arrive one second apart (R1 at t=0, R2 at t=1, and so on).
- t=0: R1 to A (A busy until 10).
- t=1: R2 to B (least: B=0, C=0, tie to B; B busy until 2).
- t=2: R3. B just finished. A=1, B=0, C=0, tie to B. B busy until 3.
- t=3: R4 (10 s). A=1, B=0, C=0, goes to B. B busy until 13.
- t=4: R5. A=1, B=1, C=0, goes to C.
- t=5: R6. C finished at 5; A=1, B=1, C=0, goes to C.
Round robin in the same timeline would have put R4 on A, behind R1, giving A 20 seconds of work. Least connections spread the two long requests across A and B. The lesson: round robin is fine when requests are uniform; least connections adapts when request durations vary.
Session persistence
Some applications store session state in server memory, so a user must keep hitting the same server. Sticky sessions (session affinity) achieve this with an LB-issued cookie or IP hashing. Stickiness hurts even distribution and loses sessions when a server dies. The better design is stateless servers with sessions in a shared store such as Redis, so any server can handle any request. The load balancing lesson in the system design track goes deeper.
Forward proxies, reverse proxies and API gateways
A proxy is an intermediary that makes requests on behalf of someone else. The difference between the two kinds is whom it represents.
FORWARD PROXY (acts for clients)
[client] --\
[client] ---> [forward proxy] ---> Internet ---> any server
[client] --/ (company or school network edge)
REVERSE PROXY (acts for servers)
any client ---> Internet ---> [reverse proxy] ---> [server]
(in front of ---> [server]
your servers) ---> [server]
Forward proxy: clients are configured to send their traffic through it. The destination server sees the proxy's IP, not the client's. Uses: corporate egress control and content filtering, caching for many users, logging, hiding client identity, reaching sites through a single allowed exit.
Reverse proxy: sits in front of servers; clients think they are talking to the real server. Uses: TLS termination, load balancing, caching, compression, serving static files, request buffering (protecting app servers from slow clients), hiding the internal topology, adding security headers, rate limiting. Nginx, HAProxy, Envoy, Caddy and Traefik are common. A CDN is a globally distributed reverse proxy; an L7 load balancer is a reverse proxy that focuses on distributing traffic.
| Forward proxy | Reverse proxy | |
|---|---|---|
| Represents | Clients | Servers |
| Configured by | Client side (browser or OS settings, PAC file) or transparently by the network | Server owner, via DNS pointing at it |
| Hides | Client identity from servers | Server topology from clients |
| Typical owner | Company IT, ISP | The website operator |
API gateways
An API gateway is a reverse proxy specialised for APIs, usually in front of many microservices. On top of routing, it handles cross-cutting concerns once instead of in every service:
- Authentication and authorisation (validate API keys, OAuth tokens or JWTs).
- Rate limiting and quotas per client or plan.
- Request routing and versioning (
/v1/ordersto one service,/v2/ordersto another). - Transformation: protocol translation (REST to gRPC), aggregating several backend calls into one response.
- Observability: logs, metrics, tracing for every call.
Examples: Kong, AWS API Gateway, Apigee, Envoy-based gateways. The risk is that the gateway becomes a bottleneck or a dumping ground for business logic; keep it to cross-cutting concerns. See API design and microservices.
NAT traversal: STUN and TURN
NAT (network address translation, from the IP addressing lesson) lets many private devices share one public IP. It works smoothly for outgoing connections, but it blocks unsolicited incoming ones: the router has no mapping for a packet nobody asked for. That is a problem for peer-to-peer applications like video calls, where two devices, each behind its own NAT, want to talk directly.
The standard toolkit, used by WebRTC (real-time audio and video in browsers):
- STUN (Session Traversal Utilities for NAT): a device asks a public STUN server "what IP and port do you see me as?". The answer is its public mapping, for example
49.36.12.8:61002. Peers exchange these addresses through a signalling server and then try sending packets directly to each other. Outgoing packets from both sides create NAT mappings that let the other side's packets in. This is called UDP hole punching. - TURN (Traversal Using Relays around NAT): when direct connection is impossible (for example symmetric NAT, which assigns a different public port for every destination, so the STUN-discovered mapping is useless for the peer), both peers send traffic through a public relay server. It always works but costs server bandwidth and adds latency.
- ICE (Interactive Connectivity Establishment): the procedure that gathers all candidate addresses (local, STUN-derived, TURN relay), tests pairs, and picks the best one that works.
In practice, most calls connect directly, and a minority fall back to TURN, which is why video-calling providers still run relay servers.
SDN and network virtualisation
Software-defined networking
A traditional router or switch contains two logical parts:
- The control plane: the "brain" that decides where traffic should go, running routing protocols (OSPF, BGP) and building tables.
- The data plane (forwarding plane): the "muscle" that moves each packet according to those tables, usually in specialised hardware at line rate.
In traditional networks, each device has its own control plane and is configured individually, often by hand through a command line. Changing network-wide behaviour means touching many boxes.
Software-defined networking (SDN) separates the two: a logically centralised controller (software) holds the network-wide view and computes forwarding rules, then programs simple switches through an API. OpenFlow was the first widely known protocol for this.
+--------------------------------+
| Applications (policy, TE, |
| firewall, monitoring) |
+---------------+----------------+
| northbound API (REST, etc.)
+---------------v----------------+
| SDN controller | control plane,
| (global network view) | centralised
+---+------------+-----------+---+
| southbound API (e.g. OpenFlow)
+---v--+ +---v--+ +--v---+
|switch|-----|switch|-----|switch| data plane:
+------+ +------+ +------+ match -> action rules
Benefits: network-wide policy from one place, programmability and automation, faster changes, and traffic engineering with a global view. Risks: the controller must be highly available, and it becomes a high-value target. Large cloud providers and data-centre operators run SDN-style designs; Google has publicly described SDN for its inter-data-centre WAN (B4).
Network virtualisation
Network virtualisation creates many isolated virtual networks on one shared physical network, the way virtual machines share one physical server (see virtualization and security).
- VLANs (from the data link layer lesson) split one switch fabric into separate broadcast domains using a 12-bit tag, so at most 4094 usable VLANs.
- Overlay networks such as VXLAN wrap a tenant's Ethernet frames inside UDP packets (port 4789) and carry them across an ordinary IP network (the underlay). VXLAN's 24-bit network identifier allows about 16 million virtual networks, enough for a public cloud with many customers.
- Cloud VPCs (virtual private clouds) give each customer their own private address space, subnets, route tables and firewall rules, all implemented in software on shared hardware. Two customers can both use
10.0.0.0/16without conflict. - Inside a host, virtual switches (Linux bridge, Open vSwitch) connect VMs and containers. Kubernetes networking (CNI plugins) builds pod networks with overlays or routed approaches.
NFV (network functions virtualisation) runs network functions that used to be dedicated appliances (firewalls, load balancers, NAT gateways) as software on ordinary servers.
Data-centre networking basics
From three-tier to leaf-spine
Traditional enterprise networks used a three-tier tree: access switches (connecting servers), aggregation switches, and core switches. This design suited north-south traffic (between users outside and servers inside). Modern applications generate mostly east-west traffic (server to server: microservices, databases, replication), and the tree's upper layers become bottlenecks. Spanning tree also disables redundant links to prevent loops, wasting capacity.
The modern design is leaf-spine, a form of Clos network:
+-------+ +-------+ +-------+ +-------+
|spine 1| |spine 2| |spine 3| |spine 4|
+-------+ +-------+ +-------+ +-------+
| \ \ \ / / | \ \ / / | \ \ / / / |
(every leaf connects to every spine)
+------+ +------+ +------+ +------+ +------+
|leaf 1| |leaf 2| |leaf 3| |leaf 4| |leaf 5|
+------+ +------+ +------+ +------+ +------+
servers servers servers servers servers
(rack) (rack) (rack) (rack) (rack)
- Leaf switches sit at the top of each rack (top-of-rack, ToR) and connect servers.
- Spine switches connect only to leaves. Every leaf connects to every spine.
- Any server reaches any other in at most leaf, spine, leaf: a predictable, small number of hops.
- All links are active. Traffic is spread across the parallel paths with ECMP (equal-cost multi-path) routing, which hashes each flow's 5-tuple to pick a path, keeping a flow's packets in order.
- To grow, add leaves for more racks or spines for more bandwidth.
The oversubscription ratio compares server-facing bandwidth to uplink bandwidth on a leaf. For example, 48 servers at 25 Gbps = 1200 Gbps down, and 4 uplinks at 100 Gbps = 400 Gbps up, gives 1200 / 400 = 3:1 oversubscription. That is fine if not every server sends across the fabric at full speed simultaneously.
Other data-centre ideas worth recognising: routing with BGP down to the rack, very low-latency, high-bandwidth links (25, 100, 400 Gbps), and RDMA (remote direct memory access) for storage and ML training clusters.
The troubleshooting toolkit
A good debugging approach is layered: confirm the lower layers work before blaming the higher ones. Is the interface up with an address? Can I reach the gateway? Does DNS resolve? Can I open TCP to the port? Does TLS succeed? What does the HTTP response say?
| Question | Tool |
|---|---|
| What are my IPs, routes and gateway? | ip addr, ip route (Linux); ifconfig, netstat -rn (macOS); ipconfig (Windows) |
| Is the host reachable, and what is the latency or loss? | ping |
| Where along the path is it slow or broken? | traceroute, mtr, tracert (Windows) |
| Which ports are listening, which connections exist? | ss, netstat, lsof -i |
| What does DNS return? | dig, nslookup |
| Does the HTTP request work, and where is the time going? | curl -v, curl -w |
| What exactly is on the wire? | tcpdump, Wireshark |
| Is a TCP port open? | nc -vz host port, curl -v telnet://host:port |
The example outputs below use documentation addresses and are trimmed for width. Exact formats differ between Linux distributions, macOS and Windows.
ip and ifconfig
$ ip -brief addr
lo UNKNOWN 127.0.0.1/8 ::1/128
eth0 UP 10.0.1.15/24 fe80::5054:ff:fe12:3456/64
$ ip route
default via 10.0.1.1 dev eth0 proto dhcp metric 100
10.0.1.0/24 dev eth0 proto kernel scope link src 10.0.1.15
What to look for:
- Is the interface UP and does it have the expected address and prefix?
- A
169.254.x.xaddress means DHCP failed (see application protocols). - Is there a default route (
default via ...)? Without one, you can reach the local subnet but nothing beyond. ip route get 203.0.113.10shows exactly which route and interface the kernel would use for that destination.ip neighshows the ARP cache (IP to MAC mappings);FAILEDentries mean ARP got no reply.
ifconfig and route/netstat -rn are the older equivalents, still standard on macOS and BSD. On Linux, the ip command from iproute2 replaced them.
ping
ping sends ICMP echo request messages and measures the time until the echo reply returns.
$ ping -c 5 203.0.113.10
PING 203.0.113.10 (203.0.113.10): 56 data bytes
64 bytes from 203.0.113.10: icmp_seq=0 ttl=54 time=21.4 ms
64 bytes from 203.0.113.10: icmp_seq=1 ttl=54 time=20.9 ms
64 bytes from 203.0.113.10: icmp_seq=2 ttl=54 time=88.3 ms
64 bytes from 203.0.113.10: icmp_seq=4 ttl=54 time=21.2 ms
--- 203.0.113.10 ping statistics ---
5 packets transmitted, 4 packets received, 20.0% packet loss
round-trip min/avg/max/stddev = 20.9/37.95/88.3/28.6 ms
What to look for:
- time: the round-trip time (RTT). Here about 21 ms normally.
- Jitter: variation in RTT. One spike to 88 ms suggests congestion or Wi-Fi interference.
- Packet loss: 20 percent (sequence 3 is missing) is serious. Even 1 to 2 percent sustained loss badly hurts TCP throughput, because TCP treats loss as congestion and cuts its window.
- ttl: the remaining hop limit. Starting values are typically 64 (Linux, macOS) or 128 (Windows), so
ttl=54suggests about 10 hops from a Linux-like host.
Common mistake
Concluding "the server is down" because ping fails. Many servers and cloud firewalls block ICMP. A failed ping only proves ICMP did not get through. Test the real service with nc -vz host 443 or curl before deciding.
traceroute and mtr
traceroute reveals the path by sending packets with increasing TTL values: TTL 1 expires at the first router, which returns an ICMP "time exceeded" message revealing itself, TTL 2 at the second, and so on. Linux traceroute sends UDP probes by default, Windows tracert uses ICMP, and traceroute -T uses TCP SYNs, which pass firewalls better.
$ traceroute -n 203.0.113.10
1 192.168.1.1 1.8 ms 1.5 ms 1.6 ms
2 100.64.0.1 6.2 ms 5.9 ms 6.4 ms
3 198.51.100.1 7.1 ms 7.3 ms 7.0 ms
4 * * *
5 198.51.100.77 19.8 ms 20.4 ms 19.9 ms
6 203.0.113.10 21.0 ms 20.8 ms 21.3 ms
Each line is one hop, with three probe RTTs. What to look for:
* * *at one hop with later hops responding normally is not a problem: that router simply does not answer probes or rate-limits ICMP.* * *from some hop onward to the end means packets stop there: a firewall or a routing failure.- A jump in latency that persists for all later hops shows where delay is added (for example a long undersea link from 20 ms to 150 ms). A spike at one hop that disappears at the next is just that router being slow to generate ICMP replies, which routers treat as low priority.
mtr combines ping and traceroute: it probes every hop continuously and shows loss and latency statistics per hop:
$ mtr -rwc 100 203.0.113.10
HOST Loss% Snt Last Avg Best Wrst StDev
1. 192.168.1.1 0.0% 100 1.6 1.7 1.3 4.2 0.4
2. 100.64.0.1 0.0% 100 6.1 6.3 5.7 9.8 0.6
3. 198.51.100.1 40.0% 100 7.2 7.4 6.9 12.1 0.8
4. 198.51.100.77 0.0% 100 20.1 20.3 19.5 25.0 0.7
5. 203.0.113.10 0.0% 100 21.2 21.1 20.6 24.8 0.6
The 40 percent loss at hop 3 is not real, because later hops show 0 percent loss: hop 3 just rate-limits its ICMP replies. Real loss carries through to the final hop. This is the single most commonly misread thing in mtr output.
ss and netstat
ss (socket statistics) lists sockets on Linux; netstat is the older tool (still used on macOS and Windows).
$ ss -tlnp
State Recv-Q Send-Q Local Address:Port Peer Address:Port Process
LISTEN 0 511 0.0.0.0:80 0.0.0.0:* users:(("nginx",pid=812,fd=6))
LISTEN 0 4096 127.0.0.1:5432 0.0.0.0:* users:(("postgres",pid=640,fd=5))
LISTEN 0 128 *:8080 *:* users:(("java",pid=1203,fd=41))
Flags: -t TCP, -u UDP, -l listening only, -n numeric (no DNS lookups), -p show process (needs root for other users' processes).
What to look for:
- Is the service listening at all, and on which address?
127.0.0.1:5432accepts only local connections;0.0.0.0or*accepts from any interface. "Connection refused" from another machine often means the service is bound to localhost only. - Which process owns a port: the classic answer to "port 8080 already in use" (also
lsof -i :8080). - For a listening socket, Recv-Q is the current accept queue and Send-Q is the backlog limit. A full accept queue means the application is not calling
acceptfast enough.
Established connections and their states:
$ ss -tan state established '( dport = :5432 )' | head -3
Recv-Q Send-Q Local Address:Port Peer Address:Port
0 0 10.0.1.15:48122 10.0.2.20:5432
0 0 10.0.1.15:48130 10.0.2.20:5432
$ ss -tan | awk 'NR>1 {print $1}' | sort | uniq -c
143 ESTAB
2 LISTEN
3870 TIME-WAIT
35 CLOSE-WAIT
- Thousands of TIME-WAIT sockets: the host opens and closes many short connections (often an HTTP client without connection pooling). It can exhaust ephemeral ports. Fix with keep-alive and connection reuse.
- Growing CLOSE-WAIT: the remote side closed, but your application never called close. This is almost always a bug: a leaked connection.
- Large Send-Q on established connections: data waiting to be sent, meaning the network or the receiver is slow.
ss -tishows TCP internals per connection: RTT, congestion window and retransmissions.
dig and nslookup
Covered in depth in the DNS lesson. Quick checks:
$ dig +short api.example.com
api.example.com.cdn-provider.net.
203.0.113.10
$ dig @8.8.8.8 api.example.com +noall +answer
api.example.com. 287 IN CNAME api.example.com.cdn-provider.net.
api.example.com.cdn-provider.net. 47 IN A 203.0.113.10
What to look for: the status (NOERROR, NXDOMAIN, SERVFAIL), whether different resolvers return different answers (stale caches or split-horizon DNS), the remaining TTL, and the Query time line in full output: a slow lookup (hundreds of ms) adds directly to page load. nslookup works on every OS, including Windows, but gives less detail.
curl -v
curl -v shows every stage of an HTTP request:
$ curl -v https://api.example.com/health
* Host api.example.com:443 was resolved.
* IPv4: 203.0.113.10
* Trying 203.0.113.10:443...
* Connected to api.example.com (203.0.113.10) port 443
* ALPN: curl offers h2,http/1.1
* TLSv1.3 (OUT), TLS handshake, Client hello (1):
* TLSv1.3 (IN), TLS handshake, Server hello (2):
* SSL connection using TLSv1.3 / TLS_AES_128_GCM_SHA256
* ALPN: server accepted h2
* Server certificate:
* subject: CN=api.example.com
* expire date: Dec 30 23:59:59 2026 GMT
* issuer: C=US; O=Example CA; CN=Example Intermediate
* SSL certificate verify ok.
> GET /health HTTP/2
> Host: api.example.com
> User-Agent: curl/8.7.1
> Accept: */*
>
< HTTP/2 200
< content-type: application/json
< x-cache: MISS
<
{"status":"ok"}
What to look for: which IP it connected to (is DNS sending you to the right place?), whether the connection succeeds (Connection refused means nothing listens; a hang then timeout suggests a firewall dropping packets), the TLS version, certificate subject, issuer and expiry (expired certificates are a frequent outage cause), the ALPN result (h2 or http/1.1), the request headers sent (>), and the status and headers received (<), including cache headers.
Useful variations: curl -I (headers only), curl -L (follow redirects), curl --resolve api.example.com:443:203.0.113.11 https://api.example.com/ (test a specific backend while keeping the right host name and SNI), and the timing format shown next.
tcpdump
tcpdump captures packets on an interface and prints them, or writes them to a .pcap file for Wireshark. It needs root.
sudo tcpdump -i eth0 -nn 'tcp port 443 and host 203.0.113.10'
sudo tcpdump -i any -nn 'udp port 53' # watch DNS queries
sudo tcpdump -i eth0 -nn 'tcp[tcpflags] & tcp-syn != 0' # SYN packets
sudo tcpdump -i eth0 -w capture.pcap port 8080 # save for Wireshark
A capture of a connection that keeps retrying:
10:15:02.100 IP 10.0.1.15.48122 > 203.0.113.10.443: Flags [S], seq 1001, win 64240, length 0
10:15:03.120 IP 10.0.1.15.48122 > 203.0.113.10.443: Flags [S], seq 1001, win 64240, length 0
10:15:05.160 IP 10.0.1.15.48122 > 203.0.113.10.443: Flags [S], seq 1001, win 64240, length 0
Flags: S = SYN, S. = SYN-ACK, . = ACK, P. = push with data, F. = FIN, R = reset.
Reading this: the client sends SYN three times with roughly doubling gaps (1 s, then 2 s: exponential backoff) and never gets a SYN-ACK. So the packets are being silently dropped: a firewall or security group blocking port 443, or a routing problem. Compare with a RST reply, which would mean the host is reachable but nothing listens on the port ("connection refused"). That distinction, timeout versus refused, is one of the most useful clues in network debugging.
Other patterns to look for: repeated retransmissions mid-connection (packet loss), zero-window advertisements (the receiver is not reading fast enough), and a healthy handshake followed by a long silence before the response (slow server).
Wireshark filters
Wireshark is a graphical packet analyser that decodes hundreds of protocols. Its display filters (applied after capture, different syntax from tcpdump's capture filters) are worth knowing:
| Filter | Shows |
|---|---|
ip.addr == 203.0.113.10 | Traffic to or from that host |
tcp.port == 443 | Traffic on port 443 |
dns | DNS packets |
http.request.method == "POST" | HTTP POST requests (plain HTTP only) |
http.response.code >= 500 | Server errors |
tcp.flags.syn == 1 && tcp.flags.ack == 0 | Connection attempts (SYNs) |
tcp.flags.reset == 1 | Resets |
tcp.analysis.retransmission | Retransmitted segments |
tcp.analysis.zero_window | Receiver window full |
tls.handshake.type == 1 | TLS ClientHello (shows SNI) |
frame.time_delta > 1 | Gaps longer than one second |
"Follow TCP stream" reassembles a whole conversation. HTTPS payloads are encrypted, but you can decrypt your own browser's traffic by setting the SSLKEYLOGFILE environment variable, which makes browsers and curl write session keys that Wireshark can load.
Interview tip
When asked how you would debug a network issue, say you go layer by layer and name a tool for each: ip for interface and routes, ping and mtr for reachability and loss, dig for DNS, nc or ss for the port, curl -v for TLS and HTTP, and tcpdump when you need ground truth. The structure matters more than any single command.
Worked walkthrough: "the site is slow"
A user in Bengaluru reports that https://shop.example.com takes several seconds to load. Here is how to investigate methodically.
Step 1: clarify and scope
Before running tools, ask questions:
- Who is affected? One user, one region, or everyone? Check monitoring dashboards and error rates.
- What is slow? The first page load, every page, only checkout, only images?
- Since when? Did it start after a deploy, a config change, or a traffic spike?
Assume: several users in India report it, starting this morning; other regions seem fine.
Step 2: break the request into phases
From a machine in the affected region, use curl's timing variables:
curl -o /dev/null -s -w \
'dns=%{time_namelookup} connect=%{time_connect} tls=%{time_appconnect} ttfb=%{time_starttransfer} total=%{time_total}\n' \
https://shop.example.com/
dns=0.012 connect=0.198 tls=0.391 ttfb=0.640 total=0.702
The values are cumulative seconds from the start. Convert to per-phase durations:
| Phase | Calculation | Duration |
|---|---|---|
| DNS | 0.012 | 12 ms |
| TCP connect | 0.198 − 0.012 | 186 ms |
| TLS handshake | 0.391 − 0.198 | 193 ms |
| Server + 1 RTT (to first byte) | 0.640 − 0.391 | 249 ms |
| Download body | 0.702 − 0.640 | 62 ms |
Interpretation: TCP connect costs one round trip, so the RTT is about 186 ms. TLS 1.3 also costs one RTT, and 193 ms matches. A Bengaluru user would normally reach a nearby CDN edge in well under 20 ms. So the problem is network distance, not the server: the user is connecting to something far away. Time to first byte minus one RTT (249 − 186 = about 63 ms) suggests the server itself is responding reasonably.
Compare with what a "slow server" would look like: connect=0.020 tls=0.040 ttfb=2.400, meaning a tiny RTT but 2.36 seconds waiting for the first byte. That would send you to application logs, slow database queries and dependency latency instead.
Step 3: check where DNS is sending users
$ dig +short shop.example.com
shop.example.com.cdn-provider.net.
198.51.100.200
Compare with a resolver elsewhere, or check the CDN's documentation for which PoP owns that address. Suppose 198.51.100.200 is the CDN's edge in Frankfurt, while users in India normally get a Mumbai or Chennai edge. Possibilities: the CDN is steering Indian traffic away from a degraded Indian PoP, the users' resolver is located abroad (some public resolvers or corporate VPNs make users look like they are elsewhere), or a recent DNS change removed the regional configuration.
curl -sI https://shop.example.com/ | grep -i -E 'cf-ray|x-served-by|x-amz-cf-pop|via' often shows a PoP code in response headers, confirming which edge served the request.
Step 4: confirm the path
$ mtr -rwc 50 shop.example.com
HOST Loss% Snt Avg Best Wrst
1. 192.168.1.1 0.0% 50 1.9 1.4 5.1
2. 100.64.0.1 0.0% 50 7.0 6.1 12.4
3. 198.51.100.9 0.0% 50 8.2 7.5 11.0
4. 198.51.100.141 0.0% 50 142.6 140.9 149.8
5. 198.51.100.200 0.0% 50 185.3 183.9 192.0
No loss anywhere, but a large jump at hop 4 that persists to the end: traffic is leaving the country on an international link. That confirms the routing story: the request is crossing continents.
Step 5: rule out the other layers
- Packet loss: none in mtr, so retransmissions are not the cause. (If there were loss carrying through to the last hop,
ss -tior tcpdump would show retransmissions and TCP throughput would collapse.) - Server: TTFB minus RTT is small; application dashboards show normal latency. Not the server.
- Page weight: in the browser's DevTools Network tab, the waterfall shows each resource's DNS, connect, TLS, waiting (TTFB) and download times. If all of them show long connect times, it is the network; if one huge unoptimised image dominates, it is the page.
Step 6: fix and verify
Here, the fix lies with the CDN's routing configuration (contact the provider, check its status page for a degraded Indian PoP, or review the DNS/geo-steering change made this morning). After the fix, rerun the same curl command and expect something like connect=0.012 tls=0.025, and confirm that dashboards for the region recover.
Finally, write up what happened and add monitoring: synthetic checks from multiple regions that measure connect time and TTFB would have caught this before users did.
A summary checklist for "slow"
1. Scope: who, what, since when, what changed?
2. curl -w timings: DNS | connect (=RTT) | TLS | TTFB | download
3. Big DNS? -> resolver issues, low TTL + cache misses
4. Big connect? -> distance or routing: dig (which edge?), mtr
5. Loss? -> mtr loss that persists to last hop, retransmits
6. Big TTFB? -> server: app logs, DB, dependencies, CPU
7. Big download? -> page weight, compression, bandwidth
8. Fix, re-measure, add monitoring.
Interview questions
Q1. What is a CDN and how does it improve performance?
A CDN is a network of edge servers close to users that cache and serve content on behalf of an origin. It reduces round-trip time for handshakes and requests, serves cached content without touching the origin, keeps warm connections to the origin for misses, and absorbs traffic spikes and DDoS attacks.
Q2. Pull versus push CDN?
A pull CDN fetches content from the origin on the first cache miss and caches it according to HTTP headers, which is low-effort and suits websites. A push CDN requires you to upload content in advance, which suits large, rarely changing files that must be available immediately.
Q3. What is the difference between L4 and L7 load balancing?
An L4 balancer routes by IP and port per connection without looking at content, so it is fast and protocol-agnostic. An L7 balancer terminates the connection and routes each HTTP request based on path, host, headers or cookies, enabling features such as path routing, retries and per-request balancing of HTTP/2 and gRPC.
Q4. Which load-balancing algorithm would you choose?
Round robin or weighted round robin for uniform requests on similar servers; least connections when request durations vary or connections are long-lived; consistent hashing when the same key should go to the same server, as for caches. "Power of two random choices" works well for large fleets.
Q5. Forward proxy versus reverse proxy?
A forward proxy acts for clients, sending their requests out to the Internet, used for filtering, caching and anonymity. A reverse proxy acts for servers, receiving client requests and forwarding them to backends, used for TLS termination, load balancing, caching and protection.
Q6. What does an API gateway do beyond a reverse proxy?
It centralises API concerns: authentication, rate limiting and quotas, routing and versioning, request transformation and aggregation, and observability. It should stay free of business logic to avoid becoming a bottleneck.
Q7. How do two devices behind NAT establish a direct connection?
They use STUN to learn their public address and port, exchange those through a signalling server, and send packets to each other so both NATs create mappings (hole punching). ICE tests candidate paths and, if direct connection fails, as with symmetric NAT, falls back to a TURN relay.
Q8. What is SDN?
Software-defined networking separates the control plane from the data plane. A logically centralised controller computes forwarding rules with a global view and programs switches through an API such as OpenFlow, making the network programmable and easier to manage.
Q9. Why do data centres use leaf-spine topologies?
Modern traffic is mostly east-west between servers. Leaf-spine gives every pair of servers a short, predictable path (leaf, spine, leaf), uses all links via ECMP rather than blocking redundant ones, and scales by adding leaves or spines.
Q10. Ping fails but the website loads. How?
The server or a firewall on the path blocks ICMP echo, while TCP 443 is allowed. Ping only tests ICMP reachability; to test a service, connect to its actual port.
Q11. How do you find which process is using port 8080?
On Linux, ss -tlnp | grep 8080 or lsof -i :8080; on macOS, lsof -i :8080; on Windows, netstat -ano | findstr 8080 and then look up the process id.
Q12. What is the difference between "connection refused" and a connection timeout?
Refused means the host replied with a TCP RST: it is reachable, but nothing is listening on that port. A timeout means no reply came at all, usually because a firewall silently drops packets or the host or route is down.
Q13. In mtr, one middle hop shows 50 percent loss but the destination shows 0 percent. Is there a problem?
No. That router rate-limits or deprioritises ICMP replies to probes directed at it, while still forwarding traffic. Real loss persists through all later hops to the destination.
Q14. Many sockets are in TIME_WAIT or CLOSE_WAIT. What does each suggest?
Many TIME_WAIT sockets mean the host actively closes lots of short-lived connections, often due to missing connection reuse, which can exhaust ephemeral ports. Growing CLOSE_WAIT means the peer closed but the local application never closed its socket, which is a connection-leak bug.
Q15. A user says the site is slow. How do you debug it?
Scope the problem first: who, what and since when. Then split the request into phases with curl -w: DNS, connect, TLS, time to first byte and download. A large connect time points to distance or routing (check dig and mtr), loss shows up in mtr and retransmissions, and a large TTFB points to the server (logs, database, dependencies). Fix, re-measure and add monitoring.
Key takeaways
- CDNs cut latency with nearby edges, cache by HTTP headers and cache keys, and come in pull and push flavours.
- L4 balancers route connections by IP and port; L7 balancers route requests by content. Pick algorithms by workload: round robin, least connections, consistent hashing.
- Forward proxies act for clients; reverse proxies (and CDNs, L7 LBs, API gateways) act for servers.
- STUN discovers public addresses for hole punching; TURN relays when direct paths fail; ICE chooses.
- SDN separates control and data planes; overlays like VXLAN and cloud VPCs virtualise networks; data centres use leaf-spine with ECMP.
- Debug layer by layer:
ip,ping,mtr,dig,ss,curl -v,tcpdump, Wireshark. - Refused means RST (nothing listening); timeout means dropped. Mid-path mtr loss that does not persist is harmless.
- For "slow",
curl -wtimings split DNS, RTT, TLS, server time and download, and point you to the right layer.
Next lesson
Continue with Computer networks interview questions.

