What the transport layer adds
IP gets a packet from one host to another, on a best-effort basis: packets can be lost, duplicated, delayed or reordered, and IP makes no promises. It also stops at the host. A server runs a web server, a database and an SSH daemon at the same time; IP has no idea which of them a packet is for.
The transport layer (layer 4) fills both gaps. It delivers data process to process using port numbers, and, depending on the protocol, adds reliability, ordering, flow control and congestion control on top of IP's unreliable service. The two protocols you must know are UDP, which adds almost nothing beyond ports and a checksum, and TCP, which builds a reliable, ordered byte stream.
This is one of the most heavily tested topics in networking interviews. You should be able to draw the TCP three-way handshake and four-way teardown with sequence and acknowledgement numbers, explain TIME_WAIT and CLOSE_WAIT, describe a SYN flood, compare Go-Back-N and Selective Repeat with numbers, and argue when UDP beats TCP. Flow control and congestion control are big enough to get their own lesson, TCP flow and congestion control.
Process-to-process delivery
Ports
A port is a 16-bit number (0 to 65,535) that identifies an endpoint for a process on a host. Ranges, as defined by IANA:
| Range | Name | Use |
|---|---|---|
| 0 to 1023 | Well-known (system) ports | Standard services; binding needs admin rights on most Unix systems |
| 1024 to 49151 | Registered ports | Applications such as MySQL 3306, PostgreSQL 5432 |
| 49152 to 65535 | Dynamic / ephemeral | Temporary client-side ports picked by the OS |
The ephemeral range actually used depends on the OS: Linux defaults to 32768 to 60999, while Windows uses the IANA range.
Ports worth remembering:
| Port | Protocol | Transport |
|---|---|---|
| 20, 21 | FTP data, control | TCP |
| 22 | SSH | TCP |
| 23 | Telnet | TCP |
| 25 | SMTP | TCP |
| 53 | DNS | UDP (and TCP) |
| 67, 68 | DHCP server, client | UDP |
| 80 | HTTP | TCP |
| 110 | POP3 | TCP |
| 123 | NTP | UDP |
| 143 | IMAP | TCP |
| 443 | HTTPS (and HTTP/3 over UDP) | TCP, UDP |
Sockets
A socket is the operating system's handle for one end of a communication: the door between a process and the transport layer. Programs read from and write to sockets, and the OS fills in the headers. A socket is identified by its IP address and port, together with the protocol.
Multiplexing and demultiplexing
- Multiplexing (at the sender): gathering data from many sockets, adding transport headers with the right port numbers, and passing segments down to IP.
- Demultiplexing (at the receiver): using header fields to deliver each arriving segment to the correct socket.
The two protocols demultiplex differently, and this is a favourite question.
UDP: connectionless demultiplexing. A UDP socket is identified by the 2-tuple (destination IP, destination port). Datagrams from any number of senders to the same IP and port land in the same socket. The application reads the sender's address from each datagram if it wants to reply.
TCP: connection-oriented demultiplexing. A TCP connection is identified by the 4-tuple (source IP, source port, destination IP, destination port). A web server listening on port 443 has one listening socket, and every accepted client gets its own connected socket. Two clients both talking to server:443 are told apart by their source IP and port.
Server 203.0.113.10, listening on :443
Client A 198.51.100.7:50001 --+
|--> socket #1 (A's 4-tuple)
Client B 198.51.100.7:50002 --+--> socket #2 (B's 4-tuple)
Client C 192.0.2.44:50001 ----+--> socket #3 (C's 4-tuple)
All three have destination :443; the source IP or port differs.
Common mistake
"A server can only handle 65,535 connections because there are only that many ports." The server uses one port (443) for all of them. Connections are distinguished by the full 4-tuple, so the limit is memory and file descriptors, not ports. The port limit applies to one client IP opening many connections to the same server IP and port.
UDP
UDP (User Datagram Protocol, RFC 768) adds only what IP lacks for process delivery: ports and an optional checksum.
UDP header
The header is just 8 bytes:
0 16 31
+----------------+----------------+
| Source port | Dest port |
+----------------+----------------+
| Length | Checksum |
+----------------+----------------+
| Data ... |
- Source port: the sender's port, so the receiver can reply (may be 0 if no reply is wanted).
- Destination port: selects the receiving socket.
- Length: header plus data, in bytes (minimum 8).
- Checksum: ones' complement sum over the header, data and a pseudo-header (source and destination IP, protocol number and UDP length) taken from the IP layer. Including IP addresses lets the receiver detect misdelivered datagrams. It is optional in IPv4 (0 means none) and mandatory in IPv6.
What UDP gives and does not give
| UDP provides | UDP does not provide |
|---|---|
| Process-to-process delivery via ports | Delivery guarantee |
| Error detection (checksum) | Ordering |
| Message boundaries (one send = one datagram) | Duplicate removal |
| No connection setup | Flow control |
| No per-connection state at the server | Congestion control |
Message boundaries matter: if you sendto 100 bytes and then 200 bytes, the receiver gets two datagrams of 100 and 200 bytes. TCP, a byte stream, could deliver them as one read of 300 bytes or three reads of 100.
Why anyone uses UDP
- No handshake: the first packet carries data. A DNS query and answer finish in one round trip.
- No head-of-line blocking: a lost packet does not hold up later ones. Live video and games prefer a slightly glitchy frame now to a perfect frame late.
- Small header (8 bytes versus TCP's 20 or more).
- No connection state, so a server can handle huge numbers of clients cheaply.
- Application control: apps can build exactly the reliability they need on top. QUIC does precisely that.
A UDP exchange in Python
import socket
server = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
server.bind(("127.0.0.1", 0)) # port 0: let the OS pick a free port
port = server.getsockname()[1]
client = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
client.sendto(b"hello", ("127.0.0.1", port)) # no connection, no handshake
data, addr = server.recvfrom(2048) # one call returns one whole datagram
print(data, "from", addr)
server.sendto(data.upper(), addr) # reply to the sender's (IP, port)
print(client.recvfrom(2048)[0]) # b'HELLO'
client.close()
server.close()
Notice there is no listen, accept or connect: the server learns who sent each datagram from recvfrom.
TCP
TCP (Transmission Control Protocol, RFC 9293, which consolidates the original RFC 793 and its updates) provides:
- Connection-oriented service: a handshake sets up state at both ends before data flows.
- A reliable, in-order byte stream: bytes arrive exactly once, in order, or the connection reports an error.
- Full duplex: both sides send at the same time over one connection.
- Flow control: the sender never overruns the receiver's buffer.
- Congestion control: the sender slows down when the network is congested.
TCP is point to point (one sender, one receiver): there is no TCP multicast.
TCP segment header
0 16 31
+---------------------+---------------------+
| Source port | Destination port |
+---------------------+---------------------+
| Sequence number |
+-------------------------------------------+
| Acknowledgement number |
+------+------+-------+---------------------+
|Offset| Rsvd | Flags | Window size |
+------+------+-------+---------------------+
| Checksum | Urgent pointer |
+---------------------+---------------------+
| Options (0 to 40 bytes) |
+-------------------------------------------+
| Data ... |
| Field | Bits | Meaning |
|---|---|---|
| Source / destination port | 16 each | Identify the sockets |
| Sequence number | 32 | Byte-stream position of the first data byte in this segment |
| Acknowledgement number | 32 | Next byte the sender of this segment expects to receive (valid when ACK is set) |
| Data offset | 4 | Header length in 32-bit words: 5 (20 bytes) to 15 (60 bytes) |
| Flags | 8 (classically 6) | CWR, ECE, URG, ACK, PSH, RST, SYN, FIN |
| Window size | 16 | Receive window: how many more bytes the receiver will accept (scaled by the window-scale option) |
| Checksum | 16 | Over header, data and IP pseudo-header, like UDP but mandatory |
| Urgent pointer | 16 | Rarely used; marks urgent data when URG is set |
| Options | variable | MSS, window scale, SACK permitted, SACK blocks, timestamps |
The flags:
- SYN: synchronise sequence numbers; used only to open a connection.
- ACK: the acknowledgement number is valid. Set on every segment after the first.
- FIN: the sender has finished sending (half-close).
- RST: reset; abort the connection immediately (for example, a segment arrived for a port with no listener).
- PSH: push buffered data to the application promptly.
- URG: urgent pointer is valid.
- ECE, CWR: explicit congestion notification signalling.
Sequence and acknowledgement numbers
TCP numbers bytes, not segments. If a segment has sequence number 1001 and carries 100 bytes, it holds bytes 1001 to 1100. The receiver replies with ack = 1101, meaning "I have everything up to 1100; send 1101 next." These are cumulative ACKs: one ACK covers all earlier bytes.
Each side picks a random initial sequence number (ISN) at connection setup. Random ISNs make it hard for attackers to inject forged segments and stop segments from an old connection being mistaken for a new one. The SYN and FIN flags each consume one sequence number even though they carry no data.
The three-way handshake
Before data flows, both sides must agree to talk and learn each other's ISN. Example: client ISN = 1000, server ISN = 5000.
Client Server
CLOSED LISTEN
| |
|---- SYN seq=1000 --------------------------->|
SYN_SENT SYN_RCVD
|<--- SYN+ACK seq=5000 ack=1001 --------------|
| |
|---- ACK seq=1001 ack=5001 ----------------->|
ESTABLISHED ESTABLISHED
| |
|---- data seq=1001 ack=5001 (100 bytes) ------>|
|<--- ACK seq=5001 ack=1101 ------------------|
Step by step:
- SYN. The client sends SYN with seq = 1000 (its ISN). It also advertises options such as MSS (maximum segment size), window scale and SACK permitted. State: SYN_SENT.
- SYN-ACK. The server allocates connection state, chooses its own ISN 5000, and replies with SYN and ACK set: seq = 5000, ack = 1001. The ack is 1001 because the SYN consumed sequence number 1000. State: SYN_RCVD.
- ACK. The client acknowledges the server's SYN: seq = 1001, ack = 5001. Both sides are now ESTABLISHED. This third segment may already carry data.
Then the client sends 100 bytes with seq = 1001 (bytes 1001 to 1100), and the server acknowledges with ack = 1101.
Why three messages and not two?
Each side must send its own ISN and hear an acknowledgement of it. That is four messages, but the server's ACK and SYN are combined into one, giving three. With only two, the server could not know the client received its ISN. Two messages would also let a delayed, duplicate SYN from an old connection open a new connection at the server with nobody on the other end; the third message proves the client is really there and wants this connection.
The handshake costs one RTT before the client can send data (1.5 RTTs before the server receives it).
Simultaneous open and TCP Fast Open
Both sides may send SYN at the same time; TCP handles this with four segments. TCP Fast Open (RFC 7413) lets a returning client send data in the SYN using a cookie from an earlier connection, saving one RTT; support varies and middleboxes sometimes interfere.
The four-way teardown
TCP is full duplex, so each direction is closed separately with a FIN. Continuing the example (client has sent up to byte 1100, server up to 5000):
Client Server
ESTABLISHED ESTABLISHED
|---- FIN seq=1101 ack=5001 ----------------->|
FIN_WAIT_1 CLOSE_WAIT
|<--- ACK seq=5001 ack=1102 ------------------|
FIN_WAIT_2 |
| (server may still send data) |
|<--- FIN seq=5001 ack=1102 ------------------|
TIME_WAIT LAST_ACK
|---- ACK seq=1102 ack=5002 ----------------->|
| CLOSED
(waits 2 x MSL)
CLOSED
- Client FIN (seq = 1101): "I have no more data to send." The client enters FIN_WAIT_1. The FIN consumes one sequence number.
- Server ACK (ack = 1102). The server enters CLOSE_WAIT and tells its application that the other side has closed (a read returns end-of-file). The client enters FIN_WAIT_2. The connection is now half-closed: the server can still send data.
- Server FIN (seq = 5001) when the server application calls
close. The server enters LAST_ACK. - Client ACK (ack = 5002). The client enters TIME_WAIT; the server, on receiving it, goes to CLOSED.
Steps 2 and 3 are often combined into one FIN+ACK when the server closes immediately, giving three segments.
A connection can also end abruptly with RST, which discards any unsent data and skips TIME_WAIT. Applications sometimes trigger this on purpose (for example with the SO_LINGER option set to zero), which avoids TIME_WAIT but risks losing data.
TCP state diagram essentials
You do not need to draw all eleven states from memory, but you must know the main paths.
+--------+
| CLOSED |
+--------+
passive open | | active open / send SYN
v v
+--------+ +----------+
| LISTEN | | SYN_SENT |
+--------+ +----------+
recv SYN / | | recv SYN+ACK / send ACK
send SYN+ACK v v
+----------+ +-------------+
| SYN_RCVD |-->| ESTABLISHED |
+----------+ +-------------+
recv ACK | |
close / send FIN | | recv FIN / send ACK
(active close) v v (passive close)
+------------+ +------------+
| FIN_WAIT_1 | | CLOSE_WAIT |
+------------+ +------------+
recv ACK | | close / send FIN
v v
+------------+ +----------+
| FIN_WAIT_2 | | LAST_ACK |
+------------+ +----------+
recv FIN / send ACK | | recv ACK
v v
+-----------+ CLOSED
| TIME_WAIT |
+-----------+
wait 2 MSL |
v
CLOSED
(There is also CLOSING, reached when both sides send FIN at the same time.)
TIME_WAIT and why it exists
The side that closes first (the active closer) waits in TIME_WAIT for 2 × MSL, where MSL (maximum segment lifetime) is the longest a segment can survive in the network. RFC 793 suggested an MSL of 2 minutes; Linux uses a fixed TIME_WAIT of 60 seconds. There are two reasons:
- Reliably finish the close. If the final ACK is lost, the other side will retransmit its FIN. The active closer must still be around to resend the ACK; otherwise it would answer with RST and the peer would see an error.
- Let old duplicates die. Delayed segments from this connection may still be wandering the network. If a new connection with the same 4-tuple started immediately, an old segment could be accepted as new data. Waiting 2 MSL ensures any stray segment in either direction has expired.
The practical problem: a busy client (or a proxy) that opens and closes many short connections to the same server accumulates thousands of sockets in TIME_WAIT, each holding a 4-tuple, and can run out of ephemeral ports. Fixes include connection reuse (HTTP keep-alive, connection pools), having the server close first only when appropriate, and on Linux the net.ipv4.tcp_tw_reuse setting, which allows reuse of TIME_WAIT sockets for new outgoing connections when timestamps show it is safe. (The older tcp_tw_recycle option broke clients behind NAT and was removed from Linux in version 4.12.)
CLOSE_WAIT and what it signals
A socket in CLOSE_WAIT has received the peer's FIN and acknowledged it, but its own application has not yet called close. TCP cannot leave CLOSE_WAIT on its own. So many sockets stuck in CLOSE_WAIT almost always mean an application bug: code that reads end-of-file but forgets to close the socket, or a leaked connection in a pool. They consume file descriptors until the process hits its limit.
Interview tip
"Many TIME_WAIT sockets" is usually normal behaviour on the side that closes first and is fixed by connection reuse. "Many CLOSE_WAIT sockets" is a bug in your own application: it is not closing sockets after the peer closed. Commands like ss -tan state close-wait show them.
SYN flood
A SYN flood is a denial-of-service attack on the handshake. The attacker sends a flood of SYNs, often with spoofed source addresses, and never completes the handshake. For each SYN, the server allocates state, replies with SYN-ACK and waits in SYN_RCVD (retransmitting the SYN-ACK a few times). The backlog queue of half-open connections fills up, and legitimate clients' SYNs are dropped.
Defences:
- SYN cookies: the server stores no state for a SYN. Instead, it encodes the connection details (a hash of the 4-tuple and a secret, a coarse timestamp and the MSS) into its own ISN. When the final ACK arrives with ack = ISN + 1, the server checks the hash and only then creates the connection. Spoofed SYNs cost the server almost nothing. Linux enables SYN cookies when the queue overflows (
net.ipv4.tcp_syncookies). The trade-off is that some TCP options are limited without timestamps. - Larger backlog queues and shorter SYN-ACK retry timeouts.
- Rate limiting and filtering at firewalls, load balancers and DDoS-protection services that absorb SYNs before they reach the server.
Reliable data transfer: building up the ideas
How do you build a reliable channel over an unreliable one? Textbooks build it in steps called rdt (reliable data transfer) versions. The mechanisms they introduce are exactly the ones TCP uses.
| Version | Channel assumption | Mechanism added |
|---|---|---|
| rdt 1.0 | Perfectly reliable | None; just send |
| rdt 2.0 | Bits can flip | Checksum, ACK and NAK, retransmit on NAK |
| rdt 2.1 | ACK/NAK can also be corrupted | Sequence numbers (0, 1) so the receiver can spot duplicates |
| rdt 2.2 | Same | NAK-free: ACK carries the sequence number of the last good packet; a duplicate ACK means "resend" |
| rdt 3.0 | Packets can also be lost | Timer: retransmit if no ACK arrives in time |
rdt 3.0 is a stop-and-wait protocol with a 1-bit sequence number, also called the alternating-bit protocol. It is correct, but slow.
The five building blocks to name in an interview: checksums (detect corruption), acknowledgements (confirm receipt), sequence numbers (detect duplicates and order), timers (detect loss), and windows / pipelining (performance).
Stop-and-wait utilisation
The sender sends one packet, then waits a full RTT for the ACK. With packet size L, link rate R and round-trip time RTT:
U_sender = (L / R) / (RTT + L / R)
Worked example. Packets of 1,500 bytes (12,000 bits) on a 1 Gbps link with RTT = 30 ms.
- Transmission time L/R = 12,000 / 10^9 = 12 µs = 0.012 ms.
- U = 0.012 / (30 + 0.012) = 0.0004, that is 0.04%.
- Effective throughput = 12,000 bits per 30.012 ms ≈ 400 kbps on a 1 Gbps link.
The sender is idle 99.96% of the time. This is why reliable protocols pipeline: send many packets before waiting for ACKs.
Pipelining
With a window of N packets in flight:
U_sender = min(1, N × (L / R) / (RTT + L / R))
Same link:
| Window N | Utilisation | Throughput |
|---|---|---|
| 1 | 0.04% | about 0.4 Mbps |
| 3 | 0.12% | about 1.2 Mbps |
| 100 | 4.0% | about 40 Mbps |
| 1,000 | 40.0% | about 400 Mbps |
| 2,501 | 100% | 1 Gbps |
To fill the link, N must be at least 1 + RTT / (L/R) = 1 + 30 / 0.012 = 2,501 packets, which is the bandwidth-delay product in packets (see networking basics).
Pipelining raises two questions: how many sequence numbers do we need, and what do we resend after a loss? The two classic answers are Go-Back-N and Selective Repeat.
Go-Back-N
In Go-Back-N (GBN):
- The sender may have up to N unacknowledged packets. The window covers sequence numbers
basetobase + N - 1. - The receiver accepts packets only in order. An out-of-order packet is discarded, and the receiver re-sends the ACK for the last in-order packet. It needs no buffer for out-of-order packets.
- ACKs are cumulative: ACK n means all packets up to n arrived.
- The sender keeps one timer, for the oldest unacknowledged packet. On timeout, it resends every packet in the window, starting from
base.
Worked example: GBN with a loss
Window N = 4, packets 0 to 7, packet 2 is lost once.
Sender Receiver
send 0 ------------------------> got 0, ACK0
send 1 ------------------------> got 1, ACK1
send 2 ----X (lost)
send 3 ------------------------> out of order, discard, ACK1
(ACK0 arrives: window slides, send 4)
send 4 ------------------------> discard, ACK1
(ACK1 arrives: window = 2..5, send 5)
send 5 ------------------------> discard, ACK1
(duplicate ACK1s ignored)
... timer for 2 expires ...
resend 2, 3, 4, 5 -------------> got 2..5 in order, ACK2..ACK5
send 6, 7 ---------------------> ACK6, ACK7
Transmissions: packets 0 to 5 once (6), then 2, 3, 4, 5 again (4), then 6 and 7 (2): 12 transmissions for 8 packets. Packets 3, 4 and 5 were delivered successfully the first time but thrown away.
Selective Repeat
In Selective Repeat (SR):
- The sender still has a window of N packets, but keeps a separate timer for each packet and resends only the packets whose timers expire.
- The receiver buffers out-of-order packets within its own window and acknowledges each packet individually. When the missing packet arrives, it delivers the whole run in order.
Same scenario: packet 2 is lost; 3, 4 and 5 arrive and are buffered and individually ACKed; only 2 times out and is resent; then the receiver delivers 2 to 5 together. Total: 9 transmissions (8 + 1 retransmission).
Sequence number space and window size
Sequence numbers live in a fixed field of k bits, so they wrap around after 2^k. The window must be small enough that the receiver can never confuse a retransmitted old packet with a new one that reuses its number.
Go-Back-N: N <= 2^k - 1
Selective Repeat: N <= 2^(k-1) (half the sequence space)
| Sequence bits k | Sequence numbers | Max GBN window | Max SR window |
|---|---|---|---|
| 2 | 0 to 3 | 3 | 2 |
| 3 | 0 to 7 | 7 | 4 |
| 4 | 0 to 15 | 15 | 8 |
Why SR needs half. With k = 2 (numbers 0 to 3) and an SR window of 3: the sender sends 0, 1, 2. The receiver gets all three, ACKs them, and moves its window to expect 3, 0, 1. Now suppose all three ACKs are lost. The sender times out and resends the old packet 0. The receiver's window contains 0 (meaning the new packet 0), so it wrongly accepts the old duplicate as new data. With a window of 2 or less, the receiver's new window never overlaps the sender's old one, and the ambiguity disappears.
Worked example: choosing window and sequence bits
A link runs at 10 Mbps, packets are 1,000 bytes and RTT is 10 ms.
- L/R = 8,000 / 10^7 = 0.8 ms.
- Stop-and-wait utilisation = 0.8 / (10 + 0.8) = 7.4%.
- Window to fill the pipe: (10 + 0.8) / 0.8 = 13.5, so N = 14.
- With N = 4: utilisation = 4 × 0.8 / 10.8 = 29.6%.
- Sequence bits for N = 14: GBN needs 2^k - 1 at least 14, so k = 4 (15). SR needs 2^(k-1) at least 14, so k = 5 (16).
GBN vs Selective Repeat
| Aspect | Go-Back-N | Selective Repeat |
|---|---|---|
| Receiver accepts | Only in-order packets | Out-of-order packets within its window |
| Receiver buffer | One packet | Up to N packets |
| ACKs | Cumulative | Individual |
| Timers | One (oldest unACKed) | One per packet |
| On loss | Resend whole window from the lost packet | Resend only the lost packet |
| Max window (k bits) | 2^k - 1 | 2^(k-1) |
| Best when | Losses are rare, receiver is simple | Losses are common or the window is large |
What TCP actually does
TCP is a hybrid. It uses cumulative ACKs like GBN, but the receiver buffers out-of-order segments like SR, and the sender usually retransmits only the missing segment. With the SACK option (selective acknowledgement), the receiver also reports exactly which blocks it holds, so the sender can fill several holes precisely. TCP keeps a single retransmission timer per connection, and uses fast retransmit: three duplicate ACKs trigger resending the missing segment without waiting for the timer. Details are in the next lesson.
Common mistake
Saying "TCP is Go-Back-N." TCP's cumulative ACK resembles GBN, but TCP receivers keep out-of-order data and senders do not blindly resend the whole window. The accurate answer is "a hybrid with cumulative ACKs, receiver buffering and, with SACK, selective retransmission."
TCP vs UDP
| Feature | TCP | UDP |
|---|---|---|
| Connection | Connection-oriented (handshake) | Connectionless |
| Reliability | Retransmits lost data | None |
| Ordering | In order | No guarantee |
| Data unit | Byte stream (no message boundaries) | Datagrams (boundaries kept) |
| Header | 20 to 60 bytes | 8 bytes |
| Flow control | Yes (receive window) | No |
| Congestion control | Yes | No (application's responsibility) |
| Setup latency | 1 RTT before data | None |
| Demultiplexing | 4-tuple | Destination IP and port |
| Broadcast / multicast | No | Yes |
| State at server | Per connection | None required |
| Typical uses | Web (HTTP/1.1, HTTP/2), email, SSH, databases, file transfer | DNS, DHCP, NTP, VoIP, video calls, games, QUIC |
When to use which
The question is never "which is better", but "which costs matter for this application".
- DNS: UDP by default. A query and answer usually fit in one datagram each, so a TCP handshake would double the latency for no benefit, and the resolver simply retries on timeout. DNS switches to TCP when a response is too large for UDP (it gets a truncated flag) and for zone transfers. DNS over TLS and DNS over HTTPS run over TCP (or QUIC) for privacy. See DNS.
- Live video calls and VoIP: UDP (usually RTP over UDP). A packet that arrives late is useless; it is better to skip it and keep playing. The application handles loss with techniques like forward error correction and adapts its bitrate itself.
- Video on demand (YouTube-style streaming): traditionally TCP (HTTP-based streaming such as HLS or DASH), because a buffer of several seconds hides retransmission delays and reliability is wanted. Increasingly delivered over HTTP/3, which runs on QUIC over UDP.
- Online games: usually UDP for position updates, where only the latest state matters, often with a custom reliable channel for important events such as purchases.
- File transfer, web pages, APIs, database connections, SSH: TCP, because every byte must arrive intact and in order.
- QUIC and HTTP/3: QUIC is a transport protocol built in user space on top of UDP. It provides reliability, congestion control and built-in TLS 1.3 encryption, with independent streams so one lost packet only stalls its own stream (no transport-level head-of-line blocking across streams), a combined transport and crypto handshake (1 RTT, or 0-RTT on resumption), and connection migration (a phone moving from Wi-Fi to mobile data keeps the connection, because connections are identified by an ID, not the 4-tuple). It uses UDP because middleboxes would block a brand-new IP protocol, and because building it in user space allows fast iteration without waiting for operating-system kernels to update.
Interview tip
For "why does DNS use UDP?" give three reasons: small single-packet messages, no handshake latency, and no per-client state on busy servers. Then add the exception: large responses and zone transfers use TCP. Interviewers often ask the exception next.
Interview questions
Q1. How does the transport layer differ from the network layer?
The network layer delivers packets between hosts on a best-effort basis. The transport layer delivers data between processes using port numbers, and TCP adds reliability, ordering, flow control and congestion control on top. IP has no idea which application a packet is for.
Q2. How does a server tell apart two clients connected to the same port?
A TCP connection is identified by the 4-tuple of source IP, source port, destination IP and destination port. Two clients to the same server port differ in source IP or source port, so each maps to its own connected socket. UDP, by contrast, delivers everything for one destination IP and port to a single socket.
Q3. Draw the three-way handshake with sequence numbers.
Client sends SYN with seq = x. Server replies SYN+ACK with seq = y and ack = x + 1. Client sends ACK with seq = x + 1 and ack = y + 1. The SYN flags consume one sequence number each, which is why the acknowledgements are ISN + 1.
Q4. Why is the handshake three-way and not two-way?
Both sides must send their initial sequence number and have it acknowledged; the server combines its ACK and SYN, so three messages suffice. A two-way handshake would let a delayed duplicate SYN create a half-open connection at the server, and the server could not confirm the client received its ISN.
Q5. Why does TIME_WAIT exist and how long does it last?
The active closer waits 2 × MSL so that it can resend the final ACK if it was lost, and so that old duplicate segments of this connection expire before the same 4-tuple can be reused. RFC 793 suggests an MSL of 2 minutes; Linux uses a fixed 60-second TIME_WAIT.
Q6. Your server shows thousands of sockets in CLOSE_WAIT. What is wrong?
The peers have closed their side, TCP has acknowledged the FIN, but the application has not called close on those sockets. It is an application bug, usually a missing close in an error path or a leaked pooled connection. They will remain until the process closes them or exits.
Q7. What is a SYN flood and how do SYN cookies stop it?
The attacker sends many SYNs, often spoofed, so the server fills its queue of half-open connections and drops real clients. With SYN cookies the server keeps no state for a SYN; it encodes the connection details and a keyed hash into its ISN. Only when a valid ACK returns does it rebuild the details and create the connection.
Q8. What is the utilisation of stop-and-wait for 1,500-byte packets on a 1 Gbps link with 30 ms RTT?
Transmission time is 12 µs, so utilisation is 0.012 / 30.012, about 0.04%, giving roughly 400 kbps of throughput. A window of about 2,501 packets would be needed to fill the link.
Q9. Compare Go-Back-N and Selective Repeat.
GBN's receiver discards out-of-order packets and uses cumulative ACKs, and the sender resends the whole window after a timeout; it is simple but wastes bandwidth on loss. SR's receiver buffers out-of-order packets and ACKs each one, and the sender resends only the lost ones; it needs more buffering and timers and a window of at most half the sequence space.
Q10. With 3-bit sequence numbers, what are the maximum GBN and SR window sizes?
GBN allows 2^3 - 1 = 7. SR allows 2^(3-1) = 4. A larger SR window lets the receiver mistake a retransmitted old packet for a new packet with the same number.
Q11. Is TCP Go-Back-N or Selective Repeat?
Neither exactly. It uses cumulative ACKs like GBN but receivers buffer out-of-order data, and with the SACK option senders retransmit only missing segments, which resembles SR. It also uses fast retransmit on three duplicate ACKs.
Q12. Why does DNS use UDP, and when does it use TCP?
Queries and answers are small, so one datagram each avoids a handshake and keeps resolvers stateless; clients just retry on timeout. TCP is used when a response is too large and comes back truncated, for zone transfers, and for encrypted DNS over TLS.
Q13. Why is QUIC built on UDP instead of being a new IP protocol?
Firewalls and NAT devices across the internet only reliably pass TCP and UDP, so a new protocol number would be blocked. UDP gives QUIC ports and a path through middleboxes, while QUIC itself implements reliability, congestion control, encryption and streams in user space, where it can be updated quickly.
Q14. What is the difference between a byte stream and a datagram service?
TCP is a byte stream: the receiver sees a sequence of bytes with no record of how the sender split its writes, so applications must frame their own messages (for example with length prefixes). UDP preserves boundaries: each send is delivered as one whole datagram or not at all.
Q15. What does an RST segment mean?
It aborts a connection immediately. It is sent when a segment arrives for a port with no listener, for a connection that no longer exists, or when an application aborts deliberately. Unlike FIN, it does not wait for data to be delivered and does not enter TIME_WAIT.
Key takeaways
- The transport layer gives process-to-process delivery with 16-bit ports; UDP demultiplexes on destination IP and port, TCP on the 4-tuple.
- UDP is an 8-byte header, a checksum and message boundaries, with no reliability, ordering or congestion control.
- TCP is a connection-oriented, reliable, ordered, full-duplex byte stream with flow and congestion control; sequence numbers count bytes and ACKs are cumulative.
- Handshake: SYN (x), SYN+ACK (y, x+1), ACK (x+1, y+1). Teardown uses a FIN and ACK in each direction.
- TIME_WAIT (2 MSL) belongs to the active closer and is normal; piles of CLOSE_WAIT mean the application is not closing sockets.
- SYN floods exhaust half-open state; SYN cookies avoid storing it.
- Stop-and-wait utilisation is (L/R) / (RTT + L/R); pipelining with a window near the bandwidth-delay product fills the link.
- GBN: cumulative ACKs, resend the window, N up to 2^k - 1. SR: individual ACKs, resend only losses, N up to 2^(k-1). TCP is a hybrid with SACK.
- Use UDP when latency beats completeness or messages are tiny (DNS, calls, games, QUIC); use TCP when every byte must arrive in order.
Next lesson
Continue with TCP flow and congestion control.

