Every time you open a website, send an email or call an API, your computer first has to answer one question: "which IP address belongs to this name?" The Domain Name System (DNS) is the distributed, hierarchical database that answers it. DNS is so reliable that most developers never think about it, until a misconfigured record takes a site offline or a stale cache sends half the users to an old server.
Interviewers like DNS because it touches many ideas at once: hierarchy and delegation, caching and consistency, UDP versus TCP, security, and load balancing. Typical probes are "walk me through what happens when you resolve www.example.com", "what is the difference between a recursive and an iterative query", "what is a CNAME and why can't you put one at the root of a domain", "what is TTL and why do DNS changes take time to propagate", and "how would an attacker poison a DNS cache". This lesson answers all of these step by step.
Why DNS exists
Computers on the Internet talk to each other using IP addresses, numbers such as 203.0.113.10 (IPv4) or 2606:2800:21f:cb07:6820:80da:af6b:8b2c (IPv6). Humans are bad at remembering numbers and good at remembering names. DNS bridges the two: it maps human-friendly domain names to machine-friendly addresses and other data.
There is a second, less obvious reason DNS matters: indirection. Because clients look up the name every time (subject to caching), the operator can change the address behind the name without telling anyone. You can move a website to a new server, add servers in another country, or fail over to a backup, and users keep typing the same name.
The early approach and why it failed
Before DNS (early 1980s), every host on the ARPANET downloaded a single file, HOSTS.TXT, from a central server. Each line mapped a name to an address. Your laptop still has a descendant of this file: /etc/hosts on Linux and macOS, C:\Windows\System32\drivers\etc\hosts on Windows.
A single central file stopped working as the network grew:
- Scale. Every host downloaded the full file, and the file kept getting bigger.
- Single point of failure and bottleneck. One server handed the file to everyone.
- Name collisions. One authority had to approve every name.
- Staleness. Changes took days to reach everyone.
DNS, designed in 1983 and described in RFC 1034 and RFC 1035, replaced the file with a distributed database. No single server knows every name. Instead, responsibility is split into a tree, and each part of the tree is managed by whoever owns it.
The namespace hierarchy
A domain name is read right to left, from the most general part to the most specific. Take www.example.com. (note the trailing dot):
www . example . com . (root)
| | | |
| | | +-- root: the top of the tree, written as "."
| | +--------- TLD (top-level domain): com
| +----------------- second-level domain: example
+------------------------ subdomain / host label: www
Each part between dots is a label. The full name, ending in the root dot, is a fully qualified domain name (FQDN). In everyday use we drop the trailing dot, but DNS software treats it as present.
The namespace is a tree:
. (root)
+---------------+----------------+
com org in
+----+-----+ | +----+----+
example google wikipedia co gov
| |
+---+---+ flipkart
www mail
Root servers
At the top sit the root name servers. They do not know the address of www.example.com. They only know which servers are responsible for each top-level domain (TLD). There are 13 named root server identities, a.root-servers.net through m.root-servers.net, run by 12 organisations. Each identity is not one machine: through anycast (many servers in different places announcing the same IP address, with routing sending you to the nearest), each is served by hundreds of physical servers worldwide.
TLD servers
Top-level domain servers are responsible for one TLD, such as .com, .org, or a country-code TLD such as .in or .uk. They do not know the final address either. They know which servers are authoritative for each second-level domain under them, for example which servers handle example.com.
Common TLD categories:
| Type | Examples | Meaning |
|---|---|---|
| Generic (gTLD) | .com, .org, .net, .dev | Open or themed registrations |
| Country code (ccTLD) | .in, .uk, .de, .jp | Assigned to countries and territories |
| Sponsored / restricted | .gov, .edu, .mil | Limited to specific organisations |
| Infrastructure | .arpa | Used for reverse DNS (in-addr.arpa) |
Authoritative servers
An authoritative name server holds the actual records for a domain and gives definitive answers about it. When you buy example.com from a registrar and host its DNS at a provider (Cloudflare, Route 53, your registrar), that provider runs the authoritative servers. The answer they give is the source of truth.
Zones and delegation
The tree is divided into zones. A zone is the part of the namespace one organisation manages. The .com zone contains a pointer saying "for example.com, ask ns1.example-dns.net and ns2.example-dns.net". That pointer is called a delegation, and it is made of NS records in the parent zone.
The owner of example.com can delegate further, for example handing eu.example.com to a different team's servers. Delegation is how DNS scales: each organisation manages its own slice without asking anyone else.
A glue record solves a chicken-and-egg problem. If example.com says its name server is ns1.example.com, a resolver would need to resolve example.com to find the server for example.com. To break the loop, the parent (.com) also stores the IP address of ns1.example.com. That extra address is the glue.
How resolution works
Several roles cooperate when your browser resolves a name:
- Stub resolver. A small library in your operating system. It does not chase answers itself. It sends one question to a configured resolver and waits.
- Recursive resolver (also called a recursive name server or full resolver). Run by your ISP, your company, or a public service such as
8.8.8.8(Google) or1.1.1.1(Cloudflare). It does the legwork: it asks root, TLD and authoritative servers until it has a final answer, and it caches what it learns. - Authoritative servers. Root, TLD and the domain's own servers, which answer only about the zones they hold.
Recursive versus iterative queries
These two words describe what the asker expects:
- A recursive query says "give me the final answer, do whatever it takes". The server must either return the answer or an error. Your stub resolver sends recursive queries to the recursive resolver.
- An iterative query says "tell me the best you know". The server answers with the record if it has it, or with a referral: "I don't know, but ask these servers". The recursive resolver sends iterative queries to root, TLD and authoritative servers.
Root and TLD servers almost never accept recursive queries. If they had to chase every lookup on the Internet to completion, they would collapse under the load. Pushing the work to recursive resolvers spreads it out and lets each resolver cache results for its own users.
Step-by-step resolution of www.example.com
Assume every cache is empty.
Your laptop Recursive Root .com TLD example.com
(stub) resolver server server authoritative
| | | | |
|-- 1. A? www.example.com (recursive) --->| | |
| | | | |
| |-- 2. A? www.example.com ---->| |
| | (iterative) | | |
| |<- 3. referral: "ask .com | |
| | servers a.gtld-servers.net ..." |
| | | |
| |-- 4. A? www.example.com ---->| |
| |<- 5. referral: "ask ns1/ns2 for example.com"
| | |
| |-- 6. A? www.example.com ------------------>|
| |<- 7. answer: www.example.com A 203.0.113.10
| | (TTL 3600) |
|<- 8. answer + caches it for up to 3600 s |
| | |
In words:
- Your browser asks the OS. The OS checks its own cache and the hosts file. On a miss, the stub resolver sends a recursive query for the A record of
www.example.comto the configured resolver, say1.1.1.1. - The resolver checks its cache. Empty, so it starts at the top. It already knows the root server addresses, because they are shipped with the software in a root hints file.
- The root server replies with a referral: the NS records for
.complus their addresses (glue). - The resolver asks a
.comserver the same question. - The
.comserver replies with a referral to the authoritative servers ofexample.com. - The resolver asks one of them.
- The authoritative server returns the A record and its TTL (time to live, explained below). The AA (authoritative answer) flag in the reply tells the resolver this came from the source of truth.
- The resolver caches every piece it learned (the
.comNS records, theexample.comNS records, and the final answer) and returns the address to your laptop.
The next lookup for mail.example.com skips steps 2 to 5 entirely, because the resolver already knows who is authoritative for example.com. This is why most lookups take a few milliseconds instead of hundreds.
Interview tip
When asked "how does DNS resolution work", name the four layers of caching before you name the servers: browser cache, OS cache (and hosts file), recursive resolver cache, then the root to TLD to authoritative walk. It shows you know that the full walk is the rare case, not the common one.
Caching and TTL
Caching is what makes DNS fast and scalable. Every record carries a TTL (time to live), a number of seconds the record may be cached. When a resolver caches a record with TTL 3600, it serves that record from memory for up to an hour, counting down. When the count reaches zero, the entry expires and the next request triggers a fresh lookup.
Caches exist at several layers:
| Layer | Example | Notes |
|---|---|---|
| Browser | Chrome's internal host cache (chrome://net-internals/#dns) | Short-lived, per browser |
| Operating system | systemd-resolved on Linux, mDNSResponder on macOS, DNS Client service on Windows | Shared by all apps on the machine |
| Recursive resolver | ISP resolver, 8.8.8.8, 1.1.1.1 | Shared by thousands or millions of users |
| Application | Java's InetAddress cache, connection pools holding open sockets | Can ignore TTL if misconfigured |
Choosing a TTL
The TTL is a trade-off between speed and agility:
- Long TTL (hours to a day). Fewer queries reach your authoritative servers, lookups are faster for users, and you survive a short outage of your DNS provider because resolvers still have the answer. But a change takes up to that long to reach everyone.
- Short TTL (30 to 300 seconds). Changes take effect quickly, which is essential for failover and DNS-based load balancing. But resolvers query you more often, and users see more cache misses.
Worked example: planning a migration
You are moving api.example.com from address 198.51.100.10 to 203.0.113.20. Its A record currently has TTL 86400 (one day).
- Two days before the move, lower the TTL to 300 seconds. Resolvers that cached the old record still hold it for up to 86400 seconds, so you must wait at least one full old TTL (one day) before every cache has picked up the new, short TTL.
- On the day, change the address to
203.0.113.20. Now the worst-case staleness is 300 seconds, so within about five minutes almost all resolvers return the new address. - Keep the old server running for a while. Some clients and resolvers ignore TTLs, and long-lived connections may still be open.
- After a day or two, raise the TTL back to a longer value.
If you skip step 1 and change the address with the one-day TTL in place, some users keep hitting the old server for up to 24 hours.
Common mistake
"DNS propagation" sounds like changes being pushed outward across the Internet. Nothing is pushed. Your authoritative servers update almost instantly. What you are waiting for is cached copies expiring at resolvers around the world, each on its own TTL clock.
Negative caching
Resolvers also cache failures. If you ask for new.example.com and it does not exist, the authoritative server returns NXDOMAIN ("non-existent domain"). The resolver caches that "no" for a period defined by the zone's SOA record (specifically, the smaller of the SOA record's own TTL and its minimum field, per RFC 2308).
This causes a classic surprise: you query a name before creating it, get NXDOMAIN, create the record, and the name still "doesn't exist" for minutes because the negative answer is cached. The fix is to wait out the negative TTL, or flush the cache of the resolver you control.
Record types
A DNS resource record has five parts: name, TTL, class (almost always IN for Internet), type, and data. For example:
www.example.com. 3600 IN A 203.0.113.10
name TTL class type data
The record types every engineer should know:
| Type | Maps | Example data | Typical use |
|---|---|---|---|
| A | name to IPv4 address | 203.0.113.10 | Point a host at a server |
| AAAA | name to IPv6 address | 2606:2800:21f::1 | Same, for IPv6 ("quad-A") |
| CNAME | name to another name (alias) | example.cdn-provider.net. | Point a subdomain at a CDN or SaaS host |
| MX | domain to mail server, with priority | 10 mail.example.com. | Where to deliver email for the domain |
| NS | zone to its authoritative name servers | ns1.example-dns.net. | Delegation |
| TXT | name to free text | "v=spf1 include:_spf.google.com ~all" | SPF, DKIM, domain-ownership proofs |
| SOA | zone to its administrative data | primary NS, admin email, serial, timers | One per zone; controls negative caching and zone transfers |
| PTR | IP address to name (reverse) | mail.example.com. | Reverse DNS, mail server reputation, logs |
| SRV | service to host and port, with priority and weight | 10 60 5060 sip1.example.com. | Service discovery (SIP, XMPP, LDAP, Kubernetes) |
| CAA | domain to allowed certificate authorities | 0 issue "letsencrypt.org" | Restrict which CAs may issue certificates |
A and AAAA
The most common records. A name can have several A records; resolvers return all of them, and clients usually try them in the order received. A dual-stack host (IPv4 and IPv6) has both A and AAAA. Modern clients use Happy Eyeballs (RFC 8305): they try IPv6 and IPv4 almost in parallel and use whichever connects first.
CNAME and its rules
A CNAME (canonical name) says "this name is an alias; look up that other name instead". If shop.example.com is a CNAME to shops.myshopify.com, the resolver follows the chain and returns the final A records.
Two rules trip people up:
- A CNAME cannot coexist with other records at the same name. If
shop.example.comhas a CNAME, it cannot also have an MX or TXT record. - So you cannot put a CNAME at the zone apex (the bare domain
example.com), because the apex must have SOA and NS records.
That is a problem when a CDN or platform gives you only a hostname, not an IP. Providers work around it with non-standard records called ALIAS, ANAME or CNAME flattening: the provider resolves the target itself and serves the result as A/AAAA records at the apex.
MX
An MX (mail exchanger) record has a priority (lower number = more preferred) and a host name:
example.com. 3600 IN MX 10 mx1.example.com.
example.com. 3600 IN MX 20 mx2.example.com.
A sending mail server tries mx1 first and falls back to mx2 if mx1 is unreachable. The MX target must be a name with A/AAAA records, not a CNAME and not an IP address.
NS and SOA
Every zone has NS records listing its authoritative servers and exactly one SOA (start of authority) record:
example.com. 3600 IN SOA ns1.example-dns.net. hostmaster.example.com. (
2026101001 ; serial: bump on every change
7200 ; refresh: how often secondaries check for changes
3600 ; retry: how long secondaries wait after a failed check
1209600 ; expire: when secondaries stop answering if primary is gone
300 ) ; minimum: used for negative-caching TTL
The admin email is written with the first @ replaced by a dot: hostmaster.example.com means hostmaster@example.com. Secondary (backup) authoritative servers copy the zone from the primary through a zone transfer (AXFR for full, IXFR for incremental), triggered when the serial number increases.
PTR and reverse DNS
Reverse DNS answers "which name belongs to this IP?". IPv4 addresses are written backwards under the special domain in-addr.arpa, so the PTR for 203.0.113.25 lives at 25.113.0.203.in-addr.arpa. Reversing the order keeps the hierarchy consistent: the most general part (the network) comes last, just like .com.
Reverse zones are controlled by whoever owns the IP block, usually your hosting provider, not by you. Mail servers check that a sending server's IP has a PTR record whose name points back to the same IP (forward-confirmed reverse DNS). Missing reverse DNS is a common reason mail lands in spam.
SRV
An SRV record tells clients the host and port for a service. Its name follows the pattern _service._protocol.domain:
_sip._tcp.example.com. 3600 IN SRV 10 60 5060 sip1.example.com.
_sip._tcp.example.com. 3600 IN SRV 10 40 5060 sip2.example.com.
| | | |
priority | port target
weight
Clients use the lowest priority first. Among equal priorities, they pick servers in proportion to weight: here sip1 gets about 60 percent and sip2 about 40 percent. Browsers do not use SRV for HTTP, but Kubernetes, SIP, XMPP, LDAP and Active Directory do.
TXT and CAA
TXT records hold arbitrary strings. They are the Swiss army knife of DNS: email authentication (SPF, DKIM, DMARC, covered in the application protocols lesson), domain verification for Google Search Console or certificate issuance, and more.
CAA (certification authority authorisation) lists which certificate authorities may issue TLS certificates for the domain. A CA must check CAA before issuing. If example.com has 0 issue "letsencrypt.org", any other CA must refuse. It limits the damage if an attacker tries to get a certificate from a careless CA.
Reading dig output
dig (domain information groper) is the standard tool for querying DNS. It ships with BIND tools on Linux and macOS. Here is a typical query with each section explained (addresses use documentation ranges):
$ dig www.example.com A
; <<>> DiG 9.18.24 <<>> www.example.com A
;; global options: +cmd
;; Got answer:
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 41523
;; flags: qr rd ra; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1
;; OPT PSEUDOSECTION:
; EDNS: version: 0, flags:; udp: 1232
;; QUESTION SECTION:
;www.example.com. IN A
;; ANSWER SECTION:
www.example.com. 300 IN CNAME example.cdn.net.
example.cdn.net. 60 IN A 203.0.113.10
;; Query time: 23 msec
;; SERVER: 192.168.1.1#53(192.168.1.1) (UDP)
;; WHEN: Sat Oct 10 10:15:02 IST 2026
;; MSG SIZE rcvd: 92
Line by line:
- status: NOERROR means the query succeeded. Other values: NXDOMAIN (name does not exist), SERVFAIL (the resolver could not get an answer, often a broken authoritative server or a DNSSEC validation failure), REFUSED (the server will not answer you, for example a non-recursive server asked a recursive question).
- flags:
qr= this is a response;rd= recursion desired (we asked for it);ra= recursion available (the server offers it). Anaaflag would mean the answer came directly from an authoritative server. Atcflag means the answer was truncated and the client should retry over TCP. - QUESTION SECTION repeats what you asked.
- ANSWER SECTION shows the records. Here
www.example.comis a CNAME to a CDN host, and dig shows the followed chain ending in an A record. - The TTL column (300, 60) is the remaining cache time. Ask again in 10 seconds and a caching resolver will show 290 and 50.
- AUTHORITY and ADDITIONAL sections carry NS records (for referrals) and extra helpful records, such as glue addresses.
- SERVER shows which resolver answered, and over which transport.
Useful variations:
dig example.com MX +short # just the data, one per line
dig @8.8.8.8 example.com # ask a specific resolver
dig example.com NS # who is authoritative?
dig @ns1.example-dns.net example.com A +norecurse # ask the source of truth
dig +trace www.example.com # do the root -> TLD -> auth walk yourself
dig -x 203.0.113.25 # reverse lookup (PTR)
dig example.com +dnssec # include DNSSEC signatures
dig +trace is the best way to see delegation with your own eyes: it starts at the root and prints every referral. Comparing the answer from your resolver with the answer from @ the authoritative server tells you instantly whether a problem is a stale cache or a wrong record.
Interview tip
If asked "a DNS change isn't showing up, how do you debug it?", say: query the authoritative server directly with dig @ns1... +norecurse to confirm the record is right, then query public resolvers to see what they cache and the remaining TTL, then check local caches. This separates "the record is wrong" from "the record is right but cached".
Transport: UDP, TCP and port 53
DNS servers listen on port 53 for both UDP and TCP.
UDP is the default for ordinary queries. A query and its answer each usually fit in a single packet, so UDP's lack of a handshake saves a round trip. If a reply is lost, the client simply retries after a timeout.
The original DNS standard limited UDP messages to 512 bytes. EDNS(0) (extension mechanisms for DNS) lets the client advertise a bigger buffer, such as 1232 or 4096 bytes. The value 1232 is a modern recommendation chosen to avoid IP fragmentation, which is unreliable on the Internet.
TCP is used when:
- The answer does not fit. The server sets the TC (truncated) flag, and the client retries the same query over TCP. Large answers come from DNSSEC signatures, many records, or long TXT records.
- Zone transfers (AXFR/IXFR) between primary and secondary servers, which move whole zones.
- Encrypted DNS such as DNS over TLS (which runs over TCP, port 853).
Common mistake
"DNS uses only UDP" is a frequently given wrong answer. Since RFC 7766, TCP support is required for all DNS implementations, and blocking TCP 53 in a firewall breaks large responses and DNSSEC.
DNS security
Classic DNS was designed for a trusting network. Messages are unauthenticated (anyone can forge a reply) and unencrypted (anyone on the path can read and log your lookups). Several attacks and defences follow from that.
Cache poisoning
In cache poisoning, an attacker tricks a recursive resolver into caching a forged record, for example bank.example.com A 198.51.100.66 pointing at an attacker's server. Every user of that resolver is then sent to the wrong place until the TTL expires.
How it works in the classic form:
- The attacker makes the resolver look up a name, for example by sending it a query.
- While the resolver waits for the real authoritative answer, the attacker floods it with forged replies pretending to come from the authoritative server.
- A forged reply is accepted if it matches the question, comes from the right source IP, arrives on the right port, and carries the right 16-bit transaction ID.
- If a forged reply arrives before the real one, it is cached.
With only 65,536 possible IDs, older resolvers were easy targets. In 2008 Dan Kaminsky showed the attack could be repeated rapidly by asking for random, non-existent subdomains (a1.bank.example.com, a2...), each giving a fresh chance, and slipping in forged NS records that hijacked the whole domain.
Defences:
- Source port randomisation. Resolvers send each query from a random UDP port, adding about 16 bits of randomness. The attacker must now guess the ID and the port.
- 0x20 encoding. Randomly mixing the letter case of the query name (
wWw.ExAmPlE.cOm) and checking that the reply echoes it exactly. - Bailiwick checks. Ignoring records in a reply that are outside the zone the server is responsible for.
- DNSSEC, the real fix, below.
DNSSEC
DNSSEC (DNS Security Extensions) adds digital signatures to DNS records so a resolver can verify that an answer really came from the zone owner and was not modified. (Signatures are explained in the TLS and network security lesson.)
The pieces:
- RRSIG: a signature over a set of records of the same name and type.
- DNSKEY: the zone's public keys, used to verify RRSIGs.
- DS (delegation signer): a hash of the child zone's key, stored in the parent zone. This links the chain.
- NSEC / NSEC3: signed proof that a name does not exist, so attackers cannot forge NXDOMAIN either.
Trust flows down from the root, whose key is built into validating resolvers:
root key (trust anchor, built into resolver)
--> signs DS record for .com in the root zone
--> DS matches .com DNSKEY, which signs DS for example.com
--> DS matches example.com DNSKEY, which signs
the RRSIG over www.example.com A 203.0.113.10
If any link fails, a validating resolver returns SERVFAIL rather than a possibly forged answer. DNSSEC provides integrity and authenticity, not confidentiality: the queries and answers are still readable by anyone on the path.
DNS over TLS and DNS over HTTPS
Privacy is a separate problem. Plain DNS reveals every site you visit to your network operator and anyone on the path. Two standards encrypt the hop between your device and the recursive resolver:
| DNS over TLS (DoT) | DNS over HTTPS (DoH) | |
|---|---|---|
| Standard | RFC 7858 | RFC 8484 |
| Transport | TLS over TCP | HTTPS (HTTP/2 or HTTP/3) |
| Port | 853 | 443 |
| Looks like | Clearly DNS traffic, easy to identify or block | Blends with ordinary web traffic |
| Common in | Android "Private DNS", system resolvers | Browsers (Firefox, Chrome), apps |
Both protect only the client-to-resolver hop. The resolver itself still sees your queries, and the resolver-to-authoritative hops are usually plain. DNSSEC and DoH/DoT solve different problems and work well together: DNSSEC proves the answer is genuine, DoH/DoT hides the question.
Interview tip
A crisp line: "DNSSEC gives authenticity, DoH and DoT give privacy. Neither replaces the other." Interviewers often check whether you confuse them.
Amplification and other DNS attacks
DNS amplification is a distributed denial of service (DDoS) technique, an attack that overwhelms a target with traffic from many sources:
- The attacker sends small DNS queries (around 60 bytes) to many open resolvers (resolvers that answer anyone on the Internet).
- The queries carry a spoofed source IP: the victim's address. UDP has no handshake, so the resolver cannot tell.
- The queries ask for large answers (
ANYqueries, or DNSSEC-signed records), perhaps 3,000 bytes. - The resolvers send those large replies to the victim.
A 60-byte query triggering a 3,000-byte reply is a 50x amplification: 1 Gbps of attacker bandwidth becomes about 50 Gbps at the victim. Defences include closing open resolvers, response rate limiting (RRL) on authoritative servers, ingress filtering by ISPs so spoofed packets never leave their networks (BCP 38), and refusing or minimising ANY responses (RFC 8482).
Other attacks worth naming:
- DNS hijacking: changing which resolver a victim uses (malware altering settings, a compromised home router) or changing a domain's records at the registrar after stealing the account. Defence: registrar locks and multi-factor authentication.
- NXDOMAIN / random-subdomain floods ("water torture"): flooding an authoritative server with queries for random non-existent names so caches cannot help.
- DNS tunnelling: hiding data inside DNS queries and answers to sneak it past firewalls, used for exfiltration and command-and-control.
- Subdomain takeover: a CNAME still points at a deleted cloud resource (say a removed storage bucket); an attacker claims that resource name and serves content on your subdomain. Defence: delete DNS records when you delete resources.
DNS-based load balancing and failover
Because DNS sits in front of every connection, it is a cheap first layer of traffic management.
Round-robin DNS
Publish several A records for one name. Resolvers rotate the order of the records in responses, and clients usually take the first, so traffic spreads across servers.
api.example.com. 60 IN A 203.0.113.10
api.example.com. 60 IN A 203.0.113.11
api.example.com. 60 IN A 203.0.113.12
Limitations:
- No health awareness. If
.11dies, plain DNS keeps handing it out. Clients that try only the first address fail. - Uneven spread. A big resolver caches one ordering and hands it to thousands of users. NAT and corporate resolvers make one "client" look like many.
- Caching delays changes, as discussed.
GeoDNS and latency-based routing
Managed DNS providers can answer differently depending on who asks. They estimate the client's location from the resolver's IP address (or the EDNS Client Subnet option, which forwards part of the user's IP) and return the address of the nearest region. AWS Route 53 calls these geolocation and latency-based routing policies. CDNs use the same idea to send you to a nearby edge server.
Health-checked failover
The provider probes your servers, for example with an HTTP request every 30 seconds. If the primary fails its checks, the provider stops returning its address and returns a backup instead. Failover time is roughly "time to detect + TTL", which is why failover records use short TTLs such as 60 seconds.
Weighted records
Return addresses in proportion to configured weights, say 90 percent to the stable fleet and 10 percent to a canary release. Useful for gradual rollouts and migrations.
| Technique | Strength | Weakness |
|---|---|---|
| Round robin | Free, simple | No health checks, uneven |
| Weighted | Gradual rollouts | Coarse, cache-affected |
| Geo / latency | Lower latency per region | Resolver location may not match user |
| Failover | Survives region outage | Recovery bounded by TTL plus detection |
DNS load balancing is coarse because of caching. Real systems pair it with load balancers (covered in modern networking and tools and load balancing) or anycast: DNS picks the region, the load balancer picks the server.
Common DNS problems and how to spot them
| Symptom | Likely cause | How to check |
|---|---|---|
| Change visible to some users, not others | Caches still holding old record | dig against several public resolvers, compare TTLs |
| New name says NXDOMAIN for minutes | Negative caching | Check SOA minimum; query authoritative directly |
| SERVFAIL from a validating resolver only | Broken DNSSEC (expired signatures, wrong DS after a provider move) | dig +dnssec, use a DNSSEC debugger, try +cd (checking disabled) |
| Works on one network, fails on another | Split-horizon DNS, corporate resolver, hosts file | dig @8.8.8.8 vs default resolver; check /etc/hosts |
| Mail rejected or marked spam | Missing PTR, SPF, DKIM or DMARC | dig -x, dig TXT example.com |
| Domain stops resolving entirely | Expired domain, wrong NS at registrar | whois, dig NS example.com +trace |
| Large responses fail | TCP 53 blocked, fragmentation | dig +tcp; check firewall |
Two definitions from that table:
- Split-horizon DNS means returning different answers to internal and external clients, for example private IPs inside the office and public IPs outside.
- Lame delegation means the parent's NS records point at a server that is not actually authoritative for the zone, a common result of moving DNS providers carelessly.
Common mistake
Changing name server providers without updating DNSSEC. If the parent zone still has a DS record for the old provider's key, every validating resolver will reject your new answers with SERVFAIL. Remove or update the DS record as part of the move.
A quick look at the DNS message format
You do not need to memorise the wire format, but knowing its shape helps you reason about attacks and dig output:
+---------------------------------------------+
| Header (12 bytes) |
| ID (16 bits) | flags: QR AA TC RD RA, |
| | opcode, RCODE |
| QDCOUNT ANCOUNT NSCOUNT ARCOUNT |
+---------------------------------------------+
| Question: name, type (A, MX...), class |
+---------------------------------------------+
| Answer: resource records |
+---------------------------------------------+
| Authority: NS records (referrals) |
+---------------------------------------------+
| Additional: glue, EDNS OPT record |
+---------------------------------------------+
The same format is used for queries and responses; the QR bit says which. The 16-bit ID is what a cache-poisoning attacker must guess. The RCODE field carries NOERROR, NXDOMAIN, SERVFAIL and so on.
Interview questions
Q1. What is DNS and why do we need it?
DNS is a distributed, hierarchical database that maps domain names to IP addresses and other data such as mail servers. Humans remember names, but networks route by address. DNS also adds a layer of indirection: operators can move or replicate services by changing records, without clients changing anything.
Q2. Explain recursive versus iterative resolution.
In a recursive query, the client asks the server to return a final answer and the server does all the work. In an iterative query, the server returns the best it knows, often a referral to another server, and the client continues. Stub resolvers send recursive queries to a recursive resolver, which then sends iterative queries to root, TLD and authoritative servers.
Q3. Walk through resolving www.example.com with empty caches.
The stub resolver asks the recursive resolver. The resolver asks a root server, which refers it to the .com TLD servers. A .com server refers it to the authoritative servers of example.com, which return the A record with a TTL. The resolver caches every referral and the answer, then returns the address to the client.
Q4. What is TTL, and how would you choose one?
TTL is the number of seconds a record may be cached. Long TTLs reduce query load and make you resilient to brief DNS outages, but slow down changes. Short TTLs allow fast changes and failover but increase load and latency. Lower the TTL well before a planned change, and wait at least one old TTL before making it.
Q5. Why can't you put a CNAME at the root of a domain?
A CNAME cannot coexist with any other record at the same name, and the zone apex must have SOA and NS records. Providers offer ALIAS, ANAME or CNAME flattening, which resolve the target themselves and return A/AAAA records at the apex.
Q6. Does DNS use TCP or UDP?
Both, on port 53. Ordinary queries use UDP for speed. TCP is used when a response is too large and comes back truncated, for zone transfers, and for DNS over TLS. TCP support is mandatory for compliant implementations.
Q7. What is negative caching?
Resolvers cache "this name does not exist" (NXDOMAIN) answers, for a time taken from the zone's SOA record. It reduces load from repeated lookups of missing names but means a newly created record can appear missing until the negative entry expires.
Q8. What is DNS cache poisoning and how is it prevented?
An attacker races the real authoritative server with forged replies, hoping the resolver caches a fake record. Mitigations are randomising the transaction ID and source port, 0x20 case randomisation, bailiwick checks, and DNSSEC, which lets resolvers verify signatures and reject forged answers.
Q9. What does DNSSEC do, and what doesn't it do?
DNSSEC signs records so resolvers can verify they come from the zone owner and were not modified, using a chain of trust from the root through DS and DNSKEY records. It does not encrypt anything, so it provides no privacy. DoH or DoT provide privacy on the client-to-resolver hop.
Q10. What is the difference between DoH and DoT?
Both encrypt DNS between client and resolver. DoT wraps DNS in TLS on a dedicated port, 853, so it is easy to identify and block. DoH sends DNS inside HTTPS on port 443, so it looks like normal web traffic.
Q11. How does DNS amplification work?
The attacker sends small queries with the victim's spoofed source IP to open resolvers, asking for large answers. The resolvers send the big replies to the victim, multiplying the attacker's bandwidth many times. Defences are closing open resolvers, response rate limiting, and ISP source-address filtering.
Q12. How can DNS be used for load balancing, and what are its limits?
Return multiple A records (round robin), weighted records, or region-specific answers (GeoDNS), with health checks to drop failed servers. Its limits come from caching: changes and failover are bounded by TTL, resolvers may not be near users, and distribution is uneven because one cached answer serves many users.
Q13. What is a PTR record used for?
It maps an IP address back to a name via the in-addr.arpa (or ip6.arpa) domain. Mail servers check it as part of spam filtering, and logs and tools like traceroute use it to show readable names.
Q14. What is glue and when is it needed?
Glue is an A or AAAA record for a name server, stored in the parent zone. It is needed when a domain's name server lives inside the domain itself (ns1.example.com for example.com), because otherwise resolving the name server would require resolving the domain first.
Q15. A user says your site works on mobile data but not on office Wi-Fi. What do you check?
Compare dig results from both networks. The office may use a corporate resolver with split-horizon records, a stale cache, a blocklist, or a firewall blocking certain DNS traffic. Also check the hosts file and the browser's DoH setting, which may bypass the office resolver entirely.
Key takeaways
- DNS is a distributed tree: root servers delegate to TLD servers, which delegate to authoritative servers that hold the real records.
- Stub resolvers send recursive queries; recursive resolvers do iterative lookups and cache everything they learn.
- TTL controls cache lifetime. "Propagation" is just caches expiring; lower TTLs before planned changes.
- Know the records: A, AAAA, CNAME, MX, NS, TXT, SOA, PTR, SRV, CAA, and that CNAMEs cannot sit at the apex.
digshows status, flags, TTLs and which server answered;dig +traceanddig @authoritativeseparate wrong records from stale caches.- DNS uses UDP 53 by default and TCP 53 for large answers and zone transfers.
- DNSSEC provides authenticity; DoH and DoT provide privacy. They solve different problems.
- DNS-based load balancing and failover are cheap but coarse, limited by caching.
Next lesson
Continue with HTTP and the web.

