Why Linux internals come up in interviews
Most servers you will work on run Linux. The earlier lessons in this track explained processes, scheduling, memory and file systems in general terms. This lesson shows how those ideas look on a real Linux machine: the process tree and PID 1, the /proc and /sys file systems, signals, file descriptors, how to read memory numbers from ps, top, free and vmstat, the page cache, the scheduler classes, and finally namespaces and cgroups, the two kernel features that make containers possible.
Backend, DevOps and SRE interviews commonly ask practical questions: "what is the difference between kill and kill -9?", "the server says it has only 1 GB free, should we worry?", "how would you find which process holds a port?", "what is a zombie process?" and "what actually is a container?". The answers all come from the material here. Example command outputs in this lesson are illustrative: your process IDs, sizes and versions will differ.
The process tree and PID 1
Every Linux process except the very first is created by another process, usually with fork() followed by exec() (see Processes). So processes form a tree: each has a parent process ID (PPID).
$ pstree -p | head
systemd(1)-+-sshd(812)---sshd(2201)---bash(2210)---vim(2315)
|-nginx(905)-+-nginx(906)
| `-nginx(907)
|-postgres(950)-+-postgres(961)
| `-postgres(962)
|-systemd-journal(311)
`-cron(640)
When the kernel finishes booting, it starts exactly one user-space program as PID 1, the init process. On most current distributions that is systemd; older systems used SysV init, and some minimal systems use alternatives such as OpenRC or BusyBox init. Everything else descends from PID 1.
PID 1 has special duties and properties:
- Starts the system: mounts file systems, starts services (daemons) in the right order, brings up networking and login prompts.
- Adopts orphans: if a parent exits before its children, the children are re-parented to PID 1 (or to a designated "subreaper" process).
- Reaps zombies: it calls
wait()on adopted children when they exit, so they do not stay as zombies. - Is protected: the kernel does not deliver signals to PID 1 for which it has not installed a handler, even
SIGKILLfrom inside its own namespace. If PID 1 exits, the kernel panics (on the host) or the whole container stops (inside a container).
Zombies and orphans
- A zombie is a process that has exited but whose parent has not yet called
wait()to collect its exit status. It uses no memory or CPU, only a process table entry, and shows stateZinps. Many zombies mean a buggy parent. You cannot kill a zombie (it is already dead); you fix or kill the parent, after which PID 1 adopts and reaps it. - An orphan is a still-running process whose parent has exited. It is adopted by PID 1 and continues normally.
Common mistake
Running an application directly as PID 1 inside a container (for example CMD ["python", "app.py"]) means it inherits PID 1's duties. It may ignore SIGTERM (no default handler for PID 1) and never reap zombie children. Tiny init programs such as tini, or docker run --init, exist for exactly this reason.
systemd in brief
systemd manages units: services (.service), mount points, timers (a cron replacement), sockets and more. Useful commands:
$ systemctl status nginx
* nginx.service - A high performance web server
Loaded: loaded (/lib/systemd/system/nginx.service; enabled)
Active: active (running) since Fri 2026-10-09 10:02:11 IST
Main PID: 905 (nginx)
Tasks: 3 (limit: 18977)
Memory: 7.9M
CPU: 1.204s
CGroup: /system.slice/nginx.service
|-905 "nginx: master process"
|-906 "nginx: worker process"
`-907 "nginx: worker process"
$ journalctl -u nginx --since "1 hour ago" # logs for one unit
$ systemctl restart nginx
Notice the CGroup: line: systemd places every service in its own cgroup, which is how it tracks all of a service's processes and applies resource limits. More on cgroups below.
/proc and /sys: the kernel as files
Linux exposes kernel state through pseudo file systems: they are mounted like ordinary file systems, but their "files" are generated by the kernel on the fly when you read them. Nothing is stored on disk.
/proc
/proc (procfs) has a directory per process, named by PID, plus system-wide files.
| Path | What it shows |
|---|---|
/proc/<pid>/status | Name, state, PPID, UIDs, memory summary (VmRSS, VmSize), thread count |
/proc/<pid>/cmdline | Command line, arguments separated by NUL bytes |
/proc/<pid>/environ | Environment variables (readable only by the owner or root) |
/proc/<pid>/fd/ | One symlink per open file descriptor |
/proc/<pid>/maps | Memory regions: address range, permissions, backing file |
/proc/<pid>/limits | Resource limits (ulimit values) |
/proc/self | Symlink to the directory of whichever process reads it |
/proc/cpuinfo | CPU model and features, one block per logical CPU |
/proc/meminfo | Detailed memory counters (what free reads) |
/proc/loadavg | Load averages and running/total task counts |
/proc/sys/ | Tunable kernel parameters (what sysctl reads and writes) |
$ grep -E 'State|PPid|Threads|VmRSS' /proc/905/status
State: S (sleeping)
PPid: 1
Threads: 1
VmRSS: 2856 kB
$ head -4 /proc/self/maps
55d0c3a00000-55d0c3a02000 r--p 00000000 103:02 1311 /usr/bin/cat
55d0c3a02000-55d0c3a07000 r-xp 00002000 103:02 1311 /usr/bin/cat
7f2b4c000000-7f2b4c021000 rw-p 00000000 00:00 0 [heap]
7ffd5e8f1000-7ffd5e912000 rw-p 00000000 00:00 0 [stack]
In maps, r-xp means readable, executable, private (copy-on-write); the code segment is executable but not writable, and the stack and heap are writable but not executable, which is the NX protection discussed in the security lesson.
Writing to files under /proc/sys changes kernel settings at run time:
$ cat /proc/sys/vm/swappiness
60
$ sudo sysctl vm.swappiness=10 # same as writing to the file
Settings that should survive reboot go in /etc/sysctl.conf or /etc/sysctl.d/.
/sys
/sys (sysfs) is newer and more structured. It exposes the kernel's device model: devices, drivers, buses, and their attributes, one value per file. Examples: /sys/block/sda/queue/scheduler (the I/O scheduler), /sys/class/net/eth0/statistics/rx_bytes, /sys/kernel/mm/transparent_hugepage/enabled, and /sys/fs/cgroup/ where cgroups are mounted.
A rough rule: /proc is about processes (plus historical system info), /sys is about devices and kernel objects.
Signals
A signal is a small asynchronous notification sent to a process: "something happened". It carries only a number. Signals are sent by the kernel (for example on an invalid memory access), by other processes (kill), or by the terminal (Ctrl+C).
| Signal | Number (x86 Linux) | Default action | Typical cause |
|---|---|---|---|
SIGHUP | 1 | Terminate | Terminal closed; by convention daemons reload config |
SIGINT | 2 | Terminate | Ctrl+C |
SIGQUIT | 3 | Terminate and core dump | Ctrl+Backslash |
SIGKILL | 9 | Terminate (cannot be caught) | kill -9, OOM killer |
SIGSEGV | 11 | Terminate and core dump | Invalid memory access |
SIGPIPE | 13 | Terminate | Writing to a pipe or socket with no reader |
SIGTERM | 15 | Terminate | kill default; polite shutdown request |
SIGCHLD | 17 | Ignore | A child stopped or exited |
SIGSTOP | 19 | Stop (cannot be caught) | kill -STOP |
SIGTSTP | 20 | Stop | Ctrl+Z |
SIGCONT | 18 | Continue if stopped | fg, bg, kill -CONT |
For each signal a process can choose one of three dispositions: the default action, ignore it, or catch it with a handler function. Two signals cannot be caught, blocked or ignored: SIGKILL and SIGSTOP. That guarantee lets an administrator always stop a runaway process.
SIGTERM versus SIGKILL
SIGTERM(15) is a request: "please shut down". The process can catch it, finish in-flight requests, flush buffers, delete temporary files, and exit cleanly. It is whatkill <pid>,systemctl stopand Kubernetes (on pod deletion) send first.SIGKILL(9) is an order to the kernel: the process is terminated immediately, never runs another instruction, and gets no chance to clean up. Open files are closed by the kernel, but unflushed user-space buffers are lost and temporary files remain.
The correct procedure is: send SIGTERM, wait a grace period, and only then send SIGKILL. systemd and Kubernetes both do exactly this (Kubernetes' default grace period is 30 seconds).
A process killed by signal N exits with status 128 + N as reported by the shell, so 137 = killed by SIGKILL (9), and 143 = killed by SIGTERM (15).
Writing a signal handler
#include <signal.h>
#include <stdio.h>
#include <string.h>
#include <unistd.h>
static volatile sig_atomic_t stop_requested = 0;
static void on_term(int signo) {
(void)signo;
stop_requested = 1; /* only set a flag: handlers must stay tiny */
}
int main(void) {
struct sigaction sa;
memset(&sa, 0, sizeof sa);
sa.sa_handler = on_term;
sigemptyset(&sa.sa_mask);
sigaction(SIGTERM, &sa, NULL);
sigaction(SIGINT, &sa, NULL); /* Ctrl+C */
printf("pid %d working; send SIGTERM to stop\n", getpid());
fflush(stdout);
while (!stop_requested) {
sleep(1); /* real work would go here */
}
printf("cleaning up: flushing buffers, closing files\n");
return 0;
}
$ ./worker &
pid 57092 working; send SIGTERM to stop
$ kill 57092
cleaning up: flushing buffers, closing files
Why only set a flag? A handler can run at any moment, even in the middle of malloc or printf in the main program. If the handler calls the same non-reentrant function, internal state can be corrupted or deadlock. Only async-signal-safe functions (such as write, _exit) may be called from a handler; printf and malloc are not on that list. The usual pattern is to set a volatile sig_atomic_t flag and act on it in the main loop. Use sigaction rather than the older signal(), whose behaviour varies across systems.
File descriptors and "everything is a file"
A file descriptor (fd) is a small integer a process uses to refer to an open file-like object. Every process starts with three: 0 = stdin, 1 = stdout, 2 = stderr. New descriptors take the lowest unused number.
"Everything is a file" is the Unix design idea that many kinds of objects are accessed through the same descriptor interface (read, write, close, poll):
- regular files and directories
- pipes and FIFOs (
|in the shell) - sockets (network connections)
- devices:
/dev/null,/dev/zero,/dev/urandom, disks, terminals (/dev/pts/0) - kernel objects:
eventfd,timerfd,signalfd,epollinstances
You can see a process's descriptors directly:
$ ls -l /proc/905/fd
lr-x------ 1 root root 64 Oct 10 10:02 0 -> /dev/null
l-wx------ 1 root root 64 Oct 10 10:02 1 -> /dev/null
l-wx------ 1 root root 64 Oct 10 10:02 2 -> /var/log/nginx/error.log
lrwx------ 1 root root 64 Oct 10 10:02 6 -> socket:[23145]
l-wx------ 1 root root 64 Oct 10 10:02 4 -> /var/log/nginx/access.log
Shell redirection is just descriptor manipulation: cmd > out.txt 2>&1 opens out.txt as fd 1, then duplicates fd 1 onto fd 2 (with dup2), so both streams go to the file. The order matters: cmd 2>&1 > out.txt sends stderr to the old stdout (the terminal), because the duplication happens first.
Each process has a limit on open descriptors (ulimit -n, often 1024 by default for interactive shells). A busy server that runs out gets EMFILE: Too many open files on accept() or open(); the fix is to raise the limit (in the systemd unit with LimitNOFILE=) and to check for descriptor leaks with ls /proc/<pid>/fd | wc -l over time.
Virtual memory in practice
Virtual memory explained demand paging. Here is how its numbers appear in tools.
VSZ versus RSS
$ ps -o pid,user,vsz,rss,stat,comm -p 950,2315
PID USER VSZ RSS STAT COMMAND
950 postgres 219040 29512 Ss postgres
2315 asha 23676 11204 S+ vim
- VSZ (virtual size, in KiB): the total virtual address space the process has mapped, including memory it has reserved but never touched, shared libraries and memory-mapped files. Often large and mostly harmless.
- RSS (resident set size, in KiB): how much of that is currently in physical RAM. Closer to real usage, but it counts shared pages in full for every process that maps them. Summing RSS over all processes overstates usage, especially for forked workers sharing code and copy-on-write memory.
- PSS (proportional set size, from
/proc/<pid>/smaps_rollup) splits each shared page among the processes that share it, so PSS values add up correctly.
The STAT column shows state: R running or runnable, S interruptible sleep (waiting for an event), D uninterruptible sleep (usually waiting on disk I/O; cannot be killed until the I/O returns), T stopped, Z zombie. Modifiers: s session leader, + foreground process group, l multi-threaded, < high priority, N low priority.
top
top - 10:41:02 up 3 days, 2:11, 2 users, load average: 1.82, 1.40, 1.12
Tasks: 213 total, 2 running, 211 sleeping, 0 stopped, 0 zombie
%Cpu(s): 21.3 us, 4.1 sy, 0.0 ni, 72.9 id, 1.4 wa, 0.0 hi, 0.3 si, 0.0 st
MiB Mem : 15872.0 total, 1120.0 free, 4310.0 used, 10442.0 buff/cache
MiB Swap: 2048.0 total, 2048.0 free, 0.0 used. 11205.0 avail Mem
PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND
3141 app 20 0 4821340 1.2g 22340 S 85.0 7.7 412:10.33 java
950 postgres 20 0 219040 29512 26880 S 6.3 0.2 12:01.88 postgres
How to read it:
- Load average (1.82, 1.40, 1.12): the average number of tasks that are running, waiting for a CPU, or in uninterruptible sleep (state
D), over the last 1, 5 and 15 minutes. Compare it with the number of CPUs: on an 8-CPU machine 1.82 is light; on a 1-CPU machine it means tasks are queueing. Because Linux countsDstate, a high load with idle CPUs often points to disk or NFS waits. - CPU line:
ususer code,sykernel code,niniced user code,ididle,waidle while waiting for I/O,hi/sihardware and software interrupts,ststeal (time a virtual machine wanted the CPU but the hypervisor ran someone else; high steal on a cloud VM means noisy neighbours or an oversubscribed host). - Memory lines: covered with
freebelow. - Per-process:
PRandNIpriority and nice value,VIRT(like VSZ),RES(like RSS),SHRshared part of RES,Sstate,%CPU(can exceed 100 percent for multi-threaded processes: 85 percent here is less than one core),TIME+total CPU time used.
htop shows the same data with per-CPU bars, a tree view and easier sorting and killing.
free
$ free -m
total used free shared buff/cache available
Mem: 15872 4310 1120 320 10442 11205
Swap: 2047 0 2047
- total: usable RAM.
- used: memory used by processes and the kernel, not counting reclaimable cache.
- free: completely unused RAM. On a healthy server this is small, and that is fine.
- buff/cache: the page cache and kernel buffers and reclaimable slabs (4310 + 1120 + 10442 = 15872, the total).
- available: the kernel's estimate of how much memory new work can get without swapping, counting free memory plus the part of the cache that can be dropped. This is the number to watch.
Interview tip
"Free memory is only 1 GB, is the server running out?" The expected answer: no, look at available (here about 11 GB). Linux deliberately uses idle RAM as page cache, and it gives that memory back the moment applications need it. Unused RAM is wasted RAM. Worry when available is low, swap-in and swap-out are active, or the OOM killer appears in dmesg.
vmstat
$ vmstat 1 3
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
r b swpd free buff cache si so bi bo in cs us sy id wa st
2 0 0 1146880 210432 10482304 0 0 12 85 1210 2304 21 4 73 2 0
3 0 0 1144320 210432 10482560 0 0 0 420 1342 2611 24 5 70 1 0
1 1 0 1139712 210436 10483072 0 0 2048 36 1288 2450 19 4 66 11 0
vmstat 1 3 prints three samples one second apart. The first line is an average since boot; read the later lines.
| Column | Meaning | Worry when |
|---|---|---|
r | Tasks runnable (running or waiting for CPU) | Consistently above the CPU count |
b | Tasks blocked in uninterruptible sleep | Consistently above 0 (I/O stalls) |
swpd, free, buff, cache | Memory in KiB | free alone is not a concern |
si, so | Swap in, swap out (KiB/s) | Non-zero and sustained means memory pressure; high both ways means thrashing |
bi, bo | Blocks read and written per second | Compare with device capability |
in, cs | Interrupts and context switches per second | Sudden large jumps |
us sy id wa st | CPU split, as in top | High wa means waiting on I/O; high st means steal |
In the third sample, b = 1, bi = 2048 and wa = 11 together show a process waiting for a burst of disk reads.
The page cache
The page cache is the kernel's cache of file contents in RAM, managed in pages. Every normal read() and write() on a file goes through it:
- Read: the kernel checks the page cache. On a hit, it copies the data from RAM. On a miss, it reads the block from disk into a cache page, and usually reads ahead the next blocks too, because it expects sequential access.
- Write: the kernel copies data into cache pages and marks them dirty, then returns. Kernel writeback threads flush dirty pages to disk later, after a few seconds or when dirty memory exceeds thresholds (
vm.dirty_background_ratio,vm.dirty_ratio).fsync()forces it immediately. - Reclaim: when processes need memory, the kernel evicts clean cache pages (cheap) and writes back then evicts dirty ones.
Consequences:
- The second time you
grepa large log file, it is much faster: the file is now cached. mmaped files share the same page cache pages, so there is one copy whether a file is read or mapped.- Databases such as PostgreSQL rely on the OS page cache in addition to their own buffer pool, while others (MySQL InnoDB is commonly configured this way) bypass it with direct I/O (
O_DIRECT) to avoid caching data twice. - You can drop clean caches for benchmarking with
sync; echo 3 | sudo tee /proc/sys/vm/drop_caches. Never do it on a production system as a "fix": it only makes the next reads slower.
Scheduler classes
Linux schedules threads (tasks), and each task belongs to a scheduling class. Classes are checked in strict priority order: a runnable task in a higher class always runs before any task in a lower class.
| Class | Policies | How it chooses | Used for |
|---|---|---|---|
| Deadline | SCHED_DEADLINE | Earliest deadline first, with a runtime budget per period | Hard periodic work (rarely used directly) |
| Real-time | SCHED_FIFO, SCHED_RR | Fixed priority 1 to 99; FIFO runs until it blocks or yields, RR adds a time slice among equal priorities | Audio, industrial control, latency-critical threads |
| Fair (normal) | SCHED_OTHER (also called SCHED_NORMAL), SCHED_BATCH | Proportional share weighted by nice value | Almost all processes |
| Idle | SCHED_IDLE | Runs only when nothing else wants the CPU | Background tasks that should never interfere |
The fair class was implemented by the Completely Fair Scheduler (CFS) from Linux 2.6.23 until it was replaced by EEVDF (earliest eligible virtual deadline first) in Linux 6.6. Both track each task's virtual runtime (CPU time used, scaled by its weight) and favour tasks that have received less than their fair share; CFS kept tasks in a red-black tree ordered by virtual runtime. See CPU scheduling for the general algorithms.
Nice values range from -20 (highest priority, least "nice" to others) to +19 (lowest). Each step changes a task's weight by about 1.25 times, so two CPU-bound tasks one nice level apart split a CPU roughly 55 to 45. Only root can lower a nice value (raise priority).
$ nice -n 10 ./backup.sh # start with nice 10
$ renice -n 5 -p 3141 # change a running process
$ chrt -f -p 50 4242 # make 4242 SCHED_FIFO priority 50 (root)
$ chrt -p 4242
pid 4242's current scheduling policy: SCHED_FIFO
pid 4242's current scheduling priority: 50
Common mistake
A SCHED_FIFO thread in an infinite loop can starve everything else on its CPU, including normal shells you would use to kill it. Linux reserves 5 percent of each second for non-real-time tasks by default (kernel.sched_rt_runtime_us = 950000 of sched_rt_period_us = 1000000) as a safety net. Use real-time classes only for short, well-understood work.
Namespaces and cgroups: how containers work
A container is not a virtual machine and not a special kernel object. It is an ordinary group of Linux processes that the kernel shows a restricted view of the system (namespaces) and gives limited resources (cgroups), usually running from its own root file system image. Docker, containerd, Podman and Kubernetes all build on these two kernel features.
Namespaces: what a process can see
A namespace wraps a global system resource so that processes inside it see their own isolated instance.
| Namespace | Isolates | Effect inside a container |
|---|---|---|
| PID | Process IDs | The container's first process is PID 1; it cannot see host processes |
| Mount (mnt) | Mount table | Its own root file system and mounts |
| Network (net) | Interfaces, IPs, routes, ports, firewall rules | Its own eth0 and localhost; can bind port 80 independently |
| UTS | Hostname and domain name | Its own hostname |
| IPC | System V IPC, POSIX message queues | Cannot use other containers' shared memory segments |
| User | User and group IDs | Root (UID 0) inside can map to an unprivileged UID outside ("rootless" containers) |
| Cgroup | View of the cgroup hierarchy | Sees its own cgroup as the root |
| Time | Some system clocks (Linux 5.6+) | Its own boot-time and monotonic clock offsets |
Namespaces are created with the clone() or unshare() system calls and joined with setns(). You can experiment from a shell:
$ sudo unshare --pid --fork --mount-proc bash
# ps aux
USER PID %CPU %MEM VSZ RSS TTY STAT START TIME COMMAND
root 1 0.0 0.0 8608 5248 pts/1 S 10:55 0:00 bash
root 8 0.0 0.0 10072 3328 pts/1 R+ 10:55 0:00 ps aux
Inside the new PID namespace, bash is PID 1 and sees only itself and its children. From the host, the same bash has an ordinary PID such as 48211. Every namespace a process belongs to is listed in /proc/<pid>/ns/.
cgroups: how much a process can use
Control groups (cgroups) organise processes into a hierarchy and apply resource limits and accounting to each group. Modern distributions use cgroup v2, a single unified hierarchy mounted at /sys/fs/cgroup.
Key controllers and their files:
| Controller | Example file | Meaning |
|---|---|---|
| cpu | cpu.max = 50000 100000 | At most 50 ms of CPU time per 100 ms period, i.e. half a CPU |
| cpu | cpu.weight | Relative share when CPUs are contended |
| memory | memory.max | Hard limit; exceeding it triggers reclaim, then the OOM killer within the group |
| memory | memory.high | Soft limit; the group is throttled and reclaimed above it |
| io | io.max | Bandwidth or IOPS limit per block device |
| pids | pids.max | Maximum number of processes or threads (stops fork bombs) |
$ cat /sys/fs/cgroup/system.slice/nginx.service/memory.current
8273920
$ cat /sys/fs/cgroup/system.slice/nginx.service/pids.current
3
When Kubernetes sets resources.limits.memory: 512Mi on a container, the runtime writes 536870912 into that container's memory.max. When it sets limits.cpu: 500m, it writes 50000 100000 into cpu.max: the container is throttled after 50 ms of CPU in each 100 ms window, which can add latency to a busy multi-threaded service even when the host has idle CPUs.
Putting a container together
container runtime (runc) does roughly:
1. clone() with new PID, mount, net, UTS, IPC (and user) namespaces
2. create a cgroup, write limits (memory.max, cpu.max, pids.max),
move the new process into it
3. mount the image layers (overlayfs) and pivot_root into them
4. set up a veth pair to connect the net namespace to a bridge
5. drop capabilities, apply seccomp and LSM profiles
6. exec() the container's entrypoint, which becomes PID 1 inside
The crucial consequence: all containers on a host share one kernel. That makes them light (start in milliseconds, no guest OS) but means isolation is only as strong as the kernel's: a kernel vulnerability can let a process escape its container. The virtualization and security lesson compares containers with virtual machines in detail.
Useful commands for interviews
ps
$ ps -ef | head -4 # every process, full format
UID PID PPID C STIME TTY TIME CMD
root 1 0 0 Oct07 ? 00:00:09 /sbin/init
root 2 0 0 Oct07 ? 00:00:00 [kthreadd]
root 311 1 0 Oct07 ? 00:00:02 /lib/systemd/systemd-journald
$ ps aux --sort=-%mem | head -3 # top memory users
USER PID %CPU %MEM VSZ RSS TTY STAT START TIME COMMAND
app 3141 85.0 7.7 4821340 1258291 ? Sl Oct07 412:10 java -Xmx1g -jar app.jar
pg 950 0.3 0.2 219040 29512 ? Ss Oct07 12:01 postgres
$ ps -eo pid,ppid,stat,comm | awk '$3 ~ /Z/' # find zombies
Names in square brackets, like [kthreadd], are kernel threads.
top and htop
Interactive keys in top: P sort by CPU, M sort by memory, 1 show each CPU, k kill, H show threads. top -H -p <pid> shows the threads of one process, useful for finding the hot thread in a JVM.
strace
strace traces the system calls a process makes, with arguments and results. It is the fastest way to see what a program is actually doing.
$ strace -e trace=openat,read -f cat /etc/hostname
openat(AT_FDCWD, "/etc/ld.so.cache", O_RDONLY|O_CLOEXEC) = 3
openat(AT_FDCWD, "/lib/x86_64-linux-gnu/libc.so.6", O_RDONLY|O_CLOEXEC) = 3
read(3, "\177ELF\2\1\1\3\0\0\0\0\0\0\0\0\3\0>\0\1\0\0\0"..., 832) = 832
openat(AT_FDCWD, "/etc/hostname", O_RDONLY) = 3
read(3, "web-01\n", 131072) = 7
read(3, "", 131072) = 0
web-01
$ strace -c -p 3141 # summary of syscalls of a running process
$ strace -f -e trace=network ./client # only network calls, follow children
Typical uses: find which config file a program reads (openat returning ENOENT), why it hangs (stuck in futex, read on a socket, connect), or permission problems (EACCES). strace slows the traced process down significantly, so be careful on production. ltrace traces library calls instead; perf and eBPF tools (such as bpftrace) give low-overhead tracing.
lsof
lsof lists open files, and since everything is a file, that includes sockets.
$ sudo lsof -i :8080 # who is using port 8080?
COMMAND PID USER FD TYPE DEVICE SIZE/OFF NODE NAME
java 3141 app 45u IPv6 58211 0t0 TCP *:8080 (LISTEN)
java 3141 app 61u IPv6 60032 0t0 TCP 10.0.0.5:8080->10.0.0.9:51234 (ESTABLISHED)
$ sudo lsof +L1 # deleted files still held open
COMMAND PID USER FD TYPE DEVICE SIZE/OFF NLINK NODE NAME
nginx 906 www 4w REG 259,2 8589934592 0 13371 /var/log/nginx/access.log (deleted)
The second example explains the classic "df says the disk is full but du cannot find the files" puzzle: a process still holds an 8 GiB deleted log open. Restart it or truncate via /proc/906/fd/4 to free the space. ss -ltnp is a faster alternative for listing listening sockets.
kill, pkill, killall
$ kill 3141 # SIGTERM (polite)
$ kill -9 3141 # SIGKILL (last resort)
$ kill -HUP 905 # ask nginx to reload its config
$ kill -0 3141 && echo alive # signal 0: check the process exists and you may signal it
$ pkill -f "python worker.py" # match on full command line
nice and renice
Shown in the scheduler section: nice -n 19 cmd starts a low-priority job; renice adjusts a running one. ionice -c3 cmd similarly lowers I/O priority (idle class) for schedulers that support it, such as BFQ.
ulimit
ulimit shows and sets per-process resource limits for the shell and its children. Each limit has a soft value (currently enforced, which a user can raise up to the hard value) and a hard value (ceiling; only root can raise it).
$ ulimit -a
core file size (blocks, -c) 0
data seg size (kbytes, -d) unlimited
max locked memory (kbytes, -l) 8192
open files (-n) 1024
max user processes (-u) 63345
stack size (kbytes, -s) 8192
cpu time (seconds, -t) unlimited
virtual memory (kbytes, -v) unlimited
$ ulimit -n 65536 # raise open-file limit (up to the hard limit)
$ ulimit -c unlimited # allow core dumps for debugging
stack size 8192 KiB is the default 8 MiB main-thread stack: deep recursion beyond it ends in SIGSEGV. For services, set limits in the systemd unit (LimitNOFILE=, LimitCORE=), since ulimit in a login shell does not affect daemons.
Other commands worth knowing: dmesg (kernel log, including OOM kills), df -h and du -sh (disk usage), iostat -x 1 (per-device I/O utilisation and latency), uptime (load averages), ss (sockets), journalctl (systemd logs).
Interview questions
Q1. What is PID 1 and what is special about it?
PID 1 is the first user-space process the kernel starts, usually systemd. It starts services, adopts orphaned processes and reaps zombies. The kernel will not deliver signals to it that it has no handler for, and if it exits the system panics or, in a container, the container stops.
Q2. What is a zombie process and how do you get rid of it?
A zombie has exited but its parent has not called wait() to read its exit status, so its process table entry remains. It cannot be killed because it is already dead. Fix or kill the parent; the zombie is then adopted by PID 1 and reaped.
Q3. What is the difference between SIGTERM and SIGKILL?
SIGTERM asks a process to terminate; it can be caught so the process can clean up, and it is the default for kill. SIGKILL cannot be caught, blocked or ignored; the kernel terminates the process immediately with no cleanup. Always try SIGTERM first, then SIGKILL after a grace period.
Q4. Why should a signal handler only set a flag?
Handlers run asynchronously, possibly in the middle of a non-reentrant function like malloc or printf. Calling such functions from the handler can corrupt state or deadlock. Only async-signal-safe functions are allowed, so the safe pattern is to set a volatile sig_atomic_t flag and handle it in the main loop.
Q5. What are /proc and /sys?
They are pseudo file systems generated by the kernel on demand. /proc exposes per-process information (status, memory maps, open descriptors, limits) and tunables under /proc/sys. /sys exposes the device model and kernel objects such as block devices, network interfaces and cgroups.
Q6. What does "everything is a file" mean?
Files, pipes, sockets, terminals, devices and many kernel objects are all accessed through file descriptors with the same read, write, close and poll calls. This lets generic tools and mechanisms (redirection, epoll, lsof) work across very different resources.
Q7. What is the difference between VSZ and RSS?
VSZ is the total virtual address space mapped by the process, including reserved but untouched memory and shared libraries. RSS is the part currently in physical RAM, counting shared pages fully for each process. PSS divides shared pages among sharers and is the better measure when summing across processes.
Q8. free shows 1 GB free on a 16 GB server. Is that a problem?
Usually not. Linux uses spare RAM as page cache, shown under buff/cache, and reclaims it instantly when needed. The available column estimates usable memory; also check swap activity (si/so in vmstat) and OOM messages in dmesg.
Q9. What does the load average measure?
The average number of tasks that are runnable or in uninterruptible sleep over 1, 5 and 15 minutes. Compare it with the number of CPUs. Because Linux includes tasks blocked on disk I/O, a high load with idle CPUs often indicates I/O waits.
Q10. What is the page cache?
The kernel's in-RAM cache of file data. Reads hit it when possible; writes land in it as dirty pages and are flushed by writeback threads or fsync. Clean pages are evicted first under memory pressure.
Q11. What scheduling classes does Linux have?
From highest priority: deadline (SCHED_DEADLINE), real-time (SCHED_FIFO and SCHED_RR, priorities 1 to 99), fair (SCHED_OTHER/SCHED_BATCH, weighted by nice value; CFS until Linux 6.6, EEVDF since), and idle (SCHED_IDLE). Higher classes always pre-empt lower ones.
Q12. What are namespaces and cgroups?
Namespaces restrict what a process can see: its own PIDs, mounts, network stack, hostname, IPC and user IDs. Cgroups restrict how much it can use: CPU, memory, I/O and process counts, with accounting. Together with a root file system image and security profiles, they make a container.
Q13. How is a container different from a virtual machine?
A container is a set of host processes isolated with namespaces and limited by cgroups, sharing the host kernel. A VM runs a full guest kernel on virtualised hardware provided by a hypervisor. Containers are lighter and start faster; VMs give stronger isolation because a guest kernel bug does not directly expose the host.
Q14. How would you find which process is listening on port 8080?
Use sudo ss -ltnp | grep 8080 or sudo lsof -i :8080; both show the PID and program. Then inspect it with ps -fp <pid> or /proc/<pid>/cmdline.
Q15. A disk is full according to df but du shows much less. Why?
Most likely a process still has deleted files open, so their space is not freed until the descriptor closes. lsof +L1 lists them. Restarting or signalling the process (for example to reopen logs) frees the space.
Key takeaways
- Processes form a tree rooted at PID 1 (usually systemd), which starts services, adopts orphans and reaps zombies.
/procexposes per-process state and kernel tunables;/sysexposes devices and kernel objects such as cgroups.SIGTERMrequests shutdown and can be handled;SIGKILLandSIGSTOPcannot be caught; exit code 128 + N means killed by signal N.- File descriptors unify files, pipes, sockets and devices; redirection is descriptor duplication.
- VSZ is mapped address space, RSS is resident memory, PSS splits shared pages; in
free, watchavailable, notfree. - The page cache uses spare RAM for file data; writes are buffered as dirty pages until writeback or
fsync. - Linux scheduling classes run in priority order: deadline, real-time, fair (EEVDF since 6.6), idle; nice ranges from -20 to 19.
- Containers are processes plus namespaces (what they see) plus cgroups (what they can use), all sharing the host kernel.
Next lesson
Continue with Virtualization and security.

