What a file system does
A disk or SSD is, to the hardware, a long array of numbered blocks. Nobody wants to remember "my essay is in blocks 88,301 to 88,342". A file system is the part of the operating system that turns raw blocks into named files organised in directories, keeps track of which blocks are free, controls who may access what, and keeps all of this consistent even if the power fails halfway through a write.
This lesson goes from the user's view (files, attributes, operations, directories, links) down to the implementation (allocation methods, inodes, free-space management, journaling), then looks at real file systems (ext4, NTFS, APFS) and how Linux supports many of them at once through the VFS layer.
Interviewers commonly ask: what is an inode, what is the difference between a hard link and a soft link, how does a file system survive a crash, and "what is the maximum file size with 12 direct pointers and one single, double and triple indirect pointer?" You will compute that last one step by step.
Files and their attributes
A file is a named collection of related data stored on secondary storage, the smallest unit of storage a user can name. To the OS, most files are just a sequence of bytes; any structure (a JPEG header, a CSV row) is interpreted by programs, not the kernel.
Every file has attributes (metadata), kept separately from its data:
| Attribute | Meaning |
|---|---|
| Name | Human-readable name. On Unix, stored in the directory, not with the file. |
| Identifier | Unique number inside the file system, such as the inode number. |
| Type | Regular file, directory, symbolic link, device, pipe, socket. |
| Location | Pointers to the blocks holding the data. |
| Size | Current length in bytes. |
| Protection | Owner, group, permission bits (read, write, execute). |
| Timestamps | Last access (atime), last data modification (mtime), last metadata change (ctime); some systems also store creation time. |
| Link count | Number of directory entries (names) pointing to the file. |
You can see most of these with stat on Linux:
$ stat notes.txt
File: notes.txt
Size: 6 Blocks: 8 IO Block: 4096 regular file
Device: 259,2 Inode: 1835021 Links: 1
Access: (0644/-rw-r--r--) Uid: ( 1000/ asha) Gid: ( 1000/ asha)
Access: 2026-10-10 09:12:40.115 +0530
Modify: 2026-10-10 09:12:38.902 +0530
Change: 2026-10-10 09:12:38.902 +0530
(Example output; your inode numbers, users and times will differ.)
File operations and the open-file table
The OS provides system calls for the basic operations: create, open, read, write, seek (reposition), close, delete (unlink on Unix), truncate, and get/set attributes.
Why is there an open at all? Looking a name up means walking directories on disk, which is slow. So you look it up once, at open(), and the kernel hands back a small integer, the file descriptor (fd) on Unix or a handle on Windows. Later calls use the fd and skip the lookup.
The kernel tracks open files in three layers:
Process A fd table System-wide In-memory
(per process) open-file table inode table
+----+-------+ +----------------+ +--------------+
| 0 | ------+--------> | offset 0 | | |
| 1 | | | mode r | ---> | inode 1835021|
| 3 | ------+---+ | refcount 1 | | size, blocks |
+----+ | | +----------------+ | open count 2 |
| +----> | offset 4096 | ---> | |
Process B | | mode rw | +--------------+
| 3 | ------+--------> | refcount 2 |
+----+ +----------------+
- The per-process file descriptor table maps small integers (0 = stdin, 1 = stdout, 2 = stderr, then 3, 4, ...) to entries in the system-wide table.
- The system-wide open-file table has one entry per
open()call. It stores the current file offset (where the next read or write happens), the access mode, and a reference count. - The in-memory inode (vnode) table holds one copy of each open file's metadata, however many times it is opened.
Consequences that interviewers like:
- Two separate
open()calls on the same file give independent offsets. - After
fork(), parent and child share the same open-file table entry, so they share the offset. If both write, their output interleaves rather than overwriting. dup2(fd, 1)makes fd 1 point at the same entry asfd. That is how a shell implements>redirection.- A file deleted while open stays on disk until the last descriptor is closed, because the inode's open count is still non-zero. This is why
dfsometimes shows space thatducannot find: a process still has a deleted log file open.
File locking coordinates access between processes. Locks can be shared (many readers) or exclusive (one writer), and advisory (processes must check voluntarily, the Unix default with flock and fcntl) or mandatory (the OS enforces them, the Windows default for files opened without sharing).
Access methods
An access method is the way a program reads and writes a file's data.
- Sequential access: read or write records in order, with an implicit position that advances. Like a tape. Most programs (compilers, log readers,
cat) work this way. - Direct (random) access: jump to any block or byte number and read it, as in
lseek(fd, 4096 * n, SEEK_SET). Databases need this. - Indexed access: an index file maps keys to block positions; look up the key in the index, then read the block directly. ISAM files and database indexes are built on top of direct access this way.
Unix offers byte-level direct access to any regular file, and you can build sequential and indexed access on top of it.
Directory structures
A directory is a table that maps names to files (on Unix, to inode numbers). Directory organisation evolved through several designs.
Single-level directory
All files of all users live in one directory. Every name must be unique system-wide; with many users, name clashes are constant. Early systems and some tiny embedded devices used it.
Two-level directory
Each user has their own user file directory (UFD), and a master file directory (MFD) lists users. Different users can now use the same filename, but a user cannot group their own files, and sharing files between users is awkward.
Tree-structured directory
Directories can contain subdirectories, to any depth. This is what every modern OS gives you.
- An absolute path starts at the root:
/home/asha/notes.txt. - A relative path starts at the process's current working directory:
notes.txtor../ravi/todo.md. - Every directory contains
.(itself) and..(its parent).
/
+------+-------+
bin home etc
+--+---+
asha ravi
+--+--+ |
notes src todo.md
|
main.c
A pure tree has exactly one path to each file, so it does not allow sharing a file under two directories.
Acyclic-graph directory
Allow a file or directory to appear in more than one directory, as long as no cycles form. This is implemented with links. Sharing is now possible, but new problems appear:
- A file has several absolute names, so tools that walk the tree (backup,
du) may count it twice. - Deletion: if one user deletes the shared file, what happens to the other names? Unix answers with link counts (hard links) or dangling pointers (soft links).
- Cycles: if links could point to ancestors, a traversal could loop forever. That is why Unix forbids hard links to directories (except the automatic
.and..), while symbolic links can point anywhere and traversal tools detect or skip them.
A general graph directory allows cycles; it needs garbage collection to reclaim unreachable files and is rarely used.
Hard links versus soft links
A hard link is an additional directory entry pointing to the same inode. All names are equal; there is no "original". The inode's link count records how many names exist, and the file's data is freed only when the link count reaches 0 and no process has it open.
A soft link (symbolic link, symlink) is a separate small file whose content is a path. Opening it makes the kernel follow the path. If the target is deleted or moved, the symlink dangles.
Hard link Soft link
dir entries inode 7742386 dir entries inode 7742389
notes.txt --+-> [links=2] soft.txt -----> [type=symlink]
hard.txt ---+ [data: hello] [data:"notes.txt"]
|
path lookup v
notes.txt --> inode 7742386
This C program demonstrates both, then deletes the original name:
#include <stdio.h>
#include <sys/stat.h>
#include <unistd.h>
static void show(const char *path) {
struct stat st;
if (lstat(path, &st) < 0) { printf("%-10s (missing)\n", path); return; }
printf("%-10s inode=%llu links=%lu size=%lld %s\n", path,
(unsigned long long)st.st_ino, (unsigned long)st.st_nlink,
(long long)st.st_size, S_ISLNK(st.st_mode) ? "symlink" : "regular");
}
int main(void) {
FILE *f = fopen("notes.txt", "w");
if (!f) { perror("fopen"); return 1; }
fputs("hello\n", f);
fclose(f);
link("notes.txt", "hard.txt"); /* new name, same inode */
symlink("notes.txt", "soft.txt"); /* new inode holding a path */
show("notes.txt"); show("hard.txt"); show("soft.txt");
unlink("notes.txt"); /* remove the original name */
printf("after unlink(notes.txt):\n");
show("hard.txt");
struct stat st;
printf("soft.txt target reachable? %s\n", stat("soft.txt", &st) == 0 ? "yes" : "no (dangling)");
return 0;
}
Output in an empty directory (inode numbers vary):
notes.txt inode=7742386 links=2 size=6 regular
hard.txt inode=7742386 links=2 size=6 regular
soft.txt inode=7742389 links=1 size=9 symlink
after unlink(notes.txt):
hard.txt inode=7742386 links=1 size=6 regular
soft.txt target reachable? no (dangling)
The symlink's size is 9 because its content is the 9-character path notes.txt. The same experiment from the shell is ln notes.txt hard.txt, ln -s notes.txt soft.txt and ls -li.
| Property | Hard link | Soft (symbolic) link |
|---|---|---|
| Points to | Inode directly | A path name |
| Own inode? | No, shares the target's | Yes |
| Across file systems? | No (inode numbers are per file system) | Yes |
| To directories? | No (prevents cycles) | Yes |
| Target deleted | Data survives while any link remains | Link dangles |
| Survives target being moved or renamed | Yes | No (path changes) |
| Created with | ln target name | ln -s target name |
Interview tip
A tidy answer: "A hard link is another name for the same inode, so the data lives until the last name and last open descriptor are gone; it cannot cross file systems or point to directories. A symlink is a tiny file containing a path; it can point anywhere but breaks if the target moves." Mention that rm actually calls unlink, which only removes a name.
On-disk layout
A typical disk is divided into partitions, each holding one file system (or swap). A Unix-style file system has roughly this layout:
+-------+-------+---------+---------+-------------+--------------+
| boot | super | free- | free- | inode table | data blocks |
| block | block | block | inode | (fixed size)| ... |
| | | bitmap | bitmap | | |
+-------+-------+---------+---------+-------------+--------------+
- Boot block: code to boot an OS, if this partition is bootable.
- Superblock: global information: block size, total blocks and inodes, free counts, the file-system type's magic number, mount state.
- Bitmaps: which blocks and which inodes are free.
- Inode table: fixed-size records, one per file.
- Data blocks: file contents and directory contents.
ext4 repeats this layout in block groups spread across the disk so that a file's inode and data are near each other, and keeps backup copies of the superblock in several groups.
A directory's data blocks simply contain (name, inode number) entries. Looking up /home/asha/notes.txt means: read the root inode (inode 2 on ext file systems), read its data to find home, read that inode, read its data to find asha, and so on. The kernel caches these lookups in the dentry cache so repeated lookups do not touch the disk.
Allocation methods
How does the file system decide which blocks hold a file's data? There are three classic methods.
Contiguous allocation
Each file occupies a run of consecutive blocks. The directory entry stores the start block and length.
dir entry: report start=14 length=3
blocks: ... [13][14 report][15 report][16 report][17] ...
- Fast sequential and direct access: block i of the file is simply
start + i, and on a hard disk the head barely moves. - External fragmentation, exactly as in contiguous memory allocation.
- Files cannot grow easily: the block after the file may be taken. You must guess the size up front or move the file.
Modern file systems use a relaxed form called extents: a file is a list of contiguous runs, each described as (start block, length). ext4, NTFS, XFS and APFS all use extents. A 1 GB file written in one go may need only a handful of extent records instead of 262,144 block pointers.
CD-ROM and DVD file systems (ISO 9660) use pure contiguous allocation, because files never change after writing.
Linked allocation
Each file is a linked list of blocks scattered anywhere. The directory stores the first (and maybe last) block; each block stores a pointer to the next.
dir entry: report first=9 last=25
[9 | ->16] --> [16 | ->1] --> [1 | ->25] --> [25 | end]
- No external fragmentation, files grow easily.
- Direct access is slow: to read block i you must follow i pointers, each a disk read.
- Pointers use space in every block (with 512-byte blocks and 4-byte pointers, only 508 bytes hold data).
- Fragile: one corrupted pointer loses the rest of the file.
FAT: linked allocation with the links in a table
The File Allocation Table (FAT) moves all the next-pointers out of the data blocks into one table at the start of the volume, with one entry per block (cluster). Entry i holds the number of the block after block i in its file, or an end-of-file marker, or 0 for free.
dir entry: report first=9
FAT index: 1 9 16 25
FAT value: 25 16 1 EOF
Following the chain now means reading FAT entries, and the FAT is small enough to cache in memory, so random access becomes fast. The FAT also doubles as the free-space list.
Size of a FAT: a 1 GiB volume with 4 KiB clusters has 2^30 / 2^12 = 2^18 = 262,144 clusters; with 4-byte FAT32 entries, the table is 262,144 x 4 = 1 MiB. FAT32 is still used on USB drives and SD cards for compatibility, but it limits a single file to just under 4 GiB, and exFAT replaces it for larger files.
Indexed allocation
Put all of a file's block pointers together in one index block. The directory entry points to the index block; entry i of the index block is the address of the file's block i.
dir entry: report index=19
index block 19: [ 9 | 16 | 1 | 25 | -1 | -1 | ... ]
| | | |
data blocks of the file, in order
- Direct access is fast: read the index block once, then jump to any block.
- No external fragmentation.
- Overhead: even a 1-byte file needs a whole index block.
- A file larger than one index block can describe needs a scheme for more pointers: linking several index blocks, a multi-level index, or the Unix combined scheme below.
The Unix inode: direct and indirect blocks
The classic Unix inode (index node) is a fixed-size record holding a file's metadata and a small array of block pointers. In the traditional design (used by ext2 and ext3) there are 15 pointers:
- 12 direct pointers, each pointing straight at a data block.
- 1 single indirect pointer, pointing at a block full of pointers to data blocks.
- 1 double indirect pointer, pointing at a block of pointers to single indirect blocks.
- 1 triple indirect pointer, one more level.
inode
+--------------+
| mode, owner, |
| size, times |
+--------------+
| direct 0 |------------------------------> data
| ... |
| direct 11 |------------------------------> data
| single ind. |--> [ptrs] ------------------> data
| double ind. |--> [ptrs] --> [ptrs] -------> data
| triple ind. |--> [ptrs] --> [ptrs] --> [ptrs] --> data
+--------------+
The design is clever because small files are fast (most files are small and need only the direct pointers, which are in the inode itself) while large files are still possible through the indirect levels.
Worked example: maximum file size
Given: block size 4 KiB (4096 bytes), block pointers 4 bytes, 12 direct, 1 single, 1 double and 1 triple indirect pointer.
Step 1: pointers per block. One block holds 4096 / 4 = 1024 pointers.
Step 2: blocks reachable at each level.
| Level | Data blocks reachable | Bytes |
|---|---|---|
| Direct | 12 | 12 x 4 KiB = 48 KiB (49,152 B) |
| Single indirect | 1024 | 1024 x 4 KiB = 4 MiB |
| Double indirect | 1024^2 = 1,048,576 | 4 GiB |
| Triple indirect | 1024^3 = 1,073,741,824 | 4 TiB |
Step 3: add them. Total blocks = 12 + 1024 + 1,048,576 + 1,073,741,824 = 1,074,791,436 blocks.
Step 4: multiply by block size. 1,074,791,436 x 4096 = 4,402,345,721,856 bytes, about 4 TiB (4 TiB + 4 GiB + 4 MiB + 48 KiB).
In real ext3 the limit with 4 KiB blocks is lower, 2 TiB, because other fields (such as the 32-bit count of 512-byte sectors in the inode) run out first. In an exam, give the pointer-based answer and mention the caveat if asked about real systems.
Second example: 1 KiB blocks, 10 direct pointers
Block size 1 KiB, 4-byte pointers, 10 direct, 1 single, 1 double, 1 triple.
- Pointers per block = 1024 / 4 = 256.
- Blocks = 10 + 256 + 256^2 + 256^3 = 10 + 256 + 65,536 + 16,777,216 = 16,843,018.
- Size = 16,843,018 x 1024 = 17,247,250,432 bytes, about 16.06 GiB.
How many disk reads to reach a byte?
Assume only the inode is in memory and nothing else is cached. To read the byte at offset 10 MiB in the 4 KiB file system above:
- Block index = 10 MiB / 4 KiB = 2560.
- Blocks 0 to 11 are direct, 12 to 1035 are single indirect, so block 2560 is in the double indirect range (1036 onwards). Its position within that range is 2560 - 1036 = 1524.
- First-level index = 1524 div 1024 = 1; second-level index = 1524 mod 1024 = 500.
- Reads: double indirect block, then the second single-indirect block it points to (index 1), then the data block (index 500): 3 disk reads.
ext4's extent tree
ext4 replaces the pointer array with an extent tree. The inode holds up to four extents directly; each extent covers up to 32,768 contiguous blocks (128 MiB with 4 KiB blocks). Larger or more fragmented files grow a small B-tree of extent blocks. ext4's maximum file size is 16 TiB with 4 KiB blocks.
Free-space management
The file system must know which blocks are free. Four standard techniques:
Bit vector (bitmap)
One bit per block: 1 = free, 0 = allocated (some systems use the opposite convention). Finding a free block means finding the first non-zero word and then its first set bit, which CPUs do with a single instruction. Easy to find contiguous free runs too.
Worked example: a 1 TiB disk with 4 KiB blocks has 2^40 / 2^12 = 2^28 blocks. The bitmap needs 2^28 bits = 2^25 bytes = 32 MiB. That is fine on disk but large to keep entirely in memory, which is one reason ext4 splits bitmaps per block group.
Linked list
Link all free blocks together, each pointing to the next, and keep a pointer to the first in the superblock. No extra space needed, but traversing it is slow (a disk read per block) and it is hard to find contiguous runs. Allocation only ever needs the first block, so that cost is fine for single-block allocation.
Grouping
Store the addresses of n free blocks in the first free block. The first n - 1 of those are actually free; the last one is another block holding the next n addresses. You can find many free blocks with one read.
Counting
Free blocks often come in runs, so store (first free block, count) pairs. This is the same idea as extents, applied to free space. XFS keeps free extents in B+ trees indexed both by start block and by length, so it can quickly find a run of a given size.
Journaling and crash consistency
Creating or appending to a file updates several on-disk structures: the data block, the inode (size and pointers), the block bitmap and maybe a directory block. The disk can only write one block at a time, so a power cut between writes leaves the file system inconsistent. For example:
- Bitmap updated but inode not: a block is marked used but belongs to no file (a space leak).
- Inode updated but bitmap not: the block is used by a file but marked free, so it may be given to another file later (corruption).
- Inode points to the new block but the data never arrived: the file contains garbage from whatever was in that block before.
The old way: fsck
Older file systems (ext2, FAT) ran a consistency checker such as fsck or chkdsk after a crash. It scans every inode, bitmap and directory to find and fix mismatches. It works, but it takes time proportional to the size of the disk, which could be hours on a large volume, and it cannot recover data whose writes were lost.
Journaling (write-ahead logging)
A journaling file system writes a description of each update to a dedicated area, the journal (log), before applying it to the main structures. The steps for one transaction:
- Journal write: write the blocks to be changed (or a description of the change) into the journal, preceded by a transaction-begin record.
- Journal commit: write a commit record. Once the commit block is safely on disk, the transaction is durable.
- Checkpoint: write the changes to their real locations in the file system.
- Free: mark the journal space as reusable.
journal: [TxBegin][inode][bitmap][dir][TxCommit]
|
after commit is on disk
v
main FS: write inode, bitmap, dir block in place (checkpoint)
After a crash, recovery is simple and fast:
- A transaction with a commit record in the journal is replayed (redone), writing its blocks to their real locations. Replaying twice is harmless, since it writes the same contents.
- A transaction without a commit record is discarded; the main file system was never touched.
Recovery time now depends on the size of the journal, not the disk: seconds instead of hours.
What gets journaled: ext4's modes
Writing all data twice (once to the journal, once in place) is expensive, so ext4 offers three modes:
| Mode | Journals | Guarantee | Cost |
|---|---|---|---|
journal | Metadata and data | Strongest: data and metadata consistent | Slowest; data written twice |
ordered (default) | Metadata only, but data blocks are written before the metadata that points to them commits | No garbage in files after a crash | Good balance |
writeback | Metadata only, no ordering for data | Metadata consistent, but a file may contain stale data after a crash | Fastest |
Ordered mode is why, after a crash, an ext4 file may be missing its last writes but will not contain another file's old data.
Copy-on-write file systems
An alternative to journaling is to never overwrite live data. ZFS, Btrfs and APFS write new versions of changed blocks to free space, then update the pointers up the tree, finishing with an atomic switch of the root pointer. A crash before the switch leaves the old tree intact; after it, the new tree. This also makes snapshots almost free: keep the old root.
Applications still need fsync
Journaling protects the file system's structures, not your application's data. Data written with write() sits in the page cache and reaches disk later. To make sure it is durable, call fsync(fd). To atomically replace a file safely, the standard pattern is: write a temporary file, fsync it, rename it over the original (rename is atomic within a file system), then fsync the directory.
Common mistake
"The file system is journaled, so my data is safe once write() returns" is wrong. write() usually returns after copying into the page cache. Without fsync() a power cut can lose recently written data on any file system.
Real file systems: ext4, NTFS, APFS
| Feature | ext4 (Linux) | NTFS (Windows) | APFS (Apple) |
|---|---|---|---|
| Metadata structure | Inodes in block groups | Master File Table (MFT), one record per file | B-trees of objects |
| Data allocation | Extents (extent tree) | Runs (extents) | Extents |
| Crash consistency | Journal (default ordered mode) | Journal ($LogFile) of metadata | Copy-on-write metadata |
| Snapshots | No (use LVM or Btrfs) | Volume Shadow Copy service | Yes, built in |
| Max file size | 16 TiB (4 KiB blocks) | Very large (practical limit set by volume size and cluster size) | Very large |
| Permissions | Unix mode bits, POSIX ACLs | Rich ACLs | Unix mode bits plus ACLs |
| Other | Delayed allocation, fsck still available | Alternate data streams, compression, encryption (EFS) | Clones (copy-on-write file copies), native encryption, space sharing between volumes |
Key points to remember:
- ext4 is the default on most Linux distributions. Its delayed allocation waits until data is flushed before choosing blocks, which leads to more contiguous extents.
- NTFS stores everything, including metadata, as files; even the MFT itself is described by an MFT record. Small files can live entirely inside their MFT record ("resident" data).
- APFS replaced HFS+ in 2017 and is optimised for SSDs. Copying a file on the same volume can create a clone that shares blocks until one copy is modified.
Others worth naming: XFS (scalable, default on Red Hat Enterprise Linux), Btrfs (copy-on-write, snapshots, checksums), ZFS (copy-on-write, checksums on every block, integrated volume management).
The virtual file system (VFS) layer
Linux supports dozens of file systems: ext4, XFS, Btrfs, FAT, NFS over the network, and pseudo file systems like /proc that have no disk at all. Programs use the same open, read, write for all of them. The virtual file system (VFS) makes that possible.
The VFS is an object-oriented layer inside the kernel. It defines common objects and operation tables; each file system supplies its own implementations.
| VFS object | Represents | Example operations |
|---|---|---|
| superblock | A mounted file system | sync, statfs |
| inode | One file (in memory) | create, lookup, link, mkdir |
| dentry | One path component, cached | name lookup, revalidate |
| file | One open instance of a file | read, write, llseek, mmap |
user program: read(fd, buf, n)
|
system call interface
|
+--------v---------+
| VFS | generic: fd -> file -> inode,
+--+-----+------+--+ page cache, dentry cache
| | |
ext4 NFS procfs file-system-specific code
| | |
block network kernel data
layer stack structures
A read() on a file reaches vfs_read(), which calls the read_iter function in the file's operation table, supplied by whichever file system owns the file. Windows has an equivalent architecture with file system drivers under the I/O manager.
Mounting
A file system must be mounted before use. Mounting attaches a file system's root directory to an existing directory, the mount point, in the single directory tree.
before: mount /dev/sdb1 /mnt/usb after
/ /
+--+---+ +--+---+
home mnt home mnt
| |
usb (empty dir) usb <- root of sdb1
+-+-+
photos docs
Steps the kernel takes:
- Read the device's superblock and check it is a valid file system of the given type.
- Create an in-memory superblock and record the mount in the mount table.
- Mark the mount point's dentry so that path lookups crossing it continue in the new file system's root.
Any files that were in the mount-point directory are hidden (not deleted) until it is unmounted. Linux's mount table is visible in /proc/mounts (or with findmnt), and file systems to mount at boot are listed in /etc/fstab. Windows uses drive letters by default but also supports mounting volumes into folders. Containers rely heavily on mounts: each container has its own mount namespace with its own view of the tree (see Linux internals).
Interview questions
Q1. What is an inode, and what does it not contain?
An inode is a per-file record holding the file's metadata: type, permissions, owner, size, timestamps, link count and pointers to its data blocks. It does not contain the file's name; names live in directory entries that map names to inode numbers. That separation is what makes hard links possible.
Q2. What is the difference between a hard link and a soft link?
A hard link is another directory entry for the same inode; the data lives until all names are removed and no process has it open. A soft link is a separate file containing a path; it can cross file systems and point to directories, but dangles if the target is removed or moved. Hard links cannot span file systems or point to directories.
Q3. Why can't you create a hard link to a directory?
It could create cycles in the directory graph, which would make traversals loop forever and make link counts unable to detect unreachable directories. Unix allows only the automatic . and .. entries. Symlinks are allowed because tools can recognise and skip them.
Q4. What happens when you delete a file that a process still has open?
rm calls unlink, which removes the directory entry and decrements the link count. If the count reaches zero but the file is still open, the inode and data stay allocated until the last descriptor is closed. This is why disk space is not freed when you delete a log file a running service is still writing.
Q5. Compare contiguous, linked and indexed allocation.
Contiguous is fastest for sequential and random access but suffers external fragmentation and makes growth hard. Linked has no fragmentation and easy growth but slow random access, pointer overhead and fragility. Indexed gives fast random access without external fragmentation at the cost of an index block per file; Unix combines direct and multi-level indirect pointers to keep small files cheap.
Q6. How does FAT improve on plain linked allocation?
It moves the next-block pointers out of the data blocks into a table at the start of the volume. The table can be cached in memory, so finding block i of a file means following pointers in RAM rather than reading i disk blocks. The table also tracks free blocks.
Q7. Compute the maximum file size with 4 KiB blocks, 4-byte pointers, 12 direct and one each of single, double and triple indirect.
Each block holds 1024 pointers. Reachable blocks = 12 + 1024 + 1024^2 + 1024^3 = 1,074,791,436. Times 4096 bytes gives 4,402,345,721,856 bytes, roughly 4 TiB.
Q8. How large is a free-space bitmap for a 1 TiB disk with 4 KiB blocks?
The disk has 2^40 / 2^12 = 2^28 blocks, one bit each, so 2^28 bits = 2^25 bytes = 32 MiB.
Q9. What is journaling and how does it speed up crash recovery?
Before changing on-disk structures, the file system writes the changes to a log and then a commit record. After a crash it replays committed transactions and discards uncommitted ones, so recovery scans only the journal instead of the whole disk. It turns multi-block updates into atomic transactions.
Q10. Explain ext4's ordered journaling mode.
Only metadata goes into the journal, but ext4 forces the related data blocks to disk before the metadata transaction commits. So after a crash, a file's metadata never points at blocks containing stale or garbage data. It is the default because it avoids writing data twice while preventing the worst inconsistencies.
Q11. Does journaling guarantee my application's writes are durable?
No. It keeps the file system's own structures consistent. Data written with write() may still be only in the page cache; call fsync() (and fsync the directory after a rename) when durability matters.
Q12. What is the VFS?
The Virtual File System is a kernel layer that defines common objects (superblock, inode, dentry, file) and operation tables. Each concrete file system implements those operations, so system calls like read work identically on ext4, NFS, FAT or /proc.
Q13. What does mounting do?
It attaches a file system's root to a directory (the mount point) in the existing tree, after reading and validating the superblock. Path lookups that reach the mount point continue in the mounted file system. Existing contents of the mount point are hidden until unmount.
Q14. What do the file offset and fork have to do with each other?
The offset lives in the system-wide open-file table entry, not in the per-process descriptor. After fork(), parent and child descriptors point to the same entry, so they share one offset and their writes append one after another. Two independent open() calls create two entries with separate offsets.
Q15. What are copy-on-write file systems?
ZFS, Btrfs and APFS never overwrite live blocks; they write modified blocks elsewhere and then atomically update the tree's root pointer. A crash leaves either the old or the new consistent tree, and snapshots come almost free by keeping old roots. Fragmentation of frequently rewritten files is a known cost.
Key takeaways
- A file system maps names to files, files to blocks, and tracks free space and permissions; names live in directories, metadata in inodes.
open()resolves the path once and returns a descriptor; offsets live in the shared open-file table entry, so forked processes share them.- Hard links share an inode and keep data alive; symlinks store a path and can dangle.
- Contiguous (extents), linked (FAT) and indexed allocation trade access speed against fragmentation and flexibility.
- Unix inodes use 12 direct plus single, double and triple indirect pointers; max size = (12 + p + p^2 + p^3) x block size, where p = block size / pointer size.
- Free space is tracked by bitmaps, lists, grouping or counting (extents); a bitmap costs one bit per block.
- Journaling replays committed transactions after a crash; ext4 defaults to ordered mode; copy-on-write file systems avoid overwriting live data.
- The VFS lets one set of system calls drive every file system, and mounting grafts a file system into the single tree.
Next lesson
Continue with Storage and I/O.

