What a filesystem actually is, from FAT to 9p
A disk is a numbered array of blocks and nothing else; every name, size and byte offset you think of as a file is bookkeeping written into those blocks.
A disk offers a computer exactly one thing: a numbered array of fixed-size blocks that can be read or written whole. Block 4,192 is 512 bytes, or 4,096, and it has no name, no owner and no relationship to block 4,193. Everything else — directories, filenames, timestamps, the idea that a file has a beginning — is bookkeeping that somebody wrote into some of those blocks and agreed to interpret the same way next time.
The four pieces of bookkeeping
Filesystems differ enormously in how they do it, but they are all answering the same four questions, and it is worth being able to name them before looking at any particular design.
- The superblock — where do I even start? A block at a known offset holding the block size, the total count, and where the other structures live. Lose it and the disk is a pile of blocks again, which is why filesystems keep backup copies of it.
- The allocation map — which blocks are free? A bitmap, one bit per block, or a table, or a tree. Somebody has to know, or two files end up in the same place.
- Directory entries — what is this thing called? A name paired with a pointer to the file's metadata. A directory is just a file whose contents happen to be a list of these.
- The content map — where are the bytes? Given a file and an offset into it, produce a block number. This is where the designs genuinely diverge, and it is what the rest of this guide is about.
FAT: a linked list kept in a table
The boot sector doubles as the superblock
FAT's superblock is the BIOS Parameter Block, sitting inside the boot sector immediately after the jump instruction and an eight-byte OEM name — which is why a FAT boot sector has so few bytes left for code, as write your own boot sector explains. Here is a freshly formatted 1.44 MB floppy, annotated. Every one of these fields is read by DOS before it can find a single file.
00000000 eb 3c 90 jmp short 0x3e / nop
00000003 6d 6b 66 73 2e 66 61 74 OEM name 'mkfs.fat'
0000000b 00 02 bytes per sector 512
0000000d 01 sectors per cluster 1
0000000e 01 00 reserved sectors 1 (the boot sector itself)
00000010 02 number of FATs 2 (the second is a copy)
00000011 e0 00 root dir entries 224
00000013 40 0b total sectors 2880
00000015 f0 media descriptor 0xF0 = 3.5 inch, 2 heads
00000016 09 00 sectors per FAT 9
00000018 12 00 sectors per track 18
0000001a 02 00 heads 2
0000001c 00 00 00 00 hidden sectors 0
00000020 00 00 00 00 large sector count 0 (unused: fits in 16 bits)
00000024 00 00 29 drive 0x00, ext boot signature 0x29
00000027 8e 5c 1b 44 volume serial number
0000002b 54 45 4d 50 4d 56 20 20 ... volume label 'TEMPMV '
00000036 46 41 54 31 32 20 20 20 filesystem type 'FAT12 '
... (boot code)
000001fe 55 aa boot signature
# 1 reserved + (2 x 9) FAT sectors + 14 sectors of root directory = 33.
# Data starts at sector 33, leaving 2880 - 33 = 2847 clusters.
# Under 4085 clusters means FAT12. That threshold is the only thing
# that distinguishes FAT12 from FAT16 from FAT32.
The table itself
The File Allocation Table is one entry per cluster, and each entry holds the number of the *next* cluster in the same file. A directory entry gives the first cluster; you look it up, find the second, look that up, and continue until you hit an end-of-chain marker — 0xFF8–0xFFF in FAT12, 0xFFF8–0xFFFF in FAT16. The allocation map and the content map are the same structure, which is elegant and cheap.
It is also why FAT is slow on large files and why fragmentation hurt so much: reaching byte 900,000,000 means walking the chain from the start, one lookup per cluster. And it is why a damaged FAT is catastrophic while a damaged directory entry is merely annoying — hence the second copy of the table that every FAT volume carries.
8.3, and the trick that gave us long names
A directory entry is 32 bytes: eight for the name, three for the extension, one attribute byte, the first cluster number, and a 32-bit size. No dot is stored — it is implied. AUTOEXEC.BAT occupies exactly the eleven name bytes with the extension right-padded, and lower case does not exist.
VFAT, shipped with Windows 95, added long names without breaking anything, by a trick worth admiring. A long name is stored in extra directory entries placed *before* the real one, each carrying 13 UTF-16 characters and an attribute byte of 0x0F — read-only, hidden, system and volume-label all at once, a combination no real file could have. Old DOS skips them as nonsense and sees only the 8.3 entry. Up to 20 such entries give 255 characters.
Why FAT is still everywhere
Because it is small enough to implement in a device that has almost nothing. FAT32 is what an EFI System Partition must be, by specification, so every UEFI machine ships with one. It is what SD cards and USB sticks arrive formatted as, what cameras and synthesisers and 3D printers write. Its famous limit is the 32-bit size field: 4,294,967,295 bytes, one byte short of 4 GiB, which is why a large video file will not copy onto a memory card.
ext2, ext3, ext4: inodes and extents
The Unix answer separates the name from the file. Every file is an inode — 128 bytes in ext2, 256 by default in ext4 — holding the mode, owner, timestamps, link count and the content map. The name lives only in a directory entry pointing at the inode number. That indirection is what makes hard links possible: two names, one inode, a link count of two.
The disk is carved into block groups, each with its own inode table and bitmaps, so a file's metadata and its data tend to sit near each other. ext2's content map is 15 block pointers: twelve direct, then one singly indirect block, one doubly indirect, one triply indirect. Small files cost nothing extra; a large file costs a lookup per level, and the map for a gigabyte is thousands of pointer blocks.
ext4 replaced that with extents: a start block plus a length, describing up to 128 MiB of contiguous space with 4 KiB blocks. Four of them fit directly in the inode's 60-byte map area; beyond that they form a tree. A contiguous 1 GB file needs eight extents rather than a quarter of a million pointers.
$ dumpe2fs -h /dev/sda1
Filesystem features: has_journal ext_attr dir_index filetype extent 64bit
flex_bg sparse_super large_file huge_file metadata_csum
Filesystem state: clean
Inode count: 65536
Block count: 262144
Free blocks: 238919
Block size: 4096
Blocks per group: 32768
Inodes per group: 8192
Inode size: 256
Journal size: 16M
$ stat notes.txt
File: notes.txt
Size: 5182 Blocks: 16 IO Block: 4096 regular file
Device: 8,1 Inode: 262146 Links: 1
# 5182 bytes occupies two 4 KiB blocks. 'Blocks: 16' counts 512-byte
# units, which is a Unix habit older than the 4 KiB block itself.
ext3 added the journal, and it is worth being precise about what that buys. Before touching the filesystem, the change is written to a reserved area and marked committed; after a crash, the journal is replayed and the structure is consistent again. What it protects is the metadata — no half-created inodes, no blocks claimed twice. In the default data=ordered mode it does not promise that the *contents* of a file being written survive. A journal saves you a two-hour fsck. It does not save your unsaved work.
ISO 9660: a filesystem that cannot change
Written for a medium that is stamped once, ISO 9660 has no allocation map at all — there is nothing to allocate. Files are contiguous runs of 2,048-byte sectors, and a directory record simply gives a start block and a length. The Primary Volume Descriptor at sector 16 begins with the characters CD001. Level 1 names are 8.3, uppercase, with a version suffix, so README.TXT;1 is the real name on the disc.
Two extensions fixed that without breaking old readers. Rock Ridge hides POSIX metadata — real names up to 255 characters, permissions, ownership, symlinks — in fields the base standard ignores, which is how a Linux install disc has a /usr/bin at all. Joliet takes the other approach and publishes a second, parallel directory tree in a supplementary volume descriptor, with names in UCS-2 up to 64 characters. A disc often carries both, and two operating systems read the same files under different names.
9p: a filesystem as a conversation
Everything above describes bytes on a disk. 9P describes a dialogue instead. It came out of Plan 9, the operating system Bell Labs built after Unix, in which every resource — a process, a network connection, the window system — is a file tree served by some program, and one protocol reaches all of them.
The 9P2000 revision has fourteen request/response pairs and no more: version, auth, attach, flush, walk, open, create, read, write, clunk, remove, stat, wstat, and an error reply. A client holds a fid, a 32-bit handle for a file it is interested in; the server answers with a qid, thirteen bytes that identify that file uniquely and for ever. walk moves a fid along up to sixteen path elements at a time. version negotiates the maximum message size and nothing else is configurable. You can implement the whole thing in an afternoon, which is precisely the point.
Because it is that small, it became the standard way to hand a host directory to a guest without a disk image: QEMU's virtio-9p does it, and so does v86. Point it at a basefs — a JSON index of every file, its size, mode and content hash — and a baseurl directory holding the contents, and the guest gets a root filesystem it can walk and open normally. Nothing is downloaded until a file is actually read, at which point the emulator fetches it over HTTP.
That is how a full Arch Linux userland boots in a tab. The alternative would be a gigabyte-scale disk image; the reality is a few hundred kilobytes of index plus whatever the boot touches. The cost is latency: every cold file is a round trip, so a find / is painfully slow, and a package install that reads thousands of small files can take longer than the compilation it precedes. Loading whole images and loading files this way are compared in disk images explained.
tmpfs, and why nothing here is saved
A tmpfs is a filesystem with no block device underneath it. It lives in the kernel's page cache, borrows pages as files are written, returns them when files are deleted, and defaults to a ceiling of half of RAM. Linux uses one for /tmp, for /dev/shm, and for the initramfs it unpacks before it has mounted anything real. It is fast because there is no disk, and empty on every boot for the same reason.
Every machine on this site behaves like that all the way down. The 9p layer holds writes in JavaScript memory and never sends them anywhere; a raw disk image is an ArrayBuffer that the guest happily writes into and that the browser discards when the tab does. There is no server, so there is nowhere for a write to go. The one way to keep a machine's state is to export it deliberately, which is snapshots and machine state.
The same job, seven ways
| Filesystem | Max file | Max volume | Journal | Case | Typical use |
|---|---|---|---|---|---|
| FAT12 | 32 MB in practice | 32 MB with 8 KB clusters | No | Insensitive, 8.3 | Floppies, tiny boot volumes |
| FAT16 | 2 GB | 2 GB (4 GB with 64 KB clusters) | No | Insensitive | DOS and early Windows |
| FAT32 | 4 GiB − 1 byte | 2 TiB with 512-byte sectors | No | Insensitive, preserved | EFI System Partitions, SD cards, USB sticks |
| ext2 | 2 TiB | 16 TiB with 4 KiB blocks | No | Sensitive | Small Linux images, flash without wear levelling |
| ext3 | 2 TiB | 16 TiB with 4 KiB blocks | Yes, metadata | Sensitive | Linux, 2001 to roughly 2010 |
| ext4 | 16 TiB with 4 KiB blocks | 1 EiB | Yes, metadata | Sensitive | The default Linux root filesystem |
| ISO 9660 | 4 GiB − 1 (levels 1 and 2) | 8 TiB | Read-only | Uppercase only at level 1 | Optical discs and install images |
| 9p | Whatever the server allows | Whatever the server allows | The server's problem | The server's rules | Handing a host directory to a guest |
| tmpfs | Limited by RAM and swap | Half of RAM by default | No | Sensitive | /tmp, initramfs, and every machine here |