Read

What a filesystem actually is, from FAT to 9p

A disk is a numbered array of blocks and nothing else; every name, size and byte offset you think of as a file is bookkeeping written into those blocks.

A disk offers a computer exactly one thing: a numbered array of fixed-size blocks that can be read or written whole. Block 4,192 is 512 bytes, or 4,096, and it has no name, no owner and no relationship to block 4,193. Everything else — directories, filenames, timestamps, the idea that a file has a beginning — is bookkeeping that somebody wrote into some of those blocks and agreed to interpret the same way next time.

The four pieces of bookkeeping

Filesystems differ enormously in how they do it, but they are all answering the same four questions, and it is worth being able to name them before looking at any particular design.

  • The superblock — where do I even start? A block at a known offset holding the block size, the total count, and where the other structures live. Lose it and the disk is a pile of blocks again, which is why filesystems keep backup copies of it.
  • The allocation map — which blocks are free? A bitmap, one bit per block, or a table, or a tree. Somebody has to know, or two files end up in the same place.
  • Directory entries — what is this thing called? A name paired with a pointer to the file's metadata. A directory is just a file whose contents happen to be a list of these.
  • The content map — where are the bytes? Given a file and an offset into it, produce a block number. This is where the designs genuinely diverge, and it is what the rest of this guide is about.

FAT: a linked list kept in a table

The boot sector doubles as the superblock

FAT's superblock is the BIOS Parameter Block, sitting inside the boot sector immediately after the jump instruction and an eight-byte OEM name — which is why a FAT boot sector has so few bytes left for code, as write your own boot sector explains. Here is a freshly formatted 1.44 MB floppy, annotated. Every one of these fields is read by DOS before it can find a single file.

00000000  eb 3c 90                      jmp short 0x3e / nop
00000003  6d 6b 66 73 2e 66 61 74       OEM name         'mkfs.fat'
0000000b  00 02                         bytes per sector      512
0000000d  01                            sectors per cluster   1
0000000e  01 00                         reserved sectors      1   (the boot sector itself)
00000010  02                            number of FATs        2   (the second is a copy)
00000011  e0 00                         root dir entries      224
00000013  40 0b                         total sectors         2880
00000015  f0                            media descriptor      0xF0 = 3.5 inch, 2 heads
00000016  09 00                         sectors per FAT       9
00000018  12 00                         sectors per track     18
0000001a  02 00                         heads                 2
0000001c  00 00 00 00                   hidden sectors        0
00000020  00 00 00 00                   large sector count    0   (unused: fits in 16 bits)
00000024  00 00 29                      drive 0x00, ext boot signature 0x29
00000027  8e 5c 1b 44                   volume serial number
0000002b  54 45 4d 50 4d 56 20 20 ...   volume label     'TEMPMV     '
00000036  46 41 54 31 32 20 20 20       filesystem type  'FAT12   '
   ...    (boot code)
000001fe  55 aa                         boot signature

# 1 reserved + (2 x 9) FAT sectors + 14 sectors of root directory = 33.
# Data starts at sector 33, leaving 2880 - 33 = 2847 clusters.
# Under 4085 clusters means FAT12. That threshold is the only thing
# that distinguishes FAT12 from FAT16 from FAT32.

The table itself

The File Allocation Table is one entry per cluster, and each entry holds the number of the *next* cluster in the same file. A directory entry gives the first cluster; you look it up, find the second, look that up, and continue until you hit an end-of-chain marker — 0xFF8–0xFFF in FAT12, 0xFFF8–0xFFFF in FAT16. The allocation map and the content map are the same structure, which is elegant and cheap.

It is also why FAT is slow on large files and why fragmentation hurt so much: reaching byte 900,000,000 means walking the chain from the start, one lookup per cluster. And it is why a damaged FAT is catastrophic while a damaged directory entry is merely annoying — hence the second copy of the table that every FAT volume carries.

8.3, and the trick that gave us long names

A directory entry is 32 bytes: eight for the name, three for the extension, one attribute byte, the first cluster number, and a 32-bit size. No dot is stored — it is implied. AUTOEXEC.BAT occupies exactly the eleven name bytes with the extension right-padded, and lower case does not exist.

VFAT, shipped with Windows 95, added long names without breaking anything, by a trick worth admiring. A long name is stored in extra directory entries placed *before* the real one, each carrying 13 UTF-16 characters and an attribute byte of 0x0F — read-only, hidden, system and volume-label all at once, a combination no real file could have. Old DOS skips them as nonsense and sees only the 8.3 entry. Up to 20 such entries give 255 characters.

Why FAT is still everywhere

Because it is small enough to implement in a device that has almost nothing. FAT32 is what an EFI System Partition must be, by specification, so every UEFI machine ships with one. It is what SD cards and USB sticks arrive formatted as, what cameras and synthesisers and 3D printers write. Its famous limit is the 32-bit size field: 4,294,967,295 bytes, one byte short of 4 GiB, which is why a large video file will not copy onto a memory card.

ext2, ext3, ext4: inodes and extents

The Unix answer separates the name from the file. Every file is an inode — 128 bytes in ext2, 256 by default in ext4 — holding the mode, owner, timestamps, link count and the content map. The name lives only in a directory entry pointing at the inode number. That indirection is what makes hard links possible: two names, one inode, a link count of two.

The disk is carved into block groups, each with its own inode table and bitmaps, so a file's metadata and its data tend to sit near each other. ext2's content map is 15 block pointers: twelve direct, then one singly indirect block, one doubly indirect, one triply indirect. Small files cost nothing extra; a large file costs a lookup per level, and the map for a gigabyte is thousands of pointer blocks.

ext4 replaced that with extents: a start block plus a length, describing up to 128 MiB of contiguous space with 4 KiB blocks. Four of them fit directly in the inode's 60-byte map area; beyond that they form a tree. A contiguous 1 GB file needs eight extents rather than a quarter of a million pointers.

$ dumpe2fs -h /dev/sda1
Filesystem features:   has_journal ext_attr dir_index filetype extent 64bit
                       flex_bg sparse_super large_file huge_file metadata_csum
Filesystem state:      clean
Inode count:           65536
Block count:           262144
Free blocks:           238919
Block size:            4096
Blocks per group:      32768
Inodes per group:      8192
Inode size:            256
Journal size:          16M

$ stat notes.txt
  File: notes.txt
  Size: 5182        Blocks: 16         IO Block: 4096   regular file
Device: 8,1  Inode: 262146      Links: 1

# 5182 bytes occupies two 4 KiB blocks. 'Blocks: 16' counts 512-byte
# units, which is a Unix habit older than the 4 KiB block itself.

ext3 added the journal, and it is worth being precise about what that buys. Before touching the filesystem, the change is written to a reserved area and marked committed; after a crash, the journal is replayed and the structure is consistent again. What it protects is the metadata — no half-created inodes, no blocks claimed twice. In the default data=ordered mode it does not promise that the *contents* of a file being written survive. A journal saves you a two-hour fsck. It does not save your unsaved work.

ISO 9660: a filesystem that cannot change

Written for a medium that is stamped once, ISO 9660 has no allocation map at all — there is nothing to allocate. Files are contiguous runs of 2,048-byte sectors, and a directory record simply gives a start block and a length. The Primary Volume Descriptor at sector 16 begins with the characters CD001. Level 1 names are 8.3, uppercase, with a version suffix, so README.TXT;1 is the real name on the disc.

Two extensions fixed that without breaking old readers. Rock Ridge hides POSIX metadata — real names up to 255 characters, permissions, ownership, symlinks — in fields the base standard ignores, which is how a Linux install disc has a /usr/bin at all. Joliet takes the other approach and publishes a second, parallel directory tree in a supplementary volume descriptor, with names in UCS-2 up to 64 characters. A disc often carries both, and two operating systems read the same files under different names.

9p: a filesystem as a conversation

Everything above describes bytes on a disk. 9P describes a dialogue instead. It came out of Plan 9, the operating system Bell Labs built after Unix, in which every resource — a process, a network connection, the window system — is a file tree served by some program, and one protocol reaches all of them.

The 9P2000 revision has fourteen request/response pairs and no more: version, auth, attach, flush, walk, open, create, read, write, clunk, remove, stat, wstat, and an error reply. A client holds a fid, a 32-bit handle for a file it is interested in; the server answers with a qid, thirteen bytes that identify that file uniquely and for ever. walk moves a fid along up to sixteen path elements at a time. version negotiates the maximum message size and nothing else is configurable. You can implement the whole thing in an afternoon, which is precisely the point.

Because it is that small, it became the standard way to hand a host directory to a guest without a disk image: QEMU's virtio-9p does it, and so does v86. Point it at a basefs — a JSON index of every file, its size, mode and content hash — and a baseurl directory holding the contents, and the guest gets a root filesystem it can walk and open normally. Nothing is downloaded until a file is actually read, at which point the emulator fetches it over HTTP.

That is how a full Arch Linux userland boots in a tab. The alternative would be a gigabyte-scale disk image; the reality is a few hundred kilobytes of index plus whatever the boot touches. The cost is latency: every cold file is a round trip, so a find / is painfully slow, and a package install that reads thousands of small files can take longer than the compilation it precedes. Loading whole images and loading files this way are compared in disk images explained.

tmpfs, and why nothing here is saved

A tmpfs is a filesystem with no block device underneath it. It lives in the kernel's page cache, borrows pages as files are written, returns them when files are deleted, and defaults to a ceiling of half of RAM. Linux uses one for /tmp, for /dev/shm, and for the initramfs it unpacks before it has mounted anything real. It is fast because there is no disk, and empty on every boot for the same reason.

Every machine on this site behaves like that all the way down. The 9p layer holds writes in JavaScript memory and never sends them anywhere; a raw disk image is an ArrayBuffer that the guest happily writes into and that the browser discards when the tab does. There is no server, so there is nowhere for a write to go. The one way to keep a machine's state is to export it deliberately, which is snapshots and machine state.

The same job, seven ways

FilesystemMax fileMax volumeJournalCaseTypical use
FAT1232 MB in practice32 MB with 8 KB clustersNoInsensitive, 8.3Floppies, tiny boot volumes
FAT162 GB2 GB (4 GB with 64 KB clusters)NoInsensitiveDOS and early Windows
FAT324 GiB − 1 byte2 TiB with 512-byte sectorsNoInsensitive, preservedEFI System Partitions, SD cards, USB sticks
ext22 TiB16 TiB with 4 KiB blocksNoSensitiveSmall Linux images, flash without wear levelling
ext32 TiB16 TiB with 4 KiB blocksYes, metadataSensitiveLinux, 2001 to roughly 2010
ext416 TiB with 4 KiB blocks1 EiBYes, metadataSensitiveThe default Linux root filesystem
ISO 96604 GiB − 1 (levels 1 and 2)8 TiBRead-onlyUppercase only at level 1Optical discs and install images
9pWhatever the server allowsWhatever the server allowsThe server's problemThe server's rulesHanding a host directory to a guest
tmpfsLimited by RAM and swapHalf of RAM by defaultNoSensitive/tmp, initramfs, and every machine here