Read

Snapshots: freezing a running computer

What the state of a running machine actually consists of, why an emulator can write all of it to a file, and what breaks when you restore one.

Ask a running computer to describe itself completely and it cannot. No instruction returns the contents of memory, the interrupt controller's mask register and the half-finished sector transfer inside the disk controller together — and any program collecting them would itself be part of the state it was recording. An emulator escapes that by not being inside the machine: everything the guest owns is a data structure the emulator holds, and a snapshot is those structures written out in order.

What a running machine is made of

The processor

The eight general registers, EIP and EFLAGS. Then the parts nobody thinks about: the segment registers plus their hidden descriptor caches, the base and limit values loaded when a selector was last written and readable by no instruction. CR0, CR3 pointing at the page directory, the GDTR and IDTR, the x87 stack with its tag word, and whether the processor sits in the one-instruction shadow after sti where interrupts stay blocked.

Memory

Guest RAM is the bulk of it, and no clever encoding is available: 512 MB of memory is 512 MB of state. Video memory is a separate array, 8 MB on a typical machine here. Caches and the TLB are the one part that can be thrown away: the architecture requires them to be transparent, so a restored guest cannot tell they were emptied.

Every emulated device

  • The 8259A interrupt controllers, master and slave: the mask register saying which IRQs are ignored, the request and in-service registers, the base vectors they were programmed with — 0x20 and 0x28 on a PC — and whether one is mid-way through its four-word initialisation sequence.
  • The 8254 timer, clocked at 1.193182 MHz: three counters, each with a mode, a reload value, a live count, and a flag saying whether the next read of the port returns the low byte or the high one.
  • The VGA: the current mode, the CRTC register file, the sequencer and graphics-controller registers, four bit-planes, the plane latches, and a 256-entry DAC palette of 6-bit RGB triples — plus which component the next write to port 0x3C9 will land on.
  • The IDE controller: which drive is selected, the LBA and sector-count registers, and — if a PIO transfer is under way — the offset within the current 512-byte sector. Frozen at byte 260, it must resume at byte 260.
  • The RTC: the CMOS bytes, the alarm and periodic-interrupt settings in status register B, and the offset between guest time and real time. Also the 8042 keyboard controller's output buffer and the serial UART's divisor latch and FIFO.

And the disks: every sector of every image attached, plus what the filesystem driver still holds in memory (what a filesystem actually is).

Why an emulator can do this and your laptop cannot

On real hardware most of that list is unreachable. The 8259's internal registers are largely write-only, the VGA palette index cannot be read back, a drive has a cache nobody can enumerate. Worse, the code doing the reading has to run, and running changes what is read.

So hibernation on a physical machine works by negotiation rather than capture: the kernel writes out its memory image and asks every driver to save what matters and reinitialise the rest on the way back. A laptop that wakes with a dead wifi card is a driver that got its half of the bargain wrong. An emulated snapshot has no such failure mode: it asks the drivers nothing.

In the emulator the interrupt controller is an object with a mask field, the timer three counters in a typed array, RAM one Uint8Array. Serialising a machine is walking a fixed list of buffers, and it comes out complete because the state has nowhere else to hide.

A frozen instant, not a cold boot

A disk image records a filesystem at rest. Restore one and the machine still has to be started: firmware, boot sector, kernel, init (the boot process explained). Whatever was in memory when the image was made is gone, because it was never in the image.

A snapshot records an instant. The guest continues from the point it was interrupted — forty instructions into a memcpy, with a half-typed command in a shell. Nothing inside can detect that anything happened, because detecting it would need state that was also restored.

The corollary is that a snapshot pins the shape of the machine: memory size, video memory and device set must match what was captured, because the file holds page tables referring to physical addresses that have to exist. Hence the machines in the catalog whose RAM is locked rather than adjustable.

What is capturedSnapshotDisk imageTemplate
CPU registers and flagsYes, to the instructionNoNo
Contents of RAMYes, all of itNoNo
Device state: PIC masks, PIT counts, VGA paletteYesNoNo
Files on diskAs of that instantAs last writtenNo, only a pointer to the image
Configuration: RAM, devices, boot orderImplicitly, and pinnedNoYes, explicitly
What you get on restoreThe same instant, mid-instructionA cold bootA machine ready to be booted

Why a restored machine appears instantly

Count what an honest boot does. Firmware counts memory and enumerates buses. The bootloader reads sectors one at a time through BIOS interrupts. The kernel decompresses itself — a bzImage is a decompressor stub with a compressed image behind it. Then device probing, some of it with timeouts measured in seconds because a controller that is not there has to be waited for. Then userspace: udev settling the device tree, a service manager building a dependency graph. Under emulation each is a stream of small operations against a cold JIT (how x86 emulation works), which is why a system that boots in eight seconds on hardware takes minutes in a tab.

A restore does none of it: allocate the buffers, copy the bytes in, set the registers, jump. The cost is proportional to the size of memory and nothing else. It is the trick behind QEMU's savevm, the .vmss and .vmem pair VMware writes on suspend, and VirtualBox's saved state — and why some bases here appear in a fraction of a second. The boot happened once, on someone else's machine, months ago.

Copy-on-write, and why chains exist

Snapshots are large in the least interesting way possible: a 512 MB machine yields roughly 512 MB of file, and ten of them are five gigabytes that are largely byte-identical. The answer is copy-on-write: keep one base read-only, record only what changed after it.

On disk that is qcow2's backing_file: reads fall through to the parent for any cluster the child lacks, writes land in the child (disk images explained). VMware calls it a linked clone; VirtualBox calls them differencing images.

# One base install, three machines, almost no extra bytes.
qemu-img create -f qcow2 -b arch-base.qcow2 -F qcow2 lab-01.qcow2
qemu-img create -f qcow2 -b arch-base.qcow2 -F qcow2 lab-02.qcow2

qemu-img info --backing-chain lab-01.qcow2
  image: lab-01.qcow2
  virtual size: 10 GiB (10737418240 bytes)
  disk size: 41.2 MiB              # only what this machine changed
  backing file: arch-base.qcow2

  image: arch-base.qcow2
  virtual size: 10 GiB
  disk size: 1.42 GiB              # shared by every child

# Fold a child into its parent; the chain is one link shorter.
qemu-img commit lab-01.qcow2

# Machine state, saved separately from the disk:
#   (qemu) savevm before-upgrade
#   (qemu) loadvm before-upgrade

The costs are real. Every read may walk the chain, so a machine four links deep is slower than one at the top, and deleting a middle link is a merge rather than a delete. Losing the base makes every descendant unreadable; modifying it by accident corrupts all of them at once. Hence bases marked read-only and never booted directly.

Three ways a restored machine will mislead you

The clock goes backwards

Restoring rewinds the guest's clock, and a great deal of software treats time as monotonic. TLS handshakes fail against certificates not valid yet, or valid at capture time and since expired. Kerberos refuses tickets outside a default five-minute clock skew. TCP connections are still marked established although the peer sent a reset hours ago, and DHCP leases have lapsed. This is why hypervisor guest agents resynchronise over NTP the moment a machine resumes.

Two machines, one identity

Clone a snapshot and you have two guests that agree about who they are: the same MAC address, so the switch sees one host in two places; the same /etc/machine-id, from which systemd derives the DHCP client identifier; the same SSH host keys, so no client can tell them apart. On Windows, the same security identifier — sysprep exists for exactly this reason.

It contains everything that was in memory

A snapshot is a complete memory dump with a friendlier name. The master key of an unlocked encrypted volume sits in the kernel keyring; the private key of any running TLS server is in its address space, as are passwords typed into a shell. Memory forensics tools read these files directly — Volatility will happily analyse a VMware .vmem. Sending a snapshot to a colleague sends all of that with it.

Snapshots on this site

v86 exposes it as two calls. save_state() returns an ArrayBuffer — a header, the CPU structure, guest memory, video memory, then each device's fields in a fixed order — and restore_state() takes that buffer back.

// Freeze. Roughly the size of guest RAM, so not something to do in a loop.
const state = await emulator.save_state();

// ... break something interesting ...

// Rewind to the instruction it was executing when you froze it.
await emulator.restore_state(state);

// Keep it beyond this tab: write the same buffer out as a file.
const blob = new Blob([state], { type: "application/octet-stream" });

An in-memory snapshot is the quick kind: take one before something risky, restore it when the risk pays off badly. A reload takes it with it. An exported state file is the same buffer written to disk, and it is the only way a machine here outlives its tab — load it back tomorrow, or on another computer, and you get the same instant. Expect roughly the size of the machine's memory; the pre-booted bases in the catalog are compressed with zstd, usually to a fraction of that.

A template in the library is a different object: it saves the machine's recipe, not its state. Which image to fetch, how much RAM, which devices, what boots first — a few hundred bytes of JSON you can also hand to the JSON builder. Templates are tiny, readable, and always produce a cold boot. Snapshots resume an instant, and are enormous, opaque and pinned to one configuration.

The workflow worth learning

  1. Start a machine from the catalog or build one at /new and let it boot properly once. This is the slow part, paid once.
  2. Get it to the state that is actually interesting: logged in, disk mounted, file loaded, editor open, JIT warmed up.
  3. Take a snapshot — on a 512 MB machine that costs roughly that much memory again — and export it if reaching this state took more than a minute.
  4. Now break it deliberately: overwrite the boot sector with dd, delete /lib, corrupt the FAT. The point of a disposable machine is doing what you would never do to a real one.
  5. Watch the failure closely — which message appears, at what stage, how far the machine gets before it stops. That is the lesson.
  6. Restore, and repeat with a different kind of damage. Terms you do not recognise are in the glossary.