Virtual machine, container or emulator: what is actually different
Three technologies people use interchangeably and shouldn't — the kernel each one shares, the instruction set each one can run, and how strong the wall really is.
Run uname -r inside a Docker container on a Linux box and you get the host's kernel version, not the container image's. That one line settles most of the confusion. A container is not a small machine running a small operating system — it is a process on your kernel, wearing a costume. A virtual machine really does run its own kernel, on your processor, in a mode the processor provides for the purpose. An emulator runs a kernel too, but on a processor that does not exist outside of software. Three different mechanisms, three different boundaries, three different bills.
A container is a process
There is no container object in the Linux kernel. There is no container_create() syscall. What exists is a set of independently developed features that, used together, make a process believe it is alone on a machine. Docker is the packaging and the ergonomics; the isolation is all kernel.
- Namespaces give a process a private view of one kind of global resource. There are eight: mount (2002), UTS, IPC, PID, network, user (3.8, 2013), cgroup (4.6) and time (5.6). A PID namespace is why the container's init is process 1; a mount namespace is why its
/is not yours. - cgroups limit and account for resources — CPU shares, memory ceilings, block I/O. Started at Google in 2006 as "process containers", merged in 2.6.24, rewritten as cgroup v2 in 4.5. Namespaces hide things; cgroups ration them.
- OverlayFS (mainline since 3.18) stacks a read-only image as lower layers with a thin writable upper layer. That is what makes a container image cheap to start and cheap to store: fifty containers from one image share one copy of the bytes.
- seccomp-bpf and capabilities narrow what the process may ask the kernel for. Docker's default profile blocks around forty syscalls outright and drops most of the root capabilities.
$ ls -l /proc/self/ns
cgroup -> cgroup:[4026531835]
ipc -> ipc:[4026531839]
mnt -> mnt:[4026531840]
net -> net:[4026531992]
pid -> pid:[4026531836]
user -> user:[4026531837]
uts -> uts:[4026531838]
# A container, by hand, with no container runtime at all:
$ sudo unshare --pid --mount --uts --net --fork --mount-proc bash
# hostname isolated && ps aux
PID COMMAND
1 bash
5 ps aux
# Same kernel underneath, though:
$ uname -r
6.8.0-45-generic
The ancestry runs back further than Docker: chroot arrived in Version 7 Unix in 1979, FreeBSD jails in 2000, Solaris Zones in 2004, LXC in 2008. Docker's contribution in 2013 was the image format and the registry, not the isolation.
Hardware virtualisation is a processor mode
A virtual machine under KVM, Hyper-V or ESXi runs its guest instructions on your actual silicon. Intel's VT-x adds two operating modes: VMX root, where the hypervisor lives, and VMX non-root, where the guest lives with its own full set of rings 0 to 3. People call root mode "ring -1"; Intel's manuals never do. A structure in memory, the VMCS, says which guest events should bounce control back to the hypervisor — a VM exit. Everything else the guest does executes directly.
The remaining cost is memory translation and devices. Nested paging — Intel EPT, AMD's RVI — lets the processor walk two levels of page tables in hardware, so a guest page fault does not have to become a VM exit. IOMMUs let a real PCI device be handed to a guest safely. The result is genuinely close to native: for CPU-bound work you are usually arguing about single-digit percentages.
The catch is the one people forget: a hypervisor cannot change the instruction set. VT-x virtualises x86 for x86 guests. Apple's Virtualization.framework virtualises ARM64 for ARM64 guests. If the guest's instructions are not instructions your processor understands, no amount of hypervisor helps, because the whole design depends on the processor executing them.
Emulation is reading a foreign machine's mail
An emulator implements the guest processor in software: fetch the next instruction, decode it, update a model of the registers and memory, repeat. Bochs does exactly this, purely and slowly, and its faithfulness is why it is still a reference implementation. Pure interpretation costs dozens of host instructions per guest instruction, most of them spent re-deciding what an instruction you have already seen a million times means.
So the fast emulators translate instead. QEMU without KVM uses TCG, the Tiny Code Generator: it converts a block of guest instructions into a small intermediate representation, then into host machine code, and caches the result. Apple's Rosetta 2 goes further and translates x86-64 binaries to ARM64 ahead of time at install, falling back to JIT only for code generated at runtime; Apple Silicon even has a hardware switch that makes the core use x86's stricter memory ordering, which is the sort of thing that turns a crippling penalty into a tolerable one. v86, the emulator TempMV runs on, translates hot x86 blocks into WebAssembly modules and hands them to the browser's own optimising compiler.
Translation buys a lot and still cannot buy everything. Faithful x86 semantics means computing flags, honouring the memory model, and getting self-modifying code right. Expect an order of magnitude off native, better on tight loops that stay in the translation cache, worse on code that keeps invalidating it. How x86 emulation works in the browser goes into the mechanism.
Paravirtualisation and the honest middle
Before VT-x existed, Xen's 2003 answer to x86's eighteen awkward instructions was to stop pretending. In paravirtualisation the guest kernel is modified: instead of executing a privileged instruction and hoping it traps, it makes an explicit hypercall to the hypervisor. Faster than binary translation, and it required no new silicon — but it required a patched guest, which is why it faded once hardware support arrived.
The idea survived where it was always sensible: devices. virtio drivers do not pretend to be an emulated Realtek card that the hypervisor must decode register by register. They are an agreed queue between guest and host, so a packet costs one notification instead of dozens of trapped I/O writes. Every serious VM today is hardware-virtualised for the CPU and paravirtualised for the devices — including the v86 machines here, which offer virtio-net, virtio-block and a 9p filesystem beside the emulated NE2000 and IDE.
Two more hybrids are worth naming. gVisor puts a user-space kernel between the container and the host, so most guest syscalls never reach the real kernel at all. Firecracker goes the other way: a genuine KVM virtual machine, but with only five devices — virtio-net, virtio-block, virtio-vsock, a serial console and a stub keyboard controller — which is how it boots in under 125 ms with under 5 MiB of overhead. Both exist because container isolation was judged too thin for running other people's code.
Where WebAssembly sits
WebAssembly is none of the three, and calling a Wasm runtime a virtual machine causes half the arguments on this subject. No instruction set is being virtualised or emulated: Wasm is a stack machine designed to be compiled, its memory is a single bounds-checked array, and its only route outward is a function the host explicitly handed it. A container escape means finding a kernel bug; a Wasm module has no ambient authority to escape from — deny it a file handle and there is no syscall to try. Smaller to reason about, and far less capable: no threads by default, no raw sockets, no fork. What WebAssembly actually is covers it properly.
The four side by side
| Approach | Isolation boundary | Kernel | Guest ISA | Startup | Overhead | Typical use |
|---|---|---|---|---|---|---|
| Container | Kernel namespaces, cgroups, seccomp. Enforced by software you also depend on. | Shared with the host. uname -r gives the host's. | Host only. Same architecture, no exceptions. | Tens of milliseconds | Effectively zero | Packaging a userland, CI, microservices |
| Hardware VM (KVM, ESXi) | VT-x/AMD-V, enforced by the processor. Surface is the VMM's device model. | Its own, complete and independent. | Host ISA only. | Seconds; ~125 ms for a microVM | A few percent on CPU work | Multi-tenant hosting, other operating systems, untrusted code |
| Emulator (QEMU TCG, v86) | The guest CPU does not exist. Escaping means breaking the emulator's own process. | Its own, complete and independent. | Any — that is the entire point. | Seconds to a minute | Roughly 10x, sometimes far worse | Cross-architecture work, retrocomputing, firmware bring-up, browsers |
| Wasm runtime | Capability-based: no imports, no reach. Memory is a bounds-checked array. | None. There is no kernel and no machine. | Wasm bytecode, compiled to the host's. | Under a millisecond | Commonly 10-50% over native | Plugins, edge functions, sandboxing one library |
Why x86 Linux will not run in a container on an ARM Mac
Because a container's binaries are executed by the host kernel on the host processor, and an M-series chip cannot decode x86 instructions. There is no translation layer in the container mechanism; that was the whole reason it is fast. Run docker run --platform linux/amd64 alpine on an Apple Silicon Mac and it appears to work anyway. What happens is that a qemu-x86_64 user-mode emulator is registered with the Linux kernel's binfmt_misc handler inside the VM, so every x86 binary you launch is silently handed to an emulator. It works. It is several times slower, and it occasionally fails on binaries that use instructions the user-mode emulator handles badly.
The phrase "inside the VM" is the other half of the answer. macOS has no Linux kernel, so it cannot host Linux namespaces. Docker Desktop on a Mac is a Linux virtual machine with a Docker daemon in it — today on Apple's Virtualization.framework or Docker's own VMM, historically on HyperKit. Docker Desktop on Windows is the same trick: the WSL 2 backend is a real Linux kernel in a lightweight Hyper-V VM. Your containers on those platforms are two boundaries deep, and the reason docker run still feels instant is that the VM was already running before you typed it.
The security boundaries, precisely
A container's boundary is the Linux syscall interface — over 350 entry points on x86-64, plus /proc, /sys, ioctl and every driver reachable through them. That is an enormous surface, defended by code that is also your host's kernel. Real escapes exist and are not exotic: CVE-2019-5736 let a container overwrite the host's runc binary through /proc/self/exe; CVE-2024-21626 leaked a file descriptor into the container's namespace. Neither needed a hypervisor bug, because there was no hypervisor.
A virtual machine's boundary is much narrower: the VMCS-configured exits and the VMM's emulated devices. Narrower is not zero. VENOM (CVE-2015-3456) was a buffer overflow in QEMU's virtual floppy controller — a device nobody used, compiled in by default, reachable from any guest. The lesson everyone drew is Firecracker's five-device model: the way to shrink a VM's attack surface is to stop emulating hardware you do not need.
An emulator in a browser tab has a third shape of boundary, and it is a good one. Your processor never executes the guest's instructions, so an exploit that depends on a real CPU errata, a real device or a real MMU has nothing to work with. Reaching your machine would take a bug in the emulator that produced a memory-safety violation in WebAssembly — where linear memory is bounds-checked and the call stack is not addressable — and then a second bug to escape the browser's process sandbox. A serious chain, though not a promise of invulnerability.
Why TempMV has to be an emulator
A tab cannot be a container: it has no kernel to share namespaces with, and JavaScript cannot call clone(CLONE_NEWPID). A tab cannot be a hypervisor: VMXON is a ring-0 instruction, /dev/kvm is a device node no web page can open, and no browser exposes hardware virtualisation to script — deliberately, since handing a web page ring -1 would be an extraordinary thing to do. That leaves one option. Build the processor in software, compile it to WebAssembly, and let the browser's JIT make the result as fast as it can.
Everything else about the site follows from that choice. The 32-bit ceiling is v86's, not the web's; the BIOS rather than UEFI is v86's firmware. The tenfold speed penalty is the price of having no VT-x, and the total absence of a server is the reward: nothing to run means nothing to trust. What is a temporary virtual machine takes that trade apart, the glossary defines the terms, and /new will boot one so you can judge the speed yourself.