Architecture
One interface, three ways to boot a microVM. Everything above the launcher is the same on all three; what is below it depends on backend: in the Boxfile.
podman and krun: libkrun
libkrun is a library that runs a process as a microVM. crun's libkrun handler uses it to boot an OCI image with its own kernel; box verifies that at build by comparing the guest's kernel with the one a plain container sees, and again before every command.
libkrun's own documentation says the guest and its VMM share one security context: the VMM does the guest's file I/O and opens its network connections. So box treats the VMM's confinement as the boundary behind a libkrun escape, and keeps only the six capabilities virtiofs and low ports need.
The krun backend writes the OCI spec itself and drives crun directly, removing podman's share of a cold boot (about 0.85 s of 1.4 s), with the same confinement.
firecracker: a snapshot per run
Firecracker is the VMM AWS built to run other people's code in Lambda and Fargate. It emulates a block device, vsock and very little else, and it does not open the guest's network connections on the host.
box boots each image once, waits for its agent, and snapshots the VM. Every run restores that snapshot instead of booting a kernel. Before anything runs, the copy is made distinct: the agent token baked into the snapshot is rotated, the guest mixes in fresh randomness and takes the host clock, and only then is /data mounted.
The VMM runs in a crun container with no capabilities at all and a read-only root holding only its binary, kernel, image, VM directory, disk and /dev/kvm: what Firecracker's jailer provides, without the root it needs.
A run, from the pool
- ClaimA waiting VM is taken by renaming its state file; two runs can never get the same one.
- CheckThe guest reports its kernel; a VM that answers with the host's is refused.
- ExecuteThe command runs through the agent, with stdout, stderr, exit code and timeout passed through.
- DiscardThe VM is destroyed. A detached tender boots its replacement in the background.
A run, from a snapshot
- Start the VMMFirecracker starts in its jail container, with the image's memory mapped copy-on-write.
- RestoreThe paused VM resumes where its agent was waiting, in tens of milliseconds.
- Make it distinctRotate the token, reseed randomness, set the clock, mount
/data. - Execute and flushRun the command, write
/dataout, then tear the VM down.
Data and forks
Everything except /data resets after each run, so a fork only has to branch one directory. That is why it is fast.
| Operation | What happens | Cost |
|---|---|---|
| fork | A new sandbox whose /data is an overlay on the parent's, mounted inside podman's user namespace without root. | ~0.25 s |
| diff | Read straight from the fork's upper layer, whiteouts and opaque directories included. No VM boots. | ms |
| apply | Merges into the parent through an os.Root, so a symlink the guest planted cannot redirect a write. | ms |
| data export | Copies /data to a host directory. Symlinks stay links, devices are skipped, setuid bits dropped. | size-bound |