Architecture

One interface, three ways to boot a microVM. Everything above the launcher is the same on all three; what is below it depends on backend: in the Boxfile.

A command flows from the box interface down each backend, through its launcher, VMM and host confinement, across KVM, into the guest. box: CLI, SDKs, local APIlauncherpodman, cruncrun, our speccrun jailVMMlibkrunlibkrunFirecrackerconfinement6 caps, seccomp6 caps, seccomp0 caps, seccompguestkernel 6.12kernel 6.12kernel 6.1KVM: own kernel per sandboxhypervisorpodmankrunfirecracker
podmanthe default
krunlibkrun, no podman
firecrackera snapshot per run
interface
box CLI, SDKs, MCP and local APIBoxfile, warm pool, up / exec, forks, data import / export, isolation checks
launcher
podmanbuilds images, runs crun
crunan OCI spec box writes
cruna jail container for the VMM
VMM
libkruninside crun's process
libkruninside crun's process
Firecracker 1.17block, vsock, little else
confinement
6 capabilitiesseccomp, no_new_privs, own netns, limits
6 capabilitiesseccomp, no_new_privs, own netns, limits
0 capabilitiesseccomp, no_new_privs, read-only jail root
hypervisor
KVMhardware virtualization: every guest runs its own kernel
guest
kernel 6.12libkrunfw, init.krun
kernel 6.12libkrunfw, init.krun
kernel 6.1box as PID 1, RAM overlay
agent channel
TCP on 127.0.0.1token-authenticated
TCP via pasta127.0.0.1, token-authenticated
vsockowner-only socket, token rotated per copy
/data
host directoryshared by virtiofs
host directoryshared by virtiofs
ext4 diskone VM mounts it at a time
security boundaryinside the guesthost side

podman and krun: libkrun

libkrun is a library that runs a process as a microVM. crun's libkrun handler uses it to boot an OCI image with its own kernel; box verifies that at build by comparing the guest's kernel with the one a plain container sees, and again before every command.

libkrun's own documentation says the guest and its VMM share one security context: the VMM does the guest's file I/O and opens its network connections. So box treats the VMM's confinement as the boundary behind a libkrun escape, and keeps only the six capabilities virtiofs and low ports need.

The krun backend writes the OCI spec itself and drives crun directly, removing podman's share of a cold boot (about 0.85 s of 1.4 s), with the same confinement.

firecracker: a snapshot per run

Firecracker is the VMM AWS built to run other people's code in Lambda and Fargate. It emulates a block device, vsock and very little else, and it does not open the guest's network connections on the host.

box boots each image once, waits for its agent, and snapshots the VM. Every run restores that snapshot instead of booting a kernel. Before anything runs, the copy is made distinct: the agent token baked into the snapshot is rotated, the guest mixes in fresh randomness and takes the host clock, and only then is /data mounted.

The VMM runs in a crun container with no capabilities at all and a read-only root holding only its binary, kernel, image, VM directory, disk and /dev/kvm: what Firecracker's jailer provides, without the root it needs.

A run, from the pool

  1. ClaimA waiting VM is taken by renaming its state file; two runs can never get the same one.
  2. CheckThe guest reports its kernel; a VM that answers with the host's is refused.
  3. ExecuteThe command runs through the agent, with stdout, stderr, exit code and timeout passed through.
  4. DiscardThe VM is destroyed. A detached tender boots its replacement in the background.

A run, from a snapshot

  1. Start the VMMFirecracker starts in its jail container, with the image's memory mapped copy-on-write.
  2. RestoreThe paused VM resumes where its agent was waiting, in tens of milliseconds.
  3. Make it distinctRotate the token, reseed randomness, set the clock, mount /data.
  4. Execute and flushRun the command, write /data out, then tear the VM down.

Data and forks

Everything except /data resets after each run, so a fork only has to branch one directory. That is why it is fast.

A sandbox forks into three trials. Two are discarded; one is applied back. agentforktrial-adiscardedtrial-btrial-cdiscardedapply/data
OperationWhat happensCost
forkA new sandbox whose /data is an overlay on the parent's, mounted inside podman's user namespace without root.~0.25 s
diffRead straight from the fork's upper layer, whiteouts and opaque directories included. No VM boots.ms
applyMerges into the parent through an os.Root, so a symlink the guest planted cannot redirect a write.ms
data exportCopies /data to a host directory. Symlinks stay links, devices are skipped, setuid bits dropped.size-bound