Benchmarks

Where every number on this site comes from, and how to reproduce it. All are medians of repeated runs, timed from the host, on one machine.

podman + kruna microVM booted per command
1,618 ms
box run, podman backendbooted per command
1,411 ms
box run, krun backendbooted per command, no podman
1,148 ms
podman + runcplain container, shares your kernel
450 ms
box run, firecrackera fresh VM from a snapshot, offline
312 ms
box run, warm poola fresh VM waiting in the pool
41 ms
box execa command in a running VM
15 ms

Against the baselines

Median of 10 runs after one warm-up. The Python workload is sum(i * i for i in range(100_000)).

PathIsolationNo-opPython workload
host processnone0.8 ms19 ms
podman + runccontainer450 ms514 ms
podman + krunmicroVM1,618 ms1,821 ms
box run, cold (podman)microVM1,411 ms1,514 ms
box run, warm poolmicroVM, fresh per run41 ms149 ms
box execmicroVM, persistent15 ms62 ms
pydantic-monty, new sessioninterpreter subprocess0.5 ms12 ms

Monty is Pydantic's sandboxed Python interpreter, listed for scale: it runs a subset of Python in microseconds, with no VM and no Linux programs. A warm-pool run is measured while the pool refills in the background; with the refill finished it measures 20 to 30 ms.

Backends, side by side

Same alpine image, 1 vCPU, 512 MiB, network: none, median of 7. podman and krun were measured with networking for up, which they need for their agent.

BackendFresh-VM runupVMM capabilities
podman1,282 ms1,612 ms6 caps
krun756 ms1,032 ms6 caps
firecracker312 ms312 ms0 caps

Method

  • x86_64 Linux under WSL2 with nested KVM, podman 5.8, crun 1.28, libkrun 1.19, Firecracker 1.17.
  • Times are from starting the command on the host to having its exit status: what an agent loop waits for.
  • One machine. Absolute numbers will differ on yours; the ratios between rows are what carries over.

Reproduce

python3 bench/bench.py -n 20
security/escape-test.sh ./box strict firecracker