Benchmarks
Where every number on this site comes from, and how to reproduce it. All are medians of repeated runs, timed from the host, on one machine.
Against the baselines
Median of 10 runs after one warm-up. The Python workload is sum(i * i for i in range(100_000)).
| Path | Isolation | No-op | Python workload |
|---|---|---|---|
| host process | none | 0.8 ms | 19 ms |
| podman + runc | container | 450 ms | 514 ms |
| podman + krun | microVM | 1,618 ms | 1,821 ms |
| box run, cold (podman) | microVM | 1,411 ms | 1,514 ms |
| box run, warm pool | microVM, fresh per run | 41 ms | 149 ms |
| box exec | microVM, persistent | 15 ms | 62 ms |
| pydantic-monty, new session | interpreter subprocess | 0.5 ms | 12 ms |
Monty is Pydantic's sandboxed Python interpreter, listed for scale: it runs a subset of Python in microseconds, with no VM and no Linux programs. A warm-pool run is measured while the pool refills in the background; with the refill finished it measures 20 to 30 ms.
Backends, side by side
Same alpine image, 1 vCPU, 512 MiB, network: none, median of 7. podman and krun were measured with networking for up, which they need for their agent.
| Backend | Fresh-VM run | up | VMM capabilities |
|---|---|---|---|
| podman | 1,282 ms | 1,612 ms | 6 caps |
| krun | 756 ms | 1,032 ms | 6 caps |
| firecracker | 312 ms | 312 ms | 0 caps |
Method
- x86_64 Linux under WSL2 with nested KVM, podman 5.8, crun 1.28, libkrun 1.19, Firecracker 1.17.
- Times are from starting the command on the host to having its exit status: what an agent loop waits for.
- One machine. Absolute numbers will differ on yours; the ratios between rows are what carries over.
Reproduce
python3 bench/bench.py -n 20 security/escape-test.sh ./box strict firecracker