Nested boxes: your machine, the confined VMM, KVM and the guest kernel. A command travels inward and runs in the guest. your machineVMM, confinedKVMguest kernel
Nested boxes: your machine, the confined VMM, KVM and the guest kernel. A command travels inward and runs in the guest.

A box for your AI.

Let your agent install, delete and break things. It does it in its own virtual machine, and nothing outside it changes.

$ curl -fsSL https://chxperiments.github.io/box/install.sh | sh

Startup times, measured.

Booting a microVM takes over a second. box boots it before the command arrives, so the command waits milliseconds.

15msa command in a running VM
41msa fresh VM from the warm pool
0.3sa fresh VM restored from a snapshot
0escapes found by the escape suite, on any backend
podman + kruna microVM booted per command
1,618 ms
box run, podman backendbooted per command
1,411 ms
box run, krun backendbooted per command, no podman
1,148 ms
podman + runcplain container, shares your kernel
450 ms
box run, firecrackera fresh VM from a snapshot, offline
312 ms
box run, warm poola fresh VM waiting in the pool
41 ms
box execa command in a running VM
15 ms

One interface, three engines.

Every call goes through box. The Boxfile's backend: picks what boots the microVM; all three give the sandbox its own kernel.

Calls from the CLI, the SDKs and MCP go to box, which reads the Boxfile's backend and boots the sandbox with podman, krun or Firecracker, each a microVM with its own kernel. CLIbox run, execSDKsPython, TS, Go, RustMCPfor AI agentspodmanlibkrun via podman1.28 skrunlibkrun, direct0.76 sfirecrackersnapshot restore0.31 sown kernel

The whole sandbox is one file.

box builds the image, checks that the sandbox really has its own kernel, and refuses to run it if not. Hover over a line to see what it does.

# agent: everything this sandbox isbase: docker.io/library/alpine:latestbackend: firecrackerisolation: strictcpus: 2ram_mib: 1024network: nonereadonly: truetimeout_seconds: 60packages: [python3, py3-pip]env: {PYTHONUNBUFFERED: "1"}

Every run, a new machine.

Branch, review, apply.

Fork a sandbox and its /data branches in about a quarter of a second. An agent can try several approaches side by side, and the original stays untouched.

Diff lists what a trial changed. Nothing reaches your data until you apply it, and you discard the rest. Agents on MCP can fork and diff, but only you can apply.

Forks in the docs

A sandbox forks into three trials. Two are discarded; one is applied back. agentforktrial-adiscardedtrial-btrial-cdiscardedapply/data

Sandboxes as a function call.

SDKs for Python, TypeScript, Go and Rust. They talk to a local server over a Unix socket only you can open, and start it when needed. For agents, box mcp serves the same tools over the Model Context Protocol.

SDK reference

from sdbox import Sandbox

# a persistent microVM, about 10 ms per command
with Sandbox("agent") as sb:
    sb.write_file("/data/task.py", code)
    r = sb.exec(["python3", "/data/task.py"])
    print(r.stdout_text, r.exit_code)

# branch /data, try something, keep it or drop it
trial = Sandbox("agent").fork("trial")
trial.run("pytest -q").check()
print(trial.diff())
trial.apply()

Not sure it fits? Ask your AI.

This prompt tells an assistant what box does, what it costs, where it stops and how to set it up. Paste it anywhere, describe what you are building, and it will tell you plainly whether box is the right tool. If it is not, the prompt asks it to say what is.

  1. Copy the prompt
  2. Paste it into your AI
  3. Answer its questions
you5.1 KB · plain text
I'm evaluating box (https://github.com/chxperiments/box) for something I'm building. Help me decide whether it fits, and if it does, how to set it up correctly. Use only the facts below about box; if something isn't covered here, say so instead of guessing.

## What box is
An open-source CLI (Go, MIT) that runs code in disposable microVMs on your own machine. Each sandbox is a real VM under KVM with its own kernel, not a container sharing the host kernel. Its purpose is to give AI agents a place to run code where a mistake, or an agent hijacked by prompt injection, cannot break anything outside the sandbox.

## How it works
- A sandbox is defined by one YAML file, the Boxfile. `box build` turns it into an image and checks that the guest kernel differs from the host's; if not, it refuses the sandbox.
- `box run <name> -- <cmd>` boots a fresh VM per command. Nothing carries over between runs except `/data`, a directory persisted on the host (~/.box/data/<name>/).
- `box up` / `exec` / `down` keeps one VM running so commands are fast and state (files, packages, processes) lasts until `down`.
- `warm: N` keeps N VMs pre-booted, so a fresh-VM `run` starts in milliseconds. Each pooled VM serves exactly one run.
- Forks: `box fork <name> <fork>` branches `/data` as an overlay. `diff` shows what changed, then `apply` merges it into the parent or `discard` throws it away. This is how an agent's changes get reviewed before they land.
- `box data import|export <name> <dir>` copies files into or out of `/data`. Exported files are treated as hostile (symlinks never followed, setuid dropped).

## Backends (`backend:` in the Boxfile)
- podman (default): podman with the krun runtime (libkrun). Supports everything. Cold run ~1.3–1.5 s.
- krun: drives crun's libkrun handler directly. Cold run ~0.75–1.1 s. `isolation: standard` only for now.
- firecracker: restores a snapshot per run, ~0.3 s, no pool needed. x86_64 Linux only; no networking, no host mounts, no forks yet.
Measured: a warm `run` is ~20–45 ms; `exec` in a running sandbox is ~15 ms for a no-op.

## Boxfile reference (unknown keys are rejected)
base (image ref), backend, cpus (max 16), ram_mib, network (bridge = internet, none = offline), readonly (read-only root; /tmp and /data stay writable), isolation (standard | strict), timeout_seconds (per run; a timed-out run exits 124), warm (0–8, needs network: bridge), warmup (commands run in each pooled VM first), packages (apk, apt or dnf, chosen from the base), run (build steps), env, mounts (host dir → guest path, ro by default; strict allows ro only), blueprint (users, write_files, runcmd, applied at build time).

## Driving it
- CLI: new [--from example], build, run, shell, up, exec, down, fork, diff, apply, discard, snapshot, restore, reset, data import/export, doctor (checks the host and prints fixes), verify, logs, destroy.
- SDKs for Python, TypeScript, Go and Rust talk to `box serve`, a local API on an owner-only Unix socket. Python example: `with Sandbox("devbox") as sb: sb.exec(["python3", "-c", "print(1)"])`. They expose run, up, exec, down, read_file, write_file, fork, diff, apply, discard.
- MCP: `claude mcp add box -- box mcp` gives an agent list_sandboxes, run, up, exec, down, read_file, write_file, fork, diff, discard. `apply` is withheld unless started with `--allow-apply`, so a human reviews forks. Creating and building sandboxes stays on the CLI.

## Security model
Two layers: the guest runs on its own kernel under KVM, and the VMM process on the host is itself confined (no_new_privs, seccomp, its own network namespace, pids and memory limits, a reduced capability set; Firecracker's VMM runs with zero capabilities). With `isolation: strict` the VMM runs as a subordinate UID instead of yours, so escaping both the VM and the VMM does not land in your account. Use strict for agents. An escape test suite ships in the repo.
Not covered: all strict sandboxes share one subordinate UID, so strict separates sandboxes from you, not from each other. The kernel check catches a fallback to a plain container but cannot catch a guest that lies about its kernel. It is not designed for hostile multi-tenant hosting.

## Requirements and limits
Linux with KVM (/dev/kvm), podman and libkrun (crun built with +LIBKRUN). macOS works through a podman machine, but up/exec, the warm pool, forks and the krun and firecracker backends are Linux-only. It runs on your machine or your server: it is not a hosted service and has no GPU support. `box shell` has a TTY; `exec` does not yet.

## What I want from you
1. Ask me what I'm building: what the agent (or other code) will do and what it must never be able to touch, how many runs and how fast they need to be, whether it needs the network, files from the host, a GPU, or my OS, and where it will run.
2. Then give a clear verdict: good fit, partial fit, or not a fit. Explain why, using the facts above.
3. If it fits, write the Boxfile I should start with (isolation: strict for agents), the commands to get running, and the SDK or MCP setup for my case. Mention the limits that will affect me.
4. If it doesn't fit, say what would serve me better and why.
Ask Claude Ask ChatGPT

Questions, answered.

Still unsure? Ask your AI with the prompt above.

01How is this different from running code in a container?

A container shares your host's kernel, so a kernel bug or a careless mount reaches your machine. box runs each sandbox as a microVM under KVM with its own kernel, and refuses to run one whose kernel matches the host's, which would mean it quietly became a plain container.

02Do I need root?

Only once, to install podman and libkrun and create the krun symlink. After that everything runs as you, with rootless podman. Your user needs write access to /dev/kvm (usually the kvm group), and isolation: strict needs a range in /etc/subuid. box doctor checks all of it and prints the fix.

03Does it work on macOS or Windows?

Linux with KVM is the main target. On macOS it runs through a podman machine, but up/exec, the warm pool, forks and the krun and firecracker backends are Linux-only. On Windows, use WSL2 with /dev/kvm available: the published benchmarks were measured that way.

04How fast is it?

A cold run takes about 1.3 s on podman, 0.75 s on krun and 0.3 s on firecracker. With warm: set, a fresh-VM run takes 20 to 45 ms, and exec into a running sandbox about 15 ms. Method and numbers.

05What is kept between runs?

Only /data, a directory on your host at ~/.box/data/<name>/. Every run is a new VM, so installed packages, files elsewhere and processes are gone. up keeps one VM alive when you want state to last across commands.

06Can a sandbox reach the internet or my machine?

With network: bridge it reaches the internet; with network: none it has no network at all. Either way it has its own network namespace and cannot reach services on your host's loopback. It sees only the host directories its Boxfile declares, read-only unless you say rw.

07Can my AI agent break anything?

Not outside its sandbox. Inside, it can do anything, even rm -rf /, and the next run starts clean. Only /data and the host folders you declare are kept, and forks let you review its changes before they reach your data. Use isolation: strict: the VMM around the agent is confined and runs as a UID that is not yours. box is built for your own agents, not for hosting strangers' code, and strict sandboxes share one UID between them. The threat model covers what it does not protect.

08Which backend should I use?

Start with podman, the default: it supports everything. krun boots faster but supports isolation: standard only, for now. firecracker is fastest to start but is x86_64-only and has no network, host mounts or forks yet.

09How do I let an agent use it?

Add the MCP server: claude mcp add box -- box mcp. The agent can run, exec, read and write files, and fork, but cannot apply a fork to real data unless you start the server with --allow-apply. Creating and building sandboxes stays with you.

10Is it a hosted service? What does it cost?

Neither. box runs on your laptop or your own server, MIT-licensed and free. Nothing is sent anywhere. It also has no GPU support, so GPU workloads need something else.

An escape suite, in the repo.

It runs what an agent gone wrong would try from inside a sandbox: reaching services on your loopback, reading host files, escaping through symlinks and ../ paths, writing through read-only mounts, reading the agent's token, a fork bomb. Then it inspects the VMM from outside: capabilities, privileges, seccomp, namespaces, limits.

It passes on every backend in both isolation modes. Run it yourself before you let an agent loose.

Read the threat model