A box for your AI.
Let your agent install, delete and break things. It does it in its own virtual machine, and nothing outside it changes.
$ curl -fsSL https://chxperiments.github.io/box/install.sh | sh
Startup times, measured.
Booting a microVM takes over a second. box boots it before the command arrives, so the command waits milliseconds.
One interface, three engines.
Every call goes through box. The Boxfile's backend: picks what boots the microVM; all three give the sandbox its own kernel.
The whole sandbox is one file.
box builds the image, checks that the sandbox really has its own kernel, and refuses to run it if not. Hover over a line to see what it does.
Every run, a new machine.
Branch, review, apply.
Fork a sandbox and its /data branches in about a quarter of a second. An agent can try several approaches side by side, and the original stays untouched.
Diff lists what a trial changed. Nothing reaches your data until you apply it, and you discard the rest. Agents on MCP can fork and diff, but only you can apply.
Sandboxes as a function call.
SDKs for Python, TypeScript, Go and Rust. They talk to a local server over a Unix socket only you can open, and start it when needed. For agents, box mcp serves the same tools over the Model Context Protocol.
from sdbox import Sandbox # a persistent microVM, about 10 ms per command with Sandbox("agent") as sb: sb.write_file("/data/task.py", code) r = sb.exec(["python3", "/data/task.py"]) print(r.stdout_text, r.exit_code) # branch /data, try something, keep it or drop it trial = Sandbox("agent").fork("trial") trial.run("pytest -q").check() print(trial.diff()) trial.apply()
import { Sandbox } from "box-sdk";
// a persistent microVM, about 10 ms per command
await new Sandbox("agent").withUp(async (sb) => {
await sb.writeFile("/data/task.py", code);
const r = await sb.exec(["python3", "/data/task.py"]);
console.log(r.stdoutText, r.exitCode);
});
// branch /data, try something, keep it or drop it
const trial = await new Sandbox("agent").fork("trial");
(await trial.run("pytest -q")).check();
console.log(await trial.diff());
await trial.apply();sb := box.New().Sandbox("agent")
if err := sb.Up(ctx); err != nil {
return err
}
defer sb.Down(ctx)
r, err := sb.Exec(ctx, box.Cmd("python3", "/data/task.py"))
fmt.Println(string(r.Stdout), r.ExitCode)
// branch /data, try something, keep it or drop it
trial, _ := sb.Fork(ctx, "trial")
trial.Run(ctx, box.Sh("pytest -q"))
changes, _ := trial.Diff(ctx)
trial.Apply(ctx)use sdbox::{Command, Sandbox};
let sb = Sandbox::new("agent")?;
sb.up()?;
let out = sb.exec(Command::new(["python3", "/data/task.py"]))?;
println!("{} {}", out.stdout_text(), out.exit_code);
sb.down()?;
// branch /data, try something, keep it or drop it
let trial = sb.fork("trial")?;
trial.run("pytest -q")?.check()?;
println!("{:?}", trial.diff()?);
trial.apply()?;
Not sure it fits? Ask your AI.
This prompt tells an assistant what box does, what it costs, where it stops and how to set it up. Paste it anywhere, describe what you are building, and it will tell you plainly whether box is the right tool. If it is not, the prompt asks it to say what is.
- Copy the prompt
- Paste it into your AI
- Answer its questions
I'm evaluating box (https://github.com/chxperiments/box) for something I'm building. Help me decide whether it fits, and if it does, how to set it up correctly. Use only the facts below about box; if something isn't covered here, say so instead of guessing.
## What box is
An open-source CLI (Go, MIT) that runs code in disposable microVMs on your own machine. Each sandbox is a real VM under KVM with its own kernel, not a container sharing the host kernel. Its purpose is to give AI agents a place to run code where a mistake, or an agent hijacked by prompt injection, cannot break anything outside the sandbox.
## How it works
- A sandbox is defined by one YAML file, the Boxfile. `box build` turns it into an image and checks that the guest kernel differs from the host's; if not, it refuses the sandbox.
- `box run <name> -- <cmd>` boots a fresh VM per command. Nothing carries over between runs except `/data`, a directory persisted on the host (~/.box/data/<name>/).
- `box up` / `exec` / `down` keeps one VM running so commands are fast and state (files, packages, processes) lasts until `down`.
- `warm: N` keeps N VMs pre-booted, so a fresh-VM `run` starts in milliseconds. Each pooled VM serves exactly one run.
- Forks: `box fork <name> <fork>` branches `/data` as an overlay. `diff` shows what changed, then `apply` merges it into the parent or `discard` throws it away. This is how an agent's changes get reviewed before they land.
- `box data import|export <name> <dir>` copies files into or out of `/data`. Exported files are treated as hostile (symlinks never followed, setuid dropped).
## Backends (`backend:` in the Boxfile)
- podman (default): podman with the krun runtime (libkrun). Supports everything. Cold run ~1.3–1.5 s.
- krun: drives crun's libkrun handler directly. Cold run ~0.75–1.1 s. `isolation: standard` only for now.
- firecracker: restores a snapshot per run, ~0.3 s, no pool needed. x86_64 Linux only; no networking, no host mounts, no forks yet.
Measured: a warm `run` is ~20–45 ms; `exec` in a running sandbox is ~15 ms for a no-op.
## Boxfile reference (unknown keys are rejected)
base (image ref), backend, cpus (max 16), ram_mib, network (bridge = internet, none = offline), readonly (read-only root; /tmp and /data stay writable), isolation (standard | strict), timeout_seconds (per run; a timed-out run exits 124), warm (0–8, needs network: bridge), warmup (commands run in each pooled VM first), packages (apk, apt or dnf, chosen from the base), run (build steps), env, mounts (host dir → guest path, ro by default; strict allows ro only), blueprint (users, write_files, runcmd, applied at build time).
## Driving it
- CLI: new [--from example], build, run, shell, up, exec, down, fork, diff, apply, discard, snapshot, restore, reset, data import/export, doctor (checks the host and prints fixes), verify, logs, destroy.
- SDKs for Python, TypeScript, Go and Rust talk to `box serve`, a local API on an owner-only Unix socket. Python example: `with Sandbox("devbox") as sb: sb.exec(["python3", "-c", "print(1)"])`. They expose run, up, exec, down, read_file, write_file, fork, diff, apply, discard.
- MCP: `claude mcp add box -- box mcp` gives an agent list_sandboxes, run, up, exec, down, read_file, write_file, fork, diff, discard. `apply` is withheld unless started with `--allow-apply`, so a human reviews forks. Creating and building sandboxes stays on the CLI.
## Security model
Two layers: the guest runs on its own kernel under KVM, and the VMM process on the host is itself confined (no_new_privs, seccomp, its own network namespace, pids and memory limits, a reduced capability set; Firecracker's VMM runs with zero capabilities). With `isolation: strict` the VMM runs as a subordinate UID instead of yours, so escaping both the VM and the VMM does not land in your account. Use strict for agents. An escape test suite ships in the repo.
Not covered: all strict sandboxes share one subordinate UID, so strict separates sandboxes from you, not from each other. The kernel check catches a fallback to a plain container but cannot catch a guest that lies about its kernel. It is not designed for hostile multi-tenant hosting.
## Requirements and limits
Linux with KVM (/dev/kvm), podman and libkrun (crun built with +LIBKRUN). macOS works through a podman machine, but up/exec, the warm pool, forks and the krun and firecracker backends are Linux-only. It runs on your machine or your server: it is not a hosted service and has no GPU support. `box shell` has a TTY; `exec` does not yet.
## What I want from you
1. Ask me what I'm building: what the agent (or other code) will do and what it must never be able to touch, how many runs and how fast they need to be, whether it needs the network, files from the host, a GPU, or my OS, and where it will run.
2. Then give a clear verdict: good fit, partial fit, or not a fit. Explain why, using the facts above.
3. If it fits, write the Boxfile I should start with (isolation: strict for agents), the commands to get running, and the SDK or MCP setup for my case. Mention the limits that will affect me.
4. If it doesn't fit, say what would serve me better and why.
Questions, answered.
Still unsure? Ask your AI with the prompt above.
01How is this different from running code in a container?
A container shares your host's kernel, so a kernel bug or a careless mount reaches your machine. box runs each sandbox as a microVM under KVM with its own kernel, and refuses to run one whose kernel matches the host's, which would mean it quietly became a plain container.
02Do I need root?
Only once, to install podman and libkrun and create the krun symlink. After that everything runs as you, with rootless podman. Your user needs write access to /dev/kvm (usually the kvm group), and isolation: strict needs a range in /etc/subuid. box doctor checks all of it and prints the fix.
03Does it work on macOS or Windows?
Linux with KVM is the main target. On macOS it runs through a podman machine, but up/exec, the warm pool, forks and the krun and firecracker backends are Linux-only. On Windows, use WSL2 with /dev/kvm available: the published benchmarks were measured that way.
04How fast is it?
A cold run takes about 1.3 s on podman, 0.75 s on krun and 0.3 s on firecracker. With warm: set, a fresh-VM run takes 20 to 45 ms, and exec into a running sandbox about 15 ms. Method and numbers.
05What is kept between runs?
Only /data, a directory on your host at ~/.box/data/<name>/. Every run is a new VM, so installed packages, files elsewhere and processes are gone. up keeps one VM alive when you want state to last across commands.
06Can a sandbox reach the internet or my machine?
With network: bridge it reaches the internet; with network: none it has no network at all. Either way it has its own network namespace and cannot reach services on your host's loopback. It sees only the host directories its Boxfile declares, read-only unless you say rw.
07Can my AI agent break anything?
Not outside its sandbox. Inside, it can do anything, even rm -rf /, and the next run starts clean. Only /data and the host folders you declare are kept, and forks let you review its changes before they reach your data. Use isolation: strict: the VMM around the agent is confined and runs as a UID that is not yours. box is built for your own agents, not for hosting strangers' code, and strict sandboxes share one UID between them. The threat model covers what it does not protect.
08Which backend should I use?
Start with podman, the default: it supports everything. krun boots faster but supports isolation: standard only, for now. firecracker is fastest to start but is x86_64-only and has no network, host mounts or forks yet.
09How do I let an agent use it?
Add the MCP server: claude mcp add box -- box mcp. The agent can run, exec, read and write files, and fork, but cannot apply a fork to real data unless you start the server with --allow-apply. Creating and building sandboxes stays with you.
10Is it a hosted service? What does it cost?
Neither. box runs on your laptop or your own server, MIT-licensed and free. Nothing is sent anywhere. It also has no GPU support, so GPU workloads need something else.
An escape suite, in the repo.
It runs what an agent gone wrong would try from inside a sandbox: reaching services on your loopback, reading host files, escaping through symlinks and ../ paths, writing through read-only mounts, reading the agent's token, a fork bomb. Then it inspects the VMM from outside: capabilities, privileges, seccomp, namespaces, limits.
It passes on every backend in both isolation modes. Run it yourself before you let an agent loose.