A coding agent running in my normal shell can turn one bad path into a colossal shit show.

Most of the time it behaves. That is not a security model.

My new friend Ron and I were talking about this on Monday. We both have been poking at the same problem from different directions: these agents are useful enough that I want them to run real shell commands and keep going after I leave them alone. I do not want the cheap model I am testing to discover my root folder as a creative writing exercise.

The obvious box was too much work

Firecracker came up first. It is the thing everyone mentions when the words “small VM” show up. It is also the wrong shape for my laptop. Firecracker uses KVM. KVM means Linux. On a Mac that means running a Linux VM so I can run a tiny Linux VM so an agent can run Linux.

That is a hat on a hat.

There is also a simpler, more boring shape: One shared VM, mount a folder, maybe use git as the handoff between the Mac and the guest... It could work. Having to use heavy apps like Parallels, UTM and go through their interface to build/maintain, or ssh after novelty wears off felt 'meh'

Docker was not the answer I wanted

I run Docker on my home server. No moral objection. I do not love Docker Desktop on a Mac. It always feels like I invited one hungry monster insatiable for memory and space.

OrbStack was better. It felt a little like Proxmox with a nicer Mac face. Fine product. Then I spun something up and it had access to my user folder. There is a file sharing issue about that behavior, which at least made me feel less insane.

That default defeats the point.

The whole reason I want a sandbox is that the agent should not see the rest of my computer. It can destroy the repo I gave it. Bad, but recoverable. It should not be able to rummage through dotfiles, local config, random archives, or whatever embarrassing folder I forgot existed.

A Mac laptop exposing one selected folder to an isolated translucent sandbox box.
The boundary I wanted: one allowed folder, not the whole laptop.

I like trying dumber models. That makes this less theoretical. The frontier models are usually careful. The smaller ones are fun because they are weird, and weird is not what I want near ~/.

So I went back to Apple containers

I had dismissed Apple’s container CLI earlier because it is new. New Apple APIs have a wet-cement feeling. Shit moves and changes because they decided it is a good idea. Some framework gets blessed at WWDC and then quietly becomes archaeology.

Still, when I am on a Mac, the Mac-native thing usually wins. Apple’s container tool is written in Swift, optimized for Apple silicon, and runs Linux containers as lightweight VMs. There is also a Containerization Swift package underneath it.

For a first version, the CLI was enough. I did not need to marry the API (I did have an entire step to make sure it satisfied the current use cases though). I needed a narrow place where I could swap it out later when Apple inevitably moves the furniture.

Matt Pocock accidentally became part of this

I have been using Matt Pocock’s skills repo lately. I love/hate grill-with-docs. It does the annoying thing a good collaborator does: it keeps asking what you mean until the words stop being mush.

The flow after that is to-prd-> to-issues-> tdd. The important part was not the ceremony but it was having to name the thing before building the thing.

That sounds small, believe me it was not. I spent more time than usual defining what a Sandbox VM is, what an Allowed Folder is, what Guest State means, and what parts of the system are allowed to know about Apple’s backend. By the time code started, the mental model had fewer loose wires.

Pocock is also big on vertical slices, which comes out of the usual software engineering canon, including The Pragmatic Programmer. Build the smallest real thing end to end. Do not build five fake layers and then pray they line up later.

I still catch agents cheating here. They say vertical slice, then sneak in fake adapters because the test becomes easier. That is where reading the issues before letting the agent loose matters, otherwise the fucking horizontal slice walks in wearing a fake mustache.

The Ousterhout shit

The thing that I am thinking about most lately is deep modules with shallow interfaces, from John Ousterhout’s A Philosophy of Software Design.

A deep module has a small surface and a lot of work hidden behind it. You ask it for one clear thing. It deals with the mess. A shallow module is the opposite. It looks harmless until every caller needs to know its internal laundry schedule.

This matters with agents because context is not only the budget, but also the smarts (The 'Smart Zone' stuff). An agent can hold a small interface in its head. It cannot keep rereading half the codebase every time it wants to run a command. If the module is deep, the agent can use it without knowing every ugly detail inside it. If the module is shallow, the ugliness leaks into every issue.

In Sand, Apple’s container command sits behind a SandboxBackend. The rest of the app talks in Sand language: create a Sandbox VM, add an Allowed Folder, run a Workload Command, open a Sandbox Session, delete the VM. The Apple-specific flags live in the adapter.

Today that adapter shells out to the Apple CLI. Later it can call the Swift package directly. Or something else. The rest of Sand should not care.

That is the bet, anyway. I have made enough leaky wrappers in my life to recognize the smell earlier now.

A small clean interface panel hiding a larger messy mechanical backend.
Small interface in front. Messy backend behind it. That is the module shape I wanted.

What the fuck is Sand?

Sand is a Swift CLI for Apple silicon Macs. It creates small named Linux environments. Each one has its own tools, shell state, package cache, and Pi login. The Host Mac only exposes folders you explicitly allow.

sand create box
sand folders add box ~/Projects/my-project rw --as /workspace
sand box run pi
sand box shell

If the sandbox is named box, the command is sand box run pi. That was not the reason for the name, but I am not going to pretend I hate it.

There is a YAML spec underneath if you want to be precise about the image, CPU, memory, and folders. The daily loop is supposed to stay dumb. Create a box. Give it one project folder. Run the agent there.

The default image is not raw Ubuntu with a blinking cursor and no tools. It is a developer-ready image with the boring stuff already installed, including Pi. I do not want every sandbox to begin with twenty minutes of apt-get archaeology.

The project page is here: onuruzunismail.com/projects/sand/. The repo is public too: github.com/onorbumbum/sand. The docs are here: onorbumbum.github.io/sand. I left the issues in there too, so the build trail is visible instead of cleaned up into a fake genius narrative.

The security caveat

I tried to get Claude and GPT to help me break out of a Sand box. Red-team framing. “We are testing this, please escape the container and read host files.” That sort of thing.

They refused. Good for them.

That does not prove Sand is safe. It proves I failed to talk two frontier models into helping me do something obviously sketchy. A weaker model may not refuse. A human who knows container escapes is not waiting for Claude’s blessing.

The isolation is Apple’s job. I did not write a hypervisor over the weekend. Sand’s job is to make the boundary easy to use: the Sandbox Guest gets its own little Linux world, and the Host Mac grants access folder by folder.

So yes, an agent can fuck up the project folder I gave it. That is still a problem. It is also a much smaller problem than an agent with my whole laptop in reach.

projects, ai, tools, programming