Next Uses: Six Places Where a Ready Machine Pays Off
The four case studies share one pattern. A machine that starts empty has to become a known, working machine before anyone can do useful work on it, and the knowledge of what “working” means lives in one reviewed file, proven by a test. Coding agents came first because they start from an empty machine many times a day, on several platforms, with nobody watching.
The pattern turns up well beyond coding agents. None of the six uses below has been tested, and none is a goal of the Beta, though some would reuse what the Beta builds. This post sets out what each would ask of preconfig, what already carries over from the Alpha, and what would be new work. It is a map for choosing what to try after the Beta, ideally with the people who run these systems.

What They Would Share
Every use below relies on the same parts of the Alpha: one spec a person can review, the targets built from it, check to catch drift, and verify to prove the machine on a clean start. What changes from one use to the next is who writes the spec, which target matters most, and who needs the proof.
A New Developer’s First Day
The first agent on every team was a person: the new hire who spends a day installing things before writing a line. The dev container target is already that person’s setup, with the same services the agents get, and verify is the proof that the README’s instructions still work.
- Carries over: the dev container, the setup script, verify.
- New work: macOS and Windows laptops without containers, where most developers actually work; a
preconfig upthat runs the setup on the developer’s own machine, with care for what is already installed there.
CI Jobs
A CI job starts from an empty runner too, and most repositories maintain a CI setup that repeats what the dev container and the agents’ files say. The Copilot target is already a GitHub Actions job; a general CI target would write the same install steps for any workflow that needs them.
- Carries over: the knowledge of GitHub’s setup actions, the services as service containers, the ready check.
- New work: a target for CI workflows beyond Copilot’s, and GitLab CI; caching that matches each provider’s rules.
Evaluation Harnesses
Benchmarks for coding agents, such as SWE-bench, need many repositories set up reliably and the same way on every run, so that a failed task means the agent failed and not the setup. A spec per repository, proven with verify, is exactly the artifact those harnesses maintain by hand.
- Carries over: detect, to draft specs for many repositories at once; verify, to prove each one; the JSON output, for tools.
- New work: runs at scale, in parallel, with results collected across repositories; pinning every version so a benchmark gives the same machine next year.
Preview Environments
Many teams start a short-lived environment for each pull request: the app, its database, sample data, a URL to click. That environment needs the same runtimes and services as everything else, and it breaks for the same reasons.
- Carries over: services, environment variables, the setup and ready steps, cloud-init for a virtual machine.
- New work: a start command for the app itself, sample data, and targets for the platforms that host previews.
Workshops and Classrooms
A workshop loses its first hour to installing. An instructor could publish one spec, and each student’s laptop, Codespace or lab machine would get the same setup, proven the night before.
- Carries over: the dev container and the setup script, and verify for the instructor.
- New work: targets for the lab platforms schools use, and plain-language messages for people who have never seen a stack trace.
Platforms That Run Other People’s Agents
Sandbox companies give agents their own machines on demand. Each customer’s environment has to be described in a form the platform can load, and each platform has its own. A spec those platforms could load directly would bring customers’ environments with them.
- Carries over: the spec, the setup script, the knowledge of what each runtime and service needs.
- New work: targets for each platform’s own format, prebuilt images, and fast starts from them.
What the Engine Would Need
| Use | New targets | New in the engine | Who writes the spec |
|---|---|---|---|
| A new developer’s first day | None | Local setup on macOS and Windows | The team |
| CI jobs | General CI workflows, GitLab CI | Caching per provider | The team |
| Evaluation harnesses | None | Runs at scale; pinned versions | The benchmark’s authors, with detect |
| Preview environments | Preview platforms | App start commands, sample data | The team |
| Workshops and classrooms | Lab platforms | Plainer messages | The instructor |
| Agent platforms | Each platform’s format | Prebuilt images | The customer, loaded by the platform |
How a New Use Would Be Tried
Each would start the way the Alpha did: a handful of real cases, the files built and checked, verify on a clean machine, every number measured and the gaps said out loud. A use goes further only if the people who run those systems want it.
Which Come First
Evaluation harnesses and agent platforms are the closest to the Alpha: they need what it already does, at a larger scale, and they are where the next wave of agents is being built and measured. A new developer’s first day is the widest audience and needs the most new work. The project would like to hear from anyone running one of these, through Contact.