The Problem Preconfiguration Solves: Coding Agents Start Every Task on a Machine Nobody Set Up
AI coding agents start each task on a blank machine. When that machine isn’t set up, they guess, fail or hand back untested code, and every vendor wants the setup written in its own file.
Agents Are Poor at Doing It Alone
If agents could reliably set up a repository by themselves, the problem would solve itself. The benchmarks say they can’t, yet:

In EnvBench, a 2025 benchmark from JetBrains Research, the best approach set up 6.69% of 329 Python repositories and 29.47% of 665 JVM repositories. In ResearchEnvBench, a 2026 benchmark of 44 research repositories, the best agent got a basic CPU run working in 17 of the 29 that have one (58.6%), and GPU runs went worse: at best 48.8% on one GPU and 37.5% on several. These are hard repositories chosen to test setup, so the numbers are not a real team’s odds. They do show that “the agent will figure it out” is a gamble.
The Vendors Say So Themselves
Cursor calls environment setup “the most important step” to improve the effectiveness of its cloud agents. GitHub warns that an agent left to find and install dependencies by trial and error is slow and unreliable, and sometimes can’t get them at all, private ones for example. And GitHub’s documentation is plain about what happens when a setup step fails: Copilot skips the remaining steps and begins working with the machine as it is.
So a broken setup doesn’t stop the agent. It starts the agent on a half-ready machine, and the failure shows up later, as tests that can’t run or code that was never tested.
One Setup, Written Several Times
Teams rarely use one agent. In The Pragmatic Engineer’s 2026 survey, 70% of engineers used two to four AI tools and 15% used five or more. Each agent platform reads its setup from its own file:
| Platform | Its setup file | Its format |
|---|---|---|
| GitHub Copilot’s cloud agent | .github/workflows/copilot-setup-steps.yml |
A GitHub Actions workflow with one job of an exact name, six settings honored and a 59-minute limit |
| Cursor’s cloud agents | .cursor/environment.json, and usually a Dockerfile |
JSON with comments but no trailing commas, and paths relative to the .cursor folder |
| Codespaces and editors | .devcontainer/devcontainer.json |
The dev container standard: images, features, compose files |
| Codex, Claude Code on the web, Jules | A setup script in each one’s environment settings | A shell script, kept in the platform rather than the repository |
| A fresh cloud server | cloud-init | YAML whose first line must be exactly #cloud-config |
So the same setup gets written three or four times, by different people at different times, from different examples.
What Breaks Today
| What breaks | Evidence |
|---|---|
| Setup copied by hand between files | A Microsoft project’s docs tell users to check, for each tool, whether the dev container and the Copilot workflow both need it, and to mirror the installation |
| One script wrapped many times | One repository calls the same setup script from Copilot’s workflow, a Claude Code start hook and devcontainer.json |
| A wrong job name | Copilot’s agent stops with an error when the workflow has no job named exactly copilot-setup-steps |
| No test before merge | Copilot uses the setup file only once it is on the main branch, so a pull request can’t prove it works unless the workflow also runs on its own |
| Formats keep moving | Cursor’s install command was previously called update; a setup copied from an older example uses a key Cursor’s schema doesn’t allow |
Each of these is small. Together they mean that a repository’s agent setup is a set of files nobody owns or tests, which break where nobody is watching.
What It Costs
Agent time is metered now. GitHub moved Copilot to usage-based billing on June 1, 2026, and a Copilot cloud agent session uses GitHub Actions minutes and AI credits. Setup runs inside that paid session, so every minute of failed or repeated setup is paid for, and a session that works on a half-ready machine pays for work that has to be done again.
How much that adds up to hasn’t been measured yet, for any team. Measuring it, minutes of setup per session and sessions lost to setup, with and without a spec, on real repositories, is one of the first jobs of the Beta.
Who Feels It Most
Platform and developer-experience teams who look after many repositories, and any team running two or more agents. For them the problem multiplies: every repository, times every agent platform, times every change to a runtime or a database.
Out of scope, on purpose: the agent’s instructions, which AGENTS.md and rules files cover; the quality of the code it writes; and sandbox security while it runs. Preconfiguration handles one thing, the machine the agent works on, and proves it.
What Preconfiguration Does About It
One spec, preconfig.yaml, says what the machine needs. preconfig writes each platform’s file from it, checks the files a repository already has against each platform’s rules, and runs the whole setup on a clean machine with the project’s own tests. The live demo shows all of it on a sample service, and how it works goes into the details.