ai agents

The Problem Preconfiguration Solves: Coding Agents Start Every Task on a Machine Nobody Set Up

AI coding agents start each task on a blank machine. When that machine isn’t set up, they guess, fail or hand back untested code, and every vendor wants the setup written in its own file.

Agents Are Poor at Doing It Alone

If agents could reliably set up a repository by themselves, the problem would solve itself. The benchmarks say they can’t, yet:

How often the best approach in each benchmark set up a repository on its own: 6.69% of 329 Python repositories and 29.47% of 665 JVM repositories in EnvBench, from JetBrains Research (2025), and a basic CPU run in 17 of the 29 research repositories that have one (58.6%) in ResearchEnvBench (2026).

In EnvBench, a 2025 benchmark from JetBrains Research, the best approach set up 6.69% of 329 Python repositories and 29.47% of 665 JVM repositories. In ResearchEnvBench, a 2026 benchmark of 44 research repositories, the best agent got a basic CPU run working in 17 of the 29 that have one (58.6%), and GPU runs went worse: at best 48.8% on one GPU and 37.5% on several. These are hard repositories chosen to test setup, so the numbers are not a real team’s odds. They do show that “the agent will figure it out” is a gamble.

The Vendors Say So Themselves

Cursor calls environment setup “the most important step” to improve the effectiveness of its cloud agents. GitHub warns that an agent left to find and install dependencies by trial and error is slow and unreliable, and sometimes can’t get them at all, private ones for example. And GitHub’s documentation is plain about what happens when a setup step fails: Copilot skips the remaining steps and begins working with the machine as it is.

So a broken setup doesn’t stop the agent. It starts the agent on a half-ready machine, and the failure shows up later, as tests that can’t run or code that was never tested.

One Setup, Written Several Times

Teams rarely use one agent. In The Pragmatic Engineer’s 2026 survey, 70% of engineers used two to four AI tools and 15% used five or more. Each agent platform reads its setup from its own file:

Platform Its setup file Its format
GitHub Copilot’s cloud agent .github/workflows/copilot-setup-steps.yml A GitHub Actions workflow with one job of an exact name, six settings honored and a 59-minute limit
Cursor’s cloud agents .cursor/environment.json, and usually a Dockerfile JSON with comments but no trailing commas, and paths relative to the .cursor folder
Codespaces and editors .devcontainer/devcontainer.json The dev container standard: images, features, compose files
Codex, Claude Code on the web, Jules A setup script in each one’s environment settings A shell script, kept in the platform rather than the repository
A fresh cloud server cloud-init YAML whose first line must be exactly #cloud-config

So the same setup gets written three or four times, by different people at different times, from different examples.

What Breaks Today

What breaks Evidence
Setup copied by hand between files A Microsoft project’s docs tell users to check, for each tool, whether the dev container and the Copilot workflow both need it, and to mirror the installation
One script wrapped many times One repository calls the same setup script from Copilot’s workflow, a Claude Code start hook and devcontainer.json
A wrong job name Copilot’s agent stops with an error when the workflow has no job named exactly copilot-setup-steps
No test before merge Copilot uses the setup file only once it is on the main branch, so a pull request can’t prove it works unless the workflow also runs on its own
Formats keep moving Cursor’s install command was previously called update; a setup copied from an older example uses a key Cursor’s schema doesn’t allow

Each of these is small. Together they mean that a repository’s agent setup is a set of files nobody owns or tests, which break where nobody is watching.

What It Costs

Agent time is metered now. GitHub moved Copilot to usage-based billing on June 1, 2026, and a Copilot cloud agent session uses GitHub Actions minutes and AI credits. Setup runs inside that paid session, so every minute of failed or repeated setup is paid for, and a session that works on a half-ready machine pays for work that has to be done again.

How much that adds up to hasn’t been measured yet, for any team. Measuring it, minutes of setup per session and sessions lost to setup, with and without a spec, on real repositories, is one of the first jobs of the Beta.

Who Feels It Most

Platform and developer-experience teams who look after many repositories, and any team running two or more agents. For them the problem multiplies: every repository, times every agent platform, times every change to a runtime or a database.

Out of scope, on purpose: the agent’s instructions, which AGENTS.md and rules files cover; the quality of the code it writes; and sandbox security while it runs. Preconfiguration handles one thing, the machine the agent works on, and proves it.

What Preconfiguration Does About It

One spec, preconfig.yaml, says what the machine needs. preconfig writes each platform’s file from it, checks the files a repository already has against each platform’s rules, and runs the whole setup on a clean machine with the project’s own tests. The live demo shows all of it on a sample service, and how it works goes into the details.