# Make your repo agent-ready: setup script, test command and image A repository is agent-ready when one command sets it up, one command tests it, and a prebuilt image starts every agent there. **Runtime (withruntime.com) turns that into a measurement: the audit below clones your repository into a fresh Firecracker microVM, which has run its first command 221 ms after the create request on Runtime's servers, and grades it on the eight things coding agents trip over.** Most failed agent runs never reach the interesting part of the task. The agent spends its first twenty tool calls guessing how to install the project, answers an interactive prompt it cannot see, or reads 4,000 lines of test output looking for the one failure. Fix the repository once and every agent, Claude Code, Codex, Cursor or your own, starts from the part that matters. ## What does "agent-ready" mean for a repository? It means an agent with no memory of your project can get from a clean machine to a trustworthy pass or fail without a person. Three artifacts get it there, and each one removes a whole class of wasted turns: 1. **A setup script** that installs everything, asks nothing, and is safe to run again. 2. **A test command** whose exit code is the truth and whose output is short when it passes and specific when it fails. 3. **An image** with the setup already done, so a new agent sandbox starts ready instead of spending minutes on installs. A short `AGENTS.md` that names the three ties them together. The rest of this post is how to write each one, and a program that checks them. ## How do you write a setup script an agent can run? Make it non-interactive, idempotent and loud only on failure. An agent runs it with no terminal attached, so anything that waits for input waits until the command times out. It will also run it again after every `git pull`, so the second run must be fast and change nothing. ```bash no-run #!/usr/bin/env bash # scripts/setup.sh: from a clean machine to a working checkout. Safe to rerun. set -euo pipefail export CI=1 DEBIAN_FRONTEND=noninteractive PIP_DISABLE_PIP_VERSION_CHECK=1 cd "$(dirname "$0")/.." # System packages, only the first time. if ! command -v pg_config >/dev/null; then sudo apt-get update -qq && sudo apt-get install -y -qq libpq-dev >/dev/null fi # Language dependencies, from the lockfiles, never newer than them. corepack enable >/dev/null 2>&1 || true pnpm install --frozen-lockfile --silent uv sync --frozen --quiet # Generated code and a local database, only when missing. [ -f src/generated/schema.ts ] || pnpm run codegen --silent [ -f .env ] || cp .env.example .env echo "setup: ok" ``` The rules behind it: - **Pin everything a guess could get wrong.** `packageManager` in `package.json`, `.python-version`, `.nvmrc`, and frozen lockfiles. An agent that runs `npm install` in a pnpm project creates a second lockfile and a diff nobody wanted. - **Guard every slow step** with a check, as above, so a rerun costs seconds. - **Never prompt.** `CI=1` and `DEBIAN_FRONTEND=noninteractive` silence most tools; `--yes`, `-y` and `--no-input` handle the rest. - **Need no secrets to set up or to test.** Put the names of optional ones in `.env.example`, with values that work offline. ## What makes a test command agent-friendly? An exit code you can trust and output that fits in a model's attention. The agent reads every byte the test prints, and pays for it in tokens and in focus. A green run should print a line or two; a red run should print the failing test's name, the assertion and the file and line, and stop. | Property | Why an agent needs it | How to get it | | ------------------------------- | --------------------------------------------------------- | ------------------------------------------------------------------ | | One command | It cannot guess between `make test`, `npm t` and `pytest` | `scripts/test.sh`, named in `AGENTS.md` | | Exit 0 means every check passed | It decides from the exit code, not by reading | `set -e`; lint, types and tests in the one script | | Quiet when green | Thousands of passing lines bury the one failure | `pytest -q`, `vitest --reporter=dot`, `go test ./... 2>&1 \| tail` | | Stops early when red | The first failure is usually the cause | `pytest -x`, `vitest --bail 1` | | A fast subset | It tests after every edit, not once at the end | `scripts/test.sh --changed`, or a path argument | | Offline after setup | Network flakes look like code failures | Fixtures and local services instead of live APIs | | Deterministic | A flaky test teaches it to retry instead of fix | Fixed seeds and clocks, no ordering between tests | On Runtime, a command's result holds at most 64 KiB of standard output; a test run that prints more should write a log file and print its summary ([run commands](/docs/javascript#run-commands)). ## How do you check a repository is agent-ready? Run the agent's first five minutes yourself, in a clean machine, and measure. This program clones a repository into a new sandbox, runs setup twice, runs the tests twice, measures how much they print, and checks the files an agent looks for. It runs as written against the Runtime API: ```ts import { Sandbox } from "withruntime"; const REPO = process.env.REPO_URL ?? "https://github.com/octocat/Hello-World.git"; const SETUP = "./scripts/setup.sh"; const TEST = "./scripts/test.sh"; const rows: { check: string; pass: boolean; detail: string }[] = []; await using sbx = await Sandbox.create({ diskMiB: 8192 }); async function timed(command: string, timeoutMs = 900_000) { const start = performance.now(); const r = await sbx.exec(command, { cwd: "/workspace/app", timeoutMs }); const seconds = (performance.now() - start) / 1000; return { ok: r.exitCode === 0, seconds, kib: (r.stdout.length + r.stderr.length) / 1024 }; } await sbx.exec(["git", "clone", "--depth", "1", REPO, "/workspace/app"], { timeoutMs: 300_000, check: true, }); const first = await timed(SETUP); rows.push({ check: "setup runs with no input", pass: first.ok, detail: `${first.seconds.toFixed(0)} s`, }); const again = await timed(SETUP); rows.push({ check: "setup is safe to rerun", pass: again.ok && again.seconds < Math.max(30, first.seconds / 3), detail: `${again.seconds.toFixed(0)} s the second time`, }); const online = await timed(TEST); rows.push({ check: "tests pass on a clean machine", pass: online.ok, detail: `${online.seconds.toFixed(0)} s`, }); rows.push({ check: "green output is short", pass: online.kib < 8, detail: `${online.kib.toFixed(1)} KiB`, }); const repeat = await timed(TEST); rows.push({ check: "tests are deterministic", pass: repeat.ok === online.ok, detail: "same result twice", }); for (const file of ["AGENTS.md", ".env.example"]) { const exists = await sbx.files.exists(`/workspace/app/${file}`); rows.push({ check: `${file} exists`, pass: exists, detail: exists ? "found" : "missing" }); } for (const row of rows) console.log(row.pass ? "PASS" : "FAIL", row.check.padEnd(32), row.detail); console.log(`${rows.filter((r) => r.pass).length}/${rows.length} agent-ready checks pass`); ``` ```python import os import time from withruntime import Sandbox REPO = os.environ.get("REPO_URL", "https://github.com/octocat/Hello-World.git") SETUP, TEST = "./scripts/setup.sh", "./scripts/test.sh" rows = [] with Sandbox.create(disk_mib=8192) as sbx: def timed(command: str, timeout_ms: int = 900_000): start = time.monotonic() r = sbx.exec(command, cwd="/workspace/app", timeout_ms=timeout_ms) return r.exit_code == 0, time.monotonic() - start, (len(r.stdout) + len(r.stderr)) / 1024 sbx.exec(["git", "clone", "--depth", "1", REPO, "/workspace/app"], check=True, timeout_ms=300_000) ok, first, _ = timed(SETUP) rows.append((ok, "setup runs with no input", f"{first:.0f} s")) ok, again, _ = timed(SETUP) rows.append((ok and again < max(30, first / 3), "setup is safe to rerun", f"{again:.0f} s the second time")) online, secs, kib = timed(TEST) rows.append((online, "tests pass on a clean machine", f"{secs:.0f} s")) rows.append((kib < 8, "green output is short", f"{kib:.1f} KiB")) repeat, _, _ = timed(TEST) rows.append((repeat == online, "tests are deterministic", "same result twice")) for name in ("AGENTS.md", ".env.example"): found = sbx.files.exists(f"/workspace/app/{name}") rows.append((found, f"{name} exists", "found" if found else "missing")) for ok, check, detail in rows: print("PASS" if ok else "FAIL", check.ljust(32), detail) print(f"{sum(ok for ok, _, _ in rows)}/{len(rows)} agent-ready checks pass") ``` Run it with `REPO_URL` set to your repository; without it, it audits GitHub's `octocat/Hello-World`, which has no setup script and fails most checks. A clone that fails stops the program with git's own error. A private repository needs a token; store it as a Runtime secret for `github.com` and the sandbox never holds it ([clone and push without the token](/use-cases/coding-agent-sandbox#push-a-branch-without-handing-over-the-token)). A command gets no input unless you pass `stdin`, exactly like an agent's tool call, so a setup step that prompts reads end of input and fails here, where you see it, instead of in the middle of an agent's task. ## How do you prove the tests need no network? Cut the sandbox off and run them again. A test that calls a live API passes on your laptop for months and then fails an agent at 3 a.m. when the API is slow, and the agent spends its turns "fixing" code that was never broken. Add this check to the audit: ```ts check import type { Sandbox } from "withruntime"; /** Runs the tests with the sandbox cut off from the internet, then puts its rules back. */ export async function testsPassOffline(sbx: Sandbox, test = "./scripts/test.sh") { const { internet, allow, deny } = await sbx.network.get(); await sbx.network.off(); try { const run = await sbx.exec(test, { cwd: "/workspace/app", timeoutMs: 900_000 }); return run.exitCode === 0; } finally { await sbx.network.set({ internet, allow, deny }); } } ``` ```python check def tests_pass_offline(sbx, test: str = "./scripts/test.sh") -> bool: """Runs the tests with the sandbox cut off from the internet, then puts its rules back.""" before = sbx.network.get() sbx.network.off() try: return sbx.exec(test, cwd="/workspace/app", timeout_ms=900_000).exit_code == 0 finally: sbx.network.set(internet=before["internet"], allow=before["allow"], deny=before["deny"]) ``` The rules are enforced on the host, outside the sandbox, so a test cannot reach the network by a route its own code forgot about, and root inside the sandbox cannot switch them back. A failure here names a test to fix or to mark as needing the network, which is information an agent can act on ([turn off sandbox internet](/how-to/turn-off-sandbox-internet)). ## What goes in AGENTS.md? The commands and the rules, in that order, and nothing a tool could find on its own. Claude Code reads `CLAUDE.md`, Codex and most others read `AGENTS.md`; keep one and point the other at it. A good one fits on a screen: ```markdown # Working in this repository - Set up: `./scripts/setup.sh` (safe to rerun after every pull) - Test everything: `./scripts/test.sh`; one area: `./scripts/test.sh src/billing` - Format: `pnpm format --changed`. Never reformat files you did not change. - Never edit `src/generated/`; run `pnpm codegen` instead. - Migrations: add a new file in `db/migrations/`; never edit an applied one. - Done means `./scripts/test.sh` exits 0. Paste its last 20 lines in the PR. ``` Each line answers a question an agent would otherwise spend turns on, or prevents a mistake you have already seen one make. ## How does an image make every agent start ready? It runs the setup once, at build time, instead of once per task. Build an image from a short Dockerfile that starts from Runtime's base image, which already has Python, Node, Bun, git and a compiler: ```dockerfile FROM runtime WORKDIR /workspace/app COPY . . RUN ./scripts/setup.sh && chown -R 1000:1000 /workspace/app ``` ```bash no-run runtime images build . -f Dockerfile.agent -t app-agent:latest runtime sandboxes create --image app-agent ``` ```ts check import { readFile } from "node:fs/promises"; import { Runtime } from "withruntime"; const runtime = new Runtime(); await runtime.images.build( { name: "app-agent", dockerfile: await readFile("Dockerfile.agent", "utf8"), contextDir: "." }, { onLog: (line) => console.log(line.text) }, ); await using sbx = await runtime.sandboxes.create({ image: "app-agent" }); const fresh = await sbx.exec("git pull -q && ./scripts/setup.sh && ./scripts/test.sh", { timeoutMs: 600_000, }); console.log(fresh.exitCode === 0 ? "ready" : fresh.stderr); ``` Commands in a sandbox from this image start in `/workspace/app`, the last `WORKDIR`, and run as the sandbox user, which is why the `chown` matters. Keep `.git` in the build context so an agent can `git pull` to today's code; the idempotent setup then installs only what changed since the build. Every build of a name is its next version, a rebuild reruns only the steps after the `COPY` that changed, and building is free with credit ([custom images](/docs/images)). Rebuild on each merge to your main branch, and agents start within a day of the current dependencies. For a machine with services already running, such as a database with seed data, a snapshot of a set-up sandbox goes one step further than an image ([snapshot or image](/use-cases/repo-onboarding#snapshot-or-image)). ## What does it cost to keep a repository agent-ready? Cents. An audit run that keeps a 2 vCPU, 4 GiB sandbox for ten minutes and uses four minutes of CPU costs $0.0067. A 3 GB image is stored for $0.24 a month. The saving is on the other side: each agent task that starts from the image skips its setup, and an agent spends nothing on install turns it never takes ([pricing](/pricing)). ## In short - A setup script that never prompts and is safe to rerun removes the most wasted agent turns. - One test command, with a trustworthy exit code and short green output, is how an agent knows it is done. - Test offline and twice in a clean machine; the audit above does both. - Bake setup into an image so every agent sandbox starts ready, and rebuild it when main changes. ## Run it on Runtime Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model ([compare costs](/how-to/compare-your-costs)). Every Runtime sandbox is its own Firecracker microVM with Ubuntu, Python, Node, Bun, git and passwordless `sudo`, so the audit runs your setup exactly as an agent would, with nothing on your laptop. CPU is billed at $0.025 per vCPU-hour actually used, and every account gets 100 hours of a 2 vCPU, 4 GB sandbox included every month: [start at withruntime.com](/docs/start).