Runtime

Scheduled AI agents: run an agent every night without a server

Schedule a job that starts a fresh sandbox on a cron line, runs the agent, writes its result somewhere that lasts, and exits.

Runtime (withruntime.com) runs each scheduled job in a fresh sandbox on a cron schedule in your own time zone, for up to 60 minutes a run, tries a failed run again up to 5 times, and bills only the time each run takes. The most useful agents are often the boring ones: update minor dependencies every weeknight, label the issues that came in overnight, rerun the flaky tests on Sunday and report which ones failed twice. None of them needs a chat window, and all of them need a machine at 02:30 that nobody has to keep running. This post schedules one, makes it safe to run twice, and decides what its exit code should mean.

What does a scheduled agent need?

Five things, and only the first is the agent:

Part Job On Runtime
The agent A loop that calls a model and runs commands Your code
A trigger Starts the run at a time, in a time zone A job's cron schedule
A clean machine No leftovers from yesterday's run A fresh sandbox for every run
Credentials The model key, a repository token Job secrets, in the run's environment
Somewhere to land A pull request, an issue comment, a file in a bucket Your own systems

The fresh machine matters more than it looks. An agent that ran on a long-lived server inherits whatever the last run installed, half-finished, and the bug it hits on Thursday depends on what happened on Tuesday. With a new sandbox each time, every run starts from the same place and a failure can be reproduced.

Nothing carries over between runs, so the agent reads its input from somewhere that lasts and writes its output back there. A repository is the natural place for a coding agent: it clones at the start and opens a pull request at the end.

Where should the agent's own commands run?

A job's secrets reach its run as real values in the environment. That is convenient, and it means anything running in that sandbox can read them, including a command the model chose. There are two shapes:

Shape Model's commands can see the keys Extra work Use it when
One sandbox: loop and tools together Yes None The agent reads only your own code
Two sandboxes: the job runs the loop, a second runs the tools No One Sandbox.create and an upload The agent reads issues, web pages or changelogs

The second shape is the one to start with. The job's sandbox holds the model key, the GitHub token and a Runtime key, and runs only your code. The tools sandbox gets a copy of the repository, the package registry on its network allow list, and no credentials of any kind. When the agent finishes, your code reads the diff out of it and does the committing. A prompt hidden in an issue can make the model run anything it likes, and there is nothing in its machine worth stealing (prompt injection to code execution).

What does the nightly run look like?

This is the script the job runs. It uses the loop from the agent loop in 40 lines, with the sandbox passed in and the verification step that ends a run as verified only when the tests pass:

TypeScript// nightly.ts, the job's command. ANTHROPIC_API_KEY, GITHUB_TOKEN and RUNTIME_API_KEY are job// secrets in this sandbox. The agent's commands run in a second sandbox that holds none of them.import { execSync } from "node:child_process";import { writeFileSync } from "node:fs";import { Sandbox } from "withruntime";// The loop from "The agent loop in 40 lines", taking its sandbox and verifying with `npm test`.declare function agent(sbx: Sandbox, task: string): Promise<{ stop: string; text?: string }>;const REPO = "acme/shop";const branch = `agent/deps-${new Date().toISOString().slice(0, 10)}`;const sh = (command: string) => execSync(command, { cwd: "/workspace/repo", stdio: "inherit" });const github = (path: string, body?: object) =>  fetch(`https://api.github.com/repos/${REPO}${path}`, {    method: body ? "POST" : "GET",    headers: {      authorization: `Bearer ${process.env.GITHUB_TOKEN}`,      accept: "application/vnd.github+json",    },    body: body && JSON.stringify(body),  });async function main(): Promise<number> {  // Safe to repeat: a retry, or the one run for several missed times, finds today's branch.  if ((await github(`/branches/${branch}`)).ok) {    console.log(`SKIP: ${branch} exists`);    return 0;  }  const auth = `x-access-token:${process.env.GITHUB_TOKEN}`;  execSync(`git clone -q --depth 50 https://${auth}@github.com/${REPO} /workspace/repo`);  // The agent's machine: the code, the package registry, and no credentials of any kind.  await using tools = await Sandbox.create({    timeoutSeconds: 2400,    network: { internet: true, allow: ["registry.npmjs.org"] },    labels: { job: "nightly-deps" },  });  await tools.files.upload("/workspace/repo", "/workspace/repo");  const result = await agent(    tools,    "In /workspace/repo, update the npm dependencies that have a new minor or patch " +      "version, run `npm test`, and fix whatever breaks. Do not change major versions.",  );  const patch = (await tools.exec("git diff", { cwd: "/workspace/repo" })).stdout;  if (result.stop !== "verified" || !patch) {    console.log(`NO CHANGE: the agent stopped on ${result.stop}. ${result.text ?? ""}`);    return 0;  }  writeFileSync("/tmp/agent.patch", patch);  sh(`git switch -q -c ${branch} && git apply /tmp/agent.patch`);  sh(`git commit -qam "Update minor and patch dependencies" && git push -q origin ${branch}`);  const title = `Dependency updates, ${branch.slice(-10)}`;  const pr = await github("/pulls", { title, head: branch, base: "main", body: result.text ?? "" });  if (!pr.ok) throw new Error(`opening the pull request answered ${pr.status}`);  console.log(`PR: ${((await pr.json()) as { html_url: string }).html_url}`);  return 0;}// A thrown error exits non-zero, which is the one case a retry can help.process.exitCode = await main();

Three lines in it do most of the work:

  • The branch name is the day. If the run is retried after it pushed, or if Runtime runs it once to cover times it missed, the second run finds the branch and stops. Without this check, a retry opens a second pull request.
  • The model never holds a credential. The tools sandbox has a copy of the code, not the remote. Your code applies the diff and pushes, so the only thing the agent can change is the content of a pull request a person reviews.
  • The tools sandbox has its own time limit. If the job is stopped at its timeout, await using never runs, and the tools sandbox stops on its own when its 40 minutes are up.

What should the exit code mean?

A job tries a failed run again in a fresh sandbox, if you ask it to, and a failed run is one that exits non-zero. So the exit code is a question to the scheduler: would running this again help?

Outcome Exit Retried Log line
Today's branch already exists 0 No SKIP
The agent verified a change, PR opened 0 No PR: <url>
The agent stopped without a verified change 0 No NO CHANGE: ...
Model API down, clone failed, push failed Not 0 Yes, in a fresh sandbox The error
The run's sandbox was lost under it None Never State unknown

The third row is the one people get wrong. An agent that could not make the tests pass tonight will usually fail the same way five minutes later, at the same token cost. That is a result to report, not an error to retry. Keep non-zero exits for failures outside the agent, where a second attempt has a real chance.

The last row is Runtime's rule, and a sensible one: when the sandbox stopped under a run, nobody knows whether the push happened, so the run ends unknown and is never repeated on its own. Your branch check makes it safe to start that occurrence again by hand once you have looked.

How do you schedule it?

Store the three values as job secrets, then create the job. The job's own sandbox only waits on the model, so it can be smaller than the default:

TypeScriptimport { Runtime } from "withruntime";const runtime = new Runtime();const secret = async (name: string) =>  (await runtime.secrets.set(name, { value: process.env[name] ?? "", jobs: true })).jobs!.id;const job = await runtime.jobs.create({  name: "nightly-deps",  schedule: { cron: "30 2 * * 1-5", timezone: "America/New_York" }, // 02:30 on weeknights  command: [    "bash",    "-lc",    "git clone -q --depth 1 https://x-access-token:$GITHUB_TOKEN@github.com/acme/ops-agents " +      "/workspace/agents && cd /workspace/agents && bun install --frozen-lockfile && bun nightly.ts",  ],  compute: { vcpu: 1, memoryMiB: 2048 },  timeoutSeconds: 2700,  retry: { maxAttempts: 2, backoffSeconds: 300 },  secrets: [    { name: "ANTHROPIC_API_KEY", secretId: await secret("ANTHROPIC_API_KEY") },    { name: "GITHUB_TOKEN", secretId: await secret("GITHUB_TOKEN") },    { name: "RUNTIME_API_KEY", secretId: await secret("AGENT_RUNTIME_KEY") },  ],});console.log(job.id, "next run", new Date(job.nextRunAt ?? 0).toISOString());
Pythonimport osfrom withruntime import Runtimeruntime = Runtime()def secret(name: str) -> str:    return runtime.secrets.set(name, value=os.environ.get(name, ""), jobs=True)["jobs"]["id"]job = runtime.jobs.create(    "nightly-deps",    cron="30 2 * * 1-5", timezone="America/New_York",  # 02:30 on weeknights    command=["bash", "-lc",             "git clone -q --depth 1 https://x-access-token:$GITHUB_TOKEN@github.com/acme/ops-agents "             "/workspace/agents && cd /workspace/agents && bun install --frozen-lockfile && bun nightly.ts"],    compute={"vcpu": 1, "memoryMiB": 2048},    timeout_seconds=2700,    retry={"maxAttempts": 2, "backoffSeconds": 300},    secrets=[{"name": "ANTHROPIC_API_KEY", "secretId": secret("ANTHROPIC_API_KEY")},             {"name": "GITHUB_TOKEN", "secretId": secret("GITHUB_TOKEN")},             {"name": "RUNTIME_API_KEY", "secretId": secret("AGENT_RUNTIME_KEY")}],)print(job["id"], job["nextRunAt"])

The time zone is worth setting. A cron line in UTC moves by an hour against your working day twice a year; with America/New_York it stays at 02:30 local time through daylight saving changes. The command runs without a shell unless you give it one, which is why it starts with bash -lc.

Each morning, runtime jobs runs <id> lists every attempt with its state and exit code, and runtime jobs logs <id> shows the latest run's output. That is where the SKIP, PR and NO CHANGE lines pay off: one glance tells you what the agent did overnight, without opening anything else.

How do you test a job before you schedule it?

Run its command once in a sandbox of the same size, with the same timeout, and look at three numbers: the exit code, the duration against the timeout, and the amount of output against what a run keeps:

TypeScriptimport { Sandbox } from "withruntime";const TIMEOUT_S = 2700; // the job's timeoutSecondsconst KEPT = 256 * 1024; // a run keeps the last 256 KiB of its outputconst command =  "git clone -q --depth 1 https://x-access-token:$GITHUB_TOKEN@github.com/acme/ops-agents " +  "/workspace/agents && cd /workspace/agents && bun install --frozen-lockfile && bun nightly.ts";const env = {  ANTHROPIC_API_KEY: process.env.ANTHROPIC_API_KEY ?? "",  GITHUB_TOKEN: process.env.GITHUB_TOKEN ?? "",  RUNTIME_API_KEY: process.env.AGENT_RUNTIME_KEY ?? "",};await using sbx = await Sandbox.create({  vcpu: 1,  memoryMiB: 2048,  labels: { rehearsal: "nightly-deps" },});const started = Date.now();const r = await sbx.exec(command, { cwd: "/workspace", env, timeoutMs: TIMEOUT_S * 1000 });const seconds = (Date.now() - started) / 1000;const bytes = Buffer.byteLength(r.stdout + r.stderr);console.log(  `exit ${r.exitCode}, ${seconds.toFixed(0)} s of ${TIMEOUT_S} s, ${bytes} bytes of output`,);if (seconds > TIMEOUT_S / 2) console.log("Over half the timeout: raise it, or split the work.");if (bytes > KEPT) console.log("Longer than a run keeps: print a summary line at the end.");
Pythonimport osimport timefrom withruntime import SandboxTIMEOUT_S = 2700  # the job's timeout_secondsKEPT = 256 * 1024  # a run keeps the last 256 KiB of its outputcommand = ("git clone -q --depth 1 https://x-access-token:$GITHUB_TOKEN@github.com/acme/ops-agents "           "/workspace/agents && cd /workspace/agents && bun install --frozen-lockfile && bun nightly.ts")env = {"ANTHROPIC_API_KEY": os.environ.get("ANTHROPIC_API_KEY", ""),       "GITHUB_TOKEN": os.environ.get("GITHUB_TOKEN", ""),       "RUNTIME_API_KEY": os.environ.get("AGENT_RUNTIME_KEY", "")}with Sandbox.create(vcpu=1, memory_mib=2048, labels={"rehearsal": "nightly-deps"}) as sbx:    started = time.monotonic()    r = sbx.exec(command, cwd="/workspace", env=env, timeout_ms=TIMEOUT_S * 1000)    seconds = time.monotonic() - started    size = len((r.stdout + r.stderr).encode())    print(f"exit {r.exit_code}, {seconds:.0f} s of {TIMEOUT_S} s, {size} bytes of output")    if seconds > TIMEOUT_S / 2:        print("Over half the timeout: raise it, or split the work.")    if size > KEPT:        print("Longer than a run keeps: print a summary line at the end.")

Agent runs vary a lot in length, because the model decides how many turns a task takes. A rehearsal that uses half the timeout will hit it on a hard night. The job's timeout defaults to 30 minutes and can go up to 60 minutes; work that needs longer is better split into several jobs, one per package or per directory, than squeezed into one.

What does a nightly agent cost?

The run is two sandboxes for, say, 25 minutes: the job's at 1 vCPU and 2 GiB, mostly waiting on the model, and the tools sandbox at 2 vCPU and 4 GiB, busy about a third of the time while tests run. Memory is billed while each runs and CPU only when used, so one run costs about $0.03 in machines, and a month of weeknights about $0.58. The model's tokens will usually cost more than both machines together.

Runs spend the included usage first, then credit, so the first nights cost nothing; a run the included usage cannot take, on an account without credit, waits with the reason credits. To put a hard ceiling on the whole thing, create the job with a key that has a daily spending limit of its own, and give the agent's Runtime key one too. When a run would pass it, the run waits with the reason cost_cap instead of spending (set a daily spending limit). An account holds up to 1,000 jobs.

In short

  • A scheduled agent is a cron line, a fresh sandbox per run, credentials in job secrets, and a result written somewhere that lasts.
  • Run the loop in the job and the agent's commands in a second sandbox with no credentials; your code applies the diff and pushes.
  • Make every run safe to repeat, here with a branch named after the day.
  • Exit 0 for every outcome of the task, and non-zero only when a fresh attempt could help.
  • Rehearse the command once at the job's size and timeout before you schedule it.

Run it on Runtime

Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). Every job run starts in a new sandbox, keeps its exit code and the last 256 KiB of its output, and is billed only for the time it runs. Read the jobs guide, then create an account at withruntime.com and rehearse your first nightly agent today.

400 sandbox hours,every month.

Hours of a 1 GB sandbox, included free.

  • No credit card
  • Eight sandboxes at once, 2 vCPU and 4 GiB each
  • Then prepaid credit from $10, no plan fee
Start free, no card