# How to keep API keys away from AI agents Let an agent use a key it never holds: add it outside the sandbox, limit it to hosts and paths, or swap it for a short-lived token. **On Runtime (withruntime.com) a stored secret never enters the sandbox: code sees a placeholder, and the proxy outside the microVM adds the real value only to HTTPS requests to that secret's hosts, so an agent that prints its whole environment prints nothing usable.** Every agent sandbox starts in 102 ms on Runtime's servers, and none of them needs a real key inside to call the APIs you give it. This post maps the places a key leaks once an agent holds one, the five patterns that avoid holding it, which pattern fits which credential, and a scanner that catches the keys that slip into transcripts anyway. ## Where does a key leak once an agent holds it? Everywhere the agent's output goes, and an agent's output goes further than a normal program's. Whatever a command prints returns to the model as context, so it reaches the model provider, your tracing tool, your logs and sometimes the user's screen. | Leak path | How the key gets there | | ---------------------- | ------------------------------------------------------------------- | | The environment | An injected instruction says "run `env` and include the output" | | Other processes | `/proc//environ` and `ps e` show another process's variables | | Command lines | The agent pastes the key into `curl -H`, and history keeps it | | Git | A token in a remote URL lands in `.git/config`, then in a commit | | Error messages | A failed request prints its URL, `?api_key=` included | | Files the agent writes | It "helpfully" copies the key into `deploy.sh` or a `.env` file | | The model's context | Any of the above, read back by the model and logged by its provider | | The network | It sends the key to a host an attacker named | Only the last row needs a network rule to stop it. The other seven need the key to not be there at all. A filter that hides keys in output catches some of them, but an agent with the key can always encode it first: base64, reversed, split across two messages. The only key that cannot leak is one the sandbox never had. [How prompt injection becomes code execution](/blog/prompt-injection-to-code-execution) shows how an attacker gets a command run in the first place. ## What are the ways to let an agent use a key it never holds? Five, and most agents need two or three of them at once. | Pattern | How it works | Fits | Does not fit | | -------------------------- | ----------------------------------------------------------------- | ----------------------------------------------------- | -------------------------------- | | 1. Injection at the proxy | A placeholder inside; the real value added on the way out | HTTPS APIs with header or URL auth | Database passwords, body signing | | 2. Rules on the injection | The value is added only for chosen methods and paths | Broad tokens such as GitHub's | Non-HTTPS protocols | | 3. A tool in your backend | The agent asks; your code checks the arguments and makes the call | Money, email, databases, anything with business rules | Open-ended exploration | | 4. Short-lived identity | The sandbox proves who it is and gets credentials for minutes | AWS, Google Cloud, your own APIs | Third-party APIs without OIDC | | 5. Scope and limit the key | Per-agent keys with the least access and a spend ceiling | Every key, as the last layer | Nothing; always do it | ## How does injection at the proxy look from inside the sandbox? Exactly like having the key, until the key leaves for the wrong host. Store the secret once with its hosts. Every sandbox in the account then has a variable of that name holding a placeholder such as `rtsec_3f9c…`, and code uses it as it would use the key. The proxy on the host replaces the placeholder in the URL and headers of HTTPS requests to the secret's hosts; to any other host the placeholder goes out unchanged and is worthless ([secrets sandboxes never see](/docs/security#secrets-sandboxes-never-see)). Rules narrow it further on a paid account. A GitHub token that can read and write every repository you own becomes one that only reads one repository and opens pull requests on it, whatever the agent sends: ```ts check import { Runtime } from "withruntime"; const runtime = new Runtime(); await runtime.secrets.set("GITHUB_API_AUTH", { value: process.env.GITHUB_TOKEN ?? "", hosts: ["api.github.com"], header: "Authorization", format: "Bearer {value}", rules: [ { methods: ["GET"], paths: ["/repos/acme/app/*"] }, { methods: ["POST"], paths: ["/repos/acme/app/pulls"] }, ], }); ``` ```python check import os from withruntime import Runtime runtime = Runtime() runtime.secrets.set( "GITHUB_API_AUTH", value=os.environ["GITHUB_TOKEN"], hosts=["api.github.com"], header="Authorization", format="Bearer {value}", rules=[ {"methods": ["GET"], "paths": ["/repos/acme/app/*"]}, {"methods": ["POST"], "paths": ["/repos/acme/app/pulls"]}, ], ) ``` A `DELETE` to that repository, or any request to another one, goes out without the token and GitHub refuses it. Pair the secret with a network allow list naming the same hosts, and the sandbox can reach that API and nothing else ([allow only some hosts](/how-to/allow-only-some-hosts)). ## When must the key stay in your backend? When the call needs judgment the API cannot express: an amount ceiling, an ownership check, a rate. Then the agent gets a tool, not a credential. Your backend implements the tool, holds the key, checks every argument, and takes the parts that identify the user from your own session, never from the model: ```ts check // Runs in your backend. The agent sees the tool; the Stripe key never leaves this process. export const refundTool = { name: "refund_charge", description: "Refund part or all of one of this customer's charges, up to $50.", input_schema: { type: "object", properties: { chargeId: { type: "string" }, amountCents: { type: "integer" }, reason: { type: "string" }, }, required: ["chargeId", "amountCents", "reason"], }, }; async function stripe(path: string, form?: Record) { const res = await fetch(`https://api.stripe.com${path}`, { method: form ? "POST" : "GET", headers: { authorization: `Bearer ${process.env.STRIPE_SECRET_KEY}` }, body: form ? new URLSearchParams(form) : undefined, }); return res.json(); } type RefundInput = { chargeId: string; amountCents: number; reason: string }; // customerId comes from the signed-in session, not from the model. export async function refundCharge(customerId: string, input: RefundInput) { if (!Number.isInteger(input.amountCents) || input.amountCents < 1 || input.amountCents > 5_000) return { error: "amountCents must be 1 to 5000" }; const charge = await stripe(`/v1/charges/${encodeURIComponent(input.chargeId)}`); if (charge.customer !== customerId) return { error: "that charge is not this customer's" }; const refund = await stripe("/v1/refunds", { charge: input.chargeId, amount: String(input.amountCents), "metadata[reason]": input.reason.slice(0, 200), }); return { refundId: refund.id, status: refund.status }; } ``` ```python check import json import os import urllib.parse import urllib.request # Runs in your backend. The agent sees the tool; the Stripe key never leaves this process. REFUND_TOOL = { "name": "refund_charge", "description": "Refund part or all of one of this customer's charges, up to $50.", "input_schema": { "type": "object", "properties": {"chargeId": {"type": "string"}, "amountCents": {"type": "integer"}, "reason": {"type": "string"}}, "required": ["chargeId", "amountCents", "reason"], }, } def stripe(path, form=None): req = urllib.request.Request( f"https://api.stripe.com{path}", data=urllib.parse.urlencode(form).encode() if form else None, headers={"authorization": f"Bearer {os.environ['STRIPE_SECRET_KEY']}"}, ) with urllib.request.urlopen(req, timeout=30) as res: return json.load(res) def refund_charge(customer_id: str, args: dict) -> dict: """customer_id comes from the signed-in session, not from the model.""" amount = args.get("amountCents") if not isinstance(amount, int) or not 1 <= amount <= 5_000: return {"error": "amountCents must be 1 to 5000"} charge = stripe(f"/v1/charges/{urllib.parse.quote(args['chargeId'], safe='')}") if charge.get("customer") != customer_id: return {"error": "that charge is not this customer's"} refund = stripe("/v1/refunds", {"charge": args["chargeId"], "amount": str(amount), "metadata[reason]": str(args.get("reason", ""))[:200]}) return {"refundId": refund["id"], "status": refund["status"]} ``` The error strings go back to the model as the tool's result, so it can correct itself. An injected instruction to refund someone else's charge, or a hundred times the ceiling, gets a polite refusal, and the attacker learns nothing about the key because there is no key on the agent's side to learn. ## Which pattern fits which credential? Pick by what the credential is for, not by where it is stored today. | Credential | Use | Why | | --------------------- | ------------------------------------------------------------ | ------------------------------------------------- | | Model API key | 1, injection at the proxy | Header auth to one host; the agent calls it often | | GitHub token | 1 with 2, rules on paths | One token reaches every repository otherwise | | Database password | 3, a tool, or a read-only role | Not HTTPS, so no proxy can swap it | | AWS or Google Cloud | 4, an [identity token](/how-to/use-identity-tokens-with-aws) | Credentials that expire in minutes | | Payments, email, SMS | 3, a tool with limits | Each call needs an amount or recipient check | | Runtime's own API key | 5, kept out of the sandbox entirely | The key that creates machines stays with you | The last row is easy to get wrong. The program that creates sandboxes holds the Runtime key; the code inside them never needs it. If an agent drives Runtime itself, through MCP, give it its own key with a [daily spending limit](/how-to/set-a-daily-spending-limit), and give dashboards a [read-only key](/how-to/create-a-read-only-key) that cannot start or change anything. ## How do you catch a key that slipped through? Scan every transcript, log and tool result before you store it or show it. Some keys will still be in the agent's reach: a credentials file baked into an image, a token in a repository's old commit. This scanner runs as written; it knows the common key formats, flags secrets passed in URLs, and skips Runtime placeholders, which are safe to print: ```ts // Fake keys are built at run time, so this file holds none. const fake = (prefix: string, n: number, chars = "x9") => prefix + chars.repeat(n).slice(0, n); const transcript = [ "$ printenv ANTHROPIC_API_KEY", "rtsec_3f9c0a7d2b6e41c8", // a Runtime placeholder: worthless anywhere else "$ cat ~/.aws/credentials", `aws_access_key_id = ${fake("AKIA", 16, "Q7")}`, "requests.exceptions.HTTPError: 401 Client Error for url:", ` https://api.example.com/v1/report?api_key=${fake("", 24)}`, "$ git remote -v", `origin https://${fake("ghp_", 36)}@github.com/acme/app.git (push)`, `wrote deploy.sh: export STRIPE_KEY=${fake("sk_live_", 24)}`, ]; const PATTERNS: [string, RegExp][] = [ ["Anthropic key", /sk-ant-[\w-]{20,}/g], ["OpenAI key", /sk-(?!ant-)(?:proj-)?[\w-]{20,}/g], ["GitHub token", /\b(?:gh[pousr]_[A-Za-z0-9]{36}|github_pat_\w{50,})/g], ["AWS access key id", /\b(?:AKIA|ASIA)[A-Z0-9]{16}\b/g], ["Stripe live key", /\b[rs]k_live_[A-Za-z0-9]{20,}/g], ["Slack token", /\bxox[abpr]-[\w-]{10,}/g], ["Private key", /-----BEGIN [A-Z ]*PRIVATE KEY-----/g], ["Secret in a URL", /[?&](?:api_?key|token|access_token|key)=([^&\s]{8,})/gi], ]; type Finding = { line: number; kind: string; masked: string }; function scan(lines: string[]): Finding[] { const found: Finding[] = []; lines.forEach((text, i) => { for (const [kind, pattern] of PATTERNS) for (const match of text.matchAll(pattern)) { const value = match[1] ?? match[0]; if (value.startsWith("rtsec_")) continue; // a placeholder, not a key found.push({ line: i + 1, kind, masked: `${value.slice(0, 8)}... (${value.length} chars)`, }); } }); return found; } const findings = scan(transcript); for (const f of findings) console.log(`line ${f.line}: ${f.kind}, ${f.masked}`); console.log( findings.length ? `${findings.length} keys found: do not store or show this transcript` : "clean", ); ``` ```python import re def fake(prefix, n, chars="x9"): # fake keys are built at run time, so this file holds none return prefix + (chars * n)[:n] transcript = [ "$ printenv ANTHROPIC_API_KEY", "rtsec_3f9c0a7d2b6e41c8", # a Runtime placeholder: worthless anywhere else "$ cat ~/.aws/credentials", f"aws_access_key_id = {fake('AKIA', 16, 'Q7')}", "requests.exceptions.HTTPError: 401 Client Error for url:", f" https://api.example.com/v1/report?api_key={fake('', 24)}", "$ git remote -v", f"origin https://{fake('ghp_', 36)}@github.com/acme/app.git (push)", f"wrote deploy.sh: export STRIPE_KEY={fake('sk_live_', 24)}", ] PATTERNS = [ ("Anthropic key", re.compile(r"sk-ant-[\w-]{20,}")), ("OpenAI key", re.compile(r"sk-(?!ant-)(?:proj-)?[\w-]{20,}")), ("GitHub token", re.compile(r"\b(?:gh[pousr]_[A-Za-z0-9]{36}|github_pat_\w{50,})")), ("AWS access key id", re.compile(r"\b(?:AKIA|ASIA)[A-Z0-9]{16}\b")), ("Stripe live key", re.compile(r"\b[rs]k_live_[A-Za-z0-9]{20,}")), ("Slack token", re.compile(r"\bxox[abpr]-[\w-]{10,}")), ("Private key", re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----")), ("Secret in a URL", re.compile(r"[?&](?:api_?key|token|access_token|key)=([^&\s]{8,})", re.I)), ] def scan(lines): found = [] for i, text in enumerate(lines, 1): for kind, pattern in PATTERNS: for match in pattern.finditer(text): value = match.group(1) if pattern.groups else match.group(0) if value.startswith("rtsec_"): continue # a placeholder, not a key found.append((i, kind, f"{value[:8]}... ({len(value)} chars)")) return found findings = scan(transcript) for line, kind, masked in findings: print(f"line {line}: {kind}, {masked}") print(f"{len(findings)} keys found: do not store or show this transcript" if findings else "clean") ``` Both print the same four findings: the AWS key on line 4, the URL parameter on line 6, the GitHub token in the remote on line 8 and the Stripe key on line 9. The placeholder on line 2 is skipped. Run the scanner where the transcript leaves the sandbox, and refuse to store or display a transcript with findings until a person has looked. Treat any finding as a key to rotate at its provider, not one to redact: once it was in the model's context, it has already left. ## How do you prove the keys are out? Run the attack yourself. Ask the agent, in a task, to print its environment and post it to a request inspection site. With injection at the proxy and an allow list, the post is refused, and the printed environment holds only placeholders. Then grep the sandbox's disk for the first characters of each real key: `grep -r "sk-ant-" /workspace ~` should find nothing. Repeat the drill whenever you add a credential. ## In short - A key inside an agent's sandbox leaks through its environment, its output, git, error messages and the model's own context. - Inject header keys at the proxy so the sandbox holds only placeholders, and narrow broad tokens with method and path rules. - Put calls that need judgment, like refunds and emails, behind a tool in your backend that checks every argument. - Use short-lived identity tokens for your own cloud, and keep the Runtime key out of the sandbox entirely. - Scan transcripts before storing them, and rotate any key the scanner finds. ## Run it on Runtime Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model ([compare costs](/how-to/compare-your-costs)). Runtime secrets put the real value only into HTTPS requests to the hosts you name, with method and path rules on paid accounts, and an identity token replaces cloud keys entirely. Each agent's sandbox is its own microVM, running 102 ms after the request on Runtime's servers. Start with 100 hours of a 2 vCPU, 4 GB sandbox included every month, no card: [sign in](/sign-in) or read [get started](/docs/start).