Runtime

Build an AI data analyst that gets the numbers right on any CSV

Load the CSV into a sandboxed Python session, show the model a profile before it writes code, and reject any number it did not compute.

Runtime (withruntime.com) gives each dataset its own Firecracker microVM with pandas, NumPy and matplotlib already installed and no internet unless you allow it, and a 2 vCPU, 4 GiB session waiting for the next question costs $0.03125 an hour. A data analyst agent is easy to demo and hard to trust. The demo works on a clean file; real files have dollar signs in number columns, three date formats and a region spelled two ways. This post is about the part after the demo: making the numbers right.

Why do AI data analysts get numbers wrong?

Because the model writes code for the file it imagines, not the file it got, and then reports whatever the code printed. Here is a five-row file of the kind every business has:

Textorder_id,order_date,region,revenue,discount1001,2026-01-04,West,"$1,200.00",0.11002,2026-01-17,east,$840.50,1003,04/02/2026,West,"$2,310.00",0.051004,2026-02-11,East,$95.00,01005,2026-03-02,North,"$1,045.25",

Ask "revenue by region" and a model that has not looked will write df.groupby("region")["revenue"].sum(). That runs without an error and returns this:

TextregionEast                 $95.00North             $1,045.25West     $1,200.00$2,310.00east                $840.50

revenue was read as text, so "sum" joined the strings end to end, and east is its own region. A model that reads this output sometimes notices. Sometimes it reports West as "$1,200.00$2,310.00", or worse, quietly does the sum in its head. Each mistake has a cause you can design against:

What goes wrong Why The defense
Text columns summed as strings Currency symbols and thousands separators Show column types and examples before any code
One category split in two Inconsistent case or spacing Show distinct values for low-cardinality text
Wrong rows in a date range Mixed date formats parsed differently Show examples, and parse dates explicitly
A column name that does not exist The model guessed from the question Give the exact column list
Rows silently dropped dropna() or a failed parse Report row counts in every answer
A number nobody computed Mental arithmetic on printed output Check every number in the answer against output

What should the model see before it writes code?

A profile: the row count, and for every column its type, empty count, distinct count, range and a few example values. It costs a few hundred tokens and prevents most of the table above. Compute it in the sandbox, from the real file, before the model sees the question:

TypeScriptimport { Sandbox } from "withruntime";const CSV = `order_id,order_date,region,revenue,discount1001,2026-01-04,West,"$1,200.00",0.11002,2026-01-17,east,$840.50,1003,04/02/2026,West,"$2,310.00",0.051004,2026-02-11,East,$95.00,01005,2026-03-02,North,"$1,045.25",`;const PROFILE = `import json, sysimport pandas as pddf = pd.read_csv(sys.argv[1], low_memory=False)print(f"{len(df)} rows, {len(df.columns)} columns")for name in df.columns:    values = df[name]    column = {"name": name, "type": str(values.dtype), "empty": int(values.isna().sum()),              "distinct": int(values.nunique())}    if pd.api.types.is_numeric_dtype(values) and values.notna().any():        column["min"], column["max"] = float(values.min()), float(values.max())    column["examples"] = [str(v) for v in values.dropna().unique()[:4]]    print(json.dumps(column))`;await using sbx = await Sandbox.create({ network: { internet: false } });await sbx.files.write("/workspace/data/sales.csv", CSV);await sbx.files.write("/workspace/profile.py", PROFILE);const profile = await sbx.exec("python3 /workspace/profile.py /workspace/data/sales.csv", {  timeoutMs: 120_000,});console.log(profile.stdout);
Pythonfrom withruntime import SandboxCSV = """order_id,order_date,region,revenue,discount1001,2026-01-04,West,"$1,200.00",0.11002,2026-01-17,east,$840.50,1003,04/02/2026,West,"$2,310.00",0.051004,2026-02-11,East,$95.00,01005,2026-03-02,North,"$1,045.25","""PROFILE = """import json, sysimport pandas as pddf = pd.read_csv(sys.argv[1], low_memory=False)print(f"{len(df)} rows, {len(df.columns)} columns")for name in df.columns:    values = df[name]    column = {"name": name, "type": str(values.dtype), "empty": int(values.isna().sum()),              "distinct": int(values.nunique())}    if pd.api.types.is_numeric_dtype(values) and values.notna().any():        column["min"], column["max"] = float(values.min()), float(values.max())    column["examples"] = [str(v) for v in values.dropna().unique()[:4]]    print(json.dumps(column))"""with Sandbox.create(network={"internet": False}) as sbx:    sbx.files.write("/workspace/data/sales.csv", CSV)    sbx.files.write("/workspace/profile.py", PROFILE)    profile = sbx.exec("python3 /workspace/profile.py /workspace/data/sales.csv", timeout_ms=120_000)    print(profile.stdout)

For the file above, the profile says revenue has type object with examples like $1,200.00, region has four distinct values including both East and east, and order_date mixes 2026-01-04 with 04/02/2026. A model that reads those three lines writes the cleaning code first. One line per column, as JSON, keeps a 60-column file under a few thousand tokens.

The file goes in with files.write and the sandbox has no internet, so the data never leaves the machine that analyzes it. For files too large to send in one call, files.upload streams them, and a bucket can be mounted as a directory instead (storage).

How do you load the data once?

Load it in a stateful session before the first question, as cell zero, from your code rather than the model's. Then every question works on the same df, and nobody re-parses a 2 GB file per question:

TypeScriptimport { Sandbox } from "withruntime";// The dataset's own sandbox, which has held /workspace/data/sales.csv since the upload.const sbx = await Sandbox.getOrCreate("analyst-dataset-7", {  memoryMiB: 8192,  network: { internet: false },});const loaded = await sbx.interpreter.run(  "import pandas as pd\n" +    "df = pd.read_csv('/workspace/data/sales.csv', low_memory=False)\n" +    "print(len(df), 'rows loaded')",  { timeoutMs: 600_000 },);console.log(loaded.status, loaded.stdout);
Pythonfrom withruntime import Sandbox# The dataset's own sandbox, which has held /workspace/data/sales.csv since the upload.sbx = Sandbox.get_or_create("analyst-dataset-7", memory_mib=8192, network={"internet": False})loaded = sbx.interpreter.run(    "import pandas as pd\n"    "df = pd.read_csv('/workspace/data/sales.csv', low_memory=False)\n"    "print(len(df), 'rows loaded')",    timeout_ms=600_000,)print(loaded["status"], loaded["stdout"])

From here the model's tool is the stateful run_python from building a code interpreter tool, and the system prompt carries the profile and the rules:

TextYou are a data analyst. The user's file is loaded as `df` in a Python session.Its profile is below.- Before computing, check the columns you need in the profile and clean them  in code: strip currency symbols and separators, parse dates explicitly,  normalize case and spacing.- Compute every number in your answer with run_python and print it. Do no  arithmetic yourself, including percentages, ratios and differences.- Say how many rows the answer is based on, and what you excluded and why.- If the question is ambiguous, say which reading you chose.- For a chart, draw it with matplotlib and do not describe numbers it does  not print.Profile:<the profile output>

Pandas needs several times the CSV's size in memory, more with text columns, so size the sandbox to the file: up to 64 GiB on one sandbox. Memory is billed per GiB-hour while the session runs, so a large session that pauses between questions is still cheap.

How do you know every number was computed?

Check the answer against what the cells printed. Any number in the answer that no output contains, at the answer's precision, was typed by the model. Send it back with a request to compute it in code, or show it to the user marked as unverified:

TypeScriptconst NUMBER = /(?<![\w.,-])-?\d[\d,]*(?:\.\d+)?/g;const numbers = (text: string) => (text.match(NUMBER) ?? []).map((n) => n.replaceAll(",", ""));/** Numbers in the answer that no cell printed: the model worked them out itself. */function untraced(answer: string, question: string, printed: string[]): string[] {  const seen = printed.flatMap(numbers).map(Number);  const asked = new Set(numbers(question));  return numbers(answer).filter((n) => {    if (asked.has(n) || /^-?\d$/.test(n)) return false; // "top 3" is not a result    const places = n.split(".")[1]?.length ?? 0;    return !seen.some((value) => Number(value.toFixed(places)) === Number(n));  });}const question = "Which region sold the most, and by how much over the next one?";const printed = ["region\nWest     3510.00\nNorth    1045.25\nEast      935.50"];const answer = "West sold $3,510.00, about 3.4 times North's $1,045.25, across 5 orders.";console.log("untraced:", untraced(answer, question, printed)); // [ "3.4" ]
Pythonimport reNUMBER = re.compile(r"(?<![\w.,-])-?\d[\d,]*(?:\.\d+)?")def numbers(text: str) -> list[str]:    return [n.replace(",", "") for n in NUMBER.findall(text)]def untraced(answer: str, question: str, printed: list[str]) -> list[str]:    """Numbers in the answer that no cell printed: the model worked them out itself."""    seen = [float(n) for text in printed for n in numbers(text)]    asked = set(numbers(question))    missing = []    for n in numbers(answer):        if n in asked or re.fullmatch(r"-?\d", n):  # "top 3" is not a result            continue        places = len(n.split(".")[1]) if "." in n else 0        if not any(round(value, places) == float(n) for value in seen):            missing.append(n)    return missingquestion = "Which region sold the most, and by how much over the next one?"printed = ["region\nWest     3510.00\nNorth    1045.25\nEast      935.50"]answer = "West sold $3,510.00, about 3.4 times North's $1,045.25, across 5 orders."print("untraced:", untraced(answer, question, printed))  # ['3.4']

The ratio 3.4 is flagged: no cell printed it, so the model divided in its head. It happens to be close, which is the danger: a check that only catches wrong numbers would let it through, and the next one might not be close. A single retry with "compute 3.4 in code and print it" fixes nearly all of these. Keep it a flag rather than a hard failure, since an answer can legitimately repeat a year or an id from the question.

How do you show the work to the user?

Show the answer first and the work one click away. People trust an analyst who can show the query behind a number, and an AI analyst can show more than most, because every step it took is a cell you already have.

Keep three things with every answer:

  • The cells that produced it, in order, with their printed output. That is a complete, rerunnable record of how the number was reached.
  • The row count and the exclusions. "Based on 48,210 of 50,000 orders; 1,790 had no region" tells a reader more than any confidence score.
  • The assumption, if the question had two readings. "Best month" by revenue or by order count is a choice the user should see, not discover.

Save the cells to the sandbox as a script as well. When the same question comes back next month on a new file, rerun the script on it instead of asking the model again: the answer is reproducible, and it costs no tokens.

Say no when the data cannot answer. A file with no cost column cannot give a profit figure, and a model asked for one will sometimes invent a proxy. Add one line to the system prompt: if the columns needed are not in the profile, say which are missing and stop. A clear "this file has no cost data" is a better answer than a confident guess, and users remember which one they got.

What does a question cost?

Very little, if the session pauses between questions. Take a session that answers one question in 90 seconds of running time, with 20 CPU-seconds of pandas work, on 2 vCPU and 4 GiB, and then waits 60 seconds before it pauses itself:

TextMemory:  (90 + 60) s / 3,600 × 4 GiB × $0.0075 = $0.00125CPU:     20 / 3,600 × $0.025                         = $0.00014A question:                                           $0.00139

While paused, the session keeps df in its saved memory and pays $0.08 per decimal GB per 30-day month of what it alone stores. The next question wakes it, and df is still there (pause, don't delete).

In short

  • The model writes code for the file it imagines; show it the real one first with a compact profile.
  • Load the data once, from your code, into a stateful session the model's tool shares.
  • Tell the model to clean in code and to compute and print every number.
  • Check every number in the answer against the printed output, and send untraced ones back.
  • Keep the data offline in its own sandbox, and pause between questions.

Run it on Runtime

Runtime costs 42% to 88% less than fourteen other sandbox providers for an agent that mostly waits on a model (compare costs). Every Runtime sandbox has pandas, NumPy and matplotlib installed, a stateful interpreter that returns charts as images, and no internet unless you allow it. A new sandbox runs its first command 221 ms after the create request on Runtime's servers, and CPU bills at $0.025 per vCPU-hour of actual use. Start with 100 free hours, no card: sign in, or read get started and the data analysis agent guide.

Your first 100 hoursare on us.

  • No credit card
  • Eight sandboxes at once, 2 vCPU and 4 GiB each
  • Then prepaid credit from $10, no plan fee
Claim 100 hours free