Skip to main content
Back to Blog
AI/MLEnterpriseData Analysis
2 October 202617 min readUpdated 4 October 2026

The Agent Said It Was Done. The Database Disagreed.

The Agent Said It Was Done. The Database Disagreed. Published October 3, 2026 Microsoft ThinkingBox evaluates AI agents by examining the records and side effects they leave behi...

By Software Development Team

The Agent Said It Was Done. The Database Disagreed.

Published October 3, 2026

Microsoft ThinkingBox evaluates AI agents by examining the records and side effects they leave behind, rather than only the text they generate. It also tests whether they can produce the correct result repeatedly. The benchmark is available through Hugging Face.

A customer contacts support about a $745 kitchen appliance that has been stuck in a courier "exception" at a Nashville distribution center for fifteen days beyond its estimated delivery date.

The AI agent performs nine apparently careful tool calls. It retrieves the order, checks tracking, looks up the customer profile, searches the refund policy twice, confirms that no support ticket exists, opens one, documents the timeline, and reads the policy correctly. The customer's account segment does not qualify for late-delivery compensation.

The agent then closes the ticket as resolved and replies:

Since your query is resolved, is there anything I may assist you with?

Two problems remain. The carrier exception is still open, so the required ticket state is on hold, pending resolution. The customer also never receives a meaningful answer to the actual question.

A grader examining tool calls would see nine valid-looking operations. A grader checking the database would see that the required state was not reached. The database is what disagrees.

That gap is the focus of ThinkingBox. Across 507 stateful business workflows, each run 20 times against various LLM models, it evaluates terminal backend state and side effects. This article summarizes the findings, examines the cost of consistency, and explains how to run the benchmark through OpenEnv.

The example is adapted from the benchmark task sandbox_external_retail_group1.py:test_case_ST003_006. The executable check that fails is a single field: the ticket status is solved, while the required final state is hold. The complete trace appears in Appendix D.4, Case 3 of the ThinkingBox paper.

A tool call is not an outcome

Final responses and valid tool calls are only indirect indicators of success. An agent can sound correct while writing the wrong value, changing the wrong record, or creating an unintended side effect. The records left behind provide the decisive evidence.

The gap is substantial. In a common-set ablation containing 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. The executable checks nevertheless found incorrect field values in 77.61% of those failures, unintended extra effects in 43.30%, and missing required effects in 25.36%. These findings overlap.

A trajectory is a claim. Database state is the evidence. Repetition is the trust test.

One success is not reliability

An agent that processes a refund correctly once and mishandles it during the next four attempts is not a reliable refund agent. For that reason, every task runs 20 independent times, beginning from an identical clean backend. ThinkingBox reports three metrics:

MetricWhat it measuresWhat it answers
pass@1Share of all attempts that succeededHow does it usually perform?
pass@20Share of tasks solved at least once in 20 triesCan it ever do this? This measures breadth.
Observed 20/20Tasks that passed all 20 recorded attemptsCan it always be correct?

The observed 20/20 value used here is the literal number of the 507 tasks that passed all 20 attempts. It is not an estimator and uses no smoothing.

Single-attempt results

The following table reports pass@1, the single-attempt score estimate, by domain. Each model was evaluated on every task for 20 repeated trials. Bold values mark the group leader, and underlined values mark the runner-up. Standard errors for the single-attempt estimates are provided in Table 4 of the ThinkingBox paper.

ModelRetail (98)Auto insurance (100)Travel (104)Neobank (104)Consulting (101)Overall, task-weighted (507)
Proprietary models
Claude Opus 5.580.9768.4054.2871.2561.5867.16
Claude Opus 580.7165.8049.9570.6266.1966.50
GPT-5.476.3362.6568.1265.3454.6065.36
GPT-5.6 Sol67.6565.3060.3459.0957.5261.91
Claude Sonnet 4.672.3554.4058.9456.3954.3159.19
GPT-6 Astra71.7346.5555.8760.8756.8358.31
GPT-5.270.2022.4053.7051.1534.0646.28
Claude Opus 4.668.628.3021.1135.6727.8232.09
o3-pro37.702.9517.3124.2814.6019.31
Grok-4.343.932.6015.141.789.5514.38
Open-weight models
Kimi-K382.2450.8061.8341.3551.6357.37
Qwen3.8-27B64.0347.8553.4147.8845.6951.70
DeepSeek-V4-Pro68.2129.6543.1344.8631.0443.26
Kimi-K2.653.7224.5039.5233.6537.3337.66
GLM-5.158.6725.7035.4313.2734.0633.19
Qwen3.6-27B43.1129.0046.3927.8418.3732.94
Qwen3.5-9B19.900.704.711.152.335.65
Mistral-Large-311.281.308.991.150.744.66

Claude Opus 5.5 leads overall at 67.16%, two-thirds of a percentage point above Claude Opus 5. Kimi-K3 is the strongest open-weight model, within one point of GPT-6 Astra. Domain also has a major effect. Claude Opus 4.6 scores 68.62% on retail but only 8.30% on auto insurance.

A single successful run shows that a model can complete a task. It does not show that the model will repeat that result. Running every task 20 times reveals how much of the score survives.

Only three models retain most of their pass@1 scores across 20 repetitions. GPT-6 Astra retains 78% of its single-attempt rate, while Claude Opus 5.5 and Claude Opus 5 each retain 71%. At the other end, GLM-5.1, Kimi-K2.6, and DeepSeek-V4-Pro each retain about 8%.

The difference between what a model can do once and what it does every time is central to evaluating agents that modify business records.

Can you depend on the model behind your agent?

Kimi-K3 has the broadest coverage of any tested model. It solves 93.89% of the benchmark at least once, or 476 of 507 tasks. Only 31 tasks defeat it entirely, the lowest count in the field. On retail workflows, it leads with an 82.24% pass@1 score, ahead of every proprietary model.

Kimi-K3 is also among the least consistent. Only 68 of 507 tasks, or 13.41%, succeed in all 20 attempts.

Claude Opus 5 shows the opposite pattern. It solves fewer tasks at least once, 79.09%, with 106 tasks defeating it entirely, but completes 47.53% of the benchmark in every attempt.

A newer model does not automatically solve the consistency problem. Claude Opus 5.5 scores higher than Claude Opus 5 on the overall single-attempt measure, 67.16% versus 66.50%, and solves more tasks at least once. However, both models pass exactly 241 tasks on all 20 attempts. The half-point increase in headline accuracy produces no additional dependable tasks in this comparison.

  • Kimi-K3 solves 75 more tasks at least once than Claude Opus 5.
  • Claude Opus 5 solves 173 more tasks consistently than Kimi-K3.

For workflows that touch real records, pass@20 is therefore not the only column that matters. The observed 20/20 count shows how many tasks the model completed reliably across the entire test sequence.

What consistency costs

Capability comparisons often stop at accuracy. For deployment, another question is relevant: what does a successful unit of work cost? ThinkingBox measures this as cost per successful task attempt. The term "task attempt" matters because every benchmark task is run repeatedly and cost is incurred per attempt, making pass@1 the matching quality denominator.

The calculation uses each model's recorded token usage from its complete 507 by 20 campaign. It prices that usage at undiscounted list rates available on OpenRouter, reverses promotional discounts, and excludes endpoints that declare quantization. Input, output, and cache rates all come from one provider endpoint per model.

The calculation is:

Cost per successful task attempt = estimated cost for 507 attempts, one per task ÷ (507 × pass@1)

This is a comparative efficiency index, not an invoice or the price of serving one production request. It measures the cost of individual successes rather than consistency.

For example, GPT-5.4 costs $43.49 for 507 attempts and has a 65.36% pass@1 score:

$43.49 ÷ (507 × 0.6536) = $0.131 per successful task attempt

Pareto cost frontier

A model is on the cost frontier when no other model is both no more expensive and at least as accurate. Three models qualify.

The frontier has three steps:

  • GPT-5.6 Sol has the lowest cost per success, at $0.127.
  • GPT-5.4 increases pass@1 by 3.45 percentage points for $0.004 more per success.
  • Claude Opus 5.5 adds another 1.80 percentage points at $0.276 per success.

Each remains on the frontier because no cheaper model matches its pass@1 score.

Claude Opus 5 provides a clear example of a dominated model. At $0.475 per successful attempt and 66.50% pass@1, it is both more expensive and less accurate than Claude Opus 5.5, which costs $0.276 per successful attempt and scores 67.16%.

Pricing consistency

Cost per success rewards a model that is inexpensive and often correct. It does not reward a model that is correct every time. To measure that property, ThinkingBox also calculates cost per dependable task:

Cost per dependable task = estimated cost of 20 runs of 507 attempts ÷ tasks passing 20/20

For example, GPT-6 Astra costs 20 × $86.03, or $1,720.60, for the full campaign and passes 231 tasks on every attempt:

$1,720.60 ÷ 231 = $7.45 per dependable task

ModelTasks passing 20/20Estimated cost, 20 runsCost per dependable task
GPT-5.4128 (25.25%)$869.80$6.80
GPT-6 Astra231 (45.56%)$1,720.60$7.45
Claude Opus 5.5241 (47.53%)$1,880.77$7.80
GPT-5.6 Sol82 (16.17%)$800.00$9.76
Claude Opus 5241 (47.53%)$3,206.00$13.30
Claude Sonnet 4.6102 (20.12%)$1,587.60$15.56
GPT-5.244 (8.68%)$878.00$19.95
Kimi-K368 (13.41%)$1,406.40$20.68
Qwen3.8-27B38 (7.50%)$925.80$24.36

When ranked by consistency, GPT-5.4 is the least expensive at $6.80 per dependable task, although only 128 tasks meet the 20/20 standard. GPT-6 Astra reaches 231 dependable tasks at $7.45, while Claude Opus 5.5 reaches the joint-highest count of 241 at $7.80.

None of these three models dominates the others. Each additional dependable task costs more. Claude Opus 5 also passes 241 tasks, but at $13.30, so Claude Opus 5.5 dominates it in this comparison. GPT-5.6 Sol, the least expensive model per individual success at $0.127, costs $9.76 per dependable task.

The least expensive way to obtain a correct answer is not necessarily the least expensive way to obtain a dependable one.

Failure signatures

ThinkingBox assigns each failed trace one deterministic diagnostic signature. Across the ablation study in Table 5 of the paper, the distribution is:

Failure signatureShare of failures
Tool usage79.9%
Wrong state updates10.3%
Incomplete user resolutions7.0%
No state-changing action2.9%

These are unweighted averages of per-model shares and observable labels, not unique causal explanations.

The practical pattern is that agents usually get far enough to attempt the workflow, then fail to recover from tool errors, failed preconditions, or empty lookups. In these traces, retry and error recovery are problems before they are model-capability problems.

Difficulty also varies by domain. Across the models listed in the table above, retail averages 59.52% pass@1, while auto insurance averages 33.83%.

The 20/20 rate can be treated as a design input rather than only as a verdict. The same signal that the benchmark grades is available in production: verify the terminal state before committing, rather than relying on the model's summary.

Tool and system errors can be classified so that retries target recoverable conditions. The tool surface can be limited to the operations required by the workflow. Changes that are difficult to reverse can require human approval. ThinkingBox has not measured the improvement from these interventions on this benchmark, leaving that question available for future testing.

How it works

ThinkingBox is the agent sandbox, while ThinkingBox-Bench is the dataset benchmark used to evaluate agents.

Each task defines:

  • A starting backend state
  • A user goal
  • The available MCP tools
  • The relevant domain policy
  • Executable checks over the terminal state

A simulated user holds private context, such as a booking reference, preference, or date of birth, and releases it only when asked.

Every attempt receives an isolated MCP session with freshly initialized state. Two attempts of the same task never share a database row or cached tool state, making 20-trial comparisons meaningful.

At the end of an attempt, a side-effect extractor identifies what actually changed. Deterministic judges compare those changes with the required final state. Any trajectory that produces the correct outcome is accepted, while wrong, missing, or extra effects are rejected.

For requirements without a clean database value, such as whether the agent disclosed that something was not guaranteed, a narrow binary rubric question handles the semantics. Of the 507 tasks, 477 are graded on state alone and 30 also include response rubrics.

The trust boundary keeps the golden state, assertions, grading internals, and credentials on the evaluator side. The model sees the task, dialogue, and tool schemas.

Run it yourself

ThinkingBox is available through Hugging Face as both the harness and the ThinkingBox-Bench dataset. The benchmark uses the OpenEnv interface, and each completed episode returns a binary pass or fail reward. The released adapter is intended for evaluation, while separate non-benchmark scenarios can use the same interface in training workflows.

Before you start

The setup was tested on Linux and WSL with Python 3.11 or later, uv, and Docker. You also need a checkout of thinkingbox-data at the pinned release and model endpoints for the agent, simulated user, and judge. One endpoint can serve all three roles.

The OpenEnv image starts only the OpenEnv API. The other services must be started separately.

Install

## 1. OpenEnv and the ThinkingBox environment
git clone https://github.com/huggingface/OpenEnv
cd OpenEnv
uv sync --project envs/thinkingbox_env --frozen

## 2. The executable benchmark at the pinned release
git clone https://github.com/microsoft/thinkingbox-data
git -C thinkingbox-data checkout thinkingbox-bench-v1.0

## 3. The ThinkingBox CLI, which provides `tb`
uv tool install "thinkingbox @ git+https://github.com/microsoft/thinkingbox"

Start Typesense

In a second terminal, start Typesense 30.1 and wait for its health check:

mkdir -p .typesense-data
docker run --rm -d --name thinkingbox-typesense \
  -p 8108:8108 \
  -v "$PWD/.typesense-data:/data" \
  typesense/typesense:30.1 \
  --data-dir /data --api-key=Fake --enable-cors
until curl -fsS http://127.0.0.1:8108/health; do sleep 1; done

Start the MCP servers

In a third terminal, start the Session Proxy and MCP servers:

cd OpenEnv
tb mcp-start --host 127.0.0.1 --port 7111 \
  --servers "$PWD/thinkingbox-data/servers/servers.yaml"
curl -fsS http://127.0.0.1:7111/health

Start the OpenEnv server

In the first terminal, start the OpenEnv server with a ThinkingBox YAML configuration that names the three models. The configuration guide provides the required format.

OPENENV_TB_CONFIG="$PWD/thinkingbox.yaml" \
uv run --project envs/thinkingbox_env --frozen server

Check readiness

Check readiness before running an evaluation. The endpoint returns 503 until its observable data, configuration, and Session Proxy checks pass. It cannot observe Typesense or live-probe every model endpoint, so those services must be checked separately.

curl -sS http://127.0.0.1:8000/ready

Score an episode

The example_usage.py script resets the environment and lists tools. For agent actions, effects, and assertions, use the packaged evaluator:

echo "- sandbox_external_retail_group1.py:test_case_ST002_001" > one_task.yaml
uv run --project envs/thinkingbox_env thinkingbox-eval \
  one_task.yaml \
  --config "$PWD/thinkingbox.yaml" \
  --output results.jsonl \
  --errors-output errors.jsonl \
  --repeat 1 --message-timeout 1800

The OpenEnv adapter writes operational failures to an errors sidecar so they can be rerun instead of being silently combined with model outcomes. A canonical result must resolve or explicitly account for those attempts. In the reported results, system errors were counted as unsuccessful trials.

Runs are gated on a pinned framework commit, a pinned data release, and a bundle hash, making a canonical result verifiable.

Where this goes next

The central contribution of this work is the evaluation environment, not only the pass@1 leaderboard.

For an agent that modifies real records:

  1. Inspect a failure. Find a run that terminated cleanly but still failed, then examine what actually changed in the database.
  2. Reproduce one task. Run it through OpenEnv with another model.
  3. Report a repeat metric and define it. State which value of k the use case requires, whether the result is best-of-k or every-of-k, and how it was calculated.

Further resources include:

ThinkingBox code is MIT-licensed. The benchmark data uses CDLA-Permissive-2.0, and the OpenEnv environment uses the OpenEnv BSD-3-Clause license.

Every task in the public benchmark is a synthetic reconstruction. The workflows and policies model real AI agent enterprise patterns, but the customers are not real.