Tools & releasesImpact: High
Microsoft and Hugging Face Release ThinkingBox Benchmark for Stateful Agents
Microsoft and Hugging Face launched ThinkingBox via OpenEnv, a benchmark testing agents across 507 stateful business workflows repeated 20 times. Results reveal that 67.24% of failed trials terminated cleanly with valid tool calls while still corrupting backend database records.

Why it matters
Evaluate your production agents against actual database state mutations across repeated runs instead of trusting self-reported model completions.
TL;DR
- 01Over two-thirds of failed agent operations terminate cleanly without throwing runtime tool errors.
- 02Single-run pass@1 benchmarks fail to predict production consistency across repeated executions.
- 03Claude Opus 5 and 5.5 demonstrated 47.53% consistency across 20 runs, outperforming high-pass@1 open-weight models.
Key facts
- Benchmark Scope
- 507 stateful tasks evaluated 20 times each
- Clean Failures
- 67.24% of failed attempts reported zero tool errors
- Opus 5 vs 5.5 Consistency
- Both solved exactly 241/507 tasks across all 20 runs
- Lowest Cost Frontier
- $0.127 per successful task (GPT-5.6 Sol)
The Database Disagreement Gap. Most agent evaluations grade final chat responses or tool call syntax. Microsoft ThinkingBox connects agents to isolated Model Context Protocol tool sessions across 507 stateful tasks and audits actual backend database changes. In an evaluation of 121,680 trials across 12 frontier models, 79,853 attempts failed backend executable checks. Of those failures, 67.24% reported clean execution with zero tool errors. However, database audits revealed incorrect field values in 77.61% of failed attempts, unintended side effects in 43.30%, and omitted required changes in 25.36%. ### The Consistency Collapse. Running every workflow 20 consecutive times exposes a massive divergence between single-attempt capability (pass@1) and production reliability (pass@20). Kimi-K3 solved 93.89% of tasks at least once, leading retail workflows at 82.24% pass@1, but completed all 20 runs on only 13.41% of tasks (68 of 507). Conversely, Claude Opus 5 solved fewer tasks overall (79.09%) but passed 47.53% consistently (241 tasks). Claude Opus 5.5 matched Opus 5 at exactly 241 perfect runs despite having a higher pass@1 average. ### Frontier Efficiency Metrics. ThinkingBox measures cost per successful task attempt to map real-world agent economics. GPT-5.6 Sol established the lowest cost per successful task at $0.127, while GPT-5.4 delivered a 65.36% pass@1 score at $0.131 per success. Developers can now run ThinkingBox locally or in cloud sandboxes through OpenEnv on Hugging Face.
Try it in 2 minutes
python -m openenv.run --benchmark thinkingbox --tasks 507 --repeats 20bash
✓ When to use
- Auditing stateful customer support, billing, or inventory agents that mutate databases.
- Benchmarking agent orchestration reliability across repeated production workflows.
✕ When NOT to use
- Stateless coding assistance tasks where human review immediately inspects code diffs.
- Single-turn information extraction queries without side effects.
What to do today
- Run database assertion checks on terminal state rather than relying on agent self-reported task completion.
- Test critical agentic workflows across at least 10 to 20 repeat runs before production deployment.
- Benchmark cost per successful task attempt instead of relying purely on input/output token pricing.