← The Daily Chart
Agent Reliability

A Top AI Agent Passes 65% of Attempts but Only 25% of Tasks Every Time

Microsoft ThinkingBox-Bench: GPT-5.4 on 507 business tasks

A Top AI Agent Passes 65% of Attempts but Only 25% of Tasks Every Time

A Top AI Agent Passes 65% of Attempts but Only 25% of Tasks Every Time
Share this chartXLinkedIn

Why it matters: On Microsoft’s ThinkingBox-Bench, leading model GPT-5.4 averaged 65.36% pass@1 across 507 executable business tasks, but passed every one of 20 trials on only 25.25% of tasks (pass^20), exposing a steep discovery-versus-reliability gap for production agents.

View source →
Source
Microsoft
Unit
% of 507 tasks
Data as of
August 19, 2026
Scope
Microsoft ThinkingBox-Bench evaluation of GPT-5.4 across five business domains