Benchmarks

Measure the decisions that matter.

Models are easy to impress in a demo. Our benchmarks test whether they can make defensible decisions when information is incomplete and the cost of error is real.

Sequential clinical triage reasoning

TriageBench

Can a model ask the right question, stop at the right time and recommend the right level of care before the full case is known?

Leaderboard
RankModelScore
01GPT 5.6 Sol (Max reasoning)45.0%
02GPT 5.6 Sol (default)39.5%
03Gemini 3.5 Flash35.9%