Benchmarks
Measure the decisions that matter.
Models are easy to impress in a demo. Our benchmarks test whether they can make defensible decisions when information is incomplete and the cost of error is real.
Sequential clinical triage reasoning
TriageBench
Can a model ask the right question, stop at the right time and recommend the right level of care before the full case is known?
Leaderboard
RankModelScore
01GPT 5.6 Sol (Max reasoning)45.0%
02GPT 5.6 Sol (default)39.5%
03Gemini 3.5 Flash35.9%