Benchmarks

Measure the decisions that matter.

Models are easy to impress in a demo. Our benchmarks test whether they can make defensible decisions when information is incomplete and the cost of error is real.

Sequential clinical triage reasoning

TriageBench

Can a model ask the right question, stop at the right time and recommend the right level of care before the full case is known?

Public v0.2
RankModelScore
01GPT 5.5 (xHigh reasoning)50.6%
02Claude Opus 4.8 (default)45.0%
03Qwen 3.5 Plus44.0%