← Benchmarks

Sequential clinical triage reasoning

TriageBench

Can a model ask the right question, stop at the right time and recommend the right level of care before the full case is known?

Public v0.2

Public leaderboard

Exact-decision accuracy on the transparent 100-task development set. The random-choice baseline is 25%.

RankModel configurationProviderScoreCases evaluated
01GPT 5.5 (xHigh reasoning)Azure50.6%83
02Claude Opus 4.8 (default)Amazon Bedrock45.0%100
03Qwen 3.5 PlusAlibaba44.0%100
04Claude Opus 5 (default)Amazon Bedrock43.0%100
05Gemini 3.1 ProGoogle42.0%100
06GPT 5.6 Sol (default)Azure42.0%100
07Qwen 3.7 PlusAlibaba41.0%100
08Qwen 3.8 MaxAlibaba41.0%100
09Claude Opus 4.7 (default)Amazon Bedrock40.0%100
10Claude Sonnet 5 (default)Amazon Bedrock40.0%100
11GPT 5.6 Sol (Max reasoning)Azure40.0%100
12Muse Spark 1.1 (xHigh reasoning)Meta40.0%100
13GPT 5.4 (default)Azure39.0%100
14GPT 5.5 (default)Azure39.0%100
15GPT 5.6 Luna (default)Azure39.0%100
16GPT 5.6 Terra (Max reasoning)Azure39.0%100
17Gemini 3.5 Flash-Lite (default)Google38.0%100
18Grok 4.5 (default)xAI38.0%100
19Mistral Large 3Mistral38.0%100
20Muse Spark 1.1 (default)Meta37.0%100
21Muse Spark 1.2 (default)Meta37.0%100
22Gemini 3.6 Flash (default)Google36.0%100
23Grok 4.3 (default)xAI36.0%100
24Kimi K3 (Max reasoning)Moonshot AI36.0%100
25GPT 5.6 Terra (default)Azure35.0%100
26Gemini 3.5 Flash (default)Google34.0%100
27GPT 5.6 Luna (Max reasoning)Azure34.0%100
28Muse Spark 1.2 (xHigh reasoning)Meta34.0%100
29Kimi K2.5Moonshot AI31.0%100
30Kimi K2.6Moonshot AI26.0%100
31InklingDeepInfra20.0%20

One deterministic run per configuration. Scoring uses exact option match, with no model judge and no partial credit. Cases evaluated states the number of benchmark tasks completed by each configuration. View the complete run metadata ↗.

Public release

100 tasks. 50 question decisions. 50 assessment decisions.

The public release includes ten matched pairs for testing whether one changed patient answer changes the expected next action.

Research article

Why clinical triage needs a sequential benchmark.

Read the methodology and examples →