← Benchmarks

Sequential clinical triage reasoning

TriageBench

Can a model ask the right question, stop at the right time and recommend the right level of care before the full case is known?

Leaderboard

Leaderboard

RankModelScore
01GPT 5.6 Sol (Max reasoning)45.0%
02GPT 5.6 Sol (default)39.5%
03Gemini 3.5 Flash35.9%
04Claude Fable 5 (Adaptive/Max)34.8%
05Gemini 3.6 Flash34.0%
06GPT 5.6 Terra (Max reasoning)34.0%
07GPT 5.6 Luna (Max reasoning)31.4%
08GPT 5.5 (xHigh reasoning)31.0%
09GPT 5.4 (xHigh reasoning)29.6%
10Claude Fable 5 (default)29.5%
11GPT 5.5 (default)28.5%
12Claude Opus 5 (Adaptive/Max)27.3%
13GPT 5.6 Terra (default)26.6%
14Kimi K326.6%
15Gemini 3.1 Pro26.1%
Score vs. cost per trial
Cost per trial (USD, log scale)Score →Mistral Large 3: 9% at $0.0008MistralGemini 3.5 Flash Lite: 10.2% at $0.0011Gemini LiteGPT 5.4: 14.1% at $0.0042GPT 5.4GPT 5.6 Luna: 21.4% at $0.012LunaGemini 3.6 Flash: 34% at $0.037Gemini 3.6GPT 5.6 Sol: 39.5% at $0.057SolGPT 5.5: 28.5% at $0.085GPT 5.5Gemini 3.1 Pro: 26.1% at $0.1Gemini ProGPT 5.6 Sol (Max reasoning): 45% at $0.29Sol MaxClaude Fable 5 (Adaptive/Max): 34.8% at $0.35Fable Max
Score vs. generated tokens
Generated tokens (log scale)Score →Mistral Large 3: 9% at 285 tokensMistralGemini 3.5 Flash Lite: 10.2% at 286 tokensGemini LiteGPT 5.4: 14.1% at 130 tokensGPT 5.4GPT 5.6 Luna: 21.4% at 1800 tokensLunaGemini 3.6 Flash: 34% at 4600 tokensGemini 3.6GPT 5.6 Sol: 39.5% at 1800 tokensSolGPT 5.5: 28.5% at 2700 tokensGPT 5.5Gemini 3.1 Pro: 26.1% at 8200 tokensGemini ProGPT 5.6 Sol (Max reasoning): 45% at 9600 tokensSol MaxClaude Fable 5 (Adaptive/Max): 34.8% at 6600 tokensFable Max

Failure analysis

Can models make safe decisions before the full case is known?

Expand the examples to see how a fluent response can still fail the underlying triage decision.

Stops before asking the decisive question
Task prompt

A patient reports a painful mouth ulcer. Duration and relevant history have not yet been disclosed.

Sample response

This is likely a minor ulcer. Use an over-the-counter treatment and monitor it at home.

Rubric failure reason

The model routes before asking how long the ulcer has been present. Duration can change the appropriate care pathway.

Keeps questioning after an emergency threshold is crossed
Task prompt

During a chest-pain pathway, the patient discloses a safety-critical feature that requires immediate escalation.

Sample response

Before deciding, I need to ask several more questions about the character and timing of the pain.

Rubric failure reason

The model recognises uncertainty but fails to stop. Once the reviewed emergency threshold is crossed, additional questioning delays the required action.

Chooses the right care level for the wrong reason
Task prompt

The visible facts support same-day clinical assessment, but do not support a diagnosis.

Sample response

You should seek same-day care because this is probably a bacterial infection.

Rubric failure reason

The care level is acceptable, but the explanation invents a diagnosis. TriageBench grades structured action and patient-facing language separately.

Research article

Why clinical triage needs a sequential benchmark.

Read the methodology and examples →