HealthBench Hard Leaderboard
Last updated September 8, 2026 · 16 models evaluated
Muse Spark leads the HealthBench Hard leaderboard at 0.428, ahead of GPT-6 Astra (0.363) and GPT-5 (0.347). HealthBench Hard is the hardest slice of OpenAI's HealthBench benchmark: the 1,000 health conversations frontier models handled worst, graded against physician-written rubrics on a 0 to 1 scale. The best score at the benchmark's May 2025 release was o3's 0.320.
Leaderboard
Sources1,000-conversation set, scores as published by each document. Updated September 8, 2026.
Muse Spark
GPT-6 Astra
GPT-5
GPT-5.2
GPT-5.6 Sol
GPT-5.6 Terra
GPT-5.6 Luna
GPT-5.5
GPT-5.6 Sol (August)
GPT OSS 120B
GPT-5.4
GPT-5.6 Luna (August)
GPT-5.3 Chat
GPT-5.1
GPT-5.5 Instant
GPT OSS 20B
best frontier score at launch 0.320
Full ranking
Sources| # | model | score | size | context | cost in / out per 1M | license | |
|---|---|---|---|---|---|---|---|
| 1 | Muse Spark Meta | 0.428 | — | 1.0M | $1.25 / $4.25 | proprietary | |
| 2 | GPT-6 Astra OpenAI | 0.363 | — | 1.05M | $10.00 / $50.00 | proprietary | |
| 3 | GPT-5 OpenAI | 0.347 | — | 400K | $1.25 / $10.00 | proprietary | |
| 4 | GPT-5.2 OpenAI | 0.343 | — | 400K | $1.75 / $14.00 | proprietary | |
| 5 | GPT-5.6 Sol OpenAI | 0.331 | — | 1.1M | $5.00 / $30.00 | proprietary | |
| 6 | GPT-5.6 Terra OpenAI | 0.327 | — | 1.1M | $2.00 / $12.00 | proprietary | |
| 7 | GPT-5.6 Luna OpenAI | 0.320 | — | 1.1M | $0.20 / $1.20 | proprietary | |
| 8 | GPT-5.5 OpenAI | 0.315 | — | 1.05M | $5.00 / $30.00 | proprietary | |
| 9 | GPT-5.6 Sol (August) OpenAI | 0.314 | — | 1.1M | $5.00 / $30.00 | proprietary | |
| 10 | GPT OSS 120B OpenAI | 0.300 | 117B | 131K | — | open | |
| 11 | GPT-5.4 OpenAI | 0.291 | — | 1.05M | $2.50 / $15.00 | proprietary | |
| 12 | GPT-5.6 Luna (August) OpenAI | 0.287 | — | 1.1M | $0.20 / $1.20 | proprietary | |
| 13 | GPT-5.3 Chat OpenAI | 0.259 | — | 128K | $1.75 / $14.00 | proprietary | |
| 14 | GPT-5.1 OpenAI | 0.254 | — | 400K | $1.25 / $10.00 | proprietary | |
| 15 | GPT-5.5 Instant OpenAI | 0.229 | — | 400K | $5.00 / $30.00 | proprietary | |
| 16 | GPT OSS 20B OpenAI | 0.108 | 21B | 131K | — | open | |
Scores are rubric credit across the 1,000-conversation set, copied exactly as the document under each model name prints them. The line under a name says which document, and links to its entry on the sources page; "source pending" marks a row whose document has not been located or verified yet. Prices are vendor list rates per 1M tokens, checked September 8, 2026. GPT-5.6 models charge higher rates above 272K input tokens. GPT OSS models are open weights without vendor list pricing. Every model name links through to its full result.
What the benchmark measures
HealthBench Hard keeps only the failures. OpenAI's HealthBench grades models on 5,000 realistic health conversations, from emergency triage to global health, against 48,562 rubric criteria written by 262 physicians across 60 countries. The Hard subset is the 1,000 conversations on which frontier models scored worst at release, which makes it the part of HealthBench that still separates current models. Two open-weights models, GPT OSS 120B and GPT OSS 20B, sit on the board alongside the closed frontier. The full construction is on the benchmark page.
How to read the scores
A score is rubric points earned over maximum positive points, clipped to 0 to 1. Low numbers are the design, not a flaw: the subset exists because models fail on it, the field's best result is 0.428, and 1.0 is nowhere in sight. The numbers here are compiled, not measured by this site: each row is read from a published document, most often the vendor's own system card, and the sources page shows the quote it came from. Two documents can grade the same model differently, for reasons the methodology page lays out, so a gap between two rows read from different documents carries that caveat.
Head to head
The pairings worth deciding between get full pages; the compare page holds the complete score-difference matrix.
- 0.428 vs 0.331 · Muse Spark by 0.097
- 0.331 vs 0.327 · GPT-5.6 Sol by 0.004
- 0.320 vs 0.300 · GPT-5.6 Luna by 0.020
- 0.300 vs 0.108 · GPT OSS 120B by 0.192
- 0.347 vs 0.331 · GPT-5 by 0.016
- 0.320 vs 0.259 · GPT-5.6 Luna by 0.061
Data
Everything on this page ships as JSON and CSV under CC BY 4.0, at URLs that never move. The data page has the citation format and version history.