HealthBench Hard

HealthBench Hard Leaderboard

Last updated September 8, 2026 · 16 models evaluated

Muse Spark leads the HealthBench Hard leaderboard at 0.428, ahead of GPT-6 Astra (0.363) and GPT-5 (0.347). HealthBench Hard is the hardest slice of OpenAI's HealthBench benchmark: the 1,000 health conversations frontier models handled worst, graded against physician-written rubrics on a 0 to 1 scale. The best score at the benchmark's May 2025 release was o3's 0.320.

Leaderboard

Sources

1,000-conversation set, scores as published by each document. Updated September 8, 2026.

  1. 0.428
    Meta logoMuse Spark
  2. 0.363
    OpenAI logoGPT-6 Astra
  3. 0.347
    OpenAI logoGPT-5
  4. 0.343
    OpenAI logoGPT-5.2
  5. 0.331
    OpenAI logoGPT-5.6 Sol
  6. 0.327
    OpenAI logoGPT-5.6 Terra
  7. 0.320
    OpenAI logoGPT-5.6 Luna
  8. 0.315
    OpenAI logoGPT-5.5
  9. 0.314
    OpenAI logoGPT-5.6 Sol (August)
  10. 0.300
    OpenAI logoGPT OSS 120B
  11. 0.291
    OpenAI logoGPT-5.4
  12. 0.287
    OpenAI logoGPT-5.6 Luna (August)
  13. 0.259
    OpenAI logoGPT-5.3 Chat
  14. 0.254
    OpenAI logoGPT-5.1
  15. 0.229
    OpenAI logoGPT-5.5 Instant
  16. 0.108
    OpenAI logoGPT OSS 20B

best frontier score at launch 0.320

Full ranking

Sources
#modelscoresizecontextcost in / out per 1Mlicense
1Meta logoMuse Spark Meta0.428—1.0M$1.25 / $4.25proprietary
2OpenAI logoGPT-6 Astra OpenAI0.363—1.05M$10.00 / $50.00proprietary
3OpenAI logoGPT-5 OpenAI0.347—400K$1.25 / $10.00proprietary
4OpenAI logoGPT-5.2 OpenAI0.343—400K$1.75 / $14.00proprietary
5OpenAI logoGPT-5.6 Sol OpenAI0.331—1.1M$5.00 / $30.00proprietary
6OpenAI logoGPT-5.6 Terra OpenAI0.327—1.1M$2.00 / $12.00proprietary
7OpenAI logoGPT-5.6 Luna OpenAI0.320—1.1M$0.20 / $1.20proprietary
8OpenAI logoGPT-5.5 OpenAI0.315—1.05M$5.00 / $30.00proprietary
9OpenAI logoGPT-5.6 Sol (August) OpenAI0.314—1.1M$5.00 / $30.00proprietary
10OpenAI logoGPT OSS 120B OpenAI0.300117B131K—open
11OpenAI logoGPT-5.4 OpenAI0.291—1.05M$2.50 / $15.00proprietary
12OpenAI logoGPT-5.6 Luna (August) OpenAI0.287—1.1M$0.20 / $1.20proprietary
13OpenAI logoGPT-5.3 Chat OpenAI0.259—128K$1.75 / $14.00proprietary
14OpenAI logoGPT-5.1 OpenAI0.254—400K$1.25 / $10.00proprietary
15OpenAI logoGPT-5.5 Instant OpenAI0.229—400K$5.00 / $30.00proprietary
16OpenAI logoGPT OSS 20B OpenAI0.10821B131K—open

Scores are rubric credit across the 1,000-conversation set, copied exactly as the document under each model name prints them. The line under a name says which document, and links to its entry on the sources page; "source pending" marks a row whose document has not been located or verified yet. Prices are vendor list rates per 1M tokens, checked September 8, 2026. GPT-5.6 models charge higher rates above 272K input tokens. GPT OSS models are open weights without vendor list pricing. Every model name links through to its full result.

16
models evaluated
2
labs represented
1,000
hardest conversations
5,000
conversation source set
262
rubric-writing physicians
60
countries represented

What the benchmark measures

HealthBench Hard keeps only the failures. OpenAI's HealthBench grades models on 5,000 realistic health conversations, from emergency triage to global health, against 48,562 rubric criteria written by 262 physicians across 60 countries. The Hard subset is the 1,000 conversations on which frontier models scored worst at release, which makes it the part of HealthBench that still separates current models. Two open-weights models, GPT OSS 120B and GPT OSS 20B, sit on the board alongside the closed frontier. The full construction is on the benchmark page.

How to read the scores

A score is rubric points earned over maximum positive points, clipped to 0 to 1. Low numbers are the design, not a flaw: the subset exists because models fail on it, the field's best result is 0.428, and 1.0 is nowhere in sight. The numbers here are compiled, not measured by this site: each row is read from a published document, most often the vendor's own system card, and the sources page shows the quote it came from. Two documents can grade the same model differently, for reasons the methodology page lays out, so a gap between two rows read from different documents carries that caveat.

Head to head

The pairings worth deciding between get full pages; the compare page holds the complete score-difference matrix.

Data

Everything on this page ships as JSON and CSV under CC BY 4.0, at URLs that never move. The data page has the citation format and version history.