HealthBench Hard

FAQ

The questions this leaderboard actually gets, answered in a few sentences each. For the detail behind any answer, use the pages linked at the bottom.

Which AI model scores highest on HealthBench Hard?

Muse Spark scores highest at 0.428 as of August 16, 2026, ahead of GPT-5.6 Sol at 0.331 and GPT-5.6 Terra at 0.327. 9 models are on the leaderboard.

What is HealthBench Hard?

HealthBench Hard is the hardest slice of HealthBench, the health benchmark OpenAI released in May 2025. From 5,000 realistic health conversations it keeps the 1,000 that frontier models handled worst, each graded against a physician-written rubric on a 0 to 1 scale.

How is HealthBench Hard different from HealthBench?

It is the same conversations, rubrics, and grading, restricted to the hardest fifth. Five frontier models scored every conversation in the parent set, and the 1,000 with the lowest average score became the Hard subset. Scores on the parent set run far higher; the Hard slice is the part that still separates current models.

How are the scores graded?

A model grader (GPT-4.1) judges each rubric criterion as met or not met. Criteria carry physician-assigned weights from -10 to +10, the score is earned points over maximum positive points, and the overall number is the average across conversations, clipped to 0 to 1.

Why are the scores so low?

By construction. The subset keeps only conversations the frontier failed on: the best score at release was o3's 0.320, three then-current models scored exactly zero, and the current leader stands at 0.428. Low absolute numbers are the point; they are the remaining headroom.

What is the best open-weights model on HealthBench Hard?

GPT OSS 120B, at 0.300 and rank 5 of 9. It is an Apache 2.0 release with 117B parameters, which makes it the strongest option for deployments that require local weights.

Which model gives the most score per dollar?

GPT-5.6 Luna, at 0.320 for $0.20 per million input tokens and $1.20 per million output tokens, scores within 0.108 of the leader at a fraction of the price of every model above it.

How often is the leaderboard updated?

Whenever a frontier model ships or changes materially, the board is re-run. Every change gets a dated entry on the updates page; the current table of 9 models is from August 16, 2026.

Can I download the results?

Yes: JSON and CSV files live at stable URLs under CC BY 4.0, and the data page carries a suggested citation format and the version history.

Start with the leaderboard for the full ranking. The benchmark page covers how the subset was built, the methodology page covers the run protocol, and the data page has the downloads.