HealthBench Hard

FAQ

The questions this leaderboard actually gets, answered in a few sentences each. For the detail behind any answer, use the pages linked at the bottom.

Which AI model scores highest on HealthBench Hard?

Muse Spark scores highest at 0.428 as of September 8, 2026, ahead of GPT-6 Astra at 0.363 and GPT-5 at 0.347. 16 models are on the leaderboard.

What is HealthBench Hard?

HealthBench Hard is the hardest slice of HealthBench, the health benchmark OpenAI released in May 2025. From 5,000 realistic health conversations it keeps the 1,000 that frontier models handled worst, each graded against a physician-written rubric on a 0 to 1 scale.

How is HealthBench Hard different from HealthBench?

It is the same conversations, rubrics, and grading, restricted to the hardest fifth. Five frontier models scored every conversation in the parent set, and the 1,000 with the lowest average score became the Hard subset. Scores on the parent set run far higher; the Hard slice is the part that still separates current models.

How are the scores graded?

A model grader (GPT-4.1 in the benchmark paper) judges each rubric criterion as met or not met. Criteria carry physician-assigned weights from -10 to +10, the score is earned points over maximum positive points, and the overall number is the average across conversations, clipped to 0 to 1.

Where do the numbers on this leaderboard come from?

From published documents: OpenAI's system cards and update reports, other vendors' model cards, and named leaderboards. Arcophos does not currently run HealthBench Hard itself. Each row links to the document it was read from on the sources page, with the quote that carries the number; 16 of 16 rows have a located document as of September 8, 2026, and the rest are marked "source pending".

Why does a score here differ from the one in another report?

Because documents grade under different settings. Length adjustment, reasoning effort, deployment setting (API against consumer product), and grader version each move a HealthBench Hard number, and a document does not always state all four. The sources page shows the exact quote behind each row so the setting can be checked.

Why are the scores so low?

By construction. The subset keeps only conversations the frontier failed on: the best score at release was o3's 0.320, three then-current models scored exactly zero, and the current leader stands at 0.428. Low absolute numbers are the point; they are the remaining headroom.

What is the best open-weights model on HealthBench Hard?

GPT OSS 120B, at 0.300 and rank 10 of 16. It is an Apache 2.0 release with 117B parameters, which makes it the strongest option for deployments that require local weights.

Which model gives the most score per dollar?

GPT-5.6 Luna, at 0.320 for $0.20 per million input tokens and $1.20 per million output tokens, scores within 0.108 of the leader at a fraction of the price of every model above it.

How often is the leaderboard updated?

Whenever a new system card, model card, or leaderboard entry reports a HealthBench Hard number, or a document is located for a pending row. Every change gets a dated entry on the updates page; the current table of 16 models is from September 8, 2026.

Can I download the results?

Yes: JSON and CSV files live at stable URLs under CC BY 4.0, and the data page carries a suggested citation format and the version history.

Start with the leaderboard for the full ranking. The benchmark page covers how the subset was built, the methodology page covers how numbers are read from their documents, the sources page lists those documents, and the data page has the downloads.