HealthBench Hard

OpenAI logoGPT-5 on HealthBench Hard

rank 9 of 9 · updated August 16, 2026

On the 9-model HealthBench Hard board, GPT-5 holds rank 9 with a score of 0.016. OpenAI's August 2025 flagship, still served at unchanged prices but superseded by the GPT-5.6 series. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.

Score and API facts

rank9 of 9
score0.016
labOpenAI
context window400K tokens
API price per 1M tokens$1.25 in / $10.00 out
licenseproprietary
released2025-08-07

Where it sits

Muse Spark tops the board at 0.428, which puts GPT-5 0.412 off the lead. One place up is GPT OSS 20B at 0.108. Every row on the board comes out of one run, so a gap between two models is measured on the same conversations under the same grader.

What does GPT-5 score on HealthBench Hard?

As of August 16, 2026, GPT-5 scores 0.016 on HealthBench Hard, 9 of 9 models on the board.

What does GPT-5 cost per million tokens?

OpenAI lists GPT-5 at $1.25 per million input tokens and $10.00 per million output tokens.

Head to head

The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.

How the conversations are graded is on the methodology page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.