HealthBench Hard

OpenAI logoGPT-5.6 Luna on HealthBench Hard

rank 4 of 9 · updated August 16, 2026

On the 9-model HealthBench Hard board, GPT-5.6 Luna holds rank 4 with a score of 0.320. OpenAI's high-volume 5.6 tier, cut to $0.20 per million input tokens in the July 30, 2026 repricing. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.

Score and API facts

rank4 of 9
score0.320
labOpenAI
context window1.1M tokens
API price per 1M tokens$0.20 in / $1.20 out
licenseproprietary
released2026-07-09

Where it sits

Muse Spark tops the board at 0.428, which puts GPT-5.6 Luna 0.108 off the lead. One place up is GPT-5.6 Terra at 0.327. One place down is GPT OSS 120B at 0.300. Every row on the board comes out of one run, so a gap between two models is measured on the same conversations under the same grader.

What does GPT-5.6 Luna score on HealthBench Hard?

As of August 16, 2026, GPT-5.6 Luna scores 0.320 on HealthBench Hard, 4 of 9 models on the board.

What does GPT-5.6 Luna cost per million tokens?

OpenAI lists GPT-5.6 Luna at $0.20 per million input tokens and $1.20 per million output tokens.

Head to head

The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.

How the conversations are graded is on the methodology page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.