GPT-5.6 Luna (August) on HealthBench Hard
rank 12 of 16 · updated September 8, 2026via GPT-5.6 - August Updates (system card addendum)
On the 16-model HealthBench Hard board, GPT-5.6 Luna (August) holds rank 12 with a score of 0.287. OpenAI's high-volume 5.6 tier, cut to $0.20 per million input tokens in the July 30, 2026 repricing. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.
Score and API facts
| rank | 12 of 16 |
|---|---|
| score | 0.287 |
| source | GPT-5.6 - August Updates (system card addendum) (p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)), vendor-reported |
| configuration | ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars) |
| lab | OpenAI |
| context window | 1.1M tokens |
| API price per 1M tokens | $0.20 in / $1.20 out |
| license | proprietary |
| released | 2026-07-09 |
Where it sits
Muse Spark tops the board at 0.428, which puts GPT-5.6 Luna (August) 0.141 off the lead. One place up is GPT-5.4 at 0.291. One place down is GPT-5.3 Chat at 0.259. Rows on this board are compiled from published documents, so a gap between two models is exact only when both numbers came from the same document under the same settings; the sources page shows which document each row came from.
What does GPT-5.6 Luna (August) score on HealthBench Hard?
As of September 8, 2026, GPT-5.6 Luna (August) scores 0.287 on HealthBench Hard, 12 of 16 models on the board. The number is read from the GPT-5.6 - August Updates (system card addendum) (p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)), vendor-reported.
What does GPT-5.6 Luna (August) cost per million tokens?
OpenAI lists GPT-5.6 Luna (August) at $0.20 per million input tokens and $1.20 per million output tokens.
Head to head
The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.
- GPT-5.6 Luna (August) vs Muse Spark0.287 vs 0.428 · Muse Spark by 0.141
- GPT-5.6 Luna (August) vs GPT-6 Astra0.287 vs 0.363 · GPT-6 Astra by 0.076
- GPT-5.6 Luna (August) vs GPT-50.287 vs 0.347 · GPT-5 by 0.060
- GPT-5.6 Luna (August) vs GPT-5.20.287 vs 0.343 · GPT-5.2 by 0.056
- GPT-5.6 Luna (August) vs GPT-5.6 Sol0.287 vs 0.331 · GPT-5.6 Sol by 0.044
- GPT-5.6 Luna (August) vs GPT-5.6 Terra0.287 vs 0.327 · GPT-5.6 Terra by 0.040
- GPT-5.6 Luna (August) vs GPT-5.6 Luna0.287 vs 0.320 · GPT-5.6 Luna by 0.033
- GPT-5.6 Luna (August) vs GPT-5.50.287 vs 0.315 · GPT-5.5 by 0.028
- GPT-5.6 Luna (August) vs GPT-5.6 Sol (August)0.287 vs 0.314 · GPT-5.6 Sol (August) by 0.027
- GPT-5.6 Luna (August) vs GPT OSS 120B0.287 vs 0.300 · GPT OSS 120B by 0.013
- GPT-5.6 Luna (August) vs GPT-5.40.287 vs 0.291 · GPT-5.4 by 0.004
- GPT-5.6 Luna (August) vs GPT-5.3 Chat0.287 vs 0.259 · GPT-5.6 Luna (August) by 0.028
- GPT-5.6 Luna (August) vs GPT-5.10.287 vs 0.254 · GPT-5.6 Luna (August) by 0.033
- GPT-5.6 Luna (August) vs GPT-5.5 Instant0.287 vs 0.229 · GPT-5.6 Luna (August) by 0.058
- GPT-5.6 Luna (August) vs GPT OSS 20B0.287 vs 0.108 · GPT-5.6 Luna (August) by 0.179
How numbers are read from their documents is on the methodology page, the document for this row is on the sources page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.