GPT-5.6 Luna vs GPT-5.3 Chat on HealthBench Hard
updated August 16, 2026
A gap of 0.061 separates these two on the 1,000 hardest HealthBench conversations: GPT-5.6 Luna at 0.320, GPT-5.3 Chat at 0.259. The table adds what each one costs at list rates.
Side by side
| score | 0.320 | 0.259 |
|---|---|---|
| rank | 4 of 9 | 6 of 9 |
| context window | 1.1M | 128K |
| price per 1M tokens, in / out | $0.20 / $1.20 | $1.75 / $14.00 |
| 1,000-exchange workload | $1.24 | $13.30 |
| released | 2026-07-09 | 2026-03-05 |
| license | proprietary | proprietary |
The workload row prices 1,000 exchanges of 2,000 input and 700 output tokens each, at the list rates current on August 16, 2026. GPT-5.6 models charge higher rates above 272K input tokens. GPT OSS models are open weights without vendor list pricing.
Reading the matchup
A generation cleanup: Luna scores 0.061 higher, costs about a ninth as much per input token, and GPT-5.3 Chat is already deprecated, with OpenAI pointing API users to the 5.6 series. The only reason to run 5.3 Chat on this workload is an integration that has not migrated yet.
Which scores higher on HealthBench Hard, GPT-5.6 Luna or GPT-5.3 Chat?
GPT-5.6 Luna. On the 1,000-conversation set it scores 0.320 to GPT-5.3 Chat's 0.259, a margin of 0.061 as of August 16, 2026.
Which is cheaper to run, GPT-5.6 Luna or GPT-5.3 Chat?
GPT-5.6 Luna. The same workload of 1,000 exchanges (2,000 input and 700 output tokens each) comes to $1.24 on it and $13.30 on the other.
Related comparisons
- 0.320 vs 0.300
Each model's full page: GPT-5.6 Luna and GPT-5.3 Chat. Every other pairing lives in the score-difference matrix.