GPT-5 vs GPT-5.6 Sol on HealthBench Hard
updated September 8, 2026
A gap of 0.016 separates these two on the 1,000 hardest HealthBench conversations: GPT-5 at 0.347, GPT-5.6 Sol at 0.331. The table adds what each one costs at list rates.
Side by side
| score | 0.347 | 0.331 |
|---|---|---|
| rank | 3 of 16 | 5 of 16 |
| context window | 400K | 1.1M |
| price per 1M tokens, in / out | $1.25 / $10.00 | $5.00 / $30.00 |
| 1,000-exchange workload | $9.50 | $31.00 |
| released | 2025-08-07 | 2026-07-09 |
| license | proprietary | proprietary |
The workload row prices 1,000 exchanges of 2,000 input and 700 output tokens each, at the list rates current on September 8, 2026. GPT-5.6 models charge higher rates above 272K input tokens. GPT OSS models are open weights without vendor list pricing.
Reading the matchup
The pairing that changed most when this board moved to documented sources. GPT-5 now sits 0.016 ahead of GPT-5.6 Sol, and both numbers come from the same table in OpenAI's GPT-5.6 system card, read under the same length-adjusted protocol, so the comparison is as clean as a compiled board gets. A gap that size on 1,000 conversations is within measurement variation: treat the two as level on this benchmark and decide on price, where GPT-5 lists at a quarter of Sol's input rate. An earlier version of this page showed GPT-5 at 0.016; that figure was the model's hallucination rate from a different table, not a HealthBench Hard score, and the updates page records the correction.
Which scores higher on HealthBench Hard, GPT-5 or GPT-5.6 Sol?
GPT-5. On the 1,000-conversation set it scores 0.347 to GPT-5.6 Sol's 0.331, a margin of 0.016 as of September 8, 2026.
Which is cheaper to run, GPT-5 or GPT-5.6 Sol?
GPT-5. The same workload of 1,000 exchanges (2,000 input and 700 output tokens each) comes to $9.50 on it and $31.00 on the other.
Related comparisons
- 0.428 vs 0.331
- 0.331 vs 0.327
Each model's full page: GPT-5 and GPT-5.6 Sol. Every other pairing lives in the score-difference matrix.