HealthBench Hard

Compare models

updated September 8, 2026

Matchups that decide deployments get a full page each, with the score gap, list prices, and a costed usage workload. The matrix below covers every remaining pairing.

Head-to-head pages

Score-difference matrix

Each cell is the row model's score minus the column model's, so positive means the row wins. Column numbers follow rank.

model12345678910111213141516
1. Muse Spark·+0.065+0.081+0.085+0.097+0.101+0.108+0.113+0.114+0.128+0.137+0.141+0.169+0.174+0.199+0.320
2. GPT-6 Astra-0.065·+0.016+0.020+0.032+0.036+0.043+0.048+0.049+0.063+0.072+0.076+0.104+0.109+0.134+0.255
3. GPT-5-0.081-0.016·+0.004+0.016+0.020+0.027+0.032+0.033+0.047+0.056+0.060+0.088+0.093+0.118+0.239
4. GPT-5.2-0.085-0.020-0.004·+0.012+0.016+0.023+0.028+0.029+0.043+0.052+0.056+0.084+0.089+0.114+0.235
5. GPT-5.6 Sol-0.097-0.032-0.016-0.012·+0.004+0.011+0.016+0.017+0.031+0.040+0.044+0.072+0.077+0.102+0.223
6. GPT-5.6 Terra-0.101-0.036-0.020-0.016-0.004·+0.007+0.012+0.013+0.027+0.036+0.040+0.068+0.073+0.098+0.219
7. GPT-5.6 Luna-0.108-0.043-0.027-0.023-0.011-0.007·+0.005+0.006+0.020+0.029+0.033+0.061+0.066+0.091+0.212
8. GPT-5.5-0.113-0.048-0.032-0.028-0.016-0.012-0.005·+0.001+0.015+0.024+0.028+0.056+0.061+0.086+0.207
9. GPT-5.6 Sol (August)-0.114-0.049-0.033-0.029-0.017-0.013-0.006-0.001·+0.014+0.023+0.027+0.055+0.060+0.085+0.206
10. GPT OSS 120B-0.128-0.063-0.047-0.043-0.031-0.027-0.020-0.015-0.014·+0.009+0.013+0.041+0.046+0.071+0.192
11. GPT-5.4-0.137-0.072-0.056-0.052-0.040-0.036-0.029-0.024-0.023-0.009·+0.004+0.032+0.037+0.062+0.183
12. GPT-5.6 Luna (August)-0.141-0.076-0.060-0.056-0.044-0.040-0.033-0.028-0.027-0.013-0.004·+0.028+0.033+0.058+0.179
13. GPT-5.3 Chat-0.169-0.104-0.088-0.084-0.072-0.068-0.061-0.056-0.055-0.041-0.032-0.028·+0.005+0.030+0.151
14. GPT-5.1-0.174-0.109-0.093-0.089-0.077-0.073-0.066-0.061-0.060-0.046-0.037-0.033-0.005·+0.025+0.146
15. GPT-5.5 Instant-0.199-0.134-0.118-0.114-0.102-0.098-0.091-0.086-0.085-0.071-0.062-0.058-0.030-0.025·+0.121
16. GPT OSS 20B-0.320-0.255-0.239-0.235-0.223-0.219-0.212-0.207-0.206-0.192-0.183-0.179-0.151-0.146-0.121·

On a 1,000-conversation set, anything within 0.01 sits inside the variation between two measurements of the same model; read those cells as ties. Rows read from different documents carry the extra caveats on the methodology page; the document behind each row is on the sources page.

Pricing and context windows for every model are on the models page.