HealthBench Hard

Models

9 models evaluated · updated August 16, 2026

The whole board in one table, together with the API facts that determine whether a score matters in practice: context window, list price per token, license. Every model links through to a page with its full result and head-to-head comparisons.

Ranking

#modelscoresizecontextcost in / out per 1Mlicense
1Meta logoMuse Spark Meta0.4281.0M$1.25 / $4.25proprietary
2OpenAI logoGPT-5.6 Sol OpenAI0.3311.1M$5.00 / $30.00proprietary
3OpenAI logoGPT-5.6 Terra OpenAI0.3271.1M$2.00 / $12.00proprietary
4OpenAI logoGPT-5.6 Luna OpenAI0.3201.1M$0.20 / $1.20proprietary
5OpenAI logoGPT OSS 120B OpenAI0.300117B131Kopen
6OpenAI logoGPT-5.3 Chat OpenAI0.259128K$1.75 / $14.00proprietary
7OpenAI logoGPT-5.5 Instant OpenAI0.229400K$5.00 / $30.00proprietary
8OpenAI logoGPT OSS 20B OpenAI0.10821B131Kopen
9OpenAI logoGPT-5 OpenAI0.016400K$1.25 / $10.00proprietary

Model pages