HealthBench Hard

OpenAI logoGPT-5.5 Instant on HealthBench Hard

rank 15 of 16 · updated September 8, 2026via GPT-5.5 Instant System Card

On the 16-model HealthBench Hard board, GPT-5.5 Instant holds rank 15 with a score of 0.229. The speed-focused variant that took over as ChatGPT's default in May 2026. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.

Score and API facts

rank15 of 16
score0.229
sourceGPT-5.5 Instant System Card (Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT), vendor-reported
configurationlength-adjusted (21.3 unadjusted, 1,794 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
labOpenAI
context window400K tokens
API price per 1M tokens$5.00 in / $30.00 out
licenseproprietary
released2026-05-05

Where it sits

Muse Spark tops the board at 0.428, which puts GPT-5.5 Instant 0.199 off the lead. One place up is GPT-5.1 at 0.254. One place down is GPT OSS 20B at 0.108. Rows on this board are compiled from published documents, so a gap between two models is exact only when both numbers came from the same document under the same settings; the sources page shows which document each row came from.

What does GPT-5.5 Instant score on HealthBench Hard?

As of September 8, 2026, GPT-5.5 Instant scores 0.229 on HealthBench Hard, 15 of 16 models on the board. The number is read from the GPT-5.5 Instant System Card (Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT), vendor-reported.

What does GPT-5.5 Instant cost per million tokens?

OpenAI lists GPT-5.5 Instant at $5.00 per million input tokens and $30.00 per million output tokens.

Head to head

The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.

How numbers are read from their documents is on the methodology page, the document for this row is on the sources page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.