HealthBench Hard

OpenAI logoGPT-5.3 Chat on HealthBench Hard

rank 6 of 9 · updated August 16, 2026

On the 9-model HealthBench Hard board, GPT-5.3 Chat holds rank 6 with a score of 0.259. The API alias of the non-reasoning ChatGPT default from spring 2026, now deprecated in favor of GPT-5.6. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.

Score and API facts

rank6 of 9
score0.259
labOpenAI
context window128K tokens
API price per 1M tokens$1.75 in / $14.00 out
licenseproprietary
released2026-03-05

Where it sits

Muse Spark tops the board at 0.428, which puts GPT-5.3 Chat 0.169 off the lead. One place up is GPT OSS 120B at 0.300. One place down is GPT-5.5 Instant at 0.229. Every row on the board comes out of one run, so a gap between two models is measured on the same conversations under the same grader.

What does GPT-5.3 Chat score on HealthBench Hard?

As of August 16, 2026, GPT-5.3 Chat scores 0.259 on HealthBench Hard, 6 of 9 models on the board.

What does GPT-5.3 Chat cost per million tokens?

OpenAI lists GPT-5.3 Chat at $1.75 per million input tokens and $14.00 per million output tokens.

Head to head

The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.

How the conversations are graded is on the methodology page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.