HealthBench Hard

OpenAI logoGPT OSS 20B on HealthBench Hard

rank 16 of 16 · updated September 8, 2026via gpt-oss-120b & gpt-oss-20b Model Card

On the 16-model HealthBench Hard board, GPT OSS 20B holds rank 16 with a score of 0.108. The smaller Apache 2.0 open-weights release, sized to run locally in 16GB of memory. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.

Score and API facts

rank16 of 16
score0.108
sourcegpt-oss-120b & gpt-oss-20b Model Card (Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-20b high), vendor-reported
configurationraw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
labOpenAI
context window131K tokens
API price per 1M tokensno published list pricing
licenseopen
parameters21B
released2025-08-05

Where it sits

Muse Spark tops the board at 0.428, which puts GPT OSS 20B 0.320 off the lead. One place up is GPT-5.5 Instant at 0.229. Rows on this board are compiled from published documents, so a gap between two models is exact only when both numbers came from the same document under the same settings; the sources page shows which document each row came from.

What does GPT OSS 20B score on HealthBench Hard?

As of September 8, 2026, GPT OSS 20B scores 0.108 on HealthBench Hard, 16 of 16 models on the board. The number is read from the gpt-oss-120b & gpt-oss-20b Model Card (Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-20b high), vendor-reported.

What does GPT OSS 20B cost per million tokens?

OpenAI publishes no list pricing for GPT OSS 20B.

Head to head

The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.

How numbers are read from their documents is on the methodology page, the document for this row is on the sources page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.