GPT OSS 120B on HealthBench Hard
rank 10 of 16 · updated September 8, 2026via gpt-oss-120b & gpt-oss-20b Model Card
On the 16-model HealthBench Hard board, GPT OSS 120B holds rank 10 with a score of 0.300. OpenAI's most capable open-weights model, an Apache 2.0 mixture-of-experts release that fits on a single 80GB GPU. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.
Score and API facts
| rank | 10 of 16 |
|---|---|
| score | 0.300 |
| source | gpt-oss-120b & gpt-oss-20b Model Card (Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-120b high), vendor-reported |
| configuration | raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board. |
| lab | OpenAI |
| context window | 131K tokens |
| API price per 1M tokens | no published list pricing |
| license | open |
| parameters | 117B |
| released | 2025-08-05 |
Where it sits
Muse Spark tops the board at 0.428, which puts GPT OSS 120B 0.128 off the lead. One place up is GPT-5.6 Sol (August) at 0.314. One place down is GPT-5.4 at 0.291. Rows on this board are compiled from published documents, so a gap between two models is exact only when both numbers came from the same document under the same settings; the sources page shows which document each row came from.
What does GPT OSS 120B score on HealthBench Hard?
As of September 8, 2026, GPT OSS 120B scores 0.300 on HealthBench Hard, 10 of 16 models on the board. The number is read from the gpt-oss-120b & gpt-oss-20b Model Card (Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-120b high), vendor-reported.
What does GPT OSS 120B cost per million tokens?
OpenAI publishes no list pricing for GPT OSS 120B.
Head to head
The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.
- GPT OSS 120B vs Muse Spark0.300 vs 0.428 · Muse Spark by 0.128
- GPT OSS 120B vs GPT-6 Astra0.300 vs 0.363 · GPT-6 Astra by 0.063
- GPT OSS 120B vs GPT-50.300 vs 0.347 · GPT-5 by 0.047
- GPT OSS 120B vs GPT-5.20.300 vs 0.343 · GPT-5.2 by 0.043
- GPT OSS 120B vs GPT-5.6 Sol0.300 vs 0.331 · GPT-5.6 Sol by 0.031
- GPT OSS 120B vs GPT-5.6 Terra0.300 vs 0.327 · GPT-5.6 Terra by 0.027
- 0.300 vs 0.320 · GPT-5.6 Luna by 0.020
- GPT OSS 120B vs GPT-5.50.300 vs 0.315 · GPT-5.5 by 0.015
- GPT OSS 120B vs GPT-5.6 Sol (August)0.300 vs 0.314 · GPT-5.6 Sol (August) by 0.014
- GPT OSS 120B vs GPT-5.40.300 vs 0.291 · GPT OSS 120B by 0.009
- GPT OSS 120B vs GPT-5.6 Luna (August)0.300 vs 0.287 · GPT OSS 120B by 0.013
- GPT OSS 120B vs GPT-5.3 Chat0.300 vs 0.259 · GPT OSS 120B by 0.041
- GPT OSS 120B vs GPT-5.10.300 vs 0.254 · GPT OSS 120B by 0.046
- GPT OSS 120B vs GPT-5.5 Instant0.300 vs 0.229 · GPT OSS 120B by 0.071
- 0.300 vs 0.108 · GPT OSS 120B by 0.192
How numbers are read from their documents is on the methodology page, the document for this row is on the sources page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.