GPT OSS 120B vs GPT OSS 20B on HealthBench Hard
updated August 16, 2026
A gap of 0.192 separates these two on the 1,000 hardest HealthBench conversations: GPT OSS 120B at 0.300, GPT OSS 20B at 0.108. The table adds what each one costs at list rates.
Side by side
| score | 0.300 | 0.108 |
|---|---|---|
| rank | 5 of 9 | 8 of 9 |
| context window | 131K | 131K |
| price per 1M tokens, in / out | — | — |
| 1,000-exchange workload | — | — |
| released | 2025-08-05 | 2025-08-05 |
| license | open | open |
The workload row prices 1,000 exchanges of 2,000 input and 700 output tokens each, at the list rates current on August 16, 2026. GPT-5.6 models charge higher rates above 272K input tokens. GPT OSS models are open weights without vendor list pricing.
Reading the matchup
The two open-weights releases, 0.192 apart, the price of dropping from 117B to 21B parameters. The 120B model needs a single 80GB GPU; the 20B model runs in 16GB of memory. For hard health conversations the small model keeps about a third of its sibling's score, so it is a fit for constrained hardware, not a substitute.
Which scores higher on HealthBench Hard, GPT OSS 120B or GPT OSS 20B?
GPT OSS 120B. On the 1,000-conversation set it scores 0.300 to GPT OSS 20B's 0.108, a margin of 0.192 as of August 16, 2026.
Which is cheaper to run, GPT OSS 120B or GPT OSS 20B?
GPT OSS 120B carries no vendor list price, which rules out a like-for-like cost comparison.
Related comparisons
- 0.320 vs 0.300
Each model's full page: GPT OSS 120B and GPT OSS 20B. Every other pairing lives in the score-difference matrix.