GPT-6 Astra on HealthBench Hard
rank 2 of 16 · updated September 8, 2026via GPT-6 Astra System Card
On the 16-model HealthBench Hard board, GPT-6 Astra holds rank 2 with a score of 0.363. OpenAI's September 2026 flagship, the first of the GPT-6 line, priced at $10 per million input tokens with a 1.05M context window. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.
Score and API facts
| rank | 2 of 16 |
|---|---|
| score | 0.363 |
| source | GPT-6 Astra System Card (p. 19, sec. 6.1 (Table 6 cell, column 'gpt-6 Astra': 36.3 (37.8, 2192))), vendor-reported |
| configuration | length-adjusted, max reasoning effort (37.8 unadjusted, 2,192 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'. |
| lab | OpenAI |
| context window | 1.05M tokens |
| API price per 1M tokens | $10.00 in / $50.00 out |
| license | proprietary |
| released | 2026-09-03 |
Where it sits
Muse Spark tops the board at 0.428, which puts GPT-6 Astra 0.065 off the lead. One place down is GPT-5 at 0.347. Rows on this board are compiled from published documents, so a gap between two models is exact only when both numbers came from the same document under the same settings; the sources page shows which document each row came from.
What does GPT-6 Astra score on HealthBench Hard?
As of September 8, 2026, GPT-6 Astra scores 0.363 on HealthBench Hard, 2 of 16 models on the board. The number is read from the GPT-6 Astra System Card (p. 19, sec. 6.1 (Table 6 cell, column 'gpt-6 Astra': 36.3 (37.8, 2192))), vendor-reported.
What does GPT-6 Astra cost per million tokens?
OpenAI lists GPT-6 Astra at $10.00 per million input tokens and $50.00 per million output tokens.
Head to head
The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.
- GPT-6 Astra vs Muse Spark0.363 vs 0.428 · Muse Spark by 0.065
- GPT-6 Astra vs GPT-50.363 vs 0.347 · GPT-6 Astra by 0.016
- GPT-6 Astra vs GPT-5.20.363 vs 0.343 · GPT-6 Astra by 0.020
- GPT-6 Astra vs GPT-5.6 Sol0.363 vs 0.331 · GPT-6 Astra by 0.032
- GPT-6 Astra vs GPT-5.6 Terra0.363 vs 0.327 · GPT-6 Astra by 0.036
- GPT-6 Astra vs GPT-5.6 Luna0.363 vs 0.320 · GPT-6 Astra by 0.043
- GPT-6 Astra vs GPT-5.50.363 vs 0.315 · GPT-6 Astra by 0.048
- GPT-6 Astra vs GPT-5.6 Sol (August)0.363 vs 0.314 · GPT-6 Astra by 0.049
- GPT-6 Astra vs GPT OSS 120B0.363 vs 0.300 · GPT-6 Astra by 0.063
- GPT-6 Astra vs GPT-5.40.363 vs 0.291 · GPT-6 Astra by 0.072
- GPT-6 Astra vs GPT-5.6 Luna (August)0.363 vs 0.287 · GPT-6 Astra by 0.076
- GPT-6 Astra vs GPT-5.3 Chat0.363 vs 0.259 · GPT-6 Astra by 0.104
- GPT-6 Astra vs GPT-5.10.363 vs 0.254 · GPT-6 Astra by 0.109
- GPT-6 Astra vs GPT-5.5 Instant0.363 vs 0.229 · GPT-6 Astra by 0.134
- GPT-6 Astra vs GPT OSS 20B0.363 vs 0.108 · GPT-6 Astra by 0.255
How numbers are read from their documents is on the methodology page, the document for this row is on the sources page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.