# HealthBench Hard Leaderboard, 2026-08-16 Muse Spark leads the HealthBench Hard leaderboard at 0.428, ahead of GPT-5.6 Sol (0.331) and GPT-5.6 Terra (0.327). 9 models evaluated. Source: https://healthbenchhard.ai ## Ranking | rank | model | lab | score | context | $ in / 1M | $ out / 1M | license | |---|---|---|---|---|---|---|---| | 1 | Muse Spark | Meta | 0.428 | 1.0M | $1.25 | $4.25 | proprietary | | 2 | GPT-5.6 Sol | OpenAI | 0.331 | 1.1M | $5.00 | $30.00 | proprietary | | 3 | GPT-5.6 Terra | OpenAI | 0.327 | 1.1M | $2.00 | $12.00 | proprietary | | 4 | GPT-5.6 Luna | OpenAI | 0.320 | 1.1M | $0.20 | $1.20 | proprietary | | 5 | GPT OSS 120B | OpenAI | 0.300 | 131K | - | - | open | | 6 | GPT-5.3 Chat | OpenAI | 0.259 | 128K | $1.75 | $14.00 | proprietary | | 7 | GPT-5.5 Instant | OpenAI | 0.229 | 400K | $5.00 | $30.00 | proprietary | | 8 | GPT OSS 20B | OpenAI | 0.108 | 131K | - | - | open | | 9 | GPT-5 | OpenAI | 0.016 | 400K | $1.25 | $10.00 | proprietary | Prices are vendor list rates per 1M tokens as of 2026-08-16. GPT-5.6 models charge higher rates above 272K input tokens. GPT OSS models are open weights without vendor list pricing. ## About the benchmark HealthBench Hard is the hardest slice of HealthBench, the health benchmark OpenAI released in May 2025 (arXiv 2505.08775). From 5,000 realistic health conversations, it keeps the 1,000 that frontier models handled worst: five frontier models scored every conversation and the 1,000 with the lowest average score became the subset. Rubrics were written by 262 physicians across 60 countries, 48,562 criteria in the parent set. A model grader (GPT-4.1) judges each physician-written rubric criterion as met or not met. Criteria carry weights from -10 to +10, the score is earned points over maximum positive points, averaged and clipped to 0 to 1. The best score at the benchmark's release was o3's 0.320. ## How these numbers were produced Each model was run under default sampling, one response per conversation, no tools and no scaffolding, graded with the benchmark's default grader configuration. API models run through their vendor's public endpoint; open-weights models run from the released weights. Scores are comparable within this table; they are not comparable against numbers graded under other configurations, including OpenAI's in-product results. ## Models - Muse Spark (Meta): 0.428, rank 1 of 9. Meta's frontier multimodal reasoning model, sold through the Meta Model API, with weights announced to open later in 2026. Released 2026-08-05. https://healthbenchhard.ai/models/muse-spark - GPT-5.6 Sol (OpenAI): 0.331, rank 2 of 9. OpenAI's flagship 5.6-series model, built for its heaviest reasoning and agentic work. Released 2026-07-09. https://healthbenchhard.ai/models/gpt-5.6-sol - GPT-5.6 Terra (OpenAI): 0.327, rank 3 of 9. The middle of OpenAI's 5.6 lineup, sitting between Sol and Luna on price and capability. Released 2026-07-09. https://healthbenchhard.ai/models/gpt-5.6-terra - GPT-5.6 Luna (OpenAI): 0.320, rank 4 of 9. OpenAI's high-volume 5.6 tier, cut to $0.20 per million input tokens in the July 30, 2026 repricing. Released 2026-07-09. https://healthbenchhard.ai/models/gpt-5.6-luna - GPT OSS 120B (OpenAI): 0.300, rank 5 of 9. OpenAI's most capable open-weights model, an Apache 2.0 mixture-of-experts release that fits on a single 80GB GPU. Released 2025-08-05. https://healthbenchhard.ai/models/gpt-oss-120b - GPT-5.3 Chat (OpenAI): 0.259, rank 6 of 9. The API alias of the non-reasoning ChatGPT default from spring 2026, now deprecated in favor of GPT-5.6. Released 2026-03-05. https://healthbenchhard.ai/models/gpt-5.3-chat - GPT-5.5 Instant (OpenAI): 0.229, rank 7 of 9. The speed-focused variant that took over as ChatGPT's default in May 2026. Released 2026-05-05. https://healthbenchhard.ai/models/gpt-5.5-instant - GPT OSS 20B (OpenAI): 0.108, rank 8 of 9. The smaller Apache 2.0 open-weights release, sized to run locally in 16GB of memory. Released 2025-08-05. https://healthbenchhard.ai/models/gpt-oss-20b - GPT-5 (OpenAI): 0.016, rank 9 of 9. OpenAI's August 2025 flagship, still served at unchanged prices but superseded by the GPT-5.6 series. Released 2025-08-07. https://healthbenchhard.ai/models/gpt-5 ## Reuse Data: CC BY 4.0. Cite as: Arcophos. HealthBench Hard Leaderboard. 2026-08-16 snapshot. https://healthbenchhard.ai. Machine-readable: https://healthbenchhard.ai/data/leaderboard.json