# HealthBench Hard Leaderboard > Frontier language models ranked on HealthBench Hard, the 1,000 hardest conversations in OpenAI's HealthBench benchmark. Muse Spark leads 9 models at 0.428. Updated 2026-08-16. ## Leaderboard - [Leaderboard](https://healthbenchhard.ai/): the live ranking, 9 models with scores, context windows, and list prices - [Full results, JSON](https://healthbenchhard.ai/data/leaderboard.json): every result in machine-readable form, CC BY 4.0 - [Full results, CSV](https://healthbenchhard.ai/data/leaderboard.csv): spreadsheet-ready, one model per row - [Plain-text leaderboard](https://healthbenchhard.ai/llms-full.txt): all of the site's numbers in a single document ## Background - [What is HealthBench Hard](https://healthbenchhard.ai/benchmark): how the subset was selected and how it differs from HealthBench, Consensus, and Professional - [Methodology](https://healthbenchhard.ai/methodology): how the runs work, the grader configuration, and what the scores cannot say - [FAQ](https://healthbenchhard.ai/faq): direct answers on scores, grading, and the best open-weights model ## Models - [Muse Spark](https://healthbenchhard.ai/models/muse-spark): 0.428, rank 1 of 9 - [GPT-5.6 Sol](https://healthbenchhard.ai/models/gpt-5.6-sol): 0.331, rank 2 of 9 - [GPT-5.6 Terra](https://healthbenchhard.ai/models/gpt-5.6-terra): 0.327, rank 3 of 9 - [GPT-5.6 Luna](https://healthbenchhard.ai/models/gpt-5.6-luna): 0.320, rank 4 of 9 - [GPT OSS 120B](https://healthbenchhard.ai/models/gpt-oss-120b): 0.300, rank 5 of 9 - [GPT-5.3 Chat](https://healthbenchhard.ai/models/gpt-5.3-chat): 0.259, rank 6 of 9 - [GPT-5.5 Instant](https://healthbenchhard.ai/models/gpt-5.5-instant): 0.229, rank 7 of 9 - [GPT OSS 20B](https://healthbenchhard.ai/models/gpt-oss-20b): 0.108, rank 8 of 9 - [GPT-5](https://healthbenchhard.ai/models/gpt-5): 0.016, rank 9 of 9 ## Optional - [Compare models](https://healthbenchhard.ai/compare): the score-difference matrix plus head-to-head pages for key matchups - [Best AI for healthcare](https://healthbenchhard.ai/best-ai-for-healthcare): a pick for each constraint, including open weights - [Updates](https://healthbenchhard.ai/updates): every change to the board, dated