Muse Spark on HealthBench Hard
rank 1 of 16 · updated September 8, 2026via Introducing Muse Spark: Scaling Towards Personal Superintelligence
On the 16-model HealthBench Hard board, Muse Spark holds rank 1 with a score of 0.428. Meta's frontier multimodal reasoning model, sold through the Meta Model API, with weights announced to open later in 2026. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.
Score and API facts
| rank | 1 of 16 |
|---|---|
| score | 0.428 |
| source | Introducing Muse Spark: Scaling Towards Personal Superintelligence (Launch-post benchmark table image, HEALTH section, row HealthBench Hard, column Muse Spark Thinking; identical table in the Eval Methodology PDF p. 5; protocol p. 2: HealthBench Hard: This is a subset of OpenAI's HealthBench benchmark, containing 1000 prompts. We used the same implementation as in the OpenAI’s official simple-evals repo, with GPT-4.1-genai as the LLM-as-judge model.), vendor-reported |
| configuration | raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table. |
| lab | Meta |
| context window | 1.0M tokens |
| API price per 1M tokens | $1.25 in / $4.25 out |
| license | proprietary |
| released | 2026-04-08 |
Where it sits
Nothing ranks above it: Muse Spark holds first place with 0.065 of clear air over GPT-6 Astra. Rows on this board are compiled from published documents, so a gap between two models is exact only when both numbers came from the same document under the same settings; the sources page shows which document each row came from.
What does Muse Spark score on HealthBench Hard?
As of September 8, 2026, Muse Spark scores 0.428 on HealthBench Hard, 1 of 16 models on the board. The number is read from the Introducing Muse Spark: Scaling Towards Personal Superintelligence (Launch-post benchmark table image, HEALTH section, row HealthBench Hard, column Muse Spark Thinking; identical table in the Eval Methodology PDF p. 5; protocol p. 2: HealthBench Hard: This is a subset of OpenAI's HealthBench benchmark, containing 1000 prompts. We used the same implementation as in the OpenAI’s official simple-evals repo, with GPT-4.1-genai as the LLM-as-judge model.), vendor-reported.
What does Muse Spark cost per million tokens?
Meta lists Muse Spark at $1.25 per million input tokens and $4.25 per million output tokens.
Head to head
The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.
- Muse Spark vs GPT-6 Astra0.428 vs 0.363 · Muse Spark by 0.065
- Muse Spark vs GPT-50.428 vs 0.347 · Muse Spark by 0.081
- Muse Spark vs GPT-5.20.428 vs 0.343 · Muse Spark by 0.085
- 0.428 vs 0.331 · Muse Spark by 0.097
- Muse Spark vs GPT-5.6 Terra0.428 vs 0.327 · Muse Spark by 0.101
- Muse Spark vs GPT-5.6 Luna0.428 vs 0.320 · Muse Spark by 0.108
- Muse Spark vs GPT-5.50.428 vs 0.315 · Muse Spark by 0.113
- Muse Spark vs GPT-5.6 Sol (August)0.428 vs 0.314 · Muse Spark by 0.114
- Muse Spark vs GPT OSS 120B0.428 vs 0.300 · Muse Spark by 0.128
- Muse Spark vs GPT-5.40.428 vs 0.291 · Muse Spark by 0.137
- Muse Spark vs GPT-5.6 Luna (August)0.428 vs 0.287 · Muse Spark by 0.141
- Muse Spark vs GPT-5.3 Chat0.428 vs 0.259 · Muse Spark by 0.169
- Muse Spark vs GPT-5.10.428 vs 0.254 · Muse Spark by 0.174
- Muse Spark vs GPT-5.5 Instant0.428 vs 0.229 · Muse Spark by 0.199
- Muse Spark vs GPT OSS 20B0.428 vs 0.108 · Muse Spark by 0.320
How numbers are read from their documents is on the methodology page, the document for this row is on the sources page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.