Sources
16 of 16 rows documented · 12 documents · updated September 8, 2026
This page lists the document behind every score on the leaderboard. Scores are compiled from published documents: OpenAI system cards and update reports, other vendors' model cards, and named leaderboards. Arcophos does not currently run this benchmark itself. Entries are grouped by document. Under each document title sit the rows read from it, each with the score as that document prints it, the sentence or table cell it came from, where in the document it sits, who reported it, and how far the reading has been verified. A row that a second document repeats or disputes appears under that document too, so one model can show up more than once. Rows with no located document are listed at the end. The reading rules are on the methodology page.
GPT-5 System Card
system cardpublished by OpenAIfirst party
- GPT-5OpenAI46.2length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking.
“State of the art on HealthBench Hard improves from 31.6% for OpenAI o3 to 46.2% for gpt-5-thinking.”
p. 18, section 3.10 Health, Figure 6 (HealthBench Hard, raw score %)disagrees with the rowboard shows 0.347 from its primary source
GPT-5.3 Instant System Card
system cardpublished by OpenAIfirst partypublished Mar 2, 2026retrieved Sep 7, 2026
- GPT-5.3 ChatOpenAI25.9%raw score (no length adjustment), GPT-5.3 Instant system card Table 3, column GPT-5.3-INSTANT; OpenAI later cards print 20.2 length-adjusted (17.8 unadjusted) for the re-run model.
“Hard 26.8% 25.9%”
Section 4.1 HealthBench, Table 3: HealthBench, row Hard, column GPT-5.3-INSTANTsource of the rowvendor-reportedas of 2026-03confidence: verified
GPT-5.5 Instant System Card
system cardpublished by OpenAIfirst party
- GPT-5.3 ChatOpenAI20.2raw score (no length adjustment), GPT-5.3 Instant system card Table 3, column GPT-5.3-INSTANT; OpenAI later cards print 20.2 length-adjusted (17.8 unadjusted) for the re-run model.
“HealthBench Hard 21.6 (23.0, 2,181) 23.3 (23.5, 2,022) 20.2 (17.8, 1,693) 22.9 (21.3, 1,794)”
Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.3 INSTANTdisagrees with the rowboard shows 0.259 from its primary source - GPT-5.5 InstantOpenAI22.9length-adjusted (21.3 unadjusted, 1,794 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
“HealthBench Hard 21.6 (23.0, 2,181) 23.3 (23.5, 2,022) 20.2 (17.8, 1,693) 22.9 (21.3, 1,794)”
Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANTsource of the rowvendor-reportedas of 2026-05confidence: verified
GPT-5.5 System Card
system cardpublished by OpenAIfirst party
- GPT-5OpenAI34.7length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking.
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289)”
Section 5 Health, Table 7, column GPT-5corroborates the rowboard shows 0.347 from its primary source
GPT-5.6 - August Updates (system card addendum)
system cardpublished by OpenAIfirst partypublished Aug 6, 2026retrieved Sep 7, 2026
- GPT-5.6 Sol (August)OpenAI31.4ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (27.1 unadjusted, 1,450 chars)
“HealthBench Hard 20.2 (17.8, 1,693) 22.9 (21.3, 1,794) 31.4 (27.1, 1,450) 28.7 (24.9, 1,523)”
p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)source of the rowvendor-reportedas of 2026-08confidence: verified - GPT-5.6 Luna (August)OpenAI28.7ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars)
“HealthBench Hard 20.2 (17.8, 1,693) 22.9 (21.3, 1,794) 31.4 (27.1, 1,450) 28.7 (24.9, 1,523)”
p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)source of the rowvendor-reportedas of 2026-08confidence: verified - GPT-5.3 ChatOpenAI20.2raw score (no length adjustment), GPT-5.3 Instant system card Table 3, column GPT-5.3-INSTANT; OpenAI later cards print 20.2 length-adjusted (17.8 unadjusted) for the re-run model.
“HealthBench Hard 20.2 (17.8, 1,693) 22.9 (21.3, 1,794) 31.4 (27.1, 1,450) 28.7 (24.9, 1,523)”
p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.3 Instantdisagrees with the rowboard shows 0.259 from its primary source - GPT-5.5 InstantOpenAI22.9length-adjusted (21.3 unadjusted, 1,794 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
“HealthBench Hard 20.2 (17.8, 1,693) 22.9 (21.3, 1,794) 31.4 (27.1, 1,450) 28.7 (24.9, 1,523)”
p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instantcorroborates the rowboard shows 0.229 from its primary source
GPT-5.6 Preview System Card
system cardpublished by OpenAIfirst party
- GPT-5OpenAI34.7length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking.
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5corroborates the rowboard shows 0.347 from its primary source - GPT-5.2OpenAI34.3length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (38.9 unadjusted, 2585 chars)
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2corroborates the rowboard shows 0.343 from its primary source - GPT-5.6 SolOpenAI33.1length-adjusted, max reasoning effort (31.1 unadjusted, 1,751 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 490.
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOLcorroborates the rowboard shows 0.331 from its primary source - GPT-5.6 TerraOpenAI32.7length-adjusted, max reasoning effort (34.3 unadjusted, 2,199 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRAcorroborates the rowboard shows 0.327 from its primary source - GPT-5.6 LunaOpenAI32.0length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491.
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNAcorroborates the rowboard shows 0.320 from its primary source - GPT-5.5OpenAI31.5length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (33.8 unadjusted, 2289 chars)
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5corroborates the rowboard shows 0.315 from its primary source - GPT-5.4OpenAI29.1length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars)
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4corroborates the rowboard shows 0.291 from its primary source - GPT-5.1OpenAI25.4length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (41.4 unadjusted, 4049 chars)
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1corroborates the rowboard shows 0.254 from its primary source
GPT-5.6 System Card
system cardpublished by OpenAIfirst partypublished Jul 9, 2026retrieved Sep 7, 2026
- GPT-5OpenAI34.7length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking.
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5source of the rowvendor-reportedas of 2026-06confidence: verified - GPT-5.2OpenAI34.3length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (38.9 unadjusted, 2585 chars)
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2source of the rowvendor-reportedas of 2026-06confidence: verified - GPT-5.6 SolOpenAI33.1length-adjusted, max reasoning effort (31.1 unadjusted, 1,751 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 490.
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOLsource of the rowvendor-reportedas of 2026-06confidence: verified - GPT-5.6 TerraOpenAI32.7length-adjusted, max reasoning effort (34.3 unadjusted, 2,199 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRAsource of the rowvendor-reportedas of 2026-06confidence: verified - GPT-5.6 LunaOpenAI32.0length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491.
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNAsource of the rowvendor-reportedas of 2026-06confidence: verified - GPT-5.5OpenAI31.5length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (33.8 unadjusted, 2289 chars)
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5source of the rowvendor-reportedas of 2026-06confidence: verified - GPT-5.4OpenAI29.1length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars)
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4source of the rowvendor-reportedas of 2026-06confidence: verified - GPT-5.1OpenAI25.4length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (41.4 unadjusted, 4049 chars)
“HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)”
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1source of the rowvendor-reportedas of 2026-06confidence: verified
GPT-6 Astra System Card
system cardpublished by OpenAIfirst partypublished Sep 3, 2026retrieved Sep 8, 2026
- GPT-6 AstraOpenAI36.3length-adjusted, max reasoning effort (37.8 unadjusted, 2,192 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.
“Astra has a length-adjusted HealthBench Professional score of 63.4 (+2.9 relative to GPT-5.6 Sol), HealthBench score of 58.1 (+1.1), HealthBench Hard score of 36.3 (+3.2), and HealthBench Consensus score of 95.8 (+0.3).”
p. 19, sec. 6.1 (Table 6 cell, column 'gpt-6 Astra': 36.3 (37.8, 2192))source of the rowvendor-reportedas of 2026-09confidence: verified
GPT-6 Astra System Card - HealthBench (Deployment Safety Hub)
system cardpublished by OpenAIfirst party
- GPT-6 AstraOpenAI36.3length-adjusted, max reasoning effort (37.8 unadjusted, 2,192 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.
“HealthBench Hard score of 36.3 (+3.2)”
section 6.1, HTML rendering of the same cardrepeats the rowboard shows 0.363 from its primary source
gpt-oss-120b & gpt-oss-20b Model Card
model cardpublished by OpenAIfirst partypublished Aug 5, 2025retrieved Sep 7, 2026
- GPT OSS 120BOpenAI30.0raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
“HealthBench Hard 22.8 26.9 30.0 9.0 12.9 10.8”
Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-120b highsource of the rowvendor-reportedas of 2025-08confidence: verified - GPT OSS 20BOpenAI10.8raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
“HealthBench Hard 22.8 26.9 30.0 9.0 12.9 10.8”
Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-20b highsource of the rowvendor-reportedas of 2025-08confidence: verified
Introducing Muse Spark: Scaling Towards Personal Superintelligence
launch postpublished by Metafirst partypublished Apr 8, 2026retrieved Sep 7, 2026
- Muse SparkMeta42.8raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table.
“HealthBench Hard | Open-Ended Health Queries | 42.8 | 14.8 | 20.6 | 40.1 | 20.3 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT-5.4 Xhigh, Grok 4.2 Reasoning)”
Launch-post benchmark table image, HEALTH section, row HealthBench Hard, column Muse Spark Thinking; identical table in the Eval Methodology PDF p. 5; protocol p. 2: HealthBench Hard: This is a subset of OpenAI's HealthBench benchmark, containing 1000 prompts. We used the same implementation as in the OpenAI’s official simple-evals repo, with GPT-4.1-genai as the LLM-as-judge model.source of the rowvendor-reportedas of 2026-04confidence: verified
Muse Spark Eval Methodology
model cardpublished by Metafirst party
- Muse SparkMeta42.8raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table.
“HealthBench Hard | Open-Ended Health Queries | 42.8 | 14.8 | 20.6 | 40.1 | 20.3 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT-5.4 Xhigh, Grok 4.2 Reasoning)”
p. 5 results table image, row HealthBench Hard; protocol p. 2corroborates the rowboard shows 0.428 from its primary source
Rows without a documented source
Every row on the board currently has a located document.
The same fields ship in the JSON download: a source object on every model plus the sources list this page is rendered from. Corrections go to Arcophos.