Methodology
Scores are compiled from published documents: OpenAI system cards and update reports, other vendors' model cards, and named leaderboards. Arcophos does not currently run this benchmark itself. Every row on the board links to the document its number was read from on the sources page, and a row whose document has not been located or verified yet is marked "source pending" wherever it appears. As of September 8, 2026, 16 of 16 rows have a documented source, read from 12 documents. This page covers how a number gets from a document into the table; what the benchmark measures and how the subset was chosen is on the benchmark page.
How numbers are read
| Document | First-party documents come first: the vendor's system card or model card for a vendor-reported number, the benchmark maintainers' board or paper for an official number. A launch post ranks below a system card, a named third-party leaderboard below that, and a page that repeats another document's number is used only as corroboration. |
|---|---|
| Value | The score is copied exactly as the document prints it, with the sentence or table cell it came from kept as a quote and the page, table, or anchor kept as a locator. No rounding, no rescaling, no averaging across documents. |
| Configuration | When a document reports several configurations for one model, the row records which one was taken. A variant note under the model name on the sources page carries that caveat. |
| Disagreement | When two documents give different numbers for the same model, the higher-ranked document supplies the row and the other is listed beside it on the sources page with its own quote. The row is marked disputed until the difference is explained. |
| Confidence | Verified means the number was read from a first-party document and checked against the quote. Partial means the document is located but a detail, such as the exact configuration, is still open. Unverified means no document has been located yet. Disputed means located documents disagree. |
| Pricing and specs | The context and price columns are read from vendor documentation and current price sheets on each update date; the rates shown were checked on September 8, 2026. |
Why documents disagree
Two documents can report different HealthBench Hard numbers for the same model, and both can be correct for what they measured. Four settings account for most of the spread. Length adjustment: some reports penalize long answers and some do not, and a model that writes at length scores differently under the two. Reasoning effort: a model run at low, medium, or high reasoning effort is three different results, and a document does not always say which one it prints. Deployment setting: a number measured through the API with no system prompt differs from one measured inside a consumer product with its own instructions and tools. Grader version: the rubric grader is itself a model (GPT-4.1 in the paper), and a document graded with a later grader is not on the same scale as one graded with the original. The sources page records the quote for each row so a reader can check which of these applies.
Comparability
Two rows read from the same document, an OpenAI system card that reports several of its own models for instance, were graded under one setup and can be read against each other directly. Two rows read from different documents carry the caveats above, and a gap of a few hundredths between them may be a difference in setup rather than in the models. The sources page groups rows by document so the two cases are easy to tell apart. Comparing a row here against a number graded under a configuration no document describes is not meaningful.
Limitations
A compiled board inherits every choice its documents made. Vendors pick which configurations to report, and a system card usually shows a model at its best setting. Independent measurement would close that gap; this site does not currently provide it. The grader is a model too, and grader bias remains an open problem for rubric evaluations. The subset itself was frozen against the May 2025 frontier, so nothing released since had any hand in the selection. This site does not break scores out by theme. Above all, a benchmark number certifies nothing clinically: it records rubric adherence across 1,000 hard conversations and stops there.
Updates
A new system card, model card, or leaderboard entry that reports HealthBench Hard triggers an update, as does a located document for a row that was pending. The updates page dates every change down to price corrections, and the data page carries each release of the results in full, source fields included. As of September 8, 2026 the table covers 16 models.