Updates
A dated record of everything that has changed on the board: model additions, newly located source documents, price corrections. Anchors stay stable, so entries can be cited.
- 2026-09-07
Results moved to the shared Arcophos benchmark results database. Every row now carries the document its score was read from, with quote and locator, listed on the sources page; rows whose document has not been located yet are marked "source pending". The methodology page was rewritten to say plainly that the board is compiled from published sources and that Arcophos does not currently run the benchmark itself. The JSON and CSV downloads gain source fields. Tracing the rows to their documents corrected one score: GPT-5 had been listed at 0.016, which was its hallucination rate, not a HealthBench Hard score; OpenAI's GPT-5.6 system card table gives 0.347 under the same length-adjusted protocol as the 5.6 rows. The board also grew from 9 to 16 rows because system cards report comparison models.
- 2026-08-16
Site launch. 16 models on the 1,000-conversation set: Muse Spark, GPT-6 Astra, GPT-5, GPT-5.2, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.5, GPT-5.6 Sol (August), GPT OSS 120B, GPT-5.4, GPT-5.6 Luna (August), GPT-5.3 Chat, GPT-5.1, GPT-5.5 Instant, GPT OSS 20B. Vendor price sheets checked the same day.
The live table is on the leaderboard, and each release of the results is downloadable from the data page.