← AI Switchboard
AI Switchboardby Waggle
MODELS · September 24, 2026
Sep 22

HLE-Diamond lands, and tools are worth more than twenty points to every model

A 1,000-question closed-book subset of Humanity's Last Exam. GPT-6 Astra leads at 60.6% without tools and 82.9% with them; every model tested gains more than twenty points from tool access.

The Center for AI Safety and Scale AI released HLE-Diamond on Tuesday: 1,000 questions, 500 reasoning and 500 knowledge, designed to be answerable closed-book. The leaderboard without tools reads GPT-6 Astra 60.6%, Claude Opus 5.5 55.0%, Claude Fable 5.1 51.3%, Claude Opus 5 38.6%, Gemini 3.8 Flash 34.3%, GPT-6 Sol 33.8%, GPT-5.6 Sol 31.2% and Muse Spark 1.3 25.4%.

Give the same models web and code tools and the whole board moves up by more than twenty points: Astra 82.9%, Opus 5.5 73.9%, Fable 5.1 72.4%, Opus 5 69.1%, GPT-6 Sol 64.9%, Gemini 3.8 Flash 60.3%, GPT-5.6 Sol 56.5%, Muse Spark 1.3 55.5%. The ordering barely changes; the gap between the best and worst model narrows sharply. Opus 5 gains thirty points. Muse Spark gains thirty.

That compression is the finding worth carrying. On a closed-book test the models are separated by roughly 35 points. Hand them a browser and a code interpreter and the spread collapses to 27, with the weakest tool-using model beating the fourth-best closed-book one. For anyone buying rather than benchmarking, that suggests the harness around the model is doing more work than the model choice — which is the same conclusion McKinsey reached from the invoice side this week.

One awkward pairing: Meta's Muse Spark 1.3 finishes last of the eight, on both measures, in the same week that Muse the product topped the app charts in two countries and moved a basket of financial stocks.

  • Confirmed HLE-Diamond is 1,000 closed-book questions, 500 reasoning and 500 knowledge, from the Center for AI Safety and Scale AI. Humanity's Last Exam
  • Confirmed Without tools: GPT-6 Astra 60.6%, Claude Opus 5.5 55.0%, Claude Fable 5.1 51.3%, down to Muse Spark 1.3 at 25.4%. Humanity's Last Exam
  • Confirmed With web and code tools: Astra 82.9%, Opus 5.5 73.9%, and Muse Spark 1.3 up thirty points to 55.5% — every model gains more than twenty. Humanity's Last Exam

Models & releasesScience & research

This week in the September 24, 2026 edition · front page