Humanity's Last Exam
Humanity's Last Exam is A refined subset of Humanity's Last Exam, following a year-long process of cleaning and refinement with input from research communities.. It is ranked #191 on the AI Radar, in Frontier models, first seen 2 h ago and shared in 1 post (14.2k views).
Introducing HLE-Diamond | Humanity's Last Exam Humanity's Last Exam Introducing HLE-Diamond & Dataset load_dataset(" cais/hle-diamond ") Center for AI Safety and Scale AI September 22, 2026 We are releasing HLE-Diamond , a refined subset of Humanity’s Last Exam (HLE) question collection, following a year-long process of cleaning and refinement with input from research communities. HLE-Diamond consists of 1,000 questions . Main results. We compare the performance of current models on HLE-Diamond without tools. HLE-Diamond GPT-6 Astra 60.6 % Claude Opus 5.5 55.0 % Claude Fable 5.1 51.3 % Claude Opus 5 38.6 % Gemini 3.8 Flash 34.3 % GPT-6 Sol…
What people said about Humanity's Last Exam on X
We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE), following a year-long process of cleaning and refinement with input from various research communities. https://lastexam.ai/blog/hle-diamond w/ @ScaleAILabs
— @CAIS, 2 h ago · 310 likes · see the post
Alternatives to Humanity's Last Exam
- Intelligence Benchmarking — Detailed intelligence benchmarking methodology for LLM quality evaluations.
- Qwen Studio — Qwen Studio offers comprehensive functionality spanning chatbot, image and video understanding, image generation,…
- Claude — Claude is Anthropic's AI, built for problem solvers. Tackle complex challenges, analyze data, write code, and think…
- 阶跃星辰 — Model page
- www-cdn.anthropic.com —
- Qwen3.8-Omni-Flash — Qwen’s next-generation native omni-modal model supports context lengths of up to 1M tokens and natively accepts text,…
Humanity's Last Exam in numbers
- Rank on the AI Radar: #191 of 1868
- Shared in 1 post by 1 account: @CAIS
- 14.2k views on those posts
- First seen 2 h ago, last shared 2 h ago
- Pricing seen by Jev: not stated
- Market: Frontier models
FAQ
What is Humanity's Last Exam?
A refined subset of Humanity's Last Exam, following a year-long process of cleaning and refinement with input from research communities. It was first shared on X 2 h ago and is ranked #191 on the AI Radar.
Is Humanity's Last Exam free?
Pricing is not stated on the page we read.
Who shared Humanity's Last Exam?
1 account on X, including @CAIS, in 1 post totalling 14.2k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 44.7k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 17:53 UTC. Full method.