Humanity's Last Exam
Humanity's Last Exam (HLE) is a language model benchmark consisting of 2,500 questions across a broad range of subjects. It was created jointly by the Center for AI Safety and Scale AI.
Creation
[edit]Stanford HAI's AI Index 2025 Annual Report cites Humanity's Last Exam as one of the "more challenging benchmarks" developed in response to the popular AI benchmarks having reached "saturation".[1] The test has been described as the brainchild of Dan Hendrycks, a machine learning researcher and the director of the Center for AI Safety, who stated that he was inspired to create the test after a conversation with Elon Musk, who thought the existing language model benchmarks, such as the MMLU, were too easy. Hendrycks worked with Scale AI to compile the questions.[2] The questions were crowdsourced from subject matter experts from various institutions across the world.[3][4] The questions were first filtered by the leading AI models; if the models failed to answer the question or did worse than random guessing on the multiple-choice questions, they were reviewed by human experts in two rounds and approved for inclusion in the dataset. The submitters of the top-rated questions were given prize money from a pool of 500,000 U.S. dollars—$5,000 for each of the top 50 questions and $500 for the next 500. After the initial release, a "community feedback bug bounty program" was opened to "identify and remove major errors in the dataset".[4]
Composition
[edit]The benchmark consists of 2,500 questions in the publicly released set. The questions "typically require graduate-level expertise or test knowledge of highly specific topics." The paper classifies the questions into the following broad subjects: mathematics (41%), physics (9%), biology/medicine (11%), humanities/social science (9%), computer science/artificial intelligence (10%), engineering (4%), chemistry (7%), and other (9%). Around 14% of the questions require the ability to understand both text and images, i.e., multi-modality. 24% of the questions are multiple-choice; the rest are short-answer, exact-match questions. A private set is also maintained to test for benchmark overfitting.[4]
An example question:[2]
Hummingbirds within Apodiformes uniquely have a bilaterally paired oval bone, a sesamoid embedded in the caudolateral portion of the expanded, cruciate aponeurosis of insertion of m. depressor caudae. How many paired tendons are supported by this sesamoid bone? Answer with a number.
An independent investigation by FutureHouse, published in July 2025, suggested that around 30% of the HLE answers for text-only chemistry and biology questions could be incorrect; the benchmark's team partially replicated the findings, and said they hope to institute a continuous revision process.[5] The team subsequently launched HLE-Rolling, a version of the benchmark that is "regularly updated to address community feedback and integrate new questions".[4]
Results
[edit]| Organization | Model | Accuracy (%) |
|---|---|---|
| Anthropic | Claude Fable 5 | 53.3 |
| OpenAI | GPT-5.6 Sol | 47.2 |
| Meta Superintelligence Labs | Muse Spark 1.1 | 45.1 |
| Google DeepMind | Gemini 3.1 Pro Preview | 44.7 |
| Moonshot AI | Kimi K3 | 44.3 |
| xAI | Grok 4.5 | 40.3 |
| Z.ai | GLM-5.2 | 40.1 |
| Alibaba Cloud | Qwen 3.7 Max | 38.1 |
| MiniMax | MiniMax-M3 | 37.1 |
| DeepSeek | DeepSeek-V4-Pro | 35.9 |
| Xiaomi | MiMo-V2.5-Pro | 33.8 |
| Thinking Machines Lab | Inkling | 29.7 |
| Nvidia | Nemotron 3 Ultra | 26.6 |
References
[edit]- ↑ Maslej, Nestor; et al. (April 2025). The AI Index 2025 Annual Report (PDF) (Report). Institute for Human-Centered AI. pp. 141–142.
- 1 2 Roose, Kevin (23 January 2025). "When A.I. Passes This Test, Look Out". New York Times. Retrieved 24 January 2025.
{{cite web}}: CS1 maint: deprecated archival service (link) - ↑ Dastin, Jeffrey; Paul, Katie (16 September 2024). "AI experts ready 'Humanity's Last Exam' to stump powerful tech". Reuters. Retrieved 24 January 2025.
{{cite web}}: CS1 maint: deprecated archival service (link) - 1 2 3 4 Center for AI Safety; Scale AI; HLE Contributers Consortium (2026). "A benchmark of expert-level academic questions to assess AI capabilities". Nature. 649: 1139–1146. doi:10.1038/s41586-025-09962-4. PMC 12851929.
- ↑ Skarlinski, Michael; Laurent, Jon; Bou, Albert; White, Andrew (16 September 2025). "About 30% of Humanity's Last Exam chemistry/biology answers are likely wrong". FutureHouse. Retrieved 15 October 2025.