NeurIPS 2026 Competition

AIMO Interpretability Challenge

Can your method tell which AI truly understands the problem
and which one is faking it?
Distinguish robust from spurious reasoning in frontier AI.

Starting end of July 2026 -- register for updates!
💰 $10,000 in prizes 📐 Unseen olympiad problems 🧠 Frontier reasoning models from AIMO 3
▷ Join the Challenge 📄 NeurIPS 2026 Proposal How to participate

Why Interpretability for AIMO?

AIMO (AI Mathematical Olympiad) is an annual competition of AI systems in solving unseen frontier mathematical problems. Thanks to its importance and large prizes ($2.2M USD this year), AIMO attracts huge community attention — last year, the competition was covered by global media The Wall Street Journal and Bloomberg, and participated in by over 4,000+ teams.

The ambition that we have in the Fields Model Initiative is to turn this tremendous engineering effort into scientific knowledge — laying robust stepping stones that will accelerate progress in future AIMO years and in AI as a whole.

We believe that interpretability research can play a key role in helping us achieve this goal — allowing us to understand the mechanisms underlying SOTA reasoning in AI systems and their practical robustness. More broadly, AIMO Interpretability Challenge aims to steer the focus towards robustness and actionability — that interpretability needs to make a real-world impact!

Robust or Spurious?

The AIMO Interpretability Challenge will require participants to submit a system that decides which LLM provides an answer to a given AIMO problem robustly. The full problem formulation, data collection, and evaluation protocol are described in our competition proposal paper.

For each problem, we provide one of two LLMs:

🛡️

Robust Model

The highest-ranked LLM submission to AIMO that passes all our robustness checks.

⚠️

Spurious Model

The highest-ranked LLM for which we verify its reliance on at least one spurious pattern by counterfactual evaluation.

Submitted methods will be evaluated for their accuracy in classifying if a given model responds to a given problem robustly.

📋

Competition submissions will be submitted as Codabench submission bundles that comply with the unified interface defined in the baselines repository. Submissions will be evaluated on our servers without internet access. The provided validation set will cover all types of models contained in the test set.

Submission Types

🏆

Main Track

Includes the full scale of top-performing models from AIMO 3 — no restrictions on model size.

🔬

Small Models Track

Subsets the evaluation to the best-performing models below the 10-billion-parameter scale, providing a comparable setup for compute-heavy methods such as Sparse Autoencoders or Transcoders.

How to Participate

1️⃣

Register for updates

Fill in the registration form for email updates, or join our Discord channel.

2️⃣

Make your first submission

Start from one of the reference baselines and submit it as-is to see your submission in the Codabench leaderboard.

3️⃣

Implement and package your new shiny method

Wrap your method in the Codabench interface specified in the baselines repository. No internet access is available at evaluation time.

4️⃣

Compete for the prizes

Iterate on the leaderboard until the final deadline on October 25, 2026 — with $10,000 in prizes across the Main and Small Models tracks.

Need compute support? Submit a brief proposal in the Fields Model Initiative to request access to compute for your participation.

Key Dates

mid-July: competition start & warmup phase
Competition start: release of train+validation data and baseline submission bundles, two-week warm-up phase
July 31 – October 25, 2026
Main competition phase
October 25, 2026
Final submission deadline — closing of the submission interface
October 25 – October 30, 2026
Participant technical report submission window
November 1 – November 15, 2026
Validation of results, technical reports review and rules compliance check, result analysis
December 11, 2026
Competition workshop at NeurIPS 2026

Frequently Asked Questions

1. Who can participate?

The challenge is open to everyone — academic researchers, independent researchers, and industry practitioners alike. There are no restrictions on team size or affiliation.

2. Do I need to be a participant of AIMO to join?

No. The AIMO Interpretability Challenge is a separate competition. You will be given access to AIMO model submissions as part of the provided environment — you do not need to submit to AIMO yourself.

3. How do I get access to the models and compute?

Apply through the Fields Model Initiative by submitting a one-page research proposal. Accepted participants will receive access to H200 GPU clusters.

4. What is the difference between the Main Track and the Small Models Track?

The Main Track covers the full scale of top-performing models from AIMO 3 with no restriction on model size. The Small Models Track subsets the evaluation to the best-performing models below the 10-billion-parameter scale, offering a comparable setup for methods that are more compute-heavy to train — such as Sparse Autoencoders or Transcoders — where analysing very large models may be infeasible.

5. What runtime environment does my submission need?

Submissions must be packaged as Codabench submission bundle and must conform to the unified interface defined in the baselines repository. Containers are evaluated on our servers without internet access. Detailed technical documentation and a starter kit will be released mid-July: competition start & warmup phase.

6. Will there be tutorial material and a starter kit?

Yes. A starter kit with baseline implementations, worked examples, and full documentation will be released alongside the validation data (mid-July: competition start & warmup phase). The problem formulation and evaluation methodology are described in our competition proposal paper, available now on arXiv.

7. Can I participate in both tracks?

Yes, teams may submit to both the Main Track and the Small Models Track independently.

8. How is the winner determined?

Submissions are ranked by accuracy in correctly identifying the robust model from each pair on the held-out test set. The Main Track and Small Models Track are ranked separately.

9. What if I have more questions?

Reach out to us at [email protected] or open a discussion on the GitHub repository. We aim to respond within 48 hours and will update this FAQ regularly.

Who We Are

The organizing team combines expertise in evaluation, interpretability, and data collection with top-tier olympiad-level mathematicians — a combination that lets us curate the high-quality data this challenge is built on. Several of us are also closely involved in organizing AIMO, bringing first-hand experience running competitions at scale.

Michal Štefánik
Michal Štefánik
NII, Japan
Lead organizer, website, communication. Researcher in evaluation, interpretability, and the robustness of language models.
Philipp MondorfPM
Philipp Mondorf
LMU Munich
Core organizer, baselines, leaderboard. PhD researcher on reasoning and generalization in language models.
Andreas WaldisAW
Andreas Waldis
Univ. of Tübingen
Core organizer, baselines, beta testing. PostDoc and junior group lead working on the reliability of language models.
Qianying LiuQL
Qianying Liu
NII, Japan
Math expertise, evaluation. Researcher in multilingual reasoning and interpretability, with a PhD from Kyoto University.
Chuan YangCY
Chuan Yang
Fuzhou University
Math expertise, data collection. Associate Professor of mathematics and optimization.
Michal SpiegelMS
Michal Spiegel
Masaryk University
Data collection, robustness labels. Researcher in mechanistic interpretability of reasoning and length generalization.
Josef KuchařJK
Josef Kuchař
Masaryk University
Infrastructure, Codabench. Researcher bridging software engineering and machine learning.
Marek KadlčíkMK
Marek Kadlčík
Masaryk University
Symbolic chains, permutations. PhD student in reasoning and generalization of neural models.
Adam Vawda-OomerjeeAV
Adam Vawda-Oomerjee
NII / UCL
Infrastructure, Codabench, beta testing. Researcher in reasoning, reinforcement learning, and world models.
Chaoran LiuCL
Chaoran Liu
NII, Japan
Infrastructure, participant support. Project Associate Professor working on large language models.
Simon FriederSF
Simon Frieder
Oxford / AIMO
Competition design, AIMO liaison. AIMO Prize Manager and founder of the Benchmarks & Baselines non-profit.
Barbara PlankBP
Barbara Plank
LMU Munich
Evaluation, data collection. Professor and Chair for AI & Computational Linguistics, leading the MaiNLP lab.
Fazl BarezFB
Fazl Barez
Univ. of Oxford
Interpretability methods. Senior Researcher leading the Technical Safety & Governance lab.
Pontus StenetorpPS
Pontus Stenetorp
UCL / NII
Evaluation, adversarial data, promotion. Professor of NLP and Deputy Director of the UCL Centre for AI.

Get in Touch

For any questions about the challenge, compute access proposals, or technical issues, please reach out — we're happy to help.

✉️

Primary contacts: [email protected] and the AIMO-Interp Discord.

We aim to respond within 48 hours. For technical questions, please also consider opening an issue or discussion on the corresponding GitHub repository. This FAQ is regularly updated.