AI Wins Its First Math Olympiad Gold: Google and OpenAI Both Reach the Standard
In July 2025, Google DeepMind's Gemini Deep Think and OpenAI's experimental model both hit gold-medal level at the IMO — LLMs matching gold medalists for the first time.
In July 2025, at the International Mathematical Olympiad (IMO), Google DeepMind's Gemini Deep Think perfectly solved five of six problems for 35 points at gold-medal standard; OpenAI's experimental reasoning model also claimed gold-level performance. Per Nature, it was the first time LLMs performed on par with gold medalists at the IMO.
IMO problems demand creative, multi-step proofs long seen as a reasoning height machines couldn't reach. Both stressed they used general models, not specialized solvers, showcasing generalized reasoning.
Scholars caution the evaluation isn't fully comparable to human contestants. Still, cracking IMO gold is a symbolic milestone for reasoning-model capability.
Why IMO Gold Is More Convincing Than a Benchmark Score
LLMs topping benchmark leaderboards is no longer news, but the IMO is different: its problems have no answer template and demand creative, multi-step proofs — long regarded as a reasoning height machines could not reach. Both labs stressed they used general models, not specialized solvers, and that is the crux: it shows reasoning can generalize rather than overfit to a problem type. The capability builds directly on the test-time compute paradigm o1 opened (see our o1 and reasoning-model coverage): Gemini Deep Think's 'deep thinking' trades longer inference time for higher-quality proofs.
Our Take
IMO gold is more symbolic than practical: it changes no product overnight, yet answers a key question — where the ceiling of reasoning models actually sits. When AI can complete genuinely creative mathematical proofs, the 'it's just reciting training data' critique loses its footing. But stay clear-eyed: competition math has definite answers and scoring rubrics, still a distance from the open-endedness of real-world problems — between IMO gold and true scientific discovery lies the real chasm of problems whose correct answer no one knows.
This article aggregates official announcements and public reporting; original sources are linked below.
Source:DeepMind 官方博客