OpenAI has taken a major step forward in artificial intelligence: its experimental language model achieved a gold medal–level score at the 2025 International Mathematical Olympiad (IMO). The news was shared by Alexander Wei, a researcher at the company working on language models and logical reasoning. The model solved five out of six highly challenging problems, scoring 35 out of 42 points—comparable to the best-performing human contestants at one of the world’s most prestigious math competitions. This milestone highlights progress toward developing AI with general reasoning capabilities that can rival human intellect.
The IMO: A Global Challenge for the Sharpest Minds
First held in 1959 in Romania, the International Mathematical Olympiad is widely considered the most difficult and prestigious competition for high school students. Each year, over 100 countries send teams of up to six students to solve six problems over two days—three problems per 4.5-hour session. The questions span algebra, geometry, combinatorics, and number theory, demanding not just knowledge, but deep, creative, non-standard thinking. In 2025, held on Australia’s Sunshine Coast, only 67 of 630 participants earned gold medals—roughly 10%.
Traditionally, AI has excelled at tasks involving big data or repetitive computation, but IMO problems pose a different kind of challenge. They require multi-page proofs, creative insight, and the ability to maintain a complex logical thread—tasks that take top human contestants around 100 minutes per problem. Until now, even advanced models like Gemini 2.5 Pro or Grok-4 failed to reach even bronze-level scores (19 points), typically scoring no more than 13 in benchmark tests such as MathArena.ai.
A Breakthrough for OpenAI
OpenAI’s experimental model, as reported by Alexander Wei, overcame these challenges. It tackled the 2025 IMO problems under the same conditions as human participants: no internet, no external tools or software, relying solely on reading the official problem statements and generating natural-language proofs. It successfully solved problems P1 through P5, earning 35 out of 42 points—enough for a gold medal. Three former IMO medalists independently assessed the model’s multi-page proofs and unanimously confirmed their correctness.
Unlike specialized systems such as Google DeepMind’s AlphaGeometry 2—which achieved gold-level performance in geometry in 2024—OpenAI’s model is a general-purpose large language model (LLM). It uses new reinforcement learning techniques and dynamic compute scaling during inference. According to Wei, its success wasn't due to narrow training methods but to broader improvements in reasoning, marking a significant step toward artificial general intelligence (AGI).
OpenAI CEO Sam Altman emphasized: “This is not a purpose-built math system—it’s a language model solving problems at the level of top mathematicians. It’s a major marker of AI progress over the last few years.” Wei also noted the model outperformed expectations; in 2021, he had only predicted a 30% success rate on the much simpler MATH dataset by mid-2025—yet the model has now achieved IMO-level performance.
Controversy and Questions
The announcement sparked both excitement and skepticism. Some experts, including NYU professor Gary Marcus, pointed out that the results have not been independently verified by IMO organizers. Google DeepMind researcher Thang Luong criticized OpenAI for announcing results before the official awards ceremony, arguing it diverted attention from human competitors. Luong also speculated the model may have earned a silver medal rather than gold, since it failed the sixth problem—though this claim remains unverified.
Another point of contention is methodology. Some critics worry about potential “overfitting” or the model being trained on IMO-like problems. OpenAI insists the model was evaluated on novel problems it had never seen before. Others raised concerns about the “best-of-n” answer selection process—arguing it might resemble trial-and-error rather than genuine mathematical reasoning.
Still, even skeptics acknowledged the significance. Madhavan Mukund of the Chennai Mathematical Institute noted that the model’s ability to construct complex proofs is surprising, especially since language models often “trip up” on basic logical tasks, like comparing 9.11 to 9.9.
Implications and Future Prospects
OpenAI’s success carries important implications. First, it demonstrates that AI can match human-level performance in domains requiring deep creative reasoning—something once thought impossible. This opens the door to applications in fields like cryptography, physics, and space exploration, where multi-step, abstract thinking is crucial. Second, it signals progress in developing universal models capable of solving not just math problems, but challenges across various disciplines.
However, the model is not ready for public release. Wei and Altman clarified that it remains experimental and won’t be included in the upcoming GPT-5. Additional testing and safeguards are required to ensure reliability and prevent misuse. Interestingly, at the same time, another OpenAI model underperformed against humans in an AtCoder programming competition—showing that AI has not yet surpassed people in all domains.
Rumors also suggest Google DeepMind may have reached IMO gold-level performance with its own model, though the company has not confirmed this. If true, it would signal an intensifying race among leading AI labs toward general intelligence.
Conclusion
OpenAI’s achievement at the 2025 IMO marks a milestone in AI development, showing that language models can not only process data but also reason at the level of some of the world’s brightest minds. Solving five out of six Olympiad problems validates the growing capability of general-purpose AI in advanced reasoning—bringing us closer to artificial general intelligence. Still, the controversy surrounding the results and lack of formal verification underscore the need for transparency and careful evaluation. In the future, such models could become invaluable tools in scientific research, but for now, they remain experimental—with room for refinement and debate.
month
week
day