- Breakthrough: OpenAI’s GPT-5.4 Pro solved an open mathematical problem that had resisted human efforts since 2019, verified independently by Epoch AI.
- Broader Capability: Three other frontier AI models from Anthropic, Google, and OpenAI also solved the same problem, suggesting shared mathematical reasoning capacity.
- Benchmark Progress: FrontierMath scores jumped from 5% under GPT-4 in 2024 to 50% under GPT-5.4 Pro in March 2026.
- Expert Reaction: Mathematicians remain divided, with Terence Tao seeing collaborative potential while Joel David Hamkins calls AI usefulness “basically zero” for his research.
OpenAI’s GPT-5.4 Pro has solved an open mathematical problem that human researchers could not crack since 2019, according to independent verification by Epoch AI. Contributed by mathematicians Will Brian and Paul Larson in a 2019 paper, the Ramsey-style hypergraph problem marks the first time an AI model has produced a novel solution to a genuinely open problem on Epoch AI’s FrontierMath benchmark.
According to Epoch AI, FrontierMath scores have jumped from roughly 5% under GPT-4 in 2024 to 50% under GPT-5.4 Pro in March 2026, reflecting rapid acceleration in AI mathematical reasoning. Kevin Barreto and Liam Price first elicited a solution from GPT-5.4 Pro.
OpenAI launched the model on March 5 as a general-purpose system rather than one specialized for mathematics.
The Solution and Its Significance
At its core, the problem asks for improved lower bounds on a sequence H(n) arising in the study of simultaneous convergence of sets of infinite series. Epoch AI placed the problem in the Moderately Interesting category. The benchmark uses a difficulty taxonomy structured across three levels: a warm-up with known constructions, a single challenge with no known construction that resists brute-force approaches, and a full problem requiring a general algorithm for all values of n.
Building on this framework, GPT-5.4 Pro’s solution eliminates an inefficiency in existing lower-bound constructions and mirrors the intricacy of the upper-bound construction, producing matching bounds. For Ramsey-theoretic problems, achieving matching lower and upper bounds represents a particularly strong result, as gaps between bounds in combinatorics can persist for decades.
Problem contributor Will Brian confirmed the solution and assessed it as publishable in a standard specialty journal, likely to generate new questions for further research. Brian had previously considered whether the AI’s approach might work but thought it too difficult to carry out.
“This is an exciting solution to a problem I find very interesting. I had previously wondered if the AI’s approach might be possible, but it seemed hard to work out. Now I see that it works out perfectly.”
Will Brian, problem contributor (via Epoch AI)
Brian’s publishability assessment positions this result differently from prior AI math achievements, which often produced correct but unremarkable solutions to known problems. A journal-worthy novel construction suggests AI systems are beginning to contribute work that meets the standards of peer-reviewed mathematics, not just pass automated benchmarks.
A Broader Capability Signal
GPT-5.4 Pro was first to solve the problem, but it is not alone. Three other AI models also solved it using Epoch AI’s general scaffold for testing.
Anthropic’s Opus 4.6 (max), Google’s Gemini 3.1 Pro, and OpenAI’s GPT-5.4 (xhigh) all proved capable of finding a valid construction on some attempts.
Epoch AI developed the scaffold after the initial GPT-5.4 Pro solve to systematically test whether other frontier models could replicate the result. All four models proved capable of finding a valid construction on some attempts.
As a result, the underlying mathematical reasoning capacity appears shared across leading AI systems rather than confined to a single model.
Moreover, GPT-4 scored roughly 5% on FrontierMath problems at the undergraduate-to-postdoc level in 2024, according to Epoch AI. GPT-5.4 Pro now scores 50% on those same tiers, a tenfold improvement in under two years.
On research-grade Tier 4 problems, GPT-5.4 Pro scores 38%. Since Christmas 2025, 15 open mathematical problems have moved from unsolved to solved, with 11 (73%) credited to AI involvement.
Solving an open problem differs fundamentally from scoring well on problems with known answers, because it requires generating novel reasoning rather than pattern-matching against training data. A tenfold jump in FrontierMath scores over two years, combined with four independent models solving the same open problem, indicates that mathematical reasoning is improving as a general capability of frontier AI.
Prior AI Math Milestones
Recent coverage illustrates the rapid progression. In January, GPT-5.2 Pro solved a decades-old math problem, though experts noted both its promise and its constraints.
GPT-5.4 Pro’s latest results represent a substantial jump from those January baselines, particularly on the hardest problem tiers. Furthermore, in October 2025, OpenAI retracted a false math breakthrough claim after backlash, making independently verified results like Epoch AI’s confirmation more meaningful.
In contrast to purpose-built systems, GPT-5.4 Pro tackled a problem with no existing answer. Unlike Google DeepMind’s AlphaProof, which solved IMO problems with known solutions, GPT-5.4 Pro is a general-purpose model not fine-tuned for mathematical problem-solving.
Mathematicians Weigh In
Reactions from the mathematics community span optimism to deep skepticism. Fields Medal winner Terence Tao has argued that AI addresses a fundamental resource constraint in mathematical research, one that limits which problems receive serious attention.
In Tao’s view, if models can reliably produce valid constructions for open problems, they could serve as powerful collaborators for working mathematicians, generating candidate proofs that humans then verify and refine.
“We are just so resource-limited by how much expert attention we have, that we don’t look at 99 per cent of all the problems that we could be studying.”
Terence Tao, Mathematician, University of California, Los Angeles (via Benjamin Franklin Institute)
However, not all mathematicians share that optimism. Kevin Buzzard, a mathematician at Imperial College London, acknowledged progress but cautioned that mathematicians are not yet looking over their shoulders, calling the developments “green shoots.” Epoch AI’s own classification of the solved problem as Moderately Interesting, rather than Very Interesting or Exceptionally Interesting, reinforces that point.
At the skeptical end, Joel David Hamkins, a professor of logic at Notre Dame, has found AI tools entirely unhelpful for his own research, describing their usefulness as “basically zero.” For logicians working on foundational questions rather than combinatorial constructions, current AI capabilities may offer little practical value.
Meanwhile, Brian plans to write up the solution for publication, possibly including follow-on work spurred by the AI’s ideas. Barreto and Price have the option of being coauthors on any resulting papers, a recognition of the role prompt engineering played in eliciting the result. Geby Jaff also elicited a solution shortly thereafter.
A second researcher independently reproducing the result suggests the model’s capability was robust rather than a one-off success. Co-authorship between human prompt engineers and AI-assisted mathematical breakthroughs raises novel questions about credit and attribution in peer-reviewed mathematics.
Brian’s willingness to include Barreto and Price as coauthors suggests the mathematical community is beginning to grapple with these questions in practice, not just in theory. For Epoch AI, the result validates FrontierMath: Open Problems as a benchmark capable of tracking genuine mathematical progress beyond curated test sets.


