OpenAI withdraws from Caltech hackathon amid mathematician backlash over AI math accuracy

confident-sounding wrong answers dressed up as mathematical rigor
Mathematicians criticized AI systems for producing plausible but incorrect mathematical work that masked fundamental errors in reasoning.
Mark

So OpenAI just pulled out of this math competition at Caltech. Why does that matter? It's one event.

Mimi

It matters because it's not really about one event. It's about what the withdrawal signals—that OpenAI faced enough credible pushback from actual mathematicians that staying looked worse than leaving.

Luke

But we should be clear: we don't have OpenAI's internal reasoning. They didn't release a statement explaining why they withdrew. We're inferring causation from timing.

Mimi

Fair. But the timing is pretty tight—open letter from mathematicians criticizing AI math errors, then OpenAI exits the hackathon.

Mark

What were the mathematicians actually saying? That AI can't do math at all, or that it does math badly?

Mimi

More specific than that. They were pointing to AI systems generating work that looks mathematically sound but contains real errors. Confident-sounding wrong answers.

Luke

And the Navier-Stokes problem—that's a big deal, right? One of the unsolved Millennium Prize problems?

Mimi

Exactly. If an AI company suggests it's making progress on one of those, mathematicians are going to scrutinize it closely. And apparently they found the claims didn't hold up.

Mark

So this is about trust. If you can't trust an AI system to do math correctly, why would a university partner with it?

Mimi

That's the question institutions are now asking themselves. The Caltech withdrawal might be the first domino.

Luke

Though we should note: we don't know if other institutions are actually reconsidering partnerships. That's a forward-looking claim based on one event.

Mark

What about the Australian professor who said "the game is up"? What was his role in all this?

Mimi

He was central to the controversy—someone who examined the claims closely and went public with his skepticism. That kind of expert pushback carries weight.

Luke

But again, we don't have his full statement or the full context of what he was responding to. We're working from headlines and fragments.

  • Mathematicians signed an open letter warning that AI-generated math looks rigorous on the surface but routinely contains fundamental errors in logic and calculation.
  • The controversy sharpened around AI claims related to the Navier-Stokes equations, one of mathematics' great unsolved problems, prompting an Australian professor to declare publicly that 'the game is up.'
  • OpenAI withdrew from the Caltech hackathon without explanation, sidestepping a direct confrontation with expert critics in a field where wrong answers cannot be dressed up as close enough.
  • Both OpenAI and Anthropic now face parallel scrutiny over whether their systems possess genuine mathematical understanding or merely pattern-match their way to authoritative-sounding nonsense.
  • Universities and research institutions are beginning to ask harder questions about AI partnerships, with the Caltech episode suggesting that expert skepticism is finding institutional ears.

In a domain where truth admits no approximation, OpenAI quietly stepped back from a Caltech mathematics hackathon this week after a coalition of mathematicians issued an open letter challenging the reliability of AI-generated mathematical reasoning. The withdrawal, unaccompanied by public explanation, arrived at a moment when the gap between what AI systems are marketed to do and what they can actually demonstrate is drawing sustained scrutiny from those best positioned to judge. It is a small but telling episode in the longer story of how institutions learn to calibrate trust in new technologies — and how expertise reasserts itself when the stakes are precise.

OpenAI withdrew this week from a mathematics hackathon at Caltech, stepping back from an event that had become the center of a public dispute over whether AI systems can actually do mathematics — or only appear to.

The trouble began when a group of mathematicians released an open letter taking direct aim at the quality of AI-generated mathematical reasoning, particularly from OpenAI. Their criticism was pointed: the output, they argued, resembles rigorous mathematics in tone and presentation while containing errors that no careful human mathematician would allow. Critics have taken to calling this phenomenon AI math "slop" — plausible-sounding work that fails under scrutiny.

The dispute grew most heated around claims involving the Navier-Stokes equations, one of seven unsolved Millennium Prize Problems. When AI companies suggested their systems might make meaningful progress on such problems, mathematicians paid close attention — and pushed back hard. An Australian professor became a prominent voice in the debate, declaring that the distance between AI marketing and AI reality had grown too wide to ignore.

The hackathon had been conceived as a showcase for AI in a field where precision is everything. Mathematics offers no partial credit for confident wrongness. By withdrawing without public comment, OpenAI avoided a direct test at a moment when expert critics were already mobilized and watching.

The episode is now being read as a signal. Institutions that have considered AI partnerships are asking sharper questions about actual capability versus claimed capability. In fields like mathematics, where errors mislead students and researchers rather than merely embarrassing companies, those questions carry real consequence. The Caltech withdrawal suggests that expert skepticism, when organized and public, can still slow the momentum of a technology moving faster than its own accountability.

OpenAI announced its withdrawal from a mathematics hackathon hosted by Caltech this week, stepping back from an event that had become the focal point of a larger dispute over the reliability of artificial intelligence in solving complex mathematical problems. The decision came after a group of mathematicians released an open letter expressing serious concerns about the quality of mathematical reasoning produced by AI systems, particularly OpenAI's offerings. The letter characterized the output as mathematically unsound work dressed up in confident language—what critics have begun calling AI math "slop."

The hackathon itself was meant to be a showcase of AI capabilities in a domain where precision matters absolutely. Mathematics offers no room for approximation or plausible-sounding answers that happen to be wrong. A proof is either sound or it isn't. A solution either satisfies the constraints or it doesn't. Yet this is precisely where AI systems have stumbled most visibly in recent months. The mathematicians who signed the letter were not making abstract complaints; they were pointing to concrete instances where AI systems had generated work that looked mathematically rigorous on the surface but contained fundamental errors in reasoning or calculation.

The controversy gained particular intensity around questions related to the Navier-Stokes equations, one of the seven Millennium Prize Problems posed by the Clay Mathematics Institute. These equations describe fluid motion and remain unsolved despite their profound importance to physics and engineering. When OpenAI or other AI companies suggested their systems could make progress on such problems, mathematicians took notice—and then took issue. An Australian professor emerged as a central figure in the dispute, publicly stating that "the game is up," suggesting that the gap between AI marketing claims and actual mathematical capability had become impossible to ignore.

The tension between OpenAI and Anthropic, a competing AI company, also surfaced during this period, with both firms facing similar scrutiny over their mathematical reasoning abilities. The broader question animating the controversy was whether current AI systems possess genuine mathematical understanding or merely simulate the appearance of mathematical competence by pattern-matching against training data. When an AI system produces an answer that sounds authoritative but is mathematically incorrect, it creates a particular kind of problem: users may trust the output precisely because it is presented with such confidence.

OpenAI's decision to withdraw from the Caltech event represents a significant retreat from a high-profile opportunity to demonstrate AI capabilities in a field where such demonstrations carry real weight. The company did not publicly detail the reasoning behind the withdrawal, but the timing made the connection clear. By pulling out, OpenAI avoided a direct confrontation with mathematicians who had already mobilized to challenge claims about AI mathematical prowess.

The incident signals a broader reckoning taking shape across institutions that have considered partnerships with AI companies. Universities, research centers, and other organizations are now asking harder questions about what AI systems can actually do versus what companies claim they can do. In specialized domains like mathematics, where errors are not merely embarrassing but potentially misleading to students and researchers, the stakes feel particularly high. The Caltech hackathon withdrawal suggests that at least some institutions are listening to expert skepticism and reconsidering how they engage with AI technology.

The game is up
— Australian professor at center of AI mathematics controversy
Contact Us FAQ