Back to feed
News Story
SSignal86
DeepTech深科技
1 sources

AI Refutes Century-Old Math Conjectures but Lacks Ability to Invent New Theories; World Models May Be Key

Recent AI systems like GPT-5.6 Sol and Claude Fable 5 have helped mathematicians refute long-standing conjectures such as Maxwell's and Jacobian's, sparking debate about AI's role in scientific discovery. However, DeepMind researcher Tom Zahavy's position paper argues that large language models lack abductive reasoning, hindering their ability to invent new theories, and suggests building physically consistent world models to bridge this gap.

SynthePulse Insight · AI deep reading

AI Falsifies Century-Old Conjectures but Struggles to Invent New Theories: Can World Models Bridge the Abductive Gap?

Version 1 · 1 source

As AI successively overturns Maxwell's conjecture and the Jacobian conjecture, Google DeepMind researcher Tom Zahavy's old paper 'Large Language Models Cannot Jump' has reignited debate: AI excels at deduction and induction but lacks the abductive ability to propose new explanatory hypotheses. Can world models be the key to crossing this chasm?

  • On July 29, three American mathematicians used GPT-5.6 Sol to construct a counterexample, overturning Maxwell's conjecture by finding at least 24 non-degenerate equilibrium points, far exceeding the upper bound of 16.
  • On July 20, an Anthropic number theorist announced the refutation of the Jacobian conjecture with Claude Fable 5, with Terence Tao subsequently verifying it using ChatGPT.
  • In a January 2026 paper, Tom Zahavy argued that large language models lack abductive reasoning—the ability to infer new explanations from anomalous phenomena—which is crucial for inventing new theories.
  • Zahavy believes that building physically consistent world models is key to bridging this gap, mentioning the Genie series as a possible path.
  • Three paths (robotics, world models, evolutionary search) are converging, and world models may become the medium connecting virtual and real worlds.
Open section navigationAI Consecutively Falsifies Century-Old Conjectures

AI Consecutively Falsifies Century-Old Conjectures

On July 29, three American mathematicians uploaded a paper titled 'Maxwell's Conjecture Is Wrong' to arXiv, constructing a counterexample that overturns this classical physics conjecture. The acknowledgments show that the core idea came from GPT-5.6 Sol. They constructed a 'five-charge triangular bipyramid' and found at least 24 non-degenerate equilibrium points, far exceeding the upper bound of 16 set by the conjecture.

Earlier, on July 20, an Anthropic number theorist announced the refutation of the Jacobian conjecture with Claude Fable 5, and Terence Tao subsequently used ChatGPT to follow up and publicly verify the process. On July 24, at the International Congress of Mathematicians, he warned that the mathematical community is entering an era of 'proof surplus' rather than 'proof scarcity.'

The Argument of 'Large Language Models Cannot Jump'

Amid the wave of AI-assisted scientific discovery, a position paper titled 'Large Language Models Cannot Jump' by Tom Zahavy, a Google DeepMind researcher and core developer of AlphaProof, written in January 2026, has been widely circulated. He uses Peirce's trichotomy to divide reasoning into deduction, induction, and abduction, pointing out that abduction is the only form of reasoning that can generate new premises and is currently a shortcoming of large models.

He cites Einstein's 1907 thought experiment of imagining free fall to propose the equivalence principle as an example, explaining that such 'embodied simulation' relies on bodily sensations, whereas large language models are like 'high-dimensional Chinese rooms,' capable of processing physical language but lacking a sensory channel for physical objects. Therefore, he argues that AI needs to build physically consistent world models to achieve the 'jump.'

Are Counterexamples an AI Invention?

Taking the falsification of Maxwell's conjecture as an example, AI proposed a geometric configuration based on existing frameworks, but it did not invent a new theory or question the system of assumptions. Therefore, according to Zahavy's classification, it still falls within the category of efficient exploration under deduction. However, AI's suggestion was a direction that the mathematical community had not attempted in 150 years, so it cannot be simply reduced to symbolic search.

Zahavy himself acknowledges that induction, deduction, and abduction are intertwined in practice. He responded that his article does not claim 'LLMs are a dead end,' but rather explores what AI still lacks to achieve Einstein-level discoveries.

World Models and the Convergence of Three Paths

The 'Einstein Test' proposed by Google DeepMind chairman Hassabis shows that today's systems cannot independently derive general relativity under limited knowledge. Zahavy's paper proposes 'embodied simulation' approaches such as world models as a solution.

Regarding the sim-real gap, there are three paths: robotics provides real physical experience, world models provide high-fidelity digital simulations, and AlphaEvolve uses evolutionary code generation to bypass 'physical intuition.' These three paths are converging: the robotics route uses world models as training environments, and evolutionary search also uses world models to replace offline experiments.

Credibility boundary

This article's information primarily comes from DeepTech's analysis, which includes transcriptions of papers, tweets, and speeches. The specific details of Maxwell's conjecture falsification (such as the 24 equilibrium points) come from the paper, but the description of AI's contribution is based on the paper's acknowledgments; the Jacobian conjecture refutation is an announcement by an Anthropic number theorist and has not been independently verified; Zahavy's paper's viewpoints are directly quoted.

Insight takeaway

AI has made significant progress in deduction and induction, but abductive reasoning remains a shortfall. World models may become the medium connecting virtual and real worlds, but whether they can enable AI to pose questions humans have not thought of remains unknown.

Primary report

DeepTech深科技

Primary source