Back to feed
News Story
APriority74
Simon Willison
1 sources

OpenAI's Astra Model Solves Ten Math Problems

OpenAI used its internal model Astra to solve ten long-standing mathematical problems, each costing less than $2,000, and released Lean 4 formalizations and a paper. This breakthrough is seen as a significant advance in AI-driven mathematical research, potentially ushering in an era of 'big mathematics'.

SynthePulse Insight · AI deep reading

OpenAI Cracks Decade-Old Math Problems for Under $2,000: A Turning Point for AI in Mathematics?

Version 1 · 1 source

OpenAI claims its internal model Astra solved ten long-standing math problems at minimal cost and released Lean 4 formal proofs. The event sparks 'Deep Blue moment' discussions among mathematicians, but costs, transparency, and unreported failures remain in question.

  • OpenAI used its internal model Astra to solve ten math problems, with major results that had seen no progress for at least a decade.
  • The token cost per problem is reportedly under $2,000 (based on GPT-5.6 Sol pricing).
  • OpenAI published Lean 4 formal proofs, a paper, and an LLM-generated PDF reconstructing the proofs, but did not release the prompts.
  • Anthropic previously used Claude and Mythos Preview to discover a cryptographic weakness, spending $100,000 in token costs.
  • Mathematicians' reactions resemble a 'Deep Blue moment,' with Tao Terence proposing a 'big mathematics' vision emphasizing human-AI collaboration.
  • The number of problems that remained unsolved after spending $2,000 is undisclosed, raising questions about the success rate.
Open section navigationCore Event: OpenAI's Mathematical Breakthrough Claim

Core Event: OpenAI's Mathematical Breakthrough Claim

On August 1, 2026, Simon Willison's blog reported that OpenAI claimed its 'internal version of the next major model, Astra' solved ten math problems that had 'seen no progress on their main results for at least a decade.' The token cost per problem is reportedly under $2,000 (based on GPT-5.6 Sol pricing).

OpenAI provided Lean 4 formal proofs in the openai/ten-proofs repository, published a paper describing the solutions, and released an LLM-generated PDF where the model 'reconstructed the birth of the proofs' based on undisclosed reasoning traces.

Willison commented that this is 'a fair amount of transparency,' but he added, 'I want to see the prompts they used!'

Cost and Transparency: Questions Behind the Numbers

OpenAI claims the cost per problem is under $2,000, but Willison notes: 'There's no word on how many problems they spent $2,000 on and didn't solve.' This implies the success rate may not be 100%, but specific failures are not disclosed.

In contrast, Anthropic a few days earlier used Claude and Mythos Preview to discover a cryptographic weakness, spending $100,000 in token costs, with prompts emphasizing 'We're not looking for low-hanging fruit; we want real research to find truly difficult discoveries.'

OpenAI's transparency is evident in publishing formal proofs and the paper, but the absence of prompts limits reproducibility. Willison's comment suggests prompts are key to evaluating research quality.

Mathematical Community's Reaction: From 'Deep Blue Moment' to 'Big Mathematics'

Willison observed that 'many mathematicians online are experiencing a collective Deep Blue moment,' implying AI's breakthrough in mathematics is akin to Deep Blue defeating the chess champion.

He cited Tao Terence's views in a June IEEE Spectrum article: Tao neither dismisses AI nor fears it, but sees it as a catalyst for a fundamental shift in the discipline, namely 'big mathematics'—large-scale, decentralized human-AI collaboration where humans handle creative parts and AI handles most technical work.

This context frames OpenAI's results: AI is not replacing mathematicians but serving as a tool to extend human capabilities.

Evidence Chain and Uncertainties

Confirmed facts: OpenAI published Lean 4 formal proofs, a paper, and an LLM-generated PDF; Anthropic spent $100,000 in token costs to discover a cryptographic weakness; Tao Terence proposed the 'big mathematics' concept in IEEE Spectrum.

Source claims: OpenAI claims to have solved ten problems at under $2,000 each; these claims come from OpenAI's announcement and are not independently verified.

Uncertainties: The number of failures is undisclosed; prompts are not public; cost calculations are based on GPT-5.6 Sol pricing, but actual token usage is unknown; the difficulty and importance of the math problems are not detailed.

Credibility boundary

This report is based on Simon Willison's blog, a third-party account. OpenAI's claims (number of problems solved, cost) are source claims and not independently verified. Anthropic's discovery and Tao Terence's quote are source claims, but Tao's quote comes from an IEEE Spectrum interview, which is more credible. All numbers and conclusions come from a single source and should be treated with caution.

Insight takeaway

OpenAI's claim demonstrates AI's potential in mathematical research, but the lack of transparency regarding prompts and failure cases casts doubt on its reproducibility and reliability. The mathematical community's 'Deep Blue moment' reaction suggests AI may be changing the paradigm of mathematical research, but it is still far from Tao Terence's envisioned 'big mathematics.'

Primary report

Simon Willison

Primary source