Back to feed
News Story
APriority79
机器之心
1 sources

Meta AI Wins Gold in Five STEM Olympiads with Pure Reasoning

Meta announced that its AI model achieved gold or gold-level results in five STEM olympiads, including perfect scores on two physics theory exams, all without using any tools to test pure reasoning. The model is an internal version of the Muse Spark series, utilizing multi-agent orchestration and parallel reasoning. This marks a potential resurgence for Meta in the AI race.

SynthePulse Insight · AI deep reading

Meta's Pure Reasoning Wins Five Olympiad Golds: A Revival Signal or a Broken Benchmark?

Version 1 · 1 source

Meta announced that its AI model achieved gold or gold-level results in five STEM Olympiads, including perfect scores in two theoretical physics exams, all without using tools. Does this report card still shine in 2026, or is the Olympiad as a reasoning benchmark losing its validity?

  • Meta announced on 𝕏 that its AI model achieved gold or gold-level results in five STEM Olympiads, with perfect scores in two theoretical physics exams.
  • Meta emphasized pure reasoning: no search, coding, or calculators, to test true reasoning ability.
  • In the report card, IMO is listed as gold, while IChO and RMM are listed as 'gold-level performance,' implying they were not officially evaluated.
  • Meta researchers added that the model comes from an internal version of the Muse Spark series, using multi-agent orchestration and parallel reasoning, and it answered a question correctly where the official answer was wrong.
  • At the 2026 IMO, Huawei's Celia and Xiaohongshu's dots-note-3.0 achieved perfect scores, making Meta's gold less prominent in the rankings.
  • The Olympiad as a reasoning benchmark is failing: within a year, it went from 'unattainable' to 'multiple labs can achieve full marks.'
Open section navigationSubtle Differences in the Report Card and Wording

Subtle Differences in the Report Card and Wording

Meta announced on 𝕏 that its AI model achieved gold or gold-level results in all five STEM Olympiads, with perfect scores in the theoretical exams of the Asian Physics Olympiad (APhO) and the International Physics Olympiad (IPhO). In the report card, IMO is explicitly listed as 'gold,' while IChO and RMM are listed as 'gold-level performance.' This wording typically implies that the results were not officially evaluated by the competition committees but were self-assessed against the year's gold medal cutoff scores. Meta also thanked the contestants and committees for their support, suggesting that participation in some events was coordinated with the committees, which is more formal than OpenAI's self-scoring approach in 2025.

Meta specifically emphasized that to test pure reasoning ability, the model was disabled from using all tools: no search, no coding, no calculators. This constraint is significant because in the past year, many top Olympiad results relied on agent pipelines of 'strong model + verifier + multi-round refinement.' For example, a researcher used Gemini 3.1 Pro to build a simple agent that achieved perfect scores on IPhO 2025 theoretical problems five times, but the author also noted that data contamination could not be ruled out. Meta deliberately abandoned this path, highlighting its pure reasoning positioning. However, physics competitions include experimental exams, and Meta's report likely covers only the theoretical portion.

From Llama to Muse: Meta's Road to Reconstruction

2025 was a difficult year for Meta AI: Llama 4 Maverick scored only 18 on the Artificial Analysis intelligence index, Behemoth was delayed due to insufficient capability, and Zuckerberg restructured by bringing in Alexandr Wang and the Scale AI team to form Meta Superintelligence Labs (MSL), rebuilding the pretraining, data, and RL post-training systems from scratch. On April 8, 2026, MSL released Muse Spark, natively multimodal with 262k context, and the intelligence index jumped from 18 to 52. On July 9, Muse Spark 1.1 launched, focusing on agentic and coding, becoming Meta's first paid API model. This week, Muse Code, a terminal coding agent based on Muse Spark 1.2, entered beta. This Olympiad report card is more like a preview of the next-generation Muse model.

The 2026 Coordinate: Gold Medals Are No Longer Scarce

In July 2025, when OpenAI and Google DeepMind both won IMO gold medals (35/42, solving five of six problems), OpenAI researchers called it a 'moonshot moment.' But a year later, the coordinate system has changed: after the July 2026 IMO in Shanghai, Huawei's Celia and Xiaohongshu's dots-note-3.0 announced perfect scores of 42/42, with Xiaohongshu claiming that no large model had ever achieved a perfect score under official IMO evaluation. Additionally, there are reports that Anthropic's Claude Opus 5 passed all six problems in one go without using agent frameworks or tools. Therefore, Meta's IMO gold is not particularly prominent in 2026; what truly carries weight are the two perfect scores in theoretical physics—provided they can be independently verified.

More concerning is that the Olympiad as a reasoning benchmark is failing. Meta's stated reason is 'to understand whether we are making real progress in reasoning,' but when the same competitions go from 'unattainable' to 'multiple labs can achieve full marks' within a year, the information content of this benchmark decays sharply. Terence Tao pointed out in 2025 that what AI can achieve depends heavily on the testing method itself. Olympiad problems have standard answers, clear scoring criteria, and ample historical problem banks, which are conditions that real-world research problems do not have.

Has Meta Stepped Up?

From Muse Spark to Muse Code and now this report card, Meta is at least no longer a company that only maintains its presence through open-source narratives. But to claim a return to the top tier, it needs to deliver more, especially a product that can actually be used. Currently, Meta has not released a technical blog or research report, and independent verification of the results remains key.

Credibility boundary

This article is primarily based on Machine Heart's retelling of Meta's official 𝕏 statement and researchers' tweets. In the report card, IMO is listed as gold, while IChO and RMM are listed as 'gold-level performance,' which is not officially evaluated and is Meta's self-assessment. The perfect scores in physics refer only to the theoretical portion; the experimental portion is not mentioned. Meta has not released a technical report, and independent verification is lacking.

Insight takeaway

Meta's pure reasoning Olympiad report card demonstrates progress in its model's capabilities, but in the context of the failing Olympiad benchmark in 2026, its true value depends on whether it can be independently verified and translated into practical products.

Primary report

机器之心

Primary source