Back to feed
News Story
SSignal86
量子位
1 sources

Open-Source Models Close Gap to Opus 5 with Prime Intellect's Multi-Agent Harness

Prime Intellect introduced a multi-agent harness that enables open-source models to approach top-tier closed-source models in a nanoGPT optimizer speedrun. In experiments, Kimi K3 with the harness achieved 2930 steps, surpassing GPT-5.6 Sol and coming close to Opus 5, showing that open-source models can catch up through efficient research processes.

SynthePulse Insight · AI deep reading

Open-Source Models + Harness Close the Gap with Closed-Source Flagships: The Trial-and-Error Throughput Revolution in AI Research

Version 1 · 1 source

Prime Intellect's latest experiments show that the open-source model Kimi K3, aided by an Agent Harness, achieved 2930 steps, approaching Claude Opus 5's 2920 steps and surpassing GPT-5.6 Sol. This challenges the narrative that the strongest closed-source models will achieve RSI first, suggesting that AI research capability may depend more on trial-and-error throughput than on single-model intelligence.

  • Prime Intellect had 18 frontier models speedrun a nanoGPT optimizer task; Kimi K3 with Harness achieved 2930 steps, surpassing GPT-5.6 Sol's 3042 steps and only 10 steps behind Opus 5's 2920.
  • The experiment included 153 fully autonomous runs, with the longest exceeding 8 days, but no experiment proposed a truly novel method; all effective techniques had similar ideas in existing research.
  • The real differentiator was trial-and-error throughput: stronger models were better at handling noise, retrying with different seeds, and building their own tools, rather than coming up with unique ideas.
  • Fable 5 ran the furthest with 2726 steps, from a 3290-step baseline to a 2600-step human record, consuming 564 steps of optimization space, about 82%.
  • Prime Intellect proposes a Multi-Agent Harness approach: use cheap open-source models for monitoring and implementation, while strong models only make key judgments, enabling more experiments at lower cost.
  • Kimi K3's Harness provided a persistent IPython kernel, allowing it to actively overturn its own hypotheses, seen as evidence of the importance of research infrastructure in RSI.
Open section navigationExperimental Design: nanoGPT Optimizer Speedrun

Experimental Design: nanoGPT Optimizer Speedrun

Prime Intellect designed a nanoGPT optimizer speedrun experiment, where 18 frontier models (including Fable 5, Opus 5, GPT-5.6 Sol, Kimi K3, etc.) autonomously completed an optimization task on 8 H200s. Each agent received a code repository, rulebook, and objective instructions, but did not know the specific improvement direction and had to explore on its own.

The goal was to reduce the validation loss of a 124M-parameter GPT to below 3.28 using as few training steps as possible, with a baseline of 3290 steps and a human best of 2600 steps. The experiment was fully offline and locked in a sandbox to prevent models from directly copying existing answers.

Results: No New Methods, but Differences in Trial-and-Error

In the end, 153 fully autonomous experiments were conducted, with the longest lasting over 8 days, but no experiment proposed a truly novel method. Effective techniques such as normalization, learning rate adjustments, and weight averaging all had similar ideas in existing research.

However, the gap between models did not come from ideas but from trial-and-error throughput. Stronger models did not easily give up on an idea; they would change seeds, re-run ablations, and even retest old solutions after recipe changes. For example, Opus 5 re-tuned beta2 to set a new record, and Fable 5 re-explored old solutions to save 31 steps.

Stronger models were better at handling noise, judging whether small improvements were real or random fluctuations, and even discovering that GPU non-determinism caused loss variations when repeating the same seed, leading them to redo the experimental process around this phenomenon.

Kimi K3 and the Harness Breakthrough

Kimi K3, with the Prime Agent Harness, achieved 2930 steps, surpassing GPT-5.6 Sol's 3042 steps and only 10 steps behind Opus 5's 2920. This shows that open-source models with a Harness can approach or even surpass some top closed-source models.

The Harness provided Kimi K3 with a persistent IPython kernel, enabling it to actively overturn its own hypotheses and improve efficiency. This is seen as evidence of the importance of research infrastructure in RSI.

Trial-and-Error Throughput: The New Bottleneck in AI Research

The experiments show that ideas themselves may not be scarce; what really affects performance is the ability to try more directions, quickly eliminate errors, capture weak signals, and stack them. Research capability increasingly resembles a trial-and-error throughput problem.

Prime Intellect is therefore shifting focus to a Multi-Agent Harness approach: let cheap open-source models handle monitoring and implementation, while strong models only make key judgments, enabling more experiments at lower cost. This suggests that improving AI research capability does not necessarily depend on stronger models; improving the 'trial-and-error machine' is equally effective.

Challenges to the RSI Narrative

This result challenges the classic RSI narrative that the strongest closed-source model will cross the AI-developing-AI critical point first and then winner-take-all. Systems of open-source models plus Harness can also reach near the top tier.

Prime Intellect plans to extend Speedrun to more parts of the training stack and scale up the experiments. This points to a possibility: the acceleration of AI developing AI may not come from a single model's intelligence explosion, but from increasingly cheaper and faster AI research machines.

Credibility boundary

This article is based on QbitAI's report on Prime Intellect's research; all data comes from that report and has not been confirmed by primary sources or official confirmation. Some conclusions (such as 'ideas are not scarce') are inferences from the report and should be treated with caution.

Insight takeaway

Improving AI research capability may no longer depend on the intellectual breakthrough of a single model, but on trial-and-error throughput—through cheaper models and more efficient Harnesses, open-source systems can approach or even surpass closed-source flagships, providing a new development path for AI4AI.

Primary report

量子位

Primary source