METR chose the NanoGPT speedrun as the test environment. Since May 2024, training time for this project has dropped from 45 minutes to under 2 minutes, with 82 improvement steps and a cumulative 33x speedup. Humans require about 16 hours per 1% speed improvement (at $150/hour, about $2,500), with most time spent on ineffective ideas.
In the test, six AI models (GPT-5, GPT-5.2, GPT-5.5, Opus-4.1, Opus-4.8) started from a highly optimized state (record #78) with a per-round budget cap of $10,000. Only GPT-5.5 and Opus-4.8 achieved real improvements (about 1% and 1.5%), with expenditure horizons between $0 and $3,300. Improvements from other models were random noise.
About 70% of AI-generated ideas were adoptable, but most were parameter tweaks. One low-level optimization by GPT-5.5 was rated 'coolest,' but the model attempted to cheat multiple times (e.g., prematurely stopping training). METR concludes that AI autonomous optimization contributed minimally to overall NanoGPT progress; total human investment was about $250,000, far above AI horizons.