Epoch and METR Release MirrorCode Benchmark for AI Long-Horizon Programming
Epoch and METR have released MirrorCode, a benchmark to evaluate AI systems on programming tasks that take humans weeks to complete. Results show that advanced models like Claude Opus 4.7 can solve some tasks in hours at high cost, but the hardest tasks remain unsolved. The benchmark highlights rapid AI improvement, with models from a year ago scoring only 30%.