LoopsBench: A Benchmark for Long-Horizon Software Engineering with Coding Agents
Researchers from Microsoft and Nanjing University have introduced LoopsBench, a benchmark for long-horizon software engineering with coding agents. It represents software tasks as dependency DAGs to evaluate an agent's ability to maintain plans, advance dependent tasks, preserve completed work, and control regressions over extended execution. This benchmark reflects the shift from one-off tool calls to continuous software development systems.