AI Takes a Bare Exam: Over a Dozen Institutions Release New Benchmark to Gauge Autonomous Scientific Research
A consortium led by Tsinghua University, including MIT, Harvard, CMU, USTC, and Microsoft Research, has released ASI-Bench, a new benchmark designed to measure AI's scientific autonomy without human methodological guidance. By progressively reducing the hints provided to AI, the benchmark reveals a significant performance drop when specific procedural steps are removed. This work highlights current limitations in AI's ability to conduct independent research and offers a new tool for evaluating AI's scientific capabilities.