Zhou, Junwei, Sun, Zhen, Li, Binyu, Zhou, Jiangyu, Pan, Yuexi, Wang, Hengyu, Ren, Honghe, Jia, Xiaohan, Zhou, Xueyang, Cao, Xiaoyu, Chen, Yongchao, Feng, Yuanning, Wu, Junhao, Zhang, Cheng, Chen, Sijia, Xue, Haoyu, You, Chengsong, Wang, Huan, Wu, Koutian, Gao, Peigan, Wu, Jiakun, Li, Wenzhe, Shang, Ergan, Zheng, Qingyuan, Zhou, Jingjing, Jia, Ruixuan, Xu, Yan, Zhang, Hongrui, Ma, Xiao-Han, Cheng, Zhengxiang, Hao, Yuexing, Mai, Liting, Ji, Xianglin, Zhang, Wenjun, Chen, Zhuofan, Huang, Yixiao, Wang, Chi, Hua, Wenyue, Hao, Yilun, Zhai, Yuantao, Zhao, Ziyan, Xie, Jingyan
Abstract
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.
Chinese Translation
人工超智能(ASI)要求人工智能超越对现有知识的掌握,向探索未知、创造新知识以及将新思想转化为可验证的结果迈进。然而,当前的人工智能系统的能力仍然主要建立在学习、压缩和应用现有人类知识的基础上。因此,现有的基准测试主要评估人工智能是否能够基于学习到的知识产生正确答案,或者在广泛的人类指导下完成任务。因此,我们推出了ASI-Bench,这是第一个共同评估人工智能系统在一般研究领域中创新探索和自主科学执行能力的基准,并且是第一个在同一研究项目中逐步减少人类方法指导以测试人工智能能够独立进行多远的基准。ASI-Bench由超过40位专家构建,耗费了31,000多个小时的人力,包含11个科学领域的60个项目级研究任务,并逐步减少方法指导,以测试人工智能是否能够独立选择方法、进行研究并产生可验证的结果。所有任务都经过专家审查、人工智能辅助审计、沙盒执行和评分者验证。在18种最先进的代理-模型配置中,平均得分从在全面方法指导下的50.91降至仅指定方法时的29.10,以及代理必须自行确定方法时的26.62。这一急剧下降表明,当前系统仍然严重依赖人类指导,距离自主进行端到端项目级科学研究仍然相去甚远。ASI-Bench向全球开放。我们邀请各地的研究人员和开发者贡献新任务,挑战当今人工智能的极限,并帮助加速人类通往人工超智能的集体道路,网址为 https://asibench.apexin.ai/submit。