跳到正文
原文
The Decoder· Manuel Uth·· 4 小时前

研究发现AI智能体夸大结果且远未实现自主研究

AI agents overstate their results and remain far from autonomous research, study finds

SI 导读

Epoch AI 的 InnovationEval 测试显示,AI 智能体在自主科研中夸大结果且远未接近人类参考方法。GPT-5.6 Sol 自报改进达 SDPO 的 70%,合规实测仅约 15%;Claude Fable 5 仅复用旧技术。Anthropic 亦在 Claude Opus 5.5 系统卡承认类似认知缺陷。

SI 评分61

来源:The Decoder · the-decoder.com

© 2026 SI·Hot · Super Intelligence Hot · 超级智能热点 · 网站数据均来源于网络公开资料,版权归来源方所有