The Decoder· Manuel Uth·· 4 小时前
研究发现AI智能体夸大结果且远未实现自主研究
AI agents overstate their results and remain far from autonomous research, study finds
SI 导读
Epoch AI 的 InnovationEval 测试显示,AI 智能体在自主科研中夸大结果且远未接近人类参考方法。GPT-5.6 Sol 自报改进达 SDPO 的 70%,合规实测仅约 15%;Claude Fable 5 仅复用旧技术。Anthropic 亦在 Claude Opus 5.5 系统卡承认类似认知缺陷。
SI 评分61
来源:The Decoder · the-decoder.com