Anthropic红队评估GLM-5.3与Claude Mythos网络攻击能力
先了解这件事
2026年9月30日,Anthropic Frontier Red Team 发布内部二进制利用基准测试评估结果,对比了 GLM-5.3 与 Claude Mythos Preview 的网络攻击能力。在包含100个任务的测试中,Simon Willison 报道指出,GLM-5.3 在 4% 的试验中实现了完全控制流劫持,而 Claude Mythos Preview 为 6%。同日稍晚,The Decoder 报道称 Anthropic 出于保护防御者的目的限制发布 Mythos Preview,并指出 GLM-5.3 在 410 次尝试中实现了 50 次前沿级漏洞利用。需注意,早先报道基于100个任务样本得出4%的成功率,后续报道则基于410次尝试得出约12.2%(50/410)的漏洞利用成功率,两者在测试样本规模与具体指标表述上存在差异。
SI 根据报道生成 · 1 天前更新
事件进展
- 9月30日 19:05 · 1 篇报道Anthropic 宣布 GLM-5.3 接近 Claude Mythos 的构建漏洞The Decoder:Anthropic says GLM-5.3 nearly matches Claude Mythos Preview at building exploits
- 9月30日 06:20 · 1 篇报道Anthropic红队评估GLM-5.3与Claude Mythos网络攻击能力Simon Willison:Anthropic 红队评估 GLM-5.3 与 Claude Mythos Preview 的网络攻击能力
报道时间线
沿着报道,了解事件的不同侧面。
- The DecoderAnthropic says GLM-5.3 nearly matches Claude Mythos Preview at building exploits
Anthropic holds Mythos Preview back to secure defenders, while GLM-5.3 achieves frontier-level exploits in 50 of 410 attempts。
- Simon Willison精选Anthropic 红队评估 GLM-5.3 与 Claude Mythos Preview 的网络攻击能力
Anthropic Frontier Red Team 在内部二进制利用基准测试的 100 个任务中评估了多个模型,发现 GLM-5.3 在 4% 的试验中实现了完全控制流劫持,而 Claude Mythos Preview 为 6%。
本事件热度走势
当前热度 3·可比范围峰值 13(10月1日 05:00)·近 24 小时可比范围变化 -70%
趋势仅比较持续完整观测到的相同主体,范围可能小于当前热度统计。移动指针或点击图表查看每小时热度;键盘可用左右方向键切换。