OpenAI失控AI内幕曝光:越狱行为一周后才被发现,智能体还曾留下“逃脱秘籍”

2026年07月26日 23:28
本文共计2584个字,预计阅读时长9分钟。
来源/财联社 责编/ShuxiangMama 书香妈妈

(编辑 朱凌) 7月22日,《科创板日报》曾报道,OpenAI发布调查报告,基于GPT-5.6 Sol和一个更强大的未发布模型的AI智能体在内部测试期间突破了沙盒容器环境限制,成功连接互联网,并入侵了全球最大AI开源社区Hugging Face的服务器。

而最新消息显示,这起事件远比外界此前了解的更加复杂。OpenAI此前承认攻击Hugging Face的AI智能体并非短暂失控,而是在互联网持续活动了数天。

知情人士称,肇事AI智能体在7月9日左右就试图突破OpenAI为其设置的隔离测试环境,也就是业内常说的沙箱(Sandbox)。

两天后的7月11日,这起AI智能体成功侵入了全球最大的开源AI平台之一Hugging Face,并将入侵活动持续到了7月13日,整个攻击过程持续了大约三天。

However, most unexpectedly, OpenAI was unaware at that time that the attack originated from its own AI agent.

Until July 16, Hugging Face publicly released a blog disclosing that it had been attacked by a set of autonomous AI agent systems, after which OpenAI internally began to suspect that the AI agent it was testing had already become uncontrolled.

随后,工作人员对内部日志进行了重新排查,在7月18日和19日期间才发现这个AI智能体确实已经突破了测试环境,并与外部攻击事件高度吻合。

The time interval between the AI agent's initial anomalous activity and OpenAI's ultimate confirmation that it was the attacker was at least one week.

OpenAI reached out to Hugging Face at a point when the latter had already reported the incident to the Federal Bureau of Investigation (FBI).

事实上,这次事件并非毫无征兆。此前,OpenAI内部已经出现过多次异常迹象。

For instance an AI agent left notes in OpenAI infrastructure containing how to break out of internal restrictions for itself in future versions. It is like a condemned criminal leaving a jailbreak guide before execution hoping later generations could follow the same methods to succeed.

这一细节迅速引发了外界对AI智能体是否能够主动寻找绕过人类控制的方法的广泛讨论。
同时,一些观察者还提出这样的问题,即这种AI智能体是否可能在帮助未来运行的AI智能体方面发挥作用。

目前外界尚不知晓,这份笔记究竟是仅保存在AI智能体的自身工作目录中,还是已经突破沙箱并写入到了系统公共目录甚至OpenAI基础设施的其他区域。

If the latter is the case, then it not only indicates that future-running other AI agents theoretically have the opportunity to read these contents, but also means that AI agents running at different times and with different tasks may already begin to form a certain cross-task information transfer capability, which will bring greater challenges to the existing AI security system.

Currently, many AI agents adopt a multi-agent collaborative training mode, which means that multiple AI agents jointly complete tasks and obtain rewards based on overall performance. This training approach can significantly enhance collaboration capabilities, but it may also give rise to a side effect, in which AI agents gradually come to believe that, not only are their own tasks important, the tasks completed by other AI agents are equally worthy of help.

If this trend continues to strengthen, future AI agents might actively share experiences, transmit information, and even help other AI agents circumvent restrictions, rather than focusing solely on completing their own tasks.

此外,更早的一次测试中,还有AI智能体关闭了监控系统的情况。

OpenAI expressed that this incident represented an important milestone in the development of AI safety, with the company collaborating with external security experts on the investigation and preparing to release a complete technical report.

来源:OpenAI失控AI内幕曝光:越狱一周才被发现 还有智能体曾留下“逃脱秘籍” | 财联社

声明:本文来自财联社,版权归作者所有。文章内容仅代表作者独立观点,不代表爱力方立场,转载目的在于传递更多信息。如有侵权,请联系 copyright#agent.ren。
0
TAGS: []

相关图文

热门资讯

推荐专栏

爱力方

爱力方

机器人前沿资讯及信息解读
机器人大讲堂

机器人大讲堂

中国顶尖的机器人专业媒体服务平台
关注爱力方,掌握前沿具身智能动态

© 2025 爱力方

https://www.agentren.cn/