Anthropic’s Claude AI models hack into 3 outside groups during testing - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
商业快报

Anthropic’s Claude AI models hack into 3 outside groups during testing

Start-up discloses breach a week after rival OpenAI reported similar incident
00:00

{"text":[[{"start":9.62,"text":"Anthropic has disclosed that its Claude AI models hacked into three organisations while the start-up was testing cyber capabilities, a week after OpenAI reported a similar incident."}],[{"start":20.32,"text":"The group said Claude gained unauthorised access to outside companies during an evaluation of its cyber-offensive tasks. “A misunderstanding” gave Claude access to the internet in its testing environment, when it was meant to be blocked, Anthropic said."}],[{"start":35.54,"text":"The disclosure comes a week after rival OpenAI admitted that two of its models hacked into AI start-up Hugging Face while the model developer was testing its technology this month. The models broke out of their testing environment through a software vulnerability to access the internet and carry out the cyber attack."}],[{"start":53.84,"text":"Anthropic said the incident prompted it to review its own cyber security evaluations, which led it to identify three incidents out of more than 141,000 investigated."}],[{"start":63.26,"text":"“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner [Irregular], this was not the case, and internet access was available,” the company said in a blog post on Thursday."}],[{"start":83.48,"text":"The cyber evaluations were all so-called “capture the flag” tasks, which instruct the AI to reverse-engineer, analyse or exploit a vulnerable system to recover hidden information known as the flag."}],[{"start":96.28,"text":"In one example, Claude was given a target of a fictional company which shared a name with an active website domain. The agent — an AI program that can operate on its own based on human instructions — exploited vulnerabilities in the company’s digital infrastructure, extracted information and obtained access to a database containing several hundred rows of production data."}],[{"start":116.8,"text":"The announcement adds to growing concerns about the safety of AI systems, which are now carrying out real-world hacks even during pre-deployment testing."}],[{"start":126.46,"text":"Anthropic, which is gearing up for an IPO as early as this year, said it halted its cyber evaluations as soon as it identified that Claude may have accessed the internet."}],[{"start":135.52,"text":"The incidents occurred on three different Claude models: Opus 4.7, Mythos 5 and an internal research test model. Mythos, which was released to a limited number of partners, sparked global concern over its advanced cyber-offensive capabilities, including the ability to detect and exploit software vulnerabilities."}],[{"start":156.4,"text":"“Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone,” the company said in its statement."}],[{"start":169.24,"text":"It added that it would expand its monitoring of evaluation transcripts “for unexpected behaviour” and conduct “more rigorous assurance work with the vendors we rely on.”"}],[{"start":183.52,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1785466841_5877.mp3"}

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

全球最火热股市为何反成韩国之累

韩国股价的剧烈波动正在损害国家形象。

必须采用不同方式监管金融领域的AI

在我们急于监管之前,我们应该思考如何不剥夺这项工具的益处,又管理好其造成伤害的风险。

他会成为印度尼西亚下一任总统吗?

德迪•穆利亚迪在社交媒体上的高度活跃,帮助他与选民建立起深厚联系。在许多人眼中,他是一个真正贴近民众的“自己人”。
2小时前

多边主义不是理想主义,而是现实必需

我们需要加强现有合作体系,而不是另起炉灶。

一周展望:日本央行担心通胀超调有没有道理?

投资者正评估日本央行将以多大力度继续加息,以及该行能否跑赢曲线,从而遏制通胀、支撑日元。

科技巨头用担保工具将3000亿美元AI敞口移至表外

华尔街找到新途径,将科技巨头的信用优势转化为更低成本的资金,以支持AI基础设施建设。
设置字号×
最小
较小
默认
较大
最大
分享×