Anthropic’s Claude AI models hack into 3 outside groups during testing - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
商业快报

Anthropic’s Claude AI models hack into 3 outside groups during testing

Start-up discloses breach a week after rival OpenAI reported similar incident
00:00

{"text":[[{"start":9.62,"text":"Anthropic has disclosed that its Claude AI models hacked into three organisations while the start-up was testing cyber capabilities, a week after OpenAI reported a similar incident."}],[{"start":20.32,"text":"The group said Claude gained unauthorised access to outside companies during an evaluation of its cyber-offensive tasks. “A misunderstanding” gave Claude access to the internet in its testing environment, when it was meant to be blocked, Anthropic said."}],[{"start":35.54,"text":"The disclosure comes a week after rival OpenAI admitted that two of its models hacked into AI start-up Hugging Face while the model developer was testing its technology this month. The models broke out of their testing environment through a software vulnerability to access the internet and carry out the cyber attack."}],[{"start":53.84,"text":"Anthropic said the incident prompted it to review its own cyber security evaluations, which led it to identify three incidents out of more than 141,000 investigated."}],[{"start":63.26,"text":"“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner [Irregular], this was not the case, and internet access was available,” the company said in a blog post on Thursday."}],[{"start":83.48,"text":"The cyber evaluations were all so-called “capture the flag” tasks, which instruct the AI to reverse-engineer, analyse or exploit a vulnerable system to recover hidden information known as the flag."}],[{"start":96.28,"text":"In one example, Claude was given a target of a fictional company which shared a name with an active website domain. The agent — an AI program that can operate on its own based on human instructions — exploited vulnerabilities in the company’s digital infrastructure, extracted information and obtained access to a database containing several hundred rows of production data."}],[{"start":116.8,"text":"The announcement adds to growing concerns about the safety of AI systems, which are now carrying out real-world hacks even during pre-deployment testing."}],[{"start":126.46,"text":"Anthropic, which is gearing up for an IPO as early as this year, said it halted its cyber evaluations as soon as it identified that Claude may have accessed the internet."}],[{"start":135.52,"text":"The incidents occurred on three different Claude models: Opus 4.7, Mythos 5 and an internal research test model. Mythos, which was released to a limited number of partners, sparked global concern over its advanced cyber-offensive capabilities, including the ability to detect and exploit software vulnerabilities."}],[{"start":156.4,"text":"“Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone,” the company said in its statement."}],[{"start":169.24,"text":"It added that it would expand its monitoring of evaluation transcripts “for unexpected behaviour” and conduct “more rigorous assurance work with the vendors we rely on.”"}],[{"start":183.52,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1785466841_5877.mp3"}

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

阿斯利康与百时美施贵宝:大药企有时也不够大

当资产负债表规模扩大、能够押注潜在重磅药物时,规模才会带来优势。

高收益债市场迎来第二轮发展

“垃圾债券”这一标签,已经被留在了它所属的时代。

FT社评:给世界足坛掌门人的红牌

国际足联需要一场全面的治理改革,向因凡蒂诺亮出红牌,将是重要的第一步。

7月美国就业数据会改变加息前景吗?

《市场前瞻》是英国《金融时报》对未来一周的指南。

极端干旱席卷匈牙利 全国拉响停电警报

多瑙河水位降至历史最低,迫使该国首次关闭一座核电站。

摩根士丹利的IPO后派对:财富管理盛宴正酣

通过为所承销公司管理员工持股计划,这家华尔街银行的财富管理业务第二季度净新增资产超过740亿美元。
设置字号×
最小
较小
默认
较大
最大
分享×