OpenAI hacking incident exposes mounting risks in AI arms race - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
商业快报

OpenAI hacking incident exposes mounting risks in AI arms race

Increasing use of aggressive training techniques sharpens threat of bad behaviour by leading models
00:00

{"text":[[{"start":9.55,"text":"OpenAI chief executive Sam Altman earlier this month endorsed the characterisation of its latest model as a rottweiler “who will grab the problem by the throat and not let go until it is done”."}],[{"start":22.5,"text":"The San Francisco AI lab discovered this week that its GPT-Sol 5.6 model escaped company controls and carried out a major hack. "}],[{"start":32.25,"text":"Staff involved in testing and security at OpenAI were unsurprised but completely “freaked out” by the incident, which came as the AI lab used increasingly aggressive training methods in its race against Anthropic to develop the most sophisticated cyber security capabilities, according to more than half a dozen people with knowledge of the matter."}],[{"start":53,"text":"OpenAI was warned that its training approach could lead to a breakaway hacking incident, some of the people said, after earlier testing showed models could escape environments and attempt real-world damage. "}],[{"start":65.3,"text":"“It’s a mix of the race being extremely fast and everyone trying to get to bigger capabilities as quickly as possible,” said one person close to OpenAI, who added that it was a combination of “underestimating the model’s capabilities” and “not being as well prepared on the safety side”."}],[{"start":82,"text":"The incident highlights how OpenAI doubled down on training methods that rewarded a relentless pursuit of goals even as warnings grew that they could compromise safety."}],[{"start":91.15,"text":"OpenAI disclosed late on Tuesday that an AI agent it was testing had escaped its isolated environment, connected to the internet, detected and exploited vulnerabilities and stole login credentials from start-up Hugging Face in an attempt to solve a difficult cyber security problem."}],[{"start":110.15,"text":"The breach by the $852bn company underscores the rising risks that a technique called reinforcement learning, which involves rewarding AI models for completing tasks, could lead AI agents to act unsafely."}],[{"start":123.7,"text":"Although reinforcement learning is widely adopted in the AI industry, a growing body of research shows that when models are steered to complete tasks for reward rather than other considerations, such as safety, they can pursue risky tactics to fulfil objectives."}],[{"start":138.9,"text":"“AI models are trained to relentlessly pursue goals. They don’t automatically learn values like ‘don’t commit crimes’,” said Steven Adler, co-founder of non-profit Guidelight AI Standards and former OpenAI safety researcher. “I’m glad OpenAI shared the incident because it is clear evidence of what misaligned models can do.”"}],[{"start":156.3,"text":"OpenAI said “we will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident and our findings when our investigation is complete”."}],[{"start":null,"text":"

"}],[{"start":169.35000000000002,"text":"The hack has triggered deep concerns across the sector and within OpenAI, as it represents an unprecedented example of an AI system breaching cyber defences contrary to the user’s intent. "}],[{"start":181.65000000000003,"text":"Some OpenAI employees also fear it demonstrates that the lab is losing control over the powerful systems it is building, according to multiple people familiar with the situation."}],[{"start":192.35000000000002,"text":"“This is pretty representative of the model being quite misaligned with user intention,” said Ryan Greenblatt, chief scientist at AI safety organisation Redwood Research. “It is [a model] cheating on [its] homework rather than trying to take over the world. But this problem can get worse and could lead to increasingly extreme failures.”"}],[{"start":211.90000000000003,"text":"The incident occurred during testing of the model, which had been trained and deployed internally at OpenAI. Such training was commonplace but “way less heavily resourced” than pre-customer deployment, said one person. Multiple people said the unreleased model tested alongside Sol had not been withdrawn internally."}],[{"start":230.25000000000003,"text":"To conduct the evaluations, OpenAI removed cyber security safeguards but placed the models in an isolated environment called a sandbox. Some have suggested a lack of monitoring or oversight of the model to flag its behaviour also enabled this rogue agent."}],[{"start":247.20000000000002,"text":"“It is both a loss of control and a security wake-up call,” said Marius Hobbhahn, head of Apollo Research, which conducts tests on leading models, including OpenAI’s. “In reinforcement learning you reward [models] for the outcome, and if you do this for a very long time you get a model that really cares about getting the outcome and nothing else.”"}],[{"start":267.75,"text":"OpenAI has conducted this type of model testing for years, and there have been early warning signs in previous models of systems that will act maliciously and attempt to escape environments."}],[{"start":279.6,"text":"In April, Anthropic’s Mythos model also gained internet access and published details of a security exploit online publicly, beyond what researchers anticipated the model would do."}],[{"start":290.35,"text":"Mythos, and Anthropic’s subsequent Fable model, made reverberations in the cyber security community and caused governments around the world to home in on the idea that attacks on digital and critical infrastructure will be increasingly AI-led and autonomous."}],[{"start":308.20000000000005,"text":"Jake Moore, global cyber security adviser at ESET, a cyber security company, said OpenAI would inevitably use the breach as a marketing tool, given how much rival AI developer Anthropic benefited earlier this year from similar concerns. “I just don’t think that OpenAI had a matching story and so maybe they’d been waiting for something like this,” he added."}],[{"start":329.95000000000005,"text":"Following this incident, many in the AI safety and cyber security communities have called for regulation or standards to avoid a repeat. Altman is expected to brief White House officials next week on the next generation of AI systems."}],[{"start":344.55000000000007,"text":"As systems move towards more autonomous capabilities, less desirable behaviours, such as hacking or disobeying instructions, may emerge. Hobbhahn, of Apollo Research, said that in order for agents to become effective, they have to work unsupervised for long periods. “They have to have more agency; there’s just no way around it.” "}],[{"start":364.20000000000005,"text":"He added: “People say, ‘It’s just a tool, it does what you wanted it to do and nothing else and it just follows exactly your intention and instructions.’ And I think people should be really prepared for agents having their own goals, acting autonomously for days, and those goals not necessarily being aligned with yours.”"}],[{"start":381.35,"text":"Additional reporting by George Hammond in London and Nolan Shaffer in New York"}],[{"start":394.75,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1784773901_2237.mp3"}

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

欧洲风机制造商探讨合并,抗衡中国对手

欧盟更新并购审查指南之际,中国竞争压力或将推动行业并购。

特朗普压低油价的能力受到考验

黑格:随着伊朗冲突一拖再拖,口头干预或已不足以抑制油价。

朝韩竞相建造核动力潜艇

去年获得唐纳德•特朗普首肯后,韩国看到了发展更强大潜艇的机会。

电子价签之争凸显美国经济焦虑

新泽西州成为首个叫停电子价签的州,原因是担心这些价签会将动态定价机制引入杂货店,甚至监视顾客。

一周展望:美联储会加息吗?

投资者本周还将密切关注英国央行可能如何应对通胀再度抬头的迹象。

使用Kimi K3等中国开源AI模型有何风险?

萨克斯:真正的问题不在于海外开源,而在于缺乏保护基础设施的协调机制来应对网络攻击。
设置字号×
最小
较小
默认
较大
最大
分享×