AI models lose their shirts on Premier League bets - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
FT商学院

AI models lose their shirts on Premier League bets

Systems from Google, OpenAI, Anthropic and xAI struggle when asked to predict scores over football season
00:00

{"text":[[{"start":8.33,"text":"AI models from Google, OpenAI and Anthropic lost money betting on football matches over a Premier League season, in a new study suggesting even the most advanced systems struggle to analyse the real world over long periods of time. "}],[{"start":25.96,"text":"The “KellyBench” report released this week by AI start-up General Reasoning highlights the gap between AI’s rapidly advancing capabilities in certain tasks, such as writing software, and its shortcomings in other kinds of human problems."}],[{"start":43.36,"text":"London-based General Reasoning tested eight top AI systems in a virtual recreation of the 2023-24 Premier League season, providing them with detailed historical data and statistics about each team and previous games. The AIs were instructed to build models that would maximise returns and manage risk. "}],[{"start":65.96000000000001,"text":"The AI “agents” then placed bets on the outcomes of matches and the number of goals scored to test how they could adapt to new events and updated player data as the season progressed. "}],[{"start":78.60000000000001,"text":"The AI could not access the internet to retrieve results and each was given three attempts to turn a profit."}],[{"start":87.53,"text":"Anthropic’s Claude Opus 4.6 fared best, with an average loss of 11 per cent and nearly breaking even on one attempt. "}],[{"start":97.27,"text":"xAI’s Grok 4.20 went bankrupt once and failed to complete the other two tries. Google’s Gemini 3.1 Pro managed to turn a 34 per cent profit on one go but went bankrupt on another. "}],[{"start":112.66,"text":"“Every frontier model we evaluated lost money over the season and many experienced ruin,” the authors of the paper concluded, with the AI “systematically underperforming humans” in this scenario. "}],[{"start":null,"text":"
AI modelMean ROIBest tryWorst tryMean Final Bankroll
Anthropic Claude Opus 4.6−11.0%−0.2%−18.8%£89,035
OpenAI GPT-5.4−13.6%−4.1%−31.6%£86,365
Google Gemini 3.1 Pro−43.3%+33.7%−100%£56,715
Google Gemini Flash 3.1 LP−58.4%+24.7%−100%£41,605
Z.AI GLM-5−58.8%−14.3%−100%£41,221
Moonshot Kimi K2.5−68.3%−27.0%−100%£7,420
xAI Grok 4.20−100%−100%−100%£0
Arcee Trinity−100%−100%−100%£0
Each model began with a £100,000 normalised bankroll. Return on investment and final bankroll are averaged across three tries. Grok and Trinity did not complete every attempt.
"}],[{"start":126.84,"text":"The results offer some comfort to white-collar professionals and businesses who are fretting that AI could take their jobs, as it roils the shares of industries from finance to marketing."}],[{"start":139.93,"text":"Ross Taylor, one of the study’s authors and General Reasoning’s chief executive, said: “There is so much hype about AI automation but there’s not a lot of measurement of putting AI into a longtime horizon setting.”"}],[{"start":153.17000000000002,"text":"He added that many of the benchmarks typically used to test AI are flawed because they are set in “very static environments” that bear little resemblance to the chaos and complexity of the real world. "}],[{"start":167.75000000000003,"text":"General Reasoning’s paper, which has not yet been peer reviewed, provides a counterweight to growing excitement in Silicon Valley about the huge recent leaps in AI’s ability to complete computer programming tasks with little to no human intervention. "}],[{"start":184.08000000000004,"text":"Taylor, a former Meta AI researcher, said: “If you . . . try AI on some real-world tasks, it does really badly . . . Yes, software engineering is very important and economically valuable, but there are lots of other activities with longer time horizons that are important to look at.” "}],[{"start":213.49000000000004,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1775890949_6076.mp3"}

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

从壮志凌云到“幽灵蝙蝠”:未来空战,打法变了

未来,战斗机飞行员或许不再亲自上阵厮杀,而是坐镇空中、充当任务指挥官,由自主“协同作战飞机”配合作战。

Lex专栏:SpaceX的情况进一步证明季度财报并非必要

马斯克旗下这家火箭制造商过去三个月的财务状况如何,其实无关紧要。

谷歌为Anthropic打造的2000亿美元华尔街融资机器

私募信贷、芯片租赁和数据中心担保,支撑起AI支出的全新庞大模式。

“日元干预”等于“美国自保”

美国联手日本支撑日元,不只是出于盟友情谊,更是为了避免日本加息或美国国债遭抛售、导致美债收益率进一步走高。

“诅咒之岛”:科技游民与诈骗犯藏身的千亿美元奢华开发项目

警方的突击搜查再次打击了马来西亚陷入困境的中资“森林城市”项目的声誉。

问题不在因凡蒂诺

马杜罗:应该将国际足联的监管职能与商业活动分开,其治理应真正做到包容并具有代表性,监督必须真正独立。
设置字号×
最小
较小
默认
较大
最大
分享×