{"text":[[{"start":7.9,"text":"This article is an on-site version of our The AI Shift newsletter. Premium subscribers can sign up here to get the newsletter delivered every Thursday. Standard subscribers can upgrade to Premium here, or explore all FT newsletters"}],[{"start":22.15,"text":"Welcome back to the AI Shift, our weekly exploration of how models and agents are reshaping work. For this edition, we’re taking a look at recent developments in ‘fine-tuning’ cheaper AI models to develop professional domain specialism, and the possibly diverging implications for top knowledge economy firms and the big US AI companies."}],[{"start":43.099999999999994,"text":"John writes"}],[{"start":44.849999999999994,"text":"For several years now we’ve been hearing claims that domain-specific AI models in areas like medicine, education and law could perform similarly to frontier models from OpenAI and Anthropic at tasks within their specialist field. The specialist models are typically created by fine-tuning a cheaper and less generally capable model — in giving it a crash course in medical diagnosis, for example, and tweaking its internal workings so it becomes more likely to produce correct responses."}],[{"start":75,"text":"These experiments have certainly been interesting, and occasionally impressive, but upon closer inspection the results have generally flattered to deceive. In many cases, purported matching of bigger and more expensive models’ capabilities was only true when comparing to earlier non-frontier generations of the big US labs’ products."}],[{"start":95.35,"text":"But we’re now seeing some interesting developments that suggest there may be some cases where specialist fine-tuning — especially when done at the level of individual firms rather than whole professional domains — may not only match but outperform frontier models, and perhaps durably."}],[{"start":110.5,"text":"Until recently, most leading fine-tuning exercises looked like one from earlier this year by the legal AI firm Harvey AI, which worked with the platform Fireworks AI to train a specialist legal AI agent. In practical terms, they took the cheap Chinese model Kimi 2.6, had it answer questions from a legal AI agent benchmark test while monitoring what was going on under the hood, discarded the instances where it gave the wrong answer and used the correct performances to retrain the model to produce more results like that. The process resulted in almost a 40 per cent improvement, enough to bring it in line with frontier model performance, at just an 11th of the cost."}],[{"start":150.75,"text":"But the results ultimately fall into the category of impressive but ephemeral. Anthropic released a new version of its Opus model days before the results were published, setting new records on the same legal tests that would most likely have pushed the fine-tuned model back into second place. The cost savings, though, remain hugely impressive. There will surely be many low-stakes business cases where the trade-off of comparable performance but vastly lower expense is worth it, raising questions of how much frontier use is truly needed."}],[{"start":182.6,"text":"But a new case last month took things a step further, when Ray Dalio’s investment firm Bridgewater Associates partnered with AI platform company Thinking Machines Lab (founded by former OpenAI CEO Mira Murati) to fine-tune a model based specifically on how its own investment managers do their work. As with the legal example, the results significantly outperformed frontier models, this time at a 14th of the cost. Crucially, however, use of the firm’s own proprietary records and its highly specialist staff’s knowhow may make these gains more durable."}],[{"start":218.2,"text":"Bridgewater had its own experts write bespoke prompts that framed questions in a way that guided the models to the correct answers. This produced solid gains, but they still topped out below 80 per cent accuracy."}],[{"start":231.25,"text":"A prompt only imparts the expertise a professional is able to put into words — what is much better is to learn from their actions. For the fine-tuning step they put together a set of tasks drawn from their own investors’ daily workflows, and crucially also had staff ensure that the ideal responses which would guide the model’s training did not just represent ‘correct answers’ but ‘exactly how our investment professionals would approach this’."}],[{"start":256.95,"text":"At the end of the process they had a bespoke model whose behaviour had been tuned towards Bridgewater’s own assessment of excellence, taking it up to 85 per cent accuracy — an almost 30 per cent reduction in errors compared to the frontier models — at a tiny fraction of the cost."}],[{"start":275.45,"text":"The prospect of gains in performance and cost being long lasting is tantalising. To date, one reason specialist medical models have only outperformed frontier models until they didn’t is that newer generalist models had been exposed to more training data and could better reason their way to the correct answers. But in cases like Bridgewater’s, the performance boost came from proprietary information and judgment that the models cannot access."}],[{"start":301.15,"text":"This doesn’t mean the expert fine-tuning advantage will not fade eventually, but it will surely take longer (we should note these results have only been reported by the companies involved and not independently verified). It also shows why many AI firms are hiring knowledge workers on temporary contracts to pipe exactly this type of specialised professional expertise into their training data."}],[{"start":324.2,"text":"I am left pondering several things."}],[{"start":327.05,"text":"Could this approach fundamentally change what it means to have AI perform tasks within a business setting? Frontier models are often likened to a smart and indefatigable intern, but specialist models trained in-house strike me as more akin to a trainee who has spent a couple of years absorbing the firm’s best practices."}],[{"start":347.25,"text":"Will the next winners be ‘model as a service’ companies that help firms create and maintain bespoke models?"}],[{"start":353.5,"text":"Given it is corporate usage that has propelled the massive revenues and valuations of AI companies, is this a huge threat to Anthropic and OpenAI’s business model?"}],[{"start":363.85,"text":"Sarah, what’s your sense of where this might be headed?"}],[{"start":366.65000000000003,"text":"Sarah writes"}],[{"start":368.20000000000005,"text":"I think you’re right, John, that this development would be bad news for companies like Anthropic and OpenAI if it really took off. But I think it would be extremely good news for those businesses which have the technical talent and wherewithal to take advantage of it. It’s not just the fact that it could be much more cost-effective. It would also go some way to addressing people’s uneasiness about the prospect of vast numbers of companies becoming dependent on a handful of proprietary black-box models from a few providers. "}],[{"start":397.35,"text":"Microsoft’s Satya Nadella made this point recently in a post, writing that “the last thing any of us want is a world where every company across every sector is ceding value to a few models that eat everything they see. If all the value is accrued by only a few models, the political economy will simply not tolerate it.” "}],[{"start":415.1,"text":"A world in which companies fine-tune smaller models, based on the expertise of their own staff, sounds preferable. It still raises a lot of interesting questions, though. Some of the experiments you mention, John, used Chinese base models (Bridgewater used Qwen3-235B, for example). I wonder if that opens the possibility of companies getting caught up in a new type of geopolitical risk (although as we saw recently when the White House put export controls on Anthropic’s Fable 5 and Mythos, American models aren’t immune to political risk either)."}],[{"start":450.20000000000005,"text":"Secondly, it’s clear from the Bridgewater example that the tacit and institutional knowledge of the company’s own investment professionals was key to improving the performance of the model. But will employees feel it is in their interest to help fine-tune models for their employers, based on the expertise they have spent decades building up inside their own heads? Will they trust their employers not to extract that knowledge, and then dispense with their services altogether? "}],[{"start":476.20000000000005,"text":"Or will they see it as an opportunity to help create some tools that are less generic than the ones they’re getting from companies like OpenAI and Anthropic, and actually better suited to their needs? I don’t know the answer to this one, which will probably depend to a large extent on the balance of trust and power in each workplace. Readers, do email us and let us know what you think. The email address aishift@ft.com reaches us both."}],[{"start":501.25000000000006,"text":"One last thing . . . "}],[{"start":503.15000000000003,"text":"The AI Shift is now officially a “multi-award-winning” newsletter after we picked up three more gongs last week at the Newsletter Publisher Awards, including the big one (“best overall newsletter”). We are very chuffed! Many thanks to our editors Elaine Moore, Srinidhi Balakrishnan and Alice Fishburn, together with our producer Georgina Quach, who deserve much of the credit."}],[{"start":525.8000000000001,"text":"Recommended reading"}],[{"start":527.45,"text":"This New Yorker piece about how an American family interacts with AI was immersive and thought-provoking. You might want to save it for your commute home as it’s quite long. (Sarah)"}],[{"start":538.0500000000001,"text":"Harvard’s David Deming has a great essay on humanity’s big advantage over AI in how efficiently we learn from (and about) one another (John)"}],[{"start":null,"text":"