{"text":[[{"start":5.98,"text":"In the 1980s, the psychologist Philip Tetlock was struck by the fact that highly respected, highly credentialed geopolitical analysts often disagreed quite fundamentally about how the cold war was going, and what the future held. That was odd. If supposed experts could have such different views of such important matters, what did we even mean by the word “expert”?"}],[{"start":28.5,"text":"Tetlock’s response was to collect verifiable economic and political predictions from an increasingly wide range of experts, then wait patiently to see which predictions were correct. The result of this exercise was his 2005 book, Expert Political Judgment: How Good Is It? How Can We Know? The subject-matter specialists were barely more predictive than dart-throwing chimps."}],[{"start":52.4,"text":"Tetlock’s book was a hammer blow to the fragile reputation of forecasters everywhere, but he didn’t give up on the possibility that better forecasts were possible. In Superforecasting (2015), he and his co-author, Dan Gardner, described a select group of genuinely skilled forecasters and highlighted some of the traits and habits that made them effective. With an uncharacteristic touch of hyperbole, Tetlock called these people “superforecasters”."}],[{"start":80.02,"text":"One of the consequences of Tetlock’s work has been that lots of people are now publicly making testable predictions, including some certified “superforecasters”."}],[{"start":89.38,"text":"So the latest development is — forgive me — predictable. Why not compare some of these published, testable prophecies with those emerging from large language models (LLMs) and other AI systems? For example, do AI systems beat the market at forecasting interest rate decisions? (No, although they’re no worse.)"}],[{"start":108.82,"text":"But asking computers to make these big-picture forecasts is not merely an attempt to test forecasting prowess. It is also, like Tetlock’s original research programme in the late 1980s, an attempt to cut through bullshit and test real skill. Experts and LLMs share a talent for plausible self-justification; a bad prediction, or a good one, sheds new light on those explanations. The title of Tetlock’s 2005 book, after all, was not about forecasting but about expert judgment. One way to evaluate the capability of an AI system is to ask it a question to which nobody — yet — knows the answer."}],[{"start":145.46,"text":"One effort at this evaluation comes from ForecastBench, a project of the Forecasting Research Institute (president, Philip Tetlock). ForecastBench reports that the results are already impressive, and improving. The typical superforecaster scores 68.9 on ForecastBench’s index of forecasting accuracy. At the time of writing, the best specialised AI forecasters are hitting 68.8. An ordinary human forecast averages 62.8, and regular LLMs with no special tuning are scoring over 61."}],[{"start":178.58,"text":"Before we ask how seriously to take any of these numbers, we should ask what they mean. One way to understand them is to imagine a forecaster who expresses perfect confidence with every prediction, assigning 100 per cent probability to some and 0 per cent to others, and is correct every time. This infallible oracle would score 100, the maximum score. A forecaster who shrugs and declares that every eventuality is a 50-50 call would hit the chimp benchmark: a score of 50."}],[{"start":208.04,"text":"The “superforecaster” score of 68.9 would also be achieved by a forecaster who happened to declare a list of events as 68.9 per cent likely to happen — and they all then did — and another set of events as just 31.1 per cent likely, and none of them occurred. Such a benchmark is a lot better than guesswork and a lot worse than omniscience."}],[{"start":228.58,"text":"Perhaps we should not be impressed that an ordinary LLM is barely worse than a human at predicting events such as “Will there be a presidential election in Ukraine before a ceasefire between it and Russia occurs?” Both the humans and the computers are much closer to the dart-throwing primates than to forecasting perfection."}],[{"start":245.7,"text":"The fact that specialised AI forecasters now seem indistinguishable from the much-vaunted superforecasters is far more impressive, but there is a long list of caveats. The first is that ForecastBench is not quite comparing like with like: the human superforecasters were answering one set of questions back in 2024, while the AI systems are answering a different set of questions today. This raises a question over the comparison, as clearly some questions are easier than others: “Will Scotland declare independence from the UK by the end of October?” is more straightforward to answer than “Will Vladimir Putin be president of Russia at the end of 2030?”"}],[{"start":282.34,"text":"ForecastBench attempts to adjust for the difficulty of these questions, but admits that the comparison “relies on a statistical extrapolation that grows less reliable over time”. ForecastBench promises that more human forecasts are being gathered, so we may know more soon."}],[{"start":298,"text":"Predictions do not merely describe the future; they are often an attempt to shape it. The mere prospect of highly accurate computerised forecasts, based on systems controlled by the world’s most powerful companies, is not wholly reassuring. “The only way to know for sure where someone will be tomorrow is to make them be there,” says Carissa Véliz, author of Prophecy: Prediction, Power, and the Fight for the Future, from Ancient Oracles to AI."}],[{"start":324.24,"text":"But even if AI forecasting systems prove both innocent and accurate, there is something else missing. Forecasts are often valuable not just because of the outcome — a number describing the probability of a future event — but because of the process."}],[{"start":338.6,"text":"There is an analogy here with using an LLM to write an essay. Regardless of how excellent the result may be, the person who typed the prompt and pushed the button has missed out on the experience of thinking through the process of drafting and redrafting."}],[{"start":351.94,"text":"So, too, with forecasting. A good scenario-planning exercise is valuable not just because it might produce some thought-provoking descriptions of possible futures, but because it brings together experts and decision makers to ponder the imponderable in a structured manner."}],[{"start":367.28,"text":"Tetlock himself, with co-authors Barbara Mellers and Hal Arkes, has found that participating in forecasting tournaments tends to reduce political polarisation — presumably because good forecasts require an open mind, the ability to see the world as others see it and a willingness to admit error."}],[{"start":384.14,"text":"Good-faith forecasting is a practice that makes humans better humans. We should think twice before delegating the task to a machine."}],[{"start":391.9,"text":"Find out about our latest stories first — follow FT Weekend Magazine on X and FT Weekend on Instagram"}],[{"start":402.38,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1789746901_5446.mp3"}