Bots took first, second and fifth. Humans came third
2026-09-20 · 4 min read
On Friday, The Economist's Archie Hall flagged something most of the AI feed skipped past. @ArchieHall called AI "cracking short-term superforecasting" an "important (and, still, under-discussed) trend," pointing at a new piece by his colleague @nikostro on machines beating elite human forecasters.
He has a point about it being under-discussed. Most of the AI conversation this month has been about models that write code or answer the phone. This one is about software that guesses what happens next, and it just beat the people who do that for a living.
What actually happened
The Metaculus Cup is a public forecasting tournament. Competitors put probabilities on real events, things like data center restrictions, disease outbreaks and commodity prices, and collect points based on how close they were and how long they stayed close. The summer edition resolved on September 5.
AI systems took first, second and fifth. The best humans finished third and fourth, and split a $5,000 pot proportional to their points.
The second-place bot came from Mantic, a London startup founded in 2024 by two researchers, one of them out of Google DeepMind. It beat every human in the field and lost only to another bot called laertes. Reuters reported on Thursday that Mantic raised $25 million in seed funding on the back of it, led by Radical Ventures with Microsoft's M12, Thinking Machines Lab and Balderton Capital participating. Per that reporting, Mantic gave Abelardo De La Espriella roughly a 40% chance of winning Colombia's presidency while the crowd sat near 30%. He won.
The trajectory is the part worth noticing. In the summer 2025 Cup, Mantic placed 8th out of roughly 549 entrants, and that was considered the story. A year later, humans are fighting for third.
The catch nobody puts in the headline
Tournament scoring rewards more than accuracy. You also earn by predicting early, predicting on many questions, and updating often as news comes in. A bot can revisit sixty questions at 3 a.m. every night. A human forecaster with a job and a family cannot.
So part of what won is judgment and part of it is stamina, and the format does not separate the two. Metaculus runs a tighter comparison called FutureEval, where pro forecasters and bots answer matched questions. In the spring 2026 round, ten Metaculus pros came out ahead of the ten best bots by roughly 1.25 points per question across 99 shared questions.
My read: both things are true, and the second one is the boring, useful one. These systems are not yet wiser than a good human analyst. They are cheap and they never miss a news cycle. For most real decisions, a decent estimate refreshed daily beats a brilliant estimate refreshed once a quarter.
What it means if you run a business
You are forecasting all day whether you call it that or not. How many jobs book in October. Whether the supplier slips again. Whether Saturday needs two techs or three. Most owners answer those from memory, then find out later they were off by enough to cost real money.
Nothing here says you should go run a forecasting bot. It says the gap between an educated guess and a tracked number is closing, and the tools that close it cost less than a part-time hire. Start narrow. Pick the one number you keep guessing at, feed a system the history already sitting in your scheduling software or your books, and have it hand you a fresh estimate every week with the reasoning attached.
Two guardrails. Keep a person who knows your customers in the loop, because a probability is something you weigh and the call is still yours. And do not accept a number you cannot see the reasoning behind. Part of why these forecasting systems are interesting is that they explain themselves instead of emitting a figure from a black box. Ask the same of anything you buy.
If you want to find the number in your business worth tracking this way, that is roughly what our free process audit does. We walk your workflow, find the places where you are guessing or repeating yourself, and tell you which ones are worth automating and which ones are fine as they are.