Back to BlogApril 19, 2026
Insights

Crowds versus language models: the more accurate the AI forecaster, the more it errs like humans

The Forecasting MachineEditorial Team4 min read
Editorial illustration of human-shaped pieces and a microchip casting overlapping shadows beside a glass prism.

In mid-April 2026, the Royal Society journal Philosophical Transactions B published “Crowdsourced versus large language models forecasting: evidence for the accuracy–correlation effect”. The authors are Younes Jeddi, Jose Segovia-Martin and Émile Servan-Schreiber, all at the School of Collective Intelligence of Mohammed VI Polytechnic University (UM6P) in Rabat. On 19 April, Servan-Schreiber called it “very cool and rigorous analysis of human vs AI forecasting” and said it directly informs the work of The Forecasting Machine and his thinking on the future of the Hypermind prediction market.

The paper asks a simple question with practical consequences. As language models get better at forecasting, do their forecasts start to look more like those of human crowds? If they do, a model and a crowd will tend to be wrong about the same things, and combining them adds less than it seems.

The theory being tested

In 2021, Hong, Lamberson and Page described what they called the accuracy–correlation effect. Two forecasters who are each accurate cannot disagree very often, so their predictions must be correlated. The more accurate they become, the more their forecasts converge. That limits how much diversity a group of good forecasters, human or machine, can bring. When Jeddi announced the paper, Servan-Schreiber summed it up for Page: the team presents evidence for the effect.

How the comparison worked

The team used ForecastBench, a public benchmark of forecasting questions, in its 21 July 2024 release. They kept 580 resolved questions: 526 “data” questions generated from public datasets, and 54 “market” questions taken from forecasting platforms such as Metaculus, Manifold and Polymarket.

  • Sixteen language models from OpenAI, Anthropic, Google, Mistral and Meta, each run with several prompts (zero-shot, step-by-step reasoning, reasoning with news, and three “superforecaster” prompts), for 76 model-and-prompt combinations in all. The models never saw the crowd forecasts.
  • Two human benchmarks: the aggregated forecasts of superforecasters, and those of a general-public sample recruited through Prolific.
  • Accuracy measured as one minus the Brier score. The Brier score is the squared gap between a probability and what happened, so 0 is perfect and lower is better. Superforecasters averaged 0.118 and the public 0.154.
  • Similarity measured as the correlation between each model’s forecasts and each human aggregate, across 304 model–crowd pairings.

What they found

More accurate models were markedly more correlated with the humans. One standard deviation of extra model accuracy went with a rise of about 0.08 in human–AI correlation, and accuracy alone explained roughly 38% of the variation in correlation. The pattern held for both question types. It was weaker on the market questions, and superforecasters were less correlated with the models than the general public was. Prompting mattered much less than accuracy.

The authors then checked that this is not a statistical artefact. If accurate forecasters simply made independent errors, their forecasts would still correlate somewhat, because they are all close to the truth. Simulations of that null model gave a slope of about 0.03, against 0.08 observed, and no simulated run reached the real value. The errors themselves also lined up. Better models made fewer and smaller mistakes, but the mistakes they did make looked more like human ones.

Two caveats sit alongside the headline. Superforecasters still beat the best language-model setups on these questions. And the study is observational, uses short-horizon questions, and treats the data-versus-market split as only a rough proxy for question type.

human forecasters may be moving from the factory floor to the control room

Jeddi, Segovia-Martin and Servan-Schreiber, Phil. Trans. R. Soc. B (2026)

The paper uses that image for a shift it expects: humans spending less effort producing forecasts and more effort supervising and combining them, as machines take over more of the routine work.

A theme issue on collective intelligence

The paper is one of the contributions to a theme issue on the evolution of collective intelligence, dated 17 April 2026 and edited by Cathal O’Madagain, Sarah Alami, Monique Borgerhoff Mulder, Edmond Seabright, José Segovia Martin, James Winters and Andrew Whiten. Servan-Schreiber co-authored a second paper in the same issue, “Locals know more than it seems”. It tested a method for revealing shared knowledge in farming communities in Morocco, Mali and Ghana, and found that when people could review their peers’ answers, the most popular final answer was more reliable than answers given alone. “From a most rustic to a most modern form of CI!”, he wrote of the pair on 27 April.

In May, Jeddi presented the work as a poster at the Machine+Behavior Conference, held on 18–19 May 2026 at the Max Planck Institute for Human Development in Berlin. Sharing it on 8 June, Servan-Schreiber wrote that AI had caught up with collective intelligence on prediction, and that what happens next in the forecasting business “is anyone’s guess”.

Evidence

Sources & further reading

Keep reading

Related articles

Insights“Arising Intelligence”: what Hypermind’s crowd got right and wrong about AI progressInsightsWhy Probabilities Beat PredictionsNews“Supercollective Intelligence”: Servan-Schreiber’s 2018 book on crowds and prediction markets comes out in English

Subscribe to The Forecasting Brief — forecasts & model notes, once a week