Back to BlogMarch 19, 2025
Insights

How humans and an early forecasting AI saw 2030: the OECD's Hypermind pilot

The Forecasting MachineEditorial Team4 min read
Editorial illustration of a brass balance holding a seedling and a glass-encased microchip above forecast folios marked 2030.

In June 2024, Hypermind invited forecasters to estimate the benefits and risks that artificial intelligence might create by 2030. The exercise was designed with the OECD's network of AI experts and supported by Longview Philanthropy. Nine months later, Émile Servan-Schreiber shared the resulting OECD background note: an 82-page study of global AI governance with a seven-page forecasting annex.

The exercise put aggregate human forecasts beside an early LLM-based forecasting system across 25 questions. Its purpose was not to declare a winner. The human forecasts were the primary evidence; the AI system supplied a baseline and additional rationales. That distinction is central to reading the results.

From four scenarios to measurable questions

The OECD Global Strategy Group organized its discussion around four futures for AI governance: Inclusive Governance, Shared Prosperity; Cautious and Controlled; Fast and Furious; and Breakthroughs Backfire. The scenarios varied along two dimensions—safety and responsibility, and inclusion and access—and examined innovation, labour, social cohesion, markets and geopolitics.

Scenarios help decision-makers imagine coherent worlds. Forecasts do a different job: they turn selected parts of those worlds into questions that can be assigned probabilities or numerical estimates. The pilot asked about scientific breakthroughs, climate patents, business adoption, health, work, trust, democratic threats, cyberattacks, catastrophic incidents, regulation and AI research. Each question received an average of 161 human forecasts.

What the human aggregate expected

The human panel expected AI's contribution to science and business to grow substantially, but not uniformly. It forecast that 40% of the breakthroughs on Science's annual lists for 2030–2032 would involve a substantial AI contribution, and that 32% of climate patents would contain AI components. Among organizations already using AI, it expected 34% to cut costs by more than 10% and 44% to increase revenue by more than 10%.

The social picture was less optimistic. The panel expected 36% of people to use AI in support of healthcare by 2030, but only 30% to trust political organizations using AI. It forecast that the share of workers saying AI improved their enjoyment at work would fall to 48% in finance and 31% in manufacturing. These were dated estimates for a 2030 horizon, not observed outcomes.

On high-impact risks, the crowd assigned 7% to an AI incident causing more than one million deaths by 2030 and 37% to an AI-driven cyberattack severe enough to activate the European Union's Integrated Political Crisis Response arrangements. A small probability attached to a catastrophic outcome should not be read as reassurance: expected impact depends on both probability and consequence.

Where the machine disagreed

The early forecasting tool used language models, specialized prompts, recent news, base rates and other inputs to produce probabilities and rationales. The OECD cautioned that its output was likely to reflect an online baseline consensus rather than an authoritative forecast. It was included to expose differences and demonstrate an emerging capability.

  • On one third of the questions, the AI and human estimates differed by less than 20%.
  • On another third, they differed by 20% to 50%.
  • On the remaining third, they differed by more than 50%.

Those bands measure disagreement, not accuracy. Most questions will not resolve until 2030, and several ask for quantities whose final measurement will itself require care. The report found that the AI was generally more optimistic about benefits such as cost savings, revenue, climate patents, healthcare and trust. The direction on risk was mixed: for example, it put 0.9% on a million-fatality AI incident versus the crowd's 7%, but expected a larger role for AI in threats to democracy.

The limitations are part of the result

The OECD explicitly described the exercise as experimental and indicative. It identified three open problems: whether validated crowd methods transfer to horizons beyond two years, whether they work equally well for emerging-technology impacts, and how their output should enter political decisions. Financial incentives and aggregation can improve care, but they do not solve uncertain definitions, correlated evidence or a distant resolution date.

That caution makes the pilot more useful, not less. It established a time-stamped set of human and machine expectations that can later be scored. It also showed a practical relationship between strategic foresight and forecasting: scenarios widen the field of possibilities, while probability estimates make selected assumptions explicit enough to revisit.

Evidence

Sources & further reading

Keep reading

Related articles

NewsHow the OECD’s foresight unit fits crowd and AI forecasting into policy planningInsights“Arising Intelligence”: what Hypermind’s crowd got right and wrong about AI progressInsightsCrowds versus language models: the more accurate the AI forecaster, the more it errs like humans

Subscribe to The Forecasting Brief — forecasts & model notes, once a week