Back to Blog

How independent errors become a crowd forecast: Hypermind’s four aggregation steps

2 min read
Navy cover with cream serif text “Hypermind’s four steps for aggregating crowd forecasts,” coral “Hypermind · 2023,” a coral rule and Émile Servan-Schreiber on LinkedIn · 2023-07-06.

On 6 July 2023, Émile Servan-Schreiber reshared a Smarter Together preview of crowd forecasting. Its central claim was that combining different human judgements can reduce error while accumulating information. The preview mentioned a four-step algorithm without listing its steps. A companion case study, credited to Servan-Schreiber and Camille Larmanou, supplies them.

That distinction matters. Gathering a crowd is only the input. An aggregation rule determines which forecasts count, how much each contributes and how the combined number is adjusted. The preview’s examples ranged from elections to AI benchmarks and infectious disease; none came with a new dated probability in this post.

The four operations

  • Weighting: the case study gives greater influence to forecasters with stronger accuracy records and more frequent updates.
  • Culling: retain a fraction of the most recent individual forecasts, rather than combining every forecast regardless of age.
  • Averaging: combine the retained estimates using their assigned weights.
  • Extremizing: sharpen the combined probabilities to compensate for collective underconfidence, as the authors describe it.

Why different mistakes matter

The July preview emphasised independence. If people bring different information and make different errors, combining their estimates can cancel some mistakes. If they all repeat the same mistaken premise, gathering more answers will preserve that mistake. Diversity is useful when it contributes information or different ways of assessing it.

The operations address distinct problems. Weighting uses the record of past judgements; culling deals with stale estimates; averaging combines information. Extremizing addresses a further possibility: a combined forecast may be too cautious. That last correction has to be assessed against resolved questions, because making a probability more emphatic does not automatically make it more accurate.

Test the aggregation rule against outcomes

The method separates the forecasting poll from a market. In a poll, participants submit estimates and an algorithm combines them; in a market, trades determine prices. Both can collect information, but their outputs arise through different processes. The four operations describe one way to combine submitted estimates, rather than a universal recipe for every crowd.

The journal’s Johns Hopkins forecasting study covers the health application and its results. The broader design lesson is to distinguish diversity in the inputs from adjustments to the output. An algorithm may reward updating or sharpen probabilities, but it still needs information that participants have assessed separately. Testing against resolved questions shows whether those choices improve the combined forecast.

Evidence

Sources & further reading

Keep reading

Related articles

InsightsCrowdfunders and experts largely agreed on theatre projects — with different funding thresholdsInsightsAn hourly COVID economic dashboard, and Émile’s idea of hybrid collective intelligenceInsightsHugo, Galton and the clues a crowd can put together

Subscribe to The Forecasting Brief — forecasts & model notes, once a week