On 6 July 2023, Émile Servan-Schreiber reshared a Smarter Together preview of crowd forecasting. Its central claim was that combining different human judgements can reduce error while accumulating information. The preview mentioned a four-step algorithm without listing its steps. A companion case study, credited to Servan-Schreiber and Camille Larmanou, supplies them.
That distinction matters. Gathering a crowd is only the input. An aggregation rule determines which forecasts count, how much each contributes and how the combined number is adjusted. The preview’s examples ranged from elections to AI benchmarks and infectious disease; none came with a new dated probability in this post.
The four operations
- Weighting: the case study gives greater influence to forecasters with stronger accuracy records and more frequent updates.
- Culling: retain a fraction of the most recent individual forecasts, rather than combining every forecast regardless of age.
- Averaging: combine the retained estimates using their assigned weights.
- Extremizing: sharpen the combined probabilities to compensate for collective underconfidence, as the authors describe it.
Why different mistakes matter
The July preview emphasised independence. If people bring different information and make different errors, combining their estimates can cancel some mistakes. If they all repeat the same mistaken premise, gathering more answers will preserve that mistake. Diversity is useful when it contributes information or different ways of assessing it.
The operations address distinct problems. Weighting uses the record of past judgements; culling deals with stale estimates; averaging combines information. Extremizing addresses a further possibility: a combined forecast may be too cautious. That last correction has to be assessed against resolved questions, because making a probability more emphatic does not automatically make it more accurate.
Test the aggregation rule against outcomes
The method separates the forecasting poll from a market. In a poll, participants submit estimates and an algorithm combines them; in a market, trades determine prices. Both can collect information, but their outputs arise through different processes. The four operations describe one way to combine submitted estimates, rather than a universal recipe for every crowd.
The journal’s Johns Hopkins forecasting study covers the health application and its results. The broader design lesson is to distinguish diversity in the inputs from adjustments to the output. An algorithm may reward updating or sharpen probabilities, but it still needs information that participants have assessed separately. Testing against resolved questions shows whether those choices improve the combined forecast.

