Back to BlogNovember 21, 2021
Insights

What 15 months of crowd forecasting taught Johns Hopkins about disease outbreaks

The Forecasting MachineEditorial Team5 min read
Editorial illustration of a petri dish holding coral and ivory beads beside research cards and a timer.

On 21 November 2021, Émile Servan-Schreiber shared a paper written by researchers at the Johns Hopkins Center for Health Security together with his team at Hypermind. “Using prediction polling to harness collective intelligence for disease forecasting”, by Tara Kirk Sell and seven co-authors, including Servan-Schreiber and Maurice Balick of Hypermind, had appeared in BMC Public Health the day before. It reports on a crowd-forecasting project on infectious disease outbreaks that ran for 15 months, from January 2019 to March 2020. The last weeks of that period covered the start of COVID-19.

Screenshot of the Johns Hopkins Center for Health Security “Disease Prediction” forecasting platform, powered by Hypermind, showing the question “How many WHO member states will report more than 1000 confirmed cases of COVID-19 on or before April 2, 2020?” with answer ranges, a probability chart and red callout labels explaining parts of the interface.
The Johns Hopkins Disease Prediction platform, powered by Hypermind.

How the experiment worked

The project used prediction polling. Participants do not trade contracts as they would in a prediction market. Instead, each one gives a probability for every possible answer and updates it as news arrives. An algorithm then combines all the forecasts into one crowd forecast. The authors chose this format partly because many people, including many public-health experts, find trading unfamiliar.

Over the 15 months, 562 people forecast on at least one of 61 questions about 19 diseases, including Ebola, measles, influenza and COVID-19. The questions had 217 possible answers between them. A typical question asked for a range, such as how many WHO member states would pass a given number of COVID-19 cases by a set date, and stayed open for about a month. In all, the platform collected 10,750 forecasts.

Just over half of the participants (54%) were public-health professionals, and 15% worked in other health fields. The other 31% reported no health background. Among them were 132 forecasters already vetted on Hypermind for their general forecasting skill; only five of these were public-health professionals. Prizes went to the top of the final leaderboard, and the platform was Hypermind’s Prescience software.

The crowd forecast was not a simple average. The algorithm gave more weight to people who updated often and had been accurate before. It kept only the most recent 30% of forecasts, averaged them, and then made the result more extreme to correct for a crowd’s habit of hedging. Participants saw only a simpler version of the crowd forecast, so they could not just copy the best answer.

What the study found

  • Calibration. The crowd was well calibrated. Its reliability scores, a standard calibration measure borrowed from weather forecasting where 0 is perfect, were .0043 and .0015 at the two levels of precision tested.
  • Accuracy. Across all 61 questions, the crowd’s average Brier score was 0.238, against 0.460 for chance. The paper describes this as 48% more accurate than chance. The crowd did worse than chance on only 6 questions.
  • Crowd against individuals. On the 54 questions where the comparison was fair, the mean and median individual forecaster scored no better than chance. Only 6 people beat the plain average of everyone’s forecasts, only 3 beat the version shown on screen, and none beat the full crowd forecast (0.245).
  • Timeliness. On the median question, the crowd beat chance for 97% of the time the question was open, and the correct answer was its favourite for 77% of that time. The crowd settled for good on the right answer, on the median question, 42% of the way through.

crowd forecasts aggregated using best-practice adaptive algorithms are well-calibrated, accurate, timely, and outperform all individual forecasters.

Sell et al., BMC Public Health, 2021

The authors also list the limits. Writing good questions was hard, and in several cases the answer ranges should have gone higher. Some questions had to be voided. Crowd forecasting also depends on surveillance data, both to inform forecasters and to settle the questions. The paper presents it as an addition to surveillance and modelling, not a replacement. Two of the authors are Hypermind partners, and the company supplied both the platform and 23.5% of the participants. The paper declares these interests; the research was funded by Open Philanthropy.

Experts or skilled forecasters?

Servan-Schreiber had previewed a second analysis of the same data in March 2021, in a short talk at the DIMACS Workshop on Forecasting: From Forecasts to Decisions. In the recording, he called it a natural experiment with design flaws, not a controlled trial. His preliminary findings, as presented there:

  • Individually, most public-health professionals forecast no better than chance, but as a group they were very good.
  • The group of skilled general-purpose forecasters did as well as the group of experts, and adding them to the experts improved the result.
  • The skilled forecasters made more than twice as many forecasts as other participants, and kept forecasting when the experts became busy with COVID-19.
  • Including people with neither kind of expertise did no harm, and some of them proved to be good forecasters.

His advice from the talk was to trust experts as a group rather than any one of them, and to use skilled forecasters to add to an expert panel or stand in for it. The BMC paper cites this comparison of skill and domain expertise as a manuscript in preparation. It was not yet peer reviewed, so its findings are Servan-Schreiber’s and not the published study’s.

Evidence

Sources & further reading

Keep reading

Related articles

NewsCrowd forecasts of COVID-19 reach the US Congress, March 2020InsightsWhy Probabilities Beat PredictionsInsightsEleven and a half years of Hypermind forecasts: does a 70% price come true 70% of the time?

Subscribe to The Forecasting Brief — forecasts & model notes, once a week