On 11 July 2022, Émile Servan-Schreiber published the first scorecard for “Arising Intelligence”, a Hypermind contest that asked forecasters to predict the state of the art in machine learning. A year after the forecasts were made, the results showed AI moving much faster than the crowd expected on some skills, and not at all on another.
How the contests worked
Hypermind’s first AI contest ran from 16 February to 9 April 2021. It posed 26 questions on hardware and supercomputing, benchmark performance, research trends, and economic and financial impact, with $7,000 in prizes and support from Open Philanthropy, as ActuIA reported at the time.
In the summer of 2021 came a larger contest, built with Jacob Steinhardt of UC Berkeley. Steinhardt wrote that his group commissioned six questions, each with a $5,000 prize pool, $30,000 in total, funded by Open Philanthropy. Hypermind ran the competition and produced the crowd forecasts. Four questions tracked AI capability benchmarks and two tracked the computing power used to train the largest models, including whether the United States or China would lead. Forecasters predicted the state of the art on 30 June of each year from 2022 to 2025.
- MATH: competition mathematics problems.
- MMLU: multiple-choice questions on a wide range of academic subjects, a test of language understanding.
- Something Something v2: recognising actions and objects in short videos.
- Adversarial CIFAR-10: classifying images that have been deliberately altered to fool the model, a test of robustness.
Forecasters gave full probability distributions for each benchmark score, not single numbers, so the result could be judged by how much probability the crowd had put near the actual outcome.
Year one: faster than expected
Hypermind’s chart compared the crowd’s June 2022 forecasts with the state of the art a year on.

- Mathematical problem solving (MATH) went from 6.9% to 50.3%, against a forecast of 12.7%.
- Language understanding (MMLU) went from 48.9% to 67.5%, against 57.1%.
- Action and object identification went from 69.0% to 75.3%, against 73.0%.
- Robust image classification stayed at 66.6%, against a forecast of 70.4%.
Servan-Schreiber wrote that progress had outpaced expectations almost everywhere, and that the jump on MATH had needed more and better data and far more compute rather than a new algorithmic paradigm. He called the lack of progress on robustness worrying: models were getting better in ideal conditions while staying as easy to fool as a year before. He also noted that the United States had widened its lead over China in training compute.
Steinhardt’s own one-year review was blunt. He judged the forecasters “not very good”: two of the four outcomes fell outside the crowd’s 90% credible intervals. The 50.3% MATH result was announced on the very day the question resolved. He also noted that the forecasters had still done better than he had himself, and probably better than the median machine-learning researcher, and he flagged limits of the exercise: roughly 60 to 70 participants, small average payouts, and an interface that constrained how wide a distribution could be.
Year two: a mixed scorecard
Hypermind closed its forecasts for June 2023 on 14 August 2022. When they were scored, MMLU had reached 86.4%, far above the crowd’s 73.2%. On MATH, the crowd’s 65.0% came close to the actual 69.6%. Video recognition reached 77.3% against a forecast of 78.7%, and robust image classification 70.7% against 75.1%, so the crowd overestimated progress on both. In July 2023, Servan-Schreiber called robustness progress “disappointingly slow”, bad news, he wrote, in an age of deep fakes.
Steinhardt compared several forecasting groups on the 2023 MATH and MMLU questions. His ranking put Metaculus and himself first, then the AI experts and superforecasters of the Existential Risk Persuasion Tournament, with Hypermind’s forecasters last. Hypermind’s big miss was MMLU, which GPT-4 pushed far outside the crowd’s predicted range. On MATH, he judged Hypermind’s forecast better than his own and Metaculus’s, because its range was narrower and the actual result fell inside it. He also stressed calibration: one answer in two landing far outside the predicted interval “is not good”, all the more because the same had happened the year before.
What the exercise shows
The contests are a reminder that forecasting AI is hard, even for informed crowds. The errors were not random. Forecasters underestimated progress driven by scale and overestimated progress on robustness. A crowd that is wrong in the same direction as the experts it learns from gains little from its numbers, a theme that later research co-authored by Servan-Schreiber examines in comparing human crowds with language models.
In July 2023, Hypermind opened a new round for the June 2024 and June 2025 horizons with a $21,000 prize.
