Back to Blog

The US–China compute race in Hypermind’s September 2021 AI contest

3 min read
“Hypermind’s 2021 US–China AI Compute Contest” sits beside a poster of facing head silhouettes over Chinese and US flags, reading “Hypermind Forecasts” and “China vs USA: The Race for AI Computing Power”.

On 22 and 23 September 2021, Émile Servan-Schreiber reshared French and English invitations to a Hypermind AI forecasting contest. Both said six questions had been developed with experts from UC Berkeley, invited entries before 30 September, and advertised $12,000 to be shared among the best forecasters. The accompanying artwork framed one part of the exercise as a China-versus-USA race for AI computing power.

The baseline behind the invitation

The posts supplied a June 2021 starting point. Hypermind described Huawei’s PanGu-α as China’s largest documented training experiment, at 583 petaflop/s-days, and OpenAI’s GPT-3 as the US counterpart, at 3,640 petaflop/s-days. It characterised the resulting ratio as a little over six to one in favour of the United States. These are the invitation’s published figures and its classification of the largest documented experiments at that date.

A petaflop/s-day measures training work: a processing rate sustained for a day. The comparison therefore concerns model-training compute, rather than parameter counts or a country’s total hardware capacity. Both values are reported here in the common unit used by Hypermind’s French invitation.

The GPT-3 and PanGu-α research papers establish the underlying models and their training research. They do not, by themselves, prove that the invitation’s choices exhausted all experiments conducted in each country. In particular, a comparison of documented experiments depends on what laboratories disclose.

What the six questions tried to measure

Jacob Steinhardt’s own August 2021 account confirms the broader collaboration. His group commissioned six questions: two concerned training compute, and four concerned future performance on machine-learning benchmarks. Hypermind ran the competition and produced the aggregate forecasts, with funding from Open Philanthropy.

One compute question compared the largest Chinese experiment with the largest US experiment; the other concerned the largest experiment outside specified incumbent organisations and China. The capability questions covered mathematics, broad academic knowledge, adversarially robust image classification and video recognition. Forecasters supplied probability distributions for outcomes in 2022, 2023, 2024 and 2025, rather than simply choosing which nation would win.

Steinhardt described $30,000 in funding across the wider exercise, at $5,000 per question. The September invitation advertised $12,000. The two sources describe different prize amounts whose allocation relationship is unspecified. The advertised entry deadline likewise does not mean every forecast resolved on 30 September.

From a national rivalry to a checkable forecast

Hypermind’s promotional text connected AI leadership to national security. That was its framing of the stakes. The more useful forecasting move was to specify a measurable quantity and a dated baseline. Steinhardt also warned that laboratories might stop disclosing compute, making the geopolitical measures harder to interpret.

The journal’s Arising Intelligence retrospective examines later benchmark results. This September recruitment round records a different step: turning a broad rivalry into questions forecasters could answer, while exposing the limits of the evidence used to define the race.

Evidence

Sources & further reading

Keep reading

Related articles

Insights“Arising Intelligence”: what Hypermind’s crowd got right and wrong about AI progressInsightsHow humans and an early forecasting AI saw 2030: the OECD's Hypermind pilotNewsHypermind’s Africa of the future contest: Côte d’Ivoire first, a wider agenda next

Subscribe to The Forecasting Brief — forecasts & model notes, once a week