An AI forecast and a prediction-market price may both say 60%, but they arrive there differently. One is produced by an evidence and model pipeline; the other emerges from trading. A useful comparison asks whether they concern the same event, what each knew at the time, and how their probabilities perform after resolution. It does not start by declaring humans or machines the permanent winner.
What AI forecasting contributes
For event forecasting, an AI system can retrieve information, organize evidence, generate probability estimates and aggregate model judgments. That is distinct from a language model writing a plausible narrative about the future. A forecast needs a specified outcome, deadline, evidence cutoff and probability that can eventually be assessed. A polished explanation cannot substitute for those requirements.
The Forecasting Machine is positioned as an AI forecasting tool, not a trading venue. Agenda Pública's description of its use emphasizes aggregation of open sources to help readers prioritize scenarios. That is a decision-support role: make uncertainty explicit, examine the evidence and update the assessment. It is not evidence that the product beats every market or that every output is calibrated.
What a prediction market contributes
Markets give participants a shared contract and a mechanism for expressing disagreement. A trader may use public news, specialist knowledge, a model or another market. The resulting price can reflect information that an analyst has not considered. But prices also depend on the order book and the willingness of participants to trade. Our Polymarket odds explainer distinguishes a displayed probability from an executable bid or ask.
A missing or thinly traded market creates a coverage problem, not proof of an impossible question. An organization may also need an assessment of a narrowly defined decision that has no public contract. Conversely, a popular public event may have a rich market history that is valuable evidence. Choose tools around the question, rather than assuming one method is appropriate for every situation.
A five-step comparison workflow
- Align the outcome. Copy the actual resolution criteria, deadline and measurement source; similar headlines are insufficient.
- Align the information cutoff. Record the AI forecast and market observation before subsequent news or the eventual outcome is known.
- Separate independent estimates from market-informed estimates. If the AI has seen the market price, its agreement is not independent confirmation.
- Record the explanation for disagreement. Identify evidence, assumptions or contract details that could account for the difference.
- Score all resolved questions using the same convention and lead time, including inconvenient misses. Compare with a declared baseline.
An illustrative disagreement makes this concrete. Suppose an AI assigns 35% to a policy taking effect before year-end while a market displays 55%. Before averaging them, check whether the market concerns announcement rather than implementation. Then ask whether one estimate predates a legislative vote. Only after those checks is the remaining disagreement a useful research question. These percentages are hypothetical, not a current recommendation.
Do AI models beat human forecasters?
A defensible answer names a benchmark and its date. The ForecastBench paper introduces prospective questions whose answers were unknown at submission and compares model forecasts with human estimates. In that study, expert forecasters outperformed the strongest tested language model. This is evidence about a particular evaluation, not a permanent ranking of all current models, market participants or products.
The separate accuracy–correlation study adds an important concern: accurate human and model forecasts can share errors. Combining several estimates does not automatically create several independent sources of evidence. A panel of models repeating the same underlying information may sound diverse while adding little new signal.
Measure forecasting, not just confident narratives
Use Brier scores and calibration checks across comparable resolved questions. Keep the sample size, selection rules and forecast horizon visible. A model scored one hour before resolution cannot be fairly compared with a market scored three months earlier. A collection of easy outcomes can make an ordinary system look exceptional unless the baseline faces the same questions.
Trading profit is another metric, not a synonym for forecast quality. Entry prices, spreads, fees, position sizes and execution determine financial results. A good probability estimate can be unprofitable at the available price, while a fortunate trade does not establish a reliable forecasting process. An analyst should report these quantities separately.
The goal is a better decision record: what was believed, why, with what uncertainty, and what changed. Markets and AI can both contribute to that record. Neither eliminates the need for clear resolution criteria, human judgment or an honest track record.
