A probabilistic forecast is not graded simply by whether its favorite won. A 60% outcome can fail; a 5% outcome can happen. The Brier score measures the squared distance between a probability and the resolved outcome. Calibration asks a different question: across many comparable forecasts, do outcomes occur about as often as the stated probabilities imply? Use both to assess a forecasting process, rather than judging it from one memorable success.
The binary Brier score formula
For a binary event, use a probability p between 0 and 1 and an outcome y equal to 1 if the event happened and 0 if it did not. The single-outcome Brier score is (p − y)². Average this value over the resolved questions being evaluated. Lower is better: 0 is a perfectly correct certain prediction, and 1 is a certainly wrong one. A 50% forecast scores 0.25 regardless of which binary outcome occurs.
- 70% forecast, event occurs: (0.70 − 1)² = 0.09.
- 70% forecast, event does not occur: (0.70 − 0)² = 0.49.
- 20% forecast, event occurs: (0.20 − 1)² = 0.64.
- 20% forecast, event does not occur: (0.20 − 0)² = 0.04.
For a small worked example, suppose four forecasts are 70%, 20%, 80% and 40%, and the events resolve yes, no, yes and no. Their scores are 0.09, 0.04, 0.04 and 0.16. The mean is 0.0825. These are invented teaching examples, not product results. Four questions are far too few to establish a robust performance claim.
Check which Brier convention a report uses
Some reports sum squared error across every mutually exclusive outcome. For a binary question that counts both yes and no, the score is twice the single-outcome value: a 50/50 forecast scores 0.50 rather than 0.25. A multi-outcome question uses a probability vector whose entries sum to one. Do not compare published numbers until the formula and normalization match. A lower-looking number may simply use a different scale.
The Brier score is also not the only proper scoring rule. Metaculus explains how proper rules reward sincere probabilities in expectation and uses log-based scores. A log score and a Brier score cannot be read as though they were the same metric. State the evaluation rule before attaching an accuracy label to a leaderboard.
What does calibration mean?
Imagine a sufficiently large set of comparable forecasts, each assigned about 70%. If roughly 70% of those events occur, that group is calibrated. This is a statement about frequencies across a set, not a demand that any one 70% forecast resolve yes. Calibration plots usually group probabilities into bins and compare their average estimates with observed event rates. Sparse bins and dependent questions require caution.
Calibration alone does not make a system informative. If events in a dataset occur half the time, always issuing 50% can be calibrated while distinguishing none of the easy questions from the difficult ones. A useful forecaster tries to identify when the odds genuinely differ, while remaining honest about uncertainty. Our Hypermind calibration case study concerns an observed track record, rather than the worked examples here.
What counts as a good Brier score?
There is no universal threshold detached from the questions. Compare with a baseline evaluated on the same outcomes and at the same lead time. A 50% baseline is a transparent starting point for binary questions, but a base-rate forecast may be more appropriate when almost all events resolve the same way. A score of 0.10 could be impressive on hard questions and unremarkable on a dataset full of near-certainties.
If reporting skill relative to a baseline, make the definition explicit: one common form is 1 minus the system's mean Brier score divided by the baseline's mean Brier score. With a nonzero baseline, positive values indicate improvement over it; zero indicates parity; negative values indicate worse performance. This expresses a comparison, not a guarantee about the next event.
Compare AI and prediction markets on equal terms
- Use identical questions and resolution rules, with a recorded evidence cutoff.
- Choose the observation time before seeing outcomes: for example, thirty days before the deadline.
- State whether scores use one snapshot or a time-weighted sequence of updates.
- Include all eligible resolved questions, not only favorable examples.
- Report sample size, category mix, baseline and uncertainty; flag correlated outcomes.
- Separate forecast scoring from trading returns, fees and execution.
Brier.fyi's methodology discusses matching markets and forecast timing. Those choices matter when comparing Polymarket, Kalshi and forecaster panels. Our play-money versus real-money benchmark report documents why unmatched question sets limit the conclusions. For model comparisons, the AI-versus-market workflow adds an independence check: an AI that has read a market price is not a wholly separate signal.
For decision-makers, the practical aim is not the smallest unexplained number. It is a process whose probabilities can be understood, tested and improved. A clear scoring record makes it possible to learn from both correct calls and surprises without pretending that uncertainty has disappeared.
