01 / Frontier models
Compared to frontier LLMs
- 36% lower Brier scorevs average individual LLM
- 63% lower calibration error(Expected Calibration Error)
Internal benchmark; not independently audited. TFM was compared with Claude Opus 4.8, Claude Sonnet 4.6, Gemini 3.1 Pro, and GPT 5.4 across 38 shared questions in politics, geopolitics, economics, and sports. Forecasts were made between September 2025 and June 2026; the sample contains 1,265 predictions and 3,147 event probabilities. Brier score and Expected Calibration Error are lower-is-better measures. Relative reductions use the average individual LLM as the baseline.









