THE FUTURE IS PROBABILISTIC. EXPLORE THE POSSIBILITIES.Experimental platform

FIELD NOTES

A gap is a question, not an answer.

How to read disagreement between AI consensus and a reference market without confusing it with evidence of an edge.

Start with a shared question

A comparison is only useful when both probabilities describe the same event, outcome, and time horizon. A home win after normal time is different from qualifying after penalties. Check the resolution criteria before reading the numbers.

Keep the units honest

A sample model consensus of 63% against a sample market probability of 54% creates a gap of 9 percentage points. The gap says nothing by itself about the quality of the input data or the reliability of either estimate.

Look for the explanation, then its evidence

Models may disagree because of different sources, assumptions, data freshness, or errors. Correlated models may share the same blind spot. A credible comparison makes the source snapshots and timestamps visible.

The next test is history

Repeated, pre-recorded forecasts and independently resolved outcomes let us investigate whether apparent disagreements contain useful information. Forecast Arena’s current examples are synthetic, so no empirical advantage can be inferred.

Sources and scope

This note describes our product design and sample framework; it reports no empirical findings. Read the full methodology for definitions, assumptions, and primary metric references. All numerical examples are illustrative.