What was scored
Working Paper No. 5 scored 620 forecasts from Jump Trading’s Probability Cup on the 2026 World Cup by Brier score (Brier, 1950). A fixture is priced from one view of one match, and the margin does not survive that.
| Measure | Value |
|---|---|
| Settled forecasts scored | 620 |
| Mean Brier score | 0.2309 |
| Share of questions resolving yes | 44.2% |
| Constant forecast at that rate | 0.2466 |
| Edge a question | 0.0157 |
| Constant 50% benchmark, withdrawn | 0.25 |
| Agreement with the published calibration table | within 0.001 |
| t-statistic if draws were independent | 2.2 |
| Markets a fixture priced from one view | fifteen to twenty |
| Correlation assumed within a fixture | 0.10 |
| t-statistic at that correlation | 1.4 |
| Resolution, the part earned by telling outcomes apart | 0.020 |
| Reliability, the penalty for drift from observed frequencies | 0.006 |
| Contrarian win rate | 48% |
Where it scored
The three best categories turn on how a match is played, not who wins it. They were picked after the fact, and at seven calls a win rate carries a standard error near seventeen percentage points.
| Category | Forecasts | Displayed figure |
|---|---|---|
| Halftime tied | 11 | +17.4 |
| Penalty and red card | 12 | +10.5 |
| Team corners | 15 | +6.9 |
| All three | 38 | of 620 forecasts |
Where it lost
The platform’s own buckets carry an ordering and a sign. They show where ground was given back, not that departure carried negative expectancy.
| Category type | Forecasts | Figure | Verdict |
|---|---|---|---|
| Match Outcome | 179 | −6.1 | largest book; ground lost |
| Goals & Scoring | 145 | −15.9 | largest give-back; units unresolved |
| Set Pieces & Possessions | 63 | −2.8 | small negative |
| Cards & Discipline | 45 | +0.1 | flat |
A documented benchmark
Rescored against the best constant forecast for each category, its own base rate b, which scores b(1 − b), the book is positive in aggregate.
| Category | n | Mean Brier | Base rate | Benchmark | Edge |
|---|---|---|---|---|---|
| Halftime state | 28 | 0.2203 | 0.500 | 0.2500 | +0.0297 |
| Match outcome | 35 | 0.2073 | 0.629 | 0.2335 | +0.0262 |
| Cards | 43 | 0.2073 | 0.349 | 0.2271 | +0.0198 |
| Corners | 42 | 0.2284 | 0.405 | 0.2409 | +0.0125 |
| Penalty and red card | 34 | 0.1676 | 0.235 | 0.1799 | +0.0123 |
| Goals and scoring | 165 | 0.2433 | 0.455 | 0.2479 | +0.0047 |
| Shots | 134 | 0.2380 | 0.403 | 0.2406 | +0.0026 |
| Offsides | 28 | 0.2518 | 0.429 | 0.2449 | −0.0069 |
| Timing | 18 | 0.2385 | 0.333 | 0.2222 | −0.0163 |
| All 555 | 555 | 0.2299 | 0.422 | 0.2439 | +0.0139 |
The schemes do not partition the same forecasts. This book lost ground where the crowd was sharpest, which both records allow and neither establishes.
| Measure | Platform | Reconstruction |
|---|---|---|
| Forecasts sorted | 432 of 620 | 555 of 620 |
| Scale | undocumented, reconciles to no overall figure | Brier against base rate b |
| Categories | four types | nine, 527 sorted and 28 unclassified |
| Match outcome, forecasts | 179 | 35 |
| Match outcome, standing | one of the two worst | second best of nine |
| Goals and scoring, standing | largest give-back | positive |
| Categories positive | n/a | seven of nine, and in aggregate |
| Share of the book in those two labels | 52% | n/a |
The free tilt
A flat forecast captures much of the edge. The account did not take that tilt deliberately, hedging towards 50%, which No. 5 called underconfidence.
| Measure | Value |
|---|---|
| Share resolving yes | 42.2% |
| Flat 42% forecast, mean Brier | 0.2439 |
| Flat 50% forecast, mean Brier | 0.2500 |
| Gain a question from the tilt alone | 0.0061 |
| This book’s edge over the same benchmark | 0.0139 |
| Tilt as a share of that edge | 44% |
| Band No. 5 recorded as underconfident | 45–55% |
What this does not claim
No. 5 did not claim the account out-thought the field. The finish invited that reading, retired here with the constant 50% benchmark. The contrarian win rate sat below half, on an undisclosed denominator and at no significant distance from half. Nothing makes departure the source of the edge, nor calibration.
One tournament, one crowd, one settlement. The two disclosed calls fit no constant against a consensus benchmark but fit exactly against one built from forecaster means, so no claim against the consensus is made here. Counts by category are small, labels came after the fact, and the account chose its markets, so the mix is itself an output. This shows where the account scored, not where an edge exists.
The paper is kept because it withdraws a claim this site published, then one of its own. A collection that only adds is not a ledger.
References
- Brier, G. W. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, 78(1), 1–3.