What was scored

Working Paper No. 5 scored 620 forecasts from Jump Trading’s Probability Cup on the 2026 World Cup by Brier score (Brier, 1950). A fixture is priced from one view of one match, and the margin does not survive that.

Table A. The scored record, from Working Paper No. 5.
MeasureValue
Settled forecasts scored620
Mean Brier score0.2309
Share of questions resolving yes44.2%
Constant forecast at that rate0.2466
Edge a question0.0157
Constant 50% benchmark, withdrawn0.25
Agreement with the published calibration tablewithin 0.001
t-statistic if draws were independent2.2
Markets a fixture priced from one viewfifteen to twenty
Correlation assumed within a fixture0.10
t-statistic at that correlation1.4
Resolution, the part earned by telling outcomes apart0.020
Reliability, the penalty for drift from observed frequencies0.006
Contrarian win rate48%

Where it scored

The three best categories turn on how a match is played, not who wins it. They were picked after the fact, and at seven calls a win rate carries a standard error near seventeen percentage points.

Table B. The platform’s three highest categories, on 38 of the 620 forecasts.
CategoryForecastsDisplayed figure
Halftime tied11+17.4
Penalty and red card12+10.5
Team corners15+6.9
All three38of 620 forecasts

Where it lost

The platform’s own buckets carry an ordering and a sign. They show where ground was given back, not that departure carried negative expectancy.

Table 1. Relative Brier figures by category type across the 432 of 620 forecasts the platform buckets.
Category typeForecastsFigureVerdict
Match Outcome179−6.1largest book; ground lost
Goals & Scoring145−15.9largest give-back; units unresolved
Set Pieces & Possessions63−2.8small negative
Cards & Discipline45+0.1flat

A documented benchmark

Rescored against the best constant forecast for each category, its own base rate b, which scores b(1 − b), the book is positive in aggregate.

Table 2. Mean Brier score by category against the best constant forecast for that category, reconstructed from 555 settled forecasts.
CategorynMean BrierBase rateBenchmarkEdge
Halftime state280.22030.5000.2500+0.0297
Match outcome350.20730.6290.2335+0.0262
Cards430.20730.3490.2271+0.0198
Corners420.22840.4050.2409+0.0125
Penalty and red card340.16760.2350.1799+0.0123
Goals and scoring1650.24330.4550.2479+0.0047
Shots1340.23800.4030.2406+0.0026
Offsides280.25180.4290.2449−0.0069
Timing180.23850.3330.2222−0.0163
All 5555550.22990.4220.2439+0.0139

The schemes do not partition the same forecasts. This book lost ground where the crowd was sharpest, which both records allow and neither establishes.

Table C. The two partitions compared.
MeasurePlatformReconstruction
Forecasts sorted432 of 620555 of 620
Scaleundocumented, reconciles to no overall figureBrier against base rate b
Categoriesfour typesnine, 527 sorted and 28 unclassified
Match outcome, forecasts17935
Match outcome, standingone of the two worstsecond best of nine
Goals and scoring, standinglargest give-backpositive
Categories positiven/aseven of nine, and in aggregate
Share of the book in those two labels52%n/a

The free tilt

A flat forecast captures much of the edge. The account did not take that tilt deliberately, hedging towards 50%, which No. 5 called underconfidence.

Table D. The threshold tilt, on 555 reconstructed forecasts.
MeasureValue
Share resolving yes42.2%
Flat 42% forecast, mean Brier0.2439
Flat 50% forecast, mean Brier0.2500
Gain a question from the tilt alone0.0061
This book’s edge over the same benchmark0.0139
Tilt as a share of that edge44%
Band No. 5 recorded as underconfident45–55%

What this does not claim

No. 5 did not claim the account out-thought the field. The finish invited that reading, retired here with the constant 50% benchmark. The contrarian win rate sat below half, on an undisclosed denominator and at no significant distance from half. Nothing makes departure the source of the edge, nor calibration.

One tournament, one crowd, one settlement. The two disclosed calls fit no constant against a consensus benchmark but fit exactly against one built from forecaster means, so no claim against the consensus is made here. Counts by category are small, labels came after the fact, and the account chose its markets, so the mix is itself an output. This shows where the account scored, not where an edge exists.

The paper is kept because it withdraws a claim this site published, then one of its own. A collection that only adds is not a ledger.

References

  1. Brier, G. W. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, 78(1), 1–3.