1. The competition

Jump Trading’s Probability Cup is a public forecasting competition on the 2026 World Cup. Entrants state probabilities on match questions, whether a match produces 22 or more shots, whether Spain puts five shots on target, and each is graded by the Brier score, the squared error between stated probability and realized outcome (Brier, 1950). The rule pays nothing for a plausible story. I entered an automated agent.

Scorecard. As scored by the organizer.
MeasureValue
Settled forecasts620
Mean Brier score0.2309
Score from always answering 50%0.25
Edge a question, as first reportedroughly two cents
Edge a question, revised in No. 170.0157
Base-rate baseline, revised in No. 170.2466 at a realized base rate of 44.2%
Share of forecasts in the 40–60% band40%
Calls of high convictionnone
Final standingtop 6%

The edge a question was small; the finish was not. A small honest edge taken 620 times separates a forecaster from a field that is largely either guessing or confident and wrong. Grinold’s fundamental law states the same mechanism for portfolios, information ratio equals information coefficient times the square root of breadth (Grinold, 1989); this leaderboard is additive in forecasts submitted, not square root.

2. Calibration

Calibration asks whether events called at 40% happen 40% of the time. Table 1 reports every band.

Table 1. Calibration by band of stated probability: realized frequency, forecast count, and verdict.
Stated probabilityRealizedForecastsVerdict
0–20%24.1%54slightly overconfident on longshots
20–35%32.1%187calibrated
35–45%40.5%158calibrated
45–55%59.6%104underconfident; edge left on the table
55–65%60.9%69calibrated
65–80%71.8%39calibrated
80–100%55.6%9overconfident in the tail (small n)

The two failures are opposite errors. In the middle band the agent knew more than it said, reporting near 51% while outcomes realized near 60%; hedging toward 50% is a tax rather than prudence, and that band held most of the book. In the tail the reverse held: on the nine occasions the agent was nearly certain it was right barely half the time, the error about tail risk documented in this site’s LTCM case study.

3. What transfers

4. What this does not claim

Three claims in the first version of this note are withdrawn or restated in Working Paper No. 17. The calibration table survives that review.

Standing record. First claim against the current reading.
ClaimWhat the record says
Coin-flip baseline of 0.25withdrawn for 0.2466, the score of answering the realized base rate of 44.2% every time
Edge of two cents a questionrestated at 0.0157 a question
Finish read through Grinold’s √breadth termwithdrawn; the leaderboard is additive in forecasts submitted
The edge is realnot statistically significant once forecasts are treated as correlated within fixtures
Calibration tablestands; No. 17 decomposes it into resolution of 0.020 against reliability of 0.006
The 80–100% bandnine forecasts, too few to carry weight on their own
Retained because this site’s argument is that tested breadth under an incorruptible scoreboard beats brilliance under a flattering one. The Cup tested it against a live field, under a referee that forecasts for a living, and it held.

References

  1. Brier, G. W. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, 78(1), 1–3.
  2. Grinold, R. C. (1989). The Fundamental Law of Active Management. Journal of Portfolio Management, 15(3), 30–37.
  3. Newey, W. K., and West, K. D. (1987). A Simple, Positive Semi-Definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix. Econometrica, 55(3), 703–708.