Get in touchReach out on LinkedIn

620 Probabilities: Calibration and Compounded Breadth in a
Brier-Scored World Cup Forecasting Competition

BlueShip Research
Working Paper No. 5 · 20 July 2026

In one line: I made 620 World Cup forecasts, each barely better than a coin flip. That finished in the top 6%.

Correction, 27 July 2026. Three claims below are revised in Working Paper No. 17. The 0.25 coin-flip baseline is withdrawn in favour of 0.2466, the score of a forecaster who answers the realized base rate of 44.2% every time, which puts the edge at 0.0157 a question rather than the two cents quoted here. The reading of the finish through Grinold’s √breadth term is withdrawn, the competition board being additive in forecasts submitted. And the edge is not statistically significant once the forecasts are treated as correlated within fixtures. The calibration table itself stands, and No. 17 decomposes it into resolution of 0.020 against reliability of 0.006.

Abstract. I report the performance of a systematic forecasting agent entered in Jump Trading’s Probability Cup, a Brier-scored competition on the 2026 World Cup. Across 620 settled forecasts the agent’s mean Brier score was 0.2309, against a coin-flip baseline of 0.25. That roughly two-cent per-question edge compounded to a top-6% finish with no single high-conviction call. A calibration audit isolates two opposite errors: in the 45–55% band the agent was underconfident, with outcomes realizing 59.6% of the time, and in the 80–100% band it was overconfident, realizing 55.6% across nine forecasts. I interpret the result through Grinold’s fundamental law: measured skill accrues through the number of honest, independent attempts rather than the magnitude of any single call.

Keywords: forecasting, calibration, Brier score, proper scoring rules, fundamental law of active management, prediction markets.

1. Introduction

Jump Trading’s Probability Cup is a public forecasting competition on the 2026 World Cup. Entrants submit probabilities on match questions, for example whether a match will produce 22 or more shots or whether Spain will put five shots on target, and every forecast is scored by the Brier score, the squared error between the stated probability and the realized outcome (Brier, 1950). The rule admits no narrative and awards no partial credit for a plausible story. This study entered an automated agent for that reason: the aim was to have the pipeline’s research habits graded by a scoreboard that cannot be argued with. The final standing was the top 6% of entrants, reached with no high-conviction call.

2. A Proper Scoring Rule and Compounded Breadth

Always answering 50% yields a Brier score of 0.25, the coin-flip baseline. Across 620 settled forecasts the agent’s mean Brier score was 0.2309, a roughly two-cent edge over the coin per question. A proper scoring rule compounds that edge the way a return stream compounds: a small, genuine edge taken 620 independent times separates a forecaster from a field that is largely either guessing, at 0.25, or confident and wrong, above it.

This is the mechanism of Grinold’s fundamental law of active management, IR = IC × √breadth (Grinold, 1989): the skill realized in the information ratio scales with the number of independent, honest attempts, not with the conviction attached to any one. The same law governs the collection’s equity pipeline and its Medallion case study; the Probability Cup supplies an out-of-sample instance in a separate arena.

3. Calibration

3.1 The calibration table

Calibration asks whether a stated probability matches its realized frequency: when the agent says 40%, does the event occur 40% of the time? The full table is reported below, including the two bands in which calibration failed.

Table 1. Calibration by stated-probability band: realized frequency, forecast count, and verdict.
Stated probabilityRealizedForecastsVerdict
0–20%24.1%54slightly overconfident on longshots
20–35%32.1%187calibrated
35–45%40.5%158calibrated
45–55%59.6%104underconfident; edge left on the table
55–65%60.9%69calibrated
65–80%71.8%39calibrated
80–100%55.6%9overconfident in the tail (small n)

3.2 Two opposite errors

The two failing bands are opposite errors. In the middle band the agent knew more than it stated: when it reported near 51%, outcomes realized near 60%. Hedging toward the coin is not prudence but a tax, and 40% of all forecasts fell in the 40–60 band. In the tail the reverse held: on the nine occasions the agent was nearly certain, it was right barely half the time. Certainty was rare and still overpriced, the same tail-risk error documented in the collection’s LTCM case study.

4. Transfer to Equity Research

  1. Breadth beats heroics. A top-6% finish came with no memorable single prediction. The pipeline rests on the same premise: 182 strategies tested, with promotion decided by accumulated honest attempts rather than by any one backtest.
  2. Use a scoreboard that cannot be charmed. The Brier score is a proper scoring rule; its maximum is reached only by reporting one’s true belief. The promotion gates in the equity work, Newey-West t-statistics, margin-of-safety reruns, and multiple-testing deflators, are the attempt to build the same property into research.
  3. The market is the prior. The agent’s edge came from disciplined base rates updated with public information, not from out-thinking the field from scratch. Prediction markets are being quantified by large firms for this reason: they are young markets, and young markets reward the early and the systematic.
  4. Publish misses alongside rank. The two failing bands cost leaderboard places and are the most useful rows in the table: the next cycle sharpens the mid-range forecasts and tempers the near-certain ones. A forecast log that cannot be audited is a highlight reel.
This study is retained because the collection’s argument is that tested breadth under an incorruptible scoreboard beats brilliance under a flattering one. The Probability Cup tested that argument in an external arena, against a live field, under a referee run by a firm that forecasts for a living, and the argument held.

References

  1. Brier, G. W. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, 78(1), 1–3.
  2. Grinold, R. C. (1989). The Fundamental Law of Active Management. Journal of Portfolio Management, 15(3), 30–37.
  3. Newey, W. K., and West, K. D. (1987). A Simple, Positive Semi-Definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix. Econometrica, 55(3), 703–708.