Get in touchReach out on LinkedIn

Where the Edge Was: A Category-Level Decomposition
of 620 Brier-Scored World Cup Forecasts

BlueShip Research
Working Paper No. 17 · 27 July 2026

In one line: I scored best in small markets nobody prices carefully. In the two biggest I lost ground, and it turns out my forecasts there were fine. The crowd was just better.

Abstract. I decompose by category the 620 settled forecasts of Working Paper No. 5, entered in a Brier-scored competition on the 2026 World Cup. Three thin process markets score +17.4, +10.5 and +6.9 on 11, 12 and 15 forecasts, against Match Outcome at −6.1 on 179 and Goals and Scoring at −15.9 on 145. Its units and its benchmark are undocumented, so nothing is claimed beyond the ordering and the sign. The contrarian win rate was 48%, below half, on an undisclosed denominator and at a distance from half that no subset size makes significant. I therefore retire the reading that the edge came from out-thinking the field, and read the same calibration table, resolution 0.020 against reliability 0.006, as discrimination rather than calibration. Rescoring 555 of the forecasts against a benchmark that is documented, the best constant forecast in each category, reverses the sign on both of the platform’s worst categories and puts match outcome second of nine, which locates the deficit in the strength of the competition rather than in the forecasts. Of those 555, 42.2% resolved yes, and a flat 42% forecast would have earned 44% of this book’s edge over the same benchmark knowing nothing else.

Keywords: forecasting, Brier score, category decomposition, shrinkage, market-implied priors, prediction markets.

1. Introduction

Working Paper No. 5 reported 620 settled forecasts entered in Jump Trading’s Probability Cup, a competition on the 2026 World Cup scored by squared error against the outcome (Brier, 1950). The mean Brier score was 0.2309. Against 0.2466 for a forecaster who answers the realized base rate of 44.2% every time, the edge is 0.0157 a question, and the constant 50% forecast at 0.25 that No. 5 quoted is withdrawn as a benchmark. The 0.2309 reconciles with the calibration table published there to within 0.001 by Murphy decomposition, reliability 0.006 less resolution 0.020 plus uncertainty 0.247. It is the only number here checked against a second source. It is also not significant: independent draws would give a t-statistic near 2.2, but fifteen to twenty markets a fixture were priced from one view of the same match, and at an intra-fixture correlation of 0.10 the effective sample falls to about 250 and the statistic to about 1.4.

Two aggregates were disclosed and one of them is only a rank. The platform reports +2.7 a forecast, which it renders as better than 79% of forecasters, alongside a top ~6% finish on cumulative points against an undisclosed entrant count. That board is additive in forecasts submitted, so the distance between the two percentiles is participation. Working Paper No. 5 also read the finish through Grinold’s √breadth term (Grinold, 1989), which governs confidence in an edge and not rank on an additive board, and that reading is withdrawn. What none of these figures says is where in the book the edge came from.

2. The Category Record

2.1 Where the book scored

Three categories scored highest, on 38 of the 620 forecasts between them. The platform displays halftime tied at +17.4 on 11 forecasts, penalty and red card at +10.5 on 12, and team corners at +6.9 on 15. All three settle on how a match is played rather than on who wins it, and on the account’s reading they are where a fitted base rate met the least competition.

The evidence is thinner than the mechanism is appealing. Read as the per-forecast averages the platform labels them, those three come to +420.9 points against −3,397.4 for the two largest negative categories, an eighth of what was given back. They were selected on the outcome from an undisclosed field, their contrarian sub-samples run to 7, 2 and 3 calls, and on seven calls the standard error on a win rate is roughly seventeen percentage points.

2.2 Where the book lost

The platform also groups the book into four category types, reported below as displayed. It does not state whether the figures are per-forecast averages or bucket totals, the four types cover 432 of the 620 settled forecasts, and neither reading reconciles to the account’s overall figure. The ordering and the sign are unambiguous under either.

Table 1. Relative Brier figures by category type across the 432 of 620 forecasts the platform buckets: settled forecasts, figure, and verdict.
Category typeForecastsFigureVerdict
Match Outcome179−6.1largest book; ground lost
Goals & Scoring145−15.9largest give-back; units unresolved
Set Pieces & Possessions63−2.8small negative
Cards & Discipline45+0.1flat

Match Outcome was the account’s largest category at 179 forecasts, 29% of everything it submitted, and Goals and Scoring at 145 gave back more. Together they are 52% of the book and the two worst of the four displayed types, with 188 forecasts in no displayed bucket at all. They are also the two most likely to carry a continuous public line, though the platform documents its benchmark in no category. The figures do not separate departures from calibration, so the record shows where ground was given back and not that departure there carried negative expectancy.

The fixture tails point the same way, on unequal lists. The four largest positive fixtures sum to +991.0 and the five largest negative to −2,258.2, so the comparison that holds is per fixture rather than sum against sum. Three of the five worst involved France, which reached the semi-final and supplied more fixtures than most sides. Without France’s fixture count or a test against it, that concentration is a hypothesis for the next cycle and not a diagnosis.

2.3 The same book against a documented benchmark

Everything above is the platform’s scale, whose benchmark it does not publish. That limitation is removable. The settled log carries a stated probability and a Brier score for every forecast, and squaring back recovers each outcome, so the book can be scored against a benchmark that is written down. The benchmark used here is the best constant forecast available in each category, its own realized base rate b, which scores b(1 − b). It represents a forecaster who knows how often that kind of question resolves yes and nothing else whatever. The reconstruction covers 555 of the 620 settled forecasts, the last dump taken before the semi-finals, and categories are assigned by question wording rather than by the platform’s own grouping.

Table 2. Mean Brier score by category against the best constant forecast for that category, reconstructed from 555 settled forecasts. Edge is benchmark less mean Brier, so positive is better.
CategorynMean BrierBase rateBenchmarkEdge
Halftime state280.22030.5000.2500+0.0297
Match outcome350.20730.6290.2335+0.0262
Cards430.20730.3490.2271+0.0198
Corners420.22840.4050.2409+0.0125
Penalty and red card340.16760.2350.1799+0.0123
Goals and scoring1650.24330.4550.2479+0.0047
Shots1340.23800.4030.2406+0.0026
Offsides280.25180.4290.2449−0.0069
Timing180.23850.3330.2222−0.0163
All 5555550.22990.4220.2439+0.0139

The two readings disagree, and the disagreement is the finding. Match outcome and goals and scoring are the platform’s two worst categories and are positive here, match outcome second best of the nine. Both statements can hold at once, because they are measured against different things. Beating the base rate and beating a well-informed crowd are separate achievements, and the categories where this book lost ground are the ones where the crowd was sharpest, not the ones where it forecast badly. Section 2.2 read a deficit against an undocumented benchmark as a fault in the forecasts. On the evidence here that reading is too strong, and what the deficit locates is competition rather than error.

Two results survive both readings. Halftime state is the strongest category under either, and its base rate is 0.500 to three figures, so none of its edge is the tilt described below and all of it is discrimination. Timing is negative under both, which is the category the hydration call sits in. Offsides is negative here on 28 forecasts and was not separately displayed by the platform. The counts remain small and the categories were assigned after the fact, so this table reorders the account’s own book and does not establish where an edge exists in the venue.

2.4 The tilt nobody had to model

Of the 555 reconstructed forecasts, 42.2% resolved yes. The question set was written as thresholds, four or more cards, three or more goals, a penalty awarded, a named player scoring, and threshold questions of that kind resolve no more often than not. A forecaster who answered a flat 42% to every question, knowing nothing about football, scores 0.2439. A forecaster who answered a flat 50% scores 0.2500. The difference, 0.0061 a question, is available for recognising the shape of the question set and is 44% of this book’s 0.0139 edge over the same benchmark. It was not deliberately taken. The account hedged toward 50% rather than toward 42%, which Working Paper No. 5 recorded as underconfidence in the 45–55% band without identifying its cause.

3. What the Record Does Not Establish

3.1 Departure was not shown to be the edge

The tempting reading of a strong finish is that the account out-thought the field. Working Paper No. 5 did not claim it, the finish invited it, and it is retired here. The contrarian win rate was 48%, below half, on a denominator the platform does not disclose and at a distance from half that no plausible subset size makes significant. The claim is not that departure was reliably wrong. It is that nothing in the record makes departure the source of the edge.

If the departures added anything it came through magnitude, the frequency being under half. The magnitudes do not support that either, the best single call gaining 78.0 points and the worst losing 83.4, with the fixture tails running the same way. Nor was the edge calibration. The published table decomposes into resolution 0.020 against reliability 0.006, so calibration is a small net cost, and its largest well-populated miss is the 45–55% band, 104 forecasts realizing 59.6%.

Two accounts were entered. The systematic one is described above; the second, where I picked the questions and the numbers by hand, finished in the bottom ~6%. One person, one tournament and one period, but not one question set, and the hand-picked account’s settled count is undisclosed on a board that ranks cumulatively. The pairing was not designed as an experiment, and that is both its value and its limit. A low rank on few forecasts is equally consistent with less volume and with worse judgment, so it settles nothing about which process was better. It is reported because it was the evidence in front of me when the next cycle was planned.

3.2 The cost of unshrunk conviction

The roughest call was a timing question. I posted 82% that a goal would be scored before the first hydration break, against a displayed benchmark of 48%. No goal was scored, and the call cost 83.4 relative Brier points. It sits in the 80–100% band, which Working Paper No. 5 later reported as an overconfident band, nine forecasts realizing 55.6%, one of two such tails alongside 0–20%. That table is a retrospective on the same 620 forecasts and did not exist when the call was made.

A standard rule posts pfinal = pmarket + k(pownpmarket), with k the fraction of the departure retained. At pmarket = 0.48 and pown = 0.82 the Brier scores are 0.6724, 0.5402, 0.4225, 0.3192 and 0.2304 at k = 1, 0.75, 0.50, 0.25 and 0, so k = 0.75 is shrinking a quarter of the way to the anchor. Measured over the anchor’s own 0.2304, shrinking a quarter of the way avoids 29.9% of the excess loss, half of the way 56.5%, and three quarters 79.9%. Restraint is cheapest at the start. Shrinking toward a market-implied prior is a public technique whose finance-native form is Black and Litterman (1992), and the estimator is neither new nor mine.

Three things the ladder does not establish. The 82% is what was submitted rather than a recorded unshrunk output. The 48% is displayed rounded, so this is the geometry of shrinkage on one call and not a reconstruction of what the system would have posted. And the call was chosen for being the worst, which guarantees that the rule looks good on it. Applied uniformly the same rule shrinks the best call too, 32% against a benchmark of 56% on whether Messi would score or assist, which also did not occur, surrendering 43.2% of that gain at k = 0.50. The honest test is the weight fitted across all 620 forecasts, and I have not run it. Nothing here claims that shrinkage would have improved the result.

A concurrent entrant built a system around anchoring on the consensus estimate and departing from it only where the evidence justified the departure; that work is unpublished, a fuller comparison is deferred pending permission, and nothing here rests on it.

4. Limitations

Seven, recorded during the decomposition rather than after it. First, this is one tournament, one crowd and one settlement, and the forecasts are not 620 independent draws. Second, the categories that scored highest rest on the small counts given in Section 2.1 and were selected on the outcome from an undisclosed field. Third, the relative Brier scale is the platform’s and its benchmark is not documented. The two disclosed calls cannot be fitted by one constant against a consensus benchmark but fit exactly against a mean-of-individual-forecasters benchmark, and those two constructions imply opposite answers to whether this book beat the consensus, so no consensus-relative skill claim is made anywhere here.

Fourth, the two accounts in Section 3.1 were not a designed experiment, and that section states what the omission costs. Fifth, the category-type units are undisclosed and no reading of them reconciles to the account’s overall figure. Sixth, the markets were chosen by the account rather than assigned, so the category mix is itself an output of the process and not a grid the process was tested on. Seventh, the rescoring in Section 2.3 covers 555 of the 620 forecasts, the categories there are assigned by question wording rather than by the platform, and each outcome is recovered by matching the reported Brier score to the nearer of the two values the stated probability admits. That recovery is exact wherever the two differ, which is everywhere except a stated 50%. What can be read from this decomposition is where the account scored, not where an edge exists.

5. What Changes Next Cycle

  1. Write the consensus estimate down first. The crowd forecast is a quantity to be measured before any view attaches to it, and this account never recorded it as a number. Next cycle states one for every question before the model is consulted.
  2. Shrink toward it, and set the weight by evidence quality. Retain most of a departure where the view is a fitted base rate on a market carrying no public line, almost none where it is a second opinion on a price professionals quote all day. That weight has not been fitted, and the honest estimate is the least-squares optimum over all 620 forecasts.
  3. Price the process markets with the distributions they need. Timing questions want a hazard rate rather than a flat probability, and the hydration call was priced without one. Cards and corners are counts that may be overdispersed relative to Poisson, which this book has not tested.
  4. No departures on well-covered outcome markets. Where a liquid public line exists the job is to match it, and the risk budget goes where the line is absent. Section 2.3 sharpens rather than softens this. Those forecasts were sound against a base rate and still lost ground, which is what being outclassed looks like, and the remedy is to stop competing there rather than to forecast harder.
  5. Measure the base rate first, and hedge toward it. The question set resolved yes 42.2% of the time and the account hedged toward 50%. Next cycle computes the running base rate over settled questions of the same type and treats that, not the coin, as the answer given in the absence of a view.
This study is retained because it withdraws a claim the collection published, and then withdraws one of its own. The earlier paper let a strong finish stand as evidence of judgment. The platform’s category record appeared to say the largest books forecast badly, and rescoring them against a benchmark that is written down says instead that they were outclassed, which is a different fault with a different remedy. A collection that only ever adds papers is not keeping a ledger, and one that will not correct the paper it published this morning is not keeping one either.

References

  1. Black, F., and Litterman, R. (1992). Global Portfolio Optimization. Financial Analysts Journal, 48(5), 28–43.
  2. Brier, G. W. (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, 78(1), 1–3.
  3. Grinold, R. C. (1989). The Fundamental Law of Active Management. Journal of Portfolio Management, 15(3), 30–37.