Abstract. The museum edition of the companion working paper, prepared for the Prediction Markets Conference at Yale School of Management (23 October 2026). The weather paper worked because the government publishes a probabilistic forecast of the settled quantity. This paper is about what happens when it does not. [Read the full paper (PDF, 6 pages).](../papers/Dimopoulos_MissingBenchmark_YalePM2026.pdf)
1. The benchmark that does not exist
The Bureau of Labor Statistics produces the truth and publishes no forecast of it. No federal body publishes a probability distribution over any quantity Kalshi settles monthly. So inflation markets cannot be graded the way the weather market was, and anyone claiming otherwise has substituted a benchmark of their own choosing. This paper does that substitution in the open, with the only free, vintaged, official-sector forecast of the exact settled quantity: the Cleveland Fed’s inflation nowcast.
2. The comparison that can be made
| Measure | Value |
|---|---|
| Contracts, releases | 1,047 across 53 monthly CPI releases |
| Brier, market one day before release | 0.089 |
| Brier, Cleveland Fed nowcast | 0.137 |
| Difference | 0.049 (release-clustered t = 3.45) |
| Dispersion sensitivity, 0.6x to 2.0x realised error | t between 3.3 and 5.3 |
| Headline monthly / year-over-year | t = 3.3 / 4.4 |
| Core year-over-year | t = 0.55, no advantage |
The market is the better forecast where the attention is, and shows nothing where it is not; the paper reports both.
3. The measurement trap, which is the more useful result
These contracts settle on the first print. BLS re-estimates seasonal factors every February and revises five prior years, so the seasonally adjusted series available today is not the series the contracts were graded against. Of 155 months in the archive, 89, or 57 per cent, fall in a different one-tenth-of-a-point bin under the current vintage than under the first print. December 2021 first printed at +0.470 and reads +0.691 today: a contract on “above 0.5 per cent” settled no and would be scored yes. Any study grading these contracts from a current data pull is silently mis-scoring more than half its sample. That warning holds regardless of every modelling choice in section 2.
4. What this does not claim
The nowcast is a point forecast, so a dispersion assumption was required; every setting tested leaves the conclusion standing, but the assumption is named. Fifty-three releases is a small sample and is treated as one.
- Method
- Every number above is regenerated from the lab record by a strict checker before the PDF builds; 124 numbers, none unbacked. Vintage archive and replication are in the paper.
- Honesty
- The author has never held an account or position on Kalshi or any other prediction-market venue. All data are public.
- Related
- Working Paper No. 38 (the companion weather paper), No. 16 (calibration), the forecast wing (RN 5, No. 23).