Abstract. The museum edition of a full working paper, prepared for the Prediction Markets Conference at Yale School of Management (23 October 2026). This page carries the claim and the headline numbers; the paper carries the method, the robustness and the references. [Read the full paper (PDF, 9 pages).](../papers/Dimopoulos_WeatherVsNOAA_YalePM2026.pdf)
1. The question
Prediction market prices are read as probabilities, but they are rarely scored against a professional forecast of the same quantity, because for most events no such forecast exists. Daily maximum temperature is the exception. The National Weather Service publishes, through the National Blend of Models, an operational probabilistic forecast of exactly the quantity Kalshi’s temperature ladders settle on, at the same named stations, archived and free. So the comparison nobody usually gets to make is available here in full: a retail market against the government’s own probabilistic guidance, on 26,583 contracts across 6,023 city-days in eight US cities, August 2021 to November 2025.
2. The answer
The market wins, in all eight cities, and not narrowly.
| Measure | Market | The Blend |
|---|---|---|
| Brier score, 24 hours before close | 0.1355 | 0.1564 |
| Difference, event-clustered | 0.0209 (t = 22.2; bootstrap CI 0.0191 to 0.0228) | |
| Encompassing regression coefficient | 0.98 (t = 42.1) | 0.16 |
| Murphy reliability (lower is better) | 0.00027 | 0.00206 |
| Murphy resolution (higher is better) | 0.03396 | 0.01521 |
Two results give the finding its content. The market is not relaying the public bulletin: with both forecasts on the right-hand side, the market enters at essentially one and the guidance collapses to 0.16. And the advantage is not sharpness traded against calibration; the market is better on both terms.
3. The mechanism, which is not the obvious one
The guidance is not underdispersed. Standardising outcomes by its own stated standard deviation gives a variance of 1.002 on 4,407 events. What it cannot do is vary its spread with the day: where it states one degree of uncertainty the realised dispersion is 1.44, and where it states four degrees, 0.82. Its uncertainty is close to regime-invariant, so it fails to separate easy days from hard ones. That is a resolution deficit, not a calibration one, and it is the kind a market full of people looking out of the window can exploit.
4. What this does not claim
This is not a demonstration that crowds beat meteorologists. The Blend is automated statistical guidance, not the official human forecast; the market trades after the guidance publishes and may use it freely; and the result is one settled quantity on one venue. The paper states each limit and the timing in full.
- Method
- Every number above is regenerated from the lab record by a strict checker before the PDF builds; 167 numbers, none unbacked. Data, provenance and replication are in the paper.
- Honesty
- The author has never held an account or position on Kalshi or any other prediction-market venue. All data are public.
- Related
- Working Paper No. 39 (the companion, on inflation markets and the benchmark that does not exist), No. 16 (calibration), No. 23 and RN 5 (the forecast wing).