Abstract. Three pre-registered claims about prediction-market pricing, tested on the resolved Kalshi record. Two are falsified. Prices are not calibrated, the favourite-longshot bias is confirmed and economically large, and my claim that one-sided order flow marks crowd error is falsified in sign. Nothing is promoted to a tradable strategy.
1. Venue and data
Prediction markets resolve to a known truth, which makes calibration measurable rather than inferred. The source is a public dataset of Kalshi and Polymarket records. I use Kalshi only, a deviation recorded before running anything, because its rows carry price, taker side and size as their own fields while Polymarket trades need blockchain decode. Sample, price convention and my own graded record in these markets are in Table A.
| Measure | Value |
|---|---|
| Resolved Kalshi markets | 250,858 |
| Distinct events | 59,748 |
| Base rate of YES resolution | 0.381 |
| Price measure | volume-weighted average over the final 24 hours of trading, minimum five trades |
| Brier score | 0.0927 |
| Murphy decomposition | reliability 0.0193 − resolution 0.1617 + uncertainty 0.2358 |
| Realised YES, 0.30 to 0.40 band | 17.0%, so a buyer pays roughly twice what the contract is worth |
| Realised YES, 0.60 to 0.70 band | 87.8% |
| One market per event, 59,748 observations, 0.30 to 0.40 | +0.140 |
| One market per event, 0.60 to 0.70 | −0.214 |
| By duration, under six hours to beyond two days, longshot half | +0.104 to +0.124 |
| By duration, favourite half | −0.138 to −0.162 |
| Author’s graded record, Jump Trading Probability Cup | top 6%, 620 forecasts, mean Brier 0.231 against a 0.25 coin baseline |
2. What the measurement actually is
Correction, 19 August 2026. The price measure deviates from the pre-registration, which fixed a snapshot at “a FIXED horizon before resolution, not at the last trade”. The loader instead averages a window ending at the close, and for most of the sample that window is the market’s whole life, the failure the pre-registration was written to prevent. The headline magnitude therefore belongs to the price convention and not to the market. The within-event estimate survives every convention tested and is the durable finding here.
| Price convention | What it gives |
|---|---|
| 24-hour volume-weighted average, used throughout this note | slope of the miss 0.215 |
| Final traded price | slope of the miss 0.037, and the 0.30 to 0.40 band misses by roughly four cents per dollar of price |
| Within-event estimate, every convention tested | +0.038 to +0.052, t above 31 |
| Markets whose entire history sits inside the 24-hour window | 170,282 of 250,858, or 67.9% |
3. Results
The Murphy decomposition splits the Brier score into reliability, resolution and uncertainty. Prices sort winners from losers, but reliability, which perfect calibration would drive to zero, is not near zero. The miss is systematic and monotonic across the price range (Table 1). Longshots are dear and favourites cheap.
| Price band | n | Mean price | Realised | Error |
|---|---|---|---|---|
| 0.0-0.1 | 48,237 | 0.049 | 0.005 | +0.044 |
| 0.1-0.2 | 37,974 | 0.148 | 0.024 | +0.124 |
| 0.2-0.3 | 29,894 | 0.248 | 0.077 | +0.171 |
| 0.3-0.4 | 23,821 | 0.348 | 0.170 | +0.178 |
| 0.4-0.5 | 20,510 | 0.449 | 0.347 | +0.103 |
| 0.5-0.6 | 19,243 | 0.549 | 0.684 | -0.135 |
| 0.6-0.7 | 18,009 | 0.650 | 0.877 | -0.228 |
| 0.7-0.8 | 17,163 | 0.750 | 0.952 | -0.202 |
| 0.8-0.9 | 17,397 | 0.851 | 0.980 | -0.130 |
| 0.9-1.0 | 18,610 | 0.948 | 0.995 | -0.048 |
I predicted that markets with taker flow heavily on one side would be worse calibrated, on the reasoning that concentrated flow marks a crowd about to be wrong. The data say the opposite. One-sided flow is the best calibrated tercile (Table 2), because concentrated flow marks an outcome that has become obvious. That hypothesis stays rejected.
| Flow | n | Mean absolute error | Brier |
|---|---|---|---|
| balanced | 83,620 | 0.2754 | 0.1052 |
| middle | 83,619 | 0.2593 | 0.0971 |
| one-sided | 83,619 | 0.1992 | 0.0758 |
| Claim | What the record says |
|---|---|
| Prices are calibrated | Falsified. Realised frequency misses price in every band |
| A favourite-longshot bias is present | Confirmed, and economically large |
| One-sided taker flow marks crowd error | Falsified in sign. One-sided flow is the best calibrated tercile |
4. Robustness
Correlated contracts and intraday churn could explain the bias away, and neither does (Table A).
5. Why this is not tradable
A calibration gap measured on traded prices is a measurement, not an edge. Execution happens against a quoted book, and this dataset contains none. A commercial history of quotes at the top of the book would settle it, and the provider holds none for any market in this sample. A finding this large on a regulated exchange should raise suspicion, not confidence, until that check is made.
References
- Murphy, A. H. (1973). A New Vector Partition of the Probability Score. Journal of Applied Meteorology, 12(4), 595–600.
- Thaler, R. H., and Ziemba, W. T. (1988). Anomalies: Parimutuel Betting Markets. Journal of Economic Perspectives, 2(2), 161–174.
- Snowberg, E., and Wolfers, J. (2010). Explaining the Favorite-Longshot Bias. Journal of Political Economy, 118(4), 723–746.