Abstract. Three pre-registered claims about prediction-market pricing, tested on the resolved Kalshi record. Two are falsified. Prices are not calibrated, the favourite-longshot bias is confirmed and economically large, and my claim that one-sided order flow marks crowd error is falsified in sign. Nothing is promoted to a tradable strategy.

1. Venue and data

Prediction markets resolve to a known truth, which makes calibration measurable rather than inferred. The source is a public dataset of Kalshi and Polymarket records. I use Kalshi only, a deviation recorded before running anything, because its rows carry price, taker side and size as their own fields while Polymarket trades need blockchain decode. Sample, price convention and my own graded record in these markets are in Table A.

Table A. Sample, design and headline results.
MeasureValue
Resolved Kalshi markets250,858
Distinct events59,748
Base rate of YES resolution0.381
Price measurevolume-weighted average over the final 24 hours of trading, minimum five trades
Brier score0.0927
Murphy decompositionreliability 0.0193 − resolution 0.1617 + uncertainty 0.2358
Realised YES, 0.30 to 0.40 band17.0%, so a buyer pays roughly twice what the contract is worth
Realised YES, 0.60 to 0.70 band87.8%
One market per event, 59,748 observations, 0.30 to 0.40+0.140
One market per event, 0.60 to 0.70−0.214
By duration, under six hours to beyond two days, longshot half+0.104 to +0.124
By duration, favourite half−0.138 to −0.162
Author’s graded record, Jump Trading Probability Cuptop 6%, 620 forecasts, mean Brier 0.231 against a 0.25 coin baseline

2. What the measurement actually is

Correction, 19 August 2026. The price measure deviates from the pre-registration, which fixed a snapshot at “a FIXED horizon before resolution, not at the last trade”. The loader instead averages a window ending at the close, and for most of the sample that window is the market’s whole life, the failure the pre-registration was written to prevent. The headline magnitude therefore belongs to the price convention and not to the market. The within-event estimate survives every convention tested and is the durable finding here.

Table C. Sensitivity of the miss to the price convention. Rerun of 18 August 2026.
Price conventionWhat it gives
24-hour volume-weighted average, used throughout this noteslope of the miss 0.215
Final traded priceslope of the miss 0.037, and the 0.30 to 0.40 band misses by roughly four cents per dollar of price
Within-event estimate, every convention tested+0.038 to +0.052, t above 31
Markets whose entire history sits inside the 24-hour window170,282 of 250,858, or 67.9%

3. Results

The Murphy decomposition splits the Brier score into reliability, resolution and uncertainty. Prices sort winners from losers, but reliability, which perfect calibration would drive to zero, is not near zero. The miss is systematic and monotonic across the price range (Table 1). Longshots are dear and favourites cheap.

Table 1. Realised outcome frequency by price bucket. Error is mean price minus realised frequency.
Price bandnMean priceRealisedError
0.0-0.148,2370.0490.005+0.044
0.1-0.237,9740.1480.024+0.124
0.2-0.329,8940.2480.077+0.171
0.3-0.423,8210.3480.170+0.178
0.4-0.520,5100.4490.347+0.103
0.5-0.619,2430.5490.684-0.135
0.6-0.718,0090.6500.877-0.228
0.7-0.817,1630.7500.952-0.202
0.8-0.917,3970.8510.980-0.130
0.9-1.018,6100.9480.995-0.048

I predicted that markets with taker flow heavily on one side would be worse calibrated, on the reasoning that concentrated flow marks a crowd about to be wrong. The data say the opposite. One-sided flow is the best calibrated tercile (Table 2), because concentrated flow marks an outcome that has become obvious. That hypothesis stays rejected.

Table 2. Calibration error by tercile (three equal buckets) of taker flow imbalance. Lower is better.
FlownMean absolute errorBrier
balanced83,6200.27540.1052
middle83,6190.25930.0971
one-sided83,6190.19920.0758
Table B. The three pre-registered claims and their verdicts.
ClaimWhat the record says
Prices are calibratedFalsified. Realised frequency misses price in every band
A favourite-longshot bias is presentConfirmed, and economically large
One-sided taker flow marks crowd errorFalsified in sign. One-sided flow is the best calibrated tercile

4. Robustness

Correlated contracts and intraday churn could explain the bias away, and neither does (Table A).

5. Why this is not tradable

A calibration gap measured on traded prices is a measurement, not an edge. Execution happens against a quoted book, and this dataset contains none. A commercial history of quotes at the top of the book would settle it, and the provider holds none for any market in this sample. A finding this large on a regulated exchange should raise suspicion, not confidence, until that check is made.

A study that confirms everything its author expected is usually measuring the author.

References

  1. Murphy, A. H. (1973). A New Vector Partition of the Probability Score. Journal of Applied Meteorology, 12(4), 595–600.
  2. Thaler, R. H., and Ziemba, W. T. (1988). Anomalies: Parimutuel Betting Markets. Journal of Economic Perspectives, 2(2), 161–174.
  3. Snowberg, E., and Wolfers, J. (2010). Explaining the Favorite-Longshot Bias. Journal of Political Economy, 118(4), 723–746.