Priced Wrong on Purpose: Favourite-Longshot Bias and the
Failure of Two Pre-Registered Hypotheses in Prediction Markets
In one line: Contracts priced near 35 cents pay out 17% of the time, and two of my three predictions were wrong.
Abstract. I test three pre-registered claims about prediction-market pricing on 250,858 resolved Kalshi markets spanning 59,748 distinct events. The claims were written and committed before the dataset finished downloading. Two of the three are falsified. Prices are not calibrated: the reliability component of a 0.0927 Brier score is 0.0193, and the reliability curve departs sharply from the diagonal. The favourite-longshot bias is confirmed and is economically large, with contracts trading near 0.35 resolving YES 17.0% of the time and contracts near 0.65 resolving 87.8%. My third claim, that one-sided order flow marks crowd error, is falsified in SIGN: one-sided books are better calibrated than balanced ones, not worse. The bias survives restriction to one market per event and holds across every contract-duration bucket. I do not promote it, because the dataset carries no order book and a gap measured on traded prices is not an edge until the spread is shown not to consume it.
1. Why this venue
Every equity study in this collection reaches the same verdict. Signals derived from price in large-capitalisation US equities are picked over, and nothing in the published factor zoo survives an honest bar. Prediction markets are a different population. They resolve to a known truth, which makes calibration directly measurable rather than inferred, and they are traded by a different set of participants. I have a graded record in them, having finished in the top 6% of Jump Trading's Probability Cup over 620 forecasts at a mean Brier score of 0.231 against a 0.25 coin baseline. That record is what makes the question interesting to me and is also why the pre-registration matters.
2. Data and method
The source is a public dataset of Kalshi and Polymarket market and trade records. I use Kalshi only, and recorded that deviation before running anything: Kalshi trade rows carry price, taker side and size as first-class fields and market rows carry the resolution, whereas Polymarket trades are raw blockchain events requiring decode, which would inject avoidable error into a calibration measurement. For each resolved market I take the volume-weighted average price over the final 24 hours of trading, a fixed window so that fast-resolving and slow-resolving contracts are treated alike, and require at least five trades in that window. The base rate of YES resolution is 0.381.
3. Results
3.1 Calibration, and the size of the miss
The Murphy decomposition gives Brier 0.0927 = reliability 0.0193 − resolution 0.1617 + uncertainty 0.2358. The resolution term is large, meaning prices do separate winners from losers, but the reliability term is not near zero and the direction of the miss is systematic rather than noisy.
| Price band | n | Mean price | Realised | Error |
|---|---|---|---|---|
| 0.0-0.1 | 48,237 | 0.049 | 0.005 | +0.044 |
| 0.1-0.2 | 37,974 | 0.148 | 0.024 | +0.124 |
| 0.2-0.3 | 29,894 | 0.248 | 0.077 | +0.171 |
| 0.3-0.4 | 23,821 | 0.348 | 0.170 | +0.178 |
| 0.4-0.5 | 20,510 | 0.449 | 0.347 | +0.103 |
| 0.5-0.6 | 19,243 | 0.549 | 0.684 | -0.135 |
| 0.6-0.7 | 18,009 | 0.650 | 0.877 | -0.228 |
| 0.7-0.8 | 17,163 | 0.750 | 0.952 | -0.202 |
| 0.8-0.9 | 17,397 | 0.851 | 0.980 | -0.130 |
| 0.9-1.0 | 18,610 | 0.948 | 0.995 | -0.048 |
3.2 The favourite-longshot bias
The pattern in Table 1 is the classical one and it is steep. Contracts trading between 0.30 and 0.40 resolve YES 17.0% of the time, so a buyer pays roughly twice what the contract is worth. At the other end, contracts trading between 0.60 and 0.70 resolve 87.8% of the time. Longshots are dear and favourites are cheap, monotonically across the book.
3.3 Crowding, falsified in sign
I predicted that markets with heavily one-sided taker flow would be worse calibrated, on the reasoning that concentrated flow marks a crowd about to be wrong. The data say the opposite.
| Flow | n | Mean absolute error | Brier |
|---|---|---|---|
| balanced | 83,620 | 0.2754 | 0.1052 |
| middle | 83,619 | 0.2593 | 0.0971 |
| one-sided | 83,619 | 0.1992 | 0.0758 |
One-sided books are the best calibrated of the three groups. The natural reading is that concentrated flow usually marks an outcome that has become obvious, not a crowd in error. My hypothesis had the sign backwards, and recording it in advance is what makes saying so cheap.
4. Robustness
Two objections could explain the bias away, and neither does. First, correlated contracts: a single event often carries many strikes, so treating markets as independent would overstate the sample. Restricting to one market per event, 59,748 independent observations, leaves the bias essentially unchanged at +0.140 in the 0.30 to 0.40 band and −0.214 in the 0.60 to 0.70 band. Second, intraday churn: Kalshi lists many short-duration contracts that might behave differently. Splitting by contract duration, the longshot-half error runs +0.104 to +0.124 and the favourite-half error −0.138 to −0.162 across contracts lasting under six hours, six to forty-eight hours, and beyond two days. The effect is not an artifact of either.
5. Why this is not promoted
A twenty-point gap invites the conclusion that this is money. It is not, on this evidence, and the reason was written into the pre-registration before any number existed: a calibration gap measured on traded prices is a measurement, not an edge. Execution happens against a quoted book, and this dataset contains no order book. If a market showing a mid of 0.65 quotes 0.60 bid against 0.72 ask, the gap is consumed before a position exists. I attempted to close this with a commercial top-of-book history and found the provider holds no data for any market in this sample, so the question remains open rather than answered. A finding this large on a regulated exchange with professional participants should raise suspicion, not confidence, until that check is made.
References
- Murphy, A. H. (1973). A New Vector Partition of the Probability Score. Journal of Applied Meteorology, 12(4), 595–600.
- Thaler, R. H., and Ziemba, W. T. (1988). Anomalies: Parimutuel Betting Markets. Journal of Economic Perspectives, 2(2), 161–174.
- Snowberg, E., and Wolfers, J. (2010). Explaining the Favorite-Longshot Bias. Journal of Political Economy, 118(4), 723–746.