Get in touchReach out on LinkedIn

Priced Wrong on Purpose: Favourite-Longshot Bias and the
Failure of Two Pre-Registered Hypotheses in Prediction Markets

BlueShip Research
Working Paper No. 16 · 26 July 2026

In one line: Contracts priced near 35 cents pay out 17% of the time, and two of my three predictions were wrong.

Abstract. I test three pre-registered claims about prediction-market pricing on 250,858 resolved Kalshi markets spanning 59,748 distinct events. The claims were written and committed before the dataset finished downloading. Two of the three are falsified. Prices are not calibrated: the reliability component of a 0.0927 Brier score is 0.0193, and the reliability curve departs sharply from the diagonal. The favourite-longshot bias is confirmed and is economically large, with contracts trading near 0.35 resolving YES 17.0% of the time and contracts near 0.65 resolving 87.8%. My third claim, that one-sided order flow marks crowd error, is falsified in SIGN: one-sided books are better calibrated than balanced ones, not worse. The bias survives restriction to one market per event and holds across every contract-duration bucket. I do not promote it, because the dataset carries no order book and a gap measured on traded prices is not an edge until the spread is shown not to consume it.

Keywords: prediction markets, calibration, favourite-longshot bias, Brier decomposition, pre-registration, order flow.

1. Why this venue

Every equity study in this collection reaches the same verdict. Signals derived from price in large-capitalisation US equities are picked over, and nothing in the published factor zoo survives an honest bar. Prediction markets are a different population. They resolve to a known truth, which makes calibration directly measurable rather than inferred, and they are traded by a different set of participants. I have a graded record in them, having finished in the top 6% of Jump Trading's Probability Cup over 620 forecasts at a mean Brier score of 0.231 against a 0.25 coin baseline. That record is what makes the question interesting to me and is also why the pre-registration matters.

2. Data and method

The source is a public dataset of Kalshi and Polymarket market and trade records. I use Kalshi only, and recorded that deviation before running anything: Kalshi trade rows carry price, taker side and size as first-class fields and market rows carry the resolution, whereas Polymarket trades are raw blockchain events requiring decode, which would inject avoidable error into a calibration measurement. For each resolved market I take the volume-weighted average price over the final 24 hours of trading, a fixed window so that fast-resolving and slow-resolving contracts are treated alike, and require at least five trades in that window. The base rate of YES resolution is 0.381.

3. Results

3.1 Calibration, and the size of the miss

The Murphy decomposition gives Brier 0.0927 = reliability 0.0193 − resolution 0.1617 + uncertainty 0.2358. The resolution term is large, meaning prices do separate winners from losers, but the reliability term is not near zero and the direction of the miss is systematic rather than noisy.

Table 1. Realised outcome frequency by price bucket. Error is mean price minus realised frequency, so positive means the market priced the contract too high.
Price bandnMean priceRealisedError
0.0-0.148,2370.0490.005+0.044
0.1-0.237,9740.1480.024+0.124
0.2-0.329,8940.2480.077+0.171
0.3-0.423,8210.3480.170+0.178
0.4-0.520,5100.4490.347+0.103
0.5-0.619,2430.5490.684-0.135
0.6-0.718,0090.6500.877-0.228
0.7-0.817,1630.7500.952-0.202
0.8-0.917,3970.8510.980-0.130
0.9-1.018,6100.9480.995-0.048

3.2 The favourite-longshot bias

The pattern in Table 1 is the classical one and it is steep. Contracts trading between 0.30 and 0.40 resolve YES 17.0% of the time, so a buyer pays roughly twice what the contract is worth. At the other end, contracts trading between 0.60 and 0.70 resolve 87.8% of the time. Longshots are dear and favourites are cheap, monotonically across the book.

3.3 Crowding, falsified in sign

I predicted that markets with heavily one-sided taker flow would be worse calibrated, on the reasoning that concentrated flow marks a crowd about to be wrong. The data say the opposite.

Table 2. Calibration error by taker-flow imbalance tercile. Lower is better.
FlownMean absolute errorBrier
balanced83,6200.27540.1052
middle83,6190.25930.0971
one-sided83,6190.19920.0758

One-sided books are the best calibrated of the three groups. The natural reading is that concentrated flow usually marks an outcome that has become obvious, not a crowd in error. My hypothesis had the sign backwards, and recording it in advance is what makes saying so cheap.

4. Robustness

Two objections could explain the bias away, and neither does. First, correlated contracts: a single event often carries many strikes, so treating markets as independent would overstate the sample. Restricting to one market per event, 59,748 independent observations, leaves the bias essentially unchanged at +0.140 in the 0.30 to 0.40 band and −0.214 in the 0.60 to 0.70 band. Second, intraday churn: Kalshi lists many short-duration contracts that might behave differently. Splitting by contract duration, the longshot-half error runs +0.104 to +0.124 and the favourite-half error −0.138 to −0.162 across contracts lasting under six hours, six to forty-eight hours, and beyond two days. The effect is not an artifact of either.

5. Why this is not promoted

A twenty-point gap invites the conclusion that this is money. It is not, on this evidence, and the reason was written into the pre-registration before any number existed: a calibration gap measured on traded prices is a measurement, not an edge. Execution happens against a quoted book, and this dataset contains no order book. If a market showing a mid of 0.65 quotes 0.60 bid against 0.72 ask, the gap is consumed before a position exists. I attempted to close this with a commercial top-of-book history and found the provider holds no data for any market in this sample, so the question remains open rather than answered. A finding this large on a regulated exchange with professional participants should raise suspicion, not confidence, until that check is made.

Two of three pre-registered claims failed, and one failed in the direction opposite to my reasoning. That is the return on writing the predictions down first. A study that confirms everything its author expected is usually measuring the author.

References

  1. Murphy, A. H. (1973). A New Vector Partition of the Probability Score. Journal of Applied Meteorology, 12(4), 595–600.
  2. Thaler, R. H., and Ziemba, W. T. (1988). Anomalies: Parimutuel Betting Markets. Journal of Economic Perspectives, 2(2), 161–174.
  3. Snowberg, E., and Wolfers, J. (2010). Explaining the Favorite-Longshot Bias. Journal of Political Economy, 118(4), 723–746.