1. The shop that started it
In the mid-1980s Morgan Stanley ran a small quantitative group under Nunzio Tartaglia: buy the laggard, short the leader, wait for the snap back. Its alumni seeded modern quantitative finance, David Shaw among them, and Gatev, Goetzmann and Rouwenhorst later published the definitive academic test. I managed risk at the same firm decades later. The published record since is a staircase going down, in their numbers, not mine.
| Source | What the record says |
|---|---|
| Gatev, Goetzmann and Rouwenhorst, Review of Financial Studies 2006 | 1.44% per month on the twenty closest pairs; 0.895% per month if you wait one day before trading, a gap implying roughly 1.6% of round-trip friction |
| Do and Faff, by era | 0.86% per month through 1988, 0.37% through 2002, 0.24% after that; negative once realistic costs are charged |
| Do and Faff, the one after-cost survivor | Carefully matched same-industry pairs, around 0.3% per month |
| Engelberg, Gao and Jagannathan | Profits fade within days of a divergence and live mostly in hard-to-trade, illiquid stocks |
| Zhu 2024, not peer-reviewed | Screening out illiquid names removes roughly 60% of the gross profit |
| Farago and Hjalmarsson, JFQA 2019 | The tight statistical relationship the textbook assumes between paired stocks mostly does not exist |
2. Five tests, a rising bar
Testing five variants of one idea gives five chances to get lucky, so the evidence required must rise with each attempt, and my pipeline raises the required t-statistic automatically. Most tools do the opposite: try ten versions, show the best, grade it as though it were the only one.
The first three tests ranked stocks by distance from their usual partner. The ranking holds real information, is far too weak to trade, and what profit it held sat in the least liquid third of the market. My pipeline’s reviewer, a second agent that attacks results before publication, then ruled that ranking stocks is not trading pairs, so I built a faithful replication of the original trade. It also caught a sign error that had the replication running backwards; every verdict below comes from the corrected run.
| Measure | Value |
|---|---|
| Tests, registered as variants of one idea | Five |
| Required t-statistic, first test to fifth | 2.50 rising to 3.03 |
| Ranking signal, formal test | t-statistic of 2.52, returns in order from best-ranked bucket to worst |
| Best risk-adjusted return of the ranking | Near 0.2, where a strategy worth running scores closer to 1 |
| Share of ranking profit from the least liquid third | 94 to 96%, in two of the three tests |
| Replication design | Twenty closest pairs, two-standard-deviation entry, convergence exit, one-day execution delay |
| Replication scale | 247 overlapping portfolios, 8,051 round trips, average holding period 54 days |
| Convergence earns, before costs | Roughly 0.3% a year; 0.30% on the corrected run |
| Conservative trading costs | Roughly 0.7% a year; 0.67% on the corrected run |
| Cost charged | 5 basis points per traded dollar, plus borrow |
| Classic strategy, after costs | Loses 0.37% a year |
| Slow-pairs version, after costs | Loses 0.57% a year |
| Slow pairs, before costs | Approximately zero |
| Passing bar on the fifth test | 3.03, and both replication results are negative before any statistical standard applies |
3. Not wrong, just too expensive
Diverged twins still converge, so the strategy is not wrong about the market. It simply costs about twice what it earns on the only stocks liquid enough to matter at size. In smaller names the published research still finds gross profit and so do I: the profit sits where trading is dearest, so it was always a fee for supplying liquidity, not information about prices.
4. What remains
The one new result is the slow-pairs test: if delay destroys fast-converging pairs, restrict to pairs that converge over weeks. The literature records the damage of delay but never tested that rescue. I tested it twice, and the slow pairs carry no edge to protect. What survives is the literature’s own after-cost pocket, their claim and untested here; the illiquid pocket, too small to absorb capital; and merger arbitrage, the pairs trading textbook’s last third, where trading is forced by the deal and needs a merger database, a data problem rather than a modelling one. A sixth test must bring a new mechanism, not another formula.
5. What this does not establish
The caveats below are stated in full; the largest of them make the rejection stronger, not weaker.
| Caveat | Effect on the verdict |
|---|---|
| Universe | Today’s large-cap survivors, which flatters this strategy specifically: a pair that diverged because one company was collapsing towards delisting never enters the sample, so losses from divergences that never came back are undercounted. Every gross profit figure above is an upper bound, which makes the rejections stronger, not weaker |
| Costs, the one that runs the other way | The 5 basis points per traded dollar is venue-blind and carries no legging cost; a pairs order that fills the liquid leg and misses the other is, in the words of a market-design practitioner, being “handed back” with the hedge on and the position off. Modelling it would make the cost-death verdict stronger still (added 21 August 2026) |
| Scope | One pipeline, one universe, daily prices. Nothing here speaks to intraday versions of the trade run with serious execution infrastructure, which is a different business |
| The literature’s surviving pockets | Reported as their claims, not replicated here |
| Dead variants | Each is recorded in a database with its cause of death, so none is retested here by accident. That is the entire point of publishing failures |
Provenance, bs-prov/1.0
- Method
- Five registered tests of one idea, judged 11 August 2026, with the required t-statistic rising from 2.50 to 3.03 as the test count grew. Pair-level replication: 252-day formation window, 126-day trading window, two-standard-deviation entry, convergence exit, one-day execution delay, 5 basis points per traded dollar plus borrow costs, capital counted as committed whether deployed or not; 247 overlapping portfolios, 8,051 round trips. The independent reviewer’s two interventions (the design ruling, and the sign error caught before publication) are documented in the pipeline’s cycle report.
- Sources
- Gatev, Goetzmann and Rouwenhorst, Review of Financial Studies 2006; Do and Faff 2010 and 2012; Engelberg, Gao and Jagannathan; Zhu 2024 (not peer-reviewed); Farago and Hjalmarsson, JFQA 2019; Vidyamurthy, Pairs Trading: Quantitative Methods and Analysis, 2004. History per the 2006 paper’s account and standard references.
- Honesty
- The test universe is today’s large-cap survivors, which flatters this strategy specifically; every gross profit figure is an upper bound and the rejections are robust in the only direction that matters.
- Related
- Working Paper No. 13 (the deflated bar), No. 22 (who has to trade), No. 25 (what daily bars cannot see), No. 28 (the constraint audit), No. 29 (the ninety-second quant).