1. The shop that started it

In the mid-1980s Morgan Stanley ran a small quantitative group under Nunzio Tartaglia: buy the laggard, short the leader, wait for the snap back. Its alumni seeded modern quantitative finance, David Shaw among them, and Gatev, Goetzmann and Rouwenhorst later published the definitive academic test. I managed risk at the same firm decades later. The published record since is a staircase going down, in their numbers, not mine.

Table 1. Forty years of shrinking profits, as reported in the published literature.
SourceWhat the record says
Gatev, Goetzmann and Rouwenhorst, Review of Financial Studies 20061.44% per month on the twenty closest pairs; 0.895% per month if you wait one day before trading, a gap implying roughly 1.6% of round-trip friction
Do and Faff, by era0.86% per month through 1988, 0.37% through 2002, 0.24% after that; negative once realistic costs are charged
Do and Faff, the one after-cost survivorCarefully matched same-industry pairs, around 0.3% per month
Engelberg, Gao and JagannathanProfits fade within days of a divergence and live mostly in hard-to-trade, illiquid stocks
Zhu 2024, not peer-reviewedScreening out illiquid names removes roughly 60% of the gross profit
Farago and Hjalmarsson, JFQA 2019The tight statistical relationship the textbook assumes between paired stocks mostly does not exist

2. Five tests, a rising bar

Testing five variants of one idea gives five chances to get lucky, so the evidence required must rise with each attempt, and my pipeline raises the required t-statistic automatically. Most tools do the opposite: try ten versions, show the best, grade it as though it were the only one.

The first three tests ranked stocks by distance from their usual partner. The ranking holds real information, is far too weak to trade, and what profit it held sat in the least liquid third of the market. My pipeline’s reviewer, a second agent that attacks results before publication, then ruled that ranking stocks is not trading pairs, so I built a faithful replication of the original trade. It also caught a sign error that had the replication running backwards; every verdict below comes from the corrected run.

Table 2. Five registered tests of one idea, judged on the corrected run. Large United States stocks, daily prices.
MeasureValue
Tests, registered as variants of one ideaFive
Required t-statistic, first test to fifth2.50 rising to 3.03
Ranking signal, formal testt-statistic of 2.52, returns in order from best-ranked bucket to worst
Best risk-adjusted return of the rankingNear 0.2, where a strategy worth running scores closer to 1
Share of ranking profit from the least liquid third94 to 96%, in two of the three tests
Replication designTwenty closest pairs, two-standard-deviation entry, convergence exit, one-day execution delay
Replication scale247 overlapping portfolios, 8,051 round trips, average holding period 54 days
Convergence earns, before costsRoughly 0.3% a year; 0.30% on the corrected run
Conservative trading costsRoughly 0.7% a year; 0.67% on the corrected run
Cost charged5 basis points per traded dollar, plus borrow
Classic strategy, after costsLoses 0.37% a year
Slow-pairs version, after costsLoses 0.57% a year
Slow pairs, before costsApproximately zero
Passing bar on the fifth test3.03, and both replication results are negative before any statistical standard applies

3. Not wrong, just too expensive

Diverged twins still converge, so the strategy is not wrong about the market. It simply costs about twice what it earns on the only stocks liquid enough to matter at size. In smaller names the published research still finds gross profit and so do I: the profit sits where trading is dearest, so it was always a fee for supplying liquidity, not information about prices.

4. What remains

The one new result is the slow-pairs test: if delay destroys fast-converging pairs, restrict to pairs that converge over weeks. The literature records the damage of delay but never tested that rescue. I tested it twice, and the slow pairs carry no edge to protect. What survives is the literature’s own after-cost pocket, their claim and untested here; the illiquid pocket, too small to absorb capital; and merger arbitrage, the pairs trading textbook’s last third, where trading is forced by the deal and needs a merger database, a data problem rather than a modelling one. A sixth test must bring a new mechanism, not another formula.

5. What this does not establish

The caveats below are stated in full; the largest of them make the rejection stronger, not weaker.

Table 3. What this paper does not claim.
CaveatEffect on the verdict
UniverseToday’s large-cap survivors, which flatters this strategy specifically: a pair that diverged because one company was collapsing towards delisting never enters the sample, so losses from divergences that never came back are undercounted. Every gross profit figure above is an upper bound, which makes the rejections stronger, not weaker
Costs, the one that runs the other wayThe 5 basis points per traded dollar is venue-blind and carries no legging cost; a pairs order that fills the liquid leg and misses the other is, in the words of a market-design practitioner, being “handed back” with the hedge on and the position off. Modelling it would make the cost-death verdict stronger still (added 21 August 2026)
ScopeOne pipeline, one universe, daily prices. Nothing here speaks to intraday versions of the trade run with serious execution infrastructure, which is a different business
The literature’s surviving pocketsReported as their claims, not replicated here
Dead variantsEach is recorded in a database with its cause of death, so none is retested here by accident. That is the entire point of publishing failures

Provenance, bs-prov/1.0

Method
Five registered tests of one idea, judged 11 August 2026, with the required t-statistic rising from 2.50 to 3.03 as the test count grew. Pair-level replication: 252-day formation window, 126-day trading window, two-standard-deviation entry, convergence exit, one-day execution delay, 5 basis points per traded dollar plus borrow costs, capital counted as committed whether deployed or not; 247 overlapping portfolios, 8,051 round trips. The independent reviewer’s two interventions (the design ruling, and the sign error caught before publication) are documented in the pipeline’s cycle report.
Sources
Gatev, Goetzmann and Rouwenhorst, Review of Financial Studies 2006; Do and Faff 2010 and 2012; Engelberg, Gao and Jagannathan; Zhu 2024 (not peer-reviewed); Farago and Hjalmarsson, JFQA 2019; Vidyamurthy, Pairs Trading: Quantitative Methods and Analysis, 2004. History per the 2006 paper’s account and standard references.
Honesty
The test universe is today’s large-cap survivors, which flatters this strategy specifically; every gross profit figure is an upper bound and the rejections are robust in the only direction that matters.
Related
Working Paper No. 13 (the deflated bar), No. 22 (who has to trade), No. 25 (what daily bars cannot see), No. 28 (the constraint audit), No. 29 (the ninety-second quant).