Get in touchReach out on LinkedIn

The Factor Zoo Meets an Honest Bar: 158 Published
Technical Alphas, Zero Survivors of a Deflated,
Cost-Stressed Promotion Test

BlueShip Research
Working Paper No. 13 · 26 July 2026

In one line: I tested 158 published trading signals. 29 looked good at first. None survived honest math and realistic costs.

Abstract. I run all 158 technical alphas of Microsoft Qlib’s Alpha158 library as a single test family on roughly 21 years of US large-cap equities, through the museum’s fixed gates. Admitting both sign choices yields a family of 316 hypotheses. Twenty-nine of the 158 clear a naive significance bar of t ≥ 2.5. None survives a multiplicity-deflated promotion bar of t ≥ 4.11 once a margin-of-safety re-audit halves gross return and doubles costs. The strongest survivors are uniformly short-term reversal and overbought signals whose measured gross edge the doubled-cost stress removes; the volatility features collapse hardest under that stress. I read the published technical-factor zoo as short-term reversal and a volatility risk premium in disguise, and record zero promotions. The universe holds only current S&P constituents, so it flatters even these figures, which are therefore upper bounds.

Keywords: factor zoo, multiple testing, Alpha158, short-term reversal, volatility risk premium, transaction costs, survivorship bias.

1. Introduction

The published cross-section of expected returns is crowded with factors, and the count is itself the problem. Harvey, Liu, and Zhu (2016) argue that the sheer number of tests demands a significance bar well above the conventional t of 2, and Harvey and Liu (2019) catalog the resulting zoo. This paper takes that argument as an operating rule rather than a commentary. I adopt an off-the-shelf technical-factor library, run every one of its signals as one declared family, and ask a single question: does anything clear an honest bar that is deflated for multiplicity and re-audited under stress? The answer, on this universe, is nothing.

2. Data and method

The signals are the 158 technical alphas of Microsoft Qlib’s Alpha158 feature set, computed on roughly 21 years of US large-cap prices. The universe is current S&P constituents and is therefore survivorship-biased. Only names that survived to the present are counted, which lifts the apparent edge of any signal that rewards persistence. Because a signal and its negation are distinct hypotheses, the family is scored at size 316, both sign choices of all 158 features, and the multiplicity correction is set against that count before any result is read.

Three bars are applied in sequence. The first is a naive significance threshold of t ≥ 2.5. The second is a deflated promotion threshold of t ≥ 4.11, set in the Bonferroni and Taleb spirit for a family of this size. The third is a margin-of-safety re-audit in which gross return is halved and transaction costs are doubled, so that a claim must remain significant after its economics are deliberately worsened.

3. Results

Twenty-nine of the 158 alphas clear the naive t ≥ 2.5 bar. Zero clear the deflated and stressed bar. No sign choice, and no member of the 316-hypothesis family, is promoted.

The signals that come closest are, without exception, short-term reversal and overbought measures. The strongest is a days-since-20-day-high signal (IMAX20), with a Newey–West t of 4.46 that falls to a stressed t of 3.68 once gross is halved and costs are doubled; the same pattern holds for the short-term reversal and range signals near the top of the list (RSV20, RANK20, MIN5). These are real gross effects, but they trade quickly, and doubling the cost assumption vaporizes them. The volatility features collapse hardest under the same stress, consistent with their edge being compensation for bearing risk rather than a mispricing.

Table 1. Promotion outcome across the three bars, Alpha158 as one family of 158 signals (316 with both signs).
BarThresholdClearingOutcome
Naive significancet ≥ 2.529 of 158none promoted
Deflated promotion (Bonferroni/Taleb)t ≥ 4.110 of 158none promoted
Deflated + stressed (gross halved, costs doubled)0 of 158none promoted
Table 2. Strongest survivor under the naive bar and its behaviour under the cost stress.
SignalNW tStressed tOutcome
IMAX20 (days since 20-day high)4.463.68not promoted

4. Discussion

Read together, the survivors describe the zoo rather than escape it. The technical alphas that carry any measured edge on this universe are short-term reversal and a volatility risk premium in disguise: fast, overbought-and-revert signals whose profit lives inside the transaction costs, and volatility exposures that pay for risk. None of it clears a bar that is honestly deflated for the size of the search and then re-audited under doubled costs. The result is a null, and it is reported as one. Nothing here is a tradable claim; the figures are tombstones and risk observations, not a signal.

5. Limitations

Two, both of which cut in the same direction. First, the universe is current S&P constituents and so survivorship-biased; the naive counts and the survivor t-statistics are inflated by that bias, and are best read as upper bounds. A universe that included the delisted names would not raise these numbers. Second, the study covers one technical-factor library, not the whole zoo; it demonstrates that a widely used off-the-shelf set fails an honest bar, not that every published factor does. Both limitations make the conclusion more conservative, not less. The honest bar is already unmet on figures that flatter the signals.

This tombstone is retained because Alpha158 is exactly the sort of library a desk mines for product, and twenty-nine names cross the bar a casual test would use. Recording that none of them survives multiplicity and cost is what keeps the few signals that do pass worth believing.

References

  1. Harvey, C. R., and Liu, Y. (2019). A Census of the Factor Zoo. SSRN Working Paper 3341728.
  2. Harvey, C. R., Liu, Y., and Zhu, H. (2016). …and the Cross-Section of Expected Returns. Review of Financial Studies, 29(1), 5–68.

Provenance

Asserts about
Hypothesis d57232a184d38ec1, status rejected, and the Alpha158 feature family of 158 members. The features were reimplemented in this system’s own point-in-time registry from Qlib’s public definitions; Qlib’s data layer and backtester were not used.
Data veins
Adjusted OHLCV and dollar volume (Yahoo), roughly 425 current S&P constituents, 2005 to July 2026, one trading day of information lag. Ken French FF5 plus momentum stood behind the residual-alpha stage, which no member reached. Both veins are free; neither is snapshotted.
Gate regime
G2, in force 12 to 25 July 2026: deflated significance, deflated bootstrap, a 30% limit on in-sample to out-of-sample degradation, and the margin-of-safety rerun. The verdict was rendered 25 July 2026. Two engine changes dated 26 July, a small-sample guard and a fix to the factor-stage risk-free treatment, postdate this verdict and neither moves it.
Family
316, declared before testing as 158 features by two sign choices, giving a deflated bar of t ≥ 4.11. The store records a family as a count, not as a membership; the 158 names survive only in the run artifact, not in the hypothesis record.
Binding gate
Out-of-sample decay, not multiplicity. The strongest member, IMAX20, cleared the deflated bar at t = 4.46 and cleared the margin-of-safety rerun at a stressed t = 3.68 against a base bar of 2.50. What stopped it was a 55% Sharpe degradation from in-sample to out-of-sample against a 30% limit. Per-member rejection reasons were not recorded by the runner, so this identification covers the strongest member only.
Retracts if
  • a walk-forward re-audit shows that the 55% degradation is an artifact of the single fixed 70/30 split, which would carry IMAX20 to the factor stage and make the headline one of 158 rather than zero;
  • point-in-time constituents replace this current-membership universe and the naive count rises above 29 instead of falling, contradicting Section 5;
  • the reimplemented definitions are shown to diverge materially from Qlib’s published Alpha158 definitions, in which case the family tested is not the family named.
Family size is not a retraction lever here. The strongest member cleared the deflated bar unaided, so the null does not rest on the multiplicity correction.
Reproducibility
Not pinned. Prices are refetched nightly and the universe is resolved from the current index list, so neither is frozen at publication, and the runner recorded no data window. A re-run on 27 July 2026 returned t = 4.25 and a stressed t = 3.56 for IMAX20 against the 4.46 and 3.68 reported above. The verdict and the binding gate were unchanged.