Abstract. Every technical alpha in Microsoft Qlib’s Alpha158 library was run as one declared family, through this site’s fixed statistical hurdles. Some clear a naive bar. None is promoted.
1. Introduction
Harvey, Liu, and Zhu (2016) argue that the number of published tests demands a significance bar well above convention, and Harvey and Liu (2019) catalogue the resulting zoo. This paper treats that as an operating rule: a library is taken off the shelf, every signal declared as one family before testing, and the family judged whole.
2. Method
A family is a batch of related tests counted together. Each added test is another chance to get lucky, so the required evidence rises with the count, and a signal and its negation count as two. Significance uses Newey-West t-statistics, so that autocorrelated returns do not overstate a result. Three bars follow: a naive threshold; a promotion bar deflated in the Bonferroni and Taleb spirit, dividing the allowed false-positive rate by the number of tests; and a margin-of-safety re-audit halving gross return and doubling costs.
3. Results
The closest signals are, without exception, measures of reversal over short horizons and of overbought levels: days since the high (IMAX20) is the strongest, then RSV20, RANK20 and MIN5. Their gross edge is real and survives the doubled costs by a clear margin (Table 3). What stops them is the deflated bar a family this size must clear. The volatility features collapse hardest under the same stress: compensation for bearing risk, not a mispricing.
| Bar | Threshold | Clearing | Outcome |
|---|---|---|---|
| Naive significance | t ≥ 2.5 | 29 of 158 | none promoted |
| Deflated promotion (Bonferroni/Taleb) | t ≥ 4.11 | 1 of 158 (IMAX20, t = 4.46) | none promoted |
| Deflated + stressed (gross halved, costs doubled) | 1 of 158 (IMAX20, t = 4.46) | none promoted |
| Signal | NW t | Stressed t | Outcome |
|---|---|---|---|
| IMAX20 (days since the high over 20 days) | 4.46 | 3.68 | not promoted |
| Measure | Value |
|---|---|
| Sample | roughly 21 years of US large caps |
| Hypotheses declared before testing | 316 (158 features, two sign choices) |
| Conventional significance bar | t = 2 |
| Base bar the margin-of-safety rerun must clear | t = 2.50 |
| Stressed t, strongest signals (IMAX20, RSV20, RANK20, MIN5) | 2.8 to 3.7 (2.81 to 3.68) |
| Promotions | zero |
4. Verdict
Read together, the survivors describe the zoo rather than escape it. Fast signals reverting from overbought levels keep their profit inside the transaction costs, and volatility exposures are paid for risk, not insight. The published technical zoo is reversal over short horizons and a volatility risk premium in disguise. The null is recorded as a null, and nothing here is tradable.
5. Limitations
Two, both cutting the same way. The universe holds current index members only, so the naive count and the survivor t-statistics are inflated and read as upper bounds; the delisted names would not raise them. And one library is not the whole zoo: a widely used set fails an honest bar, which is no proof that every published factor does.
References
- Harvey, C. R., and Liu, Y. (2019). A Census of the Factor Zoo. SSRN Working Paper 3341728.
- Harvey, C. R., Liu, Y., and Zhu, H. (2016). …and the Cross-Section of Expected Returns. Review of Financial Studies, 29(1), 5–68.
Provenance
- Asserts about
- The Alpha158 excavation: a single registered hypothesis covering all 158 alphas as one
declared family, judged on 25 July 2026 and rejected. Its record in the
pipeline’s database of every idea tested here is
d57232a184d38ec1. That identifier is the one to quote if you want to audit this paper against the pipeline, because it is what that database answers to. The 158 features were reimplemented in this system’s own point-in-time registry, in which each feature value uses only information available on its date, from Qlib’s public definitions. Qlib’s data layer and backtester were not used. - Data sources
- Adjusted open, high, low and close prices with volume (OHLCV) and dollar volume (Yahoo), roughly 425 current S&P constituents, 2005 to July 2026, one trading day of information lag. Ken French’s Fama-French five-factor model (FF5) plus momentum was in place for the pipeline’s later residual-alpha stage, the check of return left over after controlling for known factors, which no member reached. Both sources are free; neither is snapshotted.
- Statistical hurdles
- Regime G2, in force 12 to 25 July 2026: deflated significance, a deflated bootstrap resampling check, a 30% limit on degradation from in-sample to out-of-sample, that is, from the data the tests were fit on to data they never saw, and the rerun for margin of safety. The verdict was rendered 25 July 2026. Two pipeline code changes dated 26 July, a guard for small samples and a fix to the risk-free treatment at the pipeline’s factor-model stage, postdate this verdict and neither moves it.
- Family
- 316, declared before testing as 158 features by two sign choices, giving a deflated bar of t ≥ 4.11. The database of tested ideas records a family as a count, not as a membership; the 158 names survive only in the run artifact, not in the hypothesis record.
- Binding hurdle
- Out-of-sample decay, not multiplicity. The strongest member, IMAX20, cleared the deflated bar at t = 4.46 and cleared the rerun for margin of safety at a stressed t = 3.68 against a base bar of 2.50. What stopped it was a 55% degradation in Sharpe ratio, return per unit of risk, from in-sample to out-of-sample against a 30% limit. Rejection reasons for each member were not recorded by the pipeline’s run script, so this identification covers the strongest member only.
- Retracts if
- a walk-forward re-audit, retesting on a series of rolling windows rather than one split, shows that the 55% degradation is an artifact of the single fixed 70/30 split of the data into fit and evaluation periods, which would carry IMAX20 to the factor-model stage and make the headline one of 158 rather than zero;
- point-in-time constituents replace this universe of current members and the naive count rises above 29 instead of falling, contradicting Section 5;
- the reimplemented definitions are shown to diverge materially from Qlib’s published Alpha158 definitions, in which case the family tested is not the family named.
- Reproducibility
- Not pinned. Prices are refetched nightly and the universe is resolved from the current index list, so neither is frozen at publication, and the run script recorded no data window. A re-run on 27 July 2026 returned t = 4.25 and a stressed t = 3.56 for IMAX20 against the 4.46 and 3.68 reported above. The verdict and the binding gate were unchanged.