1. The count
| The record, as of 4 August 2026 | Count |
|---|---|
| Ideas filed | 316 |
| Tested to a verdict | 184 |
| Survived | 2 |
| Died | 182 |
| Honest pass rate, 2 of 184 | 1.1% |
| Judged under this museum's own gates | 69 |
| Judged inside a vendor's simulator, and labelled as such | 115 |
| Errored before producing a verdict; roughly three-quarters blocked on vendor data available only for a fee, most of the rest plain compute timeouts. Excluded from every rate. | 128 |
| Still queued, also excluded | 4 |
| Judged as of 18 August 2026 | 272 |
| Judged as of 19 August 2026 | 277 |
| New trading survivors since 4 August | 0 |
2. Three libraries, no survivors
| Lane | Ideas judged | Survived |
|---|---|---|
| WorldQuant BRAIN candidates | 114 | 0 |
| Microsoft Qlib Alpha158, tested as one family | 158 alphas | 0 |
| Sector and industry grid | 32 | 0 |
| Neural forecasters | 3 | 0 |
| CFTC futures positioning | 3 | 0 |
Microsoft's alphas are variants of one idea, judged together, because many variants mean many chances at luck and the required evidence rises with each attempt.
| Microsoft Qlib Alpha158, judged as one family | Record |
|---|---|
| Technical alphas in the library, meaning rule-based trading signals | 158 |
| Clear a naive t-statistic of 2.5, t being the standard measure of whether a result is distinguishable from luck | 29 |
| Survive the stress that halves the gross return and doubles the costs | 16 of 29 |
| Best alpha, days since the high of the last 20 days | t = 4.46, and 3.68 stressed |
| Its Sharpe ratio, return per unit of risk, from the data used to build it to data it never saw | a 55% collapse, against a 30% limit |
| Promoted | 0 |
| What the library decomposes to | reversal at short horizons and a volatility risk premium, expressed 158 ways |
| What kills it | out-of-sample decay and the rising evidence bar, not costs |
3. The two that lived
The gamma signal predicts the next day's range, not its direction, and it survives the practitioner's first objection, that this is volatility wearing a costume, but only as an increment to VIX. The bubble screen clears its gate only just, on the honest denominator. Neither survivor is a price factor and neither was clever.
| The two survivors | What the record says |
|---|---|
| Dealer gamma classifier, mechanism decades old, published signal dated 2018 | Dealers who are long gamma hedge against the move and compress the day's range. Dealers who are short gamma chase it and expand the range. |
| Range, bottom gamma quintile against top | 2.59 times |
| Pin rate, top quintile against unconditional | 79.7% against 61%, a lift of 1.31 rather than the naked 80% a reader might price |
| Standing alone | t of 11.14 across 3,565 trading days |
| Independent windows | The regressor is a 252-day rolling percentile, so nearer fourteen than 3,565; the t should be read accordingly |
| Correlation with VIX | −0.47 |
| Controlling for VIX | t falls to 7.36, not to nothing; VIX itself carries t = 12.3, so the control dominates and gamma is the increment |
| Stressed | 3.68. This lane has no cost model, so stressed here means the effect halved against the unstressed standard error, not the gross-halved and costs-doubled rerun of Section 2. |
| Family size | Recorded as one, which is wrong. This shop has judged at least seven dealer-gamma looks, so the bar is t ≥ 3.13 rather than 2.50. The result clears the corrected bar, so the verdict stands, but as an unauditable pass rather than a clean one. |
| Bubble screen, mechanism decades old, published signal dated 2019 | Names that double inside two years crash more often than a control drawn from the same month. |
| Crash rate by name-month | 23.0% against a 15.1% base, a ratio of 1.51 |
| Crash rate by episode, one flag per run-up, the count the lab itself calls the honest N | 232 of 1,192, or 19.5%, against a 15.2% base: a ratio of 1.283 against a gate set at 1.25 |
| Stressed z | 5.34 |
| Episode-level binomial p | 3.7e-05, not the 2.7e-80 the name-month count produces. The gate was evaluated on name-months; the honest margin is the episode one, and it is thin. |
| Clustering | Run-ups arrive in cohorts, so effective independent observations are far below 1,189 and the p-value is optimistic by an amount I have not bounded |
4. The question that separates them
Positioning against price is too loose a distinction. I believed it, and Commitments of Traders died with the sign the wrong way. Somebody has to be forced. A short-gamma dealer does not decide whether to hedge: the hedge follows from the book he holds, on a schedule the market can see. A speculator long S&P futures holds an opinion, and an opinion poll with a lag is not a mechanism. Hayek got there first. The question is who has to trade tomorrow whether they want to or not.
| Positioning and forced-flow mechanisms rejected here | Verdict |
|---|---|
| Commitments of Traders, the Commodity Futures Trading Commission (CFTC) report of futures positioning and the canonical positioning dataset | 3 looks; best t of 1.24 against a required 2.87, sign the wrong way |
| The index effect | rejected 24 July |
| Tax-loss harvesting | rejected 13 July at a negative t |
| The tax-loss-selling bounce | rejected 24 July |
| Dividend-cut forced selling | rejected 28 July at the margin-of-safety rerun |
| Levered ETF end-of-day rebalancing | rejected 14 August, see Section 7 |
| Mechanisms Section 5 originally named as predictions | dealer hedging, index reconstitution, selling driven by taxes, liquidation driven by redemptions, margin calls |
5. What this predicts
Corrected 19 August 2026. This section once named forced-flow mechanisms that should clear the bar faster than price signals. Four of them had already been rejected here before publication, which makes them examples rather than predictions. Sentiment and survey data should keep failing, because an opinion imposes no deadline. A sentiment signal clearing a properly raised bar under cost stress would falsify that, and it is cheap to test.
6. What is wrong with this paper
The universe is survivorship biased, today's large caps, so raw statistics are upper bounds rather than estimates. Section 4 is a pattern read off two successes, a selection rule for what to test next rather than a finding. The lanes were not chosen at random. The survivors have not been re-tested on data that postdates their passing, the test that matters most; it is scheduled, not done.
7. Update, 19 August 2026
Sections 1 to 6 are a 4 August snapshot, and what followed is worse for the paper. Levered ETF rebalancing is the purest instance of Section 4's definition, mechanical and publicly scheduled, and it died a cost death rather than a null. Both arms of the differential prediction returned zero, and two zeroes distinguish nothing. Section 4 has survived nothing.
| Since the snapshot | Record |
|---|---|
| Judged between 4 and 18 August | 88 |
| WorldQuant BRAIN expressions, price signals killed inside that vendor's own post-simulation checks | 79 |
| Pairs-family rejections | 5 |
| Compelled-counterparty mechanisms tested, the thing Section 4 predicts should survive at a higher rate | 3, all rejected |
| Study-lane pass, 12 August: the museum's own forecast-calibration self-test, not a tradable signal, outside the vault count | 1 |
| Out-of-sample forced-flow arm | 3 tested, 0 passed |
| Out-of-sample price arm | 84 tested, 0 passed |
| Levered ETF rebalancing, ticket 2aa66023b9be729f, rejected 14 August | Newey-West t of −0.03, bootstrap P(Sharpe ≤ 0) of 0.53, in-sample to out-of-sample degradation of 701% |
| The same flow scaled by SEC-reported fund assets | t = 0.89, also dead |
| The critic's addendum of 14 August | A cost death, not a no-effect: gross return +7.5% a year at a diagnostic t of 2.02, destroyed by the margin-of-safety rerun to t = −3.14. The mechanism leaves a gross footprint and cannot be traded through this engine's cost assumption. Same verdict as Working Paper No. 25, reached from the opposite direction. |
| Rejection funnel, of 274 rejections | 199 carry the stage label brain-checks, which fires only after a completed vendor simulation, chiefly on low Sharpe, low fitness and self-correlation; 38 died at this engine's own validation stage; 37 are scattered. A census of one vendor lane under a foreign rulebook, not a transferable lesson about where to spend compute. |
| Comparison standard | On the house standard the arms would need adjudicating in the hundreds before the difference could speak at all |
| Method | What was done |
|---|---|
| Source | The pipeline's database of every idea tested here on 4 August 2026, read through the read-only query layer, except the sector and industry grid, whose 32 cells were judged inside its own lab artifact and never filed as database rows. |
| Judged | A test that reached a verdict of passed, rejected or decayed. Errored and queued tests are excluded from every rate. |
| Main automated lane | A Newey-West t-statistic, adjusted so that autocorrelation in returns does not overstate significance, against a bar raised by a Bonferroni correction for the number of variants of the idea; a ten thousand draw block bootstrap, resampling the return series in blocks to see how often chance alone matches the result; an in-sample to out-of-sample degradation limit; and a rerun for margin of safety with the gross return halved and the costs doubled. |
| Research lane | A subset, stated in each study's own paper. The studies of 13F crowding (the 13F is the Securities and Exchange Commission form on which large managers disclose US equity holdings each quarter) and of disclosure text were judged on a Newey-West t-statistic against a bar raised for the number of attempts, with no bootstrap, no out-of-sample split and no cost model. Claiming the full stack of hurdles applied to all 184 would be false. |
| Computation | Every statistic is computed by deterministic code and none by a language model. |
Revision, 19 August 2026. An adversarial review against the house lens files (Lopez de Prado, Goetzmann, Harvey, Simons) with every count re-read from the pipeline database found six defects that changed a claim, four of them introduced by the 18 August revision. All are corrected above rather than removed. (i) Section 7 claimed the post-4-August window was out-of-sample evidence for Section 4; it is not, and the levered-ETF rebalancing ticket rejected on 14 August is a direct falsification of the thesis, now reported as such. (ii) Section 7 read the rejection funnel backwards: brain-checks fires after a completed vendor simulation, not before, so those 199 are finished backtests rather than cheap screens. (iii) The count as of 18 August is 272, not 277; 277 is today's. (iv) The bubble screen's published 23.0% was the name-month rate; the episode rate the lab calls honest is 19.5%, ratio 1.283. (v) The gamma signal's recorded family size of one is wrong, the corrected bar at n = 7 is 3.13, and its stressed figure is not the cost-based rerun Section 2 describes. (vi) The deck's “judged under this museum's gates” was true of 69 of the 184, not all of them. Section 5's prediction list has also been corrected: four of the mechanisms it named had already been rejected here before publication. No survivor verdict changed.
Revision, 18 August 2026. Added Section 7 and the survivor statistics in Section 3. Every figure is read from the same database rows that carry each verdict, not restated from memory. No verdict changed and no 4 August count was altered; that snapshot stands as published.
Revision, 16 August 2026. An adversarial re-derivation of every count against the database corrected five things: the as-of date is 4 August, not 3 (the published counts are that day's snapshot); the sector grid row is its lab's 32 cells, not 12; the CFTC row is the ticket's 3 declared looks, not 4; the Alpha158 cost sentence now states the record (16 of 29 naive hits survive the doubled-cost stress; decay and the rising bar do the killing); and the casualty and survivor descriptions were tightened. No verdict changed.
Educational research only. Not investment advice.