Who Has To Trade Tomorrow
In one line: I have put 183 trading ideas through the same gates and two came out alive. The 181 that died include three entire published factor libraries, and the survivors share one thing the casualties do not: somebody is forced to trade against them on a known schedule.
1. The count
Three hundred and three ideas filed. One hundred and eighty-three actually tested. Two survived.
The gap between 303 and 183 is not a rounding error and I will not hide it. One hundred and seventeen tickets errored before they produced a verdict, almost all of them blocked on data I could not obtain for free, and three are still queued. A ticket that never ran tested nothing, so it earns no place in the denominator and buys no multiplicity. The honest pass rate is two of 183, or 1.1%.
That number is the point of this paper, but not in the way it first reads. A 1.1% pass rate is not a claim about how hard I work. It is a claim about how much of what is publicly described as an edge does not survive being charged for the number of times someone looked.
2. Three libraries, no survivors
The deaths are not evenly distributed. They cluster, and the clusters are informative.
| Lane | Ideas judged | Survived |
|---|---|---|
| WorldQuant BRAIN candidates | 113 | 0 |
| Microsoft Qlib Alpha158, tested as one family | 158 alphas | 0 |
| Sector and industry grid | 12 | 0 |
| Neural forecasters | 3 | 0 |
| CFTC futures positioning | 4 | 0 |
The Alpha158 result is the cleanest of these and the one I would point a skeptic at first. Microsoft publishes a library of 158 technical alphas. I ran the whole thing as a single family, which means the significance bar rises with the number of attempts rather than staying at the level appropriate for one. Twenty-nine of the 158 clear a naive t of 2.5. Zero clear the deflated bar of 4.1, and zero survive the cost stress on top of it. What the library actually contains, once decomposed, is short-horizon reversal and a volatility risk premium wearing 158 different costumes. The gross edge is real. Doubling the costs takes all of it.
That is the general shape. These lanes do not fail because the ideas are stupid. They fail because they are all measuring the same two or three things from price, everyone has them, and the honest bar for the best of many correlated tries is much higher than the bar for one.
3. The two that lived
A dealer gamma classifier. Options dealers who are long gamma hedge against the move and compress the day's range. Dealers who are short gamma chase it and expand the range. The signal reads where dealer positioning sits and predicts the next day's range, not its direction.
A bubble screen. Names that double inside two years go on to crash more often than a same-month control, at a rate that clears the bar but only just, and I have published exactly how narrowly.
Neither is a price factor. Neither was clever. Both are decades old and publicly documented, which is worth sitting with, because the 129 price-derived candidates that died included many that were far more sophisticated.
4. The question that separates them
The distinction I want to draw is not positioning against price. I believed that for a while and this month proved it too loose.
CFTC Commitments of Traders is positioning data. It is the canonical positioning dataset, published weekly since the 1980s, and I tested it three ways this week against a bar of 2.87. The best look reached a t of 1.24, with the sign pointing the wrong way. Positioning data, dead on arrival.
So what do the two survivors have that COT does not?
Somebody is forced. An options dealer who is short gamma does not get to decide whether to hedge. The hedge is a mechanical consequence of the book he already holds, it happens on a schedule the market can see, and it moves the very price it is responding to. That is a forcing function.
A speculator with a large net long position in S&P futures is not forced to do anything. He has an opinion. He may hold it for a year. Commitments of Traders is an opinion poll with a lag, and an opinion poll is not a mechanism.
The bubble screen sits between the two and survives on the weaker version of the same logic: a crowd that has doubled its money in two years is not mechanically forced to sell, but the position is fragile in a way a quiet position is not, and the unwind when it comes is not optional.
So the operative question is not what the crowd thinks. It is who has to trade tomorrow whether they want to or not.
5. What this predicts
A rule that only explains the past is decoration. This one makes forward claims, and they are the reason to write it down.
It predicts that mechanisms with a compelled counterparty and a visible schedule will keep clearing the bar at a higher rate than anything derived from price alone: dealer hedging, index reconstitution, tax-driven selling, redemption-driven liquidation, margin. It predicts that sentiment and survey data will keep failing, because holding an opinion imposes no deadline. And it predicts that adding a 159th technical alpha to a library of 158 does nothing, because the family bar rises faster than the new candidate's evidence.
It would be falsified by a survey or sentiment signal that clears a properly deflated bar with a cost stress applied, or by a forced-flow mechanism that fails repeatedly under those same conditions. I would like to know either way, and both are cheap to test.
6. What is wrong with this paper
The universe is survivorship-biased. It is today's large caps, so names that died are absent from every test above, and raw statistics are upper bounds rather than estimates.
Two survivors is a sample of two. Everything in section 4 is a pattern read off two successes and 181 failures, and patterns read off two successes are exactly the kind of thing this engine exists to be suspicious of. I am describing a selection rule for what to test next, not a finding I would defend as established.
The lanes were not chosen at random. I tested BRAIN and Alpha158 because they were available, not because they were representative of price factors in general, and a different sample of libraries could give a different rate.
And the survivors have not been re-tested on data that did not exist when they passed. That is the test that matters most and it is scheduled, not done.
Method. All counts read from the engine's hypothesis store on 3 August 2026 via the read-only query layer. Judged means a ticket reached a verdict of passed, rejected or decayed; errored and queued tickets are excluded from every rate. Gates are a Newey-West t against a Bonferroni-deflated family bar, a ten thousand draw block bootstrap, an in-sample to out-of-sample degradation limit, and a margin-of-safety rerun with the gross halved and the costs doubled. Every statistic is computed by deterministic code and none by a language model.
Educational research only. Not investment advice.