Abstract. Three revisions of a model that forecasts volatility from options dealer positioning, and the gates that retired the first. The signal is real. The state chain adds nothing over a competent ensemble.
1. The claim
A signal that beats implied volatility need not beat a competent econometric ensemble. A process that cannot retire its founder’s favourite model is not a process.
| Version | Claim | Statistic | Outcome |
|---|---|---|---|
| v1 (state chain) | Nine states of dealer positioning × realized vol forecast the range of the next day | t = 4.34 vs implied vol; lost to HAR | Retired after gate tightening (stressed t = 2.17) |
| v1.5 (neural emissions) | An MLP on the same states beats the linear ladder | t = −4.37 (walk-forward, an ensemble of 3 seeds) | Buried; capacity is not information at daily frequency |
| v3 (residual corrector, pre-registered) | Shrunk state corrections improve a log HAR+VIX+GJR backbone at a horizon of 5 days | t = −1.07 | Buried, name and all |
2. The error
Version 1 entered its nine states as least squares dummies, level shifts that cannot repair where a baseline errs, and it fought the naive benchmark rather than the incumbent structure in volatility: persistence (HAR), leverage asymmetry (GJR-GARCH) and implied volatility. Beating the benchmark while losing to HAR says the information was old under a new label, as with anomalies that die under a Fama-French plus momentum decomposition.
3. The gates
Two promotion criteria were added midway, proposed by the pipeline’s independent critic and approved by the human owner; thresholds are never changed by machine. The first deflates the significance bar by the number of variants already tried; the second re-tests every survivor at half its edge and double its costs. Gates that spare the past are decorative, so the vault was re-audited. The tulip mania screen’s counting unit was wrong, since one mania is counted many times; on the honest count it sits on top of its threshold rather than above it. The namesake model was retired, without exception.
4. Version 3
Version 3 inverted the architecture: the strictest baseline became the backbone, and the states were left one job, correcting its residuals. Pre-registration is more than hygiene under the deflator: every extra variant raises the bar for the whole group. The corrections did not improve the backbone. A state-level result read off after a failed gate is a hypothesis for a future mechanism, not a specification to refit, and is logged as an open question.
| Measure | Value |
|---|---|
| Dealer gamma exposure (GEX) classifier vs implied volatility, standing in the vault | t = 7.4; stressed re-audit 3.68 |
| v1 vs implied volatility | t = 4.34 |
| v1 under the margin-of-safety rerun | stressed t = 2.17; retired |
| v1.5, neural emissions | t = −4.37 |
| v3 residual corrector, decisive pass/fail test | t = −1.07 |
| v3 forecast error vs backbone, mean absolute error (MAE) | 34.1 vs 33.7 bps |
| v3 decomposition, state by state | one high-gamma state +3.5 bps; one stressed state −4.5 bps |
| v3 backbone | log HAR plus implied plus GJR, refit monthly by expanding ordinary least squares (OLS) |
| v3 shrinkage toward zero, James-Stein flavour | k = 60, across three conditionings: the nine GEX×realized volatility (RV) cells, the sign of the Cboe Volatility Index (VIX) term structure, a flag for state arrival |
| v3 target horizon and leakage guards | range over 5 days; training truncated at the refit boundary minus the horizon; heteroskedasticity and autocorrelation consistent (HAC) lags ≥ 10; Duan smearing applied symmetrically to both models |
| Multiple-testing deflator | Bonferroni: the required significance level is divided by the number of variants of the idea already tried, tracked in the database of every idea tested here, so an idea tried thirty ways is judged as one draw of thirty |
| Margin-of-safety rerun | mean halved against the original HAC standard error, costs doubled; halving the whole return series instead leaves t unchanged, since t does not move with scale |
| Tulip mania screen, counts by name and month | z = 5.3 |
| Same screen, one count per episode | 1.28 against its own 1.25 bar; 1.18 on a different data pull |
| Bar a version 4 would face under the deflator | near 3.1 |
5. What stands
The classifier stays in the vault; the chain does not. Dealer positioning carries information relative to implied volatility alone, but none over an ensemble of persistence, leverage and implied volatility at daily or weekly horizons, under criteria that require the edge to survive at half strength. A version 4 is warranted only if it brings a mechanism the ensemble cannot carry: transition rows conditioned on FOMC meetings and CPI releases are the standing candidate.
References
- Carhart, M. M. (1997). On Persistence in Mutual Fund Performance. Journal of Finance, 52(1), 57–82.
- Corsi, F. (2009). A Simple Approximate Long-Memory Model of Realized Volatility. Journal of Financial Econometrics, 7(2), 174–196.
- Duan, N. (1983). Smearing Estimate: A Nonparametric Retransformation Method. Journal of the American Statistical Association, 78(383), 605–610.
- Fama, E. F., and French, K. R. (1993). Common Risk Factors in the Returns on Stocks and Bonds. Journal of Financial Economics, 33(1), 3–56.
- Glosten, L. R., Jagannathan, R., and Runkle, D. E. (1993). On the Relation between the Expected Value and the Volatility of the Nominal Excess Return on Stocks. Journal of Finance, 48(5), 1779–1801.
- James, W., and Stein, C. (1961). Estimation with Quadratic Loss. Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, 1, 361–379.