The Light Was Green: Out-of-Sample Backtesting of Three
One-Day Value-at-Risk Models on Twenty Years of the S&P 500
In one line: I tested three standard risk models over twenty years. All understated crash losses. The official safety check still says green.
Abstract. I backtest three standard one-day Value-at-Risk models (Historical Simulation, a Gaussian parametric model, and the RiskMetrics EWMA) on 4,922 out-of-sample forecasts for the S&P 500 ETF (SPY), December 2006 to July 2026. All three under-cover the 1% and 5% tails and reject Kupiec’s unconditional-coverage test at both confidence levels; the Gaussian is worst, realizing a 2.95% exception rate against a 1% budget (145 breaches where 49.2 were expected). Christoffersen’s conditional-coverage test rejects all three. The breaches cluster rather than scatter. Restricting to crisis windows reverses the full-sample ranking. EWMA, whose variance reacts within days, took 6 breaches in the 2008 window against 28 for the Gaussian and 20 for Historical Simulation. Because the Basel traffic light reads only the trailing 250 days, all three currently show green. The models’ honest failure, rather than any single winner, is the result.
1. Introduction
Value-at-Risk states a promise with a number attached: on 99 days out of 100, tomorrow’s loss will not exceed a stated threshold. A risk number is worth exactly what its backtest certifies, so any VaR must be graded before it is trusted. This paper backtests three standard one-day models on SPY, the S&P 500 ETF, using twenty years of daily returns and grading them strictly out of sample. Every VaR quoted for a day is estimated only from returns before that day, then checked against a realized return the model never observed.
2. Data and method
Each model estimates the 1% and 5% tail of the next day’s return and is re-estimated daily on a rolling 500-day window (approximately two years). Historical Simulation reads the empirical quantile directly off the window. The Gaussian model assumes normality and uses the window’s mean and standard deviation. EWMA follows the RiskMetrics specification, a zero-mean normal whose variance decays at λ = 0.94 and therefore discounts old calm quickly. All three are graded on an identical evaluation sample of 4,922 one-day forecasts from December 2006 to July 2026, so the comparison is like for like. An exception is a day whose realized loss breached the VaR. Coverage is assessed by Kupiec’s unconditional test (is the exception rate correct?) and Christoffersen’s conditional-coverage test (is the rate correct, and are the exceptions independent or do they cluster?). Deterministic code computes every figure reported below.
3. Results
3.1 Full-sample coverage
| Method | Conf. | Expected | Observed | Rate | Kupiec p | Christoffersen p | Basel (last 250) |
|---|---|---|---|---|---|---|---|
| Historical Sim | 99% | 49.2 | 86 | 1.75% | 0.0000 | 0.0000 | green |
| Gaussian | 99% | 49.2 | 145 | 2.95% | 0.0000 | 0.0000 | green |
| EWMA (RiskMetrics) | 99% | 49.2 | 119 | 2.42% | 0.0000 | 0.0000 | green |
| Historical Sim | 95% | 246.1 | 278 | 5.65% | 0.0408 | 0.0000 | — |
| Gaussian | 95% | 246.1 | 285 | 5.79% | 0.0130 | 0.0000 | — |
| EWMA (RiskMetrics) | 95% | 246.1 | 300 | 6.10% | 0.0006 | 0.0027 | — |
All three models reject Kupiec’s unconditional-coverage test at both 99% and 95%. Every one under-covers the tail. The Gaussian is worst. It budgeted for 49 breaches at 99% and realized 145, a 2.95% exception rate against a promised 1%, roughly three times the losses it claimed to insure against. Historical Simulation is the least severe on the full sample at 1.75%, and EWMA lies between at 2.42%. Christoffersen’s conditional-coverage test is rejected for all three, so the breaches are not evenly scattered but arrive in clusters. A VaR that fails in bursts is worse than one that fails at random, because the bursts coincide with the conditions the measure exists to cover.
3.2 Crisis-window coverage
The full-sample rate conceals where the misses occur. Restricting the 99% breaches to two crisis windows sharpens the picture.
| Window | Days | Method | Expected | Observed |
|---|---|---|---|---|
| 2008 GFC | 253 | Historical Sim | 2.53 | 20 |
| 2008 GFC | 253 | Gaussian | 2.53 | 28 |
| 2008 GFC | 253 | EWMA (RiskMetrics) | 2.53 | 6 |
| 2020 COVID (Feb–Apr) | 62 | Historical Sim | 0.62 | 10 |
| 2020 COVID (Feb–Apr) | 62 | Gaussian | 0.62 | 12 |
| 2020 COVID (Feb–Apr) | 62 | EWMA (RiskMetrics) | 0.62 | 6 |
The ranking now reverses. In 2008 the Gaussian took 28 breaches against 2.5 expected and Historical Simulation took 20, both because a trailing window is slow to register that conditions have changed, and a 99% number built on the prior year’s calm is fiction once the calm breaks. EWMA, whose variance reacts within days, took 6. The 2020 COVID crash is the harder case. Even EWMA took 6 against 0.6 expected, because no one-day model anticipates an overnight gap, while the window methods took 10 and 12. The model with the best full-sample rate, Historical Simulation, is not the model to hold in a crisis.
4. The regulatory blind spot
The regulatory reading diverges from the twenty-year record. Basel’s traffic light examines only the most recent 250 days. That window is currently quiet, so all three models show green, a clean bill of health from the test a supervisor reads first. Over twenty years the same three models under-cover their tails, and the 2008 column shows two of them coming apart precisely when coverage matters most. A green light describes the last twelve months; it says nothing about the model. Only a backtest spanning the crises reveals what a position actually carries.
5. Discussion
No one-day VaR examined here covered its tail. All three reject Kupiec at 99% and 95%, and the Gaussian fails worst. EWMA holds up best through the crisis windows, and it is not the model the traffic light currently rewards. None of these results would earn a promotion under the pipeline’s standards; they are retained as a documented, honest failure.
6. Limitations
Three, stated plainly. One instrument. The study covers SPY alone, a single deeply liquid index with continuous history, which is the easy case for VaR. A book of individual names, some since delisted, is strictly harder and its VaR would look worse; the wider price panel is survivorship-biased, and this study avoids that only by testing the index itself. The normal tail. Both parametric models assume a bell curve, which understates fat tails even when the volatility estimate is correct. This accounts for most of the Gaussian and EWMA under-coverage and is a modelling choice rather than an artifact of this sample. The window length. Five hundred days is a choice; a shorter window reacts faster to a regime change but estimates the 1% quantile from fewer tail points, and no single window length wins in both calm and crisis.
References
- Basel Committee on Banking Supervision (1996). Supervisory Framework for the Use of “Backtesting” in Conjunction with the Internal Models Approach to Market Risk Capital Requirements. Bank for International Settlements.
- Christoffersen, P. F. (1998). Evaluating Interval Forecasts. International Economic Review, 39(4), 841–862.
- J.P. Morgan/Reuters (1996). RiskMetrics—Technical Document, 4th ed. New York.
- Kupiec, P. H. (1995). Techniques for Verifying the Accuracy of Risk Measurement Models. Journal of Derivatives, 3(2), 73–84.