The test
Value-at-Risk makes a promise: on 99 days out of 100, tomorrow’s loss will not exceed a stated threshold. The promise is worth what its backtest certifies. I graded three standard one-day models on SPY, the S&P 500 exchange-traded fund, strictly out of sample.
| Item | Value |
|---|---|
| Instrument | SPY, the S&P 500 ETF |
| Evaluation sample | December 2006 to July 2026 |
| Forecasts graded, one day each | 4,922 |
| Rolling estimation window | 500 days, about two years |
| Tails estimated | 1% and 5% |
| Historical Simulation | empirical quantile of the window |
| Gaussian | normal; window mean and standard deviation |
| EWMA (RiskMetrics) | normal, mean zero; variance decays at λ = 0.94 |
| Coverage tests | Kupiec, unconditional; Christoffersen, conditional |
| Basel traffic light | trailing 250 days; 99% measure only |
| Gaussian at 99%, full sample | 145 breaches against 49 budgeted, a 2.95% rate |
| 2008 GFC at 99%, 2.5 expected | Gaussian 28, Historical Simulation 20, EWMA 6 |
| 2020 COVID at 99%, 0.6 expected | Gaussian 12, Historical Simulation 10, EWMA 6 |
Coverage over the full sample
All three models reject Kupiec’s test of unconditional coverage at 99% and at 95%. Every one under-covers the tail, and the Gaussian is worst, taking roughly three times the breaches it budgeted for. Christoffersen’s test of conditional coverage rejects all three, so breaches cluster rather than scatter.
| Method | Conf. | Expected | Observed | Rate | Kupiec p | Christoffersen p | Basel (last 250) |
|---|---|---|---|---|---|---|---|
| Historical Sim | 99% | 49.2 | 86 | 1.75% | 0.0000 | 0.0000 | green |
| Gaussian | 99% | 49.2 | 145 | 2.95% | 0.0000 | 0.0000 | green |
| EWMA (RiskMetrics) | 99% | 49.2 | 119 | 2.42% | 0.0000 | 0.0000 | green |
| Historical Sim | 95% | 246.1 | 278 | 5.65% | 0.0408 | 0.0000 | |
| Gaussian | 95% | 246.1 | 285 | 5.79% | 0.0130 | 0.0000 | |
| EWMA (RiskMetrics) | 95% | 246.1 | 300 | 6.10% | 0.0006 | 0.0027 |
Inside the crises
Restricting the 99% breaches to two crisis windows reverses the ranking. A trailing window is slow to register that conditions have changed, and a 99% threshold built on the prior year’s calm is fiction once the calm breaks. EWMA, whose variance reacts within days, is the exception. COVID is the harder case, since no model forecasting one day ahead anticipates an overnight gap.
| Window | Days | Method | Expected | Observed |
|---|---|---|---|---|
| 2008 GFC | 253 | Historical Sim | 2.53 | 20 |
| 2008 GFC | 253 | Gaussian | 2.53 | 28 |
| 2008 GFC | 253 | EWMA (RiskMetrics) | 2.53 | 6 |
| 2020 COVID (Feb–Apr) | 62 | Historical Sim | 0.62 | 10 |
| 2020 COVID (Feb–Apr) | 62 | Gaussian | 0.62 | 12 |
| 2020 COVID (Feb–Apr) | 62 | EWMA (RiskMetrics) | 0.62 | 6 |
The regulatory blind spot
Basel’s traffic light reads only the most recent trading year. That window is quiet, so all three models currently show green. Over twenty years the same three under-cover their tails, and two of them come apart precisely when coverage matters most. A green light describes the last twelve months. It says nothing about the model.
Verdict
None of the three covered its tail. EWMA holds up best through the crises, and it is not the model the traffic light rewards. None of these results would pass the statistical hurdles of the research pipeline behind this site. They are kept as a documented failure, because most Value-at-Risk figures in circulation are quoted and never backtested. A statistic that fails a visible test is more informative than one that passes a test nobody ran.
Limitations
Three, stated plainly.
| Limitation | What it means for the result |
|---|---|
| One instrument | SPY alone, a single deeply liquid index with continuous history, is the easy case for VaR. A portfolio of individual stocks, some since delisted, is strictly harder and its VaR would look worse. The wider individual-stock database behind this site is survivorship biased, over-representing the companies that survived; testing the index itself avoids that. |
| The normal tail | Both parametric models assume a bell curve, which understates fat tails even when the volatility estimate is correct. This accounts for most of the Gaussian and EWMA under-coverage, and is a modelling choice rather than an artefact of this sample. |
| The window length | 500 days is a choice. A shorter window reacts faster to a regime change but estimates the 1% quantile from fewer tail points, and no single window length wins in both calm and crisis. |
References
- Basel Committee on Banking Supervision (1996). Supervisory Framework for the Use of “Backtesting” in Conjunction with the Internal Models Approach to Market Risk Capital Requirements. Bank for International Settlements.
- Christoffersen, P. F. (1998). Evaluating Interval Forecasts. International Economic Review, 39(4), 841–862.
- J.P. Morgan/Reuters (1996). RiskMetrics Technical Document, 4th ed. New York.
- Kupiec, P. H. (1995). Techniques for Verifying the Accuracy of Risk Measurement Models. Journal of Derivatives, 3(2), 73–84.