The test

Value-at-Risk makes a promise: on 99 days out of 100, tomorrow’s loss will not exceed a stated threshold. The promise is worth what its backtest certifies. I graded three standard one-day models on SPY, the S&P 500 exchange-traded fund, strictly out of sample.

Design. Instrument, models, tests.
ItemValue
InstrumentSPY, the S&P 500 ETF
Evaluation sampleDecember 2006 to July 2026
Forecasts graded, one day each4,922
Rolling estimation window500 days, about two years
Tails estimated1% and 5%
Historical Simulationempirical quantile of the window
Gaussiannormal; window mean and standard deviation
EWMA (RiskMetrics)normal, mean zero; variance decays at λ = 0.94
Coverage testsKupiec, unconditional; Christoffersen, conditional
Basel traffic lighttrailing 250 days; 99% measure only
Gaussian at 99%, full sample145 breaches against 49 budgeted, a 2.95% rate
2008 GFC at 99%, 2.5 expectedGaussian 28, Historical Simulation 20, EWMA 6
2020 COVID at 99%, 0.6 expectedGaussian 12, Historical Simulation 10, EWMA 6

Coverage over the full sample

All three models reject Kupiec’s test of unconditional coverage at 99% and at 95%. Every one under-covers the tail, and the Gaussian is worst, taking roughly three times the breaches it budgeted for. Christoffersen’s test of conditional coverage rejects all three, so breaches cluster rather than scatter.

Table 1. Coverage tests over the full evaluation sample of 4,922 forecasts, one day each.
MethodConf.ExpectedObservedRateKupiec pChristoffersen pBasel (last 250)
Historical Sim99%49.2861.75%0.00000.0000green
Gaussian99%49.21452.95%0.00000.0000green
EWMA (RiskMetrics)99%49.21192.42%0.00000.0000green
Historical Sim95%246.12785.65%0.04080.0000
Gaussian95%246.12855.79%0.01300.0000
EWMA (RiskMetrics)95%246.13006.10%0.00060.0027
Kupiec and Christoffersen p-values below 0.05 reject the model. A reported p of 0.0000 is the code’s printout to four decimals of a value too small to display, not an exact zero. The Basel traffic light is defined only for the 99% test.

Inside the crises

Restricting the 99% breaches to two crisis windows reverses the ranking. A trailing window is slow to register that conditions have changed, and a 99% threshold built on the prior year’s calm is fiction once the calm breaks. EWMA, whose variance reacts within days, is the exception. COVID is the harder case, since no model forecasting one day ahead anticipates an overnight gap.

Table 2. 99% exceptions within two crisis windows (GFC: global financial crisis).
WindowDaysMethodExpectedObserved
2008 GFC253Historical Sim2.5320
2008 GFC253Gaussian2.5328
2008 GFC253EWMA (RiskMetrics)2.536
2020 COVID (Feb–Apr)62Historical Sim0.6210
2020 COVID (Feb–Apr)62Gaussian0.6212
2020 COVID (Feb–Apr)62EWMA (RiskMetrics)0.626

The regulatory blind spot

Basel’s traffic light reads only the most recent trading year. That window is quiet, so all three models currently show green. Over twenty years the same three under-cover their tails, and two of them come apart precisely when coverage matters most. A green light describes the last twelve months. It says nothing about the model.

Verdict

None of the three covered its tail. EWMA holds up best through the crises, and it is not the model the traffic light rewards. None of these results would pass the statistical hurdles of the research pipeline behind this site. They are kept as a documented failure, because most Value-at-Risk figures in circulation are quoted and never backtested. A statistic that fails a visible test is more informative than one that passes a test nobody ran.

Limitations

Three, stated plainly.

Limitations. What this study does not establish.
LimitationWhat it means for the result
One instrumentSPY alone, a single deeply liquid index with continuous history, is the easy case for VaR. A portfolio of individual stocks, some since delisted, is strictly harder and its VaR would look worse. The wider individual-stock database behind this site is survivorship biased, over-representing the companies that survived; testing the index itself avoids that.
The normal tailBoth parametric models assume a bell curve, which understates fat tails even when the volatility estimate is correct. This accounts for most of the Gaussian and EWMA under-coverage, and is a modelling choice rather than an artefact of this sample.
The window length500 days is a choice. A shorter window reacts faster to a regime change but estimates the 1% quantile from fewer tail points, and no single window length wins in both calm and crisis.

References

  1. Basel Committee on Banking Supervision (1996). Supervisory Framework for the Use of “Backtesting” in Conjunction with the Internal Models Approach to Market Risk Capital Requirements. Bank for International Settlements.
  2. Christoffersen, P. F. (1998). Evaluating Interval Forecasts. International Economic Review, 39(4), 841–862.
  3. J.P. Morgan/Reuters (1996). RiskMetrics Technical Document, 4th ed. New York.
  4. Kupiec, P. H. (1995). Techniques for Verifying the Accuracy of Risk Measurement Models. Journal of Derivatives, 3(2), 73–84.