Get in touchReach out on LinkedIn

The Light Was Green: Out-of-Sample Backtesting of Three
One-Day Value-at-Risk Models on Twenty Years of the S&P 500

BlueShip Research
Working Paper No. 10 · 25 July 2026

In one line: I tested three standard risk models over twenty years. All understated crash losses. The official safety check still says green.

Abstract. I backtest three standard one-day Value-at-Risk models (Historical Simulation, a Gaussian parametric model, and the RiskMetrics EWMA) on 4,922 out-of-sample forecasts for the S&P 500 ETF (SPY), December 2006 to July 2026. All three under-cover the 1% and 5% tails and reject Kupiec’s unconditional-coverage test at both confidence levels; the Gaussian is worst, realizing a 2.95% exception rate against a 1% budget (145 breaches where 49.2 were expected). Christoffersen’s conditional-coverage test rejects all three. The breaches cluster rather than scatter. Restricting to crisis windows reverses the full-sample ranking. EWMA, whose variance reacts within days, took 6 breaches in the 2008 window against 28 for the Gaussian and 20 for Historical Simulation. Because the Basel traffic light reads only the trailing 250 days, all three currently show green. The models’ honest failure, rather than any single winner, is the result.

Keywords: Value-at-Risk, backtesting, Kupiec test, Christoffersen conditional coverage, RiskMetrics EWMA, Basel traffic light, tail risk.

1. Introduction

Value-at-Risk states a promise with a number attached: on 99 days out of 100, tomorrow’s loss will not exceed a stated threshold. A risk number is worth exactly what its backtest certifies, so any VaR must be graded before it is trusted. This paper backtests three standard one-day models on SPY, the S&P 500 ETF, using twenty years of daily returns and grading them strictly out of sample. Every VaR quoted for a day is estimated only from returns before that day, then checked against a realized return the model never observed.

2. Data and method

Each model estimates the 1% and 5% tail of the next day’s return and is re-estimated daily on a rolling 500-day window (approximately two years). Historical Simulation reads the empirical quantile directly off the window. The Gaussian model assumes normality and uses the window’s mean and standard deviation. EWMA follows the RiskMetrics specification, a zero-mean normal whose variance decays at λ = 0.94 and therefore discounts old calm quickly. All three are graded on an identical evaluation sample of 4,922 one-day forecasts from December 2006 to July 2026, so the comparison is like for like. An exception is a day whose realized loss breached the VaR. Coverage is assessed by Kupiec’s unconditional test (is the exception rate correct?) and Christoffersen’s conditional-coverage test (is the rate correct, and are the exceptions independent or do they cluster?). Deterministic code computes every figure reported below.

3. Results

3.1 Full-sample coverage

Table 1. Coverage tests over the full evaluation sample of 4,922 one-day forecasts.
MethodConf.ExpectedObservedRateKupiec pChristoffersen pBasel (last 250)
Historical Sim99%49.2861.75%0.00000.0000green
Gaussian99%49.21452.95%0.00000.0000green
EWMA (RiskMetrics)99%49.21192.42%0.00000.0000green
Historical Sim95%246.12785.65%0.04080.0000
Gaussian95%246.12855.79%0.01300.0000
EWMA (RiskMetrics)95%246.13006.10%0.00060.0027
Kupiec and Christoffersen p-values below 0.05 reject the model. A reported p of 0.0000 is the code’s four-decimal printout of a value too small to display, not an exact zero. The Basel traffic light is defined only for the 99% test.

All three models reject Kupiec’s unconditional-coverage test at both 99% and 95%. Every one under-covers the tail. The Gaussian is worst. It budgeted for 49 breaches at 99% and realized 145, a 2.95% exception rate against a promised 1%, roughly three times the losses it claimed to insure against. Historical Simulation is the least severe on the full sample at 1.75%, and EWMA lies between at 2.42%. Christoffersen’s conditional-coverage test is rejected for all three, so the breaches are not evenly scattered but arrive in clusters. A VaR that fails in bursts is worse than one that fails at random, because the bursts coincide with the conditions the measure exists to cover.

3.2 Crisis-window coverage

The full-sample rate conceals where the misses occur. Restricting the 99% breaches to two crisis windows sharpens the picture.

Table 2. 99% exceptions within two crisis windows.
WindowDaysMethodExpectedObserved
2008 GFC253Historical Sim2.5320
2008 GFC253Gaussian2.5328
2008 GFC253EWMA (RiskMetrics)2.536
2020 COVID (Feb–Apr)62Historical Sim0.6210
2020 COVID (Feb–Apr)62Gaussian0.6212
2020 COVID (Feb–Apr)62EWMA (RiskMetrics)0.626

The ranking now reverses. In 2008 the Gaussian took 28 breaches against 2.5 expected and Historical Simulation took 20, both because a trailing window is slow to register that conditions have changed, and a 99% number built on the prior year’s calm is fiction once the calm breaks. EWMA, whose variance reacts within days, took 6. The 2020 COVID crash is the harder case. Even EWMA took 6 against 0.6 expected, because no one-day model anticipates an overnight gap, while the window methods took 10 and 12. The model with the best full-sample rate, Historical Simulation, is not the model to hold in a crisis.

4. The regulatory blind spot

The regulatory reading diverges from the twenty-year record. Basel’s traffic light examines only the most recent 250 days. That window is currently quiet, so all three models show green, a clean bill of health from the test a supervisor reads first. Over twenty years the same three models under-cover their tails, and the 2008 column shows two of them coming apart precisely when coverage matters most. A green light describes the last twelve months; it says nothing about the model. Only a backtest spanning the crises reveals what a position actually carries.

5. Discussion

No one-day VaR examined here covered its tail. All three reject Kupiec at 99% and 95%, and the Gaussian fails worst. EWMA holds up best through the crisis windows, and it is not the model the traffic light currently rewards. None of these results would earn a promotion under the pipeline’s standards; they are retained as a documented, honest failure.

6. Limitations

Three, stated plainly. One instrument. The study covers SPY alone, a single deeply liquid index with continuous history, which is the easy case for VaR. A book of individual names, some since delisted, is strictly harder and its VaR would look worse; the wider price panel is survivorship-biased, and this study avoids that only by testing the index itself. The normal tail. Both parametric models assume a bell curve, which understates fat tails even when the volatility estimate is correct. This accounts for most of the Gaussian and EWMA under-coverage and is a modelling choice rather than an artifact of this sample. The window length. Five hundred days is a choice; a shorter window reacts faster to a regime change but estimates the 1% quantile from fewer tail points, and no single window length wins in both calm and crisis.

This study is retained because most Value-at-Risk figures in circulation are quoted but never backtested: a number on a risk report with no scorecard behind it. A statistic that fails a visible test is more informative than one that passes a test nobody ran.

References

  1. Basel Committee on Banking Supervision (1996). Supervisory Framework for the Use of “Backtesting” in Conjunction with the Internal Models Approach to Market Risk Capital Requirements. Bank for International Settlements.
  2. Christoffersen, P. F. (1998). Evaluating Interval Forecasts. International Economic Review, 39(4), 841–862.
  3. J.P. Morgan/Reuters (1996). RiskMetrics—Technical Document, 4th ed. New York.
  4. Kupiec, P. H. (1995). Techniques for Verifying the Accuracy of Risk Measurement Models. Journal of Derivatives, 3(2), 73–84.