QuanterLab produced this study: it wasn’t written up afterwards. Registered hypothesis and search record in Appendix A2.

A note on AI. QuanterLab is a quantitative finance research platform, and every number in this study comes from a run on the platform. The hypothesis, the parameter choices, the validation design and the conclusions belong to the author. Runs execute on point-in-time data with walk-forward validation, and each study ships with its methodology and logs, so a reader can reconstruct the result instead of trusting it. I use AI to edit and structure the prose; it does not generate results, produce numbers, or decide what a study concludes.

As seen on Quantocracy
All research
Registered atlas · the regime shelf, audited

The Regime Atlas - ten ways to read the market, twenty years, one scoreboard

What this is. A registered atlas: the metric set below was frozen in a version-controlled registration (commit d63acc33) before the one and only compute ran. No arm is raced, no P-value is claimed, and nothing was re-run. Every figure and every number is generated live from the frozen atlas record, 2006-01-03 to 2026-08-19.

Every quant platform sells a way to read the market's mood, and so do we: ten of them, from a hidden Markov model to a moving average. So we audited our own shelf. Twenty years of the S&P 500, every classifier at its shipped defaults, every label produced only from data it could have seen, every metric registered before the compute ran. The forecast that works is volatility, not return. The winners are embarrassingly simple. And one respected input reads nothing at all.

A regime has no ground truth. Nobody stamps the days calm, choppy and stressed; a classifier's label is just a claim about what kind of day tomorrow will be, so the only honest test is to hold each label against what the market did next. That is the whole design. Ten classifiers, four families: statistical models that fit distributions (two hidden Markov models, a Gaussian mixture, a GARCH), indicators that read one number (VIX bands, price against its 200-day average), market structure read across the index's most liquid names, about a hundred for breadth and average correlation and sixty for the turbulence index, resolved from each year's point-in-time membership, and a macro classifier built from credit and rate proxies. All ten ship on our canvas with the same three-word vocabulary, all ten ran at shipped defaults, and the fitted models were refit monthly, point in time, by the same engine path every user's forward test uses. Their labels can never peek. The metric set was frozen in a version-controlled registration before the single compute ran, so nothing below was selected after seeing the answer.

SPY, log scale · red shading = share of classifiers saying stressed20062008201020122014201620182020202220242026HMM on returnsHMM on volatilityGaussian mixtureGARCH conditional volVIX bands200-day trendBreadthCorrelationTurbulenceMacro proxies
Figure 1. Twenty years, ten opinions. SPY on a log scale; the red shading behind it is the share of classifiers calling stress that week. Below, one lane per classifier:  calm ·  choppy ·  stressed. Dashed verticals mark the six registered crisis onsets. Rendered live from the frozen record.

The wall of bands is the study at a glance. In 2008 the shelf speaks with one voice: every lane goes red early and stays red, because the stress had been building since January and every kind of measurement could see it. Then look at 2022, where the lanes argue: breadth and GARCH went red in early October 2021, three months before the top, because the average stock was already breaking down while the index made highs on a handful of names. The VIX went red in December on the omicron spike. The trend lane never called the onset at all, because the 200-day average itself did not bend until spring. Full agreement is rare enough to be an event: on just 45 days in twenty years did every classifier with an opinion say stressed at once, and every one of those days but a single 2019 stray sits inside the 2008 crash or the covid spring. On one day in five, not one classifier says stressed. The interesting days, and there are thousands of them, are the disagreements, and they cluster exactly where the money is decided: at turning points.

The scoreboard. Registered metrics only. Separation is the difference in what SPY did next, stressed days minus calm days: forward 21-bar total return, forward 21-bar realized volatility (annualized), forward 63-bar maximum drawdown. Labels are attributed from the next bar; a classifier earns separation only from what it could not yet see.
ClassifierFamilyFlips/yrDwell % stressedRet sep (pp)Vol sep (pp) Drawdown sep (pp)
HMM on returns refit statistical 6.8 23 17.7 +0.8 +10.8 -2.6
HMM on volatility refit statistical 2.3 64 27.9 +0.1 +7.7 -3.2
Gaussian mixture refit statistical 3.0 45 8.8 +0.9 +10.6 -3.9
GARCH conditional vol refit statistical 5.1 23 30.7 +0.2 +10.8 -3.9
VIX bands indicator 5.9 25 16.7 +1.6 +15.9 -4.0
200-day trend indicator 2.8 31 16.1 -0.3 +13.9 -5.4
Breadth structure 6.0 25 39.5 -0.0 +6.7 -2.7
Correlation structure 4.2 35 32.6 +0.3 +8.1 -1.8
Turbulence structure 4.5 41 41.6 -0.2 +3.2 -1.1
Macro proxies macro 9.6 23 33.9 +0.5 -0.4 +0.9

The scoreboard settles what a label is actually worth. Read the volatility column first: every lane but one separates future turbulence, meaning its stressed days really are followed by wilder markets than its calm days, and the best separator is the simplest thing on the shelf. VIX bands at 16 and 25, two thresholds a retail reader could apply from a phone, beat every statistical model we own at forecasting volatility, by 15.9 annualized points between red days and green days. The 200-day trend line comes second at 13.9, and it wins the column that costs real money: its stressed days precede 63-day drawdowns 5.4 points deeper than its calm days, the widest damage gap on the shelf. The fear gauge forecasts weather; the trend line forecasts wreckage. Now read the return column and notice there is almost nothing in it. The best return separation on the shelf is 1.6 points and several are negative; after red days the market pays, if anything, slightly more than after green ones, which is the oldest bargain in equities dressed in new labels. Regime classification is a risk tool. It is not an alpha tool, and any product page that implies otherwise is selling you the wrong thing. One more result, and we let it stand because we measured it on ourselves: our macro classifier, credit spreads and curve slope by ETF proxy at shipped defaults, separates nothing. Minus 0.4 points of volatility separation, the wrong sign on drawdowns, and the most restless lane on the wall at nearly ten flips a year. As shipped, it reads nothing, and this atlas is the reason we now know that. It stays on the shelf marked with this finding, its inputs go back to the shop, and a future atlas run says whether it earned its place back. One more thing about this scoreboard: every line of it ships as a canvas node, at exactly the defaults measured here. Nothing in this atlas was scored in a lab you cannot enter.

The six crises. Days from the registered onset to each classifier's first stressed call, inside a window of 95 days either side. "early" means it was already stressed when the window opened; a blank means it never called stress inside the window at all.
Classifier200820112015201820202022
HMM on returns -76 +89 -54 +29 +7 +29
HMM on volatility early +38 +29 early
Gaussian mixture -45 +8 +61 +7
GARCH conditional vol early +28 -54 +29 +7 -94
VIX bands -62 +6 +3 +82 +4 -28
200-day trend early +27 +18 +44 +31
Breadth early -2 early early +7 -91
Correlation early +6 early +62 +7
Turbulence early -10 early early -21 -63
Macro proxies early -56 -54 +44 early early

One figure was added after the registration, and we label it so you never have to wonder: it was drawn from the frozen labels, one fixed rule, nothing searched. The rule is the bluntest use of a regime call that exists. Hold the index; when a lane says stressed, hold cash; act on the next bar; pay ten basis points per switch. Every gate pays. The index compounded 11.3 percent over this stretch and not one of ten gates kept up; the toll runs from seven tenths of a point a year to six and a half. What the toll buys is the product: the returns-reading Markov model cut the worst drawdown from 55 percent to 25 for 0.7 points a year, the cheapest insurance on the shelf, and the trend line bought 22 points of drawdown for about one. And the shelf's best forecaster is one of its worst gates: the VIX lane gave up 4.4 points a year and still rode a 36 percent drawdown, because the fear gauge stays elevated long after prices turn and a gate built on it keeps you out of recoveries. Reading the weather and driving the car are different jobs. Which single voice deserves the wheel, and whether a committee of all ten drives better, is the sealed question the next study answers on a real book.

HMM on returnsgated 10.6%/yr, dd -25 · SPY 11.3%/yr, dd -55 · in market 82% · 58 switchesHMM on volatilitygated 7.7%/yr, dd -24 · SPY 11.3%/yr, dd -55 · in market 71% · 18 switchesGaussian mixturegated 10.4%/yr, dd -49 · SPY 11.3%/yr, dd -55 · in market 91% · 24 switchesGARCH conditional volgated 8.0%/yr, dd -19 · SPY 11.3%/yr, dd -55 · in market 68% · 42 switchesVIX bandsgated 6.9%/yr, dd -36 · SPY 11.3%/yr, dd -55 · in market 83% · 52 switches200-day trendgated 10.4%/yr, dd -34 · SPY 11.3%/yr, dd -55 · in market 83% · 18 switchesBreadthgated 7.2%/yr, dd -19 · SPY 11.3%/yr, dd -55 · in market 61% · 68 switchesCorrelationgated 7.9%/yr, dd -20 · SPY 11.3%/yr, dd -55 · in market 67% · 36 switchesTurbulencegated 6.8%/yr, dd -35 · SPY 11.3%/yr, dd -55 · in market 57% · 71 switchesMacro proxiesgated 4.9%/yr, dd -39 · SPY 11.3%/yr, dd -55 · in market 66% · 118 switches
The cash gate, priced. Added after registration and labeled so: one fixed rule over the frozen labels, nothing searched.  holder who goes to cash while that lane reads stressed ·  SPY total return, same span. Next-bar execution, ten basis points per switch, cash earns zero. Log scale per panel.

The crisis table is where reputations get specific. Nobody missed 2008; the slowest classifier on the shelf was still six weeks early, and most were red before the window even opened. 2020 belongs to the turbulence index, which called stress three weeks before the fastest crash in history, while the trend line needed a month after the top, exactly as a 200-day average must in a vertical fall. 2011 and 2015 belong to the structure family, breadth and turbulence red before the index broke. 2022 belongs to breadth and GARCH, red in the autumn of 2021 while the index was still printing highs. And four of the ten, two fitted models and the trend line among them, never flagged 2022's onset inside our window at all. The pattern worth keeping: no family owns crisis detection. The statistical models were never first. The market's own internals, breadth and correlation and turbulence, front-ran three of the six crises, and the two simple indicators carried the rest.

Who agrees with whom. Pairwise Cohen's kappa on overlapping days, three states. Darker green is stronger agreement; 1.0 would be identical calendars, 0 is chance.
HMM on returnsHMM on volatilityGaussian mixtureGARCH conditional volVIX bands200-day trendBreadthCorrelationTurbulenceMacro proxies
HMM on returns · 0.12 0.18 0.34 0.19 0.23 0.24 0.17 0.06 0.06
HMM on volatility 0.12 · 0.32 0.37 0.20 0.25 0.14 0.33 0.03 0.03
Gaussian mixture 0.18 0.32 · 0.30 0.32 0.37 0.09 0.26 0.12 0.02
GARCH conditional vol 0.34 0.37 0.30 · 0.37 0.32 0.30 0.37 0.15 -0.03
VIX bands 0.19 0.20 0.32 0.37 · 0.27 0.15 0.26 0.13 0.03
200-day trend 0.23 0.25 0.37 0.32 0.27 · 0.21 0.31 0.14 0.02
Breadth 0.24 0.14 0.09 0.30 0.15 0.21 · 0.25 0.24 0.01
Correlation 0.17 0.33 0.26 0.37 0.26 0.31 0.25 · 0.21 0.03
Turbulence 0.06 0.03 0.12 0.15 0.13 0.14 0.24 0.21 · 0.00
Macro proxies 0.06 0.03 0.02 -0.03 0.03 0.02 0.01 0.03 0.00 ·

Family means: statistical with statistical 0.27; indicator with statistical 0.28; statistical with structure 0.19; macro with statistical 0.02; indicator with indicator 0.27; indicator with structure 0.20; indicator with macro 0.03; structure with structure 0.24; macro with structure 0.01.

The agreement matrix explains why a committee might beat a champion. The strongest pairwise agreement on the whole shelf is a kappa of 0.37, between GARCH and average correlation; most pairs sit near 0.2. These are ten genuinely different opinions, not ten copies of one signal, and diversity is the raw material ensembles are made of. The macro classifier agrees with nothing, which is what you would expect from a lane that measures nothing; a vote can carry a dead member as long as it is outnumbered. Whether the vote actually beats the best single voice is not a question an atlas can answer, because it is a portfolio question, and we have already sealed it: the next study races the shelf's best single classifier, VIX bands by the rule registered before this scoreboard existed, against the shipped ensemble of all ten, driving the same sector rotation machine our earlier series certified, on real books with real costs. The atlas does not tell you which gate to trade. It tells you which claims about reading the market survive twenty years of being checked. Fewer than advertised. More than zero.

Data coverage note: the cross-sectional classifiers read each year's point-in-time S&P 500 membership, priced through a vendor that no longer carries most delisted names. Names actually priced per year: Breadth 96 to 100; Correlation 96 to 100; Turbulence 60 to 60.

How to read this atlas

The metric set, classifier list, crisis dates and the next study's selection rule were frozen in a version-controlled registration before the single compute ran; the registration commit is printed at the top of this page and nothing was re-run. Labels were extracted through the engine's own causal path in calendar-year chunks from 2006 through August 2026: trailing-rule classifiers emit per-bar labels that use only trailing data, and the four fitted models were refit monthly on expanding as-of dates, each fit seeing only data at or before its date, with the same delay-confirm smoothing every forward test applies. Outcomes are measured on SPY rebuilt to total return with dividends on their ex-dates, and a day's label is credited only with outcomes that begin the next bar. Outcome attribution begins in August 2006, where the platform's twenty-year data window stood at compute time; earlier label days appear on the wall but score nothing. One product fix predates the registration and is disclosed: the trend classifier received the same five-bar minimum-duration de-noise every other classifier already had, after an engineering probe measured its raw series flickering at a median dwell of six bars. The cross-sectional classifiers resolve each year's point-in-time index membership, survivorship-honest at the membership level, priced through a vendor that no longer carries most delisted names; the names actually priced each year are printed above. The macro classifier's ETF proxies begin trading in 2007, so its record starts there. No multiple-testing correction is claimed anywhere on this page for the simplest reason: nothing was searched. One design, registered, computed once. The cash-gate figure was added after registration and says so where it stands: one fixed rule over the frozen labels, the same ten basis points every study charges, cash earning zero by the platform's stated rule, and no classifier re-run.

QuanterLab · Research atlas 902030d2259e · 2026-08-20. A registered, descriptive audit of the platform's ten shipped regime classifiers; its metric set was frozen before its single compute ran, and its figures render live from the frozen record. Educational research, not investment advice: every result shown is simulated, and nothing here is a recommendation to buy or sell any security.
Ten classifiers, one canvas

Every classifier in this atlas is a node users drag onto a QuanterLab canvas, at the same defaults audited here. The atlas record stays frozen; the shelf is live.

All research Follow the research

Run a study like this one

Everything above was produced inside QuanterLab, the registration, the walk, the statistics and the paper itself. Build the circuit on a canvas, register the hypothesis before you score it, and the platform enforces the rest.

The lab is in private beta and opens in September 2026. Reading the research needs no account, follow it and we'll tell you when the next study publishes.

As seen on Quantocracy

A note on AI. QuanterLab is a quantitative finance research platform, and every number in this study comes from a run on the platform. The hypothesis, the parameter choices, the validation design and the conclusions belong to the author. Runs execute on point-in-time data with walk-forward validation, and each study ships with its methodology and logs, so a reader can reconstruct the result instead of trusting it. I use AI to edit and structure the prose; it does not generate results, produce numbers, or decide what a study concludes.