QuanterLab produced this study — it wasn’t written up afterwards. Sealed hypothesis and search record in Appendix A2.
Open the Lab
QuanterLab · Research

Almost nothing in common, the same return: gross profitability vs. ROE across twenty sealed windows of the S&P 500

Universe · S&P 500 (point-in-time constituents)
Method · Comparative — Arm A vs Arm B
Manipulated variable · Gross Profitability vs. ROE
Step size · 1 year per forward window
In-sample · 2 years before each anchor
Out-of-sample span · 2006-01-03 → 2025-12-31
Compiled · August 05, 2026
Search record · none (size unknown — see §2.3)
Abstract

Two stock screens that agree on fewer than three names in twenty earned the same return for twenty years — and the cheaper one finished ahead once trading costs were counted. That is the finding of this study, and none of it was visible in advance.

The design is a controlled comparison of two definitions of quality. Arm A ranks the point-in-time S&P 500 on Robert Novy-Marx's gross profitability — trailing-twelve-month gross profit over total assets. Arm B ranks the same universe on return on equity, the quality metric every retail screener ships with. Both hold their top twenty equal-weighted, re-selected quarterly, across twenty sealed one-year windows from 2006 through 2025, benchmarked in-run against RSP, the equal-weight S&P 500 ETF — a benchmark that matches how the books are built, so the comparison never rests on the weighting mismatch that distorts concentrated portfolios measured against a cap-weighted index.

The screens select almost disjoint portfolios: mean overlap 2.8 names of twenty per rebalance, and one of the ninety-five rebalances produced entirely disjoint books. They trade at very different rates: gross profitability replaced 13% of its book per quarter, ROE 39%. And they arrive together: 9.43% against 9.76% CAGR, a paired Sharpe of 0.02, and a block bootstrap over 5,011 paired trading days that puts the probability gross profitability beats ROE at 53.9% — a coin flip. At 17 basis points per trade, the threefold turnover difference flips the ordering: 9.26% net against 9.23%. Both beat the equal-weight index by roughly two points a year.

Is either screen crisis insurance? Both defended when the index fell — ROE's six-window summary statistics are the better ones, gross profitability's excess concentrated in the deepest drawdowns — and the interpretation section takes care to separate what the statistics say from what the composition suggests. What the study can say without interpretation: the choice between the two most common definitions of quality decided which companies you owned, how much you traded, and which crises hurt — and not, over twenty years, what you earned.

Author’s note

I published this study on the second attempt, and the first attempt is the reason to trust the second — the full account is in the note at the top of the paper. Two things changed my read between the drafts. With the equal-weight benchmark sealed into the run, both quality screens beat the index rather than trailing the market, which retired the story I expected to write about post-publication decay. And on the clean data the two definitions became fully indistinguishable on gross returns while remaining almost disjoint in holdings — which moved the paper's centre of gravity from the horse race I designed to the tie I found, and to the one difference that survives it: the trading bill. The title question I started with — is this crisis insurance? — gets its honest, partial answer in the discussion, with the statistics and the interpretation kept deliberately apart.

Read this first: price returns, and a discarded first compilation

Two notes before the figures. First: every number in this paper — both arms and the benchmark — is a price return without dividends. Do not compare any figure here against published total-return index performance; within the paper, all comparisons are like-for-like.

Second, on why this record deserves trust: this study's first compilation was discarded. During review, the gross-profitability metric was found to be silently mis-scoring a class of companies — the data vendor reports a zero cost of revenue for some filers (health insurers above all), which made gross profit equal revenue and promoted those names to the top of the ranking on a number that measured the wrong thing. The metric was repaired to exclude what it cannot measure, all twenty windows were re-run from scratch, and the sealed record published here was verified name-by-name at every rebalance: it contains zero such names. The defect, the fix and the re-run are disclosed rather than hidden because the platform's entire premise is that the record you read is the record that ran.

1  Methodology

This is a pre-registered comparative study: two arms, one manipulated variable, twenty sealed one-year windows anchored on the first of January 2006 through 2025. Each step's hypothesis was registered and hashed before its out-of-sample window was scored, and the anchor advances only forward — the walk cannot revisit a window or close a step with a result registered for a different one. "Walk-forward" here means the anchor itself walks: nothing in this design is re-fitted from step to step, and everything it computes — rankings, risk projections — is estimated at each anchor from data available before that anchor only.

The universe is the point-in-time S&P 500, reconstructed at every anchor and every rebalance from the exchange's dated constituent change-log, so names later removed or delisted compete on the dates they actually traded: roughly 350 names at the 2006 anchor, rising to about 500 by 2025. Fundamentals are quarterly statements carrying SEC acceptance timestamps; prices are daily bars. All figures platform-wide are price returns without dividends, benchmark included, so no number here should be compared against published total-return indices. The benchmark is RSP, the equal-weight S&P 500 ETF, sealed into the run for both arms — chosen because the books are equal-weighted, and a cap-weighted comparison would confound the factor with the weighting scheme.

Arm A ranks on gross profitability: trailing-twelve-month gross profit divided by the total assets of the most recent quarter in that sum. The four quarters must be fiscally contiguous — joined on calendar year and period rather than on dates, so 52/53-week filers survive — with restated quarters deduplicated in favour of the latest filing. A quarter becomes visible only when both halves of its filing pair have been accepted: the income statement and the balance sheet each carry their own acceptance date, and the later of the two gates the pair. Where the vendor reports a zero cost of revenue and gross profit merely restates revenue, the name is excluded from the ranking rather than scored on a number that measures the wrong thing; this coverage boundary runs from 67% of the fundamentals universe at the 2006 anchor to 92% by 2020 and is discussed in the limitations. Arm B ranks on return on equity as the vendor reports it. In each arm, only that arm's single metric is switched on inside the quality family — every other quality metric is explicitly off in both arms — so the family score is the metric itself, and the composite weights quality at one hundred percent.

Scores are cross-sectional z-scores, winsorized at the 1st and 99th percentiles, computed fresh at every rebalance. Each arm holds its top twenty names at equal weight and re-selects quarterly, re-ranking the full point-in-time universe as of the rebalance date. Fills are same-bar daily closes. A holding that loses price coverage mid-window is dropped and the book renormalises over the remaining names; the appendix rows note each occurrence.

Inside each window the two arms' daily returns are inner-joined date by date, and the paired difference is the pre-declared object under test — one contrast, sealed before any window was scored, which is why no multiple-testing deflation is applied and none is claimed. The probability that one arm beats the other comes from a stationary block bootstrap over the 5,011 paired trading days (block length 10, 2,000 resampled paths, seed 1234). Sharpe figures throughout are raw-return Sharpe — mean over standard deviation of the portfolio's own returns, with no risk-free deduction. Risk projections are per-rebalance Monte Carlo cones (P5–P95) and daily value-at-risk, fitted exclusively on data preceding the segment they are scored against; their realised calibration is reported in the appendix rather than assumed. Trading costs and taxes are zero in the sealed record, deliberately, and are added back as disclosed arithmetic in the closing section rather than by re-running the study.

Transaction costs are not modelled in this study; all results are gross of costs.

2  Results

2.1  Headline

Arm A — pooled Sharpe
0.54
5012 OOS bars
Arm B — pooled Sharpe
0.59
5012 OOS bars
P(Arm A beats Arm B)
53.9%
5011 paired bars · CAGR gap (Arm A − Arm B) -0.3 pp
Out-of-sample equity — normalised growth (1.00x = break even)-0.02x3.59x7.20x2006200920122015201820212024
Figure 1. Both arms stitched through the identical windows —  Arm A (+507.0%),  Arm B (+542.2%), benchmark grey (+325.6%). Dotted verticals mark the step boundaries; the dashed horizontal is break-even.
Out-of-sample equity — normalised growth (1.00x = break even)0.56x3.24x5.91x20102012201420162018202020222024
Figure 2. The same walk, re-based to 1.00x at the first window starting in 2010 — 16 of the 20 windows above.  Arm A (+417.8%),  Arm B (+430.7%), benchmark grey (+362.5%). This is a subset of Figure 1, not a correction to it. The study is anchored before the 2007–09 crisis on purpose: a method that only works in calm markets should be caught doing it. But one crisis window and its equally singular recovery set the vertical scale for the whole of Figure 1, and everything after 2010 is compressed into the bottom of it. This figure shows the same windows, same method, same data, with that period outside the frame — so the post-crisis years can be read at their own scale. Neither figure is the honest one on its own; the full record is what the study claims, and the era rows below put a number on how much of the gap came from which period.

2.2  Per-step results

Table 1. One row per step — raw out-of-sample results.
#Out-of-sample window Arm A SR Arm B SR
1 2006-01-03 → 2006-12-29 1.10 0.70
2 2007-01-03 → 2007-12-31 0.13 0.97
3 2008-01-02 → 2008-12-31 -0.47 -0.84
4 2009-01-02 → 2009-12-31 1.12 1.64
5 2010-01-04 → 2010-12-31 1.27 0.63
6 2011-01-03 → 2011-12-30 0.41 0.41
7 2012-01-03 → 2012-12-31 0.39 0.91
8 2013-01-02 → 2013-12-31 1.72 2.37
9 2014-01-02 → 2014-12-31 0.55 1.31
10 2015-01-02 → 2015-12-31 -0.84 0.06
11 2016-01-04 → 2016-12-30 0.66 0.27
12 2017-01-03 → 2017-12-29 1.99 2.10
13 2018-01-02 → 2018-12-31 0.07 -0.41
14 2019-01-02 → 2019-12-31 1.45 2.54
15 2020-01-02 → 2020-12-31 0.79 0.57
16 2021-01-04 → 2021-12-31 2.01 1.87
17 2022-01-03 → 2022-12-30 -0.48 -0.69
18 2023-01-03 → 2023-12-29 0.85 0.53
19 2024-01-02 → 2024-12-31 1.19 2.58
20 2025-01-02 → 2025-12-31 0.68 0.21
Out-of-sample equity — normalised growth (1.00x = break even)0.53x1.00x1.47xbars into the window →
Figure 3. Arm A — every step's out-of-sample curve overlaid, each rebased to 1× at its own start. Read alongside Table 1: consistent shape across steps is the walk-forward's evidence; a single lucky leg is not.
Out-of-sample equity — normalised growth (1.00x = break even)0.48x1.01x1.54xbars into the window →
Figure 4. Arm B — the same windows, the other arm. Compare shape-for-shape with the previous figure: the two arms trade the identical out-of-sample legs.

2.3  Search accounting

No search record exists for this design. It was not promoted from a recorded evolving search, so the number of alternatives tried before it — on paper, in another tool, or in the author's head — is unknown. Unknown is a different fact from one: a study with no lineage is not a strategy with one trial, it is a strategy with an unrecorded number of them. Accordingly this paper claims no deflated Sharpe and no trial count; the honest statement is the raw out-of-sample result plus this disclosure. The registered per-step record below (§4) still guarantees each window's hypothesis was sealed before that window was scored.

2.4  The comparison

Both arms trade the same sealed windows, so their returns can be PAIRED: inside each window the two return series are inner-joined date by date and the difference rArm A − rArm B is the object under test. Because this is ONE pre-declared contrast — sealed before any window was scored — the paired statistic needs no multiple-testing deflation, and the per-arm pooled numbers above carry no deflated Sharpe either — this design has no recorded search to deflate against (§2.3). The paired contrast is the one statistic here that a missing search record does not weaken: it was declared in advance, and it is scored on the difference rather than on either arm's level.

Table 2. Window-by-window paired comparison. Δ is the growth gap (Arm A − Arm B) over the window's paired dates.
#WindowPaired bars Arm AArm B ΔLeader
1 2006-01-04 → 2006-12-29 250 +13.7% +8.9% +4.8 pp Arm A
2 2007-01-04 → 2007-12-31 250 +0.8% +15.1% -14.3 pp Arm B
3 2008-01-03 → 2008-12-31 252 -25.9% -33.4% +7.5 pp Arm A
4 2009-01-05 → 2009-12-31 251 +37.9% +44.9% -7.0 pp Arm B
5 2010-01-05 → 2010-12-31 251 +25.1% +9.1% +16.0 pp Arm A
6 2011-01-04 → 2011-12-30 251 +6.9% +6.6% +0.4 pp Arm A
7 2012-01-04 → 2012-12-31 249 +4.9% +11.7% -6.7 pp Arm B
8 2013-01-03 → 2013-12-31 251 +24.4% +31.5% -7.0 pp Arm B
9 2014-01-03 → 2014-12-31 251 +6.1% +16.6% -10.5 pp Arm B
10 2015-01-05 → 2015-12-31 251 -13.2% -0.3% -12.9 pp Arm B
11 2016-01-05 → 2016-12-30 251 +8.9% +2.7% +6.3 pp Arm A
12 2017-01-04 → 2017-12-29 250 +20.8% +17.5% +3.3 pp Arm A
13 2018-01-03 → 2018-12-31 250 -0.2% -8.0% +7.8 pp Arm A
14 2019-01-03 → 2019-12-31 251 +22.4% +38.1% -15.8 pp Arm B
15 2020-01-03 → 2020-12-31 252 +23.1% +14.3% +8.9 pp Arm A
16 2021-01-05 → 2021-12-31 251 +28.7% +25.4% +3.3 pp Arm A
17 2022-01-04 → 2022-12-30 250 -14.5% -17.8% +3.3 pp Arm A
18 2023-01-04 → 2023-12-29 249 +12.0% +7.1% +5.0 pp Arm A
19 2024-01-03 → 2024-12-31 251 +17.5% +35.5% -18.0 pp Arm B
20 2025-01-03 → 2025-12-31 249 +12.5% +2.2% +10.3 pp Arm A

Paired Sharpe of the difference track: 0.02 · block bootstrap (2000 paths, block 10, seed 1234): P(Arm A beats Arm B) = 53.9%.

Window win-rate. Arm A led 12 of 20 windows (60.0%), Arm B led 8 — yet the mean window gap runs the other way: -0.79 pp toward Arm B. Arm A wins more often and smaller; Arm B wins less often and larger. The count and the mean answer different questions, and neither settles the comparison by itself. Widest single window: 2024 at -18.0 pp.

Table 3. The same comparison split at 2010. Averaging across a crisis and a decade of calm hides which one the difference came from.
PeriodWindows Arm AArm B Mean gapArm A led
All windows 20 +10.59% +11.38% -0.79 pp 12/20
Before 2010 4 +6.62% +8.88% -2.25 pp 2/4
2010 onward 16 +11.59% +12.01% -0.43 pp 10/16
All windowsn=20 · Arm A led 12+10.6%+11.4%-0.79 ppBefore 2010n=4 · Arm A led 2+6.6%+8.9%-2.25 pp2010 onwardn=16 · Arm A led 10+11.6%+12.0%-0.43 ppgap
Figure A1 — mean window return per period. Arm A above, Arm B below, with the gap at right. The pooled bar and the post-2010 bar are the same comparison over different periods.

3  The circuit

The strategy is a circuit of platform primitives, frozen when the study is registered. Below is the circuit as wired on the canvas, the objective it encodes and how the search runs through it, followed by the mathematics each primitive actually computes — the same formulas the execution engine runs. The complete parameterisation is preserved in the study ledger (Appendix A).

The hypothesis under test

A COMPARATIVE study — Arm A vs Arm B, walked on the same sealed out-of-sample windows. Arm A: S&P 500, rebalanced quarterly across the selected basket, and run out-of-sample from the anchor — this design has no tunable in-sample parameters, so nothing is re-fitted step to step and the walk is a sequence of sealed out-of-sample windows; its disposition is the realized forward path versus the benchmark. Arm B: S&P 500, rebalanced quarterly across the selected basket, and run out-of-sample from the anchor — this design has no tunable in-sample parameters, so nothing is re-fitted step to step and the walk is a sequence of sealed out-of-sample windows; its disposition is the realized forward path versus the benchmark. The arms differ in: Quality Factor — gp_to_assets: high → off; Quality Factor — roe: off → high. The contrast under test: whether Arm A generates better risk-adjusted returns than Arm B over the identical out-of-sample windows.

The frozen circuit — data flows left to rightuniverse — click for detailsuniverseprice loader — click for detailsprice loaderfactor loader — click for detailsfactor loaderfactor quality — click for detailsfactor qualityfactor composite — click for detailsfactor compositefactor top tier — click for detailsfactor top tierportfolio backtest — click for detailsportfolio backtestportfolio forward autopsy — click for detailsportfolio forward autopsyuniverse — click for detailsuniverseprice loader — click for detailsprice loaderfactor loader — click for detailsfactor loaderfactor quality — click for detailsfactor qualityfactor composite — click for detailsfactor compositefactor top tier — click for detailsfactor top tierportfolio backtest — click for detailsportfolio backtestportfolio forward autopsy — click for detailsportfolio forward autopsyArm AArm Bshared
Figure 5. The frozen circuit — every node a primitive, every wire a typed data-flow; the two arms are colour-coded (Arm A green, Arm B blue, shared feeds neutral). Each box is one step of the strategy; data flows along the wires left to right, and no box can see data dated later than the box feeding it. The whole diagram was frozen when the hypothesis was registered. Click any node to open what that step ran with and what it produced.

Envelopes show counts, ratios, dates, and the parameters the author chose. Full price and per-name data series are not republished: the underlying market data is licensed to QuanterLab. Point figures quoted in the prose — a named holding's return over a stated span — are summary facts derived from public market prices, not redistributed series.

What each part does
Universe — The starting set of tickers — resolved point-in-time from the index change-log, so names delisted or removed later still compete on the dates they traded.
Price Loader — Bulk OHLCV fetch for the whole universe — point-in-time, no future bars.
Factor Loader — Point-in-time fundamentals — never let the user see a number before the SEC did.
Factor Quality — Quality — is this a strong, profitable, well-financed business?
Factor Composite — The weighting console — blend Value, Quality, Momentum, Growth into one 0–100 score.
Factor Top Tier — The cut out of the factor lane — keep the top-ranked names.
Portfolio Backtest — Replay the portfolio forward — rebalanced, point-in-time. No transaction-cost overlay is wired in this circuit, so these results are gross of costs.

The objective and the search

Arm A — S&P 500, rebalanced quarterly across the selected basket, and run out-of-sample from the anchor — this design has no tunable in-sample parameters, so nothing is re-fitted step to step and the walk is a sequence of sealed out-of-sample windows; its disposition is the realized forward path versus the benchmark.

UniverseS&P 500 index constituents.
Validation & out-of-sampleportfolio forward test (buy-and-hold book) (1y horizon from the anchor, quarterly rebalance).
Other componentsFactor models: Factor Composite, Factor Select, Fundamentals Loader (PIT), Quality Factor.

Arm B — S&P 500, rebalanced quarterly across the selected basket, and run out-of-sample from the anchor — this design has no tunable in-sample parameters, so nothing is re-fitted step to step and the walk is a sequence of sealed out-of-sample windows; its disposition is the realized forward path versus the benchmark.

UniverseS&P 500 index constituents.
Validation & out-of-sampleportfolio forward test (buy-and-hold book) (1y horizon from the anchor, quarterly rebalance).
Other componentsFactor models: Factor Composite, Factor Select, Fundamentals Loader (PIT), Quality Factor.

What differs between the arms — one manipulated variable, expressed as 2 paired settings on one node:

  • paramQuality Factor — gp_to_assets: high → off
  • paramQuality Factor — roe: off → high

Everything else is held identical, so an out-of-sample gap between the arms is attributable to this one change.

No transaction-cost elements are wired into this circuit; results are gross of costs.

Show the mathematics — 7 primitives, formulas and parity notes

3.1  Universe

The starting set of tickers — resolved point-in-time from the index change-log, so names delisted or removed later still compete on the dates they traded.

Before any math, you need a list of stocks. An index preset (S&P 500, Nasdaq-100, Dow 30) is reconstructed as it stood ON your anchor date by replaying the historical add/drop change-log backwards — so a 2018 backtest sees the 2018 membership, not today's winners.

Point-in-time membership

Start from today's constituents and un-apply every membership change after the anchor t:

\mathcal{U}(t) = \mathcal{U}_{\text{now}} \;\ominus\; \{\text{adds after } t\} \;\oplus\; \{\text{drops after } t\}
Constituents resolved from the index change-log; the same point-in-time set the factor + screening modules use.

3.2  Price Loader

Bulk OHLCV fetch for the whole universe — point-in-time, no future bars.

Momentum, volatility, trend — every price-based metric needs history. This loads open/high/low/close/volume for all names in parallel, clipped so nothing after the anchor can leak in. The lookback window is derived automatically from the deepest metric you wired.

The window is derived, not guessed

It loads exactly enough history for the hungriest downstream metric plus a warm-up buffer:

W = \max_k(\text{lookback}_k) + \text{buffer}, \qquad \text{bars} \le \text{anchor } t

3.3  Factor Loader

Point-in-time fundamentals — never let the user see a number before the SEC did.

Loads ~23 fundamental metrics (valuation, quality, growth) per name, but with one inviolable rule: a financial statement becomes visible only on or after its SEC acceptedDate. A 2020 backtest sees only what was actually filed by 2020 — no look-ahead, ever.

The PIT gate
\text{visible}(f, t) \iff \text{acceptedDate}(f) \le t
Missing acceptance dates fall back to filingDate, else statement date + 45 days.
Byte-identical to FM101FBKT (shared_libs/factor_core). US indexes only (SEC reliability).

3.4  Factor Quality

Quality — is this a strong, profitable, well-financed business?

Blends profitability (ROE, ROA, ROIC, margins) and balance-sheet strength (debt-to-equity inverted, current ratio, interest coverage). Same z-score, winsorize, importance-weight recipe as every factor family.

Family score
\text{Quality}_i = \frac{\sum_k \omega_k\,z_{i,k}}{\sum_k \omega_k}
Debt metrics enter inverted (less leverage = higher quality).
Byte-identical to FM101FBKT (shared_libs/factor_core).

3.5  Factor Composite

The weighting console — blend Value, Quality, Momentum, Growth into one 0–100 score.

Where the four factor families become a single ranking. Each family score is standardized across the universe, blended with your slider weights (or the radar's suggested tilt), and min-max scaled to 0–100. Winsorizing tames outliers; z-score or percentile normalization is your choice.

Cross-sectional standardize + winsorize
z_{i,f} = \frac{x_{i,f} - \bar x_f}{s_f}\quad(\text{clipped at the 1st / 99th percentile})
Weighted blend, scaled to 0–100
C_i = \sum_f W_f\,z_{i,f}, \qquad \text{score}_i = 100\cdot\frac{C_i - \min_j C_j}{\max_j C_j - \min_j C_j}
W = your four slider weights (total 100) OR the Regime Tilt radar's suggestion. Needs ≥ 10 names, ≥ 3 valid metrics each.
Byte-identical to FM101FBKT ranking (shared_libs/factor_core.rank_stocks_at_date).

3.6  Factor Top Tier

The cut out of the factor lane — keep the top-ranked names.

Takes the composite-ranked factor set and keeps the best N, carrying the composite score, the four family scores and the point-in-time market cap for each survivor. Feed 10–20 to a direct portfolio, or 30–100 as an optimizer pool.

Rank cut
\{\, i : \operatorname{rank}(C_i) \le N\,\}, \quad C_i = \text{composite score}
Byte-identical to FM101FBKT ranking (shared_libs/factor_core).

3.7  Portfolio Backtest

Replay the portfolio forward — rebalanced, point-in-time.

Holds the basket and rebalances on schedule, re-selecting and re-optimizing point-in-time at each rebalance (so it only ever uses information available then), and reports the equity curve, Sharpe, drawdown and trade stats — optionally net of cost and risk overlays.

Compounded equity
E_t = E_{t-1}\big(1 + \mathbf w_{t}^{\top}\mathbf r_t - \text{costs}_t\big)
Drawdown
\text{DD}_t = \frac{E_t}{\max_{\tau\le t}E_\tau} - 1, \qquad \text{MaxDD} = \min_t \text{DD}_t

4  Projection calibration, pooled across the walk

Every rebalance carried a Monte Carlo cone and a 95% VaR estimated before the segment it is scored against. Two questions, pooled over the whole study: did realized outcomes land inside the band as often as the band claims, and were VaR breaches as frequent as 5%?

This section is produced by the forward tester itself: every portfolio backtest fits the cone and the VaR estimate at each rebalance and scores them against the segment that followed. It does not require — and this circuit does not contain — a Monte Carlo primitive; that primitive is a separate, standalone analysis.

Arm A71 of 80 inside the 90% band-36%+4%+44%in band20062007200820092010201120122013201420152016201720182019202020212022202320242025Arm B71 of 80 inside the 90% band-36%+4%+44%in band20062007200820092010201120122013201420152016201720182019202020212022202320242025
Figure A2 — projected range versus what occurred, at each of 160 scored rebalance segments, pooled across both arms. Segments too short to form an estimation window are not scored, which is why this count can sit below the raw rebalance totals in the table beneath. Each vertical bar is that rebalance's P5–P95 Monte Carlo cone with the median ticked; the dot is the realized return of the segment that followed. Filled green = the outcome landed inside its own cone; red = it did not. The strip beneath repeats that as one mark per rebalance, so a run of misses in one period is visible as a run. Every cone was fitted only on data prior to the segment it is scored against.
Arm Steps Rebalances In band Coverage Expected VaR days Breach rate Expected
Arm A 20 95 71 / 80 88.8% ±3.53 90.0% 4951 5.59% ±0.327 5.0%
Arm B 20 95 71 / 80 88.8% ±3.53 90.0% 4951 5.88% ±0.334 5.0%

± values are binomial standard errors on the estimate. A coverage figure below the expected band means the projection was over-confident; a breach rate above 5% means the same of the risk model. Both forecasts used only data prior to the segment scored.

5  Discussion

5.1  Findings

1 — The two definitions of quality select almost entirely different companies. Mean per-rebalance overlap across ninety-five rebalances is 2.77 names of twenty, the maximum ever reached is six, and one rebalance produced entirely disjoint books. Window-level Jaccard similarity averages 0.105 (range 0.051 to 0.155). Whatever these two screens are measuring, they disagree about which S&P 500 companies embody it.

2 — Their twenty-year results are statistically indistinguishable. Arm A (gross profitability) compounded at 9.43% a year, Arm B (ROE) at 9.76% — a gap of 0.33 points (cumulative paths in Figure 1). The compiled comparison reports a paired Sharpe of 0.02 and P(A beats B) = 53.9% from a block bootstrap on 5,011 paired trading days (block length 10, 2,000 paths, seed 1234) — as close to a coin flip as twenty years of daily data can produce. Pooled daily Sharpe is 0.537 for A and 0.591 for B; mean one-year window Sharpe 0.73 against 0.89.

3 — Both screens beat the equal-weight index. Against RSP's 7.50% CAGR over the same sealed windows, Arm A earned +1.93 points a year and Arm B +2.26. The benchmark is equal-weight by design — the books are twenty names at equal weight, so cap-weighted comparisons would confound the factor with the weighting scheme. This choice was sealed into the run, not applied afterwards.

4 — Both defend when the index falls, and ROE's summary numbers are the better ones. In the six down-benchmark windows (2007, 2008, 2011, 2015, 2018, 2022), Arm B's mean excess over RSP was +5.60 points (SE 2.78), beating the index five times of six with 53% downside capture; Arm A's was +4.22 (SE 3.58), four of six, 65% capture. In the fourteen up windows both roughly matched the index (102% and 105% capture). The composition of the defense differs — A's excess concentrated in the deepest drawdowns (+14.3 in 2008, +10.1 in 2011, +10.0 in 2018, against a −9.0 failure in 2015), while B's largest single defense (+15.4) came in 2007, a window in which the benchmark fell just 0.3%, and its 2008 excess was +6.8 — but on the summary statistics alone, B defended more reliably. What to make of that difference in shape is treated as interpretation in the discussion, not as a finding.

5 — The turnover gap is a factor of three, and it is the mechanism made visible. Gross profitability replaced 12.8% of its book per quarterly rebalance (0.51× annualised one-way); ROE replaced 39.2% (1.57×) — a 3.07× ratio, stable across all twenty windows. Gross profit over assets is a persistent property of a business; reported ROE is not. This is Novy-Marx's own argument — that gross profit sits above the accounting discretion and leverage effects that contaminate earnings-based measures — replicating in the holdings even though the return premium does not.

6 — Net of trading costs, the cheaper screen wins. At 17 basis points per trade, turnover costs Arm A roughly 0.17 points a year and Arm B 0.53. Applied to the gross records, the 0.33-point gap does not merely narrow — it changes sign: 9.26% against 9.23%. The margin is three basis points and carries no statistical weight; the point is directional and structural. Two screens the bootstrap cannot separate gross are separated net by nothing but their trading, and the ordering favours the one that trades a third as much at any cost assumption.

7 — The era pattern is concentration, not clean decay. Pre-publication windows (2006-2012) gave A +3.5 and B +3.4 points of annual excess; the three crisis windows alone (2007-2009) gave +7.4 and +10.0. Post-publication (2013-2025), A's excess is +0.28 points — statistically zero — while B's +1.17 is dominated by two windows: 2019 (+11.6) and 2024 (+24.5, a year in which its book held NVIDIA through a +179% run). Twenty windows cannot separate publication decay from regime concentration; both readings fit, and the paper claims neither.

8 — Ranking coverage is the study's honest asterisk. Gross profitability requires four contiguous quarters of both income and balance-sheet filings with SEC acceptance dates, and the vendor cannot state gross profit for every company — where its cost-of-revenue field is empty and gross profit merely echoes revenue, the name is excluded from the GP/A ranking rather than scored on a wrong number. Coverage rises from 67% of the fundamentals universe at the 2006 anchor to 92% by 2020, and the exclusions cluster in health insurers and media. Arm B's ROE ranking does not share this constraint. Every reported number sits downstream of this coverage boundary.

5.2  Interpretation

The question this study set out to answer is whether the better definition of quality pays a retail investor. The answer the twenty windows give is precise: the definition changed almost everything about the portfolio and almost nothing about the gross return — and the one durable difference it did produce, the trading bill, favours gross profitability.

Start with the tie, because it is stranger than it looks. These two screens agree on fewer than three names in twenty at any moment. Their books turn over at rates three times apart. They lose money in different years, recover in different years, and their worst names come from different industries. Yet after twenty sealed years they sit 0.33 points of CAGR apart with a bootstrap that cannot tell them apart at all. A study reporting only the ending wealth would conclude the choice of quality metric is irrelevant. The holdings say the opposite: the choice decides which companies you own, how often you trade, and which crisis hurts you. It is the rare case where the process differs enormously and the outcome does not — and once costs enter, the process difference is the outcome: 9.26 against 9.23, the cheaper screen ahead.

Second, both screens beat the equal-weight index, and the study is built so that sentence means something. The books are twenty names at equal weight; the benchmark is the equal-weight S&P 500, sealed into the run. An earlier draft benchmarked against the cap-weighted index and needed a section explaining why that comparison misleads; building the control into the experiment removed the excuse in advance. Against the weight-matched index: +1.9 and +2.3 points a year, both arms, twenty years.

Third, the down-market record — reported straight, then interpreted. The summary statistics favour ROE: across the six windows in which the benchmark fell, B's mean excess was +5.60 to A's +4.22, it beat the index five times to A's four, and its downside capture was 53% to A's 65%. On those numbers, ROE was the more reliable defensive screen, and its 2008 book — for all its 33.4% loss — still fell 6.8 points less than the index. The interpretation, and it is offered as nothing more, concerns shape: A's defense sits in the deepest, most systemic drawdowns (2008 by 14.3 points, 2011 by 10.1, 2018 by 10.0), while nearly half of B's six-window mean comes from 2007, a window the sign convention counts as "down" because the benchmark fell 0.3% — a commodity-boom year in which B's book held Freeport and Amazon, not a stress year in which anything was defended. Strip that single near-flat window and the two screens' defensive records converge toward the same handful of points. A reader is free to weight the summary statistics, the composition, or both; the statistics side with B, the deepest-crisis record with A, and 2015 — where A lost nine points to a flat index while B gained four — is the counterexample that keeps either story honest. Separately, and outside the down-window ledger entirely: A's best absolute year against the index was 2020 (+13.2), an up-benchmark window in which its asset-light book (IDEXX +89%, Cadence +91%) compounded through the recovery.

Fourth, the mechanism. Novy-Marx's argument for gross profitability was never merely empirical: gross profit sits at the top of the income statement, above the depreciation schedules, leverage effects and accrual choices that make reported earnings noisy. If that is true, a GP/A ranking should be more stable than an ROE ranking — and it is, by a factor of 3.07 in turnover, consistent in every one of the twenty windows. The signal's persistence replicated perfectly; the return premium did not. Post-2013, A's excess over equal weight is 0.28 points a year, indistinguishable from zero, while ROE's apparent post-publication edge rests substantially on holding NVIDIA through 2024. Whether that is the market pricing the anomaly away after publication or three windows of regime doing the talking, twenty observations cannot say.

Fifth, what the tie means for practice. If two screens earn the same gross return, the cheaper and calmer one wins net — and here that is not a homily but the record: 9.26 against 9.23 at 17 basis points, before the tax asymmetry the closing section quantifies, which also runs in the low-turnover screen's favour. The practical reading of this study is not that gross profitability is the better anomaly — the data refuse to say that — but that it is the cheaper implementation of the same tie, and cheap compounds.

No search record exists for this study: the design was not promoted from a recorded evolving search, so the number of alternatives tried before it is UNKNOWN — which is a different fact from one. No deflated Sharpe is claimed; the honest statement is the raw out-of-sample result plus this disclosure. The out-of-sample windows are historical.

Costs and taxes: the frictions the sealed record left out

The sealed record is gross of trading costs and taxes, deliberately. This section adds both back as arithmetic on the frozen results — disclosed, approximate, and not a re-run.

Costs first, because they change the paper's answer. One-way turnover annualises to 0.51× for gross profitability and 1.57× for ROE; each unit is a sale plus a purchase. At 17 basis points per trade — the platform's default for S&P 500 large-caps (10bp spread, 5bp slippage, 2bp impact) — the drag is roughly 0.17 points a year for Arm A and 0.53 for Arm B. Applied to the gross records that is 9.26% against 9.23%: the 0.33-point gross gap does not narrow, it changes sign. The margin is three basis points — nothing statistically — but its direction is structural: it exists at any positive cost assumption, it triples at 50 basis points, and it is the only difference between these two screens that twenty years of data actually delivered. Neither arm's margin over the index is threatened; the index pays its own rebalancing spread.

Taxes are the larger friction, and one thing must be said plainly first: the index holder pays them too. Nobody in this comparison escapes capital-gains tax — a buy-and-hold position is taxed in full when it is finally sold. What the index holder gets is not exemption but timing: gains that stay untaxed for twenty years keep compounding before the one payment at the end, while a quarterly re-selected book hands the tax office a slice of each year's gain and loses the compounding on every slice.

The arithmetic, per 100 invested over the twenty windows at a 25% capital-gains rate, with both sides taxed. Untaxed, Arm A's 9.43% would grow 100 into about 606 and the index's 7.50% into about 425. Taxing Arm A's gain every year leaves about 392 — roughly 214 of terminal wealth surrendered to taxes and their lost compounding. Taxing the index holder once at the final sale leaves about 344 — a tax bill of roughly 81. So the screen pays over two and a half times as much to the tax office as the index holder pays, on the same starting capital — and still finishes about 49 ahead per 100 invested (392 against 344; net CAGRs 7.07% against 6.37%). At the German Abgeltungsteuer rate of 26.375% the same accounting gives 383 against 339 — still about 44 ahead. The screen's edge survives the least favourable honest comparison: itself taxed annually, the alternative taxed only at exit.

Against the cap-weighted alternative nobody in this study traded, the ordering flips. SPY earned 8.47% over the identical twenty windows — computed from the same price source as supplementary arithmetic, outside the sealed record — which with exit-only taxation nets about 406 per 100 (7.26%), slightly ahead of the annually-taxed screen. A taxable retail investor's honest summary: the quality screen beat its own index after every friction including its heavier tax bill, and the decade's cap-weighted mega-cap run beat both. One further symmetry: dividends, which this platform does not model, are taxed in the year received for screen and index holder alike — deferral only ever applied to price gains, so adding dividends would narrow the timing advantage, not widen it.

QuanterLab Circuit Diagram
QuanterLab Circuit Diagram

5.3  Limitations

This is a large-capitalisation, long-only study and Novy-Marx's was neither. His premium is estimated across the full cross-section of US equities, long minus short, and is widest among small and illiquid names. This study holds the long leg only, inside the S&P 500 only. The absence of a return premium here is a statement about large-cap retail implementation, not about the anomaly in the population he measured.

The ranking universe is not complete, and the incompleteness is informative. Gross profitability can only be computed where the vendor's cost-of-revenue field is genuinely populated; where it is empty and gross profit merely restates revenue, the name is excluded from the GP/A ranking rather than scored on a number that would place it — wrongly — at the top of the cross-section. Coverage runs from 67% of the fundamentals universe at the 2006 anchor to 92% by 2020, and the excluded names cluster in health insurance and media. Arm B's ROE ranking has no such exclusion. An earlier compilation of this study was discarded when this exact defect was found live; the record described here contains zero such names at any rebalance of any window, verified name by name.

Both arms compute price returns without dividends, benchmark included, so every level understates a total-return equivalent and no figure should be compared against published total-return index numbers; the comparisons inside the study are internally consistent. Trading costs and taxes are zero in the sealed record — deliberately, to read the raw signal — and the closing section adds them back as disclosed arithmetic rather than as a re-run. The SPY figure quoted there (8.47%) is likewise supplementary arithmetic: computed from the same price source over the identical twenty windows, outside the sealed record. Membership is point-in-time via the vendor's constituent history; fills are same-bar closes on daily bars.

Twenty windows is twenty observations. The six down-benchmark windows carry standard errors near 3 points on their mean excess, and one of the six (2007) contributes a +15.4 observation to Arm B in a window whose benchmark fell 0.3% — the interpretation section discusses, and deliberately does not resolve, how much weight that observation deserves. Era splits rest on three to thirteen windows. Nothing in this study clears conventional significance except the bootstrap's refusal to separate the arms, which is powered by 5,011 paired days. Era claims are descriptions of what happened, not estimates of what recurs.

The risk model's projections are mildly optimistic: 88.8% band coverage against a 90% target (within one standard error) and VaR breach rates of 5.59% and 5.88% against 5% expected (about two standard errors hot) on 4,951 scored days. Fifteen of ninety-five rebalance segments were too short to score.

Finally, this study compares two definitions of quality in isolation. It says nothing about other quality formulations, about combining quality with value or momentum, about other holding counts or cadences, or about the short leg where much of the academic premium lives.

References

QuanterLab reference architecture
  1. Gelman, A., & Loken, E. (2013). The garden of forking paths: Why multiple comparisons can be a problem, even when there is no “fishing expedition.” Working paper, Columbia University.
  2. Harvey, C. R., Liu, Y., & Zhu, H. (2016). … and the Cross-Section of Expected Returns. Review of Financial Studies, 29(1), 5–68. doi:10.1093/rfs/hhv059
  3. Lo, A. W. (2002). The Statistics of Sharpe Ratios. Financial Analysts Journal, 58(4), 36–52. doi:10.2469/faj.v58.n4.2453
Author’s references?
  1. Novy-Marx, R. (2013). The Other Side of Value: The Gross Profitability Premium. Journal of Financial Economics, 108(1), 1-28.
  2. Novy-Marx, R., & Velikov, M. (2016). A Taxonomy of Anomalies and Their Trading Costs. Review of Financial Studies, 29(1), 104-147.
  3. Ball, R., Gerakos, J., Linnainmaa, J. T., & Nikolaev, V. (2015). Deflating Profitability. Journal of Financial Economics, 117(2), 225-248.
  4. Fama, E. F., & French, K. R. (2015). A Five-Factor Asset Pricing Model. Journal of Financial Economics, 116(1), 1-22.
  5. Asness, C. S., Frazzini, A., & Pedersen, L. H. (2019). Quality Minus Junk. Review of Accounting Studies, 24(1), 34-112.
  6. Hou, K., Xue, C., & Zhang, L. (2015). Digesting Anomalies: An Investment Approach. Review of Financial Studies, 28(3), 650-705.
  7. McLean, R. D., & Pontiff, J. (2016). Does Academic Research Destroy Stock Return Predictability? Journal of Finance, 71(1), 5-32.
  8. Harvey, C. R., Liu, Y., & Zhu, H. (2016). ...and the Cross-Section of Expected Returns. Review of Financial Studies, 29(1), 5-68.
  9. Arnott, R., Harvey, C. R., Kalesnik, V., & Linnainmaa, J. T. (2019). Alice's Adventures in Factorland: Three Blunders That Plague Factor Investing. Journal of Portfolio Management, 45(4), 18-36.
  10. Politis, D. N., & Romano, J. P. (1994). The Stationary Bootstrap. Journal of the American Statistical Association, 89(428), 1303-1313.

Appendix A  Reproducibility in QuanterLab

Each step is backed by a frozen run report. The study is re-derivable from the ledger below.

#CommitReportAnchorOOS window
1 0111e6623273 506 2006-01-01 2006-01-03 → 2006-12-29
2 054768c5e1cc 507 2007-01-01 2007-01-03 → 2007-12-31
3 86b63597dfb5 508 2008-01-01 2008-01-02 → 2008-12-31
4 27afb6b1ec86 509 2009-01-01 2009-01-02 → 2009-12-31
5 1bbaac16ed7f 510 2010-01-01 2010-01-04 → 2010-12-31
6 a7d8c4143614 511 2011-01-01 2011-01-03 → 2011-12-30
7 926794e481c8 512 2012-01-01 2012-01-03 → 2012-12-31
8 19cb35977f3e 513 2013-01-01 2013-01-02 → 2013-12-31
9 23ac850f2294 514 2014-01-01 2014-01-02 → 2014-12-31
10 f918f5bc975d 515 2015-01-01 2015-01-02 → 2015-12-31
11 667c55f142f9 516 2016-01-01 2016-01-04 → 2016-12-30
12 29dea9a651dd 517 2017-01-01 2017-01-03 → 2017-12-29
13 7d117988a4bc 518 2018-01-01 2018-01-02 → 2018-12-31
14 da30ac451cc5 520 2019-01-01 2019-01-02 → 2019-12-31
15 6edd71343baf 521 2020-01-01 2020-01-02 → 2020-12-31
16 59492f1368c7 522 2021-01-01 2021-01-04 → 2021-12-31
17 c735fea1ba13 523 2022-01-01 2022-01-03 → 2022-12-30
18 f230c39f70d2 524 2023-01-01 2023-01-03 → 2023-12-29
19 dd00fd1b023e 525 2024-01-01 2024-01-02 → 2024-12-31
20 3691d8d22b36 526 2025-01-01 2025-01-02 → 2025-12-31

Appendix A2  Sealed-hypothesis record

What this record does and does not establish. Every window in this study is historical: the data existed before the study began, so this is sequential sealing on past windows, not pre-registration in the clinical-trial sense, and no procedure could make it so. What the platform does enforce is order — each step's specification was frozen and hashed before that step was scored, and the walk cannot advance past a step that was never run or close one with a result registered for a different window. The two timestamp columns below are the evidence: read them together and each seal precedes its own run, and each run precedes the next seal. A study whose seals all post-date its runs would show it here. Wall-clock spacing between seals varies with the author's schedule and queue latency; the ordering, not the tempo, is the claim.

“A COMPARATIVE study — Arm A vs Arm B, walked on the same sealed out-of-sample windows. Arm A: S&P 500, rebalanced quarterly across the selected basket, and run out-of-sample from the anchor — whatever the design estimates from history — covariance, expected returns, rankings — is re-estimated at each anchor from pre-anchor data only, and the walk advances through sealed out-of-sample windows; its disposition is the realized forward path versus the benchmark. Arm B: S&P 500, rebalanced quarterly across the selected basket, and run out-of-sample from the anchor — whatever the design estimates from history — covariance, expected returns, rankings — is re-estimated at each anchor from pre-anchor data only, and the walk advances through sealed out-of-sample windows; its disposition is the realized forward path versus the benchmark. The arms differ in one manipulated variable, expressed as 2 paired settings on one node — Quality Factor — gp_to_assets: high → off; Quality Factor — roe: off → high. The contrast under test: whether Arm A generates better risk-adjusted returns than Arm B over the identical out-of-sample windows.”

The same hypothesis was sealed independently at every step — registered before each step's out-of-sample window was scored:

Table 4. Registration audit — one row per sealed step, with the time each specification was frozen and the time its window was scored. The hypothesis is identical on every row by design: it was sealed once and re-sealed unchanged at each anchor. Rows that differ would mean the specification moved mid-walk, which is the thing this record exists to rule out. The timestamps are the separate claim: each seal precedes its own run, and each run precedes the next seal.
#AnchorSealed at (UTC)Run completed (UTC)
1 2006-01-012026-08-04 17:05:29 2026-08-04 17:30:51
2 2007-01-012026-08-04 17:30:56 2026-08-04 17:48:59
3 2008-01-012026-08-04 17:49:04 2026-08-04 18:01:26
4 2009-01-012026-08-04 18:01:31 2026-08-04 18:17:12
5 2010-01-012026-08-04 18:17:18 2026-08-04 18:32:39
6 2011-01-012026-08-04 18:32:44 2026-08-04 18:48:48
7 2012-01-012026-08-04 18:48:53 2026-08-04 19:02:34
8 2013-01-012026-08-04 19:02:39 2026-08-04 19:40:21
9 2014-01-012026-08-04 19:40:27 2026-08-04 19:57:09
10 2015-01-012026-08-04 19:57:14 2026-08-04 20:14:36
11 2016-01-012026-08-04 20:14:41 2026-08-04 20:29:23
12 2017-01-012026-08-04 20:29:29 2026-08-04 20:47:30
13 2018-01-012026-08-04 20:47:35 2026-08-04 21:06:57
14 2019-01-012026-08-04 21:07:02 2026-08-04 21:25:24
15 2020-01-012026-08-04 21:25:29 2026-08-04 21:40:50
16 2021-01-012026-08-04 21:40:55 2026-08-04 21:59:37
17 2022-01-012026-08-04 21:59:42 2026-08-04 22:18:27
18 2023-01-012026-08-04 22:18:32 2026-08-04 22:36:56
19 2024-01-012026-08-04 22:37:02 2026-08-04 22:52:23
20 2025-01-012026-08-04 22:52:28 2026-08-04 23:11:11

Appendix B  Per-step diagnostics

What each step's run actually did beyond its return: capital allocation across lanes and regimes, the portfolio book's rebalancing and cost drag, and how positions were sized. Harvested from the frozen run reports — present where the circuit produced them.

Step 1 · 2006-01-03 → 2006-12-29

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 2.83% of 247 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2006-01-01 -7.2068% 3.7438% 17.177% 2.3001%yes 1.3719% 2 / 61
2006-04-01 -6.8797% 3.498% 15.103% 0.2157%yes 1.2497% 2 / 62
2006-07-01 -8.2072% 2.7496% 15.0897% 4.8631%yes 1.3976% 2 / 62
2006-10-01 -8.0001% 3.2125% 15.8691% 7.7786%yes 1.3972% 1 / 62
2007-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 4.86% of 247 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2006-01-01 -6.48% 3.7814% 16.2712% 3.9808%yes 1.3708% 3 / 61
2006-04-01 -9.3167% 2.7366% 16.4765% -2.0914%yes 1.5854% 4 / 62
2006-07-01 -6.5832% 4.7678% 17.5763% 0.0299%yes 1.4049% 4 / 62
2006-10-01 -6.3724% 4.0688% 15.7456% 6.1546%yes 1.284% 1 / 62
2007-01-01 no segment follows this rebalance — not scored

Step 2 · 2007-01-03 → 2007-12-31

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 11.74% of 247 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2007-01-01 -7.7965% 3.3105% 14.5573% 5.566%yes 1.3156% 3 / 60
2007-04-01 -7.7807% 3.1342% 15.4161% 2.4276%yes 1.3353% 2 / 62
2007-07-01 -6.9528% 4.1256% 16.5993% 1.344%yes 1.3059% 14 / 62
2007-10-01 -9.3713% 3.4243% 16.6643% -11.1291%no 1.5505% 10 / 63
2008-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 19 names held · selection: reselect · 0.0% in cash · 1 name dropped from the held union of 43 (weights renormalised onto the rest)

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 10.93% of 247 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2007-01-01 -6.1863% 4.818% 15.9308% 6.1857%yes 1.2833% 3 / 60
2007-04-01 -7.8607% 2.1103% 13.228% 9.6428%yes 1.234% 2 / 62
2007-07-01 -6.1969% 3.3966% 14.0358% -1.047%yes 1.1317% 11 / 62
2007-10-01 -7.5532% 3.1395% 13.9674% -4.4387%yes 1.2366% 11 / 63
2008-01-01 no segment follows this rebalance — not scored

Step 3 · 2008-01-02 → 2008-12-31

Arm A

Portfolio book — rebalanced quarterly · 4 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 20.88% of 249 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2008-01-01 -11.9754% 1.8021% 16.1804% -0.9684%yes 1.7094% 14 / 60
2008-04-01 -14.0518% 1.9882% 19.2152% 1.2542%yes 1.9278% 6 / 63
2008-07-01 -16.2434% 0.8136% 19.3835% -2.6018%yes 2.1555% 12 / 63
2008-10-01 -19.6596% -0.801% 20.2352% -24.402%no 2.5116% 20 / 63

Arm B

Portfolio book — rebalanced quarterly · 4 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 50.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 19.28% of 249 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2008-01-01 -8.6533% 3.1548% 15.201% -10.5135%no 1.502% 14 / 60
2008-04-01 -11.7291% 2.5432% 17.5642% -2.2315%yes 1.7619% 5 / 63
2008-07-01 -11.2887% 2.3962% 16.7111% -2.0215%yes 1.8056% 7 / 63
2008-10-01 -17.8386% 1.0594% 22.0629% -23.935%no 2.4167% 22 / 63

Step 4 · 2009-01-02 → 2009-12-31

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 5.65% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2009-01-01 -29.7463% -4.0999% 27.2349% -12.887%yes 3.186% 11 / 60
2009-04-01 -30.1142% -6.11% 26.3557% 21.9718%yes 3.5266% 3 / 62
2009-07-01 -30.2698% -3.3306% 30.223% 18.7156%yes 3.699% 0 / 63
2009-10-01 -28.6441% -0.9708% 33.5324% 9.894%yes 3.6588% 0 / 63
2010-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 3.23% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2009-01-01 -24.1403% -1.4825% 24.9202% -8.8456%yes 2.6915% 7 / 60
2009-04-01 -26.4935% -1.9605% 30.9795% 29.046%yes 3.6191% 1 / 62
2009-07-01 -26.7706% -1.5768% 28.8913% 17.2722%yes 3.3581% 0 / 63
2009-10-01 -26.3829% 0.2041% 32.7464% 4.7124%yes 3.5143% 0 / 63
2010-01-01 no segment follows this rebalance — not scored

Step 5 · 2010-01-04 → 2010-12-31

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 0.81% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2010-01-01 -26.7002% -0.1125% 32.3197% 10.7185%yes 3.4903% 0 / 60
2010-04-01 -24.1719% 2.2452% 38.1053% -11.9223%yes 3.534% 2 / 62
2010-07-01 -26.1487% 1.1988% 34.886% 9.8219%yes 3.6378% 0 / 63
2010-10-01 -25.294% 3.2903% 38.8023% 14.8046%yes 3.5139% 0 / 63
2011-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 1.21% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2010-01-01 -23.1672% 0.658% 28.653% 5.2916%yes 3.2405% 1 / 60
2010-04-01 -19.9464% 1.235% 28.1955% -9.5129%yes 2.8991% 2 / 62
2010-07-01 -22.651% 0.671% 28.0249% 7.6825%yes 3.0608% 0 / 63
2010-10-01 -22.1008% 1.3765% 28.9098% 5.808%yes 3.0766% 0 / 63
2011-01-01 no segment follows this rebalance — not scored

Step 6 · 2011-01-03 → 2011-12-30

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 6.05% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2011-01-01 -15.5413% 6.9114% 38.2845% 5.0486%yes 2.729% 0 / 61
2011-04-01 -8.7729% 10.2022% 33.2704% 2.7941%yes 2.0879% 1 / 62
2011-07-01 -8.0808% 8.1225% 25.38% -15.2359%no 1.8523% 12 / 63
2011-10-01 -12.7767% 3.5214% 22.9873% 16.6181%yes 2.107% 2 / 62
2012-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 8.06% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2011-01-01 -14.2254% 4.8397% 30.5181% 8.4582%yes 2.4069% 0 / 61
2011-04-01 -8.0114% 7.4718% 25.6746% 2.7777%yes 1.8312% 0 / 62
2011-07-01 -6.9821% 6.916% 21.3958% -15.3336%no 1.6557% 12 / 63
2011-10-01 -10.898% 3.3292% 19.9312% 12.5422%yes 1.8441% 8 / 62
2012-01-01 no segment follows this rebalance — not scored

Step 7 · 2012-01-03 → 2012-12-31

Arm A

Portfolio book — rebalanced quarterly · 4 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 1.63% of 246 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2012-01-01 -11.4078% 4.9879% 26.3665% 12.8984%yes 2.2254% 0 / 61
2012-04-01 -11.3118% 5.7514% 26.2264% -7.9609%yes 2.1532% 2 / 62
2012-07-01 -14.6947% 2.3222% 22.8635% 0.1551%yes 2.2909% 1 / 62
2012-10-01 -11.143% 3.8361% 23.0827% -1.0344%yes 2.017% 1 / 61

Arm B

Portfolio book — rebalanced quarterly · 4 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 2.03% of 246 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2012-01-01 -9.1984% 4.6036% 22.0743% 12.3109%yes 1.7982% 0 / 61
2012-04-01 -9.5927% 5.5877% 23.4283% -4.6146%yes 1.8886% 3 / 62
2012-07-01 -13.4654% 1.875% 20.0486% 5.1018%yes 1.9283% 0 / 62
2012-10-01 -10.7102% 3.6782% 22.0427% -2.5315%yes 1.809% 2 / 61

Step 8 · 2013-01-02 → 2013-12-31

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 1.21% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2013-01-01 -11.816% 4.41% 20.6852% 9.4802%yes 1.9637% 0 / 59
2013-04-01 -13.4287% 4.025% 22.9959% 6.9947%yes 2.0294% 2 / 63
2013-07-01 -12.6952% 4.4921% 23.1031% 3.6714%yes 2.0067% 0 / 63
2013-10-01 -11.9162% 5.2522% 23.8138% 2.4129%yes 1.9808% 1 / 63
2014-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 2.82% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2013-01-01 -10.5225% 3.7106% 17.7084% 10.9029%yes 1.6711% 2 / 59
2013-04-01 -10.8236% 3.6649% 18.9234% 4.1354%yes 1.6018% 4 / 63
2013-07-01 -11.6407% 3.3111% 19.1446% 5.02%yes 1.6778% 0 / 63
2013-10-01 -5.494% 6.3074% 18.352% 7.6394%yes 1.2513% 1 / 63
2014-01-01 no segment follows this rebalance — not scored

Step 9 · 2014-01-02 → 2014-12-31

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 2.82% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2014-01-01 -6.9042% 6.1131% 19.5113% -2.2527%yes 1.4465% 3 / 60
2014-04-01 -8.0555% 3.9866% 17.69% -2.8692%yes 1.5377% 1 / 62
2014-07-01 -9.7968% 3.0684% 16.3965% 0.3238%yes 1.5556% 1 / 63
2014-10-01 -6.7184% 4.3226% 15.5289% 11.0952%yes 1.335% 2 / 63
2015-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 4.84% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2014-01-01 -6.1151% 7.5542% 21.6912% 4.4808%yes 1.4235% 4 / 60
2014-04-01 -5.9499% 4.5203% 16.2274% 4.1727%yes 1.3763% 2 / 62
2014-07-01 -4.9747% 5.9072% 16.9157% 0.6099%yes 1.2429% 1 / 63
2014-10-01 -4.9284% 5.7243% 16.4781% 5.7215%yes 1.1523% 5 / 63
2015-01-01 no segment follows this rebalance — not scored

Step 10 · 2015-01-02 → 2015-12-31

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 9.68% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2015-01-01 -6.773% 3.9189% 14.6922% 0.5532%yes 1.2247% 7 / 60
2015-04-01 -6.3949% 3.8466% 15.2782% -4.4494%yes 1.2372% 1 / 62
2015-07-01 -7.9035% 2.9375% 13.9346% -8.5168%no 1.274% 9 / 63
2015-10-01 -9.8813% 1.2281% 12.5511% -1.4383%yes 1.2907% 7 / 63
2016-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 8.06% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2015-01-01 -5.7042% 5.1069% 16.0% 1.6074%yes 1.2523% 5 / 60
2015-04-01 -5.3844% 5.391% 17.4673% -3.9011%yes 1.3676% 2 / 62
2015-07-01 -7.5904% 3.4565% 14.6801% -7.8509%no 1.3241% 12 / 63
2015-10-01 -10.818% 1.2417% 13.658% 9.7835%yes 1.5491% 1 / 63
2016-01-01 no segment follows this rebalance — not scored

Step 11 · 2016-01-04 → 2016-12-30

Arm A

Portfolio book — rebalanced quarterly · 4 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 5.24% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2016-01-01 -11.1945% -0.0042% 11.3803% 10.9912%yes 1.3684% 7 / 60
2016-04-01 -8.8164% 3.1951% 15.5255% -2.4902%yes 1.4085% 3 / 63
2016-07-01 -10.5874% 1.8704% 14.741% -3.0906%yes 1.4719% 3 / 63
2016-10-01 -10.6953% 1.082% 14.495% 4.7607%yes 1.4728% 0 / 62

Arm B

Portfolio book — rebalanced quarterly · 4 constructions · 19 names held · selection: reselect · 0.0% in cash · 1 name dropped from the held union of 35 (weights renormalised onto the rest)

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 4.03% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2016-01-01 -9.1833% 2.4991% 14.4108% 2.8834%yes 1.3399% 4 / 60
2016-04-01 -8.6512% 3.7008% 16.4181% -0.5006%yes 1.4369% 3 / 63
2016-07-01 -9.5284% 2.5446% 14.9561% 0.8327%yes 1.3939% 2 / 63
2016-10-01 -11.4149% 1.0459% 15.3476% -0.0296%yes 1.54% 1 / 62

Step 12 · 2017-01-03 → 2017-12-29

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 1.21% of 247 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2017-01-01 -10.4684% 0.2946% 13.5248% 3.4104%yes 1.4309% 0 / 61
2017-04-01 -10.2512% 0.445% 12.4895% -0.6219%yes 1.3603% 2 / 62
2017-07-01 -11.1648% 0.1487% 12.9818% 3.5846%yes 1.4652% 1 / 62
2017-10-01 -10.4657% 1.1273% 14.3022% 13.0214%yes 1.4872% 0 / 62
2018-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 0.4% of 247 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2017-01-01 -10.5047% 0.5789% 14.2494% 3.2304%yes 1.5383% 0 / 61
2017-04-01 -10.9479% 1.6033% 16.0123% 2.1346%yes 1.6672% 1 / 62
2017-07-01 -9.1324% 2.2929% 15.2341% 5.7118%yes 1.4899% 0 / 62
2017-10-01 -8.2047% 2.3852% 14.2692% 5.6766%yes 1.283% 0 / 62
2018-01-01 no segment follows this rebalance — not scored

Step 13 · 2018-01-02 → 2018-12-31

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 9.72% of 247 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2018-01-01 -8.2291% 2.4059% 13.1328% -1.0302%yes 1.2826% 5 / 60
2018-04-01 -7.8873% 3.3416% 14.773% 10.4455%yes 1.2188% 1 / 63
2018-07-01 -6.9348% 3.7094% 15.6438% 11.751%yes 1.2793% 2 / 62
2018-10-01 -5.1842% 4.9901% 16.3251% -15.686%no 1.1852% 16 / 62
2019-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 13.36% of 247 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2018-01-01 -5.9872% 3.7057% 13.3757% 0.8156%yes 1.0956% 10 / 60
2018-04-01 -5.2902% 4.851% 15.045% 2.387%yes 1.0152% 4 / 63
2018-07-01 -5.6813% 4.5699% 16.0048% 6.2471%yes 1.244% 1 / 62
2018-10-01 -5.2416% 4.4729% 15.2487% -13.8829%no 1.133% 18 / 62
2019-01-01 no segment follows this rebalance — not scored

Step 14 · 2019-01-02 → 2019-12-31

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 4.44% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2019-01-01 -8.06% 3.9342% 16.1827% 16.7851%no 1.4043% 2 / 60
2019-04-01 -7.6029% 4.5345% 18.3511% -3.8457%yes 1.4681% 4 / 62
2019-07-01 -8.5372% 5.056% 19.2089% -3.2835%yes 1.5777% 5 / 63
2019-10-01 -8.9726% 4.6101% 18.7589% 11.8946%yes 1.8312% 0 / 63
2020-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 4.44% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2019-01-01 -9.0295% 2.5316% 14.3042% 21.3455%no 1.5478% 3 / 60
2019-04-01 -6.7168% 4.3643% 16.8379% 3.9041%yes 1.4743% 2 / 62
2019-07-01 -7.5051% 5.1292% 18.1522% -0.0731%yes 1.532% 4 / 63
2019-10-01 -8.7851% 4.1144% 17.4646% 9.0611%yes 1.5689% 2 / 63
2020-01-01 no segment follows this rebalance — not scored

Step 15 · 2020-01-02 → 2020-12-31

Arm A

Portfolio book — rebalanced quarterly · 4 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 50.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 8.84% of 249 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2020-01-01 -8.6103% 4.0596% 19.9023% -19.4125%no 1.8888% 14 / 61
2020-04-01 -18.5421% -0.3384% 22.0765% 32.2098%no 2.2724% 6 / 62
2020-07-01 -15.3537% 4.4385% 26.5008% 11.2757%yes 2.2984% 1 / 63
2020-10-01 -15.3499% 4.1875% 25.9148% 8.7124%yes 2.2338% 1 / 63

Arm B

Portfolio book — rebalanced quarterly · 4 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 50.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 8.03% of 249 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2020-01-01 -8.1432% 3.416% 17.6998% -20.7242%no 1.7177% 14 / 61
2020-04-01 -17.1632% 0.1217% 21.1467% 26.396%no 2.2353% 5 / 62
2020-07-01 -19.9022% 3.1152% 29.8323% 6.7353%yes 2.7582% 0 / 63
2020-10-01 -17.5412% 3.1053% 26.4142% 12.3724%yes 2.5332% 1 / 63

Step 16 · 2021-01-04 → 2021-12-31

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 0.81% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2021-01-01 -14.2845% 5.4955% 27.3973% 2.9795%yes 2.3345% 0 / 60
2021-04-01 -13.1614% 5.5063% 28.3321% 7.9135%yes 1.9403% 1 / 62
2021-07-01 -13.3806% 7.1276% 30.0397% -0.5581%yes 2.2395% 0 / 63
2021-10-01 -13.1725% 6.6586% 28.6716% 15.0346%yes 1.9951% 1 / 63
2022-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 1.61% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2021-01-01 -16.93% 4.3412% 28.3543% 10.1695%yes 2.4514% 1 / 60
2021-04-01 -15.8582% 4.5823% 30.1528% 3.731%yes 2.0729% 1 / 62
2021-07-01 -14.8608% 5.7881% 28.96% -0.8405%yes 2.0205% 1 / 63
2021-10-01 -12.6203% 7.9097% 30.8141% 8.2858%yes 1.96% 1 / 63
2022-01-01 no segment follows this rebalance — not scored

Step 17 · 2022-01-03 → 2022-12-30

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 8.1% of 247 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2022-01-01 -10.799% 8.0852% 33.2892% -8.8311%yes 2.0727% 1 / 61
2022-04-01 -12.6918% 7.225% 34.1853% -19.5602%no 2.2956% 10 / 61
2022-07-01 -11.5222% 6.9809% 27.2124% -4.8104%yes 2.1869% 4 / 63
2022-10-01 -13.2385% 2.7206% 21.7347% 17.3439%yes 2.1824% 5 / 62
2023-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 75.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 8.5% of 247 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2022-01-01 -11.7741% 6.6641% 31.2152% -9.2821%yes 2.0021% 3 / 61
2022-04-01 -14.0748% 6.1974% 33.8218% -15.39%no 2.2826% 9 / 61
2022-07-01 -10.9272% 4.8399% 21.6429% -5.8393%yes 1.8775% 4 / 63
2022-10-01 -11.7273% 3.1851% 20.7263% 10.7555%yes 1.8482% 5 / 62
2023-01-01 no segment follows this rebalance — not scored

Step 18 · 2023-01-03 → 2023-12-29

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 0.41% of 246 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2023-01-01 -12.3037% 3.918% 25.0681% 3.6002%yes 2.245% 1 / 61
2023-04-01 -13.554% 2.2695% 22.8652% 3.7454%yes 2.1102% 0 / 61
2023-07-01 -15.2352% 1.1997% 20.9457% -9.9171%yes 2.0384% 0 / 62
2023-10-01 -16.8274% -0.6939% 18.6913% 16.9554%yes 2.0388% 0 / 62
2024-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 2.03% of 246 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2023-01-01 -11.05% 4.5015% 24.5958% 3.5725%yes 1.8883% 3 / 61
2023-04-01 -12.7229% 2.6296% 22.4858% -2.109%yes 2.0403% 2 / 61
2023-07-01 -14.4504% 1.9127% 21.529% -2.7971%yes 2.0652% 0 / 62
2023-10-01 -15.1354% 1.0134% 20.3569% 10.3048%yes 2.0043% 0 / 62
2024-01-01 no segment follows this rebalance — not scored

Step 19 · 2024-01-02 → 2024-12-31

Arm A

Portfolio book — rebalanced quarterly · 4 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 1.61% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2024-01-01 -16.557% 1.3132% 20.8463% 10.3944%yes 2.1434% 0 / 60
2024-04-01 -14.0836% 3.453% 24.7036% -3.0048%yes 2.0284% 0 / 62
2024-07-01 -12.4508% 5.3242% 24.6661% 9.2459%yes 1.8203% 3 / 63
2024-10-01 -8.844% 6.0945% 21.8449% 3.6414%yes 1.6058% 1 / 63

Arm B

Portfolio book — rebalanced quarterly · 4 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 1.21% of 248 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2024-01-01 -15.6242% 1.6826% 20.4657% 13.9173%yes 2.0934% 0 / 60
2024-04-01 -11.3566% 2.8172% 19.3601% 3.6478%yes 1.6945% 0 / 62
2024-07-01 -9.978% 5.1413% 21.1356% 11.7919%yes 1.6516% 2 / 63
2024-10-01 -7.3301% 6.085% 20.0074% 3.8945%yes 1.4394% 1 / 63

Step 20 · 2025-01-02 → 2025-12-31

Arm A

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 8.13% of 246 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2025-01-01 -5.9109% 6.7003% 18.8549% -4.2452%yes 1.2906% 7 / 59
2025-04-01 -7.5819% 4.0437% 18.4088% 7.8519%yes 1.6113% 9 / 61
2025-07-01 -9.5397% 5.899% 22.2673% 4.2581%yes 1.6991% 3 / 63
2025-10-01 -9.5239% 7.0031% 24.6966% 3.8113%yes 1.8416% 1 / 63
2026-01-01 no segment follows this rebalance — not scored

Arm B

Portfolio book — rebalanced quarterly · 5 constructions · 20 names held · selection: reselect · 0.0% in cash

Projection accuracy — realized outcome fell inside the P5–P95 cone in 100.0% of 4 scored rebalances (an honest 90% band would contain ~90.0%) · 95% VaR breached on 8.54% of 246 days (expected ~5.0%)

Rebalance P5 Median P95 Realized In band VaR 95 (1d) Breaches
2025-01-01 -5.1597% 7.1332% 18.9383% -3.4488%yes 1.338% 10 / 59
2025-04-01 -5.0759% 4.5827% 16.2506% 5.513%yes 1.248% 8 / 61
2025-07-01 -6.5312% 5.8387% 18.5436% 2.0593%yes 1.1008% 0 / 63
2025-10-01 -6.1507% 7.9904% 22.7389% -1.5772%yes 1.3423% 3 / 63
2026-01-01 no segment follows this rebalance — not scored
QuanterLab · Study d3890b5b87ab · compiled August 05, 2026. Point-in-time constituents and hypothesis-registration timestamps are enforced by the platform; transaction costs are not modelled in this study. This report is generated from the frozen study artifact and is reproducible from the ledger above. Educational research, not investment advice: every result on this page is simulated, and nothing here is a recommendation to buy or sell any security.

Run a study like this one

Everything above was produced inside QuanterLab — the sealed registration, the walk, the statistics and the paper itself. Build the circuit on a canvas, register the hypothesis before you score it, and the platform keeps you honest about the rest.

The lab is in private beta and opens in September 2026. Reading the research needs no account — follow it and we'll tell you when the next study publishes.