Use your browser's “Save as PDF” destination for a paper-ready file.
QuanterLab
QuanterLab Research · Pre-registered Walk-Forward Study

2 arms

A pre-registered, point-in-time walk-forward investigation with a multiple-testing (deflated Sharpe) correction.
Universe · Nasdaq 100 (point-in-time constituents)
Method · Comparative — Arm A vs Arm B
Manipulated variable · Regime
Step size · 1 year per forward window
In-sample · 2 years before each anchor
Out-of-sample span · 2011-01-03 → 2025-12-31
Compiled · July 24, 2026
Abstract

We pre-register and walk a strategy forward across Nasdaq 100 in 1-year steps, each tested out-of-sample on data the strategy had never touched, with survivorship-bias-free constituents reconstructed as of every anchor. This study walks the SAME sealed windows on two arms — Arm A and Arm B — identical in every respect except one declared variable: Regime. Paired date by date inside each window (3758 common out-of-sample observations across 15 windows), Arm A compounds at 17.4% a year against 15.4% for Arm B — a gap of 2.0 pp. A seeded block bootstrap of the paired return differences puts the probability that Arm A genuinely beats Arm B at 69.8%. Every attempt made along the way — 15 registered circuits, pruned candidates included — is counted in the correction, so the headline number reflects the true cost of the search.

1  Methodology

The study is sealed before it starts: the universe, the step size and the method are fixed at registration and cannot be changed while the walk is running, and the anchor advances only forward. This removes the two commonest ways a backtest flatters itself — moving the test window until the numbers look good, and changing the rules with hindsight.

At each step the strategy is fitted on the 2 years of history ending at the anchor and then tested on the subsequent out-of-sample window it had never seen. Index membership is reconstructed as of the anchor date from the exchange's dated constituent change-log, so names that were later removed or delisted still compete on the dates they actually traded and there is no survivorship bias.

The inferential statistic is the Deflated Sharpe Ratio (Bailey and López de Prado, 2014), which asks: given that 15 configurations were tried, what is the probability that the observed Sharpe is genuinely positive rather than the best of many noisy draws? The number of trials is frozen at compile time and taken as the larger of the registered circuit count and the number of logged runs, so it can never understate the search.

2  Results

2.1  Headline

Arm A — pooled deflated
76%
SR 0.68 · 3759 OOS bars
Arm B — pooled deflated
82%
SR 0.69 · 3759 OOS bars
P(Arm A beats Arm B)
69.8%
3758 paired bars · CAGR gap +2.0 pp
-0.48x7.42x15.32x
Figure 1. Both arms stitched through the identical windows —  Arm A (+985.3%),  Arm B (+736.4%), benchmark grey (+626.6%). Dotted verticals mark the step boundaries; the dashed horizontal is break-even.

2.2  Per-step results

Table 1. One row per surviving step. The deflated Sharpe uses the cumulative trial count at that step, so it falls as the search widens — the decay is the point.
#Out-of-sample windowN Arm A SRdefl. Arm B SRdefl.
1 2011-01-03 → 2011-12-30 1 0.90 83% 1.01 84%
2 2012-01-03 → 2012-12-31 2 -0.55 13% 0.62 54%
3 2013-01-02 → 2013-12-31 3 2.23 94% 1.57 84%
4 2014-01-02 → 2014-12-31 4 0.46 28% -0.03 14%
5 2015-01-02 → 2015-12-31 5 1.58 66% 0.55 26%
6 2016-01-04 → 2016-12-30 6 1.12 46% 0.27 15%
7 2017-01-03 → 2017-12-29 7 1.35 49% 0.89 30%
8 2019-01-02 → 2019-12-31 8 2.72 90% 2.35 90%
9 2019-01-02 → 2019-12-31 9 2.72 89% 2.35 89%
10 2020-01-02 → 2020-12-31 10 1.41 44% 1.06 30%
11 2021-01-04 → 2021-12-31 11 -0.14 4% -0.64 1%
12 2022-01-03 → 2022-12-30 12 -0.12 4% -0.60 1%
13 2023-01-03 → 2023-12-29 13 2.48 79% 3.24 97%
14 2024-01-02 → 2024-12-31 14 1.84 55% 1.23 31%
15 2025-01-02 → 2025-12-31 15 0.03 4% 1.11 25%
0.28x1.73x3.18x
Figure 2. Every step's out-of-sample curve overlaid, each rebased to 1× at its own start. Read alongside Table 1: consistent shape across steps is the walk-forward's evidence; a single lucky leg is not.

2.3  Trial accounting

The search registered 15 circuits in total (pruned candidates included); the correction uses this registered-circuit count, N = 15. The pooled deflated Sharpe is the inferential headline; the per-step deflated track in Table 1 is illustrative, since near-identical variants are correlated and per-step deflation over-penalises.

2.4  The comparison

Both arms trade the same sealed windows, so their returns can be PAIRED: inside each window the two return series are inner-joined date by date and the difference rArm A − rArm B is the object under test. Because this is ONE pre-registered contrast — declared before any window was scored — the paired statistic needs no multiple-testing deflation; the per-arm pooled numbers above are still deflated by the trial count as usual.

Table 2. Window-by-window paired comparison. Δ is the growth gap (Arm A − Arm B) over the window's paired dates.
#WindowPaired bars Arm AArm B ΔLeader
1 2011-01-04 → 2011-12-30 251 +28.2% +21.0% +7.3 pp Arm A
2 2012-01-04 → 2012-12-31 249 -37.3% +9.9% -47.2 pp Arm B
3 2013-01-03 → 2013-12-31 251 +22.1% +22.9% -0.7 pp Arm B
4 2014-01-03 → 2014-12-31 251 +8.4% -1.7% +10.1 pp Arm A
5 2015-01-05 → 2015-12-31 251 +35.6% +7.1% +28.5 pp Arm A
6 2016-01-05 → 2016-12-30 251 +24.2% +3.4% +20.8 pp Arm A
7 2017-01-04 → 2017-12-29 250 +22.9% +19.1% +3.8 pp Arm A
8 2019-01-03 → 2019-12-31 251 +45.4% +12.4% +33.0 pp Arm A
9 2019-01-03 → 2019-12-31 251 +45.4% +12.4% +33.0 pp Arm A
10 2020-01-03 → 2020-12-31 252 +52.5% +32.9% +19.7 pp Arm A
11 2021-01-05 → 2021-12-31 251 -16.3% -29.7% +13.4 pp Arm A
12 2022-01-04 → 2022-12-30 250 -17.6% -29.9% +12.3 pp Arm A
13 2023-01-04 → 2023-12-29 249 +56.8% +183.1% -126.3 pp Arm B
14 2024-01-03 → 2024-12-31 251 +49.0% +39.6% +9.3 pp Arm A
15 2025-01-03 → 2025-12-31 249 -5.1% +20.9% -26.0 pp Arm B

Paired Sharpe of the difference track: 0.13 · block bootstrap (2000 paths, block 10, seed 1234): P(Arm A beats Arm B) = 69.8%.

3  Strategy specification and evolution

Each step's strategy is a circuit of platform primitives, frozen at registration. The table lists the full component set per step and — where the hypothesis evolved — exactly what changed against the previous step: components added, removed, or re-tuned.

Table 3. The frozen component specification per step. Parameters shown are the reproduction-relevant ones; the complete circuit is preserved in the study ledger (Appendix A).
#Components (primitives & key parameters)Evolution vs previous step
1 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto)
2 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
3 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
4 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
5 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
6 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
7 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
8 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
9 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
10 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
11 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
12 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
13 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
14 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward
15 signal module · static optimizer · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · composite score · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · top n · universe (preset=nasdaq100) · price loader · filter hurst (direction=keep_below, threshold=0.5) · filter ou halflife (direction=keep_below, threshold=15) · filter adf (direction=keep_below, threshold=0.05) · composite score · top n · signal module · walkforward rolling · backtest validator (horizon=1y, sizing=equal, sizing_mode=kelly, kelly_variant=full) · per regime optimizer · regime gmm (n_regimes=auto) unchanged — carried forward

4  Pre-registration record

The integrity of a walk-forward rests on registering each hypothesis before its out-of-sample window is scored. The order below is the order in which the hypotheses were sealed.

Table 4. The registered hypothesis for each step, with its registration timestamp.
#Registered hypothesisAnchorRegistered at
1 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2011-01-01 2026-07-24
2 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2012-01-01 2026-07-24
3 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2013-01-01 2026-07-24
4 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2014-01-01 2026-07-24
5 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2015-01-01 2026-07-24
6 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2016-01-01 2026-07-24
7 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2017-01-01 2026-07-24
8 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2018-01-01 2026-07-24
9 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2019-01-01 2026-07-24
10 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2020-01-01 2026-07-24
11 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2021-01-01 2026-07-24
12 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2022-01-01 2026-07-24
13 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2023-01-01 2026-07-24
14 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2024-01-01 2026-07-24
15 “NASDAQ 100, selected by statistical / factor criteria, traded via enter long when RSI (length=7) crosses below 30; exit when RSI (length=7) crosses above 70, conditioned on the wired regime classifier, with parameters tuned in-sample to sharpe, and validated out-of-sample via walk-forward — the rule set is re-optimized on a rolling in-sample window and tested on the unseen window after it, is expected to generate positive risk-adjusted returns over the forward test window.” 2025-01-01 2026-07-24

5  Discussion

A garden of forking paths (Gelman and Loken, 2013) is not cheating — it is what honest research feels like from the inside. At each step there were several defensible things to try, and trying them is the work. What separates research from a hunt for a flattering backtest is that the paths not taken are still counted. Here they are: the trial count stands at 15; a pooled inferential number awaits a fully-logged walk.

Pooled deflated Sharpe is the inferential number; the per-step DSR track is illustrative (near-identical variants are correlated, so per-step N over-deflates). N is a conservative upper bound on the multiple-testing penalty, frozen at compile time and taken as the larger of the lineage trial count and the explored-runs count so it cannot understate the search. Integrity rests on pre-registration: each step's submitted_at is the registration order, auditable against when its OOS was scored.

References

  1. Bailey, D. H., & López de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality. Journal of Portfolio Management, 40(5), 94–107.
  2. Gelman, A., & Loken, E. (2013). The garden of forking paths: Why multiple comparisons can be a problem, even when there is no “fishing expedition.” Working paper, Columbia University.
  3. Harvey, C. R., Liu, Y., & Zhu, H. (2016). … and the Cross-Section of Expected Returns. Review of Financial Studies, 29(1), 5–68.
  4. Lo, A. W. (2002). The Statistics of Sharpe Ratios. Financial Analysts Journal, 58(4), 36–52.

Appendix A  Reproducibility

Each step is backed by a frozen run report. The study is re-derivable from the ledger below.

#CommitReportAnchorOOS window
1 66ff2b275849 268 2011-01-01 2011-01-03 → 2011-12-30
2 b62735fc09f0 269 2012-01-01 2012-01-03 → 2012-12-31
3 7aede40fa0dc 270 2013-01-01 2013-01-02 → 2013-12-31
4 b448d206a540 271 2014-01-01 2014-01-02 → 2014-12-31
5 0c51cee24dce 272 2015-01-01 2015-01-02 → 2015-12-31
6 bcafb4f5cb51 273 2016-01-01 2016-01-04 → 2016-12-30
7 69b71a995d79 274 2017-01-01 2017-01-03 → 2017-12-29
8 caf74981ddf2 276 2018-01-01 2019-01-02 → 2019-12-31
9 403f59583278 278 2019-01-01 2019-01-02 → 2019-12-31
10 23aa21afe954 279 2020-01-01 2020-01-02 → 2020-12-31
11 f82540f9abd0 280 2021-01-01 2021-01-04 → 2021-12-31
12 a52d694ce260 281 2022-01-01 2022-01-03 → 2022-12-30
13 b66d0e51cd52 282 2023-01-01 2023-01-03 → 2023-12-29
14 92bba7cc69e9 283 2024-01-01 2024-01-02 → 2024-12-31
15 e6ab5dcff61b 284 2025-01-01 2025-01-02 → 2025-12-31

Appendix B  Per-step diagnostics

What each step's run actually did beyond its return: capital allocation across lanes and regimes, the portfolio book's rebalancing and cost drag, and how positions were sized. Harvested from the frozen run reports — present where the circuit produced them.

Step 1 · 2011-01-03 → 2011-12-30

Position sizing — sizing: full_kelly

Step 2 · 2012-01-03 → 2012-12-31

Position sizing — sizing: full_kelly

Step 3 · 2013-01-02 → 2013-12-31

Position sizing — sizing: full_kelly

Step 4 · 2014-01-02 → 2014-12-31

Position sizing — sizing: full_kelly

Step 5 · 2015-01-02 → 2015-12-31

Position sizing — sizing: full_kelly

Step 6 · 2016-01-04 → 2016-12-30

Position sizing — sizing: full_kelly

Step 7 · 2017-01-03 → 2017-12-29

Position sizing — sizing: full_kelly

Step 8 · 2019-01-02 → 2019-12-31

Position sizing — sizing: full_kelly

Step 9 · 2019-01-02 → 2019-12-31

Position sizing — sizing: full_kelly

Step 10 · 2020-01-02 → 2020-12-31

Position sizing — sizing: full_kelly

Step 11 · 2021-01-04 → 2021-12-31

Position sizing — sizing: full_kelly

Step 12 · 2022-01-03 → 2022-12-30

Position sizing — sizing: full_kelly

Step 13 · 2023-01-03 → 2023-12-29

Position sizing — sizing: full_kelly

Step 14 · 2024-01-02 → 2024-12-31

Position sizing — sizing: full_kelly

Step 15 · 2025-01-02 → 2025-12-31

Position sizing — sizing: full_kelly

Appendix C  The mathematics of the circuit

Every primitive this study wired, with the mathematics it actually computes — the same formulas the execution engine runs, as documented in the platform's per-primitive "Explain the math". Nothing here is illustrative; it is the calculation.

C.1  Signal Module

The entry / exit rule — turn indicators into a per-bar trade signal.

Composes indicators (RSI, moving averages, …) with comparison and logic operators into a rule that says enter, exit, or hold each bar. The rule is emitted as a portable config the optimizer tunes and the walk-forward validates — so what you design is exactly what gets traded.

Boolean rule → position state
\text{signal}_t = \begin{cases} +1 & \text{entry rule true} \\ 0 & \text{exit rule true} \\ \text{hold} & \text{otherwise}\end{cases}
e.g. enter when RSI < 30, exit when RSI > 50.

C.2  Static Optimizer

Grid-search one best parameter set over the whole in-sample window.

Sweeps a grid of parameter combinations, backtests each on the in-sample data, and keeps the single combination that scores best on your objective (Sharpe by default). One rule for the whole period — no time variation.

Argmax over the grid
\boldsymbol\theta^* = \arg\max_{\boldsymbol\theta\in\text{grid}} \;\mathcal O\big(\text{backtest}(\boldsymbol\theta)\big)
Objective 𝒪 ∈ {Sharpe, Calmar, Sortino, total return, profit factor, win rate}.
Default objective — Sharpe
\text{Sharpe} = \frac{\bar r - r_f}{\sigma_r}\,\sqrt{252}

C.3  Walkforward Rolling

Walk-forward with a sliding window — fixed-width, always recent.

Same out-of-sample discipline, but the training window is a fixed width that slides forward — each fold trains on the SAME amount of data, just more recent. Better when old regimes hurt and only recent behaviour matters.

Sliding folds
\text{fold}_k:\quad [\,\text{split}_k - W,\;\text{split}_k\,]\ \text{train} \;\to\; [\,\text{split}_k,\;\text{end}_k\,]\ \text{test}
Fixed train width W slides forward. Non-overlapping tests by default (step = test).

C.4  Backtest Validator

Forward-test the winning rule on unseen, out-of-sample data.

Takes the wired rule config (the Walk-Forward validated config wins, else the optimized config, else the raw signal config) and trades it FORWARD on the out-of-sample window to the right of the anchor — data it never saw during optimization — re-deriving the regime as-of each bar. It produces the true out-of-sample equity curve, trades and statistics: the signal-path twin of the Portfolio Forward Test, not an in-sample replay.

Apply the frozen rule forward (OOS)
E_t = E_{t-1}\,(1 + r_t),\qquad \text{Sharpe} = \frac{\bar r - r_f}{\sigma_r}\sqrt{252}
Config frozen from optimization / walk-forward, then replayed bar-by-bar on the forward window it has never seen, with cost + risk overlays applied.

C.5  Universe

The starting set of tickers — resolved point-in-time so there is no survivorship bias.

Before any math, you need a list of stocks. An index preset (S&P 500, Nasdaq-100, Dow 30) is reconstructed as it stood ON your anchor date by replaying the historical add/drop change-log backwards — so a 2018 backtest sees the 2018 membership, not today's winners.

Point-in-time membership

Start from today's constituents and un-apply every membership change after the anchor t:

\mathcal{U}(t) = \mathcal{U}_{\text{now}} \;\ominus\; \{\text{adds after } t\} \;\oplus\; \{\text{drops after } t\}
Constituents resolved from the index change-log; the same point-in-time set the factor + screening modules use.

C.6  Price Loader

Bulk OHLCV fetch for the whole universe — point-in-time, no future bars.

Momentum, volatility, trend — every price-based metric needs history. This loads open/high/low/close/volume for all names in parallel, clipped so nothing after the anchor can leak in. The lookback window is derived automatically from the deepest metric you wired.

The window is derived, not guessed

It loads exactly enough history for the hungriest downstream metric plus a warm-up buffer:

W = \max_k(\text{lookback}_k) + \text{buffer}, \qquad \text{bars} \le \text{anchor } t

C.7  Filter Hurst

The Hurst exponent — is this series trending, random, or mean-reverting?

Rescaled-range (R/S) analysis measures how the spread of a series grows as you look over longer windows. A random walk spreads like √n; trends spread faster, mean-reversion slower. The exponent H captures which.

Rescaled range scales as a power of the window
\mathbb{E}\!\left[\tfrac{R(n)}{S(n)}\right] \sim c\,n^{H} \;\;\Longrightarrow\;\; H = \frac{\log\!\big(R/S\big)}{\log n}
R = range of the cumulative deviation, S = standard deviation, over log-spaced windows n (10 → min(N/4, 200)).
Bias correction
H = \operatorname{clip}\big(\text{slope} - 0.06,\ 0,\ 1\big)
The R/S estimator runs slightly high on finite samples; the −0.06 correction (clamped to [0,1]) matches the Indicator-Strategies scanner exactly — the same ticker reads the same H in both.
Reading it

H < 0.5 → mean-reverting · H ≈ 0.5 → random walk · H > 0.5 → trending. The metric is attached to each stock; ranking + the cut happen in Composite Σ / Top-N.

The exact compute_hurst of the Indicator-Strategies MR scanner (shared scanner_metrics) — same algorithm, same bias correction, same number.

C.8  Composite Score

The composite — turn many metrics into one 0–100 score per stock.

Each wired metric is ranked across all stocks into a 0–100 percentile (you choose whether high or low is "good"), then the percentiles are weight-averaged. Ranking instead of raw values means no single unit dominates and outliers can't blow it up. It scores; it does not drop.

Per-metric cross-sectional percentile
\text{pct}_k(i) = 100 \cdot \frac{\operatorname{rank}_k(i)}{N}
Direction-aware: "low is good" (e.g. Hurst) inverts the rank.
Weighted blend
\text{score}_i = \frac{\sum_k w_k\,\text{pct}_k(i)}{\sum_k w_k} \in [0,100]
Metrics fanned in PARALLEL all contribute; a name missing a metric just omits that term.

C.9  Filter Ou Halflife

Ornstein–Uhlenbeck half-life — how many days a deviation takes to decay by half.

First remove the long-term drift (an OLS trend line fitted to log-price), then fit an AR(1) to what remains. The autoregressive coefficient β says how fast deviations from trend get pulled back; convert it to a half-life in days. Roughly 5–40 days is the tradeable sweet spot for mean reversion.

Detrend log-price first
\log P_t = a + b\,t + \varepsilon_t \quad\Longrightarrow\quad x_t = \log P_t - (a + b\,t)
Without detrending, a drifting stock looks like it never reverts — the AR(1) must see deviations from trend, not the trend itself.
AR(1) on the detrended residual
x_t = \alpha + \beta\,x_{t-1} + \varepsilon_t
Half-life from the decay rate
\text{half-life} = \frac{\ln 2}{\lvert \ln \beta \rvert}
β close to 1 → very slow reversion (long half-life); small β → fast. No mean reversion detected (β outside (0,1)) reports 999 — it ranks last and fails any "keep below" gate.
The exact compute_halflife of the Indicator-Strategies MR scanner (shared scanner_metrics) — detrended log AR(1), same number in both modules.

C.10  Filter Adf

Augmented Dickey–Fuller — a statistical test for stationarity (mean reversion).

Regress the change in price on its lagged level. If the level coefficient is significantly negative, deviations get pulled back — the series is stationary (mean-reverting). A low p-value rejects the "random walk" null.

The test regression
\Delta x_t = \gamma\,x_{t-1} + \sum_{i=1}^{p}\delta_i\,\Delta x_{t-i} + \varepsilon_t
Test H₀: γ = 0 (unit root / random walk) vs γ < 0 (stationary).
Reading it

p < 0.05 → reject the random walk → mean-reverting. The metric carried is the p-value (or the ADF statistic).

C.11  Top N

Keep the best N — rank, then cut.

Sort the survivors by the Composite Σ (or, if none is wired, the last metric in the chain) and keep the top (or bottom) N. The final narrowing from a scored list to a committed basket.

Order statistic cut
\text{Top-}N = \{\, i : \operatorname{rank}(\text{score}_i) \le N \,\}
"Keep highest" for momentum; "keep lowest" for e.g. Hurst (mean reversion).

C.12  Per Regime Optimizer

A separate best parameter set for each regime.

Partitions the in-sample window by a wired regime classifier and optimizes independently within calm, choppy and stressed. The strategy then switches parameters as the regime switches — different behaviour for different weather.

Argmax per regime
\boldsymbol\theta^*_g = \arg\max_{\boldsymbol\theta}\;\mathcal O\big(\text{backtest}(\boldsymbol\theta)\mid \text{regime}=g\big), \quad g\in\{\text{calm},\text{choppy},\text{stressed}\}

C.13  Regime Gmm

Gaussian mixture — cluster days into regimes, count chosen by BIC.

Treats each day as a point in (return, volatility) space and fits a Gaussian mixture; the Bayesian Information Criterion picks how many regimes the data actually support. States are vol-sorted and short runs de-noised. The rigorous detector the Per-Regime and Regression optimizers were designed around.

Mixture density + model selection
p(\mathbf x) = \sum_{k=1}^{K}\pi_k\,\mathcal N(\mathbf x\mid\boldsymbol\mu_k,\boldsymbol\Sigma_k), \qquad K^* = \arg\min_K \text{BIC}(K)
Features x = (log-return, realized-vol). K ∈ {2,3,4} or auto.
QuanterLab · Study 85dcdc0f76ea · compiled July 24, 2026. Point-in-time constituents, transaction costs and pre-registration timestamps are enforced by the platform. This report is generated from the frozen study artifact and is reproducible from the ledger above.
Run a study like this yourself

This paper was produced end-to-end in QuanterLab: the pre-registration, the sealed walk, the point-in-time data, the statistics and the document you are reading. The platform is in closed beta.

Request access More research RSS