Monte Carlo · Methodology

Rolling-Origin Validation

This is the page that explains why every probability on this platform is described as conditional. A research harness scores the return generators out of sample against published thresholds. It has run, it works, and its published verdict is negative: no generator has been accepted.

The published result, stated plainly

mechanism_verified
true
statistically_qualified
false
product_approved
false
independent_review_status
pending

The accepted and default-eligible sets are empty. That is not the same as every candidate having failed: under the current policy the best-calibrated arms carry no rejecting gate at any horizon and are held short of acceptance by the panel's measurement power, not by their own performance. The three states below are the distinction, and collapsing them is the most common way to misread this page.

The stationary block bootstrap is no longer the default and is no longer selectable. It was retired because the study measured it as rejected at every horizon, over-dispersing by 27% to 56%. The current default is a variance-targeted GJR-GARCH arm, chosen for the best measured calibration available, which is a change of default and not a promotion: it is insufficient_evidence, not accepted.

This is a completed validation study with a negative promotion result. It is not a calibrated-forecast claim, and it is published precisely so that the absence of one is visible rather than inferred from silence.

Three verdicts, and why the middle one matters

Every gate returns one of three states. They are not degrees of the same thing, and reading the middle one as failure is wrong about exactly the arms that perform best.

accepted
The gate passed at that horizon. The evidence was sufficient and the model met the threshold.
insufficient_evidence
The panel cannot resolve the gate. This says nothing about the model: a perfectly calibrated generator and a badly miscalibrated one return the same state when there is not enough independent data to tell them apart. It is a statement about the measurement.
rejected
The model failed the gate. The evidence was sufficient and the threshold was not met.

Long horizons are permanently evidence-bounded

At one day the study has 5,900 independent observations and full power, which is why the h1 verdicts discriminate: the retired default is rejected there on calibration alone. At one year it has 269, and the tail gates need thousands.

This does not improve by waiting. The number of independent observations at a given horizon is the calendar span divided by that horizon, so a one-year gate needs roughly 5,000 years of history to resolve its 1% expected-shortfall floor. Collecting another decade moves 269 to about 279. Any long-horizon number on this platform is therefore reported as conditional and always will be, and no future run will change that by itself.

That is why an arm can carry zero rejecting gates and still not be accepted, and why absence of rejection is never read as acceptance anywhere in this study.

What the harness does

The harness answers one question at four horizons, 1, 21, 63 and 252 trading sessions: standing at a historical date and knowing only what was knowable then, how well does each generator's forecast distribution describe what actually happened next?

  • It freezes and digest-verifies an approved research panel of Yahoo adjusted closes, and records complete data provenance so a replay is reproducible rather than merely repeatable.
  • It builds causal expanding-window folds, so every fit sees only data from before its origin.
  • It routes candidate paths through the production simulation and planning ledgers, not a research reimplementation. A generator is scored as the product would run it.

The developer study contains 628 folds over 161 origins, 3,140 scored model-horizon forecasts with zero failures, and all 3,725 expected planning replay rows with zero failures. The deterministic artifacts, a JSON evidence file and a 104-page PDF, are digest-pinned.

Leakage safety

A backtest that lets any future information reach a fit produces optimistic results that nothing downstream can undo. Two mechanisms prevent it, and they address different failures.

Exact information-interval purging

A training observation is dropped when its information interval overlaps the test target's. For a horizon- target evaluated over , any training label whose own interval intersects that range carries information about the test outcome and is removed. This is exact rather than approximate: overlapping labels are the mechanism by which information leaks in overlapping-horizon designs.

One-horizon embargo

After the test interval, a further sessions are embargoed from training. Purging handles direct overlap; the embargo handles serial dependence that survives the boundary, because a volatility cluster does not stop at the edge of a labelling window.

The scoring surface

Every score is strictly proper: it is minimized in expectation only by the true forecast distribution, so a generator cannot improve its score by hedging or by widening its intervals.

The continuous ranked probability score is the workhorse. For a forecast distribution with CDF and an observed outcome :

Estimated from a finite ensemble it is biased, and the bias shrinks with ensemble size, so a candidate run with fewer paths would score artificially well. The harness therefore uses the fair(unbiased) form throughout:

ScoreMeasures
Fair CRPSDistributional accuracy of the terminal outcomeThe finite-ensemble bias correction, so a candidate cannot look better simply by drawing fewer paths.
Censored integrated-variance CRPSRealized variance of the daily-rebalanced portfolio returnNormalized from training data only, and causally censored so the target cannot see past the origin.
FZ0 VaR and ES lossJoint quantile and expected-shortfall accuracyThe Fissler-Ziegel strictly consistent scoring function for the (VaR, ES) pair, which has no single-quantity elicitable form.
Fair energy scoreComplete-path multivariate accuracyThe multivariate generalization of CRPS. Scores the whole simulated path, not the endpoint.
Graph variogram scoreDependence structure across assets and timeSensitive to correlation misspecification in a way that marginal scores are not.
Drawdown and downside scoresThe path shape a planner actually cares aboutMaximum drawdown and downside deviation scored as distributions, with training-only normalization.

QLIKE is retained as a reported diagnostic rather than a confirmatory score. Its deviance form , with , diverges as the realized proxy approaches zero. At a one-session horizon the realized proxy is a single squared return, which is a chi-square with one degree of freedom, so the statistic is unbounded below in a way that no amount of resampling fixes.

Dependence-aware inference

Rolling origins at a 21-session stride with a 252-session horizon produce heavily overlapping targets. Treating those rows as independent would understate every standard error by a large factor, so the whole inference layer is built around not doing that.

  • Origin-row dependence floors. A gate cannot be evaluated below a declared minimum number of effective non-overlapping blocks, whatever the raw row count says.
  • Corrected automatic Politis-White block selection for the paired stationary bootstrap, with an audited structural overlap fallback that is published whenever its guard fires rather than silently substituted.
  • Intersection-union non-inferiority tests. A candidate must be non-inferior on every component to pass, so the composite claim inherits the level without an extra correction. This is the conservative direction by construction.
  • Authoritative Holm adjustment across the family of comparisons, controlling the family-wise error rate.
  • Romano-Wolf sensitivity is explicitly pending until its finite-sample size behaviour on this design is validated. It is not quietly used and not quietly dropped.

Selection stability and backtest overfitting

Scoring well on average is not enough. A generator that wins on average but is never selected by a train-only rule is not usable, because in production you must choose before you see the outcome.

Combinatorial purged cross-validation,

Six groups, two held out at a time, giving splits and five recombined out-of-sample paths. A train-only selector picks a candidate on each split, and the analysis measures how stable that choice is. Vetoes are field-invariant and pairwise against the incumbent, so adding an unrelated candidate cannot change another candidate's verdict.

Guarded CSCV probability of backtest overfitting

PBO estimates how often the in-sample best performer lands in the bottom half out of sample. The panels here are phase-separated, carry enforceable target-interval non-overlap, publish a complete trial registry, and require at least 16 non-overlapping origins per CSCV block. A PBO number computed on overlapping trials is not a PBO number.

Attribution and auxiliary evidence

Results are attributed by leakage-safe NIFTY regime label, named Indian market event, and a grid of Indian transaction-cost levels. Regime slices are training-only, named-event rows are descriptive, and the paired cost cells carry fail-closed sample guards.

Auxiliary diagnostics include ESR and VaR/ES severity, and residual raw, absolute and squared dependence tests admitted under a complete registry with per-test Benjamini-Hochberg FDR control and a declared max-absolute-ACF effect-size floor. The effect-size floor matters: at a long training window, a statistically significant autocorrelation of 0.01 is not a modelling defect.

Every eligible annual planning origin is replayed per candidate and cost level through the production accounting ledger, so a generator is scored on the planning quantities the product actually reports, not only on returns. Any single non-overlapping phase is retained as a secondary sensitivity rather than as the headline.

Conditional synthetic-path strategy stress is kept structurally separate from validation. Stressing a strategy on paths from a generator says something about the strategy; it says nothing about whether the generator is right, and mixing the two would let one borrow credibility from the other.

Sensitivity studies published alongside

Four completed studies used the same harness to test specific hypotheses. Every arm in them is validation-only and stays outside the production generator enum, so none of them changed what a user can select.

Block-length sensitivity for the stationary bootstrap

The automatic selector was run beside fixed 5, 21, 63 and 126-session arms over the frozen 7,160-row India panel, four horizons, 1,000 paths per fold and 9,999 paired resamples. Six annual-horizon cells crossed the governed 5% normalized-loss materiality bar: 21, 63 and 126-session blocks lowered mean QLIKE deviance by 7.8%, 13.4% and 17.3%, while raising maximum-drawdown CRPS by 8.3%, 10.0% and 9.3%. Those cells do not have 20 effective non-overlapping annual blocks, so the finding is descriptive rather than dependence-qualified. The automatic selector remains the default.

Dynamic conditional correlation

A validation-only FHS-EWMA arm replaced complete-row residual resampling with per-asset empirical residual marginals joined by a pathwise DCC(1,1) Gaussian copula. It never layers DCC over complete historical rows, so contemporaneous dependence has one source rather than being counted twice. All 2,216 paired forecast jobs completed with zero failures.

DCC improved the complete-path variogram after Holm control at all four horizons and lowered correlation Fisher-z RMSE at 21, 63 and 252 sessions. It still failed the predeclared activation gate, which required every proper multivariate score and correlation-RMSE cell to be non-inferior within a 5% relative margin. The 21-session complete-path energy point estimate was 7.6% lower, but its one-sided upper confidence bound was 12.8% higher, so non-inferiority was not demonstrated. Production activation is rejected, not deferred.

Conditional residual-row stationary bootstrap

An arm that holds the conditional volatility fit fixed and changes only residual resampling: complete cross-asset residual rows follow a Politis-Romano stationary bootstrap in observed order, with the expected block length selected per fit by the median per-asset Politis-White rule and clamped to 2 through 21 sessions. The lower bound prevents an IID-equivalent arm; the monthly upper bound limits long historical sequence replay.

Regime-FHS and MSGARCH activation

The soft-state regime-FHS/MSGARCH arm completed all 375 registered forecasts with zero failures across 19 annual origins, four horizons, four declared baselines, 1,000 paths per fold and 9,999 dependence-aware resamples. Activation is rejected, not deferred.

Every one of the 190 state and asset fits used a fallback, against a 10% aggregate ceiling. All 19 fit folds had a fallback, against a 5% ceiling. Five of 96 score cells were not estimable, because FZ0 is undefined when its required tail inputs are unavailable, and those cells fail the complete-paired-registry gate rather than being post-hoc subsetted out. The reference cost gate passed, at 249.0 seconds and 800,882,688 peak process bytes for the largest measured fit, but cost cannot override a failed statistical gate.

regime_fhs therefore remains absent from the selectable return process enum, from billing, and from every product surface. It is implemented and unit-tested, and it is deliberately unreachable. The all-or-nothing state is the safety property.

What this means for the numbers you see

Monte Carlo probabilities are conditional on the selected model, the frozen optimizer-time inputs, the effective configuration, the engine version and the seed. They are not calibrated forecasts, and this page is the evidence for why that wording is used rather than a disclaimer.

No generator has passed out-of-sample validation. No rebalancing policy has been shown to perform better than another. The model-risk cohort is empty, which is why the result panel and the PDF render a withheld-claim disclosure instead of a cross-generator range.

None of that makes the output useless. A conditional distribution is the right tool for asking whether a plan survives a bad decade, how much a goal probability moves when the assumed mean drops two percentage points, or what a stress arm costs. It is the wrong tool for asserting how likely any of it is.

References

  • Gneiting, T., & Raftery, A. E. (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. JASA, 102(477), 359–378.
  • Fissler, T., & Ziegel, J. F. (2016). Higher Order Elicitability and Osband's Principle. Annals of Statistics, 44(4), 1680–1707. The joint (VaR, ES) scoring family.
  • Scheuerer, M., & Hamill, T. M. (2015). Variogram-Based Proper Scoring Rules for Probabilistic Forecasts of Multivariate Quantities. Monthly Weather Review, 143(4), 1321–1334.
  • Ferro, C. A. T. (2014). Fair scores for ensemble forecasts. Quarterly Journal of the Royal Meteorological Society, 140(683), 1917–1923. The finite-ensemble bias correction.
  • López de Prado, M. (2018). Advances in Financial Machine Learning. Wiley. Purging, embargo and combinatorial purged cross-validation.
  • Bailey, D. H., Borwein, J., López de Prado, M., & Zhu, Q. J. (2017). The Probability of Backtest Overfitting. Journal of Computational Finance, 20(4), 39–69.
  • Politis, D. N., & White, H. (2004). Automatic Block-Length Selection for the Dependent Bootstrap. Econometric Reviews, 23(1), 53–70.
  • Holm, S. (1979). A Simple Sequentially Rejective Multiple Test Procedure. Scandinavian Journal of Statistics, 6(2), 65–70.
  • Patton, A. J. (2011). Volatility Forecast Comparison Using Imperfect Volatility Proxies. Journal of Econometrics, 160(1), 246–256. Why QLIKE needs a well-behaved realized proxy.

Not investment advice. Past performance is not indicative of future results.