Monte Carlo · Methodology
Model Risk and Generator Disagreement
One generator produces one number. Several credible generators produce a range, and the width of that range is a measurable statement about how much the answer depends on a modelling choice rather than on the portfolio. This page describes how that range is computed, and why the product currently withholds it.
The production state
The accepted cohort is empty. Every generator was rejected by the rolling-origin validation harness under the frozen policy, so no cross-generator range is published.
Every Monte Carlo probability is therefore presented with the withheld-claim disclosure and remains conditional on the selected model. The mechanism is built, tested and wired to the result panel, the PDF and the Excel export. It is waiting on evidence, not on engineering.
Why disagreement is measured rather than assumed away
Two generators fitted on the same history can produce materially different goal probabilities. A stationary block bootstrap resamples observed days and inherits whatever crises the window contains. A filtered historical simulation conditions on today's volatility level and forecasts forward from it. Neither is obviously right, and the gap between them is not sampling noise that more paths would remove.
For a metric and an effective cohort , the published point envelope is simply
with the generators attaining each end named, so a reader can see which modelling assumption drives each bound. Where the backend can compute them, conservative simultaneous bounds on the true spread are published alongside the point spread, because each is itself estimated and the range of estimates is a biased estimate of the range.
This is not a confidence interval
The envelope has no coverage property. It is the spread of point estimates across a small, discrete, deliberately chosen set of models. A wider envelope means the answer is more sensitive to model choice; a narrow one means the accepted models happen to agree, which is not the same as being right together. The surfaces carry a server-owned caveat sentence saying exactly this, rendered verbatim rather than paraphrased per surface.
The six cohort states
The reason a range is missing is itself information, so the states are kept distinct rather than collapsed into one absence. “No cohort exists” and “a cohort exists and one member did not finish” are different facts, and a surface that renders both as an empty space has lost the second one.
| Status | Meaning |
|---|---|
unavailable_no_statistically_qualified_generatorsrenders as unavailable | No generator has cleared the rolling-origin thresholds, so no cohort exists to disagree. This is the production state today. |
unavailable_no_product_approved_generatorsrenders as unavailable | Generators cleared statistically but have not received product and model-risk approval. Statistical qualification is necessary, not sufficient. |
unavailable_no_runnable_generatorsrenders as unavailable | An approved cohort exists but none of its members can run this particular request, for example because the calibration window is too short to fit them. |
not_estimable_single_generatorrenders as single | The cohort has exactly one member. Disagreement is undefined, not zero, and a zero range would be a false statement of agreement. |
incomplete_expected_generator_missingrenders as incomplete | Two or more generators were expected, and at least one produced no comparable result. Any partial range is not the range, so the envelope is withheld and marked diagnostic only. |
availablerenders as available | Two or more generators completed with admissible results. The envelope is rendered as a qualified range. |
A run produced by an engine older than the model-risk block reports absent, which is distinct again from a modern run whose cohort is empty. Backfilling a state onto a run that never computed one would be inventing evidence.
Governance failures are surfaced, not swallowed
Each metric view reports, separately: the generators in the effective set with their values, expected generators that produced nothing, generators that ran but were excluded together with the precision status that excluded them, and qualified generators that cannot be run for this particular request.
That last category is the one most easily lost. A generator can be accepted in general and still be unusable here, typically because the portfolio's calibration window is too short to fit it. Reporting it as simply missing would suggest a failure where there is a requirement.
The incomplete state is treated as the most dangerous, because it is the case where a surface is most tempted to show the numbers it does have. The compact projection withholds the partial envelope so it cannot, and marks the presentation diagnostic_only so a caller reading the fuller artifact block still gets the right label.
The governance pin
Cohort governance is a mutable file. Without protection, a run admitted while the cohort was empty could execute after an activation and quietly perform comparison work that was never metered or reserved.
Each run therefore carries a server-owned model_risk_pin, stamped by the submission path and replayed by the worker. A client-supplied value is overwritten at admission. The effect is that a run always executes under the governance state it was admitted under, and a later activation does not retroactively change what a queued run does.
Where it appears
Cross-model disagreement is rendered wherever a Monte Carlo probability is presented: the result panel, the PDF report and the Excel workbook. The disclosure sentence and the envelope caveat are server-owned strings rendered verbatim by each surface, so the three cannot drift into three subtly different claims.
The same rule governs the comparison arms under adaptive precision. A model-risk comparison interval divides the error budget across metrics and generators, so it can remain inadmissible after the selected generator has met its own tier. Comparison arms continue independently to the next checkpoint until those stricter simultaneous intervals meet their targets or reach the ceiling, rather than inheriting a stop decision they did not earn.
References
- Board of Governors of the Federal Reserve System (2011). Supervisory Guidance on Model Risk Management, SR 11-7. The source of the separation between model development, validation and approval used here.
- Derman, E. (1996). Model Risk. Goldman Sachs Quantitative Strategies Research Notes.
- Cont, R. (2006). Model Uncertainty and its Impact on the Pricing of Derivative Instruments. Mathematical Finance, 16(3), 519–547. On measuring model uncertainty as a range over a set of admissible models.
- Danielsson, J. (2002). The Emperor Has No Clothes: Limits to Risk Modelling. Journal of Banking & Finance, 26(7), 1273–1296.