Backtest Accounting and Policy Comparison
A walk-forward backtest says what a strategy did. It does not say which part of the result came from keeping an allocation in place and which part came from changing it. This page states the accounting conventions of the backtest, then the three-arm comparison that separates those two effects, and the cost model that prices them.
Accounting Within a Hold Period
At rebalance date the method is fitted on prices up to and including the close of . The hold period runs from the next session to the close of . No price from the hold period enters the fit, and the risk-free rate for the fold is a trailing mean that ends at .
Inside the hold period the walk-forward arm applies the fold weights to every day:
Constant weights through time is a rebalance back to the fold weights every day. The backtest charges nothing for that daily rebalance. A book that is actually left alone between rebalance dates drifts with prices instead, and earns a different return. The two control arms below are built on drift, so the difference between them and the walk-forward arm contains this convention as well as the re-optimization. The chained hold periods are the only evaluation window: no training-window return enters the summary metrics.
Two Turnover Bases
Turnover is one-way: the fraction of the book bought, which equals the fraction sold. Each period reports it on two bases, and the field names say which.
turnover: previous target to new target
This is the long-standing figure, and every published result already means it. It ignores drift, so a method that re-optimizes to an unchanged target reports zero turnover although it traded back to that target.
turnover_drifted: drifted book to new target
The pre-trade weights are where the previous target drifted to with prices over the previous period. This is the trade a book that was left alone would need. At the entry trade there is no previous book, and the two bases agree. When the previous period cannot be priced in full, the drifted figure falls back to the target-to-target figure rather than reporting a zero.
The Never-Rebalanced Control Arm
A benchmark index differs from a strategy in its holdings as well as in its behaviour. The control arm differs only in behaviour: it is the same kind of portfolio, chosen once at the anchor date and then never traded. It holds fixed units, so its value needs no weight state at all:
Because nothing writes a weight after the anchor, no path can rebalance it. The allocation comes from one of three named sources, and the result records which:
first_fold(the default): the method's own first-fold weights. The comparison then isolates the value of continuing to intervene.supplied_allocation: weights you supply, for example the parent run's weights. Without weights the request is refused rather than given a different allocation.equal_weight: equal weight over the securities tradable at the first origin.
The arm reports a value_of_intervention block: the walk-forward total return minus the control total return, their terminal wealth ratio, the metric differences, and turnover_avoided. That last figure counts rebalance-date turnover only. An equal-weight method can therefore show a large value of intervention beside zero turnover avoided, because its intervention is the daily rebalance inside each period, which turnover does not count.
The arm is dropped, with a named reason, when it cannot cover every hold date of the walk-forward arm. It is never compounded over a shorter span.
Three Policies on One Evidence Set
The policy comparison adds a middle arm. All three start from the same book on the same anchor date, read the same prices over the same hold dates, and differ only in when they trade.
| Arm | When it trades | Note |
|---|---|---|
| Buy and hold | Trades once, at the anchor date, and never again. | Holds a fixed number of units. Its weights drift the most of the three. |
| Rebalance to original target | Trades back to the anchor weights on every rebalance date. | Its target never changes. The book it holds drifts between rebalance dates. |
| Walk-forward re-optimize | Re-estimates a new target on every rebalance date and trades to it. | The primary backtest arm described on the walk-forward page. |
The middle arm is not "static weights". Its target is static, but the weights it holds drift between rebalance dates, which is why it trades. Between two rebalance dates it is buy and hold; at each date its weights reset to , and its turnover is measured from the drifted pre-trade weights.
The effect decomposition
With , and the total returns of the three arms, each effect is the left arm minus the right arm:
The two effects sum to the total by construction, because they are adjacent differences of the same three numbers. The identity is published with the figures so that a reader can check it. A pair with a missing arm is not published: a gap with one side missing is not a smaller gap.
What the re-optimization effect contains
Two things, and the result names both: the estimation of a new target, and the daily rebalance back to the fold weights inside each period that the walk-forward accounting performs. A method that re-estimates the same target every period, such as EquiWeighted, therefore still shows a small re-optimization effect. The effect is not a pure measure of the value of re-estimation.
The Cost Model
Costs are a flat rate in basis points of notional traded, zero by default. A one-way turnover trades a notional of (the buys plus the sells). The cost is charged multiplicatively at each trade , so a cost paid early also loses the growth it would have earned:
The reported transaction_cost_drag is . The entry trade from cash counts as one trade on every arm, so buy and hold pays it too. Gross figures never change; net figures appear beside them.
The turnover basis is not uniform across arms, and this matters for net figures. The two control arms are charged on drifted pre-trade turnover. The walk-forward arm is charged on its target-to-target turnover, which ignores drift and understates its trading. Its net figure is therefore optimistic, and a net difference that involves the walk-forward arm is a lower bound on its cost, not a measurement of it. The daily rebalance inside each walk-forward period is not charged at all.
Every figure is pre-tax. The backtest holds no tax ledger, so tax avoided and slippage avoided are reported as unavailable rather than as zero. With no cost configured, the frictions the backtest omits fall almost entirely on the arms that trade, so a reported advantage of trading is an upper bound.
The Evaluation Window
A policy comparison is a statement about one cutoff, so the cutoff is something you set and the result records. cutoff_date is the first evaluation session: the allocation is formed from data through that session's close and held from it, so the cutoff becomes the first origin rather than being rounded to the next calendar boundary. end_date is the last evaluation session.
The formation history before the cutoff stays available to every training window; the cutoff restricts the evaluation, not the data. Both dates resolve onto real trading sessions and the result keeps the requested and resolved pairs, so a date that fell on a holiday cannot mean a different span on a later run. A cutoff and end that cannot form an evaluation window fail the backtest with an error that names the dates; the window is never widened to make it fit. Separately, a standalone backtest whose parent history cannot form a train and hold pair fails with BACKTEST_INSUFFICIENT_HISTORY.
Configuration
The fields sit in the rolling_backtest block of POST /jobs and in the body of POST /runs/{run_id}/backtest. See Backtests API. The MCP tool submit_backtest does not take them, so a backtest submitted through an AI assistant uses the defaults: both arms on, first_fold, no cost, and no explicit evaluation window.
| Field | Type | Description |
|---|---|---|
shadow_account.enabled | boolean, default true | Build the never-rebalanced control arm. The policy comparison needs it. |
shadow_account.allocation_source | first_fold, supplied_allocation or equal_weight; default first_fold | Which single allocation the control arm holds. Stored with the result. |
shadow_account.weights | object of weights, optional | Required with supplied_allocation and refused with any other source. |
policy_comparison.enabled | boolean, default true | Build the rebalance-to-original arm and the effect decomposition. |
policy_comparison.transaction_cost_bps | number from 0 to 500, default 0 | Flat cost in basis points of notional traded. Zero leaves net equal to gross. |
evaluation.cutoff_date | date, optional | The first evaluation session. Resolved forward to a trading session. |
evaluation.end_date | date, optional | The last evaluation session. Resolved backward to a trading session. Must be after the cutoff. |
{
"rebalance_frequency": "quarterly",
"window_type": "expanding",
"shadow_account": { "enabled": true, "allocation_source": "first_fold" },
"policy_comparison": { "enabled": true, "transaction_cost_bps": 15 },
"evaluation": { "cutoff_date": "2020-01-01", "end_date": "2024-12-31" }
}Each method's backtest result carries shadow_account, policy_comparison and evaluation_window. The workbook shows them on the Control Arm and Policy Comparison sheets.
What This Can and Cannot Show
- One path. Every figure is one realized historical path for one universe and one cutoff. It shows what each policy did here. It is not a forecast, and it is not evidence that any policy performs better in general.
- The universe is chosen today. The securities in a backtest are the ones you hold or chose now, so each of them survived to the present. A historical backtest over them carries survivorship and selection bias that no accounting convention removes.
- No market impact. The flat rate does not grow with trade size, and it does not model the spread or the price a trade actually gets.
- The difference between two arms is not significance. One difference of total returns carries no standard error. The Deflated Sharpe Ratio on the walk-forward arm is where the multiple-testing question is addressed.
References
- Perold, A. F., & Sharpe, W. F. (1988). "Dynamic Strategies for Asset Allocation." Financial Analysts Journal, 44(1), 16-27. doi:10.2469/faj.v44.n1.16.
- DeMiguel, V., Garlappi, L., & Uppal, R. (2009). "Optimal Versus Naive Diversification: How Inefficient Is the 1/N Portfolio Strategy?" The Review of Financial Studies, 22(5), 1915-1953. doi:10.1093/rfs/hhm075.
- Bailey, D. H., Borwein, J. M., Lopez de Prado, M., & Zhu, Q. J. (2014). "Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance." Notices of the American Mathematical Society, 61(5), 458-471. doi:10.1090/noti1105.
- Brown, S. J., Goetzmann, W., Ibbotson, R. G., & Ross, S. A. (1992). "Survivorship Bias in Performance Studies." The Review of Financial Studies, 5(4), 553-580. doi:10.1093/rfs/5.4.553.
Not investment advice. Past performance is not indicative of future results.