Estimation Conventions
Every optimizer reads an estimate, not a fact. This page states the estimators FolioLab uses before any method runs: which dates enter the sample, how a return is defined, how the expected-return vector and the covariance matrix are built, and how the figures printed on a result are computed. Two runs can only be compared when they share these conventions.
The Sample: One Common Window, No Filled Prices
A run uses adjusted daily closes for every security and for the benchmark. The sample is the longest window in which every security and the benchmark have a price on the same session. A date that lacks a price for any one of them is removed from the whole panel. No price is carried forward, because a filled price manufactures a zero return on a day the security did not trade, and that zero understates both its volatility and its correlation with the rest of the portfolio.
The consequence is that the youngest security sets the start of the sample for all of them. A portfolio that adds one recent listing to nine old ones is estimated on the recent listing's history, not on the history of the nine. The run reports the start and end of the window it used.
Four checks run before estimation, and each one names the security it affects:
- A security needs at least 150 valid daily observations. With
ignore_short_historyoff, a shorter one fails the run, and the error names each such security with its count; with it on, the security is dropped and listed with the reasoninsufficient_history. - A security with a non-positive close anywhere in the window is dropped, because a return below -100% is not a price movement. The benchmark has no such exemption: a non-positive benchmark close refuses the run, since every relative measure depends on it.
- A security whose last price is more than five sessions older than the rest of the panel is dropped as stale, so one delisted name cannot end the sample early for every other name.
- A security that no source can price stops the run and is named. It is not silently removed so that a smaller portfolio is optimized in its place.
Returns and the Trading Year
Every estimator starts from the simple daily return of each security:
Simple returns, not log returns, because a portfolio return is then the exact weighted sum of its holdings' returns. A year is 252 trading sessions for every annualisation on this page. Each daily portfolio return applies the same weights to every day of the window:
Constant weights through time is a rebalance to the target every day. It is the standard convention for in-sample evaluation, and it is not the return a portfolio left to drift would have earned. The backtest accounting page states where this convention also applies out of sample.
The Expected-Return Vector
The methods that read an expected-return vector receive the geometric annualised historical mean of each security over the sample, which is the security's own compound annual growth rate:
This is not the arithmetic mean times 252. For a security with annual volatility , the geometric mean sits below the arithmetic mean by approximately , so a volatile security reads a lower expected return under this convention than under the arithmetic one. The difference is large enough to change a portfolio: at 40% annual volatility it is about 8 percentage points.
A historical mean is the noisiest input an optimizer reads. Its standard error over years is about , so ten years of a security with 30% volatility still leave a standard error near 9.5 percentage points (Merton, 1980). Methods that need no expected return at all, such as minimum variance, the hierarchical methods and inverse volatility, are immune to this error, which is one reason to compare them against the return-seeking methods. See Expected Returns for the general estimators; this page states the one FolioLab passes to the mean-variance family.
Two methods, MaximumDiversification and StackingOptimization, compare the risk-free rate against the arithmetic annualised mean in their feasibility check instead of the geometric one. On a short, volatile panel the two bases can differ by more than a percentage point, which straddles typical Indian risk-free levels, so the basis is fixed per method rather than left to vary.
The Covariance Matrix: Ledoit-Wolf Shrinkage
The sample covariance of daily returns is unbiased but unstable when the number of securities is not small relative to the number of days. Its smallest eigenvalues are biased toward zero, and an optimizer that inverts it loads on exactly those directions. FolioLab replaces it on every run with the Ledoit-Wolf (2004) estimator, a convex combination of the sample matrix and a structured target:
Here is the maximum-likelihood sample covariance of the demeaned daily returns (divisor ), and the target is the identity scaled by the average sample variance: every security has the same variance and no security co-moves with another. The shrinkage intensity is not chosen by hand. It is the closed-form value that minimises the expected Frobenius distance between the estimate and the true covariance, estimated from the same sample. More securities and fewer days give a larger .
Annualisation
Scaling by 252 assumes daily returns are serially uncorrelated. Positive autocorrelation, common in thinly traded small caps, makes the annualised figure too low.
What the target costs
Shrinkage toward a scaled identity pulls every correlation toward zero and every variance toward the cross-sectional average. The estimate is well conditioned and the weights are more stable, but the matrix understates the common market factor in an equity book whose securities are all positively correlated. A volatility read from is therefore not the realised volatility of the same weights, and the two are reported separately below.
Two numerical safeguards
- A security whose daily returns have a standard deviation of 1e-12 or less is left out of the fit. It receives the median daily variance of the other securities (at least 1e-6) on the diagonal and zero covariance with every other security, so the matrix stays invertible. The hierarchical methods do not take this path: they refuse a portfolio whose correlation matrix is undefined.
- Every diagonal entry receives 1e-10 in daily units before annualisation. This is far below any real variance and only guards the inversion.
These two inputs feed the mean-variance family. A method that builds its own estimator, for example an exponentially weighted covariance, a regime model or a scenario set for a tail-risk objective, states that estimator on its own method page. The spectral decomposition on the risk attribution page uses the same shrunk matrix.
The Risk-Free Rate in Each Formula
A run carries one annual risk-free rate , resolved as the risk-free rates page describes. Two forms of it appear in the formulas:
- The annual rate is subtracted from an annualised return, as in the Sharpe ratio.
- A daily excess return, as in the downside deviation of the Sortino ratio and in the CAPM regression, uses the compounded daily equivalent:
The rate is constant over the full-run window. A walk-forward backtest uses a different rate at each rebalance date: the mean of the daily rate over the 365 calendar days that end on that date, and never earlier than the start of the training window. A fold therefore reads no rate from its own hold period.
Measured Figures and Model Figures
A result carries two kinds of figure for each method, and they answer different questions. Read the field name before you compare two numbers.
Measured: the realised daily series
Every method, without exception, reports these from the daily portfolio return over the fitted window. One field name therefore means one quantity across all 31 methods and in the out-of-sample backtest. Sortino, Treynor and M-squared read the same expected_return.
| Field | Formula | Meaning |
|---|---|---|
expected_return | Arithmetic mean of the daily portfolio return, annualised by 252. | |
volatility | Sample standard deviation (divisor T - 1) of the daily portfolio return, annualised by the square root of 252. | |
sharpe | The two figures above with the annual risk-free rate subtracted from the annualised mean. | |
cagr | Compound annual growth of the cumulative value V, with T the number of daily returns. |
Model: the optimizer's own estimate
A method that returns its own performance estimate keeps it under separate names. Reports label these figures "Not a measurement". They exist only for the methods that produce such an estimate.
| Field | Formula | Meaning |
|---|---|---|
model_expected_return | The solved weights applied to the geometric expected-return vector the optimizer read. | |
model_volatility | The solved weights applied to the annualised Ledoit-Wolf covariance. | |
model_sharpe | The ratio of the two model figures above, after the risk-free rate. |
Why the two disagree
- The model return is geometric per security and then weighted; the measured return is the arithmetic mean of the weighted daily series. For a diversified portfolio the measured figure is usually the higher of the two.
- The model volatility reads the shrunk covariance; the measured volatility reads the realised series. Shrinkage toward the scaled identity usually makes the model figure the lower of the two for a long-only equity book.
- Both are in-sample. The weights were fitted on the same window that measures them, so neither is evidence of out-of-sample performance. The walk-forward backtest and the Deflated Sharpe Ratio exist for that question.
A Monte Carlo request that names no method selects the method with the highest measured Sharpe ratio, not the highest model Sharpe ratio.
Limitations
- Stationarity is assumed. Both estimators weight every day of the window equally. A structural break inside the window, for example a change of business or of index membership, is averaged in rather than detected.
- The window is the youngest security's history. A short common window gives a larger shrinkage intensity and a noisier mean, whatever the history of the older securities.
- The universe is chosen today. A security that is listed now survived to be chosen. Its historical mean carries that selection, so the in-sample figures of any portfolio built from today's list are biased upward relative to what an investor choosing at the start of the window could have held.
References
- Ledoit, O., & Wolf, M. (2004). "A Well-Conditioned Estimator for Large-Dimensional Covariance Matrices." Journal of Multivariate Analysis, 88(2), 365-411. doi:10.1016/S0047-259X(03)00096-4.
- Merton, R. C. (1980). "On Estimating the Expected Return on the Market: An Exploratory Investigation." Journal of Financial Economics, 8(4), 323-361. doi:10.1016/0304-405X(80)90007-0.
- Jagannathan, R., & Ma, T. (2003). "Risk Reduction in Large Portfolios: Why Imposing the Wrong Constraints Helps." The Journal of Finance, 58(4), 1651-1683. doi:10.1111/1540-6261.00580.
- Elton, E. J., Gruber, M. J., & Blake, C. R. (1996). "Survivorship Bias and Mutual Fund Performance." The Review of Financial Studies, 9(4), 1097-1120. doi:10.1093/rfs/9.4.1097.
Not investment advice. Past performance is not indicative of future results.