Risk and Asset Allocation
Portfolio construction sits on a chain of decisions, and almost every one of them is usually taken silently. What is the modellable quantity inside a price series? Over what horizon does a risk number apply? How much of an estimate is signal and how much is sampling noise? What happens to an optimiser when it is fed estimates rather than truths? This arc works through that chain from the ground up, in the framework of Meucci's Risk and Asset Allocation, building each piece from scratch and checking it in more than one language. The organising idea is that estimation error is not a footnote but a first-class modelling problem — an allocation derived from point estimates inherits every bit of their uncertainty, and usually amplifies it. Everything begins with the first two steps below, because a risk model that has not identified its invariant, or has quietly scaled a daily number to a yearly one, is answering a question nobody asked. From there the arc estimates what the optimiser needs — shrinking the sample mean and covariance, then replacing point estimates with posteriors — before turning to allocation itself: Black–Litterman for blending views with equilibrium, and robust optimisation for the worst case inside an uncertainty set. The last three projects change the question from how should I allocate? to what can this actually lose? — coherent risk measures and extreme-value theory, copulas for the dependence a correlation cannot express, and finally copula-GARCH, where volatility and dependence are both allowed to move. A theme sharpens as it goes: the arc keeps finding that the honest error bars are wider than the ones the models report. Parameter uncertainty is the easy part; model choice moves a joint-crash probability several times further than any credible interval, and the equity–bond correlation that anchors a 60/40 portfolio does not merely need better estimation — it changed sign in 2023.
What an optimiser does, and why it is an error-maximiser. The recipe underneath almost every page here is Markowitz’s: given expected returns for each asset and a covariance matrix saying how they move together, find the weights with the best return for a given risk. As arithmetic it is unimpeachable and takes one line of linear algebra. The difficulty is that it needs two inputs nobody observes, and it is actively hostile to errors in them. An asset whose average return happened to come out flattered by sampling noise looks, to the optimiser, exactly like a genuinely good asset — so it buys it, and buys more of it the noisier the estimate is. Nothing in the procedure distinguishes signal from luck; it simply maximises, and what it maximises includes the error. Hence the name: an error-maximiser. Every technique in this arc is a different way of not being one — shrink the inputs, replace them with posteriors, anchor them to what the market already implies, or distrust them inside the objective itself.
How a portfolio is scored. Four measures recur and it is worth having them in one place. The Sharpe ratio is return per unit of risk — average excess return divided by its standard deviation — and it is the standard single number for comparing strategies of different sizes. The certainty equivalent is the more demanding one, and it carries the headline result on two pages below: the guaranteed return an investor would accept in place of the risky portfolio, so it prices return and risk together in a single figure and can be negative when a portfolio is bad enough. Because it can be computed against a known truth on synthetic data, it is what makes the cost of estimation error measurable rather than rhetorical. Turnover is how much of the book has to be traded to get from last period’s weights to this period’s — a direct proxy for trading cost, and one of the places shrinkage wins unconditionally. And gross leverage is the sum of the absolute weights: 100% means a fully-invested long-only book, while 518% means the optimiser has taken enormous offsetting long and short positions, which is the signature of a fit driven by noise.
N/T, the number that decides how much of this matters. One ratio determines the verdict on nearly every page: the count of assets N against the number of observations T. It matters because the covariance matrix has N(N+1)/2 distinct entries to estimate, so the quantity to be learned grows as the square of the universe while the data grows only in proportion to its length. At N/T ≈ 0.1 — ten sectors and a couple of hundred weeks — there is little estimation error to fight, and the honest finding on several pages is that the sophisticated methods buy nothing over the naive ones. At N/T = 0.8 the sample covariance is nearly singular, its inverse is nonsense, and the same methods are the difference between a usable portfolio and a 199%-volatility disaster. Where a result below looks like a contradiction between two pages, this ratio is usually the reason. The recurring benchmark is 1/N — simply holding every asset in equal weight — which is famous for beating optimisers, and does so here too at low N/T, for the instructive reason that it estimates nothing at all and so pays no estimation cost.
The two findings the arc is organised around
The left pair is estimation risk, run here on the ten sector ETFs the projects use: their sample mean and covariance are treated as the truth, a year of weekly data is drawn from that truth 300 times, and each draw is optimised as a manager would optimise it. Three curves follow. The claimed frontier is what the manager reports from their own estimates. The true frontier is what was actually available. And actually delivered is where the claimed portfolios really land when the truth is applied to them. Claimed sits optimistically up and to the left; delivered sits below and to the right — worse on both axes at once.
At the middle of the frontier the manager claims 0.35%/week and receives 0.18% — 48% of the promise — while running 2.20% risk against the 1.81% advertised. The middle panel puts both failures on one scale, as percentages of what was promised, and the two behave quite differently. The return shortfall grows with ambition: it is negative at the very lowest targets, where the optimiser accidentally over-delivers, crosses zero around 0.16%/week, and reaches 64% at the most aggressive end — the more you reach for, the smaller the fraction that arrives. The risk excess does not vary at all. It sits flat at 21% across the entire frontier: wherever you choose to sit, you are taking about a fifth more risk than you were told, and the number is the same whether you are cautious or aggressive. Neither is sampling noise that would average out. The mechanism is the one named above — the optimiser selects whatever the sampling error flattered, so the errors it picks up are precisely the ones biased in its favour.
The right panel is the limitation the arc runs into, and it is a different kind of problem. Everything to its left treats the covariance matrix as a fixed thing to be estimated better from noisy data. This is the 500-day rolling correlation between equities and long Treasuries — the relationship a 60/40 portfolio is built on — computed from GARCH residuals so that shared volatility is not doing the work. It runs at −0.56 in 2011 and reaches +0.13 by the end of 2024, and it is negative in 88% of windows, which is exactly what makes the drift easy to miss. The zero crossing is not one clean event either: it crosses in the first days of January 2023, falls back, crosses again in February, and only stays positive from July. No amount of shrinkage repairs an estimate of a parameter that is not constant, which is why the arc ends by letting volatility and dependence move rather than by estimating them better.
The chain, in order
Unusually for this collection, these eight are a genuine sequence rather than a menu — each one opens by acting on what the previous one established, and the last two arrive at a problem none of the earlier ones can fix. Read top to bottom.
What are we modelling?
1 Invariance & horizonfind the quantity that repeats — then project it to the horizon you actually invest overEstimating what the optimiser needs
2 Shrinkagethe sample mean and covariance are too noisy to use raw — and it reports when shrinking stops paying 3 Bayesian estimationmake the noise explicit: every weight is a distribution, and most straddle zeroAllocating anyway
4 Black–Littermanstop estimating returns from history; start from what prices already imply, then add views 5 Robust allocationoptimise the worst case inside an uncertainty set rather than the best case at a pointWhat can it lose?
6 Coherent risk & EVTVaR, Expected Shortfall, and a demonstration that VaR punishes diversification 7 Copulasa correlation cannot say whether assets crash together; a copula can 8 Copula-GARCHlet volatility and dependence both move — and find the one relationship the arc assumed was constantTwo things to watch for on the way down. The first is that the arc keeps disagreeing with itself productively: shrinkage beats the sample estimator, then stops beating it once there is enough data; equal weighting beats every optimiser on ten sectors and loses to them on forty-eight. Almost every one of those reversals is the N/T ratio changing, not a method failing. The second is that the honest error bars keep coming out wider than the models report — and in the last two projects the reason stops being parameter uncertainty. Choosing the wrong copula family moves a joint-crash probability roughly six times further than the posterior interval spans, and the equity–bond correlation that anchors a 60/40 portfolio was never a parameter to be estimated better in the first place.
The Invariance Quest & Horizon Projection
The foundation the rest of the arc stands on, and two questions usually skipped. First, what do we model? Not the stock price — it trends forever and today's value is essentially yesterday's — but the invariant hidden inside it: the increment whose distribution repeats identically and independently. The hunt is settled empirically across three markets, with lag-1 autocorrelation separating persistent levels (equity price +0.999, 10-year yield +0.997, VIX +0.965) from their i.i.d. increments (−0.099, −0.015, −0.071). Notably the notebooks then test the assumption rather than assume it, and report that it does not strictly hold: a Ljung-Box statistic of 217 on 20 degrees of freedom says returns carry real dependence — in the volatility, not the mean — which is exactly what motivates GARCH elsewhere in the collection. Second, over what horizon? Because the invariant is i.i.d., the horizon distribution is its T-fold convolution, computed three independent ways (FFT characteristic function, moments plus Cornish–Fisher, and simulation) that agree. That reveals the central-limit drift — skew decaying like 1/√T, kurtosis like 1/T — and why the industry's square-root-of-time VaR rule understates losses by around a percentage point at practical horizons: it assumes a normality that has not arrived yet. A PyMC engine closes the loop by putting error bars on the risk number itself, showing the Bayesian predictive VaR is more extreme than the plug-in because integrating parameter uncertainty fattens the tail.
View example →Shrinkage Estimation of Mean and Covariance
Every allocation recipe needs an expected-return vector and a covariance matrix, and neither is ever observed — 50 assets means estimating 1,325 numbers, usually from a few hundred observations. The estimates are unbiased but noisy, and a mean–variance optimiser is an error-maximiser that loads up on whatever was over-estimated by chance. Shrinkage trades a little bias for a lot of variance. For the mean that is Stein's paradox — for N ≥ 3 the sample mean is provably inadmissible — demonstrated by Monte Carlo rather than cited, cutting squared error 23% when data is scarce. For the covariance it is Ledoit–Wolf, whose closed-form intensity is derived and then checked against the notebook's own implementation, pulling the condition number from 13,174 to 117 on 48 stocks. What lifts this above a standard demo is that it reports when shrinkage loses: the sample minimum-variance portfolio pays a 65% volatility penalty when T barely exceeds N, but at a 52-week window it is perfectly fine and shrinking toward the identity buys nothing. Dimensionality is the amplifier — on the Fama–French 48 industries the sample portfolio realises 22.7% annualised volatility against 12–13% for shrinkage, with eight times the turnover. A Bayesian companion closes the loop by showing the shrinkage formula is a posterior mean.
View example →Bayesian Estimation & Estimation Risk
Shrinkage ended by noting that its formula has the shape of a posterior mean. This project takes that seriously, estimating expected returns and covariance with a conjugate Normal-Inverse-Wishart prior — and the point is not a better point estimate but that the output is a distribution, which turns "estimation risk" from a worry into a measurable quantity. The machinery is validated first: the closed-form posterior is reproduced by a from-scratch Metropolis sampler, and again by an independent Gibbs sampler in R. Then the central demonstration — plot the frontier a manager claims from sample estimates, the true frontier, and what those claimed portfolios actually deliver. The claimed curve sits optimistically above and left; reality sits below and right, worse on both risk and return, systematically, because the optimiser selects whatever sampling error flattered. The cost is then priced: at T=16 the sample optimiser gives up 13.2 units of certainty equivalent against 0.17 for the Bayesian allocation. Dimensionality decides how much it matters — on the Fama–French 48 industries at N/T=0.8 the Bayesian portfolio reaches Sharpe 1.01 against 0.30 for the sample optimiser, and beats even the formidable 1/N benchmark at 0.83. A PyMC companion makes it visual: 60% of 48 positions have a posterior weight interval straddling zero, so the data cannot determine even the sign — something a plug-in optimiser hides completely.
View example →The Black–Litterman Model
The practitioner's answer to everything the previous projects diagnosed, and its insight is to change the starting point. Rather than estimating expected returns from history, reverse optimisation asks what returns the market must already believe for today's capitalisation weights to be optimal, and treats those as a prior. The motivation is shown rather than asserted: feeding historical means into an optimiser on the same ten sectors produces positions from −109% to +188% at 518% gross leverage, a book nobody would hold. Views are then stated in the model's grammar — a pick matrix, a claim, and a confidence — here one relative (Tech beats Staples) and one absolute and bearish (Energy below equilibrium), each compared against what equilibrium already implies so the direction of the bet is explicit. The posterior does exactly what the design intends: viewed sectors tilt, and sectors carrying no view stay at equilibrium. The confidence dial makes it tangible — sweep Ω from near-certainty to near-ignorance and the portfolio slides continuously back to the market itself. A PyMC companion shows the famous master formula is nothing exotic: it is the posterior mean of an ordinary Bayesian model, with equilibrium as prior and views as observations.
View example →Robust Bayesian Allocation
If portfolio weights are really distributions, optimise accordingly. Robust allocation maximises against the worst case inside an uncertainty set rather than trusting a point estimate — asking what is best if the estimate is as wrong as it plausibly could be. The construction is verified against both its limits: at q = 0 it reproduces mean–variance exactly, and as q grows it converges on minimum-variance, since enough distrust of the return estimates leaves only risk worth optimising. What it buys is insurance: on a known-truth market, raising q lifts the worst-case outcome from −10.31 to +0.12 while the spread of results collapses more than tenfold. The most useful section puts two treatments of the same disease side by side — Bayesian shrinks the inputs, robust distrusts them in the objective — and the PyMC engine shows they are one idea, with robust weights from the posterior dispersion matching the analytic ellipsoid to four decimals. The out-of-sample race is refreshingly honest: on ten sectors 1/N beats every optimiser, because at N/T ≈ 0.1 there is little estimation error to fight and equal weighting pays no estimation cost at all. Push to 48 stocks and the Fama–French industries and it inverts — sample mean–variance degenerates to 199% volatility while Bayesian and robust hold Sharpe near 1.0 at 12% and overtake 1/N.
View example →Coherent Risk: VaR, Expected Shortfall & EVT
The arc turns from building portfolios to measuring what they can lose — and to the ways the industry-standard measure misleads. Estimating VaR four ways on the same returns shows the disagreement plainly: the Gaussian understates every tail (3.28% against a historical 6.12% at the 0.1% level) while Cornish–Fisher explodes to 14.18%, because the moment expansion is only valid for mild non-normality and this series has excess kurtosis near 12. Where the sample runs out, extreme-value theory takes over: a Generalized Pareto fit to 378 threshold exceedances gives shape ξ = 0.155, a genuinely power-law tail, and extrapolates where historical estimation simply has no observations left. Then the mathematical objection, demonstrated rather than asserted — a constructed case where VaR(A) + VaR(B) = −4.0 but the diversified VaR(A+B) = 49.0, so VaR punishes diversification, while Expected Shortfall stays coherent. The backtests close it honestly: Gaussian VaR fails on rate, historical VaR passes on rate, and both fail independence because breaches arrive in clusters — a diagnosis that points at dynamic volatility models rather than pretending the static measure suffices. A Bayesian EVT companion adds the quietly devastating number: the 1-in-10,000-day VaR carries a credible interval spanning a factor of 1.9, all of which a single plug-in figure conceals.
View example →Copulas and Tail Dependence
Every model in the arc so far summarised co-movement with a covariance matrix. A correlation cannot answer the question that matters — not do these assets move together but do they crash together. Sklar's theorem splits the joint distribution into marginals and a copula, and five copulas calibrated to the same rank correlation make the point unmistakable: identical by correlation, completely different in the corners. Fitted to five cross-asset ETFs the Student-t copula wins on AIC in both engines. Two things are worth stating precisely. First, tail dependence is a limit, not a probability at any measurable threshold — at q = 0.10 the fitted Gaussian gives 0.47 for SPY–HYG against 0.52 in the data, so its failure is asymptotic rather than the infinite gap that comparing against its limit of zero implies; the page reports both side by side. Second, and sharper: a Bayesian Clayton fit returns a tight posterior for the joint-crash coefficient, 0.63 with an interval of [0.59, 0.66] — while the t-copula puts the same quantity at 0.185. Model choice moves the number six times further than the posterior interval does, and the observed conditional-crash curve decays with depth where Clayton's runs flat. A tight credible interval is not a well-known quantity.
View example →Copula-GARCH: Time-Varying Volatility Meets Dependence
The capstone, and the model that answers a diagnosis the arc left open: both static VaR methods failed the independence backtest because their breaches arrived in clusters. Copula-GARCH splits returns into each asset's time-varying volatility and the dependence of the standardised residuals, estimated separately. That decomposition immediately pays: fitting a t-copula to raw returns and then to residuals, SPY–HYG tail dependence falls from 0.38 to 0.22 — part of what looked like a tendency to crash together was never dependence at all, but shared volatility. The Bayesian engine confirms the drop is real, with the entire posterior of the difference above zero. The reassembled one-day-ahead VaR rises the moment volatility spikes instead of drifting up late and lingering, and it passes both Kupiec and Christoffersen where the static VaR fails independence at p = 0.000. That comparison originally gave the winner full-sample parameters against a rolling-window baseline, so it is now also run the hard way — everything fitted on a 1,000-day burn-in and frozen — and the verdict survives. Honest throughout about what the marginals do not achieve: GARCH whitens three of the five series, not all five. And it ends on the sharpest limitation in the arc. SPY–TLT — the equity–bond relationship a 60/40 portfolio leans on hardest — changed sign. GARCH does not explain it (raw −0.298 becomes residual −0.282), but the rolling correlation drifts from −0.56 in 2011 to +0.16 in 2024, crossing zero in January 2023. Shrinkage, Bayesian estimation, Black–Litterman and robust allocation all treat the covariance matrix as a fixed object to estimate better, and no amount of shrinkage repairs an estimate of a parameter that is not constant. The Bayesian engine makes the point from the other side: its posterior for that pair sits entirely below zero, tight and emphatic, about an average of a quantity that reversed inside the sample window.
View example →