How to Compare Five Portfolios Without Letting the Backtest Pick the Winner
The fixed rules behind Portfolio Lab: Live — and seven questions to ask before you believe any backtest.
The fixed rules behind Portfolio Lab: Live - and seven questions to ask before you believe any backtest.
You do not need to invent a single price to make a backtest tell a different story.
Move the starting date past a crash. Replace an investable ETF with a longer but different gold series. Ignore trading costs. Change the score from ending wealth to a risk-adjusted ratio. Each choice can be defensible. Each can also change the apparent winner.
That is why the first question is not, “Which portfolio won?”
It is, “Which decisions were fixed before anyone looked at the result?”
Portfolio Lab: Live compares five hypothetical portfolios under one public protocol. The commitment is simple: no changing the portfolio rules after seeing how they performed.
Quick note: this is educational model tracking, not investment advice, an account record, or a recommendation. The Lab uses hypothetical total-return models. Historical drawdowns are report cards, not floors; a future decline can be larger.
A backtest can be honest and still answer the wrong question
A polished chart can be numerically correct while leaving its most important choices invisible.
Suppose one study begins after an asset’s worst collapse, while another includes it. Suppose one uses adjusted prices with distributions reinvested, while another uses price alone. Suppose one portfolio is rebalanced for free every month, while another pays modeled trading costs once a year.
The outputs may all be calculated correctly. They are still not answering the same question.
The Lab therefore treats the method as part of the result. The clock, vehicles, data treatment, cash-flow rule, trading assumptions, metrics, and correction policy all belong beside the performance numbers.
Seven questions to ask before you believe any backtest
1. What is the exact window?
The start date decides which crashes, recoveries, inflation regimes, and interest-rate cycles enter the record. “Long term” is not a date. A useful backtest states the first and last observation and explains why that boundary exists.
2. What could an investor actually have owned?
A spot index, futures series, mutual fund, and ETF can represent the same broad asset while producing different histories. Vehicle choice affects availability, distributions, fees, tracking, and the earliest fair comparison date.
3. Is this price return or total return?
A price chart can omit dividends and other distributions. The Lab uses adjusted-close data so distributions are reinvested in the model. That does not create an actual account; it makes the return definition explicit.
4. Are outside cash flows mixed into performance?
Deposits and withdrawals can change an account balance without saying anything about the portfolio’s underlying return. The Lab starts each model with the same hypothetical amount and adds no later cash flows. It reports time-weighted performance.
5. How are rebalancing and trading costs modeled?
Rebalancing can create trades, and trades are not frictionless. Frequency, drift rules, commissions, slippage, and market impact can all alter the path. A strategy rule without an execution rule is incomplete.
6. Which score decides the verdict?
Ending wealth, annualized return, maximum drawdown, volatility, Sharpe, and Calmar do not measure the same thing. If the winning metric changes after the results appear, the backtest can select its own champion.
7. Can the author quietly rewrite history?
A credible live record needs dated inputs, reproducible outputs, and visible corrections. Method changes should create a new version. They should not overwrite the old receipt.
The five fixed lines
The Lab is not a search for one universally best portfolio. It is a fixed bench of five deliberately different designs.
100% QQQ - QQQ 100%. The growth extreme.
100% S&P 500 - SPY 100%. The broad U.S. equity benchmark.
Stock-Gold 60/40 - SPY 48%, QQQ 12%, GLD 40%. The Lab’s case-study allocation from its backtest research - not a textbook standard and not a recommendation.
Classic 60/40 - SPY 60%, TLT 40%. The conventional stock-and-long-Treasury reference.
Harry Browne’s Permanent Portfolio - SPY 25%, TLT 25%, SHY 25%, GLD 25%. The textbook four-part design published in 1981, spread across growth, deflation, cash-like stability, and gold.
The single-asset lines are buy and hold. The three multi-asset lines rebalance annually, subject to a 0.1% drift threshold. The roster and weights remain fixed from one weekly issue to the next.
The differences are the point. QQQ makes growth concentration visible. SPY supplies an equity control. Stock-Gold asks what happens when gold supplies the defensive sleeve. Classic 60/40 represents a familiar stock-bond balance. The Permanent Portfolio pushes diversification further.
Why the common clock begins in 2005
The hypothetical start is January 1, 2005, with the first common trading observation on January 3.
That boundary is not chosen because 2005 makes every portfolio look good. It is the first year-start after all investable ETFs used by the five public models existed. GLD sets the limiting date.
The choice creates a fair common clock, but it also creates a limitation: QQQ’s 2000-02 crash sits outside the tracker. This is a recovery-era QQQ sample, not QQQ’s complete history.
The same window contains a strong period for gold and an unusually severe bond bear market. Earlier studies that use spot gold or gold futures reach further back, but they are not directly interchangeable with this GLD-based tracker. GLD is an investable ETF with an approximately 0.40% expense ratio and tracking difference; those vehicle costs are part of why the histories are not interchangeable.
There is no neutral start date. There is only a disclosed one, with consequences the reader can see.
The frozen protocol
Every line begins with a hypothetical $10,000. No deposits or withdrawals follow.
Data. Daily adjusted close, with distributions reinvested. Tiingo is the primary source and yfinance is used as a cross-check.
Execution. Single-asset portfolios remain buy and hold. Multi-asset portfolios use annual rebalancing with a 0.1% drift threshold.
Modeled trading costs. Commission is $0.0005 per share with a $1.99 minimum, plus 0.010% slippage and 0.0050% market impact.
Return windows. One week, year to date, and one year are cumulative returns. Three years and since inception are annualized returns; since inception also reports the cumulative multiple.
Drawdown. Maximum drawdown measures the deepest observed fall from a previous peak. It does not predict the deepest possible future loss.
Volatility. Annualized variability of daily model returns. It describes dispersion, not the lived length of a recovery.
Calmar. Annualized return divided by the absolute maximum drawdown. No risk-free-rate assumption enters that calculation.
The protocol excludes taxes, investor-specific cash flows, account constraints, and individual execution. Those omissions narrow what the model can claim.
Why Calmar is the fixed verdict column
The weekly scoreboard needs one stable risk-adjusted comparison. The Lab uses Calmar because its ingredients are visible: annualized growth and the deepest historical drawdown.
Sharpe remains in the underlying data contract, using a constant 3% risk-free rate. It is not the fixed public verdict column. Its ranking can change with the chosen risk-free rate, measurement window, and asset vehicle - especially when gold histories differ.
Calmar is not “the best” measure. It is the declared measure. That distinction prevents the score from changing simply because another ratio produces a more convenient winner.
One dated example: different questions, different winners
In the frozen record through July 31, 2026, QQQ produced the highest ending value and annualized return. The Permanent Portfolio produced the highest Calmar ratio. Stock-Gold 60/40 sat between those two lines on ending value, annualized return, maximum drawdown, and volatility.
That is not a contradiction. Ending wealth rewards the largest terminal value. Calmar asks how much annualized growth appeared relative to the deepest observed fall. A one-week leaderboard asks something else again.
The example is frozen so the methodology page and portfolio dossiers share one evidence base. Current results continue in the Portfolio Lab: Live weekly series.
Frozen W31 evidence. Same hypothetical $10,000 start; common observation period January 3, 2005 to July 31, 2026; 2026-W31 snapshot; weekly receipt SHA-256.
f6a135b4671fc8468d0de922538ccfd3a1ece69f4d752c496a8d626170e45500
100% QQQ — 20.2× ending multiple | 15.0% CAGR | −53.0% maximum drawdown | 0.28 Calmar
100% S&P 500 — 9.1× ending multiple | 10.8% CAGR | −54.7% maximum drawdown | 0.20 Calmar
Stock-Gold 60/40 — 11.2× ending multiple | 11.9% CAGR | −35.1% maximum drawdown | 0.34 Calmar
Classic 60/40 — 5.6× ending multiple | 8.3% CAGR | −28.4% maximum drawdown | 0.29 Calmar
Harry Browne’s Permanent Portfolio — 4.5× ending multiple | 7.2% CAGR | −18.4% maximum drawdown | 0.39 Calmar
The benchmark’s maximum drawdown was −54.7%, slightly deeper than QQQ’s −53.0% in this window. That does not make QQQ safer: its 2000-02 crash is outside the common clock. It is a useful reminder that a benchmark, a growth extreme, and a full-history risk claim are three different questions.
The public audit trail
Every completed weekly run produces a dated receipt. Publication began with 2026-W31 and continues weekly. The public receipt is the SHA-256 hash of the exact bytes of that week’s latest.json. During the sealed period, the snapshot file remains private; the public record exposes the hash only. One hash proves little - the continuous series is the evidence. The goal is not to make the process look cryptographic. It is to make silent revision harder.
Changes fall into three layers:
Method changes the experiment - for example, a weight, vehicle, data treatment, rebalancing rule, or cost assumption. A method change requires a new version and a new frozen reference snapshot.
Framing changes the question asked of the same experiment - for example, adding a new diagnostic window, changing an axis from linear to logarithmic, or changing aggregation. A framing change must not shrink declared data coverage or hide any in-window sequence or extreme. It must be labeled without pretending the underlying portfolio changed.
Presentation changes how the same result is displayed - for example, improving a chart label. It must not alter the numbers.
Corrections are additive. The old receipt remains available, the correction states what changed, and a new artifact receives a new hash.
The sealed line: a public pre-registration
The Lab currently has one and only one sealed line outside the five public portfolios. It is an eight-asset, annually rebalanced experiment with a max-Calmar objective. Its optimization window is 2003-11-07 through 2022-12-31, with 2023-01-01 onward treated as out of sample. The specification is pre-registered under SHA-256.
e3063e6874bd97767df5368a0ddd5846f0f6fed027e0faee1de96e78dec157de
Until its release condition is met: no peeking, no tweaking, no performance quotes. It does not enter the public leaderboard, and no sealed performance number appears here.
Any future sealed sample must publish its own specification hash before its shadow run begins. When this seal is lifted, a reader can hash the published specification and check the match.
What this method cannot tell you
The Lab can compare five models consistently. It cannot tell any individual reader which portfolio fits their goals, liabilities, taxes, account rules, time horizon, or ability to stay invested.
It cannot turn one historical window into a law of nature. The 2005-2026 record includes particular equity, gold, inflation, and rate regimes. Another era can reorder the results.
It cannot make historical maximum drawdown a safety boundary. Future losses can be deeper, recoveries can take longer, and real execution can differ from modeled execution.
And it cannot reduce portfolio choice to one number. Return measures the destination. Drawdown measures the deepest hole. Recovery time measures how long the old peak stayed out of reach. Volatility measures variation. Each leaves something out.
How the Lab answers the seven questions
The checklist is useful only when it reaches the actual experiment:
Window: January 1, 2005 inception; first common trading observation January 3, 2005; current frozen example through July 31, 2026.
Vehicle: investable ETF vehicles; GLD sets the common boundary and carries its expense ratio and tracking difference.
Return: daily adjusted close with distributions reinvested; below one year cumulative, three years and since inception annualized, with since-inception multiple also shown.
Cash flows: identical hypothetical $10,000 start and no later deposits or withdrawals; the models report time-weighted performance.
Execution: single-asset buy and hold; diversified lines rebalance annually at a 0.1% drift threshold with explicit modeled trading costs.
Verdict: Calmar is declared in advance and never re-chosen; Sharpe remains available with a constant 3% risk-free rate but is not the public verdict column.
Audit: exact weekly
latest.jsonbytes are committed by SHA-256; method, framing, and presentation changes are labeled; corrections are additive.
If another backtest cannot answer these questions, its conclusion may still be useful. You just know where its claim becomes weaker.
Follow the experiment, not a prediction
The methodology page stays fixed while the weekly evidence keeps moving. Follow Portfolio Lab: Live and subscribe to watch the same five rules through new weeks, drawdowns, and leadership changes - without rewriting the experiment after the fact.
Aim right. Stay long.
UnclAlpha | Quant-trained. Simply explained.
Protocol record: v1.0 · Last updated August 4, 2026 · Change layer: n/a - initial release. Future changes will be labeled method, framing, or presentation and appended to this Protocol record; prior receipts remain available.
For educational purposes only. Not investment advice, not actual holdings, and not a recommendation to buy, sell, or hold any security. The Stock-Gold 60/40 is a case-study allocation from the Lab’s backtest research, not a recommendation. Model results are hypothetical and exclude taxes, investor cash flows, account constraints, and individual execution. Historical performance does not guarantee future results. Full disclaimer: https://unclalpha.com/disclaimer
