Skip to content

Research

Research Methodology

BestFolio aggregates published tactical and fixed asset allocation strategies and backtests them under a single, transparent framework. This page is the technical reference for everything you see on a strategy page: which strategies we accept into the catalog, how their backtests are constructed, how transaction costs and regime stress are modeled, how we validate robustness, and how we decide when not to endorse a strategy.

Our research and code are developed with the help of AI tools and reviewed before publication.

New to all this? Start with the plain-English map: where should your money go? It walks through the foundations and the main routes before you dig into the methodology here.

1. Inclusion Criteria

Not every strategy we encounter makes it into the catalog. Before a strategy is implemented, it has to clear four gates. Strategies that fail at any gate are either set aside or, if we implement them anyway for comparison, published in our rejection log with candid reasoning.

  1. Published, documented rules. The strategy must come from an academic paper, a book, a practitioner whitepaper, or a well-documented public post. We do not implement “secret” or proprietary strategies we cannot reproduce from public information.
  2. Mechanical, unambiguous signal definition. The rules must be specified tightly enough that two independent implementers would produce the same signal on the same data. Strategies requiring discretionary judgment or human interpretation are rejected.
  3. Implementable with publicly tradable instruments. Every ticker used must be investable by a retail account — ETFs, mutual funds, or index funds. Strategies that rely on institutional-only instruments, OTC derivatives, or unlisted alternatives are excluded.
  4. Minimum 10-year backtest window. Either the underlying ETFs have enough history, or our documented proxy chain can extend the series to a statistically meaningful window. Strategies that cannot be validated over at least one full business cycle are not admitted.

Note on editorial control. Passing the inclusion gates does not equal endorsement. Several strategies in the catalog — including the Kelly Signal family — are flagged as educational-only and documented in our rejection log with the reasons we do not recommend running them with real capital.

2. Data Sources

  • Price data: Sourced from an institutional-grade market data provider, using adjusted close prices that account for stock splits and dividend distributions.
  • Total return: All returns assume dividends are reinvested at the time of distribution.
  • Macro data: Select strategies use FRED series (unemployment, CPI, yield curves) as signal inputs. All FRED pulls are release-date-aware to prevent look-ahead on revised series.
  • Frequency: Daily prices are fetched and resampled to monthly frequency for signal computation. Daily granularity is retained for drawdown and NAV calculations.
  • Refresh cycle: Price data is refreshed daily. Signals are recomputed at the end of each month.

Credits and notices

Prices of US stocks, ETFs and mutual funds: Data sourced by Tiingo.

This product uses the FRED® API but is not endorsed or certified by the Federal Reserve Bank of St. Louis.

Series retrieved from FRED, Federal Reserve Bank of St. Louis:

  • Source: U.S. Bureau of Labor Statistics via FRED: the unemployment rate (UNRATE) and US consumer prices (CPIAUCSL, CPIAUCNS).
  • Source: Board of Governors of the Federal Reserve System via FRED: the federal funds rate (DFF, FEDFUNDS), Treasury yields (DGS1, DGS2, DGS10, GS1, GS10) and the dollar-euro exchange rate (DEXUSEU).
  • Source: Federal Reserve Bank of St. Louis via FRED: breakeven inflation (T5YIE, T5YIFR).
  • Source: Moody's via FRED: the Baa corporate bond yield (DBAA, BAA), used as a signal input. © 2017, Moody's Corporation, Moody's Investors Service, Inc., Moody's Analytics, Inc. and/or their licensors and affiliates (collectively, "Moody's"). All rights reserved.
  • Source: Chicago Board Options Exchange via FRED: the VIX (VIXCLS), used as a signal input. Copyright, 2016, Chicago Board Options Exchange, Inc. Reprinted with permission.
  • Source: OECD via FRED: the euro exchange rate before 1999 (CCUSMA02EZM618N), national consumer price indices (DEUCPIALLMINMEI, FRACPIALLMINMEI, ITACPIALLMINMEI, ESPCPIALLMINMEI) and, when the OECD's own service is down, its composite leading indicators.
  • Source: Eurostat via FRED: euro-area consumer prices (CP0000EZ19M086NEST). Copyright, European Union.

Other sources:

  • OECD (2026), Composite leading indicators, OECD Data Explorer, https://data-explorer.oecd.org/ (accessed on the date of each data refresh).
  • OECD (2026), Main Economic Indicators, OECD Data Explorer, https://data-explorer.oecd.org/ (accessed on the date of each data refresh).
  • Fama-French factors and portfolios: Ken French Data Library, Tuck School of Business at Dartmouth.

S&P 500 is a registered trademark of S&P Dow Jones Indices LLC or its affiliates, and Russell 2000 is a trademark of FTSE Russell. Other index names are trademarks of their owners. BestFolio is not affiliated with or endorsed by them.

What you may do with third-party data is set out in section 3.5 of the Terms of Service.

3. Backtest Mechanics

  • Rebalancing: Monthly, on the last trading day of each calendar month (unless the strategy’s original rule specifies a different frequency, in which case we preserve it).
  • Signal timing: Signals are computed using end-of-month closing prices. The resulting allocation is applied to the following month. A signal dated Feb 28 uses complete February data and determines what you hold during March.
  • Between signals: The backtest buys the target weights at each signal and holds those positions until the next one, so the weights drift with prices, as they do for a follower who only trades when a signal arrives. A monthly signal trades the drifted positions back to its target, even when the target has not changed, and the cost of that trade is measured from the drifted positions (Section 4). Daily strategies only send a signal when their allocation changes. Portfolios hold their strategies the same way between their rebalances: a fixed-weight portfolio moves back to its target weights at each month end, a walk-forward portfolio at each of its own rebalances.
  • No look-ahead bias: Signals use only data that was available at the time of computation. Macro inputs such as the unemployment rate use first-release values from the FRED ALFRED vintage archive; where the archive is shorter than the backtest, only the earlier segment falls back to revised history, and each signal records which vintage it used. No future data leaks into past decisions.
  • Starting NAV: Normalized to $100 at the beginning of each backtest.

4. Transaction Costs & Stress-Adjusted Slippage

All backtests include realistic transaction cost modeling. Costs are not set to zero — they are deducted from portfolio returns at every rebalance.

  • Base cost: 10 basis points (0.10%) per one-way trade, applied proportionally to portfolio turnover at each rebalance. A full 100% portfolio rotation incurs 0.10% in costs; a partial 30% rotation incurs 0.03%.
  • Stress-adjusted slippage: During periods of elevated volatility, spreads widen and execution quality degrades. Our engine dynamically scales transaction costs by up to 3× the base rate when recent 20-day realized volatility exceeds the baseline (~16% annualized). Backtests already account for the higher friction you would experience during market stress (e.g., March 2020, Q4 2018).
  • Turnover tracking: Every backtest reports total transaction costs, annual turnover, and average trades per year in the Summary Statistics panel. Friction impact is directly auditable.
  • Taxes: backtests are gross of tax. Every metric we publish (CAGR, Sharpe, Sortino, MaxDD, Calmar, turnover, etc.) is pre-tax. All 98+ strategies use the same gross-return engine, so cross-strategy comparisons are apples-to-apples regardless of when each was added. Jurisdictional after-tax estimates live in the US Tax Impact panel on each strategy detail page (US federal rates, holding-period aware), and in the TBSZ analysis for Hungarian users. Other jurisdictions are on the roadmap.
  • What is not included: Margin costs, fund expense ratios (already reflected in ETF NAV), and behavioural factors (delayed rebalancing, panic selling) are not modeled. These vary by investor and brokerage.

5. Survivorship & Selection Bias

  • Faithful implementation. We implement each strategy as described by its original author(s). No proprietary optimization or curve-fitting is applied on top of published rules.
  • Selection transparency. The catalog contains well-known, publicly documented strategies. This is a form of selection bias — we are showing strategies that gained attention, which may correlate with strong historical performance. We try to offset it by publishing the rejection log, which documents strategies we implemented but do not endorse. The Robustness score (Section 11) attacks the same problem quantitatively: it deflates every strategy’s Sharpe for the number of strategy variants we have tested, including the ones we never released.
  • Ticker survivorship. Our proxy fallback chain (Section 6) uses mutual funds and index series that predate modern ETFs. The fund chain is visible in the Price Build Log on every strategy page.

6. Synthetic / Proxy Data (Fallback Chain)

Many ETFs have limited history. To produce longer backtests, we use proxy data for periods before an ETF existed:

  • Fallback chain: Each ticker has a documented hierarchy of proxies. When the primary ETF has no data for a given date, the engine walks down the chain until it finds a valid source. Price continuity is maintained by scaling at each join point.
  • Transparency: The full fallback chain for every ticker in a strategy is visible on the strategy detail page under the “Price Build Log” section.

Key substitution chains (100+ tickers covered):

  • S&P 500: SPY (1993) → VFINX (1987) → S&P 500 total return, daily (1976: the index's daily closes plus each month's dividends from the monthly series) → S&P 500 total-return monthly series (1900). Until 28 September 2026 VFINX filled 1976-1986 too, but as the price providers serve it that fund carries no distribution before 1980 and none of its capital-gains payments of 1980-1986, so those years trailed the S&P 500 by 2 to 9 points a year (1985: +22.7% against +31.7%).
  • Nasdaq 100: QQQ (1999) → ^NDX index (1985, price plus a 0.6% dividend yield) → Fama-French large-growth monthly (1926). Until 28 September 2026 a Vanguard growth fund (VWUSX) filled 1960-1985; its provider history is month-end prints with no dividends and none of its capital-gains payments, which read as one-day drops of 7% to 14% each September, so it was dropped.
  • US Small Cap: IJR (2000) → NAESX (1990); IWM (2000) → ^RUT index (1987) → Fama-French small-cap monthly (1926)
  • US Mid Cap: IJH (2000) → VIMSX (1998) → Fama-French mid-cap monthly (1926, the value-weighted average of NYSE size deciles 6 to 8). The ^MID index tier listed between them has no data source wired up, so until September 2026 every mid-cap chain stopped at 1998 while small caps and large caps already reached the 1920s. The Fama-French tier is a reconstruction of mid-cap behaviour, not the S&P MidCap 400 itself; over the 1998-2026 overlap it tracks IJH at 0.98 monthly correlation with a 3.7% annual tracking error.
  • US REITs: VNQ (2004) → VGSIX (1996) → FRESX (1986) → NAREIT monthly (1971)
  • Int'l Developed: EFA (2001) → VGTSX (1996) → PRITX (1989, T. Rowe Price International Stock) → MSCI EAFE monthly (1969). Until September 2026 the 1989-1996 tier pointed at PRTIX by mistake. PRTIX is T. Rowe Price's US Treasury index fund, so for those seven years every developed-international ticker (EFA, VEA, VEU, VGK, IEFA, SCZ, VXUS) carried a Treasury fund's returns: +32% in 1990, the year MSCI EAFE lost 23%. PRITX covers only the years PRTIX used to cover. Before 1985 its price history is month-end prints only, and from 1985 to 1989 it trailed the EAFE index by about 10 points a year, so the index keeps those years.
  • Emerging Markets: EEM (2003) → VEIEX (1994) → FEMKX (1990) → reconstructed EM monthly series (1925; no investable EM index existed before 1988, so this tail is research-only)
  • Long Treasuries: TLT (2002) → VUSTX (1986) → US 30-year Treasury monthly (1920)
  • Intermediate Treasuries: IEF (2002) → VFITX (1991) → FRED 10-year yield, synthesized to total return (1962) → US 10-year Treasury monthly (1920)
  • Short Treasuries: SHY (2002) → VFISX (1991) → FRED 2-year yield, synthesized to total return (1976) → T-bill monthly (1900)
  • T-Bills: BIL (2007) → ^IRX yield (1970) → T-bill monthly (1900)
  • TIPS: TIP (2003) → VIPSX (2000) → PRTNX (1998) → synthesized TIPS monthly series (1972). TIPS were first issued in 1997, so nothing before 1998 is a traded TIPS price: that tail is a research reconstruction, and the Price Build Log names it whenever it serves.
  • Gold: GLD (2004) → GC=F futures (2000) → London fixing monthly (1968) → gold spot monthly (1920; an official fixed price before August 1971, not a free-floating asset)
  • Commodities: DBC (2006) → ^SPGSCI (1990) → GSCI monthly (1969, energy-heavier than DBC's basket)
  • Aggregate Bonds: AGG (2003) → VBMFX (1986) → US 10-year Treasury monthly (1920; rates only, no credit spread)
  • High Yield: HYG (2007) → VWEHX (1980)
  • Mutual-fund tiers are total return. Where a provider's record of a fund's distributions starts late, the missing payments come from a second provider's record and are reinvested when the history is built: VWEHX, VWESX and VUSTX before December 1989, VBMFX before 1990, FRESX before March 1990, and a handful of early payments of the Fidelity Select, NAESX and Vanguard index funds. Until 28 September 2026 those years were price only, which left out 8 to 15 points a year of bond-fund income and 6 of REIT dividends.
  • Managed Futures: KMLM (2020) → RYMFX (2007, scaled 1.2x) → Barclay BTOP50 monthly (a real index from 1987; the series before that is reconstructed)
  • Leveraged ETFs: Synthetic daily-leveraged returns from underlying, with expense ratio and borrowing cost deductions

The deepest tiers are monthly series (institutional indexes, Fama-French portfolios and, for TIPS, emerging markets and mid caps, research reconstructions), densified to daily and spliced like any other proxy. A strategy only reaches them when its own start date allows, and the Price Build Log on each strategy page names the tier that served every period, so you can see when a backtest is running on a reconstruction rather than a traded price.

Portfolio leverage before a fund existed

The Leverage setting on a portfolio moves part of each broad stock fund it holds into a 2x or 3x version of that fund, and its results show on the portfolio page and in the Leverage impact card. From a fund's first trading day, that swap follows the fund's own traded prices. Before a fund started trading (SPY in 1993, QQQ in 1999, IWM in 2000, EFA in 2001, EEM in 2003, USMV in 2011, and likewise for any other fund), leverage runs on the same stand-in price history the backtests use, the fallback chains above, as far back as that history goes, and the 2x or 3x version is modelled from it with its borrowing cost and fund fee.

A portfolio whose history starts earlier therefore shows the effect of leverage in the market falls before those dates too. Where the stand-in history has only monthly prices, it is turned into daily prices as in the backtests, so a 2x or 3x version built on it misses the daily volatility drag within those months.

7. Walk-Forward Validation

  • Availability: Offered for strategies whose rules include tunable parameters.
  • Method: A rolling-window approach — train on N months of history, test on the next M months, then slide the window forward and repeat.
  • Purpose: Prevents overfitting to in-sample data by validating that a strategy’s edge persists out-of-sample.
  • Identification: Strategies with walk-forward validation results are marked accordingly on their detail pages.

8. Walk-Forward Portfolio Optimization

Beyond validating a single strategy, BestFolio offers walk-forward optimization for multi-strategy portfolios. Given a basket of strategy variants, the engine rolls a historical window forward and re-optimizes the sleeve weights at each rebalance date using only data available at that point in time. The resulting track record is genuinely out-of-sample — the weights held in any given month were never fit on that month’s returns.

Available criteria. Seven objective functions can drive the optimization. Two of them implement classical Markowitz mean-variance optimization (Markowitz, 1952, Portfolio Selection) using mean and covariance estimates from the rolling window. The remaining criteria reflect alternative or post-modern frameworks that address known limitations of the mean-variance approach.

  • Max Sharpe Ratio — The tangency portfolio on the Markowitz efficient frontier. Maximizes the ratio of expected return to portfolio volatility, with the same zero cash reference as the Sharpe ratio shown across the site.
  • Min Variance — The Global Minimum Variance (GMV) portfolio on the Markowitz efficient frontier. Minimizes portfolio variance without regard to expected return.
  • Max CAGR — Directly maximizes realized compound annual growth rate over the lookback window. Closer in spirit to growth-optimal / Kelly sizing than to mean-variance.
  • Min Max-Drawdown — Minimizes the worst peak-to-trough decline over the window. A path-dependent objective, distinct from variance-based risk.
  • Max Sortino Ratio — Substitutes downside deviation for total volatility. A post-modern portfolio theory (PMPT) variant that does not penalize upside volatility.
  • Risk Parity — Inverse-volatility weighting (equal risk contribution). Developed explicitly as an alternative to mean-variance optimization, which is known to be sensitive to estimation error in expected returns.
  • Equal Weight — A 1/N baseline. Often surprisingly competitive with optimization in out-of-sample studies (DeMiguel, Garlappi & Uppal, 2009).

Constraints & solver. All optimization is long-only with a user-configurable maximum weight per sleeve (default 40%) and weights summing to 100%. The problem is solved via sequential least-squares programming (SLSQP) at each rebalance date.

9. Parameter Sensitivity

Any strategy with tunable parameters (lookback windows, thresholds, top-N selection) or a choice of input series is at risk of being fragile: a tiny change to a parameter can flip a great backtest into a disaster. We address this at three levels.

  • Defaults match the source. The default parameters published on each strategy page match the original author’s specification. We do not re-tune the defaults to make the headline numbers look better.
  • Walk-forward covers weights, not parameters. Our walk-forward engine re-estimates sleeve weights on the trailing window at each rebalance date, so no weight is ever fit on data that includes the test period. It never re-tunes a strategy’s own lookbacks or thresholds, so it is not evidence about how robust those are.
  • Sensitivity surfaces (roadmap). We are building a UI to display the full parameter sensitivity surface (e.g., CAGR and max drawdown as a function of lookback window in months) so readers can see how rigid or fragile a strategy’s published parameters are before committing capital.
  • Input-series choice. A strategy that reads an economic series can depend more on which series it reads than on any lookback. Inflation Compass is the clearest case: over 2003–2026, with everything else pinned, its published breakeven gauge and a trailing-CPI gauge differ by about 6 points of CAGR and 22 points of maximum drawdown. Where this applies, the strategy page states the magnitude, not only the direction.

10. Regime Analysis

Aggregate metrics (CAGR, Sharpe, max drawdown over the full backtest) hide an important question: when did the strategy do well or badly? A tactical strategy that shines only in the 2009–2021 QE regime is not the same animal as one that held up in stagflation, the dot-com bust, and 2022 simultaneously.

  • Full-period display. Every strategy page shows the equity curve and drawdown curve across the full available history, not a cherry-picked window. The backtest-period line on each page tells you exactly which years are in the sample.
  • Stress windows. The Summary Statistics panel reports performance during well-known stress regimes (2000–2002 dot-com, 2007–2009 GFC, Q4 2018, Q1 2020 COVID, 2022 dual drawdown) where the data supports it. Strategies that survived only one of those regimes are clearly visible. The stress windows table puts every strategy and variant side by side through the same five windows.
  • Regime-split metrics (roadmap). Our research pipeline includes a planned regime-tagging layer (growth/recession, expansion/contraction, high-vol/low-vol, rising/falling rates) to split every backtest into sub-period metrics. This is on the roadmap; ETA in a monthly release note.

11. Robustness Score (Deflated Sharpe Ratio)

Sections 7–10 stress each strategy on its own terms. The Robustness score answers a question aimed at us rather than at any single strategy: after backtesting an entire catalog, how well does this strategy’s Sharpe ratio hold up against the best that pure luck would have produced across everything we tried? The score is the Deflated Sharpe Ratio (DSR) of Bailey & López de Prado (2014), a number from 0 to 1. It is a historical statistical assessment, conditional on the trials we counted and on the moments estimated from each track record; it is not a probability that the strategy will keep working. It appears as the Robustness column on the leaderboard, as a Robustness badge on strategy pages, and as Deflated Sharpe in the Metrics Comparison table on the Performance page. All three surfaces share one computation, so the numbers always agree.

A catalog is a multiple-testing machine: run enough backtests and the best one will look brilliant by chance alone. A raw Sharpe ratio ignores that. The DSR deflates it with three corrections:

  • Selection breadth. N is every strategy variant we have ever backtested and kept a result for, released or not and retired or not: the leaderboard rows, the reference families we keep as baselines (the single-asset 200-day trend variants, for example), the research strategies that never made it to the site, and variants we have since retired. A trial is a trial whether or not we shipped it or kept it, so N is larger than what the leaderboard shows and retiring a variant never lowers it; the live value, with the number of retained variants and of usable trials, is in the Robustness tooltips. From the spread of Sharpe ratios across those N trials we compute the Sharpe the best one would be expected to show by pure luck, and a strategy’s Sharpe has to clear that bar rather than zero. Because N grows as we test more, scores drift down when we add strategies, published or not. That is the metric working as intended, not a bug.
  • Track-record length. The same Sharpe is worth more over 50 years of monthly returns than over 8. Shorter histories earn lower scores.
  • Non-normality. The computation uses each backtest’s actual monthly skewness and kurtosis rather than a normal assumption. Negative skew and fat tails, exactly the return profile that makes a raw Sharpe overstate how dependable a strategy is, reduce the score.

The result is the estimated probability, on the backtest alone, that the strategy’s true Sharpe exceeds the expected best-of-N-by-luck benchmark. Mechanically, the DSR is the Probabilistic Sharpe Ratio evaluated at that deflated benchmark, computed on per-period (monthly) Sharpe as in the original formulation. Scores below 0.90 are flagged as fragile. Most long-history strategies in the catalog score close to 1.0, so the cutoff keeps the warning on the genuinely fragile minority instead of colouring the whole field. A fragile flag does not mean a strategy is broken; it means its track record cannot yet statistically separate skill from luck, given how many strategies we tried.

Two honest caveats. The score is a property of the full track record, so it does not change with the period filter selected on the leaderboard. And even this N is a floor: parameter variations we tried and discarded are not logged as separate trials, and no N we compute covers every strategy the industry has ever tried, so treat a high Robustness score as necessary evidence, not sufficient proof. Reference: Bailey, D. H. & López de Prado, M. (2014), “The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality”, The Journal of Portfolio Management, 40(5). Our returns are monthly and their lag-1 autocorrelation has a median of -0.03 across the catalog, so the autocorrelation-aware standard error of Lopez de Prado, Lipton and Zoonekynd (2026) moves no score by more than 0.02; the 2014 formula stands.

12. Sharpe Stability Ratio

The Sharpe Stability Ratio (SSR) measures how consistently a strategy’s rolling Sharpe stayed above zero. Two strategies can have the same full-history Sharpe even when one earned its returns steadily and the other depended on a few good stretches. The Stability column on the leaderboard makes that difference visible over each selected period. The badge on a strategy page uses that variant’s full history in USD.

We take simple monthly returns from the last NAV observation in each calendar month, in the selected currency, and compute a rolling 36-month Sharpe: the arithmetic mean divided by the sample standard deviation, with a cash reference of zero and no annualisation. We divide the mean of those rolling Sharpes by their long-run standard deviation, estimated with Newey-West, Bartlett (triangular) weights and bandwidth 35. Each autocovariance is divided by the total number of rolling windows. The benchmark Sharpe is zero. We require at least 120 monthly returns (ten years) and show n/a if any rolling deviation or the long-run variance is zero. Stability is recomputed for every leaderboard period.

The source paper lets a data-driven rule choose how many lags the long-run variance uses (15 in its 29-year hedge-fund sample) and reports that the ranking barely moves with that choice. We fix it at 35 lags, the full overlap of a 36-month window, so every strategy is measured with the same estimator. That makes our values lower than the paper’s method would give, so compare orderings across sources, not levels.

A value near 0 typically means the Sharpe swung between positive and negative stretches. Around 1 suggests steady risk-adjusted performance across the regimes covered. Values far above 1 deserve suspicion rather than applause: Fairfield Sentry, a Madoff feeder fund, scored 1.55 in the Portfolio Optimizer example linked below. A high score is a reason to examine the record, not proof that the returns are dependable.

The period matters. In our September 5, 2026 catalog check, the S&P 500 scored 0.30 since 2000, and 156 of the 169 variants with data over that window beat it. Starting in 2008, the S&P 500 scored 0.74 and beat two thirds of them. A window dominated by a long bull market can make buy-and-hold look especially steady. Compare the same period and currency, and try several start dates before drawing a conclusion.

Reference: Bajo Traver, M. and Rodriguez Dominguez, A., The Sharpe Stability Ratio: Temporal Consistency of Risk-Adjusted Performance (SSRN 6344658). Thanks to Roman Rubsamen of Portfolio Optimizer for the suggestion and for the explanation of the ratio that prompted this addition.

The catalog-wide study behind this column is published as a working paper by BestFolio Research: The Sharpe Stability Ratio in Tactical and Static Asset Allocation Backtests: Evidence from 173 Published Strategy Variants (DOI: 10.2139/ssrn.7441662). The paper covers 173 published variants and reports findings unfavourable to the statistic as well as favourable ones. The replication archive holds the code and result tables behind every number in the paper.

13. Metrics Computation

All metrics are computed from the monthly NAV series unless otherwise noted.

MetricDefinition
CAGRCompound Annual Growth Rate. Annualized geometric return derived from the full NAV series.
Sharpe Ratio(Mean monthly return × 12) / (Sample standard deviation of monthly returns × √12). The cash reference is zero: no risk-free rate is subtracted, so this is the zero-reference Sharpe. Computed on month-end NAV in the displayed currency, and one definition on every page and every period, including the leaderboard windows, the comparison tables and portfolios. It is not CAGR divided by the volatility column, which is a daily-basis figure.
Sortino RatioSame numerator as Sharpe (mean monthly return × 12) over the annualized downside deviation, so upside volatility is not penalized. Downside deviation is the root mean square shortfall below a minimum acceptable return of zero, measured across every month rather than only the losing ones: months at or above zero contribute zero. A track with no losing month has no Sortino rather than an infinite one. Same monthly basis and same definition on every page and period as the Sharpe beside it.
VolatilityAnnualized volatility: the standard deviation of returns, scaled to a year. Strategy pages, portfolio backtests and the leaderboard's strategy rows use daily returns (times √252). Library and shared portfolio pages, the walk-forward page and model portfolio rows on the leaderboard's Full History use month-end returns (times √12). It is not the Sharpe ratio's denominator, which is always monthly. Where a backtest runs on the monthly series of Section 6, the days inside each month are filled in at a constant rate, which smooths daily returns, so daily-basis volatility over those years reads lower than a traded fund's would.
Max DrawdownLargest peak-to-trough decline in NAV, as a percentage of the peak. Which NAV depends on the view: strategy pages, portfolio backtests and the leaderboard's Full History view in USD use daily NAV, while windowed and EUR leaderboard views, strategy cards, the library and shared portfolio pages use month-end NAVs, which miss troughs inside a month and so read shallower.
RecoveryHow long the worst drawdown took to regain its previous high: from the last day at that high before the fall to the first day back at or above it, shown in years. It is measured on the same NAV as Max Drawdown in each view, so depth and duration describe one episode. A strategy still below that high shows the time so far, marked as still underwater. Longest DD is a different figure: the longest time below a previous high in any episode, often a shallower one.
Calmar Ratio (MAR)CAGR divided by the absolute value of max drawdown, both over the same period and on the same NAV as the Max Drawdown shown beside it. A path-dependent risk-adjusted measure. Over a strategy's whole record this is the MAR ratio (compound annual return since inception over the maximum drawdown since inception), so the leaderboard labels it MAR on Full History and Calmar on trailing and custom periods. The original Calmar ratio looks back 36 months only.
Keller RatioThe return adjusted for drawdown that Wouter Keller and Jan Willem Keuning introduced in their 2017 paper Breadth Momentum and Vigilant Asset Allocation (VAA), later named after Keller: K = R × (1 − D / (1 − D)) when R ≥ 0 and D ≤ 50%, and K = 0 otherwise, where R is the CAGR and D the maximum drawdown as a positive fraction. D / (1 − D) is the gain that recovers a drawdown of D, so the cut grows faster than the drawdown and takes the whole return at 50%, which needs a 100% gain to recover. K reads like a return: 10% CAGR with a 20% drawdown gives 10% × (1 − 0.25) = 7.5%. The authors measure D on month-end values. The leaderboard uses the CAGR and Max Drawdown of the same row, the pair Calmar divides, so on Full History in USD, where that drawdown is measured on daily NAV, K can read slightly lower than on their convention. Keller's later papers also report a stricter K25 = R × (1 − 2D / (1 − 2D)), zero at a 25% drawdown; the leaderboard shows the original.
SWRSafe Withdrawal Rate. The maximum constant, inflation-adjusted annual withdrawal rate that survives every rolling 30-year window in the backtest. Withdrawals are scaled by recorded inflation in the display currency. In USD that is US CPI-U (CPIAUCNS from 1913, CPIAUCSL from 1947), so century-scale backtests are tested against the 1929 deflation and the 1940s and 1970s inflation rather than a flat rate. In EUR it is euro-area HICP from December 1996, chained back to 1955 on a synthetic euro-area CPI: the OECD consumer price indices of Germany, France, Italy and Spain, combined with fixed weights taken from Eurostat's 1999 euro-area HICP country weights (41%, 25%, 23% and 11% after normalising over the four). A window that starts before its inflation series (1913 in USD, 1955 in EUR) assumes 3% a year instead. Before September 25, 2026, EU mode used HICP alone, which covers no full 30-year window, so every EUR rate rested on that 3% assumption.
PWRPerpetual Withdrawal Rate. The maximum withdrawal rate that preserves real (inflation-adjusted) capital over every rolling 30-year window, on the same inflation series as SWR in each currency.
Ulcer IndexThe root mean square of the drawdown below the running high, taken at every observation, so a decline weighs more the deeper it goes and the longer it lasts. Lower is better. Strategy and portfolio backtests measure it on daily NAV; the leaderboard and the library on month-end NAVs. Every page and CSV download shows it as a percentage, the unit of the drawdowns it averages (6.79%, for example).
UPIUlcer Performance Index. CAGR divided by the Ulcer Index (drawdown depth and duration). Higher values indicate better risk-adjusted performance with emphasis on drawdown pain.
Robustness (DSR)Deflated Sharpe Ratio. Probability (0 to 1) that the Sharpe reflects a real edge rather than the luckiest pick among everything we tested, corrected for track-record length, fat tails, and the number of variants tried. See Section 11.

14. Strategy Variants: SmartStack & SmartLeverage

Many strategies on BestFolio expose more than one tab in the variant selector (Standard, Leveraged 2x, SmartStack, SmartLeverage, and so on). These overlays keep the strategy’s underlying signal logic unchanged and modify only the execution step, i.e. which tickers and weights the base allocation is translated into. Two of those overlays are proprietary: SmartStackand SmartLeverage.

SmartStack (return-stacking overlay)

SmartStack layers a gold + managed-futures diversifier on topof the base allocation rather than funding it by selling equities. The engine walks each position in the base strategy’s signal output and applies a priority ladder:

  1. Prefer a return-stacked ETF if one exists for the position (for example, LQD can be replaced by RSBT, which embeds bonds plus managed futures in a single instrument).
  2. Otherwise, substitute a 3x leveraged ETF at 1/3 weight (SPY to UPRO, QQQ to TQQQ, TLT to TMF). This preserves full economic exposure to the original asset while freeing 2/3 of capital.
  3. If no 3x product is available, substitute a 2x ETF at 1/2 weight (VNQ to URE, VEA to EFO). This frees 1/2 of capital.
  4. For assets without a suitable leveraged product (DBC, AGG, BND), keep the full weight and layer an equal-sized managed-futures (KMLM) sleeve on top. This piece is managed futures only, with no gold counterpart.
  5. Cash-like positions (BIL, SHV, SGOV) pass through unchanged.

The split between gold and managed futures is not a fixed 50/50. Only the capital freed by the leverage substitutions (steps 2 and 3) is split 50/50 between GLDM (gold) and KMLM (managed futures). The step-4 overlay adds KMLM on its own, with no matching gold. So whenever the base signal holds a commodity or aggregate-bond position, managed futures carries more weight than gold; when it holds none, the two come out close to 50/50. Because the base strategy rotates its holdings every month, the realized gold-to-managed-futures ratio shifts month to month, and gold never exceeds managed futures by construction. The overlay is funded only by the capital the substitutions free. When step 4 asks for more managed futures than that (any month holding DBC, AGG or BND), the overlay is scaled down to fit and the strategy legs are never touched, so the base strategy always keeps 100% of its own exposure. Until September 2026 the engine scaled every position instead, strategy legs included, which left the base strategy at 80% of its own exposure in about a quarter of all months; the five SmartStack variants were restated when that changed.

Worked example. Suppose the base signal holds four assets at 25% each: QQQ, IWM, VEA (equities) and DBC (commodities). Steps 2 and 3 convert the three equity sleeves into fractional leveraged positions (TQQQ 8.3%, TNA 8.3%, EFO 12.5%) and free roughly 46% of capital, which would split into about 23% GLDM and 23% KMLM. DBC stays at full weight (step 4) and asks for an equal 25% KMLM sleeve. The overlay would then need 71% of capital against the 46% actually free, so it is scaled by about 0.65 and lands at GLDM 15% and KMLM 31%, while the four strategy legs keep their full exposure. The 16-point gap between gold and managed futures is the commodity sleeve after the cap.

The net result: the investor keeps full exposure to the base strategy and adds a diversifying return stream, all inside a standard brokerage or UCITS account with no margin and no futures.

Capital weight is different from exposure. A 3x substitution frees two thirds of that sleeve, so gold and managed futures can become the majority of the account. For example, a 100% IEF defensive signal becomes roughly 33.3% TYD, 33.3% GLDM and 33.3% KMLM: about 1.67× estimated product exposure. A 100% BIL signal stays 100% BIL. The defensive label describes the base signal; it does not remove leveraged Treasury or alternatives risk. Daily leverage targets do not promise the same multiple over a longer holding period.

Each SmartStack variant shows the latest, median and maximum ETF capital weights and estimated product exposure from its full published backtest, plus how often defensive observations held a majority directly in gold and managed futures. Gold holdings can include a base-strategy position, so they are not labelled entirely as an added overlay. RSBT embeds managed futures as well as bonds; it counts as 2× product exposure and is shown separately from direct GLD/GLDM and KMLM holdings. These estimates include cash at 1× and do not measure the gross notional of derivatives inside the funds.

Why we built it. Classic diversification funds gold or managed futures by cutting equities, which in a strong equity decade turns diversification into a drag. Institutional investors solve this with futures-based portable alpha. SmartStack reproduces the effect using only retail-accessible leveraged and return-stacked ETFs.

SmartLeverage (fractional leverage overlay)

SmartLeverage is a leverage dial from 1.0x to 3.0x that amplifies the base strategy’s total exposure while keeping total notional at 100%. It does not naively swap every ticker for its leveraged equivalent. Instead, the engine uses a two-pass greedy algorithm:

  1. Pass 1 (prefer 2x): upgrade positions from 1x to 2x ETFs first. 2x products have less volatility decay and lower expense ratios than 3x, so they are the preferred lever.
  2. Pass 2 (escalate to 3x): if 2x cannot absorb enough notional to reach the target leverage, selectively upgrade positions from 2x to 3x. A maximum-leverage calculator determines the theoretical ceiling for any given allocation.
  3. Cash buffer: any capital not absorbed by leveraged positions is parked in a cash-like ETF (default BIL). Cash-like positions in the base allocation pass through unchanged.

Example: a 1.2x leverage on a 50/50 SPY/IEF portfolio produces roughly SSO 30%, IEF 60%, BIL 10% — 1.2x exposure with only 90% capital deployed and the remainder in cash.

Difference at a glance

  • SmartStack changes what you hold: it adds a gold + managed-futures overlay on top of the base allocation.
  • SmartLeverage changes how much you hold: it amplifies the base allocation’s total exposure without changing its asset mix.
  • Both are execution-layer transforms. The underlying signal, rebalancing cadence, and research logic of the base strategy are untouched.
  • Both are opt-in. Toggling between Standard and a Smart* variant on a strategy page simply swaps the execution overlay and re-runs the backtest. The Asset Universe panel updates to show the actual tickers the variant trades.

15. Research Pipeline & Changelog

BestFolio is a living research platform. We publish new research every month: new strategies and variants once they pass the inclusion criteria, and revisions to existing implementations whenever a source paper is updated or a bug is reported. There is no fixed number of new strategies per month. A candidate that fails review goes to the rejection log, not into the catalog.

Current catalog: 98 strategies, 77 tactical + 21 fixed, 6 free forever.

Recently added. Latest additions to the catalog (most recent first):

AddedStrategyAuthor
2026-04-11Dynamic Macro Allocation (Sadek)Bill Sadek
2026-04-11Alpha-One MomentumBill Sadek
2026-04-01Composite MomentumBestFolio Research
2026-04-01Growth-Inflation Sector TimingInspired by David Varadi (CSS Analytics)
2026-03-27Golden RatioBogleheads Community

See the full strategy catalog for the complete list, filters, and sort order.

16. Rejection Log

Any research platform that only publishes strategies it likes is cherry-picking. We maintain a public rejection log documenting strategies we implemented honestly, backtested under this same framework, and then decided we could not recommend with real money. Each entry includes the strategy rules, our backtest metrics, the specific failure modes, and our verdict.

Read the rejection log →

17. Behaviour Check (live vs backtest)

A backtest is a claim about how a strategy behaves. Once it runs live, that claim becomes checkable, but not on returns. At 15% annual volatility, telling a 2 percentage-point difference in annual return apart from noise takes roughly 225 years of live data. Ten months of live returns says nothing at all.

Behaviour is different. How often a strategy trades, how much of the book it moves, how often it sits in cash: these have a far smaller spread relative to their average than returns do, and implementation faults show up in them as a consistent offset rather than as noise. So behaviour answers in months what returns need decades to answer. This is what the Behaviour Check on each strategy page measures.

How a strategy is scored

We take the strategy's last six rebalances and compare them against every comparable six-rebalance stretch in its own history. Comparable means the market looked like today: windows are matched on the broad equity benchmark's trailing 12-month return and volatility, using fixed thresholds so nothing about the future leaks into the definition of a normal market.

Comparing whole windows rather than individual months is the point. A strategy sitting in cash for six months is not six independent coin flips, and scoring it that way overstates how unusual it is. When we tested the naive month-by-month version over 87 strategies and 176 months it flagged more than a third of the catalogue in a typical month, and the alerts in its lower, “unusual” tier typically lasted a single month. The window method flags about 3.5% of the strategies it can score in a typical month, and a flag has to hold for two rebalances before it is shown at all.

The same 15-year test shows the check's main limit. In about one strategy-month in five there was not enough comparable history to score it at all, because the market was in a state the strategy's own past had rarely seen. The check says so rather than guessing, and that gap is widest when markets are unusual: most of 2020, 2022 and 2023 fell into it. Of the months that could be scored, 95% read “in line”, a strategy scored two months running kept the same verdict 97% of the time, a typical flag lasted two months, and 21 of the 87 strategies were never flagged.

When our own history is wrong

A second, separate check re-runs each recently published signal through today's backtest engine and compares the result with the signal that is in your account. These are different questions: the first is about the strategy, this one is about us.

Some strategies read economic data that is revised after publication. When that happens, a backtest computed today can place the strategy somewhere different from where the live signal placed it at the time, purely because the underlying data has since changed. Where we detect this, the strategy carries a History under review notice naming how many rebalances are affected. In those months, the signals you received are the record; the backtest curve should be treated as provisional.

We publish this rather than quietly correcting it. A ticker that has merely been renamed to its canonical equivalent is not counted, and neither is a single isolated month: a notice appears only when the disagreement persists and changes the actual risk posture.

What it does not tell you

  • It is not a return forecast. A flag predictsdifference from the backtest's typical behaviour, in either direction. It says nothing about whether a strategy is about to do well or badly.
  • It does not prove the edge is real. It checks that the live version is doing what the tested version did, which is a narrower and more answerable question than whether the backtest was overfit in the first place.
  • It cannot score every strategy. Coverage is currently 174 of 225 variants. Annual strategies do not rebalance often enough to fill a six-rebalance window and are not scored at all; quarterly and event-driven strategies are scored only once they have traded enough times. Any strategy also goes unscored while the market is in a state its own history has rarely seen, so the check has least to say when markets are least like the past.

18. Limitations & Disclaimers

  • Hypothetical results: All backtested performance is hypothetical. It does not represent actual trading and was not achieved with real capital.
  • No guarantee: Past performance — backtested or live — does not guarantee future results.
  • Remaining frictions: While transaction costs and stress-adjusted slippage are modeled (Section 4), taxes, margin costs, fund expense ratios, and behavioural factors (panic selling, delayed rebalancing) are not captured.
  • Execution assumption: the backtest fills every rebalance at the closing price of the signal day (day D), so it earns the move from that close to the next. The schedule we publish for followers is the next session’s open (09:30 ET), which does not capture that move. This is a fill-timing gap, not look-ahead; the rebalance sensitivity card on each strategy page carries a “delayed close” line that trades one session later and bounds how much of the return rides on it. Next-open fills are not modeled: adjusted open prices are not in our price history.
  • No universal winner: No strategy is guaranteed to outperform a simple buy-and-hold approach in all market environments. Tactical allocation involves active decisions that can underperform passive benchmarks for extended periods.

Questions about our methodology? Check the FAQ or reach out at [email protected].

See the methodology in action

Explore 98+ strategies with full backtest results, drawdown charts, and signal history. 6 are free forever.