Skip to content

Research / Rejection Log

What we tested and chose not to endorse

Every strategy we publish is backtested under one framework and shown with its full record. This page is the other side of that ledger: the strategies and ideas we put through the same tests and decided not to stand behind. Some, like the Kelly Signal family below, stay listed for comparison. Others we tested and never listed at all, because they did not beat what we already ship. Either way, the reasoning and the numbers are here.

This log is the human-readable companion to our Robustness score, which counts every strategy we have tested, released or not, when it judges whether a track record is real skill or luck. If you disagree with a verdict, the data is on the strategy page or in the numbers below. Challenge it.

Why publish a rejection log at all?

  • Credibility. Any platform that only publishes strategies it likes is cherry-picking. A rejection log is the public record of what we looked at and decided not to endorse, with reasons you can check.
  • Education. The rejected strategies often illustrate the exact failure modes we want readers to recognise in their own portfolios: closed-system assumptions, look-ahead bias, signal lines that never reset, path-dependent leverage.
  • Honesty about authorship. BestFolio is built by one person. If we hide the strategies we no longer believe in, we are lying by omission. This page is the antidote.
Educational onlyAdded 2026-04-11

The Kelly Signal family (3Sig / 6Sig / 9Sig)

Jason Kelly · 2015

Quarterly value-averaging strategies that pursue a fixed 3%, 6%, or 9% quarterly target on a stock ETF while using a bond ETF as a cash buffer. The book's headline claim is that the strategy forces you to buy low and sell high, quarter after quarter. The backtest says otherwise.

How the strategy works (as written)

  1. Pick a target quarterly growth rate (3%, 6%, or 9%) and a stock/bond ETF pair.
  2. Start at a fixed stock/bond split (e.g., 80/20 for 3Sig, 50/50 for 6Sig).
  3. Each quarter, compute the target stock-sleeve value by compounding the growth rate from the starting value.
  4. If the stock sleeve is below target, buy stocks with bond money. If above, sell stocks and park the proceeds in bonds.
  5. Apply the "30 Down" rule: if the stock ETF has fallen 30%+ from its all-time high, move all bonds into stocks immediately.

Our backtest, honest implementation

VariantCAGRMax DDSharpeVol
3Sig (IJR / BND)
1987-04 to 2026-04 (38.9 years)
10.6%-48.9%0.7016.5%
6Sig (MVV 2x mid-cap / AGG)
1987-04 to 2026-04 (38.9 years)
16.8%-84.2%0.6234.7%
9Sig (TQQQ 3x Nasdaq / AGG)
1987-04 to 2026-04 (38.9 years)
26.9%-94.3%0.7059.0%

For context, our Classic 60/40 benchmark (SPY/AGG) returns 8.7% CAGR at -34.7% max drawdown (Sharpe 0.78) over its full 46-year history, and a plain S&P 500 buy-and-hold returns 8.6% CAGR at -56.8% max drawdown (Sharpe 0.47) over its 36-year history. All three Kelly variants have a lower Sharpe than 60/40, and the 6Sig and 9Sig drawdowns are far deeper than SPY's worst.

Why we do not recommend it

Closed-system backtest flaw (the critical one)

The book's own examples assume periodic cash infusions from the investor's salary to replenish the bond sleeve. A faithful closed-system backtest, no new money, just rebalancing, is NOT what the book describes. Once the bond sleeve is depleted after a severe drawdown, there is nothing left to buy the dip with, and the strategy degenerates into leveraged buy-and-hold. Our backtests reflect that degenerate regime faithfully.

The signal line never resets downward

The target value for the stock sleeve compounds forward at 3/6/9% per quarter regardless of what the market has actually done. After a crash, the target sits far above reality, the rebalancer goes all-in on stocks, and, if there is no fresh cash, the strategy is now 100% stocks at the worst possible time.

The 30 Down rule is a capitulation trigger in disguise

By design, the 30 Down rule dumps the entire bond sleeve into stocks once the ETF is 30%+ below its all-time high. That can be a fine rule if you have dry powder coming in next month. In a closed system it is a one-way valve: all future upside is levered, all future downside is fully exposed.

Leveraged variants (6Sig/9Sig) are path-dependent traps

MVV is 2x daily mid-cap. TQQQ is 3x daily Nasdaq. Both suffer volatility decay on top of the mechanical flaws above. Our 9Sig backtest shows -94.3% max drawdown over 38 years. An investor who held through it would have needed to watch $100,000 turn into $5,700 and then wait years to break even, a level of drawdown that is psychologically and practically unsurvivable.

It does not beat simple buy-and-hold on risk-adjusted terms

Our Classic 60/40 benchmark (SPY/AGG) produces 8.7% CAGR with a -34.7% max drawdown and 0.78 Sharpe over its full 46-year history. A plain S&P 500 buy-and-hold produces 8.6% CAGR with a -56.8% max drawdown and 0.47 Sharpe over its 36-year history. All three Kelly variants have a LOWER Sharpe than 60/40, and 6Sig / 9Sig have drawdowns that dwarf even SPY's 2008 experience. For 3Sig specifically: similar CAGR to buy-and-hold but with a longer peak-to-recovery path and an added layer of mechanical complexity. There is no version of the Kelly family where the risk-adjusted outcome is worth the complexity, the tax drag, or the behavioural hazard.

Verdict

We keep the three Kelly variants listed as 'Educational only' because we implemented them honestly and we want readers to be able to compare them directly to the strategies we do endorse. But we have never run them with real money, and we recommend that you don't either. If you like value-averaging as an idea, use it for contributions (adding new money to a target), not as a rebalancing rule for a closed portfolio.

Source

Jason Kelly, "The 3% Signal: The Investing Technique That Will Change Your Life" (Plume, 2015) and later variants.

Related

Tested, not adoptedAdded 2026-07-08

Managed futures as the defensive leg (instead of bonds)

Reader submission, modeled in-house · 2026

A reader suggested reversing the defensive leg of one of our dual-momentum strategies: instead of moving to bonds when the momentum gate turns risk-off, move to a pure managed-futures sleeve (a 70/30 KMLM/DBMF blend). The thesis is that managed futures are crisis alpha, so they should beat bonds exactly when the strategy plays defense. We modeled it back to 1987. It loses on both return and drawdown.

How the strategy works (as written)

  1. Run the base strategy's momentum gate unchanged (risk-on when the trend and absolute momentum are both positive).
  2. Risk-on: hold the offensive sleeve as normal.
  3. Risk-off: instead of rotating to Treasuries or T-bills, hold 70% KMLM / 30% DBMF (pure managed futures).
  4. Reassess monthly.

Our backtest, honest implementation

VariantCAGRMax DDSharpeVol
Baseline: bonds risk-off (shipped HAA-Simple RSST)
1987 to 2026 (monthly)
16.4%-18.8%1.1314.5%
Managed futures risk-off (70/30 KMLM/DBMF)
1987 to 2026 (monthly)
15.5%-26.8%0.9816.0%
Diversified risk-off (bonds + managed futures + gold)
1987 to 2026 (monthly)
16.3%-21.3%1.11--

Deep managed-futures history before 2020 is a shared proxy chain (KMLM to RYMFX to BTOP50 to an academic trend series), so the 70/30 split is cosmetic pre-2020. Drawdowns are monthly and therefore shallower than our daily production engine. Neither caveat changes the ranking.

Why we do not recommend it

Managed futures as defense is mistimed by construction

A momentum gate flips defensive only after the trend has already turned, typically one to four months late. Managed futures pay at the onset of a trend, not in its mature or reversing phase, so the gate holds you in them through exactly the whipsaws that hurt them most. 2022 is the tell: it was a banner year for trend-following, yet the reversed strategy earned +0.2% across its defensive months versus +2.4% for bonds, because the gate missed the front-loaded January-to-April gains and held through the second-half reversals.

Bonds simply won the risk-off months

Isolating only the months the gate was defensive, a best-of Treasuries/T-bills leg returned +6.1% per year at 5% volatility. The managed-futures sleeve returned +4.2% per year at 13.5% volatility. Higher return at under 40% of the risk: when the strategy actually needs its defensive leg, bonds delivered more for less.

The crisis-alpha reputation does not hold in the crises that matter

In 2008 the managed-futures leg returned +7.4%, but Treasuries returned +10.7% and gold +16.7%. Managed futures did win the slow trending bears (the 1987 crash, the 2000-2002 dot-com grind), and lost 2008, 2013-14 and 2015-16. As a general-purpose defensive leg it is the wrong tool.

Verdict

We did not ship it. Managed futures earn their place always-on inside a diversified sleeve, which is exactly where they already live in our RSST strategies, not bolted onto a lagging defensive trigger. A reader who wants managed futures in the defense can get parity, not an edge, from a diversified bonds-plus-managed-futures-plus-gold basket. Nothing in this study came close to the baseline's risk-adjusted profile.

Source

Proposed by a BestFolio reader; modeled against our shipped HAA-Simple RSST strategy over 1987 to 2026 (monthly, with realistic switching costs).

Related

Tested, not adoptedAdded 2026-07-07

A credit-strength sensor on top of trend and volatility

BestFolio Research · 2026

A recurring idea is that credit spreads lead equities, so a high-yield credit-strength signal (a HYG-versus-Treasuries ratio, or the high-yield option-adjusted spread) should sharpen a market-timing rule. We added a credit z-score on top of a rule that already used a price-trend gate and a volatility gate, then measured what the credit signal contributed on its own. The answer is nothing, slightly worse than nothing.

How the strategy works (as written)

  1. Score three regime sensors on the S&P 500: a price-trend gate, a volatility-term gate, and a credit-strength gate (high-yield versus intermediate Treasuries).
  2. Combine them into a single risk-on/risk-off decision.
  3. Hold equities risk-on, cash risk-off, reassess monthly.

Why we do not recommend it

Adding the credit sensor was marginally negative

In a leave-one-out test on the S&P 500 from 2008 to 2026, adding the high-yield credit z-score on top of the trend and volatility gates moved CAGR by -0.54 percentage points, left the Sharpe essentially unchanged, made the worst drawdown 0.7 points deeper, and added 19% turnover. On net it subtracted a little while adding trades.

Credit strength is already inside price trend

By the time high-yield spreads widen enough to flip a credit gate, the equity trend gate has usually already moved. The information the credit signal carries is largely priced into the thing you are already watching, so stacking it on adds trades without adding foresight.

The simpler rule we ship already wins

The full three-sensor rule managed 8.8% CAGR at a 0.39 Calmar ratio. Our plain 200-day trend rule did 9.4% at a 0.44 Calmar and a -21.5% drawdown, using roughly a quarter of the trades. More sensors read as more sophisticated; here they were strictly worse.

Verdict

Not adopted. The sensor ranking from the study is worth stating plainly: the trend gate is essential (removing it deepens drawdowns by more than 40 points), the volatility gate is valuable, and the credit gate is harmful. A candidate that differs from something we already ship by one added sensor has to prove that sensor is additive before it earns a slot. This one made things worse.

Source

In-house research candidate; leave-one-out ablation on the S&P 500, 2008 to 2026, every sensor computed literally on real data.

Tested, not adoptedAdded 2026-06-17

The 20% leveraged-ETF backtest that is really 13%

BestFolio Research · 2026

A popular strategy holds 3x S&P 500 and 3x long Treasuries, volatility-targeted to a fixed level, and reports around 20% CAGR on a backtest that runs to 1987. The long history comes from tripling the daily returns of old index funds. That shortcut is where the 20% comes from. Model the leverage the way real 3x ETFs actually behave and the same rules return about 13%.

How the strategy works (as written)

  1. Hold 60% 3x S&P 500 / 40% 3x long Treasury.
  2. Each month, scale total exposure so the blend's trailing realized volatility hits a fixed target (around 14%), parking the remainder in cash.
  3. For history before the 3x ETFs existed, synthesize them from the underlying index.

Our backtest, honest implementation

VariantCAGRMax DDSharpeVol
As marketed (naive 3x, tripled daily returns)
as published, from 1987
20.2%-26.0%----
Same rules, calibrated leverage costs
1987 to 2026
12.6%-38.8%0.8715.0%

Our synthetic-leverage cost model is not a guess. Regressed against real UPRO, SPXL, TQQQ and TMF over 16 to 18 years, it reproduces real UPRO to within 0.3 percentage points of CAGR per year (realized beta 2.99, R-squared 0.997). The gap above is the cost of borrowing to run 3x leverage, which the naive method ignores.

Why we do not recommend it

Tripling daily returns is not what a 3x ETF does

A real 3x ETF borrows to lever up, and that borrowing costs money every day: the financing rate plus a spread. Backtests that just multiply an index's daily return by three skip that cost entirely. Over decades it compounds into 5 to 7 percentage points of CAGR per year of pure illusion. This is the same trap as the leveraged Kelly variants above.

The drawdown is understated too

The marketed version reports a -26% max drawdown. With realistic leverage costs and daily resolution, the same rules take -38.8% through the dot-com bust. Cheaper-than-real leverage flatters the drawdown as well as the return.

Volatility targeting is real, the headline is not

To be fair to the idea, volatility targeting genuinely helps: it contained the 2022 bond crash far better than a static leveraged blend, and the honest 13% version still beats a plain S&P buy-and-hold on return, Sharpe and drawdown. The strategy is not junk. The 20% is.

Verdict

We did not adopt the strategy as published, because the headline depends on leverage that does not exist at that price. We did keep what actually works: our own leveraged strategies are trend-gated and volatility-targeted, and every one is modeled with the calibrated leverage costs above, so the numbers you see are the numbers you would have lived. If a leveraged backtest reports a number that looks too good, check whether its leverage is free.

Source

A widely-shared volatility-targeting strategy on synthetic 3x equity and Treasury ETFs, marketed at roughly 20% CAGR; re-modeled in-house with empirically calibrated leverage costs, 1987 to 2026.

Related

Tested, not adoptedAdded 2026-06-16

Gold as the risk-off leg (instead of T-bills)

BestFolio Research · 2026

Some regime strategies park their risk-off capital in gold instead of cash, on the logic that gold is a crisis hedge. We tested a gold defensive leg against our default of T-bills, on the actual signal one of our regime strategies uses. Over the window we can honestly test, it added almost no return, hurt the Sharpe, and deepened the drawdown in every crisis, because gold itself crashed in 2008.

How the strategy works (as written)

  1. Run the strategy's regime composite unchanged.
  2. Risk-off: hold gold when gold is above its 200-day average, otherwise T-bills (instead of always holding T-bills).
  3. Reassess on the strategy's normal cadence.

Why we do not recommend it

Gold crashed in the crisis it was supposed to hedge

In the 2008 liquidity panic gold fell hard as investors sold everything they could. A strategy that fled to gold took a -38.9% drawdown through the financial crisis versus -22.3% for the same strategy holding T-bills. A defensive leg whose job is to cushion a crash cannot itself crash in the crash.

Almost no return for the extra risk

Over 2007 to 2026 on the live signal, the gold leg added 0.3 percentage points of CAGR (13.1% versus 12.8%), while lowering the Sharpe (0.62 versus 0.64) and roughly doubling turnover. The one number that moved in gold's favor is inside the margin of noise; everything that matters moved against it.

The impressive backtest is a gold-regime artifact

A naive 25-year isolation of a gold defensive leg looks spectacular, but almost all of that edge is the 2001-2007 gold bull market, a period our credit-dependent composite cannot even trade. Strip the window back to what the real signal can act on and the edge disappears. T-bills do not crash, and that turns out to be the whole point.

Verdict

Not adopted. Our regime strategies hold T-bills when they play defense, and they will keep doing so unless someone is making an explicit inflation-regime bet and has separately solved the 2008-style gold-crash exposure. A defensive leg is judged by what it does in the worst months, and on that test gold failed.

Source

In-house study comparing a gold risk-off leg against our shipped T-bill default, run on the live six-signal regime composite over 2007 to 2026.

Tested, not adoptedAdded 2026-07-13

The 3x-Nasdaq trend timer whose backtest starts in 2009

Reader submission, modeled in-house · 2026

A reader shared a popular strategy that holds 3x Nasdaq (TQQQ) while a stack of price and momentum rules on the Nasdaq-100 read as an uptrend, and long-duration Treasuries (ZROZ) otherwise. Its headline is a growth chart worth almost +4,000%. The catch is the start date: the backtest begins in late 2009, at the bottom of the financial crisis, so it has never once been tested through a real technology crash. Run the same rules back through the dot-com bust and the picture changes.

How the strategy works (as written)

  1. Compute a set of trend and momentum signals on the Nasdaq-100: price versus its 10, 200 and 250-day moving averages, the 20, 60 and 120-day returns, and the 14-day RSI.
  2. If any one of about twenty hand-picked combinations of those signals is satisfied, hold 100% TQQQ (3x Nasdaq).
  3. Otherwise hold 100% ZROZ (25+ year zero-coupon Treasuries).
  4. Re-check daily.

Our backtest, honest implementation

VariantCAGRMax DDSharpeVol
The rules, the marketed window
2010 to 2026 (daily)
30.2%-64.9%0.8640.5%
The same rules, extended through the dot-com crash
1999 to 2026 (daily)
26.4%-64.9%0.7643.6%
The disciplined 200-day-SMA core of the same idea
1999 to 2026 (daily)
18.3%-93.4%0.5853.8%

TQQQ before its 2010 inception is our calibrated synthetic 3x model, regressed against real TQQQ and UPRO to within a few tenths of a point of CAGR per year; ZROZ before its 2009 inception is proxied from long Treasuries. Reconstructing someone else's roughly twenty-clause rule set is necessarily approximate, and our version runs a few points hot on return versus the original chart, so if anything the drawdowns shown here are conservative.

Why we do not recommend it

The +4,000% is a start-date artifact

TQQQ launched in 2010 and this backtest starts with it, at the exact bottom of the 2008-09 crash. It rode the entire 2010-2021 Nasdaq bull and has never been shown through the 2000-2002 dot-com collapse, when a daily-reset 3x Nasdaq fund would have lost essentially everything. A leveraged backtest that begins after the last crash is not evidence that it survives the next one.

The worst drawdown is recent, not ancient

You do not need the dot-com era to see the risk. On the flattering 2010-start window, our reconstruction still draws down about 65%, and it happened in 2022: the trend rules were late leaving TQQQ off the January-2022 top, and the defensive leg made it worse rather than better.

The ZROZ safe leg is a second bet, not a hedge

When the Nasdaq breaks trend, the strategy parks in ZROZ, which is 25+ year zero-coupon Treasuries, one of the most rate-sensitive assets in existence. In 2022 that leg was falling at the same time as the Nasdaq, so the defensive position deepened the loss instead of cushioning it. A leg whose job is to protect you cannot itself crash in the same event; a T-bill would have done the job it was there to do.

The clean version of the idea still drops 93%

Strip the twenty tuned rules down to the honest core, hold 3x Nasdaq while the index is above its 200-day average and step aside otherwise, and run it through the dot-com bust: a 93% max drawdown. The stack of extra rules is not robustness, it is curve-fitting to the one history the chart happens to show. The underlying idea has a catastrophic tail either way.

Verdict

We did not adopt it. Trend-timing a 3x ETF is a real family, and we ship several honest versions of it, trend-gated, drawdown-labeled and scored for robustness against every variant we have ever tested. What we will not do is publish a 3x backtest that starts after the last crash and call the result a track record. When a leveraged strategy shows you a spectacular chart, the first question is always: what start date is it hiding behind?

Source

Shared by a BestFolio reader from a strategy-sharing site; the exact rule set re-modeled in-house on daily data back to 1999, with our calibrated synthetic-leverage costs.

Related

Think a strategy on BestFolio should be in this log?

We take candidate rejections seriously. If you can point to a structural flaw, a look-ahead bias, or a failure mode in one of our listed strategies, email [email protected]. If we agree with the analysis, we will add the strategy here with credit to whoever flagged it.