Skip to content
Important: BestFolio provides information for educational purposes only. Nothing on this site constitutes investment advice. Past performance does not guarantee future results. Read full disclaimer
·9 min read·BestFolio Research Team

9Sig Backtest: TQQQ and Extended History

Last week we published a piece arguing that Jason Kelly’s 9Sig system could expose a closed portfolio to near-total-loss risk, citing a 99.3% max drawdown in our backtest. The post did what posts like that tend to do on r/LETFs and Reddit’s Kelly-adjacent communities: it generated a lot of replies, some agreeing, many explaining very specifically where our simulation had gotten the rules wrong. We owe those readers a direct answer, so here it is.

Our original 9Sig implementation was not a faithful rendition of the Kelly Letter rulebook. It was close enough to be misleading and wrong enough to overstate the drawdown. This post explains what the faithful rules actually are, what our corrected backtest produces, and why our conclusion about the closed-system risk still stands, just with more honest numbers.

What we had wrong

The version we shipped on April 11 used five assumptions that do not match Kelly’s published rules. In the spirit of showing our receipts rather than quietly fixing the code:

  • Starting allocation 80/20 instead of 60/40. The canonical 9Sig base allocation and post-reset target is 60% TQQQ / 40% AGG, per both Kelly’s own Substack posts and the community reference simulator at 9sig.networthcast.com. Starting at 80/20 over-weighted the leveraged sleeve from day one and inflated the drawdown that followed.
  • Thirty-Down lookback was an all-time high. Kelly defines the trigger as SPY closing at or below 70% of its highest close over the past two years, a rolling high, not an all-time high. Using an all-time high means the rule almost never triggers during a multi-year grind, because the reference point gets stuck at the pre-crash peak. The rule was designed to respond to crashes relative to recent conditions, not to ancient history. One engine choice to flag: we evaluate the 70%-of-two-year-high test on each plan's own stock ETF rather than on SPY, which makes the trigger fire earlier in leveraged-fund crashes.
  • Thirty-Down effect was “dump all bonds.” The actual rule is far less aggressive: when triggered, the plan simply skips upcoming sell signals while buys happen normally. The original book rule skipped the next four sells; Kelly later cut that to one for 6Sig and 9Sig, in a 2020 revision documented on his site. Our engine skips the next two, between the two published values. Either way, the plan does not throw the reservoir at the bottom.
  • No buying-power throttle or bond floor. Kelly’s canonical quarterly buy is capped at 90% of the current bond balance. We also add a hard 10% bond floor relative to total portfolio value, so the sleeve cannot be drained to zero in a closed system. Our previous code did neither, which meant the first serious crash burned the reservoir completely.
  • No spike reset, no containment valve. Kelly's published spike reset forces 9Sig back to a 60% stock allocation when TQQQ gains 100% or more in a single quarter. Our corrected engine also adds a second valve of our own design: when bonds grow past ~30% of the portfolio after a sell quarter, snap back to 60/40 and reset the signal line. That base reset is our containment rule, not one of Kelly's; nothing like it appears in his published material. Neither valve was in our previous engine.

The corrected engine

Our rewritten engine implements the rules as follows. Where we knowingly deviate from Kelly's published rules (trigger index, skip count, base reset), the item says so:

  • 60% TQQQ / 40% AGG starting and base-reset allocation.
  • Signal line grows 9% per quarter, compounding, never adjusted down.
  • Quarterly rebalance: sell surplus above signal, buy shortfall below.
  • Buys capped at 90% of current bond value (Kelly throttle) and further capped so bonds never fall below 10% of portfolio value.
  • Thirty-Down: evaluated against an 8-quarter rolling high of the plan's own stock ETF (Kelly's published trigger watches SPY); trigger at 70% of that high; skips the next 2 sell signals (Kelly's current rule for 6Sig and 9Sig skips 1, the original book rule 4); window exits after 2 skips, on recovery above the threshold, or after 8 quarters (forced base reset).
  • Base reset when bonds exceed 30% after a sell: snap to 60/40, reset the signal line. This valve is our addition, not a published Kelly rule.
  • Spike reset (9Sig only): if TQQQ returns 100% or more in a single quarter and the plan is not in a Thirty-Down window, snap to 60/40.

That last point is the subtle one. The rules taken together are not arbitrary. They are what keeps the closed system from blowing up: the throttle and floor prevent bond depletion in a single bad quarter, the 8-quarter safety exit prevents the plan from being trapped indefinitely, and the base reset prevents the sleeve from becoming a dead-weight cash position after a parabolic run. Skip any one of them and the math gets worse than what Kelly himself publishes.

What the corrected backtest actually shows

Each variant is tested with its canonical Kelly asset pair: 3Sig runs IJR (small-cap 1x) with BND, 6Sig runs MVV (mid-cap 2x) with AGG, and 9Sig runs TQQQ (Nasdaq-100 3x) with AGG. Our earlier post leaned on TQQQ for all three; that was a simplification we should not have taken, and it is fixed in the corrected engine.

We re-ran each variant on two windows. The first uses real ETF data only, starting when both the stock leg and the bond leg have a live track record (IJR + BND from 2007-04, MVV + AGG from 2006-06, TQQQ + AGG from 2010-02). The second extends history with synthetic proxies for the leveraged legs (2x IJH for MVV pre-2006, 3x QQQ for TQQQ pre-2010) and a 4%-per-year proxy for AGG pre-2003. The second window includes the dot-com crash and is the stress test.

Variant Assets Window CAGR Max DD Sharpe
3Sig v2 IJR + BND 2007-04 to 2026-04 (real) 9.0% -52.3% 0.31
6Sig v2 MVV + AGG 2006-06 to 2026-04 (real) 11.7% -78.3% 0.23
9Sig v2 TQQQ + AGG 2010-02 to 2026-04 (real) 39.4% -72.1% 0.82
3Sig v2 IJR + BND 2000-09 to 2026-04 (extended) 9.17% -46.32% 0.61
6Sig v2 MVV + AGG 2000-09 to 2026-04 (extended, synth MVV pre-2006) 11.31% -82.75% 0.47
9Sig v2 TQQQ + AGG 1999-07 to 2026-04 (extended, synth TQQQ pre-2010) 8.28% -99.73% 0.44

The "extended" rows above use the same engine as our strategy pages at /strategies/kelly-3sig, /strategies/kelly-6sig, and /strategies/kelly-9sig. Those pages have since been re-run with synthetic history extended back to 1985, a longer window than this table, so the CAGR figures there print higher while the max drawdowns land on the same values. The "real" rows are a companion analysis from our validation script that restricts each variant to its earliest common real-ETF data (no synthetic leverage extrapolation). We report both so that skeptical readers can see how much of the headline result depends on the synthetic pre-inception history.

Two things stand out. First, the real-data 9Sig returns 39% CAGR in a closed system over 16 years, which is in the right ballpark for what Kelly subscribers report once you account for the fact that their simulated portfolios also include monthly contributions. That is the faithful rules doing their job over a single, very Nasdaq-friendly decade and a half.

Second, the real-data 6Sig (MVV + AGG) shows a -78% drawdown despite being marketed as the moderate variant. Mid-cap 2x leverage through 2008-2009 and 2020 is enough to produce that, even with the signal rebalancing and the bond sleeve. The moderate label belongs on 3Sig (IJR + BND), not 6Sig. 3Sig’s -52% drawdown is comparable to holding an unlevered equity index through the same period, which is the honest way to describe its risk.

The extended-history column is where the thesis holds up. Run 9Sig through the 2000-2002 Nasdaq crash with no contributions and the closed-system drawdown is 99.7%. Our earlier post reported -99.3% on the same window; the corrected engine is in the same catastrophic neighborhood, just with the rules faithfully applied. The earlier figure was arrived at through a wrong path (no 8-quarter lookback, no bond floor, no skip-2-sell-signals) but landed close to the right number for the wrong reason. The reason is physics: 3x daily leverage on QQQ through a 78% peak-to-trough decline compounds to something very close to total capital loss, and the bond sleeve is not large enough to refill a leg that has lost 99.9% of its value. 6Sig fares better through the same period (-83%) because MVV is 2x instead of 3x. 3Sig is effectively unharmed (-53%) because it is not leveraged at all.

How fragile is that 99% drawdown? A block-bootstrap stress test

A fair objection to any single backtest is that history only happened once. The 99.7% drawdown landed right on top of the dot-com collapse, so maybe it was a one-in-a-lifetime ordering and a different sequence of the same returns would have been survivable. We tested that directly with a moving-block bootstrap, the same method that powers our in-app Monte Carlo. We resample contiguous 12-month blocks of the faithful v2 9Sig's own monthly returns (TQQQ + AGG, synthetic TQQQ before 2010), reassemble 2,000 alternate 27-year histories, and read off the dispersion. Twelve-month blocks preserve the serial correlation that makes leverage decay and drawdowns compound, which plain return-shuffling would destroy.

The result is not reassuring. Across 2,000 resampled histories of the strategy's own returns:

OutcomeMax drawdownTerminal multipleCAGR
Luckiest 5%-82%3,665x+35%
Median-98%7.7x+7.8%
Worst 5%-100%~0x-19%
9Sig (TQQQ+AGG) max drawdown across 2,000 block-bootstrapped histories. The realized backtest drawdown of -99.6% sits at the worst end of the distribution, near the median.
Max drawdown across 2,000 resampled 27-year histories. The realized -99.6% is near the median, not the tail.

These are per-metric percentiles across the 2,000 paths, not a single path. A favorable ordering pairs a milder drawdown with a higher ending value, an adverse one pairs a near-total drawdown with a near-zero ending value, which is why the columns line up the way they do.

The realized -99.7% is not a tail event. It is essentially the median. A 98% drawdown is the typical outcome when you reshuffle this strategy's own history in year-long blocks, and even the luckiest 5% of reorderings still draw down 82%. There is no ordering of these returns that is gentle. The wide terminal spread (a median of 7.7x but a 5th percentile of zero) is the signature of a closed-system leveraged sleeve: most paths compound, a meaningful fraction go to nearly nothing, and which one you get is mostly path luck.

Conditioning on the dot-com window makes it starker. Resampling blocks drawn only from 2000 to 2002 into a 24-month episode, the median drawdown is 93% and the median ending value is about 9 cents on the dollar. One caveat we want to be explicit about: a within-regime resample of a 27-month window is a small, autocorrelated sample, so treat the dot-com-only bands as illustrative of the regime rather than a precise forecast.

This is the same point our longest-backtest work keeps making from the other direction. A strategy with clean data only since 2010 has never been shown a 2000 or a 1973. When you synthesize the leverage back through the exact bust it was designed without, the closed-system risk stops looking like a fat tail and starts looking like the center of the distribution.

Why the networthcast numbers look better

The community simulator at 9sig.networthcast.com starts a $10,000 portfolio in Q1 1999 and models a monthly $500 contribution, 100% to the bond sleeve. That changes the mathematics completely. Every month, new cash refills the reservoir and lifts the signal line by half of the contribution, exactly as Kelly describes in his Letter. Over 27 years, the contributor adds $162,000 of real money to the system. The strategy survives the dot-com crash not because its rules are self-healing but because its operator is constantly bailing the boat.

That is the real distinction between a strategy backtest and a financial plan. A backtest measures what the rules produce in isolation. A plan measures what a disciplined investor produces across decades of contributions, withdrawals, tax events, and emergencies. 9Sig passes as a plan for someone who can guarantee monthly contributions and never needs to touch the portfolio during a 99% drawdown. It fails as a strategy for anyone else, because a closed-system ruin of 99.7% is a real possibility in a real bear market.

What we are keeping and what we are changing

We are keeping the core decision: 9Sig, 6Sig, and 3Sig stay unlisted on BestFolio. Their strategy pages are public so anyone can audit the numbers, but they appear in no leaderboard or ranking. We are not going to show a Pro subscriber a ranked leaderboard entry whose downside can credibly be total loss. The corrected CAGR numbers do not change the risk profile enough to warrant promotion.

We are restating the headline claim in our previous post. The “-99.3% drawdown in a 9Sig backtest” statement was based on a flawed implementation that happened to land near the right answer. The faithful-rule number, over the full window, is -99.7%. Over the real-TQQQ-only window, it is -72.1%. Both numbers are catastrophic in a closed system. Neither number is what Kelly’s subscribers experience because they are adding new money every month.

We are also publishing the corrected engine’s source code and the rule parameters in our rejection log, so anyone who wants to reproduce the backtest with different assumptions (different lookback, different floor, different contribution schedule) can do so against the same code we used.

What we learned

Getting the rules wrong was the lesson. We wrote a post arguing that a popular strategy could expose a closed portfolio to near-total-loss risk, and we based it on numbers that did not faithfully reproduce the strategy. The critics who pointed this out on Reddit and in email were correct, and the fix has been in our tracker since April 19.

The broader point about closed-system leveraged value averaging does not go away with the corrected engine. It gets sharper. 9Sig, run without the monthly contributions that Kelly’s Letter assumes, drew down 99.7% through the dot-com crash in our corrected simulation. 9Sig, run with real TQQQ data after 2010 and no contributions, drew down 72% through the 2022 Nasdaq selloff. Either of those is enough to end a retail portfolio that cannot tolerate a multi-year underwater period. The strategy works for someone with a 30-year contribution horizon and iron discipline. For everyone else, the quiet assumption that the bond sleeve refills itself is still the flaw, and the faithful rules do not repair it. They just make the failure mode more interesting to read about.

Thank you to the r/LETFs readers who flagged the specific errors. The follow-up comment on the original thread will point here.

Frequently Asked Questions

What is Jason Kelly's 9Sig strategy?

9Sig is the aggressive variant of Jason Kelly's Signal system. You hold a leveraged Nasdaq fund (TQQQ) plus a bond fund and target 9% stock-sleeve growth per quarter: sell the excess into bonds when you overshoot, pull money out of bonds to buy more stock when you undershoot. The bond sleeve acts as a reservoir that feeds the stock sleeve during drawdowns. The 3Sig and 6Sig variants apply the same rule at 3% and 6% quarterly targets.

What did the observed TQQQ study show?

From TQQQ's February 2010 launch through April 2026, the faithful 9Sig rules produced 39.4% CAGR with a -72.1% maximum drawdown in this fixed historical study. That observed window starts after the dot-com crash and covers one unusually strong Nasdaq regime, so it does not show how the rules would have behaved before the fund existed.

How was pre-2010 TQQQ history modeled?

TQQQ did not exist before 2010. The extended July 1999 to April 2026 study therefore uses a synthetic series based on three times QQQ's daily return minus modeled borrowing costs, while AGG before September 2003 is represented by a 4% annual return proxy. Those pre-inception results are simulations, not observed fund performance.

What changed in the corrected 9Sig engine?

The corrected engine starts 9Sig at 60% TQQQ and 40% AGG, caps a quarterly buy at 90% of the bond sleeve, keeps a 10% bond floor, evaluates Thirty-Down against the prior eight quarter-end closes, skips two sell signals while that window is active, and implements the documented base and spike resets.

What does the 9Sig block-bootstrap test measure?

The study resamples 12-month blocks from the strategy's own historical monthly returns to construct 2,000 alternate 27-year sequences. It preserves some return clustering and shows how sequence order affected the tested payoff profile, but it reuses one historical sample and is neither a probability forecast nor a model of future market regimes.

Do the 9Sig study results include investor contributions?

No. Both the observed and extended studies are closed-system backtests with no deposits or withdrawals. Kelly's community reference simulation includes regular cash contributions to the bond sleeve, which changes the cash-flow plan and is not directly comparable with the closed-system figures reported here.

Share this article

Try these strategies on BestFolio

Browse 62 tactical allocation strategies with monthly signals, walk-forward validation, and portfolio blending. Free to start.

Create free account

BestFolio Monthly Briefing

Liked this post? Get a free monthly recap of TAA strategy signals, performance rankings, and market regime updates. No spam, unsubscribe anytime.