Skip to content
Important: BestFolio provides information for educational purposes only. Nothing on this site constitutes investment advice. Past performance does not guarantee future results. Read full disclaimer
·7 min read·BestFolio Research Team

9Sig Backtest: Kelly's TQQQ Rules Since 1999

This article also covers what was previously published at /blog/kelly-signal-danger-v2.

Note (24 September 2026): figures in this post are dated. Later data corrections moved some of them; each strategy page carries the current numbers.

Revised 2026-09-11. This article now carries the corrected 9Sig analysis in one place. The version we published on 2026-04-11 used a simplified engine that did not follow Jason Kelly's published rules. Readers on r/LETFs caught it, we rebuilt the engine, and on 2026-04-23 we published the rerun as a separate follow-up. That follow-up now redirects here, and every number below is the corrected one. The superseded April 11 figures are kept at the end for the record.

The short version

  • Jason Kelly's Signal system (3Sig, 6Sig, 9Sig) targets fixed quarterly growth in a stock sleeve (3%, 6% or 9%) and rebalances against a bond fund that acts as a reservoir: sell the overshoot into bonds, buy the shortfall out of bonds.
  • Run with Kelly's published rules and no new money, 9Sig (TQQQ + AGG, synthetic TQQQ before 2010, 1999 to 2026) returns 8.3% a year with a 99.7% maximum drawdown and a 0.44 Sharpe. Index-fund returns for near-total-loss risk.
  • On real TQQQ data only (2010 to 2026, close to an uninterrupted bull market) the same rules show 39.4% a year and a 72.1% drawdown. The start date does the marketing.
  • Sharpe doesn't improve with aggression: 0.61 for 3Sig, 0.47 for 6Sig, 0.44 for 9Sig on the extended windows.
  • A block bootstrap across 2,000 resampled histories puts the median 9Sig drawdown at 98%. The luckiest 5% of paths still lose 82% somewhere along the way.
  • Kelly's own subscribers survive the same crashes because the Letter assumes monthly contributions. A savings plan carries the strategy, and that is the distinction the marketing skips.
  • We don't rank the Kelly variants. Their strategy pages are public for study, and the family sits in our rejection log with the full metrics.

What Kelly's Signal actually is

Every few years someone discovers the Signal system and tells us we should list it. The pitch is easy to like: a mechanical rebalancing rule, a stock sleeve plus a bond sleeve, and backtests that look spectacular.

The rule itself is simple. You hold a stock ETF and a bond ETF. Each quarter you check whether the stock sleeve grew by the target (3%, 6% or 9% depending on the variant). Grew more than that, you sell the excess into bonds. Grew less, or dropped, you pull money from the bond sleeve and buy stocks. The bond sleeve is the reservoir that feeds the stock sleeve through drawdowns.

On paper this is value averaging with a leveraged fund in the stock sleeve, and it forces you to buy low and sell high. The question is what happens to the reservoir when "low" keeps going lower for 10 quarters.

What our first backtest got wrong

The April 11 version deviated from Kelly's rulebook in 5 ways. Commenters on the r/LETFs thread (thanks u/Gehrman_JoinsTheHunt, u/IllPoem4426 and u/Fee-Massive) walked us through each one, and we validated the rebuilt engine against the community simulator at 9sig.networthcast.com.

  • Starting allocation. We started 9Sig at 80/20. The canonical base allocation and post-reset target is 60% TQQQ / 40% AGG. Starting at 80/20 over-weighted the leveraged sleeve from day one and inflated the drawdown that followed.
  • The 30 Down lookback. Kelly's trigger is SPY closing at or below 70% of its highest close over the past 2 years, a rolling high. We used the all-time high, which almost never fires during a multi-year grind because the reference point stays stuck at the pre-crash peak.
  • The 30 Down effect. We had it dump the entire bond sleeve into stocks. The real rule is far milder: when triggered, the plan skips upcoming sell signals while buys happen normally. The book skipped the next 4 sells; Kelly cut that to 1 for 6Sig and 9Sig in a 2020 revision documented on his site.
  • No buying-power throttle. Kelly caps each quarterly buy at 90% of the current bond balance. Our old code had no cap, so the first serious crash drained the reservoir completely.
  • No spike reset. Kelly's published spike reset snaps 9Sig back to 60% stocks when TQQQ gains 100% or more in a single quarter. Our old engine had nothing like it.

The engine we run now

These are the rules in the corrected engine, which also powers the 3Sig, 6Sig and 9Sig strategy pages. Where we knowingly deviate from Kelly's published rules, the item says so.

  • 60% stock ETF / 40% bond ETF as the starting and base-reset allocation.
  • The signal line grows 3%, 6% or 9% per quarter, compounding, never adjusted down.
  • Quarterly rebalance: sell the surplus above the signal line, buy the shortfall below it.
  • Buys are capped at 90% of the current bond value (Kelly's throttle). We add a floor of our own: bonds never fall below 10% of portfolio value, so the sleeve can't be drained to zero in a closed system.
  • 30 Down: evaluated against an 8-quarter rolling high of the plan's own stock ETF (Kelly's published trigger watches SPY, so ours fires earlier in leveraged-fund crashes). Trigger at 70% of that high. Skip the next 2 sell signals (Kelly's current rule skips 1, the original book rule 4). The window exits after 2 skips, on recovery above the threshold, or after 8 quarters with a forced base reset.
  • Base reset when bonds exceed 30% of the portfolio after a sell quarter: snap back to 60/40 and reset the signal line. This valve is our addition; nothing like it appears in Kelly's published material.
  • Spike reset (9Sig only): if TQQQ returns 100% or more in a single quarter and the plan is outside a 30 Down window, snap to 60/40.

Taken together, these rules are what keeps the closed system from blowing up in a single quarter. The throttle and the floor stop the sleeve from emptying on one bad print, the 8-quarter exit stops the plan from being trapped, and the base reset stops the bond sleeve from becoming dead weight after a parabolic run. Drop any one of them and the math gets worse than what Kelly himself publishes.

What the corrected backtest shows

Each variant runs with its canonical Kelly pair: 3Sig on IJR (small caps, unleveraged) with BND, 6Sig on MVV (2x mid caps) with AGG, 9Sig on TQQQ (3x Nasdaq-100) with AGG. Our April 11 version had leaned on TQQQ for all 3, a shortcut we shouldn't have taken.

2 windows per variant. The first uses real ETF data only, starting when both legs have a live track record: IJR + BND from 2007-04, MVV + AGG from 2006-06, TQQQ + AGG from 2010-02. The second extends history with synthetic proxies for the leveraged legs (2x IJH for MVV before 2006, 3x QQQ for TQQQ before 2010) and a 4%-a-year proxy for AGG before 2003. The second window includes the dot-com crash, and it's the stress test. No contributions in either window.

Variant Assets Window CAGR Max drawdown Sharpe
3SigIJR + BND2007-04 to 2026-04 (real)9.0%-52.3%0.31
6SigMVV + AGG2006-06 to 2026-04 (real)11.7%-78.3%0.23
9SigTQQQ + AGG2010-02 to 2026-04 (real)39.4%-72.1%0.82
3SigIJR + BND2000-09 to 2026-04 (extended)9.17%-46.32%0.61
6SigMVV + AGG2000-09 to 2026-04 (extended, synthetic MVV before 2006)11.31%-82.75%0.47
9SigTQQQ + AGG1999-07 to 2026-04 (extended, synthetic TQQQ before 2010)8.28%-99.73%0.44
Required recovery gains: 3Sig 109.6%, 6Sig 360.8%, 9Sig 258.4%, calculated from the real-data drawdowns in the article.
The real-data drawdowns of 52.3%, 78.3% and 72.1% imply gains of 109.6%, 360.8% and 258.4% from their respective troughs to regain the preceding peak. This arithmetic does not estimate how long recovery takes, and the three history windows differ.

2 things stand out. Real-data 9Sig returns 39% a year in a closed system over 16 years, which is in the right ballpark for what Kelly subscribers report once you remember their simulated portfolios also take monthly contributions. That's the faithful rules doing their job across one very Nasdaq-friendly decade and a half.

Real-data 6Sig (MVV + AGG) draws down 78% despite being sold as the moderate variant. Mid-cap 2x leverage through 2008-2009 and 2020 is enough for that, even with the signal rebalancing and the bond sleeve. The moderate label belongs on 3Sig. Its 52% drawdown is about what an unlevered equity index did over the same period, which is what its risk actually looks like.

The extended window is where the thesis holds. Run 9Sig through the 2000-2002 Nasdaq crash with no contributions and the closed-system drawdown is 99.7%. Our April 11 version reported 99.3% on a similar window with the wrong rules, so it landed near the right number for the wrong reasons. The reason is arithmetic: 3x daily leverage on QQQ through the 2000 to 2002 collapse compounds to something very close to total loss, and a 40% bond sleeve is nowhere near large enough to refill a leg that has lost 99.9% of its value. 6Sig fares better through the same period (an 83% drawdown) because MVV is 2x. 3Sig comes through with a 46% drawdown because it isn't leveraged at all.

One note on the strategy pages. They run the same engine from 1985-05 with synthetic history before each fund's inception, a longer window than this table, so the CAGR figures there print higher: as of 2026-09-10, 9.7% a year for 3Sig, 15.1% for 6Sig and 19.7% for 9Sig, against maximum drawdowns of 46.3%, 82.7% and 99.7% that match the extended rows above to a tenth of a percent. The "real" rows come from our validation script, which restricts each variant to its earliest common real-ETF data. We report both so you can see how much of the headline result depends on synthetic pre-inception history.

How fragile is that 99% drawdown?

A fair objection to any single backtest is that history only happened once. The 99.7% landed right on top of the dot-com collapse, so maybe a different ordering of the same returns would have been survivable. We tested that with a moving-block bootstrap, the same method behind our in-app Monte Carlo. We resample contiguous 12-month blocks of the faithful 9Sig's own monthly returns (TQQQ + AGG, synthetic TQQQ before 2010), reassemble 2,000 alternate 27-year histories, and read off the dispersion. 12-month blocks preserve the serial correlation that makes leverage decay and drawdowns compound; plain return shuffling would destroy it.

Outcome Max drawdown Terminal multiple CAGR
Luckiest 5%-82%3,665x+35%
Median-98%7.7x+7.8%
Worst 5%-100%~0x-19%

These are per-metric percentiles across the 2,000 paths, not a single path. A favorable ordering pairs a milder drawdown with a higher ending value, an adverse one pairs a near-total drawdown with a near-zero ending value, which is why the columns line up the way they do.

9Sig (TQQQ and AGG) maximum drawdown across 2,000 block-bootstrapped 27-year histories; the realized backtest drawdown sits near the median of the distribution
Max drawdown across 2,000 resampled 27-year histories. The realized -99.6% is near the median, not the tail.

The realized drawdown sits at the median of the resampled paths (99.6% on the monthly series the bootstrap uses, 99.7% on daily data). A 98% drawdown is the typical outcome when you reshuffle this strategy's own history in year-long blocks, and even the luckiest 5% of reorderings still draw down 82%. No ordering of these returns is gentle. The wide terminal spread, a median of 7.7x against a 5th percentile of zero, is the signature of a closed-system leveraged sleeve: most paths compound, a meaningful fraction go to nearly nothing, and which one you get is mostly path luck.

Conditioning on the dot-com window makes it starker. Resampling blocks drawn only from 2000 to 2002 into a 24-month episode, the median drawdown is 93% and the median ending value is about 9 cents on the dollar. One caveat: a within-regime resample of a 27-month window is a small, autocorrelated sample, so treat the dot-com-only bands as illustrative of the regime rather than a forecast.

Why the reservoir runs dry

In a short, sharp correction, March 2020 or December 2018, the rule pulls money from bonds to buy the dip, stocks rebound, the gains refill the sleeve, and the round trip completes. This is the pattern that dominates any backtest that starts in 2010.

In a long, grinding bear market, 2000 to 2002 or 1973 to 1974, the rule keeps pulling from bonds quarter after quarter as stocks drop another leg. The throttle and the floor slow the bleed, and they are what keeps the corrected drawdown at 99.7% instead of a flat zero, but they can't create a refill. Refills only come from sell signals, and deep in a bear market there are no surpluses to sell. The 30 Down rule then deliberately skips the first sell signals of the recovery, so the sleeve stays thin through exactly the stretch where the plan needs it most.

Kelly's answer is new cash. The community simulator at 9sig.networthcast.com starts a $10,000 portfolio in Q1 1999 and adds $500 a month, all of it to the bond sleeve. Every month that refills the reservoir and lifts the signal line by half the contribution, exactly as the Letter describes. Over 27 years the contributor adds $162,000 of real money. The strategy survives the dot-com crash because its operator keeps bailing the boat.

That's the difference between a strategy backtest and a savings plan. A backtest measures what the rules produce in isolation. A plan measures what a disciplined saver produces across decades of contributions, withdrawals, tax events and emergencies. 9Sig passes as a plan for someone who can guarantee monthly contributions and never needs to touch the portfolio during a 99% drawdown. It fails as a strategy for anyone else.

Why the headline numbers look amazing anyway

The real-data 39% CAGR is the arithmetic of a window that starts in 2010, includes the 2009 to 2020 bull market and the 2020 to 2021 rally, and never meets a 3-year Nasdaq bear. Compounding a 3x fund through that decade produces eye-watering numbers. Extend the window to 1999 and the same rules produce 8.3% a year, because the portfolio spends most of the 2000s digging out of a 99.7% hole.

A 99% drawdown is one margin call, one job loss or one medical bill away from being permanent. Anyone who needed cash in 2002 or 2009 sold at the bottom and turned the paper drawdown into a realized loss. The backtest quietly assumes you never needed a single euro during those years. It's a thought experiment dressed up as a retirement plan.

What we list instead

The strategies on our leaderboard, GEM, HAA, VAA, KDA, AlphaOne, Dynamic Macro Allocation, share a property the Signal system lacks: they have an exit. When the trend turns they move to cash or bonds and wait. They don't pull from a reservoir that can empty, and they don't double down into oblivion. Their full-history drawdowns are on their strategy pages, and none of them is anywhere near 99%.

Those strategies have lower CAGR than 9Sig in a window that starts in 2010. They also keep you solvent in 1974, 2002, 2008 and whichever bear market lasts longer than a few quarters. That trade-off is the point of tactical allocation.

What we're keeping and what changed

The core decision stands: 3Sig, 6Sig and 9Sig stay unlisted. Their strategy pages are public so anyone can audit the numbers, but they appear in no leaderboard or ranking, and we won't show a Pro subscriber a ranked entry whose downside is credibly total loss. The corrected returns don't change that risk profile.

The headline claim of the April 11 version is restated. "A 99.3% drawdown in a 9Sig backtest" came from an implementation that didn't follow the rules and happened to land near the right answer. The faithful-rule number over the full window is 99.7%; over the real-TQQQ window it's 72.1%. Both are catastrophic in a closed system. Neither is what Kelly's subscribers experience, because they add new money every month.

The rule set above is the complete parameter list we run, and the rejection log entry carries the same rules with the current engine numbers. If you think we're wrong about Kelly, email us and we'll reopen the case with new evidence. What we won't do is put a 99% drawdown into a subscriber's portfolio and call it research.

What we learned

Getting the rules wrong was the lesson. We wrote a post arguing that a popular strategy would ruin retail investors, and we based it on numbers that didn't faithfully reproduce the strategy. The readers who pointed this out were right, and the fix has been in our tracker since April 19.

The broader point about closed-system leveraged value averaging survives the correction and gets sharper. Without the monthly contributions the Letter assumes, 9Sig drew down 99.7% through the dot-com crash. With real TQQQ data after 2010 and no contributions, it drew down 72% through the 2022 Nasdaq selloff. Either one ends a retail portfolio that can't sit underwater for years. The strategy works for someone with a 30-year contribution horizon and iron discipline. For everyone else, the quiet assumption that the bond sleeve refills itself is still the flaw, and the faithful rules don't repair it. They just make the failure mode more interesting to read about.

If a backtest looks too good to be true, ask what assumption is being held constant that would break in reality. For Kelly it's the infinite reservoir. For other popular systems it's survivor bias in the asset universe, look-ahead in the signal, unrealistic costs, or the quiet exclusion of 2000 to 2002 from the window. Our methodology page lists the 4 gates every strategy clears before it reaches the leaderboard.

For the record: the superseded April 11 numbers

The first version of this article reported these figures from the simplified engine (TQQQ for all 3 variants, an 80/20 start, an all-time-high 30 Down trigger that dumped the bond sleeve, no throttle, no spike reset). They stay here so the correction can be audited. Don't quote them.

Variant (April 11 engine) CAGR Max drawdown Sharpe
3Sig10.4%-52.9%0.64
6Sig16.3%-85.5%0.59
9Sig21.1%-99.3%0.62

Thanks again to the r/LETFs readers who flagged the specific errors.

Revision history

  1. Consolidated with the April 23 follow-up

    The article now carries the corrected engine, the corrected results and the bootstrap stress test in one place; the April 11 numbers are kept at the end as superseded, and the follow-up URL redirects here.

  2. Graphs and research review — September 15, 2026

    Add a figure with accessible labels and the study’s original data boundaries; preserve existing prose.

  3. Readable data tables on mobile

    Added horizontal scrolling for wide data tables on small screens. All figures, source notes and article prose are unchanged.

Past performance does not guarantee future results. Backtested results are hypothetical and do not represent actual trading.

Written with the help of AI tools and reviewed before publication.

Frequently Asked Questions

What is Jason Kelly's 9Sig strategy?

9Sig is the aggressive variant of Jason Kelly's Signal system. You hold a leveraged Nasdaq fund (TQQQ) plus a bond fund and target 9% stock-sleeve growth per quarter: sell the excess into bonds when you overshoot, pull money out of bonds to buy more stock when you undershoot. The bond sleeve acts as a reservoir that feeds the stock sleeve during drawdowns. The 3Sig and 6Sig variants apply the same rule at 3% and 6% quarterly targets.

What did the observed TQQQ study show?

From TQQQ's February 2010 launch through April 2026, the faithful 9Sig rules produced 39.4% CAGR with a -72.1% maximum drawdown in this fixed historical study. That observed window starts after the dot-com crash and covers one unusually strong Nasdaq regime, so it does not show how the rules would have behaved before the fund existed.

How was pre-2010 TQQQ history modeled?

TQQQ did not exist before 2010. The extended July 1999 to April 2026 study therefore uses a synthetic series based on three times QQQ's daily return minus modeled borrowing costs, while AGG before September 2003 is represented by a 4% annual return proxy. Those pre-inception results are simulations, not observed fund performance.

What changed in the corrected 9Sig engine?

The corrected engine starts 9Sig at 60% TQQQ and 40% AGG, caps a quarterly buy at 90% of the bond sleeve, keeps a 10% bond floor, evaluates Thirty-Down against the prior eight quarter-end closes, skips two sell signals while that window is active, and implements the documented base and spike resets.

What does the 9Sig block-bootstrap test measure?

The study resamples 12-month blocks from the strategy's own historical monthly returns to construct 2,000 alternate 27-year sequences. It preserves some return clustering and shows how sequence order affected the tested payoff profile, but it reuses one historical sample and is neither a probability forecast nor a model of future market regimes.

Do the 9Sig study results include investor contributions?

No. Both the observed and extended studies are closed-system backtests with no deposits or withdrawals. Kelly's community reference simulation includes regular cash contributions to the bond sleeve, which changes the cash-flow plan and is not directly comparable with the closed-system figures reported here.

Share this article

Data and method

Study dates and assumptions are documented in the article and its revisions. Our current methodology explains the platform's data sources, proxy histories, trade timing and inflation treatment.

  • Seven Leveraged Tactical Strategies: How They Work, What Breaks Them

    Seven new strategies in the catalog today, six of them leveraged. They span the named names in the r/LETFs community (RNAProf, u/Wongkok, u/Low-Initiative-1327), a published author (David Alan Carter), and two r/LETFs-classic ideas (200d-SMA TQQQ trend and RSI/SMA dip-buying) that BestFolio has stabilized into shipping form. The point of this writeup is not a sales pitch. It is the rule set, what each one does well, and what breaks it. Kelly 3sig/6sig/9sig deliberately held back: the 9sig 99.73% dot-com drawdown makes that family unsuitable for a wider release.

  • A wider trend band can create a deeper drawdown

    A volatility-scaled buffer gave a 200-day TQQQ gate room in calm markets and tightened the exit when turbulence was already high.

  • DCA Into a Leveraged ETF: The Return Path Matters More Than the Average

    Regular contributions helped a synthetic 2x S&P 500 position across long historical windows, but they never removed sequence risk. In 5,000 block-resampled 15-year paths, the leveraged route still trailed 24.74% of the time.

Explore with tools

Try these strategies on BestFolio

Browse 77 tactical allocation strategies with monthly signals, walk-forward validation, and portfolio blending. Free to start.

Create free account

BestFolio Monthly Briefing

Liked this post? Get a free monthly recap of TAA strategy signals, performance rankings, and market regime updates. No spam, unsubscribe anytime.