Note (24 September 2026): figures in this post are dated. Later data corrections moved some of them; each strategy page carries the current numbers.
Revised 2026-09-11. This article now carries the corrected 9Sig analysis in one place. The version we published on 2026-04-11 used a simplified engine that did not follow Jason Kelly's published rules. Readers on r/LETFs caught it, we rebuilt the engine, and on 2026-04-23 we published the rerun as a separate follow-up. That follow-up now redirects here, and every number below is the corrected one. The superseded April 11 figures are kept at the end for the record.
The short version
- Jason Kelly's Signal system (3Sig, 6Sig, 9Sig) targets fixed quarterly growth in a stock sleeve (3%, 6% or 9%) and rebalances against a bond fund that acts as a reservoir: sell the overshoot into bonds, buy the shortfall out of bonds.
- Run with Kelly's published rules and no new money, 9Sig (TQQQ + AGG, synthetic TQQQ before 2010, 1999 to 2026) returns 8.3% a year with a 99.7% maximum drawdown and a 0.44 Sharpe. Index-fund returns for near-total-loss risk.
- On real TQQQ data only (2010 to 2026, close to an uninterrupted bull market) the same rules show 39.4% a year and a 72.1% drawdown. The start date does the marketing.
- Sharpe doesn't improve with aggression: 0.61 for 3Sig, 0.47 for 6Sig, 0.44 for 9Sig on the extended windows.
- A block bootstrap across 2,000 resampled histories puts the median 9Sig drawdown at 98%. The luckiest 5% of paths still lose 82% somewhere along the way.
- Kelly's own subscribers survive the same crashes because the Letter assumes monthly contributions. A savings plan carries the strategy, and that is the distinction the marketing skips.
- We don't rank the Kelly variants. Their strategy pages are public for study, and the family sits in our rejection log with the full metrics.
What Kelly's Signal actually is
Every few years someone discovers the Signal system and tells us we should list it. The pitch is easy to like: a mechanical rebalancing rule, a stock sleeve plus a bond sleeve, and backtests that look spectacular.
The rule itself is simple. You hold a stock ETF and a bond ETF. Each quarter you check whether the stock sleeve grew by the target (3%, 6% or 9% depending on the variant). Grew more than that, you sell the excess into bonds. Grew less, or dropped, you pull money from the bond sleeve and buy stocks. The bond sleeve is the reservoir that feeds the stock sleeve through drawdowns.
On paper this is value averaging with a leveraged fund in the stock sleeve, and it forces you to buy low and sell high. The question is what happens to the reservoir when "low" keeps going lower for 10 quarters.
What our first backtest got wrong
The April 11 version deviated from Kelly's rulebook in 5 ways. Commenters on the r/LETFs thread (thanks u/Gehrman_JoinsTheHunt, u/IllPoem4426 and u/Fee-Massive) walked us through each one, and we validated the rebuilt engine against the community simulator at 9sig.networthcast.com.
- Starting allocation. We started 9Sig at 80/20. The canonical base allocation and post-reset target is 60% TQQQ / 40% AGG. Starting at 80/20 over-weighted the leveraged sleeve from day one and inflated the drawdown that followed.
- The 30 Down lookback. Kelly's trigger is SPY closing at or below 70% of its highest close over the past 2 years, a rolling high. We used the all-time high, which almost never fires during a multi-year grind because the reference point stays stuck at the pre-crash peak.
- The 30 Down effect. We had it dump the entire bond sleeve into stocks. The real rule is far milder: when triggered, the plan skips upcoming sell signals while buys happen normally. The book skipped the next 4 sells; Kelly cut that to 1 for 6Sig and 9Sig in a 2020 revision documented on his site.
- No buying-power throttle. Kelly caps each quarterly buy at 90% of the current bond balance. Our old code had no cap, so the first serious crash drained the reservoir completely.
- No spike reset. Kelly's published spike reset snaps 9Sig back to 60% stocks when TQQQ gains 100% or more in a single quarter. Our old engine had nothing like it.
The engine we run now
These are the rules in the corrected engine, which also powers the 3Sig, 6Sig and 9Sig strategy pages. Where we knowingly deviate from Kelly's published rules, the item says so.
- 60% stock ETF / 40% bond ETF as the starting and base-reset allocation.
- The signal line grows 3%, 6% or 9% per quarter, compounding, never adjusted down.
- Quarterly rebalance: sell the surplus above the signal line, buy the shortfall below it.
- Buys are capped at 90% of the current bond value (Kelly's throttle). We add a floor of our own: bonds never fall below 10% of portfolio value, so the sleeve can't be drained to zero in a closed system.
- 30 Down: evaluated against an 8-quarter rolling high of the plan's own stock ETF (Kelly's published trigger watches SPY, so ours fires earlier in leveraged-fund crashes). Trigger at 70% of that high. Skip the next 2 sell signals (Kelly's current rule skips 1, the original book rule 4). The window exits after 2 skips, on recovery above the threshold, or after 8 quarters with a forced base reset.
- Base reset when bonds exceed 30% of the portfolio after a sell quarter: snap back to 60/40 and reset the signal line. This valve is our addition; nothing like it appears in Kelly's published material.
- Spike reset (9Sig only): if TQQQ returns 100% or more in a single quarter and the plan is outside a 30 Down window, snap to 60/40.
Taken together, these rules are what keeps the closed system from blowing up in a single quarter. The throttle and the floor stop the sleeve from emptying on one bad print, the 8-quarter exit stops the plan from being trapped, and the base reset stops the bond sleeve from becoming dead weight after a parabolic run. Drop any one of them and the math gets worse than what Kelly himself publishes.
What the corrected backtest shows
Each variant runs with its canonical Kelly pair: 3Sig on IJR (small caps, unleveraged) with BND, 6Sig on MVV (2x mid caps) with AGG, 9Sig on TQQQ (3x Nasdaq-100) with AGG. Our April 11 version had leaned on TQQQ for all 3, a shortcut we shouldn't have taken.
2 windows per variant. The first uses real ETF data only, starting when both legs have a live track record: IJR + BND from 2007-04, MVV + AGG from 2006-06, TQQQ + AGG from 2010-02. The second extends history with synthetic proxies for the leveraged legs (2x IJH for MVV before 2006, 3x QQQ for TQQQ before 2010) and a 4%-a-year proxy for AGG before 2003. The second window includes the dot-com crash, and it's the stress test. No contributions in either window.
| Variant | Assets | Window | CAGR | Max drawdown | Sharpe |
|---|---|---|---|---|---|
| 3Sig | IJR + BND | 2007-04 to 2026-04 (real) | 9.0% | -52.3% | 0.31 |
| 6Sig | MVV + AGG | 2006-06 to 2026-04 (real) | 11.7% | -78.3% | 0.23 |
| 9Sig | TQQQ + AGG | 2010-02 to 2026-04 (real) | 39.4% | -72.1% | 0.82 |
| 3Sig | IJR + BND | 2000-09 to 2026-04 (extended) | 9.17% | -46.32% | 0.61 |
| 6Sig | MVV + AGG | 2000-09 to 2026-04 (extended, synthetic MVV before 2006) | 11.31% | -82.75% | 0.47 |
| 9Sig | TQQQ + AGG | 1999-07 to 2026-04 (extended, synthetic TQQQ before 2010) | 8.28% | -99.73% | 0.44 |

2 things stand out. Real-data 9Sig returns 39% a year in a closed system over 16 years, which is in the right ballpark for what Kelly subscribers report once you remember their simulated portfolios also take monthly contributions. That's the faithful rules doing their job across one very Nasdaq-friendly decade and a half.
Real-data 6Sig (MVV + AGG) draws down 78% despite being sold as the moderate variant. Mid-cap 2x leverage through 2008-2009 and 2020 is enough for that, even with the signal rebalancing and the bond sleeve. The moderate label belongs on 3Sig. Its 52% drawdown is about what an unlevered equity index did over the same period, which is what its risk actually looks like.
The extended window is where the thesis holds. Run 9Sig through the 2000-2002 Nasdaq crash with no contributions and the closed-system drawdown is 99.7%. Our April 11 version reported 99.3% on a similar window with the wrong rules, so it landed near the right number for the wrong reasons. The reason is arithmetic: 3x daily leverage on QQQ through the 2000 to 2002 collapse compounds to something very close to total loss, and a 40% bond sleeve is nowhere near large enough to refill a leg that has lost 99.9% of its value. 6Sig fares better through the same period (an 83% drawdown) because MVV is 2x. 3Sig comes through with a 46% drawdown because it isn't leveraged at all.
One note on the strategy pages. They run the same engine from 1985-05 with synthetic history before each fund's inception, a longer window than this table, so the CAGR figures there print higher: as of 2026-09-10, 9.7% a year for 3Sig, 15.1% for 6Sig and 19.7% for 9Sig, against maximum drawdowns of 46.3%, 82.7% and 99.7% that match the extended rows above to a tenth of a percent. The "real" rows come from our validation script, which restricts each variant to its earliest common real-ETF data. We report both so you can see how much of the headline result depends on synthetic pre-inception history.
How fragile is that 99% drawdown?
A fair objection to any single backtest is that history only happened once. The 99.7% landed right on top of the dot-com collapse, so maybe a different ordering of the same returns would have been survivable. We tested that with a moving-block bootstrap, the same method behind our in-app Monte Carlo. We resample contiguous 12-month blocks of the faithful 9Sig's own monthly returns (TQQQ + AGG, synthetic TQQQ before 2010), reassemble 2,000 alternate 27-year histories, and read off the dispersion. 12-month blocks preserve the serial correlation that makes leverage decay and drawdowns compound; plain return shuffling would destroy it.
| Outcome | Max drawdown | Terminal multiple | CAGR |
|---|---|---|---|
| Luckiest 5% | -82% | 3,665x | +35% |
| Median | -98% | 7.7x | +7.8% |
| Worst 5% | -100% | ~0x | -19% |
These are per-metric percentiles across the 2,000 paths, not a single path. A favorable ordering pairs a milder drawdown with a higher ending value, an adverse one pairs a near-total drawdown with a near-zero ending value, which is why the columns line up the way they do.
The realized drawdown sits at the median of the resampled paths (99.6% on the monthly series the bootstrap uses, 99.7% on daily data). A 98% drawdown is the typical outcome when you reshuffle this strategy's own history in year-long blocks, and even the luckiest 5% of reorderings still draw down 82%. No ordering of these returns is gentle. The wide terminal spread, a median of 7.7x against a 5th percentile of zero, is the signature of a closed-system leveraged sleeve: most paths compound, a meaningful fraction go to nearly nothing, and which one you get is mostly path luck.
Conditioning on the dot-com window makes it starker. Resampling blocks drawn only from 2000 to 2002 into a 24-month episode, the median drawdown is 93% and the median ending value is about 9 cents on the dollar. One caveat: a within-regime resample of a 27-month window is a small, autocorrelated sample, so treat the dot-com-only bands as illustrative of the regime rather than a forecast.
Why the reservoir runs dry
In a short, sharp correction, March 2020 or December 2018, the rule pulls money from bonds to buy the dip, stocks rebound, the gains refill the sleeve, and the round trip completes. This is the pattern that dominates any backtest that starts in 2010.
In a long, grinding bear market, 2000 to 2002 or 1973 to 1974, the rule keeps pulling from bonds quarter after quarter as stocks drop another leg. The throttle and the floor slow the bleed, and they are what keeps the corrected drawdown at 99.7% instead of a flat zero, but they can't create a refill. Refills only come from sell signals, and deep in a bear market there are no surpluses to sell. The 30 Down rule then deliberately skips the first sell signals of the recovery, so the sleeve stays thin through exactly the stretch where the plan needs it most.
Kelly's answer is new cash. The community simulator at 9sig.networthcast.com starts a $10,000 portfolio in Q1 1999 and adds $500 a month, all of it to the bond sleeve. Every month that refills the reservoir and lifts the signal line by half the contribution, exactly as the Letter describes. Over 27 years the contributor adds $162,000 of real money. The strategy survives the dot-com crash because its operator keeps bailing the boat.
That's the difference between a strategy backtest and a savings plan. A backtest measures what the rules produce in isolation. A plan measures what a disciplined saver produces across decades of contributions, withdrawals, tax events and emergencies. 9Sig passes as a plan for someone who can guarantee monthly contributions and never needs to touch the portfolio during a 99% drawdown. It fails as a strategy for anyone else.
Why the headline numbers look amazing anyway
The real-data 39% CAGR is the arithmetic of a window that starts in 2010, includes the 2009 to 2020 bull market and the 2020 to 2021 rally, and never meets a 3-year Nasdaq bear. Compounding a 3x fund through that decade produces eye-watering numbers. Extend the window to 1999 and the same rules produce 8.3% a year, because the portfolio spends most of the 2000s digging out of a 99.7% hole.
A 99% drawdown is one margin call, one job loss or one medical bill away from being permanent. Anyone who needed cash in 2002 or 2009 sold at the bottom and turned the paper drawdown into a realized loss. The backtest quietly assumes you never needed a single euro during those years. It's a thought experiment dressed up as a retirement plan.
What we list instead
The strategies on our leaderboard, GEM, HAA, VAA, KDA, AlphaOne, Dynamic Macro Allocation, share a property the Signal system lacks: they have an exit. When the trend turns they move to cash or bonds and wait. They don't pull from a reservoir that can empty, and they don't double down into oblivion. Their full-history drawdowns are on their strategy pages, and none of them is anywhere near 99%.
Those strategies have lower CAGR than 9Sig in a window that starts in 2010. They also keep you solvent in 1974, 2002, 2008 and whichever bear market lasts longer than a few quarters. That trade-off is the point of tactical allocation.
What we're keeping and what changed
The core decision stands: 3Sig, 6Sig and 9Sig stay unlisted. Their strategy pages are public so anyone can audit the numbers, but they appear in no leaderboard or ranking, and we won't show a Pro subscriber a ranked entry whose downside is credibly total loss. The corrected returns don't change that risk profile.
The headline claim of the April 11 version is restated. "A 99.3% drawdown in a 9Sig backtest" came from an implementation that didn't follow the rules and happened to land near the right answer. The faithful-rule number over the full window is 99.7%; over the real-TQQQ window it's 72.1%. Both are catastrophic in a closed system. Neither is what Kelly's subscribers experience, because they add new money every month.
The rule set above is the complete parameter list we run, and the rejection log entry carries the same rules with the current engine numbers. If you think we're wrong about Kelly, email us and we'll reopen the case with new evidence. What we won't do is put a 99% drawdown into a subscriber's portfolio and call it research.
What we learned
Getting the rules wrong was the lesson. We wrote a post arguing that a popular strategy would ruin retail investors, and we based it on numbers that didn't faithfully reproduce the strategy. The readers who pointed this out were right, and the fix has been in our tracker since April 19.
The broader point about closed-system leveraged value averaging survives the correction and gets sharper. Without the monthly contributions the Letter assumes, 9Sig drew down 99.7% through the dot-com crash. With real TQQQ data after 2010 and no contributions, it drew down 72% through the 2022 Nasdaq selloff. Either one ends a retail portfolio that can't sit underwater for years. The strategy works for someone with a 30-year contribution horizon and iron discipline. For everyone else, the quiet assumption that the bond sleeve refills itself is still the flaw, and the faithful rules don't repair it. They just make the failure mode more interesting to read about.
If a backtest looks too good to be true, ask what assumption is being held constant that would break in reality. For Kelly it's the infinite reservoir. For other popular systems it's survivor bias in the asset universe, look-ahead in the signal, unrealistic costs, or the quiet exclusion of 2000 to 2002 from the window. Our methodology page lists the 4 gates every strategy clears before it reaches the leaderboard.
For the record: the superseded April 11 numbers
The first version of this article reported these figures from the simplified engine (TQQQ for all 3 variants, an 80/20 start, an all-time-high 30 Down trigger that dumped the bond sleeve, no throttle, no spike reset). They stay here so the correction can be audited. Don't quote them.
| Variant (April 11 engine) | CAGR | Max drawdown | Sharpe |
|---|---|---|---|
| 3Sig | 10.4% | -52.9% | 0.64 |
| 6Sig | 16.3% | -85.5% | 0.59 |
| 9Sig | 21.1% | -99.3% | 0.62 |
Thanks again to the r/LETFs readers who flagged the specific errors.