Skip to content
Important: BestFolio provides information for educational purposes only. Nothing on this site constitutes investment advice. Past performance does not guarantee future results. Read full disclaimer
·4 min read·BestFolio Research Team

Is that Sharpe real? Inside the Robustness score

Someone asked under one of my Reddit posts a few weeks back how I actually calculate the robustness score shown on each strategy. Fair question. I went to grab the methodology link and realized I'd never written the answer down anywhere public. I shipped the Robustness column months ago and never documented it. That's fixed, the methodology page now has a full section on it. But I think this number does the most honest job on the whole site, so it gets its own post.

Start with the uncomfortable part. When this post was published on 4 August 2026, I'd backtested 202 strategy variants that were still alive in my database. 150 sat on the public leaderboard; the other 52 belonged to research strategies I never released. Imagine all 202 had zero real edge, just random monthly returns dressed up as strategies. Rank them by Sharpe and look at the winner. It won't look random at all. Given the actual spread of Sharpe ratios across everything I'd tested, the luckiest of 202 no-edge strategies should still show an annualized Sharpe around 0.56. From pure noise. These counts are a dated snapshot; the live values change as research is added or published.

That number bothers me. It means a positive Sharpe is close to meaningless when it comes out of a large menu of backtests. And every strategy platform, mine included, is a large menu of backtests.

What the score actually is

The Robustness score is the Deflated Sharpe Ratio, from a 2014 paper by David Bailey and Marcos López de Prado. It's a probability from 0 to 1: how likely is it that this strategy's Sharpe reflects a real edge rather than being the luckiest pick among everything I tested?

Three corrections do the work.

Catalog size. I count N as every variant I've backtested and still track, whether or not it ever made it to the site. N was 202 in the 4 August snapshot. The unreleased research pile counts, and reference rows like the single-asset 200-day trends count, because every one of them was a chance to get lucky. From the Sharpe spread across those trials I compute that best-of-N-by-luck benchmark, and each strategy's Sharpe has to clear it, which sits well above zero.

Track length. A 0.8 Sharpe over 50 years of monthly data is a much stronger claim than 0.8 over 8 years. Fewer months, lower score.

Return shape. I plug in each backtest's actual monthly skewness and kurtosis instead of assuming a bell curve. Negative skew and fat tails make a raw Sharpe look more dependable than it is, so they pull the score down.

Scores above 0.90 print as a quiet gray number. Anything below 0.90 I flag amber as fragile, right on the leaderboard.

What it says about my own catalog

In the 4 August snapshot, the median score on the leaderboard was 0.99 and about 3 in 10 variants rounded to 1.0. Those figures move whenever the trial set changes. The current values are always on the live leaderboard data.

In that same snapshot, 25 of the 150 public variants sat below the 0.90 line, with 8 more fragile trials in the unreleased pile. The lowest visible scores were 0.48, 0.51, and 0.58. They were all flagged amber, by me, on my own product. Treat those values as historical context, not a current leaderboard extract.

A fragile flag means the track record can't yet separate the strategy's edge from the luck of the draw, given how many things I tested. Single-asset trend rules carry modest Sharpes by construction, so they sit close to the noise benchmark. My honest reading is "this might work, the data can't prove it yet". I'd rather print that than hide it.

The property that sold me on it

Every variant I test raises N, which raises the luck benchmark, which drags every score down a little. Published or not. I widened N from the 150 public variants to all 202 tested ones while writing this post, and watched every single score drop and 5 borderline strategies pick up the fragile flag. My own research output works against my own headline numbers, which is exactly the direction the incentives usually don't run.

As far as I know, nobody else in the retail strategy space deflates their published numbers as a function of their own catalog size. I'd honestly like to be wrong about that, it would be good for everyone.

The caveats, because there are always caveats

I compute the score on monthly Sharpe over the full track record, so flipping the leaderboard to a 5-year or 10-year window leaves it untouched. I built it that way on purpose. It answers one question: is this Sharpe likely real? How the strategy did last year is a different question, and the period columns already answer it.

Even this N is a floor. The parameter variations I tried and threw away before publishing aren't logged as separate trials, and the industry at large has run millions of backtests, so the true luck benchmark out there is higher than anything I can compute. A high Robustness score gets a strategy a hearing. It can't win the case alone. A 1.0 doesn't guarantee the next decade behaves, and a low score mostly tells me to size the position like the evidence is thin, because it is.

If you want to poke at it, sort the leaderboard by the Robustness column and hover the header for the short version. The full math, the 0.90 cutoff, and the paper reference are in the methodology, section 11. And if you're wondering how I handle the cousin problem, overfitting inside a single strategy's parameters, that's the walk-forward deep dive.

Past performance does not guarantee future results. Backtested results are hypothetical and do not represent actual trading.

Share this article

Data and method

Study dates and assumptions are documented in the article and its revisions. Our current methodology explains the platform's data sources, proxy histories, trade timing and inflation treatment.

Explore with tools

Try these strategies on BestFolio

Browse 73 tactical allocation strategies with monthly signals, walk-forward validation, and portfolio blending. Free to start.

Create free account

BestFolio Monthly Briefing

Liked this post? Get a free monthly recap of TAA strategy signals, performance rankings, and market regime updates. No spam, unsubscribe anytime.