V I S O R

Strategies · intermediate · 10 min

Why Most Strategies Fail the Gate

Across this track we have taken six of the most famous, most heavily marketed strategies in retail trading, written each as a precise mechanical rule, and run it on real SPY data through Visor's robustness gates. This lesson puts all six verdicts in one place — because the pattern across them is the real lesson, worth more than any single result.

Backtest Results showing the robustness verdict that most strategies fail: strong headline numbers, but no edge left after the gate

Here is the pattern, stated up front: almost all of them made money, and every single one failed. That is not a contradiction. It is the most important thing a new trader can understand, and it is the thing every course selling these strategies is structurally unable to show you.

The six verdicts

Each card below is the honest robustness report, computed once at build time by the same engine the app runs on every backtest. Read How to read a robustness report if the four gates are unfamiliar.

The golden cross — the SMA(50)/SMA(200) crossover. Made about +60% on SPY over five years:

Robustness verdict — SPY · 5y · as of 2026-07-13
Failed robustness testingThis backtest did NOT survive robustness testing. The headline numbers overstate it.
Backtest return 60.37%Buy & hold 72.73%Trades 2

Random control

Beat 87% of 500 randomly-timed versions of itself (real 60.37% vs random average 36.89%).

FAIL

Out-of-sample

Too few trades to compare (1 in-sample, 1 out).

NOT PROVEN

Significance

Too few trades to compute a t-statistic.

NOT PROVEN

Deflated Sharpe

Not enough trades to estimate a Sharpe.

NOT PROVEN

These checks describe how much of this backtest survives statistical scrutiny. Past simulated performance is not a guide to future results, and nothing here is a recommendation to trade.

A moving-average trend follower — the faster SMA(20)/SMA(50) version:

Robustness verdict — SPY · 5y · as of 2026-07-13
Failed robustness testingThis backtest did NOT survive robustness testing. The headline numbers overstate it.
Backtest return 12.41%Buy & hold 72.73%Trades 12

Random control

Beat 31.6% of 500 randomly-timed versions of itself (real 12.41% vs random average 25.67%).

FAIL

Out-of-sample

Too few trades to compare (9 in-sample, 3 out).

NOT PROVEN

Significance

t = 0.546 against a 3.5 threshold.

FAIL

Deflated Sharpe

0.7056 — Sharpe of 0.1577.

FAIL

These checks describe how much of this backtest survives statistical scrutiny. Past simulated performance is not a guide to future results, and nothing here is a recommendation to trade.

Order blocks — buying the impulse away from the last opposing candle. Made about +54%, and even survived out-of-sample:

Robustness verdict — SPY · 5y · as of 2026-07-13
Failed robustness testingThis backtest did NOT survive robustness testing. The headline numbers overstate it.
Backtest return 53.54%Buy & hold 72.73%Trades 47

Random control

Beat 80.2% of 500 randomly-timed versions of itself (real 53.54% vs random average 35.12%).

FAIL

Out-of-sample

In-sample 28.07% · held-out 19.89%.

PASS

Significance

t = 1.671 against a 3.5 threshold.

FAIL

Deflated Sharpe

0.9503 — Sharpe of 0.2438.

PASS

These checks describe how much of this backtest survives statistical scrutiny. Past simulated performance is not a guide to future results, and nothing here is a recommendation to trade.

Fair value gaps — trading the three-bar imbalance. Made about +45%:

Robustness verdict — SPY · 5y · as of 2026-07-13
Failed robustness testingThis backtest did NOT survive robustness testing. The headline numbers overstate it.
Backtest return 44.73%Buy & hold 72.73%Trades 54

Random control

Beat 63.8% of 500 randomly-timed versions of itself (real 44.73% vs random average 41.81%).

FAIL

Out-of-sample

In-sample 24% · held-out 16.73%.

PASS

Significance

t = 1.396 against a 3.5 threshold.

FAIL

Deflated Sharpe

0.9188 — Sharpe of 0.19.

FAIL

These checks describe how much of this backtest survives statistical scrutiny. Past simulated performance is not a guide to future results, and nothing here is a recommendation to trade.

Liquidity sweeps — fading the false breakdown. The one that fails hardest, in two ways at once:

Robustness verdict — SPY · 5y · as of 2026-07-13
Failed robustness testingThis backtest did NOT survive robustness testing. The headline numbers overstate it.
Backtest return 7.07%Buy & hold 72.73%Trades 46

Random control

Beat 32% of 500 randomly-timed versions of itself (real 7.07% vs random average 18.25%).

FAIL

Out-of-sample

In-sample 17.37% · held-out -8.78%.

FAIL

Significance

t = 0.384 against a 3.5 threshold.

FAIL

Deflated Sharpe

0.6493 — Sharpe of 0.0566.

FAIL

These checks describe how much of this backtest survives statistical scrutiny. Past simulated performance is not a guide to future results, and nothing here is a recommendation to trade.

Mean reversion — buying RSI-oversold dips:

Robustness verdict — SPY · 5y · as of 2026-07-13
Failed robustness testingThis backtest did NOT survive robustness testing. The headline numbers overstate it.
Backtest return 6.8%Buy & hold 72.73%Trades 8

Random control

Beat 66.4% of 500 randomly-timed versions of itself (real 6.8% vs random average 3.06%).

FAIL

Out-of-sample

Too few trades to compare (4 in-sample, 4 out).

NOT PROVEN

Significance

t = 0.555 against a 3.5 threshold.

FAIL

Deflated Sharpe

0.6976 — Sharpe of 0.1964.

FAIL

These checks describe how much of this backtest survives statistical scrutiny. Past simulated performance is not a guide to future results, and nothing here is a recommendation to trade.

What the pattern is telling you

Line the six up and three failure modes account for all of them:

  1. They were long a market that went up. SPY roughly doubled over the window. Almost any rule that spends time long makes money — the return is mostly participation, not skill. Four of the six made a positive return and still lost to simply buying and holding. "It made money" was never the question.

  2. The timing didn't beat its own scramble. The random control re-runs each rule 500 times with its entry timing shuffled, keeping the trade count, exit, and market identical. Every one of the six was beaten by a meaningful share of its randomly-timed twins. When a coin-flip on when to enter does about as well as your signal, the signal's timing — its entire claim to skill — was not doing measurable work.

  3. The ones that looked best in-sample broke out-of-sample. The liquidity-sweep rule made +17% on the first 70% of the timeline and lost on the held-out 30% — the textbook overfitting fingerprint. A rule that works on the data you can see and stops on the data you held back has fit the past, not found the future.

The honest conclusion — and the honest caveat

This is not a claim that technical analysis is worthless, that these patterns are fake, or that no one can trade. It is a narrower, harder, more useful claim: the simple, mechanical, first-taught version of every famous strategy, on this instrument and window, was indistinguishable from luck once tested properly — and the reason nobody selling these strategies tells you that is that the test is easy to hide. Show the in-sample line going up; don't show the held-out half or the scrambled twins.

Two honest caveats keep this fair. First, these are specific mechanical versions on one instrument over one five-year window — a discretionary trader reading the same pattern in context is doing something these rules don't capture, and a different instrument or period could look different. Second, and more important: none of this tells you what will happen. It tells you what these rules did, and how to check whether a result is signal or luck. That check — the random control and the out-of-sample split — is the single most valuable habit in trading, and it is the one thing Visor insists on that a £997 course never will.

What to read next