Every strategy we tested. Including the failures.
Each entry below was run on 3 years of real 1-minute futures data with an out-of-sample split, executable fills, and slippage charged — the same standard applied to everyone else's strategies and to our own. Most of them lose. Two survive. Two we published and then withdrew when a re-test killed them.
Note the two columns most strategy libraries leave out: OUT-OF-SAMPLE and PER-YEAR. A number that only exists in the sample it was found in is not a result. Where we never measured something, it shows a dash rather than a guess.
+0.114R · positive every year · zero tuned parameters
⚠️ §59 PROVENANCE FLAG (2026-09-10), unresolved: these NQ figures do not reproduce from the cited script. scripts/crabel-graduation.py today returns NQ n=782, +0.086R, PF 1.18, maxDD 17.6R (pre-dedupe it returned +0.088R / maxDD 16.8R), and its 2026 row moves +0.00 → −0.00. The ES and GC rows in lib/stats/crabel-stat.ts DO match that script exactly, so the NQ row is the odd one out — and its trades: 719 is byte-identical to GC's 719. The +0.114R / PF 1.23 pair matches a different computation (breadth-lab.py's unconditional stretch set, n=717, now +0.113R). Nothing here is invented to replace it: the marketed number needs one canonical script before it is quoted again. Still positive at 4× slippage. ES marginal (+0.062R → +0.059R, cost-fragile), GC rejected (−0.070R, unchanged — gold has no duplicate bars). Sizing warning: the drawdown breaches a $2k trailing-DD account at $250/trade.
+0.62R · 65% win · PF 3.62 — but n=43
Strongest expectancy in the corpus and the lowest drawdown, but ~14 signals/yr means the sample is thin. Being proven in public — the forward ledger is on this page.
+0.33R · positive every year · n=114
The broader, thinner half of the AND-gate. Kept as a separate lead because it fires ~38×/yr instead of 14.
+$19/trade per MGC · 56.5% win · n=85 of 110 signals · IS +$14 / OOS +$29 · every year positive · max DD $803
IN A PUBLIC FORWARD TEST from 2026-09-09 — see /research. CANDIDATE, not a lead: ~350 cells were examined. Two errors were found and corrected while building the forward test, both published: (1) the first headline version ('enter at the reclaim close', +$23) was LOOKAHEAD — it selected the day using a confirmation close that happens AFTER the entry; an independent re-implementation gave −$14/trade. (2) We called it a 'breaker block', but the code never implemented one — all 110 qualifying days were same-bar sweep-and-reclaim. Entering at the confirmation bar at market LOSES (−$26). Only the limit retest survives. Pieter's preferred 20-pt target is the weakest cell here (+$9, OOS +$1); 30 and 40 are better. ~28 trades/yr ≈ $540/yr on one MGC before ~$2/trade commissions.
+0.353R · 53% win · PF 1.85 · IS +0.356 / OOS +0.347 · every year positive
~49 trades/yr. Still +0.305R at 4× slippage; 64% of months positive; 0% gap-through so the fill is honest; max DD 6.8R would NOT breach a $2k trailing-DD eval at $250/trade. CAVEAT: ~20 cells were examined across §54–55, so the placebo p≈0.015 does not survive a multiplicity correction. Earns a forward test, not a product claim. It is credible mainly because the same direction (alone/divergent beats confirmed) appears independently in a second, mechanically different setup family.
~68% directional / +1.29R when 4/4 — on n≈60 out-of-sample
A conviction/stand-down filter, not a number to size off. The confidence band is roughly ±12pp and it needs about twice the sample before it earns anything stronger than “beta”.
Published at +0.84R. Honest number: −0.19R.
Our former flagship. A fill-model fix (entries were filling at the level even when price gapped through it) turned it negative. Withdrawn publicly in June 2026.
Published at +0.64R. Breakeven-to-negative at realistic slippage.
Same fill correction. The levels remain useful CONTEXT — the trade does not survive costs. The dashboard card was relabelled rather than deleted.
~86% fill at some point — ~56% as a tradeable RTH fill, ≈breakeven
The famous 86% is a base rate (“fills eventually”), not a trade win rate. Relabelled site-wide once we understood the difference.
Size decides: NQ small 87% / medium 58% / large 32% (base 59%) · direction is a coin flip
Holds on ES (91/65/30%) and GC (76/37/10% — gold's base rate is 41%, not 86%). Yesterday's unfilled gap adds only +5–7pp. After a gap the session closes up ~51–59% either way: a gap tells you about the FILL, never the DIRECTION. Untested as a trade; the prior (§49) says a small gap is also a small target. Now the ONE gap-fill number the site shows.
51–55% win but ≈breakeven: +$5–6/trade on MGC before ~$2 commissions · fading it loses every year
DIRECTION IS SETTLED: continuation beats fade, and fade is negative at every stop size in every year (−$15 to −$22). The constraint is geometry, not direction — winners' MAE is median 6.2 pts but p90 17.6, so a 10-pt stop cuts 32% of winners and the smallest sensible stop (~20 pts) forces 1:1 on a 20-pt target. Timeframe is not a lever (1-min +$7, 5-min +$6, 15-min +$6). The one condition that moves it: a QUIET prior day (range below the 32-pt median) → 54.3%, +$14/trade, OOS +$38, vs −$4 above median. Needs a forward test. DO NOT run this rule on NQ: −$134 to −$217 per trade at every stop tested.
55% of PDH/PDL trades gap through · +0.737R naive → +0.052R executable · stretch gap-through 0.0%
63% of PDH/PDL breaks trigger on the 09:30 bar, which frequently opens already beyond yesterday's level. This reproduces our June withdrawal from scratch and explains it. It also explains why the Crabel stretch survived the same fix: the stretch is derived from today's open (0 of 717 gap-throughs), so a stop order there is always fillable. Rule for all future work: a level the session can OPEN beyond needs an executable-fill model; a level derived from the session's own open does not.
+2,122 pts total: Asia +1,704 (80%) · London +689 · post-close +950 · NY −1,219
Asia positive in both halves and after trimming 5% tails. NY's median session is +0.35 pts — flat — but its big selloffs happen in New York hours (trimmed sum −37, both halves negative). Structural skew (mean/σ ≈ 0.08), not a trade. Also: NY sets the day's high 38% / low 35%; London only 16% / 14%. Range: 39 pts median 3y, 107 pts last 60 sessions.
Breaks on 92% of days — direction is a 51% coin flip
Something really does happen every night: 92% break rate, median excursion 55% of the Asia range. But 27% trap back inside, 23% whipsaw, and no cross-asset filter sorts it. Asia range size predicts London CHARACTER (small Asia → 1.6× expansion, 32% whipsaw), never direction.
Replicates at 90.27% — but 54% of those days had never left the level, and the trade loses
The baseline he never publishes: 40.43% (n=470) when Pre-London did NOT sweep. The lift is real but over half of it is a re-touch tautology — 54.0% of swept days were still beyond the level at 02:00 and retake it 100% of the time by arithmetic. Excluding those: 78.83% vs 40.43%, and distance-matched the residual lift is unstable (+19pp at 0–10 pts, −2pp at 10–25 pts). Same archetype as the 86% gap fill, the 83% IB midpoint and Lien's 87% overnight break. His descriptive marginals replicate almost exactly on our data (both-swept 1.69% vs his 1.7%), so the data is honest; the inference is the problem.
adds nothing over entering at the reclaim close · ASIA/IFVG negative across its entire target×stop grid
Tested across 4 levels (prior day, Asia, London, overnight) × 2 sessions × a 5×5 target/stop grid. The whole grid is published. Only the breaker survived, and the matched placebo showed even that works as a day filter rather than an entry trigger. Also logged: the first run of this grid contained a LOOKAHEAD bug (trading the London range during the London window) that produced a 97.8% win rate — a reminder that a level cannot be traded before its defining window closes.
with executable fills: 0-of-3 confirm +0.574R vs 3-of-3 confirm −0.120R · ES-confirms −0.097R
§59 RE-VERIFIED (2026-09-10): unchanged by the cache dedupe — 0-of-3 +0.574R, 3-of-3 −0.120R, ES-confirms −0.097R reproduce to three decimals on deduplicated data, and an independent clean-room rebuild lands on the same three cells. Until §59 NO committed script reproduced this ladder: breadth-verify.py enters AT the level (the naive fill) and prints +0.847R for the same n=352 cell. The executable-fill half was never saved. It is now scripts/breadth-fillcheck.py. ⚠️ That audit also found a sub-minute lookahead inherited from breadth-verify.py and cross-asset-confirm.ts — the breadth gate reads the reference index's bar AT the entry minute; correcting it collapses the ES-confirms population 352→71 and makes the cell WORSE (−0.471R), so the sign holds, but the AND-gate's own +0.84R shares that defect and needs its own pass. Monotone across breadth (0→3 confirm: +0.574 / +0.039 / −0.085 / −0.120R) and NOT a fill effect: among clean non-gapped fills, ES-confirms is −0.338R vs +0.102R unconditional. Mechanism: everyone beyond their level at once IS the gap-through morning. ⚠️ Our own shipped AND-gate rests on this premise and its live forward test is at −0.06R vs a +0.62R backtest — re-verification pending with the executable-fill model.
+0.027R · 56% win · PF 1.08 — breakeven; continuation reaches 1× range only 49.5%
Breaks in 81% of sessions. Big Asian ranges +0.054R, small ranges −0.095R. Claimed 60–70% win rates are the break base rate, not a trade.
−0.423R · 19% win · PF 0.60 · negative every year — the sweep is not the trap, fading it is
Returns to the opposite side only 35% of the time after a break. Same verdict as Turtle Soup / Judas on NQ. Also: London sets the day's high/low only 16% / 14% of the time — the least of any session.
continues a further 0.5× range 48.8% of the time, not 87% · trade −0.072R · PF 0.80
Our NQ overnight-direction filter does not port to gold (with-direction −0.077R, against −0.060R). The 87% is a 'moves further at some point' base rate over one year.
the 77% single-break stat is REAL; the trade is +0.051R · PF 1.10 · 2023 negative — not +$105k
Single break 76.8%, double 7.2%, none 16%. Break-at-close variants −0.011R / −0.006R. The IB-50 lesson again: the level is real, the geometry eats it. Settings were re-optimised three times in six months by its own authors.
50.6% · 47.8% · 50.3% — three coin flips; trading the 08:30 candle −0.248R
DST-correct fix times (London and New York shift on different dates — most scripts get this wrong). AM→PM drift mean −0.25 pts, n=763. 10:00→11:00 reverses 08:20→10:00 51.4%.
The filter is genuinely predictive (+22pp) and the trade still loses
Same-day fill odds rise from 56.7% to 78.7% after the 70.5% close — real information. But the stop sits far wider than the target, so a 68%-win setup is still negative. And by the time the signal fires, the nearer targets are already behind price.
−0.36R over 3 years (NQ)
Re-run at 3 years it is worse than the original 1-year −0.08R. Adding an SMT/divergence filter lifts it only to breakeven.
−0.34R out-of-sample
One of five independent attacks on the London window, all negative.
Breakeven-to-negative
Fits the corpus law: on index futures the sweep is a continuation signal, not a reversal.
−0.15 to −0.20R (NQ), −0.20 to −0.28R (ES)
Blows a $2,000-drawdown account in 13 trades at 1:1 with a 47% win rate. The inverted (continuation) version also loses — the bare M15 swing sweep has no edge in either direction.
−0.41R (ES) / −0.46R (NQ), negative every year
Retracement-to-imbalance entries have no standalone edge. What pays is a trend-aligned pullback at a watched, once-a-day level — not at every 3-candle gap.
−0.28R across ~72,000 signals
Fires ~95×/day — it is noise, and small swing stops get eaten by slippage. Both the reversal and continuation arms lose. Note that the “order block” taught elsewhere is this same object by another name.
All 44 tested variants negative out-of-sample
The T-spot limit entry makes it WORSE (adverse selection — the limit fills on setups that keep running through you). The CISD gate makes it worse still. Even inverted it is ≤ breakeven.
Their 65% sweep stat is real — it's actually 87% — and the trade still loses
The opposite side gets swept on 87% of NQ days and 90% of ES days — better than they claim. The 1:1 trade is −0.30R (NQ) / −0.48R (ES). Their “aggressive selling” qualifier makes it worse, because you are fading an already-extended range.
Great in-sample, dies out-of-sample; thresholds disagree across instruments
NQ needs RVOL>1.3 to look good, ES needs >1.5, and both die out-of-sample. When the sweet spot moves between instruments and vanishes on unseen data, that is parameter shopping.
88.8% replicates exactly — and it's a range-size tautology
Their number is right (we get 88.8% vs their 88.4%). But it is the maximum of 135 candidates, the whole 04:45–06:45 block scores 88–91%, and double-break rate simply tracks range size: smallest windows 83.8%, largest 59.5%. Their 05:30 range is 17.5pt; their “baseline” 09:30 range is 82.2pt — 4.7× larger.
The 74% directional stat is real — the trade is breakeven once the fill is modelled
On the strongest days price never returns to the midpoint to fill you — and those are exactly the easy winners. Model the fill and 74% becomes ~46%; at 1:1 that is breakeven.
≈ breakeven once fills are modelled
Modest positive on NQ before honest fills; not survivable after. Continuation only pays at major, watched, once-a-day levels.
83% touch rate replicates — the do-nothing baseline is 92%
We reproduce NQ 07:00 at 83.3% vs their claimed 83.4%. But the unconditional probability of touching the midpoint within 3 hours is HIGHER at every hour — 92.2% at 07:00. The breakout is negative information, and the stop-100%/target-mid geometry loses on all three instruments.
Rejected; PM windows are the worst of the day
The windowed flagship tested −0.51R in the PM. The corpus supports ONE quality decision point per day, not three — more windows means more cost events, not more edge.
Six independent attacks on London, all negative
Continuation mildly negative, fade a disaster (−0.28 to −0.46R). The overnight-extreme histogram also kills the popular “06:30 SAST reversal”: that window is the QUIETEST of the night (7.6% of extremes vs a 24% baseline).
Positive only in the 2026 regime — flat 2023–25 at honest fills
The best community-sourced result we have tested — and it still fails. The profit is entirely 2026; through-fills flatten 2023–25 to zero; ES loses in all 16 configurations. Confirming-close entries lose everywhere: the edge, such as it is, lives in the resting limit.
Rejected — the “self-aware” filter makes it worse
Trend-following by continuous band flip loses. Edges live in cross-asset agreement and trend-aligned limit pullbacks, not in chasing every flip.
The paper's window does not replicate on 2023–26
The most credible external candidate we tested. ES is net negative and both instruments flip sign in 2025 — a regime artifact. The overnight-vs-RTH split IS real (NQ +9.4pts/day held overnight, positive every year) but that is long beta while asleep, not an edge, and it is incompatible with prop drawdown rules.
Dead — no exploitable dislocation at 1-min resolution
The relationship is far too efficient at retail cost structures.
−1.05 pts/trade unconditional; his own headline leg (high sweep → up break) is −0.25 with OOS −3.98
760 cells examined; 15 clear our promotion bar, but shuffling the (Asia size, Pre-London action) labels 400 times produces a median of 5 and a p90 of 14 passing cells — P(null ≥ 15) = 0.083, so the survivors are inside the noise, and 5 of the top 6 go negative at 4× slippage. His own 2×2 is half refuted: continuation beats fade in BOTH Asia buckets, so the large-Asia 'flip' claim fails and the size split carries no information — swap our trailing median for his fixed 70.9-pt constant and the winning bucket inverts (S/CONT +2.78 → +0.47). The OR break itself is a big loser (market at the break close: −5.3 to −8.2 pts); the limit retest is only a better fill, and it skips the 15% of breaks that run furthest (median 66.8 pts). The best pre-specified cell, S/CONT 20/30 (+2.78, n=90, positive all four years), fails on OOS retention (52% of IS), a negative neighbour and a bootstrap CI of [−1.47, +7.29]. Verified by an independent re-implementation (scripts/london-playbook-verify.py) — numbers match exactly.
Cannot be mechanized as taught — documented, not dismissed
Every level depends on first subjectively identifying the manipulation leg, confirmed by CISD and MSS. There is no mechanical definition, so there is no backtest — and a framework that cannot be falsified cannot be validated either. That is the finding.
Every number traces to a script in the repository and to the master research log. Market context and research transparency — not financial advice, and not a signal service. Back to research →