Verified Track Record

Every signal we published, graded against what price actually did.

Not a win rate. Not a screenshot. The R-multiple of every setup, measured on 5-minute bars with pessimistic fills, scored against the no-skill baseline — what a coin flip would have earned on the exact same entry, stop and target. Every trade is entered at the open of the next session after the signal, the earliest moment a signal provably could not have known. When the number is bad, it stays up.

Last graded 31 Aug 2026 · 684 signals · 38 tickers · 216 flag episodes
Expectancy
−0.11R
per closed trade · n=493
Win rate
57.2%
vs 60.5% no-skill baseline
95% interval
−0.18 … −0.04
bootstrap, 20k resamples
Clustered by episode
t = −1.7
not distinguishable from zero

August 2026 lost money, and the honest answer is that it was close to a coin flip.

The scan's published levels returned −0.1123R per closed trade across 493 trades. A driftless market grading these same levels returns 0.00R by construction. Counting each trade as its own bet, the shortfall clears three standard errors — but 684 signals are only 216 distinct flag episodes over 17 usable scan days, and clustered that way it is t = −1.7: real-looking, not yet provable. We publish the number, the interval and the caveat together, because a track record you only show in good months is marketing rather than evidence.

The direction test

Throw away entry, stop and target. Buy or sell at the next session's open and ask only whether the direction was right. A signal carrying no information scores zero.

HorizonSigned returnRightt (per signal)t (per episode)t (per day)
1 session−0.34%45.5%−2.24+0.02−0.88
3 sessions−0.54%44.9%−2.24−1.12−1.28
9 sessions−0.94%47.1%−2.29−2.13−1.13

684 signals are not 684 independent bets. The same flag is re-published on about three consecutive days, so the real sample is 216 episodes; a dozen of the 38 tickers are the same semiconductor trade. Cluster the errors that way and nothing here clears significance. On this month the direction call is a mild negative tilt and otherwise indistinguishable from a coin.

An earlier draft of this page said 35.4%, and it was wrong

The first pass anchored each trade on the signal's captured_at timestamp. Rows are upserted after the close on a (ticker, trade_date) key, so spot_price is the close of the signal's own trade date while captured_at keeps the earlier pre-market write. The "forward" window therefore contained a session the row already knew the end of. Re-anchoring to the next session's open roughly halved the apparent effect and erased its significance. The error was in our measurement, not in the scan — and this is the sort of thing the page exists to catch.

Two samples, opposite signs

August is not the only month on record. A second, larger sample exists in the backfilled scan — 29 April to 17 July, 47 mostly mega-cap names, 580 flag episodes over 55 scan days, verified to contain no data past its own trade date. Running the identical test on it gives the opposite answer.

SampleEpisodesDaysDriftMarket-neutral excesst
29 Apr – 17 Jul (backfill)58055+0.39%+0.894%+3.82
3–28 Aug (live)21617−1.07%−1.296%−3.21

Both figures are market-neutral: each is measured against the drift on the very names traded that day, so neither is explained by the tape. Longs and shorts both beat drift in the first sample; both lost to it in the second. Pooling the 66 scan days, the excess shows no significant relationship to market direction (r = +0.14, t = +1.1) — so "it only works in a trending market" is a story we cannot yet support either.

Why we weight the negative sample more heavily than the positive one

The backfill covers 29 April to 17 July but was generated between 6 July and 13 August — the same weeks the scanner's parameters were being chosen. screener.py still carries the comment that the dealer-gamma weight was raised "because it is the strongest structural predictor in the track record." A period you tuned on is not evidence. August is the only stretch that is both live and after the parameters were frozen, and August is the one that lost. Until several genuinely out-of-sample months exist, the honest position is that no stable edge has been demonstrated in either direction.

What the 684 signals became
OutcomeCountShare
Target hit24335.5%
Stopped out18727.3%
Timed out at the hold limit639.2%
Still open568.2%
Price passed target before entry was possible537.7%
Never triggered446.4%
No session left in the data (28 Aug signals)385.6%
Why a 54% win rate is a losing number

The target sits closer than the stop

Median planned reward:risk is 0.82 : 1, and 56% of setups are below 1.0. At that shape a coin flip reaches the target 60.5% of the time. A 57.2% win rate is three points below knowing nothing. This is the one finding the anchor fix did not move — it is arithmetic, not measurement.

Winners pay less than losers cost

Average winner +0.50R. Average loser −0.93R. Adverse fills shrink the reward and widen the risk on the same trade, so a target hit pays less than the plan promised while a stop still costs a full R.

How it is measured

Every fill assumption is deliberately pessimistic, because the interesting failure mode of a backtest is flattering itself. Fills: a setup fills only once a bar trades through the trigger, and if the bar gapped past it the fill is the bar's open, not the trigger. No lookahead: every trade is entered at the open of the session after its trade date, and earlier bars are physically absent from the data. We deliberately do not trust the captured_at column for this — it is written on insert while the row is upserted again after the close, so it understates the signal's true vintage. Ambiguity: if one bar touches both stop and target the stop is assumed first, and the row is flagged. Survivorship: unfilled and untradeable setups are reported, never dropped. Trades whose hold window has not closed are shown as open, not marked to market. Excluded: 1,747 backfilled setups, generated weeks after their trade dates — grading those would be hindsight.

Why this page and Graded Trades disagree

This page reports −0.11R. The Graded Trades page reports about −0.07R. Both are honest, and they are not the same measurement. Publishing one and hiding the other would be the easy way out, so here is the difference in full.

 This page (audited)Graded Trades
Bars5-minuteDaily
Populationevery signal, no score filterscore ≥ 60 only
Entry anchor09:30 ET open of the next sessionfirst daily bar that touches the trigger
Hold9 calendar days10 sessions, then a time-stop
Clusteringby flag episodeper trade
Universe38 tickers on all 18 scan daysevery scanned name

Daily bars cannot see an intraday stop and target hit in the same session, so they resolve some trades a coin-flip differently. A score floor keeps the better-shaped setups. A fixed universe removes names that drift in and out. Each choice moves the number a little, and none of them is the “real” one — a single number was never available to publish. What matters is that two measurements built independently, on different data at different resolutions, both land negative and close to zero. Until 31 August 2026 they disagreed by 0.32R and one of them was wrong; the simulator behind Graded Trades filled every entry at its trigger even when price gapped past it, and floored every loss at exactly −1R. It has been rebuilt and every trade re-graded.

If a number here ever moves, the reason moves with it.

Scope. One month, one regime: 3–28 August 2026, restricted to the 38 tickers that appear on all 18 scan days — a rule fixed before any outcome was known. Hold window 9 calendar days. Bars are 5-minute aggregates including extended hours. Signals dated 28 August have no next session in the data yet and are reported as such rather than dropped. Seventeen usable scan days is seventeen independent looks at the market: enough to say August lost money, nowhere near enough to say the signal is broken. The multi-regime replay is the work that would settle it, and it will be published here whichever way it lands.