Every signal we published, graded against what price actually did.
Not a win rate. Not a screenshot. The R-multiple of every setup, measured on 5-minute bars with pessimistic fills, scored against the no-skill baseline — what a coin flip would have earned on the exact same entry, stop and target. Every trade is entered at the open of the next session after the signal, the earliest moment a signal provably could not have known. When the number is bad, it stays up.
August 2026 lost money, and the honest answer is that it was close to a coin flip.
The scan's published levels returned −0.1123R per closed trade
across 493 trades. A driftless market grading these same levels returns 0.00R by
construction. Counting each trade as its own bet, the shortfall clears three standard errors — but
684 signals are only 216 distinct flag episodes over 17 usable scan days, and clustered that
way it is t = −1.7: real-looking, not yet provable. We publish the number, the interval and
the caveat together, because a track record you only show in good months is marketing rather than
evidence.
Throw away entry,
stop and
target. Buy or sell at
the next session's open and ask only whether the direction was right. A signal carrying no
information scores zero.
| Horizon | Signed return | Right | t (per signal) | t (per episode) | t (per day) |
|---|---|---|---|---|---|
| 1 session | −0.34% | 45.5% | −2.24 | +0.02 | −0.88 |
| 3 sessions | −0.54% | 44.9% | −2.24 | −1.12 | −1.28 |
| 9 sessions | −0.94% | 47.1% | −2.29 | −2.13 | −1.13 |
684 signals are not 684 independent bets. The same flag is re-published on about three consecutive days, so the real sample is 216 episodes; a dozen of the 38 tickers are the same semiconductor trade. Cluster the errors that way and nothing here clears significance. On this month the direction call is a mild negative tilt and otherwise indistinguishable from a coin.
An earlier draft of this page said 35.4%, and it was wrong
The first pass anchored each trade on the signal's captured_at timestamp. Rows are
upserted after the close on a (ticker, trade_date) key, so spot_price is
the close of the signal's own trade date while captured_at keeps the earlier
pre-market write. The "forward" window therefore contained a session the row already knew the end
of. Re-anchoring to the next session's open roughly halved the apparent effect and erased its
significance. The error was in our measurement, not in the scan — and this is the sort of thing
the page exists to catch.
August is not the only month on record. A second, larger sample exists in the backfilled scan — 29 April to 17 July, 47 mostly mega-cap names, 580 flag episodes over 55 scan days, verified to contain no data past its own trade date. Running the identical test on it gives the opposite answer.
| Sample | Episodes | Days | Drift | Market-neutral excess | t |
|---|---|---|---|---|---|
| 29 Apr – 17 Jul (backfill) | 580 | 55 | +0.39% | +0.894% | +3.82 |
| 3–28 Aug (live) | 216 | 17 | −1.07% | −1.296% | −3.21 |
Both figures are market-neutral: each is measured against the drift on the very names traded that day, so neither is explained by the tape. Longs and shorts both beat drift in the first sample; both lost to it in the second. Pooling the 66 scan days, the excess shows no significant relationship to market direction (r = +0.14, t = +1.1) — so "it only works in a trending market" is a story we cannot yet support either.
Why we weight the negative sample more heavily than the positive one
The backfill covers 29 April to 17 July but was generated between 6 July and 13 August —
the same weeks the scanner's parameters were being chosen. screener.py still carries the
comment that the dealer-gamma weight was raised "because it is the strongest structural predictor in
the track record." A period you tuned on is not evidence. August is the only stretch that is both
live and after the parameters were frozen, and August is the one that lost. Until several
genuinely out-of-sample months exist, the honest position is that no stable edge has been
demonstrated in either direction.
| Outcome | Count | Share | |
|---|---|---|---|
| Target hit | 243 | 35.5% | |
| Stopped out | 187 | 27.3% | |
| Timed out at the hold limit | 63 | 9.2% | |
| Still open | 56 | 8.2% | |
| Price passed target before entry was possible | 53 | 7.7% | |
| Never triggered | 44 | 6.4% | |
| No session left in the data (28 Aug signals) | 38 | 5.6% |
The target sits closer than the stop
Median planned reward:risk is 0.82 : 1, and 56% of setups are below 1.0. At that shape a coin flip reaches the target 60.5% of the time. A 57.2% win rate is three points below knowing nothing. This is the one finding the anchor fix did not move — it is arithmetic, not measurement.
Winners pay less than losers cost
Average winner +0.50R. Average loser −0.93R. Adverse fills shrink the reward and widen the risk on the same trade, so a target hit pays less than the plan promised while a stop still costs a full R.
Every fill assumption is deliberately pessimistic, because the interesting failure mode of a
backtest is flattering itself.
Fills: a setup fills only once a bar trades through the trigger, and if
the bar gapped past it the fill is the bar's open, not the trigger.
No lookahead: every trade is entered at the open of the session after
its trade date, and earlier bars are physically absent from the data. We deliberately do not trust
the captured_at column for this — it is written on insert while the row is upserted
again after the close, so it understates the signal's true vintage.
Ambiguity: if one bar touches both stop and target the stop is assumed
first, and the row is flagged.
Survivorship: unfilled and untradeable setups are reported, never
dropped. Trades whose hold window has not closed are shown as open, not marked to market.
Excluded: 1,747 backfilled setups, generated weeks after their trade
dates — grading those would be hindsight.
This page reports −0.11R. The Graded Trades page reports about −0.07R. Both are honest, and they are not the same measurement. Publishing one and hiding the other would be the easy way out, so here is the difference in full.
| This page (audited) | Graded Trades | |
|---|---|---|
| Bars | 5-minute | Daily |
| Population | every signal, no score filter | score ≥ 60 only |
| Entry anchor | 09:30 ET open of the next session | first daily bar that touches the trigger |
| Hold | 9 calendar days | 10 sessions, then a time-stop |
| Clustering | by flag episode | per trade |
| Universe | 38 tickers on all 18 scan days | every scanned name |
Daily bars cannot see an intraday stop and target hit in the same session, so they resolve some trades a coin-flip differently. A score floor keeps the better-shaped setups. A fixed universe removes names that drift in and out. Each choice moves the number a little, and none of them is the “real” one — a single number was never available to publish. What matters is that two measurements built independently, on different data at different resolutions, both land negative and close to zero. Until 31 August 2026 they disagreed by 0.32R and one of them was wrong; the simulator behind Graded Trades filled every entry at its trigger even when price gapped past it, and floored every loss at exactly −1R. It has been rebuilt and every trade re-graded.
If a number here ever moves, the reason moves with it.
Scope. One month, one regime: 3–28 August 2026, restricted to the 38 tickers that appear on all 18 scan days — a rule fixed before any outcome was known. Hold window 9 calendar days. Bars are 5-minute aggregates including extended hours. Signals dated 28 August have no next session in the data yet and are reported as such rather than dropped. Seventeen usable scan days is seventeen independent looks at the market: enough to say August lost money, nowhere near enough to say the signal is broken. The multi-regime replay is the work that would settle it, and it will be published here whichever way it lands.