Reading your first forward-test week
The trades are live, the cards are landing on your phone, and the temptation is to grade it. Resist. A week of a slow system is a handful of trades — and a handful of trades will tell you almost anything you want to hear. Here's how to read the week honestly, and what's actually worth watching.
1One week is a rumour, not a verdict
Most of the roster trades on the H4 or H1 chart. That's a few signals a week per strategy — sometimes none. So after seven days you might have four RSI trades, three Sentinel, zero RevertX. Now imagine grading a strategy on four coin-flips: three heads and you'd "prove" the coin is rigged for heads. It isn't. You just don't have enough flips.
This is the : at low trade counts, a swings wildly around its true value purely by luck. Only as the trades pile up does the estimate settle onto the real number. The picture below is the whole idea.
So the honest verdict for almost every strategy this week is not "winning" or "losing" — it's not enough evidence yet. The scorecard says exactly that, on purpose.
2What the scoreboard is telling you
The dashboard's Scoreboard tab and python -m python.live.review grade each strategy with one of five marks. Four of them are about evidence, not winning:
Why the ~15-trade floor before any pass/fail? Because a green tick on four trades would be the wobble band lying to you. "Gathering" is the system refusing to fool you — the single most valuable habit a tester can build.
3Normal noise vs a real warning
The skill this week is telling the difference between a number that looks alarming and a signal that actually is. Most scary-looking things are just small-sample noise. A few things are genuine red flags — and notice that none of them are about a losing week.
Looks scary — it's just noise
- A 3-trade losing streak on an H4 system
- Live win-rate 40% when the baseline says 55% (on n=6)
- PF 0.7 after five trades
- One strategy way ahead of the others
- A red day, or a red week, overall
Genuinely worth acting on
- A strategy that should trade and never does — check MT5 is up and the campaign window is alive
- Fills landing far from the card's entry, or stop-outs well past your stop — costs eating the edge
- The same strategy drifting a second week — now it's a pattern
- A documented kill criterion actually tripping
- The dashboard "code from…" stamp going stale — the server is running old code
The tell: the left column is all about outcomes on tiny samples. The right column is about the machine — is it running, filling cleanly, logging honestly. That's the real subject of week one.
4Week one tests the machine, not the edge
Here's the reframe that makes the whole week make sense. You cannot prove an edge in a week — the wobble band forbids it. What you can prove in a week is that the plumbing works: that signals fire on closed bars, that cards reach your phone, that approved orders actually place, that stops sit where the card said, that closes get logged with the right P/L, and that the terminal survives a reboot. That's a real, completable checklist — and it's what the first week is for.
Once a week, market closed: run the scorecard, walk the six checks, jot one or two notes, close the laptop. Change nothing unless a documented kill line tripped.
Full checklist: docs/operations/14_weekly_review.md · run it with python -m python.live.review.
If you get to the end of the week and the machine ran clean, that's a pass — regardless of whether the P/L was green. The edge is a question for month three, not day seven.
5The takeaway
1. A week is a handful of trades, and a handful proves nothing. The wobble band is why "gathering" — not a green tick — is the honest week-one verdict for most of the roster.
2. Separate the outcome from the machine. A losing week is noise; a silent strategy, a bad fill, or a stale server is a real signal. Watch the machine.
3. The discipline is doing nothing. The forward test only means something if you let it run. Notice, note, and wait — the edge reveals itself with volume, not with staring.