Stress Tests
Part B ended with a tidy theory. Part C's job was to torture it — new pairs, twisted dials, doubled costs, simulated futures. The theory cracked in the most instructive way possible.
C.1The transfer test — the drift theory meets three new rivers
Part B's conclusion: Sentinel harvests drift, and USDJPY's five-year one-way current is its edge. The obvious prediction: other yen pairs, which shared the same yen-weakness current, should work too. You exported them. Here's the exam:
| pair | net 5y drift | trades | PF | return | max DD |
|---|---|---|---|---|---|
| USDJPY · the original | 48.4% | 1,876 | 1.16 | +18.8% | 4.2% |
| EURJPY | 43.4% | 1,935 | 0.86 | -20.0% | 25.6% |
| GBPJPY | 44.2% | 1,993 | 0.86 | -27.0% | 30.5% |
| AUDJPY | 42.0% | 1,971 | 0.9 | -12.1% | 14.0% |
All three crosses drifted almost exactly as far as USDJPY (+42–44% vs +48%) — and Sentinel lost on every one of them, with drawdowns five times deeper. Drift is necessary (Part B's lake, EURUSD, still explains the stop-out gap) but it is not sufficient. Whatever makes USDJPY work is more specific than "the river flows."
C.2The attribution hunt — what makes USDJPY different?
Three suspects, each tested:
Spread The crosses pay 1.5–2.0 pips vs USDJPY's 1.0. Re-run EURJPY and GBPJPY at USDJPY's spread: losses shrink (−20%→−13.5%, −27%→−13.6%) but remain heavy. Explains roughly a third of the gap. Partial.
Roughness Maybe USDJPY's path is smoother per unit of progress? Kaufman efficiency ratio over 24h: USDJPY 0.235, EURJPY 0.226, GBPJPY 0.222, AUDJPY 0.222. Nearly identical. Acquitted.
The remainder After costs and roughness, a large gap stands unexplained by every simple aggregate we measured. The honest label is: empirical, single-pair, not yet understood.
An edge you can't explain is an edge you must watch, not worship. It may be USDJPY's particular rhythm — the policy-driven grind, the Tokyo/NY liquidity pattern Part B's clock showed — or it may be partly luck that five years of data can't distinguish. Either way, the consequence for Part D is fixed: modest size, strict kill criteria, and the forward test as judge. This finding also softens Part B's story: read the two parts together as theory → falsification → humility.
C.3The plateau maps — is the roster config a needle?
Lesson 9's question, asked three times. First the Waddah dead zone (the Round-4 "tweak"):
| dead zone (pips) | PF | return | trades |
|---|---|---|---|
| 10.0 | 1.16 | +19.0% | 1,877 |
| 20.0 | 1.16 | +18.8% | 1,876 |
| 30.0 | 1.16 | +18.9% | 1,871 |
| 45.0 | 1.15 | +17.3% | 1,844 |
Flat from 10 to 45 — a total plateau. The Round-4 "improvement" from dead-zone 20 was real but marginal; no needle risk here (and consistent with Part B's finding that this vote rarely binds).
The rudder — the SMA period behind Sentinel's most decisive vote:
| SMA period | PF | return | trades |
|---|---|---|---|
| 30 | 1.05 | +6.8% | 2,107 |
| 40 | 1.13 | +16.4% | 1,939 |
| 50 ← default | 1.16 | +18.8% | 1,876 |
| 65 | 1.16 | +18.2% | 1,784 |
| 80 | 1.11 | +12.0% | 1,772 |
A clean hilltop at 50–65 with soft shoulders: 30 is too twitchy (+6.8%), 80 too sleepy (+12.0%). The default sits on the plateau's crown — whoever chose 50 chose well.
And the risk stack (rows = stop width, columns = reward ratio):
| 1 : 2 | 1 : 3 | 1 : 4 | |
|---|---|---|---|
| ATR × 1.0 | -0.3% PF 1.0 | +5.8% PF 1.05 | +11.2% PF 1.09 |
| ATR × 1.5 | +9.6% PF 1.08 | +18.8% PF 1.16 | +15.8% PF 1.13 |
| ATR × 2.0 | +17.4% PF 1.15 | +17.9% PF 1.15 | +20.0% PF 1.17 |
The roster's ATR×1.5 / 1:3 sits in a broad healthy region — its neighbours (2.0×2, 2.0×3, 1.5×4, 2.0×4) all score within a whisker. The only bad corner is tight stops with modest targets (1.0×2 ≈ break-even). Plateau confirmed.
C.4The committee audit — firing the silent voters
Part B noticed NonLagDot and WAE-direction never cast the lone dissent. So what happens if we remove them? (Method: rebuild the unanimity vote from the indicator columns — verified to reproduce Sentinel's own signals 100.0% exactly — then delete voters.)
| committee | trades | PF | return |
|---|---|---|---|
| full committee (baseline) | 1,876 | 1.16 | +18.8% |
| drop NonLagDot | 1,876 | 1.16 | +18.8% |
| drop WAE direction | 1,876 | 1.16 | +18.8% |
| drop both | 1,901 | 1.14 | +17.2% |
| drop WAE active+rising | 2,072 | 1.1 | +14.1% |
Finding Dropping either silent voter alone changes nothing — literally identical trades. They are fully shadowed by the rest of the committee. Dropping both costs a little (+18.8%→+17.2%): together they still catch a few bad bars.
But dropping WAE active+rising hurts (+18.8%→+14.1%, PF 1.16→1.10): the energy gate genuinely earns its seat.
Verdict Keep the committee as designed. The redundancy is harmless insurance, and simplification offers no measurable gain — a genuinely good outcome for a system this ornate.
C.5The spread shock — how much toll can it bear?
| spread | PF | return | max DD |
|---|---|---|---|
| 1.0 pips | 1.16 | +18.8% | 4.2% |
| 1.5 pips | 1.1 | +12.5% | 4.5% |
| 2.0 pips | 1.05 | +6.2% | 5.3% |
| 3.0 pips | 0.95 | -6.3% | 10.6% |
Break-even sits near 2.6 pips — 2.6× the modeled cost, and far above anything Pepperstone charges on USDJPY in normal conditions. The edge is thin per trade (median hold 3 hours, 1,876 trades) but it is not a costs illusion. Guardrail for Part D: if live spreads at your entry times ever average above ~1.5 pips, the math visibly sags — the dashboard should watch this.
C.6The pain forecast — Monte Carlo
2,000 alternate futures, each 250 trades (~a year at Sentinel's pace), resampled from the real trade distribution. This is the section to re-read during the forward campaign, before checking the dashboard:
With a 35% win rate, eleven losses in a row is the NORMAL year — not a malfunction, not a dying edge, the median. Fifteen straight is merely unlucky. When it happens live (it will), this page is the reason you won't flinch — and drawdowns at this size (0.1 lots on $10k) stay in single digits even at the 99th percentile. Lesson 2's streak math, now with Sentinel's own numbers.
C.7The exit models — was the engine's simplification fair?
Part A flagged that our engine simulates fixed ATR SL/TP, not Sentinel's native exits. Comparing what we can:
| exit model | trades | PF | return |
|---|---|---|---|
| roster: ATR×1.5, 1:3 | 1,876 | 1.16 | +18.8% |
| native intent: ATR×1.5, 1:2 | 1,919 | 1.08 | +9.6% |
| signal-flip only (no SL/TP) | 2,217 | 1.12 | +17.9% |
The roster's 1:3 beats the native 1:2 intent decisively, and even exits-free signal-flip trading earns +17.9% — the alpha lives in the signals; the SL/TP mostly shapes the risk. The one exit we still can't simulate (HA-reversal-beyond-NonLagDot trailing) remains an open fidelity item, now lower-stakes.
C.8Part C verdict
Survived: the config (every dial sits on a plateau) · the committee (no beneficial simplification exists) · the costs case (break-even at 2.6× real spread) · the risk profile (single-digit drawdowns at campaign size).
Did not survive: the drift theory as a full explanation — USDJPY's edge is empirical and partly unexplained, which caps how much trust it deserves.
Handed to Part D: deploy USDJPY H1 only (transfer is disproven, not assumed) · expect streaks of 11–15 · watch live spread (>~1.5 avg = alarm) · a drift/regime gauge and explicit kill criteria, sized for an edge we respect but cannot yet fully explain.