Seven Months or Seven Centuries: Re-counting When My Gold EA Could Prove Anything

1 October 2026, 01:28
Yuki Nakayama
0
21

This is the sixth entry in the incubation diary for my gold breakout EA. The EA trades XAUUSD on M15: during 10:00 to 13:59 UTC it keeps a buy stop at the highest high and a sell stop at the lowest low of the last 128 closed bars, with an initial stop of 0.5 ATR, a trailing exit, and a maximum hold of 480 minutes. It runs on a demo account. The last entry ended with a promise: work out how long this forward test would need to run before it could say anything statistically meaningful, once trades are counted as independent breakouts rather than as individual fills.

The EA's own source header already contains an answer. It says that confirming profitability in a forward test would take about seven months when trades are counted without overlap, and that only execution quality can be checked sooner. I wrote that line in August. This entry re-derives it with the simulator that matches the EA, and the answer is that the seven-month figure was wrong twice: once in the size of the edge it assumed, and once in what "without overlap" was actually counting.

Where seven months came from

A breakout EA like this one fires many orders that are not independent. In my early research notes, under the simulator I later corrected, the strategy produced about 14.3 triggers per day but only about 1.37 distinct breakouts per day. Treating every trigger as a separate observation would have reached a t-statistic of 2 in about two weeks, which could not be right. So I used a stricter count: keep a trade only if it starts at least 480 minutes (the maximum hold time) after the last trade I kept. That gave about 1.3 trades per day and about 200 trades to reach t = 2, which is roughly 7.2 months of trading days. That is the number in the header. At that rate the rule was still keeping more than one fill per day, which should have told me it was not yet a first-fill count for one session.

Two things have changed since then. The simulator that produced it counted the same fill several times, which I described in the second entry (Seven Checks, Two False Passes). And the per-trade result under the corrected, strict simulator is +0.0068 R, not the +0.3 to +0.4 R the early work assumed (the fifth entry, One Decision, Four Rulers, has the history). A significance horizon depends on the size of the effect, so the old figure could not survive that change. What I did not expect was the second problem.

What "non-overlapping" keeps

I re-ran the strict simulator for the live configuration (Session A, 128-bar window, XAUUSD, January 2010 to July 2026) and first confirmed it reproduced the published figures: 6,576 fills, +0.0068 R per trade, and a non-overlapping t-statistic of +6.04. That last number had always sat in my summary tables next to the per-trade average. I had read it as "the same strategy, counted honestly", and +6.04 looked comfortably significant.

Then I checked which trades the rule actually keeps. Session A is four hours long and the overlap rule demands an eight-hour gap. So once a fill is kept, every other fill that day is dropped, and the next kept fill is the first one of the next session. The rule kept 2,153 fills, and there were exactly 2,153 days with at least one fill. "Non-overlapping" in my tables meant, precisely, "the first fill of each day".

That would not matter if the first fill were a typical fill. It is not:

Fill within the dayFillsMean R per trade
1st2,153+0.2734
2nd1,553-0.0278
3rd1,058-0.1102
4th732-0.2111
5th475-0.2994
6th and later605-0.1444
All fills6,576+0.0068

The 4,423 fills that the rule discarded average -0.1230 R. The 2,153 it kept average +0.2734. The t-statistic of +6.04 belongs to the 2,153 kept fills. The same calculation on all 6,576 fills gives +0.31.

The EA does not trade only the first fill. It re-places both stops on every M15 bar of the session and allows up to three positions at once. I applied that three-position cap to the simulated fills and it removed exactly one trade out of 6,576. So the strategy I am forward-testing is the bottom row of that table, and the number I had been quoting as its honest significance was measured on the top row.

Was the overlap a real problem at all?

The reason for a non-overlapping count was to stop correlated trades from inflating the sample size. It is worth asking how correlated these fills actually are. Under the strict simulator, consecutive fills on the same day have a correlation of +0.098 in R. If I sum R by day, the variance of the daily totals is 1.155 times what independent fills would produce. Only 289 of the 6,576 fills share a day, direction and exit minute with another fill.

So once the simulator stopped counting one fill several times, the remaining overlap between fills was modest. The correction I had designed for a variance problem was mostly doing something else: choosing a subset. That is the specific mistake. I treated "non-overlapping" as a more conservative version of the same measurement, when for this EA it was a different measurement.

The honest count

With the strategy defined as what the EA actually trades, here is how long a forward test would need to reach t = 2 at the historical average, using historical rates of about 33 fills, 10.8 first-of-day fills and 21.6 session days per month:

Counting methodMean per unitUnits needed for t = 2Time needed
Header claim (old simulator, 480-minute rule)+0.3 to +0.4 R (assumed)about 200 tradesabout 7.2 months
Strict simulator, first fill per day only+0.2734 R236 tradesabout 22 months
Strict simulator, all fills+0.0068 Rabout 280,000 tradesabout 700 years
Strict simulator, daily totals+0.0105 R per dayabout 210,000 daysabout 800 years

The last two rows are not meant as forecasts; nobody runs a test for seven centuries. They are a way of saying that an average this close to zero, with a per-trade standard deviation of about 1.8 R, cannot be distinguished from zero by any forward test. Even the flattering first-fill count, under the strict simulator, takes almost two years, not seven months.

What the live record says

The EA's own trade ledger starts with a trade on 18 August. By 30 September it held 47 closed trades on 20 different days, all in Session A, which is close to the simulated rate of about 33 a month. The live numbers:

MeasureValue
Closed trades47
Mean R per trade+0.0266
Median R per trade-1.003
Standard deviation1.35 R
t-statistic+0.135
Approximate 95% interval for the mean-0.36 to +0.41 R
Sum+1.25 R

The median of -1.003 is the shape of the strategy: most trades end at or slightly beyond the initial stop, and the average is carried by a few large trailing exits (the best live trade so far is +4.51 R). In the 16.5-year simulation, the top 5 percent of fills sum to +1,756 R while all fills together sum to +44.8 R, so the remaining 95 percent lose a little over 1,700 R between them.

The interval from -0.36 to +0.41 R contains zero, the long-run simulated mean, and values that would be very attractive. Forty-seven trades do not discriminate between them. To see what a significant result would mean at this sample size: if the true mean were +0.0068 R with the simulated spread of 1.8 R per trade, the chance of seeing t above 2 after 47 trades is about 2.4 percent, which is barely above the rate you would get with no edge at all. A significant early result from this forward test would be much more likely to be luck than evidence.

I also split the live trades the same way as the simulation. The first fill of each day averaged -0.020 R (20 trades) and later fills +0.061 R (27 trades), the opposite of the simulated pattern. With samples this small I do not read anything into that either way.

The tempting conclusion

The fill-order table invites an obvious change: trade only the first fill of each day. The rule uses no future information; at the moment of the first fill, the EA knows it is the first. On the 16.5-year sample it would turn +0.0068 R per trade into +0.27.

I am not making that change, and I want to be specific about why. I found the pattern by looking at the same sixteen years that produced every other number in this diary, while investigating something else. Split at February 2018, the first-fill subset averages +0.4361 R before the split and +0.1188 after it, with a t-statistic of +2.18 in the later half. That is weaker, and still positive. Since January 2025 the first-fill subset averages -0.0026 R over 178 trades. One plausible mechanism is that after the first break the 128-bar extreme sits right at the current price, so later stops are triggered by ordinary noise near the new high or low rather than by a fresh move. That is a hypothesis I have not tested, and a plausible story is not a test.

If I pursue it, the order will be: write down the rule, the measurement and the pass threshold first, then test it on data and a time period that did not suggest it, and only then touch the live EA. Changing the live parameters now would also reset the forward record, and the forward record is worth more to me than a backtest improvement I found by browsing.

What changes

Three things. First, the EA header line about seven months is wrong and will be replaced with what this entry found: under the strict simulator, a forward test cannot confirm or reject the long-run average, and what it can measure is execution. Second, when I report a non-overlapping statistic from now on, I will report which trades it kept and what the trades it dropped averaged. A t-statistic without its complement can describe a different strategy from the one being traded, and nothing in the number tells you that. Third, the forward test's job is now explicit. It cannot tell me whether the entry has an edge. It can tell me whether the live cost per trade matches the model, and at an average this close to zero, that is where the outcome will be decided. That comparison is what my free Execution Caliper utility measures.

The general lesson is uncomfortable but short. A correction that is applied to make a statistic more conservative needs to be checked for what it selects, not just for what it removes. Mine removed overlap, which was small, and selected the best fill of each day, which was large.

Next entry: the cost side. Forty-seven live trades are not enough to judge the entry, but they are already enough to compare what each trade actually cost with what the model charges.