Forward Test on MT5 Gold: One Trade Moved the Mean from +0.03R to +0.17R, and Why I Changed Nothing

10 October 2026, 10:17
Yuki Nakayama
0
18

This is the eighth entry in the incubation diary for my gold breakout EA. The EA trades XAUUSD on M15: during 10:00 to 13:59 UTC it keeps a buy stop at the highest high and a sell stop at the lowest low of the last 128 closed bars, with an initial stop of 0.5 ATR and a trailing exit. It runs on a demo account. The sixth entry counted how long a forward test like this would take to say anything about the entry, and the seventh compared live costs with my backtest cost model. Both used the 47 trades closed by 30 September. Since then the ledger has grown by five trades, and one of them moved the headline number by more than the previous 47 trades combined.

What moved

Every closed trade goes into a ledger file with its result in R, where 1R is the initial stop distance. Fills dated 7 October or later are left out of every number below, because they belong to the pre-registered test described in the last entry and I do not look at them until it is read out. That leaves 52 trades from 18 August to 6 October.

Through 30 SepThrough 6 Oct
Trades4752
Sum of R+1.25+9.02
Mean R per trade+0.0266+0.1734
Median R per trade-1.003+0.289
Winners / losers23 / 2427 / 25

The five new trades are +6.37R (2 October, filled at 12:30 UTC), then +0.71R, +0.89R, +0.93R and -1.13R on 6 October. Their sum is +7.77R, and +6.37R of that is a single trade. Without that trade the mean of the other 51 is +0.052R. Without the two best trades (+6.37R and +4.51R, the latter from 4 September) the other 50 trades sum to -1.87R, a mean of -0.037R. The two best trades are worth +10.89R against a total of +9.02R: everything else, taken together, lost money.

In the sixth entry I read the live median of -1.003R as the shape of the strategy: most trades end at the stop. Five trades later the median is +0.289R, because the win count went from 23 of 47 to 27 of 52. For a payoff that is either about -1R or a trailing exit, the median only reports which side has the majority. Neither median was describing the shape; both were describing a count. At this sample size a headline number tells me which way the last few trades went, not how the EA is doing.

Is that unusual for this strategy?

To judge that I need a reference for what 52 trades from this payoff shape look like. My simulator, the strict version of the backtest used throughout this diary (window of 128 bars, 10:00 to 13:59 UTC, 6,576 simulated trades, mean +0.0068R, standard deviation 1.80R), supplies one. I drew 52 trades at random, with replacement, from its trade list 20,000 times. The draws treat trades as independent, which ignores the small same-day correlation measured in the sixth entry, and they inherit the simulator's cost model, so they are a yardstick and not a forecast.

Question about a 52-trade sampleSimulator drawsLive ledger
Share of trades that win (simulator: all 6,576)53.3% (all 6,576)51.9% (27 of 52)
Mean R, central 95% of draws-0.43 to +0.54+0.173 (bootstrap -0.23 to +0.63)
Chance the mean is at least +0.17324%
Chance the best trade is at least +6.37R49% (a single trade is 1.3%)yes
Chance the two best trades exceed the totalabout 85%yes

The live mean of +0.173R has a t-statistic of +0.80 and an interval that contains both zero and the simulator's +0.0068R. A mean at least that high turns up in roughly one 52-trade sample in four even if the simulator is exactly right. A +6.37R trade is rare per trade (85 of 6,576, 1.3%) but close to a coin flip to appear somewhere in 52, and the two best trades outweighing the whole sum is the normal case for this payoff, not a quirk of this stretch. The ledger agrees with the simulator on the things it can check at this size: a win rate near 52 to 53 percent, and a result carried by a handful of large winners.

The losers are the other half of that shape, and they are the part I can see most clearly. Of the 25 losing trades, 24 closed between -1.0R and -1.2R (mean -1.10R across all 25), which is what an initial stop plus a little slippage should produce. The one exception is -2.07R on 11 September, where the stop-loss was filled $4.55 past its level. The spread of results comes almost entirely from the winners.

One minute carries the result

The seventh entry said I would not discuss 12:30 again until the sealed test had a result. I am going to break that for one section, and I want to say why first. The sealed test is about slippage, on fills from 7 October onward; what follows is about R, on trades that were already excluded from it, and it changes nothing in the EA or in the test. I would rather show the table than have a reader find it in the ledger and wonder why I did not.

When I sorted the 52 trades by fill time, the ledger showed something I had not set out to look for. Six trades were filled at exactly 12:30 UTC:

Fill time (UTC)DirectionR
3 Sep 12:30long+0.09
4 Sep 12:30short+4.51
10 Sep 12:30short+1.55
11 Sep 12:30short-2.07
30 Sep 12:30long-1.20
2 Oct 12:30long+6.37

Those six trades sum to +9.26R, a mean of +1.54R and a median of +0.82R. The other 46 trades sum to -0.25R, a mean of -0.005R and a median of -0.26R. In other words, the whole of the forward result sits in those six fills, and the rest of the ledger is flat to slightly negative.

The seventh entry found that the three most expensive trades of the first 47, in slippage and spread, were all filled at 12:30 UTC and closed within the same minute. They are the 3, 4 and 11 September rows above. The expensive minute and the profitable minute are the same minute. 12:30 UTC is 8:30 a.m. in New York under US daylight saving time, the usual release time for scheduled US data. Two of the six dates (4 September and 2 October) are first Fridays of the month, when the US employment report normally comes out at that time. I have not matched any of the six dates against an economic calendar, because my calendar export ends in July, so this is an association I suspect and have not verified.

I built this table because six identical fill times caught my eye, and I kept it because the sum looked odd. It was chosen after seeing the result, so it counts as a lead and not as a finding. It is six trades. The tempting reading is that 12:30 is where this EA earns its living and everything else is noise, and the equally tempting opposite reading, that 12:30 is where it pays its slippage and should be filtered out, would remove the one group that carries the profit. I have no basis to prefer either. The interval around the +1.54R mean of six trades is enormous, and the mean of the six is carried by two trades, +4.51R and +6.37R.

What I did and did not change

I changed nothing in the EA. There is no time filter, no cost adjustment and no new rule. The pre-registered news-minute slippage test from the last entry stays sealed: it will be read out once, when eight fills have landed inside the release window; if eight have not accumulated by 31 December, it stops without a verdict. It uses fills from 7 October onward only, and those fills are not in any number above.

I also withdraw one framing. The sixth entry presented the median of -1.003R as the shape of the strategy. The shape reading still holds for the losers (24 of 25 at the stop), but the median itself was never a shape statistic at this size: it sat on whichever side had one more trade, and it flipped. From here on, every ledger summary will list the trades beside the sum, so that one fill cannot carry a table without the reader seeing it.

For the forward test, another week with this many trades could move the mean by as much again, in either direction. The reading I can defend today is the one the sixth entry gave: the forward test cannot tell me whether the entry has an edge. What this entry adds is narrower: at 52 trades the ledger's shape is consistent with the simulator's, and that is all it is consistent with. The next entry waits on the sealed test and on the stop-order versus virtual-stop measurement that is still running on a demo terminal.

Source: the EA's trade ledger ( SBrk_trades.csv , rows dated before 7 October 2026) and the simulator's 6,576-trade list, resampled 20,000 times with a fixed seed. The EA is not on sale and nothing here is a performance claim; it is a record of a demo account, and the ledger is small enough that one trade rewrites it.