Three Trades, 45 Percent of the Bill: Backtest Cost Model vs Live Execution on MT5 Gold

8 October 2026, 01:38
Yuki Nakayama
1
38

This is the seventh entry in the incubation diary for my gold breakout EA. The EA trades XAUUSD on M15: during 10:00 to 13:59 UTC it keeps a buy stop at the highest high and a sell stop at the lowest low of the last 128 closed bars, with an initial stop of 0.5 ATR and a trailing exit. It runs on a demo account. The last entry concluded that a forward test of this length cannot tell me whether the entry has an edge, because the simulated average is too close to zero, but that it can tell me whether each trade costs what my backtests assume. This entry does that comparison.

Every backtest in this diary charges a round-trip cost from one formula: $0.247 + 0.0653 x ATR, where ATR is the average M15 range over the last 96 bars (24 hours), in dollars per ounce. I built that formula from tick data on the same broker before the EA went live. The constant is the median spread in the EA's own session; the ATR term was calibrated on a 64-bar configuration trading 18:00 to 21:59 UTC, an evening window with few scheduled releases, which I noted in the fifth entry. At the volatility of the last two months it charges about $0.80 per round trip. The question is what the live trades actually paid.

How the live cost is measured

No replay or simulation is involved. For every closed position, the terminal's trade history holds two orders and two deals: the entry stop order and the deal that filled it, and the exit stop order (the initial or trailed stop loss) and the deal that filled that. Each order carries the price I asked for and each deal carries the price I got. So for each trade I can read three numbers directly:

  • Entry slippage: how far the entry fill was from the stop price, counted positive when it went against me.
  • Exit slippage: how far the stop-loss fill was from the stop-loss price, positive when against me.
  • Spread: the spread the EA logged at entry. A buy enters at the ask and leaves at the bid, so a round trip measured against the mid price costs one full spread.

Live cost per trade is the sum of the three. This is the same quantity the formula is meant to represent: what a trade loses compared with filling exactly at the mid price at the level it intended. Two limitations up front. The exit spread is not logged, so I assume it equals the entry spread. And this is a demo account, so the fills are the broker's demo execution, which may not match what a real account would get. Everything below should be read with that second point in mind.

The ledger holds 47 closed trades from 18 August to 30 September, the same 47 as the last entry. All 47 matched an order and a deal in the history. One has no spread recorded, so the cost comparison uses 46. Every exit was a stop-loss order (initial or trailed); the EA has no take-profit.

The headline: cheaper usually, dearer in total

Per round trip (46 trades)Live, measuredCost model
Median$0.475$0.804
Mean$0.999$0.819
90th percentile$1.68$0.96
Lowest / highest$0.25 / $9.51$0.69 / $1.02
Total$45.95$37.67

Both readings of this table are true. The typical live trade cost about 40 percent less than the model charges: the median ratio of live to model cost is 0.60, and 35 of the 46 trades came in under the model. And the live total is 1.22 times the model total, because a few trades cost far more than the model can ever charge.

That last point is structural, not a matter of tuning. The formula is a constant plus a multiple of the 24-hour ATR. Across these 46 trades, ATR ranged from $6.72 to $11.79, so the formula could only produce costs between $0.69 and $1.02. It has no way to output $9.51. A 24-hour average of bar ranges barely changes within a single afternoon, while the expensive trades were decided within a minute: the ledger records all three of them as opened and closed inside the same minute.

Where the money went

ComponentMedianMeanLargest
Spread at entry (46 trades)$0.26$0.27$0.41
Entry slippage (47 trades)$0.11$0.48$7.45
Exit slippage (47 trades)$0.07$0.24$4.55

The spread is the quiet, predictable part: between $0.20 and $0.41 on every trade. Slippage is where the dispersion lives. No entry filled better than its stop price, and one filled exactly at it. On exits, one stop filled $0.03 better than its level and three filled exactly at it. All other exits filled against me. I wrote about the entry side in the fourth entry: a stop order is executed as a market order once it triggers, so the direction of slippage carries no information. Only its size does.

The sizes are very uneven. The three most expensive trades account for $20.63 of the $45.95 total, which is 44.9 percent of all execution cost in six weeks:

Fill time (UTC)SideSpreadEntry slipExit slipTotalModelResult
3 Sep 12:30Long$0.26$7.45$1.80$9.51$0.79+0.09 R
11 Sep 12:30Short$0.41$0.66$4.55$5.62$0.88-2.07 R
4 Sep 12:30Short$0.41$4.75$0.34$5.50$0.76+4.51 R

The initial stop on the 3 September trade was $4.14 wide, so the entry alone slipped by 1.8 times the distance the trade was designed to risk. The 11 September trade shows the other way a cost tail reaches the account. Its stop-loss fill slipped $4.55 on a $4.86 stop, and the trade closed at -2.07 R instead of roughly -1 R. A stop loss limits the price at which the exit is requested, not the price at which it is filled.

Without those three trades, the remaining 43 cost $25.32 live against $35.24 in the model. On the evidence of ordinary trades, the model is too expensive. On the evidence of all trades, it is too cheap. The difference between those two statements is three trades.

The 12:30 pattern, and why I am not acting on it

All three expensive trades were filled at exactly 12:30 UTC, which is when many scheduled US data releases come out during US daylight time. That is suggestive, and it is also the kind of observation that has misled me before. Five of the 47 trades were filled at 12:30. Three of them are the three in the table, with combined slippage of $9.25, $5.21 and $5.09. The other two slipped $0.37 and $0.70. So "12:30 fills are expensive" is three cases out of five, and I chose 12:30 as the thing to look at only after seeing which trades were in the tail.

Removing the session's last hour, or pulling orders around releases, would look attractive on these 47 trades. It would also be fitted to exactly the trades that suggested it. So instead of changing the EA, I wrote the hypothesis down on 6 October, with the test fixed in advance. Every trade I had seen by then, the 47 here and a handful from early October, is excluded from it:

  • Hypothesis: fills inside a window of 2 minutes before to 5 minutes after a high-importance USD release slip more than fills outside it. Entries and exits count as separate observations; slippage is measured exactly as above, spread excluded.
  • Data: only fills from 7 October onward, the day after the hypothesis was written. The 47 trades in this entry, and the few that followed them before 7 October, are excluded because I had already seen them.
  • Test: a one-sided Mann-Whitney test at the 5 percent level, run once, when the window contains 8 observations. If 8 have not accumulated by 31 December, the test stops without a verdict.
  • What I do not look at: profit, R or win/loss by time of day. The test is on slippage alone.
  • If it passes, the only change I will consider is cancelling resting stops before a release and re-placing them after, and only if the cost-to-stop arithmetic justifies it. A pass is not permission to cut session hours.

One practical caveat: the economic calendar file I will use to define the windows currently ends in July. It has to be refreshed before any verdict, and the verdict will say so.

What this does to the backtest number

The fifth and sixth entries settled on +0.0068 R per trade as the long-run simulated average for this configuration, after the model cost has been deducted. The natural question is what happens if the model is replaced by the live cost.

Expressed as a fraction of each trade's stop width, live cost exceeded model cost by an average of +0.048 R per trade. Taken at face value, that would turn +0.0068 R into roughly -0.04 R. But a bootstrap 95 percent interval for that average runs from -0.049 R to +0.176 R. In other words, 46 trades cannot tell me whether the model undercharges or overcharges on average. The interval is more than thirty times wider than the edge it would be adjusting. The mean is dominated by the three tail trades, and resampling 46 trades sometimes includes them twice and sometimes not at all.

In stop-relative terms, the median live trade paid 11.3 percent of its stop width against 18.9 percent in the model. Summed over all trades, live cost was 22.8 percent of the total stop distance against 18.7 percent in the model. Both are far above the low single-digit percentages that I have found in long-running EAs with wide stops, and nothing here changes the conclusion: with a stop of half an ATR on M15, cost is a large fraction of every trade's risk, whichever of these numbers is closer to the truth.

What I got wrong, and what changes

I validated the cost model by its average. The proportional term was fitted to the mean cost in the tick data, so it reproduced the mean by construction, and I treated that as validation. In August a tick replay of the EA's configuration over the previous year told me the model was, if anything, slightly harsh on average, and I filed that as reassurance. The same report showed a median cost of $0.25 and a 99th percentile above $9, and I did not read those two lines as a statement about the model's shape. A model with a constant and an ATR term can only produce a narrow band. For a strategy whose profit is concentrated in a small share of its trades (the third entry found that the top 5 percent of simulated trades carry all of the net profit), the shape matters. If the expensive fills and the large trades tend to coincide, a flat cost charged to every trade will misstate exactly the trades the strategy depends on. Two of the three tail trades here were among the larger outcomes of the period, one positive and one negative. That is far too few to say anything about coincidence, but it is enough to say the question is open, and that the average-fit model cannot answer it.

Three things change. First, from now on, when I state a cost assumption, I will report its median and its tail separately alongside the mean, because a single number hides which of the two is wrong. Second, the forward test's cost check is now a standing measurement: each month I will repeat this table from the terminal history with the same script and the same definitions. Third, the 12:30 question is handed to the pre-registered test above, and this entry is the last time I will discuss it before that test has a result.

If you run a stop-entry EA on gold and want to see the cost side of your own broker before trusting a backtest, my free Execution Caliper utility measures spread and the cost-to-stop ratio on a live chart. The free version does not measure fill slippage; that needs real orders, which is what this entry used.

The general point is simple. An average cost can be right in total and wrong for almost every individual trade, and a strategy that lives on a few trades is exposed to the trades the average describes worst.