From Backtest to Validation: How I Test Whether a Trading Strategy Is Real
A profitable backtest is only the beginning.
One of the biggest mistakes in strategy development is treating a good historical result as proof that a strategy works.
The more important question is:
Does the idea continue to behave reasonably when it is exposed to data it was not designed around?
That is the main purpose of validation.
My general process is:
Develop → Freeze → Out-of-sample → Cross-market → Forward test
🎥 Short video summary
A short example of the validation process described below.
Start with one idea and one dataset
I normally begin with a clearly defined trading concept on one market.
For example, imagine developing a rule-based strategy on EURUSD.
At this stage, the purpose is not to search through hundreds of parameter combinations until the equity curve looks perfect.
The purpose is to translate the trading idea into objective rules that can actually be tested.
There will naturally be some development: defining entries, exits, stops, sessions and how ambiguous situations should be handled.
But eventually there has to be a point where the rules are frozen.
That point is important.
If I continue changing the strategy every time I see a losing period, I am no longer really testing the original strategy. I am gradually fitting it to historical data.
Freeze the rules before looking elsewhere
Once I am satisfied that the strategy definition makes sense, I stop modifying the core logic.
Only then do I move to data that did not influence the development.
This can be:
- a later period on the same market,
- another currency pair,
- or preferably both.
For example:
Develop on EURUSD → freeze the rules → test later EURUSD data → apply the same rules unchanged to USDJPY or GBPUSD.
The important word is unchanged.
If the EURUSD version uses one set of rules, I do not want to create completely different filters for USDJPY simply because the initial results look worse.
Otherwise the second market becomes another fitting exercise rather than a validation exercise.
What am I looking for in out-of-sample testing?
I am not expecting every market to produce the same result.
A strategy might perform very well on EURUSD, moderately on USDJPY and poorly on another pair.
That can still be useful information.
What interests me is whether the underlying behaviour survives.
I look beyond net profit and pay attention to things such as:
trade count, drawdown, losing streaks, profit factor, consistency across different periods, and whether the majority of the result came from only a handful of trades.
A strategy showing a strong total return may look attractive.
But if nearly all of that performance came from one unusually favourable month, I would view it differently from a strategy that produced more stable behaviour across several market conditions.
Why I like cross-pair testing
Testing the same idea on another currency pair is one of my favourite robustness checks.
The objective is not necessarily to create a multi-pair trading system.
Instead, I am asking:
Was the original behaviour specific to EURUSD historical noise, or does the same market concept appear elsewhere too?
If the exact same rules show useful behaviour across several independent markets, my confidence in the concept increases.
If the strategy completely collapses everywhere except the market where it was developed, that deserves investigation.
That does not automatically make the strategy invalid. Some market behaviours can genuinely be instrument-specific.
But it is an important warning sign.
I try not to “repair” every weak period
This is probably one of the hardest disciplines in backtesting.
Suppose a strategy performs poorly for three months.
It is very tempting to add another filter.
“Do not trade on Mondays.”
Then another:
“Only trade when volatility is above X.”
Then another:
“Except between these particular hours.”
Eventually the historical drawdown may disappear.
Unfortunately, the strategy may now contain a collection of rules whose main purpose is explaining the past.
Sometimes a weak period should simply remain a weak period.
Real strategies can have losing streaks and drawdowns.
The question is whether those periods remain compatible with the overall statistical behaviour and risk profile of the strategy.
The final test is still forward performance
No historical testing method can prove future profitability.
Backtesting helps me reject weak ideas and identify strategies that may be worth taking forward.
Out-of-sample testing and cross-market validation increase my confidence that I have not simply fitted the original historical sample.
But the final stage is still forward testing on genuinely new market data.
That is why my preferred sequence is:
Develop → Freeze → Out-of-sample → Cross-market → Forward test
The further a strategy can travel through that process without requiring major changes, the more interesting it becomes.
That does not guarantee that it will continue working.
But it gives me much more confidence than simply finding the settings that produced the best historical equity curve.


