Discussing the article: "Overcoming The Limitation of Machine Learning (Part 2): Lack of Reproducibility"

 

Check out the new article: Overcoming The Limitation of Machine Learning (Part 2): Lack of Reproducibility.

The article explores why trading results can differ significantly between brokers, even when using the same strategy and financial symbol, due to decentralized pricing and data discrepancies. The piece helps MQL5 developers understand why their products may receive mixed reviews on the MQL5 Marketplace, and urges developers to tailor their approaches to specific brokers to ensure transparent and reproducible outcomes. This could grow to become an important domain-bound best practice that will serve our community well if the practice were to be widely adopted.

For this discussion, I randomly selected two brokers I, personally, use for independent trading. In line with our community guidelines, which prohibit broker promotion, their names have been redacted and replaced with “Broker A” and “Broker B.”

Using the MetaTrader 5 Python library, I requested four years of daily historical EURUSD data from both brokers. Upon review, I noticed that the timestamps didn’t align: one broker’s data extended back to September 2019, while the other’s only reached August 2020. Nevertheless, both returned exactly 1,460 rows of daily data, correctly fulfilling our request.

Given the decentralized nature of brokers, it’s expected that their operating time zones may differ. Less obvious, however, are the effects of daylight saving time, recognized public holidays, and other subtle discrepancies, all of which can further skew timestamp alignment.

We then calculated the 10-Day EURUSD return on both brokers and found that the numerical properties of the EURUSD symbol were inconsistent with each other. The average 10-Day EURUSD return with Broker A was 0.000267 while with Broker B, the average 10-Day return was -0.000352. This represents a difference of about 232% in the expected return of the same underlying asset.

To make matters worse, it appears that the expected returns from Broker A carry 21% more risk than the expected returns from Broker B. This was suggested to us by the fact that the variance in returns between Brokers grew by the same amount, 21%. 

Author: Gamuchirai Zororo Ndawana

 
Devastating.
Thank you so much for the article.
 

I found the article main thesis very interesting, particularly the idea that an ML model can behave quite differently depending on the data feed provided by each broker.

However, while reading through the analysis, I noticed an important methodological issue that may have a significant impact on the conclusions.

The two historical datasets begin on different dates, yet they are later combined using:

pd.concat([A, B], axis=1)
Since both DataFrames have consecutive integer indexes, Pandas aligns them by index rather than by timestamp. This means that many rows from A and B may actually correspond to different trading days.

For example, if A starts in 2019 and B starts in 2020, row 0 in both datasets does not represent the same point in time. This could directly affect the reported correlation, differences in means and variances, and several of the subsequent comparisons between the two brokers.

I think the two datasets should first be synchronized by "time", using only common timestamps, for example with a "merge" or "join" on the datetime column, before drawing conclusions from these comparisons.

The main idea of the article seems valid and highly relevant, but as the experiment is currently designed, I am not sure the presented results demonstrate it conclusively.

It would be very interesting to see how the results change after repeating the analysis with both datasets properly aligned in time. 😉
 
Juan Luis De Frutos Blanco #:
Devastating.
Thank you so much for the article.
Thank you Juan Blanco. 

Do stay tuned for revisions and future discussions seeking to implement solutions to this possible danger. 
 
Miguel Angel Vico Alba #:

I found the article main thesis very interesting, particularly the idea that an ML model can behave quite differently depending on the data feed provided by each broker.

However, while reading through the analysis, I noticed an important methodological issue that may have a significant impact on the conclusions.

The two historical datasets begin on different dates, yet they are later combined using:

Since both DataFrames have consecutive integer indexes, Pandas aligns them by index rather than by timestamp. This means that many rows from A and B may actually correspond to different trading days.

For example, if A starts in 2019 and B starts in 2020, row 0 in both datasets does not represent the same point in time. This could directly affect the reported correlation, differences in means and variances, and several of the subsequent comparisons between the two brokers.

I think the two datasets should first be synchronized by "time", using only common timestamps, for example with a "merge" or "join" on the datetime column, before drawing conclusions from these comparisons.

The main idea of the article seems valid and highly relevant, but as the experiment is currently designed, I am not sure the presented results demonstrate it conclusively.

It would be very interesting to see how the results change after repeating the analysis with both datasets properly aligned in time. 😉
I love it, I love it, I love it <3

Thank you very much Miguel Alba, for highlighting that critical oversight. Indeed I agree with you, such a methodological error can have catastrophic effects on the entire analysis. I will definitely revise the experimental design, and publish a corrected analysis of the effect brought about by different data feeds. 

Indeed two heads are better than one.