Wasserstein Distance for Live Feature-Drift Detection in MQL5: Monitoring ONNX Model Inputs with Optimal Transport
Introduction
Most drift monitors for live ONNX models track only feature means and standard deviations. That is enough for simple level or volatility shifts, but it misses changes in distribution shape that preserve the first two moments — for example, when a unimodal feature becomes bimodal while the mean remains nearly unchanged. A model trained on the old regime then keeps operating on input space it no longer recognizes.
This article builds a detector that looks at the whole distribution instead of a couple of summary statistics, using the one-dimensional Wasserstein-1 distance (Earth Mover's Distance) borrowed from optimal transport theory. Given a frozen reference sample and a live rolling sample, it measures the minimum "work" needed to reshape one into the other. With equal-size samples, W1 has a simple closed form based on sorted values, making it cheap enough to run per feature inside a live MetaTrader 5 Expert Advisor without turning OnTick into a bottleneck.
Where this fits in the series: unlike the earlier ADWIN/Page-Hinkley article, this method compares empirical distributions rather than tracking a change point in a single statistic. The two approaches are complementary.
The implementation includes an 8-feature builder, a per-feature Wasserstein detector with manifest-driven thresholds and actions, an ONNX directional model, ATR-based risk sizing, and an offline Python pipeline for training and synthetic validation.
Why Distribution Shape Needs Optimal Transport
For two samples of equal size n, drawn from distributions P and Q, the one-dimensional Wasserstein-1 distance has an exact discrete form. Sort both samples ascending, pair up the i-th smallest of each, and average the absolute differences:
// W1(P, Q) for equal-size sorted samples W1 = (1/n) * sum_i | P_sorted[i] - Q_sorted[i] |
That is the entire computation. There is no numerical optimization, no simplex solver, no iterative Sinkhorn regularization the way there would be for multivariate optimal transport — the 1D case reduces to sorting because the optimal transport plan between two ordered lists is always the "monotone" one that matches ranks to ranks. That is also why the detector below insists that the reference window and the live window are exactly the same size: the pairing only makes sense when both sides have the same number of order statistics.
It's worth being explicit about why this metric is used rather than KL divergence or the Kolmogorov–Smirnov (KS) statistic. KL divergence becomes undefined whenever the live window contains a value outside the reference window's support — common with financial features once a live window has moved into a new price regime. Wasserstein-1 stays finite for any two finite samples regardless, since it is fundamentally a "cost to move mass" rather than a ratio of densities. The KS statistic only looks at the single largest gap between the two empirical CDFs, missing a distribution that shifted broadly but modestly everywhere — a common real-world pattern. W1 integrates the gap across the entire distribution, so both a broad modest shift and a narrow severe one register, proportionally to how much probability mass actually moved.
Two design choices sit on top of that base metric. First, raw W1 is in the feature's native units, so a log-return near zero and a volume z-score in the tens are not comparable — each feature's W1 gets divided by that feature's reference-window interquartile range, which puts every feature on a similar footing before they are combined. Second, a single blended score can hide a problem: if seven features are calm and one has completely decoupled from its reference distribution, a weighted average can still sit under threshold. The detector tracks both a weighted composite and a max-statistic across features, and treats either one crossing the critical band as a signal.
System Architecture
The pipeline is a straight line from closed bar to trade decision, with the drift detector sitting alongside the model rather than in front of it — it observes the same feature vector the model consumes, so it always describes exactly the input space the model is actually seeing, not a proxy for it.
Component and Data Flow

Fig. 1. Feature vector construction feeds both the drift detector and the ONNX model in parallel; the manifest's thresholds and the risk sizer both feed into the EA's trade-decision layer.
Four include files carry the logic, and the EA mainly orchestrates them. FeatureBuilder.mqh builds the 8-dimensional vector once per closed bar. WassersteinDetector.mqh manages reference/live windows and computes drift. ManifestConfig.mqh loads thresholds and enforces a three-way feature-count consistency check. RiskSizing.mqh converts an ATR-based stop distance into position size using account-specific risk-per-lot. A companion MQL5 script exports historical features and labels to CSV, and two Python scripts train the ONNX classifier and validate the drift detector against synthetic shifts offline.
Everything that is state rather than pure logic is namespaced per symbol and timeframe. The reference window file is named refwindow_<SYMBOL>_<PERIOD>.csv, so attaching the same EA to XAUUSD M5 on one chart and XAUUSD M15 on another never lets one instance's frozen baseline leak into the other's. The manifest is deliberately the exception — thresholds and weights are treated as policy that applies across whatever the EA is attached to, so there is a single drift_manifest.json rather than one per symbol/timeframe pair, unless your instruments differ enough in volatility character to warrant splitting it.
Feature Vector Design
Every feature the model sees also has to be something the drift detector can meaningfully track, so the vector is deliberately compact — eight scalars, each with a specific reason for existing rather than "whatever the indicator library had lying around."
//+------------------------------------------------------------------+ //| CFeatureBuilder - builds the fixed 8-dimensional feature vector | //+------------------------------------------------------------------+ class CFeatureBuilder { private: string m_symbol; // symbol this builder is bound to ENUM_TIMEFRAMES m_timeframe; // timeframe this builder is bound to int m_h_rsi; // handle: RSI(14) int m_h_atr; // handle: ATR(14) int m_h_bands; // handle: Bollinger Bands(20,2) int m_h_macd; // handle: MACD(12,26,9) public: CFeatureBuilder(void); ~CFeatureBuilder(void); bool Init(const string symbol,const ENUM_TIMEFRAMES tf); void Deinit(void); bool BuildFeaturesAt(const int shift,double &features[]); };
The class owns four indicator handles (RSI, ATR, Bollinger Bands, MACD), created once in Init and reused for the EA's lifetime rather than recreated per bar. BuildFeaturesAt turns a bar shift into the fixed 8-value vector; this same class is reused verbatim by the EA, the reference-window bootstrap, and the offline exporter, so the live and training feature vectors can never silently drift apart.
//+------------------------------------------------------------------+ //| CFeatureBuilder::Init | //+------------------------------------------------------------------+ bool CFeatureBuilder::Init(const string symbol,const ENUM_TIMEFRAMES tf) { m_symbol = symbol; m_timeframe = tf; m_h_rsi = iRSI(m_symbol,m_timeframe,FB_RSI_PERIOD,PRICE_CLOSE); m_h_atr = iATR(m_symbol,m_timeframe,FB_ATR_PERIOD); m_h_bands = iBands(m_symbol,m_timeframe,FB_BB_PERIOD,0,FB_BB_DEVIATION,PRICE_CLOSE); m_h_macd = iMACD(m_symbol,m_timeframe,FB_MACD_FAST,FB_MACD_SLOW,FB_MACD_SIGNAL,PRICE_CLOSE); //--- any INVALID_HANDLE means the terminal could not create the indicator instance //--- (bad symbol, bad params, or history not ready yet) - fail loudly instead of //--- silently returning zeros for every feature downstream if(m_h_rsi==INVALID_HANDLE || m_h_atr==INVALID_HANDLE || m_h_bands==INVALID_HANDLE || m_h_macd==INVALID_HANDLE) { Print("CFeatureBuilder::Init failed - one or more indicator handles invalid"); return(false); } return(true); }
//+------------------------------------------------------------------+ //| CFeatureBuilder::BuildFeaturesAt | //+------------------------------------------------------------------+ bool CFeatureBuilder::BuildFeaturesAt(const int shift,double &features[]) { //--- the array is resized here rather than by the caller so BuildFeaturesAt is the //--- single source of truth for "how many features exist" - the 3-way contract check //--- in OnInit calls this once and compares ArraySize() against NUM_FEATURES and the //--- manifest's "num_features" field ArrayResize(features,NUM_FEATURES); double close_buf[],rsi_buf[],atr_buf[],upper_buf[],lower_buf[],macd_buf[],signal_buf[]; long vol_buf[]; ArraySetAsSeries(close_buf,true); ArraySetAsSeries(rsi_buf,true); ArraySetAsSeries(atr_buf,true); ArraySetAsSeries(upper_buf,true); ArraySetAsSeries(lower_buf,true); ArraySetAsSeries(macd_buf,true); ArraySetAsSeries(signal_buf,true); ArraySetAsSeries(vol_buf,true); //--- request one bar more than the deepest lookback so the "5-bar log return" (f1) //--- and the volume z-score window (f5) always have enough closed bars behind `shift` int need = MathMax(FB_VOL_PERIOD,5) + shift + 2; if(CopyClose(m_symbol,m_timeframe,0,need,close_buf) < need) return(false); if(CopyBuffer(m_h_rsi,0,0,need,rsi_buf) < need) return(false); if(CopyBuffer(m_h_atr,0,0,need,atr_buf) < need) return(false); if(CopyBuffer(m_h_bands,1,0,need,upper_buf) < need) return(false); if(CopyBuffer(m_h_bands,2,0,need,lower_buf) < need) return(false); if(CopyBuffer(m_h_macd,0,0,need,macd_buf) < need) return(false); if(CopyBuffer(m_h_macd,1,0,need,signal_buf) < need) return(false); //--- bug-class guard: CopyTickVolume() requires a long[] destination, never double[] if(CopyTickVolume(m_symbol,m_timeframe,0,need,vol_buf) < need) return(false); double atr_now = atr_buf[shift]; //--- ATR sits in every denominator below (f3, f6, f7) - a flat/illiquid opening print //--- can make it exactly zero, which would turn every ratio into +-inf; clamp to a //--- tiny epsilon so the feature degrades gracefully instead of poisoning the window if(atr_now <= 0.0) atr_now = _Point; //--- f0: 1-bar log return features[0] = MathLog(close_buf[shift] / close_buf[shift+1]); //--- f1: 5-bar log return features[1] = MathLog(close_buf[shift] / close_buf[shift+5]); //--- f2: RSI rescaled from [0,100] to [-1,1] so its magnitude is comparable to the //--- other bounded features instead of dominating the composite drift score features[2] = (rsi_buf[shift] - 50.0) / 50.0; //--- f3: ATR expressed as a fraction of price - makes volatility comparable across //--- the very different absolute price levels XAUUSD trades at over time features[3] = atr_now / close_buf[shift]; //--- f4: Bollinger %B - position of price within the band, deliberately left //--- unclamped (can exceed [0,1] on a squeeze breakout, which is itself informative) double band_width = upper_buf[shift] - lower_buf[shift]; if(band_width <= 0.0) band_width = _Point; features[4] = (close_buf[shift] - lower_buf[shift]) / band_width; //--- f5: tick-volume z-score over FB_VOL_PERIOD bars double vol_sum = 0.0, vol_sq_sum = 0.0; for(int i=shift; i<shift+FB_VOL_PERIOD; i++) { vol_sum += (double)vol_buf[i]; vol_sq_sum += (double)vol_buf[i] * (double)vol_buf[i]; } double vol_mean = vol_sum / FB_VOL_PERIOD; double vol_var = (vol_sq_sum / FB_VOL_PERIOD) - (vol_mean*vol_mean); double vol_std = (vol_var > 0.0) ? MathSqrt(vol_var) : 1.0; // avoid /0 on a dead tape features[5] = ((double)vol_buf[shift] - vol_mean) / vol_std; //--- f6: MACD histogram (main - signal), normalized by ATR so it is scale-invariant //--- across instruments/regimes rather than carrying raw price units features[6] = (macd_buf[shift] - signal_buf[shift]) / atr_now; //--- f7: bar range relative to ATR - flags abnormally wide or narrow bars vs. the //--- prevailing volatility regime, independent of direction double high_val = iHigh(m_symbol,m_timeframe,shift); double low_val = iLow(m_symbol,m_timeframe,shift); features[7] = (high_val - low_val) / atr_now; return(true); }
The eight features: f0 is the 1-bar log return and f1 the 5-bar log return, giving immediate momentum plus a slightly longer view. f2 is RSI(14) rescaled to [-1,1]. f3 is ATR(14) divided by close, volatility as a fraction of price. f4 is Bollinger %B, left unclamped since a squeeze breakout is itself informative. f5 is a 20-bar tick-volume z-score computed by hand, since no built-in indicator provides one. f6 is the MACD histogram divided by ATR, scale-invariant across regimes. f7 is bar range divided by ATR, flagging unusually wide or narrow bars.
ATR-dependent features (f3, f6, f7) clamp the denominator to a tiny epsilon to avoid a flat opening bar producing a silent inf; the volume z-score falls back to a unit standard deviation on the same kind of dead tape.
Honest note: predictive performance of this feature set was effectively random — holdout ROC-AUC ~0.50 for logistic regression and 0.4936 for random forest, across five training-size/horizon combinations. The features are retained because this article's focus is the drift layer, not directional alpha.
The Wasserstein-1 Drift Detector
This is the core of the article, and it is built around two windows per feature: a frozen reference window established once (from history, at EA startup, or loaded from a prior run's saved file) and a rolling live window that is always the most recent live_window_size closed bars.
//+------------------------------------------------------------------+ //| CDriftDetector - per-feature Wasserstein-1 drift monitor | //+------------------------------------------------------------------+ class CDriftDetector { private: int m_num_features; // number of monitored features (must equal NUM_FEATURES) int m_ref_size; // reference window length (fixed, frozen at bootstrap) int m_live_size; // live window length - MUST equal m_ref_size (see note in Init) double m_ref_sorted[]; // flattened [feature*m_ref_size + i], pre-sorted ascending, frozen double m_ref_iqr[]; // per-feature IQR of the reference window, used to normalize W1 double m_live_buf[]; // flattened circular buffer [feature*m_live_size + slot] int m_live_head; // next slot index to overwrite on push int m_live_count; // samples pushed since last reset, caps at m_live_size bool m_reference_ready; // true once the reference window has been frozen or loaded void SortSlice(double &arr[],const int offset,const int count); double Wasserstein1D(const double &sorted_a[],const int offset_a, const double &sorted_b[],const int offset_b,const int n); double PercentileFromSorted(const double &sorted[],const int offset,const int n,const double p); public: CDriftDetector(void); bool Init(const int num_features,const int ref_size,const int live_size); bool LoadOrBootstrapReference(const string file_path,CFeatureBuilder &builder); void BootstrapLiveWindow(CFeatureBuilder &builder); void PushLiveSample(const double &features[]); bool IsLiveWindowFull(void) const { return(m_live_count>=m_live_size); } bool IsReferenceReady(void) const { return(m_reference_ready); } void ComputeDriftScores(double &per_feature_w1[],double &per_feature_normalized[]); double ComputeComposite(const double &normalized[],const double &weights[],double &max_stat); };
The reference window is stored pre-sorted (sorted once, when frozen, not on every score). The live window is a flat circular buffer indexed as feature * live_size + slot, avoiding the cost of shifting an array on every new sample.
Cost: computing one feature's score means sorting a 250-element live slice (O(n log n), MQL5's native ArraySort) plus a linear pass against the pre-sorted reference. Across eight features that is eight sorts, done once every recompute_every_n_bars closed bars rather than every tick — cheap enough that the recompute cadence is a freshness knob, not a performance necessity, at this window size and feature count.
//+------------------------------------------------------------------+ //| CDriftDetector::Wasserstein1D | //+------------------------------------------------------------------+ double CDriftDetector::Wasserstein1D(const double &sorted_a[],const int offset_a, const double &sorted_b[],const int offset_b,const int n) { //--- closed-form discrete W1 for two equal-size, already-sorted 1D samples: //--- the optimal transport plan simply matches rank-i of A to rank-i of B, so the //--- distance collapses to the mean absolute difference between paired order statistics double sum_abs_diff = 0.0; for(int i=0;i<n;i++) sum_abs_diff += MathAbs(sorted_a[offset_a+i] - sorted_b[offset_b+i]); return(sum_abs_diff / n); }
//+------------------------------------------------------------------+ //| CDriftDetector::PercentileFromSorted | //+------------------------------------------------------------------+ double CDriftDetector::PercentileFromSorted(const double &sorted[],const int offset,const int n,const double p) { //--- linear-interpolation percentile (same convention as numpy's default) so the IQR //--- used for normalization matches what the Python validation script computes double rank = p*(n-1); int lo = (int)MathFloor(rank); int hi = (int)MathCeil(rank); if(lo==hi) return(sorted[offset+lo]); double frac = rank-lo; return(sorted[offset+lo]*(1.0-frac) + sorted[offset+hi]*frac); }
PercentileFromSorted() is used only to compute the reference IQR with the same interpolation convention as NumPy, ensuring offline and in-terminal consistency.
//+------------------------------------------------------------------+ //| CDriftDetector::BootstrapLiveWindow | //+------------------------------------------------------------------+ void CDriftDetector::BootstrapLiveWindow(CFeatureBuilder &builder) { //--- pre-fill the live window from the most recent m_live_size closed bars (shift 0..m_live_size-1) //--- so the detector is immediately usable after OnInit instead of forcing a silent //--- multi-hundred-bar warm-up period before the first drift score can be produced double feat[]; for(int shift=m_live_size-1; shift>=0; shift--) { if(builder.BuildFeaturesAt(shift,feat)) PushLiveSample(feat); } }
//+------------------------------------------------------------------+ //| CDriftDetector::PushLiveSample | //+------------------------------------------------------------------+ void CDriftDetector::PushLiveSample(const double &features[]) { for(int f=0; f<m_num_features; f++) m_live_buf[f*m_live_size + m_live_head] = features[f]; //--- classic ring-buffer advance: wrap the write head and cap the fill count at //--- m_live_size so IsLiveWindowFull() flips true exactly once the window is dense m_live_head = (m_live_head+1) % m_live_size; if(m_live_count < m_live_size) m_live_count++; }
Bootstrapping the live window at startup avoids a long warm-up period before the first score becomes available.
//+------------------------------------------------------------------+ //| CDriftDetector::ComputeDriftScores | //+------------------------------------------------------------------+ void CDriftDetector::ComputeDriftScores(double &per_feature_w1[],double &per_feature_normalized[]) { ArrayResize(per_feature_w1,m_num_features); ArrayResize(per_feature_normalized,m_num_features); double live_slice[]; ArrayResize(live_slice,m_num_features*m_live_size); //--- copy the circular buffer out in plain chronological-irrelevant order - Wasserstein-1 //--- is a distribution distance, so sample ORDER inside the window does not matter, only //--- the set of values does; this lets us skip unwinding the ring buffer's wrap point for(int f=0; f<m_num_features; f++) for(int i=0;i<m_live_size;i++) live_slice[f*m_live_size+i] = m_live_buf[f*m_live_size+i]; for(int f=0; f<m_num_features; f++) { SortSlice(live_slice,f*m_live_size,m_live_size); double w1 = Wasserstein1D(m_ref_sorted,f*m_ref_size,live_slice,f*m_live_size,m_live_size); per_feature_w1[f] = w1; //--- normalize by the reference IQR so features with very different natural scales //--- (a log-return near zero vs. a volume z-score in the tens) contribute comparably //--- to the composite score instead of the largest-magnitude feature dominating it double iqr = m_ref_iqr[f]; if(iqr <= 0.0) iqr = 1e-8; // degenerate reference (constant feature) - avoid /0 per_feature_normalized[f] = w1 / iqr; } }
//+------------------------------------------------------------------+ //| CDriftDetector::ComputeComposite | //+------------------------------------------------------------------+ double CDriftDetector::ComputeComposite(const double &normalized[],const double &weights[],double &max_stat) { double weighted_sum = 0.0, weight_total = 0.0; max_stat = 0.0; for(int f=0; f<m_num_features; f++) { weighted_sum += weights[f]*normalized[f]; weight_total += weights[f]; //--- track the single worst-drifting feature separately from the blended average - //--- a weighted mean can stay below threshold even while one feature has fully //--- decoupled from its reference distribution, which the max-statistic catches if(normalized[f] > max_stat) max_stat = normalized[f]; } if(weight_total <= 0.0) weight_total = 1.0; return(weighted_sum/weight_total); }
To show this is not just a nice formula that happens to work on a whiteboard, the Python companion script synthetic_shift_validation.py runs the same computation offline against three synthetic shift types — a mean shift, a variance shift, and a bimodal mixture shift constructed to leave the mean and variance almost unchanged — and compares detection rates against a naive mean/std detector, with both thresholds calibrated to roughly the same 5% false-positive rate on unshifted data.
Wasserstein-1 vs. Naive Mean/Std Detection Rate

Fig. 2. Detection rate at a matched ~5% false-positive rate. On an actual run of this script: mean shift 100% vs 100%, variance shift 100% vs 20%, bimodal mixture 100% vs 1%. The naive detector is nearly blind to a shift that changes shape without moving the mean.
That last column is the entire argument for this article in one number. A mean/std monitor caught the bimodal mixture shift on 1 out of 100 trials. The Wasserstein-1 detector caught it on all 100.
Composite Drift Score Timeline

Fig. 3. A synthetic composite-score timeline: quiet noise around the reference level, then a regime break ramps the score through the warning band and into the critical band, where HandleDriftAction fires.
//+------------------------------------------------------------------+ //| CDriftDetector::LoadOrBootstrapReference | //+------------------------------------------------------------------+ bool CDriftDetector::LoadOrBootstrapReference(const string file_path,CFeatureBuilder &builder) { //--- prefer an existing frozen reference so a live EA restart does not silently //--- re-baseline itself against whatever regime happens to be live at restart time if(FileIsExist(file_path)) { int handle = FileOpen(file_path,FILE_READ|FILE_TXT|FILE_ANSI); if(handle==INVALID_HANDLE) { Print("LoadOrBootstrapReference: FileOpen failed for ",file_path); return(false); } //--- FileReadString() on FILE_TXT silently returns only the first line if you call //--- it once and stop - read every line explicitly with FileIsEnding() instead string header = ""; if(!FileIsEnding(handle)) header = FileReadString(handle); string parts[]; int n_parts = StringSplit(header,',',parts); if(n_parts!=2 || (int)StringToInteger(parts[0])!=m_num_features || (int)StringToInteger(parts[1])!=m_ref_size) { Print("LoadOrBootstrapReference: header mismatch in ",file_path, " - deleting stale file and re-bootstrapping"); FileClose(handle); FileDelete(file_path); } else { for(int f=0; f<m_num_features && !FileIsEnding(handle); f++) { string line = FileReadString(handle); string vals[]; int n_vals = StringSplit(line,',',vals); if(n_vals!=m_ref_size) { Print("LoadOrBootstrapReference: row ",f," has ",n_vals," values, expected ",m_ref_size); FileClose(handle); return(false); } for(int i=0;i<m_ref_size;i++) m_ref_sorted[f*m_ref_size+i] = StringToDouble(vals[i]); } FileClose(handle); for(int f=0; f<m_num_features; f++) { double q1 = PercentileFromSorted(m_ref_sorted,f*m_ref_size,m_ref_size,0.25); double q3 = PercentileFromSorted(m_ref_sorted,f*m_ref_size,m_ref_size,0.75); m_ref_iqr[f] = (q3-q1); } m_reference_ready = true; return(true); } } //--- bootstrap path: no frozen file yet, so build the reference from the ref_size //--- closed bars immediately BEHIND the live window (shift = live_size .. live_size+ref_size-1), //--- keeping reference and live disjoint and time-ordered so there is no leakage between them double feat[]; for(int i=0;i<m_ref_size;i++) { int shift = m_live_size + i; if(!builder.BuildFeaturesAt(shift,feat)) { Print("LoadOrBootstrapReference: BuildFeaturesAt failed at shift ",shift, " - not enough history loaded yet"); return(false); } for(int f=0; f<m_num_features; f++) m_ref_sorted[f*m_ref_size+i] = feat[f]; } for(int f=0; f<m_num_features; f++) { SortSlice(m_ref_sorted,f*m_ref_size,m_ref_size); double q1 = PercentileFromSorted(m_ref_sorted,f*m_ref_size,m_ref_size,0.25); double q3 = PercentileFromSorted(m_ref_sorted,f*m_ref_size,m_ref_size,0.75); m_ref_iqr[f] = (q3-q1); } //--- freeze to disk as plain comma-separated text; DoubleToString is used explicitly on //--- every value written (never the default %s cast) to sidestep locale-dependent decimal //--- separators when this file is later read back on a machine with different Windows regional settings int out_handle = FileOpen(file_path,FILE_WRITE|FILE_TXT|FILE_ANSI); if(out_handle==INVALID_HANDLE) { Print("LoadOrBootstrapReference: could not create ",file_path," - continuing without persistence"); } else { FileWriteString(out_handle,IntegerToString(m_num_features)+","+IntegerToString(m_ref_size)+"\r\n"); for(int f=0; f<m_num_features; f++) { string row = ""; for(int i=0;i<m_ref_size;i++) { row += DoubleToString(m_ref_sorted[f*m_ref_size+i],8); if(i<m_ref_size-1) row += ","; } FileWriteString(out_handle,row+"\r\n"); } FileClose(out_handle); } m_reference_ready = true; return(true); }
The reference window draws from the ref_size bars immediately behind the live window (shifts live_size through live_size + ref_size - 1), keeping the two time-disjoint by construction — no bar is ever compared against a window that also contains it.
Manifest-Driven Configuration
Every threshold, weight and window size lives in a JSON file read at OnInit, not in #define constants — tightening a threshold after a week of live warnings should not require recompiling.
//+------------------------------------------------------------------+ //| SDriftManifest - runtime-tunable parameters loaded from JSON | //+------------------------------------------------------------------+ struct SDriftManifest { int num_features; // must equal NUM_FEATURES (3-way contract check) int reference_window_size; // bars in the frozen reference window int live_window_size; // bars in the rolling live window (must equal reference_window_size) int recompute_every_n_bars; // cost control - drift score recomputed every N closed bars double feature_weights[]; // per-feature weight in the composite weighted-sum score double warn_threshold; // composite/normalized score above this -> warning log double critical_threshold; // composite/normalized score above this -> action fires int action_mode; // 0 = log only, 1 = halt new entries, 2 = flag-for-retrain file drop double risk_percent; // percent equity risked per trade (feeds RiskSizing.mqh) };
A full JSON parser is unnecessary machinery for a flat, known-schema file — the loader just locates each expected key by name and lifts out the value substring, zero dependencies beyond MQL5's own string functions.
//+------------------------------------------------------------------+ //| ExtractJsonNumber | //+------------------------------------------------------------------+ double ExtractJsonNumber(const string &text,const string key,const double default_value) { string pattern = "\""+key+"\""; int key_pos = StringFind(text,pattern); if(key_pos<0) return(default_value); int colon_pos = StringFind(text,":",key_pos); if(colon_pos<0) return(default_value); int start = colon_pos+1; int len = StringLen(text); int end = start; //--- scan forward until a delimiter that can never appear inside a bare JSON number while(end<len) { ushort ch = StringGetCharacter(text,end); if(ch==',' || ch=='}' || ch=='\n' || ch=='\r' || ch==']') break; end++; }
//+------------------------------------------------------------------+ //| LoadManifest | //+------------------------------------------------------------------+ bool LoadManifest(const string file_path,SDriftManifest &manifest) { //--- FILE_COMMON is deliberately NOT used here: the manifest is meant to be edited //--- per-terminal/per-tester-agent without needing to touch a shared common folder, //--- and it is small enough that duplicating it per install costs nothing int handle = FileOpen(file_path,FILE_READ|FILE_TXT|FILE_ANSI); string full_text = ""; bool file_found = (handle != INVALID_HANDLE); if(!file_found) { //--- the Strategy Tester's sandboxed Files folder is isolated from the terminal's //--- own Files folder (the same reason OnnxModel.mqh embeds the model as a #resource //--- instead of loading it by path) - a Tester agent will routinely NOT have this //--- file even when it is sitting exactly where it should on the live terminal. //--- Rather than abort OnInit over a missing config file, fall through with an empty //--- full_text: every field below already has a documented default via //--- ExtractJsonNumber's default_value argument, so an empty string simply means //--- every field takes its default - the exact same values the shipped //--- drift_manifest.json ships with. This function still returns true in that case, //--- since the manifest struct ends up fully and validly populated either way. Print("LoadManifest: could not open ",file_path," - using built-in defaults ", "(expected in Strategy Tester agents; copy drift_manifest.json into the ", "agent's own Files folder if you need non-default settings during testing)"); } else { //--- read the whole file into one string rather than parsing line-by-line, since //--- JSON keys can legally be split across lines and a line-oriented reader would miss that while(!FileIsEnding(handle)) full_text += FileReadString(handle) + "\n"; FileClose(handle); //--- a UTF-16 BOM (or a stray one from a text editor) surviving into the first key //--- would break the very first StringFind() match, so strip it defensively StringReplace(full_text,"\uFEFF",""); } manifest.num_features = (int)ExtractJsonNumber(full_text,"num_features",NUM_FEATURES); manifest.reference_window_size = (int)ExtractJsonNumber(full_text,"reference_window_size",250); manifest.live_window_size = (int)ExtractJsonNumber(full_text,"live_window_size",250); manifest.recompute_every_n_bars = (int)ExtractJsonNumber(full_text,"recompute_every_n_bars",10); manifest.warn_threshold = ExtractJsonNumber(full_text,"warn_threshold",1.5); manifest.critical_threshold = ExtractJsonNumber(full_text,"critical_threshold",2.5); manifest.action_mode = (int)ExtractJsonNumber(full_text,"action_mode",1); manifest.risk_percent = ExtractJsonNumber(full_text,"risk_percent",0.5); int n_weights = ExtractJsonNumberArray(full_text,"feature_weights",manifest.feature_weights); if(n_weights != manifest.num_features) { //--- fall back to equal weighting rather than aborting - a malformed weights array //--- (or, as above, a genuinely empty full_text) is a recoverable config error, //--- unlike a feature-count mismatch below Print("LoadManifest: feature_weights length (",n_weights, ") does not match num_features (",manifest.num_features,") - using equal weights"); ArrayResize(manifest.feature_weights,manifest.num_features); for(int i=0;i<manifest.num_features;i++) manifest.feature_weights[i] = 1.0; } return(true); }
The manifest file format, exactly as consumed by LoadManifest:
| Field | Type | Enforced how |
|---|---|---|
| num_features | int | runtime-checked against compiled NUM_FEATURES and BuildFeaturesAt() output length; mismatch aborts OnInit |
| reference_window_size / live_window_size | int | runtime-checked equal to each other; mismatch aborts OnInit |
| recompute_every_n_bars | int | documentation-only cost control; any positive value accepted |
| feature_weights | array of double, length num_features | falls back to equal weights (all 1.0) if the array length does not match num_features |
| warn_threshold / critical_threshold | double | documentation-only; no ordering check, so critical < warn is technically legal (and a config mistake) |
| action_mode | int (0/1/2) | unrecognized value logged and treated as 1 (halt) by HandleDriftAction |
| risk_percent | double | documentation-only; passed straight into CalcLotSize |
//+------------------------------------------------------------------+ //| ValidateManifest | //+------------------------------------------------------------------+ bool ValidateManifest(const SDriftManifest &manifest,const int actual_feature_count) { //--- the 3-way feature contract: the compile-time NUM_FEATURES constant, the manifest's //--- declared num_features, and BuildFeaturesAt()'s actual output length must all agree. //--- Any mismatch means the code was edited (a feature added/removed) without updating //--- the deployed manifest or vice versa - continuing would silently misalign the //--- reference window against a live window built from a different feature set if(manifest.num_features != NUM_FEATURES) { Print("ValidateManifest: manifest num_features=",manifest.num_features, " does not match compiled NUM_FEATURES=",NUM_FEATURES); return(false); } if(actual_feature_count != NUM_FEATURES) { Print("ValidateManifest: BuildFeaturesAt() produced ",actual_feature_count, " values, expected NUM_FEATURES=",NUM_FEATURES); return(false); } if(manifest.reference_window_size != manifest.live_window_size) { Print("ValidateManifest: reference_window_size must equal live_window_size"); return(false); } return(true); }
ValidateManifest enforces the three-way contract: NUM_FEATURES, manifest.num_features, and the length returned by BuildFeaturesAt() must match before OnInit succeeds. This exists because the failure mode it prevents is silent and expensive: add a ninth feature to FeatureBuilder.mqh without updating the deployed manifest, and without this check the EA would happily keep running with a reference window built against the old 8-feature shape while the live window silently carries 9 — the drift score would still produce a number, it would just be meaningless rather than throwing an obvious error.
ONNX Inference and the Two-Output Export Trap
The directional signal is intentionally simple — a logistic regression trained offline, exported to ONNX, and embedded into the compiled EA as a resource rather than loaded from a runtime file path.
//+------------------------------------------------------------------+ //| CDriftGatedModel - ONNX session wrapper for directional signal | //+------------------------------------------------------------------+ class CDriftGatedModel { private: long m_session; // ONNX runtime session handle matrix<float> m_output; // pre-sized output buffer, required by ONNX_NO_CONVERSION public: CDriftGatedModel(void); ~CDriftGatedModel(void); bool Load(void); void Release(void); bool Predict(const double &features[],double &prob_up); };
//+------------------------------------------------------------------+ //| CDriftGatedModel::Load | //+------------------------------------------------------------------+ bool CDriftGatedModel::Load(void) { m_session = OnnxCreateFromBuffer(OnnxModelBytes,ONNX_DEFAULT); if(m_session==INVALID_HANDLE) { Print("CDriftGatedModel::Load - OnnxCreateFromBuffer failed, error ",GetLastError()); return(false); } //--- fix the input shape explicitly to [1, NUM_FEATURES]; the model was exported with //--- a dynamic batch axis, so the batch dimension of 1 must be pinned here or OnnxRun //--- will reject a mismatched tensor shape at inference time ulong input_shape[] = {1,NUM_FEATURES}; if(!OnnxSetInputShape(m_session,0,input_shape)) { Print("CDriftGatedModel::Load - OnnxSetInputShape failed, error ",GetLastError()); return(false); } //--- ONNX_NO_CONVERSION demands the caller's buffer type match the model's declared //--- output type exactly (float32 here, not MQL5's default double) AND be pre-sized - //--- the runtime will not allocate or convert it on our behalf ulong output_shape[] = {1,2}; if(!OnnxSetOutputShape(m_session,0,output_shape)) { Print("CDriftGatedModel::Load - OnnxSetOutputShape failed, error ",GetLastError()); return(false); } m_output.Init(1,2); return(true); }
Model deployment: embedding via #resource instead of OnnxCreate(file_path) sidesteps a real Strategy Tester limitation — a tester agent's sandboxed Files folder is isolated from the terminal's own, so a path-based load that works live can fail, or silently load a stale file, inside an optimizer or tester agent. A #resource embedded at compile time travels with the compiled .ex5 with no runtime file lookup at all.
//+------------------------------------------------------------------+ //| CDriftGatedModel::Predict | //+------------------------------------------------------------------+ bool CDriftGatedModel::Predict(const double &features[],double &prob_up) { //--- MQL5's bare "matrix" is matrix<double>; the ONNX model's input tensor was exported //--- as float32, so under ONNX_NO_CONVERSION the input buffer must be matrix<float> or //--- OnnxRun rejects the call with a type-mismatch error matrix<float> input_matrix; input_matrix.Init(1,NUM_FEATURES); for(int i=0;i<NUM_FEATURES;i++) input_matrix[0][i] = (float)features[i]; if(!OnnxRun(m_session,ONNX_NO_CONVERSION,input_matrix,m_output)) { Print("CDriftGatedModel::Predict - OnnxRun failed, error ",GetLastError()); return(false); } //--- output tensor "probabilities" is [P(down), P(up)] after stripping skl2onnx's //--- default ZipMap wrapper at export time (see train_export_onnx.py) - column 1 is //--- the probability of the "up" class prob_up = (double)m_output[0][1]; return(true); }
Tensor type/shape: input must be [1, NUM_FEATURES] float32 — matrix<float> explicitly, since MQL5's bare matrix is matrix<double> and will not satisfy OnnxRun under ONNX_NO_CONVERSION. Output is preallocated as [1, 2] float32, since that mode requires the caller to own the allocation.
A bug worth calling out: skl2onnx's classifier conversion emits two outputs even with zipmap=False. Output 0 is the int64 label (argmax), and output 1 is the probability tensor. MQL5's OnnxRun binds output buffers by positional index, and CDriftGatedModel::Predict above only supplies one output buffer, at index 0 — so without an extra step, the EA would bind its float probability buffer against the int64 label output and fail at inference time, not at compile time.
Export trap fix: the label output is explicitly stripped from the ONNX graph before saving, leaving probabilities as the sole — and therefore index-0 — output, in the export script rather than in MQL5.
def main(): df = load_training_data(CSV_PATH) X = df[FEATURE_COLS].to_numpy(dtype=np.float32) y = df["label"].to_numpy(dtype=np.int64) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.25, shuffle=False # time series - no shuffling across the split ) # diagnostic block: an exact 0.5000 ROC-AUC almost always means the model is # predicting the SAME probability for every row rather than genuinely finding no # signal - these three checks distinguish "no usable signal" from "something upstream # is degenerate" (a single-class training split, a constant feature, a failed fit) print(f"Rows: total={len(df)} train={len(y_train)} test={len(y_test)}") print(f"Train label balance: {dict(zip(*np.unique(y_train, return_counts=True)))}") print(f"Test label balance: {dict(zip(*np.unique(y_test, return_counts=True)))}") # deliberately simple model by default: this article's focus is the drift-detection # layer, not squeezing out extra accuracy from the directional signal itself. # StandardScaler matters for the logistic branch specifically: f0/f1 (log returns) # sit around 1e-4 while f5 (a volume z-score) sits around 1-3, and under L2 # regularization a coefficient large enough to make a tiny-scale feature matter gets # penalized just as hard as one on a well-scaled feature, so without scaling the model # can end up effectively blind to the smallest-magnitude features. Wrapping scaler + # classifier in a Pipeline also means skl2onnx exports BOTH steps into the same graph, # so OnnxRun() applies identical preprocessing at inference time automatically. if MODEL_TYPE == "logistic": model = Pipeline([ ("scaler", StandardScaler()), ("classifier", LogisticRegression(max_iter=2000, C=1.0)), ]) elif MODEL_TYPE == "random_forest": # trees split on raw feature values, so scaling is a no-op for what the model can # learn - the scaler step is left out here rather than kept as dead weight in the # exported graph. max_depth and min_samples_leaf are kept conservative on purpose: # a diagnostic experiment on ~15-20k noisy financial rows should be biased toward # NOT finding fake patterns (overfitting) rather than toward finding real ones - # an unconstrained forest can memorize noise and report a misleadingly high AUC model = RandomForestClassifier( n_estimators=300, max_depth=5, min_samples_leaf=50, random_state=42, n_jobs=-1, ) else: raise ValueError(f"Unknown MODEL_TYPE: {MODEL_TYPE!r}") model.fit(X_train, y_train) preds = model.predict(X_test) probs = model.predict_proba(X_test)[:, 1] print(f"Holdout accuracy: {accuracy_score(y_test, preds):.4f}") print(f"Holdout ROC-AUC: {roc_auc_score(y_test, probs):.4f}") n_unique_probs = len(np.unique(np.round(probs, 6))) print(f"Unique predicted probabilities on test set: {n_unique_probs}") print(f"Predicted probability range: [{probs.min():.4f}, {probs.max():.4f}]") if MODEL_TYPE == "logistic": clf = model.named_steps["classifier"] print(f"Classifier coefficients (scaled feature space): {np.round(clf.coef_[0], 4).tolist()}") else: importances = np.round(model.feature_importances_, 4).tolist() print(f"Feature importances ({', '.join(FEATURE_COLS)}): {importances}") if n_unique_probs <= 2: print("WARNING: the model is predicting almost the same probability for every row.") print(" This points to a degenerate fit (near-constant feature, extreme class") print(" imbalance, or a convergence problem), not just 'no signal in the data'.") # export with an explicit float32 input type matching the [1, NUM_FEATURES] shape. # OnnxModel.mqh pins at load time via OnnxSetInputShape. For the logistic branch this # converts a Pipeline (scaler + classifier both become graph nodes); for the forest # branch it converts the RandomForestClassifier directly. onnx_model = to_onnx( model, initial_types=[("float_input", FloatTensorType([None, len(FEATURE_COLS)]))], options={id(model): {"zipmap": False}}, # skl2onnx wraps classifier outputs in a ZipMap (label -> {class: prob} dict) by # default; MQL5's OnnxRun cannot consume that structure under ONNX_NO_CONVERSION, # so zipmap=False forces a plain [N, 2] float32 probability tensor instead ) # skl2onnx's classifier conversion emits TWO graph outputs even with zipmap=False: # "label" (the argmax class, int64) and "probabilities" (the [N,2] float tensor we # actually want). MQL5's OnnxRun binds output buffers by POSITIONAL index, and # OnnxModel.mqh only supplies one output buffer at index 0 - so the unwanted "label" # output must be removed from the graph here, leaving "probabilities" as the sole # (and therefore index-0) output. Without this step the EA would silently bind its # float probability buffer to the int64 label output and fail at OnnxRun. keep_outputs = [o for o in onnx_model.graph.output if o.name == "probabilities"] del onnx_model.graph.output[:] onnx_model.graph.output.extend(keep_outputs) with open(ONNX_OUT, "wb") as f: f.write(onnx_model.SerializeToString()) print(f"Wrote {ONNX_OUT}") print(f"Input tensor name: float_input, shape [None, {len(FEATURE_COLS)}], float32") print("Output tensor name: probabilities, shape [None, 2], float32 (column 1 = P(up))") if __name__ == "__main__": main()
Verified directly: the naive export showed OUTPUT idx 0 label, OUTPUT idx 1 probabilities under onnx.checker; after the fix, a single OUTPUT idx 0 probabilities, shape [None, 2], and onnx.checker.check_model passes clean.
StandardScaler is essential for logistic regression because feature magnitudes differ by orders of magnitude (f0/f1 around 1e-4 vs. f5 around 1-3, which L2 regularization would otherwise penalize unevenly). Exporting scaler and classifier as a single Pipeline ensures identical preprocessing in Python and MQL5.
Risk-Normalized Position Sizing
Sizing is deliberately decoupled from the drift and inference layers — it only needs a stop distance and an account risk percentage, and it does not care where the stop distance came from.
//+------------------------------------------------------------------+ //| GetAtrStopDistance | //+------------------------------------------------------------------+ double GetAtrStopDistance(const string symbol,const ENUM_TIMEFRAMES tf,const int atr_handle) { double atr_buf[]; ArraySetAsSeries(atr_buf,true); if(CopyBuffer(atr_handle,0,0,1,atr_buf) < 1) { Print("GetAtrStopDistance: CopyBuffer failed, falling back to a fixed point distance"); //--- a hard fallback keeps the EA from placing a zero-distance stop (which most //--- brokers reject outright) if the indicator buffer is briefly unavailable return(200*_Point); } return(atr_buf[0]*RS_ATR_STOP_MULTIPLIER); }
//+------------------------------------------------------------------+ //| CalcLotSize | //+------------------------------------------------------------------+ double CalcLotSize(const string symbol,const double risk_percent,const double stop_distance_price, const ENUM_ORDER_TYPE direction) { double equity = AccountInfoDouble(ACCOUNT_EQUITY); double risk_amount = equity * (risk_percent/100.0); double entry_price = (direction==ORDER_TYPE_BUY) ? SymbolInfoDouble(symbol,SYMBOL_ASK) : SymbolInfoDouble(symbol,SYMBOL_BID); double stop_price = (direction==ORDER_TYPE_BUY) ? entry_price-stop_distance_price : entry_price+stop_distance_price; //--- SYMBOL_TRADE_TICK_VALUE can misreport on some brokers/instruments (notably for //--- non-USD-denominated or synthetic symbols), so the actual monetary loss for a //--- 1.0-lot move is measured directly via OrderCalcProfit instead of derived from //--- tick value arithmetic double loss_per_lot = 0.0; if(!OrderCalcProfit(direction,symbol,1.0,entry_price,stop_price,loss_per_lot)) { Print("CalcLotSize: OrderCalcProfit failed, falling back to broker minimum lot"); return(SymbolInfoDouble(symbol,SYMBOL_VOLUME_MIN)); } loss_per_lot = MathAbs(loss_per_lot); if(loss_per_lot <= 0.0) return(SymbolInfoDouble(symbol,SYMBOL_VOLUME_MIN)); double raw_lots = risk_amount / loss_per_lot; //--- snap to the broker's allowed lot step and enforce min/max so the sized order is //--- actually accepted rather than rejected for an invalid volume double lot_step = SymbolInfoDouble(symbol,SYMBOL_VOLUME_STEP); double lot_min = SymbolInfoDouble(symbol,SYMBOL_VOLUME_MIN); double lot_max = SymbolInfoDouble(symbol,SYMBOL_VOLUME_MAX); double snapped = MathFloor(raw_lots/lot_step)*lot_step; snapped = MathMax(lot_min,MathMin(lot_max,snapped)); return(snapped); }
Monetary loss per lot is measured with OrderCalcProfit rather than SYMBOL_TRADE_TICK_VALUE arithmetic, since tick value can misreport on non-USD or synthetic instruments — OrderCalcProfit asks the trade server directly what a 1.0-lot move from entry to stop would cost. The rest (lot-step snapping, min/max clamping) just keeps the sized order broker-acceptable.
Putting It Together: The EA Control Flow
Everything above is assembled by a small amount of orchestration code in the EA itself.
//+------------------------------------------------------------------+ //| OnInit | //+------------------------------------------------------------------+ int OnInit(void) { //--- LoadManifest always returns true - a missing or unreadable file falls through to //--- built-in defaults internally (see ManifestConfig.mqh) rather than failing OnInit, //--- since this is the expected situation in a Strategy Tester agent's sandboxed Files //--- folder. The call is still made unconditionally because it is what populates //--- g_manifest in the first place. LoadManifest(BuildManifestPath(),g_manifest); if(!g_builder.Init(_Symbol,(ENUM_TIMEFRAMES)Period())) { Print("OnInit: CFeatureBuilder::Init failed"); return(INIT_FAILED); } //--- everything below this point in the ORIGINAL design ran here too: a probe //--- BuildFeaturesAt() call, the 3-way contract check, detector/reference/live-window //--- bootstrap, all inside OnInit. That was too strict - at the exact instant OnInit //--- runs (particularly the very first tick of a Strategy Tester run), the terminal or //--- tester agent may not yet have enough history warmed up for CopyClose/CopyBuffer to //--- satisfy even a 20-30 bar lookback, even though the full history genuinely exists //--- and becomes available moments later. Hard-failing OnInit in that situation aborts //--- the entire run over what is really just a timing race, not a real problem. All of //--- that work is deferred to TryLazyInit(), called from OnTick and retried on every //--- closed bar until it succeeds or LAZY_INIT_MAX_ATTEMPTS is exhausted. if(!g_model.Load()) { Print("OnInit: ONNX model load failed"); return(INIT_FAILED); } g_atr_handle = iATR(_Symbol,(ENUM_TIMEFRAMES)Period(),FB_ATR_PERIOD); if(g_atr_handle==INVALID_HANDLE) { Print("OnInit: ATR handle for RiskSizing failed"); return(INIT_FAILED); } g_last_bar_time = 0; g_bars_since_check = 0; g_trading_halted = false; g_lazy_init_done = false; g_lazy_init_attempts = 0; return(INIT_SUCCEEDED); }
OnInit does deliberately little: load the manifest, create the feature builder, load the ONNX session, create the ATR handle for risk sizing — nothing that depends on having a working window of historical price data lives here anymore. An earlier version of this EA did the three-way contract probe and the full detector/reference-window bootstrap inside OnInit too, and that was too strict: at the exact instant OnInit runs, especially the first tick of a fresh Strategy Tester run, history can momentarily not be warmed up enough to satisfy even a 20-30 bar lookback, even though it genuinely exists and becomes available moments later. Hard-failing the whole run over a timing race rather than a real problem was the wrong trade-off.
//+------------------------------------------------------------------+ //| TryLazyInit | //+------------------------------------------------------------------+ bool TryLazyInit(void) { //--- 3-way feature contract check: build one live feature vector and compare its //--- actual length against both the compiled NUM_FEATURES constant and the manifest's //--- declared num_features before trusting any of the drift-detection machinery below. //--- Returning false here (rather than failing hard) just means "not enough history //--- YET" - OnTick will call this again on the next closed bar. double probe_features[]; if(!g_builder.BuildFeaturesAt(1,probe_features)) return(false); if(!ValidateManifest(g_manifest,ArraySize(probe_features))) { //--- unlike a history-timing gap, a genuine feature-count mismatch will never fix //--- itself by waiting - this is a real configuration error, so it is reported once //--- and then left to keep failing (the caller's attempt counter will eventually stop //--- the retries) rather than silently retrying forever on a problem more history //--- cannot solve Print("TryLazyInit: feature contract validation failed - check manifest num_features"); return(false); } if(!g_detector.Init(g_manifest.num_features,g_manifest.reference_window_size,g_manifest.live_window_size)) return(false); if(!g_detector.LoadOrBootstrapReference(BuildReferenceFilePath(),g_builder)) return(false); g_detector.BootstrapLiveWindow(g_builder); return(true); }
Everything needing price history is deferred to TryLazyInit, called from OnTick and retried every closed bar until it succeeds — the contract probe, detector init, and window bootstrap, now retried rather than attempted once. A genuine configuration error still gets reported and keeps failing every retry, since more history cannot fix a contract mismatch; LAZY_INIT_MAX_ATTEMPTS is the safety valve for when it really is broken rather than just slow.
//+------------------------------------------------------------------+ //| OnTick | //+------------------------------------------------------------------+ void OnTick(void) { //--- everything in this EA - feature construction, drift scoring, inference and sizing - //--- is defined on CLOSED bars, so intrabar ticks are ignored entirely; this also keeps //--- CopyBuffer/CopyClose calls cheap since they only run once per bar close if(!IsNewBar()) return; //--- lazy setup: retried on every closed bar until it succeeds, rather than being a //--- one-shot check inside OnInit - see the comment on TryLazyInit() for why. Once it //--- succeeds it is never called again for the lifetime of this EA instance. if(!g_lazy_init_done) { g_lazy_init_attempts++; if(TryLazyInit()) { g_lazy_init_done = true; if(InpVerboseLogging) PrintFormat("TryLazyInit succeeded after %d closed bar(s)",g_lazy_init_attempts); } else { if(g_lazy_init_attempts >= LAZY_INIT_MAX_ATTEMPTS) Print("OnTick: lazy init still failing after ",LAZY_INIT_MAX_ATTEMPTS, " bars - this is no longer a timing issue, check history/manifest"); return; // not ready yet - try again on the next closed bar } } double features[]; //--- shift=1 is the most recently CLOSED bar; shift=0 would still be forming and would //--- make every feature (and therefore the drift score) non-deterministic mid-bar if(!g_builder.BuildFeaturesAt(1,features)) { Print("OnTick: BuildFeaturesAt failed, skipping this bar"); return; } g_detector.PushLiveSample(features); g_bars_since_check++; if(g_detector.IsReferenceReady() && g_detector.IsLiveWindowFull() && g_bars_since_check >= g_manifest.recompute_every_n_bars) { g_bars_since_check = 0; RunDriftCheck(); } if(!g_trading_halted) ExecuteTradeLogic(features); }
Every closed-bar feature vector, once built, is pushed into the drift detector's live window unconditionally — that bookkeeping happens regardless of whether a drift check or a trade decision fires this bar. The detector is only asked to actually compute scores every recompute_every_n_bars closed bars, which keeps the sort-and-sum cost of eight Wasserstein-1 computations off the hot path most of the time, freeing the manifest's cadence setting to act purely as a freshness knob rather than a performance necessity.
//+------------------------------------------------------------------+ //| RunDriftCheck | //+------------------------------------------------------------------+ void RunDriftCheck(void) { double per_feature_w1[], per_feature_normalized[]; g_detector.ComputeDriftScores(per_feature_w1,per_feature_normalized); double max_stat = 0.0; double composite = g_detector.ComputeComposite(per_feature_normalized,g_manifest.feature_weights,max_stat); if(InpVerboseLogging) PrintFormat("DriftCheck: composite=%.3f max_stat=%.3f (warn=%.2f critical=%.2f)", composite,max_stat,g_manifest.warn_threshold,g_manifest.critical_threshold); //--- the composite AND the max-statistic are both checked against the critical band: //--- either one alone can miss a real regime break (a broad mild shift raises the //--- composite without any single feature standing out; a narrow severe shift in one //--- feature can hide inside a low weighted average) double worst = MathMax(composite,max_stat); if(worst >= g_manifest.critical_threshold) HandleDriftAction(composite,max_stat); else if(worst >= g_manifest.warn_threshold) { Print("RunDriftCheck: WARNING band - composite=",DoubleToString(composite,3), " max_stat=",DoubleToString(max_stat,3)); g_trading_halted = false; } else g_trading_halted = false; }
//+------------------------------------------------------------------+ //| HandleDriftAction | //+------------------------------------------------------------------+ void HandleDriftAction(const double composite,const double max_stat) { switch(g_manifest.action_mode) { case 0: //--- log-only mode: record the breach but let the model keep trading - useful //--- while calibrating thresholds before trusting the gate to halt live entries Print("HandleDriftAction: CRITICAL drift (log-only mode) composite=", DoubleToString(composite,3)," max_stat=",DoubleToString(max_stat,3)); break; case 1: //--- halt mode: stop opening new positions until the next recompute cycle shows //--- the score back under the warning band; existing open positions are left alone //--- deliberately, since flattening on a drift signal is a separate policy decision Print("HandleDriftAction: CRITICAL drift - halting new entries"); g_trading_halted = true; break; case 2: { //--- flag-for-retrain mode: drop a sentinel file the offline retraining pipeline //--- can poll for, rather than trying to trigger retraining from inside MQL5 itself int h = FileOpen("WassersteinDriftGuard\\retrain_flag_"+_Symbol+".txt",FILE_WRITE|FILE_TXT|FILE_ANSI); if(h!=INVALID_HANDLE) { FileWriteString(h,TimeToString(TimeCurrent())+" composite="+DoubleToString(composite,3)+ " max_stat="+DoubleToString(max_stat,3)+"\r\n"); FileClose(h); } Print("HandleDriftAction: CRITICAL drift - retrain flag written"); g_trading_halted = true; break; } default: Print("HandleDriftAction: unknown action_mode ",g_manifest.action_mode," - defaulting to halt"); g_trading_halted = true; break; } }
Composite and max-statistic are both checked against the critical band, either crossing it triggers HandleDriftAction. The three action modes trade aggressiveness for simplicity: log-only for calibration; halt stops new entries but leaves existing positions alone (a separate policy decision); flag-for-retrain drops a sentinel file for an external pipeline to poll, rather than retraining from inside MQL5 itself.
//+------------------------------------------------------------------+ //| ExecuteTradeLogic | //+------------------------------------------------------------------+ void ExecuteTradeLogic(const double &features[]) { if(PositionSelect(_Symbol)) return; // one position at a time - keeps sizing/risk math unambiguous double prob_up; if(!g_model.Predict(features,prob_up)) return; //--- symmetric 0.5 threshold: a fresh, well-calibrated binary classifier centers its //--- decision boundary at 0.5, so no separate "confidence cutoff" input is needed here ENUM_ORDER_TYPE direction; if(prob_up > 0.55) direction = ORDER_TYPE_BUY; else if(prob_up < 0.45) direction = ORDER_TYPE_SELL; else return; // inside the dead zone - no edge, stay flat double stop_distance = GetAtrStopDistance(_Symbol,(ENUM_TIMEFRAMES)Period(),g_atr_handle) + InpFixedStopBuffer; double lots = CalcLotSize(_Symbol,g_manifest.risk_percent,stop_distance,direction); if(lots <= 0.0) return; double price = (direction==ORDER_TYPE_BUY) ? SymbolInfoDouble(_Symbol,SYMBOL_ASK) : SymbolInfoDouble(_Symbol,SYMBOL_BID); double sl = (direction==ORDER_TYPE_BUY) ? price-stop_distance : price+stop_distance; double tp = (direction==ORDER_TYPE_BUY) ? price+stop_distance*1.5 : price-stop_distance*1.5; if(direction==ORDER_TYPE_BUY) g_trade.Buy(lots,_Symbol,price,sl,tp,"WassersteinDriftGuard"); else g_trade.Sell(lots,_Symbol,price,sl,tp,"WassersteinDriftGuard"); }
The 0.55/0.45 dead zone around the model's raw probability output means a coin-flip prediction produces no trade at all rather than a low-conviction one — there is no separate "confidence" input to tune because a freshly calibrated binary classifier's decision boundary already sits at 0.5 by construction.
The Offline Pipeline: Export, Train, Validate
Three pieces run outside MetaTrader entirely, and all three depend on the same CFeatureBuilder class the live EA uses, so there is exactly one implementation of "what a feature vector is" in this whole project.
//+------------------------------------------------------------------+ //| OnStart | //+------------------------------------------------------------------+ void OnStart(void) { CFeatureBuilder builder; if(!builder.Init(_Symbol,(ENUM_TIMEFRAMES)Period())) { Print("ExportFeatureHistory: CFeatureBuilder::Init failed"); return; } int handle = FileOpen(InpOutputFileName,FILE_WRITE|FILE_TXT|FILE_ANSI); if(handle==INVALID_HANDLE) { Print("ExportFeatureHistory: could not create ",InpOutputFileName); return; } //--- header row documents the exact CSV column layout the Python trainer expects: //--- f0..f7 are the 8 features in the same order BuildFeaturesAt() produces them, //--- label is 1 if price rose over InpLabelHorizon bars, else 0 string header = "f0,f1,f2,f3,f4,f5,f6,f7,label"; FileWriteString(handle,header+"\r\n"); double feat[]; int written = 0; //--- start at InpLabelHorizon+1 so every row has both a full feature lookback behind it //--- AND a full forward window ahead of it to compute the label from - starting any //--- closer to shift=0 would need unclosed/future bars that do not exist yet for(int shift=InpLabelHorizon+1; shift<InpBarsToExport+InpLabelHorizon+1; shift++) { if(!builder.BuildFeaturesAt(shift,feat)) break; // ran out of history double close_now = iClose(_Symbol,(ENUM_TIMEFRAMES)Period(),shift); double close_future = iClose(_Symbol,(ENUM_TIMEFRAMES)Period(),shift-InpLabelHorizon); int label = (close_future > close_now) ? 1 : 0; string row = ""; for(int i=0;i<NUM_FEATURES;i++) row += DoubleToString(feat[i],8) + ","; row += IntegerToString(label); FileWriteString(handle,row+"\r\n"); written++; } FileClose(handle); builder.Deinit(); PrintFormat("ExportFeatureHistory: wrote %d rows to %s",written,InpOutputFileName); }
The exported CSV has a documented, fixed column layout: f0,f1,f2,f3,f4,f5,f6,f7,label, one row per historical bar, where label is 1 if price closed higher InpLabelHorizon bars later and 0 otherwise. The loop starts at InpLabelHorizon + 1 bars back. Each row needs a full lookback for features and a full forward window to compute the label; starting closer would require future bars.
def load_training_data(path: str) -> pd.DataFrame: # ExportFeatureHistory.mq5 opens its output file with FILE_ANSI explicitly # (FileOpen(..., FILE_WRITE|FILE_TXT|FILE_ANSI)), which forces MQL5 to write # plain single-byte text instead of its UTF-16 default - that flag is a # deliberate choice here, not an oversight, precisely so this file does NOT # need UTF-16 decoding on the Python side. Since every value in this CSV is # plain ASCII (digits, '.', ',', and the fixed column names), UTF-8 decodes # ANSI-written ASCII bytes identically - reading as "utf-16" here would (and # did, if you hit this before the fix) raise UnicodeDecodeError immediately, # because ANSI text has no UTF-16 byte-order-mark for pandas to detect. df = pd.read_csv(path, encoding="utf-8", sep=None, engine="python") # the delimiter MQL5 used follows the Windows regional setting at export time # (comma vs semicolon) - sep=None with the python engine lets pandas sniff it # instead of hard-coding an assumption that breaks on a different machine # defensive only: an ANSI-written file has no BOM, so this is normally a no-op, # but it costs nothing to keep in case the export path ever changes back to a # UTF-16 FileOpen mode df.columns = [c.replace("\ufeff", "") for c in df.columns] return df
Two quirks follow from the file being written by MQL5, not a general-purpose CSV writer. ExportFeatureHistory.mq5 uses FILE_ANSI deliberately, forcing plain single-byte text instead of MQL5's UTF-16 default — so encoding="utf-8" reads it correctly, while "utf-16" raises UnicodeDecodeError since ANSI text carries no byte-order-mark. The delimiter follows the exporting machine's regional settings, so sep=None, engine="python" lets pandas sniff it rather than assuming.
Training itself is a plain logistic regression — the point of this article is the drift layer, not squeezing extra accuracy out of the directional model — followed by the export and output-stripping shown above. Once WassersteinModel.onnx exists, it gets copied into MQL5\Files\WassersteinDriftGuard\ and the EA is recompiled so the #resource directive picks up the new weights.
def wasserstein_1d(sample_a: np.ndarray, sample_b: np.ndarray) -> float: # mirrors CDriftDetector::Wasserstein1D() exactly: sort both equal-size samples, # then take the mean absolute difference between paired order statistics a_sorted = np.sort(sample_a) b_sorted = np.sort(sample_b) return float(np.mean(np.abs(a_sorted - b_sorted)))
def naive_mean_std_score(reference: np.ndarray, live: np.ndarray) -> float: # the common alternative this script benchmarks against: how many reference # standard deviations has the live mean moved, ignoring shape entirely ref_std = reference.std() if ref_std <= 0: ref_std = 1e-8 return float(abs(live.mean() - reference.mean()) / ref_std)
def make_shift(kind: str) -> np.ndarray: if kind == "none": return RNG.normal(loc=0.0, scale=1.0, size=WINDOW_SIZE) if kind == "mean_shift": return RNG.normal(loc=0.6, scale=1.0, size=WINDOW_SIZE) if kind == "variance_shift": return RNG.normal(loc=0.0, scale=1.8, size=WINDOW_SIZE) if kind == "bimodal_mixture": # construct a mixture with mean ~0 and variance ~1 by design, so a naive # mean/std detector sees almost nothing - the shift is purely in SHAPE half = WINDOW_SIZE // 2 lobe_a = RNG.normal(loc=-1.4, scale=0.55, size=half) lobe_b = RNG.normal(loc=1.4, scale=0.55, size=WINDOW_SIZE - half) mixture = np.concatenate([lobe_a, lobe_b]) RNG.shuffle(mixture) return mixture raise ValueError(kind)
def calibrate_thresholds(): # calibrate both detectors' thresholds off the null ("no shift") distribution so # each is compared at roughly the SAME false-positive rate (~5%) - otherwise a # "more sensitive" detector could just be a looser one _, _, w1_null, naive_null = run_trials("none", threshold_w1=1e9, threshold_naive=1e9) threshold_w1 = float(np.percentile(w1_null, 95)) threshold_naive = float(np.percentile(naive_null, 95)) return threshold_w1, threshold_naive
The bimodal mixture case is constructed on purpose to defeat a moment-based detector: two Gaussian lobes at +/-1.4 with matching spread, mixed and shuffled so the combined sample's mean sits near zero and its variance sits near one — almost indistinguishable from the reference by the numbers a naive detector looks at, and completely different in shape. Thresholds for both detectors are calibrated off the null (no-shift) distribution's 95th percentile before any of the shift trials run, so the comparison in Fig. 2 is apples-to-apples at matched false-positive rates rather than one detector simply being set to fire more often.
Edge Cases and Pitfalls
Equal window sizes are mandatory. The discrete 1D W1 formula assumes equal-length sorted samples; CDriftDetector::Init rejects a reference_window_size/live_window_size mismatch at initialization rather than silently computing a meaningless number.
ATR can be zero. A flat opening bar can zero out ATR; every ATR-dividing feature (f3, f6, f7) and the ATR-derived stop clamp the denominator to a tiny epsilon to avoid a silent inf poisoning the window.
Tester file sandboxing matters, in two places. The reference window bootstraps fresh per run (correct backtest behavior, not a workaround). The manifest needs #property tester_file "WassersteinDriftGuard\\drift_manifest.json" at the top of the EA, or a Tester run fails immediately with LoadManifest: could not open ... even though the same file exists in the live terminal.
CopyTickVolume needs a long[] destination. A double[] fails silently; FeatureBuilder.mqh declares vol_buf as long[] and casts to double only at the point of use.
This is a feature-space detector only. It is not a full explanation of why a model's predictions might be degrading — it cannot catch target drift where features stay in-distribution but the true relationship shifts underneath them. Pairing it with the ADWIN/Page-Hinkley sentinel from the earlier article, which watches the model's own error signal directly, covers more of the failure surface than either alone.
The detector is univariate, not multivariate. Each feature is monitored independently; a shift visible only in the correlation between two features — both individually staying in-distribution while their joint relationship changes — will not trip it. Catching that needs a multivariate optimal-transport formulation (Sinkhorn-regularized, since exact multivariate OT lacks this article's closed form), left as a deliberate scope boundary here, not an oversight.
The manifest is read once, at OnInit. Editing drift_manifest.json mid-run has no effect until the EA is re-attached — a deliberate simplicity trade-off, since hot-reload would mean re-validating the feature contract and reconciling a changed window size against the frozen reference.
OnInit should not hard-fail on a history-warmup timing race. A Tester run can abort immediately with insufficient history loaded even on a symbol with years of data, since history can momentarily lag at the exact instant OnInit runs. See TryLazyInit in the EA control flow section.
Action modes are conservative about existing exposure. Halt and flag-for-retrain both leave open positions alone — flattening on a drift signal is a separate risk-management decision this EA does not make for you, though it is a small addition to HandleDriftAction if you want it.
Testing in the Strategy Tester
The comparison that actually validates this architecture is a baseline run (the ONNX model trading unmonitored, action_mode = 0) against a drift-gated run (the same model, action_mode = 1, halting new entries whenever the composite or max-statistic crosses the critical threshold) over the identical historical window.
| Setting | Value |
|---|---|
| Symbol / Timeframe | XAUUSD / M5 |
| Test period | 2026.01.01–2026.08.25 (8 months, 100% history quality) |
| reference_window_size / live_window_size | 250 / 250 |
| recompute_every_n_bars | 10 |
| warn_threshold / critical_threshold | 1.5 / 2.5 |
| risk_percent | 0.5 |
| Modeling | Every tick based on real ticks |
Both runs used the same random-forest model from the feature-design section above — this tests whether the gate reduces harm, not whether it improves alpha.
| Metric | Baseline | Drift-gated |
|---|---|---|
| Total Trades | 2,983 | 2,749 |
| Net Profit | -$3,892.17 | -$3,429.13 |
| Profit Factor | 0.95 | 0.95 |
| Max Balance Drawdown | 55.57% | 54.66% |
| Expected Payoff / trade | -$1.30 | -$1.25 |
| Sharpe Ratio | -2.01 | -1.82 |
| Win Rate | 38.65% | 38.92% |
Baseline vs. Drift-Gated Strategy Tester Results

Fig. 4. Actual Strategy Tester output over the same 8-month XAUUSD M5 window: net profit and max balance drawdown, baseline vs. drift-gated.
The gate did not create profitability — profit factor remained 0.95 in both runs — but it reduced trading frequency (234 fewer trades), net loss (11.9% smaller), drawdown, and Sharpe deterioration. This is consistent with a drift gate limiting exposure rather than manufacturing alpha.
Conclusion
This article presented a practical feature-space drift detector for live MQL5 systems based on one-dimensional Wasserstein-1 distance. Because the 1D case reduces to sorting and rank-wise pairing, the method is cheap enough to run per feature inside an EA while remaining sensitive to shape-only distribution shifts that mean/std monitors miss — proven synthetically at 100% vs. 1% detection on the bimodal mixture shift.
The implementation adds a manifest-driven control layer around an ONNX model, with persistent reference windows, startup-safe lazy initialization, and configurable actions. Strategy Tester results showed modest harm reduction, not alpha creation — exactly what should be expected given the weak directional model reported honestly throughout this article rather than glossed over.
The main limitation is scope: this is a univariate feature-space detector, not a multivariate or target-drift solution. Still, the central lesson is simple: unchanged averages do not imply unchanged distributions.
| File | Type | Description |
|---|---|---|
| WassersteinDriftGuard_EA.mq5 | Expert Advisor | Main EA: wires the feature builder, drift detector, ONNX model and risk sizer together; OnTick control flow. |
| FeatureBuilder.mqh | Include | Builds the fixed 8-dimensional feature vector shared by the EA, the reference-window bootstrap and the CSV exporter. |
| WassersteinDetector.mqh | Include | Per-feature Wasserstein-1 drift detector: reference/live window management, composite scoring. |
| ManifestConfig.mqh | Include | Minimal JSON manifest loader and the three-way feature-contract validation. |
| RiskSizing.mqh | Include | ATR-derived stop distance and OrderCalcProfit-based lot sizing. |
| OnnxModel.mqh | Include | #resource-embedded ONNX session wrapper for the directional signal. |
| ExportFeatureHistory.mq5 | Script | Exports historical feature vectors + forward-return labels to a training CSV. |
| drift_manifest.json | Config | Runtime-tunable thresholds, weights and action mode read at OnInit. |
| WassersteinModel.onnx | Model | Trained/exported ONNX classifier embedded into the EA at compile time (placeholder weights — retrain on your own exported data). |
| train_export_onnx.py | Python | Trains the logistic regression classifier and exports it to ONNX with the label-output stripped. |
| synthetic_shift_validation.py | Python | Offline validation comparing Wasserstein-1 vs. naive mean/std detection on synthetic distribution shifts; produces Fig. 2. |
Warning: All rights to these materials are reserved by MetaQuotes Ltd. Copying or reprinting of these materials in whole or in part is prohibited.
This article was written by a user of the site and reflects their personal views. MetaQuotes Ltd is not responsible for the accuracy of the information presented, nor for any consequences resulting from the use of the solutions, strategies or recommendations described.
Working with ONNX Models in MQL5 (Part 2): Drawing the Model Graph on an Interactive Chart Panel
Position Management: Deriving a Self-Calibrating Exit Ladder From Historical MFE in MQL5
Competitive Swarm Optimizer (CSO)
Rough Volatility: Building a Roughness Index Feature from the RFSV Model for ML Trade Filtering
- Free trading apps
- Over 8,000 signals for copying
- Economic news for exploring financial markets
You agree to website policy and terms of use