preview
Certified Robustness Radius: Knowing How Much Noise Your ONNX Trading Model Can Survive

Certified Robustness Radius: Knowing How Much Noise Your ONNX Trading Model Can Survive

MetaTrader 5 — Machine learning |
220 0
Adedayo David Gbadebo
Adedayo David Gbadebo

Introduction

Every ONNX-driven trading system I've built eventually runs into the same silent problem: the model outputs a class, you trust the class, and you never ask how close it was to being a different class entirely. A softmax score of 0.51 for "long" gets treated the same as 0.98 — both open the same position at the same size. But one signal is a coin flip from flipping to "short," and the other is a genuinely confident read. Softmax confidence isn't calibrated as a distance to the decision boundary, so it doesn't reliably tell you which situation you're in.

This article builds a native MQL5 system that answers a different question: not "what did the model predict," but "how much could the input change before the prediction would flip." That's a certified robustness radius, from a technique originally built to defend image classifiers against adversarial attacks — randomized smoothing — repurposed here for a more mundane threat: noisy live feeds, tick jitter, spread spikes, and stale-bar artifacts that can flip a fragile prediction.

By the end you'll have a working EA that runs every trade decision through certification before acting, sizes positions by how robust the signal is, and logs a rolling radius history so you can review a model's confidence margins over time.

For traders, the payoff is a new input independent of confidence score: identical softmax outputs can have wildly different certified radii, and it's the radius that tells you which prediction is statistically robust to noise. That's narrower than "safe to trade" — robustness says nothing about whether the direction is correct. For programmers, it's a self-contained native pattern using only MQL5's standard library — no external DLLs, no Python round-trips, no ALGLIB. The takeaway for both: raw model output conflates confidence, robustness, and correctness. Randomized smoothing measures one of those — robustness — directly, instead of reading it off an uncalibrated probability.

A few situations where the certified radius changes what a careful trader would do. The first two below describe the mechanism as designed, not a separately tested claim — no dedicated news-release or feed-gap test was run here, only the drift-test and equity-comparison results described later:

Around high-impact news releases. Spreads widen and ticks arrive erratically right after an NFP or CPI print. A raw-signal EA has no way to tell a genuine breakout from noise — it just trades whatever class comes out. The certified radius typically collapses in that window, so the gate skips or downsizes the trade, without ever being told an economic release is happening.

After a feed gap or reconnect. A stale or partially-filled bar produces a feature vector that's quietly off-distribution — nothing crashes or errors, the numbers just don't reflect real market structure. A raw-signal EA classifies and trades that vector anyway. The certifier is more likely to abstain here, since noised passes around it disagree — a sign of local instability, not proof the trade would have been bad, but a reasonable hypothesis for why abstains cluster around this kind of event.

Filtering out genuinely unreliable predictions. In a drift test on XAUUSD H1 across 2,996 bars, the certifier abstained on 1,210 as too fragile; on the 1,786 it committed to, accuracy came out to 50.8%. That's conditional accuracy on a self-selected subset, not overall accuracy. The one-in-three comparison against blind guessing also assumes balanced classes — the actual split isn't reported, so treat it as a loose sanity check, not a rigorous baseline. What the test shows cleanly: accuracy on engaged bars was meaningfully above chance for a 3-class problem — not a claim about the model overall or about profitability.

Comparing signal quality across instruments. Running the same certifier against EURUSD H1 and XAUUSD M15 side by side lets you allocate size by which symbol's current prediction the model can actually stand behind — a capital-allocation input a plain classification score can't give you.

Scope note: this is not a proof against a malicious adversary crafting worst-case perturbations. It's a practical robustness gate against the ordinary noise every live feed produces — feed gaps, requote jitter, and volatility bursts — using the same mathematical machinery that formal adversarial robustness research relies on.


Contents

  1. Why a Point Prediction Isn't Enough
  2. Randomized Smoothing and the Certified Radius
  3. Architecture Walkthrough
  4. Trade Gating and Position Sizing
  5. Edge Cases and Pitfalls
  6. The Monitor Indicator
  7. Training, Export, and Validation Tooling
  8. Testing in the Strategy Tester
  9. Conclusion


Why a Point Prediction Isn't Enough

Think about what actually happens between a tick arriving and an ONNX model producing a class label. The feature engine reads indicator buffers, computes normalized returns, volatility ratios, and z-scores, packs them into a fixed-width vector, and hands that vector to the model. Every step in that chain is a source of small, unavoidable noise: a requote shifts the spread feature by a tick, a late price update changes the volume z-score slightly, a broker feed hiccup rounds a bar differently than it would a second later.

None of this noise is adversarial. Nobody is attacking your feature pipeline. But the effect on a fragile prediction is the same as if someone were: a feature vector that sits close to the model's decision boundary can flip its predicted class from a perturbation that's well within the normal noise floor of the instrument. The model has no way of telling you this. Softmax probabilities are not a distance to the decision boundary — a model can be badly miscalibrated and still emit a confident-looking 0.85 for a prediction that flips under noise a fraction of the size of one bar's normal range.

What we actually want is a geometric answer: given the current feature vector x, how large a perturbation ball around x can we guarantee the predicted class won't change? If that radius is large relative to the instrument's typical feature-space noise, the signal is robust and worth acting on at full size. If it's small, the signal is fragile — mathematically, a puff of ordinary feed noise could have produced the opposite prediction, and sizing into it at full confidence is a bet on noise, not on the model.

Fragile vs. robust prediction near a decision boundary

Fig. 1. Two feature vectors both classified "long," but one sits a hair from the boundary and the other sits well inside the long region — the certified radius is what tells them apart.


Randomized Smoothing and the Certified Radius

Randomized smoothing sidesteps the need to analyze the model's internals directly — which matters a lot here, since the base classifier is an opaque ONNX graph exported from scikit-learn, and we have no gradient access or architecture-specific certification machinery available natively in MQL5. Instead of certifying the base classifier f directly, we certify a smoothed version of it, g, defined as:

g(x) = argmax_c P( f(x + ε) = c ), where ε ~ N( 0 , σ²I)

In plain terms: instead of asking the model once, we ask it many times, each time with independent Gaussian noise added to the feature vector at scale σ, and we let the class that wins the most votes be the smoothed prediction. This is exactly what CRobustnessCertifier::Certify() does in the code — it runs N noised inference passes and tallies how often each class wins.

The interesting part isn't the voting itself, it's what the vote proportions let us prove. Cohen, Rosenfeld, and Kolter's 2019 result gives a closed-form certified radius based on a statistical lower bound on the winning class's true win probability under the noise distribution. We don't get to observe that true probability directly — we only ran a finite number of noisy trials — so we compute a Clopper-Pearson lower confidence bound, p_A, on it:

p_A = Beta^-1( α; n_A, N - n_A + 1 )

where n_A is the vote count for the winning class out of N total passes, and α is our significance level (0.001 by default — a 99.9% one-sided confidence bound). This is implemented natively with MQL5's MathQuantileBeta() from the standard statistics library, no external dependency required. Once we have p_A, the certified radius follows directly:

R = σ · Φ^-1( p_A )

where Φ-1 is the inverse standard normal CDF, computed with MathQuantileNormal(). If p_A ≤ 0.5, the winning class didn't clear a bare statistical majority under noise, and the certifier abstains rather than report a meaningless near-zero radius. Precisely what this distinguishes: "confident under noise" from "not confident" — nothing checks whether the winning class was correct. A model can be confidently and consistently wrong and still certify with a large radius every time; that robustness-vs-correctness distinction is the main limitation to keep in mind throughout.

One wrinkle: this formula is the binary simplification of the general multi-class bound, comparing the winner against a fixed 0.5 threshold rather than the specific runner-up. It's more conservative than the tightest possible bound but simpler to implement correctly — and conservative is the right direction to err on when real money sizes off the number.

Two more points worth stating plainly. First, what gets certified is the smoothed classifier g, not the ONNX model f directly — "the prediction won't change" is loose shorthand for "the smoothed vote g(x) won't change." Second, the standard scheme (Cohen et al.) uses two independent sample sets — one to select the class, a separate one for the confidence bound — to keep the steps independent. This implementation reuses the same N passes for both, cheaper but meaning the stated confidence level is a looser approximation of the textbook guarantee, not the guarantee itself.

//--- Clopper-Pearson lower bound + normal inverse -> certified radius
int err_code = 0;
double lower = MathQuantileBeta(alpha, (double)nA, (double)(n_passes - nA + 1), err_code);

if(lower <= 0.5)
  {
   predicted_class = CERT_CLASS_ABSTAIN;   // no certifiable margin at this sigma/N/alpha
   return true;
  }

int qerr = 0;
double z = MathQuantileNormal(lower, 0.0, 1.0, qerr);
radius = sigma * z;


Architecture Walkthrough

The system splits cleanly into three responsibilities, mirrored by three include files. CFeatureEngine owns indicator handles and builds the fixed 10-wide feature vector every bar. CRobustnessCertifier owns the ONNX session and the smoothing/certification math. The EA's OnTick() just wires the two together and translates a certified prediction into a trade decision.

The feature contract is the most important part to get right, and the easiest to break silently. CERT_FEATURE_COUNT, FeatureEngine::Build()'s column count, and the ONNX model's declared input width must all agree exactly — a mismatch doesn't crash, it produces plausible-looking predictions from misaligned features, which is worse. RobustnessCertifier::Init() enforces this by binding shapes explicitly and failing loudly on mismatch, and the Python export script runs the same check before writing the .onnx file, so drift gets caught at training time, not on a live account.

Data flow from tick to certified trade decision

Fig. 2. Feature engine builds the vector once per closed bar; the certifier runs N noised passes through the same ONNX session before the EA acts on the result.

Inside Certify(), each of the N passes draws independent Gaussian noise via MathRandomNormal(), adds it to the base feature vector, and runs one OnnxRun() call — every pass is a full inference, so N and model complexity set your per-bar compute budget directly. For the lightweight MLP here, 100 passes on an M15 bar comfortably fits inside tick-to-tick timing, but measure it on your own hardware; the testing section includes explicit timing numbers.

for(int pass = 0; pass < n_passes; pass++)
  {
   int cls = InferOnce(features, sigma);
   if(cls >= 0 && cls < (int)m_num_classes)
      counts[cls]++;
  }

Two non-obvious details in InferOnce(). First, the ONNX input tensor is float32 (4 bytes/element), but MQL5's bare matrix type is matrix<double> (8 bytes) — feeding that into OnnxRun() with ONNX_NO_CONVERSION silently mismatches the input size; the fix is declaring matrix<float> for both input and output. Second, ONNX_NO_CONVERSION requires the output matrix pre-sized to the model's output shape — it won't auto-allocate. Both fail with unhelpful low-level errors if you don't know to look for them.

CertDefines.mqh is the smallest file in the project but arguably the most load-bearing one, since it's where the feature contract actually lives as a compile-time constant rather than a magic number scattered across files:

//+------------------------------------------------------------------+
//| Feature contract constants                                       |
//+------------------------------------------------------------------+
#define CERT_FEATURE_COUNT   10    // Width of the feature vector fed to ONNX
#define CERT_NUM_CLASSES     3    // short / flat / long

Every other file includes this one and builds against these two constants — the EA, the indicator, every diagnostic script, and the Python training script all assert against CERT_FEATURE_COUNT/NUM_CLASSES rather than hardcoding 10 and 3 independently. That's deliberate: if you ever add or remove a feature, there's exactly one place to change it, and every consumer either picks up the new value automatically or fails loudly at init time instead of quietly computing garbage.

FeatureEngine::Build() is where the 10 features actually get computed, once per closed bar. It's mostly indicator-buffer plumbing, but a couple of the feature choices are worth walking through directly rather than leaving as a bulleted list:

//--- 4: MACD_hist_norm -- histogram normalized by ATR, not by price,
//--- so the feature scale stays comparable across low- and high-priced
//--- instruments (a raw MACD histogram value means something very
//--- different on EURUSD than on XAUUSD; dividing by ATR fixes that).
double hist = macd_main[0] - macd_sig[0];
features[4] = hist / MathMax(atr14, CERT_EPS);

//--- 6: Volume_zscore -- how unusual is this bar's volume relative
//--- to its own recent history, rather than an absolute volume number
//--- that would need per-instrument calibration to mean anything.
double mean_v = 0.0, sd_v = 0.0;
for(int i = 0; i < 20; i++) mean_v += (double)vol[i];
mean_v /= 20.0;
for(int i = 0; i < 20; i++) sd_v += MathPow((double)vol[i] - mean_v, 2);
sd_v = MathSqrt(sd_v / 20.0);
features[6] = ((double)vol[19] - mean_v) / MathMax(sd_v, CERT_EPS);

Every feature follows this same pattern — divided or z-scored against something local, never an absolute number. That matters for randomized smoothing: noise is added in this normalized space at a fixed σ, so if features weren't roughly comparable in scale, the same σ would mean different things for different features, and the certified radius would stop being a meaningful geometric quantity.

Finally, the EA's own control flow — OnInit() does the one-time setup, and OnTick() is where the actual per-bar decision happens, gated on a single new-closed-bar check so the certification cost (N inference passes) is paid once per bar, not once per tick. OnInit() is worth showing in full because it's where the feature-contract check actually runs, inside RobustnessCertifier::Init():

//+------------------------------------------------------------------+
//| Init                                                             |
//+------------------------------------------------------------------+
bool Init(const string model_path, const ulong feature_count, const ulong num_classes)
  {
   m_feature_count = feature_count;
   m_num_classes   = num_classes;

   m_onnx_handle = OnnxCreate(model_path, ONNX_DEFAULT);
   if(m_onnx_handle == INVALID_HANDLE)
    {
      PrintFormat("OnnxCreate failed for '%s', err=%d", model_path, GetLastError());
      return false;
    }

//--- Consistency check: declared ONNX input shape vs feature contract
   if(!OnnxGetInputCount(m_onnx_handle))
    {
      Print("RobustnessCertifier: model reports zero inputs.");
      Release();
      return false;
    }

//--- Bind the model's declared input/output shapes explicitly. If the
//--- model's actual input width doesn't match feature_count, or its
//--- output width doesn't match num_classes, this fails HERE, loudly,
//--- at startup -- instead of silently computing on misaligned data
//--- once the EA is already live.
   const long in_shape_expected[]  = {1, (long)feature_count};
   const long out_shape_expected[] = {1, (long)num_classes};
   if(!OnnxSetInputShape(m_onnx_handle, 0, in_shape_expected) ||
     !OnnxSetOutputShape(m_onnx_handle, 0, out_shape_expected))
    {
      PrintFormat("RobustnessCertifier: shape binding failed -- model input/output width does not match CERT_FEATURE_COUNT=%d / CERT_NUM_CLASSES=%d. err=%d",
              feature_count, num_classes, GetLastError());
      Release();
      return false;
    }

   m_ready = true;
   return true;
  }

The EA's own OnInit() calls this, plus sets up the rolling radius history buffer used for the gate's baseline, and opens the CSV log:

//+------------------------------------------------------------------+
//| Expert initialization function                                   |
//+------------------------------------------------------------------+
int OnInit()
  {
   if(!g_features.Init(_Symbol, _Period)) return INIT_FAILED;

   if(!g_certifier.Init(InpModelFileName, CERT_FEATURE_COUNT, CERT_NUM_CLASSES))
    {
      Print("certifier init / feature-contract validation failed.");
      return INIT_FAILED;
    }

   ArrayResize(g_radius_history, InpBaselineLookback);
   ArrayInitialize(g_radius_history, -1.0); // -1 = not yet populated
   g_radius_ptr = 0;

   g_trade.SetExpertMagicNumber(InpMagicNumber);

   if(InpLogToFile)
    {
      g_file_handle = FileOpen("CertifiedONNXRadius\\cert_log.csv", FILE_WRITE | FILE_CSV | FILE_COMMON);
      if(g_file_handle != INVALID_HANDLE)
        FileWrite(g_file_handle, "time", "predicted_class", "pA_lower", "radius", "baseline", "gate_ratio", "lot", "equity", "action");
    }

   return INIT_SUCCEEDED;
  }

The rolling baseline itself is a plain trailing median over the last InpBaselineLookback certified radii, skipping abstained bars (stored as -1 sentinels) rather than treating them as zero:

//+------------------------------------------------------------------+
//| ComputeBaseline                                                  |
//+------------------------------------------------------------------+
double ComputeBaseline(void)
  {
   double valid[];
   int n = 0;
   for(int i = 0; i < InpBaselineLookback; i++)
      if(g_radius_history[i] >= 0.0) n++;

   if(n < 5) return -1.0; // not enough history yet -- gate stays closed

   ArrayResize(valid, n);
   int j = 0;
   for(int i = 0; i < InpBaselineLookback; i++)
      if(g_radius_history[i] >= 0.0) valid[j++] = g_radius_history[i];

   ArraySort(valid);
   return (n % 2 == 1) ? valid[n/2] : (valid[n/2 - 1] + valid[n/2]) / 2.0;
  }

And OnTick() itself — this is the actual, complete decision path, not a trimmed excerpt. Everything the Trade Gating section below describes in prose is the second half of this function:

//+------------------------------------------------------------------+
//| Expert tick function                                             |
//+------------------------------------------------------------------+
void OnTick()
  {
   datetime cur_bar_time = iTime(_Symbol, _Period, 0);
   if(cur_bar_time == g_last_bar_time) return; // once per closed bar
   g_last_bar_time = cur_bar_time;

   double features[];
   if(!g_features.Build(1, features)) return; // shift=1: last fully closed bar

   int predicted_class = CERT_CLASS_ABSTAIN;
   double radius = 0.0, pA_lower = 0.0;
   if(!g_certifier.Certify(features, InpSigma, InpNPasses, InpAlpha, predicted_class, radius, pA_lower))
    {
      Print("Certify() failed this bar -- skipping, pointer still advances.");
    }

   g_radius_history[g_radius_ptr] = (predicted_class == CERT_CLASS_ABSTAIN) ? -1.0 : radius;
   g_radius_ptr = (g_radius_ptr + 1) % InpBaselineLookback;

   double baseline = ComputeBaseline();
   double gate_ratio = (baseline > 0.0) ? radius / baseline : 0.0;

   string action = "SKIP_ABSTAIN";
   double lot = 0.0;

   bool has_position = PositionSelect(_Symbol);

   if(predicted_class != CERT_CLASS_ABSTAIN && baseline > 0.0)
    {
      if(gate_ratio < InpMinRadiusMultiple)
        {
          action = "SKIP_LOW_RADIUS";
          if(has_position)
            {
              g_trade.PositionClose(_Symbol);
            action = "CLOSE_LOW_RADIUS";
          }
        }
      else
        {
          double lot_multiple = MathMin(gate_ratio, InpMaxLotMultiple);
          lot = NormalizeDouble(InpBaseLots * lot_multiple, 2);

          if(predicted_class == CERT_CLASS_LONG)
            {
              if(has_position && PositionGetInteger(POSITION_TYPE) == POSITION_TYPE_SELL)
                g_trade.PositionClose(_Symbol);
            if(!PositionSelect(_Symbol))
              {
                g_trade.Buy(lot, _Symbol);
              action = "OPEN_LONG";
            }
          }
          else if(predicted_class == CERT_CLASS_SHORT)
            {
              if(has_position && PositionGetInteger(POSITION_TYPE) == POSITION_TYPE_BUY)
                g_trade.PositionClose(_Symbol);
            if(!PositionSelect(_Symbol))
              {
                g_trade.Sell(lot, _Symbol);
              action = "OPEN_SHORT";
            }
          }
          else // FLAT
          {
            if(has_position)
            {
              g_trade.PositionClose(_Symbol);
            action = "CLOSE_FLAT_SIGNAL";
            }
          else
            action = "STAY_FLAT";
        }
    }
  }

   if(InpLogToFile && g_file_handle != INVALID_HANDLE)
    FileWrite(g_file_handle, TimeToString(cur_bar_time), predicted_class, pA_lower, radius, baseline, gate_ratio, lot,
        AccountInfoDouble(ACCOUNT_EQUITY), action);
  }

Note the shift=1 in Build(), not shift=0 — the EA always certifies the last fully closed bar. Within a single Certify() call the base feature vector is built once and held fixed; only the noise changes across passes. The real reason to avoid shift=0: an unclosed bar's price keeps updating, so consecutive OnTick() calls on the same still-forming bar would build different feature vectors, making "the same bar's" certification inconsistent across calls. The direction-flip logic also always closes an opposing position before opening a new one, matching how EquityCompareBacktest.mq5's simulation (covered later) accounts for direction changes.

Three honest gaps in the OnTick() above. First, sizing by gate_ratio only happens on a new position open — an already-open position in the same direction isn't resized against a fresh radius. Second, PositionSelect(_Symbol) and PositionClose(_Symbol) aren't filtered by magic number, so on a symbol shared with another EA this logic can touch positions it didn't open — add a magic-number check before acting on any position found this way. Third, Certify()'s return value is checked and logged on failure, but doesn't change downstream behavior — predicted_class stays CERT_CLASS_ABSTAIN either way, so an internal failure and a genuine abstain land on the same path; a distinct fail state would be a meaningful improvement. None of these break the certification math, but all three are worth fixing before live money.


Trade Gating and Position Sizing

The full OnTick() code above already shows how this works — this section covers the reasoning behind the two constants that drive it. The rolling baseline is adaptive rather than hardcoded: what counts as "robust" differs by instrument, so the EA learns its own recent normal from a rolling window. InpMinRadiusMultiple (0.75 default) sets how far below baseline a radius must fall before the gate calls it too fragile; InpMaxLotMultiple caps how far above baseline a radius can push position size.

Radius gating decision flow

Fig. 3. The gate compares this bar's certified radius against the rolling baseline before any trade action is taken.

Gate logic in one line: no certified class → abstain and hold. Certified but radius below 0.75× baseline → flat, close any open position. Certified and radius at or above baseline threshold → trade, sized by how far above baseline the radius sits, capped at 1.5× base lot.

Worth flagging honestly: "abstain and hold" means an already-open position stays open through a bar the certifier couldn't vouch for — a real asymmetry against the protective framing, since a fresh signal is refused but an existing position isn't closed on an inconclusive bar. Reacting to every abstain by flattening would be more conservative, but also more reactive to noise in the certification process itself.

One design choice worth naming, visible in ComputeBaseline() above: the ring buffer's write pointer always advances every bar regardless of outcome, storing abstains as a -1 sentinel rather than skipping them — a pointer that only advanced on successful certifications would stall during a run of abstains, freezing the baseline exactly when it should keep adapting.


Edge Cases and Pitfalls

The most important limitation before running this live: during the first InpBaselineLookback bars after attach, fewer than 5 valid readings exist, ComputeBaseline() returns -1, and the gate stays closed until the buffer warms up — safer than trading ungated, but budget roughly InpBaselineLookback closed bars before the first possible trade.

Sigma needs tuning in the same normalized feature units the model trained on, not raw price units. Too small and every prediction trivially "certifies" (votes nearly unanimous, radius uninformative); too large and even strong signals fail the 0.5 majority floor and abstain constantly. The σ sweep in the testing section exists to find where the radius actually discriminates rather than saturating at one extreme.

The Clopper-Pearson bound is also sensitive to N — a small N (say 20) under-certifies (more abstains, smaller radii than the model's true robustness would justify) rather than over-certifies, the safer failure mode but costlier in trade opportunities. N=100-200 is a reasonable middle ground here; heavier models trade N against latency.

Finally, what this system does and doesn't protect against. It certifies robustness to noise shaped like the training-time Gaussian assumption — not a genuine adversary crafting worst-case perturbations, and not a model that's wrong in a way that's robust to noise: a confidently, consistently wrong model certifies beautifully. It's also a guarantee about a specific noise model, not real broker or feed noise, which doesn't necessarily follow an isotropic Gaussian — perturbations added directly to engineered features can land on combinations no real OHLCV state would produce. A single scalar σ compounds this by treating every feature dimension as equally scaled and independent; a diagonal-covariance noise model would represent real feature heterogeneity more faithfully, at the cost of another parameter to tune.

A few methodological caveats about the testing here, stated plainly. There's no described train/validation/test split, so the drift-test result may be partially or fully in-sample. The 50.8% figure isn't compared against the same raw model on the same decided bars, so "accuracy lift" describes the certified subset, not an isolated, proven contribution from certification. N-sensitivity is discussed qualitatively but not quantified against abstain rate, radius, latency, and PnL. During long abstain streaks the baseline is built only from whatever valid radii remain, so it can become unrepresentative or the gate can simply stay closed. The threshold experiment below (0.75 vs. 0.4) demonstrates this parameter matters a great deal, not a validated way to pick it. And everything measured here comes from one model, one instrument, one window, one drift test — real, specific evidence about that run, not a general claim across markets or regimes.


The Monitor Indicator

CertRadiusMonitor.mq5 is a separate-window indicator that plots the certified radius and rolling baseline historically, so you can review a model's confidence margins over any period without running a script or reading a CSV. It reuses the same CFeatureEngine and CRobustnessCertifier classes as the EA — nothing about the certification math is reimplemented — the only new logic is redrawing efficiently across historical bars inside OnCalculate():

//+------------------------------------------------------------------+
//| Custom indicator iteration function                              |
//+------------------------------------------------------------------+
int OnCalculate(const int rates_total, const int prev_calculated, ...)
  {
   if(rates_total < InpBaselineLen + 5) return 0;

//--- Only re-certify newly-appeared bars on each redraw, not the whole
//--- chart history every time -- prev_calculated tracks how far the
//--- indicator already got on the last call.
   int start = (prev_calculated == 0) ? rates_total - InpBaselineLen - 1
                                       : rates_total - prev_calculated + 1;

   for(int shift_from_end = MathMax(start, 1); shift_from_end >= 1; shift_from_end--)
    {
      double features[];
      if(!g_features.Build(shift_from_end, features)) continue;

      int predicted_class = CERT_CLASS_ABSTAIN;
      double radius = 0.0, pA_lower = 0.0;
      g_certifier.Certify(features, InpSigma, InpNPasses, InpAlpha, predicted_class, radius, pA_lower);
      //--- write radius into BufRadius[], trailing median into BufBaseline[]
    }
   return rates_total;
  }

This exists as a separate indicator rather than relying on the EA's CSV log because reviewing radius behavior on the price chart is a much faster sanity check than scrolling a spreadsheet. The tradeoff is cost — InpNPasses defaults to 100, same as the EA, and redrawing a long history means running full certification on every visible bar, so the defaults are tuned lighter than a live-trading pass would need.


Training, Export, and Validation Tooling

Everything so far covers the live trading path, which depends on a trained ONNX model existing in the first place. The real numbers in the testing section below came from a specific set of scripts that train the model, export it correctly, and honestly check whether the certification mechanism does what this article claims. Six files exist for this purpose — offline tooling, none of them run during live trading.

RawONNXSignal_EA.mq5 exists purely as a baseline. It shares the same feature engine and the same trained model as the certified EA, but skips certification entirely — a single un-noised inference per bar, fixed lot size, no gate:

//+------------------------------------------------------------------+
//| RawPredict                                                       |
//+------------------------------------------------------------------+
int RawPredict(const double &features[])
  {
   matrix<float> in_matrix;
   in_matrix.Init(1, CERT_FEATURE_COUNT);
   for(int i = 0; i < CERT_FEATURE_COUNT; i++)
      in_matrix[0][i] = (float)features[i];

   matrix<float> out_matrix;
   out_matrix.Init(1, CERT_NUM_CLASSES); // must pre-size -- see note above
   if(!OnnxRun(g_onnx_handle, ONNX_NO_CONVERSION, in_matrix, out_matrix)) return -1;

   int best_idx = 0;
   float best_val = out_matrix[0][0];
   for(int c = 1; c < CERT_NUM_CLASSES; c++)
      if(out_matrix[0][c] > best_val) { best_val = out_matrix[0][c]; best_idx = c; }
   return best_idx;
  }

Its OnTick() trades that class directly, every bar, at a fixed lot — no smoothing, no radius, no gate. Its only purpose is giving the certified EA something honest to be measured against.

ExportTrainingData.mq5 is a script — run once on a chart — that walks backward through historical bars and writes out the same indicator values FeatureEngine::Build() computes, alongside a label derived mechanically from what actually happened next in price:

double ret = (c_future - c_now) / c_now;
int label;
if(ret > InpLabelThreshold)       label = 2; // long
else if(ret < -InpLabelThreshold) label = 0; // short
else                           label = 1; // flat

c_future looks InpLabelHorizon bars ahead of c_now — valid only offline, on history that's already happened. This matches the label logic used later in the drift test, so training labels and drift-test "ground truth" stay comparable. The script prints a label split as a sanity check — a training set that's 90% "flat" won't produce a model that learns much, and it warns if that happens.

train_export_model.py trains a small MLP classifier (scikit-learn) and exports it to ONNX. One wrinkle: skl2onnx exports sklearn classifiers with two outputs by default — a class-label tensor and a probability tensor — but the EA, indicator, and every script here bind only one. Left alone, that mismatch produces a working-looking ONNX file that fails at runtime. The fix forces a plain-tensor probability output and discards the label output before saving:

onnx_model = convert_sklearn(
    clf, initial_types=initial_type,
    options={id(clf): {"zipmap": False}},
)
# keep only the [N, num_classes] probability tensor, drop the label output
prob_output = None
for out in onnx_model.graph.output:
    if len(out.type.tensor_type.shape.dim) == 2:
        prob_output = out
del onnx_model.graph.output[:]
onnx_model.graph.output.append(prob_output)

This same script re-runs the feature contract check from the Python side — input width against FEATURE_COUNT, output width against NUM_CLASSES — mirroring the runtime check RobustnessCertifier::Init() does in MQL5.

DriftTestLogger.mq5 produced the real Fig. 5 numbers, close to the live EA's certification loop but run backward over history: since it's offline, it can look InpLabelHorizon bars into the future to compute the "correct" class, then log predicted-vs-actual for every bar — something the live EA can never do:

double c_now    = iClose(_Symbol, _Period, shift);
double c_future = iClose(_Symbol, _Period, shift - InpLabelHorizon);
double ret = (c_future - c_now) / c_now;

int actual_class;
if(ret > InpLabelThreshold)       actual_class = CERT_CLASS_LONG;
else if(ret < -InpLabelThreshold) actual_class = CERT_CLASS_SHORT;
else                           actual_class = CERT_CLASS_FLAT;

bool abstained = (predicted_class == CERT_CLASS_ABSTAIN);
int  correct   = (!abstained && predicted_class == actual_class) ? 1 : 0;
FileWrite(file_handle, TimeToString(bar_time), shift, predicted_class, actual_class,
      correct, (abstained ? 1 : 0), radius, pA_lower);

That per-bar CSV — time, predicted class, actual class, correct/abstained flags, radius — is the raw material for the rolling accuracy and rolling radius series plotted in Fig. 5. The label threshold and horizon here are deliberately the same values ExportTrainingData.mq5 uses, so "ground truth" means the same thing during training and during this validation pass.

EquityCompareBacktest.mq5 produced the real Fig. 4 numbers. Two separate Strategy Tester passes proved unreliable — each agent has an isolated sandbox that doesn't share files with the terminal and can reset between runs. So this script simulates both strategies in one pass over the same history, marking each bar to market using the previous bar's direction to avoid lookahead bias:

equity_gated += prev_gated_dir * (c_now - c_prev) * prev_gated_lot;
equity_raw   += prev_raw_dir   * (c_now - c_prev) * InpBaseLots;

Equity is in price-equivalent units (price difference × lot), not simulated account currency — spread, commission, and swap aren't modeled, so it's a simplified proxy, not broker-accurate, though genuine and computed from real data.

analyze_drift_test.py never runs inside MetaTrader — a standalone Python script that reads DriftTestLogger's CSV and checks, honestly, whether the certified radius leads the rolling accuracy trend at any lag, rather than eyeballing a chart:

best_lag, best_corr = 0, -2.0
for lag in range(-30, 31):
    shifted_radius = clean['rolling_radius'].shift(-lag)
    valid = shifted_radius.notna() & clean['rolling_accuracy'].notna()
    if valid.sum() < 30:
        continue
    corr = np.corrcoef(shifted_radius[valid], clean['rolling_accuracy'][valid])[0, 1]
    if corr > best_corr:
        best_corr, best_lag = corr, lag

A positive best_lag with a strong best_corr would mean the radius genuinely moves ahead of accuracy — what a real leading-indicator effect would look like. What actually came back was a best-fit correlation under 0.1 at every lag tested, so the article states plainly that no such pattern was found.


Testing in the Strategy Tester

The backtest compares two EAs built from the identical trained model and entry logic — raw ONNX signal, no gating, against the certified-radius-gated version — isolating what certification changes about the outcome. One note before the table: the sigma/N sweep and the EURUSD/XAUUSD comparison below describe the testing plan as originally scoped. What was actually run and reported with real numbers here is narrower — a single configuration (the table's defaults) on XAUUSD H1, plus one threshold variant (0.4× instead of 0.75×) discussed below. Treat the sweep ranges and EURUSD row as scope, not results.

Setting
Value
Instrument/timeframe actually tested
XAUUSD/H1
Originally scoped secondary instrument (not run)
EURUSD/H1
Sigma sweep range (planned, not run)
0.01 - 0.10 (feature-normalized units)
N sweep values (planned, not run)
50/100/200 passes
Confidence level (α) actually used
0.001 (Clopper-Pearson, one-sided)
Baseline lookback actually used
50 bars
Min radius gate multiple actually used
0.75× (default run), 0.4× (variant, see below)

Here's the honest result from running both strategies over the same real history: on this XAUUSD H1 window (2,999 bars — a separate EquityCompareBacktest.mq5 run from the 2,996-bar drift-test window above, since each script trims history slightly differently), raw finished at roughly 207 price-equivalent units of cumulative PnL against roughly 90 for gated at the default 0.75× threshold. Neither number is a claim about live-account PnL — no spread, commission, slippage, or swap modeled — they isolate the effect of gating on the same raw price moves. Gated didn't win. Worth understanding why.

The gate sat out 64% of bars at the default threshold. This window was a persistently trending stretch for gold, and the raw signal — always positioned — caught most of that trend simply by being in the market. The gated version's conservative threshold kept it flat through much of the same trend, and even on bars it did trade, the average outcome was slightly negative. Loosening to 0.4× recovered some ground (time in market to 45.5%, active-bar PnL turned slightly positive, final equity to roughly 103), but the gap to raw stayed wide.

The honest conclusion: the certifier's accuracy lift on decided bars doesn't automatically translate into a P&L advantage over a simpler always-in-market signal, particularly in a trending regime where time-in-market matters more than selectivity. Radius is a robustness measure, not an accuracy measure, and neither is automatically a profitability measure. Gating is useful for identifying statistically fragile predictions and controlling size around that fragility, but it's not a guaranteed edge over a naive baseline in every regime — this test is direct evidence of that.

Real cumulative PnL, gated vs. raw signal, XAUUSD H1, 2,999 bars — shows raw outperforming gated in this trending test window

Fig. 4. Real cumulative PnL from the same real price history and the same trained model — certified-radius-gated (default 0.75× threshold) against a raw ungated signal, XAUUSD H1, 2,999 bars. The raw signal outperformed in this trending window; see the discussion above for why.

Separately, a genuine drift test on XAUUSD H1 over 2,996 bars: the certifier abstained on 1,210, and on the 1,786 it committed to, accuracy came out to 50.8% — above the rough one-in-three you'd expect from blind guessing if classes were balanced (see the earlier caveat). That's conditional accuracy on the certifier's self-selected subset, not overall accuracy, and it isn't isolated against the raw model on the same subset — read it as "engaged predictions did meaningfully better than a naive guess," not "certification improves accuracy by this much."

What that test did not show is a clean leading-indicator relationship — correlation between rolling radius and rolling accuracy across a wide range of lags showed no consistent pattern (best-fit correlation under 0.1, not even in the "radius drops first" direction). The radius is a solid real-time confidence filter; treat it as that, not an early-warning system for model staleness.

Real rolling certified radius vs. rolling accuracy, XAUUSD H1, 2,996 bars — shows no consistent lead-lag relationship

Fig. 5. Real drift-test output on XAUUSD H1 (2,996 bars) — rolling certified radius against rolling accuracy on decided bars. The two series move somewhat independently; no consistent lead-lag relationship was found in this window.

Measure actual per-bar certification cost with GetMicrosecondCount() around the Certify() call before committing to a live N value — on the lightweight MLP used here, N=100 stayed well inside the tick budget for both H1 and M15 testing, but that's architecture-dependent and worth re-measuring for a meaningfully different model.


Conclusion

A softmax score tells you what the model decided. A certified radius tells you how robust that decision is to live-feed noise — separate from whether it's correct, and separate again from whether acting on it makes money. That gap is where fragile predictions hide, invisible until you go looking with something like randomized smoothing. This was practical to implement natively because MQL5's statistics library already has the two functions the method rests on, MathQuantileBeta() and MathQuantileNormal(), sitting in the standard library. What this article's testing showed, plainly: the certifier's accuracy lift on engaged bars is real and measured, but didn't translate into a cumulative-PnL advantage over a simpler raw signal in the one trending window tested — robustness, correctness, and profitability are three different things, and demonstrating one doesn't demonstrate the others.

The natural next extension is tightening the multi-class bound beyond the binary simplification used here, or extending the noise model beyond isotropic Gaussian to something that better reflects the actual per-feature noise characteristics of a given instrument — spread noise and volume noise don't really behave the same way, and a diagonal covariance smoothing distribution would capture that more faithfully than a single scalar σ.

File
Type
Description
CertDefines.mqh
Include
Feature contract constants and default certification parameters.
FeatureEngine.mqh
Include
Builds the fixed 10-wide feature vector each closed bar.
RobustnessCertifier.mqh
Include
Randomized smoothing, Clopper-Pearson bound, certified radius calculation.
CertifiedONNXRadius_EA.mq5
Expert Advisor
Live trading logic: certifies each bar, gates trade action and sizing on radius.
CertRadiusMonitor.mq5
Indicator
Plots certified radius and rolling baseline in a separate chart window.
RawONNXSignal_EA.mq5
Expert Advisor
Ungated baseline: single un-noised inference per bar, fixed lot, no certification. Exists purely for comparison.
ExportTrainingData.mq5
Script
Exports historical bars + indicator values + forward-return labels to a CSV for model training.
DriftTestLogger.mq5
Script
Offline diagnostic: certifies historical bars and compares predictions against realized outcomes. Produced Fig. 5.
EquityCompareBacktest.mq5
Script
Simulates gated vs. raw signal equity curves bar-by-bar against real history, no lookahead. Produced Fig. 4.
train_export_model.py
Python Script
Trains base classifier, exports ONNX, enforces feature-contract consistency check.
analyze_drift_test.py
Python Script
Reads DriftTestLogger's output and checks radius-vs-accuracy lead/lag via cross-correlation.
Attached files |
MQL5.zip (26.45 KB)
Features of Custom Indicators Creation Features of Custom Indicators Creation
Creation of Custom Indicators in the MetaTrader trading system has a number of features.
Neural Networks in Trading: From Transformers to Spiking Neurons (Conclusion) Neural Networks in Trading: From Transformers to Spiking Neurons (Conclusion)
Neural networks are already changing the way we analyze markets, and new architectures are opening up even more possibilities. In this article, we wrap up our work with the SpikingBrain framework, which opens up new possibilities for us.
Features of Experts Advisors Features of Experts Advisors
Creation of expert advisors in the MetaTrader trading system has a number of features.
MiniRocket: A Deterministic Time-Series Classifier and What It Finds in Seven Classic Setups MiniRocket: A Deterministic Time-Series Classifier and What It Finds in Seven Classic Setups
This article delivers a native MQL5 MiniRocket: 84 fixed convolution kernels yield 9,996 features quickly and deterministically, requiring no training loop and no external runtime. We verify the port against sktime and a float64 reimplementation, then run a reproducible audit of seven classic setups; planted and coin‑flip controls confirm correctness, and a 25%‑flipped control sets the detection threshold.