Discussing the article: "MetaTrader 5 Machine Learning Blueprint (Part 22): Auditing the Selection Criterion — Measuring Overfit in Hyperparameter Search"

Check out the new article: MetaTrader 5 Machine Learning Blueprint (Part 22): Auditing the Selection Criterion — Measuring Overfit in Hyperparameter Search.

We show that nested cross-validation removes optimistic evaluation bias but does not stop hyperparameter search from overfitting a finite sample, and we add a diagnostic for when the criterion cannot resolve candidates from noise. We implement OffsetPurgedKFold (boundary-shifted repartitioning), a criterion variance audit (resolution ratio and selection stability), a selection optimization curve, a corrected one-standard-error rule, noise-floor early stopping, and an effective trial count for correlated trials. The result is a dispersion report per candidate and an effective count for the Deflated Sharpe Ratio.

In Part 16 we moved hyperparameter search inside every outer fold and reported the resulting nested score instead of the search's own best value. That closed a specific hole: the number we publish is now an estimate of what the pipeline does on data it has never touched, rather than the maximum of a quantity the search spent hundreds of trials maximizing.

It closed one hole and left a larger one open. Nested cross-validation makes the reported figure honest about the damage; it does not reduce the damage, and — importantly — it does not isolate the damage attributable to selection overfitting specifically. Nested CV evaluates the entire pipeline end-to-end. If the reported score is poor, the cause could be selection overfitting, a regime shift between folds, label noise, an unstable feature set, or any combination of them. To attribute the shortfall to selection overfitting alone, a counterfactual benchmark is needed: an oracle that knows the best hyperparameters, or a fixed-hyperparameter pipeline scored on the same folds, or an independent resampling of the selection procedure. This article builds the diagnostic tools, but the reader should be clear from the start that the resolution ratio below measures the local discriminating power of the criterion, not the total damage of the pipeline. Those are related but distinct quantities. The search still tunes hyperparameters against a finite sample, still finds whatever peculiarities that sample contains, and still hands back a model shaped partly by noise. Nothing in the pipeline asks the prior question: on this sample, with this criterion, is the search capable of distinguishing one hyperparameter vector from another at all?

This article answers that question with a number. We implement afml.cross_validation.selection_overfit. It measures the criterion's dispersion across legitimate repartitions of the same data, compares it with the dispersion across candidates, and reports their ratio — specifically, the ratio of candidate spread to noise floor (signal-to-noise, not noise-to-signal). In the study below the ratio is 0.49. The criterion's noise is about twice the spread it is expected to resolve. Under that condition the search selected the candidate that is best on average on one partition in eight, and on this particular sample its held-out score was beaten by twenty-six of the sixty candidates it rejected. (As Section 6 explains, that last number does not by itself prove that uniform random selection would have done better; it shows only that the specific candidate chosen by the argmax was not the one that generalized best.)


Author: Patrick Murimi Njoroge