← Back to blog

Research method /

A leakage-safe evaluation checklist for machine learning

How to keep information from the test set out of preprocessing, model selection, and reported results.

Updated

A model can appear strong because the evaluation pipeline quietly gave it information it would not have in the real world. This is data leakage. It can happen even when the final classifier was trained only on rows labeled “training.” Feature selection, scaling, repeated records, and decisions made after looking at test results can all leak information.

1. Define the unit that must stay separate

In medical data, several scans may belong to one person. If one scan appears in training and another in testing, the model may learn subject-specific patterns. The split should therefore respect the subject, not just the row. The same principle applies to documents from one source, multiple visits to one site, or repeated measurements from one device.

Write down the prediction setting first: what information will exist at the moment of prediction, and what unit counts as genuinely unseen?

2. Split before learning transformations

Do not estimate normalization parameters, select features, impute missing values, or tune a threshold on the complete dataset before splitting it. Any fitted transformation belongs inside the training fold. The test fold receives only the transformation learned from training data. A pipeline helps enforce that order during cross-validation.

3. Keep model selection away from the final test

Choose features, hyperparameters, prompts, thresholds, and stopping rules using training and validation data. Reserve the final test for a single, pre-specified assessment. Repeatedly checking the test set and revising the model makes that set part of development, even if no test row is passed to fit().

4. Check the boundary again

Before reporting a result, verify that group identifiers do not cross the split, that duplicated records are handled, and that preprocessing was fitted only where intended. Report the split strategy and uncertainty, not just one score. If an independent cohort exists, evaluate it separately and explain any distribution differences.

In my cross-cohort research, the central question is whether Alzheimer’s disease classifiers generalize between ADNI and OASIS under subject-level separation and leakage-safe preprocessing. The general lesson extends well beyond medical imaging: evaluation is part of the model, not paperwork around it.

Keep a record of each modeling decision

When several cohorts are available, reversing training and evaluation roles can reveal whether generalization is directional. Keep each learned transformation inside the training fold, reserve the external cohort for final validation, and record each modeling decision in a reproducible pipeline. These controls determine whether a reported result reflects generalization rather than information passed across the evaluation boundary.

References: scikit-learn: common pitfalls and data leakage, scikit-learn: cross-validation with groups.

Contact

Elgün · AI Researcher

Let's talk

Good questions welcome. Citations optional.

Academic CV