Trading the Breaking

Trading the Breaking

Research

[WITH CODE] Feature selection: Filter-based methods

Supervised approches

Jun 16, 2026
∙ Paid

Before you begin, remember that you have an index with the newsletter content organized by clicking below.

INDEX


From raw market data to scalable statistical filters

Filter-based feature selection evaluates the intrinsic statistical properties of candidate variables independent of downstream model training. By screening out noisy inputs before fitting complex algorithms, statistical filters reduce dimensional overreach, limit execution friction, and protect systematic strategies from temporal leakage.

What’s inside:

  1. Defining the filter paradigm: Filters evaluate features through direct statistical tests against target labels outside the model training loop, providing a fast screening layer against high-dimensional noise.

  2. Controlling selection risks: Ranking variables on statistical scores can mislead researchers if a feature increases turnover, inflates execution costs, or relies on unavailable forward bar data.

  3. Measuring Information Gain: Discretizing continuous metrics into quantile bins isolates entropy reduction, revealing non-linear regime thresholds that standard linear filters overlook.

  4. Class separation with Fisher Score and ANOVA: Comparing between-class variance against within-class dispersion highlights features with distinct conditional distributions, while Welch’s ANOVA corrects for fat tails and non-constant variance.

  5. Capturing non-linear signals through Mutual Information: Applying continuous nearest-neighbor estimators detects non-linear dependencies, preserving valuable predictive signals such as U-shaped volume reversals.

  6. Uncovering joint dependencies with ReliefF: Evaluating instances against their nearest hits and misses in multi-dimensional space isolates localized feature interactions without training a full model.

  7. Preventing collinearity with mRMR: Optimizing target relevance against pairwise mutual information keeps feature sets diverse and stops models from clustering around identical risk drivers.

Sample
386KB ∙ PDF file
Download
Download

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Quant Beckman · Publisher Privacy ∙ Publisher Terms
Substack · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture