DOI: 10.3390/e28080933 ISSN: 1099-4300

Parameter-Independent Feature Ranking with Volume-Integrated Sharma–Mittal Entropy: Kernel-Based Estimation, Theoretical Properties and Empirical Validation

Nida Oruç Ünal, Muzaffer Göztaş, Doğan Yıldız

Feature selection is a critical step in regression problems where a large number of continuous explanatory variables explain the same target through different dependency structures. Classical filters may remain sensitive to a single form of dependence, a single scale, or a specific discretization scheme; generalized entropy measures, on the other hand, typically require the parameters to be fixed at a single point. This study proposes a framework that evaluates the Sharma–Mittal entropy volumetrically across a two-dimensional parameter region rather than for a single parameter pair. For the continuous target and explanatory variables, the marginal, joint, and conditional densities are obtained using a Gaussian kernel density estimation; the conditional entropy and information gain surfaces are integrated across the region Ω = [0.05, 0.95]2 in the α-β plane to define three indices: PICSME, which measures the conditional uncertainty volume; PIGSME, which measures the gain volume; and NIGSME, which is the ratio of this gain to the total entropy volume of the target. The method is supported by bandwidth consistency and the renormalization of conditional densities; thus, the issue of negative gain that can occur in the continuous variables is resolved, yielding positive and interpretable scores across all six datasets. It is formally demonstrated that the fact that the three indices produce the same ranking is not an empirical observation but rather the result of a monotonicity relationship valid under a fixed target entropy volume. The method is compared with Pearson and Spearman correlations, the Shannon information gain, mutual information, and random forest variable importance across six regression datasets (Airfoil Self-Noise, AirQualityUCI, BodyFat, Meteorology, Concrete, and WineQualityWhite) that differ in their sample size, dimensions, and application domain. The evaluation is not limited to ranking consistency; the out-of-sample prediction performance is measured using least-squares models on the top-k subsets, with rankings calculated from the training partition. The findings show that NIGSME exhibits a performance comparable to that of built-in filters, outperforms them on the Concrete and Meteorology datasets, and never ranks as the weakest method on any dataset. The results demonstrate that volumetric entropy metrics defined across the entire parameter space provide a feature-ranking tool that is independent of parameter selection for continuous variables.

More from our Archive