AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.4 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Open Access

Adaptive feature selection method for high-dimensional imbalanced data classification

Jianzhen WU1Zhen XUE1,2( )Liangliang ZHANG1Xu YANG1
School of Mathematics, North University of China, Taiyuan 030051, China
Department of Mathematics, City University of Hong Kong, Kowloon 999077, Hong Kong, China
Show Author Information

Abstract

Data collected in fields such as cybersecurity and biomedicine often encounter high dimensionality and class imbalance. To address the problem of low classification accuracy for minority class samples arising from numerous irrelevant and redundant features in high-dimensional imbalanced data, we proposed a novel feature selection method named AMF-SGSK based on adaptive multi-filter and subspace-based gaining sharing knowledge. Firstly, the balanced dataset was obtained by random under-sampling. Secondly, combining the feature importance score with the AUC score for each filter method, we proposed a concept called feature hardness to judge the importance of feature, which could adaptively select the essential features. Finally, the optimal feature subset was obtained by gaining sharing knowledge in multiple subspaces. This approach effectively achieved dimensionality reduction for high-dimensional imbalanced data. The experiment results on 30 benchmark imbalanced datasets showed that AMF-SGSK performed better than other eight commonly used algorithms including BGWO and IG-SSO in terms of F1-score, AUC, and G-mean. The mean values of F1-score, AUC, and G-mean for AMF-SGSK are 0.950, 0.967, and 0.965, respectively, achieving the highest among all algorithms. And the mean value of G-mean is higher than those of IG-PSO, ReliefF-GWO, and BGOA by 3.72%, 11.12%, and 20.06%, respectively. Furthermore, the selected feature ratio is below 0.01 across the selected ten datasets, further demonstrating the proposed method’s overall superiority over competing approaches. AMF-SGSK could adaptively remove irrelevant and redundant features and effectively improve the classification accuracy of high-dimensional imbalanced data, providing scientific and technological references for practical applications.

References

【1】
【1】
 
 
Journal of Measurement Science and Instrumentation
Pages 612-624

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
WU J, XUE Z, ZHANG L, et al. Adaptive feature selection method for high-dimensional imbalanced data classification. Journal of Measurement Science and Instrumentation, 2025, 16(4): 612-624. https://doi.org/10.62756/jmsi.1674-8042.2025059

1471

Views

59

Downloads

0

Crossref

0

CSCD

Received: 14 April 2025
Revised: 30 May 2025
Accepted: 09 July 2025
Published: 01 December 2025
© The Author(s) 2025.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits use, distribution and reproduction in any medium, provided the original work is properly cited.