AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.5 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Improved chi-square feature selection for robust heart disease data classification

Heba Nayl1Elkhateeb S. Aly2,3( )Amira Rezk4M. E. Fares1
Department of Mathematics, Faculty of Science, Mansoura University, Mansoura, Egypt
Department of Mathematics, College of Science, Jazan University, P.O. Box 114 Jazan 45142, Kingdom Saudi Arabia
Nanotechnology research unit, College of Science, Jazan University, P.O. Box 114 Jazan 45142, Kingdom Saudi Arabia
Department of Information Systems, Faculty of Computer and Information Sciences, Mansoura University, Mansoura, Egypt
Show Author Information

Abstract

Early diagnosis of heart disease is vital for reducing mortality and improving patient outcomes; yet, accurate prediction remains a significant challenge owing to the complexity and high dimensionality of medical data. Data preprocessing is essential for overcoming these issues by cleaning, transforming, reducing, and balancing data to provide reliable inputs for feature selection and classification. This study introduces an improved chi-square ( χ 2 ) feature selection framework combined with multiple classifiers to enhance predictive performance. Our method was applied to Cleveland heart disease and diabetes datasets, where numeric attributes were discretized into categorical values, enabling χ 2 to select the most informative features while eliminating redundancy. Several classifiers, including support vector machine (SVM), logistic regression (LR), K-nearest neighbors (KNN), and naive Bayes (NB), were trained using both the reduced subset and the complete feature set. Results show that the preprocessing include χ 2 feature selection, achieved the highest performance. On the Cleveland dataset, the model attained a mean accuracy of 93.72%, precision of 94.01%, recall of 93.72%, F1-score of 93.74%, and an area under the curve(AUC) of 97.87%, while on the diabetes dataset, it achieved mean values of 93.55% accuracy, 94.23% precision, 93.55% recall, 93.48% F1-score, and an AUC 93.53%. The main contribution of this work lies in integrating discretization with χ 2 based selection to produce a compact and discriminative feature subset. With a minimal number of selected features, the proposed approach delivers robust, accurate, and computationally efficient heart disease prediction, outperforming existing methods.

CLC number: 97R40, 97P50, 62F07

References

【1】
【1】
 
 
AIMS Mathematics
Pages 2682-2701

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Nayl H, Aly ES, Rezk A, et al. Improved chi-square feature selection for robust heart disease data classification. AIMS Mathematics, 2026, 11(1): 2682-2701. https://doi.org/10.3934/math.2026108

183

Views

3

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 28 September 2025
Revised: 07 January 2026
Accepted: 14 January 2026
Published: 27 January 2026
©2026 the Author(s), licensee AIMS Press.

This is an open access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0)