AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (23.3 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

ELM-APDPs: An Explainable Ensemble Learning Method for Accurate Prediction of Druggable Proteins

Mujeebu Rehman1Qinghua Liu1Ali Ghulam2Tariq Ahmad3Jawad Khan4( )Dildar Hussain5( )Yeong Hyeon Gu5
School of Information and Communication Engineering, Guilin University of Electronic Technology, Guilin, 541004, China
Information Technology Centre, Sindh Agriculture University, Tandojam, 70060, Pakistan
School of Electrical and Information Engineering, Hunan University, Changsha, 410082, China
School of Computing, Gachon University, Seongnam, 13120, Republic of Korea
Department AI and Data Science, Sejong University, Seoul, 05006, Republic of Korea
Show Author Information

Abstract

Identifying druggable proteins, which are capable of binding therapeutic compounds, remains a critical and resource-intensive challenge in drug discovery. To address this, we propose CEL-IDP (Comparison of Ensemble Learning Methods for Identification of Druggable Proteins), a computational framework combining three feature extraction methods Dipeptide Deviation from Expected Mean (DDE), Enhanced Amino Acid Composition (EAAC), and Enhanced Grouped Amino Acid Composition (EGAAC) with ensemble learning strategies (Bagging, Boosting, Stacking) to classify druggable proteins from sequence data. DDE captures dipeptide frequency deviations, EAAC encodes positional amino acid information, and EGAAC groups residues by physicochemical properties to generate discriminative feature vectors. These features were analyzed using ensemble models to overcome the limitations of single classifiers. EGAAC outperformed DDE and EAAC, with Random Forest (Bagging) and XGBoost (Boosting) achieving the highest accuracy of 71.66%, demonstrating superior performance in capturing critical biochemical patterns. Stacking showed intermediate results (68.33%), while EAAC and DDE-based models yielded lower accuracies (56.66%–66.87%). CEL-IDP streamlines large-scale druggability prediction, reduces reliance on costly experimental screening, and aligns with global initiatives like Target 2035 to expand action-able drug targets. This work advances machine learning-driven drug discovery by systematizing feature engineering and ensemble model optimization, providing a scalable workflow to accelerate target identification and validation.

References

【1】
【1】
 
 
Computer Modeling in Engineering & Sciences
Pages 779-805

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Rehman M, Liu Q, Ghulam A, et al. ELM-APDPs: An Explainable Ensemble Learning Method for Accurate Prediction of Druggable Proteins. Computer Modeling in Engineering & Sciences, 2025, 145(1): 779-805. https://doi.org/10.32604/cmes.2025.067412

371

Views

8

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 02 May 2025
Accepted: 09 September 2025
Published: 30 October 2025
© The Author 2024.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.