Sort:
Open Access Medical Psychology Issue
SMOTEENN hybrid sampling combined with stacking ensemble model achieves optimal performance in predicting depression tendency among college students: a comparison of 84 models based on class imbalance handling
Journal of Army Medical University 2026, 48(15): 2214-2225
Published: 15 August 2026
Abstract PDF (2.2 MB) Collect
Downloads:1
Objective

To address the dilemma of insufficient recognition of the minority class caused by class imbalance in predicting depression tendency among college students, and systematically compare the predictive performance of 7 oversampling, undersampling, and hybrid sampling methods combined with 12 machine learning models on a class-imbalanced dataset of music APP listening habits and depression tendency in them, providing a reference for method selection in handling class-imbalanced data.

Methods

This study adopted a cross-sectional study design. Data were collected from questionnaire surveys of college students from 29 provinces, autonomous regions, and municipalities in China in 2023, with college students primarily being undergraduate and postgraduate students aged 18 to 24 years. Univariate analysis was used to screen 10 features significantly associated with depression tendency in the participants. Three oversampling methods including synthetic minority over-sampling technique (SMOTE), adaptive synthetic sampling approach (ADASYN), and synthetic minority over-sampling technique for regression with Gaussian noise (SMOGN), 2 undersampling methods including edited nearest neighbors (ENN) and Tomek Links, and 2 hybrid sampling methods including SMOTEENN and ADASYN-Tomek were applied to process the raw data. The raw data and processed data were respectively used to construct prediction models with 12 classification algorithms, including logistic regression (LR), support vector machine (SVM), random forest (RF), decision tree (DT), Light Gradient Boosting Machine (LightGBM), K-Nearest Neighbors (KNN), eXtreme Gradient Boosting (XGBoost), Lasso logistic regression (LASSO), ridge regression (Ridge), elastic net (ENet), multilayer perceptron (MLP), and stacking ensemble model (Stacking). Area under the receiver operating characteristic curve (AUC), Recall, F1-score, and balanced accuracy were selected as evaluation metrics to assess model performance. The optimal model was screened by comparing various combinations of class-imbalance handling techniques and machine learning strategies. SHapley Additive exPlanations (SHAP) analysis was applied for interpretability analysis for the optimal model.

Results

Results Univariate analysis identified 10 features significantly associated with depression tendency in college students: gender, grade, major category, preference for Children’s songs, approximate length since music listening habits, preference for Chinese style songs, usual music styles, daily listening duration, period of listening during a day, and frequency of comments on music APP. The comparison of 84 combined models with the 12 models based on raw data showed that hybrid sampling outperformed single sampling methods among the 7 resampling strategies; among the 12 classification algorithms, the stacking ensemble model demonstrated better comprehensive performance than individual classifiers. The combination of SMOTEENN hybrid sampling with Stacking achieved the best performance (AUC=0.95, F1=0.89, Recall=0.88, balanced accuracy=0.89). SHAP analysis showed that “usual listening style” contributed the highest in predicting depression tendency among college students.

Conclusion

Based on the comparison of 84 combination models, the SMOTEENN+stacking ensemble model can accurately predict depression tendency in college students. The comparative analysis provides a methodological basis and practical reference for method selection in class-imbalanced data scenarios and for early identification of mental health risk based on music listening behavior data.

Open Access Monographic Report Issue
LightGBM demonstrates optimal performance in predicting 28-day mortality risk in patients with sepsis-associated acute kidney injury: development and validation of an interpretable machine learning model based on MIMIC-Ⅳ and eICU databases
Journal of Army Medical University 2026, 48(15): 2129-2138
Published: 15 August 2026
Abstract PDF (1.7 MB) Collect
Downloads:1
Objective

Sepsis-associated acute kidney injury (SA-AKI) is a life-threatening condition with high mortality. This study aims to develop and validate an interpretable machine-learning model for assessing mortality risk in SA-AKI patients, and facilitate early identification of high-risk patients to support clinical decision-making.

Methods

A retrospective cohort study was conducted on the clinical data derived from the Medical Information Mart for Intensive Care Ⅳ (MIMIC-Ⅳ) and the eICU Collaborative Research Database (eICU-CRD). A total of 24487 patients from MIMIC-Ⅳ and 13757 ones from eICU-CRD were included for analysis. Predictive variables were selected through univariate analysis and least absolute shrinkage and selection operator (LASSO) regression. The MIMIC-Ⅳ dataset was randomly divided into training and internal testing sets in a ratio of 8:2. Ten machine-learning algorithms were used to construct mortality risk prediction models, including logistic regression (LR), random forest (RF), light gradient boosting machine (LightGBM), K-nearest neighbors (KNN), eXtreme gradient boosting (XGBoost), ridge regression (Ridge), elastic net (ENet), decision tree (DT), stacking ensemble model (Stacking), and multilayer perceptron (MLP). Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, specificity, and with calibration curve and decision curve analyses. SHapley Additive exPlanations (SHAP) were adopted to analyze the contribution and directional effect of each variable on prediction. External validation was performed using SA-AKI patients meeting inclusion criteria from the eICU-CRD database.

Results

A total of 31 predictive variables were selected through univariate analysis combined with LASSO regression, and 10 machine-learning models for mortality risk prediction were constructed. The LightGBM model achieved an AUC value of 0.817 (95%CI: 0.803 to 0.830) in the internal testing set and 0.735 (95%CI: 0.725 to 0.745) in the external validation set based on the eICU-CRD database, outperforming other models. After comprehensive comparison of discrimination, calibration, and decision curve analyses across all models, the LightGBM model demonstrated optimal performance, suggesting favorable predictive efficiency and clinical applicability in mortality risk prediction for SA-AKI patients. SHAP analysis revealed that the most important variables for 28-day mortality risk included sequential organ failure assessment (SOFA) score, Charlson comorbidity index, Glasgow coma scale (GCS) score, mean body temperature, mean blood urea nitrogen, mean serum sodium, mean respiratory rate, mean creatinine, Oxford acute severity of illness score (OASIS), and red blood cell distribution width (RDW).

Conclusion

For SA-AKI patients, the mortality risk prediction model based on the LightGBM algorithm demonstrates optimal performance. The SOFA score, Charlson comorbidity index, and other variables are important indicators for predicting 28-day mortality in these patients, providing a reference for early identification of high-risk patients.

Open Access Public Health and Preventive Medicine Issue
Correlation between music APP listening habits and depression tendency in college students based on SMOTEENN algorithm
Journal of Army Medical University 2024, 46(23): 2670-2680
Published: 15 December 2024
Abstract PDF (671.7 KB) Collect
Downloads:0
Objective

To investigate the influencing factors for tendency towards depression in college students having music listening habits with music APP, and develop a prediction model and further optimize it.

Methods

A total of 1157 college students were subjected with convenient sampling and surveyed with questionaires between April and May 2023. Univariate analysis and logistic regression analysis were employed to identify the influencing factors. Then a prediction model was constructed based on these factors. SMOTEENN over-sampling algorithm was utilized to enhance the dataset and construct the prediction model.

Results

Logistic regression analysis revealed that female (OR=1.730, 95%CI: 1.257~2.396), senior grade (OR=2.649, 95%CI: 1.198~7.506), postgraduate grade (OR=2.041, 95%CI: 1.231~ 3.885), major in Science(OR=1.573, 95%CI: 1.052~2.350), listening for a duration of 0.5~2 h (OR=1.661, 95%CI: 1.011~2.695), music style of melancholy (OR=2.668, 95%CI: 1.701~4.226) and of nostalgia (OR=1.751, 95%CI: 1.086~2.837), and frequency of comments on 0~5% of songs (OR=2.938, 95%CI: 1.018~8.417) were independent risk factors for depressive tendency. Time since listening to music for 1~3 years (OR=0.547, 95%CI: 0.347~0.872), listening to music from 14:00 to 18:00 (OR= 0.375, 95%CI: 0.167~0.845) and 18:00 to 21:00 (OR=0.313, 95%CI: 0.148~0.671), and preference for Chinese style songs (OR=0.711, 95%CI: 0.541~0.941) were independent protective factors. The logistic early warning model based on SMOTEENN algorithm demonstrated optimal predictive performance with an AUC value of 0.923.

Conclusion

Our constructed logistic regression model has identified 9 independent influencing factors associated with depression tendency among college students.The early warning model based on SMOTEENN algorithm can predict the depression tendency more accurately for college students.

Total 3