AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (5.2 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Biomedical Engineering | Publishing Language: Chinese | Open Access

Pregnancy probability prediction models based on 5 machine learning algorithms and comparison of their performance

Chao REN1,2Huan YANG1Niya ZHOU3,4Qing CHEN1Wenzheng ZHOU5Tong WANG1Xi LING1Lei SUN1Peng ZOU1Zhuoyue LIANG1Lin AO1Jinyi LIU1Jia CAO1( )
Institute of Toxicology, Faculty of Military Preventive Medicine, Army Medical University (Third Military Medical University), Chongqing
Department of Urology, General Hospital of Western Theater Command, Chengdu, Sichuan
Department of Scientific Research Management and Foreign Affairs, Chongqing Health Center for Women and Children, Women and Children's Hospital of Chongqing Medical University, Chongqing
Chongqing Research Center for Prevention and Control of Maternal and Child Diseases and Public Health, Chongqing
Department of Quality Control Center, Clinical and Public Health Research, Chongqing, China
Show Author Information

Abstract

Objective

To construct 5 machine-learning models and compare their performance in predicting the associations between pre-pregnancy socio-psycho-behavioral exposures of both spouses and preconception outcomes.

Methods

Based on Chongqing Preconception Reproductive Health and Birth Outcome Cohort of volunteers recruited from Chongqing Health Center for Women and Children during January 2019 and March 2022, 5447 couples were recruited and surveyed through interviewer-interview for the demographic and social-psychological-behavioral data of both spouses (221 variables). According to the inclusion and exclusion criteria, 4097 couples were finally included, and randomly assigned into a training set (n=2867 spouses) and a validation set (n=1230 spouses) at a ratio of 7∶3. Feature analysis and collinear screening were applied to select the potential exposure factors. In consideration of difficulty to carry out semen parameters analysis in primary healthcare institutions, feature Set 1 including sperm parameters and feature Set 2 excluding semen parameters were constructed by including or excluding sperm quality simultaneously in the training set and the validation set. Five algorithms, that is, Logistic Regression, Naive Bayes, Random Forest, Gradient Boosting Machine, and Support Vector Machine, were used to construct preconception outcome prediction models, and the parameters of each model were optimized using random search combined with grid search. The predictive performance of each model was compared using precision, recall, F1 score, area under the receiver operating characteristic curve (AUC), and calibration curve. The optimal model was then selected by comparing the changes in the predictive ability of the questionnaire data for fertility outcomes with or without semen parameters.

Results

There were 24 variables screened out in feature Set 1, and 16 variables in feature Set 2. In feature Set 1, the gradient boosting machine performed better, with a relatively higher AUC value (0.651) and better F1 score (0.61). The logistic regression model performed stably (AUC value =0.647) and was suitable as the reference model. The random forest (AUC value=0.641), Naive Bayes (AUC value=0.641), and support vector machine (AUC value=0.634) performed second-best. By utilizing the gradient boosting machine, comparable results were found between the predictions from feature sets with or without semen parameters, as in feature Set 1, the AUC value of its validation set was 0.651(95%CI: 0.629~0.681), the prediction accuracy was 0.63, the recall rate was 0.65,and the average precision value F1 was 0.61; and in feature Set 2, the AUC value of its validation set was 0.649(95%CI: 0.624~0.663), and both the calibration curves were close to the ideal curve. The prediction results indicated that in feature Set 1, the features highly negatively correlated with preconception outcomes were female age, male age, and no pregnancy within 1 year without contraception, while the features highly positively correlated with preconception outcomes were female pregnancy history, total sperm vitality, and use of contraceptive measures before enrollment.

Conclusion

Among the 5 machine-learning algorithms performed in this cohort data, the gradient boosting machine shows slightly better performance. There are 24 factors being associated with preconception outcomes in both spouses, and the performance of the simplified model excluding semen parameters is not significantly declined. It is feasible to use machine-learning methods to predict human preconception outcomes through social-psychological-behavioral questionnaires.

CLC number: R195.1; R319; R714.1 Document code: A

References

【1】
【1】
 
 
Journal of Army Medical University
Pages 1376-1387

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
REN C, YANG H, ZHOU N, et al. Pregnancy probability prediction models based on 5 machine learning algorithms and comparison of their performance. Journal of Army Medical University, 2025, 47(12): 1376-1387. https://doi.org/10.16016/j.2097-0927.202502013

1

Views

0

Downloads

0

Crossref

0

Scopus

0

CSCD

Received: 06 February 2025
Revised: 17 May 2025
Published: 30 June 2025
© 2025 Journal of Army Medical University

This is an open access article under the CC BY license (https://creativecommons.org/licenses/by/4.0/).