Predicting small for gestational age (SGA) and large for gestational age (LGA) newborns is crucial for preventing adverse pregnancy outcomes and improving neonatal health. Existing purely data-driven methods have achieved promising performance in SGA-LGA newborn prediction, but most of them lack model interpretability, raising doubts in clinical auxiliary diagnosis and failing to provide targeted interventions. To address this challenge, this paper proposes an obstetric knowledge-driven large language model, termed Obstet-LLM. Specifically, Obstet-LLM is pre-trained on the content from professional knowledge bases in the field of obstetrics. This pre-training enables the model to understand and integrate domain-specific knowledge and clinical insights. Subsequently, the model undergoes prompt engineering and instruction fine-tuning to learn precise predictions of neonatal growth outcomes. To enhance model interpretability, a causal learning paradigm is designed, allowing Obstet-LLM to generate clear and understandable explanations. Experimental results show that Obstet-LLM achieves an accuracy rate of over 90% in predicting SGA-LGA newborns, outperforming existing data-driven models. Moreover, by integrating domain-specific knowledge, Obstet-LLM provides actionable insights for clinicians, thereby enhancing the quality of prenatal care and reducing the risk of adverse pregnancy outcomes.
- Article type
- Year
- Co-author
Open Access
Issue
Nonnegative Matrix Factorization (NMF) is one of the most popular feature learning technologies in the field of machine learning and pattern recognition. It has been widely used and studied in the multi-view clustering tasks because of its effectiveness. This study proposes a general semi-supervised multi-view nonnegative matrix factorization algorithm. This algorithm incorporates discriminative and geometric information on data to learn a better-fused representation, and adopts a feature normalizing strategy to align the different views. Two specific implementations of this algorithm are developed to validate the effectiveness of the proposed framework: Graph regularization based Discriminatively Constrained Multi-View Nonnegative Matrix Factorization (GDCMVNMF) and Extended Multi-View Constrained Nonnegative Matrix Factorization (ExMVCNMF). The intrinsic connection between these two specific implementations is discussed, and the optimization based on multiply update rules is presented. Experiments on six datasets show that the effectiveness of GDCMVNMF and ExMVCNMF outperforms several representative unsupervised and semi-supervised multi-view NMF approaches.
Open Access
Issue
The lack of labeled image data poses a serious challenge to the application of artificial intelligence (AI) in medical image diagnosis. Medical image notes contain valuable patient information that could be used to label images for machine learning tasks. However, most image note texts are unstructured with heterogeneity and short-paragraph characters, which fail traditional keyword-based techniques. We utilized a deep learning approach to recover missing labels for medical image notes automatically by using a combination of deep word embedding and deep neural network classifiers. Bidirectional encoder representations from transformers trained on medical image notes corpus (MinBERT) were proposed. We applied the proposed techniques to two typical classification tasks: Medical image type identification and clinical diagnosis identification. The two methods significantly outperformed baseline methods and presented high accuracies of 99.56
京公网安备11010802044758号