Sort:
Open Access Original Article Issue
Multimodal artificial intelligence predicts PIK3CA mutation in breast cancer from digital pathology and clinical data: a multicenter study
Cancer Biology & Medicine 2026, 23(3): 430-450
Published: 01 March 2026
Abstract PDF (2.6 MB) Collect
Downloads:10
Objective

Accurate detection of PIK3CA mutations is essential for guiding PI3K-targeted therapies in breast cancer, yet sequencing is not universally accessible, and single-modality prediction models have limited performance. This study developed a multimodal deep learning framework integrating whole-slide imaging (WSI) and structured clinical data to improve mutation prediction.

Methods

A total of 1,047 patients from TCGA and 166 patients from 3 external centers were included. The histopathology model used a transformer-based pretrained encoder (H-optimus-0) and a clustering-constrained attention multiple instance learning (CLAM-SB MIL) classifier to generate WSI-level representations. The clinical model incorporated engineered clinical variables and an extreme gradient boosting (XGBoost) model. A decision-level late fusion strategy (Multimodal PIK3CA Model, MPM) combined probabilistic outputs from both branches. Performance was evaluated with the area under the curve (AUC) and secondary metrics. Interpretability was assessed via attention heatmaps and shapley additive explanations (SHAP) analysis.

Results

MPM outperformed single-modality models. It achieved an AUC of 0.745 on TCGA and maintained stable performance across external cohorts (0.695, 0.690, and 0.680). SHAP analysis identified molecular subtype as the most influential clinical feature, whereas attention maps highlighted mutation-associated morphological regions.

Conclusions

The developed multimodal framework effectively integrates complementary morphological and clinical information, and provides a robust and generalizable method for predicting PIK3CA mutation status. Strong multicenter adaptability and biological interpretability support its potential use as a clinical decision-support tool and an accessible alternative to molecular testing.

Issue
Multimodal Interpretable Model for Predicting Pathological Complete Response After Neoadjuvant Therapy in Lung Cancer: a Multicenter Study
Medical Journal of Peking Union Medical College Hospital 2026, 17(4): 943-953
Published: 16 June 2026
Abstract PDF (3.8 MB) Collect
Downloads:1
Objective

To develop an multimodal interpretable model that integrates whole slide images(WSI) and clinical features, and to validate its efficacy in predicting pathological complete response(pCR) in lung cancer patients following neoadjuvant therapy.

Methods

The clinicopathologic data who received neoadjuvant therapy of patients as well as hematoxylin and eosin stained sections were retrospectively collected between March 2015 and March 2025. For the WSI branch, the predictive performance of five pathology foundation models(CTransPath, Virchow2, H-optimus-0, Phikon-Ⅴ2, and UNI-Ⅴ2) was compared within the CLAM-SB framework to identify the optimal feature extractor. The clinical branch was constructed using the extreme gradient boosting(XGBoost) model. Subsequently, a multimodal prediction model for pathologic complete response(MP-pCR) was established through decision-layer logistic regression fusion. Model performance was evaluated using the area under the receiver operating characteristic curve(AUC), accuracy, sensitivity, specificity, F1-score, and Brier score. Interpretability analysis was performed using attention heatmaps and the Shapley Additive Explanations(SHAP) algorithm.

Results

A total of 728 patients were enrolled in this study. Among them, a total of 676 patients from the Fourth Hospital of Hebei Medical University were selected as internal dataset, which was randomly divided into training set(n=536), validation set(n=70), and internal test set(n=70) at a ratio of 8∶1∶1. Additionally, 32 patients from the Affiliated Hospital of Hebei University and 20 from Handan First Hospital were selected as external validation set 1 and external validation set 2, respectively. In the internal testing set, UNI-Ⅴ2 emerged as the top-performing feature extractor for the WSI branch, achieving an AUC of 0.774(95% CI: 0.688-0.861). The resulting MP-pCR multimodal model yielded an AUC of 0.812(95% CI: 0.725-0.899) in the internal testing set, outperforming both the WSI-only model(0.774) and the clinical-only model(0.780), with an accuracy of 81.8%, sensitivity of 70.6%, specificity of 86.5%, and a Brier score of 0.125. In external validation set 1, the MP-pCR model achieved an AUC of 0.746(95% CI: 0.598-0.898) and an accuracy of 75.0%; in external validation set 2, it achieved an AUC of 0.722(95% CI: 0.538-0.918) and an accuracy of 80.0%. Attention heatmaps revealed that the model-focused regions were primarily concentrated within the tumor parenchyma and treatment-related peritumoral areas. SHAP analysis indicated that preoperative treatment regimen, pathological diagnosis, tumor-to-stroma ratio, age, and smoking history were the top five contributors to pCR prediction.

Conclusions

MP-pCR model developed and validated based on multicenter data outperforms single-modality models in predicting pCR for lung cancer. The interpretability results demonstrate good consistency with clinicopathologic knowledge, suggesting that the model holds promise as a potential auxiliary reference for the clinical evaluation of treatment response and the optimization of individualized therapeutic decisions.

Total 2