TY - JOUR AU - Wang, Yinuo AU - Chen, Kai AU - Zeng, Yue AU - Meng, Cai AU - Pan, Chao AU - Tang, Zhouping PY - 2025 TI - Zero-shot multi-modal large language models v.s. supervised deep learning: A comparative analysis on CT-based intracranial hemorrhage subtyping JO - Brain Hemorrhages SN - 2589-238X SP - 323 EP - 330 VL - 6 IS - 6 AB - ObjectiveAccurate identification of intracranial hemorrhage (ICH) subtypes on non-contrast CT is crucial for prognosis and treatment but remains challenging due to low contrast and blurred boundaries. This study evaluates the zero-shot performance of multi-modal large language models (MLLMs) versus traditional deep learning in ICH detection and subtyping.MethodsUsing 192 NCCT volumes from the RSNA dataset, we compared MLLMs (GPT-4o, Gemini 2.0 Flash, Claude 3.5 Sonnet V2) with deep learning models (ResNet50, Vision Transformer). MLLMs were prompted for ICH presence, subtype, localization, and volume estimation.ResultsTraditional deep learning models outperformed MLLMs in both ICH detection and subtyping. For subtyping, MLLMs showed lower accuracy, with Gemini 2.0 Flash achieving a macro-averaged precision of 0.41 and F1 score of 0.31.ConclusionWhile MLLMs offer enhanced interpretability through language-based interaction, their accuracy in ICH subtyping remains inferior to deep learning networks. Further optimization is needed to improve their utility in three-dimensional medical imaging. UR - https://doi.org/10.1016/j.hest.2025.10.004 DO - 10.1016/j.hest.2025.10.004