AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.7 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

Applicability Assessment and Comparison of Clinical Reasoning Outputs from General-Purpose Large Language Models in Medical Education

He XU1,2,3, Linxuan ZHAI3, Yalun ZHU4, Hengfu CUI4, Zuyi ZHU4( ), Yang JIAO2( )
Department of Neurology & Innovation Center for Neurological Disorders, Xuanwu Hospital, Capital Medical University, National Center for Neurological Disorders, Beijing 100053, China
Department of General Practice (General Internal Medicine), Peking Union Medical College Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College, Beijing 100730, China
4+4 MD Program, Chinese Academy of Medical Sciences & Peking Union Medical College, Beijing 100730, China
Baichuan AI Technology Co., Ltd, Beijing 100084, China
Show Author Information

Abstract

Objective

To compare the diagnostic accuracy and differential diagnosis comprehensiveness of six general-purpose large language models (LLMs) on standardized clinical cases, and to assess their applicability for clinical reasoning education.

Methods

A total of 74 standardized clinical cases published in the BMJ "Endgames" column between March 2023 and February 2025 were included. Six general-purpose LLMs (ChatGPT-4, ChatGPT-4o, DeepSeek-V3, DeepSeek-R1, Gemini-2.0, and Kimi-k1.5) were prompted with identical text-only case information via their application programming interfaces (APIs). Diagnostic accuracy and differential diagnosis comprehensiveness were evaluated and compared across models. For statistical analysis, Cochran's Q test was used for overall comparisons of diagnostic accuracy, with post-hoc pairwise comparisons performed using McNemar tests; the Friedman test was used for overall comparisons of differential diagnosis comprehensiveness, with post-hoc pairwise comparisons performed using Wilcoxon signed-rank tests; all P-values were adjusted using the Bonferroni method. Grade distribution differences were analyzed using Pearson's chi-square test.

Results

Diagnostic accuracy ranged from 62.2% to 78.4% across the six models, with a statistically significant overall difference among models (P=0.022); however, no pairwise comparison remained significant after Bonferroni correction (all P > 0.05). Stratified analysis showed that accuracy was higher for cases without images (63.3%-81.6%) than for those with images (52.0%-72.0%), with no significant differences among models in either subgroup (all P > 0.05). Regarding differential diagnosis comprehensiveness, DeepSeek-R1 achieved the highest mean coverage rate (61.2%), while ChatGPT-4o scored the lowest (50.3%), with a statistically significant overall difference among models (χ2=16.6, P=0.005). Post-hoc pairwise comparisons revealed that ChatGPT-4o was significantly inferior to Gemini-2.0 (adjusted r=0.436, P=0.009). Regarding grade distribution, approximately 70%-80% of cases achieved only "moderate" or "limited" comprehensiveness, with no significant difference among models (χ2=17.90, P=0.268).

Conclusions

Current general-purpose LLMs can achieve a moderate level of diagnostic accuracy on standardized clinical cases, but demonstrate insufficient differential diagnosis comprehensiveness. Given that diagnostic accuracy and differential diagnosis comprehensiveness are fundamental elements of clinical reasoning education, these findings suggest that the quality of LLMs-generated outputs is not yet sufficient for direct application in teaching scenarios and should be used with caution under teacher supervision.

CLC number: R5;G642 Document code: A Article ID: 1674-9081(2026)05-1250-06

References

【1】
【1】
 
 
Medical Journal of Peking Union Medical College Hospital
Pages 1250-1255

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
XU H, ZHAI L, ZHU Y, et al. Applicability Assessment and Comparison of Clinical Reasoning Outputs from General-Purpose Large Language Models in Medical Education. Medical Journal of Peking Union Medical College Hospital, 2026, 17(5): 1250-1255. https://doi.org/10.12290/xhyxzz.2026-0716

5

Views

0

Downloads

0

Crossref

0

Scopus

0

CSCD

Received: 25 May 2026
Accepted: 18 August 2026
Published: 02 September 2026
© 2026 Medical Journal of Peking Union Medical College Hospital