Publications
Sort:
Issue
Applicability Assessment and Comparison of Clinical Reasoning Outputs from General-Purpose Large Language Models in Medical Education
Medical Journal of Peking Union Medical College Hospital 2026, 17(5): 1250-1255
Published: 02 September 2026
Abstract PDF (1.7 MB) Collect
Downloads:0
Objective

To compare the diagnostic accuracy and differential diagnosis comprehensiveness of six general-purpose large language models (LLMs) on standardized clinical cases, and to assess their applicability for clinical reasoning education.

Methods

A total of 74 standardized clinical cases published in the BMJ "Endgames" column between March 2023 and February 2025 were included. Six general-purpose LLMs (ChatGPT-4, ChatGPT-4o, DeepSeek-V3, DeepSeek-R1, Gemini-2.0, and Kimi-k1.5) were prompted with identical text-only case information via their application programming interfaces (APIs). Diagnostic accuracy and differential diagnosis comprehensiveness were evaluated and compared across models. For statistical analysis, Cochran's Q test was used for overall comparisons of diagnostic accuracy, with post-hoc pairwise comparisons performed using McNemar tests; the Friedman test was used for overall comparisons of differential diagnosis comprehensiveness, with post-hoc pairwise comparisons performed using Wilcoxon signed-rank tests; all P-values were adjusted using the Bonferroni method. Grade distribution differences were analyzed using Pearson's chi-square test.

Results

Diagnostic accuracy ranged from 62.2% to 78.4% across the six models, with a statistically significant overall difference among models (P=0.022); however, no pairwise comparison remained significant after Bonferroni correction (all P > 0.05). Stratified analysis showed that accuracy was higher for cases without images (63.3%-81.6%) than for those with images (52.0%-72.0%), with no significant differences among models in either subgroup (all P > 0.05). Regarding differential diagnosis comprehensiveness, DeepSeek-R1 achieved the highest mean coverage rate (61.2%), while ChatGPT-4o scored the lowest (50.3%), with a statistically significant overall difference among models (χ2=16.6, P=0.005). Post-hoc pairwise comparisons revealed that ChatGPT-4o was significantly inferior to Gemini-2.0 (adjusted r=0.436, P=0.009). Regarding grade distribution, approximately 70%-80% of cases achieved only "moderate" or "limited" comprehensiveness, with no significant difference among models (χ2=17.90, P=0.268).

Conclusions

Current general-purpose LLMs can achieve a moderate level of diagnostic accuracy on standardized clinical cases, but demonstrate insufficient differential diagnosis comprehensiveness. Given that diagnostic accuracy and differential diagnosis comprehensiveness are fundamental elements of clinical reasoning education, these findings suggest that the quality of LLMs-generated outputs is not yet sufficient for direct application in teaching scenarios and should be used with caution under teacher supervision.

Issue
Artificial Intelligence in General Practice: A Bibliometric Analysis
Medical Journal of Peking Union Medical College Hospital 2025, 16(3): 687-696
Published: 30 May 2025
Abstract PDF (5.3 MB) Collect
Downloads:37
Objective

To analyze the current status, research hotspots, and evolving trends in the application of artificial intelligence (AI) in general practice over the past two decades using bibliometric methods, thereby providing insights for future research.

Methods

A systematic search was conducted in the Web of Science Core Collection (WoSCC) to identify English-language literature on AI applications in general practice published between January 1, 2004, and December 31, 2024. VOSviewer 1.6.19 was used for co-occurrence analysis of countries/regions (≥5 publications) and authors (≥3 publications), while Scimago Graphica 1.0.46 was employed to visualize collaboration networks among countries/regions. CiteSpace 6.2.R2 was utilized for institutional co-occurrence analysis (≥5 publications), keyword co-occurrence, and clustering analysis.

Results

A total of 394 relevant articles were included (307 original research articles and 87 reviews). The annual publication output showed a gradual increase, with a notable surge after 2018. The United States and the United Kingdom ranked first in terms of publication volume (81 articles, 20.56%) and total citations (2077 citations), respectively. The University of London was the most prolific institution (20 articles, 5.08%), while Kueper from Western University was the most productive author (6 articles, 1.52%). Journal of Medical Internet Research had the highest number of publications (12 articles, 3.05%) and citations (317 citations) in this field. High-frequency keywords included "diagnosis" (31 occurrences), "management" (29), "risk" (28), "electronic health records" (24), and "care" (23), which were categorized into 7 clusters and further summarized into 4 major research themes. AI applications in health management, diagnostic assistance, and general practitioner training emerged as core research areas, while technological acceptability remained a key challenge.

Conclusions

AI technologies are being progressively integrated into general practice, shifting from fulfilling basic functional needs toward enhancing personalized and intelligent healthcare experiences. Future research should focus on improving AI acceptance among general practitioners and patients through strengthened international and institutional collaboration.

Total 2