AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.4 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Review | Open Access

Large Language Model Benchmarks in Medical Tasks

Lawrence K. Q. Yan1, Qian Niu2( ), Ming Li3, Yichao Zhang4, Caitlyn Heqi Yin5, Cheng Fei6, Benji Peng7, Ziqian Bi8, Pohsun Feng9, Keyu Chen3, Tianyang Wang10, Yunze Wang11, Silin Chen12, Ming Liu13, Junyu Liu2, Xinyuan Song14, Riyang Bao14 , Zekun Jiang15 , Ziyuan Qin14
Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology, Hong Kong, China
Graduate School of Medicine, Kyoto University, Kyoto, Japan
School of Computer Science, Georgia Institute of Technology, Atlanta, Georgia, USA
Department of Physics, The University of Texas at Dallas, Richardson, Texas, USA
Department of Statistics, University of Wisconsin–Madison, Madison, Wisconsin, USA
Department of Computer Science, Cornell University, Ithaca, New York, USA
AppCubic, Miami, Florida, USA
Department of Computer Science, Purdue University, West Lafayette, Indiana, USA
Department of Electrical Engineering, National Taiwan Normal University, Taipei, China
School of Computer Science and Informatics, University of Liverpool, Liverpool, UK
Biomedical Sciences, The University of Edinburgh, Edinburgh, UK
Department of Respiratory and Critical Care Medicine, The Second Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, China
Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, Indiana, USA
Department of Computer Science, Emory University, Atlanta, Georgia, USA
West China Biomedical Big Data Center, West China Hospital, Sichuan University, Chengdu, China
Show Author Information

Abstract

With the increasing application of large language models (LLMs) in the medical domain, evaluating these models' performance using benchmark datasets has become crucial. This paper presents a comprehensive survey of various benchmark datasets used in medical LLM tasks. These datasets span multiple modalities including text, image, and multimodal benchmarks, focusing on various aspects of medical knowledge such as electronic health records, doctor–patient dialogues, medical question answering, and medical image captioning. The survey categorizes the datasets by modality and examines their significance, data structure, and roles in model development and evaluation across tasks such as diagnostic support, report generation, and predictive decision support. Representative resources include Medical Information Mart for Intensive Care Ⅲ (MIMIC‐Ⅲ), MIMIC‐Ⅳ, BioASQ, PubMedQA, and CheXpert, which provide data and evaluation settings for research in clinical NLP, medical question answering, and chest‐radiograph interpretation. This paper summarizes the challenges and opportunities in leveraging these benchmarks for advancing multimodal medical intelligence, emphasizing the need for datasets with a greater degree of language diversity, structured omics data, and innovative approaches to synthesis. This synthesis is intended to inform future research on the applications of LLMs in medicine and medical artificial intelligence.

Graphical Abstract

This survey provides a comprehensive review of benchmark datasets for evaluating large language models in medical tasks, spanning text, image, and multimodal modalities, and covering key resources including MIMIC‐Ⅲ/Ⅳ, BioASQ, PubMedQA, and CheXpert. We identify critical limitations in current benchmarks—such as insufficient language diversity, demographic representation gaps, and model hallucination risks—and propose actionable recommendations alongside a sustainability framework for long‐term benchmark development to bridge the gap between academic evaluation and real‐world clinical deployment. BioASQ, biomedical semantic indexing; MIMIC, Medical Information Mart for Intensive Care; PubMed question answering.

References

【1】
【1】
 
 
Medicine Advances
Pages 316-341

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Yan LKQ, Niu Q, Li M, et al. Large Language Model Benchmarks in Medical Tasks. Medicine Advances, 2026, 4(3): 316-341. https://doi.org/10.1002/med4.70085

7

Views

0

Downloads

0

Crossref

Received: 08 November 2025
Revised: 10 February 2026
Accepted: 31 March 2026
Published: 19 September 2026
© 2026 The Author(s). Tsinghua University Press.

This is an open access article under the terms of the Creative Commons Attribution‐NonCommercial‐NoDerivs License, which permits use and distribution in any medium, provided the original work is properly cited, the use is non‐commercial and no modifications or adaptations are made.