TY - JOUR AU - Yan, Lawrence K. Q. AU - Niu, Qian AU - Li, Ming AU - Zhang, Yichao AU - Yin, Caitlyn Heqi AU - Fei, Cheng AU - Peng, Benji AU - Bi, Ziqian AU - Feng, Pohsun AU - Chen, Keyu AU - Wang, Tianyang AU - Wang, Yunze AU - Chen, Silin AU - Liu, Ming AU - Liu, Junyu AU - Song, Xinyuan AU - Bao, Riyang AU - Jiang, Zekun AU - Qin, Ziyuan PY - 2026 TI - Large Language Model Benchmarks in Medical Tasks JO - Medicine Advances SN - 2834-4391 SP - 316 EP - 341 VL - 4 IS - 3 AB - With the increasing application of large language models (LLMs) in the medical domain, evaluating these models' performance using benchmark datasets has become crucial. This paper presents a comprehensive survey of various benchmark datasets used in medical LLM tasks. These datasets span multiple modalities including text, image, and multimodal benchmarks, focusing on various aspects of medical knowledge such as electronic health records, doctor–patient dialogues, medical question answering, and medical image captioning. The survey categorizes the datasets by modality and examines their significance, data structure, and roles in model development and evaluation across tasks such as diagnostic support, report generation, and predictive decision support. Representative resources include Medical Information Mart for Intensive Care Ⅲ (MIMIC‐Ⅲ), MIMIC‐Ⅳ, BioASQ, PubMedQA, and CheXpert, which provide data and evaluation settings for research in clinical NLP, medical question answering, and chest‐radiograph interpretation. This paper summarizes the challenges and opportunities in leveraging these benchmarks for advancing multimodal medical intelligence, emphasizing the need for datasets with a greater degree of language diversity, structured omics data, and innovative approaches to synthesis. This synthesis is intended to inform future research on the applications of LLMs in medicine and medical artificial intelligence. UR - https://doi.org/10.1002/med4.70085 DO - 10.1002/med4.70085