AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.3 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Intelligent Ophthalmology | Open Access

Bilingual evaluation of large language models for patient education in refractive surgery

Tsung-Hsien Tsai1,2Chin-Ling Tsai3Jui-Hung Hsu4Ching-Hsi Hsiao2,5Hung-Chi Chen2,5( )
Department of Ophthalmology, Chang Gung Memorial Hospital, Keelung 204, Taiwan, China
School of Medicine, College of Medicine, Chang Gung University, Taoyuan 333, Taiwan, China
Shanghai Aier Eye Hospital, Shanghai 200031, China
Department of Medical Education, Chang Gung Memorial Hospital, Keelung 204, Taiwan, China
Department of Ophthalmology, Chang Gung Memorial Hospital, Linkou 333, Taiwan, China

Co-first Authors: Tsung-Hsien Tsai and Chin-Ling Tsai

Show Author Information

Abstract

AIM

To evaluate the ability of six advanced large language models (LLMs)—in providing accurate, comprehensive, and readable patient education on corneal refractive surgeries [laser in-situ keratomileusis (LASIK), keratorefractive lenticule extraction (KLEx), and photorefractive keratectomy (PRK)] in both English and Chinese.

METHODS

This is a cross-sectional, comparative study. Twenty-six questions, compiled from authoritative ophthalmologic sources and covering four domains (procedure basics and eligibility; safety, risks and long-term stability; recovery and postoperative experience; and practical concerns), were administered in both English and Chinese via fresh chat sessions with each LLM, respectively. Five performance metrics were evaluated: accuracy, comprehensiveness, word count, readability, and reproducibility, using appropriate statistical tests.

RESULTS

OpenAI o1 and DeepSeek-R1 consistently achieved the highest accuracy and most comprehensive responses, significantly outperforming ChatGPT-4o, Gemini Advanced, Claude Sonnet, and Tongyi Qwen (Friedman P<0.001). Although overall accuracy and comprehensiveness were similar across languages, Chinese responses were significantly longer. Readability varied among the models, with Claude Sonnet generally producing the most readable English texts. Reproducibility analysis revealed moderate consistency, reflecting inherent variability in outputs to identical prompts.

CONCLUSION

Reasoning-augmented LLMs, particularly OpenAI o1 and DeepSeek-R1, demonstrate superior performance in delivering bilingual patient education for corneal refractive surgery, with high accuracy and comprehensiveness. However, variations in response length, readability, and reproducibility indicate that further refinement is necessary before these tools can be reliably integrated into clinical practice.

References

【1】
【1】
 
 
International Journal of Ophthalmology
Pages 1019-1027

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Tsai T-H, Tsai C-L, Hsu J-H, et al. Bilingual evaluation of large language models for patient education in refractive surgery. International Journal of Ophthalmology, 2026, 19(6): 1019-1027. https://doi.org/10.18240/ijo.2026.06.01

253

Views

12

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 15 November 2025
Accepted: 28 January 2026
Published: 18 June 2026
© 2026 International Journal of Ophthalmology Press

This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).