AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Review | Open Access

Efficient multimodal large language models: a survey

Yizhang Jin1,2, Jian Li2, Tianjun Gu3, Yexin Liu4Bo Zhao4Jinxiang Lai5Zhenye Gan2 Yabiao Wang2 Chengjie Wang2 Xin Tan3 Lizhuang Ma1 ( )
Shanghai Jiao Tong University, Shanghai, China
Tencent, Shanghai, China
East China Normal University, Shanghai, China
Beijing Academy of Artificial Intelligence, Beijing, China
Hong Kong University of Science and Technology, Hong Kong, China

Equal contributors

Show Author Information

Abstract

In the past years, multimodal large language models (MLLMs) have demonstrated remarkable performance in tasks such as visual question answering and visual understanding and reasoning. However, the extensive model size and high training and inference costs have hindered the widespread application of MLLMs in academia and industry. Thus, studying efficient and lightweight MLLMs has enormous potential, especially in edge computing scenarios. In this survey, we provide a comprehensive and systematic review of the current state of efficient MLLMs. Specifically, this survey summarizes the timeline of representative efficient MLLMs, the current state of research in structures and strategies, and the applications. Finally, the limitations of current efficient MLLM research and promising future directions are discussed.

References

【1】
【1】
 
 
Visual Intelligence
Article number: 27

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Jin Y, Li J, Gu T, et al. Efficient multimodal large language models: a survey. Visual Intelligence, 2025, 3: 27. https://doi.org/10.1007/s44267-025-00099-6

1453

Views

41

Crossref

Received: 24 April 2025
Revised: 19 November 2025
Accepted: 20 November 2025
Published: 28 February 2026
© The Author(s) 2025.

This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.