AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.1 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Review | Open Access

A Review on Vision-Language-Based Approaches: Challenges and Applications

Huu-Tuong Ho#,1Luong Vuong Nguyen#,1Minh-Tien Pham1Quang-Huy Pham1Quang-Duong Tran1Duong Nguyen Minh Huy2Tri-Hai Nguyen3( )
Department of Artificial Intelligence, FPT University, Danang, 550000, Vietnam
Department of Business, FPT University, Danang, 550000, Vietnam
Faculty of Information Technology, School of Technology, Van Lang University, Ho Chi Minh City, 70000, Vietnam

#These authors contributed equally to this work

Show Author Information

Abstract

In multimodal learning, Vision-Language Models (VLMs) have become a critical research focus, enabling the integration of textual and visual data. These models have shown significant promise across various natural language processing tasks, such as visual question answering and computer vision applications, including image captioning and image-text retrieval, highlighting their adaptability for complex, multimodal datasets. In this work, we review the landscape of Bootstrapping Language-Image Pre-training (BLIP) and other VLM techniques. A comparative analysis is conducted to assess VLMs’ strengths, limitations, and applicability across tasks while examining challenges such as scalability, data quality, and fine-tuning complexities. The work concludes by outlining potential future directions in VLM research, focusing on enhancing model interpretability, addressing ethical implications, and advancing multimodal integration in real-world applications.

References

【1】
【1】
 
 
Computers, Materials & Continua
Pages 1733-1756

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Ho H-T, Nguyen LV, Pham M-T, et al. A Review on Vision-Language-Based Approaches: Challenges and Applications. Computers, Materials & Continua, 2025, 82(2): 1733-1756. https://doi.org/10.32604/cmc.2025.060363

344

Views

6

Downloads

11

Crossref

5

Web of Science

14

Scopus

Received: 30 October 2024
Accepted: 20 January 2025
Published: 28 February 2025
© The Author 2024.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.