AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (450.8 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Review | Open Access

A Comprehensive Survey of Recent Transformers in Image, Video and Diffusion Models

Dinh Phu Cuong Le1,2Dong Wang1Viet-Tuan Le3( )
College of Computer Science and Electronic Engineering, Hunan University, Changsha, 410082, China
Faculty of Information Technology, Yersin University of Da Lat, Da Lat, 66100, Vietnam
Faculty of Information Technology, Ho Chi Minh City Open University, Ho Chi Minh City, 722000, Vietnam
Show Author Information

Abstract

Transformer models have emerged as dominant networks for various tasks in computer vision compared to Convolutional Neural Networks (CNNs). The transformers demonstrate the ability to model long-range dependencies by utilizing a self-attention mechanism. This study aims to provide a comprehensive survey of recent transformer-based approaches in image and video applications, as well as diffusion models. We begin by discussing existing surveys of vision transformers and comparing them to this work. Then, we review the main components of a vanilla transformer network, including the self-attention mechanism, feed-forward network, position encoding, etc. In the main part of this survey, we review recent transformer-based models in three categories: Transformer for downstream tasks, Vision Transformer for Generation, and Vision Transformer for Segmentation. We also provide a comprehensive overview of recent transformer models for video tasks and diffusion models. We compare the performance of various hierarchical transformer networks for multiple tasks on popular benchmark datasets. Finally, we explore some future research directions to further improve the field.

References

【1】
【1】
 
 
Computers, Materials & Continua
Pages 37-60

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Le DPC, Wang D, Le V-T. A Comprehensive Survey of Recent Transformers in Image, Video and Diffusion Models. Computers, Materials & Continua, 2024, 80(1): 37-60. https://doi.org/10.32604/cmc.2024.050790

779

Views

5

Downloads

18

Crossref

23

Web of Science

27

Scopus

Received: 17 February 2024
Accepted: 17 May 2024
Published: 18 July 2024
© The Author 2024.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.