AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.6 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Adjusted Reasoning Module for Deep Visual Question Answering Using Vision Transformer

Christine Dewi1,3Hanna Prillysca Chernovita2Stephen Abednego Philemon1Christian Adi Ananta1Abbott Po Shun Chen4( )
Department of Information Technology, Satya Wacana Christian University, Salatiga, 50711, Indonesia
Department of Information Systems, Satya Wacana Christian University, Salatiga, 50711, Indonesia
School of Information Technology, Deakin University, Burwood, VIC 3125, Australia
Department of Marketing and Logistics Management, Chaoyang University of Technology, Taichung City, 413310, Taiwan
Show Author Information

Abstract

Visual Question Answering (VQA) is an interdisciplinary artificial intelligence (AI) activity that integrates computer vision and natural language processing. Its purpose is to empower machines to respond to questions by utilizing visual information. A VQA system typically takes an image and a natural language query as input and produces a textual answer as output. One major obstacle in VQA is identifying a successful method to extract and merge textual and visual data. We examine “Fusion” Models that use information from both the text encoder and picture encoder to efficiently perform the visual question-answering challenge. For the transformer model, we utilize BERT and RoBERTa, which analyze textual data. The image encoder designed for processing image data utilizes ViT (Vision Transformer), Deit (Data-efficient Image Transformer), and BeIT (Image Transformers). The reasoning module of VQA was updated and layer normalization was incorporated to enhance the performance outcome of our effort. In comparison to the results of previous research, our proposed method suggests a substantial enhancement in efficacy. Our experiment obtained a 60.4% accuracy with the PathVQA dataset and a 69.2% accuracy with the VizWiz dataset.

References

【1】
【1】
 
 
Computers, Materials & Continua
Pages 4195-4216

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Dewi C, Chernovita HP, Philemon SA, et al. Adjusted Reasoning Module for Deep Visual Question Answering Using Vision Transformer. Computers, Materials & Continua, 2024, 81(3): 4195-4216. https://doi.org/10.32604/cmc.2024.057453

172

Views

6

Downloads

1

Crossref

1

Web of Science

0

Scopus

Received: 18 August 2024
Accepted: 01 November 2024
Published: 31 December 2024
© The Author 2024.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.