AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

Multi-Task Visual Semantic Embedding Network for Image-Text Retrieval

School of Computer Science and Technology, Dalian University of Technology, Dalian 116024, China
School of Computer Science, Shaanxi Normal University, Xi’an 710119, China
School of Computer Engineering, Weifang University, Weifang 261061, China
Guangxi Colleges and Universities Key Laboratory of Intelligent Industry Software, Wuzhou University, Wuzhou 543002 China
Show Author Information

Abstract

Image-text retrieval aims to capture the semantic correspondence between images and texts, which serves as a foundation and crucial component in multi-modal recommendations, search systems, and online shopping. Existing mainstream methods primarily focus on modeling the association of image-text pairs while neglecting the advantageous impact of multi-task learning on image-text retrieval. To this end, a multi-task visual semantic embedding network (MVSEN) is proposed for image-text retrieval. Specifically, we design two auxiliary tasks, including text-text matching and multi-label classification, for semantic constraints to improve the generalization and robustness of visual semantic embedding from a training perspective. Besides, we present an intra- and inter-modality interaction scheme to learn discriminative visual and textual feature representations by facilitating information flow within and between modalities. Subsequently, we utilize multi-layer graph convolutional networks in a cascading manner to infer the correlation of image-text pairs. Experimental results show that MVSEN outperforms state-of-the-art methods on two publicly available datasets, Flickr30K and MSCOCO, with rSum improvements of 8.2% and 3.0%, respectively.

Electronic Supplementary Material

Download File(s)
JCST-2401-14125-Highlights.pdf (157.4 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 811-826

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Qin X-Y, Li L-S, Tang J-Y, et al. Multi-Task Visual Semantic Embedding Network for Image-Text Retrieval. Journal of Computer Science and Technology, 2024, 39(4): 811-826. https://doi.org/10.1007/s11390-024-4125-1

974

Views

6

Crossref

5

Web of Science

7

Scopus

1

CSCD

Received: 16 January 2024
Accepted: 20 June 2024
Published: 20 September 2024
© Institute of Computing Technology, Chinese Academy of Sciences 2024