AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research | Open Access

Uncertainty-aware coarse-to-fine alignment for text-image person retrieval

Yifei Deng1Zhengyu Chen2Chenglong Li2( )Jin Tang1
Information Materials and Intelligent Sensing Laboratory of Anhui Province, Anhui Provincial Key Laboratory of Multimodal Cognitive Computation, School of Computer Science and Technology, Anhui University, Hefei, 230601, China
Information Materials and Intelligent Sensing Laboratory of Anhui Province, Anhui Provincial Key Laboratory of Security Artificial Intelligence, School of Artificial Intelligence, Anhui University, Hefei, 230601, China
Show Author Information

Abstract

Text-to-image person retrieval, a fine-grained cross-modal retrieval problem, aims to search for person images from an image library that match a given textual caption. Existing text-to-image person retrieval methods usually use fixed-point embedding to express the semantics of the two modalities and perform multi-granularity alignment between modalities in the embedding space. However, owing to the inherent mutual one-to-many correspondence between images and texts, it is often difficult for fixed-point embedding methods to adequately capture this relationship, leading to erroneous retrieval results. To address this problem, we propose a novel uncertainty-aware coarse-to-fine alignment method, which first maps fixed-point embedding to probability distributions and then aligns two modalities in terms of distributions and sampling points at a coarse-to-fine granularity, for accurate text-to-image person retrieval. Specifically, we first introduce two contrastive learning tasks of distribution contrast learning and point contrast learning, to achieve coarse-grained inter-modal alignment with uncertainty-aware. The distribution contrast learning task ensures that distributions with the same identity are as similar as possible across modalities through distribution-based contrastive learning. The point contrast learning task performs the contrastive learning of inter-modal and intra-modal sampling points, which not only models rich and diverse cross-modal associations, but also optimizes the learning of distributions. For the fine-grained association requirements of text-to-image person retrieval, we design the task of uncertainty-aware attribute masking language reconstruction, which achieves fine-grained alignment by randomly masking attribute words in the text and reconstructing them via inter-modal sample point interactions. Extensive experiments on two public datasets demonstrate the superior performance of our method.

References

【1】
【1】
 
 
Visual Intelligence
Article number: 6

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Deng Y, Chen Z, Li C, et al. Uncertainty-aware coarse-to-fine alignment for text-image person retrieval. Visual Intelligence, 2025, 3: 6. https://doi.org/10.1007/s44267-025-00078-x

1483

Views

12

Crossref

Received: 05 September 2024
Revised: 15 April 2025
Accepted: 16 April 2025
Published: 29 April 2025
© The Author(s) 2025.

This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.