AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.8 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

Fine-grained semantic-enhanced cross-modal image-text retrieval method for civil aviation

Shuyan LIU1,2, Liu HE1,2,3( ), Jianghui ZENG1,2
Department of Standard and Data Technology Research,China Aero-Polytechnology Establishment,Beijing 100028,China
Industrial Internet Application Innovation Center,China Aero-Polytechnology Establishment,Beijing 100028,China
School of Computer Science,Fudan University,Shanghai 200438,China
Show Author Information

Abstract

In the fields of low-altitude economy and civil aviation, security assurance tasks heavily rely on the efficient correlation of cross-modal information such as images and texts. However, while mainstream cross-modal retrieval models perform well on general datasets, they underperform in these areas, which require high levels of fine-grained semantic understanding. Based on existing civil aviation datasets, a cross-modal retrieval approach with fine-grained semantic augmentation is suggested as a solution to this problem, creating a whole pipeline that includes data processing, model development, and training. First, text descriptions are optimized and enhanced based on a large multimodal model to construct a cross-modal retrieval dataset containing rich semantic information. Second, a model is created using a popular cross-modal retrieval framework. To improve the model's ability to express fine-grained semantic features, techniques such class supervision, key semantic information masking, and a fine-grained feature extraction module are introduced. Experimental results on two datasets verify the effectiveness of the proposed method, providing a reference technical path for cross-modal retrieval in low-altitude economy scenarios.

CLC number: V219;TP391.3 Document code: A Article ID: 1001-5965(2026)09-3075-14

References

【1】
【1】
 
 
Journal of Beijing University of Aeronautics and Astronautics
Pages 3075-3088

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
LIU S, HE L, ZENG J. Fine-grained semantic-enhanced cross-modal image-text retrieval method for civil aviation. Journal of Beijing University of Aeronautics and Astronautics, 2026, 52(9): 3075-3088. https://doi.org/10.13700/j.bh.1001-5965.2025.0549

1

Views

0

Downloads

0

Crossref

0

Scopus

0

CSCD

Received: 06 August 2025
Published: 09 October 2025
© Journal of Beijing University of Aeronautics and Astronautics