AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.2 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

Large model attribute parsing-based aerial-ground pedestrian retrieval method

Quange TAN1, Rong WANG1, Mancheng LIAO2, Xin LI1( )
College of Information Network Security,People’s Public Security University of China,Beijing 100038,China
College of Criminal Investigation,People’s Public Security University of China,Beijing 100038,China
Show Author Information

Abstract

Synergistic analysis of drone-captured imagery and fixed surveillance video enables continuous tracking of targets in inspection tasks, achieving cross-view person re-identification in areas where surveillance probes are sparsely distributed. The transferability of image retrieval algorithms created for conventional surveillance situations to aerial-ground cross-view environments is limited by the notable discrepancy between the horizontal view of ground surveillance and the bird’s-eye view of drones. Existing aerial-ground pedestrian retrieval methods primarily focus on mitigating the appearance discrepancies caused by cross-view variations, while the mining and analysis of person attribute characteristics remain insufficiently explored. To address these issues, this paper proposes a aerial-ground pedestrian retrieval method based on large model attribute parsing. A person parsing module is constructed based on a multi-modal large model to generate fine-grained semantic attributes, and an attribute triplet loss is designed to achieve cross-view semantic alignment. A view decoupling architecture is introduced to separate view-specific features through hierarchical subtraction, with an orthogonal loss applied to constrain feature independence. A multi-scale dilated Transformer is incorporated, combining multi-scale dilated attention with global self-attention to optimize the balance between computational complexity and receptive field, thereby reducing model parameters. The effectiveness of the methodology is confirmed by experiments on the AG-ReID.v1, AG-ReID.v2, and CARGO datasets, which show that the suggested strategy successfully increases performance on Rank-1, mAP, and mINP metrics in aerial-ground pedestrian retrieval tasks.

CLC number: TP391.4;V279 Document code: A Article ID: 1001-5965(2026)09-3108-09

References

【1】
【1】
 
 
Journal of Beijing University of Aeronautics and Astronautics
Pages 3108-3116

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
TAN Q, WANG R, LIAO M, et al. Large model attribute parsing-based aerial-ground pedestrian retrieval method. Journal of Beijing University of Aeronautics and Astronautics, 2026, 52(9): 3108-3116. https://doi.org/10.13700/j.bh.1001-5965.2025.0834

1

Views

0

Downloads

0

Crossref

0

Scopus

0

CSCD

Received: 02 December 2025
Published: 03 April 2026
© Journal of Beijing University of Aeronautics and Astronautics