Synergistic analysis of drone-captured imagery and fixed surveillance video enables continuous tracking of targets in inspection tasks, achieving cross-view person re-identification in areas where surveillance probes are sparsely distributed. The transferability of image retrieval algorithms created for conventional surveillance situations to aerial-ground cross-view environments is limited by the notable discrepancy between the horizontal view of ground surveillance and the bird’s-eye view of drones. Existing aerial-ground pedestrian retrieval methods primarily focus on mitigating the appearance discrepancies caused by cross-view variations, while the mining and analysis of person attribute characteristics remain insufficiently explored. To address these issues, this paper proposes a aerial-ground pedestrian retrieval method based on large model attribute parsing. A person parsing module is constructed based on a multi-modal large model to generate fine-grained semantic attributes, and an attribute triplet loss is designed to achieve cross-view semantic alignment. A view decoupling architecture is introduced to separate view-specific features through hierarchical subtraction, with an orthogonal loss applied to constrain feature independence. A multi-scale dilated Transformer is incorporated, combining multi-scale dilated attention with global self-attention to optimize the balance between computational complexity and receptive field, thereby reducing model parameters. The effectiveness of the methodology is confirmed by experiments on the AG-ReID.v1, AG-ReID.v2, and CARGO datasets, which show that the suggested strategy successfully increases performance on Rank-1, mAP, and mINP metrics in aerial-ground pedestrian retrieval tasks.
Publications
- Article type
- Year
Year
Issue
Journal of Beijing University of Aeronautics and Astronautics 2026, 52(9): 3108-3116
Published: 03 April 2026
Downloads:0
Total 1
京公网安备11010802044758号