AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.1 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

Multimodal feature interaction and semantic guided fusion for RGB-T population counting

Yong CHEN1,2( )Jiaojiao ZHANG1Ke DONG1
School of Electronic and Information Engineering,Lanzhou Jiaotong University,Lanzhou 730070,China
Gansu Provincial Engineering Research Center for Artificial Intelligence and Graphics & Image Processing,Lanzhou 730070,China
Show Author Information

Abstract

RGB-T mode crowd counting is designed to take advantage of the complementarity of visible RGB and thermal infrared image to achieve crowd counting. Aiming at the problems of insufficient information interaction between modes and insufficient feature fusion in the feature extraction of the RGB-T multimodal crowd counting method, an RGB-T crowd counting method based on multi-modal feature interaction and semantic guided fusion is proposed. Firstly, a stacked small scale convolution kernel is designed as a branch of the backbone network to extract the coarse features of each single mode. Secondly, in order to address the limited information interaction between the modes, a multi-modal feature interaction module is suggested. This module will extract the features of each RGB and thermal infrared mode and actualize the interactive features of the mode information. Then, a semantic-guided fusion module is designed to enhance the semantic relevance of multi-modal crowd features through global and local feature-guided fusion, so as to fully integrate multi-context information and improve the recognition ability of the target population. Finally, the regression head is used to generate the population density map and output the counting results. Experimental results demonstrate that the proposed method outperforms the comparison algorithms on the open RGBT-CC dataset, with a 31.12% reduction in the root-mean-square error value compared to the CMCRL method and higher accuracy for crowd counting under various scenarios.

CLC number: TP391.4 Document code: A Article ID: 1001-5965(2026)01-0028-10

References

【1】
【1】
 
 
Journal of Beijing University of Aeronautics and Astronautics
Pages 28-37

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
CHEN Y, ZHANG J, DONG K. Multimodal feature interaction and semantic guided fusion for RGB-T population counting. Journal of Beijing University of Aeronautics and Astronautics, 2026, 52(1): 28-37. https://doi.org/10.13700/j.bh.1001-5965.2023.0735

473

Views

8

Downloads

0

Crossref

0

Scopus

0

CSCD

Received: 08 November 2023
Published: 12 March 2024
© Journal of Beijing University of Aeronautics and Astronautics