In recent years, visual relocalization has emerged as a pivotal task in the domains of 3D computer vision and learning based methodologies, witnessing substantial advances due to the evolution of learning based methodologies. This article reviews the current landscape and recent developments in visual relocalization research. It methodically discusses visual relocalization tasks, delineates fundamental solution methodologies, categorizes existing studies, and outlines research objectives within this field. By systematically organizing and elucidating the stateof-the-art of visual relocalization through the lens of learning based methodologies, this paper aims to provide a comprehensive analysis to aid researchers to swiftly grasp the essence of the research problem. It offers a lucid overview of the specific advances in various research directions, thereby facilitating effective applications and further investigation. Additionally, this article anticipates future research trajectories to address visual relocalization challenges.
- Article type
- Year
- Co-author
Open Access
Review Article
Issue
Open Access
Research Article
Issue
Crowd counting provides an important foundation for public security and urban management. Due to the existence of small targets and large den-sity variations in crowd images, crowd counting is a challenging task. Mainstream methods usually apply convolution neural networks (CNNs) to regress a density map, which requires annotations of individual persons and counts. Weakly-supervised methods can avoid detailed labeling and only require counts as annotations of images, but existing methods fail to achieve satisfactory performance because a global perspective field and multi-level information are usually ignored. We propose a weakly-supervised method, DTCC, which effectively combines multi-level dilated convolution and transformer methods to realize end-to-end crowd counting. Its main components include a recursive swin transformer and a multi-level dilated convolution regression head. The recursive swin trans-former combines a pyramid visual transformer with a fine-tuned recursive pyramid structure to capture deep multi-level crowd features, including global features. The multi-level dilated convolution regression head includes multi-level dilated convolution and a linear regression head for the feature extraction module. This module can capture both low- and high-level features simultaneously to enhance the receptive field. In addition, two regression head fusion mechanisms realize dynamic and mean fusion counting. Experiments on four well-known benchmark crowd counting datasets (UCF_CC_50, ShanghaiTech, UCF_QNRF, and JHU-Crowd++) show that DTCC achieves results superior to other weakly-supervised methods and comparable to fully-supervised methods.
京公网安备11010802044758号