In crowd counting, traditional methods often struggle to balance preserving local details and capturing global context due to the high variability in crowd density and occlusion levels within scenes, leading to performance degradation in highly congested or severely perspective-distorted images. To address this issue, this paper proposes a dense-scene crowd counting method based on deep spectral adaptive modulation (DSAMNet), aiming to simultaneously enhance the spatial resolution of density map estimation and global counting consistency. Specifically, the method first employs a depth-aware module to jointly extract multi-scale texture features and coarse depth estimates, utilizing predicted depth to guide perspective correction and achieve geometrically consistent feature sampling. Subsequently, a spatially adaptive high-frequency encoding module is introduced, which dynamically modulates coordinate frequencies based on local depth during the encoding process, thereby enhancing the positional representation capacity in geometrically sensitive regions and improving the network's adaptability to scale and structural variations. Finally, an implicit density decoding module integrates visual and geometric representations through a cross-domain attention mechanism and performs point-wise density regression via a hierarchically conditioned multi-layer perceptron, achieving high-quality reconstruction of continuous density fields. Experimental results demonstrate that the proposed method achieves superior performance on several mainstream crowd counting datasets, exhibiting strong robustness and generalization capability in complex perspective scenarios.
Publications
- Article type
- Year
- Co-author
Year
Open Access
Issue
Journal of Northwest University (Natural Science Edition) 2026, 56(3): 572-582
Published: 25 June 2026
Downloads:0
Total 1
京公网安备11010802044758号