Abstract
Vehicle-infrastructure cooperative perception is critical for advanced autonomous driving but faces challenges: heavy reliance on ground-truth (GT) vehicle-infrastructure poses for spatial synchronization and limitations of traditional object detection in dynamic scene modeling. This paper proposes a differentiable pose-based spatial alignment (DPSA) module and explores occupancy flow output for cooperative perception. The DPSA module eliminates GT pose requirements by estimating 3-degree-of-freedom relative poses via feature fusion, spatial average pooling, and fully connected layers, balancing accuracy and practicality. Evaluations on DAIR-V2X and V2X-Seq datasets show DPSA outperforms misaligned fusion methods in key metrics, maintains robustness under low-location-precision scenarios, and reduces inference time. Notably, occupancy flow demonstrates superior dynamic modeling capability compared to traditional object detection by capturing spatial occupancy changes and motion flows (instead of bounding boxes), enabling high-accuracy future occupancy prediction and occlusion resistance via multi-source temporal context. This work advances practical cooperative perception through an efficient synchronization solution and highlights occupancy flow advantages, bridging algorithmic innovation and engineering deployment.
京公网安备11010802044758号
Comments on this article