Collaborative perception allows real-time inter-agent information exchange and thus offers invaluable opportunities to enhance the perception capabilities of individual agents. However, limited communication bandwidth in practical scenarios restricts the inter-agent data transmission volume. This implies a trade-off between perception performance and communication cost. To address this issue, we propose Which2comm, a novel multi-agent 3D object detection framework leveraging object-level sparse features. By integrating semantic information of objects into detection boxes, we introduce semantic detection boxes (SemDBs). Innovatively transmitting these object-level sparse features among agents not only significantly reduces the demanding communication volume, but also improves object detection performance. Moreover, an adaptive strategy is further proposed to select only safety-critical connected and automated vehicles (CAVs) for collaborative perception when there are multiple CAVs available, thereby maintaining stable communication costs. To validate the proposed method, a large-scale, multi-modal dataset, Multi-V2X, is established for vehicle-to-everything (V2X) perception tasks with various CAV penetration rates. Multi-V2X comprises 146k frames with over 4.2 million 3D annotations, featuring high agent density (up to 31 agents) to evaluate perception robustness in complex traffic environments. Extensive experiments demonstrate that Which2comm consistently outperformed other state-of-the-art methods on both detection performance and communication cost, exhibiting superior robustness to real-world latency.
Publications
- Article type
- Year
- Co-author
Article type
Year
Open Access
Research Article
Just Accepted
Communications in Transportation Research
Available online: 17 August 2026
Downloads:16
Total 1
京公网安备11010802044758号