Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Collaborative perception allowed real-time interagent information exchange and thus offered invaluable opportunities to enhance the perception capabilities of individual agents. However, limited communication bandwidth in practical scenarios restricted the interagent data transmission volume. This implied a trade-off between perception performance and communication cost. To address this issue, we proposed Which2comm, a novel multiagent three-dimensional (3D) object detection framework leveraging object-level sparse features. By integrating semantic information of objects into detection boxes, we introduced semantic detection boxes (SemDBs). Innovatively transmitting these object-level sparse features among agents not only significantly reduces the demanding communication volume but also improves object detection performance. Moreover, an adaptive strategy was further proposed to select only safety critical connected and automated vehicles (CAVs) for collaborative perception when there were multiple CAVs available, thereby maintaining stable communication costs. To validate the proposed method, a large-scale, multimodal dataset, Multi-V2X, is established for vehicle-to-everything perception tasks with various CAV penetration rates. Multi-V2X comprises 1.46 ×105 frames with over 4.2 × 106 3D annotations, featuring high agent density (up to 31 agents) to evaluate perception robustness in complex traffic environments. Extensive experiments demonstrate that Which2comm consistently outperforms other state-of-the-art methods on both detection performance and communication cost, exhibiting superior robustness to real-world latency.

This is an open access article under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0 http://creativecommons.org/licenses/by/4.0/).
Comments on this article