AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (8.2 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Cross-view remote sensing and street-level data fusion for intelligent traffic congestion analysis

Inzamam Mashood Nasir1( )Hend Alshaya2Sara Tehsin3Wided Bouchelligua2
Human-Environment-Technology (HET) Systems Centre, Mykolas Romeris University, Vilnius 08303, Lithuania
Applied College, Imam Mohammad Ibn Saud Islamic University (IMSIU), Riyadh 11432, Saudi Arabia
Faculty of Informatics, Kaunas University of Technology, 51368 Kaunas, Lithuania
Show Author Information

Abstract

The issue of urban traffic congestion is a persistent problem for the sustainable management of cities through transportation systems, as there is a need for models that integrate and analyze heterogeneous sources to yield accurate, interpretable outcomes. This paper introduces the cross-view fusion network (CVF-Net), a new multimodal deep learning framework for analyzing congestion across entire cities by combining remote-sensing imagery (drone aerial views), street-view camera images, and graph-structured sensor data into a single model. This model is introduced through a very novel architecture that includes a hierarchical attention fusion transformer (HAFT), which fuses cross-view attention (CVA) between the aerial and ground view, a temporal graph neural network (TGNN) that uses a spatio-temporal dynamic, and a graph refinement (GR) network for consistency relative to the graph topology. Extensive experiments across three benchmarks (CityFlowV2, METR-LA, PEMS-BAY) demonstrate that CVF-Net consistently outperforms other recent state-of-the-art methods, reducing forecasting error (MAE) by 9.3% and increasing tracking continuity (IDF1) by 7.0%. Ablation studies suggest that hierarchical fusion and temporal modeling improve accuracy and stability, while sensitivity analyses show that attention maps capture congestion and causal temporal patterns, which are real symptoms of congestion. The model also shows strong cross-dataset generalizability and robustness to sensor noise, which extends its performance in the real world. Unlike existing spatio-temporal GNNs and multimodal Transformers that rely on flat feature aggregation or implicitly assume cross-view alignment, the proposed framework introduces a hierarchical, alignment-aware fusion strategy that explicitly integrates aerial visual context with graph-temporal traffic dynamics.

CLC number: 68T07

References

【1】
【1】
 
 
AIMS Mathematics
Pages 1547-1589

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Nasir IM, Alshaya H, Tehsin S, et al. Cross-view remote sensing and street-level data fusion for intelligent traffic congestion analysis. AIMS Mathematics, 2026, 11(1): 1547-1589. https://doi.org/10.3934/math.2026065

338

Views

5

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 13 November 2025
Revised: 17 December 2025
Accepted: 22 December 2025
Published: 19 January 2026
©2026 the Author(s), licensee AIMS Press.

This is an open access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0)