Accurate and efficient land-use and land-cover (LULC) classification from remote sensing imagery remains challenging. This is because it requires capturing long-range spatial dependencies while maintaining computational scalability. Recent transformer-based models improve global context modeling. However, they suffer from quadratic complexity and are limited in applicability to high-resolution imagery. We introduce Mamba-RSI: a linear-time, state-space deep learning framework using selective recursion, hierarchical multi-scale feature extraction, and lightweight global representations. Mamba-RSI captures both fine-grained spectral/texture information and coarse structural patterns with significantly less computational overhead than existing quadratic self-attention transformers. Extensive experimentation on EuroSAT and NWPU-RESISC45 demonstrated that Mamba-RSI achieves state-of-the-art performance. It achieved 99.72% accuracy on EuroSAT and 96.84% on RESISC45. This represents a +0.40% improvement over the strongest transformer baseline, ATMformer, on EuroSAT, a +0.29% improvement on RESISC45, and more than +0.53% over ViT-B on EuroSAT. Robustness tests under severe Gaussian noise (
- Article type
- Year
Open Access
Research Article
Issue
Open Access
Research Article
Issue
The issue of urban traffic congestion is a persistent problem for the sustainable management of cities through transportation systems, as there is a need for models that integrate and analyze heterogeneous sources to yield accurate, interpretable outcomes. This paper introduces the cross-view fusion network (CVF-Net), a new multimodal deep learning framework for analyzing congestion across entire cities by combining remote-sensing imagery (drone aerial views), street-view camera images, and graph-structured sensor data into a single model. This model is introduced through a very novel architecture that includes a hierarchical attention fusion transformer (HAFT), which fuses cross-view attention (CVA) between the aerial and ground view, a temporal graph neural network (TGNN) that uses a spatio-temporal dynamic, and a graph refinement (GR) network for consistency relative to the graph topology. Extensive experiments across three benchmarks (CityFlowV2, METR-LA, PEMS-BAY) demonstrate that CVF-Net consistently outperforms other recent state-of-the-art methods, reducing forecasting error (MAE) by 9.3% and increasing tracking continuity (IDF1) by 7.0%. Ablation studies suggest that hierarchical fusion and temporal modeling improve accuracy and stability, while sensitivity analyses show that attention maps capture congestion and causal temporal patterns, which are real symptoms of congestion. The model also shows strong cross-dataset generalizability and robustness to sensor noise, which extends its performance in the real world. Unlike existing spatio-temporal GNNs and multimodal Transformers that rely on flat feature aggregation or implicitly assume cross-view alignment, the proposed framework introduces a hierarchical, alignment-aware fusion strategy that explicitly integrates aerial visual context with graph-temporal traffic dynamics.
京公网安备11010802044758号