AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (25.7 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

3D Enhanced Residual CNN for Video Super-Resolution Network

Weiqiang Xin#,1,2,3Zheng Wang#,4Xi Chen1,5Yufeng Tang1Bing Li1Chunwei Tian2,5( )
School of Software, Northwestern Polytechnical University, Xi’an, 710072, China
Shenzhen Research Institute of Northwestern Polytechnical University, Northwestern Polytechnical University, Shenzhen, 518057, China
State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, 210023, China
School of Interdisciplinary Studies, Lingnan University, Hong Kong, 999077, China
Yangtze River Delta Research Institute, Northwestern Polytechnical University, Taicang, 215400, China

#These authors contributed equally to this work

Show Author Information

Abstract

Deep convolutional neural networks (CNNs) have demonstrated remarkable performance in video super-resolution (VSR). However, the ability of most existing methods to recover fine details in complex scenes is often hindered by the loss of shallow texture information during feature extraction. To address this limitation, we propose a 3D Convolutional Enhanced Residual Video Super-Resolution Network (3D-ERVSNet). This network employs a forward and backward bidirectional propagation module (FBBPM) that aligns features across frames using explicit optical flow through lightweight SPyNet. By incorporating an enhanced residual structure (ERS) with skip connections, shallow and deep features are effectively integrated, enhancing texture restoration capabilities. Furthermore, 3D convolution module (3DCM) is applied after the backward propagation module to implicitly capture spatio-temporal dependencies. The architecture synergizes these components where FBBPM extracts aligned features, ERS fuses hierarchical representations, and 3DCM refines temporal coherence. Finally, a deep feature aggregation module (DFAM) fuses the processed features, and a pixel-upsampling module (PUM) reconstructs the high-resolution (HR) video frames. Comprehensive evaluations on REDS, Vid4, UDM10, and Vim4 benchmarks demonstrate well performance including 30.95 dB PSNR/0.8822 SSIM on REDS and 32.78 dB/0.8987 on Vim4. 3D-ERVSNet achieves significant gains over baselines while maintaining high efficiency with only 6.3M parameters and 77 ms/frame runtime (i.e., 20× faster than RBPN). The network’s effectiveness stems from its task-specific asymmetric design that balances explicit alignment and implicit fusion.

References

【1】
【1】
 
 
Computers, Materials & Continua
Pages 2837-2849

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Xin W, Wang Z, Chen X, et al. 3D Enhanced Residual CNN for Video Super-Resolution Network. Computers, Materials & Continua, 2025, 85(2): 2837-2849. https://doi.org/10.32604/cmc.2025.069784

199

Views

2

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 30 June 2025
Accepted: 15 August 2025
Published: 23 September 2025
© The Author 2024.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.