AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (9.4 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

A survey on deep learning techniques for image and video feature aggregation

Md Majedul Islam1( )Sai Prakash Reddy Rao1Selena He1( )
Department of Computer Science, Kennesaw State University, Marietta 30060, GA, USA
Show Author Information

Abstract

Nowadays, deep learning has demonstrated impressive performance in the area of computer vision and pattern recognition, such as objects recognition, videos classification and image segmentation. In particular, convolutional neural networks (CNNs) an advanced deep-learning technique have achieved strong performance in image and video analysis owing to their powerful feature-extraction capabilities.Based on that, current research also make a breakthrough in image and video recognition via aggregating features extracted from various layers of CNNs. Over the last decade, feature fusion strategies have developed from conventional schemes restricted to conditions we need to control over advanced ones which can be achieved with a large number of images or videos. However, due to the rapid development in this field, it is challenging to track and systematically analyze recent advancements. This has inspired us to offer a comprehensive survey of the significant steps taken towards feature aggregation strategies. To better organize and facilitate understanding, we first introduce preliminary knowledge on deep learning. Then this paper focuses on categorizing and reviewing the current strategies from two main aspects: feature aggregation on images and feature aggregation on videos. We further highlight the comparative strengths and limitations among these strategies. Finally, we point out the challenges in this area and motivate further work via proposing future directions on feature aggregation techniques for both images and videos.

References

【1】
【1】
 
 
CAAI Artificial Intelligence Research
Article number: 9150054

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Islam MM, Rao SPR, He S. A survey on deep learning techniques for image and video feature aggregation. CAAI Artificial Intelligence Research, 2025, 4: 9150054. https://doi.org/10.26599/AIR.2025.9150054

3508

Views

284

Downloads

0

Crossref

Received: 02 January 2025
Revised: 07 October 2025
Accepted: 18 December 2025
Published: 14 January 2026
© The author(s) 2025.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).