Nowadays, deep learning has demonstrated impressive performance in the area of computer vision and pattern recognition, such as objects recognition, videos classification and image segmentation. In particular, convolutional neural networks (CNNs) an advanced deep-learning technique have achieved strong performance in image and video analysis owing to their powerful feature-extraction capabilities.Based on that, current research also make a breakthrough in image and video recognition via aggregating features extracted from various layers of CNNs. Over the last decade, feature fusion strategies have developed from conventional schemes restricted to conditions we need to control over advanced ones which can be achieved with a large number of images or videos. However, due to the rapid development in this field, it is challenging to track and systematically analyze recent advancements. This has inspired us to offer a comprehensive survey of the significant steps taken towards feature aggregation strategies. To better organize and facilitate understanding, we first introduce preliminary knowledge on deep learning. Then this paper focuses on categorizing and reviewing the current strategies from two main aspects: feature aggregation on images and feature aggregation on videos. We further highlight the comparative strengths and limitations among these strategies. Finally, we point out the challenges in this area and motivate further work via proposing future directions on feature aggregation techniques for both images and videos.
Publications
- Article type
- Year
- Co-author
Article type
Year
Open Access
Article
Issue
CAAI Artificial Intelligence Research 2025, 4: 9150054
Published: 14 January 2026
Downloads:284
Total 1
京公网安备11010802044758号