Low-light image/video analysis is essential for various applications, e.g., night surveillance and photography, high-speed imaging, and autonomous vehicles. Under such conditions, cameras suffer from low signal-to-noise ratio, which degrades image quality severely and poses challenges for downstream tasks such as object detection. Data-driven methods have achieved enormous success for normal-light image/video restoration and high-level vision tasks. However, the lack of a high-quality benchmark dataset with accurate semantic annotations for low-light images and especially videos greatly hinders research progress. In this paper, we contribute the first multi-illuminance, multi-camera, low-light dataset, DarkVision, serving both image/video enhancement and object detection applications. We provide bright and dark pairs with pixel-wise registration, in which the bright counterpart provides a reliable reference for enhancement and annotation. This dataset comprises 13,455 images of 900 static scenes with objects from 15 categories, and 89,411 frames of 32 dynamic scenes with 4 categories of objects. For each scene, images/videos were captured at 5 illuminance levels using three cameras of different quality grades; average photon numbers can be reliably estimated from the calibration curves for quantitative studies. The static images and dynamic videos respectively contain around 7344 and 320,667 object instances in total. With DarkVision, we establish baselines for image/video enhancement and object detection by representative algorithms. To demonstrate an exemplary application of DarkVision, we propose two simple yet effective approaches to improve the performance of video enhancement and object detection respectively by exploiting temporal cues. Furthermore, we study the relationship between image enhancement and object detection. We believe DarkVision can help to advance the state-of-the-art in both low-light image/video enhancement and object detection, as well as benefiting cross-task studies.
- Article type
- Year
- Co-author
Open Access
Research Article
Issue
Open Access
Issue
Human action recognition and posture prediction aim to recognize and predict respectively the action and postures of persons in videos. They are both active research topics in computer vision community, which have attracted considerable attention from academia and industry. They are also the precondition for intelligent interaction and human-computer cooperation, and they help the machine perceive the external environment. In the past decade, tremendous progress has been made in the field, especially after the emergence of deep learning technologies. Hence, it is necessary to make a comprehensive review of recent developments. In this paper, firstly, we attempt to present the background, and then discuss research progresses. Secondly, we introduce datasets, various typical feature representation methods, and explore advanced human action recognition and posture prediction algorithms. Finally, facing the challenges in the field, this paper puts forward the research focus, and introduces the importance of action recognition and posture prediction by taking interactive cognition in self-driving vehicle as an example.
京公网安备11010802044758号