Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
High dynamic range (HDR) imaging aims to faithfully capture, reconstruct, and reproduce the luminance and color information of real-world scenes whose intensity range far exceeds the recording and display capabilities of conventional imaging sensors and standard dynamic range systems. By overcoming the intrinsic problems of highlight saturation and shadow detail loss, HDR imaging has become a key enabling technology in computational photography, computer vision, video processing, and intelligent imaging systems. Over the past decades, the field has evolved from early radiometric calibration and exposure fusion to a comprehensive technical framework covering acquisition, reconstruction, compression, transmission, display, and quality assessment. With the rapid expansion of application scenarios, HDR imaging is now playing an increasingly important role in autonomous driving, virtual reality/augmented reality, ultra-high-definition video, mobile photography, and other vision-centric systems. Despite this progress, dynamic-scene HDR reconstruction remains fundamentally difficult because exposure variation is often entangled with camera motion, object motion, occlusion, saturation, and noise. These factors make artifact-free and high-fidelity reconstruction extremely challenging, and they continue to limit the practical deployment of HDR imaging in real-world environments. Against this background, a systematic review of datasets, evaluation protocols, technical routes, and open challenges is of both scientific and engineering significance.
Research on HDR imaging has shown a clear transition from traditional handcrafted methods to deep learning-driven and generative paradigms. Early methods mainly relied on multiple exposure stacks and were designed around motion detection, image alignment, and patch-based optimization. These methods laid the foundation for HDR deghosting, but they were often sensitive to inaccurate registration, large motion, occlusion, and saturated regions, which limited their robustness in complex dynamic scenes. Meanwhile, the classical radiometric reconstruction pipeline based on camera response recovery remains influential, not only historically but also because many later datasets still depend on it for label generation. However, this pipeline is subject to several limitations, including restrictive physical assumptions, strict acquisition requirements, noise sensitivity, color inconsistency, and limited sampling efficiency.
In recent years, the field has been reshaped by deep neural networks. CNN-based methods introduced end-to-end HDR reconstruction and significantly improved feature extraction, fusion quality, and deghosting performance. Representative strategies include optical-flow-based alignment, direct feature concatenation, correlation-guided implicit alignment, and image-translation-based exposure transformation. Among them, optical-flow-based methods provide explicit motion compensation, yet they often suffer from alignment errors, high computational cost, and poor deployability under large motion. Feature-concatenation methods reduce the dependency on explicit motion estimation, but they may accumulate errors and struggle with severe saturation or occlusion. Correlation-guided methods, especially those based on attention mechanisms, have become particularly influential because they enable implicit alignment in feature space and selectively suppress unreliable regions. This line of work has further benefited from the integration of Transformer architectures, whose self-attention mechanism enhances global context modeling and improves cross-region luminance consistency in challenging scenes. More recently, diffusion models have been introduced to hallucinate or compensate for missing content in saturated or occluded regions, offering new possibilities for detail recovery and realism enhancement. However, diffusion-based methods still face practical issues such as slow inference, semantic inconsistency, and the trade-off between generative flexibility and faithful reconstruction.
Progress has also been made in datasets and evaluation. Existing HDR datasets now cover single-frame HDR reconstruction, multi-frame HDR reconstruction, HDR video, and multi-exposure fusion, with differences in exposure settings, scene diversity, resolution, realism, and label availability. Nevertheless, current benchmark construction still inherits some weaknesses from the traditional HDR generation pipeline, such as inaccurate exposure metadata, unavoidable dynamic background interference, scene distortion, and insufficient coverage of night scenes. On the evaluation side, commonly used metrics include PSNR, SSIM, and HDR-VDP-2, which respectively measure pixel-level fidelity, structural similarity, and perceptual visibility under HDR conditions. Recent benchmarking further compares state-of-the-art methods from the perspectives of reconstruction quality, cross-dataset generalization, computational complexity, memory consumption, and inference time. These comparisons show that while some recent methods achieve leading reconstruction scores, methods with stronger domain generalization and better deployment efficiency are still relatively limited. In particular, architectures that balance fidelity, robustness, and efficiency appear more promising for practical applications than methods that only optimize in-domain benchmark performance.
Overall, the survey indicates that the core of multi-frame HDR reconstruction lies in reliable alignment and effective fusion under exposure variation, motion, saturation, and occlusion. Traditional explicit alignment pipelines are increasingly being replaced or complemented by implicit correlation-guided feature aggregation, which offers greater flexibility and robustness in dynamic scenes. Even so, current methods remain vulnerable to unseen domains, saturated-occluded regions, heavy motion, real sensor noise, and computational constraints. Looking ahead, future research can be organized into three major directions. The first concerns fundamental challenges, including robust image or feature alignment, the construction of larger and more realistic HDR datasets with trustworthy labels, and stronger out-of-domain generalization through transfer learning, semi-supervised learning, or broader data priors. The second concerns key performance optimization, especially real-time deployment, high-fidelity deghosting, and temporally consistent HDR video reconstruction for high-resolution and resource-constrained scenarios. This direction calls for lightweight architectures, low-bit quantization, efficient temporal modeling, and deployment-oriented design. The third concerns frontier exploration, including cross-modal HDR reconstruction with event cameras, multimodal or vision-language-driven HDR reasoning, unified frameworks spanning HDR reconstruction and tone mapping, and the integration of HDR imaging with NeRF, 3D Gaussian representations, and large foundation models. These emerging directions suggest that HDR imaging is moving from a narrowly defined reconstruction problem toward a broader intelligent imaging paradigm that unifies physical fidelity, perceptual quality, scene understanding, and real-world usability.
This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Comments on this article