AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (8.7 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access | Just Accepted

Pixel-Level Registration Method and Dataset Construction for Multimodal Image Data in Intelligent Cockpits

Xin Ning1,2,3Kai Zhao1,2Dan Li1,2Jin Ning1,4,5Weijun Li1,3,6,7Baoli Lu1( )

1 Institute of Semiconductors, Chinese Academy of Sciences, Beijing 100083, China.

2 College of Materials Science and Opto-Electronic Technology, University of Chinese Academy of Sciences, Beijing 100049, China.

3 School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing 101408, China.

4 Center of Materials Science and Optoelectronics Engineering, University of Chinese Academy of Sciences, Beijing 100049, China.

5 School of Electronic, Electrical and Communication Engineering, Uni-versity of Chinese Academy of Sciences, Beijing 100049, China.
6 School of Integrated Circuits, University of ChineseAcademy of Sciences, Beijing 100049, China.

7 Zhongguancun Academy, Beijing 100094, China.

Show Author Information

Abstract

Driver behavior monitoring is fundamental to safety assurance and interaction optimization in intelligent cockpits. To address key challenges in existing research, including limited modality diversity, cross-view spatial mis-alignment, and insufficient robustness of multimodal perception, we construct a multimodal intelligent-cockpit image dataset consisting of synchronously captured RGB, infrared (IR), and depth images. We also propose a two-stage registration method for dataset preprocessing that combines geometric mapping and mask-guided inpainting to achieve pixel-level alignment across modalities. Using the constructed dataset, we evaluate the impact of different modality combinations on object detection. Results on two baseline multimodal detectors show that the tri-modal setting significantly outperforms RGB-only input in both accuracy and robustness, particularly under challenging lighting conditions, with mAP@0.5:0.95 gains of 36.8% and 36.3%, respectively. These findings verify the effectiveness of the proposed dataset for multimodal object detection and driver behavior understanding. The proposed dataset provides a valuable benchmark for multimodal perception in intelligent cockpits and facilitates future research on driver monitoring and active safety.

References

【1】
【1】
 
 
Tsinghua Science and Technology

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Ning X, Zhao K, Li D, et al. Pixel-Level Registration Method and Dataset Construction for Multimodal Image Data in Intelligent Cockpits. Tsinghua Science and Technology, 2026, https://doi.org/10.26599/TST.2026.9010048
Part of a topical collection:

365

Views

26

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 03 April 2026
Revised: 07 May 2026
Accepted: 12 May 2026
Available online: 14 May 2026

© The author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).