AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (10.5 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access | Online First

Reverse the Auditory Processing Pathway: Coarse-to-Fine Audio Reconstruction from Human Brain Activity

State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China, and also with School of Future Technology, University of Chinese Academy of Sciences, Beijing 100049, China
State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China, and also with School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
Show Author Information

Abstract

Drawing inspiration from the hierarchical processing of the human auditory system, which transforms sound from low-level acoustic features to high-level semantic understanding, we introduce a novel Coarse-to-Fine (C2F) audio reconstruction method. Leveraging non-invasive functional Magnetic Resonance Imaging (fMRI) data, our approach first utilizes Contrastive Language-Audio Pretraining (CLAP) to decode fMRI signals coarsely into a semantic space, followed by a semantically guided fine-grained decoding into the Audio Mask Autoencoder (AudioMAE) latent space. These fine-grained neural features then serve as conditions for high-fidelity audio reconstruction through a Latent Diffusion Model (LDM). Extensive validation on three public fMRI datasets demonstrates the superiority of our C2F decoding method over conventional fine-grained approaches, achieving state-of-the-art performance across metrics including Fréchet Distance (FD), Fréchet Audio Distance (FAD), and Kullback–Leibler divergence (KL). Furthermore, reconstruction quality in challenging scenarios is enhanced through an innovative semantic prompting mechanism. This framework holds potential for advancing brain-computer interfaces and assistive technologies, such as improved hearing aids and neural communication systems for those with auditory or speech impairments. Reconstructed results are available at https://neurofusex.github.io/c2f-ldm/.

References

【1】
【1】
 
 
Tsinghua Science and Technology

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Liu C, Du C, Chen X, et al. Reverse the Auditory Processing Pathway: Coarse-to-Fine Audio Reconstruction from Human Brain Activity. Tsinghua Science and Technology, 2026, https://doi.org/10.26599/TST.2026.9010004

1179

Views

46

Downloads

2

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 30 August 2025
Revised: 18 November 2025
Accepted: 22 December 2025
Published: 29 September 2026
© The author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).