Drawing inspiration from the hierarchical processing of the human auditory system, which transforms sound from low-level acoustic features to high-level semantic understanding, we introduce a novel Coarse-to-Fine (C2F) audio reconstruction method. Leveraging non-invasive functional Magnetic Resonance Imaging (fMRI) data, our approach first utilizes Contrastive Language-Audio Pretraining (CLAP) to decode fMRI signals coarsely into a semantic space, followed by a semantically guided fine-grained decoding into the Audio Mask Autoencoder (AudioMAE) latent space. These fine-grained neural features then serve as conditions for high-fidelity audio reconstruction through a Latent Diffusion Model (LDM). Extensive validation on three public fMRI datasets demonstrates the superiority of our C2F decoding method over conventional fine-grained approaches, achieving state-of-the-art performance across metrics including Fréchet Distance (FD), Fréchet Audio Distance (FAD), and Kullback–Leibler divergence (KL). Furthermore, reconstruction quality in challenging scenarios is enhanced through an innovative semantic prompting mechanism. This framework holds potential for advancing brain-computer interfaces and assistive technologies, such as improved hearing aids and neural communication systems for those with auditory or speech impairments. Reconstructed results are available at https://neurofusex.github.io/c2f-ldm/.
Publications
- Article type
- Year
- Co-author
Article type
Year
Open Access
Research Article
Online First
Tsinghua Science and Technology
Published: 29 September 2026
Downloads:46
Total 1
京公网安备11010802044758号