Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Drawing inspiration from the hierarchical processing of the human auditory system, which transforms sound from low-level acoustic features to high-level semantic understanding, we introduce a novel Coarse-to-Fine (C2F) audio reconstruction method. Leveraging non-invasive functional Magnetic Resonance Imaging (fMRI) data, our approach first utilizes Contrastive Language-Audio Pretraining (CLAP) to decode fMRI signals coarsely into a semantic space, followed by a semantically guided fine-grained decoding into the Audio Mask Autoencoder (AudioMAE) latent space. These fine-grained neural features then serve as conditions for high-fidelity audio reconstruction through a Latent Diffusion Model (LDM). Extensive validation on three public fMRI datasets demonstrates the superiority of our C2F decoding method over conventional fine-grained approaches, achieving state-of-the-art performance across metrics including Fréchet Distance (FD), Fréchet Audio Distance (FAD), and Kullback–Leibler divergence (KL). Furthermore, reconstruction quality in challenging scenarios is enhanced through an innovative semantic prompting mechanism. This framework holds potential for advancing brain-computer interfaces and assistive technologies, such as improved hearing aids and neural communication systems for those with auditory or speech impairments. Reconstructed results are available at https://neurofusex.github.io/c2f-ldm/.
The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).
Comments on this article