AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

Audio Enhancement for Computer Audition—An Iterative Training Paradigm Using Sample Importance

Chair of Embedded Intelligence for Health Care and Wellbeing, University of Augsburg, Augsburg 86159, Germany
Chair of Health Informatics, München rechts der Isar, Technical University of Munich, Munich 81675, Germany
Munich Center for Machine Learning, Munich 80333, Germany
Huawei Technologies, Munich, Munich 80992, Germany
Munich Data Science Institute, Garching 85748, Germany
Group on Language, Audio and Music, Imperial College London, London SW7 2AZ, U.K.
Show Author Information

Abstract

Neural network models for audio tasks, such as automatic speech recognition (ASR) and acoustic scene classification (ASC), are susceptible to noise contamination for real-life applications. To improve audio quality, an enhancement module, which can be developed independently, is explicitly used at the front-end of the target audio applications. In this paper, we present an end-to-end learning solution to jointly optimise the models for audio enhancement (AE) and the subsequent applications. To guide the optimisation of the AE module towards a target application, and especially to overcome difficult samples, we make use of the sample-wise performance measure as an indication of sample importance. In experiments, we consider four representative applications to evaluate our training paradigm, i.e., ASR, speech command recognition (SCR), speech emotion recognition (SER), and ASC. These applications are associated with speech and non-speech tasks concerning semantic and non-semantic features, transient and global information, and the experimental results indicate that our proposed approach can considerably boost the noise robustness of the models, especially at low signal-to-noise ratios, for a wide range of computer audition tasks in everyday-life noisy environments.

Electronic Supplementary Material

Download File(s)
JCST-2210-12934-Highlights.pdf (370.3 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 895-911

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Milling M, Liu S, Triantafyllopoulos A, et al. Audio Enhancement for Computer Audition—An Iterative Training Paradigm Using Sample Importance. Journal of Computer Science and Technology, 2024, 39(4): 895-911. https://doi.org/10.1007/s11390-024-2934-x

895

Views

4

Crossref

4

Web of Science

4

Scopus

1

CSCD

Received: 26 October 2022
Accepted: 30 June 2024
Published: 20 September 2024
© Institute of Computing Technology, Chinese Academy of Sciences 2024