AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.7 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

Translation Optimization Technology of Automatic Speech Recognition Based on Industry-Specific Vocabulary

Xiaoliang MA1,2,3Lingling AN1Congjian DENG1,3,4Dequan DU2,3Guoxin ZHANG5
Guangzhou Institute of Technology, Xidian University, Guangzhou 510555, Guangdong, China
Guangzhou Branch of China Telecom Co., Ltd., Guangzhou 510620, Guangdong, China
Ma Xiaoliang’s Model Worker and Innovative Craftsman Workshop, Guangzhou 510620, Guangdong, China
Guangzhou Yunqu Information Technology Co., Ltd., Guangzhou 510665, Guangdong, China
Guangdong Branch of China Telecom Co., Ltd., Guangzhou 510080,Guangdong, China
Show Author Information

Abstract

Automatic speech recognition (ASR) technology has been developed relatively mature, and general ASR engines have been widely used in transportation, medical, communication and other industries. However, due to non-independent homology of industry-specific vocabulary in the large-scale training corpus, there comes to low recognition accuracy of industry-specific vocabulary when the general ASR engines are applied to various subdivisions of industries. As compared with 16 kHz audio sampling rate in Internet environment, narrowband low sampling (8 kHz) of call center may result in more significant decrease of recognition accuracy of ASR. In order to improve the accuracy of speech recognition of industry-specific words, this paper proposes a translation optimization technology of ASR based on industry-specific vocabulary. Specifically, first, convolutional neural network model and deep neural network BERT model are used to predict word for corpus text data, and an industry-specific error correction vocabulary is generated. Next, in the production environment, a general ASR engine is used to perform initial transcription of telephone call voice data. Then, the transcribed text is corrected by using the Soft-Masked BERT model combined with the industry-specific error correction vocabulary, thus improving the accuracy of speech recognition. Finally, by using 12345 hotline customer service call voice data for modeling and testing, the proposed translation optimization technology is proved efficient in improving the accuracy of general ASR recognition by 10 percentage points with high error correction speed and good applicability.

CLC number: TP391.1 Article ID: 1000-565X(2023)08-0118-08

References

【1】
【1】
 
 
Journal of South China University of Technology (Natural Science Edition)
Pages 118-125

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
MA X, AN L, DENG C, et al. Translation Optimization Technology of Automatic Speech Recognition Based on Industry-Specific Vocabulary. Journal of South China University of Technology (Natural Science Edition), 2023, 51(8): 118-125. https://doi.org/10.12141/j.issn.1000-565X.220740

416

Views

7

Downloads

0

Crossref

0

Web of Science

2

Scopus

1

CSCD

Received: 10 November 2022
Published: 25 August 2023
© Journal of South China University of Technology(Natural Science Edition)