AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (373.7 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline

A Unified Framework for Multilingual Text-to-Speech Synthesis with SSML Specification as Interface

Zhiyong WU1,2( )Guangqi CAO1M. Helen MENG1,2Lianhong CAI2
Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong, Shatin, N.T., Hong Kong SAR, China
Tsinghua-CUHK Joint Research Center for Media Sciences, Technologies and Systems, Graduate School at Shenzhen, Tsinghua University, Shenzhen 518055, China
Show Author Information

Abstract

This paper describes the design of a unified framework for a multilingual text-to-speech (TTS) synthesis engine – Crystal. The unified framework defines the common TTS modules for different languages and/or dialects. The interfaces between consecutive modules conform to the speech synthesis markup language (SSML) specification for standardization, interoperability, multilinguality, and extensibility. Detailed module divisions and implementation technologies for the unified framework are introduced, together with possible extensions for the algorithm research and evaluation of the TTS synthesis. Implementation of a mixed-language TTS system for Chinese Putonghua, Chinese Cantonese, and English demonstrates the feasibility of the proposed unified framework.

References

【1】
【1】
 
 
Tsinghua Science and Technology
Pages 623-630

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
WU Z, CAO G, MENG MH, et al. A Unified Framework for Multilingual Text-to-Speech Synthesis with SSML Specification as Interface. Tsinghua Science and Technology, 2009, 14(5): 623-630. https://doi.org/10.1016/S1007-0214(09)70127-0

104

Views

4

Downloads

8

Crossref

N/A

Web of Science

12

Scopus

19

CSCD

Received: 26 March 2009
Revised: 10 June 2009
Published: 01 June 2009
© Tsinghua University Press 2009