AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (20.5 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

CODS: An Audio-Text Aligned Dataset for Cantonese Opera Vocal Synthesis

Yue LI1( )Yihan HUANG1Zhengwei PENG2Jixuan XIE1Yuye DU1
School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, Guangdong, China
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou 510006, Guangdong, China
Show Author Information

Abstract

As one of the traditional Chinese arts, Chinese opera culture has unique musical expressiveness. Cantonese opera, as one of the main Chinese opera genres and an important carrier of Lingnan culture, has been indexed in the World Intangible Cultural Heritage List. In recent years, generative artificial intelligence technology has demonstrated its powerful capabilities in the field of content creation. For example, singing synthesis technology can synthesize natural singing based on specified music scores. This provides a new idea for the digital protection and innovation of Cantonese opera. However, the collection and organization of opera data faces problems such as poor audio quality and complex dialect annotation, resulting in an extreme shortage of high-quality opera data sets. Based on this, this paper applied the singing synthesis technology in the field of pop music to the field of Cantonese opera vocal synthesis, and proposed the first Cantonese opera vocal synthesis dataset with phoneme-level annotation and audio-text alignment. Firstly, this paper constructed the CODS dataset through a systematic process. This dataset was derived from 29 original works by four famous performers with a total length of 3.81 hours, which provides important support for the research and digitization of Cantonese opera. Using this dataset, this paper conducted experiments with a deep learning-based method for Cantonese opera voice synthesis, realizing controllable generation in terms of lyrics, timbre, and melody. Finally, this paper established a comprehensive evaluation framework for Cantonese opera synthesis. Both objective and subjective evaluations reached a satisfactory level within the domain, further validating the usability of the proposed dataset. The CODS dataset constructed in this paper successfully filled the gap in artificial intelligence in the field of Cantonese opera vocal synthesis, and strongly promoted the inheritance and innovation of this traditional art.

CLC number: TP39 Article ID: 1000-565X(2025)09-0001-10

References

【1】
【1】
 
 
Journal of South China University of Technology (Natural Science Edition)
Pages 1-10

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
LI Y, HUANG Y, PENG Z, et al. CODS: An Audio-Text Aligned Dataset for Cantonese Opera Vocal Synthesis. Journal of South China University of Technology (Natural Science Edition), 2025, 53(9): 1-10. https://doi.org/10.12141/j.issn.1000-565X.250134

90

Views

3

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 06 May 2025
Published: 25 September 2025
© Journal of South China University of Technology (Natural Science Edition)