AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.5 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

A Single-Channel Speech Separation Model Based on Time-Domain Comprehensive Attention Mechanism

Junmei YANG( )Bangcheng ZHANGLu YANGDelu ZENG
School of Electronic and Information Engineering, South China University of Technology, Guangzhou 510640, Guangdong, China
Show Author Information

Abstract

Single-channel speech separation aims to extract clean target speaker speech from a mixed audio signal recorded by a single microphone, with significant application value in scenarios such as smart homes, conference systems, and hearing aids. With the rapid development of deep learning technology, self-attention network-based approaches to single-channel speech separation have achieved remarkable progress. While self-attention networks excel at capturing contextual information in long sequence, they still exhibit limitations in capturing detailed features such as temporal/spectral continuity, spectral structure, and timbre in real-world speech scenarios. Moreover, existing separation architectures based on a single attention paradigm struggle to achieve effective multi-scale feature fusion. To address these challenges, this paper proposed a Temporal Comprehensive Attention Network (TCANet), which addresses the aforementioned issues through a synergistic design of local and global attention modules. Local modeling employs an S& C-SENet-enhanced Conformer structure to capture short-term features such as spectral structure and timbre in detail, while global modeling incorporates a modified Transformer module with relative position embedding to explicitly learn long-term speech dependencies in speech. Furthermore, TCANet achieves cross-scale fusion of intra-block local features and inter-block global correlations through a dimension transformation mechanism. Experimental results on three benchmark datasets—LRS2-2Mix, Libri2Mix, and EchoSet—demonstrate that the proposed method outperforms existing end-to-end speech separation approaches in terms of scale-invariant signal-to-noise ratio improvement (SI-SNRi) and signal-to-distortion ratio improvement (SDRi).

CLC number: TP391 Article ID: 1000-565X(2026)01-0070-13

References

【1】
【1】
 
 
Journal of South China University of Technology (Natural Science Edition)
Pages 70-82

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
YANG J, ZHANG B, YANG L, et al. A Single-Channel Speech Separation Model Based on Time-Domain Comprehensive Attention Mechanism. Journal of South China University of Technology (Natural Science Edition), 2026, 54(1): 70-82. https://doi.org/10.12141/j.issn.1000-565X.250054

250

Views

1

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 03 March 2025
Published: 01 January 2026
© Journal of South China University of Technology(Natural Science Edition)