AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (11.1 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

A Prosody-Guided Multi-Stream Framework for Universal Detection of AI-Synthesized Speech across Codec and Vocoder Domains

Akmalbek Abdusalomov1Mukhriddin Mukhiddinov2,3Fakhriddin Abdirazakov4Alpamis Kutlimuratov5Nodira Alimova6Ilyos Kalandarov7Ayhan Istanbullu8Rashid Nasimov9Young-Im Cho1( )
Department of Computer Engineering, Gachon University, Seongnam-si, Republic of Korea
Department of Industrial Management and Digital Technologies, Nordic International University, Tashkent, Uzbekistan
Department of Artificial Intelligence, Tashkent University of Information Technologies Named after Muhammad Al-Khwarizmi, Tashkent, Uzbekistan
Department of Computer Systems, Tashkent University of Information Technologies Named after Muhammad Al-Khwarizmi, Tashkent, Uzbekistan
Department of Applied Informatics, Kimyo International University in Tashkent, Tashkent, Uzbekistan
Department of Information Processing and Control Systems, Tashkent State Technical University, Tashkent, Uzbekistan
Department of Automation and Control, Navoi State University of Mining and Technologies, Navoi, Uzbekistan
Department of Computer Engineering, Faculty of Engineering, Balikesir University, Balikesir, Turkey
Department of Artificial Intelligence, Tashkent State University of Economics, Tashkent, Uzbekistan
Show Author Information

Abstract

Recent advancements in AI-synthesized speech have resulted in highly realistic deepfake audio, posing severe threats to authentication systems and digital media trust. Existing detection models struggle to generalize across diverse synthesis methods, especially those involving neural codec-based Audio Language Models (ALMs). In this work, we propose UniTector++, a novel prosody-aware, multi-stream detection architecture that generalizes across vocoder- and codec-based synthesis. UniTector++ incorporates three complementary streams—Whisper-based semantic embeddings, high-level prosodic features, and codec artifact representations—fused through a Multi-Domain Adaptive Graph Attention Fusion (MAGAF) module. Furthermore, an Emotion-Consistency Verification Module (ECVM) reinforces alignment between speech style and prosodic content, and a Universal Adversarial Robustness (UAR) head improves resistance against adversarial attacks. Evaluated on three benchmark datasets—ASVspoof2021, PolyFake, and Codecfake—UniTector++ achieves state-of-the-art performance with average Equal Error Rate (EER) of 0.57% under unseen synthesis scenarios, outperforming competitive baselines by a relative margin of 28%. Our results demonstrate the model’s superior generalization, interpretability, and robustness, offering a significant advancement in universal deepfake speech detection.

References

【1】
【1】
 
 
Computers, Materials & Continua

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Abdusalomov A, Mukhiddinov M, Abdirazakov F, et al. A Prosody-Guided Multi-Stream Framework for Universal Detection of AI-Synthesized Speech across Codec and Vocoder Domains. Computers, Materials & Continua, 2026, 88(1). https://doi.org/10.32604/cmc.2026.080444

4

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 09 February 2026
Accepted: 15 April 2026
Published: 08 May 2026
© The Author 2026.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.