AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (717.4 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

DeltaVLAD: An efficient optimization algorithm to discriminate speaker embedding for text-independent speaker verification

Xin Guo1Chengfang Luo2Aiwen Deng2Feiqi Deng2( )
Guangdong Communication Polytechnic, Guangzhou 510650, China
School of Automation Science and Engineering, South China University of Technology, Guangzhou 510641, China
Show Author Information

Abstract

Text-independent speaker verification aims to determine whether two given utterances in open-set task originate from the same speaker or not. In this paper, some ways are explored to enhance the discrimination of embeddings in speaker verification. Firstly, difference is used in the coding layer to process speaker features to form the DeltaVLAD layer. The frame-level speaker representation is extracted by the deep neural network with differential operations to calculate the dynamic changes between frames, which is more conducive to capturing insignificant changes in the voiceprint. Meanwhile, NeXtVLAD is adopted to split the frame-level features into multiple word spaces before aggregating, and subsequently perform VLAD operations in each subspace, which can significantly reduce the number of parameters and improve performance. Secondly, the margin-based softmax loss function and the few-shot learning-based loss function are proposed to be combined for more discriminative speaker embeddings. Finally, for a fair comparison, the experimental results are performed on Voxceleb-1 showing superior performance of speaker verification system and can obtain new state-of-the-art results.

CLC number: 00A67, 33B99

References

【1】
【1】
 
 
AIMS Mathematics
Pages 6381-6395

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Guo X, Luo C, Deng A, et al. DeltaVLAD: An efficient optimization algorithm to discriminate speaker embedding for text-independent speaker verification. AIMS Mathematics, 2022, 7(4): 6381-6395. https://doi.org/10.3934/math.2022355

0

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 29 September 2021
Revised: 29 December 2021
Accepted: 06 January 2022
Published: 15 April 2022
©2022 the Author(s), licensee AIMS Press.

This is an open access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0)