A discriminative training approach for text-independent speaker recognition

Q. Y. HONG, S. KWONG

Research output: Journal PublicationsJournal Article (refereed)peer-review

13 Citations (Scopus)

Abstract

Gaussian mixture model (GMM) has been commonly used for text-independent speaker recognition. The estimation of model parameters is generally performed based on the maximum likelihood (ML) criterion. However, this criterion only utilizes the labeled utterances for each speaker model and very likely leads to a local optimization solution. To solve this problem, this paper proposes a discriminative training approach based on the maximum model distance (MMD) criterion. We investigate the characteristics of speaker recognition and further propose a novel selection strategy of competing speakers associated with it. Experimental results based on the KING and TIMIT databases demonstrate that our training approach was quite efficient to improve the performance of speaker identification and verification. When there were three training sentences for each speaker, the verification equal error rate (EER) of 168 speakers in TIMIT could be reduced by 30.4% compared with the conventional method. © 2005 Elsevier B.V. All rights reserved.
Original languageEnglish
Pages (from-to)1449-1463
JournalSignal Processing
Volume85
Issue number7
DOIs
Publication statusPublished - Jul 2005
Externally publishedYes

Bibliographical note

This work is supported by City University Strategic Grant 7001615.

Keywords

  • Discriminative training
  • Maximum model distance
  • Speaker recognition

Fingerprint

Dive into the research topics of 'A discriminative training approach for text-independent speaker recognition'. Together they form a unique fingerprint.

Cite this