Full text
2014 58 Leibny Paola García Perera Métodos discriminativos para la optimización de modelos en la Verificación del Hablante Departamento Director/es Ingeniería Electrónica y Comunicaciones Nolazco Flores, Juan Arturo Lleida Solano, Eduardo Director/es Tesis Doctoral Autor Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA
Departamento Director/es Leibny Paola García Perera MÉTODOS DISCRIMINATIVOS PARA LA OPTIMIZACIÓN DE MODELOS EN LA VERIFICACIÓN DEL HABLANTE Director/es Ingeniería Electrónica y Comunicaciones Nolazco Flores, Juan Arturo Lleida Solano, Eduardo Tesis Doctoral Autor Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA
Departamento Director/es Director/es Tesis Doctoral Autor Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA
Universidad de Zaragoza Departamento de Ingenier´ıa Electr´onica y Comunicaciones Tesis doctoral M´etodos discriminativos para la optimizaci´on de modelos en la Verificaci´on del Hablante Discriminative Methods for model optimization in Speaker Verification Leibny Paola Garc´ıa Perera Directores de Tesis Juan Arturo Nolazco Flores, Eduardo Lleida Solano Abril 8, 2014
Any intelligent fool can make things bigger, more complex, and more violent. It takes a touch of genius – and a lot of courage – to move in the opposite direction.” Albert Einstein
iv
Acknowledgements To every expression of love... I am very thankful and fortunate to have been surrounded by amazing people during these years. My sincere and deepest thanks to each one of you. Firstly, I would like to thank Doctor Juan Arturo Nolazco Flores for his support during all this time, specially through those tough years. Thank you for always having faith in me. Secondly, I would like to thank Eduardo Lleida Solano for his kindness and support every time I needed it. It has been an awesome and fruitful experience to work with both of you. My first short stay collaboration was in Georgia Tech, where Professor Chin-Hui Lee and his research team welcomed me. I am very thankful to Professor Lee for spending time discussing my research, for sharing his knowledge and for his constant and sincere support during these years. I was very lucky to be in Gatech and learned about speech, being there helped me to foster my research skills. During my stay in CMU, I had the opportunity to work with very smart and nice people that in many ways contributed to this work. First, I would like to thank Bhiksha Raj, for rescuing me when I was about to give up. Your help and advice marked my life. There is a before Bhiksha and after Bhiksha. Thank you Rita Singh for all your support, I learned kindness from you. You are an example to all of us. My acknowledgment extends to Richard Stern for the nice discussions and the academic and personal advice. My gratitude goes to the people that anonymously make our work easier, taking care of the administrative procedures. I acknowledge the help of the secretaries in every place I have been, your help and patience to do the paperwork and to solve my administrative concerns is greatly appreciated. Special thanks to Sandra Dominguez Alanis, at the Tec de Monterrey. Friendship has been a constant in my life. Usually, colleagues transform into lifelong friends that deal with you through the up and downs in research. They do deserve more than a thank you! In Atlanta I had the opportunity to have very fruitful discussions and learn about different topics. Special thanks to Antonio Moreno for teaching me speech! Now I can not recognize what code is mine or yours. Thank you for the nice discussions and for sharing your knowledge with me and my team. I am thankful for showing me new ways of viewing problems. But most of all, for being there when I needed it. Thank you to all my friends in Atlanta, special thanks to the spouse group that gave me true friends while living in Atlanta. Pittsburgh was a watershed in my research and personal life. Firstly, I would like to thank Afsaneh Asaei (Affysita Preciosa) for her unconditional friendship and love. There are no words to express how much you mean to me. Look! There is a miracle! Apurv Tiwari, Angelito, thank you for letting me share my adventures with you. But, first, listen... do you want to know a secret?
vi ACKNOWLEDGEMENTS Akiko and Taka, thank you for being around all the time. Your awesomeness and friendship always give me peace. Antonio Juarez, thank you for existing. I always need an outlier in my life that can show me that beyond the horizon there are places full of beauty. Thank you for keeping in touch, when the probability of doing so was close to zero. Thank you Oliver Walter, for sharing your time with me. You have this strange ability to give balance and peace to my surreal world. Thank you for having faith when I get lost. Life made us meet, and it has been awesome to have you around. Tec de Monterrey, has been my home place during the last decade. My gratitude goes to the students that have been in our lab. Special thanks to Igmar Hern´andez for the very nice friendship during these years. There are no words to express my gratitude. Thank you Daniel Escobar, for being a true friend, for being part of my crazy and abnormal world over and over again, and virtually staying in touch no matter what. Thank you, Benjamin Gonz´alez for showing me other ways of thinking, and that research is not everything. I am very thankful to Roberto Aceves, for his patience and his support with the cluster along this time. It would have been almost impossible to succeed in this without your help. I extend my gratitude to Carlos M for the wonderful moments we shared, and for encouraging me to study the PhD. Zaragoza University gave me the opportunity not only to study a PhD, but to have awesome colleagues and friends. Thank you Oscar Saz for scolding me every time I need it (still now). You have been an awesome friend during these years. Special thanks to Jes´us Villalba and Rub´en Cabrejas, I really enjoyed spending time with you, dancing, parting. You are just fun! Thank you Diego Cast´an, David Mart´ınez and Julia Olcoz, I enjoyed having coffee with you, our discussions and all your help. Special thanks for my unconditional friends along this time: Jeanette Russ, Gloria Massieu, Mayra Gonz´alez, Adriana L´opez, Ra´ul Moreno, Enrique Coeto, Juan Jos´e Salas ”Fuku”. You are always in my thoughts, even when I am far away. Lastly, I would like to thank my mother, Paula Perera, my father, Guillermo Garc´ıa and my brother Guillermo Garc´ıa-Perera for being always beside me along this journey. Thank you mom, for your prayers, your immense love, your tears, but most of all for your laugh and your smile. When I am down, your laugh and happiness give me courage to continue with enthutiasm and faith. There is not other phrase to say than: “yo soy la hija de Paulita y con eso es suficiente”. Thank you Dad, for believing, for your immense love and for showing me that the word impossible is not in our dictionary. Your advices, examples, prayers and love are always with me. Thank you Memo, for everything, you are the best! Thank you for always being there, in the good, in the bad and in the very very bad. You deserve a statue somewhere. Your support and advice always make the difference in my life. Thank you grandma, wherever you are... your strength is always guiding my steps. And for those not mentioned here, all my gratitude!!! Thank you ”music” and ”poetry” just for being... If love is everything, I am sure art is the sweetest part of it. This work has been financially supported by the National Council on Science and Technology of Mexico (CONACYT) through scholarship 160892
CONTENTS xiii 6.6.1 Global parameters: Loadings estimation . . . . . . . . . . . . . . . . . . . . 59 6.6.2 Estimating Specific Parameters . . . . . . . . . . . . . . . . . . . . . . . . . 60 6.7 Noisecase......................................... 62 6.8 Comparing AUC approach with Calibration approach . . . . . . . . . . . . . . . . 63 6.9 ExperimentsandResults ................................ 64 6.9.1 On the selection of the competing samples . . . . . . . . . . . . . . . . . . . 64 6.9.2 Experiments on clean speech . . . . . . . . . . . . . . . . . . . . . . . . . . 65 6.9.3 Experiments on noisy speech . . . . . . . . . . . . . . . . . . . . . . . . . . 65 6.9.4 The effect of the optimization of the AUC of the ROC curve in the score distribution ................................... 66 6.9.5 NumberofGaussians .............................. 68 6.10Summary ......................................... 69 7 Ensemble Modeling Approach 71 7.1 Cohorts .......................................... 71 7.2 Ensemble Model for Robust Verification . . . . . . . . . . . . . . . . . . . . . . . . 72 7.3 Findingthepartitions .................................. 72 7.3.1 Supervised Partitioning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72 7.3.1.1 Environment-based partitions . . . . . . . . . . . . . . . . . . . . 73 7.3.1.2 Speaker Partitions . . . . . . . . . . . . . . . . . . . . . . . . . . . 73 7.3.2 Unsupervised Partitions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73 7.4 Training the Ensemble Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74 7.4.1 MVE ....................................... 75 7.4.2 FactorAnalysis.................................. 75 7.5 Scoringforclassification................................. 76 7.6 Ensemble example: real data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 7.7 ExperimentsandResults................................. 81 7.7.1 ExperimentalSetup ............................... 81 7.7.2 Results ...................................... 82 7.8 Summary ......................................... 84 8 AUC and Ensemble Integration 85 8.1 EnsembleModeling.................................... 85 8.2 Optimization AUC of the ROC curve . . . . . . . . . . . . . . . . . . . . . . . . . . 86 8.3 Combining Ensemble Modeling and AUC of the ROC curve optimization . . . . . 87 8.3.1 Partitioning the data space (Ensemble) . . . . . . . . . . . . . . . . . . . . 87 8.3.2 Refining the Models (AUC approach) . . . . . . . . . . . . . . . . . . . . . 88 8.3.3 ChannelMismatch................................ 89 8.3.4 NoisyEnvironment................................ 90 8.3.5 Channel Mismatch and Noisy environment . . . . . . . . . . . . . . . . . . 90 8.4 Thespeakermodeling .................................. 90 8.4.1 MVE in an Ensemble approach . . . . . . . . . . . . . . . . . . . . . . . . . 91 8.4.2 Factor Analysis in an Ensemble Approach . . . . . . . . . . . . . . . . . . . 92 8.5 ExperimentsandResults................................. 95 8.5.1 ExperimentalSetup ............................... 95 8.5.2 Results ...................................... 96 8.6 Summary ......................................... 97
xiv CONTENTS 9 Discussion 99 9.1 MOBIOdatabase..................................... 99 9.1.1 MOBIOResults ................................. 100 9.1.2 ExperimentalSetup ............................... 100 9.1.3 Results ...................................... 100 9.2 Ahumada/Gaudi databases . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102 9.2.1 ExperimentalSetup ............................... 102 9.2.2 Results ...................................... 103 9.3 YOHOdatabase ..................................... 104 9.3.1 ExperimentalSetup ............................... 105 9.3.2 Results ...................................... 106 9.4 Summary ......................................... 107 10 Conclusions and Future Work 109 10.1FutureResearch ..................................... 110 A Traditional approaches preliminaries 113 A.1 ExpectationMaximization................................ 113 A.2 SupportVectorMachine................................. 114 B Review of Factor Analysis 115 B.1 FactorAnalysis...................................... 115 B.2 Joint Factor Analysis computation in detail . . . . . . . . . . . . . . . . . . . . . . 116 B.3 i-vectors.......................................... 120 C Notes about MVE 123 C.1 Minimum Verification Error . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123 C.2 AUC optimization using FA . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125
List of Figures 1.1 Simple Architecture of a Speaker Verification System . . . . . . . . . . . . . . . . 4 1.2 Scheme of the methodology followed by this thesis . . . . . . . . . . . . . . . . . . 5 2.1 Speaker Verification architecture in detail, The big picture ............. 10 2.2 Mel Frequency Cepstral Coefficient Front End . . . . . . . . . . . . . . . . . . . . . 11 2.3 Expectation Maximization Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.4 Decision Threshold of a binary classifier based on scores . . . . . . . . . . . . . . . 18 2.5 Detection Error Tradeoff curve . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 3.1 Comparison of MAP model optimization: a) only means, b) means and covariance 25 3.2 Minimum Verification Error architecture . . . . . . . . . . . . . . . . . . . . . . . 28 3.3 Comparison of target and imposter likelihood histograms . . . . . . . . . . . . . . 29 4.1 Shows the effect of the noise in the spectral domain. a) represents the clean speech signal in time domain. b) shows the clean speech signal spectrum. c) depicts the speech signal at 0 dB SNR. d) displays the spectrum of the same speech signal at 0 dBSNR........................................... 39 4.2 log scatter plot of the first MFCC clean speech against noisy speech at 0 dB. . . . 39 4.3 Scatter plot comparison between noisy and clean features: the continuous ellipses show the GMM component of speaker model in clean conditions, the dotted ellipses represent the same recordings with added noise at 0 dB. . . . . . . . . . . . . . . 40 4.4 Example of model distribution for one speaker, a) speaker model in clean conditions b) speaker model in noisy conditions (0 dB SNR). . . . . . . . . . . . . . . . . . . 40 4.5 Scores distributions for two scenarios: a) blue line represents the clean condition, b) red line shows the noisy condition. . . . . . . . . . . . . . . . . . . . . . . . . . 41 5.1 Infraestructure used by the Speaker Verification System . . . . . . . . . . . . . . . 48 6.1 Score distributions for three different classifiers. . . . . . . . . . . . . . . . . . . . . 54 6.2 ROC and DET curves for three different classifiers. . . . . . . . . . . . . . . . . . . 56 6.3 The effect of noise in the score distribution and in the ROC curve . . . . . . . . . 63 6.4 Vulnerable area in misverification measure d...................... 64 6.5 AUC optimization: AUC results for different systems, clean conditions . . . . . . 66 6.6 AUC optimization: AUC results for different systems in Noise Conditions . . . . . 67 6.7 AUC optimization: relative improvements. . . . . . . . . . . . . . . . . . . . . . . . 67 6.8 AUC Optimization: Score Distributions for 10dB SNR, babble Noise . . . . . . . 67 6.9 AUC optimization: ROC curve for 10dB SNR, babble Noise . . . . . . . . . . . . 68 6.10 AUC optimization: DET curve for 10dB SNR, babble Noise . . . . . . . . . . . . 68 7.1 Partition ensemble Scheme . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72 7.2 Partition Ensemble: Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76
xvi LIST OF FIGURES 7.3 Partition Ensemble: A priori . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 7.4 Partition Ensemble: Best score . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77 7.5 Partition Ensemble: Combination . . . . . . . . . . . . . . . . . . . . . . . . . . . 77 7.6 Partition Ensemble: Fusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 78 7.7 Partition Ensemble Scheme, real example . . . . . . . . . . . . . . . . . . . . . . . 80 7.8 Value of d for 3 different clusters and different iterations. . . . . . . . . . . . . . . 80 7.9 Discriminant function (log-likelihood) histograms for one target and 2 cohort samples 81 7.10 Ensemble modeling: relative improvements (clean condition) . . . . . . . . . . . . . 83 7.11 Ensemble modeling: relative improvements (noise condition) . . . . . . . . . . . . . 84 8.1 Ensemble-AUC modeling scheme . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86 8.2 Ensemble-AUC modeling modeling in detail . . . . . . . . . . . . . . . . . . . . . . 87 8.3 Ensemble-AUCscheme.................................. 89 8.4 Integration approach: AUC results for babble noise condition by scoring method. . 96 8.5 Integration approach: AUC results for babble noise condition by optimization approach. ......................................... 97 8.6 Integration approach: relative improvements ( babble noise condition) . . . . . . . 97 9.1 AUC of the ROC summary results for MOBIO database. . . . . . . . . . . . . . . 101 9.2 Relative improvement summary for MOBIO database. . . . . . . . . . . . . . . . . 102 9.3 AUC of the ROC curve summary results for Ahumada database. . . . . . . . . . . 103 9.4 Relative improvement summary for Ahumada database. . . . . . . . . . . . . . . . 104 9.5 AUC of the ROC curve summary results for Gaudi database. . . . . . . . . . . . . 105 9.6 Relative improvement summary for Gaudi database. . . . . . . . . . . . . . . . . . 105 9.7 AUC of the ROC curve summary results for YOHO database. . . . . . . . . . . . . 107 9.8 Relative improvement summary for YOHO database. . . . . . . . . . . . . . . . . . 107
List of Tables 4.1 Comparison of VTS EER and relative improvement (in %), 8 dB, [1]. . . . . . . . 43 4.2 Comparison of AFA EER and relative improvement (in %), [2]. . . . . . . . . . . . 44 4.3 Comparison of LP approaches EER and relative improvement (in %),10 dB, [3]. . . 44 5.1 Table presenting the common conditions . . . . . . . . . . . . . . . . . . . . . . . . 48 5.2 Table presenting the MFCC computation . . . . . . . . . . . . . . . . . . . . . . . 50 5.3 Table presenting the final results (EER) on the Test set for NIST 2008 . . . . . . . 50 5.4 JFA training set for female . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51 5.5 JFA training set for female . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51 5.6 EER results on the Test set for NIST 2008 for JFA approach . . . . . . . . . . . . 51 5.7 EER results on the Test set for NIST 2008 for the MVE approach . . . . . . . . . 52 6.1 On the selection of the number of cohort samples for discriminative training . . . . 65 6.2 AUC optimization: EER for MVE and JFA, clean condition . . . . . . . . . . . . . 66 6.3 AUC optimization: EER of the noisy task (babble noise). . . . . . . . . . . . . . . 66 6.4 AUC optimization: EER for different number of GMM components for . . . . . . 68 7.1 Scores for 10 clusters with respect to a target speaker model. . . . . . . . . . . . . 82 7.2 Scoring selection scores with respect to a target speaker model. . . . . . . . . . . . 82 7.3 Ensemble modeling: EER and minDCF for different clusters on clean condition . 83 7.4 Ensemble modeling: EER and minDCF for different clusters (noise condition), MAP and JFA . 83 7.5 Ensemble modeling: EER and minDCF for different clusters (noise condition), MVE ...... 84 8.1 Integration approach: EER and minDCF for different clusters on noise condition, MAPandMVE. ..................................... 95 8.2 Integration approach: EER and minDCF for different clusters on noise condition, JFAandi-vector...................................... 95 9.1 Final results (EER and minDCF) on the Test set for the MOBIO database. . . . . 100 9.2 Final results (EER and minDCF) on the Test set for MOBIO database. . . . . . . 101 9.3 Final results (EER and minDCF) for Ahumada database. . . . . . . . . . . . . . . 103 9.4 Final results (EER and minDCF) for the Gaudi database. . . . . . . . . . . . . . . 104 9.5 Final results (EER and minDCF) on the Test set for the YOHO database. . . . . . 106
xviii LIST OF TABLES
List of Acronyms AUC Area Under the Curve ASR Automatic Speech Recognition CMS Cepstral Mean Subtraction CMV Cepstral Mean Variance DCT Discrete Cosine Transform EER Equal Error Rate EM Expectation Maximization FAR False Acceptance rate FA Factor Analysis FRR False Rejection Rate GMM Gaussian Mixture Model GPD Generalized Probabilistic Descent Algorithm JFA Joint Factor Analysis log Logarithm MFCC Mel Frequency Cepstral Coefficients minDCF Minimum Detection Cost Functions ML Maximum Likelihood MLLR Maximum Likelihood Linear Regression MVE Minimum Verification NIST National Institute of Standards and Technology ROC Receiver Operating Characteristic SNR Signal to Noise Ratio SV Speaker Verification SVM Support Vector Machine UBM Universal Background Model
xx LIST OF ACRONYMS WMW Wilcoxon Mann Whitney
Chapter 1 Introduction To see a world in a grain of sand, And a heaven in a wild flower, Hold infinity in the palm of your hand, And eternity in an hour. - William Blake Auguries of Innocence Over the last decade, the advances in communication technology have leaded the research community efforts to focus more and more on secure and remote transactions over the networks [4]. User authentication has attracted major attention because of the emphasis on security issues. The first idea was to use passwords and rely on the users that these passwords would not be shared and kept secure [5, 6]. However, passwords can be stolen, shared or even forgotten. The necessity of authentication systems that can obtain information from human characteristics (biometrics) was a good solution [6]. At first, the research on biometrics concentrated on fingerprint, face recognition, iris recognition, DNA and speaker authentication, among others. Nowadays, the trend is to make a fusion of them to reach the best performance. From the variety of biometric signals that can be used for authentication, speech shows several advantages. First, it is the most natural source of human communication; the users do not need any extra device to produce speech. Second, the acquisition of the signal can be performed using wellknown non-sophisticated equipment like the telephone and internet network devices (common forms of information transfer). Moreover, speech opens the possibility to perform remote transactions over those networks. Lastly, the study in areas such as speech recognition and lately in speaker recognition had consolidated the understanding of speech. Other more sophisticated applications include forensics, surveillance and speaker personalization, all of them related to information security. In forensics, the main idea is to identify a suspect given a speech sample [7]. In surveillance, the networks are monitored so that a big amount of data of several speakers is available. The challenge is to locate a target speaker within that collection [8]. The newest devices, such as smart phones, rely also on speaker recognition in their applications. The speaker personalization of the software results in a more efficient operation of the device [9]. From all of them, authentication is the most representative example of speaker recognition. It has attracted major attention due to the continuous increase of users connected to the networks. Speaker recognition arose in the last three decades [10, 11] as a solution to provide security. Speaker recognition is usually divided into two subtopics: speaker identification and speaker verification. In speaker identification the target user is distinguished from a set of possible users. In speaker verification the identity of a user is accepted or rejected by analyzing a claimed identity and a
2 Chapter 1. Introduction speech phrase. Both fields share in essence the same principles. From these two, we will study Speaker Verification (SV), which main objective is to accept or reject a prospect user with the lowest error. The speaker verification traditional schemes are based on statistical hypothesis testing, where we establish two hypothesis: a null hypothesis Hs (accepts the speaker as legitimate) and the alternative hypothesis H¯s(rejects him or her) [12, 13]. Their relation is expressed as a ratio between likelihoods with respect to two models: target l and imposter. The output ratio is then tagged as accepted or rejected user. Two types of error commonly occur: false acceptances (FA) – incorrect decision due to accept a speaker who is not actually the target speaker, and false rejections (FR) – incorrect rejection of target speakers. Ideally, the probability of both types of errors must be zero; in practice, the two occur and are considered in the system design. A traded-off between them is established so that both errors are reduced. The probability of false acceptance can be reduced at the cost of increased false rejection. Thus, for any given system, the operating point can be manipulated to obtain a desired ratio between false acceptances and false rejections. The entire range of possible operating points is characterized mainly by the Operating Receiver Characteristic (ROC) curve and detection error tradeoff (DET) curve. This thesis shows how to reduce these errors using a discriminative optimization approach from two different perspectives. First, we present a complete strategy so that the reduction of every operating point at the ROC and the DET curve are improved. Moreover, we extend our findings to the noisy condition case. Second, we searched for more specific ways to describe the speaker space based on certain attribute. We built contiguous region models that characterized a specific attribute. These region models are an aid to build more specific target models. The discriminative optimization enhances those models and improves the results of the state of the art architectures. This strategy is also extended to the noisy condition case. Finally, the merge of both metrologies improve the current results. 1.1 Thesis Motivation and Scope The success of traditional SV systems lies in computing the adequate models that can clearly classify a target speaker from an impostor. The more suited these models are to specific scenarios or target speakers, the best results we can obtain [12, 13, 14, 15]. The maximum likelihood (ML) approaches using generative modeling [12, 13] and lately factor analysis [14, 15] represented a huge improvement in the systems performance. Generative modeling relies on maximizing the likelihood of a model to a current set of data. However, the reduction of the false acceptance and false rejection errors is not taken as a primary objective but as an effect of an accurate modeling of the target speakers. Discriminative training approaches, first employed in speech recognition [16, 17, 18, 19, 20], proposed an alternative solution that included the reduction of the error by considering imposter samples in the optimization of the models. Just a few studies are found in the bibliography [21, 22, 23, 24], probably because the need of high performance computation. But nowadays, with the advances in technology, the computation is not longer and issue and the experimentation is feasible. One of the schemes that fit the SV theory is the minimum classification approach. Its main goal is to minimize the empirical classification error regardless the distribution of the data [16, 17, 25]. Hence, we considered it to be the backbone of our research and the motivation to extend its potential to SV. In this sense, the algorithm must focus on the error reduction. Moreover, the algorithm must be designed in a way that it can obtain competitive results as the ones produced by traditional approaches. Lastly, we observed that traditional methods usually measure the improvements on a single operating point, called Equal Error Rate (EER), that acts as a summary of the full range of possible outcomes. In our case, the motivation is to find a methodology to
Chapter 2 Classical Speaker Verification Systems Me di cuenta de que ten´ıa que revolucionar; aprender cosas nuevas para no quedarme atr´as. Me di cuenta y me rebel´e. - Jaime Sabines This section presents an overview of the state of the art of the Speaker Verification systems. It details the traditional way to address SV by using a Gaussian Mixture Models (GMM) framework, joint factor analysis and i-vector or discriminative approaches like maximum mutual information estimation (MMIE), support vector machines (SVM), and minimum verification error (MVE). As depicted in Figure 2.1, the SV systems have two main stages: enrollment (training) and verification (test). In the enrollment, the set of acoustic models are trained. In the verification, the speech trials evaluate the acoustic models and produce an accept/reject result. The first step of both stages is the feature extraction. It converts the speech signal into a vectorial representation that contains specific information about the speaker. In the training, we can observe that there are several options of training algorithms available. Although they have the same final goal, to obtain lower error rates, the approaches can address different aspects. We can divide them into generative, discriminative and hybrids. The generative modeling optimize the probability density functions (pdf) so that they can describe the target speakers, examples of these algorithms are MAP and Factor analysis variants. The discriminative modeling usually minimize the empirical error rate using samples of both the target speaker and impostors, examples of these approach are maximum mutual information estimation, MMIE, support vector machines, SVM, and MVE. The combination of methodologies also occurs, for instance i-vectors and SVM approach. For the purpose of our research we embed a discriminative optimization into the current generative models so that the error rates are lowered. For the purpose of this research we will just focus on FA, that is considered the state of the art and MVE as examples of both possible scenarios. In the verification, the models already trained are evaluated using unknown speech, a score (usually as likelihood ratio) is produced as an output. The score is then fed into a classifier with a predefined threshold and a decision (accept or reject a user) is made. The SV systems depending on the application are divided into text-dependent and textindependent. The text-dependent systems require that the user utters the same set of words during the training and test phases. In some cases, also belonging to text-dependent, the system can ask the speaker to utter a set of words at the test stage. Both example have predefined acoustic models depending on the text and speaker. For the text independent systems the user can speak any set of words, sometimes even in other languages, making the task harder. Usually,
10 Chapter 2. Classical Speaker Verification Systems the text-dependent systems can get better performances; and the text independent SV is still a challenge. The next sections describe how the current systems are built based on generative modeling (MLE and MAP) approach, joint factor analysis and i-vector, and discriminative approaches like MVE. normalization and decision waveform User Enrollment waveform feature extraction score computation accept reject i-vector/ PLDA target speaker i-vector models claimed Identity features scores feature extraction User Verification models UBM EM Joint Factor Analysis target speaker & channel factors MAP target speaker models MMIE target speaker models & imposter-models SVM speaker hyperplane models Minimum Verification Error target speaker models & imposter-models GENERATIVE MODELING DISCRIMINATIVE MODELING Figure 2.1: Speaker Verification architecture in detail, The big picture 2.1 Feature Extraction The proper selection of relevant information contained in the voice signal is a fundamental step. Different feature extractions have been explored by the research community: spectral, prosodic and high-level. The spectral features, which we will address in the next sections, convey the acoustic information of the vocal tract [30, 31]. The prosodic features refer to the intonation and the stress that a person uses when speaking; it depends on several factors including, for instance, emotional state and the situation the speaker is trying to communicate (sarcasm, sadness, among others)
2.1 Feature Extraction 11 [32, 33]. The high-level features extract information about the manner in which a person speaks (the lexicon), the different topics she addresses, and the style in which she uses the words [32, 34]. In this research, we will focus on the spectral feature extraction. The model used to represent the extraction process is based on a sound source – the larynx– and a filter – the vocal tract [35]. The parametrizations that work under this scheme are commonly divided into filter bank (FB) cepstral analysis and linear predicting (LP) coding [36]. For the bank filter approach the most known techniques are Mel Frequency Cepstral Coefficient (MFCC ) and Linear Frequency Cepstral Coefficient (LFCC). Both, the MFCC and LFCC are subject to the shape and the frequency of a filter bank. In the case of the MFCC, the filter banks follows a logarithmic spacing at high frequencies [30]; the LFCC adopts a linear spacing filter bank [37]. For linear predicting coding, the most representative techniques are Linear Prediction Cepstral Coefficients (LPCC) and Perceptual Linear Prediction (PLP). Both employ a linear process that can predict the speech signal at each time given previous samples. The LPCC includes an all-pole model that can describe the spectrum with a smooth envelope [38]. Finally, the PLP uses psycho-acoustic elements to make the prediction of the speech under an all-pole model similar to LPC [39]. We will briefly describe the MFCC procedure; we consider that it is the most frequently employed by the research community. 2.1.1 MFCCs computation The computation of the MFCCs is composed of several stages [30], as shown in Figure 2.2. The first stage is to pass the speech signal through a preemphasis on an overlapping hamming window, we obtain the modified, S(i), signal. The next step is to transform S(i) to its spectrum S(w), employing the short-time Fourier analysis . We can extract either the power or the magnitude of the Fourier coefficients, Sm(w). Afterwards, a filterbank (FB) transforms this signal into a smooth spectrum representation (close to the envelope). The filterbank output is then converted to the log-domain. Finally, we apply the DCT to decorrelate and produce the cepstral coefficients [31]. The filterbank can be linearly spaced, the resulting coefficients are named Linear Frequency Cepstral Coefficients. However, the most common used are the MFCCs. They follow the mel scale that resembles the way a person hears. To emphasize the dynamic features of the speech in time, the time-derivative (∆) and the time-acceleration (∆2) are usually computed. It is common to compute 12 MFCC, one Energy coefficient and its corresponding (∆) and (∆2). However, recent studies have shown that using more than 12 MFCCs can give better results [40, 28]. 1 € S i ( ) € F⋅ {} € S ω ( ) € ⋅ FB C1 C2 C3 . . C12 Energy C0 € Sm ω ( ) € log ⋅ ( ) DCT Figure 2.2: Mel Frequency Cepstral Coefficient Front End 2.1.2 Feature Normalizations Normalizations at this stage are implemented to reduce the effects of the noise and the channel distortion. For instance, the cepstral mean subtraction (CMS) [41] is a blind deconvolution that comprises the subtraction of the utterance mean of the cepstral coefficients from each feature. In 1The research in [40, 28] show that from 16 to 19 MFCCs, the performance is better than with the usual 12 MFCCs.
12 Chapter 2. Classical Speaker Verification Systems the same way, the variance normalization (CVN) [42] is also applied. Hence, the new features will fit a zero mean and variance one distribution. Another well-known feature normalization is RASTA (Relative Spectra) [43]. While CMS focus on the stationary convolution of the noise due to the channel, RASTA reduces the effect of the varying channel, it removes low and high modulation frequencies. The three of them are commonly used in the SV architecture. 2.1.2.1 Feature Warping A special normalization for SV at the feature stage is the feature warping. It belongs to the Gaussianization methods [44, 45]. The underlying concept in this normalization scheme is that every spectral attribute (cepstral coefficient in our case) is normally distributed across time, but the transmission channel distorts such distribution. The task of feature warping is to undo the distortion caused by the channel by warping each attribute’s scale so that the resulting attribute set has a normal distribution. Feature warping is accomplished by first assembling an empirical CDF (cumulative distribution function) from the ranked features after and before the current frame, and then performing the CDF-inverse at the current frame. 2.1.2.2 Feature Frame Removal Frame removal is based on the idea that low energy frames do not provide information about the identity of a person. The frames’ log-energy of each utterance are modeled by a three-component GMM. w1corresponds to the highest weight of the rightmost Gaussian, w2to the middle Gaussian, and w3to the leftmost Gaussian. According to this model every frame log-energy is labelled as high if it belongs to the rightmost Gaussian; medium if it belongs to the middle Gaussian; and low if it belongs to the leftmost Gaussian [12]. The following Equation is used to determine which frames can be extracted. N=w1+ (g∗α∗w2),(2.1) where g is a value between 0 and 1, and αis an heuristic weighting parameter. Nis the percentage (between 0 and 100) of the frames with highest energy that will be extracted. If an accurate voice activity detector (VAD) or the speech transcriptions to perform speech recognition are not available the frame removal is a suitable solution. 2.2 Speaker Modeling After obtaining the feature vectors, the aim is to design an algorithm capable to verify if a new spoken phrase belongs or not to a specific user. To successfully achieve this goal, the next stage, commonly known as training, is to extract the relevant information of each user and construct a suitable model. Hence, the system can match the new phrase with that specific model and give a decision about whether it is a speaker or not. In the early stages of SV research, the distance based methods were popular. Among them, the Dynamic Time Warping (DTW) and the Vector Quantization (VQ) were extensively studied. The DTW proposed a non-linear mapping or warping of one signal (test recording) to another (target speaker utterance) [46]. The key point is to minimize the distance between both of them and come up with a final decision. In the case of VQ, the purpose was to partition the data space for a speaker sinto non-overlapping regions and obtain a codebook that represent that space [47]. When a new recording comes into the system, a likelihood measure between the incoming features and the codebook is computed. At that stage, the experiments were performed just for a few
2.2 Speaker Modeling 13 speakers and usually for text-dependent utterances, but it was difficult to extend the methodology to text-independent databases with hundreds of users. The challenge, then, became how to produce suitable models for the new scenario. The solution came along with the advances in computing systems. The theoretical statistical methods like generative and discriminative modeling were then possible. The tendency was to construct overlapping models that can be more robust. A way to solve such a problem was to employ a statistical approach. The Gaussian Mixture Model (GMM) became that answer capable to represent a speaker in terms of a distribution with just a few parameters. We will leave the GMM approach to be fully explained in detail in the following sections and chapters, including the GMM approach that signified a forward step [48, 49]. Later, it was also shown in [50], that the previous VQ is a special case of the GMM. This thesis stands on the GMM statistical approach. We analyze SV as a classification problem that is solved using pattern recognition techniques [51]. The solution is given in terms of statistical hypothesis testing [52] and Neyman-Person Lemma [53]. The Lemma states that the relation between two competing point hypothesis models can be expressed as a ratio. The ratio is then compared to a decision threshold and an accept/reject choice is made. The null hypothesis HSaccepts the speaker as legitimate and the alternative hypothesis H¯ S rejects him/her. Under this framework, for a set of observations X=χ1, χ2, ..., χT, the ratio of two hypothesis (target-model and imposter-model) is defined as follows: θ(X) = p(HS|X) p(H¯ S|X)> τ accept HS < τ accept H¯ S,(2.2) where p(HS|X) and p(H¯ S|X) are the posterior probabilities of how likely a hypothesis HSor H¯ S is to happen given the observed data, and τis the decision threshold. The hypotheses HSand H¯ Sare described by models, Λ and ¯ Λ respectively. Including the distribution of Λ and ¯ Λ in the formulation, Equation 2.2 is defined by, θ(X) = p(HS|X, ΛS) p(H¯ S|X, Λ¯ S)=p(X|ΛS) (p(HS)) p(X|Λ¯ S) (p(H¯ S)).(2.3) where p(X|ΛS) and p(X|Λ¯ S) are the likelihood functions with respect to the target and imposter model. The likelihood denotes how probable Xis an outcome of the model Λ.2Additionaly, recall that p(HS) = 1 −p(H¯ S). Then, Equation 2.3 can be transformed in the log domain, θ(X) = log (p(X|ΛS)) −log (p(X|Λ¯ S)) + C, (2.4) where log (p(X|ΛS)) and log (p(X|Λ¯ S)) are the log-likelihoods corresponding to the target model ΛSand the imposter-model Λ¯ Sand Cis a constant. To fulfill the Neyman-Pearson Lemma the following should be considered. Both models are defined by distributions that must be known in advance to reach optimality. However, they are not available in practice and should be estimated. We also need optimal size data sets to estimate those distributions. The more training data at one’s disposal, the better estimation we can compute. However, the data sets are generally of limited size, because of the cost of collecting and organizing the utterances. Hence, the estimation of those hypotheses can give approximate solutions. The final ratio – commonly denoted as score – is given by a simplification of Equation 2.4, θ(X) = log (p(X|ΛS)) −log (p(X|Λ¯ S)) > τ accept HS < τ accept H¯ S .(2.5) The classification problem is reduced to find optimal solutions for the distributions ΛSand Λ¯ S. Two approaches are commonly used to define Λ¯ S. The first one considers a set of competing 2For further details on Bayesian modeling refer to [54, 55]
14 Chapter 2. Classical Speaker Verification Systems speakers ¯ S1, ..., ¯ SN, called cohort, for each user S[56, 57]. The second one employs just a genderdependent set for all Starget users, usually named UBM (universal background model or impostermodel). An advantage of the last one is that we just need one imposter-model for all the target speakers and has been used extensively. The next sections show various frameworks that have been effectively used to estimate the distributions of the speakers and imposters. To compute these models two branches are possible: generative modeling and discriminative training. Traditional SV systems focus mainly on generative modeling (on how to obtain an accurate model of the target speaker voice). However, discriminative training, which estimates more specific models for target and imposter data, has acquired attention recently for different approaches [24, 58, 59, 60]. Discriminative training addresses the problem from a different point of view, it optimizes the correctness of a model by formulating an objective function that penalizes the model parameters using positive and negative data examples. Some other successful architectures employ a combination of both [15, 61, 14]. In the next paragraphs we will describe the classical descriptions of both of them and emphasize the outstanding characteristics. 2.2.1 Generative Model approach The generative model is a solution to represent the distinctive characteristics of the speech of a target speaker. The only information available are the data recordings with the corresponding identity label of the speaker and a proposed statistical model to be used. The challenge is to find a distribution (model) that can describe the phenomena– a model that is able to capture the statistics of the vocal tract of each speaker. Hence, the key is to find the maximum likelihood between our data and the proposed model. 3 The Gaussian Mixture Models, GMM, showed to be a plausible statistical model in which the speech is represented just by few parameters (means, variances and weights) in a simple way. Most of the current state of the art strategies are based, in many ways, on GMM approach to compute the desired target and imposter models [49, 63, 64]. Let’s define p(X|Λ) as a GMM distribution for a D-dimendional feature vector as, p(X|Λ) = M X i=1 wiN(X|µi,Σi) (2.6) where Mare the number of components of the model, wiare the weights of each component with PM i=1 wi= 1, and N(X|µi,Σi) is the Gaussian probability density function with µimean, and Σi is the covariance diagonal matrix. The computation of the parameters of the GMM is then the challenge to address. There is not an analytical solution to find the parameters that satisfy the maximum likelihood, but the problem solution employs an iterative algorithm, called Expectation Maximization, that optimizes the model parameters. 2.2.1.1 Expectation Maximization The Expectation maximization (EM) [65] is the leading algorithm for training the GMM, under maximum likelihood criteria. The goal is to search for the maximization of the “likelihood”, p(X|Λ); i.e., the probability of the realization of Xfrom an unknown distribution, given the model Λ. In every iteration the algorithm updates the GMM model parameters, Λ = {wi,µi,Σi}, 3For further details about maximum likelihood and the algorithm used we can refer to [62]. An intuitive explanation of maximum likelihood is to find a model that best describes our data.
2.2 Speaker Modeling 15 such that, p(X|Λ∗ (n+1))≥p(X|Λ∗ (n)), where Λ∗is the model estimation at each iteration. The new model then becomes the old model until convergence is reached. Although there are several explanations to formulate the EM algorithm, in this section we briefly explain the intuition behind it and leave further details to the next chapters. To obtain a probability distribution that describes our data, EM comprises two stages expectation and maximization. First we define a “hidden variable”, Zthat helps to maximize p(X; Λ). Let us define, p(X|Λ) = Pz(X, Z|Λ). However, marginalizing Zbecomes a difficult task. EM strategy is to update the parameters of Λ∗in two steps. In the Expectation we define the function gnthat lower bounds the objective likelihood function log p(X|Λ). The analytical solution is to compute the posterior distribution p(Z|X, Λ). The initial model parameters at this stage are usually computed using VQ techniques. In the maximization step Λ(n+1) is obtained by maximizing gn(see Figure 2.3). In other words, the new Λ∗is computed by maximizing the joint distribution of Xand Z. Figure 2.3: Expectation Maximization Algorithm The purpose of the EM algorithm in SV is to produce a general model that embraces the characteristics of all speakers (sometimes, excluding the target speakers set). The GMM model at this stage is commonly known as imposter-model or Universal Background Model (UBM), and in most cases it is gender-dependent. At this point, we have solved part of the problem by estimating the imposter or UBM model, Λ∗. The straightforward idea might be to extend the idea of EM and compute the models for each target speaker. However, the availability of the data is the main constraint. The datasets usually contain recordings of just a couple of minutes [40], not suitable for ML training. Therefore, an algorithm able to perform an adaptation of the current general model to a target specific model is needed (presented in the next section). 2.2.1.2 Maximum A posteriori (MAP) ML is not able to compute accurate parameters for the target speaker models because the data available is limited. However, an algorithm proposed by [66], similar to EM, is currently used to adapt the imposter-model or UBM to each target speaker. The formulation is based on the estimation of the model parameters, Λ, that maximize the posterior, p(Λ|X). 4The new estimated model is denoted as ˆ ΛMAP = argmax p(Λ|X). The solution of this adaptation includes the prior p(Λ) in the formulation. It follows the rule, ˆ ΛMAP = argmax Λ p(X|Λ)p(Λ).(2.7) 4We recall that by Bayes p(Λ|X) = p(X|Λ)p(Λ) p(X).
16 Chapter 2. Classical Speaker Verification Systems The prior belief, p(Λ) is usually given by a previous estimation of a general model, in this case the UBM. The solution of MAP using GMM is intractable, hence, the algorithm is treated in a similar manner as EM as will be shown in next sections in detail. Note that to obtain an optimal solution the system should consider: a) the definition of the prior models, b) the appropriate estimation of those priors. 2.2.2 Discriminative Training Approach The generative modeling focuses on optimizing for a single class. In the case of the UBM, the training optimizes the parameters of a general model; in the case of MAP, it updates the parameters of a target speaker model. Hence, the classifier depends entirely on how good the models fit the data they represent (the better the representation, the better the classification). But these methods does not focus on the separation between the classes, target and imposter. In this sense, the discriminative modeling gives an alternative solution to the problem. In general, the discriminative approaches consider the opposing classes at the same time and optimize the models for both. They mainly pursue the reduction of an error, rather than improving the model likelihood with respect to the data set. However, the optimization process needs positive and negative samples for each class.The main drawback of these approaches is the amount of data needed to obtain appropriate models and the risk of over-fitting when the models become too specific to the data from which they were trained. In SV, the discriminative techniques have not been extensively used by themselves, but their properties can improve the current approaches or are competitive to the Generative Models. In this section we show some examples of them. 2.2.2.1 Maximum Mutual Information Estimation (MMIE) One of the most representative approaches of discriminative modeling is MMIE [67, 68]. The main objective of this algorithm is to maximize the mutual information between the observations X(belonging to either target or imposter user) and the model tag (Λtar, Λimp). Recall that when two random variables are dependent, the mutual information is maximized. If they are independent, the mutual information is minimal. The probability of the observation Xis given by p(X) = p(X|Λtar)p(Λtar) + p(X|Λimp)p(Λimp).(2.8) Clearly, the criterion is defined in the log domain as the sum over the posteriors of the observations, f(Λ) = X X∈Xtar logp(X|Λtar)P(Λtar) p(X)+X X /∈Xtar logp(X|Λimp)p(Λimp) p(X).(2.9) The estimated Λ∗is, Λ∗= argmax f(Λ) (2.10) It is computed, Λ∗ n+1 = Λn+η∂f ∂ΛΛ∗ n (2.11) As shown, this method tries to maximize the class-conditional probability of the observation while the weighted sum of the competing class-conditional probabilities is minimized. Hence the separation between classes is optimized. This observation guarantees that the impostor and target distributions are optimally separated and improve the error. MMIE is the preliminary of more sophisticated discriminative modeling such as MVE discriminative modeling as will be presented in the next chapters.
2.3 Decision Making 17 2.2.2.2 Support Vector Machine (SVM) Another discriminative approach that has attracted attention recently is SVM. The purpose of the SVM is to classify the multivariate data Xinto two classes tagged binary as [1,−1] [69, 70]. The data vectors X, that are not linearly separable, are mapped to a higher dimensional space via a kernel function. The solution is a hyperplane with maximal margin, meaning that the distance between a tagged vector and the hyperplane is maximal. A hyperplane is trained by using data X and its tags. In the testing stage, the hyperplane determines the class of each vector. Two ideas have mainly been developed for SVM [12]. The first one focused on the scoring. The knowledge, in the training stage, of the scores that belong to a user is employed to compute a hyperplane. In the test, the data vectors are compared to the hyperplane and a decision is produced. The final score is the average of the trial data for each target user. The second one, is to combine the generative method GMM with the SVM [71]. Instead of evaluating each vector, the whole phrase is considered, a new input feature vector is produced for the SVM, and just a final decision is obtained. This techniques have also been successfully applied to spectral, high-level and lately to prosodic features [33, 34]. The evolution of the SVM for SV has become of interest because it can obtain competitive results compared to the generative modeling approach. However, its success is mainly due to the improvements in the kernel design and its capability to refine the modeling in other techniques such as JFA and i-vectors [14]. 2.3 Decision Making After computing the models for every target and imposter speaker set (or UBM), the next step is to evaluate the system, obtain feedback from testing on a “controlled database” 5and perform further normalizations. The target models evaluate different unlabeled recordings (trials) and from each of them we compute a log-likelihood ratio with respect to a specific target model. This ratio, usually named score, computes a numerical value of the relation between two competing hypothesis. Depending on the value the system gives a decision (see Equation 2.5). The score for every trial follows a hypothesis test framework; however, the classifier is not perfect. In the process to verify if the speech signal Xbelongs to a target user S, two errors may arise: •Type I : known as false rejection, meaning that signal Xis incorrectly rejected being a target speaker. •Type II: known as false acceptance, meaning that the signal Xis incorrectly accepted being an impostor. In Figure 2.4, note the location of the errors and the position of the threshold, τ. The scores truly belonging to a target speaker follow a Gaussian distribution. The same is valid for the imposter distribution. Both overlapping distributions represent the classifier. The challenge now is to position the threshold. On one hand, an approach is to set τas high as possible. However, that will produce few false acceptances and many false rejections, meaning that it will be very secure. The target user might need to perform several trials to finally get access. On the other hand, if τ is designed to be low, there will be many false acceptances and few false rejection. Many imposters might be accredited by the system. Neither of them is desirable. A good tradeoff that depends on the specific application is appropriate. The usual technique to place an effective τbetween the two competing score distributions is to perform several experiments using a development database and harden the threshold. 5The development database from which we know that ground truth, helps us to evaluate the system in advance and tune the threshold according to specific requirements.
18 Chapter 2. Classical Speaker Verification Systems Figure 2.4: Decision Threshold of a binary classifier based on scores 2.4 Score normalization The scores exhibit variations produced by multiple targets speakers and multiple conditions that need to be compensated. Therefore, it is convenient to scale (or normalize) the scores so that they are comparable across multiple targets. Two types of normalizations are the most popular: with respect to the target data (Z-norm) and with respect to the test data (T-norm) [12]. In the following paragraphs we describe them. 2.4.1 Z-norm The Z-norm [72] or zero normalization is performed with respect to the target speaker statistics and compensates for inter-speaker variability. Equation 2.12 shows this type of normalization. θ(Xtrial, S)norm =θ(Xtrial, S)−µ(S) σ(S),(2.12) where µand σare the mean and the standard deviation of the scores computed for a certain speaker model S, given imposter data. Xtrial is the test input data. One of the advantages of this normalization is that it can be performed offline. 2.4.2 T-norm Other well known normalization is T-norm [73] or test normalization. In this case the test utterance is evaluated against imposter models as shown in Equation 2.13. θ(Xtrial, S)norm =θ(Xtrial, S)−µimpost(Xtrial) σimpost(Xtrial),(2.13) where µimpost and σimpost are the mean and the standard deviation obtained by evaluating the upcoming phrase against precomputed imposter models. The computation is performed on-line and it is preferred to have separate imposter models for each speaker. The ZT-norm is the combination of the above and can also be TZ-norm, depending on the order in which they are applied. Most of the state of art systems use one of the normalizations or a combination of them [74]. 2.5 Evaluation and Performance measures The goal of a speaker verification system is to classify the target speakers correctly and minimize the cost of the FA and miss errors. The performance of the verification systems is traditionally
3.3 Minimum Verification Error 25 explain some discriminative approaches that will aid us to understand the next chapters. −3 −2 −1 0 1 2 −3 −2 −1 0 1 2 3 −3 −2 −1 0 1 2 −3 −2 −1 0 1 2 3 Figure 3.1: Comparison of MAP model optimization: a) only means, b) means and covariance 3.3 Minimum Verification Error The theory behind minimum verification error was first conducted to solve automatic speech recognition tasks [18, 25] and was called minimum classification error (MCE). The algorithm was first conceived as a solution for a classifier optimization problem proposing a methodology to minimize the probability of misclassification 1[25, 17, 84]. The structured formulation under Bayes Decision Theory and the competitive results compared to the state of the art classifiers made it attractive. The algorithm provided an original answer to the optimization of the model parameters, regardless the nature of the distributions and focusing only in the minimization of the error. Moreover, it established an estimation criterion which leads to the update of competitive models in an iterative mode. Afterwards, the approach was embedded into the Hidden Markov Model (HMM)2theory. The purpose was to optimize the parameters of acoustic unit models such as phones or words [25]. For example, an specific phone becomes a target and the remaining ones the imposter set. Then, an iterative process takes place for that specific phone model, but also for its counter imposter model. Later in [87], MCE and MVE were applied for the combined task of utterance verification and SV. In this work, a unified scheme for both tasks was adopted for a two-pass verification system. Recently, another study presented by [88], proposed an improvement to characterize the alternative hypothesis following an optimal discrimination between the target speaker and imposters’ set. From SV point of view, the estimation of the distributions of Sand ˆ Sas computed by EM and MAP, can also be tackled using the binary version of the algorithm, known as MVE (Minimum Verification Error). Under this approach, the empirical error rate is minimized using an appropriate objective function and positive (target speaker features) and negative (imposter speaker features) examples. The procedure performs several iterations while updating the distribution parameters until reaching convergence or a specific threshold. For minimum classification, supposing that several classes are possible, the prime problem is to correctly classify χdata in Cnclasses.3The following piecewise loss function is an appropriate 1MCE was studied as an alternative solution to the novel classifiers such as Artificial Neural Networks [82] and Vector Quantization algorithms [83]. 2The complete explanation of the classical ASR process can be found in [85, 86] and is not under the scope of this thesis. However, we would like to emphasize the influence of ASR techniques and advances that are used in SV with the appropriate modifications. 3In future chapters we will use this approach to classify among attributes such as different SNRs. For now, we follow this formulation and its simplification to a SV scenario.
26 Chapter 3. Speaker Modeling option to measure the performance of such classifier. en,l =0n=l 1n6=l,(3.24) where n, l ={S, ¯ S}. This function assigns a penalty of zero for correct classification and one for incorrect classification. Moreover, we can define the conditional loss as follows, <(Cn|χ) = el,n p(Cl|χ).(3.25) If we denote C(χ) as the classifier decision, based on the observation χ, then for every χthe classifier has to be designed to achieve, <(C(χ)|χ) = min n<(Cn|χ).(3.26) The above Equations imply that, <(Cn|χ) = P(Cl|χ) = 1 −p(Cn|χ).(3.27) Then, the optimal classifier that solves for the minimum loss is stated as, C(χ) = Cnif p(Cn|χ) = max lp(Cl|χ).(3.28) Equation 3.28, based on Bayes theory, clearly transforms the classification problem into a distribution estimation problem. This Equation can be solved by MAP approach if we have complete knowledge of the distributions and enough data for each speaker. In most practical tasks, this is not the case. A useful solution is to employ a smooth representation of the objective function that is suitable for both the minimization of the error and the optimization. In [25, 17, 84], an optimization criteria is proposed based on likelihood functions and a classifier that operates with the following decision rule, C(χ) = Cnif gn(χ; Λ) = max lgl(χ; Λ) .(3.29) To accomplish this objective three elements are needed and will be discussed in the next paragraphs. Let us first define a set of discriminant functions, such as log-likelihood functions that can be plugged into the objective function. Let gn(χ; Λ) = log (p(χ; Λn)) evaluate the log-likelihood function for class Cnby observing χ. Secondly, for any speaker we define the misverification measure as, dn(χ) = −gn(χ; Λ) + Gn(χ; Λ) ,(3.30) Gn(χ; Λ) is considered the log-likelihood average of the competing classes. Gn(χ; Λ) is defined using the following expression, Gn(χ; Λ) = log {1 M−1X n6=l exp [ηgl(χ; Λ)]} 1 η,(3.31) ηis considered a slack factor to control the contribution of the discriminant functions. For the special case of SV, M= 2 (two classes are possible: target or imposter). Furthermore, if M= 2, then η= 1. Equation 3.31 can be reduced to Gn(χ; Λ) = log g(χ; Λ¯ S),(3.32)
3.3 Minimum Verification Error 27 Thirdly, let’s define the new loss function `n(χ; Λ) = `(dn), a sigmoid function that makes a soft decision centered at a threshold: `(dn(χ)) = 1 1 + exp (−γdn(χ) + θ),(3.33) where γ≥1 and usually θ= 0. If dn(χ) is less than zero, it means that not error occurred; if dn(χ) is positive, an error happened and it is penalized. With this three elements we can accomplish a smooth representation of a composite objective function. Finally, we use the indicator function 1(·) to sum the losses over the two classes (target and imposter) and review the performance: `(χ; Λ) = X n `n(χ; Λ)1(χ∈Cn).(3.34) The optimization of the parameters by minimizing the loss function is solved using the generalized probabilistic descendent (GPD) algorithm [25] as shown in Equation 3.35, Λt+1 = Λt−∇`(χ; Λ),(3.35) where is the learning step, and ∇`(·) denotes the gradient. The above procedure is applied to every element of the model parameters Λ = {µC k,ΣC k, wC k}. We perform the optimization dimension by dimension. Then, for target speaker CSand for dimension d, for instance, for µS, to avoid bias, let µ=σ˜µ. We compute the following (note that we dropped out all subindices to make the formulation clear), ˜µ(t+ 1) = ˜µ(t)−∂`(χ; Λ) ∂˜µ|Λ=Λt,(3.36) By the chain rule we compute, ∂`(χ; Λ) ∂˜µk =∂`(χ; Λ) ∂d ∂d ∂˜µk ,(3.37) The first element is then, ∂`(X; Λ) ∂d =γ`(d)(1 −`(d)),(3.38) The misverification measure for the two types of errors: ∂d ∂˜µk = −∂g(X;Λ) ∂˜µk+∂G(X;Λ) ∂˜µkX∈C(Type I) ∂g(X;Λ) ∂˜µk−∂G(X;Λ) ∂˜µkX /∈C(Type II) (3.39) ∂g(χ; Λ) ∂˜µk =1 p(χ; Λ) ∂ p(χ; Λ) ∂˜µk ,(3.40) ∂G(χ; Λ) ∂˜µk =1 p(χ;¯ Λ) ∂ p(χ;¯ Λ) ∂˜µk .(3.41) Then, consider for a specific k, dimension d, the final Equation to compute the mean is, ∂p(χ; Λ) ∂˜µk =wkN(χ|˜µk, σk) Pk0wk0N(χ|˜µk0, σk0)χ σk−˜µk.(3.42)
28 Chapter 3. Speaker Modeling The computation for Σ and wC kadopt the same procedure (for clarifications and further details refer to Appendix C.1) Once again, let ˜σ= log(σ),Then the final Equation to compute σis ∂p(χ; Λ) ∂˜σk =wkN(χ|µk,˜σk) Pk0wk0N(χ|µk0,˜σk0)(χ−µk ˜σk2 −1).(3.43) Finally, for the set of weights wkand to ensure that PK k=1 wk= 1, let, w=exp( ˜w) PK k=1 exp( ˜w). ∂p(χ; Λ) ∂˜wk =N(χ|µk, σk) Pk0˜wk0N(χ|µk0, σk0).(3.44) € X +d € X∈Ci if yes no d>0 d>0 + 1 if True acceptance 0 if false acceptance 1 if True rejection 0 if false rejection € ∑ € ∑ € L1 € L2 € L € ω 1 € ω 2 Figure 3.2: Minimum Verification Error architecture Figure 3.2 describes the formulation in detail. For the training stage and for an i-th target speaker, MVE depends upon an initial target model to evaluate the likelihood gi(X; Λ), an imposter model to compute Gi(X; Λ) and a set of target and imposter or cohort speech tokens4. Each initial target model can be obtained from the adapted model estimated by MAP or EM. The initial imposter and target models are evaluated using the data-vector χtand Equation 3.31. Afterwards, we compute the misverification distance dand map it on the loss function (sigmoid). Since the system compasses the ground truth of the data, it is able to count for the errors and penalize accordingly for FA and FR. Hence, we construct a loss function that depends on the errors `. The optimization of the model parameters with respect to the minimization of the loss function is the final calculation. The final outcome for an i-th speaker are the optimized target and imposter models under the MVE framework are described in Algorithm 3.3. In the test stage, the system evaluates the unlabeled recording employing the new model parameters for target and imposter. Although MVE is a method per se that can handle random initialization models, it has shown that if combined with techniques based on ML estimation, the models get refined and the classification more accurate [90, 88]. It is possible to analyze the effect of the MVE in terms of likelihood (see Figure 3.3). The plot depicts the output histograms of the discriminative functions gn(χ; Λ) and Gn(χ; Λ) before and after optimization. In this example, we chose a scenario where a target user and an imposter show a similar likelihood values (as depicted in Figure 3.3.a ). The top histograms show the discriminant functions with respect to a target speaker, purple for target user and pink for imposter. The next graph shows histograms with respect to the cohort or UBM. The green plot describes the distribution of the cohort likelihoods and the gray histogram belongs to the target speaker. We clearly observe the overlapping of the histograms. After 10 iterations using MVE, Figure 3.3.b shows how the histograms broaden as they also separate. This effect aids the accuracy of the classifier improving the results. 4The cohort is commonly defined according to the application [89, 57]; for instance, the selection can be based
3.4 From models to supervectors 29 −80 −70 −60 −50 −40 −30 −20 0 50 100 log likelihood sample from true speaker −80 −70 −60 −50 −40 −30 −20 0 50 100 log likelihood sample from cohort speaker (a) Discriminant function histogram before optimization −80 −70 −60 −50 −40 −30 −20 0 50 100 log likelihood sample from true speaker −80 −70 −60 −50 −40 −30 −20 0 50 100 log likelihood sample from cohort speaker (b) Discriminant function histogram after optimization Figure 3.3: Comparison of target and imposter likelihood histograms 3.4 From models to supervectors For the traditional approaches and until this point in this thesis, the trend is to build models from sets of data. Unseen data is classified according to those models. An alternative solution is to find a representation of each utterance using a single vector, usually called supervector. An approach of this kind was first presented in [91] and then reutilized by [92]. Both showed that time-averaged features contain relevant information of the speaker that can be represented as a single vector (that acts like a model). Afterwards, a distance measure between new utterances’ super vectors and the “models” is computed and a decision is taken. Other algorithms evolved in the same direction. SVMs, for instance, created a whole theory around the construction and classification of these supervectors. The sequence kernels are part of this effort. The key point is to map a set of vectors to a single one via a kernel and perform classification [93]. It is not under the scope of this thesis to have a deep description of these approaches. However, their importance is notable. The GMM components used as supervectors marked a watershed in the algorithms. The k means of all the Gaussians in a GMM are concatenated into a single vector, defined as a supervector, M= [µ1|µ2|···]. The new supervectors are the input to different state of the art systems such as SVM [94, 71], and other factor analysis based techniques (explained in next sections). on the likelihood ratio (at the score level), or in terms of similarity between two models as an alternative. There is no agreement on which is the best option [57], but the approach can be successful if built for specific conditions.
30 Chapter 3. Speaker Modeling Algorithm 3.3. Minimum Verification Error Initialization 1. Train the imposter or cohort model (usually employing EM) and the target speaker model (using MAP) ¯ Λ = {¯ WC k,¯µC k,¯ ΣC k}(3.45) Λ = {WC k, µC k,ΣC k}(3.46) 2. Compute the misverification distance and the combined loss function for the FA and FR errors using, dn(χ) = −gn(χ; Λ) + Gn(χ; Λ) (3.47) and `(χ; Λ) = X n `n(d(χ; Λ))1(χ∈Cn).(3.48) Parameter Estimation For a specific class C, a component kin the GMM, dimension d. and a set of training Xtwhere t= 1...T 1. Obtain the gradients of the sigmoid loss function for each of the model parameters: {wC k, µC k,ΣC k}and plugged them into: Λt+1 = Λt−∇`(χ; Λ),(3.49) ∇`(χ; Λ) = γ`(d)(1 −`(d))∂ p(χ; Λ) ∂Λk .(3.50) 2. Optimize the models for every class a For a specific k, dimension dand to avoid bias, let µ=σ˜µ. Then, ∂p(χ; Λ) ∂˜µk =wkN(χ|˜µk, σk) Pk0wk0N(χ|˜µk0, σk0)χ σk −˜µk.(3.51) b let, ˜σ= log(σ),Then σis ∂p(χ; Λ) ∂˜σk =wkN(χ|µk,˜σk) Pk0wk0N(χ|µk0,˜σk0)(χ−µk ˜σk2 −1).(3.52) c For wkand to ensure that PK k=1 wk= 1, let, w=exp( ˜w) PK k=1 exp( ˜w). ∂p(χ; Λ) ∂˜wk =N(χ|µk, σk) Pk0˜wk0N(χ|µk0, σk0).(3.53) Check for convergence 1. Compute the loss function with the new parameters: `new(χ; Λ) = X n `n(d(χ; Λ))1(χ∈Cn).(3.54) 2. Compare `new with `if it is higher than a threshold τiterate from Parameter Estimation. 3.5 Factor Analysis To understand the state of the art of Joint Factor Analysis and i-vectors it is useful to explain first the simplest case of factor analysis (FA). Factor analysis establishes that continuous factors can control data; meaning that this factors can model the covariance structure of the high dimensional
3.5 Factor Analysis 31 data. So we can think that the data can be generated by first, locating a point in a subspace and add noise. If that is the case, we can consider the following: X−µ=LZ +(3.55) where Xis a random variable of dimensions P×1 (observation), µis the data mean vector, L refers to the loading matrix of P×Rdimensions, Zis the latent random variable (known as factor vector) of R×1 dimensions and is a random residue with diagonal covariance matrix ψ. For a further and detailed explanation of FA, please refer to Appendix B. The key point of FA, in the context of SV, is to think about a supervector that belongs to a speaker and it is disturbed by some kind of perturbation. Under this scheme, the speaker model remains the same for all the recordings in different conditions, represented by the speaker factors. describes the contribution of the perturbation to the model. Generally, the first term, LZ, represents the contribution of any factor, but can also be extended to include multiple factors, as shown in joint factor analysis, where the model is decomposed into speaker and channel contributions. In either case, the solution of Equation 3.55 is performed in two ways, using Maximum Likelihood Estimation or Discriminative Modeling. Both techniques will be explained in detail in the following sections. 3.5.1 Joint Factor Analysis (JFA) The purpose of factor analysis is to reduce the channel and speaker variability by decomposing the distributions of the feature vectors into speaker and channel factors [15, 61]. In this framework, let us once again define Cas the number of components of a GMM, Λ, from a speaker S, that form asupervector M= [µ1|µ2|···] (here we have not explicitly shown the superscript representing the speaker for generality). The supervector MS,H representing the GMM for the distribution of data over each channel type Hby a speaker Sis assumed to be composed from a collection of factors as: MS,H =s+c,(3.56) where sis the speaker supervector and cis the channel supervector. Moreover, the distribution of scan be described by s=m+V yS+DzS,(3.57) where mis a CF ×1 supervector that represents the global mean across all speakers, Vis a low rank rectangular matrix which columns are the eigenvoices representing the subspace over which speaker-specific components of MS,H lie, ySis a normally distributed random vector (with 0 mean and unit variance) representing speaker factors specific to speaker S,Dis diagonal matrix and z is a normally distributed random vector representing residual error. In the same way, c=UxS,H,(3.58) where Uis a low rank rectangular matrix, which colummns are the eigenchannels representing the subspace over which the channel specific components of MS,H lie, and xS,H is a normally distributed vector representing channel factors specific to recordings of speaker Sover channel H. The Baum-Welch statistics for learning the loadings V,Uand D, the EM procedures for learning the factors yS,xH,zS, as well as the methods for classification with the resulting model in a manner that factors out the interfering contributions of the channel are well laid out in [61, 15]. The computation of the sufficient statistics such as 0-th, N, the 1-st, F, and 2-nd, Sare an aid to estimate la loading matrices and the factors. For verification, we match a utterance with s; that is, we assume that the supervector has the form s+UxS,H. Several scoring methods have been proposed [95]. To perform the usual likelihood ratio and integrating out the channel contribution is unfeasible. From the alternative approaches,
32 Chapter 3. Speaker Modeling the one that has attracted more attention is the linear scoring that presents a simplification that reduces the evaluation to matrix multiplication of the form, Algorithm 3.4. Joint Factor Analysis Algorithm Train the JFA matrices 1. Assume that Vand Dare zero, train the loading matrix U. M=m+Ux (a) Compute the Baum and Welch Statisitcs 0-th, N, the 1-st, F, and 2-nd, S, order statistics of speaker sand Gaussian component c employing observations T. (b) Using Baum-Welch statistics, estimate the posterior of the hidden factor y. (c) Compute statistics across speakers using the information of posterior distribution of the estimated y. (d) Compute Vand the update of the covariance. (e) Iterate from ato dbetween 15 to 25 times. In every iteration substitute the estimated V. 2. Assume Dis zero, train the loading matrix Ugiven the estimate of V. M=m+V y +Ux (a) After estimating y for each speaker, compute the Baum and Welch statistics, NHand FH, for each H(conversation side or channel) of each speaker S. (b) Calculate the shift for each speaker using R=m+V y and compute its Gaussian posterior. Subtract it from first order statistics F. The new FH(s) depends on the channel and the speaker. (c) Use the new statistics to train U. (d) Iterate 15 to 20 times to obtain Vand yusing the updated NH and FH. 3. Train the residual Dgiven the estimates of Vand U. M=m+V y +Dz (a) Compute the speaker shift of m+V ys, and the channel shift, Uxs,H . Compute the Gaussian posterior of the shifts and subtract them from the first order statistics, FS,H . (b) Estimate the initial factor z and update the statistics across the speakers. (c) Calculate D (d) Iterate form bto c15 to 20 times. Compute factors •Compute the final factors x(channel), y(speaker) and the residual z. Scoring •Compute the final score by a simplified version of the Log-likelihood ratio. Given a new utterance (referred here as tst), the loading matrices and the factors (referred here as tar), compute the product, θs,H = (V ytar +Dztar)Σ−1(Ftst −Ntst m−NtstU xtst).
3.5 Factor Analysis 33 θs,H = (V ytar +Dztar)Σ−1(Ftst −Ntst m−NtstU xtst),(3.59) where tar refers to target speaker and tst to unseen recording and Nrefers to the 0-th order, and Fto the first order statistic according to the test set. This technique has evolved in recent years obtaining good results [61]. It has adjusted to various circumstances; for instance, different training and testing channels (interview, telephone and cross-channel scenarios). 3.5.1.1 Joint Factor Analysis step by step Algorithm 3.4 details the JFA algorithm step by step. It describes the steps to estimate the parameters of the loading matrices M,V,Uand Dfor a speaker S, using Baum-Welch algorithm. Thereafter, we will show how to compute factors x,yand z. For clarity we will drop the subscript that refer to the speaker Sor channel H. For a more detailed explanation, please refer to Appendix B. 3.5.2 Front-end Factor Analysis: i-vector This training approach is based on the idea of a new low dimensional speaker and channel dependent space representation using factor analysis [14]. Whereas in JFA two spaces were defined, the speaker with its eigenvoice matrix V and the channel with its eigenchannel matrix U, in this case the new space contains both in only one space. The new space is known as ”total variability space” and encompasses the variabilities of the speaker and the channel. A total variability matrix, embracing the eigenvectors with the largest eigenvalues, is obtained from the total covariance matrix. Then, the new speaker and channel model is defined as, M=s+Tω, (3.60) where Mrepresents the combined speaker and channel-independent supervector extracted from the UBM. Tis a low rank rectangular matrix and ωis a random vector, composed of the identity vectors (i-vectors), with normal distribution N(0, I). Mis normally distributed with mean mand covariance matrix TT0. The procedure to train the matrix Tis very similar to the training of V in JFA, but in the eigenvoice training the different conversation sides of the target speakers are treated as different speakers. Hence, the new modeling is in fact a projection of a utterance onto the low-dimensional total variability space. ω, the hidden variable can be solved using Baum-Welch statistics. If we consider ωis defined by its posterior distribution, a sequence of Lframes {y1, y2, ..., yL}, and a UBM of Ccomponents in a F-dimensional space; then, Nc= L X t=1 p(c|yt,Ω) ,(3.61) Fc= L X t=1 p(c|yt,Ω) yt.(3.62) where cis the number of the Gaussian component, P(c|yt,Ω) is the posterior probability of cgiven the realization yt. Moreover, the centralized first-order Baum-Welch statistics, using the UBM, is as follows: ˆ Fc= L X t=1 p(c|yt,Ω) (yt−mc),(3.63)
34 Chapter 3. Speaker Modeling where mcis the mean of UBM component c. Then, the i-vector for a phrase is computed as, ω=I+T0Σ−1N(u)T.T0Σ−1ˆ F(u),(3.64) where N(u) is a CF ×CF diagonal matrix composed of NcIdiagonal blocks. The super vector ˆ F(u) of dimension CF ×1 is computed by concatenating all ˆ Fcfor a given phrase u. Σ is the CF ×CF diagonal covariance matrix that models the residual variability. Once the new feature vectors, i-vectors, are computed, then next step is to build an appropriate model that can discriminate among speakers. Two main branches were explored. On one hand, SVM was used on extensive studies as shown in [94, 71]5. On the other hand, the most popular one is PLDA [96]. Algorithm 3.5. I-vector Algorithm Train the Loading matrix 1. Train the loading matrix T. M=m+T w (a) As in JFA, compute the Baum and Welch statisitcs 0-th, N, the 1-st, F, and 2-nd, S, order statistics of speaker sand Gaussian component c. Collect the phrases belonging to specific channels and speaker, treat the different conversation sides of the target speakers as different speakers. (b) Using Baum-Welch statistics, estimate the posterior of the hidden factor w. (c) Compute statistics across speakers using the information of posterior distribution of the estimated w. (d) Compute Tand the update of its covariance. (e) Iterate from bto dbetween 15 to 25 times. In every iteration substitute the estimated T. Channel compensation Two approaches have shown successful results: 1. Compute the linear discriminant analysis (LDA) to reduce the dimensionality of the i-vectors. Afterwards, use Probabilistic Linear Discriminant Analysis (PLDA) as a the target trainer 2. Compute LDA of the i-vectors followed by with-In-class covariance normalization (WCCN). Scoring •Compute cosine distance scoring (CDS) on the updated i-vectors (after channel compensation). 3.5.2.1 Probabilistic Linear Discriminant Analysis (PLDA) model PLDA is a generative modeling approach to handle i-vectors. It was first used in the image recognition scenario [96] and then extended to SV [74]. Note that the formulation is similar to the one employed by JFA. The supervector rS,H representing the i-vector of a single recording over a channel Hby a speaker Sis assumed to be composed of a collection of factors, rS,H =m+V yS+UxS,H +ε, (3.65) 5It is not under the scope of this thesis to detail the SVM approaches. However, the methodology is by far interesting. Details can be found in [75]
4.4 A glance of Robust solutions 41 4.3.3 Noise in the score domain The noise not only exhibits its impact in a single operating point (for instance the EER), but also in the score distributions and along the DET curve (see Figure 4.5). The noise degrades the signal and manifests in every stage. The poor performance is clearly observed in the evaluation (last classification of the system). The targets are confused with imposter, but the opposite is likely as well. For a binary classification an intuitive explanation is that the opposing distributions overlap increasing the FA and FR and making the classification difficult. Also, we note that for the noise-score distributions the mean and variances differ. As a result, the DET curve is higher for noisy speech, including the EER. The score normalization such as Z-norm [72], T-norm [73] and ZT-norm [12] can alleviate the problem in some extent. Their main purpose of them is to eliminate the offsets caused by extraneous factors. However, to deal with low SNRs the problem is addressed by model compensation or robust feature extraction techniques. score Target speaker Impostor FR FA speaker EER 10 1 FA FR EER Operating Point Figure 4.5: Scores distributions for two scenarios: a) blue line represents the clean condition, b) red line shows the noisy condition. 4.4 A glance of Robust solutions The solution to the noise problem can be solved using well known Automatic Speech Recognition tools (as shown in this section). However, it is also useful to tackle the problem by employing embedded algorithms in the SV system. A few solutions have been given to address the noisy conditions in Speaker Verification that will be described in the next paragraphs. 4.4.1 Robustness in Automatic Speech Recognition Some researchers proposed to directly work on solutions similar to the ones employed for ASR. In the following paragraphs we want give a general view of what we consider important algorithms that are the fundament of some of the algorithms in SV. There are several solutions that have been extensively used in ASR divided into three main groups: speech enhancement, feature vector adaptation or normalization and acoustic model adaptation. In this thesis we will not discussed them in depth, but we consider that the basis of the algorithms have evolved or can be used indistinctively for either ASR or SV. The objective of speech enhancement is to clean the speech signal in the feature extraction process. Examples of the algorithms are CMS (cepstral mean subtraction) [41] and SS (spectral subtraction) [102]. The feature vector adaptation or normalization occurs in the feature space. The algorithm models the distortion between clean an noisy features by a parametric function that can later be used to compensate the mismatch in the test features. A basic example of this normalization is
42 Chapter 4. Robustness in Speaker Verification RASTA (relative spectral amplitude) [43], which is a high-pass filtering that reduces the effect of a varying channel. Another example is CDCN (codeword dependent cepstral normalization) [108]. In this case, the transformation is based on the normalization of the acoustic space, defining a universal codebook that can be transformed to the vectors in a phrase. The model adaptation compensates for the mismatch by adjusting the parameters of the acoustic models. This scheme sometimes gives better results because it can model the uncertainty of the noise. One of the main disadvantages is that a certain amount of data is needed to obtain reliable statistics. Examples of this schemes are MAP (maximum a posteriori), MLLR (maximum likelihood linear regression), PMC (parallel model compensation), and VTS (vector Taylor series). Moreover, we emphasise the Ensemble Speaker and Speaking Environment (ESSE) modeling method. Last examples are SPLICE (stereo based linear compensation environment) [109, 110] and MEMLIN (multi-environment model-based linear normalization) [111]. Although this methodologies have been used for ASR, they can be extended to SV. 4.4.2 Factor analysis and Robustness Recently, the research community presented one of the first efforts to address robustness in the SV. The first step was the construction of a noisy database to show the limitations and potentialities of the current state of the art systems [103]. The new baseline database is an extension of the clean NIST [112], switchboard [113], fisher [114] databases adding noise as well as reverberation. It also includes characteristics, already treated in previous evaluations, such as channel mismatch between the train and test, different vocal efforts. The new database was built artificially, i.e., from a set of clean data the noise and reverberation was added. The noise belongs to cocktail noisy recordings at different SNRs ( 8, 15, 20 dB). The reverberation was recorded at different room types and reverb delays (at 0.3, 0.5, 0.7 reverberation times). The system uses the current state of the art scheme; it includes the i-vector front-end, followed by the LDA for dimensionality reduction and PLDA modeling [115]. The system shows promising EERs – of less than 3%– given by clean speech even when there is a mismatch between the test and training conditions. Moreover, the current approach was also tested for different feature extractors: MFFCs, prosodic polynomial coefficients and MLLR (maximum likelihood linear regression). For this cases the degradation of the EER is from 3 to 14 times compared to the baseline. The results showed that the method does not depend on the feature extractor. It is important to note that at this stage the training of the UBM and the i-vector is computationally expensive, then a good solution is just to add noise to the PLDA and LDA part. The first attempt was to add noise to every part of the system, but by including extra noisy features to model the UBM didn’t give any significant improvements. When adding noise just to the PLDA modeling 40% improvement was achieved. These studies showed the potentiality of the current state of the art. 4.4.2.1 Vector Taylor Series VTS Another way to address the modeling is to define a mapping function that characterizes the mismatch between the training and testing models [116, 117]. This study suggests that the effects of the noise effects have been underestimated due to mathematical simplifications. It characterizes the statistics of the additive noise and the linear filtering in a transmission channel. This method can be applied to the feature vectors (including delta and doble delta) or even to the representation of those vectors. A special case of this type of modeling is the jacobian adaptation (JA). An extension of the state of the art that modifies the front-end using VTS is presented in [1]. This approach tries to clean the i-vectors, but it is not used as the usual VTS, where we search for a mapping function that characterizes the mismatch between the training and testing conditions. Under this new scheme, the utterance of the GMM, similar to what it is done in JFA, is decomposed into two distributions: the clean and the noisy. The clean distribution belongs to
4.4 A glance of Robust solutions 43 the i-vectors. Note that by including the VTS, the system can model non-linear effects in the GMM caused by the additive noise. The results of this approach show its benefits in low SNR (8 and 15 dB) and it is also robust for unseen data. Lately, in [118], the authors described a modified version of the VTS proposed by [1]. For this special case, the unscented transform, UT, is used to compute the first order VTS. The new non-linear function is applied to a set of sampled points. The resulting transformed data is used to calculate the model parameters. The approach has shown improvements for low SNRs. Table 4.1 compares the improvements obtained for 8 dB SNR. The mean and variance normalization approach is used as the baseline for the VTS i-vector approach. The VTS is then used as the baseline for the UT transform. For the clean condition experiments, the system enrollment was performed using clean data, the test is performed under noise condition. For the multi-condition scenario, the enrollment used the noisy data at the same SNR as the test phrases. VTS - ivector VTS - Unscented Transform clean multicondition clean multicondition Approach MVN VTS MVN VTS VTS UT VTS UT EER 15.5 5.2 5.9 3.2 6.3 5.9 4.2 3.6 Rel. Improvement 66 45 7 14 Table 4.1: Comparison of VTS EER and relative improvement (in %), 8 dB, [1]. 4.4.2.2 Acoustic Factor Analysis Acoustic factor Analysis [2] considers the covariance information in the feature space to obtain robust models. The state of the art factor analysis schemes employ super-vector space as the basis of the decomposition of the signal. Moreover FA uses a diagonal covariances of the GMM models believing that the MFCCs are uncorrelated. However, [119] shows that full covariances can also provide discriminative information of the speaker. In contrast, this new approach shows that the MFCCs are not fully uncorrelated and that a better estimation of the parameters can be performed under the feature space instead of the super-vector space. The new solution strategy use a mixture probabilistic principal component analysis, MPPCA [120], embedded into the FA approach. It follows the same idea as in FA, for a χ=X1, X2,··· , XT, where Xiis the ith vector in the sequence, X=LZ +µ+(4.2) where Xis a random variable of dimensions P×1 (observation), µis the mean vector, Lrefers to the loading matrix of P×Rdimensions, Zis the acoustic factor of R×1 dimensions and is the random noise with diagonal covariance matrix ψ. The latent vector Zfollows a Gaussian distribution Z∼ N(0, I ), ∼ N(0, ψ I ) and X∼ N(µ, ψ I +LLT). To capture the variations of multiple speakers and its corresponding phonemes and conditions, the MPPCA is defined as, P(X) = M X i wiN(µi, ψi I +LiLT i),(4.3) where wi, µi,ψand Liare the mixture weight, mean, noise covariance and the loading matrix of the ith PPCA model and I is the identity matrix. With this approach we can get a mixture dependent transformation for the i-vector architecture. Now, we can compute the UBM using the EM algorithm and then use it to compute the MPPCA.
44 Chapter 4. Robustness in Speaker Verification Moreover, we can decide the number of axis we require (still under investigation how many are optimal), discarding the lower dimensions. Afterwards, we can compute the loading matrices and the factors with this new scheme. Table 4.2 shows an example of the improvements obtained with respect to two baseline systems, one with full covariance and the other with diagonal covariance. qrefers to the number of dimensions retained from 60-dimensional vectors, using an i-vector/PLDA scheme. AFA shows competitive results with respect to the current baseline. Baseline AFA full cov diag cov q= 36 q= 42 q= 48 EER 2.03 2.55 2.06 1.79 1.95 Rel. Improv - - -1.1 12 4.1 Table 4.2: Comparison of AFA EER and relative improvement (in %), [2]. 4.4.3 Front-end modification The modification of the front-end signifies better models and accuracy. Although, the research in ASR can be extended to SV, there are few studies that center on robustness from a SV point of view. In this section we give an example of such a research that addresses the problem by modifying the front-end to produce new features. 4.4.3.1 Regularization of all-poles The study presented in [3] describes a modification of the traditional MFCC-based front end. The approach replaces the DFT estimation with a weighted Linear Prediction, LP, with regularization. This research shows a deep analysis of the traditional LP, stressing a comparison with the weighted LP. The experiments were performed under very noisy conditions (for example, under 10 dB factory noise and babble noise). For every case, this new modified front-end improves the baseline. The main idea of the regularized LP [121] is that it gives a penalty for rapid changes in an all-pole spectral envelopes. SV benefits from the idea that the spectral models are smoothen so that mismatch between training and test is reduced. Table 4.3 depicts an example of the results for different LP-based approaches under babble noise at 10 dB. LP approach outperforms the FFT-baseline system. Moreover, the stabilized weighted (SWLP) and the minimum variance distortionless response, MVDR,1also improve the current baseline showing the potential of the LP approaches. Baseline FFT LP SWLP MDVR EER 21.28 20.36 19.69 19.68 Rel. Improv - 4.3 7.4 7.5 Table 4.3: Comparison of LP approaches EER and relative improvement (in %),10 dB, [3]. 1Further details about this approach are described in [3]
4.5 Summary 45 4.5 Summary This chapter describes the basic background and latest approaches behind robustness from a SV point of view. We observe how noise degrades the signal process in every state, producing undesired effects in the performance. However, the research on this area is getting attention and there is an increasing number of research that try to alleviate the effect of the noise. The noise can be treated with the usual and known solutions given by ASR research field (speech enhancement, feature vector normalization and acoustic model adaptation) But the SV community has also proposed modifications to the current system algorithms that can handle the noise. We highlight the approaches based on factor analysis, which have shown to improve the performance metrics. As shown in this section, the algorithms for robust ASR motivated other studies in SV in the same directions. Our research is a fusion of both ideas, which can be seen as a vector transformation and an acoustic model adaptation. Next chapters present these ideas to improve the current known algorithms.
46 Chapter 4. Robustness in Speaker Verification
Chapter 5 Experimental Framework You’re never given a dream without also being given the power to make it true. - Richard Bach The Adventures of a Reluctant Messiah There are plenty of reasons to have NIST databases as the starting point of the SV research. Nowadays, the performance of the state-of-the-art systems compare the results in terms of curves and efficiency measures. Besides, the comparisons also takes place at several stages of the process, for instance: feature extraction, modeling and evaluation. The NIST databases provide a reliable and homogenous environment for such comparisons. Moreover, the databases have been of help to the research community when building up complex systems where the merging parts belong to different research groups. The current chapter is entirely dedicated to the NIST databases.1We start by describing the current infrastructure used to develop our SV system. Secondly, we briefly describe the characteristics of the current NIST databases. Next, this chapter shows the traditional experimental setup that we used for the baseline experiments and that will be used for the proposed approaches along this thesis. Lastly, we give the initial baseline results for this setup. 5.1 Current infrastructure The following functional cluster was used for such challenging project with a limited infrastructure (see Figure 5.1) composed of: Autonomous Beowulf cluster with 310 CPUs, with the following machines: Pentium 4 HT 3.2 GHz ; Xeon W3530 2.80GHz ; Xeon E5645 2.40GHz AMD Opteron 6272 2GHz ; AMD Opteron 6276 2.6GHz ; Core2 Duo E8400 3.00GHz ; 1Gbps LAN 12TB storage. Software: SGE, Matlab 7.11.0.584, Python 2.6.6, Perl v5.10.1, x86 64 GNU-Linux. The code was implemented in a parallelized mode in which all the computers can perform a reliable computation. 1Most of the state of the art baseline results and comparisons with other systems are performed using this set of databases [40, 99].
48 Chapter 5. Experimental Framework Figure 5.1: Infraestructure used by the Speaker Verification System 5.2 NIST database characteristics A great effort has been done by the National Standard Institute and Technology to promote the research on this field [112].2Every two years the community organizes an evaluation among participants around the world. Each evaluation addresses the lately challenges. From the beginning, the plan was to record a database for a text-independent purpose. The first tasks included only a few trials and the challenge was to generate good models for each target speaker. Over the years, the databases have increased in target speaker and trials number. Moreover, for the mismatch conditions (different conditions for the training and test) is common to include several kind of microphones and telephones. The last evaluations included different types of vocal efforts (low, high and medium) as well. The last one also included noise at different signal to noise ratio. Before 2012 the rules were strict on not using the speech signal from other target users to model a certain speaker. In the recent evaluation, the competition allowed the use of this information. To have an idea of the databases, a summary of the conditions presented is showed in Table 5.1. train condition test condition channel Interview interview same microphone Interview interview different microphone telephone telephone matched condition telephone telephone mismatched condition Interview telephone mismatch channel condition telephone interview mismatch channel condition Table 5.1: Table presenting the common conditions Lastly, every evaluation presents an increment in the number of trials as part of the concern of how real imposters try to access the verification systems. The last evaluation included a set of one million trials. For the noisy conditions we employed Aurora II database [122]. This database has been widely 2Almost every new algorithm has to be tested under the databases the organism provide to document improvements. Since 1996 the competition has tried to perfect the systems going from obtaining better models of the target speaker to include huge amount of trials (of the order of one million) and lately encompassed noise. The tasks includes harder goals to achieve.
5.3 NIST database baseline 49 employed for ASR, it includes the TI-digits in a clean scenario and TI-digits in noisy conditions. As part of this work, we just selected the babble noise to be added to the clean recordings using the noise adding tool, FaNT in [123]. The noisy recording is then employed to test the baseline and our proposed techniques. 5.3 NIST database baseline In this section we introduce the experimental setup. We show further details about the actual datasets used and the specific feature extraction. We also present the compilation of the most relevant results. We focus our research on the core-core evaluation NIST 2008, which is a defined task. We employed the NIST Speaker Evaluation 2004, 2005, 2006, 2008 and 2010 databases [40] to complete this study. The training data selection consists of building a UBM using the recordings in databases 2004, 2005,2006 and 2010. Databases 2004 to 2006 provided sufficient information to build models for telephone conditions. Database 2010 granted different microphone and mismatch conditions to our experiments. Database 2008 tested the baseline framework and the new approaches. From this database, the target speakers belonging to the core-core evaluation were used to train specific user models. Lastly, the set of trails evaluate the reliability of those models and the system. The training and testing data sets are composed of either one two-channel telephone conversation of approximately five minutes total duration, with the (target or proposed trial) channel designated previously or a microphone conversation segment of three to fifteen minutes involving both the interviewee (target speaker) and an interviewer. The data also presented various types of microphones (seven) and both conversational and interview sessions. 3The common conditions used are described in Section 5.2 and Table 5.1. For the set of experiments, we adopted the NIST evaluation restrictions (for instance, neither choosing other target model as imposter model, nor using other target data to estimate the current target model). 4Following NIST 2008 Evaluation rules, the probability of being a target, Ptarget, is 0.01 and the probability of being a impostor, Pimpostor, is 0.99. 5.3.1 Front-end For the feature extraction, a short-time 256-pt Fourier analysis is performed on a 25ms analysis window and 10ms frame rate. The feature vector (token) consists of 49 Mel Frequency Cepstral Coefficients (MFCCs), including delta and double delta coefficients, as shown in Table 5.2. The next step is to apply a feature warping normalization to undo the distortion caused by the channel. This warping is accomplished by first assembling an empirical CDF (cumulative distribution function) from the ranked features within 1.5 seconds after and before the current frame (3 seconds total), and then perform the CDF-inverse at the current frame. We included a frame removal criterion that encloses the concept of eliminating the low energy frames that do not provide information about the identity of the person. Low and 70% of the medium log-energy frames were simply discarded 5.3.2 UBM generation A gender-dependent and target-independent 512-mixture GMM UBM model was trained from the core-core of NIST-SRE 2004, 2005, 2008 and 2010 databases.5Note that databases 2008 3More information about the types of microphones can be found in [40] and in [112]. 4Although in recent evaluations the rules have become more flexible, employing the target speakers as part of the imposter models, we built imposter models independently from the target data. 5Usually, the NIST databases must be cleaned from silent recordings or recordings with any kind of problems. This process is always performed in advance.
50 Chapter 5. Experimental Framework Cep. Coeff 16 logE 1 ∆ Cep 16 ∆∆ Cep 16 Table 5.2: Table presenting the MFCC computation and 2010 training include several types of microphones that were used to test the system against mismatched conditions. EM algorithm was used to obtain the maximum likelihood estimates of the GMM parameters. For every iteration of EM, the system randomly polls 80% of the training tokens, corresponding approximately to 15 hours of speech. The UBM is first initialized using the K-means algorithm to obtain a set of 512 centroids. By using the k-means the performance of the EM was simplified, however it is always important to check that the local bounds are not very restrictive, so that EM can provide a satisfactory estimation. The EM is then repeated after the model had converged (10 iterations). 5.3.3 Speaker Modeling For each target speaker the current UBM is adapted using MAP algorithm (using 2 to 5 iterations). For the special case of the NIST dabase 2008, a recording file for each speaker were used to perform adaptation. In next sections we present the results for our approaches, but MAP models are the starting point for all of them. 5.3.4 Baseline results Table 5.3 presents the baseline results performed with the traditional techniques for database 2008. We just give and example for the simplest case using the GMM approach. 512 Gaussian components were used for this purpose. We observe that the results for mismatch condition are worse than the ones for matched conditions. Moreover, the best results are obtained by the telephone set, probably because of the amount of telephone data used in the training. Female Male Average 1Interview interview same mic 14.0 14.9 14.4 2Interview interview different mic 18.3 18.2 18.2 3Interview tel 17.9 18.8 18.3 4tel tel 12.5 13.0 12.7 Average 15.6 16.2 15.9 Table 5.3: Table presenting the final results (EER) on the Test set for NIST 2008 5.4 NIST database baseline - JFA The starting point of JFA is definitely the traditional approach. The front-end remains the same, computing 16-dimensional MFCCs including deltas and double deltas (see Table 5.2). Next, we compute a gender-dependent and target-independent model, UBM, from a pool of raw speech (NIST Speaker Evaluation 2004, 2005, 2006 and microphone recording from 2010 core database). For the JFA baseline, the speaker and channel factors were learned from a pool of recordings adapted to individual speakers. Apart from NIST databases, already described, other databases used for the training stage are Switchboard-1 (SW1) and Switchboard-2 (SW2) [113]. These two
6.5 Minimum Verification Error 57 where Lχis the number of feature vectors in χ. Note that the original AUC formulation of Equation 6.4 only considers misclassifications – both FA and FR, in keeping with the conventional formulation of MVE. The “soft” version of the AUC given by Equation 6.9 also naturally conforms to this formulation. Using the representation X χ∈H Lχ=|H| (6.10) X χ∈W Lχ=|W|,(6.11) the GPD update rule for any parameter φof the distributions is now given by φt+1 =φt− ∇φΥ(X,Λ), where ∇φΥ (X,Λ) = −1 |H||W| X χ∈H X X∈χX ˆχ∈W X ˆ X∈ˆχ γR(1 −R)∇φl(X, ˆ X, Λ) (6.12) and ∇φl(X, ˆ X, Λ) is a local gradient with respect to φat χ, ˆχand has the form given by Equation 6.13, ∇φl(X, ˆ X, Λ) = −∂θ(X) ∂φ +∂θ(ˆ X) ∂φ .(6.13) ∂θ(X) ∂φ represents the derivative of the log-likelihood-difference given by the Gaussian mixture models for the target speaker and the universal background model for vector Xwith respect to φ. The update rules for the individual parameters wS k,µS kand ΣS kare obtained by plugging in ∂θ(X) ∂wS k , ∂θ(X) ∂µS k and ∂θ(X) ∂ΣS k respectively into Equation 6.13. These Equations are relatively straightforward to derive, (please refer to Appendix C for further details). For the purpose of this chapter we will just include the final update Equations. Once again, consider µas an example, then, ˜µ=µ/σ, Let the misverification measure be ∂θ(X) ∂φk =−∂g(X; Λ) ∂φk +∂G(X; Λ) ∂φksgn(c) (6.14) where sgn(c) is +1 if it is a target and −1 if it is an imposter. Moreover, ∂g(χ; Λ) ∂Λk =1 p(χ; Λ) ∂ p(χ; Λ) ∂Λk and ∂G(χ; Λ) ∂Λk =1 p(χ;ˆ Λ) ∂ p(χ;ˆ Λ) ∂ˆ Λk .(6.15) Then for µ, a specific k, dimension dand to avoid bias, let µ=σ˜µ. Hence, ∇φΥ (χ, ˜µk) = −1 |H||W| X χ∈H X X∈χX ˆχ∈W X ˆ X∈ˆχ γR(1 −R)wkN(χ|˜µk, σk) Pk0wk0N(χ|˜µk0, σk0)χ σk −˜µk.(6.16) For σ, let ˜σ= log(σ).Then, the final Equation to compute σkis ∇φΥ (χ, ˜σk) = −1 |H||W| X χ∈H X X∈χX ˆχ∈W X ˆ X∈ˆχ γR(1 −R)wkN(χ|µk,˜σk) Pk0wk0N(χ|µk0,˜σk0)(χ−µk ˜σk2 −1).(6.17) Finally, for the set of weights wkand to ensure that PK k=1 wk= 1,let, w=exp( ˜w) PK k=1 exp( ˜w) ∇φΥ (χ, ˜wk) = −1 |H||W| X χ∈H X X∈χX ˆχ∈W X ˆ X∈ˆχ γR(1 −R)∇φl(X, ˆ X, Λ) N(χ|µk, σk) Pk0˜wk0N(χ|µk0, σk0).(6.18)
58 Chapter 6. Optimization of the area under the ROC curve Algorithm 6.1. Optimization of the ROC curve, MVE Initialization 1. Train the basis imposter or cohort model (usually employing EM) and the target speaker model (using MAP) ¯ Λ = {¯ WC k,¯µC k,¯ ΣC k}(6.19) Λ = {WC k, µC k,ΣC k}(6.20) 2. Compute the misverification distance and the combined loss function for the FA and FR errors using, dn(χ) = −gn(χ; Λ) + Gn(χ; Λ) sgn(c) (6.21) and Υ(Λ) = 1.0−Pχ∈H PX∈χPˆχ∈W Pˆ X∈ˆχR(θ(X), θ(ˆ X)) Pχ∈H LχPˆχ∈W Lˆχ (6.22) Parameter Estimation For a specific class C, a component kin the GMM, dimension d. and a set of training Xtwhere t= 1...T 1. Obtain the gradients of the sigmoid loss function for each of the model parameters: {wC k, µC k,ΣC k}and plugged them into: Λt+1 = Λt−∇Υ(χ; Λ),(6.23) ∇φΥ (X,Λ) = −1 |H||W| X χ∈H X X∈χX ˆχ∈W X ˆ X∈ˆχ γR(1 −R)∇φl(X, ˆ X, Λ) (6.24) Recall, ∇φl(X, ˆ X, Λ) = −∂g(X;Λ) ∂φk+∂G(X;Λ) ∂φksgn(c) 2. Optimize the models for every class a For a specific k, dimension dand to avoid bias, let µ=σ˜µ. Then, ∂p(χ; Λ) ∂˜µk =wkN(χ|˜µk, σk) Pk0wk0N(χ|˜µk0, σk0)χ σk −˜µk.(6.25) b let, ˜σ= log(σ),Then σis ∂p(χ; Λ) ∂˜σk =wkN(χ|µk,˜σk) Pk0wk0N(χ|µk0,˜σk0)(χ−µk ˜σk2 −1).(6.26) c For wkand to ensure that PK k=1 wk= 1, let, w=exp( ˜w) PK k=1 exp( ˜w). ∂p(χ; Λ) ∂˜wk =N(χ|µk, σk) Pk0˜wk0N(χ|µk0, σk0).(6.27) Check for convergence 1. Compute the loss function with the new parameters: Υ(Λ)new = 1.0−Pχ∈H PX∈χPˆχ∈W Pˆ X∈ˆχR(θ(X), θ(ˆ X)) Pχ∈H LχPˆχ∈W Lˆχ (6.28) 2. Compare `new with `if it is higher than a threshold τiterate from Parameter Estimation.
6.6 JFA 59 6.6 JFA In the case of JFA the set of parameters to be learned are of two kinds. The global parameters include V, the speaker loadings, U, the channel loadings, and D, the diagonal error scaling matrix. The specific parameters include yS, which is specific to a speaker Sand xS,H , which is specific to the speaker-channel combination S, H. Hence, two distinct learning problems must be addressed: learning the global loadings from a large collection of speaker recordings over a variety of channels, and learning specific factors for individual speakers. The global parameters must learn the overall characteristics of the speaker and channel subspaces. Specific parameters must be learned to customize a model to a specific speaker given training data for the speaker. The AUC objective function must be appropriately customized in each case. In all cases, the discriminant function θ() in Equation 3.1 is specified as in Equation 6.29. θ(χ) = log p(χ;V, U, D, yS(χ), xH(χ),S(χ))−log p(χ;λ¯ S, U, xH(χ),S(χ)) (6.29) Here S(χ) represents the speaker Srepresented in the recording χ.H(χ) represents the recording channel in χ. The Equation above explicitly indicates that the log-likelihood for the model of speaker Sis computed using Gaussian parameters from the supervector M=m+V yS(χ)+UxS(χ),H(χ)+Dz (6.30) whereas the parameters of the “imposter” model for any recording are obtained from M0= m+UxS(χ),H(χ)+Dz, which only considers the universal mean madjusted by the channel factors which customize them to the recording χ. The global mean mis derived from the universal background model λ¯ S. 6.6.1 Global parameters: Loadings estimation Let Xrepresent a large collection of recordings χobtained from a large number of speakers. Let Sbe the set of all speakers represented in X. Let XSrepresent the subset of Xrepresenting recordings from speaker Sand X¯ Sbe recordings from remaining speakers, i.e. X=XS∪X¯ S.X can be partitioned in this manner in as many ways as there are speakers in S. To learn global parameters, we define the AUC objective function as given in Equation 6.31, Υ(Λ) = 1.0−1 |XS||X¯ S|X S∈S X χ∈XSX χ∈X¯ S R(θS(χ), θS(ˆχ)).(6.31) As before, the GPD update rule for any global parameter φis given by φt+1 =φt−∇φΥ(X,Λ), where ∇φΥ(X,Λ) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R)∇φl(χ, ˆχ, Λ) |XS||X¯ S|(6.32) and ∇φl(χ, ˆχ, Λ) is a local gradient with respect to φat χ, ˆχ. ∇φl(χ, ˆχ, Λ) = −∂θ(χ) ∂φ +∂θ(ˆχ) ∂φ (6.33) To derive the update rules for individual parameters, it is sufficient to obtain the derivatives ∂θ(χ) ∂φ with respect to the corresponding parameters. The case was studied in [130] and it is extended in the Equations given below to the special case.2 2We use the FA formulation presented in [130]; analogous to Baum and Welch algorithm, in here we have previously derived the minimum classification approach.
60 Chapter 6. Optimization of the area under the ROC curve The formulation takes two stages (as in the usual JFA) to decouple estimation of V, U, Ψ. First, we will consider Vestimation, maintaining Ufixed, considering that Q=V V >+ Ψ, where Ψ = DD>and is diagonal. Then for µ, a specific k, dimension dand to avoid bias3, let µk=qk˜µk,Hence, ∇φΥ(χ, ˜µk) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R) |XS||X¯ S| wkN(χ|˜µk, qk) Pk0wk0N(χ|˜µk0, qk0)χ qk −˜µk.(6.34) To estimate Q, first we present the expression for Vand then Ψ; refer to Appendix C.2 that details about identities. We will also make the matrix dimensions explicit, considering that Ψ is diagonal. Then, qk, let ˜q= log(q). Then, for V, ∇φΥ(χ, ˜vij ) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R) |XS||X¯ S| wkN(χ|µk, qk) Pk0wk0N(χ|µk0, qk0)n[Q−1(x−u)]i[(x−u)>Q−1V]j−[Q−1V]ij o. (6.35) Then, for ψ, ∇φΥ(χ, ˜ ψi,j ) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R) |XS||X¯ S| wkN(χ|µk, qk) Pk0wk0N(χ|µk0, qk0)1 2Q−1 ii −[Q−1x−µ]2 i.(6.36) Now, for U, we fix the current Qand estimate accordingly. ∇φΥ(χ, ˜ Uk) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R) |XS||X¯ S| wkN(χ|µk, qk) Pk0wk0N(χ|µk0, qk0)(χ−µk ˜ Uk2 −1).(6.37) Finally, for the set of weights wkand to ensure that PK k=1 wk= 1,let, w=exp( ˜w) PK k=1 exp( ˜w) ∇φΥ(χ, ˜wk) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R) |XS||X¯ S| N(χ|µk, qk) Pk0˜wk0N(χ|µk0, qk0).(6.38) We would like to highlight that this process is replacing the current Baum and Welch estimation of the model parameters. 6.6.2 Estimating Specific Parameters To learn the specific parameters for a particular speaker S, the AUC objective is defined simply as Υ(ΛS) = 1.0−Pχ∈XSPχ∈X¯ SR(θS(χ), θS(ˆχ)) |XS||X¯ S|(6.39) Note that unlike Equation 6.31 which includes an outer summation over all speakers, Equation 6.39 only considers a single speaker S. As before, the GPD update rule for any specific parameter φis given by φt+1 =φt−∇φL(X,Λ), where ∇φΥ(X,Λ) = −Pχ∈XSPχ∈X¯ SγR(1 −R)∇φl(χ, ˆχ, Λ) |XS||X¯ S|, where ∇φl(χ, ˆχ, Λ) is defined as in Equation 6.33. 3our usual change of variables hold for this formulation
6.6 JFA 61 Algorithm 6.2. Optimization I of the ROC curve, JFA Initialization 1. Train the imposter or cohort model (EM) and the target speaker model (MAP) ¯ Λ = {¯ WC k,¯µC k,¯ ΣC k}Λ = {WC k, µC k,ΣC k}(6.40) 2. Compute the misverification distance and. dn(χ) = −gn(χ; Λ) + Gn(χ; Λ) (6.41) Parameter Estimation, Global Parameters 1. Compute the combined loss function for the FA and FR errors using, Υ(Λ) = 1.0−1 |XS||X¯ S|X S∈S X χ∈XSX χ∈X¯ S R(θS(χ), θS(ˆχ)) (6.42) 2. Obtain the gradients of the sigmoid loss function for each of the model parameters: {wC k, µC k,ΣC k}and plugged them into: Λt+1 = Λt−∇Υ(χ; Λ), ∇φΥ(X,Λ) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R)∇φl(χ, ˆχ, Λ) |XS||X¯ S|(6.43) Recall, ∇φl(X, ˆ X, Λ) = −∂g(X;Λ) ∂φk+∂G(X;Λ) ∂φksgn(c) 3. Optimize the models for every class, considering Q=V V >+ Ψ a For a specific class C, a component kin the GMM, dimension d. and a set of training Xtwhere t= 1...T ∂p(χ; Λ) ∂˜vi,j =wkN(χ|µk, qk) Pk0wk0N(χ|µk0, qk0) n[Q−1(x−u)]i[(x−u)>Q−1V]j−[Q−1V]ij o.(6.44) b For Ψ, ∂p(χ; Λ) ˜ ψi,j =wkN(χ|µk, qk) Pk0wk0N(χ|µk0, qk0)1 2Q−1 ii −[Q−1x−µ]2 i. (6.45) c For U, ∂p(χ; Λ) ˜ Uk =wkN(χ|µk, qk) Pk0wk0N(χ|µk0, qk0)(χ−µk ˜ Uk2 −1).(6.46) d For wkand to ensure that PK k=1 wk= 1, let, w=exp( ˜w) PK k=1 exp( ˜w). ∂p(χ; Λ) ∂˜wk =N(χ|µk, σk) Pk0˜wk0N(χ|µk0, σk0).(6.47) Check for convergence 1. Compute the loss function with the new parameters: Υnew(Λ) = 1.0−1 |XS||X¯ S|X S∈S X χ∈XSX χ∈X¯ S R(θS(χ), θS(ˆχ)) (6.48) 2. Compare Υnew with Υ if it is higher than a threshold τiterate from Parameter Estimation.
62 Chapter 6. Optimization of the area under the ROC curve Algorithm 6.2. Optimization II of the ROC curve, JFA Parameter Estimation, Especific Parameters Two approaches are possible: use the usual discriminative training. Alternatively, use MAP point estimate. For the discriminative approach we have the following. 1. Compute the combined loss function for the FA and FR errors using, Υ(ΛS) = 1.0−Pχ∈XSPχ∈X¯ SR(θS(χ), θS(ˆχ)) |XS||X¯ S|(6.49) 2. Obtain the gradients of the sigmoid loss function for each of the model parameters: x, y and plugged them in Λt+1 = Λt−∇Υ(χ; Λ). ∇φΥ(X,Λ) = −Pχ∈XSPχ∈X¯ SγR(1 −R)∇φl(χ, ˆχ, Λ) |XS||X¯ S|, Recall, once again, ∇φl(χ, ˆχ, Λ) = −∂g(χ;Λ) ∂φk+∂G(χ;Λ) ∂φksgn(c) 3. Optimize for the point estimate of the factors xand yin the usual discriminative form. Check for convergence 1. Compute the loss function with the new parameters: Υnew(ΛS) = 1.0−Pχ∈XSPχ∈X¯ SR(θS(χ), θS(ˆχ)) |XS||X¯ S|(6.50) 2. Compare Υnew with Υ if it is higher than a threshold τiterate from Parameter Estimation. To derive the update rules for individual parameters yS,xS,H and zS, it is sufficient to obtain the derivatives ∂θ(χ) ∂yS,∂θ(χ) ∂xS,H and ∂θ(χ) ∂zS, and employ these in the GPD update rules. In practice the speaker and channel factors yS,xH,S and zScan also be estimated using conventional EM estimate rules. Note that the estimation of global parameters also requires estimation of specific parameters, since the update rules for the former require the latter. Thus, estimation of global parameters involves estimation of both global and specific parameters for all the speakers in the training set. To learn a model for a new speaker for whom a small amount of training data have been made available, only the specific parameters need be learned employing the already-known global parameters. Although we have explained the above in terms of JFA, the same formulation can also be used to learn i-vector representations [14], which are essentially the same as the above without explicit separation of channel factors. AUC-optimized PLDA based representations can be derived similarly to the rules given above, with the modification that the GPD rules will now employ partial derivatives with respect to PLDA parameters. 6.7 Noise case Noise increases the inherent variability of the signal, potentially shifting it from any operating point a classifier may be optimized for, possibly in an unpredictable manner. We demonstrate that an operating-point-agnostic training paradigm that optimizes the entire ROC curve can result in classifiers that are significantly more robust to noise than conventional classifiers even when the effect of noise is not explicitly considered. The training formalism that optimizes the entire ROC curve by maximizing the AUC, follows the same criteria explained above.
6.8 Comparing AUC approach with Calibration approach 63 First, we will give some intuition on the role of the classifier. We can view it in the sense of the distributions and how does they are affected by the noise. Moreover, we can also observe the behavior of the ROC. As explained before, the noise can cause unpredictable effects as shown in Figure 6.3. First, we can observe random shifts on the distributions, when noise is added. The changes are more erratic when the SNR increases. The reliability of the classifier is also diminished: the EER increases (the same occurs to the DET curve) and the AUC of the ROC curve decreases. 0 0.2 0.4 0.6 0.8 1 0 0.2 0.4 0.6 0.8 1 ROC CURVE FA TP MVE−AUC MVE −3 −2 −1 0 1 2 0 0.01 0.02 0.03 0.04 Score Impostor Speaker Target Speaker Figure 6.3: The effect of noise in the score distribution and in the ROC curve To prevent this errors, we used the AUC approach in the usual way. For matched conditions, although the AUC approach does not take into account the noise, the models present improvement compared to the baseline. AUC compensates for the variations induced by the noise and focus on increasing the AUC. Note that the effects are diminished as shown also in the Figure 6.3. 6.8 Comparing AUC approach with Calibration approach In this section we point out some of the differences between the AUC approach and the Calibration approach stated in [131, 77]. The calibration has been successfully used in recent studies with successful results [132, 131, 133]. The idea behind calibration is to transform the scores into LLRs. Let, θ(χ) represent a trial score for a speaker sand phrase χ. Then, Θ(χ) = [θ1(χ)θ2(χ)...θN(χ)]. The transformation can be represented as follows, ˆ Θ(χ) = R(θ(χ),Λ) = aθ(χ) + b(6.51) where Rrepresents the linear transformation, and Λ are the parameters a, b. When optimizing these parameters, the objective logistic regression objective function is minimized [77]. On one hand, the calibration approach goal is to optimize the set of Λ parameters. The LLR function employed is, θ(χ) = log eˆ θ(χ)− N X j=1 log eˆ θ(χ).(6.52) On the other hand, the discriminative approach we propose, maximizes the separation between target and imposter models [20, 134]. It employs positive and negative samples to optimize the GMM model parameters. But the main purpose is either maximize the Area under the ROC curve, the DCF or maximize the separation between two opponent models. The main LLR function is as follows, θ(χ) = log (p(χ|ΛS)) −log (p(χ|Λ¯ S)) ,
64 Chapter 6. Optimization of the area under the ROC curve where log (p(χ|ΛS)) and log (p(χ|Λ¯ S)) are the log-likelihoods corresponding to the target model ΛSand the anti-model Λ¯ S. In this respect the goal that follow both techniques is different. However, a study of the CLLR under the discriminative approach can be for future research. 6.9 Experiments and Results We ran experiments to evaluate the proposed AUC-minimization approach. Two experiments were run. In the first we compared the performance of conventional MAP and JFA based learning with JFA optimized using the AUC criterion on speech recordings, where the noise conditions in the training and test data were matched. In the second we compare the performance of AUCminimization base MVE against conventional methods on mismatched conditions, where the test data are noisy. 6.9.1 On the selection of the competing samples In general, the discriminative methodologies use positive and negative examples to optimize the objective function (see Section 5.5 for details on the databases used). However, there are negative samples (or imposter speakers) that are more suitable to compute more accurate models. In this section we study which samples can produce impact in the correct training of the models. The different discriminative approaches described are based on the minimum classification approach. All of them depend on a misverification measure which is a difference of likelihoods of the form, d(χu,Ωu,Λu) = −g(χu,Ωu,Λu) + G(χu,Ωu,Λˆu).(6.53) In this sense, the training samples that can be highly confused as imposters or target speakers represent a higher risk, see Figure 6.4. −5 0 5 0 0.2 0.4 0.6 0.8 1 τ miss classification correct classification Figure 6.4: Vulnerable area in misverification measure d In the training, the designer has a complete control of the system. A filtering of the samples that belong to the correct or incorrect classification; i.e. the flat parts of the sigmoid function, can be granted as correct or penalized accordingly. However, the samples that appear in the slope of the sigmoid function are the ones that refine the models. Hence, after a few iterations and noting convergence, it is possible to discard samples which give high and low misverification measures and focus on the central part of the sigmoid. This observation is important for both the current algorithm where at first glance the basic approach is to perform the experiments with all the data available.
6.9 Experiments and Results 65 Table 6.1, analyzes the performance for a set of 100 users discarding 10%, 20%, 30% and 40% of the total data. By sorting the cohort sets with the lowest d, we found the sets that are more confusing for the classifier. 4We observe that the results for having the complete data and 10% reduction present marginal decay.5 Samples discarded EER original MAP 10.5 AUC-MVE 9.34 10% 9.36 20% 9.45 30% 9.87 40% 10.7 Table 6.1: On the selection of the number of cohort samples for discriminative training 6.9.2 Experiments on clean speech In the first approach all training and test recordings were noise-free, although collected over varied channels. This is the standard setup for the NIST 2008 test. We compared the baseline classification performance to that obtained with AUC-optimized JFA. For the AUC-optimized JFA, the loading matrices V, U and Dwere computed to minimize the AUC. The factors ySand xH,S were learned in a conventional way. For the test speakers, target users were enrolled conventionally, through MAP adaptation and computation of the factors: xS,yH,S and z. For the AUC-optimized training, we need: a) target and impostor tokens, b) an initial target model and the imposter. The initial target model was provided by the adapted models in the baseline (MAP). Table 6.2 shows the results obtained. They are consistent with comparisons performed by other researchers: JFA outperforms both baseline MAP learning as well MVE learning significantly. We note that the JFA uses an optimized implementation by [28] and the results too have been optimized empirically and may be considered to be competitive for the particular data setup used. All performance numbers are noted to improve as the amount of training data used to learn the base UBM increases for the first iteration and consecutively decrease in the following iterations up to 30% of the actual cohort set. More importantly, we note that the AUC-optimized classifier actually outperforms conventionally trained models both for MVE and JFA. In fact, over multiple runs of the experiment with different initializations and parameterizations, the trend of results in Table 6.2 were maintained. In addition, Figure 6.5 we show the results of the current AUC of the different approaches. 6.9.3 Experiments on noisy speech In the second experiment we compared baseline techniques to AUC-minimized training on noisy speech. Experiments were performed using speech corrupted to a variety of SNRs (10dB, 15dB, and a cocktail of 0-15dB), all of them using babble noise. The first column of Table 6.3 shows the results obtained for clean data. The results are consistent with comparisons performed by other researchers: JFA outperforms both baseline 4The experiments are based on the experimental setup in 5.5, but selecting 100 target users and performing full set optimizations for cohort and target model. 5Special care has to be given when the EER gets too low (below 1%)
66 Chapter 6. Optimization of the area under the ROC curve System EER minDCF Baseline MAP 15.95 6.7 Baseline JFA 12.07 4.2 MVE 13.51 5.8 AUC-MVE 13.21 5.4 AUC-JFA 11.93 3.9 Table 6.2: AUC optimization: EER for MVE and JFA, clean condition Figure 6.5: AUC optimization: AUC results for different systems, clean conditions MAP learning as well MVE learning significantly. In both cases, the models learned via AUCminimization somewhat outperform conventionally trained models. All performance numbers are noted to improve as the amount of training data used to learn the base UBM increases. The remaining columns of Table 6.3 compares MAP, MVE, JFA, AUC-optimized JFA and AUC-optimized MVE on noisy speech of various SNRs. We observe that AUC-optmized learning consistently outperforms conventional training in all cases. Moreover, the best results are obtained with AUC-optimized MVE. The results are consistent across all noise conditions. This contravenes the observation on clean speech, where the best performance is obtained with JFA. System clean 15 dB 10 dB 0-15dB (cocktail) MAP 15.95 18.01 17.48 35.7 MVE 13.51 17.67 17.15 28.1 JFA 12.07 17.23 16.79 27.3 JFA-AUC 11.93 16.51 16.22 24.0 MVE-AUC 13.21 15.93 15.78 22.8 Table 6.3: AUC optimization: EER of the noisy task (babble noise). Figure 6.6 shows the results for the current AUC. We can observe that for the clean scenario, JFA gets the best results; however, the MVE approaches perform better as SNR increases. Lastly, Figure 6.7 shows the relative improvement obtained from the baseline and by applying the AUC approach to both the MVE and the JFA approaches. 6.9.4 The effect of the optimization of the AUC of the ROC curve in the score distribution In this section we explore the aftermath of performing the area under the ROC curve optimization in the score domain. In here, we show an example of speech at 10dB SNR (babble noise).
7.3 Finding the partitions 73 other factors such as channel variations. Hierarchical partitioning strategies that consider multiple factors concurrently may also be used [138]. 7.3.1.1 Environment-based partitions We partition the signal space according to the SNR. In this research, we assume that partitions are formed based on the SNR of the signals. We divide the range of all possible SNR values into Pintervals. Each interval represents a partition of the signal space. Let SNRC min and SNRC max represent the minimum and maximum SNR associated with partition ΩC. A signal Xwith signal to noise ratio,SNRX, is assigned to a partition Csuch that SNRC min < SNRX≤SNRC max.(7.1) Note that partitions may also be formed based on noise type, or other known characteristics of the noise. In this research, however, we have only considered SNR. The environment partitions may also be based on the channel. For instance, we can distinguish between microphone and telephone, but also perform subcategories of them that depend on a particular feature of the device. For training purposes and in several databases (for example, NIST databases [112]) this information is available. Then, each recording associated to a certain channel type belongs to a partition ΩC. 7.3.1.2 Speaker Partitions When speaker identity is known for all recordings in the training set, partitions are obtained by clustering them by speaker. We first compute a universal background model (UBM) from unpartitioned data. We then use an agglomerative clustering procedure to cluster speakers. Initially, each speaker forms their own cluster. From all the current clusters, we select two clusters with the smallest distance and merge them.1The procedure continues until there is no cluster left. Agglomerative clustering iteratively merges the closest clusters until the desired number of clusters (and consequently, partitions) is obtained. At each stage of the clustering, the UBM is adapted via MAP adaptation to learn a model ΛC for each new cluster C. In this research the distance between any two clusters is defined by the empirical cross entropy: d(C1, C2) = 1 |χC1|log P(χC1; ΛC1) P(χC1; ΛC2)+1 |χC2|log P(χC2; ΛC2) P(χC2; ΛC1),(7.2) where χCiis the set of all recordings in cluster Ci. Other clustering mechanisms may also be employed. 7.3.2 Unsupervised Partitions When a priori knowledge about the training recordings is unavailable, partitions may be formed by clustering them using unsupervised methods. For the purpose of this research, we employed k-means as a clustering tool. The algorithm starts first deciding an appropriate number of clusters, K. In our case, we can relate the clusters according to a specific factor. For instance, if we are dealing with noisy signals at different SNR (from 0 to 20) we can decide K= 5, so that it accounts for five intervals. Later, the system can provide more granularity to these clusters. 1the smallest distance may be computed using criteria as log likelihood, cross-entropy, euclidean distance between the means, among others.
74 Chapter 7. Ensemble Modeling Approach Once the number of clusters is decided, we initialize Kinitial points, called centroids either by random values or by picking randomly Kvectors from the pool of imposter data. Afterwards, the system assigns ”memberships” to each data vector (distances from the vector to each centroid). The lowest distance, means that the vector belongs to a certain cluster. Once we account for all the data vectors, the system goes through the same process iteratively until the error function E does not change significantly, E= K X j=1 X χl∈cj (χ−cj)2.(7.3) where χis the set of all recordings in the data space, and cjis a specific centroid. The factor by which partitions are formed can be controlled by using an appropriate distance function. Generic clustering based on Euclidean distances or likelihoods may be used to cluster the data by a dominant factor. 7.4 Training the Ensemble Model Corresponding to each of the partitions Ω1,··· ,ΩPwe train a separate partition specific background model. All background models are GMMs. In principle these can be trained separately for each partition using the EM algorithm. However, we require each of the background models to be highly specific to the partition they represent, and not generalize to other partitions. In order to do so, we train all of them together using the following discriminative training procedure [138]. Let ΛCrepresent the model for a partition ΩC. Let χCrepresent all (training) recordings assigned to ΩC. For any partition ΩC, let Ω¯ C=[ ¯ C06=C Ω¯ C0(7.4) represent the complement of ΩC,i.e. the union of all partitions that are not ΩCand iis the number of the partition. Let g(χ; ΛC) = log P(χ; ΛC) represent the log-likelihood of any recording χcomputed with the distribution for partition ΩC. We can now define d(χ, ΛC), a misclassification measure for how likely it is that a data χ∈χCfrom ΩCwill be misclassified as belonging to Ω ¯ Cas d(χ, ΛC) = −g(χ; ΛC) + G(χ, Λ¯ C),(7.5) G(χ, Λ¯ C) represents the combined score obtained from a partitions in Ω ¯ C. G(χ, Λ¯ C) = log 1 |Ω¯ C|X C0:ΩC0∈Ω¯ C exp [ηg (χ, ΛC0)] 1 η .(7.6) where |Ω¯ C|is the number of partitions included in Ω ¯ C, and ηis a positive parameter. Now, we can define a new objective function for discriminative training of ΛC. This function takes the of the form, `(ΛC) = 1 |χC|X X∈χC 1 1 + exp [−γ(d(χ, ΛC) + θ)],(7.7) where |χC|represents the number of recordings in χC, and γand θare control parameters. Note that for this specific case, we consider the formulation with different competing classes (as many as the number of partitions). Hence, the optimization is performed per partition. Finally,
7.4 Training the Ensemble Model 75 the objective function in Equation 7.7 can be optimized by applying the following generalized probabilistic descent (GPD) update rule for ΛC: Λt+1 C= Λt C−∇`(ΛC)|Λt C.(7.8) Since all background models are GMMs, ΛC={wC k, µC k,ΣC k}, where wC k,µC kand ΣC kare the mixture weight, mean and covariance matrix of the k-th Gaussian of the GMM for ΛC. To obtain the update rules for these individual parameters, ∂`(ΛC) ∂wC k ,∂`(ΛC) ∂µC k and ∂`(ΛC) ∂ΣC k must respectively be plugged in for ∇`(ΛC) in the update rule of Equation 7.8. Once again, as in previous chapters we solve for the parameters. Then, lets consider the following, µ=σ˜µ, ˜σ= log(σ), w =exp( ˜w) PK k=1 exp( ˜w). For µ, a specific k, dimension dand to avoid bias, let ∇φ`(χ, ˜µk) = 1 |χC|X X∈χC γ`(1 −`)wkN(χ|µk, σk) Pk0wk0N(χ|µk0, σk0)χ σk −˜µk.(7.9) For σk, let ∇φ`(χ, ˜σk) = 1 |χC|X X∈χC γ`(1 −`)wkN(χ|µk, σk) Pk0wk0N(χ|µk0, σk0)(χ−µk ˜σk2 −1).(7.10) Finally, for the set of weights wk, ∇φ`(χ, ˜wk) = 1 |χC|X X∈χC γ`(1 −`)N(χ|µk, σk) Pk0wk0N(χ|µk0, σk0).(7.11) The procedure is employed by the MVE and JFA approaches. We obtained refined target and imposter models. 7.4.1 MVE Once the background models ΛCare obtained for all partitions, we can also train partition-specific target-speaker models, ΛC Sby fixing the background model ΛCand using a similar discriminative approach to train ΛC S. Note that we consider the usual MVE approach (the binary case), with just two classes (target and imposter), and η= 1. Hence, Equation 3.31, which is the multiclass version, can be reduced to G(χ; ΛC S) = log gχ; ΛC ¯ S.2 7.4.2 Factor Analysis In the same way, the approach can be embedded in the JFA algorithm. The basic UBM is now a set of different imposter models that are treated separately. The system trains the loading matrices V,Uand Dfor each partition. Moreover, it also computes the speaker factors for each ΩC. In the test stage, the scores for a trial are obtained against each partition.3 2refer to Chapter 3.3 for details. 3refer to Chapter 3.5.1 for details.
76 Chapter 7. Ensemble Modeling Approach 7.5 Scoring for classification Once we obtained the appropriate models for the target speaker and their corresponding imposter models, the aim is to find suitable ways to score. Given the pairs, {ΛC1,ΛC1 S},{ΛC2,ΛC2 S},···{ΛCP,ΛCP S} the set of background models for all Ppartitions and their corresponding partition-specific target speaker models for any claimed speaker S, we can compute the score θS(X) to be employed in the likelihood ratio test for any recording Xin one of several ways. Let, θS C(X) = log P(X|ΛC S)−log P(X|ΛC) (7.12) be the likelihood ratio computed in the log domain from the models for partition ΩC. The options for obtaining the final score θS(X) are: A) Partition Selection: We assign the recording to the most likely partition as ˆ C(χ) = arg max Clog P(χ|ΛC).(7.13) We then compute the score from the assigned partition: θS(X) = θS ˆ C(χ). This is a conservative score that selects the partition with signals most likely to be confused with the target speaker (see Figure 7.6). In this sense, the decision threshold must be trained with an accurate design that can guarantee a correct classification. We used Focal (described in [133, 127]) to calibrate the scores. Figure 7.2: Partition Ensemble: Selection B) A priori: If the correct partition ΩCfor Xis known a priori, then we can simply set θS(X) = θS C(X), see Figure 7.3. Figure 7.3: Partition Ensemble: A priori Although we expect to obtain the best results with this approach, the labels is not always available. So we can consider the A priori selection as an upper bound.
7.5 Scoring for classification 77 Figure 7.4: Partition Ensemble: Best score C) Best score: We select the largest score: θS(X) = maxCθS C(X), see Figure 7.4. D) Combination: Here we simply combine the scores from the different partitions (see Figure 7.5), θS(X) = X C wS CθS C(X).(7.14) In our work we trained an SVM (for details on this algorithm refer to Appendix A.2) to classify the speaker; in this case wS Care simply the weights assigned by the SVM. In a first approach we just concatenated scores from each system belonging to a target speaker, ξi= [θS 1, θS 2, ..., θS P]. Then, the labels and ξare the input data for the SVM. A second approach used the normalized log likelihoods of each trial with respect to the target model within a partition, then ξi= [gS 1, gS 2, ..., gS P]. For both, the linear kernel was employed. This is equivalent to learn the weights discriminatively. It is also comparable to the fusion approach. Figure 7.5: Partition Ensemble: Combination E) Fusion: We obtain the scores for the target speaker with respect to the Pimposter models. Then, for a trial j, target a, and Ndifferent scores, the linear fusion is given by, fj=α0+α1a1,j +α2a2,j +... +αNaN,j.(7.15) To obtain the αweights, we employ the logistic regression fusion as stated in [133, 127] (see Figure 7.6). Target speaker trial scores and imposter trial scores form two matrices: Ais a score matrix of N×K, where Nare the different classifiers (Ppartitions) and Kare the number of target trials, and Bis a score matrix of N×L, where Lare the number of imposter trials. Then, the objective logistic regression function based on the cost is C=P K K X j=1 log(1 + e−fj−logit P) + 1−P L L X j=1 log(1 + egj+logit P),(7.16) where Pis a prior probability usually set to 0.5. The target and imposter scores are given by,
78 Chapter 7. Ensemble Modeling Approach Figure 7.6: Partition Ensemble: Fusion fj=α0+ N X i=1 αiai,j gj=α0+ N X i=1 αibi,j (7.17) The complete algorithm is described in Algorithm 7.1. Algorithm 7.1. Ensemble Approach Initialization 1. Train the imposter model (UBM) in the usual way employing EM. ¯ Λ = {¯ WC k,¯µC k,¯ ΣC k}(7.18) Finding the partitions Obtain the partitions ΩPand their corresponding raw models. 1. Supervised: Assign the elements of each cluster according to some specific attribute known in advance (SNRs, similar speakers, channel). Adapt the UBM to the partition. 2. Unsupervised: Use an unsupervised clustering algorithm (like kmeans) to find the clusters. Adapt the UBM to the partition. Training the Ensemble Model 1. Compute the misverification distance for every partition with respect to the competing partitions and the loss function using, d(χ, ΛC) = −g(χ; ΛC) + G(χ, Λ¯ C),(7.19) where the combined score obtained from a partitions in Ω ¯ Cis, G(χ, Λ¯ C) = log 1 |Ω¯ C|X C0:ΩC0∈Ω¯ C exp [ηg (χ, ΛC0)] 1 η .(7.20) and the loss function, `(ΛC) = 1 |χC|X X∈χC 1 1 + exp [−γ(d(χ, ΛC) + θ)] (7.21) 2. Obtain the gradients for the model parameters: {wC k, µC k,ΣC k}and plugged them into: Λt+1 = Λt−∇`(χ; Λ),(7.22) ∇`(χ; Λ) = γ`(d)(1 −`(d))∂ p(χ; Λ) ∂Λk .(7.23)
7.6 Ensemble example: real data 79 Algorithm 7.1. Ensemble Approach, Part II Convergence 1. Compute the loss function with the new parameters: `new(χ; Λ) = X n `n(d(χ; Λ))1(χ∈Cn).(7.24) Compare `new with `if it is higher than a threshold τiterate from Parameter Estimation. Use the new Prefined models as the UBM for either MVE or JFA approach. Scoring for classification Given the pairs, {ΛC1,ΛC1 S},{ΛC2,ΛC2 S},· · · {ΛCP,ΛCP S}(7.25) 1. Compute the score for each partition, θS C(X) = log P(X|ΛC S)−log P(X|ΛC).(7.26) 2. Choose among the following scoring options. a) Partition Selection: Assign the recording to the most likely partition as, ˆ C(χ) = arg maxClog P(χ|ΛC). b) A priori: If the correct partition ΩCfor Xis known a priori, then we can simply set θS(X) = θS C(X). c) Best-score: We select the largest score, θS(X) = maxCθS C(X). d) Combination: Combine the scores from the different partitions: θS(X) = PCwS CθS C(X). in this case wS Care simply the weights assigned by the SVM. e) Fusion: Combine the scores from the partitions in a logistic regression fusion, fj=α0+PN i=1 αiai,j . 7.6 Ensemble example: real data We present an example of the ensemble approach using clean data. The intention is to provide the reader of some insight of the procedure before considering more complex experiments. In a first stage we train the target models and the UBM as usual. Secondly, we partition the data space to obtain specific models. Next, the system trains the models in a discriminative manner such that the empirical error is minimized and the model specificity is enhanced. Finally, several iterations are performed to reach convergence. The first step after computing the UBM and the target models is to obtain a set of partitions ΩPfrom the data space. We employed the unsupervised partition method to show the performance in the worst case scenario where no information about the clusters is available. Figure 7.7 shows the real partition using k-means for the data space, considering a clean dataset and an extraction of the real UBM. For each of those partitions the system builds a model, in this case an adaptation of the current UBM.4 The discriminative ensemble approach is then applied to each of the partitions. Figure 7.8 shows the misverification measure d(χ, ΛC) for some of the partitions (initial and final after 10 iterations) and a single target user. This graph clearly depicts an empirical error reduction. Note that the model parameters change along the iterations. The discriminant function (log-likelihood) distribution from a target and its corresponding 4The partition models might be an adaptation from the current UBM or an independent computation of a model. In our experience, the adaptation works better for all cases.
80 Chapter 7. Ensemble Modeling Approach −4 −3 −2 −1 0 1 2 3 −3 −2 −1 0 1 2 3 Dimension 1 Dimension 2 Figure 7.7: Partition Ensemble Scheme, real example Figure 7.8: Value of d for 3 different clusters and different iterations. cohort distributions are shown in Figure 7.9.5We can observe that the distributions spread and separate along the iterations, altering the misverification distance d. After reaching convergence for all the cohort models, the system performs the selected scoring method for every partition. The score results for the ten partitions and the specific target user model are shown in Table 7.1. Note that some of the imposter scores seem to be very close to the actual target speaker; for instance, cluster ones and four. Recall that all of these clusters represent negative samples for the target speaker. The actual best score for a set of positive samples for this speaker is of 0.91216, which is clearly higher than the imposter scores presented in Table 7.1. Moreover, if the data of the target speaker were placed in some cluster, most of the samples would be found in cluster five (approximately 60%). The imposter samples were randomly selected from the pool of imposter set, but the membership cluster is number ten (approximately 35%). The clusters that provide more information about imposters that may be confused with the target are 1 and 4, according to the likelihood results. Note that these scores were obtained using unseen data, that is the reason why the actual results from a priori ,best score, and partition selection are diferent. Several options are possible for the selection of the most appropriate score to consider. Table 7.2 describes the results for the scoring method. According to this results, the selected scoring method is the combination of scores and the fusion. 5For this specific example, we used all the cohort samples available from the imposter data set.
7.7 Experiments and Results 81 −100 −90 −80 −70 −60 −50 −40 −30 −20 0 500 1000 log likelihood sample from true speaker −100 −90 −80 −70 −60 −50 −40 −30 −20 0 500 1000 1500 log likelihood sample from cohort speaker −100 −90 −80 −70 −60 −50 −40 −30 −20 0 500 1000 log likelihood sample from true speaker −100 −90 −80 −70 −60 −50 −40 −30 −20 0 500 1000 1500 log likelihood sample from cohort speaker Figure 7.9: Discriminant function (log-likelihood) histograms for one target and 2 cohort samples This procedure is extended to all the target speakers; the results are shown in the following sections. 7.7 Experiments and Results This section includes the experimental analysis for the ensemble approach. Algorithms such as MVE and JFA are used together with the Ensemble to present its capability. We show results for both the clean and noise condition and how the partitions are performed for each case. Moreover, we present a comparison of techniques and how they are affected by our algorithm. 7.7.1 Experimental Setup We employed the NIST Speaker Evaluation 2004, 2005, 2010 and 2008 database [40] to complete this study. We conducted two experiments, one on clean data and the second on noise-corrupted data. For all experiments, we used the NIST2008 database as targets. For the noise condition experiments, babble noise, extracted from the Aurora 2 database, was added to the training and test files at different SNRS: 0,5,10,15 and 20 dB. The training data was randomly partitioned into equal parts, one for each noise condition; the same procedure was applied for the test set.6 The objective in the clean experiment was to partition the space by speaker. For the noisy data, we evaluated partitioning by noise condition. In both cases we evaluated partitions formed 6The noise was added using FaNT software [123].
82 Chapter 7. Ensemble Modeling Approach cluster score imposter samples score target samples 10.439598 0.53196 2-0.540602 0.88931 3-0.600463 0.91316 40.364815 0.44016 5-0.443482 0.87047 6-0.200627 0.84925 7-0.315752 0.85311 8-0.228535 0.83860 9-0.326056 0.86034 10 -0.510749 0.87213 Table 7.1: Scores for 10 clusters with respect to a target speaker model. Scoring method Score imposter samples Score target samples partition selection -0.443482 0.88931 a priori -0.510749 0.87047 best score -0.600463 0.91316 combination -0.538124 0.92154 fusion -0.672565 0.93768 Table 7.2: Scoring selection scores with respect to a target speaker model. from a priori information about the data, as well as by unsupervised k-means clustering. The five methods: partition selection,(PS), a priori, (AP), best score, (BS), combination, (CO), and fusion,(FU), were evaluated. We used a 256-Gaussian GMM for background models in all cases. In the clean experiment we evaluated the performance obtained with speaker-based partitions with different numbers of partitions. As a baseline we also evaluated the conventional verification framework, where we trained generic gender-dependent background models (UBM) and adapted these to the target speaker using MAP [66], Joint Factor Analysis (JFA) [15, 61] and Minimum Verification Error (MVE) [84] training. The UBM did not include data from any target speaker, not even for the data partitions and their negative/positve samples. For the noise experiment we employed 5 partitions, one corresponding to each SNR level. The proposed method only establishes a mechanism for defining background models for each partition. The method described in Section 7.4 actually uses MVE to learn partition-specific target models. Partition-specific target models were also trained via JFA or MAP. For the noise experiments these were also evaluated in addition to UBM-based baselines. 7.7.2 Results Table 7.3 shows the results for the clean experiments with different numbers of partitions. The EER and minDFC results are shown. The ensemble model improves performance in every case. Morever, partitions based on a priori knowledge consistently outperform partitions obtained from unsupervised clustering. The best result is obtained when scores from the partitions are fused; this result significantly outperforms JFA, and is the best result we have ever obtained on this test set. Increasing the number of partitions beyond 16 resulted in degradation of performance probably due to the lack of data. Discriminative training of background models to make them specific also turns out to be the key. The alternative is to simply train each of them using maximum likelihood. This was consistently worse than discriminative refinement of partition-specific background models in all experiments. Figure 7.10 presents the relative improvements for the best results (combination and fusion
8.3 Combining Ensemble Modeling and AUC of the ROC curve optimization 89 the minDCF. A certain partition ubecomes a target model, the rest of the partitions become the competing models. Let χrepresent the complete set of target and competing training instances: χ=H∪W, where Hcontains all the target instances and Wall the competing instances. Thus, the modified AUC function to optimize for a specific partition uis, Υ(χ, Λu) = 1.0−Pχ∈H Pˆχ∈W R(θu(χ), θˆu(ˆχ)) |H||W| ,(8.6) θuand θˆuare the usual scores of χwith respect to a target model Λuor competing model Λˆu, and Ris the soft representation of the sigmoid function. The modified AUC function of Equation 8.6 can then be optimized using the GPD algorithm, Λt+1 u= Λt u−∇Υ(χ, Λu) (8.7) ∇Υ(χ, Λu) = −1 |H||W| X χ∈H X ˆχ∈W γR(1 −R)∂θu(χ) ∂Λu−∂θˆu(ˆχ) ∂Λˆu.(8.8) In the above Equation Ris a short-hand notation for R(θ(χ), θ(ˆχ)), and is a learning rate parameter. The general summary process is depicted in Figure 8.3. From a full data space, the algorithm performs a partition among a specific attribute (first layer). We optimize and maximize the separation of this new models; i.e., the models are refined. In a second stage, the algorithm maximizes the separation within the current clusters, and refine them. In the final stage, for the special case of SV, we have to optimize the target speaker against the impostors. Figure 8.3: Ensemble-AUC scheme Two clear applications arise from the above methodology: channel mismatch and the noisy environments. 8.3.3 Channel Mismatch The channel among the recordings used for training and testing might be different, causing a poor system performance. To solve this issue, let us introduce the channel as the intended ensemble attribute. We start from the premise that a gender dependent UBM embraces the universe of all possible conditions and will represent the Total Space as T. Environment clustering (EC) and environment partitioning (EP) are suitable off-line options to distribute Tand to prepare the channel dependent regions. For instance, the sub-partitions belonging to Ωucan differentiate between microphone and telephone in a first level of the hierarchical structure. Afterwards, we can refine the sub-region models by increasing coverage. Using the a priori information from the channel clustering and partitioning, we can update the models by performing inter-channel training denoted as ΛQu, where Qurepresents the partition among a condition,
90 Chapter 8. AUC and Ensemble Integration followed by intra-channel training (computing ΛΩQuamong the competing models). The former will produce a spread among the channel sub-spaces, the latter will broaden the coverage. The final models are constructed accordingly. For a test vector, we estimate the model from which it was produced, stochastic matching [139] is used for this purpose. Quare estimated according to the chosen EC or EP algorithm. For both, a cluster selection is performed based on the highest likelihood to the testing data χY. The mapping of such utterance is given by, ΩQu= argmax uP(χY|R(ΩQu)) (8.9) QYu=Gφ(ΩQu),(8.10) where Yrefers to the test data, and Gφis a function that transforms the models from a training channel domain to the new test domain. 8.3.4 Noisy Environment Once again, for this approach, we take the total data space Tas our starting point. In this case, EC and EP prepare the sub-regions comprising Ωe, consistent with the noise type and SNR. According to [138], after gender separation, a second layer of the hierarchical tree divides the sub-spaces in low and high SNR. We employ this methodology with a later refinement of each sub-region using the known algorithm to increase the coverage. The inter-environment algorithm maximizes the separation among the sub-regions and computes Qe(each partition refers to a cluster with certain low or high SNR). Moreover, the coverage within each sub-region is broadened with the intraenvironment technique, obtaining ΩQe. The final on-line models, QYeare estimated according to the chosen EC or EP, considering, QYe=Gφ(ΩQe),(8.11) Gφis a mapping function that relates the training noisy condition with the test condition. 8.3.5 Channel Mismatch and Noisy environment This combined scheme presents the summary of the previous approaches. The first step is to perform a gender separation. Afterwards, EC computes Ωuusing a channel basis with the corresponding refinement. In the third layer of a hierarchical tree the noise type and SNR are included resulting in a set of Ωe csub-spaces. This final sub-spaces contain the information of channel, noise type and SNR. The next step is to increase the separation among Ωe usub-spaces, computing Qe u. Moreover, to increase the coverage, we can compute ΩQe u. The on-line models, QYe uare estimated according to QYe u=Gφ(ΩQe u),(8.12) where Gφis a mapping function that transforms the models from a training domain to the test domain. 8.4 The speaker modeling For each target speaker we perform a similar procedure. After employing any of the above analyses, it is possible to refine the speaker space by maximizing the separation among competing speakers, inter-speaker training.
8.4 The speaker modeling 91 For instance, a new objective function and misclassification measure for Uutterances from a speaker sin different conditions take the form, `(Ωs) = 1 UX U 1 1 + exp −γdχu,Ωe u,s,Λu+θ,(8.13) where χuis a utterance in the training data for partition u, Ωe s,u belongs to a spanned subspace Ω (that considers the channel and the SNR), γand θare control parameters and Λe u,s represent the GMM parameters for that specific region. Accordingly, dχu,Ωe u,s,Λu=−gχu,Ωe u,s,Λu+Gχu,Ωe ˆu,s,Λˆu,(8.14) where gχu,Ωe u,s,Λu= log p(χ; Ωe u,s,Λu), Gχu,Ωe u,s,Λˆu= log p(χ; Ωe ˆu,s,Λˆu).(8.15) Equations 8.14 and 8.15 show the condition dependance. If we just want to finish the process by increasing the coverage, we can solve by applying the (GPD) update rule for Λe u,s, Λt+1 u,e,s = Λt u,s,e −∇`(Λu,s,e)|Ωt u,s,e .(8.16) For the purpose of this work, we present a formulation using the optimization of the AUC of the ROC curve. Hence, we explore two approaches as in previous Chapters: MVE, and JFA. 8.4.1 MVE in an Ensemble approach For MVE approach we define the following objective function that represents the WMW statistic, Υ(Λu,e)=1.0−Pχ∈H PX∈χPˆχ∈W Pˆ X∈ˆχR(θu,e(X), θu,e(ˆ X)) Pχ∈H LχPχ∈W Lχ ,(8.17) where Lχis the number of feature vectors in χ. Note that the formulation dependance to the noise type and channel refined space given by Ωu,e. The scores θare computed according to this space. Using the representation X χ∈H Lχ=|H|,X χ∈W Lχ=|W|, the GPD update rule for any parameter φof the distributions is now given by φt+1 =φt− ∇φΥ(X,Λu,e), where ∇φΥ (X,Λu,e) = −1 |H||W| X χ∈H X X∈χX ˆχ∈W X ˆ X∈ˆχ γR(1 −R)∇φl(X, ˆ X, Λu,e),(8.18) and ∇φl(X, ˆ X, Λu,e) is a local gradient with respect to φat χ, ˆχand has the form given by Equation 8.19, ∇φl(X, ˆ X, Λu,e) = −∂θu,e(X) ∂φ +∂θu,e(ˆ X) ∂φ ,(8.19) ∂θu,e(X) ∂φ represents the derivative of the log-likelihood-difference given by the Gaussian mixture models for the target speaker and the universal background model for vector Xwith respect to φ. The final solution can be given in two ways: a scoring method or stochastic matching. The scoring techniques described in Section 6.5 uses a set of scores for different partitions to conform a fusion scheme. Stochastic matching is also a possible solution to decide the environment for a particular phrase without needing to perform a scoring method.
92 Chapter 8. AUC and Ensemble Integration 8.4.2 Factor Analysis in an Ensemble Approach The proposed ensemble method pursues the same goal as JFA, to extract the information of the speaker while marginalizing the contribution of any variability caused by channel or noise condition. Going further we include the ensemble approach for the i-vector and refine the current models so that the i-vectors will properly cover the space for different conditions. Recall, the utterance is defined as, M=s+Tω, (8.20) where Mrepresents the combined speaker and channel-independent supervector extracted from the UBM. Tis a low rank rectangular matrix and ωis a random vector, composed of the identity vectors (i-vectors), with normal distribution N(0, I). Mis normally distributed with mean mand covariance matrix TT0.3Hence, the new modeling is in fact a projection of a utterance onto the low-dimensional total variability space. In an ensemble scenario, the space is partitioned into several regions depending on the desired channel and/or noisy condition. Hence, let Xrepresent a large collection of recordings χobtained from a large number of speakers in a certain Ωu,e. The starting point for the training are the models Λu,e, defined by those regions. For this formulation we also assume that the score is a likelihood ratio, so that the optimization holds. Let Sbe the set of all speakers in X. Let XSrepresent the subset of Xrepresenting recordings from speaker Sand X¯ Sbe recordings from remaining speakers, i.e. X=XS∪X¯ S. Note that, X can be partitioned in this manner in as many ways as there are speakers in S. To learn global parameters, we define the following AUC objective for each region Ωu,e. Υ(Λu,e)=1.0−1 |XS||X¯ S|X S∈S X χ∈XSX χ∈X¯ S R(θu,e S(χ), θu,e S(ˆχ)) (8.21) The GPD update rule for any global parameter φu,e is given by φu,e t+1 =φu,e t−∇φu,e Υ(X,Λu,e), where ∇φu,e Υ(X,Λu,e) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R)∇φu,e l(χ, ˆχ, Λu,e) |XS||X¯ S|(8.22) and ∇φu,e l(χ, ˆχ, Λu,e) is a local gradient with respect to φu,e at χ, ˆχ. ∇φu,e l(χ, ˆχ, Λu,e) = −∂θu,e(χ) ∂φu,e +∂θu,e(ˆχ) ∂φu,e (8.23) To derive the update rules for individual parameters, it is sufficient to obtain the derivatives with respect to the corresponding loading matrix T(refer to Chapter 7 for details). Moreover, we can optimize for the i-vector in a similar manner. The next step, employing the PLDA is straightforward, for each of the subspaces Ωu,e. Consequently, we can employ a fusion scheme to decide the most appropriate scoring solution. The global approach proposed has a potential success since it structures the space in advance, giving the opportunity to use the a priori information for a better modeling. Furthermore, the technique has been presented for a SV scenario, under different real conditions. Note, that this space refinement can be properly included in other optimization schemes, such as ML-MAP approach. 3The procedure to train the matrix Tis very similar to the training a the loading Vin JFA, but in the eigenvoice training all the utterances from a speaker are treated as different speakers, refer to Appendix B.2 and B.3 for further details.
8.4 The speaker modeling 93 Algorithm 8.1. Ensemble modeling and area under the ROC curve optimization integration Initialization 1. Train the general model (UBM) in the usual way employing EM. ¯ Λ = {¯ WC k,¯µC k,¯ ΣC k}(8.24) 2. Decide the number of levels in the hierarchy and its type. Usually we get two levels: one for channel and one for noise SNR. Ensemble Modeling 1. Obtain the channel partitions Ωuand their corresponding raw models, Λu={Wu k, µu k,Σu k}(8.25) 2. Training the Ensemble Model: Choose among EC or EP. •Enviroment Clustering (EC): For a hierarchical tree of P partitions, we can denote the subspaces as: Ω = {Ω1∪ Ω2, ...Ωu, ...∪ΩP}. Each of the these clusters can be represented by a super-vector Ωrepuby a function R(·). Hence, Ωrepu= R(Ωu). •Environment Partitioning (EP): Generate from the full data space a GMM model and partition the GMM components according to similarity measures. Each subset u∈Ω of components define the new region. Once again, ΩGMMu= R(Ωu). 3. Define the loss function for Uutterances in Pdifferent conditions by, `(Ωu) = 1 UX U 1 1 + exp [−γ(d(χu,Ωu,Λu) + θ)] (8.26) where χuis a utterance in the training data for partition u, Ωubelongs to a spanned subspace Ω, γand θare control parameters and ΛC represent the GMM parameters for that specific region. 4. Define the misclassification measure d(χu,Ωu,Λu) = −g(χu,Ωu,Λu) + G(χu,Ωu,Λˆu),(8.27) and G(χu,Ωu,Λˆu) = log 1 U−1X j,j6=u exp [η×g(χu,Ωj,Λj)] 1 η . (8.28) where ηis positive, ˆuare the regions not belonging to u, and Gtakes into account the competing models contribution. 5. The solution is given by applying the (GPD) update rule for Λu, Λt+1 u= Λt u−∇`(Λu)|Ωt u.(8.29) 6. Check for convergence (a) Compute the loss function with the new parameters: `new(Λu) = 1 UX U 1 1 + exp [−γ(d(χu,Ωu,Λu) + θ)] (8.30) (b) Compare `new with `if it is higher than a threshold τ, iterate.
94 Chapter 8. AUC and Ensemble Integration Algorithm 8.1. Ensemble modeling and area under the ROC curve optimization integration Refining the Models (AUC approach) Let χrepresent the complete set target and competing training instances: χ=H ∪ W, where Hcontains all the target instances and Wall the competing instances. 1. Compute the modified AUC function to optimize for a specific partition uis, Υ(χ, Λu) = 1.0−Pχ∈H Pˆχ∈W R(θu(χ), θˆu(ˆχ)) |H||W| .(8.31) θand ˆ θare the usual scores of χwith respect to a target model Λuor competing model Λˆu, and Ris the soft representation of the sigmoid function. 2. Optimize Equation 8.31 using the GPD algorithm using the following rule, Λt+1 u= Λt u−∇Υ(χ, Λu).Therefore, ∇Υ(χ, Λu) = −1 |H||W| X χ∈H X ˆχ∈W γR(1 −R)∂θu(χ) ∂Λu −∂θˆu(ˆχ) ∂Λˆu. (8.32) 3. Check for convergence (a) Compute the loss function with the new parameters: Υnew(χ, Λu) = 1.0−Pχ∈H Pˆχ∈W R(θu(χ), θˆu(ˆχ)) |H||W| .(8.33) (b) Compare Υnew with Υ if it is higher than a threshold τ, iterate. Speaker Modeling 1. Compute the misverification function and for Uutterances for a speaker s, `(Ωs) = 1 UX U 1 1 + exp −γdχu,Ωe u,s,Λu+θ (8.34) where χuis a utterance in the training data for partition u, Ωe s,u belongs to a spanned subspace Ω (that considers the channel and the SNR), γand θare control parameters and Λe u,s represent the GMM parameters for that specific region. 2. Compute the misverification measure, dχu,Ωe u,s,Λu=−gχu,Ωe u,s,Λu+Gχu,Ωe ˆu,s,Λˆu,(8.35) where gχu,Ωe u,s,Λu= log p(χ; Ωe u,s,Λu), Gχu,Ωe u,s,Λˆu= log p(χ; Ωe ˆu,s,Λˆu) (8.36) 3. Solve by applying the (GPD) update rule for Λe u,s, Λt+1 u,e,s = Λt u,s,e −∇`(Λu,s,e)|Ωt u,s,e .(8.37) 4. Decide which method to use for optimizing: i-vector or MVE, solve accordingly.
8.5 Experiments and Results 95 This scheme has similarities with JFA and probably pursues the same goal, to marginalize the variability not belonging to the speaker. On one hand, the ensemble modeling segments the space into regions in which the classification is localized for those specific partitions. On the other hand, the AUC of the ROC curve refines the models for every operating point. But the main purpose is to be employed in noisy conditions. In this case, the target partitions are modeled so that the localized operating points are improved. 8.5 Experiments and Results This section describes the integration of both techniques and the final results up to date. Note that in this part we include experiments not only with the traditional approaches, but also with the i-vectors, considering it just as a front-end from which we can train a model, see Tables 8.1 and 8.2. 5c-MAP 5c -MVE Baseline 28.6 28.3 minDCF 15.60 15.32 Condition PS AP BS CO FU PS AP BS CO FU k-means 27.2 23.4 28.6 21.3 19.4 25.1 22.3 25.9 21.2 18.4. minDCF 15.73 12.35 17.43 10.57 8.19 11.56 8.44 12.32 8.14 6.57 prior info 23.2 21.5 24.3 20.0 17.2 21.6 19.9 22.3 18.7 16.2 minDCF 12.88 8.03 13.2 7.84 6.56 7.91 7.53 8.54 7.23 6.68 Table 8.1: Integration approach: EER and minDCF for different clusters on noise condition, MAP and MVE. 5c-JFA 5c i-vector Baseline 26.3 25.0 minDCF 14.03 13.23 Condition PS AP BS CO FU PS AP BS CO FU k-means 24.9 23.4 25.6 23.2 20.4 24.2 23.1 24.6 22.3 19.5 minDCF 13.85 13.63 15.72 12.74 7.36 12.72 12.47 14.85 11.87 6.84 prior info 23.3 22.5 23.3 21.6 19.0 22.3 21.2 22.5 20.7 18.1 minDCF 9.77 10.13 12.78 8.11 7.51 8.71 9.16 11.90 7.31 6.05 Table 8.2: Integration approach: EER and minDCF for different clusters on noise condition, JFA and i-vector. 8.5.1 Experimental Setup The setup for this experiments follows exactly the same procedure showed in the ensemble modeling. We employed the NIST Speaker Evaluation 2004, 2005, 2010 and 2008 database (see Section 7.7.1). We conducted a experiment on noise-corrupted data. For all experiments, we used the NIST2008 SRE database as targets. Babble noise, extracted from the Aurora 2 database, was added randomly to the training and test files at different SNRS: 0, 5, 10, 15 and 20 dB. Moreover, we selected the conditions and channel as attributes to explore. First, we built a hierarchical tree, which first layer focuses on the channel attribute and a second layer that deals with the noise condition. The results were obtained for 5 cluster partition
96 Chapter 8. AUC and Ensemble Integration in the first layer, plus a refinement of those models. 2 channels were considered: microphone and telephone. Not specific partition was explored within them. The second layer accommodates the noise partition for different SNR (considering this experiments were performed just using babble noise). MAP and MVE approaches were employed as explained in Chapter 7. The JFA experiments used 50 eigenchannels for U, and 400 eigenvoices for V for every region. This experiment shows the results for a simple i-vector approach as described in Chapter 3. We extracted a set of 300 i-vectors that were followed by the reduction to 100 using LDA. Afterwards, this new i-vectors are processed by the PLDA. The five methods: partition selection,(PS), a priori, (AP), best score, (BS), combination, (CO), and fusion,(FU), were evaluated. Their corresponding unsupervised (k-means) and supervised partitioning were also explored. 8.5.2 Results Tables 8.1 and 8.2 depict the results for the different optimizations and ensemble approaches; the AUC optimization is implicit in these results. We note that the best results were obtained for the combination(CO) and fusion (FU) in terms of EER and minDCF. The current ensemble approach helps the current modeling procedures such as JFA and MVE. Surprisingly, the MVE optimization obtains the best results for this specific noise condition. Figure 8.4 shows the results of the AUC of the ROC curve for all possible cases. Figure 8.4: Integration approach: AUC results for babble noise condition by scoring method. First, we want to point out the improvements in terms of the scoring selection methods. For every set of results the highest AUC were obtained by the MVE-related approaches, supporting the potentiality of the method in noise condition. Figure 8.5 shows once again the results of the AUC, but it considers the optimization methods employed. Note that the highest results are obtained for the scoring methods: combination(CO) and fusion (FU). As expected the results show that considering the prior information is of great benefit to any of the approaches. Moreover, for the unsupervised method the improvements are still noticeable. Lastly, Figure 8.6 shows the relative improvements with respect to the most successful scoring methods: combination(CO) and fusion (FU). We observe how the MVE-based approaches outperform the current baselines such as MVE and JFA. But also the ensemble modeling alone improves the results within its optimization method, specially when the prior information is known. Note also that the results considered here are under cocktail noise of 0 to 15 dB SNR.
8.6 Summary 97 Figure 8.5: Integration approach: AUC results for babble noise condition by optimization approach. Figure 8.6: Integration approach: relative improvements ( babble noise condition) 8.6 Summary In this chapter, we showed the integration of both techniques, the AUC of the ROC curve optimization and the ensemble modeling. We gave some insight of how the process takes place and gave some experimental examples. We observed that performing the integration of both can improve the actual baselines especially under noisy conditions. We concluded that by including the ensemble approach, the models become more specific. Furthermore, considering the AUC optimization the ROC curve is optimized in every point deriving in improvements for every modeling scheme.
98 Chapter 8. AUC and Ensemble Integration
9.3 YOHO database 105 Figure 9.5: AUC of the ROC curve summary results for Gaudi database. Figure 9.6: Relative improvement summary for Gaudi database. YOHO database [50] is one of the oldest databases in SV, it has the following properties: contains clean voice utterances of 138 speakers of different nationalities. It is a combination of lock phrases (for instance, ”Thirty-Two, Forty-One, Twenty-Five”), with 4 enrollment sessions per subject and 24 phrases per enrollment session; 10 verification sessions per subject and 4 phrases per verification session. Given 18768 sentences, 13248 sentences were used for training and 5520 sentences for testing. 9.3.1 Experimental Setup We used the whole database (138 users) to obtain results on our different approaches. We just used the current users of the database as target speakers and never constituted the imposter models. The basic UBM was generated using the NIST databases: 2004, 2005, 2006, 2010 (including the microphone). We tested the performance for a feature vector of 33 attributes: 16 static Cepstral, 1log Energy, and 16 delta Cepstral coefficients (same configuration as presented for MOBIO). Once again, a gender-dependent and target-independent 512-mixture GMM imposter model was used as a baseline. The initial GMM models were obtained from the NIST database plus the MOBIO data set used for training. For the test, we chose users from the database but also random users from other databases, such as NIST2004. JFA employed 50 eigenchannels for latent variable, U, and 100 for eigenvoices, V. The discriminative approaches (MVE, the optimization of the AUC of the ROC curve, and the ensemble) use the usual selection of negative samples: 500 phrases from NIST 2004, 2005,
106 Chapter 9. Discussion 2006, switchboard 1, switchboard 2 and 500 phrases from NIST 2010 microphone (3000 imposter phrases). For the clean experiments of the ensemble modeling, we limited the UBM space to five partitions. The partitions were blindly computed and the final outcome is the fusion of the scores. For the integration approach, the region was blindly partitioned in five and a final fusion score is presented. For the noise conditions experiments we used babble noise, extracted from the Aurora 2 database. It was added to the training and test files at different SNRS: 0,5,10,15 and 20 dB. For the ensemble training we used five partitions scheme. The same is applied to the integration. For every SNR-dependent region the new data space is partitioned into two new regions. 9.3.2 Results Table 9.5 shows the results for clean speech. The integration approach obtains the best results. Moreover, the system shows improvements for every approach. It is interesting that the ensemble modeling outperforms the optimized AUC for MVE, but not the optimized AUC for JFA. The JFA with the AUC optimization converges rapidly giving a good estimation of the parameters. Although the optimization of the ROC curve obtains a low misverification measure in the training stage, was not enough to compute an adequate model estimation. Table 9.5 also shows the results for babble cocktail noise. They follow the same trend as the clean experiments. Note that even for small vocabulary the proposed approaches lowered both the EER and the minDCF. Clean Noise EER minDCF EER minDCF Baseline MAP 3.29 2.1 7.24 5.8 MVE 2.93 2.1 6.53 5.2 JFA 1.54 1.1 4.39 3.9 MVE-AUC 1.64 1.2 5.58 4.8 JFA-AUC 0.81 0.5 3.23 2.1 Ensemble Modeling 1.33 0.9 3.51 2.3 Integration 0.62 0.4 2.62 1.6 Table 9.5: Final results (EER and minDCF) on the Test set for the YOHO database. Figure 9.7 shows the AUC of the ROC curve metric. Because of the few number of speakers and the high EER, the results of the AUC are very small. We can observe the best results for the Integration approach. JFA and JFA-AUC methods performed better than the MVE and MVEAUC for the clean case (following the same trend as in previous databases). The AUC values are smaller for the noise conditions. All of them seem competitive, except for the integration modeling that clearly outperforms the rest. In Figure 9.8 we show the relative improvement for the proposed approaches. The integration modeling outperforms the rest of the approaches for clean and noise condition. MVE-AUC is compared to MVE as the baseline, and JFA-AUC, ensemble and integration are compared to JFA as baseline. The ensemble modeling optimization shows the least improvements for this specific database.
9.4 Summary 107 Figure 9.7: AUC of the ROC curve summary results for YOHO database. Figure 9.8: Relative improvement summary for YOHO database. 9.4 Summary This chapter showed the final and outstanding results for the proposed optimization methods. We highlight that the optimization of the area under the ROC curve is beneficial in any case. Even for the state of the art approaches: MVE and JFA, the optimization reduces the EER and the minDCF. The ensemble approach was mainly focused to alleviate the channel mismatch problem and the noise conditions. However, it shows that it can also be used for clean conditions, obtaining fair results. Moreover, the integration of all the approaches show a significant improvement, especially when the fusion of the results is performed. The AUC of the ROC curve summary graphs show also significant results that support the contribution of the proposed approaches. Finally, the relative improvements trend emphasizes the potential of the techniques for new scenarios.
108 Chapter 9. Discussion
Chapter 10 Conclusions and Future Work Finally, from so little sleeping and so much reading, his brain dried up and he went completely out of his mind. - Miguel de Cervantes Saavedra Don Quijote The theory and experiments in the previous chapters show the improvements obtained by optimizing the modeling of different methodologies using a discriminative approach. Firstly, the AUC-based optimization is competitive with conventional training methods in the worst case (usually for a clean scenario) and can outperform them greatly under other circumstances (noise conditions). The first sign of this improvement is the decrease of the EER, minDCF results and the increase of the actual AUC of the ROC curve. Moreover, we support our initial hypothesis by providing examples where the AUC optimization actually improves the performance at all operating points. Presumably, if one were to consider the combined results at all operating points, the benefits of AUC-based learning would be even greater. Another lesson to be derived from our experiments is the direct optimization of an objective function that directly relates to the actual task being performed. This is not an epiphany – this fact has always been known; all we provide is some confirmation. Besides, if the only objective is the performance at a specific operating point (as specified by the ratio of FA and FR), then the AUC objective function could be modified to consider only that operating point. It is to be expected that the performance at that operating point will improve at the cost of performance at other operating points. The AUC optimization with noise has provided another research branch The results are consistently significant. Clearly, a learning paradigm that optimizes the entire ROC curve results in better performance than that obtained with conventional maximum-likelihood of discriminative training methods. It clearly reflects not only on the ROC curve but on the score distribution, the DET curve and every metric. What is interesting, however, is that the improvements obtained from AUC optimization are actually significantly greater on noisy speech than on clean speech. A reason is that AUC optimization naturally accounts for any shift in the data away from the intended operating point where performance is measured. More curiously, the performance of every method (JFA, MVE and i-vector), which functions over the entire space of data, benefits significantly. However, it is important to point out the outperformance of the MVE in noise conditions. Thus, the final improvement of AUC-optimized MVE on noisy data is superior to that obtained with similarly trained JFA. Moreover the gains increase with increasing noise level. We also presented an ensemble approach to partition the signal space into specific attributes. We found that partitioning and refining the space by speaker or environment both improve performance over the baseline significantly. An interesting finding was to observe how the fusion and combination of the different scoring schemes outperform the baseline. Moreover partitioning
110 Chapter 10. Conclusions and Future Work the space not only benefits a single methodology, but once again MVE and FA approaches, gain improvement when the models become specific and refined. Even more interesting is the research under noise conditions. We found that having some prior knowledge of the noise (SNR in our case) improves the current results, and that performing a blind estimation of the attribute gives competitive results. The integration of both algorithms was first thought for a robust system that can benefit from both approaches. However, the scheme can be applied to a broader range of possibilities, where not just noise is the most important factor. We also showed that by establishing a hierarchically structure, where each level acts only on a single problem improves the performance. The blind solution to obtain this structure is still on research, but if some assumptions are made, like the type of noise and the amount of channels, the challenge becomes feasible. The main drawback of this approach is the amount of data needed to practically optimize the models correctly. However, with the development of technology this is not longer a problem at the moment. We extended our research not just to the known NIST databases, but to other databases that present their own challenges. We found that our proposed approaches improved the results consistently for each one of them ( clean and noise condition), as in previous findings. Besides, it was interesting to consider databases in different languages, such as Ahumada and Gaudi and obtained competitive results. The other small databases such as YOHO and MOBIO gave us confidence on the model refinement, but also supported that the method was not only subjected to the NIST evaluations. The general models obtained from baseline databases can be adapted to new sets of data. Both of them, showed improvements for every case. The current state of the art systems were outperformed by our proposed discriminative techniques. Although discriminative training is computing expensive, we found that by discarding part of the samples that does nor provide valuable information to the models alleviate the problem. Under controlled conditions, we showed that MVE and our discriminative optimization approach performed better than the ML approaches, specifically for systems that employ less than 512 GMM models. FA-based algorithms depend on the amount of data available to perform a suitable parameter estimation. Our approaches based on discriminative training mitigate this issue, when included as the last stage of the optimization procedure. The same occurs for noise condition, where the discriminative optimization given by either the ensemble or the AUC under the ROC curve, have a better estimation of the noise models, specially for low SNRs. 10.1 Future Research We presented different approaches of discriminative model optimization. Although each of them present improvements for every shown scenario, there are still challenges that may be addressed such as: different noise types at different SNR, real channel mismatch, training with just a few seconds of speech as well as trials with just a few samples. As noted earlier, AUC-optimization can also be employed for other learning and modeling techniques, e.g. i-vector based representations [14] and PLDA [119, 74]. Generalizations can also consider multi-class classification results such as those employed for speaker identification. These too are current areas of research. For the AUC optimization would be also interesting to select a region of interest of the curve, depending on the application, and optimize accordingly. For future research we can think of optimizing the DCF directly and calibrate according to this new scheme. As part of the future research, we will also investigate other objective functions that assigns weights to the ROC/DET curve so that we can control not just the area under the curve, but the curve at each operating point. For the ensemble approach, we found that the best results were obtained when a priori knowledge of the factor in the partitioning is available. We conjecture that much of this benefit may be obtained if this information is correctly estimated in the training stage. We will investigate
10.1 Future Research 111 this in future work. We also propose to investigate methods for identifying partitions when multiple factors must be considered concurrently and particularly when they must be estimated. We will also investigate methods for formally optimizing the partitions for verification. The integration approach was investigated for the particular cases of clean speech and babble noise. As well as the other techniques, it can also be used for other type of noise types and channels. It is also interesting to find suitable techniques that can provide the appropriate hierarchical structure for each application. Finally, our findings can also be extended and applied to other fields as speaker identification or language recognition.
112 Chapter 10. Conclusions and Future Work
Appendix A Traditional approaches preliminaries Nothing is absolute. Everything changes, everything moves, everything revolves, everything flies and goes away Frida Kahlo This Appendix provide the readers with some details some of the classic algorithms used along the thesis. A.1 Expectation Maximization EM is a machine learning tool that works for point estimation. Then, for a set Xtwhere t= 1...T and a hidden (latent) variable Z, we can estimate the parameters Λ. Let us consider define the following, `(Λ) = log p(x|Λ) = log X z p(x, z|Λ) = log X z Q(z|x, Λ) p(x, z|Λ) Q(z|x, Λ) >=X z Q(z|x, Λ) log p(x, z|Λ) Q(z|x, Λ) ≡F(Q, Λ),(A.1) where Q(z|x, Λ) is the density of z. Knowing the inequality, the task is to find the lower bound of F(Q, Λ). Hence, E−step Qt+1 = arg max Qp(Q, Λt) (A.2) M−step Λt+1 = arg max Λp(Qt+1,Λ) (A.3) Let’s consider Qt+1 =p(z|x, Λt), then the distribution Qwill just depend on Λ. Under this argument, `(Λ) >=X z Q(z|x, Λ) log p(x, z|Λ) −X z Q(z|x, Λ) log Q(z|x, Λ) (A.4) >=Q(Λ|Λt) + S(Q) (A.5)
114 Chapter A. Traditional approaches preliminaries Finally, we can obtain suitable expressions for the expectation and maximization steps. E−step Q(Λ|Λt) = E(log p(x, z|Λ) (A.6) M−step Λt+1 = arg max ΛElog p(x, z|Λ) (A.7) A.2 Support Vector Machine The Support Vector Machine (SVM) is widely used in pattern recognition. In this research it is employed as a classifier that can transform feature data into a binary key. SVM was first developed by Vapnik and Chervonenkis [144]. Although it has been used for several applications, it has also been employed in biometrics [145]. Given the observation inputs and a function-based model, the goal of the basic SVM is to classify these inputs into one of two classes. Afterwards, the following set of pairs are defined {ρi, yi}; where ρi∈Rnare the training vectors and yi={−1,1}are the labels. The method relies on a linear separation of the data previously mapped in a higher dimension space H, by using φ:Rn→H;φ→φ(ρ). For generalization, let the margin between the separator hyiperplane be, {h∈H|hw,hiH+$0= 0}(A.8) and the data φ(ρ) is maximized. Moreover, The SVM learning algorithm finds an hyperplane (w, b) such that, min ρi,b,ξ 1 2w>w+C l X i=1 ξi(A.9) subject to yi(w>φ(ρi) + b)≥1−ρi(A.10) ξi≥0 where ξiis a slack variable and Cis a positive real constant known as a tradeoff parameter between error and margin. Equations A.9 and A.10 can be transformed into a dual problem represented by the Lagrange multipliers αi. l X i=1 αi−1 2 l X i=1,m=1 αiαmyiymhφ(ρi), φ(ρm)i(A.11) l X i=1 αiyi= 0, C ≥αi≥0.(A.12) αican be solved as a quadratic programming (QP) problem. The resulting values αi∈Rhave a close relation with the training points ρi. To extend the linear method to a nonlinear technique, the input data is mapped into a higher dimensional space by function φ. However, exact specification of φis not needed; instead, the expression known as kernel K(ρi, ρj)≡φ(ρi)>φ(ρm) is defined. There are different types of kernels: linear, polynomial, radial basis function (RBF) and sigmoid, among others. For our research, we just employed the linear kernel that is defined as, K(ρi, ρj)≡φ(ρi)>φ(ρm) (A.13) which is equivalent to learn the weights discriminatively.
B.3 i-vectors 121 Scoring •Compute cosine distance scoring (CDS) on the updated i-vectors (after channel compensation).
122 Chapter B. Review of Factor Analysis
Appendix C Notes about MVE Nothing is absolute. Everything changes, everything moves, everything revolves, everything flies and goes away Frida Kahlo This Appendix provides the readers with some details of Minimum Verification Approach and its extensions. It aids C.1 Minimum Verification Error To update the model parameters for a speaker Sa GMM defined as, Λ = {wC k, µC k,ΣC k;∀k},where Cis the class (target, Sor impostor speaker, ¯ S) and kis the GMM component. Let us consider a dimension optimization, where µC k= [µC kz]D z=1 ,σC k= [σC kz]D z=1 and PM k=1 wC k= 1 . Then, for µC kz we perform the following, ˜µC kz =µC kz σC kz (C.1) For simplicity, we dropped the subindex zconsidering the optimization of just one dimension (keep in mind that the optimization is performed dimension by dimension), ˜µC k(t+ 1) = ˜µC k(t)−∂`C(χ; Λ) ∂˜µC k|Λ=Λt,(C.2) where ∂`C(χ; Λ) ∂˜µC k =∂`C(χ; Λ) ∂dC ∂dC ∂˜µC k ,(C.3) ∂`C(χ; Λ) ∂dC =γ`C(dC)(1 −`C(dC)),(C.4) ∂dC ∂˜µC k =−∂gn(χ; Λ) ∂˜µC k +∂Gn(χ; Λ) ∂˜µC k χ∈Cn(C.5) ∂gC(χ; Λ) ∂µC k =1 p(χ; ΛC) ∂ p(χ; ΛC) ∂µC k ,(C.6) ∂GC(χ; Λ) ∂µC k =1 p(χ;¯ ΛC) ∂ p(χ;¯ ΛC) ∂µC k .(C.7)
124 Chapter C. Notes about MVE Consider for a specific k, dimension d, the speaker Cand to avoid bias, let µC k=σ˜ µC k. To compute the mean, ∂p(χ; Λ) ∂˜µk =wkN(χ|˜µk, σk) Pk0wk0N(χ|˜µk0, σk0)χ σk−˜µk,(C.8) The computation for ΣC kand wC kadopt the same procedure. Let ˜σC k= log(σC k), ˜σC k(t+ 1) = ˜σC k(t)−∂`C(χ; Λ) ∂˜σC k|Λ=Λt,(C.9) where ∂`C(χ; Λ) ∂˜σC k =∂`C(χ; Λ) ∂dC ∂dC ∂˜σC k ,(C.10) ∂`C(χ; Λ) ∂dC =γ`C(dC)(1 −`C(dC)),(C.11) ∂dC ∂˜σC k =−∂gn(χ; Λ) ∂˜σC k +∂Gn(χ; Λ) ∂˜σC k χ∈Cn(C.12) ∂gC(χ; Λ) ∂˜σC k =1 p(χ; ΛC) ∂ p(χ; ΛC) ∂˜σC k ,(C.13) ∂GC(χ; Λ) ∂˜σC k =1 p(χ;¯ ΛC) ∂ p(χ;¯ ΛC) ∂˜σC k .(C.14) The final solution to compute σis ∂p(χ; Λ) ∂σk =wkN(χ|µk, σk) Pk0wk0N(χ|µk0, σk0)χ−µk ˜σk−1.(C.15) For the weights, w, we proceed as follows, ˜wC k(t+ 1) = ˜wC k(t)−∂`C(χ; Λ) ∂˜wC k|Λ=Λt,(C.16) where ∂`C(χ; Λ) ∂˜wC k =∂`C(χ; Λ) ∂dC ∂dC ∂˜wC k ,(C.17) ∂`C(χ; Λ) ∂dC =γ`C(dC)(1 −`C(dC)),(C.18) ∂dC ∂˜wC k =−∂gn(χ; Λ) ∂˜wC k +∂Gn(χ; Λ) ∂˜wC k χ∈Cn(C.19) ∂gC(χ; Λ) ∂˜wC k =1 p(χ; ΛC) ∂ p(χ; ΛC) ∂˜wC k ,(C.20) ∂GC(χ; Λ) ∂˜wC k =1 p(χ;¯ ΛC) ∂ p(χ;¯ ΛC) ∂˜wC k .(C.21) Finally, for the set of weights wkand to ensure that PK k=1 wk= 1, let, w=exp( ˜w) PK k=1 exp( ˜w). ∂p(χ; Λ) ∂˜wk =N(χ|µk, σk) Pk0wk0N(χ|µk0, σk0).(C.22)
C.2 AUC optimization using FA 125 C.2 AUC optimization using FA The optimization of the parameters for FA has two stages (as in the usual JFA) to decouple estimation of V, U, Ψ. 1. Start with Vmaintaining Ufixed, considering that Q=V V >+ Ψ, where Ψ = DD> and is diagonal. The computation of µ, for a specific k, dimension dand to avoid bias, is straightforward using, µk=qk˜µk; hence, ∇φΥ(χ, ˜µk) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R) |XS||X¯ S| wkN(χ|µk, qk) Pk0wk0N(χ|µk0, qk0)χ qk −˜µk.(C.23) 2. To estimate Q, we formulate first for Vand then Ψ. We use the following identities to solve for Q. Let, ∂(log |qk|) ∂qi,j k = (qk i.j )−1(C.24) ∂(qk o.p)−1 ∂qi,j k = (qk o.i)−1(qk p.j )−1(C.25) We will also make the matrix dimensions explicit, assuming that Ψ is diagonal. Then, qk, let ˜q= log(q). Then, for V, ∇φΥ(χ, ˜vij ) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R) |XS||X¯ S|(C.26) wkN(χ|µk, qk) Pk0wk0N(χ|µk0, qk0)n[Q−1(x−u)]i[(x−u)>Q−1V]j−[Q−1V]ij o.(C.27) 3. for ψ, ∇φΥ(χ, ˜ ψi,j ) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R) |XS||X¯ S| wkN(χ|µk, qk) Pk0wk0N(χ|µk0, qk0)1 2Q−1 ii −[Q−1x−µ]2 i.(C.28) Now, for U, we fix the current Qand estimate accordingly. ∇φΥ(χ, ˜ Uk) = −X S∈S Pχ∈XSPχ∈X¯ SγR(1 −R) |XS||X¯ S| wkN(χ|µk, qk) Pk0wk0N(χ|µk0, qk0)(χ−µk ˜ Uk2 −1).(C.29) Finally, for the set of weights wkand to ensure that PK k=1 wk= 1,let, w=exp( ˜w) PK k=1 exp( ˜w)
126 Chapter C. Notes about MVE
Bibliography [1] Y. Lei, L. Burget, and N. Scheffer, “A noise robust i-vector extractor using vector taylor series for speaker recognition,” in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on. IEEE, 2013, pp. 6788–6791. [2] T. Hasan and J. Hansen, “Integrated feature normalization and enhancement for robust speaker recognition using acoustic factor analysis,” in Proc. Interspeech, 2012, pp. 1568– 1571. [3] C. Hanil¸ci, T. Kinnunen, R. Saeidi, J. Pohjalainen, P. Alku, F. Ertas, J. Sandberg, and M. Hansson-Sandsten, “Comparing spectrum estimators in speaker verification under additive noise degradation,” in Acoustics, Speech and Signal Processing (ICASSP), 2012 IEEE International Conference on. IEEE, 2012, pp. 4769–4772. [4] T. Weigold, T. Kramp, and M. Baentsch, “Remote Client Authentication,” IEEE Security & Privacy, vol. 6, no. 4, pp. 36–43, 2008. [5] L. O’Gorman, “Comparing passwords, tokens, and biometrics for user authentication,” Proceedings of the IEEE, vol. 91, no. 12, pp. 2021–2040, 2003. [6] A. Jain, A. Ross, and S. Pankanti, “Biometrics: a tool for information security,” IEEE transactions on information forensics and security, vol. 1, no. 2, pp. 125–143, 2006. [7] J. P. Campbell, W. Shen, W. M. Campbell, R. Schwartz, J.-F. Bonastre, and D. Matrouf, “Forensic speaker recognition,” Signal processing magazine, IEEE, vol. 26, no. 2, pp. 95–103, 2009. [8] P. J. Barger and S. Sridharan, “On the performance and use of speaker recognition systems for surveillance,” in Video and Signal Based Surveillance, 2006. AVSS’06. IEEE International Conference on. IEEE, 2006, pp. 109–109. [9] T. Herbig, F. Gerl, and W. Minker, “Simultaneous speech recognition and speaker identification,” in Spoken Language Technology Workshop (SLT), 2010 IEEE. IEEE, 2010, pp. 218–222. [10] Q. Li, B. Juang, C. Lee, Q. Zhou, and F. Soong, “Recent advancements in automatic speaker authentication,” IEEE Robotics & Automation Magazine, vol. 6, no. 1, pp. 24–34, 1999. [11] J. Campbell et al., “Speaker recognition: A tutorial,” Proceedings of the IEEE, vol. 85, no. 9, pp. 1437–1462, 1997. [12] D. Petrovska-Delacr´etaz, A. El Hannani, and G. Chollet, “Text-independent speaker verification: state of the art and challenges,” Progress in nonlinear speech processing, pp. 135–169, 2007.
128 BIBLIOGRAPHY [13] F. Bimbot, J.-F. Bonastre, C. Fredouille, G. Gravier, I. Magrin-Chagnolleau, S. Meignier, T. Merlin, J. Ortega-Garc´ıa, D. Petrovska-Delacr´etaz, and D. A. Reynolds, “A tutorial on text-independent speaker verification,” EURASIP journal on applied signal processing, vol. 2004, pp. 430–451, 2004. [14] N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” Audio, Speech, and Language Processing, IEEE Transactions on, vol. 19, no. 4, pp. 788–798, 2011. [15] P. Kenny, G. Boulianne, P. Ouellet, and P. Dumouchel, “Joint factor analysis versus eigenchannels in speaker recognition,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 15, no. 4, pp. 1435–1447, 2007. [16] W. Chou, “Discriminant-function-based minimum recognition error rate pattern-recognition approach to speech recognition,” Proceedings of the IEEE, vol. 88, no. 8, pp. 1201–1223, 1992. [17] B.-H. Juang and S. Katagiri, “Discriminative learning for minimum error training,” assp, pp. 1639–1641, Dec. 1992. [18] ——, “Discriminative learning for minimum error classification [pattern recognition],” Signal Processing, IEEE Transactions on, vol. 40, no. 12, pp. 3043–3054, 1992. [19] D. Povey, P. Woodland, and M. Gales, “Discriminative map for acoustic model adaptation,” in IEEE Intl. Conf. on Acoustics, Speech, and Sig. Proc. (ICASSP), vol. 1, 2003, pp. I–312. [20] D. Zhu, H. Li, B. Ma, and C. Lee, “Discriminative learning for optimizing detection performance in spoken language recognition,” in Acoustics, Speech and Signal Processing, 2008. ICASSP 2008. IEEE International Conference on. IEEE, 2008, pp. 4161–4164. [21] C. Ma and E. Chang, “Comparison of discriminative training methods for speaker verification,” in IEEE Intl. Conference on Acoustics, Speech, and Sig. Proc.(ICASSP), vol. 1, 2003, pp. 192–195. [22] M. G. Rahim, C. H. Lee, B. H. Juang, and W. Chou, “Discriminative utterance verification using minimum string verification error (msve) training,” Proc. ICASSP, vol. 6, pp. 3585– 3588, 1996. [23] L. Burget, O. Plchot, S. Cumani, O. Glembek, P. Matejka, and N. Brummer, “Discriminatively trained probabilistic linear discriminant analysis for speaker verification,” in IEEE Intl. Conf. on Acoustics, Speech and Sig. Proc. (ICASSP), 2011, pp. 4832–4835. [24] S. Cumani, N. Brummer, L. Burget, and P. Laface, “Fast discriminative speaker verification in the i-vector space,” in Acoustics, Speech and Signal Processing (ICASSP), 2011 IEEE International Conference on. IEEE, 2011, pp. 4852–4855. [25] B.-H. Juang, W. Chou, and C.-H. Lee, “Minimum classification error rate methods for speech recognition,” IEEE Trans. on Speech and Audio Processing, vol. 5, pp. 257–265, May 1997. [26] H. B. Mann and D. R. Whitney, “On a test of whether one of two random variables is stochastically larger than the other,” Annals of Mathematical Statistics, vol. 18:1, pp. 50–60, 1947. [27] C. Cortes and M. Mohri, “Auc optimization vs. error rate minimization,” Advances in neural information processing systems, vol. 16, no. 16, pp. 313–320, 2004.
BIBLIOGRAPHY 129 [28] L. Burget, M. Fapso, V. Hubeika, O. Glembek, M. Karafi´at, M. Kockmann, P. Matejka, P. Schwarz, and J. Cernock`y, “But system for nist 2008 speaker recognition evaluation.” in INTERSPEECH, 2009, pp. 2335–2338. [29] M. A. H. Huijbregts, “Segmentation, diarization and speech transcription: surprise data unraveled,” Ph.D. dissertation, Centre for Telematics and Information Technology University of Twente, 2009. [30] S. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 28, no. 4, p. 357, 1980. [31] B. Bogert, M. Healy, and J. Tukey, “The quefrency alanysis of time series for echoes: Cepstrum, pseudo-autocovariance, cross-cepstrum and saphe cracking,” in Proc. Symp. on Time Series Analysis, 1963, pp. 209–243. [32] M. Kockmann, L. Ferrer, L. Burget, E. Shriberg, and J. Cernocky, “Recent progress in prosodic speaker verification,” in Acoustics, Speech and Signal Processing (ICASSP), 2011 IEEE International Conference on. IEEE, 2011, pp. 4556–4559. [33] L. Ferrer, N. Scheffer, and E. Shriberg, “A comparison of approaches for modeling prosodic features in speaker recognition,” in Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on. IEEE, 2010, pp. 4414–4417. [34] W. Campbell, J. Campbell, D. Reynolds, D. Jones, and T. Leek, “High-level speaker verification with support vector machines,” in Acoustics, Speech, and Signal Processing, 2004. Proceedings.(ICASSP’04). IEEE International Conference on, vol. 1. IEEE, 2004, pp. I–73. [35] T. Kinnunen, “Spectral features for automatic text-independent speaker recognition,” Licenciate’s Thesis, University of Joensuu, 2003. [36] L. R. Rabiner and R. W. Schafer, Digital processing of speech signals. IET, 1979, vol. 19. [37] D. A. Reynolds, “Experimental evaluation of features for robust speaker identification,” Speech and Audio Processing, IEEE Transactions on, vol. 2, no. 4, pp. 639–643, 1994. [38] J. I. Makhoul, “Linear prediction: A tutorial review,” Proceedings of the IEEE, vol. 63, pp. 561–580, Apr. 1975. [39] H. Hermansky, “Perceptual linear predictive (PLP) analysis of speech,” Journal of the Acoustical Society of America, vol. 87, no. 4, pp. 1738–1752, Apr. 1990. [40] A. F. Martin and C. S. Greenberg, “Nist 2008 speaker recognition evaluation: performance across telephone and room microphone channels.” in INTERSPEECH, 2009, pp. 2579–2582. [41] S. Furui, “Cepstral analysis techniques for automatic speaker verification,” IEEE Trans. on Acoustics, Speech, and Signal Processing, vol. 27, pp. 254–272, Apr. 1979. [42] O. Viikki and K. Laurila, “Cepstral domain segmental feature vector normalization for noise robust speech recognition,” Speech Communication, vol. 25, no. 1-3, pp. 133–147, 1998. [43] H. Hermansky, N. Morgan, A. Bayya, and P. Kohn, “RASTA-PLP speech analysis technique,” in Proc. ICASSP, San Francisco, CA, Mar. 1992, pp. I:121–124.
130 BIBLIOGRAPHY [44] J. Pelecanos and S. Sridharan, “Feature warping for robust speaker verification,” in 2001: A Speaker Odyssey-The Speaker Recognition Workshop, 2001, pp. 213–218. [45] S. Chen and R. Gopinath, “Gaussianization,” Advances in neural information processing systems, pp. 423–429, 2001. [46] C. Myers, L. Rabiner, and A. Rosenberg, “Performance tradeoffs in dynamic time warping algorithms for isolated word recognition,” Acoustics, Speech and Signal Processing, IEEE Transactions on, vol. 28, no. 6, pp. 623–635, 1980. [47] F. Soong, A. Rosenberg, L. Rabiner, and B. Juang, “A vector quantization approach to speaker recognition,” in Acoustics, Speech, and Signal Processing, IEEE International Conference on ICASSP’85., vol. 10. IEEE, 1985, pp. 387–390. [48] J.-M. Marin, K. Mengersen, and C. P. Robert, “Bayesian modelling and inference on mixtures of distributions,” Handbook of statistics, vol. 25, pp. 459–507, 2005. [49] D. Reynolds, “A gaussian mixture modeling approach to text-independent speaker identification,” Ph.D. dissertation, Georgia Institute of Technology, 1992. [50] J. P. Campbell Jr, “Testing with the yoho cd-rom voice verification corpus,” in Acoustics, Speech, and Signal Processing, 1995. ICASSP-95., 1995 International Conference on, vol. 1. IEEE, 1995, pp. 341–344. [51] R. O. Duda, P. E. Hart et al.,Pattern classification and scene analysis. Wiley New York, 1973, vol. 3. [52] A. Wald, Sequential analysis. New York, NY: John Wiley and Sons, 1947. [53] E. Lehmann and J. Romano, Testing statistical hypotheses. Springer Verlag, 2005. [54] C. M. Bishop and N. M. Nasrabadi, Pattern recognition and machine learning. springer New York, 2006, vol. 1. [55] L. Wasserman, All of statistics: a concise course in statistical inference. Springer, 2004. [56] Y. Zigel and A. Cohen, “On cohort selection for speaker verification.” in INTERSPEECH, 2003, pp. 2977–2980. [57] T. Kinnunen, E. Karpov, and P. Fr¨anti, “Efficient online cohort selection method for speaker verification.” in INTERSPEECH, 2004, pp. 2401–2404. [58] L. Burget, O. Plchot, S. Cumani, O. Glembek, P. Matejka, and N. Brummer, “Discriminatively trained probabilistic linear discriminant analysis for speaker verification,” in Acoustics, Speech and Signal Processing (ICASSP), 2011 IEEE International Conference on. IEEE, 2011, pp. 4832–4835. [59] W.-Q. Zhang, Y. Shan, and J. Liu, “Multiple background models for speaker verification,” in Proc. of Odyssey speaker and language recognition workshop, 2010, pp. 47–51. [60] A. Sarkar and S. Umesh, “Multiple background models for speaker verification using the concept of vocal tract length and mllr super-vector,” International Journal of Speech Technology, vol. 15, no. 3, pp. 351–364, 2012. [61] P. Kenny, P. Oueleet, N. Dehak, V. Gupta, and P. Dumouchel, “A study of inter-speaker variability in speaker verification,” IEEE Trans. ASLP, vol. 16, pp. 980–988, 2008.