Full text
Ph.D. Thesis Discriminative Features for GMM and i-Vector based Speaker Diarization Author: Abraham Woubie Zewoudie Advisors: Prof. Francisco Javier Hernando Pericas Dr. Jordi Luque Serrano TALP Research Center, Speech Processing Group Department of Signal Theory and Communications Universitat Polit`ecnica de Catalunya, Barcelona, Spain June, 2017
To my family and wife iii
Abstract Speaker diarization has received several research attentions over the last decade. Among the different domains of speaker diarization, diarization in meeting domain is the most challenging one. It usually contains spontaneous speech and is, for example, susceptible to reverberation. The appropriate selection of speech features is one of the factors that affect the performance of speaker diarization systems. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used short-term speech features in speaker diarization. Other factors that affect the performance of speaker diarization systems are the techniques employed to perform both speaker segmentation and speaker clustering. In this thesis, we have proposed the use of jitter and shimmer long-term voice-quality features both for GMM and i-vector based speaker diarization systems. The voicequality features are used together with the state-of-the-art short-term cepstral and longterm speech ones. The long-term features consists of prosody and Glottal-to-Noise excitation ratio (GNE) descriptors. Firstly, the voice-quality, prosodic and GNE features are stacked in the same feature vector. Then, they are fused with cepstral coefficients at the score likelihood level both for the proposed Gaussian Mixture Modeling (GMM) and i-vector based speaker diarization systems. For the proposed GMM based speaker diarization system, independent HMM models are estimated from each set of features. In speaker segmentation, the fusion of the short-term descriptors with the long-term ones is carried out by linearly weighting the log-likelihood scores of Viterbi decoding. In speaker clustering, the fusion of the shortterm cepstral features with the long-term ones is carried out by linearly fusing the BIC scores corresponding to these feature sets. For the proposed i-vector based speaker diarization system, the feature fusion is performed exactly the same as the one in the previously mentioned GMM based system. But, the speaker clustering technique is based on the recently introduced factor analysis paradigm. Two sets of i-vectors are extracted from the speaker segmentation hypothesis. v
vi Whilst the first i-vector is extracted from short-term spectral features, the second one is extracted from the stacked voice quality, prosodic and GNE descriptors. Then, the cosine-distance and Probabilistic Linear Discriminant Analysis (PLDA) scores among ivectors are linearly weighted to obtain a unique similarity score. Finally, the final fused score is used as speaker clustering distance. We have also proposed the use of delta dynamic features for speaker clustering. The motivation for using deltas in clustering is because they capture the transitional characteristics of the speech about speaker specific information. The proposed speaker diarization system uses both the static and delta dynamic features for speaker clustering. The speaker segmentation is based only on the static MFCC feature set. The experiments have been carried out on Augmented Multi-party Interaction (AMI) meeting corpus. The experimental results show that the use of voice-quality, prosodic, GNE and delta dynamic features improve the performance of both GMM and i-vector based speaker diarization systems.
Resumen La diarizaci´on del altavoz ha recibido varias atenciones de investigaci´on durante la ´ultima d´ecada. Entre los diferentes dominios de la diarizaci´on del hablante, la diarizaci´on en el dominio del encuentro es la m´as dif´ıcil. Normalmente contiene habla espont´anea y, por ejemplo, es susceptible de reverberaci´on. La selecci´on apropiada de las caracter´ısticas del habla es uno de los factores que afectan el rendimiento de los sistemas de diarizaci´on de los altavoces. Los Coeficientes Cepstral de Frecuencia Mel (MFCC) son las caracter´ısticas de habla de corto plazo m´as utilizadas en la diarizaci´on de los altavoces. Otros factores que afectan el rendimiento de los sistemas de diarizaci´on del altavoz son las t´ecnicas empleadas para realizar tanto la segmentaci´on del altavoz como el agrupamiento de altavoces. En esta tesis, hemos propuesto el uso de jitter y shimmer caracter´ısticas de calidad de voz a largo plazo tanto para GMM y i-vector basada en sistemas de diarizaci´on de altavoces. Las caracter´ısticas de calidad de voz se utilizan junto con el estado de la t´ecnica a corto plazo cepstral y de larga duraci´on de habla. Las caracter´ısticas a largo plazo consisten en la prosodia y los descriptores de relaci´on de excitaci´on Glottal-a-Ruido (GNE). En primer lugar, las caracter´ısticas de calidad de voz, pros´odica y GNE se apilan en el mismo vector de caracter´ısticas. A continuaci´on, se fusionan con coeficientes cepstrales en el nivel de verosimilitud de puntajes tanto para los sistemas de diarizaci´on de altavoces basados en el modelo Gaussian Mixture Modeling (GMM) como en los sistemas basados en i. Para el sistema de diarizaci´on de altavoces basado en GMM propuesto, se calculan modelos HMM independientes a partir de cada conjunto de caracter´ısticas. En la segmentaci´on de los altavoces, la fusi´on de los descriptores a corto plazo con los de largo plazo se lleva a cabo mediante la ponderaci´on lineal de las puntuaciones log-probabilidad de decodificaci´on Viterbi. En la agrupaci´on de altavoces, la fusi´on de las caracter´ısticas cepstrales a corto plazo con las de largo plazo se lleva a cabo mediante la fusi´on lineal de las puntuaciones BIC correspondientes a estos conjuntos de caracter´ısticas. vii
viii Para el sistema de diarizaci´on de altavoces basado en un vector i, la fusi´on de caracter´ısticas se realiza exactamente igual a la del sistema basado en GMM antes mencionado. Sin embargo, la t´ecnica de agrupaci´on de altavoces se basa en el paradigma de an´alisis de factores recientemente introducido. Dos conjuntos de i-vectores se extraen de la hip´otesis de segmentaci´on de altavoz. Mientras que el primer vector i se extrae de caracter´ısticas espectrales a corto plazo, el segundo se extrae de los descriptores de calidad de voz apilados, pros´odicos y GNE. A continuaci´on, las puntuaciones de coseno-distancia y Probabilistic Linear Discriminant Analysis (PLDA) entre i-vectores se ponderan linealmente para obtener una puntuaci´on de similitud ´unica. Finalmente, la puntuaci´on final fusionada se utiliza como distancia de agrupaci´on de altavoces. Tambi´en hemos propuesto el uso de caracter´ısticas din´amicas delta para el agrupamiento de altavoces. La motivaci´on para usar deltas en la agrupaci´on es porque capturan las caracter´ısticas de transici´on del discurso sobre la informaci´on espec´ıfica del hablante. El sistema de diarizaci´on de altavoces propuesto utiliza tanto las caracter´ısticas din´amicas est´aticas como delta para la agrupaci´on de altavoces. La segmentaci´on del altavoz se basa ´unicamente en el conjunto de funciones MFCC est´aticas. Los experimentos se han llevado a cabo en el corpus de reuni´on de interacci´on multipartito aumentada (AMI). Los resultados experimentales muestran que el uso de calidad vocal, pros´odica, GNE y din´amicas delta mejoran el rendimiento de los sistemas de diarizaci´on de altavoces basados en GMM e i-vector.
Acknowledgements First and above all, I praise the Almighty God for providing me this opportunity and capability to complete my PhD successfully. Then, I would like to thank my thesis advisors Javier Hernando and Jordi Luque for giving me the opportunity to do my Ph.D. at the Speech Processing Research Group of The Center for Language and Speech Technologies and Applications (TALP) research center in UPC BarcelonaTech. Working with them was a great experience, and I am grateful for their time, patience, dedication and passion until the completion of my PhD. Doing Ph.D. at UPC gave me a chance to meet wonderful people who have made this journey memorable. I would like to express my gratitude to Catalan Government and UPC for funding my Ph.D study. Many thanks also to Climent Nadeu, Antonio Bonafonte, Luis Torres, Miquel and Umair Khan for their friendship and camaraderie over the past four years. Last but not the least, I would like to thank my parents Woubie Zewoudie and Azeb Manaye for their unconditional love and care. I would not have made it this far without their support. Finally, I would like to thank my wife Etsub for standing beside me throughout my four years in UPC. She has been my inspiration and motivation. She is my rock, and I dedicate this PhD dissertation to her. ix
Chapter 1 Introduction Speech technologies have been applied in automatic searching, indexing and retrieval of audio information by extracting meta-data from an audio signal. An audio segment normally consists of different speakers, music segments, noises, etc. Speaker diarization is the process of segmenting and clustering a speech recording into homogeneous regions and answers the question “Who spoke when” without any prior knowledge about the speakers [Tranter and Reynolds, 2006]. Speaker diarization needs to first classify the speech and non-speech parts of the audio signal. Then, it marks the speaker changes in the detected speech and clusters speech segments which belong to different speakers [Meignier et al., 2006]. Speaker diarization has received much attention recently [Anguera et al., 2012], and is used in automatic speech recognition, rich transcription, audio indexing and retrieval, audio archival and monitoring, speaker counting, etc. There are three major domains for speaker diarization [Reynolds and Torres-Carrasquillo, 2004]. These are broadcast news, meetings and conversational telephone speech. The broadcast news include radio and television programs over a single channel. The meeting domain includes public gatherings or lectures in which people interact in the same room. The meeting domain recordings are normally held with one or several microphones. If there is only one microphone in the meeting room, the input format is called single distant microphone (SDM). If there are more than one microphones in different locations of the meeting room, it is called multiple distant microphones (MDM). Finally, the conversational telephone speech is a telephone conversation of two more more people over a single channel. Speaker diarization systems use mostly the static Mel-frequency Cepstral Coefficients (MFCC) as acoustic signal representations. They commonly model the MFCC features 1
2Chapter 1. Introduction distribution using Gaussian Mixture Modeling (GMM) and apply Bayesian Information Criterion (BIC) methods for both speaker segmentation and clustering. The main focus of this thesis is improving the performance of the baseline HMM/GMM based speaker diarization system which is exclusively based on MFCC feature set and GMM modeling technique. Hence, this thesis has proposed the use of jitter and shimmer voice-quality features with the other long-term and short-term speech features for GMM and i-vector based speaker diarization systems. The GMM modeling technique is replaced with the recent developments in the field of speaker recognition (i.e., i-vectors). The clustering techniques are based on i-vector based cosine distance and Probabilistic Linear Discriminant Analysis (PLDA) distance metrics. This thesis has also proposed the use of dynamic delta features for speaker clustering. 1.1 Motivation and Objectives As it is mentioned in Section 1, the three major domains for speaker diarization are broadcast news, meetings and conversational telephone speech. Diarization of meeting rooms is the most challenging one since it normally contains spontaneous speech of multiple speakers with short-speaker turns. Hence, most of the recent speaker diarization researches have been on the meeting room conversations. As it is reported in [Huijbregts et al., 2012], one of the main problems in speaker diarization is the high Diarization Error Rate (DER) variation among different shows using the same speaker diarization system. One of the factors that critically affect the performance of speaker diarization approaches is the extraction of relevant speaker features. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used shortterm speech features in speaker diarization [Anguera et al., 2012]. Despite its broadly use in speech processing applications, it is reported in [Friedland et al., 2009,Zelen´ak and Hernando, 2011] that the fusion of short-term features with long-term ones provides better results for speaker diarization. This is due to the long-term feature’s provision of complimentary information to the short-term ones. The techniques used for speaker segmentation and speaker clustering have also impact on the performance of speaker diarization systems. Speaker diarization systems mostly use Gaussian Mixture Modeling (GMM) based Bayesian Information Criterion (BIC) clustering technique to merge clusters within an Agglomerative Hierarchical Clustering (AHC) approach.
1.1. Motivation and Objectives 3 Since the long-term features add complimentary information to MFCC, we are motivated to show that the average and high DER variations among different shows can be reduced by using these long-term features. Hence, we have explored the use of jitter and shimmer voice-quality features for both GMM and i-vector based speaker diarization systems. The voice-quality features are fused with the state-of-the-art short-term cepstral and longterm speech features. The long-term speech features are prosody and Glottal-to-Noise excitation Ratio (GNE). The fusion of the voice-quality features with the the longterm and short-term features is carried out at the feature and score likelihood level, respectively. The objectives of this thesis can be summarized as follows: 1. The use of Dynamic Features for Speaker Clustering Mel Frequency cepstral coefficients (MFCCs) are the most widely used short-term features for speaker diarization [Anguera et al., 2012]. Most of the state of the art speaker diarization systems use only the static MFCC for diarization. The first and second order time derivatives of the instantaneous cepstral features: delta (∆) and (∆∆) features have been successfully used in different speech applications. The delta dynamic features can be used to capture the transitional characteristics of the speech signal which contains the speaker specific information. These information are not captured by the static MFCC features. The delta dynamic features have been successfully used in speaker recognition in [Furui, 1981]. The delta features have aslso been successfully used in speaker verification [Memon et al., 2009], speaker classification [Nguyen, 2010] and speech recognition [Kumar et al., 2011]. But, the delta features are not widely used in speaker diarization experiments. For example, it is reported in [Luque, 2012] that since the delta features deteriorate the diarization results, only the static MFCC features are used in speaker diarization. It is also reported in [Yella, 2015] that delta features are not used in speaker diarization systems. Speaker clustering is highly related to speaker classification and speaker verification. Hence, we propose the use of delta dynamic features only for speaker clustering. The delta features provide new information related to each frame that can not be captured with purely static features. The main contribution of this work is the use of static and delta dynamic features in speaker clustering. The speaker segmentation is based only on the static MFCC. 2. The use of Voice quality features for HMM/GMM Speaker Diarization System
4Chapter 1. Introduction Jitter and shimmer voice-quality measurements discern variations of fundamental frequency and amplitude, respectively. Studies show that these measurements can be used to detect voice pathologies [Kreiman and Gerratt, 2005], speaking styles and emotions [Li et al., 2007], and also identify age and gender [Sadeghi Naini and Homayounpour, 2006]. For example, the authors in [Farr´us et al., 2007] report that fusing jitter and shimmer voice-quality measurements with the baseline cepstral features improve the performance of speaker recognition systems. It is also described in [Li et al., 2007] that the use of jitter and shimmer measurements together with cepstral ones improves the classification accuracy of different speaking styles more than using only the baseline cepstral features. The work of [Zhang, 2008] also reports that the fusion of voice-quality with prosodic features is able to effectively discriminate different emotions in Chinese speech emotion identification. The importance of voice-quality features in emotion identification is also discussed in [Johnstone and Scherer, 1999]. It is also shown in [Kreiman and Gerratt, 2005] that these voice-quality measurements can be used to characterize voices such as breathy, tense, harsh, whispery, creaky and hoarse. Based on these studies, we have proposed the use of jitter and shimmer voicequality measurements for speaker diarization since these features add complementary information to the baseline cepstral features. Firstly, jitter and shimmer voice quality features are extracted from the fundamental frequency (F0) contours. Then, the voice-quality features are fused with the short-term cepstral and long-term prosodic features. The fusion of features is carried out at the feature and score level, respectively. The fusion of the voice-quality features with the prosodic ones is carried out at the feature level (i.e., they are stacked in the same feature vector). The prosodic features are the extracted from the evolution in time of pitch, acoustic intensity and the first four formant frequencies. Then, the stacked long-term speech features are fused with the cepstral ones at the score likelihood level both in segmentation and clustering stages. The score fusion in segmentation is based on the log-likelihood scores corresponding to the shortand long-term speech features. The score in clustering is based on Bayesian Information Criterion (BIC) combined scores of each feature set. 3. The use of Voice quality features for i-Vector based Speaker Diarization System The techniques employed for both speaker segmentation and speaker clustering factors have also impact on the performance of speaker diarization systems, in addition to the selection of appropriate speech features. Speaker diarization systems mostly use Gaussian Mixture Modeling (GMM) based Bayesian Information
1.2. Publications from the thesis 5 Criterion (BIC) clustering technique to merge clusters within an Agglomerative Hierarchical Clustering (AHC) approach. Factor analysis techniques which are the state of the art in speaker recognition have recently been successfully applied in speaker diarization experiments [Kenny et al., 2010,Franco-Pedroso et al., 2010,Shum et al., 2011,Shum et al., 2012,Vaquero Avil´es-Casco, 2011,Senoussaoui et al., 2013]. In these works, the speech clusters generated by the segmentation are first represented by i-vectors. Then, the successive clustering stages are carried out using i-vector modeling techniques. Representing the speech clusters by i-vectors enables to reduce the large-dimensional feature vector into a small dimensional one by retaining most of the relevant information. For instance, it is reported in [Silovsky and Prazak, 2012] that modeling speech segments by i-vector and using cosine-distance clustering technique improves the performance of a diarization system more than GMM based BIC clustering technique. It is also shown in [Kenny et al., 2010,Franco-Pedroso et al., 2010,Shum et al., 2011] that i-vector based cosine-distance clustering technique has been successfully applied in speaker clustering task. Note that the above mentioned works extract i-vectors exclusively from the shortterm cepstral features for speaker clustering. The main contribution of our work is the extraction of i-vectors from the short-term cepstral, and long-term speech features. The long-term speech features are the voice-quality, prosodic and GNE features. At first, the long-term voice-quality, prosodic and GNE features are fused at the feature level (i.e., they are stacked in the same feature vector). Then, two sets of i-vectors are extracted for each segment given by the Viterbi segmentation decoding. While the first i-vector is extracted from the short-term cepstral features, the second one is extracted from the stacked long-term speech features. Finally, the cosine distance and PLDA scores of these i-vectors are fused as a distance metrics for speaker clustering. 1.2 Publications from the thesis The publications extracted from this thesis are summarized as follows: 1. Jitter and Shimmer Voice-quality Measurements for Speaker Diarization Jitter and shimmer measure fundamental frequency and amplitude variations, respectively. Previous studies have shown that these voice quality features have been successfully used in speaker recognition and emotion classification tasks. The work
6Chapter 1. Introduction in [Farr´us et al., 2007] reports that adding jitter and shimmer voice quality features to both cepstral and prosodic features improves the performance of a speaker verification system. It is also described in [Li et al., 2007] that the fusion of voice quality features together with the cepstral ones improves the classification accuracy of different speaking styles and conveys information that discriminates the different animal arousal levels. Furthermore, these voice quality features are more robust to acoustic degradation and noise channel effects [Carey et al., 1996]. Based on these studies, we propose the use of jitter and shimmer voice quality features for speaker diarization since they provide complementary information to the baseline cepstral features. The main contribution of this work is the extraction of jitter and shimmer voice quality features and their fusion with the cepstral ones in the framework of speaker diarization. The experiments have been carried out on the Augmented Multiparty Interaction (AMI) corpus . Experimental results show that incorporating jitter and shimmer measurements to the baseline cepstral features decreases the diarization error rate. The results of this work has been published in [Woubie et al., 2014]. The publication can be accessed here. 2. Using Voice-quality Measurements with Prosodic and Spectral Features for Speaker Diarization Jitter and shimmer voice-quality measurements have been successfully used to detect voice pathologies and classify different speaking styles. In this paper, we investigate the usefulness of jitter and shimmer voice measurements in the framework of the speaker diarization. The combination of jitter and shimmer voice-quality features with the long-term prosodic and short-term cepstral features is explored in a subset of the AMI corpus. The appropriate characteristics related to the human speech prosody are conveyed through intonation, rhythm and stress. Encouraged by work of [Zelen´ak and Hernando, 2011], we have extracted features related to the evolution in time of pitch, acoustic intensity and the first four formant frequencies to validate their performance in this work. Experimental results show that the best results are obtained by fusing the voice-quality features with the prosodic ones at the feature level, and then fusing them with the cepstral features at the score level. The results of this work has been published in [Woubie et al., 2015]. The publication can be accessed here. 3. Shortand Long-Term Speech Features for Hybrid HMM-i-Vector based Speaker Diarization System Recently, i-vector modeling techniques have been successfully used for speaker clustering. In this work, we propose the extraction of i-vectors from shortand
1.2. Publications from the thesis 7 long-term speech features, and the fusion of their cosine scores within the frame of speaker diarization. Firstly, two sets of i-vectors are first extracted from short-term cepstral and longterm features. The long-term features are the concatenation of voice-quality and prosodic features. Once the i-vectors are extracted from the shortand long-term speech features, the cosine scores of these two i-vectors are fused as a distance metric for speaker clustering. The experiments have been carried out on AMI corpus. Experimental results show that the extraction of i-vectors from the shortand long-term speech features, and the fusion of their cosine-distance scores provide better DER result than extracting i-vectors only from short-term cepstral features. The experimental results also show that i-vector based cosine distance clustering technique provides better results than GMM based BIC clustering technique. The results of this work has been published in [Woubie et al., 2016b]. The publication can be accessed here. 4. Improving i-Vector and PLDA based Speaker Clustering with Longterm Features Recently, i-vector modeling techniques have been successfully used for speaker clustering. In this work, we propose the extraction of i-vectors from shortand long-term speech features, and the fusion of their PLDA scores within the frame of speaker diarization. Firstly, two sets of i-vectors are first extracted from short-term cepstral and longterm features. The long-term features are the concatenation of voice-quality, prosodic and Glottal-to-Noise Excitation Ratio (GNE) features. Then, the PLDA scores of these two sets of i-vectors are fused as a distance metric for speaker clustering. The main contribution to the work in [Woubie et al., 2016b] is the use of GNE feature together with the voice-quality and prosodic features. The i-vector based cosine distance clustering technique in [Woubie et al., 2016b] is also replaced by i-vector based PLDA clustering one. Experimental results on AMI corpus show that i-vector based PLDA clustering technique provides a substantial relative DER improvement more than GMM based BIC clustering one. It also provides better DER improvement more than i-vector based cosine distance clustering technique. The addition of GNE feature to the voice-quality and prosodic features also improve the DER results. The results of this work has been published in [Woubie et al., 2016a]. The publication can be accessed here.
8Chapter 1. Introduction 1.3 Organization of the thesis The thesis is organized as follows: •Chapter 2 (State-of-the-art in Speaker Diarization): This chapter provides a brief overview of the the state of the art techniques in speaker diarization. It describes the main components of speaker diarization system. It also outlines the most widely used shortand long-term speech features in speaker diarization. The different speech and non-speech detection methods is also addressed in the chapter. It also describes the most widely used speaker segmentation and clustering techniques. Finally, it provides details about the different speaker diarization systems and evaluation metrics of speaker diarization systems. •Chapter 3 (The UPC Baseline Speaker Diarization System): This chapter describes the baseline speaker diarization system. First, the front-end processing technique is outlined. Then, the speaker segmentation and clustering techniques are discussed along with the features used. Finally, the chapter provides a short summary of the baseline speaker diarization system merging and stopping criterion techniques. •Chapter 4 (Long-term Speech Features for Speaker Diarization): This chapter describes the proposed long-term features for speaker diarization. Detailed descriptions of long-term voice-quality, prosodic and GNE features is given. The techniques and methods of the extraction of these long-term features are also outlined. •Chapter 5 (Proposed Speaker Diarization Systems): This chapter discusses about the proposed speaker diarization systems. The proposed speaker diarization architectures both for the GMM and i-vector based systems are clearly described. The feature and score fusion techniques carried out in the proposed speaker diarization systems is also discussed. The different score fusion techniques in segmentation and clustering for the proposed GMM and i-vector based speaker diarization systems are also described. •Chapter 6 (Experimental Setups and Results): This chapter explains about the experimental setups and results. It discusses about the Augmented Multi-party Interaction (AMI) meeting corpus used in the thesis. The different partitions of the AMI dataset for the training, development and test sets are clearly stated. The Universal Background Model (UBM), T-Matrix and PLDA training techniques are
1.3. Organization of the thesis 9 also outlined in the chapter. The techniques of the parameter tuning, developmental and test results are finally presented. •Chapter 7 (Conclusions and Future works): This chapter summarizes the major contributions and results obtained from the PhD thesis. It also provides future research lines that can be continued from the proposed systems.
16 Chapter 2. State-of-the-art in Speaker Diarization Glottal-to-Noise Excitation Ratio Different type of acoustic parameters have been proposed to measure different perturbations in speech signal. These parameters are usually grouped into three main categories: amplitude perturbation, frequency perturbation and noise parameters. Noise parameters can be used to provide indication of the noise content of the signal and have an extensive application in the evaluation of voice quality. GNE is an acoustic measure that can be used to assess the amount of voice excitation by vocal-fold oscillations versus excitation by turbulent noise. It indicates whether a given voice signal originates from vibrations of the vocal folds or from turbulent noise generated in the vocal tract [Michaelis et al., 1997]. Thus, it is closely related to breathiness, and it is considered a reliable measure for the relative noise level, even in the presence of strong amplitude and frequency perturbations [Michaelis et al., 1997]. The computation of GNE is independent of variations of fundamental frequency and amplitude [S´aenz Lech´on et al., 2009,Michaelis et al., 1998a]. Thus GNE is suited even to highly irregular glottal oscillations. It is reported in [S´aenz Lech´on et al., 2009] that GNE parameter has a significant potential to screen voices since it quantifies the amount of voice excitation and turbulent noise. It is also reported in [Godino-Llorente et al., 2010] that GNE provides reliable measurements for discrimination among normal and pathological voices more than other classical long-term noise measurements, such as Normalized Noise Energy and Harmonics to Noise Ratio. It has also been used successfully to screen voice disorders in [GodinoLlorente et al., 2010]. 2.3 Speech/Non-speech detection The speech/non-speech detection detects the speech and non-speech segments of a given audio signal. The errors made by speech/non-speech detection has impact on the performance of the speaker diarization system in two different ways. These are missed speech segments and false alarm speeches which directly contribute to the Diarization Error Rate (DER) in the form of missed speeches and false alarms, respectively. The falsealarm also create impurities in the acoustic models of speaker clusters [Wooters et al., 2004]. Therefore, the selection of appropriate Speech Activity Detection (SAD) is crucial since it affects the diarization evaluation metric (see Section 2.8 for DER calculations). There are two ways to detect the speech/non-speech parts of a signal in speaker diarization. These are using a Speech Activity Detection (SAD) and the manual references (Oracle SAD) of the reference files.
2.4. Speaker Modeling Techniques 17 The three widely used techniques for SAD are the following: energy based, model based and hybrid approaches. Energy based detectors: This technique uses a threshold on short-term energies to decide for speech/non-speech segments [Junqua et al., 1994,Lamel et al., 1981]. This technique does not need any training data and it is easy to implement. A constraint can be imposed on the the length of the silences to avoid false alarms. The energy-based method is mostly used in speech recognition. It is mostly used in telephone speech. Since different recordings have different channels, noise and recording conditions, this technique does not generalize to different recording scenarios. Model based detectors: This technique uses a labelled speech and non-speech data to pre-train models and classifies unlabelled speech data using pre-trained models [Zhu et al., 2008,Anguera et al., 2005,Fredouille and Senay, 2006]. A Gaussian mixture model is trained for each class and the detection of speech/non-speech segments is based on Viterbi decoding. A minimum duration of speech segments is normally constrained for each class to prevent decoding short-segments. The main problem of this approach is the amount of labelled data to train the models and their generalizability to new data. Hybrid approaches: It uses the the threshold energy and model based techniques discussed previously. The energy based detector is applied first to detect the speech segments. Then, these segments are used to train new models or adapt pre-trained models to the current recording scenario. The hybrid technique alleviates the problem of need of labelled training data. It can also overcome the problem of generalizability of pre-trained models [Anguera et al., 2006a,Wooters and Huijbregts, 2008]. The second method of detecting speech/non-speech regions is using the Oracle SAD. When Oracle SAD is used as SAD, the non-speech frames are marked. Therefore, the missed speech and false alarms have zero values in the DER computation. Since this thesis focuses on the impact of long-term speech features in GMM and i-vector based diarization systems, Oracle SAD has been used as it enables us to focus mainly on the speaker errors that occur due to segmentation and clustering. Hence, DER values reported in the experimental sections corresponds purely to speaker time confusion produced by the diarization system. 2.4 Speaker Modeling Techniques One of the crucial issues in speaker diarization is the techniques employed for speaker modeling. Several modeling techniques have been used in speaker recognition and
18 Chapter 2. State-of-the-art in Speaker Diarization speaker diarization tasks. The state-of-the-art speaker modeling techniques in speaker diarization are the following: 2.4.1 Gaussian Mixture Modeling A Gaussian Mixture Model (GMM) is a parametric probability density function represented as a weighted sum of Gaussian component densities. GMMs have been successfully used to model the speech features in different speech processing applications. A Gaussian mixture model is a weighted sum of M component Gaussian densities. Each of the components is a multi-variant Gaussian function. A GMM is represented by mean vectors, covariance matrices and mixture weights. λ={wi, µi,Σi}, i = 1, ......., C(2.1) The covariance matrices of a GMM, Σi, can be full rank or constrained to be diagonal. The parameters of a GMM can also be shared, or tied, among the Gaussian components. The number of GMM components and type of covariance matrices are often determined based on the amount of data available for estimating GMM parameters. In speaker recognition, a speaker can be modeled by a GMM from training data or using Maximum A Posteriori (MAP) adaptation [Reynolds, 2002]. While the speaker model is built using the training utterances of a specific speaker in the GMM training, the model is also usually adapted from a large number of speakers called Universal Background Model in MAP adaptation. Given a set of training vectors and a GMM configuration, there are several techniques available for estimating the parameters of a GMM [McLachlan and Basford, 1988]. The most popular and used method is the maximum likelihood (ML) estimation. The ML estimation finds the model parameters that maximize the likelihood of the GMM given a set of data. Assuming an independence between the training vectors X={xi, . . . , xN}, the GMM likelihood is typically described as : p(X|λ) = N Y t=1 p(xt|λ) (2.2) Since direct maximization is not possible on equation 2.2, the ML parameters are obtained iteratively using expectation-maximization (EM) algorithm [Dempster et al.,
2.4. Speaker Modeling Techniques 19 Figure 2.2: Example of speaker model adaptation. 1977]. The EM iteratively estimate new model parameters ¯ λbased on a given model λ such that p(X|¯ λ)≥p(X|λ). The parameters of a GMM can also be estimated using Maximum A Posteriori (MAP) estimation, in addition to the EM algorithm. The MAP estimation technique derives a speaker model by adapting from a universal background model (UBM). The “Expectation” step of EM and MAP are the same. MAP adapts the new sufficient statistics by combining them with old statistics from the prior mixture parameters. Given a prior model and training vectors from the desired class, X=x1..., xT, we first determine the probabilistic alignment of the training vectors into the prior mixture components. For mixture iin the prior model Pr(i|xt, λUBM ) is computed as the percentage of the mixture component ito the total likelihood, Pr(i|xt, λUBM ) = wig(xt|µi,Σi) PM i=1 wig(xt|µi,Σi)(2.3) Then, the sufficient statistics for the weight, mean and variance parameters is computed as follows: ni= T X t=1 Pr(i|xt, λprior)weight (2.4) Ei(x) = 1 ni T X t=1 Pr(i|xt, λprior)xtmean (2.5)
20 Chapter 2. State-of-the-art in Speaker Diarization Ei(x2) = 1 ni T X t=1 Pr(i|xt, λprior)x2 tvariance (2.6) Finally, the new sufficient statistics from the training data are used to update the prior sufficient statistics for mixture ito create the adapted mixture weight, mean and variance for mixture ias follows: wi= [αw ini/T + (1 −αw i)wi]γ(2.7) µi=αm iEi(x) + (1 −αm i)µi(2.8) µ2 i=αv iEi(x2) + (1 −αv i)(σ2 i+µ2 i)−µ2 i(2.9) The adaptation coefficients controlling the balance between old and new estimates are {αw i, αm i, αv i}for the weights, means and variances, respectively. The scale factor, γ, is computed over all adapted mixture weights to ensure they sum to unity. 2.4.2 i-Vector Different approaches have been developed recently to improve the performance of speaker recognition systems. The most popular ones were based on GMM-UBM. The Joint Factor Analysis (JFA) [Kenny et al., 2008] is then built on the success of the GMMUBM approach. JFA modeling defines two distinct spaces: the speaker space defined by the eigenvoice matrix and the channel space represented by the eigen-channel matrix. In [Dehak, 2009], it is proved that channel factors estimated using JFA, which are supposed to model only channel effects, also contain information about speakers. A new speaker verification system has been proposed using factor analysis as a feature extractor that defines only a single space, instead of two separate spaces [Dehak et al., 2011]. In this new space, a given speech recording is represented by a new vector, called total factors as it contains the speaker and channel variabilities simultaneously. Speaker recognition based on the i-vector framework [Dehak et al., 2011] is currently the state-of-the-art in the field. It is also reported in [Kenny et al., 2010,Franco-Pedroso et al., 2010,Shum et al., 2011,Shum et al., 2012,Vaquero Avil´es-Casco, 2011,Senoussaoui et al., 2013] that i-vector features can also be successfully used in speaker diarization experiments. It is also shown in [Silovsky and Prazak, 2012] that modeling the speech
2.4. Speaker Modeling Techniques 21 segments by i-vector and using cosine distance scoring improves the performance of a baseline speaker diarization system more that GMM based BIC clustering technique. Given an utterance, the speaker and channel dependent GMM supervector is defined as follows: M=m+Tw (2.10) where mis a speaker and channel independent supervector, Tis a rectangular matrix of low rank and wis a random vector having a standard normal distribution N(0,1). The components of the vector ware the total factors. These new vectors are called i-vectors. Mis assumed to be normally distributed with mean vector and covariance matrix TTt. The total factor is a hidden variable, which can be defined by its posterior distribution conditioned to the Baum–Welch statistics for a given utterance. This posterior distribution is a Gaussian distribution and the mean of this distribution corresponds exactly to i-vector. The Baum–Welch statistics are extracted using the UBM. Given a sequence of Lframes {y1, y2, ......, yn}and a UBM Ω composed of Cmixture components defined in some feature space of dimension F, the Baum–Welch statistics needed to estimate the i-vector for a given speech utterance uis given by : Nc= L X t=1 P(c|yt,Ω) (2.11) Fc= L X t=1 P(c|yt,Ω)yt(2.12) where c= 1, ...., C is the Gaussian index and P(c|yt,Ω) corresponds to the posterior probability of mixture component cgenerating the vector yt. The centralized first-order Baum–Welch statistics has also to be computed for the extraction of i-vectors as follows: ˆ Fc= L X t=1 P(c|yt,Ω)(yt−mc) (2.13) where mcis the mean of UBM mixture component c. The i-vector for a given utterance can be obtained using the following equation: w= (I+TtΣ−1N(u)T)−1. TtΣ−1ˆ F(u) (2.14)
22 Chapter 2. State-of-the-art in Speaker Diarization where Nuis a diagonal matrix of dimension CF ×CF whose diagonal blocks are NcI(c= 1, ......, C) . The supervector obtained by concatenating all first-order Baum–Welch statistics Fcfor a given utterance uis represented by ˆ F(u) which has CF ×1 dimension. The diagonal covariance matrix, Σ , with dimension CF ×CF estimated during factor analysis training models the residual variability not captured by the total variability matrix T. A clear and concise process of extraction of i-vectors is found in [Dehak et al., 2011]. After the extraction of raw i-vectors, normalization needs to be carried out on the raw i-vector to remove any useless information [Bousquet et al., 2011,Garcia-Romero and Espy-Wilson, 2011]. The normalization methods can be carried out at the feature or score level. One of the most widely used feature normalization techniques of i-vectors is length normalization [Bousquet et al., 2011,Garcia-Romero and Espy-Wilson, 2011]. Length normalization ensures that the distribution of i-vectors matches the Gaussian normal distribution and makes the distributions of i-vector more similar. It is also reported in [Jiang et al., 2012] that performing whitening before length normalization improves the performance of speaker verification systems. It is also reported in [Garcia-Romero and Espy-Wilson, 2011] that i-vector normalization improves the gaussianity of the ivectors. It reduces the gap between the underlying assumptions of the data and real distributions. It also reduces the dataset shift between development and test i-vectors. w←Σ−1 2(w−µ) ||Σ−1 2(w−µ)|| (2.15) where µand Σ are the mean and the covariance matrix of a training corpus, respectively. The data is standardized according to covariance matrix Σ and length-normalized (i.e., the i-vectors are confined to the hypersphere of unit radius). The two most widely and common intersession compensation techniques of i-vectors are Within-Class Covariance Normalization (WCCN) and Linear Discriminant Analysis (LDA). WCCN uses the within-class covariance matrix to normalize the cosine kernel functions in order to compensate for intersession variability [Dehak et al., 2011]. LDA attempts to define a reduced special axes that minimize the within-speaker variability caused by channel effects, and maximize between-speaker variability. It is shown in [Dehak et al., 2011,Dehak et al., 2010] that cosine kernel function is an effective classifier to categorize i-vectors. Cosine Distance
2.4. Speaker Modeling Techniques 23 Figure 2.3: The i-vector extraction process. Once the i-vectors are extracted from the outputs of speech clusters, cosine distance scoring tests the hypothesis if two i-vectors belong to the same speaker or different speakers. Given two i-vectors, the cosine distance among them is calculated as follows: cos(wi, wj) = wi.wj ||wi||.||wj|| Rθ(2.16) where θis the threshold value, and cos(wi,wj) is the cosine distance score between clusters iand j. The corresponding i-vectors extracted for clusters iand jare represented by wiand wj, respectively. The cosine distance scoring considers only the angle between two i-vectors, not their magnitude. Since the non-speaker information such as session and channel variabilities affect the i-vector magnitude, removing the magnitudes can increase the robustness of i-vector systems [Dehak et al., 2010]. Probabilistic Linear Discriminant Analysis The i-vector representation followed by probabilistic linear discriminant analysis (PLDA) modeling technique is the state-of-the-art in speaker verification systems [Prince and Elder, 2007]. In speaker diarization, each i-vector represents the speech of one speaker. Speaker diarization needs to determine if two i-vectors belong to the same or different speakers.
24 Chapter 2. State-of-the-art in Speaker Diarization PLDA has been successfully applied in speaker recognition experiments [Brummer et al., 2010]. It is also applied in [Jiang et al., 2012] to handle speaker and session variability in speaker verification task. It has also been successfully applied in speaker clustering since it can separate speaker and noise specific parts of an audio signal which is essential for speaker diarization [Prazak and Silovsky, 2011]. PLDA has also been successfully used in speaker clustering experiments and it is shown in the work of [Prazak and Silovsky, 2011] that PLDA-based clustering provides significance performance improvement than BICbased speaker clustering methods. It is also shown that PLDA scoring provides better speaker clustering performance than cosine scoring [Sell and Garcia-Romero, 2014]. Figure 2.4: Example of PLDA Model In PLDA, assuming that the training data consists of Ji-vectors where each of these i-vectors belong to speaker I, the j’th i-vector of the I’th speaker is denoted by: wij =µ+Fhi+Gwij + Σij (2.17) where µis the overall speaker and segment independent mean of the i-vectors in the training dataset, columns of the matrix F define the between-speaker variability and columns of the matrix G define the basis for the within-speaker variability subspace. Σij represents any unexplained data variation. The components of the vector hiare the eigenvoice factor loadings and components of the vector wij are the eigen-channel factor loadings. The term Fhidepends only on the identity of the speaker, not on the particular segment. Although the PLDA model assumes Gaussian behavior, there is empirical evidence that channeland speakereffects result in i-vectors that are non-Gaussian. It is reported in [Kenny, 2010] that the use of Student’s t-distribution, on the assumed Gaussian PLDA model, improves the performance. Since this normalization technique is complicated, a non-linear transformation of i-vectors called radial Gaussianization has been proposed
2.4. Speaker Modeling Techniques 25 in [Garcia-Romero and Espy-Wilson, 2011]. It whitens the i-vectors and performs length normalization. This restores the Gaussian assumptions of the PLDA model. A variant of PLDA model called Gaussian PLDA (GPLDA) is shown to provide better results in [Garcia-Romero and Espy-Wilson, 2011]. Because of its low computational requirements, and its performance, it is the most widely used PLDA modeling. In GPLDA model, the within-speaker variability is modeled by a full covariance residual term which allows us to omit the channel subspace. The generative PLDA model for the i-vector is represented by wij =µ+Fhi+ Σij (2.18) The residual term representing the within-speaker variability is assumed to have a normal distribution with full covariance matrix Σij. A special case of the simplified PLDA model where the speaker factors S is full-rank is termed as the two-covariance model in [Br¨ummer and De Villiers, 2010,Cumani et al., 2013]. Given two i-vectors w1and w2, the PLDA computes the likelihood ratio of the two i-vectors as follows: Score(w1, w2) = p(w1, w2|H1) p(w1|H2)p(w2|H2)(2.19) where the hypothesis H1indicates that both i-vectors belong to the same speaker and H0indicates they belong to two different speakers. log(w1, w2) = logN "w1 w2#;"µ µ#,"Σ + SSTSST SSTΣ + SST#!− logN(w1;µ, Σ + SST)−logN(w2;µ, Σ + SST) (2.20) After straightforward algebra, this turns out to be, log(w1, w2) = hwT 1wT 2i"Σ + SSTSST SSTΣ + SST#−1 hwT 1wT 2i−wT 1hwT 1wT 2i− wT 1hΣ + SSTi−1w1−wT 2hΣ + SSTi−1wt+C(2.21)
32 Chapter 2. State-of-the-art in Speaker Diarization Figure 2.8: Example of ∆BIC Values. •Generalized Likelihood Ratio(GLR) GLR has been proposed as a metric for change detection in [Willsky and Jones, 1976]. Given sequences of feature vectors Xiand Xjfrom two contiguous speech segments iand j, respectively, GLR is calculated as a likelihood ratio under the assumption of H1and H2. Therefore, two different speaker models are generated for H1and H2. In H1,θi,j is estimated with Xiand Xj. In H2, two models are estimated: θifrom Xi, and θj from Xj. The GLR is then computed as follows: GLR(H1 H2 ) = L(Xi,j|θi,j) L(Xi|θi)L(Xj|θj)(2.29) where L is the likelihood. A high value of GLR shows that the two speech segments are modeled by a single model and a low value of GLR shows that the two speech segments are modeled by two models. GLR can be used together with BIC in a two step speaker segmentation process [Delacourt and Wellekens, 2000]. In the first step, the most likely speaker change points are detected by GLR and, in the second step, BIC is used to refine the speaker change points. •Gish Distance It is a likelihood based metric obtained as a variation to the GLR. It is defined as:
2.6. Speaker Clustering 33 DGish(i, j) = −N 2log(|Si|α|Sj|(1−α) |αSi+ (1 −α)Sj|) (2.30) where Siand Sjare the covariance matrices of segments iand jand α=Ni Ni+Nj. •Information Change Rate (ICR) Information Change Rate (ICR): It is another distance measure that is recently introduced for speaker diarization [Vijayasenan et al., 2007,Vijayasenan et al., 2009,Han and Narayanan, 2008]. ICR can be used to delineate the similarity of two speech speech segments determining the variation in terms of information that would be obtained by merging them. ICR similarity is not based on model of speech segments. It is based on the distance between the segments in a space of relevance variables with maximum mutual information or minimum entropy. ICR is computationally efficient and more robust to data source variation more than BIC distance [Han and Narayanan, 2008]. The results of speaker segmentation may contain two types of error. The first type of error occurs when a true segment boundary is not detected (i.e., deletion). The second type of error occurs when a segmented boundary does not correspond to the true segment boundary in the reference (i.e., false alarm). 2.6 Speaker Clustering Speaker clustering groups speech segments that belong to a particular speaker. It has two major categories based on its processing requirements. Its two main categories are online and offline speaker clustering. In the former, speech segments are merged or split in consecutive iterations until the optimum number of speakers is acquired. Since the entire speech file is available before decision making in the later, it provides better results more than the online speaker clustering. The most widely used and popular technique for speaker clustering is Agglomerative Hierarchical Clustering (AHC). AHC builds a hierarchy of clusters, that shows relations between speech segments, and merges speech segments based on similarity. AHC approaches can be classified into bottom-up and top-down clustering. Two items need to be defined in both bottom-up and top-down clustering: 1. A distance between speech segments to determine acoustic similarity. The distance metric is used to decide whether or not two clusters must be merged (bottom-up clustering) or split (top-down clustering).
34 Chapter 2. State-of-the-art in Speaker Diarization 2. A stopping criterion to determine when the optimal number of clusters (speakers) is reached. •Bottom-up (Agglomerative): It starts from a large number of speech segments and merges the closest speech segments iteratively until a stopping criterion is met. This technique is the most widely used in speaker diarization since it is directly applied on the output of speech segments from speaker segmentation. A matrix of distances between every possible pair of clusters is computed and the pair with highest BIC value is merged. Then, the merged clusters are removed from the distance matrix. Finally, the distance matrix table is updated using the distances between the new merged cluster and all remaining clusters. This process is done iteratively until the stopping criterion is met or all pairs have a BIC value less than zero . The bottom-up approach has been used for many years in pattern classification in [Duda and Hart, 1973] but was first considered for speaker clustering in [Duda et al., 2001] and [Siegler et al., 1997]. A two pass speaker clustering has been proposed in [Chen and Gopalakrishnan, 1998]. In the first pass, the speech data is equally segmented using GLR distance matrix with agglomerative clustering until the desired number of speakers is reached. The second pass trains speaker models and iteratively decodes and trains speaker models until the total likelihood converges. •Top-down: Top-down Hierarchical Clustering methods start from a small number of clusters, usually a single cluster, that contains several speech segments, and the initial clusters are split iteratively until a stopping criterion is met. It is not as widely used as the bottom-up clustering. Figure 2.9: Bottom-Up and Top-down approaches to clustering
2.7. Approaches to Speaker Diarization System 35 2.7 Approaches to Speaker Diarization System This section describes some of the state-of-the-art speaker diarization systems. The HMM/GMM based system provides the the state-of-the-art in NIST-RT [National Institute of Standards and Technology, 2003] evaluation campaigns. The information bottleneck framework provides comparable results to that of HMM/GMM based system [Vijayasenan and Valente, 2012]. HMM/GMM system In HMM/GMM based speaker diarization system, each speaker is represented by a state of an HMM and the state emission probabilities are modeled using GMMs. The initial clustering is performed initially by partitioning the audio signal equally which generates a set of segments {si}. Let cirepresent ith speaker cluster, birepresent the emission probability of cluster ciand ftdenote a given feature vector at time t. Then, the log-likelihood logbi(st) of the feature ftfor cluster ciis calculated as follows: logbi(st) = log X (r) w(r) iN(fi, µ(r) i,Σ(r) i)(2.31) where N() is a Gaussian pdf and w(r) i, µ(r) i,Σ(r) iare the weights, means and covariance matrices of the rth Gaussian mixture component of cluster ci, respectively. The agglomerative hierarchical clustering starts by overestimating the number of clusters. At each iteration, the clusters that are most similar are merged based on the BIC distance. The distance measure is based on modified delta Bayesian information criterion [Ajmera and Wooters, 2003]. The modified BIC distance does not take into account the penalty term that corresponds to the number of free parameters of a multivariate Gaussian distribution and is expressed as: ∆BIC(ci, cj) = X ft∈{ci∪cj} logbij(ft)−X ft∈ci logbi(ft)−X ft∈cj logbj(ft)(2.32) where bij is the probability distribution of the combined clusters ciand cj. The clusters that produce the highest B IC score are merged at each iteration. A minimum duration of speech segments is normally constrained for each class to prevent decoding shortsegments. The number of clusters is reduced at each iteration. When the maximum ∆BIC distance among these clusters is less than threshold value 0, the speaker diarization system stops and outputs the hypothesis. Information bottleneck (IB) system
36 Chapter 2. State-of-the-art in Speaker Diarization Information Bottleneck (IB) system is a non-parametric system based on information theoretic principles. Its results are comparable with the HMM/GMM system [Wooters and Huijbregts, 2008]. The main advantage of IB is it requires less computation time more than HMM/GMM systems [Vijayasenan et al., 2009,Vijayasenan and Valente, 2012]. IB clustering clusters segments with similar distributions over a set of variables called relevance variables. Let X={x1, x2, ..., xn}represent the input variables to be clustered and Y={y1, y2, ..., ym} denote the relevance variables with meaningful information about clustering output C={c1, c2, ..., cr}. IB method tries to optimize the clustering process by maximizing the following equation: F=I(Y, C)−1 βI(C, X) (2.33) where βis a Lagrange multiplier, I(X, C) denotes the mutual information where X represents the speech segment set at each iteration and Crepresents the clusters, and I(Y, C) measures the mutual dependence between the relevant variables Y and the clustering partition C. The IB system uses a greedy technique to optimize the clustering process [Vijayasenan and Valente, 2012]. It starts with unique segmentation where each segment is considered as a set of input variables X. The set of relevance variables Y is components of background GMM estimated from the speech segments. Given input speech segment xi, the posterior distribution of the relevance variables for the segment xiis obtained using Bayes rule. The clustering of IB is initialized with each member of the set of speech segment Xand the two clusters with the most similar distribution are merged at each iteration. Other approaches The HMM/GMM and IB based speaker diarization systems are based on an agglomerative clustering framework. There are also other approaches to speaker diarization. They are described as follows: Top down system The top down-approach starts by modeling the entire audio signal with a single speaker model. Then, it successively generates new speaker models. The generation of new speaker models can be done using some criterion such as duration of the speech segment. A new speaker model is generated for these speech segments. This process is performed iteratively until the final number of speaker is found. Top-down approaches are not
2.8. Evaluation Metrics 37 widely used as the bottom up one. They are however computationally efficient and their performance can be improved using cluster purification as reported in [Bozonnet et al., ]. Factor analysis techniques Factor analysis techniques which are the state of the art in speaker recognition have recently been successfully used in speaker diarization [Kenny et al., 2010,Franco-Pedroso et al., 2010,Shum et al., 2011]. The speech clusters are first represented by i-vectors and the successive clustering stages are performed based on i-vector modeling. The use of factor analysis technique to model speech segments reduces the dimension of the feature vector by retaining most of the relevant information. Once the speech clusters are represented by i-vectors, cosine-distance and PLDA scoring techniques can be applied to decide if two clusters belong to the same or different speaker(s). [Dehak et al., 2011]. 2.8 Evaluation Metrics Diarization Error Rate (DER) is the metric used to measure the performance of speaker diarzation systems as described and used by NIST in the RT evaluations (NIST Fall Rich Transcription on meetings 2006 Evaluation Plan 2006). It is measured as the fraction of time that is not attributed correctly to a speaker or non-speech. A script named MD-eval-v12.pl has been used in the experiments. When DER is calculated, the hypothesized diarization output does not need to identify the speakers by name or definite ID. The speaker name or speaker id should not be the same in the hypothesis and the reference segmentation. The evaluation script first does an optimum one-to-one mapping of all speaker label ID between hypothesis and reference files. This allows the scoring of different ID tags between the two files. When evaluating DER, NIST uses a collar of 250 ms at the beginning and end of each segment boundary not to penalize slight discrepancies in the start and end times of the speech segments. The DER is calculated as follows: DER =PS s=1 dur(s).(max(Nref (s), Nhyp(s)) −Ncorrect(s)) PS s=1 dur(s).Nref (2.34) where Sis the total number of speaker segments where both reference and hypothesis files contain the same speaker pairs. Nref (s) and Nsys(s) represent the number of speaker speaking in segment s. The number of speakers that speak in segment s and are correctly
38 Chapter 2. State-of-the-art in Speaker Diarization matched between reference and hypothesis is represented by Ncorrect(s). The duration of segment s is represented by dur(s). The DER is composed of the following three errors: •Speaker Error: It is the percentage of scored time that a speaker ID is assigned to the wrong speaker. Speaker error is mainly a diarization system error (i.e., it is not related to speech/non-speech detection.) It also does not take into account the overlap speeches not detected. Speaker Error =PS s=1 dur(s).(min(Nref (s), Nhyp(s)) −Ncorrect(s)) PS s=1 dur(s).Nref (2.35) •False Alarm: It is the percentage of scored time that a hypothesized speaker is labelled as a non-speech in the reference. The false alarm error occurs mainly due to the the speech/non-speech detection error (i.e., the speech/non-speech detection considers a non-speech segment as a speech segment). Hence, false alarm error is not related to segmentation and clustering errors. False Alarm =PS s=1 dur(s).(Nhyp(s)) −Nref (s)) PS s=1 dur(s).Nref (2.36) •Missed Speech: It is the percentage of scored time that a hypothesized non-speech segment corresponds to a reference speaker segment. The missed speech occurs mainly due to the the speech/non-speech detection error (i.e., the speech segment is considered as a non-speech segment). Hence, missed speech is not related to segmentation and clustering errors. Miss Speech =PS s=1 dur(s).(Nref (s)) −Nhyp(s)) PS s=1 dur(s).Nref (2.37) •Overlap Error: When there are multiple speakers in the speech segment, the speaker diarization system has to detect and assign the segment to to all speakers. Therefore, overlap error is the percentage of scored time that some of the multiple speakers in a segment are not assigned to any speaker. The overlap error falls into one the three types of the previously mentioned diarization errors: speaker error when a speaker is detected to be present in the speech segment but the speaker is
2.8. Evaluation Metrics 39 not actually present, false alarm speech if more number of speakers are detected than the actual number of speakers in the overlapped speech segment, and missed speech when fewer speakers are detected than the actual number of speakers in the segment. Therefore, the total DER is calculated as: DER =Speaker Error +False Alarm +Miss Speech +Overlap Error (2.38) Once the DER is calculated for each show, the time weighted average is calculated among all meetings to find average DER for given set of shows. It is usual to score the diarization error rate ignoring the overlapped speech segment.
Chapter 3 The UPC Baseline Speaker Diarization System This chapter describes the techniques and implementation of the baseline speaker diarization of UPC which we use as a benchmark for the proposed systems. The baseline speaker diarization system is based on HMM/GMM systems and uses MelFrequency Cepstral Coefficients (MFCC). It uses a bottom-up agglomerative clustering approach that uses a modified version of the BIC distance in order to iteratively merge the closest clusters. Each HMM state represents a speaker whose emission probabilities are modeled using GMM. The baseline speaker diarization system follows multiple steps of agglomerative clustering and realignment (i.e., the speaker segmentation and clustering are carried out iteratively). The system is initialized first with many number of speakers. Then, the two most similar clusters are merged at each iteration. After merging, the time boundaries of segments are realigned using a Viterbi segmentation. The process is iteratively repeated until a stopping criterion is met. A detailed description about the baseline system used in the thesis is found in [Luque, 2012]. As it is shown in Figure 3.1, the baseline speaker diarization system consists of three modules. These are the feature extraction, speaker segmentation and speaker clustering. The three main stages of the baseline system diarization are as follows: •Feature extraction and removal of non-speech frames. A uniform initial clustering is performed by partitioning the data equally (see Fig. 3.1, block A). •Model complexity selection based on the amount of data per cluster and the cluster complexity ratio (CCR) is carried out to fix the amount of speech (seconds) per 41
48 Chapter 3. The UPC Baseline Speaker Diarization System each iteration. A time constraint is imposed as in [Ajmera and Wooters, 2003] on the duration of the speaker segments through a hierarchical modeling of each state as it is shown in Figure 3.3. The Viterbi decoding decisions are based on the estimation of the observation probabilities of accumulated likelihoods per cluster/state in a 3 seconds window. This procedure is carried out iteratively until the stopping criterion is reached. The stopping criterion is reached when the highest BIC distance scores among the set of clusters is less than 0. Finally, the system output the speaker segmentation outputs. Since the segmentation and clustering steps are performed iteratively in the baseline system, the errors made in the segmentation step are corrected in the clustering. 3.5 Merging and Stopping Criterion One the speech segments have been generated by the Viterbi segmentation and each segment is assigned a cluster, the speaker clustering modules merges the two closest clusters. This process is performed iteratively until there are no more clusters to merge. This techniques requires two metrics: which pairs of clusters to merge at each iteration and when to stop merging. The UPC baseline speaker diarization system uses the modified BIC algorithm described in [Ajmera and Wooters, 2003] to merge clusters and stop merging. Given two speech segments Xand Y, the modified BIC algorithm decides whether the speech segments Xand Yare uttered by the same speaker (H1) or different speakers (H2). Let Z = X∪Y. The modified BIC equation is defined as follows: ∆BIC =BIC(X, Y ) = H1−H2≶0 (3.5) which does not take into account the penalty term that corresponds to the number of free parameters of a multivariate Gaussian distribution. While model H1assumes that the speech segments are represented by the same speaker, model H2assumes that the two speech segments belong to two different speakers. The log likelihood of H1is obtained as follows: H1= nx X i=1 log p(Xi|θz) + ny X i=1 log p(Yi|θz) (3.6) where nxand nyare the number of frames in speech segments X and Y, respectively. Speech segments X and Y are are modeled by θzin H1.
3.5. Merging and Stopping Criterion 49 In the case of H2, a speaker change point exists at time Tj. The speech segments X and Y are modeled by two speaker models, which are represented by θxand θy, respectively. Then, the log likelihood H2is obtained as follows: H2= nx X i=1 log p(Xi|θx) + ny X i=1 log p(Yi|θy) (3.7) If BIC (X,Y) is greater than 0, speech segment X and Y are modeled by one speaker model, θZ. If BIC(X,Y) is less than 0, speech segment X and Y are modeled by two different speaker models, θXand θY. Equation 3.5 is similar to traditional BIC criterion except it doesn’t use the penalty term. The number of parameters in model is equal to the sum of the number of parameters in θxand θy. The segmentation obtained from the outputs of the segmentation (see Figure 3.1) defines a new set of speaker clusters and has to be trained at each iteration. We look for the the set of BIC scores for clusters Yand Ysatisfying BIC(X, Y ) ¿ 0. When there are many candidate pairs, we choose the pair of clusters with highest BIC score. This method provides a fully automatic stopping criterion that does not require the use of any tunable parameters. The stopping criterion for clustering is based on a threshold value of the BIC distances among all clusters. When the maximum BIC distance among these clusters is less than threshold value 0, the speaker diarization system stops and outputs the hypothesis.
Chapter 4 Long-term Speech Features for Speaker Diarization Feature extraction plays a significant role on the performance of speaker diarization systems. It needs to extract relevant information from the acoustic signal that can separate the different speaker present in the signal while discarding unwanted signals at the same time such as noise. It transforms raw acoustic signal into compact representation. It computes a sequence of feature vectors that represents compact speech signal. The feature vectors which are extracted from the raw signal emphasize speaker specific properties and suppress statistical redundancies. Although different types of features can be extracted from speech signals, only some of them can be used to discriminate speakers. According to [Kinnunen and Li, 2010], an ideal speech features have the following characteristics: they have large betweenspeaker variability and small within-speaker variability, they are robust against noise and distortion, they occur frequently and naturally in speech, they are easy to measure from speech signal, they are difficult to impersonate and they are not affected by the speaker’s health or long-term variations in voice. There are generally two broad categories of speech features. These are the short-term and long-term features. Short-term spectrum based features are the most widely used in speaker diarization systems since the short-term spectrum based features carry information about the vocal tract characteristics of individual speakers. They are also easy to extract and provide better performance. Short-term features are descriptors of the short-term spectral envelope which is an acoustic correlate of timbre and the resonance properties of the supralaryngeal vocal tract. 51
52 Chapter 4. Long-term Speech Features for Speaker Diarization They are extracted from either the short-term Fourier transform of the windowed speech signal or from linear prediction (LP) analysis. They are typically extracted for every 10 ms from a window size of around 30 ms. Speech signal continuously vary because of articulatory movements. Therefore, it needs to be broken down into short frames of about 20-30 milliseconds duration [Kinnunen and Li, 2010]. The speech signal is quasi-stationary within this interval and feature vectors are extracted from each frame. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used short-term acoustic features for speaker diarization [Anguera et al., 2012,Ajmera and Wooters, 2003]. They are computed with the aid of a psychoacoustically motivated filterbank, followed by logarithmic compression and discrete cosine transform (DCT). The dimensionality of MFCCs for speaker diarization is mostly around 20. Other widely used short-term spectral features used for speaker diarization include perceptual linear prediction coefficients (PLP) [Sinha et al., 2005], and linear prediction cepstral coefficients (LPCCs) [Ajmera and Wooters, 2003]. Although short-term spectral features are the most widely used speech features for different speech applications including speaker diarization, it is reported in [Friedland et al., 2009] that long-term features can be employed to reveal individual differences which can not be captured by short-term spectral features. Long-term speech features also have more discriminative power more than the short-term speech features. The selection of prosodic and other long-term features and their combination with MFCCs dramatically increases the accuracy of a state-of-the-art speaker diarization system [Friedland et al., 2009]. It is also reported in [Zelen´ak and Hernando, 2011] the performance of the stateof-the-art speaker diarization systems can be improved by combining spectral features with prosodic features to detect overlap speeches in speaker diarization system. Long-term features are extracted from portions of speech longer than one frame unlike short-term features which are extracted from a single speech frame. Long-term features capture phonetic, prosodic, lexical, syntactic, semantic and pragmatic information. They are also robust to channel variation since lexical usage or temporal patterns do not change with the change of acoustic conditions [Shriberg, 2007]. The long-term speech features used in the experiments are the delta dynamic, voicequality, prosody and GNE features.
4.1. Dynamic Features 53 4.1 Dynamic Features Mel Frequency cepstral coefficients (MFCCs) are the most widely used short-term features for speaker diarization [Anguera et al., 2012]. Most of the state of the art speaker diarization systems use only the static MFCC for diarization. However, the static MFCC features can not accurately capture the transitional characteristics of the speech signal which contains the speaker specific information. The delta dynamic features are the band-pass filtered versions of the static features that carry the temporal information of the static features. The delta features can be used to extract more detailed speech features using the time derivation of static MFCC acoustic vectors. The delta features can add dynamic information to the static MFCC features [Memon et al., 2009]. The dynamic features represent spectral changes over time and remove the time-invariant spectral information. The static cepstral features are also more adversely affected by convolutional noise (i.e. channel effect) more than the dynamic delta cepstral features. The dynamic features can be used to characterize the time trajectories of various acoustic parameters. They correspond to the slope associated with a specific parameter trajectory. The delta dynamic features provide new information to each frame that is not extracted by the static MFCC features. The MFCC feature vector describes only the power spectral envelope of a single frame. But , a speech signal has also information in the dynamics (i.e., what are the trajectories of the MFCC coefficients over time). The delta features are computed as the time differences between the adjacent vectors feature coefficients and usually appended with the static coefficients at the frame level. The extraction of the MFCC trajectories and appending them to the static MFCC features improves the performance of different speech applications. The delta features have been shown to improve the performance of speech recognition systems [Kumar et al., 2011]. It is also shown in the works of [Memon et al., 2009,Nguyen, 2010] that the delta features can be used to improve the performance of speaker verification and speaker classification, respectively. The delta features are computed by the weighted sum of feature vectors differences within a time window of 2 as follows: dt=PΘ θ=1 θ(Ci+θ−Ci−θ) 2PΘ θ=1 θ2(4.1) where dtis the delta coefficient at time tcomputed in terms of the corresponding static coefficients to Ci−θto Ci+θ. The delta window size is represented by Θ.
54 Chapter 4. Long-term Speech Features for Speaker Diarization Figure 4.1: The process of Delta feature extraction. Since clustering is more or less similar to speaker verification, we have proposed the use of dynamic features for speaker clustering task of speaker diarization. Note that for the proposed use of dynamic features for speaker clustering, the short-term static cepstral features are extracted from speech frames of 30 ms with 10ms shift. The dimension of the static cepstral features is 20. Then, the delta coefficients are computed with temporal window of 3 frames and augmented with the static coefficients to create 40 dimensional feature vector. Finally, the stacked static and dynamic features are used for speaker clustering. The speaker segmentation is based only on the static MFCC feature sets. Both the static and delta MFCC features are extracted using Hidden Markov Toolkit (HTK) [Young and Young, 1993]. 4.2 Voice-quality Voice source quality features characterize the glottal excitation signal of voiced voices such as glottal pulse shape and fundamental frequency, and carry speaker-specific information. Analysis of the voice-quality of a person is a valuable technique for speech pathology detection [Bielamowicz et al., 1996,Zwetsch et al., 2006,Wertzner et al., 2005] since the voice-disorders can be analyzed using acoustic signal parameters. Voice-quality features do not have an acoustic property that is easily distinguishable and measurable from a speech signal unlike F0. Voice quality is composed of many aspects of the speech production. It is characterized by qualitative terms such as hoarseness, whispering, creakiness, etc. The acoustic parameters can be used to detect if a person has pathological problem or not.
4.2. Voice-quality 55 The most widely used acoustic parameters used to asses the quality of a voice of a person are jitter, shimmer, Harmonics to Noise Ratio (HNR) and Glottal to Noise Excitation (GNE). However, their reliable estimation is based on an accurate measurement of the fundamental frequency which is a difficult task in the presence of certain pathologies. While fundamental frequency is determined physiologically by the number of cycles that the vocal folds do in a second, vocal intensity is affected by the amplitude of variation and tension of vocal folds [Wertzner et al., 2005]. Jitter is mainly affected by the lack of control of vocal fold vibration [Wertzner et al., 2005]. The more jitter deviates from zero, the more it correlates with erratic vibratory patterns of the vocal folds [Baken and Orlikoff, 2000]. Although the vibratory cycles of all speakers are erratic to some extent, abnormal voices are more erratic than a normal voice. While normal voices have little jitter, “hoarse” and “breathy” voices have higher degrees of jitter [Baken and Orlikoff, 2000]. It is also reported in [Wertzner et al., 2005] that patients with pathological problems often have higher values of jitter. Shimmer measures small, cycle-to-cycle changes of amplitude which occur during phonation and quantify short-term amplitude instability. It is affected mainly because of the tension and lesions of the vocal folds. It is correlated with the presence of noise emission and breathiness. Patients with pathologies have higher values of shimmer [Baken and Orlikoff, 2000]. Jitter and shimmer are used as measures to assess the micro-instability of vocal fold vibrations. The calculation of jitter and shimmer measurements is usually based on an autocorrelation method for determining the frequency and location of each cycle of vibration of the vocal folds (i.e.,pitch marks) [Rusz et al., 2011]. Studies show that these voice quality features can be used to detect voice pathologies [Wagner, 2013,Michaelis et al., 1998b,Kreiman and Gerratt, 2005]. They are normally used to measure long sustained vowels where voice-quality measurement values above a certain threshold are considered as pathological voices. In addition to this, voice quality features are related to the shape and dimension of the speaker’s vocal tract, and the way how the speech is generated by the voice production mechanism. For example, it is shown in the work of [Li et al., 2007] that jitter and shimmer measurements provide significant differences between different speaking styles. It is reported in [Linville, 1995,Schotz, 2001,Minematsu et al., 2002,Wittig and M¨uller, 2003] that jitter and shimmer are appropriate features to characterise the age and the gender of a speaker. It is also reported in [Slyh et al., 1999,Li et al., 2007] that significant differences can occur in jitter and shimmer measurements between different speaking styles, especially in shimmer measurement. Since pathological voices normally characterize a particular speaker, they can be used to discriminate different speakers.
56 Chapter 4. Long-term Speech Features for Speaker Diarization Jitter and shimmer voice quality features measure variations of the fundamental frequency and amplitude of speaker’s voice, respectively. They are very useful to describe the fluctuations of the voice signal in a qualitative way. They are given as a percentage that represents the maximum deviation from a normal frequency or amplitude. Adding jitter and shimmer voice quality features to both spectral and prosodic features improves the performance of a speaker verification system [Farr´us et al., 2007]. It is also shown in [Li et al., 2007] that fusion of voice quality features together with the spectral ones improves the classification accuracy of different speaking styles and conveys information that discriminates the different animal arousal levels such as happiness, angriness, etc. There are different type of jitter and shimmer measurements [Boersma and Weenink, 2009]. The five types of jitter measurements are Jitter (local), Jitter (local, absolute), Jitter (rap), Jitter (ppq5) and Jitter (ddp). The six kinds of shimmer measurements are Shimmer (local), Shimmer (local, dB), Shimmer (apq3), Shimmer (apq5), Shimmer (apq11) and Shimmer (ddp). Although there are different types of jitter and shimmer measurements as it is explained above, we have extracted only absolute jitter, absolute shimmer and shimmer (apq3) encouraged by previous work of [Farr´us et al., 2007]. It is reported in [Farr´us et al., 2007] that absolute jitter, absolute shimmer and shimmer apq3 measurements provide better results for speaker recognition more than the other jitter and shimmer measurements. 4.2.1 Jitter It is a measure of the periodic deviation of pitch perturbation of voice signal. Ideally, each cycle of speech has same period. Jitter measures how much one period differs from the next in the speech signal. The variations in jitter is mainly from the vocal cords where fluctuations in the opening and closing times introduce a noise that is shown as a frequency modulation in the speech signal. Jitter is a useful measure in speech pathology since pathological voices often have a higher jitter than healthy voices [Styler, 2013]. The values of jitter can be higher because a number of conditions that affect the vocal cords such as nodules, polyps, and weakness of the laryngeal muscles. •Jitter (absolute): It is a cycle-to-cycle perturbation in the fundamental frequency of the voice (i.e., the average absolute difference between consecutive periods). It is expressed as follows:
4.2. Voice-quality 57 Jitter (absolute) =1 N−1 N−1 X i=1 |Ti−Ti+1|(4.2) where Tiare the extracted pitch period lengths and Nis the number of extracted pitch periods. The Multidimensional Voice Program (MDVP) [Deliyski, 1993] calls this parameter Jita, and sets 83.2 as a threshold for pathology. Ti-1 Ti Ti+1 Figure 4.2: Jitter measurements for 3pitch periods 4.2.2 Shimmer Shimmer is similar to jitter, but instead of looking at periodicity, it measures the difference in amplitude from cycle to cycle. The shimmer changes with the reduction of glottal resistance and mass lesions on the vocal cords and is correlated with the presence of noise emission and breathiness. It is also a useful measure in speech pathology since pathological voices often have a higher shimmer than healthy voices. [Styler, 2013]. •Shimmer (absolute): This is the average absolute base-10 logarithm of the difference between the amplitudes of consecutive periods, multiplied by 20. It is expressed as follows: Shimmer (absolute) =1 N−1 N−1 X i=1 20 log Ai+1 Ai(4.3) where Aiare the extracted peak-to-peak amplitude data and Nis the number of extracted pitch periods. The Multidimensional Voice Program (MDVP) [Deliyski, 1993] calls this parameter ShdB and sets 0.350 dB as a threshold for pathology.
64 Chapter 4. Long-term Speech Features for Speaker Diarization 10 kHz sampling frequency (step 1). The inverse filtering is then carried out using the linear-prediction error signal by applying a predictor of 13th order computed by auto correlation method. It normally uses a Hanning window of 30ms length with 10ms shift in the successive frames. The calculation of the Hilbert envelopes (step 3) is done most efficiently in the frequency domain, without using a filter bank as follows: 1. Apply a real discrete Fourier transformation (DFT) on the time signal. The Fourier components at negative frequencies do not have to be calculated in a real DFT. 2. Select a frequency band from the complex spectrum and apply a Hanning window. 3. Double the length of the signal obtained from step 2 by padding zeros (i.e., setting the values at negative frequencies to zero). 4. Apply an inverse Fourier transform (IFFT). 5. Take the absolute value of the complex signal. Steps from (2) to (5) are applied to each frequency band. Since the envelopes have different phases, it is not sufficient to calculate the zero-timeshift correlation. The reason for the phase shifts might be that the maximum excitation of different frequencies does not occur at exactly the same time during glottal closure. The delay used for the correlation function between two envelopes in step 4 ranges between -3 and +3 samples (±0.3ms). The maximum within this range is picked for each correlation function (step 5). Finally in step 6, the maximal correlation is chosen as the GNE parameter. In contrast to other acoustic parameters such as jitter and shimmer, the main advantage of GNE is its computation is independent of variations of fundamental frequency and amplitude [S´aenz Lech´on et al., 2009,Michaelis et al., 1998a]. It is shown in [S´aenz Lech´on et al., 2009] that GNE parameter has a significant potential to screen voices since it quantifies the amount of voice excitation and turbulent noise. It is also reported in [Godino-Llorente et al., 2010] that GNE provides reliable measurements in terms of discrimination among normal and pathological voices more than other classical long-term noise measurements, such as Normalized Noise Energy and Harmonics to Noise Ratio. It has also been used successfully to screen voice disorders in [Godino-Llorente et al., 2010]. It is also reported in [Pop et al., 2007] that MFCC features have been used together with noise features to reliably assess normal
4.4. Glottal-to-Noise Excitation Ratio 65 and pathological voices. It is also shown in [S´aenz Lech´on et al., 2009] that GNE is reliable measurement to discriminate normal and pathological voices more than other classical long-term noise measurements found in the literature, such as Normalized Noise Energy or Harmonics to Noise Ratio. The voice-quality, prosodic and GNE features are extracted over 30ms frame length and at 10ms shift using Praat software [Boersma and Weenink, 2009]. After the computation of the actual values of the voice-quality, prosodic and GNE parameters for any given time point, suprasegmental statistical characteristics are also computed. The long-term mean statistics is computed over 500 ms windows with a 10 ms step. This is done to smooth out the feature estimation of the unvoiced frames. It is also done to synchronize the long-term features with the short-term ones. The non-speech regions are not considered when computing the statistical parameter.
Chapter 5 Proposed Speaker Diarization Systems This chapter describes the techniques and implementations of the proposed HMM/GMM and i-vector based speaker diarization systems. The main contributions of the proposed HMM/GMM system, compared with the baseline speaker diarization system, is clearly described. The proposed i-vector based clustering techniques based on i-vectors extracted from the shortand long-term speech features are also explained. After the extraction of the short-term cepstrl and long-term speech features, different types of fusion techniques have been carried out for the proposed GMM and i-vector based speaker diarization systems. The long-term features are the voice-quality, prosodic and Glottal-to-Noise Excitation Ration (GNE) features. The voice-quality features are absolute jitter, absolute shimmer and shimmer apq3. The prosodic features are the evolution in time of pitch, acoustic intensity and the first four formant frequencies. A detailed description of the long-term features used in the proposed speaker diarization systems is described in Chapter 4. Fusion techniques can be carried out at different levels. It can done at the feature extraction level, the match score level and the decision level [Sim and Lee, 2010,Zhang, 2009]. 5.1 Fusion Techniques Since fusion techniques extract multiple information from multiple sources and improve accuracy, they have been successfully used in various tasks including speaker recognition [Farr´us et al., 2007] , speaker diarization [Friedland et al., 2009,Zelen´ak and Hernando, 67
68 Chapter 5. Proposed Speaker Diarization Systems 2011,Woubie et al., 2015] and multi-biometrics [Nandakumar et al., 2009,Nandakumar et al., 2008]. The feature and score level fusion techniques carried out in the proposed GMM and i-vector based speaker diarization systems are described as follows: 5.1.1 Feature Level Fusion Feature level fusion is carried out after the extraction of features from the different sources. The extracted features can be fused in several ways. The simplest one is to stack the different features extracted from the different sources in the same feature vector. Fusion at the feature level can also be carried out in a more complex way on an algorithmic level by using other techniques. Fusion at the feature level exploits most of the information of the original data since it integrates the multi-source information at the most early stage [Xu and Zhang, 2010]. However, fusion at the feature level is susceptible that the different sources of information may not be consistent and compatible. In the proposed speaker diarization systems, the feature level fusion is carried by stacking the voice-quality, prosodic and GNE features in the same feature vector (see Figure 5.1). Figure 5.1: Example of feature level fusion. 5.1.2 Score Level Fusion Score level fusion fuses the individual scores obtained from different sources to obtain a single score. It fuses the outputs of each individual feature source using a combination algorithm. The score fusion technique provides a very high accuracy since it allows multiple scores to be independently treated and integrated [Zhang, 2009]. It uses the sum rule, maximum rule, minimum rule, and product rule to integrate the scores of different data sources
5.2. Proposed HMM/GMM Speaker Diarization System 69 [Zhang, 2009]. It is reported in [Kittler et al., 1998] that sum rule provides better result more than the other score fusion techniques. Hence, in the proposed GMM and i-vector based speaker diarization systems, the score fusion technique is carried out using the sum rule as follows: Fss = N X i=1 si(5.1) where Fss is the fused sum score, N is the number of features used and siis the score of feature i. As it is shown in Table 5.1, the fusion of the short-term spectral features with the longterm ones is carried out differently in segmentation and clustering at the score level for the proposed GMM and i-vector based speaker diarization systems. The optimum set of weight values tuned on the development data for the shortand long-term speech features have been applied for the score fusion techniques. Fused scores Diarization System Segmentation Clustering HMM/GMM BIC HMM/i-vector Cosine distance HMM/i-vector Log-likelihood scores PLDA Table 5.1: The proposed GMM and i-vector based speaker diariaztion systems and the score fusion techniques carried out in segmentation and clustering. The fusion of the shortand long-term speech features is based on the emission probabilities of GMM in speaker segmentation both in the proposed GMM and i-vector based speaker diarization systems (see Figure 5.2, block B). The fusion of the shortand long-term speech features is based on the BIC Scores in speaker clustering in the proposed GMM speaker diarization systems (see Figure 5.2, block C). The fusion of the shortand long-term speech features is based on the cosine and PLDA scores of i-vectors in the proposed i-vector based cosine distance and PLDA clustering systems (see Figure 5.3 and 5.4). 5.2 Proposed HMM/GMM Speaker Diarization System As is mentioned in Chapter 3, the UPC baseline speaker diarization system consists of three basic modules. The first module (Figure 3.1, block A) performs mainly the
70 Chapter 5. Proposed Speaker Diarization Systems feature extraction process. The second module (Figure 3.1, block B) detects speaker change points and performs Viterbi segmentation. The third module (Figure 3.1, block C) performs the bottom-up clustering and outputs the system hypothesis. Figure 5.2: Proposed HMM/GMM based speaker diarization system using shortand long-term term speech features. The highlighted boxes are the ones proposed in the HMM/GMM based speaker diarization system. The unhighlighted boxes are the same both in the baseline and proposed systems. Note that the delta features are used only in speaker clustering together with the static features.
5.2. Proposed HMM/GMM Speaker Diarization System 71 5.2.1 Feature Extraction One of the main contributions of the proposed HMM/GMM speaker diarization system is the extraction of jitter and shimmer voice-quality features, and their fusion with the prosodic and MFCC features. The prosodic features are pitch, intensity and the first four formant frequencies. Note that the baseline system is exclusively based on MFCC feature set. After the extraction of jitter and shimmer voice-quality features, the following feature fusion techniques have been carried out at the feature level (see Figure 5.2, block A): •Fusing Jitter and Shimmer Voice-quality features with Prosodic Ones •Fusing Jitter and Shimmer Voice-quality with Prosodic and GNE After the extraction of the short-term spectral features and fusion of long-term features, the speech signal is then equally partitioned equally to generate an initial number of clusters. The initial number of clusters depends on meeting duration but it is constrained the range [10,65]. This is done to solve the common issues of Agglomerative Hierarchical Clustering (AHC) such as over-clustering and its high computational cost due to the combinatorial explosion in pair-wise distance computation. Detailed description about the selection of the initial number of clusters is described in Section 3.2. The voice-quality, prosodic and GNE features are extracted over 30ms frame length with 10ms frame shift using Praat software [Boersma and Weenink, 2009]. Then, each voice-quality, prosodic and GNE feature is estimated over a 500 ms window with 10ms shift. This is done to smooth out the feature estimation of the unvoiced frames. It is also done to synchronize the long-term features with the short-term ones. The other contribution of the proposed HMM/GMM speaker diarization system is the extraction of the first order time derivatives of the instantaneous cepstral delta features for speaker clustering. At first, the static MFCC and the delta features are stacked in the same feature vector. Then, they are used for speaker clustering. Note that the speaker segmentation is based only on the static MFCC feature set. As it is shown in Figure 5.2 (see block A), the short-term and long-term features are extracted only for the speech frames. The features are extracted using Oracle SAD (true segmentation). Hence, the non-speech frames are not taken into account when the long-term statistics is calculated from the long-term speech features.
72 Chapter 5. Proposed Speaker Diarization Systems 5.2.2 Speaker Segmentation The set of acoustic features corresponding to the shortand long-term speech features are modeled independently using Hidden Markov Model(HMM). Each state of the HMM is composed of a mixture of Gaussians, fitting the probability distribution of the features by the classical expectation-maximization (EM) algorithm. The two HMM models estimated from the short-term and long-term speech features and their best paths obtained by Viterbi segmentation are fused. The number of mixtures is chosen as a function of available seconds of speech per cluster in the case of MFCC features. But, they are kept fixed for the long-term speech features. Finally, a time constraint, as in [Ajmera and Wooters, 2003], is imposed on the HMM topology. The time constraint forces the minimum duration of the speaker turn to be greater than 3 seconds which is commonly used as mean value of a speaker intervention or speaker turn [Ajmera and Wooters, 2003]. The fusion of short-term spectral features with the long-term ones is carried out at the score level in speaker segmentation as it is shown in Figure 5.2, Block B. It is carried out by fusing the log-likelihood scores corresponding to these feature sets. Given a set of input features vectors, {x}and {y}, the log-likelihood score in the proposed HMM/GMM segmentation is computed as a joint log-likelihood between features distributions as follows: log P(x,y) = αlog P(x|θix) + (1 −α) log P(y|θiy),(5.2) where log P(x,y) is the fused emission probabilities for cluster i,θix is the model of cluster ifrom MFCC feature vectors, and θiy is the model for the same cluster iusing long-term features. The weight of the spectral feature vector is αand (1 −α) is the weight of long-term speech features. The values of the αare tuned on development data set. 5.2.3 Speaker Clustering Both the baseline and proposed speaker diarization systems are based on Agglomerative Hierarchical Clustering(AHC) technique. The distance among clusters is based on the Bayesian Information Criterion (BIC). This distance measures the difference among each pair of clusters. The stopping criterion is also driven by a threshold on the same matrix of distances (see Figure 5.2, block C). A modified BIC-based metric [Ajmera and Wooters, 2003] is employed to select the set of cluster-pairs candidates with smallest distances among them. The cluster-pairs (i, j) with the highest BIC score is merged
5.3. Proposed i-Vector based Speaker Diarization System 73 at each iteration. Then, a two-step training and decoding iteration is performed again to refine the model statistics and align them with the speech recording (see Figure 5.2, block B). This process continues iteratively until the highest BIC distance score among the set of clusters is less than the threshold value of the stopping criterion. Once the speech segments are generated by the Viterbi segmentation, the speaker clustering of the proposed HMM/GMM speaker clustering system is carried out as follows: BIC(i, j) = β . BICijx + (1 −β). BICijy,(5.3) where BICijx and BICijy are the BIC distances between clusters iand jgenerated using shortand long-term speech features, respectively. The BIC score computed using the shortand long-term features set are multiplied by βand (1 −β), respectively. The values of βare tuned on development data set. Note that the long-term features in equations 5.2 and 5.3 may refer to four different possibilities: voice-quality features, prosodic features, stacked voice-quality and prosodic features, and stacked voice-quality, prosodic and GNE features. We have also proposed the use of delta dynamic features for speaker clustering. The static MFCC and the delta features (∆) are stacked first in the same feature vector. Then, the stacked features are used for speaker clustering. The proposed technique is exactly the same as in Figure 5.2 except that the BIC distance metric is computed using the stacked static and dynamic features in clustering (see Figure 5.2, block C). The speaker segmentation is based only on the static MFCC feature set. 5.3 Proposed i-Vector based Speaker Diarization System Factor analysis techniques which are the state of the art in speaker recognition have recently been successfully applied in speaker diarization experiments [Kenny et al., 2010, Franco-Pedroso et al., 2010,Shum et al., 2011,Shum et al., 2012,Vaquero Avil´es-Casco, 2011,Senoussaoui et al., 2013]. In these works, i-vectors are extracted from speech segments and the successive clustering stages are carried out using i-vector modeling techniques (i.e., the cosine distance and PLDA scores of i-vectors are used as a distance metrics for clustering). The main contribution of the proposed i-vector based speaker diarization system is the extraction of i-vectors from short-term and long-term speech features, and the fusion of
80 Chapter 6. Experimental Setups and Results 6.2 HMM/GMM based Speaker Diarization Systems As it is explained in Chapter 3, the baseline speaker diarization system uses only the short-term MFCC features (MFCC). In the proposed HMM/GMM speaker diarization system, we have proposed the use of jitter and shimmer voice quality features for speaker diarization. The fusion of the voice-quality features with the state-of-the-art long-term prosodic and short-term MFCC features is carried out at the feature and score level, respectively. The following set of experiments have been carried out in the proposed HMM/GMM based speaker diarization systems. •The use of Delta (∆) Features for Speaker Clustering Most of the state of the art speaker diarization systems use only the static MFCC for diarization. The delta dynamic features can be used to capture the transitional characteristics of the speech signal which contains the speaker specific information. These information are not captured by the static MFCC features. In this work, we propose the use of delta dynamic features for speaker clustering. Firstly, the static and the dynamic features are stacked in the same feature vector. Then, the stacked features are used for speaker clustering only. The speaker segmentation is based only on the static MFCC feature set. •The Use of Jitter and Shimmer Voice-quality Measurements for Speaker Diarization Jitter and shimmer (JS) voice quality features are first extracted from the fundamental frequency contour. Then, they are fused together with the baseline MFCC features. The fusion of the voice-quality with MFCC is carried out at the score likelihood level both in segmentation and clustering. While the fusion in segmentation is based on the log-likelihood scores of HMM models of each feature set (see equation 5.2), the fusion in clustering is based on Bayesian Information Criterion (BIC) scores of each feature set (see equation 5.3). •The Use of Prosodic Features for Speaker Diarization First, features related to the evolution in time of pitch, acoustic intensity and the first four formant frequencies are extracted. Then, they are fused with the MFCC features at the score likelihood level both in segmentation and clustering. The fusion at the segmentation level is based on the log-likelihood scores of HMM models of each feature set (see equation 5.2). The fusion at the clustering is based on BIC scores of each feature set (see equation 5.3).
6.2. HMM/GMM based Speaker Diarization Systems 81 •Using Voice-quality with Prosodic and MFCC Features for Speaker Diarization The long-term voice-quality and prosodic features are first fused at the feature level (i.e., they are stacked in the same feature vector). Then, the stacked feature is fused with the MFCC at the score likelihood level both in segmentation and clustering. The fusion at the segmentation level is based on the log-likelihood scores of HMM models of each feature set (see equation 5.2). The fusion at the clustering is based on BIC scores of each feature set (see equation 5.3). 6.2.1 Experimental Setup Manually annotated speech references have been employed to extract the speech frames and discard non-speech regions both for the development and test sets. The main reason why we are interested to use the speech references, instead of Speech Activity Detection (SAD), is we want to focus exclusively on speaker errors that occur to the diarization approach. A feature vector of 20 MFCC features is computed with 30ms frame length at 10ms frame shift. The MFCC features are extracted using the Hidden Markov Model Toolkit [Young and Young, 1993]. The voice-quality and the prosodic features are extracted over 30ms frame length and 10ms frame shift using Praat software [Boersma and Weenink, 2009]. Then, we calculate the mean of each of the voice-quality and prosodic features over a window length of 500ms with 10ms step. This is done to smooth out the feature estimation of the voice-quality and prosodic features, and also synchronize them with the MFCC features. The experiments have been developed and tested on AMI corpus, a multi-party and spontaneous speech set of recordings [AMI, 2011]. The development and test sets are based on a mono-channel audio recording. •Development set: 10 shows have been selected from IDIAP, Edinburgh, and TNO sites as a development set. The development shows include both the scenario and non-scenario recordings. These shows are used to tune the optimum parameters (i.e., optimum set of weight values for the shortand long-term speech features). The total and average duration of the development set is 284 and 28.4 minutes, respectively. The development database is based on far-field microphone array channels sampled at 16kHz. •Test set: In order to evaluate the performance of the proposed systems, the test experiments have been carried out on 120 AMI shows consisting of both scenario and non-scenario meetings from Idiap, Edinburgh and TNO sites. We have also
82 Chapter 6. Experimental Setups and Results created another test set from these recordings by chopping them into 10 minutes duration and generated another 450 chunks test sets. The selected shows are the ones recorded using the far-field microphone array channels sampled at 16KHz. The total and average duration of the test sets of the whole recording (i.e., without chunking) are 4075 minutes (about 69 hours) and 36.38 minutes, respectively. Note that optimum parameters found through experimentation on the development sets have been directly used on the test sets. The performance metric employed for assessing speaker diarization systems is the Diarization Error Rate (DER). DER represents the sum of false alarm speech, missed speech and speaker error along time. Speaker error is the percentage of scored time that a speaker ID is assigned to the wrong speaker. False alarm is the percentage of scored time that a hypothesized speaker is labelled as a non-speech in the reference. Missed speech is the percentage of scored time that a hypothesized non-speech segment corresponds to a reference speaker segment. Since speech references have been used, the rate of false alarms and missed speech have zero values in the experimental results. Hence, DER values reported in the following sections correspond purely to speaker time confusion produced by the diarization system. We have used a collar of 250ms around every speaker segment to discard any inaccuracies in the reference annotation when the DER is scored. 1 6.2.2 Delta Features Results Speaker diarization systems use mainly the static MFCC speech features extracted from short-term power spectrum. The static MFCC features represent spectral characteristics associated with the speech segment. The delta dynamic features capture the transitional characteristics of the speech signal which contains the speaker specific information. Hence, we propose the use of delta dynamic features for speaker clustering as they add dynamic information to the static MFCC features. The speaker segmentation is based only on the static MFCC feature set. Experimental Results As it is shown in Table 6.1, the baseline system of the test set has a DER of 23.97%. Note that the baseline system is based only on static MFCC feature set both for segmentation and clustering. The table shows that the use of static MFCC features in segmentation, 1The scoring tool is the NIST RT scoring used as: ./md-eval-v21.pl -1 -nafc -c 0.25 -o -R reference.rttm -S hypothesis.rttm
6.2. HMM/GMM based Speaker Diarization Systems 83 Features Segmentation Clustering DER (%) MFCC MFCC 23.97 MFCC MFCC + Delta (∆) 21.55 Table 6.1: DER of the test sets for HMM/GMM speaker diarization system using MFCC and MFCC + Delta (∆) feature set. and static MFCC and delta dynamic features in clustering reduces the DER to 21.55%. This represents a 10.01% relative DER improvement more than the baseline system. Summary We have proposed the use of delta features for speaker clustering since the delta features add dynamic information to the static cepstral features. The experimental results show that use of delta dynamic features improve the performance of speaker diarization systems by complementing the transitional characteristics of the speech signal which contains speaker specific information. 6.2.3 Jitter and Shimmer Results Jitter and shimmer (JS) measure variations in the fundamental frequency and amplitude of speaker’s voice, respectively. Due to their nature, they can be used to assess differences between speakers. Therefore, we propose the use of jitter and shimmer voice quality features for speaker diarization since they provide complementary information to the baseline MFCC features. The main contribution of this work is the extraction of jitter and shimmer voice quality features and their fusion with the MFCCs in the framework of speaker diarization. Although there are different estimations of jitter and shimmer measurements, we have extracted the following three measurements called absolute jitter, absolute shimmer and shimmer apq3 encouraged by previous work of [Farr´us et al., 2007]. It is reported in [Farr´us et al., 2007] that these three measurements provide better results for speaker recognition more than the other jitter and shimmer measurements. Jitter and shimmer voice quality measurements are first extracted from the fundamental frequency contour. Then, they are fused together with the baseline MFCC features. Features Development DER(%) Test DER(%) MFCC 20.04 23.97 MFCC + JS 18.04 22.83 Table 6.2: DER of the development and test sets for HMM/GMM speaker diarization system using MFCC, and Jitter and Shimmer (JS) feature sets.
84 Chapter 6. Experimental Setups and Results Experimental Results As it is shown in Table 6.2, the baseline system of the development and test sets show DER of 20.04% and 23.97%, respectively. Note that the baseline system is based only on MFCC feature set. The table shows that the fusion of voice-quality features with the MFCC provides better DER both in the development and test sets. It provides DER of 18.04% and 22.83% for the development and test sets, respectively. These represent a 17.58% and 4.14% relative DER improvement more than the baseline system for the development and test sets, respectively. Summary We have proposed the use of jitter and shimmer voice quality features for speaker diarization experiment as these features add complementary information to the baseline MFCC features. Jitter and shimmer voice quality features are first extracted from the fundamental frequency contour, and are then fused together with the baseline MFCC features. The fusion of the two streams in segmentation and clustering is done at the score likelihood level by weighting linearly the log-likelihoods and BIC scores of each model (see equation 5.2 and 5.3), respectively. The experimental results show that adding jitter and shimmer voice quality features to the baseline MFCC features improve the DER. 6.2.4 Prosody Results We have also carried out an experiment using prosodic features together with MFCC. We have extracted the following prosodic features encouraged by the previous work of [Zelen´ak and Hernando, 2011]. Features related to the evolution in time of pitch, acoustic intensity and the first four formant frequencies have been extracted. Then, they are fused with the MFCC. Features Development DER(%) Test DER(%) MFCC 20.04 23.97 MFCC + Prosody 19.49 23.45 Table 6.3: DER of the development and test sets for HMM/GMM speaker diarization system using MFCC and prosodic feature sets. Experimental Results As it shown in Table 6.3, the use of prosodic features together with MFCC provides a little DER improvement more than the baseline system. The use of prosodic feature together with the MFCC ones provides a DER of 19.49% in the development set. This corresponds to a 2.74% relative improvement more than the baseline system. Similarly,
6.2. HMM/GMM based Speaker Diarization Systems 85 Figure 6.1: DER of the development and test sets for HMM/GMM speaker diarization system using MFCC, JS and prosodic feature sets. the fusion of prosodic features with the MFCC ones provides a 2.74% relative DER improvement more than the baseline system for the test set. Summary The experimental results show that the extraction of selected prosodic features and their combination with the MFCC ones improves the accuracy of speaker diarization system. The fusion of the two streams in segmentation and clustering is done at the score likelihood level by weighting linearly the log-likelihoods and BIC scores of each model (see equation 5.2 and 5.3), respectively. As it is shown in Figure 6.1, the use of both voice-quality and prosodic features together with MFCC provide better results more than using only MFCC feature set. The improvements are both for the development and test sets. This shows that the use of both voice-quality and prosodic features add complimentary information to the short-term MFCC features. 6.2.5 Voice-quality and Prosody Results The main contribution of this work is the fusion of jitter and shimmer voice-quality features both with the long-term prosodic and short-term MFCC features. The fusion of voice-quality with the prosodic and MFCC features is carried out both at the feature and score likelihood level. The voice-quality features are absolute jitter, absolute shimmer and shimmer apq3. The appropriate characteristics related to the human speech prosody are conveyed through intonation, rhythm and stress. Encouraged by work of [Zelen´ak and Hernando, 2011], we have extracted features related to the evolution in time of pitch,
86 Chapter 6. Experimental Setups and Results acoustic intensity and the first four formant frequencies to validate their performance in this work. Features Development DER(%) Test DER(%) MFCC 20.04 23.97 MFCC + JS 18.04 22.83 MFCC + Prosody 19.49 23.45 MFCC + (JS + Prosody) 17.16 21.68 Table 6.4: DER of the development and test sets for HMM/GMM speaker diarization system using MFCC, JS and prosodic feature sets. The long-term voice-quality and prosodic features are first fused at the feature level (i.e., they are stacked in the same feature vector). Then, the stacked feature is fused with the MFCC at the score likelihood level both in segmentation and clustering. Experimental results At it is shown in Table 6.4, the best results in the HMM/GMM system are obtained when MFCC features are used with the voice-quality and prosodic features both in the development and test sets. The fusion of voice-quality features with the prosodic ones at the feature level and their fusion with the MFCC ones provides a 14.37% relative DER improvement more than the baseline system for the development set. It also provides a 9.55% relative DER improvement more than the baseline system for the test set. Table 6.4 also shows that the use of voice-quality and prosodic features with MFCC ones provides better DER results more than using only voice-quality and only prosodic features with the MFCC ones. Figure 6.2 shows the DER ranges of the HMM/GMM speaker diarization system using different feature sets for the development and test sets. The figure shows the minimum, lower quartile, median, upper quartile, and maximum DER of different shows, respectively. The figure shows that the use of voice-quality features together with MFCC reduces the DER variation among the different shows both for the development and test sets, compared to the system that is based only on MFCC feature set. The use of prosodic features also reduces the DER variations among the different shows both for the development and test sets. The range of the DER variations among the different becomes lowest when MFCC features are used together with the voice-quality and prosodic features both for the development and test sets. Although the combination of the different fusion systems reduce the DER error for most of the shows both in the development and test sets, the error rate increases for some recordings compared to baseline system. Reasons for this effect should be explored in the future.
6.3. i-Vector based Speaker Diarization Systems 87 Figure 6.2: Box plot of the development and test sets for HMM/GMM speaker diarization system using MFCC, JS and prosodic feature sets. Summary In this work, we have proposed the use of jitter and shimmer voice-quality measurements as complementary source of information to both the long-term prosodic and short-term MFCC features within speaker diarization task. Experimental results on AMI corpus show that the fusion of voice-quality and prosodic features at the feature level and their fusion with MFCC ones at at the score likelihood provides better DER. The results of the experiments show the usefulness of voice-quality features as complementary source of information for speaker diarization. In overall, the experimental results validate the usefulness of fusing voice-quality features with the prosodic and MFCC ones. The box plots and experimental results show that the use of voice-quality features with the prosodic and MFCC ones increase the robustness and reliability of speaker diarization systems. 6.3 i-Vector based Speaker Diarization Systems Factor analysis techniques which are the state of the art in speaker recognition have recently been successfully applied in speaker diarization experiments [Kenny et al., 2010, Franco-Pedroso et al., 2010,Shum et al., 2011,Shum et al., 2012,Vaquero Avil´es-Casco,
88 Chapter 6. Experimental Setups and Results 2011,Senoussaoui et al., 2013]. The speech clusters are first represented by i-vectors and the successive clustering stages are performed based on i-vector modeling. Note that the above mentioned works extract i-vectors exclusively from short-term MFCC features for speaker clustering. The main contribution of this work is the extraction of i-vectors from short-term MFCC and long-term speech features. The long-term features are the concatenation of voice-quality, prosodic and GNE features. Once the two sets of i-vectors are first extracted from the outputs of the Viterbi segmentation (i.e., i-vectors from the short-term and long-term features), the cosine and PLDA scores of these i-vectors are fused as a clustering distance (see equation 5.4 and equation 5.6). The fusion of short-term MFCC features with the long-term ones is carried out in speaker segmentation using the log-likelihood scores corresponding to these feature sets as in [Woubie et al., 2015]. The fusion in segmentation is carried out as it is explained in equation 5.2. The main contribution is on speaker clustering. 6.3.1 Experimental Setup The UBM and the T matrix are trained using 100 AMI shows which have duration of 60 hours. Two Universal Background Models (UBMs) of 512 Gaussians components are trained. While the first UBM is for the short-term MFCC features, the second one is for the long-term ones. The UBM of short-term MFCC features is trained on 20 cepstral co-efficients without the deltas. The UBM of long-term features is trained using the stacked voice-quality, prosodic and GNE features. A 100 and 50 dimensional raw i-vector sizes are extracted from the shortand long-term speech features, respectively. The size of the total variability matrix is 100 for the shortterm speech features and 50 for the long-term ones. The i-vector framework is carried out using ALIZE open source software [Larcher et al., 2013]. The Probabilistic Linear Discriminant Analysis (PLDA) system of the short-term and long-term speech features use a 40 and 20 dimensional speaker space. The PLDA is trained on the same data used to train the UBM and T-matrix but the audio signals are chopped into pieces of 3 second segments. The selection of threshold value for stopping criterion for the proposed i-vector based speaker diarization systems is carried out as it is shown in Figure 6.3. It is based on a data driven approach. The DER and corresponding cosine distance/PLDA score values at each iteration are compared, and λvalue that minimizes the DER value is selected. Thus, the system stops merging when the highest cosine distance/PLDA score value among all pair of clusters is less than λ. As it is shown in Figure 6.3, the DER values
6.3. i-Vector based Speaker Diarization Systems 89 first decrease for some iteration first. But, its values start to increase after some number of iterations because of over-clustering. Figure 6.3: DER and cosine-distance score per iteration for selected shows from the development set. The same development and test sets explained in Section 6.2.1 are used for the proposed i-vector based speaker diarization systems. The optimum parameters found through experimentation on the development set have been directly used on the test sets. The tuned parameters are the threshold value for stopping criterion, size of i-vectors for the shortand long-term speech features, size of eigen-voice for PLDA training, and optimum set of weight values for i-vectors extracted from the shortand long-term speech features. 6.3.2 i-Vector based Cosine Distance Clustering In the proposed i-vector based cosine distance clustering technique, two sets of i-vectors are extracted from the outputs of Viterbi segmentation (see Figure 5.3). While the first i-vector is extracted from the short-term MFCC features, the second one is extracted from the long-term speech features. After the extraction of i-vectors from the shortand long-term speech features for each cluster, the cosine-distance scores of i-vectors are linearly weighted (see equation 5.4) to obtain a single cosine distance similarity score. Finally, the fused cosine distance score is used a distance metric for clustering. The two clusters with the highest cosine distance score are merged at each iteration. The viterbi segmentation and clustering process continues iteratively until the highest cosine distance score among the set of i-vectors is less than the threshold value for stopping criterion (i.e., λvalue). The Viterbi segmentation outputs a new clustering
96 Chapter 6. Experimental Setups and Results Figure 6.7: Box Plot of the chunks test set for GMM based BIC and i-vector based PLDA clustering techniques using MFCC, JS, Prosodic and GNE features. Similarly, the boxplots in Figure 6.7 show the DER variation of GMM based BIC and i-vector based PLDA clustering using short-term and long-term features. The long-term features are the concatenation of voice-quality, prosodic and GNE features. The figure shows the DER variations of the chunk test set. The figure shows that the use of ivector based PLDA clustering technique reduces the DER variations among different shows more than GMM base BIC clustering technique. The figure also manifests that the use of long-term features help in reducing the DER variations both for the GMM based BIC and i-vector based PLDA clustering techniques. A semi-automatic threshold value λis used as a stopping criterion on the matrix of distances of clusters. When the highest PLDA score among all pair of clusters is less than λ, the merging process stops (see Figure 6.3 for more details). Summary We have proposed the use of GNE feature and i-vector based PLDA clustering technique within the framework of speaker diarization. The clustering technique is based on the fusion of PLDA scores of i-vectors extracted from shortand long-term speech features. The long-term features are the concatenation of voice-quality, prosodic and GNE features. The experimental results show that i-vector based PLDA clustering technique provides a substantial relative DER improvement more than GMM based BIC clustering one.
6.3. i-Vector based Speaker Diarization Systems 97 Experimental results also show that the extraction of i-vectors from the shortand longterm speech features provides better DER result more than extracting i-vectors only from the short-term MFCC features. Finally, the results show that the use of GNE features together with the voice-quality and prosodic ones provides better DER result more than the system that uses only the latter features for i-vector based PLDA speaker clustering technique. The results of the experiments show the usefulness of replacing GMM based BIC clustering technique with the i-vector based PLDA clustering one. The experimental results also show the usefulness of voice-quality, prosodic, GNE and delta features for speaker diarization.
Chapter 7 Conclusions and Future works This chapter provides a brief summary of the thesis. The proposed techniques are reviewed with regard to the objectives discussed in Chapter 1. Finally, suggestion for future works will be outlined. 7.1 Conclusions This thesis has proposed the use of voice-quality features for GMM and i-vector based speaker diarization systems. The proposed voice-quality features are used together with the short-term cepstral, and long-term prosodic and Glottal-to-Noise Excitation Ratio (GNE) features. The fusion of the long-term voice-quality features with the prosodic and GNE is first carried out at the feature level (i.e., they are stacked in the same feature vector). Then, the stacked long-term speech features are fused with the cepstral features at the score likelihood level both for the proposed GMM and i-vector based speaker diarization systems. The thesis has also proposed the use of delta dynamic features for speaker clustering. The delta features are stacked in the same feature vector together with the static ones, and are used for speaker clustering. The main contributions of this PhD thesis can be summarized as follows: 1. The use of Delta Features for Speaker Clustering Mel Frequency cepstral coefficients (MFCCs) are the most widely used short-term features for speaker diarization. Most of the state of the art speaker diarization 99
100 Chapter 7. Conclusions and Future works systems use only the static MFCC for diarization. The dynamic delta features capture the transitional characteristics of the speech signal which contains the speaker specific information. These information are not captured by the static MFCC features. The delta dynamic features have been successfully used in speaker recognition, speaker verification, speaker classification and speech recognition. Thus, this work assess the impact of delta features on speaker clustering since speaker clustering is related to speaker verification, identification and recognition. We have proposed the use of static and delta dynamic features for speaker clustering since the dynamic delta features add dynamic information to the static cepstral features. The speaker segmentation is based only on the static MFCC feature set. Experimental results on subset of AMI corpus show that the use of only static MFCC features in segmentation, and static MFCC features with dynamic ones in clustering provides better DER more than using only static MFCC feature set both in segmentation and clustering. 2. The use of Voice-quality Features in Speaker Diarization Jitter and shimmer voice quality features have been successfully used to characterize speaker voice traits and detect voice pathologies. Jitter and shimmer measure variations of fundamental frequency and amplitude of speaker’s voice, respectively. Due to their nature, they can be used to assess differences between speakers. Therefore, we have proposed use of jitter and shimmer voice quality features in the framework of speaker diarization as these features add complementary information to the baseline MFCC features. At fist, jitter and shimmer voice quality features are extracted from the fundamental frequency contour. Then, they fused together with the baseline MFCC features. Both sets of features are independently modeled and fused together at the score likelihood level. While the score fusion in segmentation is based on the fusion of log-likelihoods scores of the cepstral and the voice-quality features, the score fusion in clustering is based on the fusion of BIC distances of the cepstral and voice-quality features. Experimental results on subset of AMI corpus show that fusing jitter and shimmer voice quality features with the baseline cepstral features provides better DER more than the baseline system which is based on exclusively MFCC feature set. 3. Using Voice-quality Features together with Prosodic for Speaker Diarization The main contribution of this work is the fusion of jitter and shimmer voicequality features both with the long-term prosodic and short-term cepstral features.
7.1. Conclusions 101 Firstly, the voice-quality and features related to the evolution in time of pitch, acoustic intensity and the first four formant frequencies are extracted. Then, the voice-quality and prosodic features are fused at the feature level (i.e., they are stacked in the same feature vector). Finally, the stacked voice-quality and prosodic features are fused with the cepstral features at the score likelihood level both in segmentation and clustering. The score fusion in segmentation is based on the fusion of log-likelihood scores of the cepstral and the stacked voice-quality and prosodic features. The score fusion in clustering is based on the fusion of BIC distances of these feature sets. Experimental results show that the fusion of voice-quality features together with the prosodic ones at the feature level, and their fusion with the cepstral at the score level provides better DER result. It provides better DER results not only on systems that are based only on short-term cepstral features but also on systems that based on short-term cepstral and voice-quality features. Hence, the experimental results show the usefulness of voice-quality features as complementary source of information for speaker diarization systems based both on short-term cepstral and long-term prosodic features. 4. Improving i-Vector based Speaker Clustering with Long-term Features Factor analysis techniques which are the state of the art in speaker recognition have recently been successfully applied in speaker clustering. The speech clusters are first represented by i-vectors and the successive clustering stages are carried out using i-vector modeling techniques. Representing the speech clusters by i-vectors enables to reduce the large-dimensional feature vector into a small dimensional one by retaining most of the relevant information. In these works, the i-vectors are exclusively extracted from short-term cepstral features. Based on these studies, we propose the extraction of i-vectors from short-term cepstral, and long-term voice-quality, prosodic and GNE features. Thus, this work explores the the suitability of applying i-vector modeling techniques based on shortand long-term speech features within the frame of speaker diarization. Firstly, speech clusters generated by Viterbi segmentation are modeled by two sets of i-vectors. While the first i-vector represents the distribution of the commonly used short-term Mel Frequency Cepstral Coefficients, the second one depicts a selection of voice quality, prosodic and GNE features. In order to combine both the shortand long-term speech features, the cosine-distance and PLDA scores of these two i-vectors extracted from the corresponding features are linearly weighted to obtain a unique similarity score. The final fused score is used as speaker clustering distance.
102 Chapter 7. Conclusions and Future works The experimental results show the suitability of combining both sources of information within the i-vector space. Firstly, the experimental results show that i-vector based clustering techniques based on shortand long-term features provide better results more than using only the short-term features. Secondly, the results show that both i-vector based cosine and PLDA clustering techniques provide a substantial relative DER improvement more than GMM based BIC clustering. Furthermore, the results manifest that that i-vector based PLDA clustering technique provides better relative DER improvement more than i-vector based cosine clustering technique. Finally, the experimental results show the usefulness of GNE features in i-vector based PLDA clustering techniques. The addition of GNE features does not improve the results both in GMM based BIC and i-vector based cosine distance clustering techniques. The results of the experiments manifest the usefulness of i-vector based clustering technique based on shortand long-term speech features within in the framework of speaker diarization. 7.2 Future Research Lines The work performed in this thesis may be used as a guide for future research lines in speaker diarization. The possible future lines that can be continued from our work are outlined as follows: Firstly, the proposed voice-quality long-term features have been successfully applied to detect only single speaker both in the proposed GMM and i-vector based speaker diarization systems. Therefore, it is worth to explore the impact of the proposed voicequality features to detect overlapping speeches both in the proposed GMM and i-Vector based speaker diarization systems. Since speaker tracking and speaker diarization are really close to each other and generally share some key processing components, the proposed long-term features can also be applied in GMM and i-vector based speaker tracking systems. Furthermore, it is worth to explore the impact of the proposed voice-quality features in cross-show speaker diarization where reappearing speakers across shows have to be labeled with the same speaker identity. One of the main issues in speaker diarization is the substantial DER differences among different shows. One of the possible reasons is the threshold value estimated for the stopping criterion. We have suggested a semi-automatic stopping criteria that is the
7.2. Future Research Lines 103 same for all shows. It is also worth to see impact of using an automatic stopping criterion threshold value that varies per iteration and recording in the proposed systems. Finally, Deep Neural Networks (DNNs) have recently been successfully applied in speaker diarization systems. The DNNs can also be applied in the proposed system by replacing the log-likelihood scores of HMM in segmentation and the distance-metrics of clustering (BIC, cosine distance and PLDA) by the posterior probabilities the DNN.
Appendices 105
112 Bibliography [Dellwo et al., 2007] Dellwo, V., Huckvale, M., and Ashby, M. (2007). How is individuality expressed in voice? an introduction to speech production and description for speaker classification. Speaker Classification I, pages 1–20. [Dempster et al., 1977] Dempster, A. P., Laird, N. M., and Rubin, D. B. (1977). Maximum likelihood from incomplete data via the em algorithm. Journal of the royal statistical society. Series B (methodological), pages 1–38. [Desplanques et al., 2016] Desplanques, B., Demuynck, K., and Martens, J.-P. (2016). Soft vad in factor analysis based speaker segmentation of broadcast news. Odyssey Speaker and Language Recognition Workshop, pages 158–165. [Duda and Hart, 1973] Duda, R. and Hart, P. (1973). Pattern classification and scene analysis. John Wiley. [Duda et al., 2001] Duda, R. O., Hart, P. E., and Stork, D. G. (2001). “Pattern Classification”. Wiley & Sons, Inc., New York, NY, USA, 2nd edition. [Dutoit, 1997] Dutoit, T. (1997). An Introduction to Text-to-Speech Synthesis, volume 3. Springer Science & Business Media. [Farrs et al., 2006] Farrs, M., Garde, A., Ejarque, P., Luque, J., and Hernando, J. (2006). On the fusion of prosody, voice spectrum and face features for multimodal person verification. In Interspeech. [Farr´us et al., 2007] Farr´us, M., Hernando, J., and Ejarque, P. (2007). Jitter and shimmer measurements for speaker recognition. In Interspeech, pages 778–781. [Franco-Pedroso et al., 2010] Franco-Pedroso, J., Lopez-Moreno, I., Toledano, D. T., and Gonzalez-Rodriguez, J. (2010). ATVS-UAM system description for the audio segmentation and speaker diarization Albayzin 2010 evaluation. In FALA VI Jornadas en Tecnologa del Habla and II Iberian SLTech Workshop, pages 415–418. [Fredouille and Senay, 2006] Fredouille, C. and Senay, G. (2006). Technical improvements of the e-hmm based speaker diarization system for meeting records. In MLMI, volume 4299, pages 359–370. Springer. [Friedland et al., 2009] Friedland, G., Vinyals, O., Huang, Y., and Muller, C. (2009). Prosodic and other long-term features for speaker diarization. in IEEE Transactions on Audio, Speech, and Language Processing, 17(5):985–993. [Furui, 1981] Furui, S. (1981). Cepstral analysis technique for automatic speaker verification. IEEE Transactions on Acoustics, Speech, and Signal Processing, 29(2):254– 272.
Bibliography 113 [Furui, 2004] Furui, S. (2004). Fifty years of progress in speech and speaker recognition. The Journal of the Acoustical Society of America, 116(4):2497–2498. [Garcia-Romero and Espy-Wilson, 2011] Garcia-Romero, D. and Espy-Wilson, C. Y. (2011). Analysis of i-vector length normalization in speaker recognition systems. In Interspeech, volume 2011, pages 249–252. [Gauvain et al., 1999] Gauvain, J.-L., Lamel, L., Adda, G., and Jardino, M. (1999). The LIMSI 1998 Hub-4E transcription system. In Proc. DARPA Broadcast News Workshop, pages 99–104. [Ghahabi and Hernando, 2014] Ghahabi, O. and Hernando, J. (2014). Deep belief networks for i-vector based speaker recognition. In Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on, pages 1700–1704. IEEE. [Godino-Llorente et al., 2010] Godino-Llorente, J. I., Osma-Ruiz, V., S´aenz-Lech´on, N., G´omez-Vilda, P., Blanco-Velasco, M., and Cruz-Rold´an, F. (2010). The effectiveness of the glottal to noise excitation ratio for the screening of voice disorders. Journal of Voice, 24(1):47–56. [Han and Narayanan, 2008] Han, K. J. and Narayanan, S. S. (2008). Novel inter-cluster distance measure combining GLR and ICR for improved agglomerative hierarchical speaker clustering. In Acoustics, Speech and Signal Processing, 2008. ICASSP 2008. IEEE International Conference on, pages 4373–4376. IEEE. [Hinton et al., 2012] Hinton, G., Deng, L., Yu, D., Dahl, G. E., Mohamed, A.-r., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T. N., et al. (2012). Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Processing Magazine, 29(6):82–97. [Huijbregts et al., 2012] Huijbregts, M., van Leeuwen, D. A., and Wooters, C. (2012). Speaker diarization error analysis using oracle components. in IEEE Transactions on Audio, Speech, and Language Processing, 20@inproceedingssiegler1997automatic, title=Automatic segmentation, classification and clustering of broadcast news audio, author=Siegler, Matthew A and Jain, Uday and Raj, Bhiksha and Stern, Richard M, booktitle=Proc. DARPA speech recognition workshop, volume=1997, year=1997 (2):393–403. [Imseng and Friedland, 2010] Imseng, D. and Friedland, G. (2010). Tuning-robust initialization methods for speaker diarization. in IEEE Transactions on Audio, Speech, and Language Processing, 18(8):2028 –2037.
114 Bibliography [Jiang et al., 2012] Jiang, Y., Lee, K.-A., Tang, Z., Ma, B., Larcher, A., and Li, H. (2012). PLDA modeling in i-vector and supervector space for speaker verification. In Interspeech, pages 1680–1683. [Johnstone and Scherer, 1999] Johnstone, T. and Scherer, K. R. (1999). The effects of emotions on voice quality. In Proceedings of the XIVth International Congress of Phonetic Sciences, pages 2029–2032. University of California, Berkeley. [Jothilakshmi et al., 2009] Jothilakshmi, S., Ramalingam, V., and Palanivel, S. (2009). Speaker diarization using autoassociative neural networks. Engineering Applications of Artificial Intelligence, 22(4):667–675. [Junqua et al., 1994] Junqua, J.-C., Mak, B., and Reaves, B. (1994). A robust algorithm for word boundary detection in the presence of noise. in IEEE Transactions on speech and audio processing, 2(3):406–412. [Kemp et al., 2000] Kemp, T., Schmidt, M., Westphal, M., and Waibel, A. (2000). Strategies for automatic segmentation of audio data. In Acoustics, Speech, and Signal Processing, 2000. ICASSP’00. Proceedings. 2000 IEEE International Conference on, volume 3, pages 1423–1426. IEEE. [Kenny, 2010] Kenny, P. (2010). Bayesian speaker verification with heavy-tailed priors. In Odyssey, page 14. [Kenny et al., 2008] Kenny, P., Ouellet, P., Dehak, N., Gupta, V., and Dumouchel, P. (2008). A study of inter-speaker variability in speaker verification. in IEEE Transactions on Audio, Speech, and Language Processing, 16(5):980–988. [Kenny et al., 2010] Kenny, P., Reynolds, D., and Castaldo, F. (2010). Diarization of telephone conversations using factor analysis. in IEEE Journal of Selected Topics in Signal Processing, 4(6):1059–1070. [Kinnunen and Li, 2010] Kinnunen, T. and Li, H. (2010). An overview of textindependent speaker recognition: from features to supervectors. Speech communication, 52(1):12–40. [Kittler et al., 1998] Kittler, J., Hatef, M., Duin, R. P., and Matas, J. (1998). On combining classifiers. IEEE transactions on pattern analysis and machine intelligence, 20(3):226–239. [Kreiman and Gerratt, 2005] Kreiman, J. and Gerratt, B. R. (2005). Perception of aperiodicity in pathological voice. The Journal of the Acoustical Society of America, 117(4):2201–2211.
Bibliography 115 [Kubala et al., 1997] Kubala, F., Jin, H., Matsoukas, S., Nguyen, L., Schwartz, R., and Makhoul, J. (1997). The 1996 BBN byblos HUB-4 transcription system. In Proceedings of the 1997 DARPA Speech Recognition Workshop, pages 90–93. [Kumar et al., 2011] Kumar, K., Kim, C., and Stern, R. M. (2011). Delta-spectral cepstral coefficients for robust speech recognition. In Acoustics, Speech and Signal Processing (ICASSP), 2011 IEEE International Conference on, pages 4784–4787. IEEE. [Lamel et al., 1981] Lamel, L., Rabiner, L., Rosenberg, A., and Wilpon, J. (1981). An improved endpoint detector for isolated word recognition. in IEEE Transactions on Acoustics, Speech, and Signal Processing, 29(4):777–785. [Larcher et al., 2013] Larcher, A., Bonastre, J. F., Fauve, B. G. B., Lee, K., L´evy, C., Li, H., Mason, J. S. D., and Parfait, J. (2013). ALIZE 3.0 - open source toolkit for state-of-the-art speaker recognition. In Interspeech. [Lee et al., 2009] Lee, H., Pham, P., Largman, Y., and Ng, A. Y. (2009). Unsupervised feature learning for audio classification using convolutional deep belief networks. In Advances in neural information processing systems, pages 1096–1104. [Li et al., 2007] Li, X., Tao, J., Johnson, M. T., Soltis, J., Savage, A., Leong, K. M., and Newman, J. D. (2007). Stress and emotion classification using jitter and shimmer features. In Acoustics, Speech and Signal Processing, 2007. ICASSP 2007. IEEE International Conference on, volume 4, pages IV–1081. IEEE. [Linville, 1995] Linville, S. E. (1995). Vocal aging. Current Opinion in Otolaryngology & Head and Neck Surgery, 3(3):183–187. [Luque, 2012] Luque, J. (2012). Speaker diarization and tracking in multiple-sensor environments. PhD thesis, Universitat Polit`ecnica de Catalunya, Barcelona, Spain. [Luque et al., 2008] Luque, J., Segura, C., and Hernando, J. (2008). Clustering initialization based on spatial information for speaker diarization of meetings. In Interspeech, pages 383–386. [McLachlan and Basford, 1988] McLachlan, G. J. and Basford, K. E. (1988). Mixture models: Inference and applications to clustering, volume 84. Marcel Dekker. [Meignier et al., 2006] Meignier, S., Moraru, D., Fredouille, C., Bonastre, J.-F., and Besacier, L. (2006). Step-by-step and integrated approaches in broadcast news speaker diarization. Computer Speech & Language, 20(2):303–330. [Memon et al., 2009] Memon, S., Lech, M., and Maddage, N. (2009). Speaker verification based on different vector quantization techniques with gaussian mixture models. In Third International Conference on Network and System Security, pages 403–408.
116 Bibliography [Michaelis et al., 1998a] Michaelis, D., Fr¨ohlich, M., and Strube, H. W. (1998a). Selection and Combination of Acoustic Features for the description of Pathologic Voices. The Journal of the Acoustical Society of America, 103(3):1628–1639. [Michaelis et al., 1998b] Michaelis, D., Fr¨ohlich, M., Strube, H. W., Kruse, E., Story, B., and Titze, I. R. (1998b). Some simulations concerning jitter and shimmer measurement. In 3rd International Workshop on Advances in Quantitative Laryngoscopy, Aachen, Germany, pages 744–754. [Michaelis et al., 1997] Michaelis, D., Gramss, T., and Strube, H. W. (1997). Glottal-tonoise excitation ratio–a new measure for describing pathological voices. Acta Acustica united with Acustica, 83(4):700–706. [Minematsu et al., 2002] Minematsu, N., Sekiguchi, M., and Hirose, K. (2002). Automatic estimation of one’s age with his/her speech based upon acoustic modeling techniques of speakers. In Acoustics, Speech, and Signal Processing (ICASSP), 2002 IEEE International Conference on, volume 1, pages I–137. IEEE. [Mohamed et al., 2012] Mohamed, A.-r., Dahl, G. E., and Hinton, G. (2012). Acoustic modeling using deep belief networks. IEEE Transactions on Audio, Speech, and Language Processing, 20(1):14–22. [Nandakumar et al., 2008] Nandakumar, K., Chen, Y., Dass, S. C., and Jain, A. (2008). Likelihood ratio-based biometric score fusion. IEEE transactions on pattern analysis and machine intelligence, 30(2):342–347. [Nandakumar et al., 2009] Nandakumar, K., Jain, A., and Ross, A. (2009). Fusion in multibiometric identification systems: What about the missing data? Advances in Biometrics, pages 743–752. [Nguyen, 2010] Nguyen, P. T. (2010). Automatic Speaker Classification Based on Voice Characteristics. University of Canberra. [Noll, 1967] Noll, A. M. (1967). Cepstrum pitch determination. The journal of the acoustical society of America, 41(2):293–309. [Noll, 1969] Noll, A. M. (1969). Pitch determination of human speech by the harmonic product spectrum, the harmonic sum spectrum, and a maximum likelihood estimate. In Proceedings of the symposium on computer processing communications, volume 779. [Nosratighods et al., 2006] Nosratighods, M., Ambikairajah, E., and Epps, J. (2006). Speaker verification using a novel set of dynamic features. In Pattern Recognition, 2006. ICPR 2006. 18th International Conference on, volume 4, pages 266–269. IEEE.
Bibliography 117 [Pardo et al., 2007] Pardo, J., Anguera, X., and Wooters, C. (2007). Speaker diarization for multiple-distant-microphone meetings using several sources of information. IEEE Transactions on Computers, 56(9):1212–1224. [Pardo et al., 2006] Pardo, J. M., Anguera, X., and Wooters, C. (2006). Speaker diarization for multiple distant microphone meetings: mixing acoustic features and inter-channel time differences. In Interspeech. [Pelecanos and Sridhara, 2001] Pelecanos, J. and Sridhara, S. (2001). Feature warping for robust speaker verification. In International Speech Communication Association (ISCA). [Pickett and Morris, 2000] Pickett, J. and Morris, S. R. (2000). The acoustics of speech communication: Fundamentals, speech perception theory, and technology. The Journal of the Acoustical Society of America, 108(4):1373–1374. [Pop et al., 2007] Pop, P., Lupu, E., and Roman, M. (2007). Pathological voice assessment. In 1st International Conference on Advancements of Medicine and Health Care through Technology, MediTech2007. [Prazak and Silovsky, 2011] Prazak, J. and Silovsky, J. (2011). Speaker diarization using plda-based speaker clustering. In Intelligent Data Acquisition and Advanced Computing Systems (IDAACS), 2011 IEEE 6th International Conference on, volume 1, pages 347–350. IEEE. [Prince and Elder, 2007] Prince, S. J. and Elder, J. H. (2007). Probabilistic linear discriminant analysis for inferences about identity. In Computer Vision, 2007. ICCV 2007. IEEE 11th International Conference on, pages 1–8. IEEE. [Reynolds, 2002] Reynolds, D. A. (2002). An overview of automatic speaker recognition technology. In Acoustics, speech, and signal processing (ICASSP), 2002 IEEE international conference on, volume 4, pages IV–4072. IEEE. [Reynolds and Torres-Carrasquillo, 2004] Reynolds, D. A. and Torres-Carrasquillo, P. (2004). The MIT Lincoln Laboratory RT-04F diarization systems: Applications to broadcast audio and telephone conversations. Technical report, DTIC Document. [Rumelhart et al., 1988] Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1988). Learning representations by back-propagating errors. Cognitive modeling, 5(3):1. [Rusz et al., 2011] Rusz, J., Cmejla, R., Ruzickova, H., and Ruzicka, E. (2011). Quantitative acoustic measurements for characterization of speech and voice disorders in early untreated parkinson’s disease. The journal of the Acoustical Society of America, 129(1):350–367.
118 Bibliography [Sadeghi Naini and Homayounpour, 2006] Sadeghi Naini, A. and Homayounpour, M. (2006). Speaker age interval and sex identification based on Jitters, Shimmers and Mean MFCC using supervised and unsupervised discriminative classification methods. In Signal Processing, 2006 8th International Conference on, volume 1. IEEE. [S´aenz Lech´on et al., 2009] S´aenz Lech´on, N., Osma Ruiz, V., Fraile Mu˜noz, R., Godino Llorente, J. I., and G´omez Vilda, P. (2009). Screening voice disorders with the glottal to noise excitation ratio. [Schotz, 2001] Schotz, S. (2001). A perceptual study of speaker age. WORKING PAPERS-LUND UNIVERSITY DEPARTMENT OF LINGUISTICS, pages 136–139. [Schroeder, 1968] Schroeder, M. R. (1968). Period histogram and product spectrum: New methods for fundamental-frequency measurement. The Journal of the Acoustical Society of America, 43(4):829–834. [Sell and Garcia-Romero, 2014] Sell, G. and Garcia-Romero, D. (2014). Speaker diarization with PLDA i-vector scoring and unsupervised calibration. In Spoken Language Technology Workshop (SLT), 2014 IEEE, pages 413–417. IEEE. [Senoussaoui et al., 2013] Senoussaoui, M., Kenny, P., Dumouchel, P., and Stafylakis, T. (2013). Efficient iterative mean shift based cosine dissimilarity for multi-recording speaker clustering. In Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on, pages 7712–7715. IEEE. [Shriberg, 2007] Shriberg, E. (2007). Higher-level features in speaker recognition. Speaker Classification I, pages 241–259. [Shum et al., 2011] Shum, S., Dehak, N., Chuangsuwanich, E., Reynolds, D. A., and Glass, J. R. (2011). Exploiting intra-conversation variability for speaker diarization. In interspeech, volume 11, pages 945–948. [Shum et al., 2012] Shum, S., Dehak, N., and Glass, J. (2012). On the use of spectral and iterative methods for speaker diarization. In Thirteenth Annual Conference of the International Speech Communication Association. [Siegler et al., 1997] Siegler, M. A., Jain, U., Raj, B., and Stern, R. M. (1997). Automatic segmentation, classification and clustering of broadcast news audio. In Proc. DARPA speech recognition workshop, volume 1997. [Silovsky and Prazak, 2012] Silovsky, J. and Prazak, J. (2012). Speaker diarization of broadcast streams using two-stage clustering based on i-vectors and cosine distance scoring. In Acoustics, Speech and Signal Processing (ICASSP), 2012 IEEE International Conference on, pages 4193–4196. IEEE.
Bibliography 119 [Sim and Lee, 2010] Sim, K. C. and Lee, K.-A. (2010). Adaptive score fusion using weighted logistic linear regression for spoken language recognition. In Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on, pages 5018–5021. IEEE. [Sinha et al., 2005] Sinha, R., Tranter, S. E., Gales, M. J., and Woodland, P. C. (2005). The cambridge university march 2005 speaker diarisation system. In Interspeech, pages 2437–2440. [Slyh et al., 1999] Slyh, R. E., Nelson, W. T., and Hansen, E. G. (1999). Analysis of mrate, shimmer, jitter, and f/sub 0/contour features across stress and speaking style in the SUSAS database. In Acoustics, Speech, and Signal Processing, 1999. Proceedings., 1999 IEEE International Conference on, volume 4, pages 2091–2094. IEEE. [Stafylakis et al., 2012] Stafylakis, T., Kenny, P., Senoussaoui, M., and Dumouchel, P. (2012). Preliminary investigation of boltzmann machine classifiers for speaker recognition. In Odyssey Speaker and Language Recognition Workshop, pages 109–116. [Styler, 2013] Styler, W. (2013). Using praat for linguistic research. University of Colorado at Boulder Phonetics Lab. [Tranter and Reynolds, 2006] Tranter, S. E. and Reynolds, D. A. (2006). An overview of automatic speaker diarization systems. IEEE Transactions on audio, speech, and language processing, 14(5):1557–1565. [Van Leeuwen and Koneˇcn`y, 2008] Van Leeuwen, D. and Koneˇcn`y, M. (2008). Progress in the amida speaker diarization system for meeting data. Multimodal Technologies for Perception of Humans, pages 475–483. [Vaquero Avil´es-Casco, 2011] Vaquero Avil´es-Casco, C. (2011). Robust diarization for speaker characterization (Diarizaci´on robusta para caracterizaci´on de locutores). PhD thesis, University of Zaragoza, Zaragoza, Spain. [Ververidis and Kotropoulos, 2006] Ververidis, D. and Kotropoulos, C. (2006). Emotional speech recognition: Resources, features, and methods. Speech communication, 48(9):1162–1181. [Vijayasenan and Valente, 2012] Vijayasenan, D. and Valente, F. (2012). Diartk: An open source toolkit for research in multistream speaker diarization and its application to meetings recordings. In Interspeech, pages 2170–2173. [Vijayasenan et al., 2007] Vijayasenan, D., Valente, F., and Bourlard, H. (2007). Agglomerative information bottleneck for speaker diarization of meetings data. In Automatic Speech Recognition & Understanding, 2007. ASRU. IEEE Workshop on, pages 250–255. IEEE.
120 Bibliography [Vijayasenan et al., 2009] Vijayasenan, D., Valente, F., and Bourlard, H. (2009). An information theoretic approach to speaker diarization of meeting data. in IEEE Transactions on Audio, Speech, and Language Processing, 17(7):1382–1393. [Wagner, 2013] Wagner, I. (2013). A new jitter-algorithm to quantify hoarseness: an exploratory study. International Journal of Speech Language and the Law, 2(1):18–27. [Wang and Shen, 1999] Wang, X.-G. and Shen, H. C. (1999). Multiple hypothesis testing fusion method for multisensor systems. In Intelligent Robots and Systems, 1999. IROS’99. Proceedings. 1999 IEEE/RSJ International Conference on, volume 2, pages 1008–1013. IEEE. [Wertzner et al., 2005] Wertzner, H. F., Schreiber, S., and Amaro, L. (2005). Analysis of fundamental frequency, jitter, shimmer and vocal intensity in children with phonological disorders. Brazilian journal of otorhinolaryngology, 71(5):582–588. [Willsky and Jones, 1976] Willsky, A. S. and Jones, H. L. (1976). A generalized likelihood ratio approach to the detection and estimation of jumps in linear systems. in IEEE Transactions on Automatic Control, 21(1):108–112. [Wittig and M¨uller, 2003] Wittig, F. and M¨uller, C. (2003). Implicit feedback for useradaptive systems by analyzing the users’ speech. [Wooters et al., 2004] Wooters, C., Fung, J., Peskin, B., and Anguera, X. (2004). Towards robust speaker segmentation: The ICSI-SRI fall 2004 diarization system. In RT-04F Workshop, volume 23, page 23. [Wooters and Huijbregts, 2008] Wooters, C. and Huijbregts, M. (2008). The ICSI RT07s speaker diarization system. Multimodal Technologies for Perception of Humans, pages 509–519. [Woubie et al., 2014] Woubie, A., Luque, J., and Hernando, J. (2014). Jitter and shimmer measuremetns for speaker diarization. In VII Jornadas en Tecnolog´ıa del Habla and III Iberian SLTech Workshop: proceedings: November 19-21, 2014: Escuela de Ingenier´ıa en Telecomunicaci´on y Electr´onica Universidad de Las Palmas de Gran Canaria: Las Palmas de Gran Canaria, Spain, pages 21–30. [Woubie et al., 2015] Woubie, A., Luque, J., and Hernando, J. (2015). Using voicequality measurements with prosodic and spectral features for speaker diarization. In Interspeech, pages 3100–3104. [Woubie et al., 2016a] Woubie, A., Luque, J., and Hernando, J. (2016a). Improving i-vector and plda based speaker clustering with long-term features. In Interspeech
Bibliography 121 San Francisco, USA, pages 372–376. International Speech Communication Association (ISCA). [Woubie et al., 2016b] Woubie, A., Luque, J., and Hernando, J. (2016b). Short-and long-term speech features for hybrid hmm-i-vector based speaker diarization system. In Odyssey Speaker and Language Recognition Workshop. [Xu and Zhang, 2010] Xu, Y. and Zhang, D. (2010). Represent and fuse bimodal biometric images at the feature level: complex-matrix-based fusion scheme. Optical Engineering, 49(3):037002–037002. [Yao et al., 2012] Yao, K., Yu, D., Seide, F., Su, H., Deng, L., and Gong, Y. (2012). Adaptation of context-dependent deep neural networks for automatic speech recognition. In Spoken Language Technology Workshop (SLT), 2012 IEEE, pages 366–369. IEEE. [Yella, 2015] Yella, S. H. (2015). Speaker diarization of spontaneous meeting room conversations. PhD thesis, EPFL, Lausanne, Switzerland. [Yella and Stolcke, 2015] Yella, S. H. and Stolcke, A. (2015). A comparison of neural network feature transforms for speaker diarization. In Interspeech, pages 3026–3030. [Yella et al., 2014] Yella, S. H., Stolcke, A., and Slaney, M. (2014). Artificial neural network features for speaker diarization. In Spoken Language Technology Workshop (SLT), 2014 IEEE, pages 402–406. IEEE. [Young and Young, 1993] Young, S. J. and Young, S. (1993). The HTK hidden Markov model toolkit: Design and philosophy. University of Cambridge, Department of Engineering. [Zelen´ak and Hernando, 2011] Zelen´ak, M. and Hernando, J. (2011). The detection of overlapping speech with prosodic features for speaker diarization. In Interspeech, pages 1041–1044. [Zhang, 2009] Zhang, D. (2009). Advanced pattern recognition technologies with applications to biometrics. IGI Global. [Zhang, 2008] Zhang, S. (2008). Emotion recognition in chinese natural speech by combining prosody and voice quality features. Advances in Neural Networks-ISNN 2008, pages 457–464. [Zheng et al., 2001] Zheng, F., Zhang, G., and Song, Z. (2001). Comparison of different implementations of MFCC. Journal of Computer Science and Technology, 16(6):582– 589.