scieee AI-readable full text Open interactive document viewer

Repositorio Institucional de Documentos

Abstract

La motivación de esta tesis es la necesidad de soluciones robustas al problema de diarización. Estas técnicas de diarización deben proporcionar valor añadido a la creciente cantidad disponible de datos multimedia mediante la precisa discriminación de los locutores presentes en la señal de audio. Desafortunadamente, hasta tiempos recientes este tipo de tecnologías solamente era viable en condiciones restringidas, quedando por tanto lejos de una solución general. <br />Las razones detrás de las limitadas prestaciones de los sistemas de diarización son múltiples. La primera causa a tener en cuenta es la alta complejidad de la producción de la voz humana, en particular acerca de los procesos fisiológicos necesarios para incluir las características discriminativas de locutor en la señal de voz. Esta complejidad hace del proceso inverso, la estimación de dichas características a partir del audio, una tarea ineficiente por medio de las técnicas actuales del estado del arte. Consecuentemente, en su lugar deberán tenerse en cuenta aproximaciones. Los esfuerzos en la tarea de modelado han proporcionado modelos cada vez más elaborados, aunque no buscando la explicación última de naturaleza fisiológica de la señal de voz. En su lugar estos modelos aprenden relaciones entre la señales acústicas a partir de un gran conjunto de datos de entrenamiento. El desarrollo de modelos aproximados genera a su vez una segunda razón, la variabilidad de dominio. Debido al uso de relaciones aprendidas a partir de un conjunto de entrenamiento concreto, cualquier cambio de dominio que modifique las condiciones acústicas con respecto a los datos de entrenamiento condiciona las relaciones asumidas, pudiendo causar fallos consistentes en los sistemas.<br />Nuestra contribución a las tecnologías de diarización se ha centrado en el entorno de radiodifusión. Este dominio es actualmente un entorno todavía complejo para los sistemas de diarización donde ninguna simplificación de la tarea puede ser tenida en cuenta. Por tanto, se deberá desarrollar un modelado eficiente del audio para extraer la información de locutor y como inferir el etiquetado correspondiente. Además, la presencia de múltiples condiciones acústicas debido a la existencia de diferentes programas y/o géneros en el domino requiere el desarrollo de técnicas capaces de adaptar el conocimiento adquirido en un determinado escenario donde la información está disponible a aquellos entornos donde dicha información es limitada o sencillamente no disponible.<br />Para este propósito el trabajo desarrollado a lo largo de la tesis se ha centrado en tres subtareas: caracterización de locutor, agrupamiento y adaptación de modelos. La primera subtarea busca el modelado de un fragmento de audio para obtener representaciones precisas de los locutores involucrados, poniendo de manifiesto sus propiedades discriminativas. En este área se ha llevado a cabo un estudio acerca de las actuales estrategias de modelado, especialmente atendiendo a las limitaciones de las representaciones extraídas y poniendo de manifiesto el tipo de errores que pueden generar. Además, se han propuesto alternativas basadas en redes neuronales haciendo uso del conocimiento adquirido. La segunda tarea es el agrupamiento, encargado de desarrollar estrategias que busquen el etiquetado óptimo de los locutores. La investigación desarrollada durante esta tesis ha propuesto nuevas estrategias para estimar el mejor reparto de locutores basadas en técnicas de subespacios, especialmente PLDA. Finalmente, la tarea de adaptación de modelos busca transferir el conocimiento obtenido de un conjunto de entrenamiento a dominios alternativos donde no hay datos para extraerlo. Para este propósito los esfuerzos se han centrado en la extracción no supervisada de información de locutor del propio audio a diarizar, sinedo posteriormente usada en la adaptación de los modelos involucrados.<br /> <br /> Viñals Bailo, Ignacio; Ortega Giménez, Alfonso

Full text

2021 77 Ignacio Viñals Bailo Advances in Subspacebased Solutions for Diarization in the Broadcast Domain Director/es Ortega Giménez, Alfonso © Universidad de Zaragoza Servicio de Publicaciones ISSN 2254-7606 Ignacio Viñals Bailo ADVANCES IN SUBSPACE-BASED SOLUTIONS FOR DIARIZATION IN THE BROADCAST DOMAIN Director/es Ortega Giménez, Alfonso Tesis Doctoral Autor 2020 UNIVERSIDAD DE ZARAGOZA Escuela de Doctorado Programa de Doctorado en Tecnologías de la Información y Comunicaciones en Redes Móviles Repositorio de la Universidad de Zaragoza – Zaguan http://zaguan.unizar.es UNIVERSIDAD DE ZARAGOZA TESIS DOCTORAL - INGENIERÍA DE TELECOMUNICACIÓN Advances in Subspace-based Solutions for Diarization in the Broadcast Domain Author: Ignacio Viñals Bailo Supervisor: Alfonso Ortega Giménez DEPARTAMENTO DE INGENIERÍA ELECTRÓNICA Y COMUNICACIONES ESCUELA DE INGENIERÍA Y ARQUITECTURA April, 2020 A mis padres The human voice is the most perfect instrument of all. Arvo Pärt The most important questions of life are, for the most part, really only problems of probability Pierre-Simon Laplace Research is creating new knowledge. Neil Armstrong The human brain is an incredible pattern-matching machine. Jeff Bezos 4 Acknowledgements It has been five long years since I made the decision to start a PhD programme. Along all this time I have been fortunate enough to meet, collaborate and be helped as well as supported by many people, without whom this thesis would not be available today. These lines are dedicated to all of them. First and foremost, I want to dedicate some lines to my parents. I want to thank them for being on my side from the very beginning. They always offered me their support since I decided to become a researcher. For the last five years they have cheered me up during bad times and kept my feet on the ground during those limited successful occasions. This work could not be possible without their contribution. Besides, I also must thank Alfonso Ortega for giving me the chance to grow up professionally and personally. He gave me the chance to discover the researcher career when I was undergraduate, and offered me the opportunity to keep on developing myself with ViVoLAB group. For the last five years he has become a friend apart from a supervisor, who guided me along this difficult learning process. By means of our meetings he helped me to discover some of the best ideas while other times he simply made me aware that sometimes I could not see the forest for the trees. Apart from Alfonso Ortega, ViVoLAB group is also full of wonderful people who deserve their own mention. Some of my kindest memories along my PhD years include Eduardo Lleida. He always tried to build a group based on friendship relationships rather than simply professional ones. Besides, it is praiseworthy how this treatment also included ViVoLAB alumni all over the world. Another important collaborator in this thesis is Antonio Miguel. His valuable expertise in subspace models and neural networks were capital for the work done in this thesis. Besides, the theoretical discussions we had about these topics were also very enriching, opening my eyes about unseen lines of research to explore. I also want to acknowledge Johns Hopkins university staff, specially Najim Dehak and his 5 Contents 1 Introduction 1 1.1 Motivation of the work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 1.2 Objectives and Methodology . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 1.3 Thesis organization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 I Diarization Basic Knowledge 7 2 Diarization State of the Art 9 2.1 Introduction .................................... 9 2.2 Main diarization strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.2.1 Bottom-Up diarization systems . . . . . . . . . . . . . . . . . . . . . 12 2.3 Acoustic features for diarization . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.4 Audio segmentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.4.1 Metric-based segmentation . . . . . . . . . . . . . . . . . . . . . . . . 17 2.4.1.1 Bayesian Information Criterion (BIC) . . . . . . . . . . . . . 18 2.4.1.2 Kullback-Leibler Divergence (KL) . . . . . . . . . . . . . . 19 2.4.1.3 Deep Neural Networks (DNNs) . . . . . . . . . . . . . . . . 20 2.4.2 Model-based segmentation . . . . . . . . . . . . . . . . . . . . . . . . 20 2.5 Speaker characterization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 2.5.1 Early days . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 2.5.2 Model-based representations . . . . . . . . . . . . . . . . . . . . . . . 22 2.5.2.1 Gaussian Mixture Models (GMMs) . . . . . . . . . . . . . . 22 2.5.2.2 Support Vector Machines (SVM) . . . . . . . . . . . . . . . 23 2.5.2.3 Joint Factor Analysis (JFA) . . . . . . . . . . . . . . . . . . 24 iii CONTENTS 2.5.3 Embedded representations . . . . . . . . . . . . . . . . . . . . . . . . 25 2.5.3.1 I-vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 2.5.3.2 Hybrid i-vectors . . . . . . . . . . . . . . . . . . . . . . . . 26 2.5.3.3 DNN embeddings . . . . . . . . . . . . . . . . . . . . . . . 27 2.5.4 Probabilistic Linear Discriminant Analysis (PLDA) . . . . . . . . . . . 27 2.6 Clustering ..................................... 28 2.6.1 Hierarchical clustering . . . . . . . . . . . . . . . . . . . . . . . . . . 32 2.6.2 Statistical approaches . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 2.6.3 Other alternatives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 2.7 Performance metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 3 Analysis of Diarization in Broadcast Data 39 3.1 The diarization reference system . . . . . . . . . . . . . . . . . . . . . . . . . 39 3.2 Analysis of broadcast data . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 3.2.1 Multi-Genre Broadcast Challenge 2015 (MGB 2015) . . . . . . . . . . 42 3.2.2 Albayzín 2018 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43 3.2.3 Acoustic variability . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43 3.2.4 Variability in the speaker distribution . . . . . . . . . . . . . . . . . . 46 3.3 Evaluation of performance of the diarization reference system . . . . . . . . . 48 3.3.1 Evaluation of performance in MGB 2015 . . . . . . . . . . . . . . . . 48 3.3.2 Evaluation of performance in Albayzín 2018 . . . . . . . . . . . . . . 50 3.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52 3.4.1 The clustering approximation . . . . . . . . . . . . . . . . . . . . . . 52 3.4.2 The quality of the embeddings . . . . . . . . . . . . . . . . . . . . . . 52 3.4.3 The domain mismatch problem . . . . . . . . . . . . . . . . . . . . . 53 II The Clustering Problem 55 4 Clustering by means of Fully Bayesian PLDA 57 4.1 The Fully Bayesian PLDA clustering solution . . . . . . . . . . . . . . . . . . 57 4.1.1 The Fully Bayesian PLDA (FBPLDA) model . . . . . . . . . . . . . . 57 4.1.2 The clustering procedure . . . . . . . . . . . . . . . . . . . . . . . . . 60 4.1.3 Diarization using the FBPLDA model . . . . . . . . . . . . . . . . . . 62 4.2 Analysis of FBPLDA performance . . . . . . . . . . . . . . . . . . . . . . . . 65 4.2.1 Initialization impact . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 iv CONTENTS 4.2.2 Inference of the number of speakers . . . . . . . . . . . . . . . . . . . 66 4.2.3 Number of speakers vs DER . . . . . . . . . . . . . . . . . . . . . . . 69 4.2.4 Number of speakers vs ELBO . . . . . . . . . . . . . . . . . . . . . . 71 4.3 Alternative initializations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72 4.3.1 Computationally efficient initialization . . . . . . . . . . . . . . . . . 73 4.3.2 ELBO-based initialization choice criterion . . . . . . . . . . . . . . . 74 4.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 5 Uncertainty Propagation for Diarization 79 5.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 5.2 PLDA with Uncertainty Propagation (PLDAUP) . . . . . . . . . . . . . . . . . 80 5.2.1 PLDAUP in speaker recognition . . . . . . . . . . . . . . . . . . . . . 82 5.2.2 PLDAUP in speaker clustering . . . . . . . . . . . . . . . . . . . . . . 86 5.3 FBPLDA with Uncertainty Propagation (FBPLDAUP) . . . . . . . . . . . . . 89 5.4 Diarization of broadcast data with FBPLDAUP . . . . . . . . . . . . . . . . . 91 5.5 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 6 Tree-Based Clustering Approaches 97 6.1 Tree-based point of view for clustering . . . . . . . . . . . . . . . . . . . . . . 98 6.2 PLDA tree-based clustering . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100 6.2.1 PLDA-based model . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101 6.2.2 M-algorithm optimization . . . . . . . . . . . . . . . . . . . . . . . . 103 6.3 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 6.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112 III The Speaker Representation Problem 113 7 Study of embeddings for short utterances 115 7.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115 7.2 Short utterances as occluded utterances . . . . . . . . . . . . . . . . . . . . . 116 7.3 Formulation of the embedding extraction with short utterances . . . . . . . . . 118 7.3.1 General case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118 7.3.2 i-vector embeddings . . . . . . . . . . . . . . . . . . . . . . . . . . . 120 7.3.3 Short utterances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 7.4 Effects of the short utterances in i-vectors . . . . . . . . . . . . . . . . . . . . 122 v CONTENTS 7.5 Experiments & Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126 7.5.1 Experimental setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126 7.5.2 Baseline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 7.5.3 Reduction of the mismatch in α: Phonetic balance . . . . . . . . . . . 128 7.5.4 Enrollment-test distance vs log-likelihood ratio (llr) . . . . . . . . . . . 131 7.5.5 Enrollment-test distance vs performance (EER and minDCF) . . . . . . 133 7.5.6 Long-short vs Equalized Short-Short . . . . . . . . . . . . . . . . . . . 134 7.6 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136 8 DNNs embeddings for Diarization 137 8.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137 8.2 Hybrid i-vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138 8.2.1 Bottleneck Features (BNFs) . . . . . . . . . . . . . . . . . . . . . . . 139 8.2.2 Phonetic i-vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141 8.3 X-vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 142 8.4 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 144 8.4.1 Bottleneck Features (BNFs) . . . . . . . . . . . . . . . . . . . . . . . 145 8.4.2 Phonetic i-vectors & x-vectors . . . . . . . . . . . . . . . . . . . . . . 147 8.4.2.1 Speaker recognition . . . . . . . . . . . . . . . . . . . . . . 148 8.4.2.2 Broadcast diarization . . . . . . . . . . . . . . . . . . . . . 150 8.5 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153 IV The Model Adaptation Problem 155 9 Data-Efficient Domain Adaptation for PLDA Models 157 9.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157 9.2 Methods for domain mismatch reduction . . . . . . . . . . . . . . . . . . . . . 158 9.3 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 161 9.3.1 Independent unsupervised adaptation . . . . . . . . . . . . . . . . . . 162 9.3.2 Longitudinal unsupervised adaptation . . . . . . . . . . . . . . . . . . 163 9.3.3 Use of in-domain labeled data and semi-supervised adaptation . . . . . 164 9.4 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 166 vi CONTENTS V Conclusions & Future Work 167 10 Conclusions & Future work 169 10.1 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169 10.1.1 The clustering task . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169 10.1.2 The speaker characterization stage . . . . . . . . . . . . . . . . . . . . 170 10.1.3 Unsupervised domain adaptation research . . . . . . . . . . . . . . . . 171 10.2 Scientific Contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172 10.2.1 Book chapters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172 10.2.2 Papers published in journals included in the Journal Citation Reports (JCR) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172 10.2.3 Conference proceedings . . . . . . . . . . . . . . . . . . . . . . . . . 173 10.3 Future Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 173 VI Appendix 175 A Fully Bayesian PLDA with Uncertainty Propagation I A.1 Definitions ..................................... I A.2 Data ........................................ II A.3 Data conditional likelihood . . . . . . . . . . . . . . . . . . . . . . . . . . . . III A.3.1 P(Φi|yi,Xi,Θi,M). . . . . . . . . . . . . . . . . . . . . . . . . . . III A.3.2 P(Xi|yi,Θi,Φi,M). . . . . . . . . . . . . . . . . . . . . . . . . . . IV A.3.3 P(yi|Φi,Θi,M). . . . . . . . . . . . . . . . . . . . . . . . . . . . . IV A.4 Variational approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . V A.4.1 Joint probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . V A.4.2 Variational Bayes approximation . . . . . . . . . . . . . . . . . . . . . VI A.4.3 Optimal definition of q∗(Y,X). . . . . . . . . . . . . . . . . . . . . VI A.4.4 Optimal definition of q∗(Θ) . . . . . . . . . . . . . . . . . . . . . . . VII A.4.5 Optimal definition of q∗(πθ). . . . . . . . . . . . . . . . . . . . . . . VIII A.4.6 optimal definition of q∗˜ V. . . . . . . . . . . . . . . . . . . . . . . VIII A.4.7 Optimal definition of q∗(W). . . . . . . . . . . . . . . . . . . . . . . X A.4.8 Optimal definition of q∗(ε). . . . . . . . . . . . . . . . . . . . . . . XI A.4.9 Necessary Expectations . . . . . . . . . . . . . . . . . . . . . . . . . XI A.4.10 Variational Lower Bound . . . . . . . . . . . . . . . . . . . . . . . . . XIV A.5 Hyperparameter optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . XV vii CONTENTS viii List of Figures 1.1 Example of diarization results . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.2 Conceptual map of the studied topics in this Thesis . . . . . . . . . . . . . . . 5 2.1 Schematic of Bottom-Up and Top-Down diarization . . . . . . . . . . . . . . . 11 2.2 General schematic for a diarization system . . . . . . . . . . . . . . . . . . . . 12 2.3 Schematic of the MFCC extraction pipeline . . . . . . . . . . . . . . . . . . . 14 2.4 Scheme for a sliding window metric based segmentation . . . . . . . . . . . . 18 2.5 Schematic for an Agglomerative Hierarchical Clustering (AHC) performance . 32 3.1 Schematic of our baseline diarization system . . . . . . . . . . . . . . . . . . . 40 3.2 Section variability example. For 100 first embeddings from a Springwatch episode with SPLDA pairwise LLR similarity metric and Ground truth relationship. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45 3.3 Variability in the speaker distribution for MGB 2015, number of speakers per show and the proportion of speech for the most active speaker per show. . . . . 47 3.4 Variability in the speaker distribution for Albayzín 2018, number of speakers per show. and proportion of speech for the most active speaker per show. . . . . 48 3.5 Distribution of speech per speaker for two episodes: An episode with a dominant speaker and an episode with a more even speech distribution . . . . . . . . 49 4.1 Bayesian network of the Fully Bayesian PLDA . . . . . . . . . . . . . . . . . 58 4.2 Clustering schematic based on label initialization and FBPLDA resegmentation 61 4.3 Schematic for the diarization system based on the FBPLDA resegmentation . . 62 4.4 Analysis of ∆I=IORACLE −IHY P for shows in MGB 2015 with AHC and FBPLDA resegmentation diarization systems. . . . . . . . . . . . . . . . . . . 64 4.5 5-level dendrogram example. . . . . . . . . . . . . . . . . . . . . . . . . . . . 66 ix LIST OF FIGURES 4.6 Input/output relationship for the number of speakers with FBPLDA resegmentation. ....................................... 67 4.7 Lost speakers according to the relative number of speakers ∆Iin the initial partition Θ0.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69 4.8 DER (%) results for a) AHC and b) FBPLDA in terms of the relative number of speakers ∆I.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70 4.9 Distribution of the initialization with best DER in terms of the relative number of speakers. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71 4.10 Distribution of the initialization with bounded DER, a) 1% and b) 3%, in terms of the relative number of speakers. . . . . . . . . . . . . . . . . . . . . . . . . 72 4.11 Distribution of the partition with best ELBO in terms of relative speakers. . . . 73 4.12 Schematic of diarization based on the simultaneous evaluation of K different initializations. The final partition is selected by means of PELBO. . . . . . . . 75 5.1 Bayesian network for PLDA with Uncertainty Propagation (PLDAUP) . . . . . 81 5.2 DET curves with SPLDA for SRE10 corext-coreext det5 female with involved short utterances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85 5.3 DET curves with PLDAUP for SRE10 corext-coreext det5 female with involved short utterances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86 5.4 Impurity results for SPLDA and PLDAUP in SRE10 coreext-coreext det5 female chopped . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87 5.5 Impurity results for a) SPLDA and b) PLDAUP in SRE10 coreext-coreext det5 female chopped training with short utterances . . . . . . . . . . . . . . . . . . 88 5.6 Bayesian network for the Fully Bayesian PLDA with Uncertainty Propagation . 90 5.7 Histogram of DER variations between SPLDA and PLDAUP in MGB 2015 data. 92 5.8 Histogram of a) cluster and b) speaker impurities variations between SPLDA and PLDAUP initializations in MGB 2015 data. . . . . . . . . . . . . . . . . . 92 6.1 4-level tree clustering example . . . . . . . . . . . . . . . . . . . . . . . . . . 99 6.2 PLDA tree-based clustering Bayesian Network . . . . . . . . . . . . . . . . . 103 6.3 M-algorithm example for a clustering tree of depth 4 . . . . . . . . . . . . . . 104 6.4 Estimation step in a M-algorithm example for a clustering tree of depth 4 . . . 105 6.5 Maximization step in a M-algorithm example for a clustering tree of depth 4 . . 106 6.6 Analysis per show of ∆Iand DER(%) for AHC, FBPLDA and PLDA treebased clustering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 108 x LIST OF FIGURES 6.7 DER (%) results for the PLDA tree-based clustering with M-algorithm in Albayzín 2018 in terms of δ,ζand M. . . . . . . . . . . . . . . . . . . . . . . 109 6.8 DER relative results between Random order and Time order . . . . . . . . . . 111 7.1 Scenario of interest. a) Utterances red and blue in the feature domain, with the UBM components in green. b) Utterances red and blue in the i-vector domain. c) Projections of the GMM components in the i-vector domain for utterances . 123 7.2 Comparison of posterior distribution of the i-vectors with reference phoneme distribution. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124 7.3 Comparison of posterior distribution of i-vectors with modifications in the phoneme distribution α. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 7.4 Comparison of posterior distribution of i-vectors when two phonemes are not contributing and αc= 0 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126 7.5 DET curves for the scenarios Long-Long (blue), Long-Short and Short-Short Random (red continuous and dashed line respective), Long-Short and ShortShort Balanced (green continuous and dashed line respective)for SRE10 "coreextcoreext det5 female" experiment . . . . . . . . . . . . . . . . . . . . . . . . . 130 7.6 Trial score in terms of KL2 distance for the whole data pool. Represented the mean and the mean plus/minus the standard deviation . . . . . . . . . . . . . . 132 7.7 Evaluation metrics, EER (a) and minDCF (b) in terms of the KL2 distance. . . 133 7.8 DET curves for the scenarios Long-Short Random and Short-Short Equalized in SRE10 "coreext-coreext det5 female" . . . . . . . . . . . . . . . . . . . . . 135 7.9 Normalized distribution of scores for Target (blue) and Non-target(red) trials of scenarios Long-Short (continuous line) and Short-Short Equalized (dashed line).Experiment carried out with SRE10 "coreext-coreext det5 female". . . . . 135 8.1 Example of a Bottleneck Feature extractor DNN . . . . . . . . . . . . . . . . . 140 8.2 BNF pipeline from the original MFCCs up to Baum Welch statistics . . . . . . 140 8.3 Bayesian network of the phonetic i-vector . . . . . . . . . . . . . . . . . . . . 142 8.4 Phonetic i-vector pipeline from the original MFCCs up to Baum Welch statistics 143 8.5 X-vector architecture schematic . . . . . . . . . . . . . . . . . . . . . . . . . 143 8.6 DET curves for x-vectors in SRE10 with long and short utterances . . . . . . . 150 8.7 Distribution of the embedding first component in standard i-vectors, phonetic i-vectors and x-vectors for the training corpus in Albayzín 2018 . . . . . . . . 153 9.1 Schematic for the supervised and unsupervised adaptation . . . . . . . . . . . 159 xi LIST OF FIGURES 9.2 Schematic for unsupervised independent adaptation for the episodes n−1,n and n+ 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160 9.3 Schematic for unsupervised longitudinal adaptation for the episodes n−1,n and n+ 1.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160 9.4 Semi-supervised adaptation strategy based on the unsupervised independent adaptation approach for the episodes n−1,nand n+ 1.. . . . . . . . . . . . 160 9.5 Semi-supervised adaptation strategy based on the longitudinal unsupervised adaptation approach for the episodes n−1,nand n+ 1 . . . . . . . . . . . . 160 9.6 ∆DER (%) performance episode by episode for the two shows of the evaluation set. Defined as ∆DER = (DERINDEP −DERLONG). AHC refers to the Agglomerative clustering pseudo-speaker labels. . . . . . . . . . . . . . . . . . 164 A.1 Bayesian Network of the Fully Bayesian PLDA with Uncertainty Propagation . II xii Objectives and Methodology to each domain. 1.2 Objectives and Methodology The objectives of this thesis are the improvement of diarization capabilities so that systems could withstand the harmful conditions of the broadcast domain. These evolutions should be integrated in a single system, robust enough to deal with any sort of audio from the studied environment. Therefore, we should analyze possible evolutions in the previously described three lines of research. Regarding to the speaker characterization problem, we want to improve the extraction of the speaker representations, obtaining efficient and discriminative characterizations of the involved speakers. Thus, we first seek a deeper understanding about the state-of-the-art modelling techniques based on subspace projection. Once this knowledge is is acquired, it will let us explore the limitations for these technologies, as well as propose new approaches designed accordingly. With respect to the grouping task, our goal is the improvement of the clustering techniques estimating the diarization partitions. For this purpose, we make use of subspace-based techniques, specially PLDA, exploring different architectures and strategies. Finally, we also must deal with the domain variability. In this area we will try to provide tools and strategies capable of decreasing the degradation of domain mismatch in circumstances where in-domain data is scarce or unavailable. In order to reach this goal we will deal with the domain adaptation problem by exploring the inference of unsupervisedly-crafted pseudospeaker labels, obtained from the audio to diarize. These labels should be later used to specifically adapt the out-of-domain model to the evaluation audio. 1.3 Thesis organization The outline of this thesis is very oriented to the different challenges we previously described. For this reason, this work is divided in five main parts, as shown in the conceptual map in Fig. 1.2: •Basic Knowledge: This part is dedicated to present the diarization problem and an overview of the already proposed techniques in the state of the art (Chapter 2). Moreover, this part also starts the experimental activity, analyzing the characteristics of the broadcast domain and the performance of a baseline diarization system (Chapter 3). 4 Chapter 1. Introduction Diarization Basic Knowledge State of the art Baseline diarization System Model Adaptation Speaker Clustering FBPLDA Clustering Tree-based Clustering Uncertainty Propagation Conclusions & Future Work Speaker Representation Embeddings in Short Utterances DNN Embeddings Figure 1.2: Conceptual map of the studied topics in this Thesis 5 Thesis organization •Speaker Clustering: This part is focused on the different tools to improve the performance of the clustering stage. First, we analyze the performance of the Fully Bayesian Probabilistic Linear Discriminant Analysis (FBPLDA) model, dealing with its weaknesses (Chapter 4). Chapter 5updates the FBPLDA mode including the concept of Uncertainty Propagation (FBPLDAUP). Finally, in Chapter 6we present a totally independent clustering solution by means of a tree-based approach. •Speaker Representation This part of the thesis pays attention to the way speaker information is extracted from an audio utterance and compacted into a condensed representation, the embedding. First, we study the standard approximation for this information extraction, analyzing its impact on short utterances (Chapter 7). Later on, we make use of the learnt conclusions, applying them on the obtention of DNN-based embeddings for diarization (Chapter 8). •Model Adaptation: This part, consisting on Chapter 9, works on the unsupervised extraction of in-domain information, suitable for the adaptation of out-domain labels. This •Summary: This final part summarizes the conclusions for all the different parts of the thesis and proposes how this research could be followed in the future (Chapter 10) 6 Part I Diarization Basic Knowledge 7 Chapter 2 Diarization State of the Art The objective of this chapter is the revision of the state of the art in diarization. For this purpose, we take into account important reviews such as [Anguera et al., 2012][Tranter and Reynolds, 2006]. Our first goal is the identification of the main domains in which diarization has been applied. This differentiation helps understanding the evolution of diarization technologies. This knowledge allows the introduction of the two main approximations for diarization. Then, we explain in detail the functional blocks for the most popular diarization approach. Finally, the last part of the chapter includes a review about how to measure diarization performance. 2.1 Introduction The diarization task includes all the techniques and procedures needed to differentiate the contributions of speakers given an audio. In the most general case, diarization works in an unsupervised way, i.e. without prior knowledge about the involved speakers nor its number. However, diarization can get benefited by means of the knowledge of these characteristics, usually simplifying the problem. Historically, diarization research has focused on three main domains of interest: •Telephone channel domain. This environment involves the analysis of telephone conversations, characterized by the presence of few speakers, usually two, and conversational speech with short interventions. Moreover, telephone context usually considers close-tomouth microphones and restricted a priori known channel conditions. •Broadcast domain. This condition includes audios from mass media broadcasters (TV, radio, VoD, etc.). The most important feature in broadcast data is the large variability 9 Introduction of conditions. The variability in the number of speakers is almost unrestricted: from 3-4 up to 100 different speakers per hour of content, depending on the show. There is also variability in the type of speech: while some shows contain more read speech, e.g. the news, others hardly ever include it, being mainly composed of conversational speech, such as talk-shows. This characteristic has great relevance in diarization due to the length of the interventions. Whilst conversational speech usually consists of short interventions in order to maintain the conversation flow, read speech can generate longer turns due to the absence of feedback. Moreover, broadcast audio also presents variability of acoustic scenarios, such as studio and outdoors, each one with its own acoustic characteristics. Finally, except for live content, speech signal usually maintains high Signal to Noise Ratio (SNR), although very often speech is partially occluded by complementary acoustic additions such as music, and noises like canned laughter and applauses. •Meetings domain. This scenario implies recordings from meeting rooms, where an undetermined number of people is recorded from one or multiple microphones. Hence, recordings from this domain mainly include conversational speech. In this domain recording conditions are also very relevant. Despite the fact that close-to-mouth microphones can be used, more often omnidirectional microphone arrays are considered. These arrays can be located in a single point, e.g. on top of the conference table, or spread along the room. Regardless of the microphone locations, the distance between speaker and microphone cannot be ignored. This distance is responsible for noticeable channel effects in the speech propagation up to the microphones, including degradations as reverberation. Besides, the stationarity of this transmission channel cannot be guaranteed, affected by the relative movements between speaker and microphone. Finally, these channel effects usually imply power losses of the signal, making speech quality more sensitive to noises. The historic evolution of diarization originally started in the telephone domain. Due to its characteristics this domain provided the most restricted version of the diarization problem. Besides, there was a great interest for diarization solutions included in speaker recognition applications. This is why since 1996 diarization was part of NIST SRE evaluations [Przybocki and Martin, 2004]. Only after speaker recognition evolved its tools in terms of accuracy and robustness, diarization was able to export its knowledge to alternative domains, as in Rich Transcription (RT) evaluations [Garofolo et al., 2002], where alternative domains (broadcast news and meetings) complemented the conversational telephone speech. 10 Chapter 2. Diarization State of the Art SPK 1 SPK 2 SPK 3 SPK 4 SPK 1 COARSEST PARTITION FINEST PARTITION BOTTOM-UP TOP-DOWN Figure 2.1: Schematic of Bottom-Up and Top-Down diarization 2.2 Main diarization strategies Along literature several options have been proposed for the obtention of the diarization labels. However, most of these contributions can be grouped into two main conceptual approaches: Bottom-Up and Top-Down diarization strategies. Fig. 2.1 illustrates both diarization approaches in order to obtain the same diarization labels. •Bottom-Up. The given audio is first divided into individual segments, in which a single speaker is assumed to be present. Then, these segments are clustered so all blocks from the same speaker are tagged with the same label. •Top-Down. This alternative considers the opposite starting point. This approach starts considering a single speaker responsible of all the audio. Afterwards, the initial cluster is divided trying to match each final cluster with a real speaker in the audio. In spite of their opposite approach, both strategies need to solve the same two challenges: Determining whether some part of the audio contains speech from a single speaker and finding the boundaries if necessary. Despite the apparent simplicity of both tasks, their development for real applications has required several contributions in the literature. Nevertheless, both tasks are still far for being totally solved. Despite both Bottom-Up and Top-Down approaches are equally valid, they are not similarly popular. While both options have been developed along multiple publications, in recent years 11 Main diarization strategies Figure 2.2: General schematic for a diarization system the Bottom-Up strategy has gained much more awareness than the Top-Down counterpart. A reason for this popularity is the fit among the latest improvements in speaker recognition and the Bottom-Up diarization pipeline, making their inclusion straightforward. Under these circumstances, Bottom-Up diarization has taken its performance to unprecedent levels of quality. In consequence, this option has recently gained popularity becoming the standard diarization approach nowadays. 2.2.1 Bottom-Up diarization systems The popularity of Bottom-Up diarization has inspired the development of a standard architecture, which we present in Fig. 2.2. This schematic describes the standard considered blocks to transform the input raw audio into the desired final labels. The functionality of each block is described as follows: •Acoustic Feature Extraction. Raw speech audio is a very complex signal with many sorts of information. While some of them are valuable depending on the application (speaker, speech, language, etc.), others are not of interest (channel, noises, etc.) because they can alter our estimates. The feature extraction step aims to transform the raw signal into a faithful but compact representation of the acoustic information, simplifying the access to our target information and compensating those harmful degradations. •Segmentation. Generally speaking, segmentation is the task of dividing an audio into pieces according to an attribute, which should remain homogeneous along the total length of each piece. Focusing on diarization, the division attribute is the speaker identity. Thus, the goal of diarization segmentation is the division of a given audio into segments where a single speaker is present in them. This system must exploit the homogeneity of data in short periods of time to find the boundaries between speakers. An ideal segmentation step should provide the time marks for the different speaker interventions in an audio. 12 Chapter 2. Diarization State of the Art •Speaker Characterization. Speaker characterization is a high-level information extraction which collects the speaker information from the acoustic features. In order to properly do so, it requires working with audio from a single speaker. Thanks to this requirement, highly evolved techniques work along the given input segments, enhancing their speaker discriminative properties while compensating the harmful variabilities. Moreover, this process usually converts variable-length segments into fixed-dimension compact representations, more suitable for postprocessing. •Clustering. The output of the segmentation step is a set of acoustic fragments with a single speaker in each of them. However, the same speaker may have produced more than one segment. The clustering stage is responsible for grouping all those segments from the same speaker and label them with a unique tag. For this purpose, clustering takes the segment representations as input, generating the diarization labels as output. •Resegmentation. Resegmentation is an optional extra segmentation step to refine the initial segmentation boundaries. This extra border tuning takes advantage of the inferred clustering output, with an accurate knowledge about the evaluation audio. Resegmentation output may be considered as diarization labels or be fedback into the system for further refining. 2.3 Acoustic features for diarization In order to differentiate speakers, diarization systems require a subsystem capable of providing discriminative characteristics at each time step of an audio. These characteristics, also known as features, should maximize their classification capabilities. For this reason, they try to represent the audio information in a tractable manner, simplifying the information gathering while reducing harmful sorts of variability (noise, channel information, etc.) meanwhile. Moreover, feature extraction is the first diarization block, hence no assumption like number of speakers nor their identity, speaker transitions, etc. can be done. The most popular features so far are those commonly known as short-term acoustic features. Originally designed for speech recognition, these features carry out a spectral analysis of the raw signal while inspired by both the human production and perception systems. Because the speech signal is not stationary, this analysis must be performed in short analysis windows. The most popular features are the Mel Frequency Cepstral Coefficients (MFCCs), originally presented in [Davis and Mermelstein, 1980]. These features propose a short-time analysis of the 13 Audio segmentation DKL(P||Q) = X x P(x) log P(x) Q(x)(2.6) Unfortunately, its original definition is not symmetric, i.e., the KL divergence of Q with respect to P (DKL(P||Q)) may not be the same as the divergence of P with respect to Q (DKL(Q||P)). In consequence a symmetrized version, known as KL2 divergence, is used instead. This divergence for distributions P and Q is defined as: DKL2(P||Q) = DKL(P||Q) + DKL(Q||P)(2.7) Moving to diarization, this distribution is considered in segmentation in [Siegler et al., 1997] [Delacourt and Wellekens, 2000]. 2.4.1.3 Deep Neural Networks (DNNs) Thanks to the evolution of neural networks many of the tasks previously carried out by other means, such as statistics, are now performed by this technology. Regarding segmentation, some contributions have attempted the inclusion of DNNs in this task. In [Gupta, 2015] DNNs are used as classifiers. The hypothetical boundary frame is stacked along its context window, feeding a monolithic DNN consisting of feed forward layers. The final layer classifies the boundary frame as real or not. Moreover, a likelihood measure can be obtained in the process. By contrast, DNN regression capabilities can also been applied. In [Hruz and Zajic, 2017] the neural network must carry out the regression of the transition probability, softened during training. For this purpose, input data is treated by means of stacks of convolutional neural networks. In both cases, DNNs work as standalone systems. However, both architectures fit the given more general definition, where a neural network provides a metric for a fixed-length analysis window and compared against a threshold. 2.4.2 Model-based segmentation Despite the fact that metric-based segmentations are the most popular ones, other alternatives have also been proposed. Considering model-based segmentations, [Li et al., 2009] considers Hidden Markov Models (HMMs) for segmentation. The model represents each class by means of a 64-Gaussian GMM. This concept is evolved in [Diez et al., 2018], where classes are represented with tied GMMs, more suitable for speaker representation. 20 Chapter 2. Diarization State of the Art Finally, some systems [Garcia-Romero et al., 2017][Diez et al., 2019] work in terms of a coarse SbC approach. They prefer working with very short fixed-length (around 1.5 seconds) segments, not taking care for boundaries. These systems rely on the latest speaker characterization techniques, which have evolved to provide robust enough representations when working with very short segments. By doing so, they alleviate the computational costs while only introducing a small proportion of corrupted segments: as many degraded segments as real boundaries. Besides, these systems usually count with resegmentation systems to eliminate the generated distortions once speaker models are available. 2.5 Speaker characterization The nature of speech makes this information to have a sequential nature. Human beings concatenate multiple sounds to transmit the desired information. However, there is no limitation in terms of its length nor the message. It can either be a large speech or a short reply to a closed question, i.e., "yes" or "no". Moreover, it can include all the acoustic units or just a restricted set. Speaker recognition technologies should provide a tool to robustly encode the identity of the involved speaker regardless of the intra-speaker variability, i.e. the variability within all the possible utterances from the same speaker. Some reviews such as [Furui, 2004][Kinnunen and Li, 2010] provide a good overview about the evolution of these technologies. 2.5.1 Early days Some of the first successful speaker recognition systems were based on the correlation of spectrograms [Pruzansky, 1963]. This idea was later evolved to take into account the formant analysis [Doddington, 1971]. Because these techniques were not powerful enough to deal with text-independent recognition, some alternatives were explored for the following decade. Some proposals during those years are the instantaneous spectra covariance matrix [Li and Hughes, 1974], spectrum and fundamental frequency histograms [Beek et al., 1977] or linear prediction coefficients [Sambur, 1972]. The following great evolution appeared with the consideration of template models: Dynamic Time Warping (DTW) [Furui, 1981] and Vector Quantization (VQ) [Rosenberg and Soong, 1987] [Soong et al., 1985], which proposes short time feature vectors compressed in codebooks. This principle was later evolved as long as matrix quantiziers for multi-frame were also proposed [Juang, 1990]. 21 Speaker characterization 2.5.2 Model-based representations In the 80s, a great evolution in the characterization philosophy was proposed. Rather than considering speech as a deterministic process where features could be measured, state-of-the-art contributions started to define statistical models as generators for speech. Moreover, these generators were often designed only taking into account the acoustic information, not considering high-level crafted features. The generative sort of solution has many advantages. First, all segments are supposedly generated by a known distribution, a parametric solution perfectly described by a closed set of parameters, some of them speaker dependent. Besides, this sort of solution allows the same model to work with variable-length segments while providing a fixed-dimension speaker representation. Finally, statistical solutions can also provide protection against different types of randomness associated with the voice (phonetic variability, noises, etc.). When choosing the distribution to better represent speakers, Gaussian distributions are usually taken into account. Very well known among statisticians, Gaussian distributions have worthy properties. However, Gaussian distributions are too simple to properly represent all the variability and conditions in speech. Thus, combinations of them, Gaussian Mixture Models (GMMs) are considered instead. This approach is considered under the assumption that a linear combination of enough Gaussians should be able to reproduce any distribution. 2.5.2.1 Gaussian Mixture Models (GMMs) Gaussian Mixture Models are generative statistical models first introduced in speaker recognition in [Reynolds and Rose, 1995]. They are composed by the weighted sum of CGaussian components, each one with its own weight πc, mean vector µcand covariance matrix Σc, being c= 1..C. Thus, the sequence O={o1, ..., on, ..., oN}generated by a GMM has been randomly drawn as: P(O|M) = N Y n=1 C X c πcN(on|µc,Σc)(2.8) The evaluation of these systems worked in terms of a loglikelihood ratio. Two loglikelihood terms were considered, both taking into account the test audio audiotest but considering two different models: A model of the claimed enroll speaker (Menroll) and a model representing speakers except for our enrollment one (Menroll). llr = ln P(audiotest|Menroll) P(audiotest|Menroll)(2.9) 22 Chapter 2. Diarization State of the Art While Menroll was straightforward, the definition of Menroll was not so clear. Many systems worked with a pool of cohort models, chosen for each trial according to different criteria. Then, this idea was evolved in [Reynolds et al., 2000], which proposes the GMM-UBM paradigm. First, this contribution integrates the cohort of alternative speakers into a single model, responsible to represent the total variability of the acoustic data. This general model, a large GMM trained with several speakers, is known as Universal Background Model or UBM. Furthermore, instead of building from scratch individual enroll models Menroll, it proposes the option of their construction as a MAP adaptation from the UBM, specifically an adaptation of the component means. The obtained benefits are a tighter coupling between models, and faster scoring techniques. In this scenario, the proposed llr was: llr = ln P(audiotest|Menroll) P(audiotest|MUBM)(2.10) 2.5.2.2 Support Vector Machines (SVM) The GMM-UBM strategy became a milestone in speaker verification, specially regarding the way to model speakers. However, alternative scoring approaches were attempted. Within this line of research Support Vector Machines (SVMs) were proposed for speaker recognition [Campbell et al., 2006a], leading to the SVM-GMM strategy. Support Vector Machines are binary classifiers that project the input data into a highdimensional space where a hyperplane separates the two classes. The evaluation in SVMs is defined as follows: f(x) = C X c=1 ycαcK(xc, x) + b(2.11) where f(x)stands for the distance of the utterance xwith respect to the hyperplane. xc,αc and ycrepresent the support vectors, weights and labels respectively, with the restriction that PC c=1 ycαc= 0 and αc>0. The labels yctake the value +1 for one class and −1for the other one. Besides brepresents the hyperplane bias. Finally, K(·,·)stands for the kernel function, responsible for projecting the data into the high-dimension space and calculating distances terms If K(·,·)is restricted to satisfy the Mercer condition, the Kernel condition can be expressed as an inner product as: K(x, y) =< g(x), g(y)>(2.12) where g(·)is the transformation into the highly dimensional space. SVMs are trained by a maximum margin strategy. This type of training must identify a hyperplane which properly classifies the training elements while satisfying the following re23 Speaker characterization striction: the chosen hyperplane must keep the maximum distance with respect to the training populations of both classes. This request forces the hyperplane to provide the maximum margin protection against spurious data during evaluation. The inclusion of SVMs in speaker recognition [Campbell et al., 2006a] was carried out by proposing a kernel that bounds the KL divergence. In our scenario, KL divergence measures the distance between utterances Oaand Ob, modeled by GMMs Maand Mbrespectively. Both GMMs are obtained according to the GMM-UBM paradigm, hence they share the component weights πcand the component covariance matrices Σc, only differing at the component means µc. The kernel accomplishing this request is: K(Oa,Ob) = C X c=1 √πcΣ1/2 cµa c√πcΣ1/2 cµb c(2.13) In consequence, the kernel function can be interpreted as the inner product of the two GMM supervectors, a concatenation of the GMM means undergoing a diagonal scaling. Applied to our previous definition of SVMs, the enrollment supervector constitutes the set of support vectors and the test supervector plays the role of evaluated utterance, deciding whether it comes from the enrollment speaker. The GMM-SVM paradigm was complemented with the Nuisance Attribute Projection (NAP) concept. This idea, introduced in [Campbell et al., 2006b], considered the compensation of the intra-speaker variability present in the supervectors. This compensation is performed by estimating a low rank matrix U, also known as eigen-channels matrix, which defines the intraspeaker variability within the supervector space. Once modeled, a matrix P=I−UUTcan be introduced in the kernel function already seen in eq (2.13): K(Oa,Ob) = C X c=1 √πcΣ1/2 cµa cP√πcΣ1/2 cµb c(2.14) 2.5.2.3 Joint Factor Analysis (JFA) Joint Factor Analysis (JFA) [Kenny, 2005] is an evolution of the GMM representations, paying special attention to two concepts developed with the GMM-SVM approach: supervectors and subspaces for certain variabilities. Taking both concepts into account the JFA methodology evolves de the GMM-UBM paradigm decomposing the adapted GMM supervector as a sum of terms: µj=µUBM +Vyi+Uxk+Dzj(2.15) 24 Chapter 2. Diarization State of the Art where µjis the adapted supervector mean of utterance j.µUBM represents the mean supervector from the UBM model. The term Vyiis the speaker dependent term. Vis a low rank matrix describing the subspace of the inter-speaker variability while yiis a tied latent variable, i.e. a latent variable whose value is the same for all utterances from speaker i, responsible for the utterance. Similarly, we have the term Uxkor channel term. Uis a low rank matrix describing the channel variability space and xkis the tied latent variable for all utterances with the same channel k, including utterance j. Finally, we have the term Dzj, which must explain the remaining variability. For this purpose, Dis a diagonal matrix and zja latent variable unique for the utterance. All the three latent variables, yi,xkand zjare Standard Normal distributed. 2.5.3 Embedded representations The following large evolution implied the improvement of the already proposed models, but also a new methodology. On the one hand, models including latent variables to map the speaker information significantly improved the performance. On the other hand, the approach of the GMM-SVM showed that information could be extracted from the models and independently treated. The combination of both ideas created the embedding paradigm, defining models that constrain the speaker information into a restricted space where a latent variable should explain each speaker. From these latent variables we could extract compact representations, voiceprints for each speaker, also known as embeddings. Embeddings offer several advantages compared with previous approaches. Once embeddings are extracted, they can be decoupled from the original extraction method, simplifying their storage. Moreover, this decoupling makes impossible the return to the original audio, guaranteeing privacy. Finally, the obtained embeddings can be postprocessed by alternative methods, also known as backends. In fact, the current speaker recognition state of the art, from which most of these technologies are conceived, is dominated by the embedding-backend pipeline. In the following lines some of the most popular embeddings are presented, and one of the most popular backends, the Probabilistic Linear Discriminant Analysis (PLDA) is explained afterwards. 2.5.3.1 I-vectors I-vectors [Dehak et al., 2011] are a direct evolution of the JFA modeling. Rather than differentiating between speaker and channel factors, i-vector model integrates them into the total variability subspace. This fusion makes the latent variable store both speaker and channel information together. Moreover, latent variables are not linked among utterances anymore, being 25 Speaker characterization only tied along the samples from the utterance. Besides, this model no longer considers a residual variability term. In consequence, the utterance j, consisting of the sequence of frames O={o1, ..., on, ..., oN}, is now modeled by a GMM whose mean supervector µjis defined as: µj=µUBM +Twj(2.16) where µUBM again describes the UBM mean supervector. Tstands for a low rank matrix describing the total variability subspace and wjis the latent variable depending on the utterance. The mentioned model still can be evaluated in terms of likelihoods as in JFA. Nevertheless, this technology evolved to become a voiceprint extractor. The commonly used i-vector is the mean of the posterior distribution of the latent wjgiven the utterance j. This distribution is Gaussian and defined as: wj∼N (wj|µw,Σw) = Nwj|L−1 wΓw,L−1 w(2.17) Γw= C X c=1 TT cΣc Nj X n=1 γnc (on−µc) = C X c=1 TT cΣcFc(2.18) Lw=I+ C X c=1 TT c Nj X n=1 γncΣcTc=I+ C X c=1 TT cNcΣcTc(2.19) where µwrepresents the mean of the posterior distribution and Σwis its covariance. These terms are constructed in terms of Tc, the submatrix from Tdescribing the contribution of the cth Gaussian component, and Σc, the covariance matrix for the cth component in the UBM. The information of the utterance is contained in Ncand Fc, the zeroth and centered first order Baum Welch statistics for the cth component respectively. Finally, Njrepresents the total amount of samples in the utterance j. Both of them are obtained in terms of the responsibilities γnc, the probability of the nth sample onto be drawn from component cof the GMM-UBM. 2.5.3.2 Hybrid i-vectors The latest great evolution of neural networks, affecting both software and hardware, has become a milestone along most artificial intelligence tasks. This evolution also reached speech technologies [Hinton et al., 2012], including speaker characterization. In these tasks, at first, this acquisition of the new approaches was smooth, complementing existing state-of-the-art technologies. A proposed inclusion of DNNs in i-vectors was presented in [Lei et al., 2014] as hybrid ivectors. The i-vector extractor principle is the same, i.e. it explores how an utterance specific model differs from a UBM due to the unique characteristics of the utterance. However, the 26 Chapter 2. Diarization State of the Art UBM is not a GMM anymore. Now this role is played by a DNN, discriminatively trained to discern phoneme senones. This neural network is now in charge of the responsibilities γnc required to compute the Baum Welch statistics Ncand Fc, the unique input for i-vector training. However, due to the fact that no GMM-UBM is involved, γnc now represents the probability of the feature frame onto contain the the senon cinstead. An alternative proposal are phonetic i-vectors [Viñals et al., 2019d]. This proposal sets an original i-vector model in which the GMM-UBM responsibility depends on a prior activation, controlled by a DNN phoneme classifier. Under this approach, the set of C components is decomposed in multiple subsets, each one responding to individual phonemes. By doing so, particular phoneme models become more specific while reducing acoustic uncertainties. 2.5.3.3 DNN embeddings The improvements of hybrid i-vectors were outstanding, outperforming past technologies. The results in [Sadjadi et al., 2016] presented an unprecedent performance combining DNN posteriors with BNFs. However, technologies were still suffering from i-vectors flaws. The proposed evolution was a cutting-edge idea. Rather than evolving the generative ivector model, it trains a totally discriminative DNN. In [Snyder et al., 2016] x-vectors were proposed following this idea: A neural network is train to classify an audio among a closed set of speakers. The input features first undergo multiple frame-level non-linear transformations and then they are pooled into an utterance projection. This projection goes through utterancelevel non-linear transformations before its classification. The network is trained to recognize a large pool of speakers by means of cross entropy. Given a trained network, the embeddings also known as x-vectors are extracted during the forward propagation of the information, in the utterance-level transformations. The great performance of x-vectors has encouraged the community to evolve to DNNs. Now multiple alternatives to x-vectors are available, including LSTM based architectures [Wang et al., 2018], Wide Residual Network based embeddings [Villalba et al., 2019][Viñals et al., 2019d] or even expanded x-vectors [Villalba et al., 2019]. 2.5.4 Probabilistic Linear Discriminant Analysis (PLDA) PLDA is a statistical linear backend. Defined in [Prince and Elder, 2007] as a generative model, PLDA applies the subspace concept already considered in JFA, assuming the embedding φjas a sum of variability terms: φj=µ+Vyi+Uxj+ǫj(2.20) 27 Clustering where Vyirepresents the speaker variability term and Uxjthe utterance variability counterpart. Both terms consist of low rank matrices (Vand Urespectively), which define subspaces for the latent variables yiand xjrespectively. Whereas the speaker latent variable yiis tied along all utterances with the same speaker, xjis particular for each embedding j. We consider these latent variables, yiand xj, to be standard normal distributed. Additionally, the model also includes an extra variability term ǫjto explain the residual variability in each particular embedding. ǫjis modeled by means of a zero-mean Gaussian distribution and diagonal covariance matrix D−1. Finally, µis the constant speaker independent term. Although this model offers a closed-form solution, when firstly applied on embeddings (ivectors at that time), its performance was not significatively better. It requires embeddings to be Gaussian in order to properly obtain its improvement, although the extracted i-vectors were far from this distribution. The most popular solution to this issue is length normalization [Garcia-Romero and Espy-Wilson, 2011]. Embeddings, before feeding the PLDA model, are forced to reassure that its Euclidean norm is equal to one. This process projects the input embeddings into a hypersphere of radius equal to 1. Before length-normalization, embeddings should be centered and whitened. By doing so, the resulting embeddings are spread along the hypersphere rather than being concentrated in restricted regions of the hypersphere, leading to more discriminative capabilities of the systems. Multiple alternatives have appeared to the original PLDA model. The Simplified PLDA (SPLDA) fuses the channel and residual terms. Another alternative is the Discriminative PLDA [Cumani et al., 2013a], which trains the same model in a discriminative manner. An important alternative is the Heavy-Tailed PLDA (HTPLDA) [Kenny, 2010]. This model was proposed before length-normalization as a way to deal with non-Gaussian embeddings by modifying the prior distributions. However, after length-normalization its computational complexity discouraged its usage. Nevertheless, with the advent of DNN embeddings, far more nonGaussian than i-vectors, HTPLDA provides small benefits with respect to other alternatives [Brummer et al., 2018]. 2.6 Clustering The clustering stage in a Bottom-Up diarization architecture is responsible for the gathering of the acoustic fragments in terms of their speaker. This duty can alternatively be considered as a labeling task. Being the audio of Nacoustic segments represented by the set Φ={φ1, .., φN} of speaker representations or embeddings, clustering must infer a partition, a set of labels Θ = {θ1, .., θN}, so that those segments from the same speaker share a common label. 28 Chapter 2. Diarization State of the Art Table 2.1: Bell number Bin terms of the number of elements to cluster Number of segments NNumber of partitions B 1 1 22 3 5 415 552 6 203 ... ... 10 115975 ... ... 20 51724158235372 To do so, we first require a measure to determine how a certain partition Θfits the set of embeddings Φ. This metric may have multiple natures, e.g. statistical, graphs, kernels, etc. Then, given the chosen metric we must find the partition with the best metric value. Unfortunately, regardless of the metric they all share the same difficulty: The best partition is only guaranteed to be obtained if all possible partitions are analyzed, just choosing the one with the best metric. This option is usually referred as brute-force approach. Unfortunately, studies such as [Brummer and de Villiers, 2010] reveal that the number of partitions increase very fast as long as the number of segments to cluster Nrises. In fact, except for very low values of N, brute-force approaches are in general not viable. Given an audio with Nsegments, the total number of possible partitions is described by the Bell number B. This number, applicable for any clustering task, represents the total number of independent possible grouping arrangements, in our case partitions, for a set of Nelements. This number is defined by means of a recurrent relation: BN+1 = N X n=0 N nBn(2.21) B0=1 (2.22) This number increases very fast as long as the value of Nrises. In Table 2.1 values for low values of Nare shown. According to Table 2.1, even very low values of Nimply huge number of candidate partitions. If we consider even higher values of N, e.g. 100 or 200 segments for a one-hour TV show, the number of candidate hypotheses to compare becomes intractable. Fortunately, not all 29 Performance metrics •MISS ERROR (MISS). Speech audio incorrectly labeled as non-speech. This term also measures the Voice Activity Detection (VAD) performance. •FALSE ALARM ERROR (FA). Non-speech audio in which a speaker is considered to be present. Voice Activity Detection (VAD) performance is affected by this term as well. •SPEAKER ERROR (SPK). Speech misclassified as generated by an alternative speaker. •OVERLAP ERROR (OV). Periods of time when multiple speakers are simultaneously talking. This error involves the estimation of the number of speakers (underestimation when not all speakers are detected and overestimation when non-present speakers are also labeled) as well as the misclassification of the involved speakers. Regarding the four error terms, clearly two of them, Miss Error and False Alarm, are related with segmentation, specially the VAD step. With respect to the speaker and overlap errors, these terms mainly depend on the clustering stage. However, they are treated differently. According to its definition, the overlap error involves any inference error when multiple speakers are talking. Thus, it measures miss, false alarm and speaker errors for these periods of time. Hence overlap is the most challenging error term, with several proposed contributions about its detection (e.g. [Otterson and Ostendorf, 2007][Zelenák and Hernando, 2012]) but without any functional solution yet. Therefore, in certain evaluations this term is obviated for performance comparisons. Due to the fact that errors are non-overlapped, DER can be decomposed on multiple terms, each one evaluating the degradation due to each type of error. The alternative definition of DER is: DER =LMISS +LFA +LSPK +LOV Ltotal (2.31) =EMISS +EFA +ESPK +EOV (2.32) where EMISS,EFA,ESPK and EOV are the DER error terms for miss, false alarm, speaker and overlap causes respectively. Despite all its benefits regarding simplicity and decomposition of error, DER presents strong limitations. Obviating the miss error and false alarm terms, directly related with the VAD performance, no further knowledge can be inferred from DER about the speaker error. A similar score is obtained if some amount of audio is misclassified, regardless of how many speakers are affected. Besides, DER considers all the audio uniformly relevant, and errors involving 36 Chapter 2. Diarization State of the Art the same amount of audio are equally harmful. However, in real life neither the speakers nor their speech are equally valuable, invalidating this consideration. This is specially relevant when speakers do not contribute to the audio with the same amount of speech. For example, some errors may be irrelevant for the most talkative speakers but far more significant for those speakers contributing with much less speech. Therefore, some critics about DER metric are arising while the community is eager for finding an alternative score. In recent times some alternative metrics have also been proposed for diarization tasks. The Mutual Information (MI) metric was defined in DIHARD 2018, a diarization evaluation in difficult conditions. The idea behind MI is measuring how much information we have about the real labels provided our hypothesed partition. The proposed metric was proposed for study and was complemented by DER, which managed the leaderboard. The metric is defined as: MI = R X i=1 S X j=1 nij Nlog2 nijN risj (2.33) where Rrepresents the number of clusters in the reference with riduration each, and Sstands for the number of hypothesized clusters, each one with duration sj. Besides, the term nij is the amount of speech assigned to the speaker iin the reference and to the cluster jin the hypothesis. Finally, Nsymbolizes the total amount of speech. Another alternative is the Jaccard Error Rate (JER). This metric was proposed as alternative for DER in DIHARD 2019. The first step in the evaluation is a mapping among the Rclusters in the reference and Sclusters in the hypothesed partition. This mapping is carried out according to the Hungarian algorithm, so each cluster in the reference will be mapped to at most one cluster of the hypothesis and vice versa. Then for each speaker in the reference we estimate: JERref =F A +MISS T OT AL (2.34) where T OT AL represents is the amount of audio present in both the reference cluster ref and the mapped counterpart. If there was no paired cluster, its value would be the total amount of speech of speaker ref.F A stands for the total amount of speech not present in the reference cluster ref but considered as part of the paired grouping. Its value is 0if no mapping for ref was carried out. Finally, MISS is the amount of speech present in speaker ref but not included in the mapped counterpart. If ref speaker has no paired cluster, its value is equal to T OT AL. Having defined the individual terms JERref , the Jaccard error rate for a recording is the average of specific Jaccard error rates: JER =1 RX ref JERref (2.35) 37 Performance metrics Regardless of the used metrics, they do not provide any clue about the reasons for the misclassification of audio. Hence alternative metrics should be helpful to better understand the speaker error. This error is mainly generated in the clustering block, thus clustering metrics, such as the complementary clustering and speaker impurities, well described in [van Leeuwen, 2010] are suitable for this task. The cluster impurity represents how well the clusters from a hypothesized partition contain audio from a single speaker. Defined in terms of its cluster purity counterpart, cluster impurity is minimized as long as the obtained clusters contain audio from a single speaker. However, it is not obligatory that clusters from the same speaker share the same label, invalidating the metric for diarization. Similarly to the cluster impurity concept, we can also define the speaker impurity. This new concept describes how well the speech from a speaker is tagged with a single label, and it is minimized as long as more and more data from one speaker only requires a single label. This metric is also invalid for diarization because multiple speakers in the same cluster do not degrade the final score. 38 Chapter 3 Analysis of Diarization in Broadcast Data Diarization in broadcast is a complex task, composed of a large number of subtasks working together, as seen in Chapter 2. Whilst most of the previously described techniques work well in restricted conditions as the telephone domain, diarization in broadcast data requires many particularities to be taken into consideration. In this chapter we analyze the broadcast domain, emphasizing its particularities. For this purpose, we first introduce a reference diarization system. This system will serve to explore the wide variability along broadcast data afterwards. In this analysis we cover both quantitative and qualitative results, and study how results are affected by this uncertainty. Finally, according to the analysis and obtained results, we suggest the different lines of research, some of them treated along this thesis. 3.1 The diarization reference system The reference system considered for this analysis is an AHC-based diarization approach, very common in the literature as baseline system. This architecture follows a Bottom-Up approach, first dividing the raw signal into segments, which are later clustered according to their speaker identity. Fig. 3.1 illustrates the basic architecture of the system. In the following lines we explain in detail the setup for each element in our system. •Feature Extraction For the audio transformation into features, we strictly consider MFCCs, standard features in the state of the art. Our MFCC setup includes a 32-band Mel filter bank, and a final coefficient reduction, only considering coefficients C1-C20. The energy information is discarded too. The inferred stream of feature vectors does not include derivatives, and undergoes a normalization for its mean and variance (CMVN). •Segmentation The obtained stream of feature vectors are the input for the segmentation 39 The diarization reference system Figure 3.1: Schematic of our baseline diarization system stage. In the reference system the segmentation step is divided into two independent subtasks, Voice Activity Detection (VAD) to differentiate speech/non-speech and Speaker Change Point Detection (SCPD) to obtain the speaker turns. – Voice Activity Detection (VAD) The inference of the VAD mask is done by means of [Viñals et al., 2018a], using a segmentation-by-classification approach, which works in terms of DNNs. A 2-layer BLSTM DNN, with 256 neurons per layer, is taken into account. Each element in the second BLSTM output sequence is projected into a binary decisor. Thus, we infer one VAD label per input frame. This layer is trained and evaluated in 3-second analysis windows. Whenever the audio exceeds the window dimension, a sliding window analysis is carried out. This analysis implies a 3-second window and 2.5 seconds forward step. For the 0.5-second overlap period the inference works as follows. the first 0.25 seconds are obtained from the preceding window while the remaining 0.25 seconds are labeled with the inference from the following window. This overlapping design choice is made to avoid undesired windowing effects, specially near the artificial window borders. – Speaker Change Point Detection (SCPD) For the SCPD task we rely on well known techniques. In this case we opt for a SCPD hybrid solution led by a metric40 Chapter 3. Analysis of Diarization in Broadcast Data based segmentation, in particular ∆BIC considering Gaussian distributions with full covariance matrix. We operate in a sliding window regime following the description in Section 2.4. We make use of an analysis window with a minimum length of three seconds, and a 0.25-second accumulative window expansion whenever the window does not contain any boundary. Regarding the hyperparameter λ, it is adjusted according to those results obtained during the development phase. The metric-based solution is combined with a silence-based strategy, which assumes all speech/nonspeech transitions to be speaker borders. In fact, these borders are used as anchors for the ∆BIC segmentation. •Speaker Characterization The estimated segments are then converted into compact representations, each one summarizing the particular feature stream for each segment. Among the multiple options described in Section 2.5, our choice for the type of representation is the i-vector. In our system i-vectors are inferred by means of an extractor of 256 Gaussians and 100-dimension total variability matrix T. The extracted i-vectors are centered, whitened and length-normalized before feeding the clustering stage. •Clustering The clustering stage in our reference system is constructed around an AHC approach, using SPLDA pairwise log-likelihood ratio (LLR) as metric. Rather than considering the original AHC that exhaustively evaluating all similarities among clusters, we follow a simplification described in [van Leeuwen, 2010]. This simplification only requests the estimation of the initial pairwise similarity among the embeddings. Then, at each fusion iteration we approximate the exhaustive similarities by approximations considering the already estimated values. Among the different options to carry out the approximation, we opt for the UPGMA (unweighted pair group method with arithmetic mean) approach [Sokal and Michener, 1958]. Regarding the clustering stop criterion, it is done by means of a threshold experimentally adjusted during development. 3.2 Analysis of broadcast data Broadcast data is a type of domain specially characterized by its wide variability. Whenever no restrictions are applied regarding the shows of analysis, speech processing tasks must be robust enough to withstand a great range of conditions. Among the different tasks affected by this variability we must take into account diarization. There are many reasons for broadcast data to be so mutable. From different recording locations like studio and outdoor, to the considered equipment. Furthermore, extra factors should 41 Analysis of broadcast data be considered, such as the acoustic addons, i.e. acoustic artifacts like laughter or applause, that corrupt the audio signal. For diarization purposes we will pay attention to this variability along the clustering stage. Previous blocks can be interpreted as high-quality feature extractors, thus clustering must provide the knowledge to properly group together their representations in order to obtain the final labels. During clustering we must deal with two main types of variability: the one present in the acoustic representations, the acoustic variability, and the one available in the final speaker labels, the speaker distribution variability. The effects of both types of variability differ, specially due to their influence along the diarization pipeline. The acoustic variability is consequence of the different acoustic conditions along the different audios of interest. These different conditions cause embeddings from the same speaker to be less homogeneous, making them less robust for clustering purposes. In consequence, this variability may be responsible for a degradation of the diarization performance. Regarding the speaker distribution variability, we are referring to the number of speakers in an audio and how much speech they contribute with. Errors in the estimation of the number of speakers are highly important because all the speech produced by the affected orators will be misclassified. Besides, the larger is the range of possible speakers of an audio, the larger are the potential errors in this estimation and more audio is usually involved. A great factor to take into account in this estimation is the distribution of speech along the different speakers. Those talkative orators with several contributions have enough audio to be robustly modeled and small misclassifications have negligible effects on them. By contrast, those speakers with very few speech are weakly represented and small errors may cause their loss. We now present an analysis about these two types of sources of variability in broadcast data. For this purpose, we will take into consideration two large datasets: Multi-Genre Broadcast Challenge 2015[Bell et al., 2015] and Albayzín 2018 [Ortega et al., 2018]. Both datasets include a large amount of broadcast audio content covering a wide variability of shows, genres, languages and media. 3.2.1 Multi-Genre Broadcast Challenge 2015 (MGB 2015) This dataset was released for the Multi-Genre Broadcast challenge in 2015 [Bell et al., 2015]. This challenge aims at processing tasks in the broadcast domain, including ASR, alignment and diarization. The dataset consists of approximately 1600 hours of Broadcast audio collected from British Broadcasting Corporation (BBC) along four of its channels. The total amount of audio involves around 1200 episodes from 500 different shows. All this audio is divided into three subsets: train, longitudinal development and evaluation. The subset division tries not to 42 Chapter 3. Analysis of Diarization in Broadcast Data share direct knowledge among the subsets, thus all episodes from a show are placed in the same subset. Aside the audio, the three subsets were distributed with diarization labels. However, the label accuracy is not uniform among subsets. Whilst the training subset includes the originally broadcast subtitles refined by a lightly-supervised ASR alignment as metadata, development and evaluation labels are manual annotations. Finally, MGB 2015 also contains an extra subset, namely development, released for the ASR evaluation. This subset consists of 28 hours from 47 shows, with manually annotated VAD marks. 3.2.2 Albayzín 2018 Albayzín 2018 is the latest edition of the Albayzín evaluations, the attempt from Red Temática de Tecnologías del Habla (RTTH) for the evolution of speech technologies in those languages spoken in the Iberian Peninsula. Regarding diarization, 2018 is the third edition after those hold in 2010 and 2016. For the 2018 edition the evaluation consists of approximately 600 hours of data from mass media domain, covering two different languages (Spanish and Catalan) and two different mass media (TV and radio). The whole dataset was composed by three different subsets, acquired along the different editions: From 2010 edition we have available 84 labeled hours of audio from 3/24 TV channel in Catalan. These data are complemented by 2016 data: 23 hours of manually annotated audio from broadcast radio signal from Corporación Aragonesa de Radio y Televisión (CARTV) in Spanish. Finally, 2018 edition also adds around 400 hours from broadcast content from Radio Televisión Española (RTVE). The evaluation divides the pool of data as follows: For training and development both 3/24 and CARTV are available, as well as 10 hours from RTVE with manual annotations. Evaluation data consists of 40 hours from RTVE subset. 3.2.3 Acoustic variability The richness of audio content makes both MGB 2015 and Albayzín 2018 a suitable choice to analyze variability in Broadcast data, studying how different factors affect the performance of diarization systems. For the audio variability we will make use of the diarization system paradigm, specially the SPLDA model. According to the PLDA paradigm, SPLDA models the variability along its training data by projecting it in two subspaces, the inter-speaker space defined by the matrix VVTand the intraspeaker space described by matrix W−1. Moreover, due to the fact that both matrices represent covariances as well, their analysis can lead to interesting information about the variability from 43 Analysis of broadcast data Dataset tr VVTSubspace tr (W−1)Subspace Telephone SRE 0.50 0.49 Broadcast MGB 2015 0.19 0.86 Albayzín 2018 0.49 0.58 Table 3.1: Trace analysis for PLDA inter-speaker (VVT) and intra-speaker (W−1) subspaces each type, inter-speaker and intra-speaker. In Table 3.1 we carry out an analysis for both inter-speaker and intra-speaker subspaces in terms of their covariance matrices. For this analysis we study the trace of both VVTand W−1 matrices for our two broadcast datasets of interest, MGB 2015 and Albayzín 2018. This study approximates the total variability within each subspace, equivalent to add the variability along each dimension of the subspace as if they were independent. This study is complemented by a similar analysis for telephone channel data, which plays the role of baseline. This baseline analysis considers SRE data, constructing our models with excerpts from SRE04, SRE05, SRE06 and SRE08. The results illustrated in Fig. 3.1 show a great mismatch in terms of conditions between telephone and broadcast data. While telephone channel presents a similar variability contained in both subspaces, our broadcast databases show at least around 18% relative extra intra-speaker variability. This measure increases up to a 352% relative extra variability in MGB dataset. Thus, when considering the broadcast domain, we must take into account the following question: Do similar embeddings share the same speakers or just analogous acoustic conditions? Furthermore, this intra-speaker variability is not only caused by differences among shows. In fact, broadcast data presents a high within-episode variability. This sort of variability corresponds to the different conditions in which the audio is recorded, e.g. the recording location (studio, outdoors, etc.), the involved material (microphones, postprocessing, ...), and acoustic additions (laughter, applauses, etc.). Besides, the speech signal is highly affected by the presence of emotional speech, i.e. the transformation of the voice in order to transmit extra information such as shouting (wrath), whispering (fear) or whining (pain). Regardless of the nature of the variability, it is usually very correlated along time, remaining the acoustic characteristics stable during periods of time that can contain multiple interventions from different 44 Chapter 3. Analysis of Diarization in Broadcast Data ❆✁✂✄☎✆ ✝✆✞✆✟✠✡✆☎☛ ✶☞ ✷☞ ✸☞ ✹☞ ✺ ☞ ✻ ☞ ✼☞ ✽ ☞ ✾ ☞ ✶☞☞ ✶☞ ✷☞ ✸☞ ✹☞ ✺☞ ✻☞ ✼☞ ✽☞ ✾☞ ✶☞☞ ❙ ❡ ❣ ♠ ❡ ♥ t ❥ ✌✍✎✏✍✑✒ ✓ ❙✁✂✄✁☎ ✆☎✝✞✟✠ ✡☎✞☛☞ ✶✌ ✷✌ ✸✌ ✹✌ ✺ ✌ ✻ ✌ ✼✌ ✽ ✌ ✾ ✌ ✶✌✌ ✶✌ ✷✌ ✸✌ ✹✌ ✺✌ ✻✌ ✼✌ ✽✌ ✾✌ ✶✌✌ ✍ ❡ ❣ ♠ ❡ ♥ t ❥ ✎✏✑✒✏✓✔ ✕ a) Pairwise LLR b) Ground truth mask Figure 3.2: Section variability example. For 100 first embeddingsfrom a Springwatch episode a) SPLDA pairwise LLR similarity metric b) Ground truth relationship. speakers. These periods with stable conditions usually correspond to the different sections of a show. A clear example could be the news, where interventions from the news readers in studio conditions are interleaved with outdoor connections. In order to expose this variability we again make use of our reference diarization system, studying the AHC similarity matrix constructed by means of PLDA pairwise log-likelihood ratio. This matrix should contain higher values for those elements comparing embeddings from the same speaker, regardless of the acoustic conditions. In Fig. 3.2 we illustrate the acoustic similarity matrix among the 100 first detected segments from an episode of the TV show Springwatch, from MGB 2015. In this analysis we cover an approximate 25% of the total detected interventions in the episode, balancing the tradeoff between generalization and visualization capabilities. The segments are studied in chronological order for timeline comprehension. The information includes two parts, the acoustic similarity and the ground truth mask. For the acoustic similarity we make use of the PLDA pairwise LLR, where the element ij reveals how similar are the embedding iand j. Lighter colors indicate higher speaker similarities and darker colors less probability to share the same speaker. Regarding the ground truth mask, the ij position in the figure is white if both embeddings, iand j, have the same speaker label, being black otherwise. The two images shown in Fig. 3.2 reveal the capabilities do discern between speaker and acoustic conditions are limited. We first analyze the speaker labels of the last 50 embeddings. According to the ground truth, two speakers are responsible for a sequence of interleaved utterances, as in a dialog. However, the LLR scores did not realize about that, providing an 45 Conclusions 3.4 Conclusions The results obtained along the present chapter have revealed different factors for the inherent variability in broadcast data. In order to deal with the detected uncertainties, diarization systems must work in the following elements: 3.4.1 The clustering approximation Our reference system bases its diarization choices according to an agglomerative architecture. This architecture is well-known in the community due to its simplicity. Nevertheless, more evolved solutions could obtain better diarization results. The choice of an alternative clustering procedure requires that some considerations must be taken into account. In first place we must take care of the metric to determine the quality of the partition. Results in Fig. 3.2 illustrate the great influence of channel effects in the PLDA LLR. Improvements about the modelization of the intra-speaker variability should lead to great benefits. Moreover, we can also work in the partition measurement, combining the local information considered in AHC (pairwise similarity) with a more general point of view. Thus, alternative clustering approaches should take into consideration the implications for some of the clustering choices in the decision-making process. Another point to focus on is the stop criterion. Many of the alternative clusterings simultaneously work with multiple partitions, which contain a wide range or speakers. Whenever comparing hypotheses, biased measurements must be compensated prior to its comparison in order to prevent significant degradations. 3.4.2 The quality of the embeddings The observed undesired variability shown in Fig. 3.2 may not be exclusively compensated during clustering, but also during the embedding extraction. In fact, the more discriminative is the information in the embeddings, the better will be its performance during the clustering stage. Unfortunately, embeddings include more variability that is inherent to broadcast data. We are referring to more general variability terms, such as phonetic variability and short segments. Any improvement in the management of these two variabilities would lead to a general improvement in all types of diarization, as well as in speaker recognition. 52 Chapter 3. Analysis of Diarization in Broadcast Data 3.4.3 The domain mismatch problem Finally, we must also cover the domain mismatch. Even if clustering systems could cover the variability in different domains, their particular characteristics should require some individual adaptation for an optimal performance. For this reason, we can make use of the domain adaptation techniques, i.e. adapt the proposed solution to each one of the domains of interest. Another solution could be the opposite, transforming the evaluation audio to fit the training conditions. Whatever is the solution, we must face another issue: broadcast data includes several shows and genres, thus in-domain data for each of them may be limited or just unavailable. 53 Conclusions 54 Part II The Clustering Problem 55 Chapter 4 Clustering by means of Fully Bayesian PLDA The results obtained in Section 3.3 have shown the limitations of the baseline diarization system, specially concerning the clustering stage based on an AHC solution. Therefore, this poor performance motivates the search for alternative clustering options. One of the main drawbacks of AHC is the use of local decisions, i.e. decisions taking into account very little information, as the pairwise loglikelihood ratios between two single embeddings. Thus, we would rather prefer a clustering method whose metric evaluates the overall partition. Another request is that the optimization process simultaneously optimizes all labels. A solution fitting both requirements is the clustering by means of Fully Bayesian PLDA. 4.1 The Fully Bayesian PLDA clustering solution This proposal of clustering was first proposed in [Villalba and Lleida, 2014] as an unsupervised clustering for model adaptation. Along the following lines we will define the model and explain how it can be used for clustering tasks, including diarization. In this process we will pay attention to its Variational Bayes (VB) decomposition, key point in this approach. 4.1.1 The Fully Bayesian PLDA (FBPLDA) model The Fully Bayesian PLDA model [Villalba and Lleida, 2014] is a generative statistical model which describes the input embeddings in terms of latent variables, some of them tied along all embeddings from the same speaker. Based on the Simplified PLDA, the FBPLDA also describes the embedding φjfrom the ith speaker as: φj=µ+Vyi+ǫj(4.1) 57 The Fully Bayesian PLDA clustering solution µW V ε φjyi θj πθ Ni I Figure 4.1: Bayesian network of the Fully Bayesian PLDA where µstands for the speaker independent term. Vrepresents the low-dimension matrix defining the speaker subspace. yiis the speaker latent variable, standard normal distributed and common for all embeddings from the speaker i. Finally, the remaining unexplained variability in the embedding jis included by the term ǫj, which is modeled by means of a zero mean Gaussian with covariance W. The evolution of the Fully Bayesian PLDA is that, in contrast to SPLDA, speaker assignments for both training and evaluation are unknown, using latent variables instead. Thus, a set of Nembeddings Φ={φ1, ..., φj, ..., φN}is explained by a set of Icandidate speakers, each one modeled by a speaker latent variable yifrom the set Y={y1, ..., yi, ..., yI}. In order to map each embedding to its generator speaker, we consider the set of latent variables Θ = {θ1, ..., θj, ..., θN}. Each one of the θjlatent variables follows a multinomial distribution, which produces a one-hot sample with Ivalues (θj={θ1j, ..., .θij, ..., θIj}). Each one of these values θij represents the assignment of the utterance jto the ith speaker. Thus, θjwill have its component θij equal to one when the ith speaker is responsible for the embedding j, being zero otherwise. Taking this assignment into account, we can model Φin terms of Yand Θas: P(Φ|Y,Θ) = N Y j=1 I Y i=1 Nφj|µ+Vyi,W−1θij (4.2) Due to the Bayesian approach, the speaker labels Θa priori follow a multinomial distribution. This distribution is complemented by its own prior, πθ, which explains the multinomial weights according to Dirichlet distribution. Besides, the described Fully Bayesian PLDA proposes an extra evolution. Instead of considering point estimations for the model parameters (µ, Vand W), this evolution assumes them to be latent variables as well. While the mean µand the columns of the speaker matrix Vare treated with a Gaussian prior, Wis modeled in terms of a Wishart distribution. Finally, the model also includes a prior variable εfor the variable V. The Bayesian network describing the whole model is illustrated in Fig. 4.1. 58 Chapter 4. Clustering by means of Fully Bayesian PLDA The training procedure for this model is not nearly as simple as for the SPLDA. The training of the latter model works in terms of the Expectation Maximization (EM) algorithm. This algorithm requires the estimation of the posterior distribution for each one of the latent variables in the model (E step), updating the point-estimation model parameters to maximize the loglikelihood (M step). However, in the FBPLDA model a closed-form solution for each posterior is not possible, thus E step cannot be performed. Therefore, the original work [Villalba and Lleida, 2014] also proposes an alternative training strategy by means of the Variational Bayes (VB) [Attias, 1999][Bishop, 2006]. Variational Bayes is an approximation method that allows to mimic the EM algorithm by a variational equivalent. Given a model depending on the set of latent variables Z= {Z1, ..., Zh, ..., ZH}, VB approximates the posterior distribution P(Z|Φ)by a factorial distribution q(Z) = QH h=1 q(Zh). Each one of the obtained factors q(Zh)is an approximation of the real posterior distribution P(Zh|Φ). In order to obtain the best approximation following the factorial restrictions each factor q(Zi)must follow a distribution following the relationship: ln q(Zh) = E∀Zr,r6=h[ln P(Z,Φ)] (4.3) Unfortunately, the approximation by means of a factorial distribution has limitations. Despite the fact that the obtained factor distributions q(Zh)only depend on one of the latent variables Zh, they are not completely independent. Taking into account eq. (4.3), some dependencies remain, being each factor constructed on top of the expected values from the other factors. A collateral effect of the Variational Bayes approximation is that the loglikelihood of the real model is no longer a suitable metric. These type of solutions works in terms of the Evidence Lower Bound (ELBO or L). L(Φ) = Zq(Z) ln P(Φ,Z) q(Z)dZ(4.4) Both ELBO L(Φ)and loglikelihood ln P(Φ)are interconnected. In fact, the loglikelihood term is the sum of the ELBO term plus the KL divergence between the factorial distribution q(Z)and the real posterior distribution P(Z|Φ). We can express this as: ln P(Φ) = L(Φ) + KL (q(Z)||P(Z|Φ)) (4.5) In fact, the ELBO and KL terms are interconnected. The maximization of ELBO makes KL divergence to be reduced, better approximating P(Z|Φ)by means of q(Z). As long as this approximation is more accurate, our ELBO term will be a more reliable representation of the loglikelihood from the original model. 59 The Fully Bayesian PLDA clustering solution Moving to the specific case of the Fully Bayesian PLDA, the proposed decomposition of factors is described as follows: P(Y,Θ, πθ,µ,V,W,ε|Φ) = q(Y)q(Θ) q(πΘ)q(µ)q(V)q(W)q(ε)(4.6) Thus, for training purposes we can now perform an analogous alternative of the EM algorithm, now maximizing ELBO. The variational equivalent to the E step must iteratively update the different factors to obtain the posterior distributions, and the analogous M step will proceed to the point estimation update. This process is repeated until convergence. 4.1.2 The clustering procedure The clustering technique by means of the FBPLDA model proposed in [Villalba and Lleida, 2014] has a statistical background. This approach assumes that diarization labels Θdiar are those that best explain the embeddings Φ. Then, the way we should compare how partitions Θexplain the data Φis the probability P(Θ|Φ), as described in Section 2.6.2. Θdiar = arg max Θ P(Θ|Φ) = arg max ΘZP(Z′,Θ|Φ)dZ′(4.7) where Z′represents the set of all latent variables in the model (Y,πθ,µ,V,Wand ε) except for Θ. The application of this approach to the FBPLDA model is not straightforward. The same difficulties during training are present in this approach, thus we again must rely on our VB decomposition. In consequence, we must work in terms of approximations as follows: Θdiar = arg max ΘZq(Y)q(Θ) q(πθ)q(µ)q(V)q(W)q(ε)dZ′= arg max Θ q(Θ) (4.8) Despite having simplified the clustering step to the maximization of a single factor q(Θ), the same difficulties as during training remain. The proposed factors are still interconnected by means of expectations, so the optimization of the factor q(Θ) needs other factors to be optimized as well. Unfortunately, these other factors also depend on q(Θ). For this reason we work in terms of an iterative update of factors in which, starting from an initial state, we reach a maximum ELBO. This iterative process is similar to the considered EM procedure during the SPLDA training. However, in this occasion no point estimation requires update (these only affect the model parameters µ,V,Wand ε), thus we only consider the E step. 60 Chapter 4. Clustering by means of Fully Bayesian PLDA Figure 4.2: Clustering schematic based on label initialization and FBPLDA resegmentation Moreover, because we assume the model parameters µ,V,Wand εto be perfectly tuned, we exclusively reevaluate q(Y),q(Θ) and q(πΘ). This clustering procedure can be interpreted as a two-step search: The set of embeddings Φis distributed along a set of clusters Yduring the update of the factor q(θ). Then, the same clusters Yare reevaluated in terms of the recently estimated Θduring the estimation of q(Y). This iterative process may be easily understood as follows: At the beginning of each iteration, according to the current value of the speaker labels Θwe characterize each one of the considered Iclusters. This characterization is done by the reevaluation of the speaker latent variables Yduring the update of the factor q(Y). Once the clusters are redefined, each embedding is then assigned to the most likely cluster when q(Θ) is reevaluated again. The main disadvantage of this approach is the need for some initialization Θ0. This initialization can be obtained in several ways, either from some prior knowledge or more often relying on the same embeddings Φ. Hence the proposed clustering stage follows the schematic represented in Fig. 4.2. This clustering strategy can be interpreted as a two-step clustering: A first block is in charge of obtaining an initial partition Θ0, which is refined afterwards by the FBPLDA clustering approach. Apart from the benefits due to label reassignment, this reclustering by means of the FBPLDA offers another advantage: an estimation about the speaker number. q(Θ) distributes the embeddings Φalong Ia priori candidate speakers. Nevertheless, it is not obligatory that all candidates generate at least one embedding. Those candidate speakers without assigned embeddings could be eliminated as part of the stop criterion. 61 Analysis of FBPLDA performance ers in the initial partition Θ0and those obtained after the resegmentation. Whenever Θ0contains less speakers than the ground truth (positive relative speakers), the FBPLDA reclustering does not discard any single speaker, only reassigning the embeddings along the different available candidate speakers. By contrast, whenever the initialization contains more speakers than the ground truth number, the algorithm starts discarding few speakers (1-3 speakers on average) although the rejection of extra speakers is not enough, significantly overestimating the number of speakers in an audio. In consequence, a bad estimation of the speaker number is difficult to be fixed by this resegmentation. Apart from the number of speakers, diarization is affected by other factors. Another important consideration to take into account is the chance of losing real speakers. In many occasions the worthy speakers are not those who speak the most but those with few but relevant interventions. An example may be the talk shows, where the most talkative individual is the moderator despite the real valuable contributions come from the remaining speakers. Regarding the FBPLDA solution, our experimental work has revealed that it is usually reluctant to consider small sets of embeddings (sometimes a single one) as an independent speaker, opting for assuming them as spurious data from a much larger cluster. This trend of fusing small clusters with larger ones is more relevant as long as the balance of data becomes odder. By contrast, when two clusters of similar size present audio from the same speaker, the algorithm is unlikely to fuse them together. These limitations about how the VB solution handles the resegmentation specially affects to low-talkative speakers. They contribute very little to the real audio, but depending on the application, their loss is not affordable. Therefore, we study the number of lost speakers according to the initial partition. The obtained results are shown in Fig. 4.7. Again the partitions are identified in terms of relative number of speakers ∆I. The results are also shown in terms of the first, second and third quartile. According to Fig. 4.7 clear subclustering initializations (we assume up to 20 extra speakers) lead to the loss of 2-3 speakers on average. This trend seems steady after 12 extra speakers in our initial partition Θ0. This result is specially interesting when compared with Fig. 4.6, which shows a growing overestimation of the speaker number, proportional to those present in the initial partition. The combination of both sources of information leads to the conclusion that we are usually losing real speakers, and most of the overestimation of speakers is a consequence of the underclustering of the remaining ones. 68 Chapter 4. Clustering by means of Fully Bayesian PLDA ✵ ✺ ✶✵ ✶✺ ✷✵ ✷✺ ✸✵ ✲ ✷✵ ✲ ✶ ✲ ✶✷ ✲✁ ✲✂ ✵✂ ✁ ✶✷ ✶ ✷ ✵ ❘❡❧❛t ✐ ✈❡ ❙♣ ❡❛❦❡r s ✄ ■ ▲ ♦ ☎ ✆ ✝ ✞ ✟ ✠ ✡ ✟ ☛ ☎ ☞✌✍ ✎ ✏✑ ✒✓ ✔✒✕✍ ❢✌ ✕ ❱❇ ✍✌✖ ✉ ✎✗ ✌ ♥ Figure 4.7: Lost speakers according to the relative number of speakers ∆Iin the initial partition Θ0. Results obtained with Albayzín 2018. 4.2.3 Number of speakers vs DER Diarization performance does not exclusively depend on the inferred number of speakers. Actually, a good estimation about the number of speakers is not always representative for a good diarization in terms of DER. This is because DER is highly dependent on a proper identification of the largest clusters in the analysis audio. Thus, as long as really talkative speakers are properly clustered, any treatment of non-talkative speakers may be considered beneficial, even their discard. Our next experiment studies the relationship between the initialization and the DER performance measure. For this experiment we have evaluated the inferred partitions obtained from the FBPLDA reclustering of multiple levels of the AHC dendrogram. The obtained results are illustrated in Fig. 4.8, representing the obtained DER in terms of the relative number of speakers in the initialization Θ0. The represented information simultaneously analyzes the initialization AHC system (Fig. 4.8a) as well as the AHC block followed by the reclustering stage (Fig. 4.8b). The results in Fig. 4.8 show that any underestimation about the number of speakers is very harmful for both diarizations, AHC and the FBPLDA reclustering. This degradation is more severe as long as the underestimation increases. These results are reasonable from the DER perspective, because severe losses of speakers will definitely cause the misclassification of very 69 Analysis of FBPLDA performance ✵ ✶✵ ✷✵ ✸✵ ✹✵ ✺✵ ✻✵ ✼✵ ✽✵ ✲✷✵✲ ✶✻ ✲ ✶ ✷ ✲✽ ✲ ✹ ✵ ✹ ✽✶✷✶✻ ✷✵ ❘❡❧❛t✐ ✈❡ ❙♣ ❡❛❦❡rs ✁■ ❉ ❊  ✭ ✪ ✮ ✂✄☎✆✝✞ ❢♦✟ ❆❍❈ ✠♦✡✉☛☞♦♥ ✵ ✶✵ ✷✵ ✸✵ ✹✵ ✺✵ ✻✵ ✼✵ ✽✵ ✲✷✵✲ ✶✻ ✲ ✶ ✷ ✲✽ ✲ ✹ ✵ ✹ ✽✶✷✶✻ ✷✵ ❘❡❧❛t✐ ✈❡ ❙♣ ❡❛❦❡rs ✁■ ❉ ❊  ✭ ✪ ✮ ✂✄☎✆✝✞ ❢♦✟ ❆❍❈ ✰ ❱❇ ✠♦✡✉☛☞♦♥ a) AHC results b) AHC + FBPLDA results Figure 4.8: DER (%) results for a) AHC and b) FBPLDA in terms of the relative number of speakers ∆I. Results obtained from Albayzín 2018, indicating the first, second and third quartile per bin. talkative speakers. Interestingly, according to Fig. 4.8 when an overestimation about the number of speakers should happen to occur, the trend of DER is not so degraded, with small loses of performance in AHC and no noticeable degradation when the FBPLDA reclustering is applied. In real life applications this overestimation scenario may be worthy enough, specially considering semi-supervised applications. By means of automatic techniques diarization labels with multiple pure clusters per speaker could be easily obtained, only requiring little human supervision to match those clusters with a common speaker. This manual work is simpler than cleaning clusters with multiple speakers, i.e. the scenario in which an underestimation of the speaker number is done. An alternative analysis is the search of the partition that provides the best diarization result. For this analysis each episode has been diarized with multiple initial partitions Θ0, all obtained from the AHC dendrogram. The results, shown in Fig. 4.9, are compared in terms of the relative speaker of the initialization with respect to the ground truth. According to Fig. 4.9, diarization results tend to prefer an overestimation of the number of speakers, inferring more speakers than those present in the reference labels. Moreover, the overestimation can be significative, with many episodes (more than 90% of the episodes) with at least 5 extra speakers obtaining the best DER results. These results fit with those previously obtained in Fig. 4.8, which showed that an underestimation of the speakers would lead to significant degradations. The choice for the best initialization is a great challenge. The choice for the initial partition leading to the minimum DER may be difficult or even impossible. Even if this option was feasi70 Chapter 4. Clustering by means of Fully Bayesian PLDA ✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁ ✵ ✵ ✷✵ ✸✵ ✹✵ ✁✵ ✻✵ ✼✵ ❘❛♥❦ ♦❢ ❘❡ ❧ ❛t✐✈❡ ❙♣ ❡❛❦❡rs ☎■ P ✆ ✝ ✞ ✝ ✆ ✟ ✠ ✝ ✡ ✝ ☛ ❆ ✉ ❞ ✠ ✝ ☞ ✭ ✪ ✮ ✌✍✎✏✎✑✒✎ ③✑✏✎✓ ✍ ✔✓✕ ❇✖✗✏ ❉❊✘ Figure 4.9: Distribution of the initialization with best DER in terms of the relative number of speakers. Results obtained from Albayzín 2018 and presented according to 5 bins. ble, it can imply such an unaffordable computational cost. Therefore, we might sometimes seek a tradeoff, assuming certain degradations in performance if a large simplification of the systems is achieved. In Fig. 4.10 we analyze the proportion of initial partitions whose diarization output differs from the best result up to a maximum bound. This figure is composed of two different distributions, Fig. 4.10a showing the chance of a 1% DER bound and 4.10b illustrating the distribution for a 3% DER bound. Fig. 4.10 illustrates that the probability for a partition to be under a DER degradation bound is very reduced. Only an approximate 17% of the initializations reach under the 1% DER bound, being over 33% with a higher bound (3% DER). 4.2.4 Number of speakers vs ELBO In these lines we want to explore markers to determine whether we are working with a good partition. Even if a closed set of partitions is provided, e.g. the multiple levels of the AHC dendrogram, we need an unsupervised option to compare them and decide which one is our best option. Because we are taking into account a statistical solution, a fair option should be the loglike71 Alternative initializations ✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁ ✵ ✁ ✵ ✁ ✷✵ ✷✁ ✸✵ ❘❛♥❦ ♦❢ ❘❡❧❛t✐ ✈❡ ❙♣❡❛❦❡rs ☎■ P ✆ ✝ ✞ ✝ ✆ ✟ ✠ ✝ ✡ ✝ ☛ ❆ ✉ ❞ ✠ ✝ ☞ ✭ ✪ ✮ ✌✍✎✏✎✑✒✎③✑✏✎✓✍ ✔✓✕ ❇✖✗✏ ❉❊✘ ✇✎✏ ❤ ✶✙ ❇✓✚✍✛ ✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁ ✵ ✁ ✵ ✁ ✷✵ ✷✁ ✸✵ ❘❛♥❦ ♦❢ ❘❡❧❛t✐ ✈❡ ❙♣❡❛❦❡rs ☎■ P ✆ ✝ ✞ ✝ ✆ ✟ ✠ ✝ ✡ ✝ ☛ ❆ ✉ ❞ ✠ ✝ ☞ ✭ ✪ ✮ ✌✍✎✏✎✑✒✎③✑✏✎✓✍ ✔✓✕ ❇✖✗✏ ❉❊✘ ✇✎✏ ❤ ✙✚ ❇✓✛✍✜ 1% Bound 3% Bound Figure 4.10: Distribution of the initialization with bounded DER, a) 1% and b) 3%, in terms of the relative number of speakers. Results obtained from Albayzín 2018 and presented according to 5 bins. lihood. Actually, because our model is solved by means of a Variational Bayes approximation, we should consider ELBO instead. Considered as an approximation of loglikelihood, ELBO is still a good representative number about how well the partition represents the input data Φ. In the next experiment we study the reliability of ELBO as partition selection criterion, exploring which initialization obtains the best ELBO after FBPLDA reclustering is performed. This result will indicate us which results are more statistically reliable. Due to range issues, we represent our results in Fig. 4.11 as a histogram illustrating the distribution of maximum ELBO in terms of the relative speaker number of the initialization Θ0. According to the results in Fig. 4.11, ELBO is a great indicator about the number of speakers, opting for small deviations (±5 speakers) with respect the ground truth in almost 50% of the involved data. However, another 25% of partitions optimize ELBO by overclustering up to 15 speakers. This is specially undesirable when considering Fig. 4.10, which requests the opposite (underclustering) for a better diarization performance. 4.3 Alternative initializations In the previous lines we have studied the great impact of the initialization on the performance in the FBPLDA reclustering solution. Its influence extends to multiple factors such as overall quality, the estimated number of speakers, how many speakers we may lose or the overall ELBO. 72 Chapter 4. Clustering by means of Fully Bayesian PLDA ✲✁✂✄ ✲✁✂✄✂✲✁ ✲✁✂✄✂✁ ✁✂✄✂✁ ✄①✁ ✵ ✵ ✷✵ ✸✵ ✹✵ ✁✵ ✻✵ ✼✵ ❘❛♥❦ ♦❢ ❘❡❧❛t✐✈❡ ❙♣ ❡❛❦❡rs ☎■ P ✆ ✝ ✞ ✝ ✆ ✟ ✠ ✝ ✡ ✝ ☛ ❆ ✉ ❞ ✠ ✝ ☞ ✭ ✪ ✮ ❊▲❇❖ ✌❝❝✍✎✏✑✒❣ ✓✍ ✔✒✑✓✑✌✕✑③✌✓✑✍✒ Figure 4.11: Distribution of the partition with best ELBO in terms of relative speakers. Results obtained with Albayzín 2018 and represented according to 5 bins. All this acquired knowledge was extracted to better deal with the initialization issue. We expect to find an alternative initialization approach with respect to our first FBPLDA approach (Table 4.1), where a threshold determines the level of the dendrogram in the AHC, refined afterwards by the FBPLDA. According to the already seen information, we illustrate two different approaches. The first one seeks an efficient tradeoff between DER improvement and computational cost. The second option tries to reach the best possible results despite falling into more elaborated strategies with a higher computational cost. 4.3.1 Computationally efficient initialization The computationally efficient initialization approach was developed according to the information in Fig. 4.10. The illustrated information reveals that those initial partitions whose resegmentation differs from the best result below a bound are prone to contain significantly more speakers than the tuned initialization. Our first alternative simply proposes assuming an initialization whose number of speakers is guaranteed to overcome the ground truth, thus exploiting this circumstance. By doing this, we assume an initialization in which the AHC algorithm is more unlikely to have made significant errors, and exploit the stop criteria from the FBPLDA algorithm. The great benefit of this approach is that is computationally efficient, requiring as much time as our FBPLDA baseline. In Table 4.4 we analyze the impact of this approach considering different values for the upper number of speakers. The experiment includes both development and test subsets from 73 Alternative initializations Experiment Dev. DER(%) Eval. DER(%) MGB 2015 50 Speakers 26.08 42.22 75 Speakers 26.13 41.37 100 Speakers 26.55 42.25 200 Speakers 29.68 45.93 300 Speakers 29.99 44.95 Finest partition 29.24 44.76 Albayzín 2018 50 Speakers 16.48 18.36 100 Speakers 22.60 25.38 150 Speakers 25.94 28.14 Finest partition 60.48 62.54 Table 4.4: DER (%) results from AHC initialization with a maximum number of speakers. Included multiple maxima and the finest partition with one segment per cluster. Results shown for developmentand test subsets from MGB 2015 and Albayzín 2018 corpora. both MGB 2015 and Albayzín 2018. Apart from fixed number of speakers for all episodes, our results also include the case of the finest AHC partition, i.e. one embedding per candidate speaker, as limit case. The results in Table 4.4 show that a restricted number of speakers in both datasets may significantly improve the results, specially compared with Table 4.1. This may be a consequence of the different audio characteristics among shows, making thresholds on top of similarity metrics inappropiate. However, it is important to notice that not all initializations with a higher number are equally useful. As long as the value becomes higher, error ratios start appearing. This degradation can be taken to the limit considering the finest initialization. Therefore, this approach needs to be higher than the ground truth value and not too high for falling into degrading the performance. While in our approach we assume the same value for all episodes, their duration is not the same. Thus, more elaborated alternatives may adjust this value according to the audio duration. 4.3.2 ELBO-based initialization choice criterion The statistical nature of the FBPLDA clustering solution makes reasonable the use of alternative initialization choices. From a statistical point of view, the best partition ΘDIAR should be the one that best explains the given data Φ, as shown in eq. 2.28. The traditional way to measure 74 Chapter 4. Clustering by means of Fully Bayesian PLDA Figure 4.12: Schematic of diarization based on the simultaneous evaluation of K different initializations. The final partition is selected by means of PELBO. how well some model represents the given data is by means of the posterior loglikelihood. When adapting this approach to the FBPLDA model we must deal with the VB nature of our solution. This solution substitutes the original likelihood term by the ELBO term. Thus, following a similar approach our best diarization partition should be the one that maximizes the ELBO Lterm as follows: ΘDIAR = arg max ΘL(Θ,Φ)(4.9) Nevertheless, this idea is not complete yet. When considering multiple initializations, not all of them initially suppose the same number of speakers. Hence the higher the number of speakers, the more likely this data can overfit to the evaluation data. Therefore, as presented in [Viñals et al., 2018a], a penalized ELBO (PELBO) term is proposed instead. This approach is inspired in BIC, in where the likelihood of the model is penalized in terms of the modelling capabilities. Then, the best labels can be obtained as: ΘDIAR = arg max Θ PELBO(Θ,Φ) = arg max Θ (L(Θ,Φ)−λQ(Θ)) (4.10) where Q(Θ) represents the considered excess of modeling capabilities due to the total amount of speakers in the partition Θ. This term is multiplied by a finetuning parameter λ. This partition choice method can be applied right after the AHC algorithm, in order to choose a single partition. However, due to the great refining properties of the FBPLDA reclustering, we opted for choosing among reclustered partitions. Therefore, a set of Kdifferent initializations are simultaneously reclustered, choosing among them the final partition afterwards. This clustering structure is represented in Fig. 4.12. In Table 4.5 we illustrate the results obtained with this algorithm. The results include DER marks for both MGB 2015 and Albayzín 2018 datasets, including both development and test subsets. Two different results are included, a single result exclusively in terms of ELBO, and the penalized ELBO as well. 75 Conclusions Experiment Dev. DER(%) Eval. DER(%) MGB 2015 ELBO 26.82 39.12 PELBO 25.95 39.88 Albayzín 2018 ELBO 14.48 17.77 PELBO 13.90 17.79 Table 4.5: DER (%) results for ELBO and PELBO initialization choice. Results shown for development and test subsets from both MGB 2015 and Albayzín 2018 corpora The obtained results are significantly better than our baseline system, and also overcomes the previously described efficient solution. Moreover, both ELBO and penalized ELBO show very similar results, illustrating the robustness of the approach. Unfortunately, while penalized ELBO helps to improve simple ELBO during training, in evaluation slightly degrades in performance, partially illustrating the high domain mismatch between shows. 4.4 Conclusions In this chapter we have explored the addition of the FBPLDA reclustering to the baseline diarization system described in Section 3.1. According to the obtained results, the performance in both broadcast diarization datasets has been significantly improved due to this block. These improvements affect both the diarization metric DER as well as the estimation of the speaker number. Moreover, we have explored some of the FBPLDA limitations. Our analysis has revealed a great dependence of performance according to the initialization. Thus, the better the initialization, the better is the label refinement by means of FBPLDA. Besides, our study also reveals that initializations should better overestimate the number of speakers (around 10 extra speakers compared to the oracle value) so that FBPLDA provided the best labels. However, despite the improvements obtained in this overestimation scenario, we still must assume degradations as the loss of low talkative speakers. Furthermore, our analysis has determined that the Evidence Lower Bound (ELBO) seems a reasonable indicator to infer the number of real speakers within an audio. Finally, we have explored the joint collaboration of the AHC initialization and the FBPLDA reclustering rather than assuming them as independent blocks. Thus, we proposed two success76 Chapter 4. Clustering by means of Fully Bayesian PLDA ful approaches where a limited number of levels of the AHC dendrogram are refined by the FBPLDA, which is also responsible for the stop criteria. Whilst our first approach explored an efficient search by assuming a single initial partition reassuring an overestimation of its speaker number, our second strategy carries out a simultaneous reclustering of multiple initializations, opting for one of the obtained partitions according to ELBO. Both of them have demonstrated that FBPLDA reclustering is a more powerful stop criteria than AHC, obtaining significant improvements. With respect to the comparison between them, the ELBO stop criteria is able to outperform the overestimation strategy, at the cost of increasing the computational costs. 77 PLDA with Uncertainty Propagation (PLDAUP) Experiment EER (%) minDCF Long-Long 3.37 0.161 Long-Short 5.98 0.291 Short-Long 5.98 0.283 Short-Short 8.76 0.403 Table 5.1: EER (%) and minDCF results with SPLDA for SRE10 coreext-coreext det5 female with involved short utterances Short) for each role (enrollment and test). In Table 5.1 we represent the performance for each of these combinations. The performance is evaluated according to two different metrics, Equal Error Rate (EER) and minimum Detection Cost Function (minDCF). According to the obtained results, we observe a severe degradation in performance as long as short utterances are considered. These degradations are noticeable when short utterances are involved, regardless of their role. If both roles, enrollment and test, are played by short utterances the degradation is much more significative. Furthermore, the seen degradation is noticeable in both metrics, EER and minDCF. This information is complemented by DET curves, shown in Fig. 5.2. DET curves in Fig. 5.2 confirm those results obtained in Table 5.1, showing clear differences in performance when short utterances play any role in verification. Besides, this degradation is more noticeable when short utterances simultaneously play the enrollment and test roles. Finally, this behaviour is consistent along the whole curves, regardless of the operation point. The previous experiment illustrates the impact of short utterances in state-of-the-art technologies and justifies the search for alternative techniques as PLDAUP. This approach is evaluated in our next experiment, scoring the same trials with long and short utterances. However, in this experiment we restrict our evaluation to two of the previous scenarios: Long-Short (original long utterance as enrollment and short utterance as test) and Short-Short (both enrollment and test are short utterances). These two conditions are evaluated with our new PLDAUP model, undergoing two different alternatives to mimic length-normalization: scalar normalization and unscent transformation. The obtained results are shown in Table 5.2, including both EER and minDCF results. Those results in Table 5.2 illustrate the benefits due to the inclusion of Uncertainty Propagation. However, the obtained improvements do not affect the performance in the same way. While EER is clearly more affected (at least 12% relative improvements), minDCF benefits are more negligible. Besides, benefits are obtained regardless of the length-normalization approximation for matrices, although minimum extra improvements are obtained with unscent 84 Chapter 5. Uncertainty Propagation for Diarization ✵✵✵✁ ✵✵✁ ✵✁✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵ ✵✁ ✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵ ☎✵ ✆✝✞✟✠✆✝✞✟ ✆✝✞✟✠✡☛✝☞✌ ✡☛✝☞✌✠✆✝✞✟ ✡☛✝☞✌✠✡☛✝☞✌ ▼ ✐ s s P r ♦ ❜ ❛ ❜ ✐ ❧ ✐ t ② ✭ ✪ ✮ ❋✍✎ ✏❡ ❆✎ ✍✑♠ ♣✑✒✓ ✍✓✔✎ ✔ ✕✖ ✗✘✙ ❉❊❚ ❝✉✚✈✛✜ ❢✢✚ ◆■❙❚ ❙❘❊✶ ✣ Figure 5.2: DET curves with SPLDA for SRE10 corext-coreext det5 female with involved short utterances Experiment Long-Short Short-Short EER (%) minDCF EER (%) minDCF SPLDA 5.98 0.291 8.76 0.403 PLDAUP scalar 5.11 0.261 7.72 0.389 PLDAUP unscent 5.08 0.269 7.67 0.385 Table 5.2: EER (%) and minDCF results with PLDAUP for SRE10 coreext-coreext det5 female with involved short utterances 85 PLDA with Uncertainty Propagation (PLDAUP) ✵✵✵✁ ✵✵✁ ✵✁✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵ ✵✁ ✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵ ☎✵ ✆✝✞✟✠ ✝✞✟✠✡ ✞☛✝ ✝✞✟✠☛☞✌✍✎☞✏☛✝ ▼ ✐ s s P r ♦ ❜ ❛ ❜ ✐ ❧ ✐ t ② ✭ ✪ ✮ ❋✑✒✓❡ ❆✒✑✔♠ ♣✔ ✕✖✑✖✗✒ ✗ ✘✙ ✚✛✜ ❉❊❚ ❝✉✢✈✣✤ ❢ ✦✢ ◆■❙❚ ❙❘❊✶✧ ✵✵✵✁ ✵✵✁ ✵✁✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵ ✵✁ ✵✥ ✵✂ ✁ ✥ ✂ ✁✵ ✥✵ ✄✵ ☎✵ ✆✝✞✟✠ ✝✞✟✠✡ ✞☛✝ ✝✞✟✠☛☞✌✍✎☞✏☛✝ ▼ ✐ s s P r ♦ ❜ ❛ ❜ ✐ ❧ ✐ t ② ✭ ✪ ✮ ❋✑✒✓❡ ❆✒✑✔♠ ♣✔ ✕✖✑✖✗✒ ✗ ✘✙ ✚✛✜ ❉❊❚ ❝✉✢✈✣✤ ❢ ✦✢ ◆■❙❚ ❙❘❊✶✧ a) Long-Short experiment b) Short-Short Experiment Figure 5.3: DET curves with PLDAUP for SRE10 corext-coreext det5 female with involved short utterances transformations. These considerations can also be observed in Fig. 5.3, where DET curves are shown. DET curves explain the differences between minDCF and EER. PLDAUP seems to work similarly to SPLDA in those regions of the DET curve highly penalizing missing target trials. By contrast, in those regions of the curve for high false alarm is where PLDAUP obtains the highest improvements. 5.2.2 PLDAUP in speaker clustering The inclusion of the uncertainty propagation in the PLDA for speaker verification has led to small improvements, despite not being such a revolution. However, diarization is a slightly different task. Rather than making independent decisions, diarization is the result of a large set of choices depending on each other. Thus, small individual improvements may result into an accumulation of benefits. This is why we want to evaluate the uncertainty propagation capabilities in our clustering step, the stage where we can easily integrate our PLDAUP model. As a first approximation we do not work in diarization yet but in a speaker clustering task, working with short utterances. For this reason, we remain working in the telephone channel domain, using the same chopped subset previously used in speaker verification. The total amount of involved audios is 2740 utterances, containing 232 different speakers. 86 Chapter 5. Uncertainty Propagation for Diarization ✶✁ ✷✁✁ ✷✁ ✸✁✁ ✸✁ ✁  ✶✁ ✶ ✷✁ ✷ ✸✁ ✸ ✹✁ ✹ ✁ ❙✂✄☎✆ ✝✞ ✂✄☎✆P✄✟✂ ✝✞ ✂✄☎✆✠✡☛☞✌✡✍ ✟✂ ✝✞ ❙✂✄☎✆ ❙✞ ✂✄☎✆P✄✟✂ ❙✞ ✂✄☎✆✠✡☛☞✌✡✍ ✟✂ ❙✞ ◆ ✠✎✏ ✌✑ ✒✓ ☛✔ ✌✕✖✌ ✑☛ ❈❧ ✉st❡ r s ■ ♠ ♣ ✗ ✘ ✐ ✙ ② ✭ ✪ ✮ ✚✛✜✢✣✤✥✦ ♦❢ ✧★▲❉❆ ✈✩ ★▲ ❉❆❯★ Figure 5.4: Impurity results for SPLDA and PLDAUP in SRE10 coreext-coreext det5 female chopped Our experimental setup is the same as in our previous speaker verification experiment, i.e. a GMM-UBM i-vector extractor followed by a PLDA model. However, this time scores are not considered for 1vs1 trial decisions but the metric for an AHC solution, which determines the final labels. In order to provide a better overview about the potential of PLDAUP, we prefer not using any stop criterion, analyzing multiple levels of the AHC dendrogram. The results will be measured in terms of speaker and cluster impurities (SI and CI respectively) and shown in Fig. 5.4. Thick lines represent cluster impurities and dashed lines speaker impurities. The analysis involves 3 different PLDA versions: traditional SPLDA is shown in blue, PLDAUP with scalar normalization of the uncertainty matrix is shown in red and green represents PLDAUP with normalization of the uncertainty matrix by means of an unscent transformation. Our representation also includes an extra line (black) indicating the true value of speakers in this subset. The results in Fig. 5.4 show a great benefit when uncertainty propagation is applied, confirming our hypothesis of improvement accumulation. For the range of study PLDAUP cluster impurity consistently undergoes an absolute improvement within the range 5-10% with respect to SPLDA, while no evident degradations in the speaker impurity are noticed. Besides, this improvement has also reduced the bias for the Equal Impurity (EI) point in terms of the number of speakers. While SPLDA reaches the EI point at 350 speakers, PLDAUP does the same at 300 87 PLDA with Uncertainty Propagation (PLDAUP) ✶✁ ✷✁✁ ✷✁ ✸✁✁ ✸✁ ✁  ✶✁ ✶ ✷✁ ✷ ✸✁ ✸ ✹✁ ✹ ✁ ▲✂✄☎ ✆✝ ❙✞✂✟✠ ✆✝ ▲✂✄☎ ❙✝ ❙✞✂✟✠ ❙✝ ◆✡☛☞✌✟ ✂ ✍ ✎✏✌✑✒✌ ✟✎ ❈❧✉st❡r s ■ ♠ ♣ ✓ ✔ ✐ ✕ ② ✭ ✪ ✮ ✖✗✘✙✚✛✜✢ ♦❢ ✣P✤❉❆ ✇✛✜ ❤ ✤♦♥❣ ✈✥ ✣❤♦✚✜ ❯✜✜✦✚ ❛ ♥❝✦✥ ✶✁ ✷✁✁ ✷✁ ✸✁✁ ✸✁ ✁  ✶✁ ✶ ✷✁ ✷ ✸✁ ✸ ✹✁ ✹ ✁ ▲✂✄☎ ✆✝ ❙✞✂✟✠ ✆✝ ▲✂✄☎ ❙✝ ❙✞✂✟✠ ❙✝ ◆✡☛☞✌✟ ✂ ✍ ✎✏✌✑✒✌ ✟✎ ❈❧✉st❡r s ■ ♠ ♣ ✓ ✔ ✐ ✕ ② ✭ ✪ ✮ ✖✗✘✙✚✛✜✢ ♦❢ P✣❉❆❯P ✇✛✜❤ ✣♦♥❣ ✈✤ ✥❤♦✚✜ ❯✜✜✦✚❛♥❝✦✤ a) SPLDA b) PLDAUP Figure 5.5: Impurity results for a) SPLDA and b) PLDAUP with scalar normalization in SRE10 coreext-coreext det5 female chopped training with short utterances speakers. Furthermore, while speaker verification results indicated that unscent transformations were better than scalar normalization of the uncertainty matrix, those obtained for the clustering task show that the scalar normalization overcomes the performance of unscent transformations up to an absolute 2-5%. During the previous experiments we analyzed some PLDA model trained on excerpts from SRE04, 05, 06 and 08. This training pool consists of audios with large amounts of speech per utterance. Thus, training embeddings can be considered reliable enough for our standards. However, this scenario may not be so realistic in other domains, as diarization. In fact, diarization data usually consists of a combination of long and short utterances. Thus, we must analyze how our two models, SPLDA and PLDAUP, behave when short utterances are considered for model training. For this purpose, we analyze the impact of short utterances on the training pool. In this experiment we build an alternative version of the considered training pool (SRE04, SRE05, SRE06 and SRE08) by randomly chopping the original utterances guaranteeing the speech content to be within the range of 3-60 seconds. This subset will only be considered for the training of the PLDA models, both SPLDA and PLDAUP. Under these conditions we evaluate our subset with short utterances with both PLDA models, SPLDA and PLDAUP with scalar normalization of the uncertainty matrix. In Fig. 5.5 we compare how each model responds depending on the training cohort, diving between SPLDA (Fig. 5.5a) and PLDAUP (Fig. 5.5b). The illustrated results in Fig. 5.5 reveal interesting details. First, SPLDA seems to adapt well to short utterances, outperforming the version with long utterances for any operational 88 Chapter 5. Uncertainty Propagation for Diarization point. An explanation is that the training cohort suffers from shifts of the embeddings due to its length, similar to those in the evaluation subset. Consequently, these shifts can be considered as extra intra-speaker variability and are taken into account in the corresponding parameter (W). By contrast, PLDAUP seems to slightly lose some of its performance. A reason for this behaviour is that intra-speaker variability lies in a subspace controlled by Ujand W. When long reliable utterances are used to train the model most of this variability is forced to be in the Wsubpace, acting UjUT jas an addition during evaluation. However, when training involves short utterances both terms are representative and contributing, thus making decisions much noisier. 5.3 Fully Bayesian Probabilistic Linear Discriminant Analysis with Uncertainty Propagation (FBPLDAUP) The confirmation of Uncertainty Propagation beneficial capabilities motivates its evaluation in broadcast diarization. However, in this domain our best results so far have been shown in Section 4.1 by means of the FBPLDA and its Variational Bayes resegmentation. This resegmentation is in fact the key point of this best approach, fixing some of the mistakes and thus improving the performance. For this purpose, we update the FBPLDA model described in Section 4.1.1 including the new uncertainty propagation concept. The name of this new model is Fully Bayesian PLDA with Uncertainty Propagation (FBPLDAUP). This new model, as well as its predecessor, explains a set of Nembeddings Φfrom Idifferent speakers modeled by the set Y={y1, ..., yi, ..., yI}, where yiis a latent variable common for all utterances from the same ith speaker. Additionally, it also incorporates an extra latent variable xij per utterance, responsible for modeling the variability due to the utterance length. The assignment of each embedding to its responsible speaker is done in terms of θij, a latent variable taking the value of 1if the element jis generated by the ith speaker and 0 otherwise. Thus, we define the conditional distribution of Φas: P(Φ|Y,Θ,X,µ,V,W) = I Y i=1 N Y j=1 Nφj|µ+Vyi+Ujxj,W−1θij (5.11) where Nφj|µ+Vyi+Ujxij,W−1represents the distribution of the embedding φjaccording to speaker i. This modelization includes a speaker independent term µ, a speaker dependent term Vyiand the i-vector variability term Ujxij.Vis a low rank matrix explaining the speaker subspace and yiis the speaker latent variable. Ujstands for the i-vector variability subspace 89 FBPLDA with Uncertainty Propagation (FBPLDAUP) µ W V φjyi Uj xij θij πθ ε Ni I Figure 5.6: Bayesian network for the Fully Bayesian PLDA with Uncertainty Propagation full rank matrix and xij its latent variable. Finally Wis a full rank matrix explaining the within speaker subspace. Due to the fact that we are building a Fully Bayesian solution, our model parameters (µ,V and W) are distributions rather than point estimates. In fact, Vhas its own prior distribution ε, a product of gamma distributions. In addition to the model parameters, the speaker labels Θare also treated as latent variables, modeled by means of a multinomial distribution. This multinomial distribution includes a Dirichlet prior πθin order to explain the probabilities per class. The Bayesian network for the model is shown in Fig. 5.6. For diarization purposes with the FBPLDAUP we follow the statistical approach already described in Section 2.6.2 maximizing the posterior distribution P(Θ|Φ). However, the complexity of the model makes the true posterior intractable for optimization purposes. Hence, we prefer applying Variational Bayes for a more suitable solution. The applied simplification to the new model is: PY,X,Θ, πθ,˜ V,W,ε=q(Y,X)q(Θ) q(πθ), q ˜ Vq(W)q(ε)(5.12) For more information about the formulation of the different priors the formulation is included in Appendix A. Due to the relationship between FBPLDA and FBPLDAUP, both follow the diarization strategy described in Section 4.1.1, based on a Variational Bayes approximation. Our new model assumes that diarization labels Θdiar should be those which best explain the given utterances, 90 Chapter 5. Uncertainty Propagation for Diarization modeled by its mean and covariance. Thus, we must maximize P(Θ|Φ). By Variational Bayes we approximate this posterior distribution by q(Θ), whose maximum will be our solution. Unfortunately, VB factors are interconnected, requiring q(Θ) the other factors to be optimized as well for a proper solution, which also need q(Θ) adjusted as well. Therefore, we must apply an iterative reevaluation of factors (q(Y,X),q(Θ) and q(πθ)) in order to reach the best value for our hypothesis. 5.4 Diarization of broadcast data with FBPLDAUP In the following lines we present the results in broadcast diarization when Uncertainty Propagation is included. For these experiments we make use of MGB 2015 according to the configuration explained in Section 3.2.1. The applied system follows the schematic presented in Section 4.1.2, where an AHC stage to estimate partition seeds for the VB resegmentation. In our new approach, any involved PLDA is substituted by its PLDAUP counterpart, so our AHC stage works in terms of the PLDAUP score while the VB resegmentation role is now played by the FBPLDAUP. For comparison reasons both models will present the same dimension for the speaker subspace than in the original counterpart. In the first comparison we will compare the initialization labels. For this reason, we evaluate the whole AHC tree taking into account two similarity metrics, SPLDA and PLDAUP. Then, for each level on both trees we will evaluate the resulting partitions by means of DER and calculate ∆DER = DERPLDAUP −DERSPLDA. This comparison is repeated for each involved show in MGB 2015, including both development and test subsets. In Fig. 5.7 we illustrate a histogram about the relative variations of DER depending on the considered model. In Fig. 5.7 we can observe a distribution whose mean and mode are biased to negative values, indicating improvements when substituting the SPLDA model by the PLDAUP. Besides, the skewness of the distribution is also negative, showing more relevance for negatives values of ∆DER where PLDAUP outperforms SPLDA. A similar study can be performed with cluster and speaker impurities. Fig. 5.8 repeats the preceding procedure, although now we evaluate both impurities instead of DER. Similar histograms analyzing impurities for the pool of partitions are shown, differentiating between cluster (Fig. 5.8a) and speaker (Fig. 5.8b) impurities. The results in Fig. 5.8 show distributions with mean and mode close to zero for both cases. Differences arise when considering higher order moments, such as skewness when both impurities have opposite behaviour (speaker impurity is negative while cluster impurity is positive), and kurtosis, being higher in the cluster impurity distribution. The behaviour observed 91 Diarization of broadcast data with FBPLDAUP ✲✁ ✲✂ ✲✄✁ ✲✄✂ ✲✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁ ✸ ✂ ✸ ✁ ☎❉❊❘ ✭✪✮ P r ♦ ♣ ♦ r t ✐ ♦ ♥ ✆ ✝ ✞ ✟✠s✡☛✠❜✉✡✠☞✌ ☞❢ ✍✟✎✏ ❜ ❡✡✇❡❡✌ ❙✑▲✟❆ ❛✌❞ ✑▲✟❆❯✑ Figure 5.7: Histogram of DER variations between SPLDA and PLDAUP in MGB 2015 data. ✲✁ ✲✂ ✲✄✁ ✲✄✂ ✲ ✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁ ✸ ✂ ✸ ✁ ☎ ❈❧✉st❡r ■♠♣✉r✐t② ✭✪✮ P ✆ ♦ ✝ ♦ ✆ ✞ ✟ ♦ ♥ ✠ ✡ ☛ ❉☞✌✍✎☞❜✏✍☞✑✒ ✑❢ ✓ ✔✕ ❜ ✖✍✇✖✖✒ ❙✗▲❉❆ ❛✒❞ ✗▲❉❆❯✗ ✲✁ ✲✂ ✲✄✁ ✲✄✂ ✲✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁ ✂ ✁ ✄✂ ✄✁ ✂ ✁ ✸ ✂ ✸ ✁ ☎ ❙♣ ❡❛ ❦❡r ■♠♣✉r✐t②✭✪✮ P ✆ ♦ ✝ ♦ ✆ ✞ ✟ ♦ ♥ ✠ ✡ ☛ ❉☞s✌✍☞❜✎✌☞✏✑ ✏❢ ✒✓✔ ❜ ✕✌✇✕✕✑ ✔✖▲❉❆ ✗✑❞ ✖▲❉❆❯✖ a) Cluster Impurity b) Speaker Impurity Figure 5.8: Histogram of a) cluster and b) speaker impurities variations between SPLDA and PLDAUP initializations in MGB 2015 data. 92 Chapter 5. Uncertainty Propagation for Diarization Experiment DER(%) SPK Setup 40 40.61 50 39.83 75 39.72 Best FBPLDA 41.37 ELBO Setup ELBO 39.83 Best FBPLDA 39.12 Table 5.3: DER (%) results with the FBPLDAUP model in MGB 2015. in Fig. 5.8 does not match with those previously obtained in Section 5.2.2. While in those experiments benefits were obtained in the cluster impurity, remaining speaker impurity almost unaltered, in broadcast data cluster impurity shows an average 1.24% absolute extra degradation, with 70% of the partitions degrading this metric, and speaker impurities show a 2.05% absolute improvement, common for 63% of the partitions. A conclusion extracted from these results indicates that PLDAUP does no longer provide such improvements obtained in telephone channel experiments. Some causes for this degradation lie on the effect of short utterance training. In the telephone channel experiments we observed how short utterances trials were better evaluated as long as the SPLDA model considered them during training. By contrast, PLDAUP did not show any improvement but small degradations when trained with short utterances. In our current scenario we must deal with very short utterances, much shorter than those used in telephone channel experiments. Hence some reduction of the expected improvements seems reasonable. The final step is the inclusion of the new model FBPLDAUP on top of the initialization, carried out by AHC with a PLDAUP model. In Table 5.3 we include those experiments with the new model architecture for MGB 2015 evaluation subset. Two different criteria of hypothesis selection have been evaluated: prior speaker number estimation and ELBO choice. The experiments include those results obtained with the new approach as well as a line indicating those results previously obtained with the non-UP models. According to the results shown in Table 5.3, the FBPLDAUP shows potential benefits despite final results are overcome by traditional FBPLDA. When comparing setups with a fixed number of speakers our results clearly improve those obtained by FBPLDA. However, our new model does not get any benefit from the ELBO stop criterion while FBPLDA does. A possible 93 PLDA tree-based clustering 6.2 PLDA tree-based clustering The PLDA tree-based clustering is a statistical generative solution to the diarization clustering task. Hence, it follows the principle described in Section 2.6.2, identifying the target partition Θfor the set of embeddings Φas the one maximizing P(Φ,Θ). In order to do so it exploits the tree perspective of the clustering problem and the path decoding strategies. The application of statistical strategies on top of a tree structure associates a probability to each node. Regarding the leaves this probability is P(Φ,Θm) : m= 1..BN, i.e. the quality metric for each partition. In order to obtain a similar probability for the remaining nodes we make use of the product rule of probability. This rule allows the decomposition of a generic P(a1, ..., aN)as: P(a1, ..., aN) =P(a1)P(a2|a1)Pa3|a2 1...P aN|aN−1 1 = N Y j=2 Paj|aj−1 1P(a1)(6.1) where aj 1represents the set of elements {a1, ..., aj}. This definition can also be expressed in a recursive way: Paj 1=Paj|aj−1 1Paj−1 1(6.2) The application of the product rule of probability to our diarization problem is direct, substituting the jth element ajfrom the previous equation by the jth pair of variables, consisting of the embedding φjand its cluster identity label θj. Hence, the probability for any partition P(Φ,Θ) can be decomposed as: P(Φ,Θ) = N Y j=2 Pφj, θj|φj−1 1, θj−1 1P(φ1, θ1)(6.3) and its alternative recursive definition: Pφj 1, θj 1=Pφj, θj|φj−1 1, θj−1 1Pφj−1 1, θj−1 1(6.4) In consequence, each node at depth jin the clustering tree Thas the probability Pφj 1, θj 1;θj 1∈Ωθj 1associated. According to this decomposition we are assuming Φas a sequence of ordered embeddings to be clustered. Besides, these embeddings only depend on previous speaker representations of the sequence. These assumptions are reasonable in real life, where the voice evolves along time, being specially noticeable in large segments of speech. 100 Chapter 6. Tree-Based Clustering Approaches 6.2.1 PLDA-based model Along the previous lines we explored a new perspective about clustering, representing it as a tree structure to be decoded in order to obtain the best partition. Moreover, the statistical strategy associated a probability Pφj 1, θj 1to all nodes along the tree. However, the distribution for this probability has not been specified yet. Considering speaker recognition state of the art, PLDA family models seem a powerful type of solution to apply. Thus, we must keep on transforming Pφj 1, θj 1to make PLDA definition applicable. As a first transformation, we keep on applying the product rule of probability, decomposing the probability at each node into a term depending on the embeddings, the conditional distribution, and a prior distribution of the labels. This decomposition is: Pφj, θj|φj−1 1, θj−1 1=Pφj|θj,φj−1 1, θj−1 1Pθj|φj−1 1, θj−1 1(6.5) This decomposition allows to split Pφj 1, θj 1into two simpler problems. Now, we exclusively focus on the first term, the conditional distribution of the embedding jgiven its jth label as well as previous embeddings φj−1 1and labels θj−1 1. Decisions about the other term, the prior distribution of the current label θjgiven previous embeddings and decisions will be made afterwards. Unfortunately, this conditional term is still intractable to use PLDA due to the presence of label variables. Moreover, PLDA only defines dependencies among embeddings from the same speaker. Therefore, our next transformations seek separating embeddings from the labels. We take inspiration from Chapter 4, imposing Pφj|θj,φj−1 1, θj−1 1to follow a multinomial distribution on the variable θj, a one-hot sample with Ivalues (θj={θ1j, ..., θij, ..., θIj}), where Iis the number of candidate speakers. Thus, the value θij will take the value of one if the jth embedding was generated by the speaker i, being zero otherwise. Besides, we also require φj, when belonging to cluster iaccording to θj, to be exclusively explained by those embeddings already assigned to this cluster. This subset of embeddings previously assigned to cluster iat time jis denoted by Φij. Under these to conditions we can express: Pφj|θj,φj−1 1, θj−1 1= I Y i=1 Pφj|Φijθij ;j= 1..N (6.6) The definition of the term Pφj|Φijnow makes the application of PLDA principles feasible. First, we must assume the existence of a latent variable representing the speaker information yi, which allow us to redefine Pφj|Φijas: Pφj|Φij=ZPφj|yiP(yi|Φij)dyi(6.7) 101 PLDA tree-based clustering Given this definition we can now assume that the data we are dealing with is generated by a PLDA model. For this purpose, we make use of the SPLDA conditional distribution Pφj|yi,MSPLDA: Pφj|yi,MSPLDA∼ N φj|µ+Vyi,W−1(6.8) where µis the speaker independent term, Va low rank matrix describing the speaker subspace and Wa full rank matrix explaining the intra-speaker variability space. Furthermore, the second term P(yi|Φij), the posterior distribution of the latent variable given all those embeddings previously assigned to cluster i(Φij) is also modeled according to SPLDA. Thus, its definition is: P(yi|Φij,MSPLDA)∼N yi|µyi(j),L−1 yi(j)(6.9) Lyi(j) =I+VT j−1 X k=1 θkiWV (6.10) µyi(j) =L−1VTW j−1 X k=1 θki(φk−µ)(6.11) where µyi(j)and Lyi(j)represent the estimates for the mean and variance parameters respectively of the latent variable yiwhen only j−1elements were observed. As long as jincreases these estimations should get closer to the real value. Apart from well-known definitions for both distributions, the choice of the SPLDA model also provides an extra advantage. Its Gaussian nature for both Pφj|yi,MSPLDA and P(yi|Φij,MSPLDA)allows a closed form solution to the integral defining Pφj|Φij,MSPLDA. The resulting formulation for this term is: Pφj|Φij,MSPLDA∼N φj|µi(j),Σi(j)(6.12) µi(j) =µ+Vµyi(j)(6.13) Σi(j) =W−1+VL−1 yi(j)VT(6.14) After completely defining the conditional distribution of the embeddings, we now can pay attention to the label prior distribution Pθj|φj−1 1, θj−1 1. First, we assume a simplified prior distribution by eliminating the dependence with respect to the past embeddings φj−1 1. Therefore, our prior distribution will follow the form Pθj|θj−1 1. For the resulting distribution we have opted for the Distance Dependent Chinese Restaurant (DDCR) process [Blei and Frazier, 2011], already used in diarization in [Zhang et al., 2019]. This model explains the occupation of an 102 Chapter 6. Tree-Based Clustering Approaches µW V φjyi θj δ ζ Ni I Figure 6.2: PLDA tree-based clustering Bayesian Network infinite series of clusters by a sequence of elements. Then, the assignment of the element jto any cluster exclusively depends on the occupation of clusters up to this point, i.e. according to all the previous decisions, as in our decomposition. The probability of assignment of the element jto any of the already created k= 1..K clusters is proportional to its occupation at time j, namely nk. Besides, DDCR offers the possibility to create a new cluster K+ 1 proportional to γ. The mathematical formulation for DDCR is: Pθj=k|θ(j−1) 1∝(nkif k≤K ζif k=K+ 1 (6.15) DDCR deeply matches sequential ordering and assignment problem. Unfortunately, DDCR considers reasonable a continuous transition among speakers. Applied to scenarios of speaker clustering, where speaker recording can be interleaved, seems reasonable. However, in diarization we must consider the segmentation stage, which can divide any long segment into pieces of shorter length. Thus, we can add to this distribution more chances to remain in the speaker cluster. In our proposal we do so by specifically defining the situation of remaining in the current speaker cluster, with a probability proportional to δ. This addition generates the following modification of the DDCR distribution: Pθj=k|θ(j−1) 1∝     δif k=θ(j−1) nkif k6=θ(j−1) and k≤K ζif k6=θ(j−1) and k=K+ 1 (6.16) The model P(Φ,Θ), taking into account the whole set of assumptions previously described, can be represented by the Bayesian network illustrated in Fig. 6.2 6.2.2 M-algorithm optimization Once the PLDA-based model is defined, now it is time to find the way to obtain those labels Θdiar that best explain the set of embeddings Φ. Taking into account that the Viterbi algorithm 103 PLDA tree-based clustering 1 1 2 1 2 1 2 3 12 1 2 3 1 2 3 1 2 3 1 2 3 4 1st element 2nd element 3rd element 4th element Figure 6.3: M-algorithm example for a clustering tree of depth 4. 2 paths alive reach the depth 2 through the tree (green). cannot be applied, suboptimal approaches considering trustworthy paths, as the M algorithm [Jelinek and Anderson, 1971] are still applicable. The M algorithm is an iterative solution strategy. Given a scenario with a decision tree of depth N, the M algorithm tracks a subset of Msurviving paths, i.e. those paths more likely to be the solution (in our case those with higher log-likelihood). Besides, all paths must have reached depth jwithin the tree structure. Thus, the goal is the identification of those best transitions taking the Mtracked paths from depth jto depth j+ 1. In Fig. 6.3 we illustrate an example, where a clustering tree of depth 4 is analyzed by the M algorithm with M= 2. Surviving path (green lines) have reached depth 2 within the tree. The M algorithm iterative procedure is divided into two steps, estimation and maximization. The estimation step studies how the Msurviving paths in level jevolve deeper through the tree, predicting its performance in a future scenario and making decisions in consequence. For this purpose we carry out a brute-force approach, analyzing any possible transition from the M surviving paths at depth jup to a certain extra depth d. The parameter dis a design choice and responsible for a tradeoff between accuracy and computational costs. The higher d, the wider 104 Chapter 6. Tree-Based Clustering Approaches 1 1 2 1 2 1 2 3 12 1 2 3 1 2 3 1 2 3 1 2 3 4 1st element 2nd element 3rd element 4th element Figure 6.4: Estimation step in a M-algorithm example for a clustering tree of depth 4. 2 paths alive (green) reaching depth 2 are propagated to all possible nodes at depth 3 (blue) is the exploration of the tree for unseen data and hence higher accuracy might be expected, but increasing in an exponential manner the computational costs. Hence, many systems restrict d to be equal to 1. In Fig. 6.4 we represent the estimation step applied to our previous example scenario in Fig. 6.3. Each surviving path (green line) is propagated d(d= 1) levels ahead (blue lines), evaluating for each configuration the performance at this depth. The results of the estimation step provide an overview about how the tree behaves in future steps, without compromising any decision. This choice is made during the maximization step. In this step all candidate propagations are ranked, only keeping those Mwith better score. These new Mpaths now reaching depth j+ 1 are our most promising candidates so far, and those considered for the next iteration of the algorithm. This step is represented in Fig 6.5. 6.3 Experiments For the evaluation of the new clustering approach, we will make use of Albayzín 2018, as described in Section 3.2.2. For this purpose, we consider an i-vector PLDA diarization system 105 Experiments 1 1 2 1 2 1 2 3 12 1 2 3 1 2 3 1 2 3 1 2 3 4 1st element 2nd element 3rd element 4th element Figure 6.5: Maximization step in a M-algorithm example for a clustering tree of depth 4. The two surviving paths are shown in green. 106 Chapter 6. Tree-Based Clustering Approaches Experiment DER(%) Dev. Subset Eval. Subset AHC 18.88 26.36 AHC + FBPLDA 13.90 17,79 PLDA TREE-BASED CLUSTERING 13.12 17.60 Table 6.1: DER (%) results for the PLDA tree-based clustering in Albayzín 2018. Results compared with those obtained by means of AHC with and without FBPLDA resegmentation. whose setup is: A 256 Gaussian GMM-UBM followed by a 100-dimension Total Variability matrix are responsible for the i-vector extraction. The obtained embeddings undergo centering, whitening and length normalization prior to clustering, without dimensionality reduction. Finally, the new clustering approach, the PLDA tree-based clustering, uses a 100-dimension SPLDA. This setup fits in terms of dimensions with the diarization system using the FBPLDA reclustering in Chapter 4for experiments with Albayzín 2018. In our first experiment we compare the performance of FBPLDA reclustering, obtained in Chapter 4, and our new clustering approach. As a first approximation we assume the set of embeddings Φto be arranged in temporal order. We restrict hyperparameter dto be equal to 1 for computational reasons. In this experiment we consider evaluation conditions, i.e. we only present the performance of the best hyperparameter configuration (δ,ζand M) according to Albayzín 2018 development subset. The obtained results are shown in Table 6.1. According to the obtained results, the new clustering approach, working with i-vectors, provides very little improvement with respect to the FBPLDA counterpart. However, these results show benefits in the evaluation of both development and test subsets despite containing independent shows. Therefore, we can talk about limited yet consistent improvements due to our new clustering approach. Apart from the overall score for both development and test subsets, a more detailed analysis of results can also be done. For this purpose, we study the performance per show of interest with the three clustering approaches considered along this thesis: AHC, FBPLDA and PLDA tree-based clustering. For this purpose, we analyze two different metrics: On the one hand we propose the analysis of ∆I=IORACLE−IHYP, the difference in the number of speakers between our hypothesis labels and the reference. On the other hand, we analyze the DER performance. Both analyses are shown in Fig. 6.6, including all shows in Albayzín 2018. The involved shows from the development subset are millenium and La Noche en 24 Horas (LN24H). Regarding the test subset, the shows España en Comunidad (EC), Latinoamérica en 24 Horas (LA24H), La 107 Experiments ♠✁✁✂✄✄☎♠ ▲✆✝✞✟ ❊✠ ▲✡✝✞✟ ▲☛ ▲☞✝✞✟☞✂✌ ✍✝ ✷ ✍ ✶✷ ✷ ✶✷ ✝ ✷ ✸✷ ☞❚❊❊ ❋✎✏▲✑✡ ✡✟✠ * * * ◆❛✒❡ ♦❢ t ❤ ❡ ❙❤♦✇ ✓ ■ ❘✔❧✕✖✐✈✔ s♣ ✔✕❦✔rs ✗✘ ♣ ✔r s✙✚✛ ♠✁✁✂✄✄☎♠ ▲✆✝ ✞✟ ❊✠ ▲✡✝✞✟ ▲☛ ▲☞✝✞✟☞✂✌ ✶✍ ✝✍ ✸✍ ✞✍ ✺✍ ✻✍ ☞❚❊❊ ❋✎✏▲✑✡ ✡✟✠ * * * ◆❛✒❡ ♦❢ t ❤❡ ❙❤♦✇ ❉ ✓ ❘ ✭ ✪ ✮ ✔✕✖✗✘✙ ♣✚r s✛✜✢ a) ∆Ib) DER(%) Figure 6.6: Analysis per show of a) ∆Iand b) DER(%) for AHC, FBPLDA and PLDA tree-based clustering. Analysis carried out on shows from Albayzín 2018, including development and test subsets. Mañana (LM) and La Tarde en 24 Horas Tertulia (LT24HTer) are also included. Results reflect the interquartile range for each show. Those results illustrated in Fig. 6.6 show a similar behaviour of the three types of clustering per show. Thus, those more harmful shows are common for all systems. However, our new clustering approach shows a minor interquartile range per show compared to AHC and specially FBPLDA. This reduction affects both the estimation about the number of speakers and DER. Hence the performance of the system seems more consistent per individual show or domain, although small degradations might occur. This behaviour can also be extrapolated to the whole dataset, specially considering the show La Mañana (LM). While AHC and FBPLDA performances for this show are at least 100% worse than any other show in terms of DER, the PLDA tree-based clustering achieves to behave as bad as the second worst show. This improvement is also observed in the estimation of the speaker number, with a relative 25% degradation reduction. Apart from a specific setup, we can also do an analysis studying the influence for each of the model hyperparameters δ,ζand M. For this analysis we will consider the obtained scores for any possible setup. Fig. 6.7 is our chosen graphical representation to reveal the impact for the different hyperparameters. It is composed of two parts, Fig. 6.7a where we represent the relationship between ζand DER for different values of M, and Fig. 6.7b , where we represent the relationship between δ, and DER for the different values of M. In order to include all hyperparameters in each subimage, Fig. 6.7a includes some variability per measure, illustrating the interquartile range results in terms of the missing hyperparameter, δ. Similarly, measures in 108 Chapter 6. Tree-Based Clustering Approaches ✵✵✁ ✵✂ ✵✂✁ ✵✄ ✵✄✁ ✵ ☎ ✵☎✁ ✵ ✆ ✵✆✁ ✵✁ ✵ ✁✁ ✂ ✶ ✂ ✶✁ ✂ ✝ ✂ ✝✁ ✂✞ ✂✞✁ ✄✵ ✄✵✁ ✄✂ ✄✂✁ ✄✄ ▼✂ ▼✄ ▼✆ ▼✂✵ ▼✄✵ ▼✆✵ ▼✂✵✵ ✏ ❉ ❊ ❘ ✭ ✪ ✮ ✟✠✡☛☞✌ ✐♥ t❡r♠s ♦❢ ✍ ✵✵✁ ✵✂ ✵✂✁ ✵✄ ✵✄✁ ✵ ☎ ✵☎✁ ✵ ✆ ✵✆✁ ✵✁ ✵ ✁✁ ✂ ✶ ✂ ✶✁ ✂ ✝ ✂ ✝✁ ✂✞ ✂✞✁ ✄✵ ✄✵✁ ✄✂ ✄✂✁ ✄✄ ▼✂ ▼✄ ▼✆ ▼✂✵ ▼✄✵ ▼✆✵ ▼✂✵✵ ✍ ❉ ❊ ❘ ✭ ✪ ✮ ✟✠✡☛☞✌ ✐♥ t❡r♠s ♦❢ ✎ a) ζprobability b) δprobability Figure 6.7: DER (%) results for the PLDA tree-based clustering with M-algorithm in Albayzín 2018 in terms of δ,ζand M Fig. 6.7b include some variability margins indicating the first and third quartile results in terms of ζ. The information included in Fig. 6.7 reveals many important characteristics about the model. First, the results evidence the importance of M. 20% relative improvements may be obtained as long as more and more simultaneous paths are evaluated. However, this improvement is not uniform, being any increase of Mmore significant for lower values. For higher values of M, improvements are very scarce and implying large increments in the computational costs. Other detail to bear in mind is that, except for Mhyperparameter, the influence of the remaining adjustable values (δand ζ) is in general reduced (with the exception of ζfor M= 2). variations may be around 5% relative improvement/degradations, i.e. 1% absolute DER variations. Up to this point we have only mentioned three existing hyperparameters, M δ and ζ. Nevertheless, all the results were obtained by setting the embeddings Φinto a sequential order. If the analysis of the clustering tree was complete, i.e. analyzing each one of the leaves, the impact of this ordering would be null. Nevertheless, by partially exploring the clustering tree according to limited data makes this arrangement an extra factor to take into account. Therefore, while in our previous examples we exclusively applied temporal order, i.e. we can also apply different arrangements One of the key factors when using this tree-based approach is the sequence order, specially taking into account that we are exploiting the relationships between an embeddings and its predecessors in the sequence. Thus, we must explore how the ordering affects the results. While in our previous experiments we simply made use of the temporal order, this arrangement is not 109 Short utterances as occluded utterances butions, i.e. the diarization task, usually works with even shorter segments (1s-3s) to accurately deal with speaker boundaries. Hence, improvements in this scenario are becoming more and more needed. 7.2 Short utterances as occluded utterances The short utterance problem is widely known within the speaker recognition community [Poddar et al., 2017]. The evaluation of trials by means of short utterances involves a severe degradation of performance. However, there is no standard definition of short utterance in the literature. While some works have reported losses of performance with audios containing less than 30 seconds of speech, a more severe degradation is obtained considering shorter utterances (less than 10 seconds) [Mandasari et al., 2011,Kanagasundaram et al., 2011]. This short utterance problem has also been analyzed in the Speaker Recognition Evaluations (SRE), proposed by NIST. Despite traditionally considering utterances with more than 2 minutes of audio, some of the evaluations [NIST, 2008,NIST, 2010] also include a condition in which utterances contain less than 10 seconds. This loss of performance is a consequence of a higher intra-speaker variability in the estimations with short utterances. In the literature multiple contributions have been proposed to the different steps of the speaker verification pipeline, aiming to reduce the undesired variability. The feature extraction step has been studied in different ways, attempting to provide an alternative to traditional MFCCs. In [Li et al., 2015] a multi resolution timefrequency feature extraction was proposed, carrying out a multi-scaled Discrete Cosine Transform (DCT) on the spectrogram, combining the information afterwards. Alternative works like [Alam et al., 2015] fuse different features based on the amplitude and phase of the spectrum. Other contributions are focused on the modelling stage. Factor Analysis approaches were considered in [Vogt et al., 2008] to develop subspace models to better work with the short utterances. When considering i-vector representations, compensation techniques such as [Kanagasundaram et al., 2013,Kanagasundaram et al., 2014] project the obtained representations into subspaces with low variability due to short utterances. In [Sarkar et al., 2012] it is shown that systems trained on short utterances should compensate the uncertainty due to limited audio, improving the evaluation of short audios. However, when systems must deal with audios with unrestricted length, systems should be trained on long utterances for a better performance. The balance of the Baum Welch statistics, required for the extraction of i-vectors, is also proposed in [Hautamäki et al., 2013]. Besides, DNNs have also mapped short-utterance ivectors with respect to their long-utterance counterparts [Guo et al., 2017]. Other contributions 116 Chapter 7. Study of embeddings for short utterances have also worked on the backend, specially PLDA. Another technique, originally proposed in [Cumani et al., 2013b,Kenny et al., 2013] and analyzed in Chapter 5, makes the PLDA model include an extra term to compensate the uncertainty of the i-vector, which depends on the utterance length. Finally, other strategies compensate the obtained score according to reliability metrics of the involved utterances [Hasan et al., 2013,Mandasari et al., 2013], specially its duration. This idea is extended in [Viñals et al., 2018b], where the Quality Measure Function (QMF) term studies the interaction between enrollment and test utterances. In [Vogt et al., 2010] intervals of confidence are estimated, leading towards considerable accuracy. Some works such as [Ajili et al., 2016] have studied the impact of the different phonetic content in the embedding representations. According to their results, vowels and nasal phonemes are helpful for discrimination matters. By contrast, other types of phonemes, such as fricatives or plosives, can be misleading during evaluation. Our hypothesis of work applies this idea of phonemes to short utterances. The presence of certain acoustic units boosts the performance of speaker recognition systems. However, these boosting phonemes must be in both enroll and test utterances to be effective. This match in the phonetic information goes beyond the presence of certain phonemes, also requiring a match in the phonetic distribution along the utterance. In order to explain our perspective let’s make an analogy of the short utterance problem with a similar problem, face recognition with occlusions. In the best scenario, both problems contain all possible information. Working with faces we have a complete view of the person of interest, including all the face elements (two eyes, the nose, the mouth, etc.). In speaker recognition we have complete information in an utterance that contains traces for any possible phoneme and its coarticulation. As long as the utterance gets longer and longer the complete information condition is more likely to be achieved. In this scenario performance has improved more and more as long as technologies have evolved. Now we focus on short utterances. These contain much less speech, even less than a second. A simple "Yes/No" reply to a question can constitute an utterance. Hence, short utterances are very likely to lack of phonemes. In face recognition the equivalent scenario is the recognition of partial information, where some parts such as the mouth and nose are not visible. In both cases the missing information exists, but it is unavailable. Faces always have a mouth and a nose although sometimes they can be occluded, e.g. by a scarf. Regarding speaker characterization, speakers pronounce all the phonemes of a language while talking, although few of them can be missing in a specific utterance. In our hypothesis we also consider the influence of proportion. According to our analogy of face recognition, faces present a fixed set of elements (ears, nose, mouth, etc.) with a constrained size, and located in the face in specific areas. These restrictions are always the same, 117 Formulation of the embedding extraction with short utterances regardless of the person nor any occlusion. In speaker characterization the situation is slightly different. When utterances get long enough the language imposes restrictions in the phoneme distribution. These restrictions lead to a reference phoneme distribution. The longer the utterance the more its phoneme distribution tends to the reference distribution. However, short utterances contain a much shorter message, and thus its phoneme distribution can be severely distorted. In this distortion we must take into account both the missing phonemes and those present but conditioned to the message in the utterance. This distortion may lead to utterances from the same speaker with different dominant phonemes, hence complicating the evaluation. Consequently, the short utterance problem can be interpreted as an occlusion from a complete information scenario. This occlusion may be complete, where long utterances lack from certain phonemes, or partial, in which utterances have their phonemes seen in very different proportions with respect to their counterparts. The available information about the occlusion is important to be aware of. During evaluation we compare how the two speakers pronounce all the phonemes, available or not, so unbalanced information can lead to an unfair comparison. 7.3 Formulation of the embedding extraction with short utterances Current state-of-the-art speaker verification, as described in Section 2.5, relies on the pipeline embedding-backend. Utterances are first converted into compact representations, the embeddings, which feed the decision backend to obtain the score. Among all available representations, two of the most popular ones are i-vectors and x-vectors. Both have been widely tested in speaker verification obtaining great results. First, we will try to understand how we store the speaker information in these embeddings and then study its drawbacks for short utterances. 7.3.1 General case The method to compact a variable length utterance into a fixed-length representation is similar for most embedding extraction techniques. Given the utterance O, an ordered set of Nacoustic features O={o1, ..., on, ..., oN}, we transform them by function F(·), obtaining the ordered sequence F(O) = {f1, ..., fn, ..., fN}. This function maps the original feature vector oninto the speaker characteristics subspace as the projections fn. Depending on the embedding, projection fninvolves the transformation of the feature vector onas well as a small context around (approximately 0.15 seconds). By means of this mapping we attempt to highlight the speaker particularities in the features applying linear (e.g. i-vectors) or non-linear transformations (as 118 Chapter 7. Study of embeddings for short utterances in DNNs). The function F(·)is learnt from a large data pool by data analysis, e.g. by Maximum Likelihood algorithms for i-vectors or Back-Propagation [Rumelhart et al., 1986] with DNNs. Due to the fact that each one of these projections fnonly covers a small period of time, they only have information about few acoustic units. The complete characterization of a speaker requires the study of his/her particularities for all the phonemes. These acoustic units are widespread along the utterance, thus we must combine the effect of all these projections fn. The usual method to combine the projections is its temporal average. The result is the compact representation G(O), defined as: G(O) = 1 N N X n=1 fn(7.1) This embedding G(O)keeps track of the phonetic content in the utterance O. However, we can also treat each acoustic unit independently. Many state-of-the-art embeddings, such as i-vectors, can be interpreted as the sum of Crepresentations Gc(O), one per acoustic unit, each one estimated according to Ncprojections fn. According to this reasoning we can express the embedding as: G(O) = C X c=1 αcGc(O)(7.2) The obtained expression describes embeddings as a weighed sum of Cestimations Gc(O), each one representing the estimated particularities of the speaker in a single acoustic unit. Gc(O) can also be interpreted as the resulting embedding only taking into account the data related to the phoneme c. All the contributions are weighted by the term αc, the proportion of this acoustic unit in the utterance. Therefore, embeddings are conditioned to two main parts: On the one hand the stability of the distribution of weights α={α1, ..., αc, ..., αC}. On the other hand the estimations Gc(O), the particularities per phoneme. Both benefit from large utterances. Every language has its own reference phonetic distribution. Hence the longer the utterance the more its phonetic distribution becomes like this reference. Concerning the estimations Gc(O), the more available data, the less uncertain is the estimation. The average stage is the last step in which we keep track of the phoneme distribution. As a consequence, we cannot distinguish between speaker and phonetic variability afterwards. Further steps in the embedding post-processing or the backend may transform the embedding, but all phonemes are equally treated. 119 Formulation of the embedding extraction with short utterances 7.3.2 i-vector embeddings The previously described formulation also matches with the traditional i-vectors. The i-vector modeling paradigm, already described in Section 2.5, explains the utterance Oas the result of sampling from a Gaussian Mixture Model (GMM), specific for the utterance with parameters λO. This model λOis the result of the adaptation from a Universal Background Model (UBM), a large GMM that reflects all possible acoustic conditions. This adaptation process is restricted to only the UBM Gaussian means. Besides, the shift of the GMM Gaussians is tied, and explained by means of a hidden variable wO, located in the Total Variability subspace, described by matrix T. Mathematically: µO=µUBM +TwO(7.3) where µOrepresents the supervector mean, the concatenation of the GMM component means, from the target λO.µUBM is the supervector mean from the Universal Background Model (UBM), the reference model representing the average behaviour. wOis the latent variable for the utterance O, with a standard normal prior distribution and Tis a low rank matrix defining the total variability subspace. The i-vector estimation looks for the best value for the latent variable wOso as to explain the given utterance by means of the adapted model. For this purpose, we estimate the posterior distribution of the latent variable wOgiven the utterance O. The i-vector representation w corresponds to the mean of this posterior distribution. Defined in [Dehak et al., 2011], the ivector is formulated as: w= C X c=1 TT cΣ−1 cNc(O)Tc+I!−1C X c=1 TT cΣ−1 c˜ Fc(O)(7.4) =1 N(O) C X c=1 TT cΣ−1 c Nc(O) N(O)Tc+1 N(O)I!−1C X c=1 TT cΣ−1 cNc(O)˜ Fc(O)(7.5) = C X c=1 TT cΣ−1 cαcTc+1 N(O)I!−1C X c=1 αcTT cΣ−1 c˜ Fc(O)(7.6) =Ψ−1(O, α) C X c=1 αcΓc(O) = C X c=1 αcΨ−1(O, α)Γc(O) = C X c=1 αcGc(O)(7.7) where Tcrepresents the portion of the matrix Taffecting the cth component of the UBM. Σc symbolizes the covariance matrix for the cth component of the UBM. Nc(O)and ˜ Fc(O)are the 120 Chapter 7. Study of embeddings for short utterances zeroth and centered first order Baum Welch statistics for utterance O. These statistics represent the number of samples from component cand the accumulated deviation with respect to the mean of the same component respectively. N(O)symbolizes the total number of frames in the utterance O. Finally, the term ˜ Fc(O)is the average deviation per sample of the utterance for the component cof the UBM. The formulation of i-vectors offers special characteristics. First, the value of C, the number of traced acoustic units to discriminate, is fixed in the UBM. Its value is equal to the number of Gaussian components in the UBM. Therefore, Gc(O)represents the contribution per sample to the i-vector from component c, and the weight αcis the proportion of frames assumed to be sampled from same cth component. Furthermore, i-vectors have no speaker awareness in their formulation. They simply store the variations in the acoustic units within an embedding. These deviations from the average behaviour, properly treated by the backend, are responsible for the performance in speaker identification systems. 7.3.3 Short utterances Now we consider the short utterance scenario. According to the previous analysis, embeddings work well if the distribution of acoustic units αis similar to the reference distribution and the particular contributions Gc(O)are estimated with low uncertainty. These two requirements are reassured as long as the utterance contains more and more data. Concerning short utterances, their low amount of data makes them likely to have their distribution of acoustic units αfar from their reference. For the same reason short utterances may also suffer from large uncertainty in their phoneme estimations Gc(O). Hence, degradation in short utterances can be explained by the following reasons: •Errors in the contribution of phonemes. Some contributions Gc(O)were estimated with very little information. Then the uncertainty of their estimation increases. Multiple values within this uncertainty range as Gc(O)′can be estimated instead, committing the error E=Gc(O)′−Gc(O). •Mismatch in the phoneme distribution. The distribution of the weights αdoes not match the reference α, defined by language characteristics. This degradation causes the error E=PC c=1(αc−αc)Gc(O). The extreme case happens when some acoustic units are not present in the utterance, i.e. they are missing. In this situation their weight αcare equal to zero, also forcing the missing estimations Gc(O)to be set to zero, as if they were occluded. The degradation due to the mismatch in the phoneme distribution is compatible with the errors in the contribution of phonemes. 121 Effects of the short utterances in i-vectors Traditionally errors have been attributed to the contributions per acoustic unit. This is specially true when traditional embeddings, e.g. i-vectors, include an uncertainty term in its calculations. For this reason, this sort of error was the first attempted to deal with, e.g. [Kenny et al., 2013]. However, to the best of our knowledge no previous work has covered the degradation due to the phoneme distribution, which can cause similar levels of degradation. 7.4 Effects of the short utterances in i-vectors The phonetic distribution in an utterance has important implications during the embedding extraction. Embedding shifts due to incorrect contributions Gc(O)are complementary to those created by the mismatch in the phonetic distribution. In this section we illustrate their impact with i-vectors. This choice of well-known embeddings makes the study of both problems more illustrative in a simple way. For this purpose, we propose a small dimension i-vector experiment to test the effects of short utterances in some artificial controlled data. Given an evaluation UBM i-vector pipeline, we compare the i-vectors obtained from an original utterance and those obtained from the same utterance after undergoing controlled short-utterance modifications. These modifications affect both the acoustic unit distribution αand their contributions Gc(O). We make use of the following experimental setup: We first sample a large artificial data pool from a UBM i-vector pipeline. This data pool consists of more than ten thousand independent utterances, with one hundred two-dimension samples each. The UBM is a 4-Gaussian GMM whose components are located in (0,0),(0,10),(10,0) and (10,10), all of them with the identity matrix as covariance. The generative i-vector extractor has a 3-dimension hidden variable subspace. With ten thousand of these utterances we train our evaluation pipeline, an alternative UBM i-vector system. For simplicity we share the generative UBM. Regarding the i-vector extractor, we train a model with only a two-dimension latent subspace. This dimension reduction between generation and evaluation has been considered to imitate real life, where the generation of data is a too complex process that we only can approximate. From the remaining data pool we choose two extra utterances, unseen during the model training, for evaluation purposes. Because these two utterances are independent, we assume them to represent two different speakers. In Fig. 7.1 we represent them, red and blue respectively. The representation includes three parts: In the first part we show the original feature domain, i.e. the utterance set of feature vectors. Each ellipse in the figure represents the distribution of each Gaussian in their GMMs. The image also includes in green the representation of the UBM model. The second part in Fig. 7.1 represents the same red and blue utterances in 122 Chapter 7. Study of embeddings for short utterances the latent space by means of the posterior distribution of the latent variable w. The third part of Fig. 7.1 illustrates the location of the particular estimations per component Gc(O)for the two utterances in the latent space. Reddish estimations correspond to the red speaker while bluish ellipses represent the phonemes for the blue speaker. ✲ ✲✁ ✵ ✁  ✻ ✽ ✶✵ ✶✁ ✶ ✲ ✲✁ ✵ ✁  ✻ ✽ ✶✵ ✶✁ ✶ ▼❋❈❈ ✂st ❉✐♠❡♥s✐♦♥ ✄ ☎ ✆ ✆ ✷ ✝ ❞ ✞ ✟ ✠ ✡ ✝ ☛ ✟ ☞ ✝ Pr✌ ❥✍❝✎✏✌✑ ✏✑ ✒✍❛✎✉r✍ ❙♣❛❝✍ ✲ ✲✁ ✵ ✁  ✻ ✽ ✲✂ ✲✁ ✲ ✄ ✵ ✄ ✁ ✂ ■☎✈❡❝t♦r ✶st ❉✐♠❡♥s✐♦♥ ✆ ✝ ✞ ✟ ✠ ✡ ☛ ☞ ✷ ✌ ❞ ✍ ✎ ✏ ✟ ✌ ✑ ✎ ☛ ✌ P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔ ✲ ✲✁ ✵ ✁  ✻ ✽ ✲✂ ✲✁ ✲ ✄ ✵ ✄ ✁ ✂ ■☎✈❡❝t♦r ✶st ❉✐♠❡♥s✐♦♥ ✆ ✝ ✞ ✟ ✠ ✡ ☛ ☞ ✷ ✌ ❞ ✍ ✎ ✏ ✟ ✌ ✑ ✎ ☛ ✌ P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔ a) Data domain b) i-vector domain c) Components in Iv. domain Figure 7.1: Scenario of interest. a) Utterances red and blue in the featuredomain, with the UBM components in green. b) Utterances red and blue in the i-vector domain. c) Projections of the GMM components in the i-vector domain for utterancesred (reddish ellipses) and blue (bluish ellipses). Following the described setup we can carry out an analysis of degradation in short utterances. First, we illustrate the phoneme dependent estimation error due to limited data. For this reason we estimate the posterior distribution of the embeddings for multiple utterances only differing the number of samples. The distribution of phonemes αremains unaltered. Theoretically, the embeddings should not suffer any bias, but its uncertainty should get larger as long as the utterances contain less data. In Fig. 7.2 we compare the original utterances to those obtained with one fifth of the data and one tenth of the data. Fig. 7.2 illustrates the posterior distribution of the latent variable for the short utterances (dashed-line red and blue ellipses) as well as the original utterances (red and blue ellipses with continuous line respectively). The location of the ellipse represents the mean of the posterior distribution while its contour the uncertainty. As expected, the original reference utterance and their shorter versions present very reduced shifts among themselves, with almost concentric ellipses. While the blue speaker suffers almost no degradation, the red speaker biases are more noticeable. Besides, the illustration shows that the less data in the utterance, the bigger the uncertainty of the estimation. Now we study the impact of the distribution of acoustic units αon the embedding. In the reference utterances this distribution was uniform, this is, 25% of the samples came from each component. We now modify this distribution for both utterances, red and blue. In Fig. 7.3 we show the posterior distributions of the original utterances (red and blue ellipses with continuous 123 Effects of the short utterances in i-vectors ✲ ✲✁ ✵ ✁  ✻ ✽ ✲✂ ✲✁ ✲ ✄ ✵ ✄ ✁ ✂ ■☎✈❡❝t♦r ✶st ❉✐♠❡♥s✐♦♥ ✆ ✝ ✞ ✟ ✠ ✡ ☛ ☞ ✷ ✌ ❞ ✍ ✎ ✏ ✟ ✌ ✑ ✎ ☛ ✌ P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔ Figure 7.2: Comparison of posterior distribution of the i-vectors with reference phoneme distribution. Continuous line ellipse represents the original utterance while dashed-lined ellipses illustrate utterances with the limited data. line) as well as the altered short utterances (dashed-line red and blue ellipses). In the illustrated example half of the feature vectors are sampled from a single component of the GMM while the remaining data is evenly sampled along the other components. We have studied the effect with the four components in the GMM. Illustrated results in Fig. 7.3 reveal the relevance of the distribution of phonemes αfor its proper modelling. The modification of the distribution of weights makes the red speaker to offer four different representations of the same embedding. Besides, these representations are not overlapped among themselves, beyond the uncertainty region from the original utterance. Therefore, these alternative embeddings are likely to fail. Nevertheless, not all speakers behave equally. Whilst red speaker is degraded, our blue speaker has suffered the same alterations without any visible shift on his/her embeddings. The scenario with a distorted phoneme distribution can be taken to the limit. In this situation some components do not contribute to the final embedding. This scenario is the most adverse, significantly modifying the distribution of patterns αand some estimations per phoneme Gc(O) being set to zero. In this experiment we have disturbed the distribution of acoustic units α forcing two of the components to zero. In Fig. 7.4 we illustrate the six possible scenarios in terms of the non-contributing components. The results are shown for the two test speakers red and blue, with continuous line ellipses for the reference utterances and dashed-line ellipses for their altered versions. According to the representations shown in Fig. 7.4, embeddings from utterances with missing components experiment large biases with respect to the reference 124 Chapter 7. Study of embeddings for short utterances ✲ ✲✁ ✵ ✁  ✻ ✽ ✲✂ ✲✁ ✲ ✄ ✵ ✄ ✁ ✂ ■☎✈❡❝t♦ r ✶st ❉✐♠❡♥s✐♦♥ ✆ ✝ ✞ ✟ ✠ ✡ ☛ ☞ ✷ ✌ ❞ ✍ ✎ ✏ ✟ ✌ ✑ ✎ ☛ ✌ P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔ ✲ ✲✁ ✵ ✁  ✻ ✽ ✲✂ ✲✁ ✲ ✄ ✵ ✄ ✁ ✂ ■☎✈❡❝t♦ r ✶st ❉✐♠❡♥s✐♦♥ ✆ ✝ ✞ ✟ ✠ ✡ ☛ ☞ ✷ ✌ ❞ ✍ ✎ ✏ ✟ ✌ ✑ ✎ ☛ ✌ P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔ 1st Component 2nd Component ✲ ✲✁ ✵ ✁  ✻ ✽ ✲✂ ✲✁ ✲ ✄ ✵ ✄ ✁ ✂ ■☎✈❡❝t♦ r ✶st ❉✐♠❡♥s✐♦♥ ✆ ✝ ✞ ✟ ✠ ✡ ☛ ☞ ✷ ✌ ❞ ✍ ✎ ✏ ✟ ✌ ✑ ✎ ☛ ✌ P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔ ✲ ✲✁ ✵ ✁  ✻ ✽ ✲✂ ✲✁ ✲ ✄ ✵ ✄ ✁ ✂ ■☎✈❡❝t♦ r ✶st ❉✐♠❡♥s✐♦♥ ✆ ✝ ✞ ✟ ✠ ✡ ☛ ☞ ✷ ✌ ❞ ✍ ✎ ✏ ✟ ✌ ✑ ✎ ☛ ✌ P✒✓ ❥✔✕✖✗✓✘ ✗✘ ✙✚✛✔✕✖✓✒ ❙♣❛✕✔ 3rd Component 4th Component Figure 7.3: Comparison of posterior distribution of i-vectors with modifications in the phoneme distribution α embeddings. These shifts are more significant than those previously seen with less extreme distortions in the phoneme distribution α. Some of the hypothesized embeddings are far beyond the uncertainty from the original utterance. The biases suffered by the utterances are not the same for both speakers. Again, the blue speaker suffers no relevant degradation. This behaviour fits in our hypothesis because the missing components scenario is the limit case of phoneme distribution degradation. In all our experiments the red speaker has suffered from strong degradations while the blue speaker has remained almost unaltered. This different behaviour is a consequence of the locations of the phonetic estimations Gc(O)for each speaker. On the one hand, as shown in Fig. 7.1, our blue speaker has its components very close to each other, providing robustness against distribution modifications. On the other hand our red speaker has its components much further 125