scieee AI-readable full text Open interactive document viewer

Oesophageal speech: enrichment and evaluations

Raman, Sneha

Abstract

167 p.

Full text

Oesophageal Speech: Enrichment and Evaluations Doctoral thesis presented by Sneha Raman within the Language Analysis and Processing programme under the supervision of Prof. Inmaculada Hern´aez Rioja and Dr. Eva Navas Sneha Raman Doctor of Philosophy University of the Basque Country (UPV/EHU) 15 November 2021 (cc)2022 SNEHA RAMAN (cc by-nc-sa 4.0) Declaration I declare that this thesis was composed by myself and that the work contained therein is my own, except where explicitly stated otherwise in the text. (Sneha Raman) 3 Dedicated to my late father for showing me through words and actions what unconditional support really means. You were one of a kind. May more people have fathers like you. 4 Acknowledgements A PhD came to me like an unexpected visitor. I was not prepared for it, nor was I seeking it. However, it has left a very profound impact in my life, unlike anything else I have experienced before. Through this process I have gained many life skills and lessons. To manage and take care of this unexpected visitor, I was also sent a host of people and organisations to whom I express my sincerest gratitude. With your presence and participation, you have elevated my doctoral experience. First and foremost, I thank you Prof. Inma Hernaez for taking me under your wing and teaching me to fly, even at times when I was not willing to. Your direction and dedication has made me a more confident person and researcher. Thank you for pushing me whenever I have been lazy, and for patting my back whenever I have done well. Long lab lunches, being stranded at airports and a celebration of scotch whisky post a rejected conference paper are all incidents I will remember vividly for years. They have surely helped in extending our relationship outside of work and in adding some sense of humour in an otherwise very serious profession. I thank you Dr. Eva Navas for your invaluable contributions towards my PhD. The one semester course you gave us early on in the PhD was one of my finest experiences as a student. Your efficient and crystal clear advice has helped me tackle many conflicts and difficult situations throughout the PhD. I have and will always admire the elegance with which you carry yourself and your work. Thank you Dr. Axel Winneke and Dr. Clara Martin for giving me your knowledge and expertise in neuroimaging which has improved my research tremendously. Working with you has opened my eyes to a whole new world of research and research methodologies. I thank my colleagues Xabier Sarasola, Luis Serrano, Itxasne Diez, Ibon Saratxaga, Jon Sanchez, David Tavarez, Victor Garcia, Inge Salomons and Eder del Blanco for being the salt and pepper of my PhD. Without you guys this PhD would not have been so interesting and balanced. Thank you Xabier and Luis for some very important contributions to my PhD. Some of your ideas have been strong foundations for my research. Ibon, your effervescent personality i and readiness to help with all my roadblocks is much appreciated. Jon, thanks for being the go to person for any technical difficulties and ensuring that the office pantry is up to date! Itxasne, you have advised and helped me like a big sister. You never lose a chance to have a good big laugh, a reminder to all of us to take things with a good spirit! Inge, although you appeared in my life in the final stages of my PhD, your impact has been huge. Thanks for your wise company and making the thesis writing stage more enjoyable. Thank you Peio for being a strong pillar throughout this process and for keeping me physically and mentally healthy. You declared me a doctor long before I finished my PhD and brought me back into the game whenever I had thoughts of quitting. You believed that I will finish this more than myself. You pushed me to achieve goals whether it was a journal paper, 50 burpees or climbing the last hundred steps to reach the mountain top. This PhD is a result of the endurance I built with you. Thanks also to Garbi˜ne and Javi, your lovely parents who gave me so much love and affection and rushed to help me whenever I needed it. A big thank you to the ENRICH family for providing me the warmth and nourishment in this journey. All the regular update meetings and exchange of ideas helped me be on track and organised and actually finish this PhD. I will fondly remember the getting together and venting of frustrations, fun outings and fancy dinners that I shared with the ENRICH members. Amidst all the gains, I experienced a big loss too. I lost my father who was my mentor and advisor and the person from whom I have received the most affection. But he has left me with all of his positive energy which has helped me finish this thesis and will help me in my future endeavours too. I want to thank my sister Shreya for coming into my life and for being by my side to experience all the fun and tragic moments of life. Life would be so bland without you. And finally, but most importantly, I want to thank my mother for being my constant cheerleader whether I realised it or not. Thank you for all the love. ii Abstract After a laryngectomy (i.e. removal of the larynx) a patient can no more speak in a healthy laryngeal voice. Therefore, they need to adopt alternative methods of speaking such as oesophageal speech. In this method, speech is produced using swallowed air and the vibrations of the pharyngo-oesophageal segment, which introduces several undesired artefacts and an abnormal fundamental frequency. This makes oesophageal speech processing difficult compared to healthy speech, both auditory processing and signal processing. The aim of this thesis is to find solutions to make oesophageal speech signals easier to process, and to evaluate these solutions by exploring a wide range of evaluation metrics. First, some preliminary studies were performed to compare oesophageal speech and healthy speech. This revealed significantly lower intelligibility and higher listening effort for oesophageal speech compared to healthy speech. Intelligibility scores were comparable for familiar and nonfamiliar listeners of oesophageal speech. However, listeners familiar with oesophageal speech reported less effort compared to non-familiar listeners. In another experiment, oesophageal speech was reported to have more listening effort compared to healthy speech even though its intelligibility was comparable to healthy speech. On investigating neural correlates of listening effort (i.e. alpha power) using electroencephalography, a higher alpha power was observed for oesophageal speech compared to healthy speech, indicating higher listening effort. Additionally, participants with poorer cognitive abilities (i.e. working memory capacity) showed higher alpha power. Next, using several algorithms (preexisting as well as novel approaches), oesophageal speech was transformed with the aim of making it more intelligible and less effortful. The novel approach consisted of a deep neural network based voice conversion system where the source was oesophageal speech and the target was synthetic speech matched in duration with the source oesophageal speech. This helped in eliminating the source-target alignment process which is particularly prone to errors for disordered speech such as oesophageal speech. Both speaker dependent and speaker independent versions of this system were implemented. The outputs iii of the speaker dependent system had better short term objective intelligibility scores, automatic speech recognition performance and listener preference scores compared to unprocessed oesophageal speech. The speaker independent system had improvement in short term objective intelligibility scores but not in automatic speech recognition performance. Some other signal transformations were also performed to enhance oesophageal speech. These included removal of undesired artefacts and methods to improve fundamental frequency. Out of these methods, only removal of undesired silences had success to some degree (1.44 % points improvement in automatic speech recognition performance), and that too only for low intelligibility oesophageal speech. Lastly, the output of these transformations were evaluated and compared with previous systems using an ensemble of evaluation metrics such as short term objective intelligibility, automatic speech recognition, subjective listening tests and neural measures obtained using electroencephalography. Results reveal that the proposed neural network based system outperformed previous systems in improving the objective intelligibility and automatic speech recognition performance of oesophageal speech. In the case of subjective evaluations, the results were mixed - some positive improvement in preference scores and no improvement in speech intelligibility and listening effort scores. Overall, the results demonstrate several possibilities and new paths to enrich oesophageal speech using modern machine learning algorithms. The outcomes would be beneficial to the disordered speech community. iv Resumen Despu´es de una laringectom´ıa (es decir, extirpaci´on de la laringe), el paciente ya no puede hablar con una voz lar´ıngea sana. Por lo tanto, estos pacientes deben adoptar m´etodos alternativos de habla, como el habla esof´agica. En este m´etodo, el habla se produce utilizando aire tragado y las vibraciones del segmento faringoesof´agico, que introduce ruidos y efectos no deseados y produce una frecuencia fundamental anormal. Esto dificulta el procesamiento del habla esof´agica en comparaci´on con el habla sana, tanto el procesamiento auditivo como el procesamiento autom´atico de se˜nales. El objetivo de esta tesis es encontrar soluciones para lograr que las se˜nales de habla esof´agica sean m´as f´aciles de procesar tanto por parte de los humanos como de los ordenadores y evaluar dichas soluciones explorando una amplia gama de m´etricas de evaluaci´on para encontrar la m´as adecuada para este problema. En primer lugar, se realizaron algunos estudios preliminares para comparar el habla esof´agica y el habla sana. Esto revel´o una inteligibilidad significativamente menor y un mayor esfuerzo de escucha para el habla esof´agica en comparaci´on con el habla sana. Las puntuaciones de inteligibilidad fueron comparables para los oyentes familiarizados y no familiarizados con el habla esof´agica. Sin embargo, los oyentes familiarizados con el habla esof´agica expresaron menos esfuerzo en comparaci´on con los oyentes no familiares. En otro experimento, se concluy´o que el habla esof´agica supone un mayor esfuerzo de escucha en comparaci´on con el habla sana, aun para locutores con un nivel comparable de inteligibilidad. Al investigar los correlatos neuronales del esfuerzo de escucha (es decir, la potencia de la banda alfa) mediante electroencefalograf´ıa, se observ´o una potencia alfa m´as alta para el habla esof´agica en comparaci´on con el habla sana, lo que indica un mayor esfuerzo de escucha. Adem´as, los participantes con peores habilidades cognitivas (es decir, menor capacidad de memoria de trabajo) mostraron mayores valores de potencia en esta banda alfa. A continuaci´on, utilizando varios algoritmos (con enfoques tanto preexistentes como novedosos), se transform´o el habla esof´agica con el objetivo de hacerla m´as inteligible y con menor exigencia de esfuerzo de escucha. El enfoque novedoso consisti´o en un sistema de conversi´on de v 8.7 Other Activities and Achievements . . . . . . . . . . . . . . . . . . . . . . . . . . 98 8.7.1 Awards ..................................... 98 8.7.2 Workshop and Conference Attendances . . . . . . . . . . . . . . . . . . . 98 8.7.3 ResearchVisits................................. 99 8.7.4 PublicEngagement............................... 99 A 30 sentences Used in the Experiment in Chapter 4 123 B EEG Data Terminology and Procedures 125 B.1 EEG Recording Equipment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 B.1.1 CapandElectrodes ..............................125 B.1.2 Gel........................................126 B.1.3 Amplifier ....................................126 B.1.4 Recordingsoftware...............................126 B.2 Synchronisation.....................................127 B.3 EEGRecordingProcess ................................127 B.4 RawEEGData.....................................128 B.5 Cleaning Up Raw EEG Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129 B.5.1 Filtering.....................................129 B.5.2 Independent Component Analysis . . . . . . . . . . . . . . . . . . . . . . 130 B.5.3 Epoching ....................................131 B.6 Data of Interest in Clean EEG Data . . . . . . . . . . . . . . . . . . . . . . . . . 132 C Contents of the Passage Corpus Described in Section 3.3.2 133 C.1 Passage1 ........................................133 C.2 Passage2 ........................................134 C.3 Passage3 ........................................134 C.4 Passage4 ........................................135 C.5 Passage5 ........................................135 D Contents of the Words Corpus Described in Section 3.3.1 137 E Resumen 139 E.1 Descripci´on del problema . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 139 E.2 Recogida y preparaci´on de datos . . . . . . . . . . . . . . . . . . . . . . . . . . . 140 E.3 Evaluaci´on preliminar de los datos del habla esof´agica . . . . . . . . . . . . . . . 140 E.4 Enriquecimiento ....................................141 xii E.5 Evaluaci´on del habla esof´agica enriquecida . . . . . . . . . . . . . . . . . . . . . . 141 E.6 Relevancia cient´ıfica de los resultados . . . . . . . . . . . . . . . . . . . . . . . . 142 E.7 Relevancia de los resultados para la sociedad . . . . . . . . . . . . . . . . . . . . 143 xiii xiv List of Figures 2.1 Differences in the anatomy of HS and OS . . . . . . . . . . . . . . . . . . . . . . 8 2.2 Differences in the signal and spectrogram of HS and OS . . . . . . . . . . . . . . 10 2.3 Differences in the calculated fundamental frequencies of HS (top) and OS (bottom) 10 2.4 An infographic explaining the distinction of intelligibility and LE. The level of redness in the head represents the level of LE. . . . . . . . . . . . . . . . . . . . . 16 2.5 A simplified description of the voice conversion process . . . . . . . . . . . . . . . 21 2.6 Basic building blocks of a voice conversion system . . . . . . . . . . . . . . . . . 21 3.1 Comparison of the automatic labelling, manual labelling and customised automaticlabellingofOS.................................. 28 3.2 ASR Results. Mean speaker-wise Word Error Rates (in %) for ASR trained with HS and ASR trained with OS . . . . . . . . . . . . . . . . . . . . . . . . . . 29 4.1 Preliminary LE and SI Task Schematic Representation . . . . . . . . . . . . . . . 35 4.2 Mean speakerwise ‘All words’ and ‘content words only’ Word Error Rates (WER) for ‘familiar’ and ‘not familiar’ listeners. OM1, OM2, OM3, OF1 are oesophageal speakers; HM1 and HF1 are healthy speakers. Higher WER corresponds to lower intelligibility. Error bars show 95% confidence intervals. . . . . . . . . . . . . . . 37 4.3 Mean speakerwise LE for oesophageal (OM1, OM2, OM3, OF1) and healthy (HM1, HF1) speakers. On the y-axis, 1 corresponds to least effortful and 5 to most effortful. Error bars show 95% confidence intervals. . . . . . . . . . . . . . . 39 4.4 Word Error Rates (WER) for Human Speech Recognition (HSR) and Automatic Speech Recognition (ASR) for oesophageal (OM1, OM2, OM3, OF1) and healthy (HM1,HF1)speakers ................................. 39 4.5 LE Task Schematic Representation . . . . . . . . . . . . . . . . . . . . . . . . . . 42 4.6 SI Task Schematic Representation . . . . . . . . . . . . . . . . . . . . . . . . . . 42 xv 4.7 WER and LE for oesophageal (OF1) and healthy (HF1) speakers. Error bars show 95% confidence intervals. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43 5.1 WER scores and subjective LE for HS and OS . . . . . . . . . . . . . . . . . . . 51 5.2 Boxplot for average alpha frequency (8-12Hz) power for HS and OS for centroparietal channels (P3, P4, PZ, C3, CZ, C4, CP1, CP2, CP5, CP6, CPZ) . . . . . 51 5.3 Topography plot for average alpha frequency (8-12Hz) power for HS and OS. The channels O1, O2, TP1, TP2, FP1, FP2 were excluded from the analysis due to noise in those channels in the majority of participants . . . . . . . . . . . . . . 52 5.4 Scatter plots for alpha power and digit span scores for HS and OS . . . . . . . . 53 6.1 The proposed OS-HS VC system: BLSTMSS . . . . . . . . . . . . . . . . . . . . 60 6.2 ASR 1 WER and PWC scores for unprocessed OS (source), the BLSTMSS converted outputs and target SS (target). Error bars show standard errors. . . . . . 62 6.3 ASR 2 WER and PWC scores for unprocessed OS (source), the BLSTMSS converted outputs and target SS (target). Error bars show standard errors. . . . . . 63 6.4 ASR 3 WER and PWC scores for unprocessed OS (source), the BLSTMSS converted outputs and target SS (target). Error bars show standard errors. . . . . . 63 6.5 ASR scores for the multi-speaker system containing 11 OS speakers. . . . . . . . 64 6.6 ASR scores comparison for the single speaker and the multi-speaker BLSTMSS system. ......................................... 64 6.7 STOI scores for the four OS speakers and the enriched versions. Reference signal for STOI is duration-matched SS. Error bars show standard errors. . . . . . . . 65 6.8 STOI scores for the multi-speaker system containing 11 OS speakers. Reference signal for STOI is duration-matched SS. Error bars show standard deviations. . 66 6.9 STOI scores comparison for the single speaker and the multi-speaker BLSTMSS system. Reference signal for STOI is duration-matched SS. Error bars show standarderrors. .................................... 66 6.10 Histogram plots for the preference scores of the four speakers separately and All together......................................... 67 6.11 Unprocessed OS signal (bottom) and OS signal with pauses removed (top) . . . . 70 6.12WavenetSynthesis ................................... 71 7.1 LE and SI Task Schematic Representation . . . . . . . . . . . . . . . . . . . . . . 77 7.2 Word Recognition Task Schematic Representation . . . . . . . . . . . . . . . . . 78 7.3 Distribution of Alpha Power Value for all the 32 participants. . . . . . . . . . . . 79 xvi 7.4 Distribution of Alpha Power Value for a subset of 3 participants. Each participant has a separate range of alpha power values. . . . . . . . . . . . . . . . . . . . . . 80 7.5 Distribution of z-scored alpha power for a subset of 3 participants. . . . . . . . . 80 7.6 Distribution of Normalised alpha power (normalised by max value) for a subset of3participants. .................................... 81 7.7 Distribution of baseline corrected alpha power for a subset of 3 participants. . . . 81 7.8 Speech Intelligibility (SI) task performance for the three enrichment systems (BLSTMHS, PPG and BLSTMSS), HS and OS. Error bars show 95% confidence intervals.......................................... 82 7.9 SI scores for isolated words. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 82 7.10 Average Listening Effort (LE) for the three enrichment systems, HS and OS. Error bars show 95% confidence intervals. . . . . . . . . . . . . . . . . . . . . . . 83 7.11 Response times (RT) of SI and LE tasks for the three enrichment systems, HS and OS. Error bars show 95% confidence intervals. . . . . . . . . . . . . . . . . . 84 7.12 Alpha Power for the five conditions . . . . . . . . . . . . . . . . . . . . . . . . . . 85 7.13 Average Alpha power for all participants as the experiment progresses. The x axis represents the presentation order including all conditions. . . . . . . . . . . 85 7.14 Alpha power progression across blocks (time) for OS, PPG, BLSTMHS, BLSTMSS andHS ......................................... 86 7.15 WER scores for Unprocessed OS, previous systems (PPG and BLSTMHS) and BLSTMSS as calculated by ASR 3 for speaker 02M3. Error bars show standard errors. ......................................... 86 7.16 STOI scores for Unprocessed OS, previous systems (PPG and BLSTMHS) and BLSTMSS for speaker 02M3. Reference signal for STOI is duration-matched SS. Error bars show standard errors. . . . . . . . . . . . . . . . . . . . . . . . . . . . 87 B.1 EEG recording software showing impedance values of electrodes. . . . . . . . . . 127 B.2 RawEEGdata .....................................128 B.3 RawEEGdata .....................................129 B.4 Band pass filtered EEG data (1-45 Hz) . . . . . . . . . . . . . . . . . . . . . . . . 130 B.5 ICA components for EEG data from one participant . . . . . . . . . . . . . . . . 131 xvii xviii List of Tables 2.1 LE rating labels, their English translations, and the values assigned. . . . . . . . 17 4.1 Mean Word Error Rate (WER) and Listening Effort (LE) for OS and HS for familiar and not familiar listeners. . . . . . . . . . . . . . . . . . . . . . . . . . . 38 6.1 Average stimulus duration, speaking rates and intelligibility of the four OS speakers 59 6.2 ASR scores for unprocessed and enriched OS for high (02M3) and low intelligibility (16M3) OS. Text marked with * shows numbers where improvement was observed ........................................ 72 7.1 Median LE Ratings for OS, HS and the three enrichments . . . . . . . . . . . . . 83 xix xx Chapter 1 Introduction “For millions of years, mankind lived just like the animals. Then something happened which unleashed the power of our imagination. We learned to talk and we learned to listen. Speech has allowed the communication of ideas, enabling human beings to work together to build the impossible. Mankind’s greatest achievements have come about by talking, and its greatest failures by not talking. It doesn’t have to be like this. Our greatest hopes could become reality in the future. With the technology at our disposal, the possibilities are unbounded. All we need to do is make sure we keep talking.” — Stephen Hawking I would like to introduce you to this thesis by bringing some questions to your attention: What is the role of verbal communication in our lives? How important is smooth and efficient verbal communication? What is the effect of our speech on its receivers? What importance does our voice play in forming our identity? What if some key characteristics of our speech are lost? How can science and technology restore these lost characteristics? Are these restorations helpful? These questions are the starting points to the problems dealt with in this thesis. In this introductory chapter, you will read about the importance of speech communication and difficulties of disordered speech communication. I will also explain the reason for choosing to work on this problem and the possible ways to address it. 1 Figure 2.1: Differences in the anatomy of HS and OS laryngectomee with a disability to speak. In spite of the absence of vocal folds, laryngectomees can still manage to speak using alternative methods. As mentioned before, OS is one of the alternative methods of speaking after a laryngectomy. The other alternative speaking methods are Electrolaryngeal Speech (ELS) and Tracheoesophageal Speech (TOS) [61]. Both ELS and TOS require external aids—a tracheoesophageal prosthesis in the case of TOS; and electrolarynx (an external electronic substitute for the vocal folds) for ELS. The TOS prosthesis must be changed periodically (around 6 months) by a physician. In the case of OS, the vibrations in the food passage, or to be more correct, the vibrations in the pharyngo-oesophageal segment are used as a source to generate speech (See Figure 2.1) and that is the reason it is called oesophageal speech. Unlike TOS and ELS, OS does not require any external equipment. It is a skill that is developed with training from a speech therapist and requires several months of practice. It is also of poorer quality compared to TOS or ELS [133]. Nonetheless, OS has the advantage that once the skill is mastered, the laryngectomee is self-sufficient in producing speech and this makes it a promising option post-laryngectomy. Moreover, for ELS or TOS users, getting OS skills is beneficial as it can help them communicate during unexpected situations (lost or broken devices, low battery, etc.). A comparison of the anatomy of HS and OS is shown in Figure 2.1. We can observe that the HS production system has a larynx and vocal folds with which we produce speech. However, as the anatomy is altered after a laryngectomy, two separate passages are created: one for breathing (light blue line) and the other for food (dark blue lines). 8 2.1.2 Challenges Unlike HS, which is produced with vibrations from the vocal folds, OS is produced from the vibrations of the pharyngo-oesophageal segment (See Figure 2.1). Air is swallowed, inhaled, or injected and is introduced into the oesophagus, after which it is expelled with control, thereby producing vibration [133]. This generation mechanism introduces acoustic artefacts and makes OS difficult to understand [155, 92], which greatly affects communication, social activities, and hence, quality of life [90]. Moreover, these less intelligible voices are not well received by machines that are operated by speech input. An increase in the popularity of devices with voice-based interaction means that machine intelligibility is gaining importance too. We know that machine recognition, or Automatic Speech Recognition (ASR) performance, is lower compared to Human Speech Recognition (HSR) performance [73], although better ASR systems are being built increasingly, aiming towards human-like recognition abilities [132]. However, this is for HS and not for disordered speech, like OS. Figure 2.2 shows the differences in the signal (top) and spectrogram (bottom) characteristics of HS and OS for a voiced phoneme. As we can see in the top figure, HS has clear periodic lines and OS does not. This irregular periodicity in OS affects its fundamental frequency and prosody, which are important characteristics in expressive communication as well as speaker identity. Please note that the OS signal has lines (although unclear) with less frequency compared to the HS signal. This is indicative of the much lower pitch that an OS speaker appears to have. In Figure 2.3, we can see the results of the pitch estimation process by Praat [9], a software used to study and process speech signals. The figure contains the signals (top) and the pitch curves (bottom) for the word ’abrigadero’ spoken by a male HS speaker and a male OS speaker. The vertical blue lines on the speech signal (top) represent pulses and the corresponding blue curves (bottom) represent the estimated pitch curve or the f0 curve in Hertz. For the HS pitch curve, we can see a continuous non-breaking blue string around the 128.9 Hz point, indicating a continuous pitch in a normal range of male fundamental frequency (100 to 150 Hz). On the other hand, for OS, the pulses are few in number and the calculated pitch curve is around 395.3 Hz. This is incorrect as the OS speakers appear to speak in a much lower fundamental frequency and not such a high frequency as calculated. The results on this figure demonstrate the erroneous calculation of fundamental frequency of OS signals by a standard speech software. The observable differences between OS and HS characteristics in the signal and spectrogram are supported by the following studies. Cervera et. al [11] and Liu et. al [74] observed that 9 Figure 2.2: Differences in the signal and spectrogram of HS and OS Figure 2.3: Differences in the calculated fundamental frequencies of HS (top) and OS (bottom) 10 the formant frequencies were higher and the duration of vowels was longer for laryngectomees as compared to HS. Fundamental frequency, intensity, and signal-to-noise ratio of HS were significantly higher than those of OS [74]. Jitter and shimmer were significantly lower for HS compared to OS [123, 74]. On average, OS had about 10 dB less intensity compared to HS [156]. Vowel duration is longer for OS compared to HS [74, 13]. Generally speaking HS has higher speaking rate compared to OS. However, some OS speakers do have speaking rates matching that of HS [123]. 2.2 Evaluation Metrics For any task aimed at making improvements to an existing situation, it is important to conduct evaluations at the beginning and at every stage of progress. This process can tell us clearly whether we are indeed making improvements, or at least progressing in the right direction. In this section, you will find a summary of several evaluation metrics, how they have been used and on what kind of speech they have been used. This knowledge is useful to identify the metrics that are interesting and not yet explored in the context of OS enrichment or the enrichment of any other disordered speech. 2.2.1 Intelligibility Intelligibility is the most important and the most commonly evaluated property of speech. Speech is useful as a successful communication tool only if it is correctly understood. Speech intelligibility (SI) is a widely researched field, and several intelligibility measurement metrics (subjective and objective) have been explored and analysed. Subjective measures are based on the responses or opinions of the listeners. They may also be derived from listener responses. Objective measures of intelligibility are calculated using a formula, an algorithm or a software program. Intelligibility may also be categorised as HSR-based or ASR-based. Both involve getting transcriptions of the speech utterance and calculating transcription errors. A review of HSR and ASR methods can be found in [122]. Objective Measurements of Intelligibility Objective measurements of intelligibility include Speech Transmission Index (STI), Articulation index (AI), Signal-to-Noise Ratio (SNR), Harmonics-to-Noise Ratio (HNR), Short Term Objective Intelligibility (STOI) [134] and ESTOI [58]. 11 STI takes into account the noise and reverberation, and measures the reduction of performance due to speech of varying intelligibility [51]. Measuring AI involves dividing the speech spectrum into several frequency bands and a calculation of the speech-to-noise ratio in each band [66]. For both these measures, a value of 0 means poor intelligibility or no speech was heard and a value of 1 means high intelligibility or speech was perfectly heard. STOI [139] and ESTOI [58] are intrusive objective intelligibility measures which are known to be correlated with subjective intelligibility scores for noisy speech. An intrusive intelligibility measurement requires a degraded signal and an aligned reference signal. The STOI and ESTOI scores range from 0 to 1 where 0 means the least similarity to the clean reference signal (not intelligible) and 1 means totally similar to the clean reference signal (very intelligible). On the other hand non-intrusive intelligibility measures do not require a reference signal. Some examples of such measures include the non-intrusive STOI [1], methods based on Deep Neural Networks (DNN) [168, 117] and methods based on auditory-inspired filterbank analysis [35]. The main advantage of objective measurements is that they are easy to replicate and to implement. There is no need for a human subject to evaluate the speech, thereby avoiding hassles usually associated with such perceptual experiments. These metrics measure how a certain speech is received by machines or digital devices. Several of these measurements, including STOI and ESTOI, have high correlations with listening test scores [150] and hence are often used in the place of conducting and actual listening test. Other newer approaches of intelligibility measurement can be found in [1, 130, 150]. Objective measures of intelligibility can also be obtained from an ASR system. An ASR system, as the name suggests, recognises a speech utterance and gives a text output which is the transcription of the utterance. When the transcription matches the message presented in the utterance, then the ASR system is said to have recognised the utterance with 100 percent accuracy, or 0 percent errors. Such an utterance would have 100 percent intelligibility as per that specific ASR system. Currently such a high performance is not obtained with any ASR system. The lowest error (WER) so far is 2 to 3 percent [59] even with clear HS. By running speech utterances through an ASR system and then calculating the errors in transcription, we can assign an objective intelligibility score to the utterance. The errors in transcription of an ASR system are computed using a metric called Word Error Rate (WER). WER is obtained by calculating the Levenshtein distance [71] between the reference sentence and the hypothesis sentence (the sentence transcribed by the listener). The Levenshtein distance takes into account the insertions, deletions, and substitutions that are 12 observed in the hypothesis sentence. The calculation was performed with the Word Error Rate Matlab toolbox [105]. The formula used is shown in Equation 2.1. WER =Substitutions +Insertions +Deletions Total number of words in reference sentence.(2.1) Another way to evaluate the performance of an ASR system is to use the Percentage Words Correct (PWC) method. PWC is the percentage of words from the reference sentence correctly identified in the transcribed sentence. In this case it is not a measure of error, but of accuracy. Higher PWC means higher intelligibility. The PWC formula is shown in Equation 2.2. PWC =Words correctly identified in transcription Total number of words in reference sentence ×100.(2.2) There are some other objective metrics that do not measure intelligibility, but other important characteristics of intelligible speech. For example, Mel Cepstral Distortion (MCD) evaluates differences in the mel cepstra by calculating the Euclidean cepstral distance of a signal from a reference signal. It is often used to evaluate speech synthesis systems [63] and Voice Conversion (VC) systems [126, 124, 19]. See Section 2.3.1 for details on VC. The lower the MCD, more alike the signal is to the reference signal. On the other hand, Perceptual Evaluation of Speech Quality (PESQ) evaluates quality of the signal. This is done by aligning the reference and degraded signal, performing an auditory transform and extracting distortion parameters from the difference of these two signals. PESQ provides a prediction of a subjective MOS for the degraded signal [116]. Subjective Measurements of Intelligibility Subjective intelligibility involves ratings or responses from an actual human listener. It needs proper experimental design and setup and well-thought recruitment of participants to obtain reliable measures. Although this is time consuming, there is an obvious advantage with this approach as the evaluations come from an actual person which is the case in face-to-face communications in the real world. Some metrics in subjective intelligibility tests such as Mean Opinion Scores (MOS) and Speech Reception Threshold (SRT) are explained below. A MOS is a very simple and versatile metric. It is very easy to implement and useful to get a numeric grading of any characteristic that needs to be evaluated. However, MOS values are unidimensional and prone to ”misuse and misinterpretation” [137]. Despite of these limitations, MOS is a very popular method of evaluation be it for speech synthesis systems, VC systems or for other speech similarity and naturalness evaluations [154, 75, 152]. A variation of MOS is 13 the Comparative Mean Opinion Scores (CMOS) where the intelligibility (or any other speech property) is compared for two different signals. This method is useful when comparing the performance of two different systems. SRT is defined as the minimum hearing level at which 50 percent of the speech material is understood by the listener. This metric is used mainly in hearing impaired research [151]. Another popular method is to conduct a sentence recognition task. Sentence transcription tasks or ”human speech recognition” tasks [73] have been widely used for subjective intelligibility measurements [93, 62]. Yorkston et. al [167] reported the agreement of sentence transcription tasks with listeners’ estimates and ratings of intelligibility. A variation of the sentence transcription task is the sentence repetition task where instead of writing down the sentence heard by the listener, they simply have to repeat what they heard. This is less taxing for the listener. Moreover, the strengths of sentence repetition tasks are that they are “fairly simple cognitive tasks” and that they are “consistent throughout the age span” in the area of neurophysiological tests [84]. For sentence repetition or sentence transcription tasks as well, WER or PWC that was explained in the previous section can be used as an intelligibility measure. Therefore, although the responses in such tests can be subjective, in the sense that it comes from a subject, the error in transcription or recognition can be measured objectively. This method was most predominantly used to calculate SI in the experiments of this thesis. Some alternatives to sentence recognition tasks are word [17] or digit recognition tasks [159], sentence last word recognition tasks [67] and keyword recognition tasks [6]. Both last word recognition and word recognition tasks were also used for evaluations in the experiments of this thesis. Intelligibility of Disordered Speech Intelligibility measurements for the analysis of disordered speech have been explored in [76, 86, 87]. Some of them are ASR-based [76, 87], while others are not [86]. In an evaluation of disordered speech resulting from cerebral palsy and Amyotrophic Lateral Sclerosis (ALS), a high correlation was reported between STOI and subjective intelligibility ratings [55]. Intelligibility of OS was found to be lower compared to HS and intelligibility improved for OS in the combined auditory-visual mode compared to the auditory only mode [52]. In another study, Holley et. al [50] found higher intelligibility for HS compared to OS as well in the quiet condition. Some studies have been conducted in measuring the intelligibility of Spanish OS. Authors of [88] studied the voice intelligibility characteristics for Spanish OS and TOS. This HSR study was conducted for two-syllable words, and it was reported that nasal sounds resulted in most 14 transcription confusions for OS. The work in [77] describes a real time recognition system for vowel segments of Spanish OS. 2.2.2 Listening Effort While intelligibility is an important property of speech, measures of intelligibility cannot quantify how much effort was required to correctly understand the message. While poorly intelligible speech is usually difficult to understand, sometimes even highly intelligible speech is difficult to understand. For instance, if you are in a noisy room, you can still understand the person next to you, but the process is much more difficult and tiring. This is where LE is helpful as it gives an idea of the effort involved in listening to the message, and the resultant fatigue from prolonged listening to effortful speech. In literature, LE has been defined as ”the mental exertion required to attend to, and understand, an auditory message” [104]. Another related concept is that of cognitive load. The accomplishment of a listening task, especially in adverse conditions, involves the use of cognitive resources of the listener. These resources are present in the listener in the form of working memory. Moreover, each listening task has three aspects that decide the cognitive load or the listening effort of that task: intrinsic load (complex sentences, grammar, poor speech production etc.), the load on the listener for semantic processing, and extraneous load (distracting speakers or images, noise etc.). Therefore the exerted cognitive load depends on the amount of available working memory and how much of it goes into each of these kinds of loads [2]. For example, listening to a very complex sentence in a noisy environment would entail more cognitive load compared to a quieter environment. Figures 2.4 shows an infographic of the concept of LE. As you can see, the speaker is saying ’I need some water’. Both listeners understand it as ’I am some water’. Therefore, they have both understood 75 percentage of the message correctly (1 word wrongly identified out of 4). However, listener 2 has to exert more effort to decipher the message compared to listener 1, which is indicated by the redness in the head. This might depend on several factors such as the cognitive and the hearing capacities of the listener, the listening environment etc. For example, if listener 1 were to listen to the same speech in noise, he would have needed to exert more effort. The effort might also depend on the speech itself. We aim to explore all these LE factors in the context of OS and normal hearing listeners. LE has been measured in several ways: Self-reporting (questionnaires, ratings etc.); behavioural measures (performance in single tasks or multiple tasks and deriving LE from them); and physiological measures (electroencephalography, pupillometry etc.). A review of LE and 15 Figure 2.4: An infographic explaining the distinction of intelligibility and LE. The level of redness in the head represents the level of LE. various methods of measuring LE is presented by McGarrigle et. al [80]. LE has been measured in several contexts where the listeners have to put in extra investment of their neurocognitive resources. This includes understanding of distorted speech signals. Distortion can come from the speaker (e.g., foreign accent, disordered speech), the listener (e.g., hearing impairment) or from the environment or channel (e.g., noise). While there is research on LE in the context of speech in noise [115], non-native or accented speech [10, 149], and hearing impairment [46], LE of disordered speech is a less researched field, and physiological LE measurements, even less so. Self-reported Listening Effort Studies on LE often involve subjective ratings where participants are asked to indicate the perceived effort, for example, using Likert scales [112] or visual analogue scales [79, 94]. Self-reporting measures or subjective ratings are based on a set of questionnaires such as ‘Do you have to concentrate very much when listening to someone or something? ’; ‘Can you easily ignore other sounds when trying to listen to something? ’; and ‘Do you have to put in a lot of effort to hear what is being said in conversation with others?’. Based on the response of participants (say from 0 to 10), the LE can be calculated. While this method is easy and does not require much expertise, there are limitations as the responses are subjective and the threshold for effortful listening can be different for different individuals. For example, what 16 is effortful for a subject may not be as effortful as for another subject as they may have less tolerance to effort or may base their effort rating on the extent to which they could complete the task. [80] In the experiments of this thesis, two kinds of self-reported LE are used. The first one is a very simple 5point Likert scale response to the question ’How effortful was the speech to listen to?’. This method was used in the first preliminary experiment (See Section 4.2). The second method used in this thesis is the 14-point scale known as the Adaptive Categorical Listening Effort Scaling (ACALES) procedure [65]. This scale is used widely in speech-in-noise research [115, 41]. As this scale is designed for speech-in-noise experiments, the top most level is the ’only noise’ level. This is the case where the participant hears only noise and no speech. As we are not studying the effect of environmental noise while listening to OS, but the effort associated due to the disordered speech, we have not used this top level. Therefore our modified ACALES scale to measure self-reported LE is a 13-point scale that goes from ’Ning´un esfuerzo’ (No effort) to ’Much´ısimo esfuerzo’ (Extreme effort). There are 7 labels in all, but also intermediate labels (’–’) that allows participants to choose in-between options. The LE rating labels, their English translations, and their values are presented in Table 2.1. This rating scale was used in the experiments in Chapter 4, 5 and 7. Table 2.1: LE rating labels, their English translations, and the values assigned. LE Rating Labels English Translations Values Assigned Much´ısimo esfuerzo Extreme effort 13 – – 12 Mucho esfuerzo A lot of effort 11 – – 10 Esfuerzo considerable Considerable effort 9 – – 8 Esfuerzo moderado Moderate effort 7 – – 6 Poco esfuerzo Little effort 5 – – 4 Muy poco esfuerzo Very little effort 3 – – 2 Ning´un esfuerzo No effort 1 Listening Effort from Behavioural Data In the case of behavioural measures, the effort is measured by behavioural tasks, which can include single tasks or multi-tasking. In single tasks, the participant is given a listening task and their responses are recorded. In addition to the responses, the response times may also indicate LE as they are known to be slower for challenging listening conditions. In a dual task 17 24 Chapter 3 Corpus and Stimuli “Words are the coins making up the currency of sentences, and there are always too many small coins.” — Jules Renard My first introduction to this doctoral research was listening to recorded speech from several OS speakers of varying severities. As a first time listener of OS, I was both taken aback and intrigued. The following few months were spent familiarising with a database of OS speakers that we recorded in our lab. This chapter describes the database in detail and the insights gained from it which facilitated the design of the enrichment and evaluation experiments. 3.1 OS Database - Already Available The OS database contained sentences, words and sustained vowels recordings from over 30 OS speakers. The speakers are 32 laryngectomised patients who are members of ‘Association of Laryngectomees in Bilbao’. There were two categories of speakers in the database: proficient and non-proficient speakers. Proficient speakers are those who finished their speech therapy sessions months before the recording session. Non-proficient speakers on the other hand were still undergoing the speech therapy sessions. Out of these 32 speakers, 26 of them were proficient speakers, 2 of them were non-proficient speakers, 2 speakers were recorded when they were proficient as well as non-proficient and 1 other speaker was recorded in proficient OS and TOS modes. In total there were 32 speakers and 34 sets of recordings. In this thesis, we have only used the proficient speakers. The data was recorded using 4 different microphones - A studio microphone (Neumann TLM 103), an instrumentation microphone (Behringer ECM8000), a headphone microphone (DPA 4066-F), a condenser microphone (AKG C542BL) in an acoustically isolated room. However 25 the recordings used in the studies of this thesis were from the studio microphone. The protocols of recordings and the detailed description of the database is available in [39] and [123]. Some important details of the database that are relevant to this thesis are described below. 3.1.1 Sentences The database contained recordings of 100 phonetically-balanced sentences selected from a bigger corpus [121, 33]. The selection of sentences was performed with a greedy-algorithm-based tool called corpusCRT [128], with the criteria of maximised diphone coverage and a maximum of 15 words per sentence. These 100 sentences were chosen to be phonetically balanced to ensure maximum phonetic content variability. They were syntactically and semantically predictable, but had some proper nouns and many unusual words that are hard to guess. This is to be kept in mind while considering intelligibility measurements. The difficulty of these sentences make them appropriate for intelligibility and LE experiments. There are short and long sentences in the database which allows the experimenter to choose stimuli as per the need of the experiment. Some examples of the sentences are the following: ‘¿Qu´e diferencia hay entre el caucho y la hevea?’ What is the difference between rubber and hevea?, ‘Unos d´ıas de euforia y meses de aton´ıa.’ A few days of euphoria and months of despair. Words such as hevea (specific name of rubber tree) and aton´ıa (despair) are not commonly used words and hence are difficult to guess. 3.1.2 Sustained Vowels Each OS speaker recorded 4 instances of the sustained articulation of all five Spanish vowels. These data were not used in the experiments described in this thesis. However, they were great tools to understand the voice characteristics of OS. 3.1.3 Words The database contained 14 isolated words which included four words containing diphthongs. These words are useful for spoken term detection tasks. These words were not used in any of the experiments either. 3.2 HS Database Description - Already Available HS samples were obtained from an online platform [32] and hence, were recorded in variable environments. However, some of them were recorded in the aforementioned acoustically isolated 26 room, although with a different microphone. The number of speakers in the HS database keeps growing as it is an open online platform 1. There were over 35 speakers in the HS database at the time of the initial experiments. Some of the speakers in this database were used in the experiments performed in this thesis. 3.3 Additional OS Data - Newly Recorded 3.3.1 Words We recorded 150 low frequency difficult words from one OS speaker and one HS speaker. The words were chosen using an online database 2. These words were used in an experiment in Chapter 7. See the list of words in Appendix D. 3.3.2 Continuous Speech In addition to the 150 words, five continuous speech passages were recorded for the same one OS and one HS speaker. The passages were chosen from the Cervantes resources for reading for intermediate level Spanish 3. This was meant to be used in the experiment described in Chapter 7, but eventually was not used. However, these passages are useful for future OS studies with continuous speech. For one of the passages, a video recording was performed. This passage with the video recording (duration: 2 minutes and 17 seconds) was used to build an interactive demonstration of the outcomes of OS enrichment. More details on this demonstration will be presented in Chapter 6. See the contents of the passage in Appendix C. 3.4 Manual Labelling In order to create phonetic labels of the speech signals in the database, an automatic forced aligner which is part of an ASR system was used. However, using a forced aligner which was trained using HS was unfit for OS. On manual inspection, several errors were encountered in the labels likely owing to poor spectral and temporal characteristics of OS. Therefore, new acoustic models were made using the available OS recordings with the Montreal Forced Alignment tool [72]. 1https://aholab.ehu.eus/ahomytts/ 2https://www.bcbl.eu/databases/espal/idxword.php 3https://cvc.cervantes.es/aula/lecturas/intermedio/ 27 Figure 3.1: Comparison of the automatic labelling, manual labelling and customised automatic labelling of OS Manual labelling was performed for one speaker (Speaker ID:05M3) so that these manual labels can be used to evaluate the accuracy of the automatic aligner. It was not possible to perform manual labelling for all 32 speakers as it is a time consuming process. The process involved listening to each and every phoneme in the speech samples and correcting the positions of the respective phone boundary. This was done using the wavsurfer software [131]. We assessed the accuracy of the automatic labelling process by comparing the label boundary positions of the automatic labelling process to that of the manual process. Results showed that 97 percent of the errors were less than 50 ms and 83 percent of the errors were less than 5 ms [39]. Figure 3.1 shows how the the default automatic labelling and the customised automatic labelling compare to the manual alignment. These new labels were used in training an ASR system adapted to OS [39] as well as for generating Synthetic Speech (SS) with durations matching with OS (See Section 6.2). The duration matched SS was created for all the 32 speakers. This parallel SS database can be useful for future OS enhancement research. 28 3.5 Intelligibility The intelligibility of 29 proficient speakers was calculated by calculating ASR WER scores. Two different types of ASRs were used: ASR trained using HS and ASR trained with OS. Figure 3.2 shows the results of the ASR performance. It can be observed that the system trained with OS data has fewer errors compared to the system trained with HS data. There was an improvement of around 16 percentage points when trained with OS data. Figure 3.2: ASR Results. Mean speaker-wise Word Error Rates (in %) for ASR trained with HS and ASR trained with OS 3.6 Conclusions This chapter described all the data that was available to us for performing our experiments. However, not all of the data was used in the experiments described in this thesis. The experiments listed in Chapters 4, 5 6 and 7 describe in detail the chosen subsets of the database in their respective methods sections. The manual and the automatic labelling process described in this chapter form an important step in the enrichment strategies described in Chapter 6. Additionally, intelligibility of the speakers obtained from an ASR system was described. The system trained with OS data had fewer errors compared to those with HS. These WER scores are used as a criteria when selecting speakers to evaluate in the aforementioned experiments. A complete and extensive description of the database was published as a paper titled ’A 29 Spanish Multispeaker Database of Esophageal Speech’ [39]. 30 Chapter 4 Preliminary Measures of Intelligibility and Listening Effort “If you think communication is all talking, you haven’t been listening.” —Ashleigh Brilliant In the previous chapter, I described some key features of the OS database. I listed some differences in the acoustic and linguistic characteristics of OS and HS. In Section 2.1.2 we have seen that OS has a lower speaking rate, higher jitter and shimmer and lower intensity compared to HS. All these are factors that affect the ability to perceive and understand speech. To understand in what ways OS is more difficult to process compared to HS, and how much, it is necessary to conduct well designed listening experiments. In this chapter, I describe two listening experiments that collected some preliminary intelligibility and LE metrics for a small set of speakers from the database and some control healthy speakers. Both the experiments were designed to evaluate the gaps between OS and HS in intelligibility and ease of processing. Once these gaps are known, appropriate enrichment methods can be designed aimed at closing these gaps i.e. bringing enriched OS closer to the metrics of HS than unprocessed OS. The contents of this chapter have featured in previously published paper titled ’Intelligibility and Listening Effort of Spanish Oesophageal Speech’ [112]. 31 4.1 Introduction Listening to disordered speech is a challenging task, and it demands a lot of attention and effort. To quantify the challenges of listening to OS we begin by measuring its intelligibility in comparison to HS. Intelligibility measurements are common and are a useful way to quantify what percentage of the spoken message has been correctly understood (See Section 2.2.1 for more details). In this study, we have evaluated intelligibility in human–human (speaker is human, listener is a human) as well as human–machine (speaker is human, listener is a device/software) interactions. In addition to intelligibility, there is growing interest in research measuring LE and other processing load aspects of speech as it gives an additional dimension for understanding challenges in speech perception in adverse listening conditions. The motivations for measuring LE is described in detail in Section 2.2.2. In this study, we have attempted to explore LE in addition to the intelligibility measurements. In chapter 2, we presented some research were HS was found to be more intelligible than OS (Section 2.2.1) and HS was found to be more acceptable compared to OS (Section 2.2.2). These studies are the foundations for the hypotheses of this experiment. Significant positive correlations were observed in ASR and HSR for listening in adverse conditions such as age related hearing loss [37] and speech disorders [53]. This led us to hypothesise that like HSR perfomance, ASR performance will be lower for OS compared to HS. The idea of intelligibility differences between experienced and inexperienced listeners of OS was explored in [16]. The findings were that OS was ranked similarly for intelligibility by both experienced and inexperienced listeners. This was intriguing, and led us to investigate the effect of familiarity with OS on its intelligibility. In addition, as we were collecting LE ratings too, we were interested in seeing if the same was observed for LE ratings, or if they would tell a different story. We consider friends, family (spouse, siblings, children), and caretakers of OS speakers as familiar listeners. This study contains two experiments. The first experiment was web-based, and was focused on getting preliminary intelligibility and LE metrics for our data. We investigated how intelligibility (both ASR and HSR) and LE differ for the two speech types (OS and HS). We also investigated the effect of familiarity and to what extent intelligibility and LE are correlated. The second experiment (an extension of Experiment 1) was conducted in a laboratory setting, which allowed us better control of the experiment environment. The aim of this experiment was to find out if more LE is reported for OS even if the intelligibility of OS is close to that of HS. Additionally, in this experiment, we also investigated if the participants’ performance in 32 the speech perception tasks depended on their cognitive abilities. The hypotheses for Experiment 1 are: •WER is positively correlated with self-reported LE ratings. •HS is more intelligible and less effortful, compared to OS. •Listeners familiar with OS find it less effortful to process OS, compared to listeners that are not. •ASR performs worse for OS than for HS. Our hypotheses of Experiment 2 are: •For the case that intelligibility of OS is similar to that of HS, there is still more effort in understanding OS. •Listeners with better cognitive abilities have better intelligibility scores and report lesser effort. We begin by describing the materials and methods and the results of Experiment 1 in Section 4.2, followed by those of Experiment 2 in Section 4.3. Finally, a general discussion and conclusions are presented. 4.2 Experiment 1: Preliminary Word Error Rate and Listening Effort Measurements 4.2.1 Materials and Methods Experimental Design The main task for this web-based experiment was the sentence recall and transcription task. Participants listened to a sentence and then typed what they had understood. To collect LE rating measures, we asked the participants to rate the sentences for LE on a 5-point Likert scale. The sentences were played only once (to avoid any possible memory effect) and in a random order (to avoid sentence order bias). Corpus and Stimuli The corpus we used was a part of the larger database of 32 OS speakers described in detail in Section 3.1. The chosen stimuli were picked from the 100 sentences described in Section 3.1.1. 33 4.3 Experiment 2: Listening Effort for Highly Intelligible Oesophageal Speech 4.3.1 Materials and Methods Experimental Design Based on our preliminary intelligibility and self-reported LE experiment (See Experiment 1, Section 4.2), we had the chance to probe further into our data by designing a study specifically aimed at measuring LE. The aim was to investigate the differences in LE for a set of HS and OS speakers that have comparable intelligibility. As pointed out in [95], even highly intelligible OS speech was found to have different LE ratings. Therefore, this is a methodological decision in order to rule out that observed effects are due to differences in intelligibility. The experiment was designed for an EEG-based LE measurement. We aimed to record EEG data of participants while they listened to OS and HS to investigate if there are differences in the LE correlates of brain activity. Along with measuring the EEG data, we also collected subjective LE ratings from the listeners. As the participants had an EEG cap on, the stimuli were played on a loudspeaker, and not on headphones. Here, we present the LE ratings findings. EEG data acquisition process and findings are presented in Chapter 5. In addition to the LE experiment, a separate intelligibility test was conducted to replicate the results of Experiment 1 in a laboratory setting. This time we asked the participants to listen to the sentence and repeat out loud what they heard. This is less taxing for the participant as they do not have to type their responses. Also, this resulted in speedier responses and hence less effort from the listeners’ side in memorising the sentence. The advantage of oral response is that typing errors can be excluded as a confounding factor for WER. However, this involves post-processing, i.e., the task of transcription of their oral responses to text to calculate WER. In order to investigate the relationship between LE, SI and the listeners’ cognitive capacities, we conducted a Flanker task [31] and a backward digit span task [47] after the behavioural tasks. The Flanker task measures the selective attention ability [31] and the digit span task measures the working memory capacity of the listener [47]. Both of these processes are relevant in speech perception. OS has swallowing sounds and undesired artefacts which need to be ignored by the listener to selectively focus on the speech message. Working memory is also a crucial factor in speech perception (see phonological loop in [4]), especially for OS which spans a longer duration than HS. 40 Stimuli We picked a subset of one HS speaker and one OS speaker from our dataset of Experiment 1 based on intelligibility similarity. Speaker OF1 and speaker HF1 were the two speakers that had significantly similar intelligibility based on a two-sample KS test. The null hypothesis that they come from same distributions was accepted with a significance of Alpha of 0.01. All 100 sentences mentioned in Section 3.1 were used for this experiment. An intelligibility test was performed on the same 30 sentence subset described in Experiment 1. For the LE rating task, we used the other 70 sentences which were longer. This was to ensure that the participants had a sufficiently long stimulus to respond to, and that the EEG recording for each stimulus was sufficiently long to process and analyse. The sentences contained several low frequency words which made them sufficiently difficult for an LE task. For the 70 LE sentences, the number of words in each sentence ranged between 9 and 18 words (mean = 13.19, SD = 3.66). The mean duration of the OS stimuli was 8.81 seconds (SD = 1.58, min = 6.00, max = 12.55) and that of HS was 5.27 seconds (SD = 1.045, min = 3.31, max = 8.28). The lengths of OS stimuli were significantly longer (t(69)=1.66, p<0.001) than HS stimuli. Average speaking rates (syllables per second) for HS and OS were 4.32 ±1.79 and 7.36 ±3.35 respectively. Listening Test Sixteen native Spanish speakers (7 female, 9 male; age range: 19–35, mean = 26.56, SD = 4.50) participated in the study. All participants were native Spanish speakers from South America, except one who was from Spain. They were given monetary compensation for participating in the test. Ethics for conducting the experiment was approved by the local ethics committee of the University of Oldenburg. All participants had normal hearing except one participant with a 55 dB hearing loss in the left ear. The inclusion of this participant did not alter the observations of the study and hence, we chose to keep this participant. The stimuli were presented with a loudspeaker placed at a 0◦in front of the participant at distance of 1m at a comfortable listening level of 60 dB SPL. The test began with the LE task first (Figure 4.5), where 60 (30 OS and 30 HS) out of the 70 available LE sentences were played in 3 blocks of 20 sentences (randomised) each. For 15 sentences of each block, participants were prompted to provide LE ratings as per the 13-point ACALES scale (See Section 2.2.2 for more details). In the other 5 sentences (presented at random intervals), they were asked to repeat the last word of the sentence, which the experimenter scored as correct or incorrect. This was to ensure that they were attentive and actively listen41 Figure 4.5: LE Task Schematic Representation Figure 4.6: SI Task Schematic Representation ing to the stimuli. The LE task lasted for around 20-25 minutes. The average inter-stimulus interval (time between the response and onset of the next sentence) was 1.59±0.63 seconds. After the LE task, the participants got a break (approximately 10 minutes) and then they proceeded to the intelligibility task (Figure 4.6). In this task, they listened to a sentence and received a prompt on the screen to repeat the sentence that they heard. They provided oral responses for the 30 sentences. The whole session of the intelligibility test was recorded with a microphone so that it could be transcribed later. This task lasted around 15 minutes. In all, we had 45 subjective LE ratings and 15 last word recognition scores from the LE task, and 30 SI scores from the SI task for each participant. EEG data was recorded for the entire duration of the LE task and therefore were available for all the 60 LE task stimuli. Cognitive Tasks In the Flanker task, participants were presented 24 congruent (“<<<<<”), 24 incongruent (“>><<<”), and 24 neutral (“−− <−−”) stimuli. They were asked to focus on the middle symbol and correctly identify it by pressing “<” or “>” on a keyboard as quickly as possible. Their response accuracy and reaction times were recorded. The better the performance (i.e. shorter reaction times on accurate responses), the better the participant’s selective attention. In the backward digit span task, the experimenter read a list of digits and the participant was asked to repeat them in reverse (For example, experimenter: ‘3 2 9 5’; participant: ‘5 9 2 3’). The digits were read in an even tone at intervals of approximately one second. The 42 experimenter started with the set of three-digit lists. If the participant was able to correctly recall 5 three-digit lists out of 6, they graduated to the set of four-digit lists. This went on until the participant could no longer recall at least 5 lists in a set or until they reached the final nine-digit list set. The digit span score was the maximum number of digits in the list where the participant could recall 5 lists correctly. The larger the correctly recalled digit span, the larger is the participant’s working memory capacity. The cognitive tasks lasted 10-15 minutes. No EEG was recorded during the cognitive tests. 4.3.2 Analysis and Results Out of the 16 participants, transcriptions were available only for 13 participants as we could not record responses of 3 participants due to technical problems. LE ratings were not available for one other participant, also due to a technical problem with saving data. Flanker effect score was not available for a participant. For analysis purposes, these missing data were filled with the mean values of the responses of other participants. Sphericity and homogeneity checks were performed on the data with the JASP tool to ensure that assumptions of an ANOVA test are met. Intelligibility The audio responses of the sentence recognition task were transcribed by a native Spanish speaker, who was also a speech expert. WER was calculated using the same methods as elaborated in Section 2.2.1. As the WER was found to be highly correlated with all inclusive WERs in Experiment 1, we decided to proceed with all inclusive WERs only. (a) Word Error Rates (b) Self-reported LE ratings Figure 4.7: WER and LE for oesophageal (OF1) and healthy (HF1) speakers. Error bars show 95% confidence intervals. Percentage WER scores for OS was 18.88 ±5.57 and for the healthy speaker it was 11.69 ±5.07 (Figure 4.7a). ANOVA showed that speaker type had an effect on WER (F(1,15) 43 = 27.20, p<0.001, η2= 0.645). Listening Effort Ratings Mean LE (from a 13-point scale) for the OS speaker was 6.457 ±3.150 and for the healthy speaker it was 1.994 ±1.611 (Figure 4.7b). There was a difference of 6 points in median LE for OS and HS. The median LE for HS was 1 (no effort) and for OS speaker it was 7 (moderate effort). ANOVA showed that speaker type had an effect on LE (F(1,15) = 77.55, p<0.001, η2= 0.838). There were 15 responses per participant for the task of repeating the last word in the LE task. The average error made in the recognition of the last word was 1.067 response per participant. The total last word recognition error across all the 15 participants was 7 percent. We can tell, therefore, that the participants were attentive during the LE rating task. LE, WER and Cognitive Tasks The Flanker effect was calculated as shown in Equation 4.1. RT incong and RT neutral are the average reaction times taken to respond to an in-congruent trial (“>><<<”) and a neutral trial (“−− <−−”), respectively. These reactions times are calculated for correct trials only. Flanker Effect =log(RT incong)−log(RT neutral) log(RT neutral).(4.1) The mean Flanker effect score was 0.413 ±0.019 and the mean digit span score was 4.125 ±1.258. No significant correlations were found between digit span scores and Flanker scores indicating that they measure separate cognitive abilities. Correlations between Flanker effect and mean LE ratings were not significant (Spearman’s rho = −0.432, p= 0.096). Significant negative correlation (Pearson’s r = −0.554, p= 0.049) was found between digit span scores and mean WERs (average of OS and HS). 4.4 Discussion In Experiment 1, we were able to show that speaker type (OS or HS) had an effect on both LE and WER. OS speakers had poorer intelligibility compared to HS speakers and also a higher LE. The correlation between LE and WER suggests that more effort was reported as the intelligibility of the speaker worsened. Therefore, a drop in intelligibility caused an increase of LE. A further step in this direction would be to know what aspects of OS contribute more 44 to LE: Its spectral characteristics, lack of fundamental frequency, poor rhythm in speech, or a combination of these. Our findings about the effect of familiarity with the listener on intelligibility are in a similar vein to a study investigating the experience of the listener (speech expert vs. novice) on OS intelligibility [16]. However, in this study we were more interested in investigating the experience that comes from constant exposure as family members, close friends, and caretakers. We found that indeed the intelligibility scores were similar for familiar and unfamiliar listeners. However, interestingly, familiar listeners reported less LE. So LE was able to provide additional insight about listening to OS. As far as ASR is concerned, we found that ASR WER scores were higher for OS compared to HS. We compared WERs from our ASR system with HSR WERs. ASR WERs were higher compared to HSR WERs, but it could be because our ASR system was based on a unigram language model and focused only on acoustic models. The reason to choose such an ASR was to evaluate the drop in intelligibility owing to acoustic degradations, which is the case for OS. In Experiment 2, the goal was to measure LE when listening to OS and HS with similar intelligibility scores taken from Experiment 1. Although WER data in Experiment 2 indicate a higher intelligibility for HS than OS, the overall intelligibility for both OS and HS can be considered to be very high. Despite this high intelligibility for both speaker types, we observed a considerable gap in LE, and this suggests that LE is a relevant dimension to be considered in OS evaluation. The negative correlation of the digit span scores with WER suggests that participants with a poorer working memory (denoted by lower digit span scores) made more errors in recognition. This is understandable, as the ability to hold more information helps in correctly recalling and repeating the stimuli. The correlation with Flanker effect was not significant, suggesting that, in this case, selective inhibition plays a minor role to explain differences in LE. Flanker task is a measure of selective inhibition of distracting signals, such as noise added to the signals or signals with distracting speakers. However, the distractions in our stimuli are not of that nature. It is more in the form of undesired pauses and swallowing sounds that appear within and between words in the OS signal. The increase in LE was observed likely due to poorer quality of speech, rather than due to interfering information that has to be suppressed as is the case in noisy environments. On the whole, we cannot tell from these results alone whether better cognitive abilities mean better performance (low LE and low WER ) in OS speech perception. Future studies using different cognitive test batteries are necessary to help us answer that question better. 45 Finally, the familiarity effect on LE, as reported in Experiment 1, could mean that OS speakers might find it easier communicating with family and close friends as opposed to others. Although this was not investigated in this study, it would be interesting to know at what level of familiarisation does this effect show and also whether there is a ceiling effect to this familiarisation. That is, is there a point where, due to familiarity, they find OS as effortful as HS? 4.5 Conclusions We performed two different experiments to collect intelligibility and LE metrics for OS and HS. The first experiment, a web-based one, was used to collect intelligibility and self-reported LE metrics. The conclusions of this experiment were that speaker type (HS or OS) had an effect on both intelligibility and effort. There was significant correlation between WER and LE. Listeners familiar with OS fared the same for intelligibility as people who were not. However, they reported less effort in listening to OS than the not familiar listeners. The ASR intelligibility was poorer for OS compared to HS. The second experiment was to measure LE for HS and OS in a laboratory setting. The conclusions were that even if the intelligibility of OS was close to HS, there was a considerable difference in LE. LE obtained through these experiments is based on the listener’s own interpretation of ’effort involved in listening’. In the next chapter, we look deeper into LE by investigating brain activity as a physiological measure of LE, and study its relationship with this self-reported measure of LE. We have built an OS restoration system aimed at better ASR and HSR intelligibility and low LE (see Chapter 6). The methods used in this study will be used to evaluate the outputs of this system (see Chapter 7). Both HSR intelligibility and ASR intelligibility play different but important roles in OS evaluation. While improved HSR would enable better human–human interactions, an improved ASR performance would enable better human–machine interactions (e.g., digital voice assistants). Lower LE would also contribute towards improved communication with fellow humans. The evaluation of all these three metrics provides an all-round understanding of OS speech perception. Experiment 1 from this Chapter was presented at a conference [109] and the combination of both Experiment 1 and 2 was published as a paper [112]. 46 Chapter 5 Listening Effort and Oesophageal Speech: An EEG Study “Not everything that can be counted counts and not everything that counts can be counted” —Albert Einstein Recent advancements in technology has enabled us to gain in-depth understanding of the human body. One such technology is neuroimaging which allows us to have a peek into the functioning of the brain. In this chapter, I describe an experiment that explored the differences in cerebral activity when listening to HS and OS. This experiment is an attempt to answer the questions: Do listeners’ brain signals reveal any differences while processing OS and HS? What are the factors that influence these differences in the brain signals? 5.1 Introduction Speech communication requires a great deal of cognitive processing. It involves a vast network of activities such as acoustical processing, linguistic processing and emotion recognition, all performed in a very short span of time [36]. Listening to speech in (acoustically) challenging conditions increases the cognitive demand [80]. Challenging conditions can be attributed to any of the components of speech communication: sender (disordered speech, foreign accent[149]), receiver (hearing impairment [46], non-native listener [10]) or channel (reverberation, background noise [115], poor telephone connection). Moreover, listening to speech in challenging conditions for prolonged periods causes fatigue [80]. To overcome these additional challenges posed on the sensory-cognitive system of the listener, additional LE is required to understand 47 the signal of interest. For more background information on LE, refer to Section 2.2.2. We know from the previous chapter that OS is less intelligible and more effortful to listen to compared to HS. This result of LE was based on subjective ratings from listeners. As it is subjective in nature, this rating varies from person to person. Moreover, the definition of ”effortful” is very subjective. What may be effortful to one person may not be as effortful to the other person. Perhaps the other person has better cognitive abilities or they have better cognitive load bearing capacity. Our aim is to answer all these questions and to better understand the neural processes involved in speech processing and effortful listening linked to OS, with the help of EEG. The EEG activity can be decomposed into different frequency bands such as alpha (8-12Hz), beta (16-31Hz), gamma (>32Hz), theta (4-7Hz) and delta (0.5-4Hz) bands. Investigating these frequency bands has given clues to understanding problems such as speech processing in adverse conditions [153] and sensing imagined speech [28]. Alpha (8-12 Hz) power (See Section 5.2.1 for alpha power calculation procedure), particularly in parietal regions, has been found to be related to LE for speech-in-noise and is suggested to reflect the suppression of task-irrelevant information [135, 161]. Additionally, alpha power is known to increase with increasing acoustical degradation of speech [82] as well as with increasing working memory (WM) demands [98]. As OS is acoustically degraded, less intelligible and requires more LE (as per listeners’ subjective ratings), we hypothesise that this will result in a higher alpha power for OS compared to that for HS. Given the importance of cognitive functioning in speech perception, particularly of working memory, we also included individual differences in working memory functioning in the analysis. We assumed that working memory capacity could serve as buffer for LE -i.e. larger working memory capacity is associated with less LE. Some studies have looked into differences in brain activity while listening to degraded speech and control HS. Theys et al. [142] performed a study based on ERP components (See Appendix B for details on ERP components). They observed increased N100 amplitude and decreased N100 latency while listening to dysarthric speech compared to HS, indicating that the inherent degradation in dysarthric speech influenced early sensory auditory processing, and that degraded speech requires more neurophysiological resources in early processing stages. In a Near-infrared Spectroscopy (NIRS) study [114], 16 typically developing children listened to whispered speech and normally vocalised speech. A higher haemodynamic response was observed in the left ventral sensorimotor cortex (in the frontal-parietal region) for whispered speech compared to normal speech. This indicated increased cognitive effort while processing whispered speech. The degradation present in OS) has certain properties in common with 48 whispered speech such as low energy and degraded fundamental frequency. As there have not been studies on OS and brain activity, we base our experiment on the aforementioned studies conducted on other kinds of degraded speech. We expand the research on perception of degraded speech by investigating differences in the neural correlates of LE when listening to HS and OS and how they relate to subjective ratings of LE and behavioural performance (speech intelligibility) as well as the cognitive capacities of the participants. 5.2 Materials and Methods The materials and methods of this experiment were the same as ’Experiment 2’ of Chapter 4 (Section 4.3.1), in which the behavioural data results were presented. The current study presents the EEG data obtained from the same setup. Therefore this section only contains the detailed description of EEG acquisition. I will briefly summarise the experimental procedure here for the benefit of the reader. Sixteen participants listened to sentences spoken by one HS and one OS speaker. The tasks involved a SI task with 30 short sentences and a LE task with 60 longer sentences. EEG was recorded for all the 60 sentences of the LE task. In addition, cognitive tasks such as the backward digit span task and the Flanker task was conducted. For a more detailed description see Section Section 4.3.1. 5.2.1 EEG Acquisition and Analysis A continuous EEG was recorded using a 24-channel wireless Smarting EEG system (mBrainTrain, Belgrade, Serbia) at a sampling rate of 500 Hz, with a low-pass filter of 250 Hz. The 24 electrodes were attached to an elastic EEG cap (EasyCap, Herrsching, Germany) according to the International 10/20 system [57]. To record the EEG data the software Lab Streaming Layer [64] and Smarting Streamer 3.1 (mBrainTrain, Belgrade, Serbia) was used. EEGLab v.14.1.1 [18] was used offline to process and analyse the EEG data. EEG recordings were re-referenced off-line to an average of all electrodes. The EEG data was then filtered with a 0.1Hz to 45Hz bandpass filter. Excessive ocular artefacts, such as eye blinks, and other EEG artefacts were identified and corrected, using an independent component analysis as implemented in EEGLab. Epochs were extracted from the continuous EEG. The lengths of the extracted epochs were varied and was as long as the length of the entire duration of the stimulus (i.e. sentence length). 49 56 Chapter 6 Enrichment Systems “Mend your speech a little, lest it mar your fortunes.” —William Shakespeare In the previous chapters, I described characteristics and limitations of OS and the gaps in intelligibility and LE between OS and HS. In this chapter I present some experiments we performed with the aim of enriching OS. The contents of this chapter have featured in previously presented research titled: A multifaceted enrichment of oesophageal speech [110] and a previously published paper titled ’Enrichment of Oesophageal Speech: Voice Conversion with Duration-matched Synthetic Speech as Target’ [111]. 6.1 Introduction OS is less intelligible and more effortful to process compared to HS (See Chapter 4 and 5). Poor intelligibility and increased LE hinders verbal communication possibilities for OS speakers, even in non-noisy environments. Lack of intelligibility means that the OS speakers have difficulty in meeting some important needs such as telephonic conversations, calling for a medical appointment, asking for directions and ordering food. Apart from these basic needs, OS speakers find it difficult to engage in family gatherings and public speaking, and to use voice-activated digital devices. All these challenges have amplified in the COVID era with additional barriers such as wearing masks and the reduction of face-to-face interactions. Therefore, enriching OS with software interventions is a very promising aid for OS speakers to facilitate easier communication. In this chapter, I present some of my own contributions towards enrichment of OS. The methodologies ranged from simple modifications (Section 6.3) on the OS signal to elaborate DNN-based VC methods (Section 6.2). 57 6.2 Experiment 1: DNN-based OS Enrichment As stated in Section 2.3, one of the possible approaches to enrich OS is to use a VC system. The goal of a VC system is to convert the utterances of a source speaker to sound like those of a target speaker. In the OS enrichment context, utterances of an OS speaker can be mapped to a healthy speaker’s utterances, thereby having the OS acquire characteristics of HS. Like our previous approaches [126, 124], this proposed method is also based on VC. VC systems may be parallel (requires temporally aligned source target utterance pairs) or non-parallel (requires hours of speech data). Due to data limitations (100 sentences per speaker), parallel VC is best suited for our purposes. A parallel VC requires the parallel source and target sentences to be aligned for training. This is primarily done by Dynamic Time Warping (DTW) alignment which finds an optimal match based on similarities in the two sequences. The authors of [45] describe some challenges of DTW in the context of VC. One of them is the presence of silences or extra sounds in the source and not in the target. Another one is the poor estimation of end points of silences and phonemes. A third case is the many-to-one and one-to-many nature of the DTW mapping. For example, if the source contains longer durations of a phoneme, a single frame of the target may be mapped to several frames of the source. OS has undesired silences and artefacts and longer and varying durations of phonemes. These qualities make DTW challenging in the OS-HS VC task. As a workaround, in our previous attempt [126], we performed alignment at two stages: first aligning the phone boundaries and then applying DTW, anchoring the phone boundaries. In this paper, we took advantage of the available phone labels and the possibility of generating SS with explicit phone durations. This resulted in SS that matches in duration with the source OS utterances, and thus, would be a perfectly aligned target. This eliminated the need for DTW and its limitations. We hypothesise that this DTW-free VC would improve the intelligibility and quality of the enriched OS compared to our previous methods. A robust enrichment system should ideally work with OS speakers of varying speaking proficiency. Therefore, we performed enrichments for OS speakers ranging from very low to very high intelligibility. As the enrichment system is built to improve verbal communication for the OS speaker, it is important that the output of the enrichment system is preferred by listeners over the unprocessed OS. Moreover, given that voice interactions with machines are becoming more and more common, the enriched outputs should be intelligible to machines. Taking these points into consideration, we evaluated the subjective preference of the enriched system amongst human listeners as well as an objective measure of intelligibility and ASR performance. 58 To sum up, in this section, we present a novel, DTW-free, parallel VC system for OS enrichment which includes an SS target. There is a single speaker or speaker dependent system as well as a multi-speaker or speaker independent system. We evaluate its outputs for ASR performance, STOI and a preference test (for single speaker system only) in comparison with unprocessed OS. 6.2.1 Materials and Methods Data We chose four OS speakers with a wide range of intelligibility from the original corpus (See Section 3.5). In the original database, the four speakers were identified as ’02M3’, ’04M3’, ’16M3’, ’25F3’ and we continue to use these IDs. Average stimulus duration, speaking rates and intelligibility of the four speakers are presented in Table 6.1. For each speaker, we used a parallel dataset of all the 100 phonetically-balanced Spanish sentences, where the source was OS and the target was SS. The procedure followed to generate the parallel SS will be explained in the next section. As stated before, the sentences were syntactically and semantically predictable but had some low frequency words. The number of words in each sentence ranged between 9 and 18 words (mean = 13.19, SD = 3.66). Average duration per stimulus Average speaking rate ASR scores (seconds) (syllables per second) (WER in %) 02M3 7.48±1.67 4.32±1.80 56.25 04M3 9.27±2.36 3.84±1.71 74.34 16M3 12.52±3.61 2.59±1.19 90.39 25F3 7.85±2.02 4.24±1.86 43.38 Table 6.1: Average stimulus duration, speaking rates and intelligibility of the four OS speakers Proposed VC System The proposed VC system, BLSTM with SS as target (BLSTMSS), is a Neural Network based system with OS as source and SS with matching durations as target (see Figure 6.1). The procedure is described in detail in the following steps. Labelling of Oesophageal Speech Segmentation and labelling of OS is a tricky process owing to undesired artefacts, incorrect pronunciations of some consonants and unstable fundamental frequency. The forced alignment feature built into generic Spanish ASR systems such as Kaldi [107] was unsuitable for OS. Therefore, using the Montreal Forced Alignment tool [72], new models were created by using 59 Figure 6.1: The proposed OS-HS VC system: BLSTMSS OS as the training material (See Section 3.4). Automatic alignment with this forced aligner gave us the phone labels and their durations for the source OS utterances. Generating Target Synthetic Speech Using the labels, their durations and the utterance text, SS was generated by explicitly assigning these durations to the phones. The text-to-speech system used was a Hidden Markov Models (HMM) based synthesis system [34] which was originally developed for the Basque language. The Spanish version is described in [119]. This gave us equal-sized frame-by-frame aligned pairs of OS and SS. Due to constant swallowing of air to produce speech, OS contains several pauses with artefacts within utterances. During the SS generation, these pauses were replaced with silences. Voice Conversion Neural Network Voice conversion was performed with the VC recipe of the Merlin toolkit [165]. Parametrisation and resynthesis was done using the WORLD Vocoder [91]. The extracted parameters included 60 Mel Cepstral Coefficients (MCC), 1 excitation parameter (log F0), 1 Band Aperiodicity Parameter (BAP), the deltas of of the MCC, log F0 and BAP, the delta deltas of the MCC, log F0 and BAP and a voiced/unvoiced binary parameter. In all, there were 187 parameters extracted every 5 milliseconds. A matrix of size 187 X (number of 5 ms frames) of OS and SS utterances were the source and target inputs respectively. We split the 100 source-target pairs into 90 train and 10 test pairs. As the source and the target had the same number of frames, the alignment step in 60 the training process was skipped. The train parameters were normalised to 0 mean and unit variance and then fed into a 4 layered BLSTM (4 X 1024) training network. After training, the source test utterance parameters were converted using the trained model. A denormalisation of the mean and the variance was applied to the output parameters, followed by a Maximum Likelihood Parameter Generation using the variances from the training data. The resulting converted parameters were fed into the vocoder to synthesise the converted speech. A cross validation was performed 10 times, so that all the 100 sentences were available as test sentences. Multi-speaker system A multi-speaker version of the BLSTMSS method was implemented using the same 100 sentences from 11 high intelligibility OS speakers. They are the most intelligible OS speakers in the database (less than 60% WER) based on the ASR system trained with HS. These 11 speakers are rightmost 11 speakers 1in Figure 3.5 (blue bars), not including the TOS speaker 09MT. For the multi-speaker system, we combined the data from all the 11 speakers instead of performing VC for each speaker separately. Each speaker’s utterance of a certain sentence had its own different durations and hence the corresponding target signal was also different. Therefore, each utterance had a paired SS target, a total of 1100 utterances (100 sentences from 11 speakers). Ninety utterances from each speaker (the same 90 sentences from all speakers) and the corresponding target SS were put in the source and the target training set respectively. The training and conversion process was the same as that for the single speaker system described in ’Voice Conversion Neural Network’ from Section 6.2.1. 6.2.2 Evaluations and Results Evaluations involved comparing the speaker dependent BLSTMSS outputs to unprocessed OS using three ASR systems, an objective intelligibility measure and a preference test. In addition, we compared ASR scores and STOI scores of BLSTMSS with those of our previous systems. The ASR evaluation for the multispeaker system, evaluation was performed using one ASR system (ASR 3) and with STOI scores. A comparison of the speaker dependent system and the speaker independent multispeaker system is also made. ASR Evaluation for the Speaker Dependent Systems We evaluated the outputs of our proposed enrichment system using three ASR systems: the speech-to-text system from Microsoft Azure using the python azure-cognitive services-speech 1Speaker IDs: 01M3, 02M3, 03M3, 08M3, 12M3, 19M3, 22M3, 24M3, 25F3, 28F3, 29M3 61 library (ASR 1) [85], the Elhuyar speech recognition system (ASR 2) [30] and a Kaldi based system (ASR 3) [107, 126, 124]. The input files to these ASR systems were the 100 single channel speech signals sampled at 16000 Hz. The outputs were text files containing the transcriptions. The reason for using three ASR systems was to have a diverse set of evaluations. ASR 1 is a well known commercial ASR system used world wide and therefore easier for comparisons in future studies elsewhere. ASR 2 is a commercial system built locally in Spain and therefore better adapted to the speech style and vocabulary of the speakers involved in this study. ASR 3 is a customisable ASR with full control of all the components such as the language model, dictionary etc. ASR 3, which uses a limited lexicon and unigram language model was used in our previous studies [126, 124]. The advantage of this ASR is that it is not prone to updates as is the case of commercial ASRs. This allows us to make fair and accurate comparisons of our ongoing work with our previous work. We calculated two metrics from the ASR transcriptions: WER and PWC. WER and PWC were calculated using Equation 2.1 and Equation 2.2 respectively. Both the concepts are explained in Section 2.2.1. (a) Word Error Rates (b) Percentage Words Correct Figure 6.2: ASR 1 WER and PWC scores for unprocessed OS (source), the BLSTMSS converted outputs and target SS (target). Error bars show standard errors. Figure 6.2, Figure 6.3 and Figure 6.4 show mean WER and PWC scores for the 100 sentences obtained from the transcriptions of ASR 1, 2 and 3 respectively. WER scores were lower (i.e. higher intelligibility) for BLSTMSS compared to unprocessed OS for all ASRs and speakers with 2 exceptions - speaker 04M3 in ASR 1 and speaker 16M3 in ASR 2. In the case of PWC scores, a higher PWC score (i.e. higher intelligibility) was observed for the BLSTMSS samples compared to unprocessed OS samples for all speakers and ASRs. 62 (a) Word Error Rates (b) Percentage Words Correct Figure 6.3: ASR 2 WER and PWC scores for unprocessed OS (source), the BLSTMSS converted outputs and target SS (target). Error bars show standard errors. (a) Word Error Rates (b) Percentage Words Correct Figure 6.4: ASR 3 WER and PWC scores for unprocessed OS (source), the BLSTMSS converted outputs and target SS (target). Error bars show standard errors. ASR Evaluation for the Multi-speaker Systems Figure 6.5 shows the ASR scores from ASR 3 for the multi-speaker system. There was ASR improvement for four speakers (speaker 02M3, 03M3, 19M3, 22M3). For the other 7 speakers, the WER increased. Figure 6.6 shows the comparison of the ASR scores for the single speaker system and the multi-speaker system for the two OS speakers (02M3 and 25M3) that were part of both the systems. As can be observed, the multi-speaker version did not did not have better ASR scores compared to the speaker dependent system. STOI Scores for Speaker Dependent Systems We calculated STOI (See Section 2.2.1 for details) for unprocessed OS samples and converted BLSTMSS samples for the four OS speakers using the already aligned duration matched SS (target signal) as the reference signal. 63 Figure 6.5: ASR scores for the multi-speaker system containing 11 OS speakers. Figure 6.6: ASR scores comparison for the single speaker and the multi-speaker BLSTMSS system. The STOI results for the single speaker BLSTMSS method are shown in Figure 6.7. We can observe that the STOI scores have improved considerably (at least 15 percentage points) from OS to BLSTMSS for all four speakers. A high STOI score of over 62 percent was observed for all the BLSTMSS samples. STOI Scores for Multi-speaker Systems Figure 6.8 shows the STOI scores for the multi-speaker system. We can observe that the STOI scores improved for all the 11 speakers. Figure 6.9 shows the comparison of the STOI scores for the single speaker system and the multi-speaker system for the two OS speakers (02M3 and 25M3) that were part of both the systems. Both the enrichment systems have higher STOI scores, but there were no significant differences between the two systems. 64 Figure 6.7: STOI scores for the four OS speakers and the enriched versions. Reference signal for STOI is duration-matched SS. Error bars show standard errors. Subjective Tests Subjective tests were performed only for the speaker dependent systems and not for the multispeaker system. While unprocessed OS has several undesired artefacts and lacks a natural fundamental frequency, it is natural speech. On the other hand, although the BLSTMSS outputs are much clearer sounding, they are synthetically produced and may have some limitations because of that. The success of the enrichment depends majorly on whether listeners prefer to listen to the enriched version more than the unprocessed OS. Therefore, we performed a preference test to collect listeners’ opinion on whether they prefer listening to the outputs of the proposed system or the unprocessed OS. Participants listened to pairs of samples, one unprocessed OS sentence and the corresponding BLSTMSS enriched output of the same sentence. There were 10 pairs for each speaker, a total of 40 pairs. The chosen 10 pairs were the shortest sentences in the set, as that allowed us to have maximum number of evaluations while keeping the test under 20 minutes. The presentation of all the pairs, as well as the order of BLSTMSS and OS within each pair was randomised to avoid order bias. After listening to the two stimuli in each pair, the participants were asked to mark the stimulus they preferred amongst the two. The options they were given were ’Prefiero claramente la primera’ (I clearly prefer the first one), ’Prefiero la primera’(I prefer the first one), ’No percibo diferencia/Ninguna suena mejor’ (I do not perceive any difference/Neither one sounds better), ’Prefiero la segunda’(I prefer the second one), ’Prefiero claramente la segunda’(I clearly prefer the second one). 65 Speech type ASR Scores (WER in %) 02M3 16M3 Unprocessed OS 56.32 90.39 GMM VC 37.93* NA BLSTMHS VC 40.35* NA BLSTMSS VC 30.89* 61.32* Silence removal 57.99 88.95* Wavenet 56.09 91.37 Table 6.2: ASR scores for unprocessed and enriched OS for high (02M3) and low intelligibility (16M3) OS. Text marked with * shows numbers where improvement was observed 6.4 Enrichments Demonstration I created a web-based interactive demonstration simulating an OS speaker speaking with an enriched voice. For this demonstration, we recorded an OS speaker speaking a passage (See Section 3.3.2). Audio was recorded simultaneously with the same recording equipment of the original database to obtain a better quality recording. This audio was then passed through three different enrichment systems: A GMM based system (Enrichment 1), BLSTMSS (Enrichment 2) and a a pitch modified version of enrichment 2 to better suit the speaker (Enrichment 3). In the interactive demo, the user can play the video and choose amongst four options for the audio: the original speech and the 3 enrichment systems. Whenever the user chooses an audio type (by clicking the corresponding button), the video plays with the chosen audio in a synchronised manner. In this way, it is possible to visualise the OS speaker speaking in the original as well as enriched version of the speech. The demo can be found in the following link: https://aholab.ehu.eus/users/sneha/london demo/test.html. This demo was presented in the Royal Institution, London 2and as a part of a show and tell session in a virtual conference 3. 6.5 Conclusions The BLSTMSS system had better ASR scores and objective intelligibility measures compared to unprocessed OS. In recent times, communication with digital assistants and other devices is on the rise. Therefore, an improvement is this direction is desirable for efficient communication with digital devices and dialogue systems. The BLSTMSS system was preferred by listeners compared to unprocessed OS. A slight exception was observed in case of the least intelligible speaker of the set. For this speaker, 2https://www.rigb.org/whats-on/events-2020/march/public-easy-speaking-effortless-listening 3https://cmsworkshops.com/ICASSP2020/Papers/ViewPaper.asp?PaperNum=6191 72 there was not a decisive preference for the BLSTMSS method. We presume that when this less intelligible speaker gains more experience in OS with the aid of a speech therapist, their intelligibility will improve and they will gain more benefit from the enrichment. While voice conversion remains the greatest contributor in improving ASR performance, it was observed that undesired silence removal was beneficial for low intelligibility OS. Synthesis with a rich vocoder did not help in ASR improvement but it has scope for positive responses in perceptual evaluations. Some other methods to enhance OS that may benefit from further exploring are the eigenvoices method [23] and a GMM-based open source VC system called sprocket [60]. The results from the DIFFGMM algorithm in this VC system seemed promising, but a formal evaluation could not be performed. Therefore it would be useful to explore this method more and perform formal evaluations. In addition to ASR scores, STOI and preference tests, we are interested in investigating subjective and physiological LE for unprocessed OS, enriched OS and HS. This is because, while intelligibility reveals what percentage of the speech was understood correctly, it does not tell us how difficult it was to understand it. LE provides useful additional information about whether enriched OS is easier to perceive and process compared to unprocessed OS. Therefore, our future studies (Chapter 7) will focus on LE in addition to intelligibility and listener preferences. The demonstration in Section 6.4 has helped the general public, researchers and the OS speakers themselves to visualise the expected outputs of an enrichment system. The DNN-based system was published as a journal paper [111] and the light weight enrichment systems were presented in a conference [110]. 73 74 Chapter 7 Final Enrichment Evaluations “The most basic of all human needs is the need to understand and be understood. The best way to understand people is to listen to them.” —Ralph G. Nichols The final step in the OS enrichment problem is the evaluation of outputs of the enrichment processes. The aim of OS enrichment was to close the gaps in intelligibility and LE between OS and HS. In other words, we aimed to have the metrics of the enriched outputs to be as close as possible to HS. Some evaluations were presented in the previous chapter but these evaluations only compared some machine intelligibility related scores of the particular novel algorithm compared to OS. In this chapter, I describe some objective, subjective and physiological evaluations of the intelligibility and LE of all the OS enrichment tasks we have performed so far, including work from other researchers in the lab. These metrics are compared with those of unprocessed OS and clear speech (HS and SS). We shall see through the evaluations how the algorithms we adopted contributed to the enrichment of OS. 7.1 Introduction When speech is disordered, the deficiency in producing clear speech makes listening to the speech difficult. We have seen in Chapter 4 and Chapter 5 that listening to OS poses such challenges to listeners compared to HS. In Chapter 6, we described some algorithms that were developed to transform OS with the objective that listening to an OS speaker would be easier. Our findings suggest that we had success in enriching speech in some areas (ASR scores, STOI, preference scores). These results are encouraging and in many software based enrichment studies [102, 166, 12], the success of the enrichment methods are evaluated based on these objective measures or simple Likert scale based MOS scores. But these evaluations are unidimensional 75 and do not tell us the whole story. In Chapter 5, we described an in depth evaluation of OS and HS by conducting a listening experiment in a laboratory setting and measuring SI, subjective LE and EEG activity. This helped us attain a deeper understanding of the differences in HS and OS. In this chapter, I describe a similar in-depth experiment with all the major enrichment methods developed by us so far in addition to the OS and HS samples. The goal is to know which is the winning enrichment method. Ideally it would be the one that beats OS and the other enrichment methods to be more intelligible and less effortful to process. All the methods used in this experiment and the motivations to use them were explored and explained in previous chapters. This chapter is just a coming together of all the enrichments and all the types of evaluations to determine which is (if we do have one) the best enrichment system that we made. 7.2 Stimuli Five systems were evaluated in this experiment: OS, HS and three versions of enriched OS. The three versions of enriched OS were outputs of three different DNN based enrichment systems (BLSTMHS [126], PPG [124], and the new single speaker BLSTMSS method described in Chapter 6). A subjective word recognition task also included the multi-speaker version of the BLSTMSS system. In this evaluation, we have chosen samples only from speaker 02M3. Using samples from all speakers was not feasible as that would increase the conditions in the listening tests and would make them too long, making it difficult for the listeners to sustain attention. Speaker 02M3 had a high intelligibility (ASR WER: 56.25%). He was the only speaker who had performed additional recordings of words and passages which were useful for more evaluations. For all the conditions, we used all the 100 phonetically-balanced Spanish sentences described in Section 3.1. The sentences were syntactically and semantically predictable but had some difficult low frequency words. They were sufficiently difficult for a SI and LE rating task. In addition to the sentences, we used 150 low frequency difficult words for a word recognition task. The details of these words are described in Section 3.3.1 and the list of words in Appendix D. 76 Figure 7.1: LE and SI Task Schematic Representation 7.3 Experimental Procedure 32 native Spanish speakers (9 male, 23 female, Age: 25.63 ±5.11) participated in a listening test, presented with a psychopy [103] interface. Each participant listened to 100 sentences, 20 sentences in each of the 5 conditions: OS, PPG, BLSTMHS, BLSTMSS and HS. After listening to each sentence, they performed an SI task and an LE rating task. No sentences were heard more than once. All the stimuli underwent loudness normalisation in accordance with EBU R 128 Standard [29]. The 100 sentences were presented in 5 blocks of 20 sentences where each block contained stimuli from all the 5 conditions in a randomised order. The stimuli were counterbalanced across participants and conditions, such that each sentence in a particular condition was listened to an equal number of times across all participants. 7.3.1 Behavioural Tasks The behavioural tasks consisted of an SI task and an LE task with sentences and an SI task with word stimuli. In the sentences SI task, the participants repeated aloud the last word of the sentence they just heard. Their response was recorded and later checked for correctness. The SI score per stimulus was whether or not the listener got the last word right: 1 for correct, 0 for wrong. The 13-point ACALES scale (See Section 2.2.2) was used to rate LE. Unlike the previous experiments, here every stimulus had an SI as well as LE rating. Figure 7.1 shows a schematic representation of this task. An additional SI task was performed with the 150 isolated words. This was a simple test with the participant listening to the word and then repeating it. The experimenter scored the responses for correctness later: 1 for correct, 0 for wrong. Figure 7.2 shows a schematic representation of this task. In all, the test lasted 1.5-2 hours which included 30 minutes of EEG setup, 30 minutes of the combined SI-LE task, 15 minutes of the words SI task and 15-30 minutes of post test question77 Figure 7.2: Word Recognition Task Schematic Representation naires and dismounting of the EEG cap. The participants were given monetary compensation. The experiment was approved by the ethics committee of the Basque Centre on Cognition, Brain and Language. 7.3.2 EEG Acquisition and Analysis A continuous EEG was recorded while participants were performing the LE rating task. The brain activity was recorded from 32 electrode sites mounted into an elastic EEG cap (EasyCap, Herrsching, Germany) and arranged according to the International 10-20 system [57] (See Appendix B for more details). BrainVision Recorder were used to record EEG data. The EEG was recorded at a sampling rate of 1000 Hz. EEG data offline processing and analysis was conducted using EEGLab v.14 [18]. Amongst the 32 electrodes, the scalp electrodes were electrodes 1 to 27. Electrode 28 was in the right mastoid position and electrodes 29, 30, 31 and 32 were placed on the forehead and eye area to record ocular activity. An impedance of <5kOhms was ensured on each electrode during the EEG cap mounting. EEG recordings were re-referenced off-line to an average of the right mastoid electrode. The EEG data was then filtered with a 0.1Hz to 45Hz bandpass filter. Excessive ocular artefacts, such as eye blinks, and other EEG artefacts were identified and corrected, using an independent component analysis as implemented in EEGLab. Epoch extraction procedure was the same as in Section 5.2.1. Variable length epochs were extracted with durations matching the speech signals with an additional 500ms pre-stimulus as baseline. Similarly, the frequency analysis was performed as per the process described in Section 5.2.1. It involved a PSD calculation with a 99 percent overlap. Here, as the sample rate was 1000, 2000 points (corresponding to 2 seconds) were used to calculate the Fourier transform. As before, the alpha power was the mean of the power values between 8 and 12 Hz. Although electrode layouts and placements were slightly different in this experiment compared to the experiment in Chapter 5, the analysis here was also focused on the centro-parietal region 78 Figure 7.3: Distribution of Alpha Power Value for all the 32 participants. (CP1, CP2, CP5, CP6, C3, C4, CZ, P1, P2, P3, P4, P7, P8, PZ). With 32 participants, 5 conditions, 20 stimuli per condition and 12 electrodes, we had 32*5*20*12 i.e. 38400 data series. Out of these, some trials were rejected due to noisy data and we were left with 37944 data points. Alpha power is affected by individual brain activity, environmental factors, the mental and physiological state of the participant at the time of the experiment and other such factors [96]. Figure 7.3 shows the distribution of alpha power for the 32 participants for OS. Here, the ’Participant’ axis represents each of the 32 participants and the ’OS Alpha Power’ axis represents the 30 bins of alpha power values. The ’Number of Observations’ axis is the number of alpha power values observed in a particular alpha power bin. Here are two example points in the graph to help you understand the graph better. At point A, participant number 1 has 0-50 observations where alpha was in the range of 30-31 µV 2/Hz (an outlier). At point B, participant number 30 has close to 100 observations where alpha was in the range of 1-2 µV 2/Hz. As a general observation, we can see that each participant has a different alpha profile or distribution. A subset of 3 participants provides a simplified representation of the issue. Figure 7.4 shows the varying power distribution of 3 of the 32 participants. To remove all these between subjects differences and to focus only on the alpha power differences between conditions, a power normalisation was performed. There are several ways to perform normalisations for data like this such as z-scores, dividing by the max value and dividing by resting state alpha power. Each of these processes are explained in the following paragraphs. A z-scored alpha power was generated by using the mean and standard deviation of all the 79 Figure 7.4: Distribution of Alpha Power Value for a subset of 3 participants. Each participant has a separate range of alpha power values. Figure 7.5: Distribution of z-scored alpha power for a subset of 3 participants. data points from each participant. The distribution post applying the z-score is shown in Figure 7.5 for the same three participants considered before. As it can be observed, the distribution of z-scored alpha power for the three participants is overlapping. This eliminated inter subject variability. Although the z-score approach of normalising helps us get rid of the problems of intersubject variability, it assumes that the data follows a normal distribution. We can see that the alpha power data does not follow a normal distribution (See Figure 7.4). It has a slight skewness and it looks more like a log normal distribution. The second way to normalise the data is by dividing the entire dataset from a participant by the maximum alpha power of that participant. As alpha power only has positive values, this process fits all the datapoints from a participant between 0 and 1. See Figure 7.6 for the normalised alpha power distributions for the 3 participants. Another typical approach for EEG data normalisation is to divide the alpha power from one participant by the baseline resting state alpha power. The resting state corresponds to the 80 Figure 7.6: Distribution of Normalised alpha power (normalised by max value) for a subset of 3 participants. Figure 7.7: Distribution of baseline corrected alpha power for a subset of 3 participants. time when the participant is not engaged in any task. In this case, the participant was asked to focus on a cross on the screen for 60 seconds. The idea here is that this normalisation will take away the alpha power components that pertain to the participants inherent brain activity without the effect of any task performance. We divided the alpha power for each stimulus by the average alpha power during the 60 seconds of resting state EEG data. This was performed separately for each of the 12 centro-parietal electrodes. See Figure 7.7 for the normalised alpha power distributions for the 3 participants. 7.4 Results 7.4.1 Speech Intelligibility Figure 7.8 shows the averaged SI scores (% words correct or PWC scores) from the 32 participants. HS had the highest SI score (p<0.001). SI score for OS was higher than that of all the enrichments (p<0.001). Within the enrichments, the proposed system had significantly higher 81 to shorter RTs and lower LE. Our results of LE, SI and their RTs are along the same lines. The EEG activity data reiterates our findings from Chapter 5: Listening to OS entails higher alpha power compared to HS. There were no significant differences in alpha power between the OS condition and the enriched OS conditions. Although the BLSTMHS method had lower alpha compared to other enrichments, this was not significant and hence we cannot conclude if this enrichment provided any improvement in cognitive load. We observed that the alpha power increased as the experiment progressed. This may indicate increase in fatigue or decrease in alertness as was observed by Antons et. al [3] too. While shorter experiments would reduce these tendencies, it reduces the data points necessary to reach statistical significance. Therefore, such listening tests must be designed by keeping this trade off in mind. In the case of unprocessed OS, there was a drop in alpha from the fourth block to the final block. A possible explanation for this is that when the LE demands are too high, the listener often ”gives up” at the task of listening and starts exerting lesser LE. For example when you are in extreme fatigue, say, a jetlag, you cannot meet the required attentional demands and may exert lesser effort [136]. This seems to have happened with unprocessed OS in the last block, where the listener was so fatigued that they exerted a lower amount of effort. This pattern of a drop in physiological LE for stimuli above a certain level of difficulty has been often observed in many LE studies [99, 100, 164]. However, this same trend was not observed for the other conditions. A longer experiment with more blocks would be helpful in understanding this disengagement effect better. The lack of correlation between LE and alpha power, and the connection of alpha power and fatigue could mean that alpha power is an indicator of fatigue and not LE in the experiment that we performed. In general, the participant was fatigued as the experiment progressed and the fatigue was less evident for HS compared to OS and the synthetic enrichment outputs. The fatigue associated with OS reached to the extent that the participants may have disengaged from the task of listening to OS. An evaluation of systems via listening tests is time consuming and cumbersome. However, it helps us know how our improvements are received by human listeners. In the end, the aim is that the laryngectomees and the people interacting with them benefit from the enrichments. Therefore, people’s opinions in this case are crucial and valuable. In the ASR and the STOI evaluations, BLSTMSS had better scores compared to our previous systems, as well as unprocessed OS. Therefore, the enriched outputs would be useful in improving human-machine interactions. As for human-human interactions, more research 88 would be needed to develop an enrichment system that appeals more to human listeners than unprocessed OS. Nonetheless, an improvement in SI and LE compared to previous enrichment systems suggests that we have been able to move some steps forward in the OS enrichment research. 7.6 Conclusions The BLSTMSS system (developed as part of this thesis) outperformed our previous systems in STOI and ASR, as well as SI scores and subjective LE scores. When compared with OS, there was an improvement in ASR scores and STOI, but not in SI scores and subjective LE. EEG activity revealed more alpha power when listening to OS compared to HS indicating an effect of task difficulty. Additionally, alpha power increased as the experiment progressed indicating fatigue. No conclusions could be made based on the alpha power of enriched OS. Therefore, although there is an improvement in some objective measures, further improvements are needed to make the transformed OS more preferable and understandable to human listeners. The findings of this chapter are in preparation to be published as a paper titled ’Evaluation of Enriched Oesophageal Speech: an EEG Study’. 89 90 Chapter 8 Conclusions “Begin at the beginning, the King said gravely, “and go on till you come to the end: then stop.” —Lewis Carroll, Alice in Wonderland I set foot on this journey of enrichments and evaluation of OS with some very clear and straightforward aims. It is when I got into the problem that the depths and complications of this task became apparent. All the obstacles and successes in this processes have taught me some important lessons, which will be helpful for me in research as well as any problem solving activity. I cannot say that I have totally solved the problem that I aimed to solve, but surely I have progressed a bit and gained a lot of knowledge. As progress is limitless, my journey with this problem could keep going on and on. But as the King says, the journey has come to an end and I must stop. Here I summarise all the findings of my thesis and prepare my baton to be passed to future researchers who wish to take this journey forward. 8.1 A Recap of the Problem Statement and Aims OS lacks the intelligibility and quality of HS due to physical abnormalities: absence of larynx and vocal folds and separation of trachea and oesophagus. The lack of intelligibility and quality also increases the cognitive load when listening to OS. These issues impede OS speakers from social and familial engagements, expressing their needs, using voice-controlled digital devices and in general, navigating this world verbally. The main aim of this thesis was to enrich OS signals by increasing intelligibility and quality and reducing LE. The other aim was to decide and implement appropriate metrics to evaluate OS and enriched OS. 91 8.2 Methods Used and Takeaways A wide range of methodology was used to do the research in this thesis. All these methods have expanded my knowledge base and enabled me in making better research decisions. A thorough understanding of OS was gained by studying the OS database. The variety of data and some acoustical measures of that data present in the database helped me a great deal in understanding the characteristics of OS. For enrichment of OS, the main methods used were machine learning and other statistical learning tools. Machine learning is being used increasingly in a wide variety of industries such as healthcare, business, entertainment, social media, transport and other important industries. The success of machine learning in these fields is encouraging and intriguing. Through this project, I was able to gain some understanding of the inner workings of machine learning and neural networks such as learning rates, training data, normalisation, optimisation algorithms and so on. In the quest of finding a DNN system that could enrich OS, I went through several of these DNN terminologies as well as new and upcoming architectures that have shown promising results. In the case of this thesis, a DNN based method was found to be the most effective in achieving the goals of enrichment. Apart from DNNs and machine learning, a great deal of signal processing knowledge was applied too. This was essential for the understanding of speech characteristics such as fundamental frequency and formants as well as during EEG signal processing for an effective treatment of the signal. My main takeaway was that to solve any problem it is important to thoroughly understand the data in hand and have updated information about the tools available to work on that data. A great deal of expertise on software languages such as matlab and python was needed and acquired during the processes of this project. Knowing how to program in these languages as well as knowing how to use programs written by other collaborators in these languages was crucial in achieving the goals of this project. Statistical analysis was a key part of interpreting data and drawing conclusions. Providing statistical analyses made the results scientifically sound and credible. Over the course of the project I gained a lot of knowledge about various analysis tools and methods such as parametric and non-parametric tests, ANOVAs and t-tests and open source statistical tools such as JASP. I learned the importance of applying the right kind of statistical analysis on any given data set. Although I have not come out of this project as a statistics sage, I can say that I have made a few dips into the pool of statistical wisdom. And surely they have been great tools to make sense of the data that I have collected. As this is a study of speech and LE, listening tests were another major indispensable part of 92 the research. The only true and correct way to know listeners’ opinions and their response to the modifications we made was by conducting listening tests. I came across several factors that are crucial while conducting listening tests such as conducting pilot tests before the actual test. Some of these are choosing your participant demographic carefully, preparing for a no-show of participants, meticulous storage of listener data, deciding the mode of the listening test (lab or web-based), avoiding glitches in the listening test software (especially the ones that do not save data properly), maintaining anonymity of participants and ethical requirements and so on. Experiment design i.e the number and type of conditions and stimuli to be used, the manner of presentation (randomised, avoiding memory effect and priming), usage of right software tool and audio setup, length of the experiment and experiment breaks and other such factors are important in a successful listening experiment. To put it simply, conducting a listening test is like a grand orchestra where each part (small or big) plays an important role in the final result. An unexpected course in this project was the use of a biomedical tool such as EEG to measure LE. This decision came in midway through the thesis and as an engineer, I was completely unprepared to conduct an EEG experiment. However, the experience and insights I have gained from this method were mind blowing and the biggest learning points. The part objective part subjective nature of this method leaves a lot of room for continually thinking about the outcomes. Nonetheless, the fact that this data came from an actual human being makes it very interesting and valuable. Fitting a participant with an EEG cap is a time consuming and meticulous process. After testing around 50 participants, I feel I am prepared to conduct another EEG experiment. EEG signals interpretation and analysis is a vast world with a lot of possibilities. Most of the EEG research I have come across use a standard ensemble of tools to interpret and analyse EEG data. With the aid of my supervisors who are adept at signal processing, I was able to look at the EEG data with more depth. For example, I tried different window lengths and windowing techniques while extracting spectral characteristics of EEG. I found a way to analyse the EEG data corresponding to the entire length of the signal as opposed to just a fixed length of the signal, which is the usual practice. I tried to understand the inner workings of the independent component analysis process which is a sub-process of EEG data analysis. There was so much more that I could explore in the EEG experiments, but I was constrained by time. In sum, I appreciate the chance to study EEG signals in the context of disordered speech perception and gain several insights from it. I hope this tool is used in the future by researchers in this field. 93 8.3 Research Outcomes Chapters 4 to 7 described all the experiments performed as part of this thesis. Each experiment revealed some interesting insights that were useful in shaping the next study and hopefully also to the readers of this thesis for their future research ideas. Here I will summarise the outcomes of all the experiments performed in this thesis. In Chapter 4, we performed two experiments that collect intelligibility and self-reported LE metrics for OS and HS speakers. These experiments revealed the gaps in intelligibility and LE between HS and OS. OS was more effortful to process and less intelligible compared to HS. Even when the intelligibility of OS was high, there was significantly more LE associated with OS compared to HS. One interesting effect was that of listeners’ familiarity to OS. OS was not more intelligible to familiar listeners, but it was reportedly less effortful to process for familiar listeners compared to non-familiar listeners. Machine intelligibility was also poorer for OS compared to HS. In Chapter 5 we explored EEG as a way to measure LE. The findings suggested that alpha power, a neural measure known to measure LE, was higher for OS compared to HS. However, this measure was not correlated with subjective LE. The reasons for this non-correlation is an area that can be explored further. An interesting observation was that alpha power was affected by individual cognitive capacities of the participants - the participants who had poorer working memory had higher alpha power. In Chapter 6, we implemented a novel DNN-based voice conversion system which mapped OS to duration-matched SS. This system outperformed our previous systems in STOI scores and ASR scores as well as SI scores and LE obtained with a listening test. When compared with unprocessed OS, there was an improvement in ASR scores, STOI and subjective preference scores, but not in SI scores and LE. Therefore, although there is an improvement in objective measures, further improvements are needed to make the transformed OS more preferable and understandable to human listeners. Some other quick and light weight exploratory strategies provided some benefit for a low intelligibility OS speaker, but not for high intelligibility OS. As there are infinite possibilities for OS enrichment, several experiments remained either not properly explored, or not evaluated completely. Therefore, there is definitely a lot of space in this area to improve results further. In Chapter 7, three chosen enrichments were evaluated using a wide range of evaluation metrics. This included subjective metrics such as LE ratings, behavioral measurements such as SI scores and reaction times, objective measurements such as ASR scores and STOI scores and alpha power. The evaluations reveal that the proposed novel enrichment method was successful 94 in improving machine intelligibility metrics compared to previous methods. However, there was not the same level of success in improving subjective outcomes. Like the experiment in Chapter 5, a higher alpha power was observed for OS compared to HS, but nothing conclusive could be said about the enrichment methods. Overall, there was a steady increase in alpha power as the experiment progressed indicating the participants’ fatigue or reduced attention. 8.4 Future Directions The main limitation of the work described in this thesis is that the enriched outputs did not outperform unprocessed OS in subjective evaluations. Although there was some success in the preference tests described in Chapter 6, the HSR and subjective LE scores suggest a win for OS compared to the enrichments. Therefore, investigating the reasons behind this and developing enrichments that would appeal to human listeners as well would be a possible and necessary future direction. It would be interesting to explore listeners’ EEG signals further in the context of disordered speech. Some questions that can be asked are: What do the other bands of the EEG signal (other than alpha band) tell about disordered speech? Can listeners’ EEG signals reveal insights about other unexplored aspects of disordered speech such as prosody, rhythm and emotion recognition? A machine learning based analysis of EEG signals instead of the traditional ERP and frequency band analysis would also be an interesting area to explore. In the future, it would be interesting to install this enrichment system as a face-to-face communication aid in a stand-alone device or a smartphone where the device will take the unprocessed OS input and play out the enriched version in real time or with negligible delay. Another possible practical application of the proposed system could be in the form of a software plugin coupled with the microphone of the devices used by an OS speaker. This would convert any microphone input (Unprocessed OS) to an enriched version of the speech in real time or with minimum delay. Any app which requires a microphone input can use this modified speech instead. In this way, the OS speaker would be able to use the benefits of the enriched speech for telephonic conversations, zoom calls, voice commands to digital assistants and other voice based apps. 8.5 Contributions This thesis contributed to the research and development of systems that will aid laryngetomees speaking in OS and possibly speakers of other pathologies too. Some important contributions 95 are listed below. •Database and manual labelling: While automatic phonetic labelling software exist to label large speech databases, they did not work well for OS. Therefore a set of manual labelling was done for one OS speaker. When the manual labels were used to evaluate the accuracy of a customised automatic labelling process. These improved automatic labels were useful in developing an OS-friendly ASR system as well as novel methods for enrichment. •New enrichment techniques: A voice conversion based approach was used to enrich OS with a novel idea of using SS with matching durations as target. This eliminated the need for the alignment of the source OS signal to the target, which can be the cause of erroneous results. •Exploration of non VC techniques for enrichment: Apart from VC, some non-VC approaches were explored such as removing unwanted artefacts and improving the spectral characteristics of the signal. •Wavenet vocoder for Spanish: As part of one of the enrichment processes, a new high quality vocoder (Wavenet) was developed for Spanish. This vocoder can be used in future studies to generate speech from acoustic parameters of speech. •Parallel SS data for each speaker: Also as part of the enrichment process, a set of parallel SS was created using the phone labels. This synthetic speech dataset which matches OS in duration can be useful for future OS enrichment studies. •Behavioural data for OS, HS and enrichments for one speaker: An extensive set of evaluations were performed for one OS speaker and a control HS speaker. This included objective and subjective intelligibility measurements, LE measurements and preference tests. •EEG data from 12 participants when listening to an OS speaker and an HS speaker. •EEG data from 32 participants when listening to an OS speaker, an HS speaker and three OS enrichments. •ASR, STOI and preference test data for OS and enrichments of 4 speakers. •Subjective and objective LE (physiological) as an additional evaluation tool. •Code for listening tests: An interface to conduct listening tests with sentence transcription and subjective measures was created. A separate code for extraction of results and calculate intelligibility (transcription errors) is available too. 96 8.6 Publications 8.6.1 Peer-reviewed Journal Papers •Serrano, L., Raman, S., Hernaez, I., Navas, E., Sanchez, J., Saratxaga, I. (2020). A Spanish Multispeaker Database of Esophageal Speech. Computer Speech and Language, Volume 66, March 2021, 101168. DOI: 10.1016/j.csl.2020.101168 •Raman, S.; Serrano, L.; Winneke, A.; Navas, E.; Hernaez, I. Intelligibility and Listening Effort of Spanish OS. Appl. Sci. 2019, 9, 3233. DOI: 10.3390/app9163233 •Raman, S.; Sarasola, X.; Navas, E.; Hernaez, I. Enrichment of Oesophageal Speech: Voice Conversion with Duration–Matched SS as Target. Appl. Sci. 2021, 11, 5940. https://doi.org/10.3390/app11135940 8.6.2 Papers in Preparation •A paper titled ’Oesophageal Speech and Effortful Listening: an EEG Study’ based on the findings of Chapter 5. •A paper titled ’Evaluation of Enriched Oesophageal Speech: an EEG Study’ based on the findings of Chapter 7. 8.6.3 Peer-reviewed Conference Papers •Serrano, L., Raman, S., Tavarez, D., Navas, E., Hernaez, I. (2019) Parallel vs. NonParallel Voice Conversion for Esophageal Speech. Proc. Interspeech 2019, 4549-4553, DOI: 10.21437/Interspeech.2019-2194. •Serrano, L., Tavarez, D., Sarasola, X., Raman, S., Saratxaga, I., Navas, E., Hernaez, I. (2018) LSTM based voice conversion for laryngectomees. Proc. IberSPEECH 2018, 122-126, DOI: 10.21437/IberSPEECH.2018-26 •Raman, S., Hernaez, I., Navas, E., Serrano, L. (2018) Listening to Laryngectomees: A study of Intelligibility and Self-reported Listening Effort of Spanish Oesophageal Speech. Proc. IberSPEECH 2018, 107-111, DOI: 10.21437/IberSPEECH.2018-23 8.6.4 Poster Presentations •Raman, S., Hernaez, I., Navas, E., Serrano, L. A Multifaceted Enrichment of Oesophageal Speech. ICA 2019 Conference Proceedings, pp. 5739-5741. DOI: 10.18154/RWTH97 [32] Daniel Erro, Inma Hernaez, Agustin Alonso, D Garc´ıa-Lorenzo, Eva Navas, Jianpei Ye, H Arzelus, Igor Jauk, Nguyen Quy Hy, Carmen Magari˜nos, et al. Personalized synthetic voices for speaking impaired: Website and app. 16th Annual Conference of the International Speech Communication Association, 2015. [33] Daniel Erro, Inma Hern´aez, Eva Navas, Agustın Alonso, Haritz Arzelus, Igor Jauk, Nguyen Quy Hy, Carmen Magarinos, Rub´en P´erez-Ram´on, M Sul´ır, et al. Zuretts: Online platform for obtaining personalized synthetic voices. Proceedings of eNTERFACE, pages 1178–1193, 2014. [34] Daniel Erro, I˜naki Sainz, Iker Luengo, Igor Odriozola, Jon S´anchez, Ibon Saratxaga, Eva Navas, and Inma Hern´aez. Hmm-based speech synthesis in basque language using hts. Proceedings of FALA, pages 67–70, 2010. [35] Tiago H Falk, Chenxi Zheng, and Wai-Yip Chan. A non-intrusive quality and intelligibility measure of reverberant and dereverberated speech. IEEE Transactions on Audio, Speech, and Language Processing, 18(7):1766–1774, 2010. [36] R Holly Fitch, Steve Miller, and Paula Tallal. Neurobiology of speech perception. Annual Review of Neuroscience, 20(1):331–353, 1997. [37] Lionel Fontan, Isabelle Ferran´e, J´erˆome Farinas, Julien Pinquier, Julien Tardieu, Cynthia Magnen, Pascal Gaillard, Xavier Aumont, and Christian F¨ullgrabe. Automatic speech recognition predicts speech intelligibility and comprehension for listeners with simulated age-related hearing loss. Journal of Speech, Language, and Hearing Research, 60(9):2394– 2405, 2017. [38] B Garcia, Ibon Ruiz, and Amaia M´endez. Oesophageal speech enhancement using poles stabilization and kalman filtering. In 2008 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), pages 1597–1600. IEEE, 2008. [39] Luis Serrano Garc´ıa, Sneha Raman, Inma Hern´aez Rioja, Eva Navas Cord´on, Jon Sanchez, and Ibon Saratxaga. A Spanish multispeaker database of esophageal speech. Computer Speech & Language, page 101168, 2020. [40] N G´omez-Merino, F Gheller, G Spicciarelli, and P Trevisi. Pupillometry as a measure for listening effort in children: a review. Hearing, Balance and Communication, 18(3):152– 158, 2020. 104 [41] Amy Jane Hall, Axel Winneke, and Jan Rennies-Hochmuth. EEG alpha power as a measure of listening effort reduction in adverse conditions. Universit¨atsbibliothek der RWTH Aachen, 2019. [42] R Hannemann, Jonas Obleser, and Carsten Eulitz. Top-down knowledge supports the retrieval of lexical information from degraded speech. Brain Research, 1153:134–143, 2007. [43] Anne Hauswald, Anne Keitel, Ya-ping Chen, Sebastian R¨osch, and Nathan Weisz. Degradation levels of continuous speech affect neural speech tracking and alpha power differently. European Journal of Neuroscience, 2020. [44] Mark S Hawley, Phil Green, Pam Enderby, Stuart Cunningham, and Roger K Moore. Speech technology for e-inclusion of people with physical disabilities and disordered speech. In Ninth European Conference on Speech Communication and Technology, 2005. [45] Elina Helander, Jan Schwarz, Jani Nurminen, Hanna Silen, and Moncef Gabbouj. On the impact of alignment on voice conversion performance. In Ninth Annual Conference of the International Speech Communication Association, 2008. [46] Candace Bourland Hicks and Anne Marie Tharpe. Listening effort and fatigue in schoolage children with and without hearing loss. Journal of Speech, Language, and Hearing Research, 2002. [47] Sven Hilbert, Tristan T Nakagawa, Patricia Puci, Alexandra Zech, and Markus B¨uhner. The digit span backwards task. European Journal of Psychological Assessment, 2014. [48] Jens Hjortkjær, Jonatan M¨archer-Rørsted, Søren A Fuglsang, and Torsten Dau. Cortical oscillations and entrainment in speech processing during working memory load. European Journal of Neuroscience, 51(5):1279–1289, 2020. [49] Norman D Hogikyan and Girish Sethuraman. Validation of an instrument to measure voice-related quality of life (v-rqol). Journal of Voice, 13(4):557–569, 1999. [50] Sandra Cavanaugh Holley, Jay Lerman, and Kenneth Randolph. A comparison of the intelligibility of esophageal, electrolaryngeal, and normal speech in quiet and in noise. Journal of Communication Disorders, 16(2):143–155, 1983. [51] Valtteri Hongisto. A model predicting the effect of speech of varying intelligibility on work performance. Indoor Air, 15(6):458–468, 2005. 105 [52] Dee J Hubbard and Deanie Kushner. A comparison of speech intelligibility between esophageal and normal speakers via three modes of presentation. Journal of Speech, Language, and Hearing Research, 23(4):909–916, 1980. [53] Adam Jacks, Katarina L Haley, Gary Bishop, and Tyson G Harmon. Automated speech recognition in adult stroke survivors: Comparing human and computer transcriptions. Folia Phoniatrica et Logopaedica, 71(5-6):286–296, 2019. [54] Barbara H Jacobson, Alex Johnson, Cynthia Grywalski, Alice Silbergleit, Gary Jacobson, Michael S Benninger, and Craig W Newman. The voice handicap index (vhi) development and validation. American Journal of Speech-Language Pathology, 6(3):66–70, 1997. [55] Parvaneh Janbakhshi, Ina Kodrasi, and Herv´e Bourlard. Pathological speech intelligibility assessment based on the short-time objective intelligibility measure. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6405–6409. IEEE, 2019. [56] JASP Team. JASP (Version 0.8.6)[Computer software]. https://jasp-stats.org/, access date: 20th February 2018, 2018. [57] Herbert H Jasper. The international “10–20” system of the international federation. Electroencephalography and Clinical Neurophysiology, 10:371–375, 1958. [58] J. Jensen and C. H. Taal. An algorithm for predicting the intelligibility of speech masked by modulated noise maskers. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 24(11):2009–2022, 2016. [59] Joshua Y Kim, Chunfeng Liu, Rafael A Calvo, Kathryn McCabe, Silas CR Taylor, Bj¨orn W Schuller, and Kaihang Wu. A comparison of online automatic speech recognition systems and the nonverbal responses to unintelligible speech. arXiv preprint arXiv:1904.12403, 2019. [60] K. Kobayashi and T. Toda. sprocket: Open-source voice conversion software. 2018. [61] Minako Koike, Noriko Kobayashi, Hajime Hirose, and Yuki Hara. Speech rehabilitation after total laryngectomy. Acta Oto-Laryngologica, 122(4):107–112, 2002. [62] Birger Kollmeier and Matthias Wesselkamp. Development and evaluation of a german sentence test for objective and subjective speech intelligibility assessment. The Journal of the Acoustical Society of America, 102(4):2412–2421, 1997. 106 [63] John Kominek, Tanja Schultz, and Alan W Black. Synthesizer voice quality of new languages calibrated with mean mel cepstral distortion. In SLTU, pages 63–68, 2008. [64] Christian Kothe, David Medine, and Matthew Grivich. Lab streaming layer (2014). URL: https://github. com/sccn/labstreaminglayer (visited on 02/01/2019), 2018. [65] Melanie Krueger, Michael Schulte, Thomas Brand, and Inga Holube. Development of an adaptive scaling method for subjective listening effort. The Journal of the Acoustical Society of America, 141(6):4680–4693, 2017. [66] Karl D Kryter. Methods for the calculation and use of the articulation index. The Journal of the Acoustical Society of America, 34(11):1689–1697, 1962. [67] Evelyne Lagrou, Robert J Hartsuiker, and Wouter Duyck. The influence of sentence context and accented speech on lexical access in second-language auditory word recognition. Bilingualism: Language and Cognition, 16(3):508–517, 2013. [68] Sophie Landa, Lindsay Pennington, Nick Miller, Sheila Robson, Vicki Thompson, and Nick Steen. Association between objective measurement of the speech intelligibility of young people with dysarthria and listener ratings of ease of understanding. International Journal of Speech-Language Pathology, 16(4):408–416, 2014. [69] Siddique Latif, Junaid Qadir, Adnan Qayyum, Muhammad Usama, and Shahzad Younis. Speech technology for healthcare: Opportunities, challenges, and state of the art. IEEE Reviews in Biomedical Engineering, 2020. [70] Ulrike Lemke and Jana Besser. Cognitive load and listening effort: Concepts and agerelated considerations. Ear and Hearing, 37:77S–84S, 2016. [71] Vladimir I Levenshtein. Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 10, No. 8:707–710, 1966. [72] Zhen-Hua Ling, Shi-Yin Kang, Heiga Zen, Andrew Senior, Mike Schuster, Xiao-Jun Qian, Helen M Meng, and Li Deng. Deep learning for acoustic modeling in parametric speech generation: A systematic review of existing techniques and future trends. IEEE Signal Processing Magazine, 32(3):35–52, 2015. [73] Richard P Lippmann. Speech recognition by machines and humans. Speech Communication, 22(1):1–15, 1997. 107 [74] Hanjun Liu, Mingxi Wan, Supin Wang, Xiaodong Wang, and Chunmei Lu. Acoustic characteristics of mandarin esophageal speech. The Journal of the Acoustical Society of America, 118(2):1016–1025, 2005. [75] Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda, Daisuke Saito, Fernando Villavicencio, Tomi Kinnunen, and Zhenhua Ling. The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods. arXiv preprint arXiv:1804.04262, 2018. [76] Andreas Maier, Tino Haderlein, Ulrich Eysholdt, Frank Rosanowski, Anton Batliner, Maria Schuster, and Elmar N¨oth. Peaks–a system for the automatic evaluation of voice and speech disorders. Speech Communication, 51(5):425–437, 2009. [77] Alfredo Mantilla, H´ector P´erez-Meana, Daniel Mata, Carlos Angeles, Jorge Alvarado, and Laura Cabrera. Recognition of vowel segments in spanish esophageal speech using hidden markov models. 15th International Conference on Computing, pages 115–120, 2006. [78] Kenji Matsui, Noriyo Hara, Noriko Kobayashi, and Hajime Hirose. Enhancement of esophageal speech using formant synthesis. Acoustical Science and Technology, 23(2):69– 76, 2002. [79] Megan J McAuliffe, Phillipa J Wilding, Natalie A Rickard, and Greg A O’Beirne. Effect of speaker age on speech recognition and perceived listening effort in older adults with hearing loss. Journal of Speech, Language, and Hearing Research, 55(3):838–47, 2012. [80] Ronan McGarrigle, Kevin J Munro, Piers Dawes, Andrew J Stewart, David R Moore, Johanna G Barry, and Sygal Amitay. Listening effort and fatigue: What exactly are we measuring? a british society of audiology cognition in hearing special interest group ‘white paper’. International Journal of Audiology, 53(7):433–440, 2014. [81] Sharynne McLeod and Sadanand Singh. Speech sounds: A pictorial guide to typical and atypical speech. Plural Publishing, 2009. [82] Catherine M McMahon, Isabelle Boisvert, Peter de Lissa, Louise Granger, Ronny Ibrahim, Chi Yhun Lo, Kelly Miles, and Petra L Graham. Monitoring alpha oscillations and pupil dilation across a performance-intensity function. Frontiers in Psychology, 7:745, 2016. [83] Geoffrey S Meltzner, James T Heaton, Yunbin Deng, Gianluca De Luca, Serge H Roy, and Joshua C Kline. Silent speech recognition as an alternative communication device for 108 persons with laryngectomy. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 25(12):2386–2398, 2017. [84] John E Meyers, Kurt Volkert, and Anh Diep. Sentence repetition test: Updated norms and clinical utility. Applied Neuropsychology, 7(3):154–159, 2000. [85] Microsoft. Microsoft azure cognitive services speech-to-text. https: //docs.microsoft.com/en-us/azure/cognitive-services/speech-service/ get-started-speech-to-text, Accessed October 2020. [86] Catherine Middag, Tobias Bocklet, Jean-Pierre Martens, and Elmar N¨oth. Combining phonological and acoustic asr-free features for pathological speech intelligibility assessment. 12th Annual Conference of the International Speech Communication Association, 2011. [87] Catherine Middag, Jean-Pierre Martens, Gwen Van Nuffelen, and Marc De Bodt. Dia: A tool for objective intelligibility assessment of pathological speech. 6th International Workshop on Models and Analysis of Vocal Emissions for Biomedical Applications, pages 165–167, 2009. [88] Jose L Miralles and Teresa Cervera. Voice intelligibility in patients who have undergone laryngectomies. Journal of Speech, Language, and Hearing Research, 38(3):564–571, 1995. [89] Seyed Hamidreza Mohammadi and Alexander Kain. An overview of voice conversion systems. Speech Communication, 88:65–82, 2017. [90] E Ann Mohide, Stuart D Archibald, Michelle Tew, J Edward Young, and Trish Haines. Postlaryngectomy quality-of-life dimensions identified by patients and health care professionals. The American Journal of Surgery, 164(6):619–622, 1992. [91] Masanori Morise, Fumiya Yokomori, and Kenji Ozawa. World: a vocoder-based highquality speech synthesis system for real-time applications. IEICE TRANSACTIONS on Information and Systems, 99(7):1877–1884, 2016. [92] Tova Most, Yishai Tobin, and Ravit Cohen Mimran. Acoustic and perceptual characteristics of esophageal and tracheoesophageal speech production. Journal of Communication Disorders, 33(2):165–181, 2000. [93] Murray J Munro. The effects of noise on the intelligibility of foreign-accented speech. Studies in Second Language Acquisition, 20(2):139–154, 1998. 109 [94] Kathleen F Nagle and Tanya L Eadie. Perceived listener effort as an outcome measure for disordered speech. Journal of Communication Disorders, 73:34–49, 2018. [95] Kathy F Nagle and Tanya L Eadie. Listener effort for highly intelligible tracheoesophageal speech. Journal of Communication Disorders, 45(3):235–245, 2012. [96] Aljoscha C Neubauer, Andreas Fink, and Roland H Grabner. Sensitivity of alpha band erd to individual differences in cognition. Progress in Brain Research, 159:167–178, 2006. [97] Paul L. Nunez and Ramesh Srinivasan. Electroencephalogram. Scholarpedia, 2(2):1348, 2007. revision #91219. [98] Jonas Obleser, Malte W¨ostmann, Nele Hellbernd, Anna Wilsch, and Burkhard Maess. Adverse listening conditions and memory load drive a common alpha oscillatory network. Journal of Neuroscience, 32(36):12376–12383, 2012. [99] Barbara Ohlenforst, Dorothea Wendt, Sophia E Kramer, Graham Naylor, Adriana A Zekveld, and Thomas Lunner. Impact of snr, masker type and noise reduction processing on sentence recognition performance and listening effort as indicated by the pupil dilation response. Hearing Research, 365:90–99, 2018. [100] Barbara Ohlenforst, Adriana A Zekveld, Thomas Lunner, Dorothea Wendt, Graham Naylor, Yang Wang, Niek J Versfeld, and Sophia E Kramer. Impact of stimulus-related factors and hearing impairment on listening effort as indicated by pupil dilation. Hearing Research, 351:68–79, 2017. [101] Ibon Oleagordia-Ruiz and Begonya Garcia-Zapirain. Harmonic to noise ratio improvement in oesophageal speech. Technology and Health Care, 23(3):359–368, 2015. [102] Imen Ben Othmane, Joseph Di Martino, and Ka¨ıs Ouni. Enhancement of esophageal speech obtained by a voice conversion technique using time dilated fourier cepstra. International Journal of Speech Technology, 22(1):99–110, 2019. [103] Jonathan W Peirce. Psychopy—psychophysics software in python. Journal of Neuroscience Methods, 162(1-2):8–13, 2007. [104] M Kathleen Pichora-Fuller, Sophia E Kramer, Mark A Eckert, Brent Edwards, Benjamin WY Hornsby, Larry E Humes, Ulrike Lemke, Thomas Lunner, Mohan Matthen, Carol L Mackersie, et al. Hearing impairment and cognitive energy: The framework for understanding effortful listening (fuel). Ear and Hearing, 37:5S–27S, 2016. 110 [105] Eduard Polityko. Word error rate. https://www.mathworks.com/examples/matlab/community/19873word-error-rate, access date: 20th February 2018, 2018. [106] Gerasimos Potamianos and Chalapathy Neti. Automatic speechreading of impaired speech. In AVSP 2001-International Conference on Auditory-Visual Speech Processing, 2001. [107] Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al. The kaldi speech recognition toolkit. IEEE Workshop on Automatic Speech Recognition and Understanding, 2011. [108] DA Preece. Latin squares, latin cubes, latin rectangles. Wiley StatsRef: Statistics Reference Online, 2014. [109] Sneha Raman, Inma Hernaez, Eva Navas, and Luis Serrano. Listening to laryngectomees: A study of intelligibility and self-reported listening effort of spanish oesophageal speech. Proceedings of IberSPEECH 2018, pages 107–111, 2018. [110] Sneha Raman, Inma Hernaez, Eva Navas, and Luis Serrano. A multifaceted enrichment of oesophageal speech. Universit¨atsbibliothek der RWTH Aachen, 2019. [111] Sneha Raman, Xabier Sarasola, Eva Navas, and Inma Hernaez. Enrichment of oesophageal speech: Voice conversion with duration–matched synthetic speech as target. Applied Sciences, 11(13):5940, 2021. [112] Sneha Raman, Luis Serrano, Axel Winneke, Eva Navas, and Inma Hernaez. Intelligibility and listening effort of spanish oesophageal speech. Applied Sciences, 9(16):3233, 2019. [113] Shakti P Rath, Daniel Povey, Karel Vesel`y, and Jan Cernock`y. Improved feature processing for deep neural networks. 14th Annual Conference of the International Speech Communication Association, pages 109–113, 2013. [114] Gerard B Remijn, Mitsuru Kikuchi, Yuko Yoshimura, Kiyomi Shitamichi, Sanae Ueno, Tsunehisa Tsubokawa, Haruyuki Kojima, Haruhiro Higashida, and Yoshio Minabe. A near-infrared spectroscopy study on cortical hemodynamic responses to normal and whispered speech in 3-to 7-year-old children. Journal of Speech, Language, and Hearing Research, 60(2):465–470, 2017. 111 [115] Jan Rennies, Henning Schepker, Inga Holube, and Birger Kollmeier. Listening effort and speech intelligibility in listening situations affected by noise and reverberation. The Journal of the Acoustical Society of America, 136(5):2642–2653, 2014. [116] Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra. Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs. In 2001 IEEE international conference on acoustics, speech, and signal processing. Proceedings (Cat. No. 01CH37221), volume 2, pages 749– 752. IEEE, 2001. [117] Jana Roßbach, Saskia R¨ottges, Christopher F Hauth, Thomas Brand, and Bernd T Meyer. Non-intrusive binaural prediction of speech intelligibility based on phoneme classification. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 396–400. IEEE, 2021. [118] Marinela Rosso, Ljiljana ˇ Siri´c, Robert Ti´cac, Radan Starˇcevi´c, Igor ˇ Segec, and Nikola Kraljik. Perceptual evaluation of alaryngeal speech. Collegium Antropologicum, 36(2):115– 118, 2012. [119] I Sainz, D Erro, E Navas, I Hern´aez, J S´anchez, I Saratxaga, I Odriozola, I Luengo, et al. Aholab speech synthesizers for albayzin2010. Proceedings of FALA, 2010:343–348, 2010. [120] I˜naki Sainz, Daniel Erro, Eva Navas, Inma Hern´aez, Jon Sanchez, Ibon Saratxaga, and Igor Odriozola. Versatile speech databases for high quality synthesis for basque. In LREC, pages 3308–3312. Citeseer, 2012. [121] I˜naki Sainz, Daniel Erro, Eva Navas, Inmaculada Hern´aez, J Sanchez, I Saratxaga, and Igor Odriozola. Versatile Speech Databases for High Quality Synthesis for Basque. 8th international Conference on Language Resources and Evaluation (LREC), pages 3308– 3312, 2012. [122] Odette Scharenborg. Reaching over the gap: A review of efforts to link human and automatic speech recognition r‘esearch. Speech Communication, 49(5):336–347, 2007. [123] Luis Serrano. T´ecnicas para la mejora de la inteligibilidad en voces patol´ogicas. PhD thesis, University of the Basque Country (UPV/EHU), 2019. [124] Luis Serrano, Sneha Raman, David Tavarez, Eva Navas, and Inma Hernaez. Parallel vs. non-parallel voice conversion for esophageal speech. Proceedings of Interspeech 2019, pages 4549–4553, 2019. 112 [125] Luis Serrano, David Tavarez, Igor Odriozola, Inma Hernaez, and Ibon Saratxaga. Aholab system for albayzin 2016 search-on-speech evaluation. Proceedings of IberSPEECH, pages 33–42, 2016. [126] Luis Serrano, David Tavarez, Xabier Sarasola, Sneha Raman, Ibon Saratxaga, Eva Navas, and Inma Hernaez. LSTM based voice conversion for laryngectomees. In Proceedings of IberSPEECH 2018, pages 122–126, 2018. [127] Luis Serrano, David Tavarez, Xabier Sarasola, Sneha Raman, Ibon Saratxaga, Eva Navas, and Inma Hernaez. Lstm based voice conversion for laryngectomees. Proceedings of IberSPEECH, pages 122–126, 2018. [128] A. Sesma and Asunci´on Moreno. Corpuscrt 1.0: Diseno de corpus orales equilibrados. Computer Program]: http://gps-tsc. upc. es/veu/personal/sesma/CorpusCrt. ph p3, 2000. [129] Hamid Reza Sharifzadeh, Ian V McLoughlin, and Farzaneh Ahmadi. Reconstruction of normal sounding speech for laryngectomy patients through a modified celp codec. IEEE Transactions on Biomedical Engineering, 57(10):2448–2458, 2010. [130] Dushyant Sharma, Yu Wang, Patrick A Naylor, and Mike Brookes. A data-driven nonintrusive measure of speech quality and intelligibility. Speech Communication, 80:84–94, 2016. [131] K˚are Sj¨olander and Jonas Beskow. Wavesurfer-an open source speech tool. In Sixth International Conference on Spoken Language Processing. Citeseer, 2000. [132] Constantin Spille, Birger Kollmeier, and Bernd T Meyer. Comparing human and automatic speech recognition in simple and complex acoustic scenes. Computer Speech & Language, 52:123–140, 2018. [133] Smiljka ˇ Stajner-katuˇsi´c, Damir Horga, Maja Muˇsura, and Dubravka Globlek. Voice and speech after laryngectomy. Clinical Linguistics & Phonetics, 20(2-3):195–203, 2006. [134] Herman JM Steeneken. The measurement of speech intelligibility. Proceedings of Institute of Acoustics, 23(8):69–76, 2001. [135] Antje Strauß, Malte W¨ostmann, and Jonas Obleser. Cortical alpha oscillations as a tool for auditory selective inhibition. Frontiers in Human Neuroscience, 8:350, 2014. [136] Daniel J Strauss and Alexander L Francis. Toward a taxonomic model of attention in effortful listening. Cognitive, Affective, & Behavioral Neuroscience, 17(4):809–825, 2017. 113 ESTOI Extended Short Term Objective Intelligibility. 11, 12, 117 fMRI functional Magentic Resonance Imaging. 18, 117 GAN Generative Adversarial Networks. 22, 117 GMM Gaussian Mixture Models. 21, 22, 71–73, 117 HMM Hidden Markov Models. 60, 117 HNR Harmonics-to-Noise Ratio. 11, 22, 117 HS Healthy (Laryngeal) Speech. x, xv–xvii, xix, 3, 4, 8–12, 14, 20–23, 26, 27, 29, 31–41, 44–54, 57, 58, 60, 61, 70, 71, 73, 75–77, 81–84, 86, 88, 89, 91, 94–96, 117 HSR Human Speech Recognition. xv, 9, 11, 14, 23, 32, 36–39, 45, 46, 95, 117 KS Kolmogorov–Smirnov. 37, 41, 117 LE Listening Effort. xv–xvii, xix, 15–20, 23, 26, 31–33, 35–55, 57, 73, 75–78, 82–85, 87–89, 91–96, 117 LSTM Long Short-Term Memory. 22, 97, 117, 141 MCC Mel Cepstral Coefficients. 60, 117 MCD Mel Cepstral Distortion. 13, 21–23, 117 MFCC Mel-Frequency Cepstral Coefficients. 35, 117 MOS Mean Opinion Score. 13, 21–23, 75, 117 NIRS Near-infrared Spectroscopy. 48, 117 OOV Out of Vocabulary. 35, 117 OS Oesophageal Speech. ix–xi, xv–xvii, xix, 3–5, 7–11, 14, 15, 17, 19–23, 25–29, 31–41, 43–55, 57–73, 75–77, 79, 81–89, 91, 92, 94–97, 117 PESQ Perceptual Evaluation of Speech Quality. 13, 117 PET Positron Emission Tomography. 18, 117 PPG Phonetic Posteriorgrams. xvii, 22, 76, 77, 82–84, 86, 87, 117 120 PSD Power Spectral Density. 50, 78, 117 PSM Perceptual Similarity Measure. 117 PWC Percentage Words Correct. xvi, 13, 14, 62, 63, 68, 81, 117 RT Response Time. 117 SD Standard Deviation. 41, 59, 117 SI Speech Intelligibility. xv–xvii, 11, 14, 35, 40, 42, 49, 52, 54, 76, 77, 81–85, 87–89, 94, 117 SNR Signal-to-Noise Ratio. 11, 117 SPL Sound Pressure Level. 41, 117 SRT Speech Reception Threshold. 13, 14, 117 SS Synthetic Speech. xvi, xvii, 28, 58–63, 65, 66, 68, 69, 75, 87, 94, 96, 97, 117 STEM Science, Technology, Engineering and Mathematics. 117 STI Speech Transmission Index. 11, 12, 117 STOI Short Term Objective Intelligibility. xi, xvi, xvii, 11, 12, 14, 59, 61, 63–66, 68, 69, 73, 75, 87–89, 94, 96, 117 TOS Tracheoesophageal Speech. 8, 14, 20, 22, 25, 61, 117 V-RQOL Voice-related Quality of Life. 20, 117 VC Voice Conversion. xvi, 13, 20–23, 57–61, 68, 69, 71–73, 87, 96, 117 VHI Voice Handicap Index. 20, 117 WER Word Error Rate. xv–xvii, xix, 12–14, 29, 33, 36–40, 43–46, 50–52, 59, 61–63, 68, 69, 71, 72, 76, 86, 117 WM Working Memory. 48, 117 121 122 Appendix A 30 sentences Used in the Experiment in Chapter 4 1. Una fiesta en Florida Park con glamur 2. De Filadelfia vino el grupo Jud´ıos por una Paz Justa 3. Ello h ac´ıa intuir un duelo en toda regla 4. Deja mucha buena obra hecha pero me rehuye el balance 5. Hoy jueves dieciocho de julio de dos mil trece 6. Ha podido reunirse con Ch´avez tras su elecci´on 7. Quiz´a ustedes no lo advirtieron por eso lo refiero ahora 8. Abdul´a Abdul´a ministro de Exteriores va m´as all´a 9. Da igual no importa de d´onde extraiga uno la emoci´on 10. S´olo el chileno Mark Gonz´alez pon´ıa una pizca de orgullo 11. Pero usted ya conoce por dentro el mundo del cine 12. Hay voces que ya hablan de indulto ser´ıa factible 13. Lo que no cree nadie aqu´ı es que cacen a Sadam con vida 14. Unos d´ıas de euforia y meses de aton´ıa 15. Blasco Ib´a˜nez hizo alguna vez la misma cosa 123 16. Ten´ıa la voz alborotada y la amistad ruidosa 17. Apliqu´e el o´ıdo a esta rayita y percib´ı un murmullo 18. Gast´o todo el agua incluyendo el agua de las lluvias 19. Si el club no hubiera cambiado se hubiera ido 20. Goliat estuvo a punto de engullir a David 21. Tal vez fue hace siglos o acaso hace tan s´olo unas d´ecadas 22. A´un no sabemos qu´e fue a hacer a Taiwan 23. El pueblo noruego rechaz´o v´ıa refer´endum la adhesi´on 24. Con este ´album llegar´a seguro al n´umero uno de ventas 25. Mi aldea estaba a la orilla de un riachuelo como ´este 26. Fui yo por consejo del se˜nor Regueiro Souza 27. Hoy no juega al golf y el traje es azul cielo y oro 28. Occidente y el islam son dos miedos que se acechan 29. N´u˜nez ya tiene a su hijo predilecto en casa 30. Qu´e diferencia hay entre el caucho y la hevea 124 Appendix B EEG Data Terminology and Procedures B.1 EEG Recording Equipment B.1.1 Cap and Electrodes An EEG cap is a stretchable head covering with holes (for electrodes) that fit the subject’s head snugly. As the head sizes of subjects vary, there are different sizes of caps available, typically: 54cm, 56 cm, 58cm and 60cm. What size cap fits a subject is decided by measuring the head circumference with a measuring tape passing through the forehead. It is important to provide some tolerance as the cap should not be too tight. For example, a person with head circumference measuring 57cm would be fitted with a 58 cm cap. In addition, it must be ensured that the cap is placed in the right position. For this, the centre-most point of the cap is matched to the centre of the head, which is midway between the centre forehead (nasion) and the natural bump at the centre back (inion), as well as midway between the 2 ears. The caps may have 32, 64 or 128 holes depending on the type and requirement. The electrodes are arranged on an EEG cap as per the International 10-20 system for measuring EEG. It is known as the 10-20 system because the distance between any two adjacent electrodes is either 10% or 20% of the front-back or right left distance. The position and the naming convention of the electrodes follows a [brian area][direction code] system. The brain areas are Parietal (P), Frontal(F), Central(C), Temporal(T) and Occipetal(O). Intermediate or overlapping areas are named as Centro-parietal (CP), Frontcentral (FC), Tempero-parietal (TP), Front-Temporal(FT), Parietal-occipetal (PO) and so on. 125 The directions codes are z for the front-back midline area, odd numbers (1,3,5) for the left of the midline and even numbers (2,4,6) for the right of the midline. Numbers are assigned incrementally from front to back. Based on these conventions, the electrodes are named as F1 (frontal left), P6 (Parietal Right), CPz (centro-parietal midline) etc. Apart from the above-mentioned electrodes, there are reference and ground electrodes. The EEG signal is a potential difference between the scalp potential and a reference potential. This reference potential is obtained by placing an electrode away from the scalp, on a relatively neutral area such as behind the ears (mastoids) or on the nose. As these areas are not impacted by brain activity, a difference of the scalp electrode potential to this reference potential results in the potential difference associated with the brain activity in the said scalp electrode. A ground electrode is usually placed on the cap and its function is mainly to remove power line noise. In addition, some electrodes may be placed on the temples and forehead to record ocular activity. These electrodes can then later be used to remove ocular activity artefacts. B.1.2 Gel An electrolyte gel is applied (with needle-less syringes) where the electrodes come in contact with the scalp. The function of this gel is to increase the electrical connectivity of the scalp to the electrode. The gel from one electrode site must not seep through to another electrode site as it would create a short circuit which is undesirable. Sometimes the gel may contain a slightly abrasive or exfoliating component that helps in scraping off the dead skin cells in the top layer of the scalp. This also improves the connectivity. B.1.3 Amplifier The EEG amplifier amplifies the EEG signal and also converts it from digital to analogue. This is so that the recording software on the computer can receive and store the EEG signals digitally and amplified. The sampling rate of the signals are typically 250 to 2000 Hz. B.1.4 Recording software While placing electrodes and recording EEG signals, it is helpful to see the status in real time. This helps in trouble shooting and ensuring good quality of data. In a recording software (such as BrainVision), we can visualise the electrodes and their impedance while mounting the cap and electrodes (See Figure B.1). The aim in this stage is for all the electrodes to have a low impedance (preferably <5kΩ). Additionally, in the recording software interface, it is possible 126 Figure B.1: EEG recording software showing impedance values of electrodes. to see the EEG signals in real time and that can help the experimenter identify noisy channels, restless or tired participants etc. B.2 Synchronisation While conducting an EEG experiment, the experiment software sends out a trigger to the EEG recording software when the audio starts playing. However it is possible that there are a few milliseconds of delay between the experiment software sending the trigger and the EEG recording software receiving those triggers. It is necessary to ensure that the EEG markers (or triggers) and the audio stimuli start points are synchronised. This synchronisation may be achieved using hardware or manually. In the hardware method, a clock device synchronises the audio input and the EEG markers. Alternatively, it may be done manually by identifying the delay between the playing of the audio signal and the start of EEG recording and then adding the appropriate amount of delay later to synchronise the EEG and audio channels. B.3 EEG Recording Process Once the ground, reference and all the data electrodes are placed correctly with the right impedance, ocular and muscular activity are checked to unsure that the EEG data are being recorded correctly. This can be observed by spikes in the data when the participant blinks or dense activity when the participant clenches their jaw or bites. 127 Figure B.2: Raw EEG data In case there are noisy electrodes, they must be fixed by either putting more gel or correcting other connectivity issues. The signals must be constantly monitored to avoid noisy data. To avoid noisy data from the participant. they are instructed to be as still as possible and to look straight on to the screen. When the participant is ready and all the electrodes look fine, we can start recording the EEG data. Then the experiment software which displays the controls for the behavioural data and plays the audio files is run. Once the experiment is over, the EEG data recording is stopped and the EEG file is appropriately saved. There are several different formats of EEG data such as .eeg, .xdf etc. B.4 Raw EEG Data Raw EEG data is difficult to process as it is and needs to undergo several steps of preprocessing to extract meaningful data from it. To view and process these raw EEG data files, an EEG processing software is used. One such software is the EEGLab software. When the EEG file is loaded in EEGLab, we can first observe the raw EEG data. This is a time series recorded by each one of the electrodes. Figure B.2 shows the raw EEG data from a participant. The vertical axis is the electrode name and the horizontal axis is time. The vertical colourful lines are the triggers or the event markings where the participant started hearing a stimulus or performed some action. It can be observed that the data is quite noisy. Each data series is assigned a channel location which is the two-dimensional or threedimensional location of that electrode on the head. The channel locations are obtained from the manual of the EEG cap set. Because of this step, it is possible to view the electrode data 128 Figure B.3: Raw EEG data in a 2-D or 3-D topographical plot. For example, in Figure B.3, we can see the topographical representation of channel 24, which is the O2 (occipetal right) electrode. The red dot on the topographical plot shows the position of the O2 electrode on the head. Some possible artefacts and noise we can notice in the raw EEG data are noisy channels (such as channel C3 in Figure B.2). The occasional dips in the Fp1 and Fp2 channels correspond to the participant blinking. This will be clearer in the data post the filtering process. Rereferencing is the process of getting the potential difference between the electrodes and the reference electrode. Typically, the reference electrodes are the mastoid electrodes (the electrodes behind the ears). It may be one mastoid electrode or an average potential of the two mastoid electrodes. Alternatively, it is also possible to have an average reference. In this method, an average of potentials from all the electrodes is calculated and then each this average potential is subtracted from each of the electrodes. B.5 Cleaning Up Raw EEG Data B.5.1 Filtering Raw EEG data has a lot of high frequency noise and power line noise which is removed by band pass filtering. Typically the filtering is performed by band passing between 1 and 45 Hz as EEG signals of interest lie between these frequencies. Sometimes an additional 50Hz or 60Hz notch filter is added to remove power line noise. The filtered EEG signal is shown in Figure B.4. Here the ocular activity is visible in the form of abrupt dips around the 12th, 13th and 14th seconds. 129