Advances in Interactive Speech Transcription
Abstract
[ES] Novedoso sistema interactivo para la transcripción del habla que compensa el esfuerzo del usuario y el error máximo tolerado en las transcripciones resultantes.
Full text
Advances in Interactive Speech transcription ∼Valencia, June, 2012 ∼ . Master Thesis at the Pattern Recognition and Language Technologies Group by Isaías Sánchez Cortina Advisors: Dr. Alfons Juan i Ciscar Dr. J. Alberto Sanchis Navarro .
ACKNOWLEDGEMENTS Work supported by the Ministerio de Ciencia e Innovación (MICINN) under iTrans2 (TIN2009-14511) project and the the FPI Scholarship BES-2010-033005; And by the the Generalitat Valenciana under the grant GV/2010/067. 1
CONTENTS Chapter 1. Introduction 5 Chapter 2. Interactive approach to speech transcription 7 1. Foundations of the Automatic Speech Recognition 7 2. Unsupervised Learning 8 3. The Interactive Speech Recognition 8 Chapter 3. Balancing the error and supervision effort 11 1. Interactive Speech Transcription Procedure 11 2. Estimation of the Transcription Error 14 3. Confidence Measures ranking and classifier 15 4. Main contributions of our prodecure 15 Chapter 4. Performance of the interactive speech transcription system 17 1. Evaluation of the performance of an interactive transcription system 17 2. Experimental set-up 19 3. Results 21 Chapter 5. A prototype interface to interactive speech transcription 45 1. System Overview 45 2. User interface of the prototype 46 3. Demonstration 47 4. Usability 48 Chapter 6. Conclusions 49 1. On the method of balancing error and the user effort 49 2. Future Work 50 3
CHAPTER 1 INTRODUCTION Speech transcription is a crucial task in a broad range of important applications. Speech transcriptions produced by human transcribers can provide high quality results. However, the overall process is very slow and usually expensive. One way to deal with this important drawback is to produce automatically speech transcriptions based on automatic speech recognition (ASR) technology [8, 16]. However, this solution presents two main difficulties. First, the building of an ASR system implies usually an important human effort since statistical models have to be obtained. Acoustic and language models are not easily available for specific tasks and, thus, they have to be learned from manual annotated data. The second difficulty is that automatic transcriptions are still far from producing the desired quality out of some specific scenarios. Therefore, an important human effort to supervise the speech transcriptions is mandatory. Sometime, for very faulty transcriptions, the effort can even be higher than manually transcribing the whole. To deal with the first difficulty, active and unsupervised learning techniques have been applied to rapidly prototype ASR systems reducing significantly user effort [4, 19]. Upon these approaches, a little manually annotated data is used to build rapidly an ASR system. Then, this initial ASR system is used to automatically transcribe a large amount of new speech data. These new annotated data is used to improve the underlying ASR models based on active and unsupervised learning techniques. To overcome the second difficulty, an interactive paradigm has been applied to reduce significantly the supervision effort [6, 9] of the speech transcriptions. Following this paradigm, the system produces automatically speech transcriptions and the user is assisted by the system to amend output errors as efficiently as possible. In this work, we present a novel speech transcription system in which active and unsupervised learning techniques are applied along within an interactive paradigm in a tightly coupled manner. The main goal is to significantly reduce the human effort in the transcription of a speech task by allowing a maximum tolerance error in the resulting transcriptions. The supervision is performed while an estimation of the transcription errors is higher than the tolerance error. Word-level confidence measures [20] computed over the recognition word-graph are used to suggest which words should be supervised. During the transcription process, supervised and high-confidence parts of the transcriptions are used to improve incrementally the underlying statistical ASR models. This will likely improve the subsequent recognitions lowering the necessary number of supervisions. 5
6 1. INTRODUCTION This approach was successfully applied to handwriting transcription [13]; Here, we adopted it and improved the method for the estimation of the error . The empirical results confirm previous works on handwriting recognition showing that this strategy is effective to reduce the supervision effort by allowing a maximum tolerance error in the speech transcriptions. Moreover, results show that a tolerance error in the transcriptions does not affect critically on the incremental learning of acoustical models. Thus, this method can be used also for producing ASR models resulting in similar performance to those generated using fully manually transcribed corpora. In summary, this is useful approach implemented with a simple yet effective method to find an optimal balance between recognition error and supervision effort for the interactive transcription of the speech . Let us summarize the layout of this work: In chapter 2, the foundations of the ASR foundation are explained briefly, as well as the interactive speech transcription (IST). Chapter 3 details our interactive approach and the method for balancing the error and user effort and the main contribution respect to related previous works in the literature. Next, in chapter 4, a general description objective magnitudes that might be useful to describe the performance of an IST system in terms of the user effort the features, and then, an evaluation resulting from applying the method for the task of transcribing the speech of two sets of the World Street Journal Speech database of 15 and 80 hours respectively. 5 describes an shows an interface developed for our IST method. Finally, conclusions on our method and future work are discussed on chapter 6 .
CHAPTER 2 INTERACTIVE APPROACH TO SPEECH TRANSCRIPTION 1. Foundations of the Automatic Speech Recognition Speech Recognition is founded on the assumption that the natural speech can be described by means of probabilistic models in a satisfactory way. Under this framework, the facto standard is that in which the transcription is obtained as the maximum a posteriori probability (MAP) of a sequence of words (~w) given the sequence of acoustic observations ( ~ X, which are derived from the audio of the speech): (1) ˆ ~w =argmax~w∈WP(~w|~ X) = argmax~w∈WP(~ X|~w)·P(~w). The right hand term has been obtained by applying the Bayes rule. This has been the standard way to proceed over the last decades since it presents several advantages: It allows the search for the most probable transcription (string decoding) to be effectively pruned thanks to term P(~w), which is called the language model (LM). LMs can incorporate both syntactic and semantic constraints of the language and the recognition task. Also, they can be trained using the standard methods for the estimation of the n-grams probabilities from text resources independent to the speech task. On the other hand, the generative model P(~ X|~w), called the acoustic model (AM), can be estimated in a easier and more precisely way than the discriminative model (~w|~ X). However, for large vocabulary speech recognition systems, it is necessary to build statistical models for sub word speech units, build up word models from these sub word speech unit models (using a lexicon to describe the composition of words), and then postulate word sequences and evaluate the acoustic model probabilities via standard concatenation methods. The commonly used sub word units are the phonemes, or tri-phonemes. And the most accepted probabilistic for the representation of these units are the hidden Markov models (HMMs) . HMMs representations are good in matching patterns of high variability, like the realization of the phonemes. The main parameters of the HMMs are its structure and the emission distribution of the states. The structure is defined by number of internal states and the allowed connections amongst them. A state is a hidden discrete variable thanks to which the model achieves some amount of histeresis depending on the already seen observations. The emission distributions are the probability of a phoneme given the the observations and current value of of the HMMs . The standard emission distributions are the Gaussian mixtures . Additionally, the transition probabilities from state to state also play a key role 7
8 2. INTERACTIVE APPROACH TO SPEECH TRANSCRIPTION in the structure. But it has been showed that its values have little impact on the recognition performance. The training procedure of the HMMs an LMs , as well as the recognition process is beyond the scope of this introduction. But is should be noted that decoding with the Viterbi algorithm using HMMs representations can be efficiently computed in linear time with the number of states and the number of acoustic vectors (proportional to the duration of the audio signal). Furthermore, the LM will reduce the number of possible paths, and even more if some state-of-the-art LM look-ahead techniques are applied along the viterbi search. In summary, HMMs have been so successful for speech recognition because of their computational efficiency and their ability to match variable signals. 2. Unsupervised Learning The training of accurate HMMs usually requires a high amount of transcribed speech. And, usually, the training data (speech) should be similar enough to the task to automatically transcribed (in terms of the speech style, the environment, the speakers, etc.). But, as commented in the introduction, the manual process is slow and expensive, usually requiring of professional transcribers. Fortunately, HMMs can be improved even when no transcription is available using the own output of an ASR. In the context of training, this way of learning (or adapting the initial model) is known as unsupervised learning. Thus, better ASR model might be obtained with no extra human supervision than the that necessary to just bootstrap the initial models. The simplest and most extended way for the unsupervised learning confidence values is as in [19]. The confidence values can be derived using features computed over the recognition word-graphs. The word-graphs are structures containing not only the most probable decoded transcription but also other less probable paths. Also, the structure include recognition scores, time alignments, etc. The confidence values of the recognized transcription (at sentence , word, or phoneme level) serve to select pieces of the recognized samples which are likely to have been properly recognized. These selected pieces of samples and their corresponding transcriptions can added to the existing training data so the ASR model can be then re-estimated. Precisely, this is the method for unsupervised learning which has been used in this work. Further refinement can be achieved as in [4]. In that work, the pieces of the recognized transcriptions selected for being new training data, will have an reduced impact in the LM estimation. This is achieved by modulating the counts of the recognized words by their confidence values. However, it is patent that the main difficulty for the unsupervised learning in getting improved models is due to the performance of the confidence measure in predicting whether words are correctly or incorrectly recognized. Tasks in which the confidence measure yields a high classification error ratio (CER), the unsupervised learning might even worsen the initial models. 3. The Interactive Speech Recognition Interactive Speech Transcription (IST), from a general point of view, is the process of obtaining the transcription of a speech in the core of an automatic system with the help of a user. Thereby combining human accuracy with ASR efficiency. The contribution of the user can be as much as typing the full manual transcription, or as little as a few clicks or keystrokes, or any other modal interaction. On the other hand the system may provide as little as just a convenient interface; automatic segmentation; predicting the text being typed (ex. [10]); or as much
4. MAIN CONTRIBUTIONS OF OUR PRODECURE 15 (5) ˆ E−=E+ˆ N− N+ where R+and R−are the number of recognized words which have been supervised and non-supervised, respectively, up to current utterance. And E+is the accumulated edition cost of the supervised parts (up to current utterance) before corrections are made. However, a better estimation of W−can be achieved by ranking the recognized words into one of the Cgroups depending on its confidence measure [13]. Let the groups from 1to C−1refer to the word with lowest confidence, second lowest, and so on respectively, in a utterance. And let Group Crefer to the, high-confidence, rest of words not in the previous groups. With this modification (4) and (5) are expressed as follows: (6) ˆ Nc−=Nc+Rc− Rc+ (7) ˆ E−= C X c Ec+ˆ Nc− Nc+ Finally, WER of the unsupervised parts (W−) can be obtained using the estimations (6) and (7) for the unknown WER of the unsupervised parts (3) : (8) ˆ W−=PC cEc+Rc− Rc+ PC cNc+1 + Rc− Rc+ 3. Confidence Measures ranking and classifier As described before, the method needs to rank the recognized words depending on a confidence measure value ; and also for building new train samples using the confidence measure as the input of a binary classifier. The classifier may be as sophisticated as desired. However, the recognition posterior probability of the words (post-max) alone is the most important feature [14]. A simple threshold classifier can then be built using the post-max. The optimal values for combination factor and the threshold are found , for instance, simply by maximizing the area under the ROC curve ([12]). Although recent publication achieve further performance ([11]). 4. Main contributions of our prodecure In this section a brief remarks highlighting the difference with similar work published by D. Hakkani-Tur , and the work N. Serrano for handwriting transcription (HWR) in [13]. •Compared to "An Active Approach to Spoken Language Processing" by D. Hakkani-Tur in [4]: •Their approach centers on manual supervision of whole utterances, thus the effort cannot be reduced so much. •It cannot be properly implemented if the speech is not sentence-level segmented . While here the segmentation level has some influence, but this is optional. •Their unsupervised learning for the LM is cleverer, but it has a slight impact since the greatest impact is related to the performance of the confidence measures .
16 3. BALANCING THE ERROR AND SUPERVISION EFFORT •Their purpose is to ask for supervision the most informative utterances. But, in the end, their cost functions behaves much like the estimation of the WER presented here, but based on the single global confidence measure of the utterance. Thus, much less reliable. •Compared to "Balancing error and supervision effort in interactive-predictive handwriting recognition" by N. Serrano in [13]: •Our procedure can be considered a port of N. Serrano work for HWR. However, they implicitly worked like if the utterances were segmented at word-level . While, here, despite being also working in a word by word basis, calculations are made over the whole utterances under study. As a very convenient consequence, the system can jump over the high confidence rest of the utterance or even the whole utterance . See "Considering a Utterance as a word or as a sentence" in chap. 4 sec. 3 •The estimation of the error ˆ W−(eq. 8) in their work is normalized using the estimation of the corresponding reference words of the unsupervised parts ( ˆ N−not depending on the confidence measure rank as used in the numerator. Previous experiments (not shown in this work) proved the current estimation of ˆ W−is 0.5% lower in average than the older. In despite of this having little impact on the overall performance for low tolerated errors, is still a convenient improvement. •Speech transcription poses different difficulties compared to HWR. The main issue was the word-level segmentation is more error prone: Thus during the interaction it is more likely the user would not understand the uttered words in the selection of the corresponding audio. This encouraged us to build a functional prototype in order to test the system.
CHAPTER 4 PERFORMANCE OF THE INTERACTIVE SPEECH TRANSCRIPTION SYSTEM In this chapter the proposed interactive system is evaluated. First, section 1 intends to be a general discussion about the possible aspects and features useful to properly evaluate an interactive system for transcription of the speech. Section 2 describes the task to be transcribed during the evaluations. And, finally, section 3 shows the obtained results. 1. Evaluation of the performance of an interactive transcription system Here, the characteristics of the measures to be used in the results section to evaluate the system are detailed. It is clear that the key aspects to be evaluated of an interactive system like the proposed one should be: the decrease in the user effort and, the ability to learn and improve the underlying models and the quality of the resulting transcriptions: 1.1. Assesment of the user effort. An straight forward way to asses the user effort is by means of the number of times the user has been asked to supervise a word ("number of supervisions" NS for short) as in [13]. A bit more optimistic would be "the number of correct words corresponding to the parts which has been supervised" ("number of reference supervised words", #rS for short). The latter is preferred in the new results to be reported by the some authors. However, here, the first option is preferred, so the term "%Supervisions" will refer to the ratio of the number of recognized words to the number of words of the reference transcription. It can be argued that , in despite of that the definition allows for percentages greater than the 100% , the #rS may yield misleading results in other tasks: #rS might been motivated with the fact that the deletions require less effort than typing insertions, but for a task in which a user had to perform way more deletions than insertions, #rS would be lower than NS if the corrections would have been "accepting" (equal) operations instead. It is clear that effort would have been almost identical in both cases. And, even worse, yet a real effort in supervision would have existed , NSR counts will be null if the user would had to perform just deletions, . Anyway, #rS and NS in the conducted experiments, Nevertheless, It would be desirable to have another features from which could be derived a more realistic estimation of the user effort: Let them be the performed number of keystrokes (NK) and the total time of played audio (LT). These features are encouraged by the that the time for manual transcriptions strongly depends on the duration of the speech and the time expended in typing the corrections. 17
18 4. PERFORMANCE OF THE INTERACTIVE SPEECH TRANSCRIPTION SYSTEM However, it is left for future work the research on the connection of these features and the time a user really needs to fulfill the semi-supervision using the presented method. Moreover, It should be noted that in the case of an automatic evaluation of the system in that the is no real user, i some model must be assumed in order to account for the NK. In this work, it was assumed that the interface would allow to accept (equal operation) and delete with just one keystroke; Insertions and substitutions would require as many as keystrokes as the number of characters of the introduced corrections. The ratio of the NK to the relative to the total number of characters in the reference, denoted as PK, is also meaningful . Additionally, this estimation can be compared to mPK : the minimum PK necessary if our system were flawless, i.e. if it would never ask to supervise correctly recognized words. 1.2. Assesment of the improvement of the ASR system due to the interaction. The WER will be used to asses the performance of the recognitions. This is expected to strongly depend on the tolerance error, since the more supervisions, the better will the models be re-trained. Nonetheless, some improvement of the ASR models could result even when no supervisions are done when the tolerance is too high. This is because of the unsupervised learning , this is, increasing the training sets with the high-confidence parts of the automatic transcriptions. Although, in general, this could lead to worse models, it has been proved that unsupervised learning improves the AM for the WSJ when using an external LM. Moreover, in order to compare the contributions to the improvement of the acoustic and language parts , experiments using a fixed external LM or AM can be conducted , as in [4]. 1.3. Assesment of the global performance of the method. In order to depict the overall performance of our interactive system, the WER resulting from the semi-supervision of the recognized blocks versus the percentage of supervisions (PS), explained above, will be used. This is because PS strongly depends on the method (so on the tolerance W∗) , while it is not so dependent on the task . Conversely, in despite of providing a more fair picture of the real effort, PK along with LT depend absolutely on the task. Additionally, these results are to be compared with the corresponding "Fully Supervised" experiments. These experiment consists in fully transcribing a certain number of blocks, and then accepting the output of the ASR of the rest of the speech. The effort for these transcriptions is accounted as just the PS, and the residual WER as the WER resulting of the recognized blocks. This serves to compare a non-guided interactive system with ours. Moreover, it should be highlighted that an ideal interactive system would avoid the supervision of correctly recognized words. Thus, the less equal operations a user would perform, the better. For our system, too many equal operations would mean that the confidence measure cannot successfully locate the errors. In the latter case, the method will ask for to many supervisions, yielding to a residual final WER better than requested. This behavior is preferable compared to a method that may end with errors higher than tolerated. Thus, it is also convenient to show the behavior of the confidence measures along the whole process of semi-supervision. The performance of the confidence measures is usually described in term of the Classification Error Ratio (CER). Here, however, it is further more relevant the comparison or the difference between the real and the estimated edit cost of the words in each confidence measure rank. Let them be denoted by Ecand ˆ Ecrespectively. Obviously, ˆ Ec=ˆ E+ c+ˆ E−, as it can be derived from formulae in chapter 2, . Where ˆ E−, as formulated in equation 7, is the main driver of the method. And, it is also interesting the ratio of supervisions
2. EXPERIMENTAL SET-UP 19 which consisted of equal operations (NE) to the total number of supervisions (NS). Let this ratio be denoted by PE. 2. Experimental set-up Our method was evaluated for the task of transcribing the whole WSJ0 training set, as well as the WSJ0+1 training sets of the Wall Street Journal (WSJ) speech database [7]. Three different scenarios were studied: 1) Both, the AM and the LM, were trained incrementally as explained in chapter . From now on, this scenario will be referred as as incremental LM or "LM inc". 2) Only the AM was incrementally trained, while the LM was the external WSJ 3-gram LM 5k(4.986) word closed vocabulary (LM 5K). And 3) for the External WSJ 3-gram LM 20k(19.979) word open vocabulary (LM 20K). The fixed LM remained unchanged for the entire transcription process. The AM models learned using the fixed LM can be used to recognize the benchmark tests of the Nov’92 ARPA evaluations [7], in order to assess the performance of the trained HMMs . Nonetheless, it should be noted that the fixed LM experimental setups are not only motivated for the sake of the comparison to the WSJ benchmarks. It is also motivated because, in a real case scenario, it is easy to find enough data such as being able to train a fixed LM good enough along the completion of the task. While this is not true for the data needed to build a good AM.The overall characteristics of the WSJ speech database are shown in Table 2. In addition, the overall difficulty of train and test sets in terms of the perplexity is shown in Table 2 for both language models (LM-5k and LM-20k) Table 1. WSJ speech database statistics. WSJ0 WSJ1 nov’93 Hub-1 Nov’92-5k Nov’92-20k Time (h.) 15 66 0.4 0.7 0.7 Utterances 7k30k213 330 333 Words per utterance 18 ±8 17 ±7 17 ±8 16±6 16±6 Speakers 84 200 10 8 8 Vocabulary size 8k13k1k5.4k5.6k Running words 129k510k3k5.4k5.6k Table 2. Perplexity on train and test sets depending on the LM. Exteral LM WSJ0 Train nov’93 Hub-1 nov’92-5k nov’92-20k 3-gram LM-5K −115 53 − 3-gram LM-20K 146 170 −142 2.1. Initialization. •Segementation of the corpus into blocks The WSJ0 training set was split into 12 blocks, each one containing utterances from 7 different speakers. Acoustic models were completely re-trained every time the semi-supervision of a block was fulfilled. These Hidden Markov Models (HMMs) consisted of clustered word-internal 3-state triphones with a number of 16 Gaussianmixtures per state. For the larger task WSJ1+0, the WSJ1 train set was split into 29 blocks, and then the 12 blocks from the WSJ0 were added. In this case, HMMs were trained with 24 Gaussian-mixtures per state.
20 4. PERFORMANCE OF THE INTERACTIVE SPEECH TRANSCRIPTION SYSTEM The average duration of each block in WSJ0 train set was about one hour, and two hours for the WSJ1 train set. •Explored Tolerance Errors Several tolerance values were tested from W∗= 1% to 50%. On this corpus, these values correspond to roughly allow from one error every 3 utterances (W∗= 1%) up to 9 errors per utterance (W∗= 50%). Additionally, as a comparison baseline, W∗= 0%, which is equivalent to perform a full supervision was also tested. •Initial models The initial HMMs, LM, optimal parameters for the decoder, confidence threshold and parameters for ˆ Wsemi (i.e. {Ec+ r},{Rc+ r},{Rc− r},{Nc+ r}for c= 1..4) , were obtained from the initially fully supervised part of the task. This fully transcribed part is split into 2 blocks, named 0 and 1. For experiments with incremental learning of the ASR models, a pre-initial ASR system is trained with block 0. Then, the block 1 is used as a development set in order to optimize the parameters above. After that, the a new initial HMMs and LM were trained using the whole block. However, For experiments using an external LM or AM will directly use the block 1 for the optimization of the parameters. Preliminary HMMs are needed to built the initial ASR system. Also, certain ASR parameters have to be properly estimated. With this purpose, the first block was considered as fully transcribed manually and used to train initial HMMs and optimize ASR parameters. In addition, the optimal confidence threshold (Cτ) to select the high-confident parts was estimated by minimizing the confidence error rate (CER) [20] in this block. On the other hand, the balancing method updates the values of its parameters after each user interaction. Nevertheless, it still requires a good initialization ({ ˆ Ec+ 0}, {Nc+ 0}, {Rc+ 0}, {Rc− 0}) in order to perform properly from the very beginning. These initial values were those resulting of applying the method itself on the first block. 2.2. Model training. As explained before, an ASR system depends on an acoustic and a language model (AM, and LM respectively): The AM consisted of Gaussian-mixture HMMs of clustered tri-phonemes with no cross-words. The number of mixtures varied depending on the corpus sets to be transcribed: For the experiments involving just WSJ0 sets, the number of mixtures was 16. For WSJ1 or WSJ1+0 sets, it was 24. which were trained with HTK tool-kit [21]. While the LM consisted of 3-ngram models smoothed with interpolation using the Knesser-Ney discount , trained using the SRILM toolkit [15]. It should be noted that although this kind of ASR training has been the facto standard over the last decades, it has been recently superseded by other techniques. However, for the purpose of the evaluation of our interactive system, there is no need of achieving cutting-edge performance for the recognition. Experiments conducted in the fist subsection of the results section above, used the iAtros tool-kit [5] for the recognition. However, other experiments were performed using the last version of the ak recognizer [3]. The latter achieves better performance and lower computation times for the same level of pruning than iAtros, , however, it became available to us later. Consequently, only the results in the second subsection in the results section have been re-evaluated using ak . Nevertheless, the behavior of the system is quite independent of the recognizer and of their performance. 2.3. Interaction: User simulation. For the sake of speed, corrections were performed automatically by means of a simulation of a real user:
3. RESULTS 21 It was assumed that a real user can always flawlessly understand and transcribe any speech in the audio; although a couple units of WER is usually found in manual transcriptions . This way, the simulation consisted in guessing the edit path a user would apply. Moreover, the number of edit operations a trained user performs was assumed to be always minimum. Then, two different possibilities were discussed about how best mimicking the the operations a real user would do: •Straight matching using the temporary alignment: The matching would be limited to reference words spoken in the same temporal range as the recognized word. •Matching using the edit distance: The matching would be found using the Levenstein algorithm. Insertions, however, should be found over the the time range, as in the option before. It should be noted that this can yield to mismatched substrings in rare cases; Also, the edit distance is not properly defined for sub-strings since several sub-strings from the reference can match a recognized word; And that the way the matching is done varies the set of edit operations (edit path). Nevertheless, in practice, different users will likely perform identical corrections since they benefit from grasping other information that the system does not, as phonetic similarities, semantic and syntactic restrictions, and, mainly, the place of words in time. The first option may seem more fair and natural. The second, however, may yields optimistic evaluation results. However, as the implementation of the prototype interface (next chapter) allows the introduction of corrections in a utterance not necessarily corresponding to the asked word, the second option is then more realistic in our case. In any case, the simulation required a whole alignment of the text references to the audio was needed from the beginning. This alignment was carried out by training an ASR model with the full corpus, and then making a forced recognition. Moreover, in order to implement the correction convention suggested to the users (see sect. 1.3), selected reference words for matching were those which their middle time point laid inside the temporal range of the recognized word (with no audio margins). Then, the selection of the edit path followed no other criterion than choosing the first one returned by the actual implementation of the Levenstein algorithm. Despite the commented issue, this barely biased the experiment since, in almost all cases, there was just one possible edit path. This was because, thanks to the time restrictions, the matched sub-strings contained up to three words at most. Finally, in case of no recognized words in the line, full insertion of the reference text of the line was performed. 2.4. Unsupervised Learning. A brief explanation can be found on the chapter 2, section 2. 3. Results First, in subsection 3.1, it is presented the results of transcribing the smaller task, WSJ0 train set, split in 12 blocks. Several tolerance values were tested from W∗= 1% to 60% . Then, more detailed results are depicted for the task of transcribing first the WSJ1 plus the WSJ0 train sets, split into 29 and 12 blocks respectively. The results of two representative thresholds W∗= 6% to 12% are shown. All the recognitions in this subsection were performed using the iAtros toolkit using the incrementally trained HMMs of 16 and 24 gaussian mixtures, for the WSJ0 and WSJ1+0 respectively. The parameters for the recognizer and the confidence measure classifier were optimized for the first block, and then never updated. The
22 4. PERFORMANCE OF THE INTERACTIVE SPEECH TRANSCRIPTION SYSTEM number of confidence measure ranks was 4. The threshold classifier CT h was optimized in the first block but never updated. The decisions on the selected parameters are justified in the comparison in the following subsection. In the following subsection, the effect of different parameters of the method on the global performance are exemplified by means of the performance results on the smaller task of WSJ0. The necessary experiments were conducted in order to assess the following features: "Initialization of the W−parameters" , "The number of confidence measure ranks in a utterance" , "Considering a Utterance as a word or as a sentence" , "Number of blocks" and "Updating ASR and Threshold Classifier Parameters". The recognitions were performed using the ak toolkit , but for the last comparison. The task was split into 6 blocks instead of 12, but for "Number of blocks" and "Updating ASR and Threshold Classifier Parameters" comparisons. Several tolerance values were tested from W∗= 1% to 60%. 3.1. Analysis of the performance in transcribeing the WSJ train sets .•Transcription of the WSJ0 train set For clarity, only the results of the supervision effort and residual accumulated WER after the transcription of each block of WSJ0 training data using a tolerance error of W∗= 20% and 5% are depicted in Fig. 1. The rest of the lines corresponding to other tolerance values had an identical tendency to these one, from the fully supervised task (W∗= 0%) to the almost non supervised experiment (W∗= 50%). The figure shows that all the experiments behaved in a similar manner along the process. However, for LM-20k the reduction in user effort was greater. This was because LM-20k yields better recognition accuracy than LM-5k in WSJ0 training data (fig.fig:WSJ0-WER ). The incremental LM obtained intermediate results, but had a stepper rate of improvement. However, which recognition performance is better is completely dependent on this corpus and the LM: The fixed ones were estimated from a training os million of sentences. The incremental LM is from less than 1/100 that size, and the train data is pronned with errors. But, even so, the incremental is better adapted to task since the tokens (like dots, colons, etc.) were verbalized. Thus, it greatly differs from the regular english text from which were estimated the fixed LM. In the end, both effects seems to compensate. It is also important to note, that the stated before is not true for the very high tolerated errors. In these cases , almost no semi-supervision was asked to the user. Thus, the ASR models only changed because of the unsupervised learning. The incremental LM yielded catastrophic results. This is the usual when re-training a system with its own output, even the confidence measures prevents some of the worst recognized samples to be added to the new training set. However, the fixed LM the initial recognition WER surprisingly improved block after block, and the absolute difference with the results from the low W∗, or even the baseline, was considerably small. Thus, it can be stated that it is desirable stating the method with a huge initial external LM. Which, in turn, might be further adapted using the incremental training.
3. RESULTS 23 0 20 40 60 80 100 2 4 6 8 10 12 14 0 5 10 15 20 25 30 35 40 Time(h) Evolution of the %Supervisions and Residual WER for WSJ/train0 Supervisions(%) LM 20K Residual WER (W- %) 0 20 40 60 80 100 2 4 6 8 10 12 14 0 5 10 15 20 25 30 35 40 Time(h) Supervisions(%) LM inc Residual WER (W- %) 0 20 40 60 80 100 2 4 6 8 10 12 14 0 5 10 15 20 25 30 35 40 Time(h) Supervisions(%) LM 5K Residual WER (W- %) %Supervisions: W-= 5% W-=20% %W-: W-= 5% W-=20% Figure 1. User effort and final quality of the transcription of WSJ0 training data using LM-5k, LM-20k, or incremental LM: - Reduction of the user effort in terms of the number of supervisions relative to the recognized words in a block (at top). - Residual accumulated WER after semi-supervision of each one of the blocks (at bottom).
24 4. PERFORMANCE OF THE INTERACTIVE SPEECH TRANSCRIPTION SYSTEM 20 30 40 50 60 70 2 4 6 8 10 12 14 Time(h) WER(%) LM 20K 20 30 40 50 60 70 2 4 6 8 10 12 14 Time(h) Evolution of Accumulated WER of the recognized blocks of WSJ/train0 (Initial fully supervised blocks not included) WER(%) LM inc W*= 0% W*= 5% W*=20% (Unsupervised) W*=50% 20 30 40 50 60 70 2 4 6 8 10 12 14 Time(h) WER(%) LM 5K 0 10 20 30 40 50 2 4 6 8 10 12 14 Time(h) WER(%) LM 20K 0 10 20 30 40 50 2 4 6 8 10 12 14 Time(h) Delta Recognition WER of WSJ/train0 (WER Block - Baseline ) WER(%) LM inc W*= 5% W*=20% (Unsupervised) W*=50% 0 10 20 30 40 50 2 4 6 8 10 12 14 Time(h) WER(%) LM 5K Figure 2. ASR model improvement due to interaction in semisupervising the WSJ0 training set. (top) WER of the accumulated automatically recognized transcriptions. (top) Difference between the WER of each recognized block and its corrsponding baseline (W∗= 0%)
3. RESULTS 31 Table 5. Semi-supervision on block 42 (nov’931-Hub1 test) Baseline W∗= 6% W∗= 12% Supervisions 3174 3113 3026 Equal Ops. 2227 2121 2063 WER 37.46 36.30 36.59 WERsemi 0.00 3.10 6.16 Table 5 shows results on block 42. This block corresponded exactly to the WSJ’92-93 Hub-1 test partition. Speech and vocabulary in this block differs slightly from the train. For instance, in the train partition, punctuation symbols are verbalized (ex. ’(’ is spelt as ’open-paren’). While, this does not happen in test, resulting in a more natural speech. Table 6. Semi-supervision using external LM 20K on block 42 (nov’931-Hub1 test) Baseline W∗= 6% W∗= 12% Supervisions 3223 2611 1919 Equal Ops. 2631 2005 1437 WER 18.99 20.35 21.09 WER A.Sanchis 16.1 - - WERsemi 0.00 3.11 6.19 Table 6 shows results on block 42, but this time an external LM was used for recognition of this last block, instead of the learnt during the semi-supervision. All previous steps of training, semi-supervision and recognition were exactly the same as in table 5. The external 20k-words 3-gram LM was built up from a text supplied along the WSJ speech corpus, which is precisely intended for recognition on Hub-1 test. Additionally, line "WER A.Sanchis", shows the published stated-of-the-art results using iAtros. The differences are due because of the higher pruning for the decoding and the lower number of gaussian mixtures used here. In summary, •User effort is greatly reduced as it learns (fig. 4). •Recognition performance is just slightly worse than the best possible (fig. 5 top). •WER difference of each recognition between an experiments and the baseline is maintained in average (fig. 5 bottom). This means that the slight worsening in the accumulated WER throughout the task, is due only to the accumulation of errors, not a degrade in the ASR quality. •Word classifier performance behaves as expected, but its performance degrades over time (fig. 4 top). As a consequence, the number of redundant supervisions increases (fig. 4 bottom). •ˆ Wsemi insignificantly varies around requested W∗. •ˆ Wsemi is always pessimistic (which is desirable): resulting transcriptions are always better than expected. •The system effectively adapts, in terms of the number of supervisions, to the true, unknown, quality of recognition: On block 41 (table 4), 50% to 68% of the recognized words were supervised. Next block (table 5), which was poorly recognized, 98% of the words. While, only 60% to 81% when an external LM improved the recognition (Table 5).
32 4. PERFORMANCE OF THE INTERACTIVE SPEECH TRANSCRIPTION SYSTEM 3.2. Analysis of the impact of the parameters of the method. All experiments started with ASR models trained using the blocks 0 plus 1. However, the impact on the WER and PS (%supervisions) of these fully supervised blocks has been removed from the presented results. By so, eliminating the smoothing effect on the results. This way, the differences amongst the conducted experiments can be spotlighted clearer. •Initialization of the W−parameters The initialization of { ˆ Ec+ 0}, {Nc+ 0}, {Rc+ 0}, {Rc− 0} played a key role. This is because an increased number of unsupervised words R− cmay not increase W−. In fact, for R− c→ ∞,W−will tend to the WER of the supervised parts. This issue has several side effects. For instance, if W−is lower enough than the requested W∗, it is likely that no more utterances will be supervised for the remaining speech. This behavior is only desirable when W∗is much higher than the real recognition WER, and it will be still a problem if the nature of the speech would suddenly change. In order to deal with this issue, the initial values of these parameters were so that W−was initially about 100%. Then, the first block was used for tuning this initial values, as if the user would have supervised every word in the first block. In order to prevent overtraining on this initial block, or that this block would have no impact on the estimator, the initial values of R+ cand N+ cwere about the half numbers of words in the first block. Despite this initialization seem reasonable, the initial R+ chave some impact on the experiments. Fortunately, the lower W∗, the higher number of supervisions were necessary, so the lower impact on the performance. •The number of confidence measure ranks in a utterance As explained before, the estimation of the WER of the unsupervised parts, W−, has been formulated to depend on the confidence measure relative to the other confidence measures values in a utterance. To do so, words are ranked form the lowest confidence value to the highest, and they are assigned a rank value. Let c= 1..C be the number of ranks. If so, then the highest rank c=Cincludes all the highest confidence values not in the precedent ranks. It should be spotted that the nature of this ranking will overrate the expected edit cost of the words in the highest rank compared to the the real one. This overrated estimation is due to the lower confidence words in the highest ranked confidence group, because they are more likely to be asked for supervision and been wrong. To overcome this, an obvious manner seems to increase the number of ranks, or, even, to unbound the number of ranks. Unfortunately, for a high number of ranks, the estimation of the edit cost associated to the highest ranks is nether reliable because not all the utterances have so many words, and they are unlikely to be supervised.
3. RESULTS 33 0 5 10 15 20 25 30 35 40 0 20 40 60 80 100 Residual WER(%) after semi-supervision PS ( %Supervisions from block 2 to the end) Residual WER vs. %Supervisions for the WSJ0-train set (6 blocks) 1 CM levels 4 10 20 Fully Sup. Figure 8. Overal performance of the method for several W∗. Several Number of CM ranks are compared against the baseline. From figure 8, it can be stated that for this corpus the number of CM ranks level have little impact. Only for percentage of supervisions (PS) from 20% to 30% (corresponding for W∗from 20% to 30%) there was a clear optimal around the 4 ranks per utterance.
34 4. PERFORMANCE OF THE INTERACTIVE SPEECH TRANSCRIPTION SYSTEM 0.3 0.35 0.4 0.45 0.5 0.55 0.6 0.65 0.7 0.75 0.8 0 20 40 60 80 100 PE (%) PS ( %Supervisions from block 2 to the end) %Redundant Interactions vs. %Supervisions for the WSJ0-train set (6 blocks) 1 CM levels 4 20 0 5 10 15 20 25 30 35 40 0 5 10 15 20 25 30 35 40 W- (%) (from block 2 to 6) W* ( % Tolerance Error) %Residual WER vs Required WER Threshold Desiderable Behaviour 1 CM levels 4 20 10 20 30 40 50 60 70 80 90 100 0 5 10 15 20 25 30 35 40 PS (%) (from block 2 to 6) W* ( % Tolerance Error) %Supervisions vs Required WER Threshold 1 CM levels 4 20 Figure 9. Comparision of experiments with different number of CM ranks. In figure 9 all the experiments perform similarly for the percentage of equal operations. However, the number of supervisions was lower for the higher number of ranks. As a consequence, the residual WER are higher. But since all are well below the required threshold, that a better estimation of WER comes with more ranks. In table 7 is shown the relative difference between the expected and the real Edit Costs per rank : ˆ E− c−E− c E− c . Consistently, the more ranks, the lower factor for the highest confidence rank (exact estimation is factor 1) .
3. RESULTS 35 Table 7. Relative difference between the expected and the real Edit Costs per rank of confidence in a utterance 20 CM ranks: W∗Lowest 2 3 4 5 6 ... 17 18 19 Highest 5% 1.5 0.8 0.78 0.71 0.73 0.92 ... 7.7 8.31 8.87 38.29 30% 0.36 0.45 0.83 1.45 3.26 6 ... 3.77 4.2 3.44 10.39 4 CM ranks: W∗Lowest 2 3 Highest 5% 0.11 0.09 0.1 99.68 30% 0.4 0.55 0.79 98.24 Nevertheless, for the rest of experiments 4 ranks will be used since it yielded better in 8. While the performance in terms of the user effort from the 20 ranks is very small. •Considering a Utterance as a word or as a sentence In the previous chapter was remarked that our interactive method can work regardless of what a utterance is considered. For the WSJ the natural subsegmentation of the blocks are the sentences. However, it might be considered that an utterance is an only word, as it is published for handwritten recognition. 0 5 10 15 20 25 30 35 40 0 20 40 60 80 100 Residual WER(%) after semi-supervision PS ( %Supervisions from block 2 to the end) Residual WER vs. %Supervisions for the WSJ0-train set (6 blocks) Sentence level Word level Figure 10. Overal performance of the method for several W∗. The utterances of words and samples are compared against the baseline.
36 4. PERFORMANCE OF THE INTERACTIVE SPEECH TRANSCRIPTION SYSTEM 0.3 0.35 0.4 0.45 0.5 0.55 0.6 0.65 0.7 0.75 0.8 0 20 40 60 80 100 PE (%) PS ( %Supervisions from block 2 to the end) %Redundant Interactions vs. %Supervisions for the WSJ0-train set (6 blocks) Method:Sentence Method:Word , 4-CMranks 0 5 10 15 20 25 30 35 40 0 5 10 15 20 25 30 35 40 W- (%) (from block 2 to 6) W* ( % Tolerance Error) %Residual WER vs Required WER Threshold Desiderable Behaviour Method:Sentence; 4 CM ranks Method:Word; 4 CM ranks 10 20 30 40 50 60 70 80 90 100 0 5 10 15 20 25 30 35 40 PS (%) (from block 2 to 6) W* ( % Tolerance Error) %Supervisions vs Required WER Threshold Method:Sentence; 4 CM ranks Method:Word; 4 CM ranks Figure 11. Comparision of considering a every word a whole utternace and considering one sentence as a utterance. In figures 10 and 11, is shown that the sentence level performs better: the equal operations proportion is smaller and overall performance is smaller. It should be noted that the word level utterances has percentage of supervisions. This is because the method asks for high confident words as well as for the low confidence, since when the estimation is lower than the threshold the higher confidence words are not skipped, but just the next. This results in a quite random way of supervision, but the estimation is better since the supervised part have lots of high confidence words. Nevertheless, when putting together the supervisions and the resulting residual WER, the sentence-level utterances do outperform. Thus, from here all experiments will refer to the sentence-level version. •Number of blocks
3. RESULTS 37 It is expected that smaller blocks, so the more number of them, the faster will the system improve. Two conducted experiments comparing a 6 block and a 12 block partition are depicted in figures 12 and 13 0 5 10 15 20 25 30 35 40 0 20 40 60 80 100 Residual WER(%) after semi-supervision PS ( %Supervisions from block 2 to the end) Residual WER vs. %Supervisions for the WSJ0-train set (6 blocks) 4 blocks 12 blocks Figure 12. Overal performance of the method for several W∗. Segmentaton into 6 and 12 Blocks are compared against the baseline.
38 4. PERFORMANCE OF THE INTERACTIVE SPEECH TRANSCRIPTION SYSTEM 0.3 0.35 0.4 0.45 0.5 0.55 0.6 0.65 0.7 0.75 0.8 0 20 40 60 80 100 PE (%) PS ( %Supervisions from block 2 to Last) %Redundant Interactions vs. %Supervisions for the WSJ0-train set 6 Blocks 12 Blocks 0 5 10 15 20 25 30 35 40 0 5 10 15 20 25 30 35 40 W- (%) (from block 2 to Last) W* ( % Tolerance Error) %Residual WER vs Required WER Threshold Desiderable Behaviour 6 Blocks 12 Blocks 10 20 30 40 50 60 70 80 90 100 0 5 10 15 20 25 30 35 40 PS (%) (from block 2 to Last) W* ( % Tolerance Error) %Supervisions vs Required WER Threshold 6 Blocks 12 Blocks Figure 13. Comparision of segmenting the task into 6 and 12 blocks. Results, however, show the same performance for low tolerances, while a slightly worse performance for the 12 blocks experiment for higher tolerances. Nonetheless, this is because the initial partition on the 6 block segmentation contains also the double quantity of speech than the initial block of the 12-block experiment. Starting with poorer recognition performances has a more detrimental impact when tolerating a high value of errors. From here, experiments conducted for the WSJ0 train set will be split into 12 blocks. •Updating ASR and Threshold Classifier Parameters The conducted experiments were the same as in the previous section for the WSJ0 train set (split into 12 blocks, recognized with the iAtros toolkit. The update of the parameters were performed as for the initial set, explained above, but using
3. RESULTS 39 an external development set. The development set was the dev-nov’92-5K for the LM inc and LM 5K , and the dev-nov’92-20K for the LM 20K. Differences in the results were quite small, in benefit for the update version in average for the WER obtained in the recognition of the blocks of the task. However, there was no significant difference when evaluating the external tests with the HMMs models for the fixed LM experiments. This found is reasonable since the task differs considerably in the language structure to the benchmark tests. This is why all the presented results have been in the non-updated version; which requires much less computation time. Nevertheless, for other speech tasks, or when reusing the models resulting from one task to another, the update may be mandatory. In the following, all the graphs corresponding to this assessment can found together. No for further remarks on the meaning of each graph as they are identical to those presented in the previous section for the WSJ0
40 4. PERFORMANCE OF THE INTERACTIVE SPEECH TRANSCRIPTION SYSTEM 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM inc 50 55 60 65 70 75 80 85 90 95 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 5K 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 20K Balancing method Tolerance W*= 0% (Params. Not Updated) W*= 5% (Params. Not Updated) W*=20% (Params. Not Updated) 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM inc 50 55 60 65 70 75 80 85 90 95 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 5K 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 20K Balancing method Tolerance W*= 0% (Params. Not Updated) W*= 5% (Params. Not Updated) W*=20% (Params. Not Updated) 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM inc 50 55 60 65 70 75 80 85 90 95 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 5K 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 20K Balancing method Tolerance W*= 0% (Params. Not Updated) W*= 5% (Params. Not Updated) W*=20% (Params. Not Updated) 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM inc 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 5K 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 20K Balancing method Tolerance W*= 0% (Params. Not Updated) W*= 5% (Params. Not Updated) W*=20% (Params. Not Updated) 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM inc 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 5K 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 20K Balancing method Tolerance W*= 0% (Params. Not Updated) W*= 5% (Params. Not Updated) W*=20% (Params. Not Updated) 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM inc 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 5K 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 20K Balancing method Tolerance W*= 0% (Params. Not Updated) W*= 5% (Params. Not Updated) W*=20% (Params. Not Updated) Updating Not upd. Baseline W*=5% W*=20% 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM inc 50 55 60 65 70 75 80 85 90 95 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 5K 30 40 50 60 70 80 90 100 0 2 4 6 8 10 12 14 Time(h) %Supervisions for WSJ/train0 Supervisions(%) LM 20K Balancing method Tolerance W*= 0% (Params. Not Updated) W*= 5% (Params. Not Updated) W*=20% (Params. Not Updated) Figure 14. %Supervisions: Comparison between updating and not updating parameters
3. DEMONSTRATION 47 a b Figure 2. Outline of an screenshot of the textbox with the transcriptions. Shaded regions: Words recognized with low confidence measure values. Circled word: Word under supervision. It should be noted that has been published that highlighting low confidence words helps user in the detection of the errors when the confidence estimation are correct ([17]) The selected audio is automatically played once. It should be noted that the enlargement of the piece of audio to be played influences the user decisions, especially for the number of insertions to be done. Then, the user can enter the proper correction. At the same time, the user can re-play the piece of the audio as many times as required by pressing the tabulator key. Once finished, the user should press the enter key. The correction is underlined, and it assigned the highest possible confidence. After each correction, the estimator of the residual WER (W−) is updated. Then the system will ask the user to correct the next lowest confidence word in the sample, or it will jump to the next utterance if the estimator is below the tolerance error threshold (W∗). It should be noted that the required quality on the transcription was specified in the configuration file in term of W∗. This threshold can also be set by the user using a slider at any time (fig. 3). Figure 3. Screenshot of the textbox with the transcriptions. Shaded regions: Words recognized with low confidence measure values. Circled word: Word under supervision. Finally, in order to improve the recognition of remaining utterances of the task, the speech recognizer model can be improved by retraining models (fig. ). New data pairs of audio an text will be added to the training set. These pairs are the built from the contiguous subsegments not in maroon background (i.e., only those supervised and high-confidence parts): Once a new model is available, the remaining utterances can be re-recognized. This should help in reducing the number of corrections the system will ask to the user. 3. Demonstration The prototype was showcased and demonstrated on a special session during the IUI 2012 congress. A small speech corpus allowed the attendants to test our interactive speech transcription prototype. The corpus was intended to be automatically recognized in a laptop in very short time. It posed some typical recognition errors. Users were able to perform the corrections while being assisted by the system, to change the tolerance error, and to re-train the models to check how the recognition has improve and it whether it has learned new vocabulary. A video with a short overview of the system is included in the CD of this master thesis.
48 5. A PROTOTYPE INTERFACE TO INTERACTIVE SPEECH TRANSCRIPTION 4. Usability No formal tests on the usability of the interface have been conducted. However, the overall impression of the authors and other non-professional transcripts who tested the demo at the congress was satisfactory. The main concerns came from a small percentage of the words asked for supervision which segmentation was wrong. This mainly happened for the extra inserted words (i.e. when a deletion operation should be performed). Even so, although non performed deletions increase the resulting WER, they are preferable for a real user from the language understanding point of view. Also, incorrectly placed extra words are preferable for indexing and term search applications that would use the resulting transcriptions. Instead, the omission of words is a critical issue. In order to avoid the most of the chopped words an enlargement of the portion of the audio to be played of 15 ms was found to be the best option. Nevertheless, it would be desirable a cleverer way to modify the margins in order to increase the chances of understanding. In the previous chapter it was stated that a user should attach strictly to some conventions. For instance, a user should deleted a word whenever less about the half of the word is uttered even if it can be figured out. However, while this servers properly for the evaluation purposes, allowing corrections more than those for the asked words to be supervised greatly improves the user experience and the final quality of the transcriptions. This behavior was implemented by simply performing an additional estimation of which parts of the introduced corrections corresponded to which words in the utterance under supervision, instead of assuming the correction corresponds strictly to the word under supervision. The estimation is performed by means of the of the Levenstein algorithm . The update of the ˆ W−estimator is performed as if the extra corrected corresponding recognized words would have been asked for supervision, and that the corresponding corrections were introduced. It should be noted that this extra behavior is usually useful because the user can rapidly figure out some corrections just by listening the word under supervision and reading the recognized surrounding words. This way, in case that more than one recognized word was corrected, the system will skip the supervision of those recognized words if they would have been later selected for supervision. Thus, in those cases no increase of the user effort would have happened. Finally, it should be noted that for better comprehension of the speech, the more audio context the better and the more sequentially performed corrections , the better. Thus, this is an important issue that is left as a future work.
CHAPTER 6 CONCLUSIONS 1. On the method of balancing error and the user effort In summary: A simple yet effective method to find an optimal balance between recognition error and supervision effort has been applied to interactive speech transcription. Empirical results confirm previous works on handwriting recognition showing that this strategy is effective to reduce the supervision effort by allowing a maximum tolerance error in the speech transcriptions. Moreover, results show that a tolerance error in the transcriptions does not affect critically on the incremental learning of acoustical models. Thus, this method can be used also for producing ASR models resulting in similar performance to those generated using fully manually transcribed corpora. General Remarks Concerning the evaluations : •The estimation of the residual WER ( ˆ W−) is always pessimistic, which is preferable. •The pessimistic behavior is due to nature of the method, that is based solely in the supervised low-confidence parts. •The pessimistic behavior yielded to many spurious supervisions (equal operations). Nevertheless, equal operations do not require as much effort as other necessary operations; and , in turn, they help to improve the estimation which would be too pessimistic if not. •The poor performance of the confidence measures, especially at the end of the tasks, worsens the ˆ W−; Fortunately, the effect is insignificant since ˆ W−is robustly estimated throughout the previous blocks. Also, it makes the system more likely to asks for more words that it are correct. •The initialization poses an effect over the performance, but a reasonable initialization is easy to find. •The update of the parameters of the recognizer and the confidence measure classifier yield no significant improvement, but it seems that in a general scenario it would. •The procedure performs better when it works at sentence level. •In practice, badly segmented words (which is usually the cause of the wrong recognition), poses a difficulty for the user. Although it might be alleviate enlarging the margins by just 15ms. 49
50 6. CONCLUSIONS Remarks on the large task WSJ1+0: •User effort is greatly reduced as it learns (fig. 4). •Recognition performance is just slightly worse than the best possible (fig. 5 top). •WER difference of each recognition between an experiments and the baseline is maintained in average (fig. 5 bottom). This means that the slight worsening in the accumulated WER throughout the task, is due only to the accumulation of errors, not a degrade in the ASR quality. •Word classifier performance behaves as expected, but its performance degrades over time (fig. 4 top). As a consequence, the number of redundant supervisions increases (fig. 4 bottom). •ˆ W−insignificantly varies around requested W∗. •ˆ W−is always pessimistic (which is desirable): resulting transcriptions are always better than expected. •The system effectively adapts, in terms of the number of supervisions, to the true, unknown, quality of recognition: On block 41 (table 4), 50% to 68% of the recognized words were supervised. Next block (table 5), which was poorly recognized, 98% of the words. While, only 60% to 81% using an external LM improved the recognition (Table 5). 2. Future Work Although a great reduction in the effort can be achieved with this simple method, there are three main aspects that should be addressed: •The issue of the isolated words: The lack of audio context and the badly automatically segmented words makes harder for the user figuring up the proper correction. Furthermore, the system keeps jumping from one place in the sentence to another until it decides to jump to the next utterance. To alleviate this problem the driver of the method should include rules and cost functions in order to group into one larger segment words that are likely to be wrongly regonized. Also, it will be better if the segments were asked from left to the right. •The confidence measures: A bad performance of the confidence measures mislead the overall process. Thus better performance should be achieved. Our next research will precisely be centered on this topic by finding regions of low confidence instead of isolated words. This will improve the capturing of insertions on regions with no recognized words. Also it will help in using a segment driven approach, instead of the word-driven, as just proposed. •The estimation of the error: The estimation might be further refined using more features than the rank level of confidence measure. It could be modeled depending on the words itself, the relative position in the sentence, the continues value of the confidence measure, etc.
BIBLIOGRAPHY [1] Behrouz Abdolali and Hossein Sameti. A Novel Method For Speech Segmentation Based On Speakers’ Characteristics. arXiv.org, cs.AI, May 2012. [2] C. Barras, E. Geoffrois, Z. Wu, and M. Liberman. Transcriber: a Free Tool for Segmenting, Labeling and Transcribing Speech. In Proceedings of LREC, pages 1373–1376, 1998. [3] A. gimenez. AK: Adriaś kit . prhlt.iti.upv.es, pages –, 2012. [4] D. Hakkani-Tür, G. Riccardi, and G. Tur. An active approach to spoken language processing. ACM Transactions on Speech and Language Processing (TSLP), 3(3):1–31, 2006. [5] M. Luján Mares, V. Tamarit, V. Alabau, C.D. MartÄśnez-Hinarejos, M.P. i Gadea, A. Sanchis, and A.H. Toselli. iATROS: A speech and handwritting recognition system. V Jornadas en TecnologÄśas del Habla (VJTH’2008), pages 75–78, 2008. [6] Saturnino Luz, Masood Masoodian, and Bill Rogers. Interactive Visualisation Techniques for Dynamic Speech Transcription, Correction and Training. In Stuart Marshall, editor, Proceedings of CHINZ 2008, The 9th ACM SIGCHINZ Annual Conference on Computer-Human Interaction, pages 9–16, Wellington, New Zealand, 2008. ACM Press. [7] D.S. Pallett, J.G. Fiscus, W.M. Fisher, and J.S. Garofolo. Benchmark tests for the DARPA spoken language program. In Proceedings of the workshop on Human Language Technology, pages 7–18. Association for Computational Linguistics, 1993. [8] B Ramabhadran, O Siohan, and A Sethy. The IBM 2007 speech transcription system for European parliamentary speeches. In IEEE Workshop on ASRU, pages 472–477, 2007. [9] Luis Rodríguez-Ruiz, Francisco Casacuberta, and Enrique Vidal. Computer Assisted Transcription of Speech. In Proceedings of the 3rd Iberian Conference on Pattern Recognition and Image Analysis, Volume 4477 of LNCS, pages 241–248, 2007. [10] R. Sánchez Sáez, J.A. Sánchez, and J.M. Benedí. Confidence measures for error discrimination in an interactive predictive parsing framework. In Proceedings of the 23rd International Conference on Computational Linguistics: Posters, pages 1220–1228. Association for Computational Linguistics, 2010. [11] A. Sanchis, A Juan, and E Vidal. A Word-Based Naïve Bayes Classifier for Confidence Estimation in Speech Recognition. Audio, Speech, and Language Processing, IEEE Transactions on, 20(2):565–574, 2012. 51
52 6. CONCLUSIONS [12] Alberto Sanchis. Estimación y aplicación de medidas de confianza en reconocimiento automático del habla. PhD thesis, Departamento de Sistemas Informáticos y Computación, 2004. [13] N. Serrano, A. Sanchis, and A Juan. Balancing error and supervision effort in interactive-predictive handwriting recognition. In Proceedings of the 15th international conference on Intelligent user interfaces, pages 373–376. ACM, 2010. [14] G Stemmer, S Steidl, E Nöth, H Niemann, and A. Batliner. Comparison and combination of confidence measures. Text, Speech and Dialogue, pages 561–582, 2006. [15] A. Stolcke. SRILM - An Extensible Language Modeling Toolkit. In ICSLP, 2002. [16] Martin Sundermeyer, Markus Nußbaum-Thom, Simon Wiesler, Christian Plahl, Amr El-Desoky Mousa, Stefan Hahn, David Nolden, Ralf Schlüter, and Hermann Ney. The RWTH 2010 Quaero ASR evaluation system for English, French, and German. In ICASSP, pages 2212–2215. IEEE, 2011. [17] Keith Vertanen and Per Ola Kristensson. On the benefits of confidence visualization in speech recognition. In CHI ’08: Proceeding of the twenty-sixth annual SIGCHI conference on Human factors in computing systems. ACM Request Permissions, April 2008. [18] Y.Y. Wang, A. Acero, and C. Chelba. Is word error rate a good indicator for spoken language understanding accuracy. Automatic Speech Recognition and Understanding, 2003. ASRU’03. 2003 IEEE Workshop on, pages 577–582, 2003. [19] F. Wessel and H Ney. Unsupervised training of acoustic models for large vocabulary continuous speech recognition. IEEE Transactions on Speech and Audio Processing, 13(1):23–31, 2005. [20] F. Wessel, R. Schlüter, K. Macherey, and H Ney. Confidence measures for large vocabulary continuous speech recognition. IEEE Transactions on Speech and Audio Processing, 9(3):288–298, 2001. [21] S.J. Young, Woodland, P.C., and W J Byrne. HTK: Hidden Markov Model Toolkit V1.5. Cambridge Univ. Eng. Dept. and Entropic Research Labs Inc., Cambridge, UK, 1993.