scieee AI-readable full text Open interactive document viewer

Optimization algorithms for estimating modulation spectrum domain filters

Pachès Leal, Pau,Rose, R,Nadeu Camprubí, Climent

Abstract

The goal of the work described in this paper is to develop and evaluate procedures for automatic estimation of modulation spectrum filters to compensate for distortions in the modulation spectrum domain. The modulation spectrum (MS) is often used to describe the time sequence of spectral parameters (TSSPs) that are derived from the speech waveform, and is thought to be a good representation of many sources of variability in speech. These procedures will be used in the context of automatic speech recognition (ASR) applications where there is likely to be a significant mismatch in the MS characteristics that exist for system training and evaluation. Results are presented describing application of the algorithm to one task involving an artificially introduced MS distortion and to another task involving differences in speaking styles for training and testing. It is shown in the paper that these techniques are able to compensate for the effects of artificially introduced distortions that appear in testing. It is also shown that a small degree of compensation is obtained for speaking style mismatch, and this result is compared with the measured effects of the speaking style differences in the MS domain. An algorithm is presented for automatic estimation of the An algorithm to estimate automatically filters in the modulation spectrum domain. These are used to compensate for distortions in this domain or to obtain the difference coefficients that are a part of the acoustic vector handed over to the HMM-based speech model. The mathematical properties of the new algorithm are analyzed. Its performance is studied in two different experiments: in the first the goal is to alleviate an artificially simulated distortion while in the second we try to compensate for speaking rate distortions in a database which has two distinct parts differing significantly in speaking rate.

Full text

OPTIMIZATION ALGORITHMS FOR ESTIMATING MODULATION SPECTRUM DOMAIN FILTERS Pau Paches-Leal  y , z ,RichardC.Rose y , and Climent Nadeu z y AT&T Labs-Research, Florham Park, NJ, USA, z Univ. Politecnica de Catalunya, Barcelona, Spain ABSTRACT The goal of the work describ ed in this pap er is to develop and evaluate pro cedures for automatic estimation of modulation sp ectrum lters to compensate for distortions in the modulation sp ectrum domain. The mo dulation sp ectrum (MS) is often used to describ e the time sequence of sp ectral parameters (TSSPs) that are derived from the sp eechwaveform, and is thought to be a go o d representation of many sources of variability in speech. These procedures will b e used in the context of automatic sp eech recognition (ASR) applications where there is likely to b e a signicant mismatch in the MS characteristics that exist for system training and evaluation. Results are presented describing application of the algorithm to one task involving an articially introduced MS distortion and to another task involving dierences in sp eaking styles for training and testing. It is shown in the pap er that these techniques are able to comp ensate for the eects of articially intro duced distortions that appear in testing. It is also shown that a small degree of comp ensation is obtained for sp eaking style mismatch, and this result is compared with the measured eects of the sp eaking style dierences in the MS domain. 1. INTRODUCTION The time sequence of sp ectral vectors derived from the sp eech signal can b e represented by the Mo dulation Sp ectrum (MS) [5]. It has been prop osed for use in many applications. These include characterizing multipath distortions o ccurring in reverb erantenvironments [2], describing the eects of channel distortion and sp ectral estimation errors in ASR [5], and describing the eects of varying sp eaking styles for ASR [5]. When used as a representation of the long{term averaged spectrum or cepstrum parameters in ASR, the MS is dened as the p ower sp ectrum of the time sequence of the feature vectors that are input to the recognizer. In sp eech recognition feature extraction, MS lters have b een designed for the purp ose of selectively removing those p ortions of the modulation sp ectrum representing noise or channel distortions, and retaining the p ortion of the MS containing sp eech[5]. MS lters are commonly applied to ltering the sequence of cepstrum co eÆcients and also to computing the cepstrum dierence dynamic co eÆcients. It is imp ortant to note that the MS lters used for b oth the cepstrum and dynamic cepstrum coeÆcients are generally obtained empirically. As a result, it is often the case that MS lters designed to optimize performance for one task under a given set of conditions prove to be sub optimal when applied to another task. The automatic pro cedure presented here for estimating mo dulation sp ectrum lters is an attempt to improve sp eech recognition p erformance under highly mismatched recognition / training scenarios.  This researchwas conducted at AT&T Shannon Laboratory as part of P.Paches-Leal's Ph.D. thesis with the UPC The pap er is organized as follows. Section 2. describes the mo dulation compensation algorithm (MCA) and discusses implementation issues relating to the algorithm. In Section 3., the algorithm is implemented on an arti- cially intro duced distortion applied to the test data. Finally, in Section 4., the issue of automatic comp ensation for sp eaking rate mis{match using the MCA algorithm is investigated. 2. MODULATION COMPENSATION ALGORITHM The mo dulation compensation algorithm (MCA) estimates a set of lter coeÆcients to maximize the likeliho o d of the ltered observation sequence with resp ect to a given HMM model. The scenario under which the algorithm is applied is illustrated by the blo ck diagram in Figure 1. The underlying assumption in this scenario is that conditions that might aect the MS characteristics of sp eech during recognition may not b e present during HMM mo del training. Discussion of the MCA algorithm is presented in this section in two parts. First, it is introduced as an extension of a class of techniques develop ed by Chengalvarayan and Deng for simultaneous estimation of HMM and observation sequence lter parameters [1 ]. Second, the algorithm is describ ed in detail, along with algorithmic issues relating to the optimization criterion and the form of the lters that are used in the algorithm. DOMAIN SPECIFIC MODULATION SPECTRAL SHAPING Training Speech Test Speech MODULATION COMPENSATION ALGORITHM (MCA) HMM Model λ Estimated Modulation Spectrum FIlter FORWARD− BACKWARD TRAINING Figure 1. Mo dulation comp ensation algorithm (MCA) applied where mo dulation sp ectrum distortions maybe introduced in test conditions. In [1 ], a class of techniques for simultaneous estimation of HMM and observation sequence lter parameters was develop ed. This was used only for obtaining dierence cepstrum parameters from the absolute cepstrum. The comp onents of a D dimensional dynamic cepstrum vector Y t = y 1 ;:::;y D were computed from a static vector X t at time t according to Y t = f X k =  b ! k;i;m X t + k ; 1  t  T (1) where the \mo dulation" lter co eÆcients ! k;i;m are dep endent on state i and Gaussian mixture comp onent m . Equation 1 was introduced into the mo del reestimation equations for continuous Gaussian observation density HMMs and simultaneous reestimation of the lter co ef- cients and the HMM mo del parameters was p erformed. 6th European Conference on Speech Communication and Technology (EUROSPEECH’99) Budapest, Hungary, September 5-9, 1999 ISCAArchive http://www.isca-speech.org/archive Wellekens prop osed a more constrained mo dulation sp ectrum lter reestimation pro cedure aimed at optimizing only the cut{o frequency of the MS lters [7]. Simultaneous estimation of these parameters has the desirable eect of providing a closer coupling between mo del estimation and feature analysis in sp eech recognition. However, b ecause of the coupling to mo del estimation, it is unlikely that the spectral characteristics of the lter parameters estimated in this way will actually reect any meaningful structure asso ciated with the MS of sp eech. The goal in this work is to estimate the MS lter parameters separate from the HMM model facilitating the scenario depicted by the blo ck diagram in Figure 1. Since the mo del remains xed, MS mismatchbetween utterances used to train the mo del and the utterances used for testing can b e reduced. The MCA algorithm and its mathematical properties are describ ed in more detail b elow. The algorithm estimates a lter that, when applied to the absolute cepstrum co eÆcients, increases the likeliho o d of the ltered data, Y ! , with resp ect to the HMM model,  . This in theory could b e accomplished byenumerating an ensemble of MS lters  and choosing the most likely lter in the ensemble, ^ ~! = arg max ~! 2  P ( Y ! j ~!;  ) : (2) However, in practice it is very diÆcult to specify a manageable ensemble of lters that would b e suitable for representing an arbitrary set of mo dulation sp ectrum distortions and it is not clear that a maximum likeliho o d criterion would b e suitable for selecting the optimum MS lter from this ensemble. The algorithm describ ed here assumes a nite impulse response lter, whose length must be chosen, and estimates the parameters of this lter using the exp ectation maximization (EM) algorithm. It can b e applied to ltering the absolute cepstrum features or to obtaining the lter parameters used to compute the dynamic features from the absolute cepstrum. The following discussion will refer sp ecically to the former case. Given an initial HMM mo del,  , and an initial length for the MS lter, ~! , this EM based algorithm maximizes the exp ected value of the log of the ltered data likelihood, P ( Y ! j ~!;  ), with resp ect to ~! . The data to which the lter ~! is applied can be the unltered absolute data, X t , or the absolute data ltered with an initial lter. In the latter case, the overall lter that needs to b e applied to the test data so as to reduce the MS mis{matchbetween the unltered training data and the test data is the convolution of the initial lter and the lter optimized by the MCA, ~! . Unless otherwise sp ecied, no initial lter is used and the MCA is started with unltered data. The p ortion of the optimization equation that is dep endent on the ltered data is given by X i;m;t  t;i;m [ Y t  ~ x;i;m ] T r   1 x;i;m [ Y t  ~ x;i;m ] ; (3) where  t;i;m is the a p osteriori probabilityofo ccupying Gaussian mixture component m and state i at time t . The quantities ~ x;i;m and   1 x;i;m in Equation 3 are the HMM mo del means and variances for the unltered data, X t , and are not reestimated as part of this pro cedure. For the purp oses of this development, Y t in Equation 1 can b e written in vector notation as: Y t = ( X t  b ::: X t + f ) T r ( !  b;i;m :::! f ;i;m )) T r = X t + f t  b ~! i;m : (4) Note that, while Equation 4 demonstrates that it is p ossible to use MS lters ~! i;m , that are HMM state and mixture comp onent dep endent, this is not done here. By substituting the expression for Y t in Equation 3 and differentiating with respect to ~! , which does not dep end on i and m ,we obtain X l;i;m;t  t;i;m  X t + f t  b  T r   1 x;i;m X t + f t  b ~! = X l;i;m;t  t;i;m  X t + f t  b  T r   1 x;i;m ~ x;i;m (5) Finally, ~! can b e obtained by solving the matrix equation given by Equation 5 which requires the mo del means and variances, the a p osteriori probabilities computed in the forward{backward algorithm, and the unltered observations. While the ab ove pro cedure can b e iterated, using the estimated ~! obtained by solving Equation 5 as input to a following iteration, there a numb er of issues that must b e dealt with. The rst issue relates to the fact that it is the ltered data, as opp osed to the original data, that is input to the next iteration of the algorithm. As a result, the lter applied to the data for a given iteration must b e the convolution of all the lters obtained in the preceding iterations. As has b een seen, a convolution is also necessary for the case where the MCA pro cedure is started with data ltered with an initial lter. A second issue relates to artifacts that arise due to lack of energy constraints in the mo del estimation pro cedure. An empirical pro cedure for normalizing lter energies between iterations is describ ed in Section 3.. 3. COMPENSATING FOR MS DOMAIN MISMATCH The MCA algorithm was rst applied to a simulated distortion. The goal of this rst application was to implement a scenario as depicted in the block diagram in Figure 1. The form of the mo dulation sp ectrum mismatch was intended to incorp orate a lo ose approximation to existing MS mo dels of how sp eaking rate variability might b e reected in the mo dulation sp ectrum domain (e.g. [5]). These models characterize the MS of sp eechashaving a p eak at approximately four Hz with some variation in the lo cation of that peak p ossibly resulting from speaking rate dierences. The mo dulation spectrum distortion to ok the form of an FIR lter of length 7 with a p eak at 10 Hz and a strong attenuation of mo dulation frequencies b eyond 20 Hz . This lter was actually applied to the training utterances of the high SNR, noise{free TI digits database [4 ]. So the mo dulation sp ectrum shaping indicated in Figure 1 would actually b e the inverse of that sp ectrum. The HMM mo del,  , was then trained from the ltered data using the forward{ backward algorithm. The 8623 digit strings uttered bya p opulation of adult sp eakers in the training set of the TI digits database was used for training digit mo dels with a mixture of at most 16 continuous Gaussian comp onents p er HMM state. The TI digits test set was split into a 3306 utterance development set, whichwas used as input to the MCA algorithm, and a 5376 utterance test set for evaluating ASR word accuracy. The inputs to MCA are the model and a subset of the unltered development utterances and the output is a lter that reduces the mismatchbetween the utterances and the mo del. Figure 2 shows the mo dulation sp ectra for the training utterances and the unltered development utterances for cepstral co eÆcient c 4 . The lter that was applied to the training utterances is also shown. A signicant mismatch can be observed in the MS domain. The magnitude sp ectrum of the MS lter estimated using Equation 5 was found to provide a reasonably goo d approximation magnitude sp ectrum of the lter used on 0 5 10 15 20 10−1 100Modulation Spectrum for Dimension 4 dB Hz train test filter Figure 2. Mo dulation Sp ectra for the Training Utterances and for the Unltered Development Utterances, along with the Spectrum of the Filter Applied to the Training Utterances the training utterances. However, since there is no inherent constraint on the cepstrum energy levels in the MCA, signicant mismatch in the energy levels of the ltered cepstra and the original cepstra can occur. To deal with this, the lter co eÆcients are scaled with a dimension specic scale factor so that average energy of the training and adaptation utterances are normalized to the same level. The estimated lter, scaled separately for each comp onent of the input cepstrum vector, is then applied to the test utterances. Table 1 displays the recognition p erformance after the MS lter determined by the MCA was applied to the test utterances. Performance is presented as a function of the numb er of development or adaptation utterances that were used by the MCA. The algorithm was implemented here in a sup ervised mo de with the utterance transcriptions made available to the MCA. The Table shows that when the whole development set is used to determine the b est lter with which to reduce the mismatch between train and test utterances, word accuracy approaching the matched condition can be obtained. When less data is available (from 60 utterances through just one), p erformance degrades gracefully with resp ect to the whole development set and still alleviates most of the recognition rate reduction due to the MS dierences b etween training and testing. Results are given in all cases for the lter output by the rst iteration of MCA, which has length 3. Surprisingly, lters with other lengths, resulting either from running several iterations with a lter length equal to 3, or from cho osing another lter length (e.g. 5,7or9) also give goo d p erformance but short of the results when the length is 3. A length of 3 seems to strike the right balance b etween the numb er of degrees of freedom and a constrained estimation. Mo dulation Sp ectrum Adaptation %Word Comp ensation Utterances Accuracy Matched - 96.30 Mismatched - 15.04 MCA 3306 92.10 MCA 10 88.78 MCA 1 87.93 Table 1. Results for MCA with a simulated distortion MCA seems to work well for the experiments that were done with a simulated distortion. MCA was applied to estimate the lter applied to the absolute co eÆcients, no dynamic features were used. The next section describes an exp eriment where MCA was used to optimize the lter used to obtain the delta coeÆcients. 4. MCA AND SPEAKING RATE In order to test MCA in a more realistic environment, a database containing sp eech elicited at multiple speaking rates was used [6]. This database, referred to as DB1, was recorded in an anechoic ro om with a high quality microphone. It includes sp eech from 22 female, and 24 male talkers, each of which recorded 120 sentences. Both normal rate and fast rate utterances of each sentence was elicited from each speaker. Fast rate sp eechwas elicited from each sp eaker by asking sp eakers to sp eak sentences as rapidly as p ossible without gross mispronounciations. For each sp eaking rate, cross-word triphones backed o to monophones were trained from the sentences from 38 sp eakers. The remainder were reserved for testing. Each system had 4903 states and over 28000 Gaussians. In recognition, each phone may be recognized via a triphone or the corresp onding monophone at any place within a word. All 120 sentences for all the test sp eakers were recognized. Figure 3 describ es the exp eriment that was carried out. Modulation Spectrum Filter ω Fast Speaking Rate: Development Utterances λ S Speaker Dependent Model Normal Speaking Rate: MODULATION COMPENSATION ALGORITHM Fast Speaking Rate: Test Utterances ASR FILTER TIME SEQUENCE OF CEPSTRUM PARAMETERS Figure 3. MCA and sp eaking rate The goal was to study how well MCA can cop e with sp eaking rate mismatches and up to what p oint these can b e represented in the mo dulation sp ectrum. The recognition results for all 8 test speakers were analysed. The utterances from a single male sp eaker where more dramatic dierences b etween fast and normal rate sp eaking styles were observed were selected for the exp erimental study. Since MCA might improveword accuracy by p erforming sp eaker adaptation, a sp eaker adapted mo del was obtained. To do this, the gaussians in the normal-rate mo del were clustered into 4 regression classes. On the rst 60 normal-rate sentences, one Maximum Likeliho o d Linear Regression full matrix transformation (MLLR) was learned for each regression class [3]. These were applied to the means of the normal-rate normal-rate mo del, which results in a sp eaker adapted normal-rate mo del. This mo del was then input to the MCA along with the rst 60 fast sentences, regarded as development or adaptation sentences. The lter output by MCA was applied to the 60 last fast utterances. In this case, MCA was used to optimize the lter used to obtain the delta features (no other dynamic features were used apart from them). Another dierence with the previous experiment is that the lter used to obtain the delta features for the training utterances was used as the initial lter for MCA. The length of the lter output by each iteration was the initial model lter length, i.e. 5, plus 2 for each iteration. Results for this exp eriment can be found in Table 2. Test Condition %Word Acc. Sp eaker Indep. Normal Rate 57.28 Sp eaker Adapt. (SA) Normal Rate 80.67 Baseline SA Fast Rate 61.58 \Po oled" MCA Fast Rate 61.10 \Dim. Specic" MCA Fast Rate 63.96 Table 2. Results for MCA with dierent sp eaking rates Two mo des for MCA were investigated. The rst, referred to as \p o oled MCA" in Table 2, estimates the MCA lter co eÆcients according to the pro cedure outlined in Section 2. The second, termed \dimension specic MCA", estimates a separate set of MS lter co eÆcients for each dimension of the cepstrum observation vector. This represents an extension to the MS lter describ ed in Equation 1 and indep endently estimates lter parameters to optimize dimension specic ML criterion. It is clear from Table 2 that the MCA yields only a mo dest improvementin WAC over the baseline system on the fast rate for this task. 0 5 10 15 20 10−2 10−1 100Modulation Spectrum for Dimension 2 dB Hz fast normal Figure 4. Measured MS from fast and normal sp eech utterances In order to obtain some persp ective for the degree to which mo dulation sp ectrum based techniques may affect performance for sp eaking rate mismatch, the average mo dulation spectrum was measured for utterances with dierent sp eaking rates. For each cepstrum comp onent, the magnitude of the averaged modulation sp ectrum was computed over 60 sentences for normal and fast rate sp eech. Figure 4 displays the magnitude spectra for one of the comp onents of the acoustical vector. It is clear from the curves that the fast rate speech has only slightly more energy in the higher mo dulation frequencies than the normal rate sp eech. This suggests that one might exp ect only small changes in p erformance using MS domain techniques for this task, which supp orts the fairly minor improvements that were observed here. 5. CONCLUSION The mo dulation comp ensation algorithm was presented as a maximum likeliho o d technique for estimating lters for the time sequence of spectral parameters in order to reduce mismatch in the mo dulation sp ectrum. The pro cedure was describ ed as an extension to a class of techniques develop ed in [1]. It can be applied to estimating MS lters for the absolute cepstrum co eÆcients or for obtaining the dynamic cepstrum co eÆcients from the absolute coeÆcients. Two applications for the MCA were describ ed. The rst application was for a task where a simulated mo dulation sp ectrum domain distortion was introduced during testing. It was shown that the MCA can eectively reduce mismatch in the MS b etween training and testing utterances. In a second application, the MCA was applied to reducing mismatch attributable to dierences in sp eaking rate. The MCA was shown to have only a small impact on p erformance for this task. Further work is directed towards application of the algorithm to other sources of variability where distortions in the mo dulation sp ectrum are more pronounced. REFERENCES [1] Rathinavelu Chengalvarayan and Li Deng. \Use of Generalized Dynamic Feature Parameters for Sp eech Recognition". IEEE Trans. on Speech and Audio Processing , 5(3):232{242, May 1997. [2] T. Houtgast and H.J.M. Steeneken. \A review of the MTF concept in ro om acoustics and its use for estimating speech intelligibility in auditoria". JASA , 77(3):1069{1077, March 1985. [3] C.J. Leggetter and P.C. Wo o dland. \Maximum likeliho o d linear regression for speaker adaptation of continuous density hidden Markov mo dels". Computer Speech and Language , 9:171{185, 1995. [4] R.G. Leonard. \A Database for Sp eaker-Independent Digit Recognition". In Proc. ICASSP'84 , pages 42.11.1{4, 1984. [5] Climent Nadeu, Pau Paches-Leal, and Biing-Hwang Juang. \Filtering the time sequences of sp ectral parameters for sp eech recognition". Speech Communication , 22:315{332, 1997. [6] Juergen Schro eter. \A High-Quality Sp eech Database". Technical rep ort, AT&T Bell Lab oratories, Murray Hill, NJ, USA, 1992. [7] Chris J. Wellekens. \Enhanced ASR by Acoustic Feature Filtering". In Proc. ICSLP'98 , pages 2995{2998, Sydney, Australia, 1998.