scieee Science in your language
[en] (orig)

Linear prediction of the one-sided autocorrelation sequence for noisy speech recognition

Abstract

The article presents a robust representation of speech based on AR modeling of the causal part of the autocorrelation sequence. In noisy speech recognition, this new representation achieves better results than several other related techniques.

Read accessible full text

Linear prediction of the one-sided autocorrelation sequence for noisy speech recognition

Author: Hernando Pericás, Francisco Javier,Nadeu Camprubí, Climent
Year: 1997
DOI: 10.1109/89.554273
Source: https://upcommons.upc.edu/bitstream/2117/88694/1/Linear%20prediction%20of%20the%20one-sided%20autocorrelation%20sequence%20for%20noisy%20speech%20recognition.pdf
80 IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997
small. Al hough inc eased aining da a may be help ul, in he speake
dependen Manda in syllable ecogni ion p oblem, a limi ed da abase
will p obably s ill be a no mal si ua ion o some pe iod o ime in
he u u e.
VII. CONCLUSION
A new app oach is p oposed in his co espondence o ob ain
mo e elabo a e ini ial models co e ing cha ac e is ics o di e en
ones. Imp o ed s a e ansi ion opologies a e also ound o achie e
be e pe o mance compa ed wi h he simple le - o- igh model wi h
wo ansi ions. A h eshold decision app oach is u he de eloped
o imp o e he pe o mance o BS ecogni ion o he syllables
wi h he neu al one. The es esul s on e e yday Chinese show
ha a o al e o a e educ ion on he o de o 20% in he op 1
a e can be ob ained when all he concep s a e p ope ly in eg a ed.
Al hough he echniques he e a e p oposed specially o ecogni ion
o Manda in base syllables conside ing he e ec o ones, i is
ce ainly belie ed ha simila concep s a e po en ially applicable o
sol e simila p oblems in speech ecogni ion in o he languages.
REFERENCES
[1] R. He, Ed., Guoyu bao Tzdian (Manda in Chinese Daily Dic iona y).
Taipei, R.O.C.: Guoyu bao, 1976.
[2] L.-S. Lee, C.-Y. Tseng, and M. Ouh-Young, “The syn hesis ules in
a Chinese ex - o speech sys em,” IEEE T ans. Acous ., Speech, Signal
P ocessing, pp. 1309–1320, Sep . 1989.
[3] Y. R. Chao, A G amma o Spoken Chinese. Be keley, CA: Uni e si y
o Cali o nia Be keley P ess, 1968.
[4] W. J. Yang, J. C. Lee, Y. C. Chang, and H. C. Wang, “Hidden Ma ko
model o manda in lexical one ecogni ion,” IEEE T ans. Acous .,
Speech, Signal P ocessing, pp. 988–992, July 1988.
[5] F.-H. Liu, Y. Lee, and L. S. Lee, “A Di ec -conca ena ion app oach o
ain hidden ma ko models o ecognize he highly con using manda in
syllables wi h e y limi ed aining da a,” IEEE T ans. Speech Audio
P ocessing, ol. 1, no. 1, pp. 113–119, Jan. 1993.
[6] L.-S. Lee e al., “Golden manda in (I)—A eal- ime manda in speech
dic a ion machine o chinese language wi h e y la ge ocabula y,”
IEEE T ans. Speech Audio P ocessing, ol. 1, no. 2, pp. 158–179, Ap .
1993.
[7] B.-H. Juang and L. R. Rabine , “Mix u e au o eg essi e hidden Ma ko
models o speech signals,” IEEE T ans. Acous ., Speech, Signal P o-
cessing, ol. ASSP-33, no. 6, pp. 1404–1413, Dec. 1985.
[8] X. Huang e al., “The SPHNIX-II speech ecogni ion sys em: An
o e iew,” Compu . Speech Language, pp. 137–148, Feb. 1993.
Linea P edic ion o he One-Sided Au oco ela ion
Sequence o Noisy Speech Recogni ion
Ja ie He nando and Climen Nadeu
Abs ac — The aim o his co espondence is o p esen a obus
ep esen a ion o speech based on AR modeling o he causal pa o
he au oco ela ion sequence. In noisy speech ecogni ion, his new ep-
esen a ion achie es be e esul s han se e al o he ela ed echniques.
I. INTRODUCTION
Linea p edic i e coding (LPC) [1] is a spec al es ima ion ech-
nique widely used in speech p ocessing and, pa icula ly, in speech
ecogni ion. Howe e , he con en ional LPC echnique, which is
equi alen o AR modeling o he signal
x
(
n
)
, is known o be
e y sensi i e o he p esence o backg ound noise. This ac leads
o poo ecogni ion a es when his echnique is used in speech
ecogni ion unde noisy condi ions, e en i only a mode a e le el
o con amina ion is p esen in he speech signal. Simila esul s
a e ob ained wi h he well-known mel-ceps um echnique [2]. This
explains why some o he main a emp s o comba he noise
p oblem consis o inding no el acous ic ep esen a ions ha a e
mo e esis an o noise co up ion han adi ional pa ame e iza ion
echniques.
Linea p edic ion o he au oco ela ion sequence has been he
common app oach o se e al obus spec al es ima ion me hods o
noisy signals p esen ed in he pas . Fo speech ecogni ion, Mansou
and Juang [3] p oposed he sho - ime modi ied cohe ence (SMC) as a
obus ep esen a ion o speech based on ha app oach. On he o he
hand, Cadzow [4] in oduced he use o an o e de e mined se o
Yule–Walke equa ions o obus modeling o ime se ies. Al hough
Cadzow applies linea p edic ion o he signal, his me hod can also
be in e p e ed as pe o ming linea p edic ion in he au oco ela ion
domain. Bo h me hods ely, ei he explici ly o implici ly, on he ac
ha he au oco ela ion sequence is less a ec ed by b oadband noise
han he signal i sel , especially a high lag indices.
In his wo k, we conside he one-sided o causal pa o he
au oco ela ion sequence and i s ma hema ical p ope ies. As his
sequence sha es i s poles wi h he signal
x
(
n
)
, i p o ides a good
s a ing poin o LPC modeling. In his way, he new one-sided
au oco ela ion LPC (OSALPC) me hod appea s as a s aigh o wa d
esul o he app oach [5]. In addi ion, i is closely ela ed o he
SMC ep esen a ion and Cadzow’s me hod. All o hem can be
in e p e ed as AR modeling o ei he a spec al unc ion named
“en elope” o i s squa e. This in e p e a ion, which is based on he
p ope ies o he one-sided au oco ela ion, p o ides mo e insigh in o
he a ious me hods. In his co espondence, hei pe o mance in
noisy speech ecogni ion is compa ed. The op imum model o de
and ceps al li e ing ha e also been in es iga ed in noisy condi ions.
The simula ion esul s show ha OSALPC ou pe o ms he o he
echniques in se e e noisy condi ions and ob ains simila sco es o
mode a e o high SNR.
Manusc ip ecei ed Feb ua y 14, 1995; e ised No embe 9, 1995. This
wo k was suppo ed by G an nos. TIC-92-0800-C05/04 and TIC-92-1026-
C02/02. The associa e edi o coo dina ing he e iew o his pape and
app o ing i o publica ion was D . Kuldip K. Paliwal.
The au ho s a e wi h he Depa men o Signal Theo y and Communica ions,
Poly echnical Uni e si y o Ca alonia, Ba celona, Spain.
Publishe I em Iden i ie S 1063-6676(97)00766-9.
1063–6676/97$10.00 1997 IEEE
IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997 81
This co espondence is o ganized in he ollowing way. In Sec ion
II, he OSALPC echnique is in oduced, and i s ela ionship wi h he
con en ional LPC app oach and he o he pa ame e iza ions based
on AR modeling in he au oco ela ion domain is discussed. Sec ion
III epo s he applica ion o all hose pa ame e iza ion echniques
o an isola ed wo d mul ispeake ecogni ion ask, using he HMM
app oach, in o de o compa e hei pe o mances in he p esence o
addi i e whi e noise. Finally, some conclusions a e summa ized in
Sec ion IV.
II. AR MODELING IN THE AUTOCORRELATION DOMAIN
F om he au oco ela ion sequence
R
(
m
)
, we de ine he one-sided
(causal pa o he) au oco ela ion (OSA) sequence in he ollowing
way:
R
+
(
m
)=
R
(
m
)
m>
0
R
(0)
2
m
=0
0
m<
0
(1)
I s Fou ie ans o m is he complex “spec um”
S
+
(
!
)=
1
2
[
S
(
!
)+
jS
H
(
!
)]
(2)
whe e
S
(
!
)
is he eal spec um, i.e., he Fou ie ans o m o
R
(
m
)
,
and
S
H
(
!
)
is he Hilbe ans o m o
S
(
!
)
.
Due o he analogy be ween
S
+
(
!
)
in (2) and he analy ic signal
used in ampli ude modula ion, a spec al “en elope”
E
(
!
)
[6] can
be de ined as
E
(
!
)=
j
S
+
(
!
)
j
:
(3)
Due o he la ge dynamic ange o speech spec a, he en elope
E
(
!
)
s ongly enhances he highes powe equency bands wi h
espec o
S
(
!
)
[5]. Consequen ly, he noise componen s lying
ou side he enhanced equency bands a e la gely a enua ed in
E
(
!
)
wi h espec o
S
(
!
)
, and hus,
E
(
!
)
is mo e obus o b oadband
noise han
S
(
!
)
. On he o he hand, as i is well known, he OSA
sequence
R
+
(
m
)
and he signal
x
(
n
)
ha e he same poles [7].
Those wo p ope ies, i.e., obus ness o noise and pole p ese a-
ion, sugges ha AR pa ame e s o he speech signal can be mo e
eliably es ima ed om he OSA sequence
R
+
(
m
)
han di ec ly om
he signal
x
(
n
)
when
x
(
n
)
is co up ed by b oadband noise. Thus,
as he con en ional LPC echnique assumes an all-pole model o he
speech spec um
S
(
!
)
, we may apply linea p edic ion o he OSA
sequence, assuming an all-pole model o i s “spec um”
E
2
(
!
)
. This
is he basis o he one-sided au oco ela ion linea p edic i e coding
(OSALPC) pa ame e iza ion echnique [5].
A s aigh o wa d algo i hm is p oposed in [5] ha calcula es he
OSALPC ceps al coe icien s. I consis s o applying he (windowed)
au oco ela ion me hod o linea p edic ion o an es ima ion o he
OSA sequence:
a) Fi s , om he speech ame o leng h
N
, he au oco ela ion
lags un il
M
=
N=
2
a e compu ed ( his alue o
M
was
empi ically op imized o conside he well-known adeo
be ween a iance and equency esolu ion o he spec al
es ima e [8]).
b) Second, he Hamming window om
m
=0
o
M
is applied
on such es ima ed OSA sequence.
c) Thi d, i
p
is he p edic ion o de , he i s
p
+1
au oco ela ion
alues o ha OSA sequence a e compu ed om
m
=0
o
p
,
using he con en ional biased es ima o , i.e., he one ha is
commonly employed in speech p ocessing.
d) Then, hese alues a e used as en ies o he Le inson–Du bin
algo i hm o es ima e he AR pa ame e s
ak
,
k
=1
;
111
;p:
Fig. 1. Robus ness o he OSALPC ep esen a ion o addi i e whi e noise:
(a) LPC spec um and (b) OSALPC squa ed en elope o a oiced speech ame
in noise ee condi ions (solid line) and SNR equal o 0 dB (do ed line).
e) Finally, he ceps al coe icien s co esponding o he model a e
ecu en ly compu ed om hose AR pa ame e s.
The obus ness o OSALPC o addi i e whi e noise is illus a ed in
Fig. 1. As can be seen in his igu e, he OSALPC squa ed en elope
shows a p ominen i s o man , and i s whole cu e is mo e obus
o addi i e whi e noise han ha o he LPC spec um. In his case, he
con en ional biased au oco ela ion es ima o was used o compu e
he OSA sequence om he signal.
Fig. 1 also shows ha spu ious peaks may appea in he OSALPC
squa e en elope. They a e p obably due o he ac ha he OSALPC
echnique pe o ms only a pa ial decon olu ion o he speech signal
[9]. In spi e o ha , OSALPC shows a be e speech ecogni ion
pe o mance han con en ional LPC in se e e condi ions o addi i e
whi e noise, as will be seen in he nex sec ion.
The OSALPC echnique is closely ela ed o he sho - ime modi-
ied cohe ence (SMC) ep esen a ion p oposed by Mansou and Juang
in [3]. SMC is also based on AR modeling in he au oco ela ion
domain. Howe e , whe eas in he OSALPC echnique, he en ies o
he Le inson–Du bin algo i hm ( i s
p
alues o he au oco ela ion
o he OSA sequence) a e calcula ed om he OSA sequence using he
con en ional biased au oco ela ion es ima o , in he SMC ep esen-
a ion, hey a e compu ed using a squa e oo spec al shape . In ac ,
in e ms o he abo e o mula ion, ha di e ence lies in assuming
in he SMC echnique an all-pole spec al model o he en elope
E
(
!
)
ins ead o
E
2
(
!
)
. Fu he mo e,
R
+
(0)
is se o 0 in he case
o addi i e whi e noise because i is se e ely co up ed by noise.
On he o he hand, he name o he SMC ep esen a ion de i es
om he usage o a pa icula es ima o , which is e e ed o as
cohe ence in [3], o compu e he OSA sequence om he signal.
This es ima o is a mo e homogeneous measu e han he con en ional
biased au oco ela ion es ima o in he sense ha e e y es ima ed
alue is compu ed using he same numbe o signal samples, whe eas
in he con en ional es ima o , he numbe o signal samples employed
o es ima e
R
(
m
)
dec eases along he index
m
. Tha p ope y does
no ha e much ele ance in he es ima ion o he au oco ela ion
en ies o he Le inson–Du bin algo i hm since only he i s
p
+1
82 IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997
Fig. 2. Ma ix o mula ion o OSALPC and LSMYE me hods.
alues a e conside ed, and usually,
p

N
. Howe e , i may be
impo an in he es ima ion o he OSA sequence om he speech
signal since he OSA leng h conside ed in bo h OSALPC and SMC
echniques is
M
=
N=
2
and no negligible wi h espec o
N
.
The OSALPC echnique can also be easily ela ed o he o e de-
e mined se o Yule–Walke equa ions p oposed by Cadzow in
[4] o seek ARMA models o ime se ies. As an
AR
(
p
)
p ocess
con amina ed by addi i e whi e noise becomes an ARMA
(
p; p
)
p ocess, Cadzow’s me hod can be used o es ima e he pa ame e s
o his noisy AR p ocess simply by se ing he same AR and MA
o de s in he so-called leas squa es modi ied Yule–Walke equa ions
(LSMYWE’s) [8].
The ela ionship be ween he OSALPC and LSMYWE echniques
is illus a ed by he ma ix equa ion in Fig. 2, whe e
M
deno es
he highes au oco ela ion lag used, and
e
(
m
)
is he e o o
be minimized. The minimiza ion o he no m o he ull e o
ec o
e
(
m
)
g
m
=1
;
111
;M
+
p
wi h espec o he AR pa ame e s
a
k
is equi alen o he applica ion o he (windowed) au oco ela ion
me hod o linea p edic ion o he sequence
R
(
m
)
,
m
=1
;
111
;M
,
i.e., he OSALPC echnique. On he o he hand, he LSMYWE
echnique minimizes he no m o he sub ec o
e
(
m
)
g
m
=
p
+1
;
111
;M
,
and he e o e, i amoun s o applying he (unwindowed) co a iance
me hod o linea p edic ion on he same ange o au oco ela ion lags.
When
M
=2
p
, LSMYWE a e he modi ied Yule–Walke equa ions
[8] o an ARMA
(
p; p
)
p ocess. In bo h cases, only au oco ela ion
lags co esponding o he OSA sequence a e employed.
In ou compa ison, we will also conside ano he e sion o
his (unwindowed) co a iance-based app oach ha will be called
leas squa es Yule–Walke equa ions (LSYWE’s). Whe eas in he
LSMYWE echnique he i s p edic ed au oco ela ion alue is
R
(
p
+
1)
, in he LSYWE echnique, he p edic ion begins a
R
(1)
. Bo h
LSMYWE and LSYWE me hods and hei ela ionship o OSALPC
a e g aphically desc ibed in Fig. 3. As i is shown, he only di e ence
be ween he a ious echniques is he ange o au oco ela ion lags
conside ed in he minimiza ion o he e o . I is wo h no ing ha
LSYWE conside s some nega i e au oco ela ion lags ha do no
belong o he OSA sequence. In pa icula , i
M
is equal o
p
, LSYWE
a e he con en ional Yule–Walke equa ions.
As will be seen in he nex sec ion, in spi e o he simila i y be ween
all hese echniques, he OSALPC ep esen a ion ou pe o ms he
LSYWE, LSMYWE, and SMC echniques in speech ecogni ion in
se e e noisy condi ions. On he o he hand, as a as he compu a ional
complexi y o he algo i hms is conce ned, OSALPC and SMC
echniques a e much mo e e icien han LSYWE and LSMYWE
echniques because hey use he Le inson–Du bin algo i hm.
Finally, i is wo h no ing ha he OSALPC echnique may be
included in he ield o highe o de spec al es ima ion due o he
ac ha he squa ed en elope
E
2
(
!
)
is he Fou ie ans o m o he
au oco ela ion o he OSA sequence, which is a pa icula ou h-
o de momen o he signal.
Fig. 3. In e p e a ion o he (a) OSALPC, (b) LSMYWE, and (c) LSYWE
app oaches as applica ion o he au oco ela ion o co a iance me hods o
linea p edic ion o an au oco ela ion sequence in di e en lag anges.
III. SPEECH RECOGNITION EXPERIMENTS
This sec ion epo s he applica ion o all he abo e pa ame e -
iza ion echniques o ecognize isola ed wo ds in a mul ispeake
ask wi h a disc e e HMM-based sys em in o de o compa e hei
pe o mance and o gain some insigh in o he me i o he OSALPC
ep esen a ion in he p esence o addi i e whi e noise
A. Speech Da abase and Recogni ion Sys em
The da abase used in ou expe imen s consis s o 10 epe i ions o
he Ca alan digi s u e ed by se en male and h ee emale speake s
(1000 wo ds) and eco ded in a quie oom. Fi s , he sys em was
ained wi h hal o he da abase and es ed wi h he o he hal . Then,
he oles o bo h hal es we e changed, and he epo ed esul s we e
ob ained by a e aging hose wo esul s.
The analog speech signal was i s bandpass il e ed o 100–3400
Hz by an an ialiasing il e , sampled a 8 kHz and, 12 bi s quan ized.
The digi ized clean speech was manually endpoin ed o de e mine
he bounda ies o each wo d. The endpoin s ob ained in his way
we e used in all ou expe imen s, including hose in which noise
was added o he signal. Clean speech was used o aining in all he
expe imen s. Noisy speech was simula ed by adding ze o mean whi e
Gaussian noise o he clean signal so ha he SNR o he esul ing
signal becomes
1
(clean), 20, 10, and 0 dB. No p eemphasis was
pe o med.
In he pa ame e iza ion s age o he ecogni ion sys em, he signal
was di ided in o ames o 30 ms a a a e o 15 ms, and each ame
was cha ac e ized by i s ceps al pa ame e s ob ained ei he by he
con en ional LPC me hod o by any o he echniques p esen ed in
he las sec ion. Be o e en e ing he ecogni ion s age, he ceps al
pa ame e s we e ec o quan ized using bo h a codebook o 64
codewo ds and he Euclidean dis ance measu e be ween li e ed
ceps al ec o s. Each digi was cha ac e ized by a le - o- igh
disc e e hidden Ma ko model o 10 s a es wi hou skips. T aining and
es ing we e pe o med using Baum–Welch and Vi e bi algo i hms,
espec i ely.
B. Recogni ion Resul s
Fi s o all, we ca ied ou some expe imen s wi h he abo e
desc ibed speech ecogni ion sys em o op imize he model o de and
IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997 83
TABLE I
RECOGNITION RATES OF THE CONVENTIONAL LPC TECHNIQUE FOR
SEVERAL PREDICTION ORDER VALUES AND CEPSTRAL LIFTERS
TABLE II
RECOGNITION RATES OF THE CONVENTIONAL LPC, LSMYWE,
AND LSYWE TECHNIQUES FOR
p
=12
AND THE SLOPE LIFTER
TABLE III
RECOGNITION RATES OF THE CONVENTIONAL LPC, SMC, AND
OSALPC TECHNIQUES FOR
p
=12
AND THE SLOPE LIFTER
he ype o ceps al li e in he con en ional LPC echnique. In Table
I, he ecogni ion esul s o LPC model o de s
p
=
8, 12, and 16
and o he bandpass [10], in e se o s anda d de ia ion [11] (ISD),
and slope [12] li e s a e p esen ed. The ecogni ion esul s show ha
nei he he model o de no he ype o ceps al li e a e ele an o
ou ask in noise- ee condi ions. Howe e , in he p esence o noise,
he ecogni ion esul s a e e y sensi i e o bo h ac o s.
I is also clea om Table I ha he nonsymme ical li e s—slope
and ISD—ou pe o m he bandpass li e o e e y model o de . This
may be due o he ac ha in he p esence o whi e noise, he lowe
o de ceps al coe icien s a e mo e a ec ed han he highe o de
ones in he unca ed ceps al ec o .
The bes esul s o se e e noisy condi ions—10 and 0 dB o
SNR—a e ob ained using slope li e and p edic ion o de
p
equal
o 12. The con enience o his ela i ely high o de comes om he
ac ha he sensi i i y o he au oco ela ion sequence o addi i e
whi e noise ends o dec ease along he lag index. Model o de s
ha a e oo high, howe e , yield poo ecogni ion esul s since he
spec al es ima e shows spu ious peaks. Ac ually, ecogni ion a es
we e calcula ed using he slope li e o a la ge ange o alues o
he model o de , and he bes esul s we e hose ob ained o
p
=
12.
In Table II, he ecogni ion a es o con en ional LPC, LSMYWE,
and LSYWE app oaches a e p esen ed, using
M
=
N=
2
and bo h
op imum model o de and li e ob ained o he con en ional LPC
echnique, i.e.,
p
=
12 and he slope li e . Ob iously, hese a e no
he op imum condi ions o each pa ame e iza ion echnique, bu he
esul s can help o compa e hei pe o mance. As can be seen om
Table II, he con en ional LPC echnique ou pe o ms no iceably he
o he app oaches. Howe e , he excellen pe o mance o he LSYWE
app oach in noise- ee condi ions is wo h no ing.
Fig. 4. Compa ison o ecogni ion a es o he LPC, SMC, OSALPC-I and
OSALPC-II echniques.
Fig. 5. Block diag am o he calcula ion o he LPC, SMC, OSALPC-I and
OSALPC-II ceps a.
In Table III and Fig. 4, he ecogni ion a es co esponding o
he con en ional LPC echnique, he SMC ep esen a ion, and he
no el OSALPC app oach a e p esen ed, whe e we also use
M
=
N=
2
,
p
=12
, and he slope li e . The wo e sions OSALPC-I
and OSALPC-II o he OSALPC app oach co espond o he OSA
es ima o s o which we e e ed in Sec ion II: OSALPC-I uses he
con en ional biased au oco ela ion es ima o , and OSALPC-II like
SMC uses he cohe ence es ima o (and se s
R
(0)
o 0). Fig. 5 shows
84 IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, VOL. 5, NO. 1, JANUARY 1997
TABLE IV
RECOGNITION RATES FOR THE OSALPC-II TECHNIQUE FOR
SEVERAL PREDICTION ORDER VALUES AND CEPSTRAL LIFTERS
a block diag am o he calcula ion o he LPC, SMC, OSALPC-I,
and OSALPC-II ceps a ha pe mi s compa ison o hei espec i e
algo i hms.
The OSALPC and SMC ep esen a ions clea ly ou do he con-
en ional LPC echnique in se e e noisy condi ions: OSALPC-I and
OSALPC-II a es a e be e han LPC ones a 10 and 0 dB, and SMC
ou pe o ms LPC a 0 dB. Mo eo e , OSALPC-I and OSALPC-II
ep esen a ions ou pe o m he SMC echnique in all noisy condi ions.
Fo he OSALPC ep esen a ion, he use o he con en ional biased
au oco ela ion es ima o o compu ing he OSA sequence ( e sion
OSALPC-I) is con enien in se e e noisy condi ions, i.e., o an SNR
o 10 o 0 dB.
Howe e , in noise- ee condi ions, he e is a loss o ecogni ion
pe o mance in he OSALPC and SMC app oaches wi h espec o
he con en ional LPC echnique due o he impe ec decon olu ion o
he speech signal pe o med by hose echniques. This e ec seems
o be minimized by using he cohe ence es ima o o compu e he
OSA sequence, as in he case o OSALPC-II and SMC.
Finally, Table IV shows he ecogni ion a es co esponding o
OSALPC-II o he same model o de s and ceps al li e s as in
Table I. I can be no iced ha he new echnique is less sensi i e
o changes in bo h he model o de and he ype o ceps al li e
han he con en ional LPC app oach, p o ided ha he model o de
is no oo low.
IV. CONCLUSIONS
In his co espondence, se e al LPC-based echniques ha wo k
in he au oco ela ion domain a e p esen ed and compa ed in noisy
speech ecogni ion. The OSALPC echnique, which is based on
he applica ion o he (windowed) au oco ela ion me hod o linea
p edic ion o he one-sided au oco ela ion sequence, yields he bes
esul s among all he compa ed LPC-based echniques in se e e noisy
condi ions.
REFERENCES
[1] F. I aku a, “Minimum p edic ion esidual p inciple applied o speech
ecogni ion,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol.
ASSP-23, pp. 67–72, 1975.
[2] S. B. Da is and P. Me mels ein, “Compa ison o pa ame ic ep e-
sen a ions o monosyllabic wo d ecogni ion in con inuously spoken
sen ences,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol. ASSP-
28, pp. 357–366, 1980.
[3] D. Mansou and B. H. Juang, “The sho - ime modi ied cohe ence
ep esen a ion and i s applica ion o noisy speech ecogni ion,” IEEE
T ans. Acous ., Speech, Signal P ocessing, ol. 37, pp. 795–804, 1989.
[4] J. A. Cadzow, “Spec al es ima ion: An o e de e mined a ional model
equa ion app oach,” P oc. IEEE, ol. 70, pp. 907–939, 1982.
[5] J. He nando and C. Nadeu, “Speech ecogni ion in noisy ca en i on-
men based on OSALPC ep esen a ion and obus simila i y measu ing
echniques,” in P oc. ICASSP’94, Adelaide, Ap . 1994, pp. 69–72.
[6] M. A. Lagunas and M. Amengual, “Non-linea spec al es ima ion,” in
P oc. ICASSP’87, Dallas, Ap . 1987, pp. 2035–2038.
[7] D. P. McGinn and D. H. Johnson, “Reduc ion o all-pole pa ame e
es ima ion bias by successi e au oco ela ion,” in P oc. ICASSP’83,
Bos on, Ap . 1983, pp. 1088–1091.
[8] S. L. Ma ple, J ., Ed., Digi al Spec al Analysis wi h Applica ions.
Englewood Cli s, NJ: P en ice-Hall, 1987.
[9] C. Nadeu, J. Pascual, and J. He nando, “Pi ch de e mina ion using
he ceps um o he one-sided au oco ela ion sequence,” in P oc.
ICASSP’91, To on o, Canada, May 1991, pp. 3677–3680.
[10] B. H. Juang, L. R. Rabine , and J. G. Wilpon, “On he use o band-pass
li e ing in speech ecogni ion,” IEEE T ans. Acous ., Speech, Signal
P ocessing, ol. ASSP-35, pp. 947–954, 1987.
[11] Y. Tohku a, “A weigh ed ceps al dis ance measu e o speech ecog-
ni ion,” IEEE T ans. Acous ., Speech, Signal P ocessing, ol. ASSP-35,
pp. 1414–1422, 1987.
[12] B. A. Hanson and H. Waki a, “Spec al slope dis ance measu es wi h
linea p edic ion analysis o wo d ecogni ion in noise,” IEEE T ans.
Acous ., Speech, Signal P ocessing, ol. ASSP-35, pp. 968–973, 1987.
A Fas Algo i hm o Finding he Adap i e Componen
Weigh ed Ceps um o Speake Recogni ion
Mihailo S. Zilo ic, Ra i P. Ramachand an, and Richa d J. Mammone
Abs ac — In speake ecogni ion sys ems, he adap i e componen
weigh ed (ACW) ceps um has been shown o be mo e obus han he
con en ional linea p edic i e (LP) ceps um. The ACW ceps um is
de i ed om a pole-ze o ans e unc ion whose denomina o is he
p
h-o de LP polynomial
A
(
z
). The nume a o is a (
p
0
1
) h-o de
polynomial ha is up o now ound as ollows. The oo s o
A
(
z
)a e
compu ed, and he co esponding esidues ob ained by a pa ial ac ion
expansion o
1
=A
(
z
) a e se o uni y. The e o e, he nume a o is he
sum o all he (
p
0
1
) h-o de co ac o s o
A
(
z
). In his co espondence,
we show ha he nume a o polynomial is me ely he de i a i e o he
denomina o polynomial
A
(
z
). This g ea ly speeds up he compu a ion o
he nume a o polynomial coe icien s since i in ol es a simple scaling
o he denomina o polynomial coe icien s. Roo inding is comple ely
elimina ed. Since he denomina o is gua an eed o be minimum phase
and he nume a o can be p o en o be minimum phase, wo sepa a e
ecu sions in ol ing he polynomial coe icien s es ablishes he ACW cep-
s um. This new me hod, which a oids oo inding, educes he compu e
ime signi ican ly and imposes negligible o e head when compa ed wi h
he app oach o inding he LP ceps um.
I. INTRODUCTION
Speake ecogni ion is he ask o iden i ying a speake by his o he
oice [1]. A common p oblem in ealizing obus speake ecogni ion
sys ems is ha a misma ch in aining and es ing condi ions se iously
deg ades he pe o mance [2]. One o he pu sued app oaches o
Manusc ip ecei ed Feb ua y 3, 1995; e ised July 13, 1996. The associa e
edi o coo dina ing he e iew o his pape and app o ing i o publica ion
was D . Joseph Campbell.
M. S. Zilo ic is wi h Bell Communica ions Resea ch, Red Bank, NJ USA.
R. P. Ramachand an and R. J. Mammone a e wi h he CAIP Cen e ,
Depa men o Elec ical Enginee ing, Ru ge s Uni e si y, Pisca away, NJ
08855 USA.
Publishe I em Iden i ie S 1063-6676(97)00762-1.
1063–6676/97$10.00 1997 IEEE