scieee Science in your language
[en] (orig)

Integranting prosodic information into a speech recogniser

Abstract

In the last decade there has been an increasing tendency to incorporate language engineering strategies into speech technology. This technique combines linguistic and mathematical information in different applications: machine translation, natural language processing, speech synthesis and automatic speech recognition (ASR). In the field of speech synthesis, this hybrid approach (linguistic and mathematical/statistical) has led to the design of efficient models for reproducing the acoustic features of natural language. However, the incorporation of language engineering strategies into ASR is only beginning. In this paper, we present a theoretical framework for the integration of linguistic information into an ASR system. The objective is to design a model which can detect the suprasegmental features of the speech input, mainly those related to the fundamental frequency (F0) that can clarify the functionality of pauses, intonation contour, and interruptions. This specification model has been designed in the framework of a dialogue system

Read accessible full text

Integranting prosodic information into a speech recogniser

Author: López Soto, María Teresa
Year: 2001
Source: https://idus.us.es/bitstreams/fde76b4d-3838-4243-bbbb-84ea3e7d117a/download
ELIA 2, 2001
95
INTEGRATING PROSODIC INFORMATION INTO A SPEECH
RECOGNISER
Te esa López. So o
Uni e sidad de Se illa
In he las decade he e has been an inc easing endency o inco po a e
language enginee ing s a egies in o speech echnology. This echnique
combines linguis ic and ma hema ical in o ma ion in di e en applica ions:
machine ansla ion, na u al language p ocessing, speech syn hesis and
au oma ic speech ecogni ion (ASR). In he ield o speech syn hesis, his
hyb id app oach (linguis ic and ma hema ical/s a is ical) has led o he
design o e icien models o ep oducing he acous ic ea u es o na u al
language. Howe e , he inco po a ion o language enginee ing s a egies
in o ASR is only beginning. In his pape , we p esen a heo e ical
amewo k o he in eg a ion o linguis ic in o ma ion in o an ASR sys em.
The objec i e is o design a model which can de ec he sup asegmen al
ea u es o he speech inpu , mainly hose ela ed o he undamen al
equency (F0) ha can cla i y he unc ionali y o pauses, in ona ion
con ou , and in e up ions. This speci ica ion model has been designed in
he amewo k o a dialogue sys em.
1. In oduc ion
ASR sys ems gene a e a speech hypo hesis which shows an n
simila i y wi h he speech inpu . These ASR sys ems can be used in a ious
applica ions. ASR is p esen in sys ems whe e he objec i e is o iden i y he
indi idual who has u e ed ha piece o spoken language. In hese sys ems
he speech ecognise has o iden i y he gene al acous ic ea u es associa ed
wi h he speech inpu . O he applica ions in eg a e ASR in o a na u al
language p ocessing sys em, whe e he main objec i e is he ex ac ion o
he seman ic s uc u e. In hese la e applica ions, he speech ecognise has
o gene a e a li e al ansc ip ion o he speech inpu .
The ansc ip ion o he wo ds u e ed by a speake is no an easy
ask. Many di e en ac o s a ec he ecogni ion o he exac sequence:
some a e in insic o he speech signal and o he s can be said o be ex insic.
T. López So o
96
Among he so-called in insic ac o s we can men ion some o a
phonological na u e, such as junc u e and in ona ional unc ion ha ha e an
acous ic co ela e in he o m o phonemic ansi ion and F0 analysis. O he
in insic ac o s a e linguis ic: sen ence di ision, pauses, epe i ions,
anapho ic cons uc ions, e c. These in insic ac o s cha ac e ize he speech
signal, bu a e se iously dis o ed by o he “ex e nal” ac o s, elemen s which
go beyond he na u e o he speech signal i sel and which de i e di ec ly
om en i onmen al o ces. Among hese en i onmen al ac o s we can
men ion wo big g oups: en i onmen al ac o s which di ec ly a ec he
spoken si ua ion (echo, backg ound noise, coughs, low oice, e c.) and
elec onical ac o s which a ec he sys em pe o mance: ansduce
(mic ophone), obus ness, e c.
The sum o all hese ac o s makes i di icul o cons uc an
e icien model o speech ecogni ion. E en in e y obus and powe ul
ASR sys ems he e o a e may unexpec edly p e en an accu a e and
eliable ex ac ion o he meaning o he speech inpu .
To cope wi h ecogni ion e o s di e en echniques can be
in eg a ed in o he NLP module (Heeman, 1998; López-So o, 1999).
Howe e , his app oach does no eally imp o e he ASR pe o mance and
can be conside ed o be a me e ad hoc solu ion. A di e en app oach
consis s in he inco po a ion o pa sing echniques in o he ASR sys em
(wo d la ice pa sing, Van Noo d, 1998). This echnique consis s o he
applica ion o syn ac ic and mo phological in o ma ion o cons uc a wo d
hypo hesis and complemen s o he acous ic and s a is ical in o ma ion.
I seems ha he inco po a ion o language enginee ing echniques is
necessa y o imp o e ASR pe o mance, bu i is no he only solu ion:
acous ic and s a is ical in o ma ion is de ini i ely impo an o cope wi h
ac o s ha dis o he signal (backg ound noise, channel dis o ions, e c.).
Besides, he use o lexical and syn ac ic in o ma ion only is no enough o
deal wi h a wide ange o discou sal and con ex ual in o ma ion implici in
any communica i e si ua ion. We belie e ha pa o his in o ma ion can be
easily p ocessed i we ake in o accoun p osodic in o ma ion. In sec ion 2
ELIA 2, 2001
97
we desc ibe some gene al cha ac e is ics o spoken language paying special
a en ion o hose ea u es ha can be o malized using p osodic in o ma ion.
In sec ion 3 we p esen ou speci ica ion model o he desc ip ion o
p osodic in o ma ion in o an ASR sys em.
2. Some ea u es o spoken language
In his sec ion we analyze he na u e o pause and in e up ions in
spoken language.
2.1. Pause
The e a e wo main kinds o pauses: emp y pause o silence and
illed pause o lexicalized pause, which accomplish di e en unc ions:
In he case o emp y pauses o silence, he p ocessing o his in o ma ion can
de e mine wo d and sen ence di ision. This in o ma ion is e y impo an o
de e mine he meaning o he sequence.
A illed pause akes place when he speake is ying o adjus he speaking
a e o he cogni i e p ocesses ha a e going on. Filled pauses ha e a e y
complex unc ionali y in a dialogue si ua ion:
They help he speake keep a dialogue u n (“I mean”, “well”, “e ”, “um”,
“you see?”e c.). Filled pauses also help o keep he lis ene ale , p e en ing
undesi ed in e up ions.
They a e also common o exp ess emo ional s a es, such us hesi a ion (“e ”,
“um”, e c.), anxie y, ange , su p ise, e c. (“come on!”, “su e!”, “bu excuse
me!”, e c.). They a e equen ly used o exp ess cogni i e and men al s a es
( e y o en we inco po a e illed pauses when we a e looking o he
app op ia e wo d o exp ession o when we a e ying o ecall in o ma ion).
Filled pauses equen ly allow he lis ene o p edic wha comes nex , as i
happens in he ollowing example:
(1) A: “I e y seldom go ou o dinne , jus once, e ...”
B: “Once in a blue moon”
In (1) B in e up s A and inishes A’s discou se.
T. López So o
98
In dialogue sys ems, we can ind wo main app oaches o deal wi h
pauses:
In some sys ems, illed pauses a e included in he lexicon.
Some imes lexicalized pauses a e conside ed unknown e ms and do no go
beyond he ini ial le el o analysis.
In any o he cases abo e hough, we miss he in o ma ion ha hese
sequences may ha e o a co ec analysis and unde s anding o he speech
signal.
The analysis o pauses in ASR sys ems has usually aken place in he
acous ic module. The acous ic analysis o pauses will be mo e e ec i e
when hey a e lexicalized (“su e”, “well”, e c.), less so when hey a e no
lexicalized (“um”, “e ”, e c.). On he o he hand, language modeling canno
be e icien o analyse illed pauses because hey can appea a any posi ion
in he sen ence.
So he ques ion would be: Can illed pauses be analysed in a eliable
way? We can s a e ha only p osodic in o ma ion can de e mine he
occu ence o pauses in he speech signal. We can illus a e his wi h one
example. In he ollowing sen ence, he e m “well” is used as a illed pause
and as an ad e b. The meaning a ies acco ding o he pi ch change
associa ed wi h he wo d “well” when i is used as a illed pause.
“I don’ like my s eak well done” (bu maybe I like i jus done)
“I don’ like my s eak, well, done” (I like i a e)
A illed pause has no meaning in i sel , i s main unc ion is o keep
he dialogue u n. Howe e , a illed pause may a ec he gene al seman ic
s uc u e o he speech sequence. Filled pauses occu when he speaking
p ocess is in a “s and-by” s a e, he e o e a icula o y ea u es emain in ac
un il discou sal and cogni i e p ocesses a e kep equal. Ou model is based
on he analysis o F0 ansi ions and spec al de o ma ion o de ec illed
pauses. Filled pauses usually occu when he ension o he ocal olds
ELIA 2, 2001
99
emains in a iable unde cons an a icula o y pa ame e s and he F0 o
oice s ays p ac ically cons an bu he spec og aphic analysis is no ably
al e ed. The shape o he ocal ac does no a y unde cons an a icula o y
pa ame e s, no does he spec al slope de o ma ion. Wi h his sys em we can
de ec when a illed pause is in oduced in he low o he speech signal.
2.2. In e up ions
In e up ions and o he speech dis luencies (as desc ibed in Heeman,
1998) also play an impo an ole o he in e p e a ion o he discou se. In
his sec ion we will analyze he cha ac e is ics and unc ions o in e up ions
in dialogue sys ems.
In e up ions may ha e he ollowing unc ions:
They can be used o change he opic o he dialogue
Ve y o en hey con i m a message
In e up ions can epai o emphasise a message
They a e e y equen ly used o ob ain in o ma ion
In e up ions can be de ined as hose si ua ions whe e a pe son
in ends o con inue speaking bu is o ced by ano he pe son o s op
speaking, a leas empo a ily, o he con inui y o egula i y o speech is
b oken o any o he eason.
In e up ions can be classi ied in o wo gene al g oups:
Compe i i eness: The lis ene in e up s he low o speech o exp ess
u gency, deg ee o impo ance o he opic, in e es . Compe i i eness also
akes place when he lis ene wan s o exp ess s ong opinion o
disag eemen . The low o speech is di e ed a e he in e up ion.
Coope a ion: The lis ene in e up s he low o speech o con i m o
s eng hen he speake ’s poin o iew. The low o he speech emains
cons an a e he in e up ion.

T. López So o
100
The p osodic ea u es o in e up ions o class (1) e lec he
necessi y and u gency o he lis ene o ecei e in o ma ion, as well as he
necessi y o including he in e up ion in he discou se so ha i shows a high
deg ee o ele ance. On he o he hand, in e up ions o class (2) a e
o igina ed when he pe son who p oduces he in e up ion shows a g ea e
deg ee o con idence and ce ain y in he dialogue.
In e up ions mus be analysed a e conside ing he gene al
in ona ional s uc u e o he sequence, he ampli ude o he wa e o m and
he speech a e. In e up ions o class (1) a e usually associa ed wi h
unexpec ed pi ch changes (highe pi ch le els), show a g ea e ampli ude and
he speech a e accele a es. In less auma ic in e up ions (class 2) he pi ch
le el is usually low, because he pe son who is in e up ing me ely exp esses
his/he ag eemen . Howe e , he ampli ude inc eases o e lec emphasis.
The speech a e emains unal e ed.
3. A model o he speci ica ion o p osodic in o ma ion based on
p edic ions
In he p e ious sec ions we ha e desc ibed he unc ionali y o
pauses and in e up ions in a dialogue sys em. The e a e ob iously se e al
o he phenomena ha can be analysed in he gene al con ex o discou se in
o de o disco e he in o ma ion hey supply o he gene al meaning
s uc u e: epe i ions, alse s a s, e c. Howe e , in ou p esen s udy we a e
only conce ned wi h he de ec ion and analysis o pauses and in e up ions
o a Spanish co pus.
The p osodic s uc u e associa ed wi h a speech signal has a e y
speci ic unc ion. This in o ma ion is as ele an as he syn ac ic and
seman ic in o ma ion. The main obs acle we ha e o ace is ha in ona ional
ea u es a e s ongly associa ed wi h pa icula communica i e con ex s. Fo
his eason, we p opose a speci ica ion model which s udies he unc ionali y
o p osody in a dialogue sys em. In ou sys em he use can send messages o
he machine and he e exis s a limi ed in e change o ques ions and answe s
in he nego ia ion o meaning (Ál a ez e al., 1997). Tha is, we a e dealing
wi h a e y es ic ed domain whe e we can ind a limi ed numbe o
ELIA 2, 2001
101
linguis ic cons uc ions, and he e o e, o p osodic s uc u es. In he nex
sec ions we explain how he co pus was designed.
3.1. Co pus labeling
The model ha we p esen he e is based on human p edic ions
(Tamo o e al. 1999). This model is based on he e alua ion by human
labele s, who iden i y and selec he p osodic s uc u es associa ed wi h he
co pus. In his sec ion we desc ibe he co pus labeling p ocess.
To de e mine he unc ion o pauses and in e up ions in he co pus,
we ha e i s analyzed he in ona ional s uc u e in he domain aking in o
accoun pi ch ange, ampli ude, F0 alues, equency, spec al slope,
du a ion and speech a e. The esul has been a co pus labeled by ained
na i e speake s. The model has ollowed se e al s ages:
The co pus is eco ded and ansc ibed.
The ansc ip ions a e hen labeled by ained na i e speake s. The objec i e
is o selec he discou se s uc u es which cha ac e ize he co pus ollowing
he model p oposed by Ca le a e al. (1997). Tha is, each comple e
discou sal sequence is labeled acco ding o he communica i e unc ion i
shows in he gene al con ex o he dialogue.
The eco ding is il e ed o ex ac he F0. This new e sion is passed on o
he labele s which de e mine which in ona ional s uc u e co esponds o
each discou sal label acco ding o hei knowledge o he language.
3.2. P osodic labeling
Once discou sal labels had been assigned, he co pus was labeled
once mo e, his ime o show in ona ional s uc u es. P osodic labeling was
done in wo s a es:
The i s labeling is done manually, a e he labele s ha e lis ened o he
eco ding om which he F0 has been ex ac ed. These labels ollow he
no a ion desc ibed by Sil e man e al., (1992).
The second labeling is done au oma ically using Wa es TM. The
T. López So o
102
ep esen a ion o he F0 o he speech signal is ob ained and labeled
ollowing he same no a ion sys em.
Once he labeling p ocess is o e , impo an conclusions can be
aken a e con as ing he wo labeled co po a.
4. Conclusion
The model ha we p esen is based on he p edic ions made by
na i e speake s assigning in ona ional labels o a spoken co pus. A e
compa ing he wo co po a (one manually labeled, he o he one
au oma ically labeled) we can ge o mo e eliable conclusions. The co pus
ha we ha e ob ained includes in o ma ion ha is necessa y in o de o
p ocess p osodic in o ma ion, and, mo e speci ically, in o de o de ec
pauses and in e up ions. This in o ma ion is based on an analysis o he 0.
O he in o ma ion can be ex ac ed om equency, spec al slope, du a ion
and speech a e. Wi h hese esul s we a e now wo king on he de elopmen
o an au oma ic model o analyze and p ocess p osodic in o ma ion in o de
o de ec pauses and in e up ions in a spoken co pus o a dialogue sys em.
Bibliog aphical e e ences
Ál a ez, J., D. Tapias, C. C espo, I. Co áza y F. Ma ínez. 1997.
“De elopmen and E alua ion o he ATOS Spon aneous Speech
Con e sa ional Sys em”. In e na ional Con e ence on Acous ics,
Speech and Signal P ocessing (ICASSP’97), 1139-1142.
Ca le a, J., S. Isa d, G. Dohe y-Sneddon, A. Isa d, J. C. Kow ko y A. H.
Ande son. 1997. “The Reliabili y o a Dialogue S uc u e Coding
Scheme”. Compu a ional Linguis ics, 23(1):13-31.
Heeman, P. A. 1997. Speech Repai s, In ona ional Bounda ies and
Discou se Ma ke s: Modelling Speake 's U e ances in Spoken
Dialogue. PhD disse a ion. Dp . o Compu e Science, Uni e si y o
Roches e , N.Y.
López-So o, M. T. 1999. Es a egias de análisis g ama ical y semán ico
pa a un sis ema di igido po oz. PhD disse a ion. Dp o. de Lengua
Inglesa, Uni e sidad de Se illa, Se illa.
ELIA 2, 2001
103
Sil e man, K., M. Bechman, J. Pi elli, M. Os e do , C. Wigh man, P. P ice,
J. Pie ehumbe , J. Hi schbe g. 1992. “ToBI: a S anda d o Labelling
English P osody”. P oceedings, 2nd In e na ional Con e ence on
Spoken Language P ocessing, Ban /Canada:867-870
Tamo o, M., M. Kawamo i y T. Kawaba a. 1999. “In eg a ing P osodic
Fea u es in Dialogue Unde s anding”. P oc. o Eu ospeech-99. (CD-
ROM)
Van Noo d, G., G. Bouma, R. Koeling y M. Nede ho . 1998. “Robus
G amma ical Analysis o Spoken Dialogue Sys ems”. Na u al
Language Enginee ing 1(1):1-48.