Full text
Towa ds he Neu omo phic Implemen a ion o he Audi o y
Pe cep ion in he iCub Robo ic Pla o m
Daniel Gu ie ez-Galan
Uni e si y o Se illa
Se illa, Spain
[email p o ec ed]
Chia a Ba olozzi
Is i u o I aliano di Tecnologia
Genoa, I aly
Juan Ped o
Dominguez-Mo ales
Uni e si y o Se ille
Se ille, Spain
Angel Jimenez-Fe nandez
Uni e si y o Se ille
Se ille, Spain
Alejand o Lina es-Ba anco
Uni e si y o Se ille
Se ille, Spain
ABSTRACT
Hea ing can be conside ed as one o he mos impo an senses
since i plays a key ole in he audio isual lea ning p ocess. While
a lo o e o has been made o achie ing good esul s om he
adi ional app oach o audi o y pe cep ion, new ends such as
neu omo phic compu ing a e showing p omising achie emen s
in he implemen a ion o b ain s uc u es o senso y pe cep ion.
In his wo k, he design and in eg a ion o a neu omo phic e en -
based digi al model o he audi o y ascending pa hway wi hin he
iCub obo ic pla o m is p oposed. This model, which comp ises
om he cochlea up o he in e io colliculus, eplaces he adi ional
app oach o sound p ocessing al eady implemen ed on iCub and
is able o pe o m sound ecogni ion and spa ial localiza ion in eal-
ime, allowing he implemen a ion o audi o y a en ion models in
complex scena ios, like he cock ail pa y p oblem.
CCS CONCEPTS
•Compu ing me hodologies → Bio-inspi ed app oaches
;
•␣
Applied compu ing → E en -d i en a chi ec u es.
KEYWORDS
neu omo phic audi o y senso , iCub, e en -based p ocessing, sound
ecogni ion, sound localiza ion
1 INTRODUCTION
Humanoid obo s will soon be a common p esence among humans
since hei capabili ies o in e ac ion wi h he en i onmen and
humans a e s eadily inc easing, wi h objec and human ecogni ion
[
9
], and speech unde s anding [
5
]. Howe e , hese capabili ies com-
monly esul in a high compu a ional cos and a e o en deployed in
emo e compu ing sys ems, wi h he need o cons an da a ans e .
Neu omo phic pe cep ion and compu a ion, inspi ed by he compu-
a ional p inciples o biological neu al sys ems, can be used owa ds
low-powe , embedded solu ions o complex pe cep i e asks. In his
wo k, a digi al e en -based bio-inspi ed model [
6
] o he Audi o y
Ascending Pa hway (
AAP
) [
1
] on he iCub humanoid pla o m [
7
]
was implemen ed as a neu omo phic al e na i e o he adi ional
audi o y models in o de o educe la ency, powe consump ion
and compu a ional cos . This model, which has been alida ed o
sound ecogni ion and sound sou ce localiza ion, can be ex ended o
implemen mo e complex audi o y p ocessing and audio- isual u-
sion [
10
], in o de o imp o e he accu acy and lea ning capabili ies
o he sys em.
2 GIVING ICUB THE SENSE OF HEARING
The audi o y model comp ises he Neu omo phic Audi o y Com-
plex (
NAC
), implemen ed on he Zynq Field P og ammable Ga e
A ay (
FPGA
) o econ igu abili y and eal- ime signal p ocessing,
and he In e io Colliculus (
IC
) on SpiNNake , whe e la ge b ain
a eas can be implemen ed, as shown in Fig. 1(A). Thanks o he bidi-
ec ional communica ion be ween iCub and SpiNNake [
8
] h ough
Ye Ano he Robo Pla o m (
YARP
) [
2
], eal- ime, close-loop audio
applica ions can be implemen ed. The
NAC
module was in eg a ed
wi hin he iCub obo ic pla o m as an In ellec ual P ope y (
IP
)
co e connec ed o he Head P ocessing Uni (
HPU
) module. The
NAC
ecei es sound inpu om he wo mic ophones placed in
he ea s o he iCub. I is composed o h ee sub-modules: 1) a 32-
channels binau al Neu omo phic Audi o y Senso (
NAS
) [
6
], which
implemen s an e en -based digi al cochlea model; 2) he Spike-
based Supe io Oli e (
SSO
), which implemen s bo h he In e au al
Time Di e ence (
ITD
) [
3
] and he In e au al Le el Di e ence (
ILD
)
ex ac ion om he
NAS
’ ou pu e en s; and 3) he e en s moni o
oge he wi h he in e ace o he
HPU
co e, which is ca ied ou by
using he Add ess E en Rep esen a ion (
AER
) p o ocol. The sound
ea u es ex ac ed wi h he
NAC
a e combined and in eg a ed by a
mul ilaye Spiking Neu al Ne wo k (
SNN
) in SpiNNake ha mod-
els he beha io o he
IC
[
1
] o ex ac he ele an in o ma ion
ega ding he loca ion o he objec in he ho izon al axis. Then, he
ou pu o he
SNN
is sen back o he iCub o u n i s head owa ds
he sound sou ce. This close-loop sys em allows eal- ime sound
I2S
mic ophone
ZynQ
NAC
HPU Co e
SpiNNake
SNNs
Le
mic.
Le
came a
Righ
mic.
Righ
came a
Compu e
B)A) C)
Figu e 1: A) iCub-NAC in eg a ion diag am. B) Ou pu e en s om he localiza ion model. C) Sonog am om he keywo d
spo ing.
sou ce localiza ion and is he basic building block o audi o y
a en ion [4].
P elimina y esul s o he
NAC
-iCub in eg a ion show ha iCub
was success ully able o localize a sound sou ce which was shi ed
om le o igh and ice e sa in on o he iCub, as shown in
Fig. 1B). In addi ion, a da ase o spoken digi s, shown in Fig. 1C),
was di ec ly eco ded om he iCub o ain a
SNN
o pe o ming
keywo d spo ing. To he bes o ou knowledge, his is he i s
ime ha a digi al neu omo phic audi o y model has been in e-
g a ed wi hin a humanoid obo , and hese esul s open he doo
o u u e implemen a ions o senso y usion models wi h lea ning
algo i hms o sol e asks such as he cock ail pa y p oblem on a
ully neu omo phic eal- ime obo ic pla o m, bene i ing om he
compu a ional la ency and low-powe consump ion na u e o hese
sys ems.
ACKNOWLEDGMENTS
This wo k was suppo ed by he Spanish g an MINDROB (PID2019-
105556GB-C33/AEI/10.13039/501100011033), and SMALL (PCI2019-
111841-2/AEI/10.1309/501100011033 om EU CHIST-ERA p og am).
Pa o his wo k was ca ied ou du ing a esea ch in e nship
o D. G.-G. in he EDPR g oup (I alian Ins i u e o Technology,
Genoa (I aly)) suppo ed by a Fo mación de Pe sonal In es igado
Schola ship om he Spanish Minis y o Educa ion and Cul u e.
REFERENCES
[1]
Jo ge Dá ila-Chacón e al
.
2018. Enhanced obo speech ecogni ion using
biomime ic binau al sound sou ce localiza ion. IEEE ansac ions on neu al ne -
wo ks and lea ning sys ems 30, 1 (2018), 138–150.
[2]
A en Glo e e al
.
2018. The e en -d i en so wa e lib a y o YARP—Wi h
algo i hms and iCub applica ions. F on ie s in Robo ics and AI 4 (2018), 73.
[3]
Daniel Gu ie ez-Galan e al
.
2019. A neu omo phic app oach o he sound
sou ce localiza ion ask in eal- ime embedded sys ems: wo k-in-p og ess. In
P oceedings o he In e na ional Con e ence on Embedded So wa e Companion.
1–2.
[4]
Dillon A Hamb ook e al
.
2017. A Bayesian compu a ional basis o audi o y
selec i e a en ion using head o a ion and he in e au al ime-di e ence cue.
PloS one 12, 10 (2017).
[5]
Be and Higy, Alessio Me e a, Gio gio Me a, and Leona do Badino. 2018. Speech
ecogni ion o he icub pla o m. F on ie s in Robo ics and AI 5 (2018), 10.
[6]
Angel Jiménez-Fe nández e al
.
2017. A Binau al Neu omo phic Audi o y Senso
o FPGA: A Spike Signal P ocessing App oach. IEEE T ans. Neu al Ne w. Lea ning
Sys . 28, 4 (2017), 804–818.
[7]
Lo enzo Na ale e al
.
2017. icub: The no -ye - inished s o y o building a obo
child. Science Robo ics 2, 13 (2017).
[8]
Eus ace Paink as e al
.
2013. SpiNNake : A 1-W 18-co e sys em-on-chip o
massi ely-pa allel neu al ne wo k simula ion. IEEE Jou nal o Solid-S a e Ci cui s
48, 8 (2013), 1943–1953.
[9]
Giulia Pasquale, Ca lo Cilibe o, F ancesca Odone, Lo enzo Rosasco, and Lo enzo
Na ale. 2019. A e we done wi h objec ecogni ion? The iCub obo ’s pe spec i e.
Robo ics and Au onomous Sys ems 112 (2019), 260–281.
[10]
Vadim Tikhano e al
.
2010. In eg a ion o speech and ac ion in humanoid
obo s: iCub simula ion expe imen s. IEEE T ansac ions on Au onomous Men al
De elopmen 3, 1 (2010), 17–29.