scieee Science in your language
[en] (orig)

Multilayer Spiking Neural Network for Audio Samples Classification Using SpiNNaker

Abstract

Audio classification has always been an interesting subject of research inside the neuromorphic engineering field. Tools like Nengo or Brian, and hardware platforms like the SpiNNaker board are rapidly increasing in popularity in the neuromorphic community due to the ease of modelling spiking neural networks with them. In this manuscript a multilayer spiking neural network for audio samples classification using SpiNNaker is presented. The network consists of different leaky integrate-and-fire neuron layers. The connections between them are trained using novel firing rate based algorithms and tested using sets of pure tones with frequencies that range from 130.813 to 1396.91 Hz. The hit rate percentage values are obtained after adding a random noise signal to the original pure tone signal. The results show very good classification results (above 85 % hit rate) for each class when the Signal-to-noise ratio is above 3 decibels, validating the robustness of the network configuration and the training step.

Read accessible full text

Multilayer Spiking Neural Network for Audio Samples Classification Using SpiNNaker

Author: Domínguez Morales, Juan Pedro; Jiménez Fernández, Ángel Francisco; Ríos Navarro, José Antonio; Cerezuela Escudero, Elena; Gutiérrez Galán, Daniel; Domínguez Morales, Manuel Jesús; Jiménez Moreno, Gabriel
Publisher: Springer
Year: 2016
DOI: 10.1007/978-3-319-44778-0_6
Source: https://idus.us.es/bitstreams/a1909050-3dc0-4d1e-bed2-2403cd2aec29/download
Mul ilaye Spiking Neu al Ne wo k o Audio Samples
Classifica ion Using SpiNNake
Juan Ped o Dominguez-Mo ales, Angel Jimenez-Fe nandez, An onio Rios-Na a o,
Elena Ce ezuela-Escude o, Daniel Gu ie ez-Galan, Manuel J. Dominguez-Mo ales,
and Gab iel Jimenez-Mo eno
Robo ic and Technology o Compu e s Lab,
Depa men o A chi ec u e and Technology o Compu e s,
Uni e si y o Se ille, Se ille, Spain
{jpdominguez,ajimenez,a ios,ece ezuela,
dgu ie ez,mdominguez,gaji}@a c.us.es
h p://www.a c.us.es
Abs ac . Audio classifica ion has always been an in e es ing subjec o esea ch
inside he neu omo phic enginee ing field. Tools like Nengo o B ian, and ha d‐
wa e pla o ms like he SpiNNake boa d a e apidly inc easing in popula i y in
he neu omo phic communi y due o he ease o modelling spiking neu al
ne wo ks wi h hem. In his manusc ip a mul ilaye spiking neu al ne wo k o
audio samples classifica ion using SpiNNake is p esen ed. The ne wo k consis s
o diffe en leaky in eg a e-and-fi e neu on laye s. The connec ions be ween hem
a e ained using no el fi ing a e based algo i hms and es ed using se s o pu e
ones wi h equencies ha ange om 130.813 o 1396.91 Hz. The hi a e
pe cen age alues a e ob ained a e adding a andom noise signal o he o iginal
pu e one signal. The esul s show e y good classifica ion esul s (abo e 85 %
hi a e) o each class when he Signal- o-noise a io is abo e 3 decibels, ali‐
da ing he obus ness o he ne wo k configu a ion and he aining s ep.
Keywo ds: SpiNNake · Spiking neu al ne wo k · Audio samples classi ica ion ·
Spikes · Neu omo phic audi o y senso · Add ess-E en Rep esen a ion
1 In oduc ion
Neu omo phic enginee ing is a discipline ha s udies, designs and implemen s ha dwa e
and so wa e wi h he aim o mimicking he way in which ne ous sys ems wo k,
ocusing i s main inspi a ion on how he b ain sol es complex p oblems easily. Nowa‐
days, he neu omo phic communi y has a se o neu omo phic ha dwa e ools a ailable
such as senso s [1, 2], lea ning ci cui s [3, 4], neu omo phic in o ma ion fil e s and
ea u e ex ac o s [5, 6], obo ic and mo o con olle s [7, 8]. In he field o neu omo phic
senso s, di e se neu omo phic cochleae can be ound [2, 9, 10]. These senso s a e able
o decompose he audio in equency bands, and ep esen hem as s eams o sho
pulses, called spikes, using he Add ess-E en Rep esen a ion (AER) [11] o in e ace
wi h o he neu omo phic laye s. On he o he hand, he e a e se e al so wa e ools in
he communi y o spiking neu al ne wo ks (SNN) simula ion, i.e. NENGO [12] and
BRIAN [13]; o jAER [14] o eal- ime isualiza ion and so wa e p ocessing o AER
s eams cap u ed om he ha dwa e using specific in e aces [15]. Ha dwa e pla o ms
like he SpiNNake boa d [16] allows o de elop and implemen complex SNN easily
using a high-le el p og amming language such as Py hon and he PyNN [17] lib a y.
This manusc ip p esen s a no el mul ilaye SNN a chi ec u e buil in SpiNNake
which has been ained o audio samples classifica ion using a fi ing a e based algo‐
i hm. To es he ne wo k beha io and obus ness, a 64-channel binau al Neu omo phic
Audi o y Senso (NAS) o FPGA [10] has been used oge he wi h an USB-AER in e ‐
ace [15] (Fig. 1) and he jAER so wa e, allowing o p oduce diffe en pu e ones wi h
equencies a ying om 130.813 Hz o 1396.91 Hz, eco d he NAS esponse s o ing
he in o ma ion in aeda files h ough jAER and use hese files as inpu o he SNN ha
has been implemen ed in he SpiNNake boa d.
Fig. 1. Block diag am o he sys em
The pape is s uc u ed as ollows: Sec . 2 p esen s he numbe o neu ons, laye s
and connec ions o he SNN. Then, Sec . 3 desc ibes he aining algo i hm used in e e y
laye o he audio samples classifica ion. Sec ion 4 desc ibes he es scena io, including
in o ma ion abou he inpu files. Then, Sec . 5 p esen s he expe imen al esul s o he
audio samples classifica ion when using he inpu s desc ibed in Sec . 4. Finally, Sec . 6
p esen s he conclusions o his wo k.
2 Ha dwa e Se up
The s andalone ha dwa e used in his wo k consis s o wo main pa s: he 64-channel
NAS connec ed o he USB-AER in e ace o gene a ing a spike s eam o each audio
sample, and he SpiNNake o back-end compu a ion and deploymen o he SNN
classifie .
2.1 Neu omo phic Audi o y Senso (NAS)
A Neu omo phic Audi o y Senso (NAS) is used as he inpu laye o ou sys em. This
senso con e s he incoming sound in o a ain o a e-coded spikes and p ocesses hem
using Spike Signal P ocessing (SSP) echniques o FPGA [5]. NAS is composed o a
se o Spike Low-pass Fil e s (SLPF) implemen ing a cascade opology, whe e SLPF’s
co ela i e spike ou pu s a e sub ac ed, pe o ming a bank o equi alen Spikes Band-
pass Fil e s (SBPF), and decomposing inpu audio spikes in o spec al ac i i y [10].
Finally, SBPF spikes a e collec ed using an AER moni o , codi ying each spike using
he Add ess-E en Rep esen a ion, and p opaga ing AER e en s h ough a 16-bi
pa allel asynch onous AER po [11].
NAS designing is e y flexible and ully cus omizable, allowing neu omo phic engi‐
nee s o build applica ion-specific NASs, wi h di e se ea u es and numbe o channels.
In his case, we ha e used a 64-channel binau al NAS, wi h a equency esponse
be ween 20 Hz and 22 kHz, and a dynamic ange o +75 dB, syn hesized o a Vi ex-5
FPGA. Figu e 1 shows a NAS implemen ed in a Xilinx de elopmen boa d, and a USB-
AER mini2 boa d, ha implemen s a b idge be ween AER sys ems and jAER in a PC
(Fig. 2).
Fig. 2. 64-channel binau al NAS implemen ed in a Xilinx ML507 FPGA connec ed o an USB-
AER mini2 boa d.
2.2 Spiking Neu al Ne wo k A chi ec u e (SpiNNake )
SpiNNake is a massi ely-pa allel mul i-co e compu ing sys em designed o modelling
e y la ge spiking neu al ne wo ks in eal ime. Each SpiNNake chip comp ises 18
gene al-pu pose ARM968 co es, unning a 200 MHz, communica ing ia packe s
ca ied by a cus om in e connec ab ic. Packe s a e ansmi ed and hei ansmission
is b oke ed en i ely by ha dwa e, gi ing he o e all engine an ex emely high bisec ion
bandwid h. The Ad anced P ocesso Technologies Resea ch G oup (APT) [18] in
Manches e a e esponsible o he sys em a chi ec u e and he design o he SpiNNake
chip i sel .
In his wo k, a SpiNNake 102 machine was used. The 102 machine, Fig. 3, is a 4-
node ci cui boa d and hence has 72 ARM p ocesso co es, which a e ypically deployed
as 64 applica ion co es, 4 Moni o P ocesso s and 4 spa e co es. The 102 machine
equi es a 5 V 1 A supply, and can be powe ed om some USB2 po s. The con ol and
I/O in e ace is a single 100 Mbps E he ne connec ion.
Fig. 3. SpiNNake 102 machine.
3 Leaky In eg a e-and-Fi e Spiking Neu al Ne wo k
The SpiNNake pla o m allows o implemen a specific spiking neu on model and use
i in any SNN deployed on he boa d hanks o he PyNN package. Leaky In eg a e-and-
Fi e (LIF) neu ons ha e been used in a 3-laye SNN a chi ec u e o audio samples
classifica ion.
•Inpu laye . This laye ecei es he s eam o AER e en s fi ed o he audio samples
cap u ed as aeda files h ough jAER. The numbe o inpu neu ons is equal o he
numbe o channels ha he NAS has. As a 64-channel NAS (64 diffe en AER
add esses) was used in his wo k, he inpu laye consis s o 64 LIF neu ons.
•Hidden laye . The hidden laye has he same numbe o neu ons as he desi ed
numbe o classes o be classified in he ou pu laye . As an example, his laye should
consis o eigh LIF neu ons i eigh diffe en audio samples a e expec ed o be
classified.
•Ou pu laye . As he p e ious laye , his also has as many neu ons as ou pu classes.
The fi ing ou pu o he neu ons in his laye will de e mine he esul o he classi‐
fica ion.
Figu e 4 shows he SNN a chi ec u e. Connec ions be ween laye s a e achie ed using
he F omLis Connec o me hod om PyNN, meaning ha he sou ce, des ina ion and
weigh o he connec ion a e specified manually. Using o he connec o s om his
package will esul on ha ing he same weigh in all he connec ions be ween consecu i e
laye s, ins ead o a diffe en alue o each. In his a chi ec u e, each neu on in a laye
is connec ed o e e y neu on in he nex laye , and he weigh alue is ob ained om he
aining s ep, which is desc ibed in Sec . 4. The h eshold ol age o he neu ons in he
hidden laye is 15 mV, while his ol age is 10 mV in he neu ons in he ou pu laye .
Decay a e and e ac o y pe iod a e he same o bo h laye s: 20 mV/ms and 2 ms,
espec i ely.
Fig. 4. SNN a chi ec u e using an audio sample aeda file as inpu .
4 T aining Phase o Audio Classifica ion
In he p e ious sec ion, each o he h ee laye s comp ising he ne wo k we e desc ibed.
The aining phase is pe o med offline and supe ised. The main objec i e o his
aining is o ob ain he weigh alues o he connec ions be ween he inpu and he
hidden laye and be ween he hidden and he ou pu laye s o u he audio samples
classifica ion. The e o e, wo diffe en aining s eps need o be done.
The weigh s o he fi s s ep o he aining phase a e ob ained om he no malized
spike fi ing ac i i y o each NAS channel using a se o audio samples simila o hose
o be ecognized (same ampli ude, du a ion and equencies). The fi ing a e o a specific
channel (FR
channel_i
) is ob ained by di iding he numbe o e en s p oduced in ha
channel by he NAS fi ing a e (FR
T
), which is he numbe o e en s fi ed in he NAS
in a pa icula ime pe iod.
(1)
(2)
Figu e 5 shows he no malized spike fi ing ac i i y o a se o eigh pu e ones wi h
equencies ha ange om 130.813 Hz o 1396.91 Hz, loga i hmically spaced.
The weigh s o he second s ep o he aining phase a e ob ained om he fi ing
ou pu o each neu on in he hidden laye when using he se o audio samples as inpu
a e loading he weigh s calcula ed in he p e ious s ep in o he connec ions be ween
he inpu and he hidden laye . These fi ing ou pu s a e no malized by di iding each o
hem by he maximum alue. The esul s ob ained a e he weigh alues ha will be used
in he connec ions be ween he hidden and he ou pu laye .

5Tes Scena io
In his wo k, he SNN a chi ec u e and aining algo i hm p esen ed a e es ed using
eigh diffe en audio samples. These ou pu classes co espond o eigh diffe en pu e
ones wi h equencies ha ange om 130.813 Hz o 1396.91 Hz, loga i hmically
spaced (130.813, 174.614, 261.626, 349.228, 523.251, 698.456, 1046.50 and
1396.91 Hz). These samples ha e a du a ion o 0.5 s and we e gene a ed using he
audioplaye unc ion om Ma lab wi h a sampling a e o 48 KHz and a peak- o-peak
ol age alue o 1 V. A e he signal is sen o he mixe , i p opaga es he sound o
NAS inpu and sends an AER s eam o he PC h ough he AER-USB in e ace. The
jAER so wa e unning on he PC is able o cap u e his s eam and sa e i as an aeda
file. Figu e 6 shows he cochleog ams o he 130.813 Hz and he 1396.91 Hz pu e ones
a e cap u ing hem.
The fi s s ep o he aining phase can be achie ed by applying he equa ions
p esen ed in Sec . 4 o he se o eigh aeda files co esponding o each pu e one. This
will gene a e a CSV file con aining he weigh s o he 64 × 8 connec ions be ween he
inpu and he hidden laye s o he SNN based on he fi ing a e o he spike s eams o
each audio sample. As desc ibed in he p e ious sec ion, loading hose weigh s in o he
co esponding connec ions and using he eigh pu e ones as inpu will esul on a fi ing
Fig. 5. No malized spike fi ing ac i i y o each NAS channel pe audio sample.
Fig. 6. Fi s 10 ms cochleog am o he 130.813 Hz (le ) and 1396.91 Hz ( igh ) pu e ones.
ou pu on he second laye neu ons ha will be used o aining he connec ions be ween
he second and he ou pu laye s o he SNN.
A e he weigh s a e se on hese connec ions, new se s o he same pu e ones (same
equencies) a e eco ded using diffe en Signal- o-Noise Ra io (SNR) alues and es ed
on he ne wo k, calcula ing he hi a e pe cen age o each class.
6 Expe imen al Resul s
Diffe en pu e one se s wi h he same equencies and p ope ies (0.5 s and 0.5 V
ampli ude) as he ones used in his wo k we e cap u ed and used o es he ne wo k
obus ness and effec i eness. A 100 % hi a e was ob ained o e e y class when he
signal was a pu e sine wa e. Mo eo e , he ne wo k has also been es ed by adding a
noise signal consis ing o andom alues o he pu e ones o iginal signals, ob aining
audio samples wi h diffe en SNR alues ( om 35.2 dB o 0 dB). The hi a e pe cen age
o e e y class using he p e ious SNR alues a e lis ed in Table 1.
Table 1. Hi a e pe cen age o he audio samples classifica ion SNN o diffe en SNR alues.
SNR
(dB)
Pu e one equency (Hz)
130.813 174.614 261.626 349.228 523.251 698.456 1046.5 1396.91
No noise 100 % 100 % 100 % 100 % 100 % 100 % 100 % 100 %
35.1993 100 % 100 % 100 % 100 % 100 % 100 % 100 % 100 %
21.3363 100 % 83 % 96 % 100 % 100 % 100 % 100 % 100 %
13.2273 100 % 81 % 92 % 100 % 100 % 100 % 100 % 96 %
7.4733 100 % 86 % 100 % 100 % 100 % 100 % 100 % 95 %
3.0103 74 % 90 % 100 % 98 % 100 % 100 % 100 % 98 %
2 93 % 88 % 20 % 32 % 16 % 92 % 32 % 97 %
1 10 %5 %0 %0 %0 %88 %26 %94 %
0 0 %0 %0 %0 %0 %76 %22 %91 %
The esul s show e y high hi a e pe cen ages when he SNR is abo e 3 dB.
Howe e , when he SNR alls below 3 dB and app oaches ze o dB ( he ampli ude o he
pu e one is he same as he ampli ude o he noise signal) he ne wo k is no able o
classi y e e y inpu signal as i s co esponding class.
7Conclusions
In his pape , a no el mul ilaye spiking neu al ne wo k a chi ec u e o audio samples
classifica ion implemen ed in SpiNNake has been p esen ed. To achie e his goal, an
op imized aining phase o audio ecogni ion has been desc ibed and specified in wo
diffe en s eps, which allow ob aining he weigh s o he connec ions be ween he inpu
and he hidden laye s and be ween he hidden and he ou pu laye s. The ne wo k was
ained using eigh pu e ones wi h equencies be ween 130.813 Hz and 1396.91 Hz
and es ed by adding a noise signal wi h SNR alues be ween 35.1993 and 0 dB.
The hi a e alues ob ained a e many es s confi m he obus ness o he ne wo k
and he aining, which make i possible o classi y e e y pu e one wi h a p obabili y
o e 74 % e en when he SNR alue is 3 dB, ob aining almos a 100 % p obabili y o
e e y inpu when he SNR is abo e ha alue.
Finally, he SpiNNake boa d has allowed o model and de elop a leaky in eg a e-
and-fi e spiking neu al ne wo k o his pu pose in an easy, as , use - iendly and effi‐
cien way, p o ing i s po en ial, and p omo ing and acili a ing he implemen a ion o
SNNs like hese in eal ha dwa e pla o ms. The PyNN code used o es he SNN
p esen ed in his wo k is a ailable a [19].
Acknowledgemen s. The au ho s would like o hank he APT Resea ch G oup o he Uni e si y
o Manches e o ins uc ing us in he SpiNNake . This wo k is suppo ed by he Spanish
go e nmen g an BIOSENSE (TEC2012-37868-C04-02) and by he excellence p ojec om
Andalusian Council MINERVA (P12-TIC-1300), bo h wi h suppo om he Eu opean Regional
De elopmen Fund.
Re e ences
1. Lich s eine , P., Posch, C., Delb uck, T.: A 128 × 128 120 dB 15 μs la ency asynch onous
empo al con as ision senso . IEEE J. Solid-S a e Ci c. 43, 566–576 (2008)
2. Chan, V., Liu, S.C., an Schaik, A.: AER EAR: a ma ched silicon cochlea pai wi h add ess
e en ep esen a ion in e ace. IEEE T ans Ci c. Sys . I 54(1), 48–59 (2007)
3. Häflige , P.: Adap i e WTA wi h an analog VLSI neu omo phic lea ning chip. IEEE T ans.
Neu al Ne w. 18, 551–572 (2007)
4. Indi e i, G., Chicca, E., Douglas, R.: A VLSI a ay o low-powe spiking neu ons and bis able
synapses wi h spike- iming dependen plas ici y. IEEE T ans. Neu al Ne w. 17, 211–221
(2006)
5. Jiménez-Fe nández, A., Jiménez-Mo eno, G., Lina es-Ba anco, A., e al.: Building blocks
o spikes signal p ocessing. In: In e na ional Join Con e ence on Neu al Ne wo ks, IJCNN
(2010)
6. Lina es-Ba anco, A., e al.: A USB3.0 FPGA e en -based fil e ing and acking amewo k
o dynamic ision senso s. In: P oceedings o IEEE In e na ional Symposium on Ci cui s
and Sys ems, pp. 2417–2420 (2015)
7. Lina es-Ba anco, A., Gomez-Rod iguez, F., Jimenez-Fe nandez, A., e al.: Using FPGA o
isuo-mo o con ol wi h a silicon e ina and a humanoid obo . In: IEEE In e na ional
Symposium on Ci cui s and Sys ems, pp. 1192–1195 (2007)
8. Jimenez-Fe nandez, A., Jimenez-Mo eno, G., Lina es-Ba anco, A., e al.: A neu o-inspi ed
spike-based PID mo o con olle o mul i-mo o obo s wi h low cos FPGAs. Senso s 12,
3831–3856 (2012)
9. Hamil on, T.J., Jin, C., an Schaik, A., Tapson, J.: An ac i e 2-D silicon cochlea. IEEE T ans.
Biomed. Ci c. Sys . 2, 30–43 (2008)
10. Jimenez-Fe nandez, A., Ce ezuela-Escude o, E., Mi o-Ama an e, L., e al.: A binau al
neu omo phic audi o y senso o FPGA: a spike signal p ocessing app oach. IEEE T ans.
Neu al Ne wo ks Lea n. Sys . 1(0) (2016)
11. Boahen, K.: Poin - o-poin connec i i y be ween neu omo phic chips using add ess e en s.
IEEE T ans. Ci c. Sys II Analog Digi Sig. P ocess. 47, 416–434 (2000)
12. Bekolay, T., e al.: Nengo: a Py hon ool o building la ge-scale unc ional b ain models.
F on Neu oin o m. 7, 48 (2014)
13. Goodman, D., B e e, R.: B ian: a simula o o spiking neu al ne wo ks in py hon. F on
Neu oin o m. 2, 5 (2008)
14. jAER Open Sou ce P ojec . h p://jae .wiki.sou ce o ge.ne
15. Be ne , R., Delb uck, T., Ci i -Balcells, A., Lina es-Ba anco, A.: A 5 Meps $100 USB2.0
add ess-e en moni o -sequence in e ace. IEEE In e na ional Symposium on Ci cui s and
Sys ems (2007)
16. Paink as, E., e al.: SpiNNake : A 1-W 18-co e sys em-on-chip o massi ely-pa allel neu al
ne wo k simula ion. IEEE J. Solid-S a e Ci c. 48, 1943–1953 (2013)
17. Da ison, A.P.: PyNN: a common in e ace o neu onal ne wo k simula o s. F on
Neu oin o m. 2, 11 (2008)
18. SpiNNake Home Page. h p://ap .cs.manches e .ac.uk/p ojec s/SpiNNake
19. Dominguez-Mo ales, J.P.: Mul ilaye spiking neu al ne wo k o audio samples classifica ion
using Spinnake Gi hub page. h ps://gi hub.com/jpdominguez/Mul ilaye -SNN- o -audio-
samples-classifica ion-using-SpiNNake