Mul ilaye Spiking Neu al Ne wo k o Audio Samples
Classifica ion Using SpiNNake
Juan Ped o Dominguez-Mo ales, Angel Jimenez-Fe nandez, An onio Rios-Na a o,
Elena Ce ezuela-Escude o, Daniel Gu ie ez-Galan, Manuel J. Dominguez-Mo ales,
and Gab iel Jimenez-Mo eno
Robo ic and Technology o Compu e s Lab,
Depa men o A chi ec u e and Technology o Compu e s,
Uni e si y o Se ille, Se ille, Spain
{jpdominguez,ajimenez,a ios,ece ezuela,
dgu ie ez,mdominguez,gaji}@a c.us.es
h p://www.a c.us.es
Abs ac . Audio classifica ion has always been an in e es ing subjec o esea ch
inside he neu omo phic enginee ing field. Tools like Nengo o B ian, and ha d‐
wa e pla o ms like he SpiNNake boa d a e apidly inc easing in popula i y in
he neu omo phic communi y due o he ease o modelling spiking neu al
ne wo ks wi h hem. In his manusc ip a mul ilaye spiking neu al ne wo k o
audio samples classifica ion using SpiNNake is p esen ed. The ne wo k consis s
o diffe en leaky in eg a e-and-fi e neu on laye s. The connec ions be ween hem
a e ained using no el fi ing a e based algo i hms and es ed using se s o pu e
ones wi h equencies ha ange om 130.813 o 1396.91 Hz. The hi a e
pe cen age alues a e ob ained a e adding a andom noise signal o he o iginal
pu e one signal. The esul s show e y good classifica ion esul s (abo e 85 %
hi a e) o each class when he Signal- o-noise a io is abo e 3 decibels, ali‐
da ing he obus ness o he ne wo k configu a ion and he aining s ep.
Keywo ds: SpiNNake · Spiking neu al ne wo k · Audio samples classi ica ion ·
Spikes · Neu omo phic audi o y senso · Add ess-E en Rep esen a ion
1 In oduc ion
Neu omo phic enginee ing is a discipline ha s udies, designs and implemen s ha dwa e
and so wa e wi h he aim o mimicking he way in which ne ous sys ems wo k,
ocusing i s main inspi a ion on how he b ain sol es complex p oblems easily. Nowa‐
days, he neu omo phic communi y has a se o neu omo phic ha dwa e ools a ailable
such as senso s [1, 2], lea ning ci cui s [3, 4], neu omo phic in o ma ion fil e s and
ea u e ex ac o s [5, 6], obo ic and mo o con olle s [7, 8]. In he field o neu omo phic
senso s, di e se neu omo phic cochleae can be ound [2, 9, 10]. These senso s a e able
o decompose he audio in equency bands, and ep esen hem as s eams o sho
pulses, called spikes, using he Add ess-E en Rep esen a ion (AER) [11] o in e ace
wi h o he neu omo phic laye s. On he o he hand, he e a e se e al so wa e ools in
he communi y o spiking neu al ne wo ks (SNN) simula ion, i.e. NENGO [12] and
BRIAN [13]; o jAER [14] o eal- ime isualiza ion and so wa e p ocessing o AER
s eams cap u ed om he ha dwa e using specific in e aces [15]. Ha dwa e pla o ms
like he SpiNNake boa d [16] allows o de elop and implemen complex SNN easily
using a high-le el p og amming language such as Py hon and he PyNN [17] lib a y.
This manusc ip p esen s a no el mul ilaye SNN a chi ec u e buil in SpiNNake
which has been ained o audio samples classifica ion using a fi ing a e based algo‐
i hm. To es he ne wo k beha io and obus ness, a 64-channel binau al Neu omo phic
Audi o y Senso (NAS) o FPGA [10] has been used oge he wi h an USB-AER in e ‐
ace [15] (Fig. 1) and he jAER so wa e, allowing o p oduce diffe en pu e ones wi h
equencies a ying om 130.813 Hz o 1396.91 Hz, eco d he NAS esponse s o ing
he in o ma ion in aeda files h ough jAER and use hese files as inpu o he SNN ha
has been implemen ed in he SpiNNake boa d.
Fig. 1. Block diag am o he sys em
The pape is s uc u ed as ollows: Sec . 2 p esen s he numbe o neu ons, laye s
and connec ions o he SNN. Then, Sec . 3 desc ibes he aining algo i hm used in e e y
laye o he audio samples classifica ion. Sec ion 4 desc ibes he es scena io, including
in o ma ion abou he inpu files. Then, Sec . 5 p esen s he expe imen al esul s o he
audio samples classifica ion when using he inpu s desc ibed in Sec . 4. Finally, Sec . 6
p esen s he conclusions o his wo k.
2 Ha dwa e Se up
The s andalone ha dwa e used in his wo k consis s o wo main pa s: he 64-channel
NAS connec ed o he USB-AER in e ace o gene a ing a spike s eam o each audio
sample, and he SpiNNake o back-end compu a ion and deploymen o he SNN
classifie .
2.1 Neu omo phic Audi o y Senso (NAS)
A Neu omo phic Audi o y Senso (NAS) is used as he inpu laye o ou sys em. This
senso con e s he incoming sound in o a ain o a e-coded spikes and p ocesses hem
using Spike Signal P ocessing (SSP) echniques o FPGA [5]. NAS is composed o a
se o Spike Low-pass Fil e s (SLPF) implemen ing a cascade opology, whe e SLPF’s
co ela i e spike ou pu s a e sub ac ed, pe o ming a bank o equi alen Spikes Band-
pass Fil e s (SBPF), and decomposing inpu audio spikes in o spec al ac i i y [10].
Finally, SBPF spikes a e collec ed using an AER moni o , codi ying each spike using
he Add ess-E en Rep esen a ion, and p opaga ing AER e en s h ough a 16-bi
pa allel asynch onous AER po [11].
NAS designing is e y flexible and ully cus omizable, allowing neu omo phic engi‐
nee s o build applica ion-specific NASs, wi h di e se ea u es and numbe o channels.
In his case, we ha e used a 64-channel binau al NAS, wi h a equency esponse
be ween 20 Hz and 22 kHz, and a dynamic ange o +75 dB, syn hesized o a Vi ex-5
FPGA. Figu e 1 shows a NAS implemen ed in a Xilinx de elopmen boa d, and a USB-
AER mini2 boa d, ha implemen s a b idge be ween AER sys ems and jAER in a PC
(Fig. 2).
Fig. 2. 64-channel binau al NAS implemen ed in a Xilinx ML507 FPGA connec ed o an USB-
AER mini2 boa d.
2.2 Spiking Neu al Ne wo k A chi ec u e (SpiNNake )
SpiNNake is a massi ely-pa allel mul i-co e compu ing sys em designed o modelling
e y la ge spiking neu al ne wo ks in eal ime. Each SpiNNake chip comp ises 18
gene al-pu pose ARM968 co es, unning a 200 MHz, communica ing ia packe s
ca ied by a cus om in e connec ab ic. Packe s a e ansmi ed and hei ansmission
is b oke ed en i ely by ha dwa e, gi ing he o e all engine an ex emely high bisec ion
bandwid h. The Ad anced P ocesso Technologies Resea ch G oup (APT) [18] in
Manches e a e esponsible o he sys em a chi ec u e and he design o he SpiNNake
chip i sel .
In his wo k, a SpiNNake 102 machine was used. The 102 machine, Fig. 3, is a 4-
node ci cui boa d and hence has 72 ARM p ocesso co es, which a e ypically deployed
as 64 applica ion co es, 4 Moni o P ocesso s and 4 spa e co es. The 102 machine
equi es a 5 V 1 A supply, and can be powe ed om some USB2 po s. The con ol and
I/O in e ace is a single 100 Mbps E he ne connec ion.
Fig. 3. SpiNNake 102 machine.
3 Leaky In eg a e-and-Fi e Spiking Neu al Ne wo k
The SpiNNake pla o m allows o implemen a specific spiking neu on model and use
i in any SNN deployed on he boa d hanks o he PyNN package. Leaky In eg a e-and-
Fi e (LIF) neu ons ha e been used in a 3-laye SNN a chi ec u e o audio samples
classifica ion.
•Inpu laye . This laye ecei es he s eam o AER e en s fi ed o he audio samples
cap u ed as aeda files h ough jAER. The numbe o inpu neu ons is equal o he
numbe o channels ha he NAS has. As a 64-channel NAS (64 diffe en AER
add esses) was used in his wo k, he inpu laye consis s o 64 LIF neu ons.
•Hidden laye . The hidden laye has he same numbe o neu ons as he desi ed
numbe o classes o be classified in he ou pu laye . As an example, his laye should
consis o eigh LIF neu ons i eigh diffe en audio samples a e expec ed o be
classified.
•Ou pu laye . As he p e ious laye , his also has as many neu ons as ou pu classes.
The fi ing ou pu o he neu ons in his laye will de e mine he esul o he classi‐
fica ion.
Figu e 4 shows he SNN a chi ec u e. Connec ions be ween laye s a e achie ed using
he F omLis Connec o me hod om PyNN, meaning ha he sou ce, des ina ion and
weigh o he connec ion a e specified manually. Using o he connec o s om his
package will esul on ha ing he same weigh in all he connec ions be ween consecu i e
laye s, ins ead o a diffe en alue o each. In his a chi ec u e, each neu on in a laye
is connec ed o e e y neu on in he nex laye , and he weigh alue is ob ained om he
aining s ep, which is desc ibed in Sec . 4. The h eshold ol age o he neu ons in he
hidden laye is 15 mV, while his ol age is 10 mV in he neu ons in he ou pu laye .
Decay a e and e ac o y pe iod a e he same o bo h laye s: 20 mV/ms and 2 ms,
espec i ely.
Fig. 4. SNN a chi ec u e using an audio sample aeda file as inpu .
4 T aining Phase o Audio Classifica ion
In he p e ious sec ion, each o he h ee laye s comp ising he ne wo k we e desc ibed.
The aining phase is pe o med offline and supe ised. The main objec i e o his
aining is o ob ain he weigh alues o he connec ions be ween he inpu and he
hidden laye and be ween he hidden and he ou pu laye s o u he audio samples
classifica ion. The e o e, wo diffe en aining s eps need o be done.
The weigh s o he fi s s ep o he aining phase a e ob ained om he no malized
spike fi ing ac i i y o each NAS channel using a se o audio samples simila o hose
o be ecognized (same ampli ude, du a ion and equencies). The fi ing a e o a specific
channel (FR
channel_i
) is ob ained by di iding he numbe o e en s p oduced in ha
channel by he NAS fi ing a e (FR
T
), which is he numbe o e en s fi ed in he NAS
in a pa icula ime pe iod.
(1)
(2)
Figu e 5 shows he no malized spike fi ing ac i i y o a se o eigh pu e ones wi h
equencies ha ange om 130.813 Hz o 1396.91 Hz, loga i hmically spaced.
The weigh s o he second s ep o he aining phase a e ob ained om he fi ing
ou pu o each neu on in he hidden laye when using he se o audio samples as inpu
a e loading he weigh s calcula ed in he p e ious s ep in o he connec ions be ween
he inpu and he hidden laye . These fi ing ou pu s a e no malized by di iding each o
hem by he maximum alue. The esul s ob ained a e he weigh alues ha will be used
in he connec ions be ween he hidden and he ou pu laye .
5Tes Scena io
In his wo k, he SNN a chi ec u e and aining algo i hm p esen ed a e es ed using
eigh diffe en audio samples. These ou pu classes co espond o eigh diffe en pu e
ones wi h equencies ha ange om 130.813 Hz o 1396.91 Hz, loga i hmically
spaced (130.813, 174.614, 261.626, 349.228, 523.251, 698.456, 1046.50 and
1396.91 Hz). These samples ha e a du a ion o 0.5 s and we e gene a ed using he
audioplaye unc ion om Ma lab wi h a sampling a e o 48 KHz and a peak- o-peak
ol age alue o 1 V. A e he signal is sen o he mixe , i p opaga es he sound o
NAS inpu and sends an AER s eam o he PC h ough he AER-USB in e ace. The
jAER so wa e unning on he PC is able o cap u e his s eam and sa e i as an aeda
file. Figu e 6 shows he cochleog ams o he 130.813 Hz and he 1396.91 Hz pu e ones
a e cap u ing hem.
The fi s s ep o he aining phase can be achie ed by applying he equa ions
p esen ed in Sec . 4 o he se o eigh aeda files co esponding o each pu e one. This
will gene a e a CSV file con aining he weigh s o he 64 × 8 connec ions be ween he
inpu and he hidden laye s o he SNN based on he fi ing a e o he spike s eams o
each audio sample. As desc ibed in he p e ious sec ion, loading hose weigh s in o he
co esponding connec ions and using he eigh pu e ones as inpu will esul on a fi ing
Fig. 5. No malized spike fi ing ac i i y o each NAS channel pe audio sample.
Fig. 6. Fi s 10 ms cochleog am o he 130.813 Hz (le ) and 1396.91 Hz ( igh ) pu e ones.
ou pu on he second laye neu ons ha will be used o aining he connec ions be ween
he second and he ou pu laye s o he SNN.
A e he weigh s a e se on hese connec ions, new se s o he same pu e ones (same
equencies) a e eco ded using diffe en Signal- o-Noise Ra io (SNR) alues and es ed
on he ne wo k, calcula ing he hi a e pe cen age o each class.
6 Expe imen al Resul s
Diffe en pu e one se s wi h he same equencies and p ope ies (0.5 s and 0.5 V
ampli ude) as he ones used in his wo k we e cap u ed and used o es he ne wo k
obus ness and effec i eness. A 100 % hi a e was ob ained o e e y class when he
signal was a pu e sine wa e. Mo eo e , he ne wo k has also been es ed by adding a
noise signal consis ing o andom alues o he pu e ones o iginal signals, ob aining
audio samples wi h diffe en SNR alues ( om 35.2 dB o 0 dB). The hi a e pe cen age
o e e y class using he p e ious SNR alues a e lis ed in Table 1.
Table 1. Hi a e pe cen age o he audio samples classifica ion SNN o diffe en SNR alues.
SNR
(dB)
Pu e one equency (Hz)
130.813 174.614 261.626 349.228 523.251 698.456 1046.5 1396.91
No noise 100 % 100 % 100 % 100 % 100 % 100 % 100 % 100 %
35.1993 100 % 100 % 100 % 100 % 100 % 100 % 100 % 100 %
21.3363 100 % 83 % 96 % 100 % 100 % 100 % 100 % 100 %
13.2273 100 % 81 % 92 % 100 % 100 % 100 % 100 % 96 %
7.4733 100 % 86 % 100 % 100 % 100 % 100 % 100 % 95 %
3.0103 74 % 90 % 100 % 98 % 100 % 100 % 100 % 98 %
2 93 % 88 % 20 % 32 % 16 % 92 % 32 % 97 %
1 10 %5 %0 %0 %0 %88 %26 %94 %
0 0 %0 %0 %0 %0 %76 %22 %91 %
The esul s show e y high hi a e pe cen ages when he SNR is abo e 3 dB.
Howe e , when he SNR alls below 3 dB and app oaches ze o dB ( he ampli ude o he
pu e one is he same as he ampli ude o he noise signal) he ne wo k is no able o
classi y e e y inpu signal as i s co esponding class.
7Conclusions
In his pape , a no el mul ilaye spiking neu al ne wo k a chi ec u e o audio samples
classifica ion implemen ed in SpiNNake has been p esen ed. To achie e his goal, an
op imized aining phase o audio ecogni ion has been desc ibed and specified in wo
diffe en s eps, which allow ob aining he weigh s o he connec ions be ween he inpu
and he hidden laye s and be ween he hidden and he ou pu laye s. The ne wo k was
ained using eigh pu e ones wi h equencies be ween 130.813 Hz and 1396.91 Hz
and es ed by adding a noise signal wi h SNR alues be ween 35.1993 and 0 dB.
The hi a e alues ob ained a e many es s confi m he obus ness o he ne wo k
and he aining, which make i possible o classi y e e y pu e one wi h a p obabili y
o e 74 % e en when he SNR alue is 3 dB, ob aining almos a 100 % p obabili y o
e e y inpu when he SNR is abo e ha alue.
Finally, he SpiNNake boa d has allowed o model and de elop a leaky in eg a e-
and-fi e spiking neu al ne wo k o his pu pose in an easy, as , use - iendly and effi‐
cien way, p o ing i s po en ial, and p omo ing and acili a ing he implemen a ion o
SNNs like hese in eal ha dwa e pla o ms. The PyNN code used o es he SNN
p esen ed in his wo k is a ailable a [19].
Acknowledgemen s. The au ho s would like o hank he APT Resea ch G oup o he Uni e si y
o Manches e o ins uc ing us in he SpiNNake . This wo k is suppo ed by he Spanish
go e nmen g an BIOSENSE (TEC2012-37868-C04-02) and by he excellence p ojec om
Andalusian Council MINERVA (P12-TIC-1300), bo h wi h suppo om he Eu opean Regional
De elopmen Fund.
Re e ences
1. Lich s eine , P., Posch, C., Delb uck, T.: A 128 × 128 120 dB 15 μs la ency asynch onous
empo al con as ision senso . IEEE J. Solid-S a e Ci c. 43, 566–576 (2008)
2. Chan, V., Liu, S.C., an Schaik, A.: AER EAR: a ma ched silicon cochlea pai wi h add ess
e en ep esen a ion in e ace. IEEE T ans Ci c. Sys . I 54(1), 48–59 (2007)
3. Häflige , P.: Adap i e WTA wi h an analog VLSI neu omo phic lea ning chip. IEEE T ans.
Neu al Ne w. 18, 551–572 (2007)
4. Indi e i, G., Chicca, E., Douglas, R.: A VLSI a ay o low-powe spiking neu ons and bis able
synapses wi h spike- iming dependen plas ici y. IEEE T ans. Neu al Ne w. 17, 211–221
(2006)
5. Jiménez-Fe nández, A., Jiménez-Mo eno, G., Lina es-Ba anco, A., e al.: Building blocks
o spikes signal p ocessing. In: In e na ional Join Con e ence on Neu al Ne wo ks, IJCNN
(2010)
6. Lina es-Ba anco, A., e al.: A USB3.0 FPGA e en -based fil e ing and acking amewo k
o dynamic ision senso s. In: P oceedings o IEEE In e na ional Symposium on Ci cui s
and Sys ems, pp. 2417–2420 (2015)
7. Lina es-Ba anco, A., Gomez-Rod iguez, F., Jimenez-Fe nandez, A., e al.: Using FPGA o
isuo-mo o con ol wi h a silicon e ina and a humanoid obo . In: IEEE In e na ional
Symposium on Ci cui s and Sys ems, pp. 1192–1195 (2007)
8. Jimenez-Fe nandez, A., Jimenez-Mo eno, G., Lina es-Ba anco, A., e al.: A neu o-inspi ed
spike-based PID mo o con olle o mul i-mo o obo s wi h low cos FPGAs. Senso s 12,
3831–3856 (2012)
9. Hamil on, T.J., Jin, C., an Schaik, A., Tapson, J.: An ac i e 2-D silicon cochlea. IEEE T ans.
Biomed. Ci c. Sys . 2, 30–43 (2008)
10. Jimenez-Fe nandez, A., Ce ezuela-Escude o, E., Mi o-Ama an e, L., e al.: A binau al
neu omo phic audi o y senso o FPGA: a spike signal p ocessing app oach. IEEE T ans.
Neu al Ne wo ks Lea n. Sys . 1(0) (2016)
11. Boahen, K.: Poin - o-poin connec i i y be ween neu omo phic chips using add ess e en s.
IEEE T ans. Ci c. Sys II Analog Digi Sig. P ocess. 47, 416–434 (2000)
12. Bekolay, T., e al.: Nengo: a Py hon ool o building la ge-scale unc ional b ain models.
F on Neu oin o m. 7, 48 (2014)
13. Goodman, D., B e e, R.: B ian: a simula o o spiking neu al ne wo ks in py hon. F on
Neu oin o m. 2, 5 (2008)
14. jAER Open Sou ce P ojec . h p://jae .wiki.sou ce o ge.ne
15. Be ne , R., Delb uck, T., Ci i -Balcells, A., Lina es-Ba anco, A.: A 5 Meps $100 USB2.0
add ess-e en moni o -sequence in e ace. IEEE In e na ional Symposium on Ci cui s and
Sys ems (2007)
16. Paink as, E., e al.: SpiNNake : A 1-W 18-co e sys em-on-chip o massi ely-pa allel neu al
ne wo k simula ion. IEEE J. Solid-S a e Ci c. 48, 1943–1953 (2013)
17. Da ison, A.P.: PyNN: a common in e ace o neu onal ne wo k simula o s. F on
Neu oin o m. 2, 11 (2008)
18. SpiNNake Home Page. h p://ap .cs.manches e .ac.uk/p ojec s/SpiNNake
19. Dominguez-Mo ales, J.P.: Mul ilaye spiking neu al ne wo k o audio samples classifica ion
using Spinnake Gi hub page. h ps://gi hub.com/jpdominguez/Mul ilaye -SNN- o -audio-
samples-classifica ion-using-SpiNNake