Full text
AER Spiking Neu on Compu a ion on GPUs:
The F ame- o-AER Gene a ion
M.R. López-To es, F. Diaz-del-Rio, M. Domínguez-Mo ales,
G. Jimenez-Mo eno, and A. Lina es-Ba anco
Depa men o A chi ec u e and Technology o Compu e s,
A . Reina Me cedes s/n, 41012, Uni e si y o Se ille, Spain
[email p o ec ed]
Abs ac . Neu o-inspi ed p ocessing ies o imi a e he ne ous sys em and may
esol e complex p oblems, such as isual ecogni ion. The spike-based philosophy
based on he Add ess-E en -Rep esen a ion (AER) is a neu omo phic in e chip
communica ion p o ocol ha allows o massi e connec i i y be ween neu ons.
Some o he AER-based sys ems can achie e e y high pe o mances in eal- ime
applica ions. This philosophy is e y di e en om s anda d image p ocessing,
which conside s he isual in o ma ion as a succession o ames. These ames
need o be p ocessed in o de o ex ac a esul . This usually equi es e y
expensi e ope a ions and high compu ing esou ce consump ion. Due o i s ela i e
you h, nowadays AER sys ems a e sho o cos -e ec i e ools like emula o s,
simula o s, es e s, debugge s, e c. In his pape he i s esul s o a CUDA-based
ool ocused on he unc ional p ocessing o AER spikes is p esen ed, wi h he aim
o helping in he design and es ing o il e s and buses managemen o hese
sys ems.
Keywo ds: AER, neu omo phic, CUDA, GPUs, eal- ime ision, spiking
sys ems.
1 In oduc ion
S anda d digi al ision sys ems p ocess sequences o ames om ideo sou ces, like
CCD came as. Fo pe o ming complex objec ecogni ion, sequences o
compu a ional ope a ions mus be pe o med o each ame. The compu a ional
powe and speed equi ed make i di icul o de elop a eal- ime au onomous sys em.
Howe e , b ains pe o m powe ul and as ision p ocessing using millions o small
and slow cells wo king in pa allel in a o ally di e en way. Vision sensing and objec
ecogni ion in b ains a e no p ocessed ame by ame; hey a e p ocessed in a
con inuous way, spike by spike, in he b ain-co ex. The isual co ex is composed o
a se o laye s [1], s a ing om he e ina, which cap u es he in o ma ion. In ecen
yea s, signi ican p og ess has been made in he s udy o he p ocessing by he isual
co ex. Many a i icial sys ems ha implemen bio-inspi ed so wa e models use
biological-like p ocessing ha ou pe o ms mo e con en ionally enginee ed machines
[2][3][4]. Howe e , hese sys ems gene ally un a ex emely low speeds because he
models a e implemen ed as so wa e p og ams in classical CPUs. Nowadays,
a g owing numbe o esea ch g oups a e implemen ing some o hese
compu a ional p inciples on o eal- ime spiking ha dwa e h ough he so-called AER
(Add ess E en Rep esen a ion) echnology, in o de o achie e eal- ime p ocessing.
AER was p oposed by he Mead lab in 1991 [5][7] o communica ing be ween
neu omo phic chips wi h spikes. In AER, a sende de ice gene a es spikes. Each
spike ansmi s a wo d ep esen ing a code o add ess o ha pixel. In he ecei e ,
spikes a e p ocessed and inally di ec ed o he pixels whose code o add ess was on
he bus. In his way, cells wi h he same add ess in he emi e and ecei e chips a e
i ually connec ed by s eams o spikes. Usually, hese AER ci cui s a e buil using
sel - imed asynch onous logic [6]. Se e al wo ks ha e implemen ed spike-based
isual p ocessing il e s. Se ano e al. [10] p esen ed a chip-p ocesso ha is able o
implemen image con olu ion il e s based on spikes, which wo ks wi h a e y high
pe o mance (~3 GOPS o 32x32 ke nel size) compa ed o adi ional digi al ame-
based con olu ion p ocesso s [11]. Ano he app oach o sol ing ame-based
con olu ions wi h e y high pe o mances a e he Con Ne s [12][13], based on
cellula neu al ne wo ks, ha a e able o achie e a heo e ical sus ained 4 GOPS o
7x7 ke nel sizes.
One o he goals o he AER p ocessing is a la ge mul i-chip and mul i-laye
hie a chically s uc u ed sys em capable o pe o ming complica ed a ay da a
p ocessing in eal ime. Bu his pu pose s ongly depends on he a ailabili y o obus
and e icien AER in e aces and e alua ion ools [9]. One such ool is a PCI-AER
in e ace ha allows no only eading an AER s eam in o a compu e memo y and
displaying i on sc een in eal- ime, bu also he opposi e: om images a ailable in he
compu e ’s memo y, gene a e a syn he ic AER s eam in a simila manne a dedica ed
VLSI AER emi e chip [8][9][4] would do. This PCI-AER in e ace is able o each
up o 10Me en s/sec bandwid h, which allows a ame- a e o 2.5 ames/s wi h an
AER a ic load o 100% o 128x128 ames, and 25 ames/s wi h a ypical 10%
AER a ic load.
Nowadays compu e -based sys ems a e inc easing hei pe o mances exploi ing
a chi ec u al concep s like mul i-co e and many-co e. Mul i-co e is e e ed o hose
p ocesso s ha can execu e in pa allel as many h eads as co es a e a ailable on
ha dwa e. On he o he hand, many-co e compu e s consis o p ocesso s ( ha could
be mul i-co e) plus a co-p ocessing ha dwa e composed o se e al p ocesso s, like he
G aphical P ocessing Uni s (GPU). By joining a many-co e sys em wi h he
men ioned PCI-AER in e ace, a spike-based p ocesso could be implemen ed.
In his wo k we ocus on he e alua ion o GPU o un in pa allel an AER sys em,
s a ing wi h he ame- o-AER so wa e con e sion and con inuing wi h se e al usual
AER s eaming ope a ions, and a i s emula ion o a silicon e ina using CCD
came as. Fo his ask moni o ing in e nal GPU pe o mance h ough he Compu e
Visual P o ile [17] o NVIDIA® echnology was used. Compu e Visual P o ile is a
g aphical use in e ace based p o iling ool ha can be used o measu e pe o mance.
Th ough he use o his ool, we ha e ound se e al key poin s o achie e maximum
pe o mance o AER p ocessing emula ion.
Nex sec ion b ie ly explains he p os and cons o Simula ion and emula ion o
Spiking sys ems, when compa ed o eal implemen a ion. In sec ion 3, a basic
desc ip ion o N idia CUDA (Compu e Uni ied De ice A chi ec u e) a chi ec u e is
shown, ocusing on he ele an poin s ha a ec he pe o mance o ou ool. In
sec ion 4 he main pa s o he p oposed emula ion/simula ion ool a e discussed. In
sec ion 5 he pe o mance o he ope a ions o di e en so wa e o ganiza ion is
e alua ed, and in sec ion 6 he conclusions a e p esen ed.
2 Simula ion and Emula ion o Spiking Sys ems
Nowadays, he e is a ela i ely high numbe o spiking p ocessing sys ems ha use
FPGA o ano he dedica ed in eg a ed ci cui s. Needless o say ha , while hese
implemen a ions a e ime e ec i e, hey p esen many design di icul ies: ele a ed
ime and cos o digi al syn hesis, icky and cumbe some es ing, possibili y o wi ing
bugs (which o en a e di icul o de ec ), e c. Suppo ools a e hen e y help ul and
demanded by ci cui designe s. Usually he e a e se e al le els in he ield o
emula ion and simula ion ools: om he lowe (elec ical) le el o he mos
unc ional one. The e is no doub ha none o hese ools can co e he whole
“simula ion spec um”. In his pape we p esen a ool ocused on he unc ional
p ocessing o AER spikes, in o de o help in he design and es ing o il e s and
buses managemen o hese sys ems.
CPU is he mos simple and s aigh o wa d simula ion, and he e a e al eady some
o hese simula o s [25][26]. Bu simula ion imes a e cu en ly a om wha one can
wish o : a s eam p ocessing ha las s mic oseconds in a eal sys em can ex end o
minu es on hese simula o s e en when unning in high pe o mance clus e s.
Ano he hu dle ha is common in CPU emula o s is hei dependence on inpu (o
a ic) load. In a p e ious wo k [14], he pe o mance o se e al ame- o-AER
so wa e con e sion me hods o eal- ime ideo applica ions was e alua ed, by
measu ing execu ion imes in se e al p ocesso s. Tha wo k demons a ed ha o low
AER a ic loads any me hod in any mul ico e single CPU achie ed eal- ime, bu o
high bandwid h AER a ics, i depends on which me hod and CPU a e selec ed in
o de o ob ain eal- ime. This p oblem can be mi iga ed when using GPGPUs
(Gene al Pu pose GPUs).
A ew yea s ago GPU so wa e de elopmen was di icul and close o biza e [23].
Du ing he las yea s, wi h he expansion o GPGPUs, ools, unc ion lib a ies, and
ha dwa e abs ac ion mechanisms ha hide he GPU ha dwa e om de elope s, ha e
appea ed. Nowadays he e a e eliable debugging and p o iling ools like hose om
CUDA [18]. Two addi ional ones complemen hese ad an ages. Fi s , he e y low
cos pe p ocessing uni , which is supposed o pe sis , since he PC g aphics ma ke
subsidizes GPUs (only N idia has al eady sold 50 million CUDA-capable GPUs).
Secondly, he annual g ow h pe o mance a io is p edic ed o s ay e y high: by 70%
pe yea due o he con inuing minia u iza ion. The same happens o Memo y
Bandwid h. Fo example, ex an GPUs wi h 240 loa ing poin a i hme ic co es,
ealizing a pe o mance o 1 TFLOPS wi h one chip. One o he consequences o his
ad an age o i s compe i o s, is ha he op leading supe compu e (No embe 2010)
is a GPU based machine (www. op500.o g).
P e ious easons ha e pushed he scien i ic communi y o inco po a e GPUs in
se e al disciplines. In ideo p ocessing sys ems, GPUs applica ion is ob ious,
he e o e some e y in e es ing AER emula ion sys ems ha e begun o appea
[15][16].
The las ac ega ds he need o spiking ou pu came as. These came as (usually
called silicon e inas) a e cu en ly expensi e, a e and inaccu a e (due o elec ical
misma ching be ween hei cells [22]). Mo eo e , hei ac ual esolu ion is e y low,
making i impossible o wo k wi h eal ele a ed add ess s eams a p esen . This
hu dle is expec ed o be o e come in nex yea s. Howe e , cu en esea che s canno
e alua e hei spike p ocessing sys ems o high esolu ions. The e o e, a cheap CCD
came a whe e a p ep ocessing module gene a es a spiking s eaming is a e y
a ac i e al e na i e [20], e en i he eal silicon e ina speed canno be eached. And
we belie e GPUs may be he mos well-placed pla o m o do i . Emula ing a e ina
by a GPU has addi ional bene i s: on he one hand, di e en add ess coding can be
easily implemen ed [24]. On he o he hand, di e en ypes o silicon e inas could be
emula ed o e he same GPU pla o m, simply by changing he p ep ocessing il e
execu ed be o e gene a ing he spiking s eam. Nowadays implemen ed e inas a e
mainly o h ee ypes [22]: g ay le el, spa ial con as (g adien ) and dynamic ision
senso (di e en ial) e inas. The i s one can be emula ed by gene a ing spikes o
each pixel [8], he second op ion h ough a g adien and he las one, wi h a disc e e
ime de i a i e.
In his wo k he balance be ween gene ali y and e iciency o emula ion is
conside ed. Tuning simula o ou ines wi h a GPU o each a good pe o mance
in ol es a loss o gene ali y [21].
3 CUDA A chi ec u e
A CUDA GPU [18] includes an a ay o s eaming mul ip ocesso s (SMs). Each SM
consis s o 8 loa ing-poin Scala P ocesso s (SPs), a Special Func ion Uni , a mul i-
h eaded ins uc ion uni , and se e al pieces o memo y. Each memo y ype is
in ended o a di e en use; in mos cases hey ha e o be managed by he
p og amme . This makes he p og amme o be awa e o hei esou ces, achie ing
maximum pe o mance. Each SM has a “wa p schedule ” ha selec s a any cycle a
g oup o h eads o execu ion, in a ound- obin ashion. A wa p is simply a g oup o
(cu en ly) 32 ha dwa e-managed h eads. As he numbe and ypes o h eads may be
eno mous, a i e dimension o ganiza ion is suppo ed by CUDA, wi h wo impo an
le els: a g id con ains blocks ha mus no be e y coupled, while e e y block
con ains a ela i ely sho numbe o h eads, which can coope a e deeply.
I a h ead in a wa p issues a cos ly ope a ion (like an ex e nal memo y access),
hen he wa p schedule swi ches o a new wa p, in o de o hide he la ency o he
o he h ead. In o de o use he GPU esou ces e icien ly, each h ead should ope a e
on di e en scala da a, wi h a ce ain pa e n. Due o hese special CUDA ea u es,
some key poin s mus be kep in mind o achie e a good pe o mance o an AER
sys em emula o . These a e basically: I is a mus o launch housands o millions o
e y ligh h eads; he pa e n memo y access is o i al impo ance ( he CUDA
manual [18] p o ides de ailed algo i hms o iden i y ypes o coalesced/uncoalesced
memo y accesses); he e is a conside able amoun o spa ial locali y in he image
access pa e n equi ed o pe o m a con olu ion (GPU ex u e memo y is used in his
case). Any ype o bi u ca ion (b anch, loops, e c.) in he h ead code should be
a oided. The same o any ype o h ead synch oniza ion, c i ical sec ions, ba ie s,
a omic accesses, e c. ( his means ha each h ead mus be almos independen om
he o he s); Ne e heless, because GPU do no usually ha e ha dwa e-managed
caches he e will be no p oblem wi h alse dependencies (as usual in mul ico e
sys ems when e e y h ead has o w i e in he same ec o as he o he s).
4 Main Modules o he Emula ion/Simula ion Tool
An AER sys em emula ion/simula ion ool mus con ain a mos he ollowing pa s:
-Images Inpu module. I can include a p ep ocessing il e o emula e di e en
e inas.
-Syn he ic AER spikes gene a ion. I is an impo an pa since he dis ibu ion
o spikes h oughou ime mus be simila o ha p oduced by eal e inas.
-Fil e s. Con olu ion ke nels a e he basic ope a ion, bu o he s like low pass
il e s, in eg a o s, winne akes all, e c. may be necessa y.
-Buses managemen . I includes buses spli e s, me ges, and so on.
-Resul ou pu module. I mus collec he esul s in ime o de o send hem o
he CPU.
-AER Bus P obes. This module appea s necessa ily as discussed below.
The ool p esen ed he e is in ended o emula e a spiking p ocessing ha dwa e sys em.
As a esul , he algo i hms o be ca ied ou in he ool a e gene ally simple. In
GPGPU e minology, his means ha he “a i hme ic in ensi y” [18] (which is de ined
as he numbe o ope a ions pe o med pe wo d o memo y ans e ed) is going o
be e y low. The e o e, op imisa ion mus ocus on memo y accesses and ypes. As a
i s consequence some es ic ions on he numbe and size o he da a objec s we e
done. Besides, a second conclusion is p esen ed: ins ead o simula ing se e al AER
il e s in cascade (as usual in FPGA p ocessing), i will be usually be e o execu e
only a combined il e ha uses hem, in o de o sa e GPU DDRAM accesses o
CPU-GPU ansac ions. This is o be discussed in nex sec ions, acco ding o he
esul s. Finally he concep o AER Bus P obe is in oduced he e in o de o only
gene a e he in e media e AER alues ha a e s ic ly necessa y. Only when a p obe
is demanded by he use , a GPU o CPU ansac ion is inse ed o collec AER spikes
in he empo al o de . Besides, some code adjus men s ha e been in oduced o a oid
new da a s uc u es despi e o adding mo e compu a ion.
Syn he ic AER gene a ion is one o he key pieces o an AER ool. This is because
i usually las s a conside able ime and because he spike dis ibu ion should ha e a
conside able ime uni o mi y o ensu e ha neu on in o ma ion is co ec ly sen [8].
Besides, in his wo k wo addi ional easons ha e o be conside ed. Fi s ly, i has o be
demons a ed ha a high deg ee o pa allelism can be ob ained using CUDA, so ha
he mo e co es he GPU has, he less ime he ame akes o be gene a ed. And
secondly, we ha e pe o med a compa ison o execu ion imes wi h hose ob ained o
he mul ico e pla o ms used p e iously in [14].
In [14] hese AER so wa e me hods we e e alua ed in se e al CPUs ega ding he
execu ion ime. In all AER gene a ion me hods, esul s a e sa ed in a sha ed AER
spike ec o . Ac ually, spike ep esen a ion can be done in se e al o ms (in [24] a
wide codi ica ion spec um is discussed). Taking in o accoun he conside a ion o
p e ious sec ion, we ha e concluded ha he AER spike ec o ( he one used in [14])
is e y con enien when using GPUs.
One can hink o many so wa e algo i hms o ans o m a bi map image (s o ed in
a compu e ’s memo y) in o an AER s eam o pixel add esses [8]. In all o hem he
equency o appea ance o he add ess o a gi en pixel mus be p opo ional o he
in ensi y o ha pixel. No e ha he p ecise loca ion o he add ess pulses is no
c i ical. The pulses can be sligh ly shi ed om hei nominal posi ions; he AER
ecei e s will in eg a e hem o eco e he o iginal pixel wa e o m.
Wha e e algo i hm is used, i will gene a e a ec o o add esses ha will be sen
o an AER ecei e chip ia an AER bus. Le us call his ec o he “ ame ec o ”.
The ame ec o has a ixed numbe o ime slo s o be illed wi h e en add esses.
The numbe o ime slo s depends on he ime assigned o a ame ( o example
T ame=40 ms) and he ime equi ed o ansmi a single e en ( o example
Tpulse=10 ns). I we ha e an image o N×M pixels and each pixel can ha e a g ey
le el alue om 0 o K, one possibili y is o place each pixel add ess in he ame
ec o as many imes as he alue o i s in ensi y, and dis ibu e i wi h equidis an
posi ions. In he wo s case (all pixels wi h maximum alue K), he ame ec o
would be illed wi h N×M×K add esses. No e ha his numbe should be less han he
o al numbe o ime slo s in he ame ec o . Depending on he o al in ensi y o he
image he e will be mo e o less emp y slo s in he ame ec o T ame/Tpulse.
Each algo i hm would implemen a pa icula way o dis ibu ing hese add ess
e en s, and will equi e a ce ain ime. In [8] and [14] we discussed se e al algo i hms
ha we e con enien when using classical CPUs. Bu i GPUs a e o be used, we ha e
o disca d hose whe e he gene a ion o each elemen o ame ec o canno be
independen om he o he s. This clea ly happens in hose me hods based on Linea
Feedback Shi Regis e s (LFSR): as he me hod equi es calling a andom unc ion
ha always depends on i sel , he me hod canno be di ided in h eads.
To sum up, he bes -sui ed me hod o GPU p ocessing is he so-called Exhaus i e
me hod. This algo i hm di ides he add ess e en sequence in o K slices o N
×
M
posi ions o a ame o N
×
M pixels wi h a maximum g ay le el o K. Fo each slice
(k), an e en o pixel (i,j) is sen on ime i he ollowing condi ion is asse ed:
KPKPk jiji ≥+⋅ ,, mod)( and jMikMN =+⋅−+−⋅⋅ )1()1(
whe e Pi,j is he in ensi y alue o he pixel (i,j).
The Exhaus i e me hod ies dis ibu ing he e en s o each pixel in equidis an
slices. In his me hod, he e is a e y impo an ad an age when using CUDA:
elemen s o ame ec o can be sequen ially p ocessed, because he second condi ion
abo e can be implemen ed using as he coun e o he ame ec o ( ha is, he
h ead index in CUDA e minology). This means ha se e al accesses (pe o med by
di e en h eads) can be coalesced o sa e DDRAM access ime. The esul s sec ion
is based on his algo i hm.
Ne e heless, o he algo i hms could be sligh ly ans o med o adap hem o
CUDA. The mos a o able case is ha o he Random-Squa e me hod. While his
me hod equi es he gene a ion o pseudo andom numbe s (which a e gene a ed by
wo LFSR o 8 and 14 bi s), LFSR unc ions can be a oided i all he lis s o numbe s
a e s o ed in ables. Al hough his is possible, i will in ol e wo addi ional accesses
pe elemen o long ables, which p obably will eside in DDRAM. This will add a
supplemen a y delay ha is a oided wi h he exhaus i e me hod.
The es o me hods a e mo e di icul o ine- une o CUDA (namely he Uni o m,
Random and Random-Ha dwa e me hods) because o he a o emen ioned easons.
The e a e ano he g oup o so wa e me hods dedica ed o manage he AER buses.
Fo una ely, hese ope a ions a e in insically pa allel, since in ou case hey basically
consis o a p ocessing o each o he ame ec o elemen s (which plays he ole o a
comple e AER s eam).
The o he ope a ions in ol ed in AER p ocessing a e hose ha play he ole o image
il e ing and con olu ion ke nels. Nowadays, he algo i hms ha execu e hem on AER
based sys ems a e in insically sequen ial: commonly o e e y spike ha appea s in he
bus, he alues o some coun e s ela ed o his spike add ess a e changed [15][10].
The e o e, hese coun e s mus be seen as c i ical sec ions when emula ing his p ocess
h ough so wa e. This makes imp ac ical o emula e his ope a ion in a GPU. On he
con a y, he s anda d ame con olu ion ope a ion can be easily pa allelized [19] i he
image ou pu is placed in a memo y zone di e en om ha o image inpu . Execu ion o
a s anda d con olu ion gi es an eno mous speedup when compa ing o a CPU. Due o
his and conside ing ha he o he g oup o ope a ions does p esen a high deg ee o
pa allelism, in his pape con olu ions a e p ocessed in a classical ashion. None heless,
his combina ion o AER-based and classical ope a ions esul s in a good enough
pe o mance as seen in he ollowing sec ion.
5 Pe o mance S udy
In o de o analyze he pe o mance and scalabili y o CUDA simula ion and
emula ion o AER sys ems, a se ies o ope a ions ha e been coded and analyzed in
wo N idia GPUs. Table 1 summa izes he main cha ac e is ics o he pla o ms
es ed. The second GPU ha e an impo an ea u e: i can concu en ly copy and
execu e p og ams (while he i s one canno ).
Table 1. Tes ed N idia GPUs
Cha ac e is ics GeFo ce 9300 ION GTX 285
Global memo y 266010624 by es 1073414144 by es
Maximum numbe o h eads pe block 512 h eads 512 h eads
Mul ip ocesso s x Co es/MP 2 x 8 = 16 Co es 30 x 8 = 240 Co es
Clock a e 1.10 GHz 1.48 GHz
Fig. 1 depic ed a ypical AER p ocessing scena io. The i s module ‘ImageToAER’
ans o ms a ame in o an AER s eam using he Exhaus i e me hod. The ‘Spli e ’
di ides spikes in o wo buses acco ding o an add ess mask ha ep esen s he li le clea
squa e in he almos black igu e. The uppe bus is hen o a ed 90 deg ees simply by
going h ough he ‘Mappe ’ module, which changes each spike’s add ess in o ano he .
Finally a me ge be ween he o iginal image and he uppe bus gi es a new emula ed AER
bus, which can be obse ed wi h a con enien AERToImage module (which, in a ew
wo ds, makes a empo al in eg a ion). A consequence o he use o a igid size ame
ec o composed o ime slo s is ha a me ge ope a ion be ween wo ull buses canno
use pe ec ly he wo images ep esen ed in he ini ial buses. In ou implemen a ion, i
bo h buses ha e a alid spike in a ce ain ime slo , he co esponding ou pu bus slo is
going o be illed only by one o he inpu buses (which is decided wi h a simple ci cui ).
This aspec appea s also in ha dwa e AER implemen a ions, whe e an a bi e mus
decide which inpu bus “looses”, and hen i s spike does no appea in he me ged bus
[8][10].
Fig. 1. Cascade AER ope a ions benchma k
Acco ding o sec ion 4, h ee kinds o benchma ks ha e been ca ied ou . In a i s
g oup, an AER cascade ope a ions we e simula ed: in he middle o wo ope a ions,
( ha is, in an AER bus) an in e media e esul is collec ed by he CPU o check and
e i y he bus alues. In he second g oup, ope a ions a e execu ed in cascade, bu
ansi ional alues a e no “downloaded” o he CPU, hus p ese ing GPU o CPU
ansac ions. This implies ha no AER Bus P obe modules a e p esen , hus inhibi ing
inspec ion oppo uni ies. Finally a hi d collec ion is no coded in p e ious modula
ashion since he i s ou ope a ions a e g ouped oge he in one CUDA ke nel ( he
same h ead execu es all o hem sequen ially). This a oids se e al GPU DDRAM
accesses, sa ing an eno mous execu ion ime in he end. The i h ope a ion ha
ans o ms an AER s eam in o an image ame canno be easily pa allelized o he
same easons desc ibed in p e ious sec ion o con olu ion ke nels. Timing
compa ison o hese h ee g oups is summa ized in Table 2.
I is impo an o ema k ha he execu ion ime o he hi d g oup almos
coincides wi h he maximum execu ion ime o all he ope a ions in he i s g oup.
Ano he ob ious bu in e es ing ac is ha ansac ional imes a e p opo ionally
educed when CPU-GPU ansac ions a e elimina ed. And speed-up be ween GTX285
and ION 9300 is also nea o he ideal. One can conclude ha scalabili y o ou ool
is good, which means o as e upcoming GPU a sho e execu ion ime is expec ed.
Finally, con as e ina emula ion has been ca ied ou : o a ame, i s a g adien
con olu ion is done in o de o ex ac image edges, and secondly, he AER ame
ec o is gene a ed (in he same CUDA h ead). A hope ul esul is ob ained: he
mean execu ion ime o p ocess one ame is 313.3 μs, ha is, almos 3200 ames pe
second. As he execu ion imes a e small we can suppose ha hese imes could be
o e lapped wi h he ansac ion ones. The esul ing ps a io is e y much highe
(a ound 50x) han hose ob ained using mul ico e CPUs in p e ious s udies o AER
spikes gene a ion me hods [14].
Table 2. Benchma king imes o GPU GTX285 and o 9300 ION. Times in mic oseconds.
Measu ed Time Fi s G oup
GTX285
Second G oup
GTX285
Thi d G oup
GTX285
Thi d G oup
9300
CPU o GPU
ansac ions 8649.2 360.3 125.5 5815.5
CUDA ke nel
execu ion 2068.6 1982.7 781.6 15336.8
GPU o CPU
ansac ions 11762.8 2508.9 2239.2 12740.8
To al Time 22480.5 4851.9 3146.3 33893.0
Th ough hese expe imen s we ha e demons a ed ha majo ime is spen in hese
ypes o AER ools in ex e nal DDRAM GPU accesses, since da a sizes a e necessa ily
big while algo i hms can be implemen ed wi h a ew ope a ions pe CUDA h ead. This
conclusion gi es us an oppo uni y o de elop a comple ely unc ional AER simula o . A
second consequence de i ed om his is ha he size o he image can in oduce a
conside able inc emen o ime emula ion. In his wo k, he chosen size (128 × 128
pixels) esul s in a ame ec o o 8 MB (4 Me en s × 2 by es/e en ). Howe e , a
512 × 512 pixel image will ha e 4x he size image, plus wice he by es pe add ess (i no
comp ession is implemen ed). This means an 8x o al size, and hen an 8x access ime.
To sum up, elimina ing he es ic ions on he numbe and size o da a objec s can ha e
an impo an impac on he ool pe o mance.
6 Conclusions and Fu u e Wo k
A CUDA-based ool ocused on he unc ional p ocessing o AER spikes and i s i s
iming esul s a e p esen ed. I in ends o emula e a spiking p ocessing ha dwa e
sys em, using simple algo i hms wi h a high le el o pa allelism. Th ough
expe imen s, we demons a ed ha majo ime is spen in DDRAM GPU accesses, so
some es ic ions on he numbe and size o he da a objec s ha e been done. A
second esul is p esen ed: ins ead o simula ing se e al AER il e s in cascade (as
usual in FPGA p ocessing), i is be e o execu e only a combined il e ha uses
hem, in o de o sa e GPU DDRAM accesses and CPU-GPU ansac ions. Due o he
p omising iming esul s, he immedia e u u e wo k comp ises a ully emula ion o an
AER e ina using a classical ideo came a. Running ou expe imen s on a mul iGPU
pla o m is ano he demanding ex ension because o he scalabili y o ou ool.
Acknowledgmen s. This wo k was suppo ed by he Spanish Science and Educa ion
Minis y Resea ch P ojec s TEC2009-10639-C04-02 (VULCANO).
Re e ences
1. D ubach, D.: The B ain Explained. P en ice-Hall, New Je sey (2000)
2. Lee, J.: A Simple Speckle Smoo hing Algo i hm o Syn he ic Ape u e Rada Images.
IEEE T ans. Sys ems, Man and Cybe ne ics SMC-13, 85–89 (1983)
3. C immins, T.: Geome ic Fil e o Speckle Reduc ion. Applied Op ics 24 (1985)
4. Lina es-Ba anco, A., e al.: On he AER Con olu ion P ocesso s o FPGA. In: ISCAS
2010, Pa is, F ance (2010)