An MPI-CUDA implemen a ion o an imp o ed Roe
me hod o wo-laye shallow wa e sys ems
Ma c de la Asunci´ona, Jos´e M. Man asa, Manuel J. Cas ob, E. D.
Fe n´andez-Nie oc
aDp o. Lenguajes y Sis emas In o m´a icos, Uni e sidad de G anada
bDp o. An´alisis Ma em´a ico, Uni e sidad de M´alaga
cDp o. Ma em´a ica Aplicada I, Uni e sidad de Se illa
Abs ac
The nume ical solu ion o wo-laye shallow wa e sys ems is equi ed o simu-
la e accu a ely s a i ied luids, which a e ubiqui ous in na u e: hey appea
in a mosphe ic lows, ocean cu en s, oil spills, . . . Mo eo e , he implemen-
a ion o he nume ical schemes o sol e hese models in ealis ic scena ios
imposes huge demands o compu ing powe . In his pape , we ackle he acce-
le a ion o hese simula ions in iangula meshes by exploi ing he combined
powe o se e al CUDA-enabled GPUs in a GPU clus e . Fo ha pu pose,
an imp o emen o a pa h conse a i e Roe ype ini e olume scheme which
is specially sui able o GPU implemen a ion is p esen ed, and a dis ibu ed
implemen a ion o his scheme which uses CUDA and MPI o exploi he
po en ial o a GPU clus e is de eloped. This implemen a ion o e laps MPI
communica ion wi h CPU-GPU memo y ans e s and GPU compu a ion o
inc ease e iciency. Se e al nume ical expe imen s pe o med on a clus e o
mode n CUDA-enabled GPUs show he e iciency o he dis ibu ed sol e .
Keywo ds: Shallow wa e simula ion, GPU clus e compu ing, ini e
olume schemes, uns uc u ed meshes, CUDA, MPI
P ep in submi ed o Jou nal o Pa allel and Dis ibu ed Compu ing Ma ch 4, 2011
1. In oduc ion
The wo-laye shallow wa e sys em has been used as he nume ical model
o simula e se e al phenomena ela ed o s a i ied geophysical lows such
as a mosphe ic lows, ocean cu en s, oil spills o sunamis gene a ed by
unde wa e landslides. The simula ion o hese phenomena gi es o place
o e y long las ing simula ions in big compu a ional domains. Mo eo e ,
some o hese phenomena ( sunami p opaga ion o oil spills) could equi e
eal ime calcula ion. The e o e, ex emely e icien nume ical schemes and
implemen a ions a e needed o be able o analyze hose p oblems in p ac ical
execu ion imes. Since he nume ical schemes o sol e shallow wa e sys ems
usually exhibi a high deg ee o po en ial pa allelism, he de elopmen o
pa allel e sions o hese schemes o high pe o mance pla o ms seems o be
a sui able way o achie ing he equi ed pe o mance in ealis ic applica ions.
A cos e ec i e way o ob aining a subs an ially highe pe o mance in
hese applica ions consis s in using G aphics P ocesso Uni s (GPUs). These
a chi ec u es make i possible o ob ain pe o mances ha a e o de s o mag-
ni ude as e han a s anda d CPU and a e g owing in popula i y among he
scien i ic and enginee ing communi y [1, 2]. Mo eo e , se e al GPU p o-
g amming oolki s such as CUDA [3] ha e been de eloped o acili a e he
p og amming o GPUs o gene al pu pose applica ions.
The e a e p e ious p oposals o po ini e olume one-laye shallow wa e
sol e s o a GPU by using a g aphics-speci ic p og amming language [4, 5, 6],
bu cu en ly mos o he p oposals o simula e shallow lows on a single
GPU a e based on he CUDA p og amming model. A CUDA sol e o one-
2
laye sys em based on he i s o de ini e olume scheme p esen ed in [7]
is desc ibed in [8] o deal wi h s uc u ed egula meshes. The ex ension
o his CUDA sol e o wo-laye shallow wa e sys em is p esen ed in [9].
The e also exis p oposals o implemen , using CUDA-enabled GPUs, high
o de schemes o simula e one-laye sys ems [10, 11, 12] and o implemen
i s -o de schemes o one and wo-laye sys ems on iangula meshes [13].
Al hough he use o single GPU sys ems makes i possible o sa is y he
pe o mance equi emen s o se e al applica ions which a e no compu a io-
nally expensi e, his si ua ion is no equen in mos ealis ic applica ions.
Many applica ions equi e o handle huge meshes and la ge numbe o ime
s eps. Mo eo e , in some applica ions eal ime accu a e p edic ions ( o
ins ance, o app oxima e he e ec o an unexpec ed oil spill) could be e-
qui ed. The cha ac e is ics o hese applica ions sugges o combine he
powe o mul iple GPUs o sa is y he pe o mance equi emen s.
One app oach o use se e al GPUs in hese p oblems is based on p o-
g amming sha ed memo y mul i-GPU desk op sys ems. These pla o ms
ha e been used o accele a e conside ably luid dynamic [14] and shallow wa-
e [15] simula ions by combining sha ed memo y p og amming p imi i es o
manage h eads in CPU and CUDA o p og am he GPU. Al hough his isa
cos -e ec i e app oach, hese pla o ms only o e a educed numbe o GPUs
and mo e lexible sys ems a e desi able o answe o he g owing pe o mance
equi emen s o many ealis ic applica ions.
A mo e lexible app oach o ob ain he equi ed pe o mance in ol es o
use clus e s o GPU-enhanced compu e s whe e each node is equipped wi h a
single GPU o wi h a mul i-GPU sys em. The compu a ion on GPU clus e s
3
could make i possible o scale he educ ion in execu ion ime acco ding
o he numbe o GPUs (which can be easily inc eased). The e o e, his
app oach is mo e lexible han using a mul i-GPU desk op sys em and he
memo y limi a ions o a GPU-enhanced node can be o e come by sui ably
dis ibu ing he da a among he nodes, enabling us o simula e signi ican ly
la ge ealis ic models and wi h g ea e p ecision.
The use o GPU clus e s o accele a e da a in ensi e compu a ions is gai-
ning in popula i y [16, 17, 18, 19]. In [20], a scheme o sol e one-laye shallow
wa e sys ems is implemen ed on a GPU clus e o eal- ime sunami simu-
la ion. Mos o he p oposals o exploi GPU clus e s in scien i ic compu ing
use MPI [21] o implemen he communica ion among he p ocesses o he
dis ibu ed sys em and CUDA [3] o p og am he GPU (o GPUs) associa ed
o each node. A common way o educe he emo e communica ion o e head
in hese dis ibu ed implemen a ions consis s in using non-blocking commu-
nica ion MPI unc ions o o e lap he da a ans e s be ween nodes o he
clus e wi h he GPU compu a ion and he CPU-GPU da a ans e s.
This wo k deals wi h he accele a ion o he nume ical solu ion o wo-
laye shallow wa e sys ems by exploi ing he pa alleliza ion o an imp o ed
ini e olume scheme o uns uc u ed meshes on GPU clus e s. Fo ha
pu pose, a dis ibu ed implemen a ion o his scheme o a GPU clus e is
de eloped by using MPI and CUDA. This implemen a ion inco po a es an
e icien managemen o he dis ibu ed uns uc u ed mesh and mechanisms
o o e lap compu a ion wi h communica ion.
The ou line o he a icle is as ollows: he nex sec ion desc ibes he
unde lying ma hema ical model and p esen s an imp o emen o a i s o de
4
Roe ype ini e olume scheme, called IR-Roe scheme. Sec ion 3 desc ibes
a da a pa allel e sion o he IR-Roe scheme. The e iciency o se e al im-
plemen a ions o he IR-Roe scheme and he classical Roe scheme [7, 22] is
compa ed in Sec ion 4. In he wo nex sec ions we desc ibe a single and a
mul i-GPU dis ibu ed implemen a ion, espec i ely, o he me hod o ian-
gula meshes using he CUDA amewo k. Sec ion 7 shows he expe imen al
esul s ob ained when he implemen a ions a e applied o sol e an in e nal
dam b eak p oblem on a clus e o 4 NVIDIA Fe mi GPUs. Finally, Sec ion
8 summa izes he main conclusions and p esen s he u u e wo k.
2. An e icien nume ical scheme
2.1. The wo-laye shallow wa e sys em
Le us conside he sys em o equa ions go e ning he 2d low o wo
supe posed immiscible laye s o shallow luids in a subdomainΩ⊂R2:
∂W
∂ +∂F1
∂x(W)+∂F2
∂y(W)=B1(W)∂W
∂x+B2(W)∂W
∂y+S1(W)∂H
∂x+S2(W)∂H
∂y,
(1)
whe e W=!h1q1,x q1,y h2q2,x q2,y "T,
F1(W)=#q1,x
q2
1,x
h1
+1
2gh2
1
q1,xq1,y
h1
q2,x
q2
2,x
h2
+1
2gh2
2
q2,xq2,y
h2$T
F2(W)=#q1,y
q1,xq1,y
h1
q2
1,y
h1
+1
2gh
2
1q2,y
q2,xq2,y
h2
q2
2,y
h2
+1
2gh2
2$T
,
Sk(W)=!0gh1(2 −k)gh1(k−1) 0 gh2(2 −k)gh2(k−1) "T,k=1,2,
Bk(W)=
0P1,k(W)
P2,k(W)0
Pl,k(W)=
0 0 0
−ghl(2 −k)00
−ghl(k−1) 0 0
l=1,2.
5
Index 1 in he unknowns makes e e ence o he uppe laye and index 2 o
he lowe one; gis he g a i y and H(x), he dep h unc ion measu ed om
a ixed le el o e e ence; =ρ1/ρ2is he a io o he cons an densi ies o
he laye s (ρ1<ρ2) which, in ealis ic oceanog aphical applica ions, is close
o 1. Finally, hi(x, ) and qi(x, ) a e, espec i ely, he hickness and he
mass- low o he i- h laye a he poin xa ime , and hey a e ela ed
o he eloci ies ui(x, ) = (ui,x(x, ),u
i,y(x, )), i=1,2 by he equali ies:
qi(x, )=ui(x, )hi(x, ),i=1,2.
Le us de ine he ma ices Ak(W)=Jk(W)−Bk(W), k=1,2 whe e
Jk(W)=∂Fk
∂W(W) a e he Jacobians o he luxes Fk, and we assume ha
(1) is s ic ly hype bolic. Le us also ema k ha he sys em (1) e i ies he
p ope y o in a iance by o a ions. E ec i ely, le us de ine
Tη=
Rη0
0Rη
,R
η=
1 0 0
0ηxηy
0−ηyηx
,
and le us deno e Fη(W)=F1(W)ηx+F2(W)ηy,B(W) = (B1(W),B
2(W)),
and S(W)=(S1(W),S
2(W)). Then
Fη(W)=T−1
ηF1(TηW),T
ηB(W)·η=B1(TηW),T
ηS(W)·η=S1(TηW)
(2)
Mo eo e , i is easy o check ha TηW e i ies he sys em
∂ (TηW)+∂ηF1(TηW)=B1(TηW)∂ηW+S1(TηW)∂ηH+Qη⊥,(3)
whe e Qη⊥=Tη#−∂η⊥Fη⊥(W)+B(W)·η⊥∂η⊥W+S(W)·η⊥∂η⊥H$.
6
2.2. The IR-Roe Nume ical Scheme
To disc e ize (1) he compu a ional domain Ωis decomposed in o cells o
ini e olumes: Vi⊂R2. He e, i is assumed ha he cells a e iangles. Le
us deno e by L he numbe o iangles o he mesh.
Gi en a ini e olume Vi,|Vi|will ep esen i s a ea; Ni∈R2i s cen e ;
Ni he se o indexes jsuch ha Vjis a neighbo o Vi;Γij he common edge
o wo neighbo ing iangles Viand Vj, and |Γij |i s leng h; ηij =(ηij,x,ηij,y)
he no mal uni ec o a he edge Γij poin ing owa ds he iangle Vj; and
Wn
i he cons an app oxima ion o he a e age o he solu ion in he iangle
Via ime np o ided by he nume ical scheme.
Le us now b ie ly desc ibe he Roe ype scheme o sys em (1) ha can
be de ined aking in o accoun he p ope y o in a iance by o a ions [23]:
Wn+1
i=Wn
i−∆
|Vi|+
j∈Ni|Γij|FIR−ROE
ij
−(4)
whe e FIR−ROE
ij
−is de ined as ollows:
1. Le us de ine Wη=[h1q1,ηh2q2,η]T=Tη(W)[1,2,4,5], and Wη⊥=
[q1,η⊥q2,η⊥]T=Tη(W)[3,6], whe e W[i1,··· ,is]is he ec o de ined om
ec o W, using i s i1- h, . . . , is- h componen s.
2. Le Φ−
ηij be he 1D nume ical Roe lux associa ed o he 1D wo-laye
shallow-wa e sys em de ined by he 1-s , 2-nd, 4- h and 5- h equa ions
o sys em (3) whe e he e m Qη⊥
ij has been neglec ed:
Φ−
ηij =P−
ij (F1(Wηij ,j)−F1(Wηij ,i)−Bij(Wηij ,j −Wηij ,i)−Sij(Hj−Hi))
+F1(Wηij ,i).
7
whe e F1(Wηij )=F1(Tηij W)[1,2,4,5],Bij (Wηij,j −Wηij,i )=,B1,ij(Tηij (Wj−
Wi))-[1,2,4,5],Sij(Hj−Hi)=,S1,ij(Hj−Hi)-[1,2,4,5].
Finally, P−
ij =1
2Kij(I−sgn(Dij))K−1
ij , whe e Iis he iden i y ma ix,
Kij is he ma ix whose columns a e he eigen ec o s o he ma ix Aij,
and sgn(Dij) is he diagonal ma ix whose coe icien s a e he signs o
he eigen alues o Aij, being
Aij =
0 1 0 0
gh1,ij −u2
1,ηij 2u1,ηij gh1,ij 0
0 0 0 1
gh2,ij 0gh2,ij −u2
2,ηij 2u2,ηij
wi h uk,ηij =uk,ij ·ηij,k=1,2 and hk,ij =hk,i+hk,j
2, whe e uk,l,ij =
√hk,iuk,l,i+√hk,j uk,l,j
√hk,i+√hk,j
,k=1,2, l=x, y.
3. Le us de ine Φ−
η⊥
ij
=4(Φ−
ηij )[1]u∗
1,η⊥
ij (Φ−
ηij )[3]u∗
2,η⊥
ij 5T
, whe e u∗
k,η⊥
ij is
de ined as ollows
u∗
k,η⊥
ij =
qk,η⊥
ij ,i
hk,i i (Φηij )[2k−1] >0
qk,η⊥
ij ,j
hk,j o he wise
k=1,2.
Le us ema k ha Φ−
η⊥
ij
is he nume ical lux associa ed o he 3- d
and 6- h equa ions o sys em (3) whe e, again, he e m Qη⊥
ij has been
neglec ed. I s de i a ion has been done ollowing he main ideas o he
HLLC me hod o he shallow wa e sys em in oduced in [24] as qk,η⊥
ij ,
k=1,2 can be seen as a passi e scala ha is con ec ed by he low.
4. Finally, he global nume ical lux is de ined by FIR−ROE
ij
−=T−1
ηij F−
ij ,
whe e F−
ij =4(Φ−
ηij )[1] (Φ−
ηij )[2] (Φ−
η⊥
ij
)[1] (Φ−
ηij )[3] (Φ−
ηij )[4] (Φ−
η⊥
ij
)[2] 5.
8
A CFL condi ion mus be imposed o ensu e s abili y o bo h schemes:
1
2
∆
Vi+
j∈Ni|Γij%Dij%∞≤γ,1≤i≤L, wi h 0 <γ≤1.(5)
As in he case o sys ems o conse a ion laws, when sonic a e ac ion
wa es appea i is necessa y o modi y he nume ical scheme o ge en opy-
sa is ying solu ions. Fo ins ance, he Ha en-Hyman en opy ix echnique
[25] can be easily adap ed he e. Le us also ema k ha he scheme is
pa h-conse a i e in he sense in oduced by Pa es in [26] and [22]. I is
well-balanced o s a iona y solu ions co esponding o wa e a es . Mo e
gene al esul s conce ning he consis ency and well-balanced p ope ies o
Roe schemes ha e been s udied in [22] and [27].
3. Pa alleliza ion o he scheme
Figu e 1a shows a g aphical desc ip ion o he pa allel algo i hm, ob ained
om he desc ip ion o he IR-Roe nume ical scheme gi en in Sec ion 2. The
main calcula ion phases a e iden i ied wi h ci cled numbe s, and he main
sou ces o da a pa allelism a e ep esen ed wi h o e lapping ec angles.
Ini ially, he ini e olume mesh is cons uc ed om he inpu da a. Then he
ime s epping p ocess is epea ed un il he inal simula ion ime is eached.
A he (n+ 1)- h ime s ep, Equa ion (4) mus be e alua ed o upda e he
s a e o each cell.
Each o he main calcula ion phases p esen a high deg ee o pa allelism
because he compu a ion a each edge o olume is independen wi h espec
o ha pe o med a o he edges o olumes:
9
he submesh so ha all he communica ion olumes ha a e adjacen o a
pa icula submesh appea consecu i ely in he a ay. Fo example, in Figu e
2b, no e ha olumes 10-12 a e sen o he lowe submesh, while olumes 12
and 13 a e sen o he igh submesh, hus o e lapping he sendings.
In o de o pe o m his a angemen , i s ly we build a lis o commu-
nica ion olumes o each adjacen submesh. Fo example, in Figu e 2b we
would ha e wo lis s: [10, 11, 12] and [12, 13].
Now, o each communica ion olume o he submesh ha mus be sen o
wo MPI p ocesses, we build a pai (p1,p
2), meaning ha he olume mus be
sen o p ocesses p1and p2. Figu e 2c shows an example cen e ed on submesh
4, whe e all he pai s a e speci ied. Once all he pai s ha e been buil , we
pe o m a eo de ing o hem (and hei elemen s i necessa y) so ha we
ge a lis o consecu i e p ocesses. In Figu e 2c, he pai s a e eo de ed
ob aining: (0,1), (1,2), (2,3) and (5,6). This gi es he consecu i e lis o
p ocesses 0, 1, 2, 3, 5 and 6. Now we ca y ou he adequa e swaps:
1. In he lis s o ing he olumes ha a e adjacen o submesh 0, we pu
he olume sha ed wi h submesh 1 a he end.
2. In he lis s o ing he olumes ha a e adjacen o submesh 1, we pu
he olume sha ed wi h submesh 0 a he s a , and he olume sha ed
wi h submesh 2 a he end.
3. We con inue p ocessing he lis in he same way un il i inishes.
Finally we join all he lis s o communica ion olumes adequa ely o ge
he de ini i e block o communica ion olumes.
No e ha a submesh mus know he o de ing o he communica ion o-
lumes ha ecei es om ano he submesh. The e o e, a his poin all sub-
16
meshes mus send his in o ma ion o hei adjacen submeshes.
No e also ha his algo i hm does no wo k when a submesh has wo
olumes ha mus be sen o he same submeshes. This is e lec ed by a
duplica ed pai in he sequence o pai s, bu we can always pe o m a domain
decomposi ion whe e his does no occu . I nei he wo ks when a submesh
is o med by a single olume, bu his will ne e happen in a eal p oblem.
6.3. Mul i-GPU Code
We ha e implemen ed wo e sions o he mul i-GPU algo i hm: one
wi h blocking MPI sends and ecei es, and ano he one which o e laps MPI
communica ion wi h CPU-GPU memo y ans e s and ke nel compu a ion.
Algo i hm 1 shows he gene al s eps o he non-o e lapping implemen-
a ion. In lines 3-5 we send o each adjacen submesh he communica ion
olumes ha a e adjacen o i . Then, in lines 6-8 we ecei e om he same
submeshes hei communica ion olumes ha a e adjacen o ou submesh.
We ha e used bu e ized MPI send ope a ions wi h a gi en bu e o a oid
deadlocks wi h big meshes. Lines 9-10 copy he ecei ed communica ion o-
lumes o GPU memo y. In line 14 a MPI educ ion is pe o med o ob ain
he global minimum ∆ in all he MPI p ocesses. Lines 16-17 copy he new
s a es o he communica ion olumes o ou submesh om de ice o hos .
Algo i hm 2 shows he gene al s eps o he o e lapping implemen a ion.
Lines 3-5 ( he ecep ion o he communica ion olumes om he adjacen
submeshes) o e lap wi h lines 6-7 ( he copy o he new s a es o ou commu-
nica ion olumes om de ice o hos ). Then, lines 8-10 ( he sending o each
adjacen submesh o he communica ion olumes ha a e adjacen o i )
17
Algo i hm 1 Non-o e lapping mul i-GPU algo i hm
1: n←numbe o adjacen submeshes
2: while ( < end)do
3: o i= 1 o ndo
4: Send communica ion olumes o adjacen submesh i
5: end o
6: o i= 1 o ndo
7: Recei e communica ion olumes om adjacen submesh i
8: end o
9: CudaMemcpy(Laye 1 o comm. olumes om hos o de ice)
10: CudaMemcpy(Laye 2 o comm. olumes om hos o de ice)
11: p ocessEdges<<<g id, block>>>(...)
12: compu eDel aTVolumes<<<g id, block>>>(...)
13: ∆ ←ge MinimumDel aT(...)
14: MPI All educe(∆ , min ∆ ,...)
15: compu eVolumeS a es<<<g id, block>>>(...)
16: CudaMemcpy(Laye 1 o comm. olumes om de ice o hos )
17: CudaMemcpy(Laye 2 o comm. olumes om de ice o hos )
18: ← + min ∆
19: end while
o e lap wi h line 11 ( he p ocessing o he non-communica ion edges, since
hese edges do no need ex e nal da a o be p ocessed). In line 12 we wai
o he communica ion olumes o he adjacen submeshes o a i e. Once
hey ha e a i ed, in lines 13-14 we copy hem o GPU memo y. In line 15
only he communica ion edges a e p ocessed.
7. Expe imen al Resul s
In his sec ion we will es he single and mul i-GPU implemen a ions
desc ibed in Sec ions 5 and 6, espec i ely. The es p oblem and he pa-
ame e s a e he same ha we e used in Sec ion 4.
18
Algo i hm 2 O e lapping mul i-GPU algo i hm
1: n←numbe o adjacen submeshes
2: while ( < end)do
3: o i= 1 o ndo
4: Recei e comm. olumes om adjacen submesh i(Non-blocking)
5: end o
6: CudaMemcpy(Laye 1 o comm. olumes om de ice o hos )
7: CudaMemcpy(Laye 2 o comm. olumes om de ice o hos )
8: o i= 1 o ndo
9: Send comm. olumes o adjacen submesh i(Non-blocking)
10: end o
11: p ocessEdges<<<g id, block>>>(Non-communica ion edges)
12: MPI Wai all (Comm. olumes om adjacen submeshes)
13: CudaMemcpy(Laye 1 o comm. olumes om hos o de ice)
14: CudaMemcpy(Laye 2 o comm. olumes om hos o de ice)
15: p ocessEdges<<<g id, block>>>(Communica ion edges)
16: compu eDel aTVolumes<<<g id, block>>>(...)
17: ∆ ←ge MinimumDel aT(...)
18: MPI All educe(∆ , min ∆ ,...)
19: compu eVolumeS a es<<<g id, block>>>(...)
20: ← + min ∆
21: end while
We ha e used he Chaco so wa e [31] o di ide a mesh in o equally sized
submeshes, he OpenMPI implemen a ion [32] and he GNU compile . All
he p og ams we e execu ed in a clus e o med by ou In el Xeon se e s
wi h 8 GB RAM each one, connec ed wi h a Gigabi E he ne swi ch. G a-
phics ca ds used we e wo Tesla C2050 and wo GeFo ce GTX 570. Since he
GTX 570 ca d p o ides be e pe o mance o ou p og ams han he Tesla,
i is sui able o use he ou ca ds o measu e s ong scalabili y aking he
Tesla as he e e ence ca d. Table 2 shows he execu ion imes in seconds
o all he meshes and numbe o GPUs. Figu e 3a shows g aphically he
19
Tesla 2 Tesla C2050 2 Tesla + 2 GTX 570
Volumes C2050 Non-O e lap O e lap Non-O e lap O e lap
4016 0.0090 0.012 0.014 0.011 0.015
16040 0.045 0.038 0.040 0.027 0.031
64052 0.29 0.19 0.19 0.12 0.12
256576 2.10 1.18 1.16 0.69 0.63
1001898 15.63 8.40 7.98 4.60 4.08
2000608 45.37 23.34 23.00 12.08 11.66
3000948 82.84 43.52 41.83 23.34 21.23
Table 2: GPU execu ion imes in seconds o he IR-Roe me hod.
speedups ob ained wi h he single GPU p og am execu ed on a Tesla C2050
wi h espec o he CPU e sions o he IR-Roe me hod used in Sec ion 4.
Figu e 3b shows he speedups ob ained wi h he mul i-GPU implemen a ions
wi h espec o one Tesla.
As i can be seen, using a Tesla C2050, o big meshes we ha e eached
a speedup o 32 and 10 wi h espec o monoco e and quadco e CPU e -
sions, espec i ely. As expec ed, he o e lapping mul i-GPU implemen a ion
ou pe o ms he non-o e lapping e sion, and he weak and s ong scaling
eached by he o e lapping e sion a e close o pe ec o up o ou GPUs.
8. Conclusions and u u e wo k
In his pape we ha e p esen ed an imp o emen o a i s o de well-
balanced Roe ype ini e olume sol e o wo-laye shallow wa e sys em.
This nume ical scheme has p o ed o be compu a ionally mo e e icien han
he classical Roe scheme and is mo e sui able o be implemen ed in mode n
CUDA-enabled GPUs han he classical Roe scheme. A mul i-GPU dis-
20
(a) Tesla C2050 speedup wi h espec
o se ial and quadco e CPU e sions
(b) Mul i-GPU speedup wi h espec
o one Tesla C2050
Figu e 3: Speedups ob ained o one and se e al GPUs.
ibu ed implemen a ion o his scheme ha wo ks on iangula meshes has
been implemen ed using MPI and CUDA. Nume ical expe imen s ca ied ou
on a GPU clus e ha e shown he e iciency o his sol e , ob aining weak
and s ong scaling close o pe ec o up o ou GPUs by o e lapping MPI
communica ions wi h CPU-GPU memo y ans e s and GPU compu a ion.
As u he wo k, we p opose o ex end he p oposal o enable high o de
nume ical schemes and o in eg a e a dynamic load balancing s a egy (which
is necessa y in p oblems, such as lood simula ions, whe e he compu a ional
load o each spa ial subdomain could a y d ama ically).
Acknowledgemen s
The au ho s acknowledges pa ial suppo om he DGI-MEC p ojec s
MTM2008-06349-C03-03, MTM2009-11923 and MTM2009-07719.
21
Re e ences
[1] M. Rump , R. S zodka, G aphics P ocesso Uni s: New P ospec s o
Pa allel Compu ing, in: Nume ical Solu ion o Pa ial Di e en ial Equa-
ions on Pa allel Compu e s, Vol. 51 o Lec u e No es in Compu a ional
Science and Enginee ing, Sp inge , 2005, pp. 89–134.
[2] J. Owens, M. Hous on, D. Luebke, S. G een, J. S one, J. Phillips, GPU
compu ing, P oceedings o he IEEE 96 (5) (2008) 879–899.
[3] NVIDIA Co po a ion, NVIDIA CUDA C P og amming Guide 3.2, 2010.
[4] T. Hagen, J. Hjelme ik, K.-A. Lie, J. Na ig, M. O. Hen iksen, Visual
simula ion o shallow-wa e wa es, Simula ion Modelling P ac ice and
Theo y 13 (8) (2005) 716–726, P og ammable G aphics Ha dwa e.
[5] M. Las a, J. M. Man as, C. U e˜na, M. J. Cas o, J. A. Ga c´ıa-
Rod ´ıguez, Simula ion o shallow-wa e sys ems using g aphics p ocess-
ing uni s, Ma hema ics and Compu e s in Simula ion 80 (3) (2009) 598–
618.
[6] W.-Y. Liang, T.-J. Hsieh, M. T. Sa ia, Y.-L. Chang, J.-P. Fang, C.-C.
Chen, C.-C. Han, A GPU-Based Simula ion o Tsunami P opaga ion
and Inunda ion, in: P oceedings o he 9 h In e na ional Con e ence
on Algo i hms and A chi ec u es o Pa allel P ocessing, ICA3PP ’09,
Sp inge -Ve lag, Be lin, Heidelbe g, 2009, pp. 593–603.
[7] M. Cas o, J. Ga c´ıa-Rod ´ıguez, J. Gonz´alez-Vida, C. Pa ´es, A pa allel
2d ini e olume scheme o sol ing sys ems o balance laws wi h non-
22
conse a i e p oduc s: Applica ion o shallow lows, Compu e Me hods
in Applied Mechanics and Enginee ing 195 (19-22) (2006) 2788–2815.
[8] M. de la Asunci´on, J. M. Man as, M. J. Cas o, Simula ion o one-
laye shallow wa e sys ems on mul ico e and CUDA a chi ec u es, The
Jou nal o Supe compu ing (2010) 1–9.
[9] M. de la Asunci´on, J. M. Man as, M. J. Cas o, P og amming CUDA-
based GPUs o simula e wo-laye shallow wa e lows, in: Eu o-Pa 2010
- Pa allel P ocessing, Vol. 6272 o Lec u e No es in Compu e Science,
Sp inge Be lin / Heidelbe g, 2010, pp. 353–364.
[10] M. J. Cas o, S. O ega, M. de la Asunci´on, J. M. Man as, GPU com-
pu ing o shallow wa e low simula ion based on ini e olume schemes,
Bol. Soc. Esp. Ma . Apl. 50 (2010) 27–45.
[11] A. B od ko b, T. Hagen, K.-A. Lie, J. Na ig, Simula ion and isuali-
za ion o he sain - enan sys em using GPUs, Compu ing and Visuali-
za ion in Science (2011) 1–13.
[12] J. Galla do, S. O ega, M. de la Asunci´on, J. M. Man as, Two-
dimensional compac hi d-o de polynomial econs uc ions. Sol ing
nonconse a i e hype bolic sys ems using GPUs, Jou nal o Scien i ic
Compu ing (2011) 1–23.
[13] M. J. Cas o, S. O ega, M. de la Asunci´on, J. M. Man as, J. M. Ga-
lla do, GPU compu ing o shallow wa e low simula ion based on ini e
olume schemes, Comp es Rendus M´ecanique 339 (2-3) (2011) 165–184,
High Pe o mance Compu ing.
23
[14] J. Thibaul , I. Senocak, Accele a ing incomp essible low compu a-
ions wi haP h eads-CUDA implemen a ion onsmall- oo p in mul i-
GPU pla o ms, The Jou nal o Supe compu ing (2010) 1–27.
[15] M. L. Sae a, A. R. B od ko b, Shallow wa e simula ions on mul i-
ple GPUs, P oceedings o he Pa a 2010 Con e ence, Lec u e No es in
Compu e Science (2010) Accep ed o publica ion.
[16] D. Koma i schand, G. E lebache , D. G¨oddeke, D. Mich´ea, High-o de
ini e-elemen seismic wa e p opaga ion modeling wi h MPI on a la ge
GPU clus e , J. Compu . Phys. 229 (2010) 7692–7714.
[17] Z. Fan, F. Qiu, A. Kau man, S. Yoakum-S o e , GPU Clus e o High
Pe o mance Compu ing, in: P oceedings o he 2004 ACM/IEEE con-
e ence on Supe compu ing, SC ’04, IEEE Compu e Socie y, Washing-
on, DC, USA, 2004, pp. 47–.
[18] Y. Zhang, F. Muelle , X. Cui, T. Po ok, Da a-in ensi e documen clus-
e ing on g aphics p ocessing uni (GPU) clus e s, J. Pa allel Dis ib.
Compu . 71 (2011) 211–224.
[19] R. Abdelkhalek, H. Calend a, O. Coulaud, J. Roman, G. La u, Fas Seis-
mic Modeling and Re e se Time Mig a ion on a GPU Clus e , in: The
2009 High Pe o mance Compu ing & Simula ion - HPCS’09, Leipzig
Allemagne, 2009, Bes Pape Awa d a HPCS’09 To al.
[20] M. Acu˜na, T. Aoki, Real- ime sunami simula ion on mul i-node GPU
clus e , Supe compu ing (2009) [Pos e ].
24
[21] Message Passing In e ace Fo um, MPI: A Message Passing In e ace
S anda d, Uni . o Tennessee, Knox ille, Tennessee.
[22] M. Cas o, E. Fe n´andez, A. Fe ei o, A. Ga c´ıa, C. Pa ´es, High o de
ex ension o Roe schemes o wo dimensional nonconse a i e hype -
bolic sys ems, J. Sci. Compu . 39 (2009) 67–114.
[23] E. D. Fe n´andez-Nie o, Modelling an nume ical simula ion o subma ine
sedimen shallow lows: anspo and a alanches, Bol. Soc. Esp. Ma .
Apl. 49.
[24] E. D. Fe n´andez-Nie o, D. B esch, J. Monnie , A consis en in e me-
dia e wa e speed o a well-balanced HLLC sol e , Comp es Rendus
Ma hema ique 346 (13-14) (2008) 795 – 800.
[25] J. H. A. Ha en, Sel -adjus ing g id me hods o one-dimensional hype -
bolic conse a ion laws, J. Comp. Phys. 50 (1983) 235–269.
[26] C. Pa ´es, High o de ex ension o Roe schemes o wo dimensional non-
conse a i e hype bolic sys ems, SIAM J. Num. Anal. 44 (2006) 300–
321.
[27] C. Pa ´es, M. Cas o, On he well-balance p ope y o Roe’s me hod
o nonconse a i e hype bolic sys ems. applica ions o shallow-wa e
sys ems, M2AN 38 (5) (2004) 821–852.
[28] B. Chapman, G. Jos , R. an de Pas, Using OpenMP: Po able Sha ed
Memo y Pa allel P og amming, 2007.
25