scieee Open visual document viewer

An MPI-CUDA implementation of an improved Roe method for two-layer shallow water systems

Asunción, M. de la; Mantas, José Miguel; Castro Díaz, Manuel Jesús; Fernández Nieto, Enrique Domingo

Abstract

The numerical solution of two-layer shallow water systems is required to simulate accurately stratified fluids, which are ubiquitous in nature: they appear in atmospheric flows, ocean currents, oil spills, . . . Moreover, the implementation of the numerical schemes to solve these models in realistic scenarios imposes huge demands of computing power. In this paper, we tackle the acceleration of these simulations in triangular meshes by exploiting the combined power of several CUDA-enabled GPUs in a GPU cluster. For that purpose, an improvement of a path conservative Roe type finite volume scheme which is specially suitable for GPU implementation is presented, and a distributed implementation of this scheme which uses CUDA and MPI to exploit the potential of a GPU cluster is developed. This implementation overlaps MPI communication with CPU-GPU memory transfers and GPU computation to increase efficiency. Several numerical experiments performed on a cluster of modern CUDA-enabled GPUs show the efficiency of the distributed solver.

Full text

An MPI-CUDA implemen a ion o an imp o ed Roe me hod o wo-laye shallow wa e sys ems Ma c de la Asunci´ona, Jos´e M. Man asa, Manuel J. Cas ob, E. D. Fe n´andez-Nie oc aDp o. Lenguajes y Sis emas In o m´a icos, Uni e sidad de G anada bDp o. An´alisis Ma em´a ico, Uni e sidad de M´alaga cDp o. Ma em´a ica Aplicada I, Uni e sidad de Se illa Abs ac The nume ical solu ion o wo-laye shallow wa e sys ems is equi ed o simu- la e accu a ely s a i ied luids, which a e ubiqui ous in na u e: hey appea in a mosphe ic lows, ocean cu en s, oil spills, . . . Mo eo e , he implemen- a ion o he nume ical schemes o sol e hese models in ealis ic scena ios imposes huge demands o compu ing powe . In his pape , we ackle he acce- le a ion o hese simula ions in iangula meshes by exploi ing he combined powe o se e al CUDA-enabled GPUs in a GPU clus e . Fo ha pu pose, an imp o emen o a pa h conse a i e Roe ype ini e olume scheme which is specially sui able o GPU implemen a ion is p esen ed, and a dis ibu ed implemen a ion o his scheme which uses CUDA and MPI o exploi he po en ial o a GPU clus e is de eloped. This implemen a ion o e laps MPI communica ion wi h CPU-GPU memo y ans e s and GPU compu a ion o inc ease e iciency. Se e al nume ical expe imen s pe o med on a clus e o mode n CUDA-enabled GPUs show he e iciency o he dis ibu ed sol e . Keywo ds: Shallow wa e simula ion, GPU clus e compu ing, ini e olume schemes, uns uc u ed meshes, CUDA, MPI P ep in submi ed o Jou nal o Pa allel and Dis ibu ed Compu ing Ma ch 4, 2011 1. In oduc ion The wo-laye shallow wa e sys em has been used as he nume ical model o simula e se e al phenomena ela ed o s a i ied geophysical lows such as a mosphe ic lows, ocean cu en s, oil spills o sunamis gene a ed by unde wa e landslides. The simula ion o hese phenomena gi es o place o e y long las ing simula ions in big compu a ional domains. Mo eo e , some o hese phenomena ( sunami p opaga ion o oil spills) could equi e eal ime calcula ion. The e o e, ex emely e icien nume ical schemes and implemen a ions a e needed o be able o analyze hose p oblems in p ac ical execu ion imes. Since he nume ical schemes o sol e shallow wa e sys ems usually exhibi a high deg ee o po en ial pa allelism, he de elopmen o pa allel e sions o hese schemes o high pe o mance pla o ms seems o be a sui able way o achie ing he equi ed pe o mance in ealis ic applica ions. A cos e ec i e way o ob aining a subs an ially highe pe o mance in hese applica ions consis s in using G aphics P ocesso Uni s (GPUs). These a chi ec u es make i possible o ob ain pe o mances ha a e o de s o mag- ni ude as e han a s anda d CPU and a e g owing in popula i y among he scien i ic and enginee ing communi y [1, 2]. Mo eo e , se e al GPU p o- g amming oolki s such as CUDA [3] ha e been de eloped o acili a e he p og amming o GPUs o gene al pu pose applica ions. The e a e p e ious p oposals o po ini e olume one-laye shallow wa e sol e s o a GPU by using a g aphics-speci ic p og amming language [4, 5, 6], bu cu en ly mos o he p oposals o simula e shallow lows on a single GPU a e based on he CUDA p og amming model. A CUDA sol e o one- 2 laye sys em based on he i s o de ini e olume scheme p esen ed in [7] is desc ibed in [8] o deal wi h s uc u ed egula meshes. The ex ension o his CUDA sol e o wo-laye shallow wa e sys em is p esen ed in [9]. The e also exis p oposals o implemen , using CUDA-enabled GPUs, high o de schemes o simula e one-laye sys ems [10, 11, 12] and o implemen i s -o de schemes o one and wo-laye sys ems on iangula meshes [13]. Al hough he use o single GPU sys ems makes i possible o sa is y he pe o mance equi emen s o se e al applica ions which a e no compu a io- nally expensi e, his si ua ion is no equen in mos ealis ic applica ions. Many applica ions equi e o handle huge meshes and la ge numbe o ime s eps. Mo eo e , in some applica ions eal ime accu a e p edic ions ( o ins ance, o app oxima e he e ec o an unexpec ed oil spill) could be e- qui ed. The cha ac e is ics o hese applica ions sugges o combine he powe o mul iple GPUs o sa is y he pe o mance equi emen s. One app oach o use se e al GPUs in hese p oblems is based on p o- g amming sha ed memo y mul i-GPU desk op sys ems. These pla o ms ha e been used o accele a e conside ably luid dynamic [14] and shallow wa- e [15] simula ions by combining sha ed memo y p og amming p imi i es o manage h eads in CPU and CUDA o p og am he GPU. Al hough his isa cos -e ec i e app oach, hese pla o ms only o e a educed numbe o GPUs and mo e lexible sys ems a e desi able o answe o he g owing pe o mance equi emen s o many ealis ic applica ions. A mo e lexible app oach o ob ain he equi ed pe o mance in ol es o use clus e s o GPU-enhanced compu e s whe e each node is equipped wi h a single GPU o wi h a mul i-GPU sys em. The compu a ion on GPU clus e s 3 could make i possible o scale he educ ion in execu ion ime acco ding o he numbe o GPUs (which can be easily inc eased). The e o e, his app oach is mo e lexible han using a mul i-GPU desk op sys em and he memo y limi a ions o a GPU-enhanced node can be o e come by sui ably dis ibu ing he da a among he nodes, enabling us o simula e signi ican ly la ge ealis ic models and wi h g ea e p ecision. The use o GPU clus e s o accele a e da a in ensi e compu a ions is gai- ning in popula i y [16, 17, 18, 19]. In [20], a scheme o sol e one-laye shallow wa e sys ems is implemen ed on a GPU clus e o eal- ime sunami simu- la ion. Mos o he p oposals o exploi GPU clus e s in scien i ic compu ing use MPI [21] o implemen he communica ion among he p ocesses o he dis ibu ed sys em and CUDA [3] o p og am he GPU (o GPUs) associa ed o each node. A common way o educe he emo e communica ion o e head in hese dis ibu ed implemen a ions consis s in using non-blocking commu- nica ion MPI unc ions o o e lap he da a ans e s be ween nodes o he clus e wi h he GPU compu a ion and he CPU-GPU da a ans e s. This wo k deals wi h he accele a ion o he nume ical solu ion o wo- laye shallow wa e sys ems by exploi ing he pa alleliza ion o an imp o ed ini e olume scheme o uns uc u ed meshes on GPU clus e s. Fo ha pu pose, a dis ibu ed implemen a ion o his scheme o a GPU clus e is de eloped by using MPI and CUDA. This implemen a ion inco po a es an e icien managemen o he dis ibu ed uns uc u ed mesh and mechanisms o o e lap compu a ion wi h communica ion. The ou line o he a icle is as ollows: he nex sec ion desc ibes he unde lying ma hema ical model and p esen s an imp o emen o a i s o de 4 Roe ype ini e olume scheme, called IR-Roe scheme. Sec ion 3 desc ibes a da a pa allel e sion o he IR-Roe scheme. The e iciency o se e al im- plemen a ions o he IR-Roe scheme and he classical Roe scheme [7, 22] is compa ed in Sec ion 4. In he wo nex sec ions we desc ibe a single and a mul i-GPU dis ibu ed implemen a ion, espec i ely, o he me hod o ian- gula meshes using he CUDA amewo k. Sec ion 7 shows he expe imen al esul s ob ained when he implemen a ions a e applied o sol e an in e nal dam b eak p oblem on a clus e o 4 NVIDIA Fe mi GPUs. Finally, Sec ion 8 summa izes he main conclusions and p esen s he u u e wo k. 2. An e icien nume ical scheme 2.1. The wo-laye shallow wa e sys em Le us conside he sys em o equa ions go e ning he 2d low o wo supe posed immiscible laye s o shallow luids in a subdomainΩ⊂R2: ∂W ∂ +∂F1 ∂x(W)+∂F2 ∂y(W)=B1(W)∂W ∂x+B2(W)∂W ∂y+S1(W)∂H ∂x+S2(W)∂H ∂y, (1) whe e W=!h1q1,x q1,y h2q2,x q2,y "T, F1(W)=#q1,x q2 1,x h1 +1 2gh2 1 q1,xq1,y h1 q2,x q2 2,x h2 +1 2gh2 2 q2,xq2,y h2$T F2(W)=#q1,y q1,xq1,y h1 q2 1,y h1 +1 2gh 2 1q2,y q2,xq2,y h2 q2 2,y h2 +1 2gh2 2$T , Sk(W)=!0gh1(2 −k)gh1(k−1) 0 gh2(2 −k)gh2(k−1) "T,k=1,2, Bk(W)=  0P1,k(W) P2,k(W)0 Pl,k(W)=     0 0 0 −ghl(2 −k)00 −ghl(k−1) 0 0      l=1,2. 5 Index 1 in he unknowns makes e e ence o he uppe laye and index 2 o he lowe one; gis he g a i y and H(x), he dep h unc ion measu ed om a ixed le el o e e ence; =ρ1/ρ2is he a io o he cons an densi ies o he laye s (ρ1<ρ2) which, in ealis ic oceanog aphical applica ions, is close o 1. Finally, hi(x, ) and qi(x, ) a e, espec i ely, he hickness and he mass- low o he i- h laye a he poin xa ime , and hey a e ela ed o he eloci ies ui(x, ) = (ui,x(x, ),u i,y(x, )), i=1,2 by he equali ies: qi(x, )=ui(x, )hi(x, ),i=1,2. Le us de ine he ma ices Ak(W)=Jk(W)−Bk(W), k=1,2 whe e Jk(W)=∂Fk ∂W(W) a e he Jacobians o he luxes Fk, and we assume ha (1) is s ic ly hype bolic. Le us also ema k ha he sys em (1) e i ies he p ope y o in a iance by o a ions. E ec i ely, le us de ine Tη=  Rη0 0Rη  ,R η=     1 0 0 0ηxηy 0−ηyηx      , and le us deno e Fη(W)=F1(W)ηx+F2(W)ηy,B(W) = (B1(W),B 2(W)), and S(W)=(S1(W),S 2(W)). Then Fη(W)=T−1 ηF1(TηW),T ηB(W)·η=B1(TηW),T ηS(W)·η=S1(TηW) (2) Mo eo e , i is easy o check ha TηW e i ies he sys em ∂ (TηW)+∂ηF1(TηW)=B1(TηW)∂ηW+S1(TηW)∂ηH+Qη⊥,(3) whe e Qη⊥=Tη#−∂η⊥Fη⊥(W)+B(W)·η⊥∂η⊥W+S(W)·η⊥∂η⊥H$. 6 2.2. The IR-Roe Nume ical Scheme To disc e ize (1) he compu a ional domain Ωis decomposed in o cells o ini e olumes: Vi⊂R2. He e, i is assumed ha he cells a e iangles. Le us deno e by L he numbe o iangles o he mesh. Gi en a ini e olume Vi,|Vi|will ep esen i s a ea; Ni∈R2i s cen e ; Ni he se o indexes jsuch ha Vjis a neighbo o Vi;Γij he common edge o wo neighbo ing iangles Viand Vj, and |Γij |i s leng h; ηij =(ηij,x,ηij,y) he no mal uni ec o a he edge Γij poin ing owa ds he iangle Vj; and Wn i he cons an app oxima ion o he a e age o he solu ion in he iangle Via ime np o ided by he nume ical scheme. Le us now b ie ly desc ibe he Roe ype scheme o sys em (1) ha can be de ined aking in o accoun he p ope y o in a iance by o a ions [23]: Wn+1 i=Wn i−∆ |Vi|+ j∈Ni|Γij|FIR−ROE ij −(4) whe e FIR−ROE ij −is de ined as ollows: 1. Le us de ine Wη=[h1q1,ηh2q2,η]T=Tη(W)[1,2,4,5], and Wη⊥= [q1,η⊥q2,η⊥]T=Tη(W)[3,6], whe e W[i1,··· ,is]is he ec o de ined om ec o W, using i s i1- h, . . . , is- h componen s. 2. Le Φ− ηij be he 1D nume ical Roe lux associa ed o he 1D wo-laye shallow-wa e sys em de ined by he 1-s , 2-nd, 4- h and 5- h equa ions o sys em (3) whe e he e m Qη⊥ ij has been neglec ed: Φ− ηij =P− ij (F1(Wηij ,j)−F1(Wηij ,i)−Bij(Wηij ,j −Wηij ,i)−Sij(Hj−Hi)) +F1(Wηij ,i). 7 whe e F1(Wηij )=F1(Tηij W)[1,2,4,5],Bij (Wηij,j −Wηij,i )=,B1,ij(Tηij (Wj− Wi))-[1,2,4,5],Sij(Hj−Hi)=,S1,ij(Hj−Hi)-[1,2,4,5]. Finally, P− ij =1 2Kij(I−sgn(Dij))K−1 ij , whe e Iis he iden i y ma ix, Kij is he ma ix whose columns a e he eigen ec o s o he ma ix Aij, and sgn(Dij) is he diagonal ma ix whose coe icien s a e he signs o he eigen alues o Aij, being Aij =         0 1 0 0 gh1,ij −u2 1,ηij 2u1,ηij gh1,ij 0 0 0 0 1 gh2,ij 0gh2,ij −u2 2,ηij 2u2,ηij         wi h uk,ηij =uk,ij ·ηij,k=1,2 and hk,ij =hk,i+hk,j 2, whe e uk,l,ij = √hk,iuk,l,i+√hk,j uk,l,j √hk,i+√hk,j ,k=1,2, l=x, y. 3. Le us de ine Φ− η⊥ ij =4(Φ− ηij )[1]u∗ 1,η⊥ ij (Φ− ηij )[3]u∗ 2,η⊥ ij 5T , whe e u∗ k,η⊥ ij is de ined as ollows u∗ k,η⊥ ij =     qk,η⊥ ij ,i hk,i i (Φηij )[2k−1] >0 qk,η⊥ ij ,j hk,j o he wise k=1,2. Le us ema k ha Φ− η⊥ ij is he nume ical lux associa ed o he 3- d and 6- h equa ions o sys em (3) whe e, again, he e m Qη⊥ ij has been neglec ed. I s de i a ion has been done ollowing he main ideas o he HLLC me hod o he shallow wa e sys em in oduced in [24] as qk,η⊥ ij , k=1,2 can be seen as a passi e scala ha is con ec ed by he low. 4. Finally, he global nume ical lux is de ined by FIR−ROE ij −=T−1 ηij F− ij , whe e F− ij =4(Φ− ηij )[1] (Φ− ηij )[2] (Φ− η⊥ ij )[1] (Φ− ηij )[3] (Φ− ηij )[4] (Φ− η⊥ ij )[2] 5. 8 A CFL condi ion mus be imposed o ensu e s abili y o bo h schemes: 1 2 ∆ Vi+ j∈Ni|Γij%Dij%∞≤γ,1≤i≤L, wi h 0 <γ≤1.(5) As in he case o sys ems o conse a ion laws, when sonic a e ac ion wa es appea i is necessa y o modi y he nume ical scheme o ge en opy- sa is ying solu ions. Fo ins ance, he Ha en-Hyman en opy ix echnique [25] can be easily adap ed he e. Le us also ema k ha he scheme is pa h-conse a i e in he sense in oduced by Pa es in [26] and [22]. I is well-balanced o s a iona y solu ions co esponding o wa e a es . Mo e gene al esul s conce ning he consis ency and well-balanced p ope ies o Roe schemes ha e been s udied in [22] and [27]. 3. Pa alleliza ion o he scheme Figu e 1a shows a g aphical desc ip ion o he pa allel algo i hm, ob ained om he desc ip ion o he IR-Roe nume ical scheme gi en in Sec ion 2. The main calcula ion phases a e iden i ied wi h ci cled numbe s, and he main sou ces o da a pa allelism a e ep esen ed wi h o e lapping ec angles. Ini ially, he ini e olume mesh is cons uc ed om he inpu da a. Then he ime s epping p ocess is epea ed un il he inal simula ion ime is eached. A he (n+ 1)- h ime s ep, Equa ion (4) mus be e alua ed o upda e he s a e o each cell. Each o he main calcula ion phases p esen a high deg ee o pa allelism because he compu a ion a each edge o olume is independen wi h espec o ha pe o med a o he edges o olumes: 9 he submesh so ha all he communica ion olumes ha a e adjacen o a pa icula submesh appea consecu i ely in he a ay. Fo example, in Figu e 2b, no e ha olumes 10-12 a e sen o he lowe submesh, while olumes 12 and 13 a e sen o he igh submesh, hus o e lapping he sendings. In o de o pe o m his a angemen , i s ly we build a lis o commu- nica ion olumes o each adjacen submesh. Fo example, in Figu e 2b we would ha e wo lis s: [10, 11, 12] and [12, 13]. Now, o each communica ion olume o he submesh ha mus be sen o wo MPI p ocesses, we build a pai (p1,p 2), meaning ha he olume mus be sen o p ocesses p1and p2. Figu e 2c shows an example cen e ed on submesh 4, whe e all he pai s a e speci ied. Once all he pai s ha e been buil , we pe o m a eo de ing o hem (and hei elemen s i necessa y) so ha we ge a lis o consecu i e p ocesses. In Figu e 2c, he pai s a e eo de ed ob aining: (0,1), (1,2), (2,3) and (5,6). This gi es he consecu i e lis o p ocesses 0, 1, 2, 3, 5 and 6. Now we ca y ou he adequa e swaps: 1. In he lis s o ing he olumes ha a e adjacen o submesh 0, we pu he olume sha ed wi h submesh 1 a he end. 2. In he lis s o ing he olumes ha a e adjacen o submesh 1, we pu he olume sha ed wi h submesh 0 a he s a , and he olume sha ed wi h submesh 2 a he end. 3. We con inue p ocessing he lis in he same way un il i inishes. Finally we join all he lis s o communica ion olumes adequa ely o ge he de ini i e block o communica ion olumes. No e ha a submesh mus know he o de ing o he communica ion o- lumes ha ecei es om ano he submesh. The e o e, a his poin all sub- 16 meshes mus send his in o ma ion o hei adjacen submeshes. No e also ha his algo i hm does no wo k when a submesh has wo olumes ha mus be sen o he same submeshes. This is e lec ed by a duplica ed pai in he sequence o pai s, bu we can always pe o m a domain decomposi ion whe e his does no occu . I nei he wo ks when a submesh is o med by a single olume, bu his will ne e happen in a eal p oblem. 6.3. Mul i-GPU Code We ha e implemen ed wo e sions o he mul i-GPU algo i hm: one wi h blocking MPI sends and ecei es, and ano he one which o e laps MPI communica ion wi h CPU-GPU memo y ans e s and ke nel compu a ion. Algo i hm 1 shows he gene al s eps o he non-o e lapping implemen- a ion. In lines 3-5 we send o each adjacen submesh he communica ion olumes ha a e adjacen o i . Then, in lines 6-8 we ecei e om he same submeshes hei communica ion olumes ha a e adjacen o ou submesh. We ha e used bu e ized MPI send ope a ions wi h a gi en bu e o a oid deadlocks wi h big meshes. Lines 9-10 copy he ecei ed communica ion o- lumes o GPU memo y. In line 14 a MPI educ ion is pe o med o ob ain he global minimum ∆ in all he MPI p ocesses. Lines 16-17 copy he new s a es o he communica ion olumes o ou submesh om de ice o hos . Algo i hm 2 shows he gene al s eps o he o e lapping implemen a ion. Lines 3-5 ( he ecep ion o he communica ion olumes om he adjacen submeshes) o e lap wi h lines 6-7 ( he copy o he new s a es o ou commu- nica ion olumes om de ice o hos ). Then, lines 8-10 ( he sending o each adjacen submesh o he communica ion olumes ha a e adjacen o i ) 17 Algo i hm 1 Non-o e lapping mul i-GPU algo i hm 1: n←numbe o adjacen submeshes 2: while ( < end)do 3: o i= 1 o ndo 4: Send communica ion olumes o adjacen submesh i 5: end o 6: o i= 1 o ndo 7: Recei e communica ion olumes om adjacen submesh i 8: end o 9: CudaMemcpy(Laye 1 o comm. olumes om hos o de ice) 10: CudaMemcpy(Laye 2 o comm. olumes om hos o de ice) 11: p ocessEdges<<<g id, block>>>(...) 12: compu eDel aTVolumes<<<g id, block>>>(...) 13: ∆ ←ge MinimumDel aT(...) 14: MPI All educe(∆ , min ∆ ,...) 15: compu eVolumeS a es<<<g id, block>>>(...) 16: CudaMemcpy(Laye 1 o comm. olumes om de ice o hos ) 17: CudaMemcpy(Laye 2 o comm. olumes om de ice o hos ) 18: ← + min ∆ 19: end while o e lap wi h line 11 ( he p ocessing o he non-communica ion edges, since hese edges do no need ex e nal da a o be p ocessed). In line 12 we wai o he communica ion olumes o he adjacen submeshes o a i e. Once hey ha e a i ed, in lines 13-14 we copy hem o GPU memo y. In line 15 only he communica ion edges a e p ocessed. 7. Expe imen al Resul s In his sec ion we will es he single and mul i-GPU implemen a ions desc ibed in Sec ions 5 and 6, espec i ely. The es p oblem and he pa- ame e s a e he same ha we e used in Sec ion 4. 18 Algo i hm 2 O e lapping mul i-GPU algo i hm 1: n←numbe o adjacen submeshes 2: while ( < end)do 3: o i= 1 o ndo 4: Recei e comm. olumes om adjacen submesh i(Non-blocking) 5: end o 6: CudaMemcpy(Laye 1 o comm. olumes om de ice o hos ) 7: CudaMemcpy(Laye 2 o comm. olumes om de ice o hos ) 8: o i= 1 o ndo 9: Send comm. olumes o adjacen submesh i(Non-blocking) 10: end o 11: p ocessEdges<<<g id, block>>>(Non-communica ion edges) 12: MPI Wai all (Comm. olumes om adjacen submeshes) 13: CudaMemcpy(Laye 1 o comm. olumes om hos o de ice) 14: CudaMemcpy(Laye 2 o comm. olumes om hos o de ice) 15: p ocessEdges<<<g id, block>>>(Communica ion edges) 16: compu eDel aTVolumes<<<g id, block>>>(...) 17: ∆ ←ge MinimumDel aT(...) 18: MPI All educe(∆ , min ∆ ,...) 19: compu eVolumeS a es<<<g id, block>>>(...) 20: ← + min ∆ 21: end while We ha e used he Chaco so wa e [31] o di ide a mesh in o equally sized submeshes, he OpenMPI implemen a ion [32] and he GNU compile . All he p og ams we e execu ed in a clus e o med by ou In el Xeon se e s wi h 8 GB RAM each one, connec ed wi h a Gigabi E he ne swi ch. G a- phics ca ds used we e wo Tesla C2050 and wo GeFo ce GTX 570. Since he GTX 570 ca d p o ides be e pe o mance o ou p og ams han he Tesla, i is sui able o use he ou ca ds o measu e s ong scalabili y aking he Tesla as he e e ence ca d. Table 2 shows he execu ion imes in seconds o all he meshes and numbe o GPUs. Figu e 3a shows g aphically he 19 Tesla 2 Tesla C2050 2 Tesla + 2 GTX 570 Volumes C2050 Non-O e lap O e lap Non-O e lap O e lap 4016 0.0090 0.012 0.014 0.011 0.015 16040 0.045 0.038 0.040 0.027 0.031 64052 0.29 0.19 0.19 0.12 0.12 256576 2.10 1.18 1.16 0.69 0.63 1001898 15.63 8.40 7.98 4.60 4.08 2000608 45.37 23.34 23.00 12.08 11.66 3000948 82.84 43.52 41.83 23.34 21.23 Table 2: GPU execu ion imes in seconds o he IR-Roe me hod. speedups ob ained wi h he single GPU p og am execu ed on a Tesla C2050 wi h espec o he CPU e sions o he IR-Roe me hod used in Sec ion 4. Figu e 3b shows he speedups ob ained wi h he mul i-GPU implemen a ions wi h espec o one Tesla. As i can be seen, using a Tesla C2050, o big meshes we ha e eached a speedup o 32 and 10 wi h espec o monoco e and quadco e CPU e - sions, espec i ely. As expec ed, he o e lapping mul i-GPU implemen a ion ou pe o ms he non-o e lapping e sion, and he weak and s ong scaling eached by he o e lapping e sion a e close o pe ec o up o ou GPUs. 8. Conclusions and u u e wo k In his pape we ha e p esen ed an imp o emen o a i s o de well- balanced Roe ype ini e olume sol e o wo-laye shallow wa e sys em. This nume ical scheme has p o ed o be compu a ionally mo e e icien han he classical Roe scheme and is mo e sui able o be implemen ed in mode n CUDA-enabled GPUs han he classical Roe scheme. A mul i-GPU dis- 20 (a) Tesla C2050 speedup wi h espec o se ial and quadco e CPU e sions (b) Mul i-GPU speedup wi h espec o one Tesla C2050 Figu e 3: Speedups ob ained o one and se e al GPUs. ibu ed implemen a ion o his scheme ha wo ks on iangula meshes has been implemen ed using MPI and CUDA. Nume ical expe imen s ca ied ou on a GPU clus e ha e shown he e iciency o his sol e , ob aining weak and s ong scaling close o pe ec o up o ou GPUs by o e lapping MPI communica ions wi h CPU-GPU memo y ans e s and GPU compu a ion. As u he wo k, we p opose o ex end he p oposal o enable high o de nume ical schemes and o in eg a e a dynamic load balancing s a egy (which is necessa y in p oblems, such as lood simula ions, whe e he compu a ional load o each spa ial subdomain could a y d ama ically). Acknowledgemen s The au ho s acknowledges pa ial suppo om he DGI-MEC p ojec s MTM2008-06349-C03-03, MTM2009-11923 and MTM2009-07719. 21 Re e ences [1] M. Rump , R. S zodka, G aphics P ocesso Uni s: New P ospec s o Pa allel Compu ing, in: Nume ical Solu ion o Pa ial Di e en ial Equa- ions on Pa allel Compu e s, Vol. 51 o Lec u e No es in Compu a ional Science and Enginee ing, Sp inge , 2005, pp. 89–134. [2] J. Owens, M. Hous on, D. Luebke, S. G een, J. S one, J. Phillips, GPU compu ing, P oceedings o he IEEE 96 (5) (2008) 879–899. [3] NVIDIA Co po a ion, NVIDIA CUDA C P og amming Guide 3.2, 2010. [4] T. Hagen, J. Hjelme ik, K.-A. Lie, J. Na ig, M. O. Hen iksen, Visual simula ion o shallow-wa e wa es, Simula ion Modelling P ac ice and Theo y 13 (8) (2005) 716–726, P og ammable G aphics Ha dwa e. [5] M. Las a, J. M. Man as, C. U e˜na, M. J. Cas o, J. A. Ga c´ıa- Rod ´ıguez, Simula ion o shallow-wa e sys ems using g aphics p ocess- ing uni s, Ma hema ics and Compu e s in Simula ion 80 (3) (2009) 598– 618. [6] W.-Y. Liang, T.-J. Hsieh, M. T. Sa ia, Y.-L. Chang, J.-P. Fang, C.-C. Chen, C.-C. Han, A GPU-Based Simula ion o Tsunami P opaga ion and Inunda ion, in: P oceedings o he 9 h In e na ional Con e ence on Algo i hms and A chi ec u es o Pa allel P ocessing, ICA3PP ’09, Sp inge -Ve lag, Be lin, Heidelbe g, 2009, pp. 593–603. [7] M. Cas o, J. Ga c´ıa-Rod ´ıguez, J. Gonz´alez-Vida, C. Pa ´es, A pa allel 2d ini e olume scheme o sol ing sys ems o balance laws wi h non- 22 conse a i e p oduc s: Applica ion o shallow lows, Compu e Me hods in Applied Mechanics and Enginee ing 195 (19-22) (2006) 2788–2815. [8] M. de la Asunci´on, J. M. Man as, M. J. Cas o, Simula ion o one- laye shallow wa e sys ems on mul ico e and CUDA a chi ec u es, The Jou nal o Supe compu ing (2010) 1–9. [9] M. de la Asunci´on, J. M. Man as, M. J. Cas o, P og amming CUDA- based GPUs o simula e wo-laye shallow wa e lows, in: Eu o-Pa 2010 - Pa allel P ocessing, Vol. 6272 o Lec u e No es in Compu e Science, Sp inge Be lin / Heidelbe g, 2010, pp. 353–364. [10] M. J. Cas o, S. O ega, M. de la Asunci´on, J. M. Man as, GPU com- pu ing o shallow wa e low simula ion based on ini e olume schemes, Bol. Soc. Esp. Ma . Apl. 50 (2010) 27–45. [11] A. B od ko b, T. Hagen, K.-A. Lie, J. Na ig, Simula ion and isuali- za ion o he sain - enan sys em using GPUs, Compu ing and Visuali- za ion in Science (2011) 1–13. [12] J. Galla do, S. O ega, M. de la Asunci´on, J. M. Man as, Two- dimensional compac hi d-o de polynomial econs uc ions. Sol ing nonconse a i e hype bolic sys ems using GPUs, Jou nal o Scien i ic Compu ing (2011) 1–23. [13] M. J. Cas o, S. O ega, M. de la Asunci´on, J. M. Man as, J. M. Ga- lla do, GPU compu ing o shallow wa e low simula ion based on ini e olume schemes, Comp es Rendus M´ecanique 339 (2-3) (2011) 165–184, High Pe o mance Compu ing. 23 [14] J. Thibaul , I. Senocak, Accele a ing incomp essible low compu a- ions wi haP h eads-CUDA implemen a ion onsmall- oo p in mul i- GPU pla o ms, The Jou nal o Supe compu ing (2010) 1–27. [15] M. L. Sae a, A. R. B od ko b, Shallow wa e simula ions on mul i- ple GPUs, P oceedings o he Pa a 2010 Con e ence, Lec u e No es in Compu e Science (2010) Accep ed o publica ion. [16] D. Koma i schand, G. E lebache , D. G¨oddeke, D. Mich´ea, High-o de ini e-elemen seismic wa e p opaga ion modeling wi h MPI on a la ge GPU clus e , J. Compu . Phys. 229 (2010) 7692–7714. [17] Z. Fan, F. Qiu, A. Kau man, S. Yoakum-S o e , GPU Clus e o High Pe o mance Compu ing, in: P oceedings o he 2004 ACM/IEEE con- e ence on Supe compu ing, SC ’04, IEEE Compu e Socie y, Washing- on, DC, USA, 2004, pp. 47–. [18] Y. Zhang, F. Muelle , X. Cui, T. Po ok, Da a-in ensi e documen clus- e ing on g aphics p ocessing uni (GPU) clus e s, J. Pa allel Dis ib. Compu . 71 (2011) 211–224. [19] R. Abdelkhalek, H. Calend a, O. Coulaud, J. Roman, G. La u, Fas Seis- mic Modeling and Re e se Time Mig a ion on a GPU Clus e , in: The 2009 High Pe o mance Compu ing & Simula ion - HPCS’09, Leipzig Allemagne, 2009, Bes Pape Awa d a HPCS’09 To al. [20] M. Acu˜na, T. Aoki, Real- ime sunami simula ion on mul i-node GPU clus e , Supe compu ing (2009) [Pos e ]. 24 [21] Message Passing In e ace Fo um, MPI: A Message Passing In e ace S anda d, Uni . o Tennessee, Knox ille, Tennessee. [22] M. Cas o, E. Fe n´andez, A. Fe ei o, A. Ga c´ıa, C. Pa ´es, High o de ex ension o Roe schemes o wo dimensional nonconse a i e hype - bolic sys ems, J. Sci. Compu . 39 (2009) 67–114. [23] E. D. Fe n´andez-Nie o, Modelling an nume ical simula ion o subma ine sedimen shallow lows: anspo and a alanches, Bol. Soc. Esp. Ma . Apl. 49. [24] E. D. Fe n´andez-Nie o, D. B esch, J. Monnie , A consis en in e me- dia e wa e speed o a well-balanced HLLC sol e , Comp es Rendus Ma hema ique 346 (13-14) (2008) 795 – 800. [25] J. H. A. Ha en, Sel -adjus ing g id me hods o one-dimensional hype - bolic conse a ion laws, J. Comp. Phys. 50 (1983) 235–269. [26] C. Pa ´es, High o de ex ension o Roe schemes o wo dimensional non- conse a i e hype bolic sys ems, SIAM J. Num. Anal. 44 (2006) 300– 321. [27] C. Pa ´es, M. Cas o, On he well-balance p ope y o Roe’s me hod o nonconse a i e hype bolic sys ems. applica ions o shallow-wa e sys ems, M2AN 38 (5) (2004) 821–852. [28] B. Chapman, G. Jos , R. an de Pas, Using OpenMP: Po able Sha ed Memo y Pa allel P og amming, 2007. 25