scieee AI-readable full text Open interactive document viewer

DNA Sequences Alignment in Multi-GPUs: Energy Payoff on Speculative Executions

Pérez-Serrano, Jesús,Sandes, Edans,Melo, Alba,Ujaldon-Martínez, Manuel

Abstract

We present a performance per watt analysis of CUDAlign 4.0, a parallel strategy to obtain the optimal alignment of huge DNA se- quences in multi-GPU platforms using the exact Smith-Waterman method. Speed-up factors and energy consumption are monitored on different stages of the algorithm with the goal of identifying advantageous sce- narios to maximize acceleration and minimize power consumption. Ex- perimental results using CUDA on a set of GeForce GTX 980 GPUs illustrate their capabilities as high-performance and low-power devices, with a energy cost to be more attractive when increasing the number of GPUs. Overall, our results demonstrate a good correlation between the performance attained and the extra energy required, even in scenarios where multi-GPUs do not show great scalability.

Full text

DNA Sequences Alignment in Multi-GPUs: Energy Payoff on Speculative Executions Authors: J. Pérez, M. Ujaldón E. Sandes, A. Melo Computer Architecture Department Computer Science Department University of Malaga (Spain) University of Brasilia (Brazil) Presenter: M. Ujaldón Full Professor @ Computer Architecture: >100 research papers published CUDA Fellow @ NVIDIA: >100 talks/courses over the past 5 years M. Ujaldón Talk outline [20 slides] 1. Background and motivation [4] 2. DNA sequence comparison [3] 3. CUDAlign [3] 4. Experimental setup [3] 5. Experimental results [4] 6. Speculative executions [3] I. Background and motivation M. Ujaldón The Smith-Waterman algorithm (SW) SW is a well known biomedical application to compute: 1. The exact pairwise comparison of DNA/RNA sequences. 2. A protein sequence (query) to a genomic database. Fine-grained parallelism applies better to 1. Already ported to multi-GPUs and Xeon Phis. Huge data volume (“big data”): Tens-Hundreds of Million Base Pairs each sequence in our study. Several Peta-cells (250) for the dynamic programming matrix used. 4 M. Ujaldón Primary goals of this work We present a performance per watt analysis of SW using CUDAlign 4.0, identifying advantageous scenarios to maximize speed-up and minimize power consumption on GPUs. We evaluate: 1. Speed-up, scalability and power on multi-GPU systems. 2. How the workload size influences energy in data-intensive applics. 3. The energy overhead on speculative executions. 5 M. Ujaldón GPU acceleration when energy matters: Identifying good and bad scenarios Example: Acceleration versus fuel consumption in my car. When is it worth to increase 10 MPH? 6 Driving at 60 MPH: 16.66% more speed. 10% more gas. Driving at 100 MPH: 10% more speed. 25% more gas. You can save fuel up to 4x if you press the throttle wisely. M. Ujaldón Speculative executions: Evolution Back in the 80’s and 90’s: Lots of successful stories (branch prediction, prefetching, look-ahead). Because you gamble without risk, you are aggresive on bets. 7 Now that energy matters: Miss-predictions cause power consumption and no execution gains. So you are undecided to play even with good cards. II. DNA Sequence Comparison M. Ujaldón DNA Sequence Comparison A DNA sequence is represented by an ordered list of nucleotide bases, strings of the {A, C, G, T} alphabet. Score function: Example: +1 for a match. -1 for a mismatch. -2 whenever you find a gap. Smith-Waterman algorithm obtains the optimal pairwise local alignment in quadratic time and space. Two phases: Calculate the dynamic programming matrix. Obtain the alignment (traceback). 9 IV. Experimental setup M. Ujaldón Platforms used 17 Hardware resources Nvidia GPUs used Nvidia GPUs used Intel CPU Commercial model Number of cores Cores frequency Memory size and family Memory frequency Memory width Memory bandwidth Software installed GeForce GTX 980 Titan Pascal Xeon E5-2620 2048 3584 8 1126 MHz 1531 MHz 2100 MHz 4 Gbytes GDDR5 12 Gbytes GDDR5X 64 Gbytes DDR4 7 GHz 10 GHz 2.4 GHz 384 bit 384 bits 256 336 GB/s. 480 GB/s. 76.8 GB/s. CUDA 8.0 CUDA 8.0 Ubuntu 14.04 LTS 64 bits M. Ujaldón Input data set Real DNA sequences coming from the National Center for Biotechnology (NCBI) database. Comparison all human (GRCh37) and Chimpanzee (panTro4) homologous chromosomes. Among the 25 pairs of sequences, we have chosen: 18 Input seq. Input seq. Size Size Peta Cells Score Length Coverage Matches Mismatches Gaps Human Chimp. Peta Cells Score Length Coverage Matches Mismatches Gaps chr22 chr21 47M chrY 51M 50M 2.55 31.510.791 51.929.087 98.9% 88.5% 3.8% 7.7% 48M 46M 2.24 36.006.054 48.579.349 99.0% 91.9% 1.1% 7.1% 47M 33M 1.54 27.206.434 33.583.457 70.5% 94.4% 1.5% 4.1% 59M 26M 1.56 1.394.673 2.283.191 6.0% 88.1% 2.0% 10.0% M. Ujaldón Monitoring energy Beagleblone Black open-source hardware. Accelpower module with 8 sensors: 19 V. Experimental results M. Ujaldón Power, time and energy on four GTX980 GPUs 21 Sequence Average power (watts per GPU) Stage 1 Stage 2 Stage 3 Average power (watts per GPU) Average power (watts per GPU) Average power (watts per GPU) chr22 chr21 47M chrY 101.11 W. 116.26 W. 77.27 W. 102.11 W. 116.47 W. 78.89 W. 104.37 W. 117.12 W. 76.33 W. 103.25 W. 119.63 W. 0.00 W. Execution time (seconds) Execution time (seconds) Execution time (seconds) Execution time (seconds) Total time chr22 chr21 47M chrY 11161.92 s. 185.20 s. 14.25 s. 11361.38 s. 9687.36 s. 61.49 s. 11.03 s. 9759.89 s. 6694.95 s. 88.25 s. 9.05 s. 6792.26 s. 6798.12 s. 3.99 s. 0.00 s. 6802.11 s. Energy consumption (kilojules per GPU) Energy consumption (kilojules per GPU) Energy consumption (kilojules per GPU) Energy consumption (kilojules per GPU) Total energy Total cost chr22 chr21 47M chrY 1128.63 kJ. 21.53 kJ. 1.10 kJ. 1151.27 kJ. 0.1660 € 989.26 kJ. 7.16 kJ. 0.87 kJ. 997.29 kJ. 0.1440 € 698.82 kJ. 10.34 kJ. 0.69 kJ. 709.85 kJ. 0.1024 € 701.94 kJ. 0.48 kJ. 0.00 kJ. 702.42 kJ. 0.1012 € M. Ujaldón Power, time and energy for chr22 on multi-GPU 22 No. GPUs Average power (watts per GPU) Stage 1 Stage 2 Stage 3 Average power (watts per GPU) Average power (watts per GPU) Average power (watts per GPU) 4 3 2 1 101.11 W. 116.26 W. 77.27 W. 101.53 W. 108.16 W. 78.79 W. 100.30 W. 114.68 W. 76.74 W. 102.95 W. 114.44 W. 81.27 W. Execution time (seconds) Execution time (seconds) Execution time (seconds) Execution time (seconds) Total time 4 3 2 1 11161.92 s. 185.20 s. 14.25 s. 11361.38 s. 14719.32 s. 253.72 s. 17.70 s. 14990.76 s. 22080.04 s. 159.77 s. 23.17 s. 22262.99 s. 22302.24 s. 291.50 s. 46.65 s. 22640.40 s. Energy consumption (kilojules per GPU) Energy consumption (kilojules per GPU) Energy consumption (kilojules per GPU) Energy consumption (kilojules per GPU) Total energy Total cost 4 3 2 1 1128.63 kJ. 21.53 kJ. 1.10 kJ. 1151.27 kJ. 0.1660 € 1494.60 kJ. 27.45 kJ. 1.40 kJ. 1523.44 kJ. 0.1650 € 2214.77 kJ. 18.32 kJ. 1.78 kJ. 2234.88 kJ. 0.1614 € 2296.22 kJ. 33.36 kJ. 3.79 kJ. 2333.37 kJ. 0.0842 € M. Ujaldón Power consumption on 4 GTX 980 GPUs stage by stage 23 M. Ujaldón Time savings and energy penalties (chr22) 24 No. GPUs Stage 1 Stage 1 Stage 2 Stage 2 Stage 3 Stage 3 Total Total Savings (time) Penalty (energy) Savings (time) Penalty (energy) Savings (time) Penalty (energy) Savings (time) Penalty (energy) Four Three Two 49.96% 96.60% 36.47% 158.15% 69.46% 6.09% 49.82% 97.35% 34.01% 95.26% 12.97% 146.85% 62.06% 0.81% 33.79% 95.86% 1.00% 92.90% 45.20% 9.83% 50.34% -6.07% 1.67% 91.55% VI. Speculative executions