scieee AI-readable full text Open interactive document viewer

Assessing Intel OneAPI capabilities and cloud-performance for heterogeneous computing

Rodríguez Alcaraz, Silvia; Laso Rodríguez, Rubén; García Lorenzo, Óscar; López Vilariño, David; Fernández Pena, Anselmo Tomás; Fernández Rivera, Francisco

Abstract

This work presents a performance-oriented study of a heterogeneous application developed with Intel OneAPI to solve two well-known diffusion problems: heat diffusion and image denoising. We have explored CPU+iGPU and CPU+FPGA schemes, applying dynamic load balancing and conducting experiments on Intel DevCloud. The results demonstrate that the CPU+iGPU scheme outperforms the execution times achieved by the fastest device when the problem is sufficiently computationally demanding. We also found that the performance of the CPU+FPGA scheme is heavily affected by bandwidth limitations and specific strategies to manage memory efficiently are required. Moreover, it was demonstrated that dynamic workload balancing is crucial due to possible performance fluctuations in any of the implicated devices. In conclusion, Intel OneAPI provides a helpful tool for multi-platform development using a unique high-level language, DPC++. However, developing specific code for each platform is necessary to achieve optimal performance

Full text

Vol.:(0123456789) The Journal of Supercomputing https://doi.org/10.1007/s11227-024-05958-5 1 3 Assessing Intel OneAPI capabilities andcloud‑performance forheterogeneous computing SilviaR.Alcaraz1· RubenLaso1,3· OscarG.Lorenzo1,2· DavidL.Vilariño2· TomásF.Pena1,2· FranciscoF.Rivera1,2 Accepted: 3 February 2024 © The Author(s) 2024 Abstract This work presents a performance-oriented study of a heterogeneous application developed with Intel OneAPI to solve two well-known diffusion problems: heat diffusion and image denoising. We have explored CPU+iGPU and CPU+FPGA schemes, applying dynamic load balancing and conducting experiments on Intel DevCloud. The results demonstrate that the CPU+iGPU scheme outperforms the execution times achieved by the fastest device when the problem is sufficiently computationally demanding. We also found that the performance of the CPU+FPGA scheme is heavily affected by bandwidth limitations and specific strategies to manage memory efficiently are required. Moreover, it was demonstrated that dynamic workload balancing is crucial due to possible performance fluctuations in any of the implicated devices. In conclusion, Intel OneAPI provides a helpful tool for multiplatform development using a unique high-level language, DPC++. However, developing specific code for each platform is necessary to achieve optimal performance. Keywords Intel OneAPI· Intel DevCloud· GPU· FPGA· Heterogeneous computing· Cloud computing * Silvia R. Alcaraz sil[email protected] 1 Centro Singular de Investigación en Tecnoloxías Intelixentes (CiTIUS), Universidade de Santiago de Compostela, Rúa de Jenaro de la Fuente Domínguez S/N, 15782SantiagodeCompostela, Galicia, Spain 2 Departamento de Electrónica e Computación, Universidade de Santiago de Compostela, Rúa Lope Gómez de Marzoa, S/N, 15782SantiagodeCompostela, Galicia, Spain 3 Faculty ofInformatics, Research Group forParallel Computing, TU Wien, Treitlstraße 3, 1040Vienna, Austria S.R.Alcaraz et al. 1 3 1 Introduction Technological advances during the last few years have caused computer applications to become increasingly computationally demanding. This fact not only affects applications used in the scientific field but also includes applications we use on a daily basis. Simultaneously, as applications or programs demand more computing power, manufacturers have also been producing more capable devices. Nowadays, mobile phones are more powerful in memory and processing than desktop computers that existed at the beginning of the century. This evolution can be traced to the rise in powerful processors and larger memory capacities. Moreover, there has been a growth of devices harnessing the advantages offered by heterogeneous computing. This paradigm refers to utilising multiple processing units or accelerators within a unified computing system, which may have different architectures or capabilities. It aims to leverage the strengths inherent in each of these devices to attain a range of advantages, including enhanced performance and reduced energy consumption, among others. Personal computers or laptops are usually equipped with a CPU and an integrated and/or dedicated GPU. This is related to the GPGPU (General-Purpose Computing on Graphics Processing Units) concept, which is based on leveraging the computing power of GPUs for general-purpose applications, not just for rendering graphics. Nowadays, the most popular platforms to develop GPGPU applications are CUDA (Compute Unified Device Architecture) [1], to work with NVIDIA GPUs, and OpenCL (Open Computing Language) [2]. The last is an open framework that can run on various devices, including CPUs, GPUs, and other accelerators. Note that GPUs are not the only devices to be included to create heterogeneous schemes, although they are the most widely used because of the ecosystem of high-level tools to develop applications for them. Nevertheless, it is relevant to emphasise that other devices exist, such as FPGAs (Field Programmable Gate Arrays). In the past, the use of FPGAs increased due to their capacity to adapt their architecture to a specific problem based on programmable logic. Another advantage is their low power consumption, as demonstrated in several studies. Regarding this, Betkaoui et al. [3] determined the advantages of FPGA-based HPC (High Performance Computing) systems over GPUs regarding energy efficiency. Cong et al. [4] also described some performance differences between FPGAs and GPUs, where the results show that FPGAs used less power than GPUs. To make this comparison, they used the well-known benchmark for GPUs, Rodinia [5], migrating some kernels with Vivado HLS (High-Level Synthesis) C [6] to be executed on FPGAs. Both studies highlight the low power consumption of FPGAs, although they point out the limitations that bandwidth or memory management may cause in these devices. Despite the advantages mentioned about FPGAs, their popularity did not reach the magnitude of GPUs mainly because of the maturity in the software stack provided by the latter. The standard way to program FPGAs requires using hardware description languages such as VHDL or Verilog and low-level knowledge 1 3 Assessing Intel OneAPI capabilities andcloud‑performance… of hardware and circuitry to develop optimal solutions. However, FPGAs are currently experiencing a new resurgence thanks to the new high-level development tools [7]. One of the challenges faced by such tools is to provide a compromise between programming simplicity and performance portability. The concept of performance portability refers to the capacity of a code to be executed efficiently across different platforms without the necessity for significant manual modifications. Strategies for parallelising a problem or optimising its performance are different across devices. To develop efficient codes for GPUs, it is essential to understand the thread execution model to maximise the utilisation of their computing capabilities. Additionally, achieving coalescent accesses and allocating data efficiently is equally important to best use the memory. For dGPUs (dedicated GPUs), careful consideration is necessary when deciding which data to send to on-chip memory to avoid costly memory communications. For iGPUs (integrated GPUs), having some memory shared with the CPU can save data transfers, although this type of GPUs tends to have lower computational power. FPGAs are a completely different device; therefore, other techniques are used to optimise codes. Task decomposition and pipelining are essential, enabling the exploitation of parallel processing inherent in FPGAs. Furthermore, it is necessary to use the on-chip memory efficiently, coupled with a selection of appropriate memory access patterns, to minimise data movement and reduce latency. In the case of Intel, the specific tool Intel OneAPI [8] includes a software package that allows the creation of high-level, cross-platform solutions using a single programming language, DPC++ (Data Parallel C++) [9]. It offers a syntax based on C++17 and follows the heterogeneous programming model proposed by SYCL [10] to provide multi-platform programming mechanisms. The leading platforms included are CPUs, GPUs, and FPGAs, although there is also support for custom platforms. Additionally, it is possible to use this tool remotely on Intel Developer Cloud—or Intel DevCloud—where several models of CPUs, GPUs, and FPGAs manufactured by Intel are hosted. This work presents a performance-oriented study of two heterogeneous schemes using well-known iterative diffusion problems as case studies, specifically heat diffusion and image denoising. In addition, dynamic workload balancing is employed to automate the optimal distribution between the involved devices without manually tuning the parameters. The tool used for the development has been Intel OneAPI, as it allows the creation of cross-platform code for CPUs, GPUs and FPGAs. Thus, we have successfully assessed and compared the performance of the CPU+iGPU and CPU+FPGA schemes in the Intel DevCloud environment. Likewise, we have examined the tool’s usability and assessed the performance attained on each device by implementing a generic kernel code applicable to all target platforms. In summary, the main two contributions of this work are the evaluation of the Intel OneAPI tool in terms of usability, performance portability and throughput, particularly, in the Intel DevCloud environment; and a comparative performance study between two heterogeneous computing schemes, CPU+iGPU and CPU+FPGA, using dynamic load balancing. The rest of the paper is structured as follows: Sect.2 includes an overview of the related work; Sect.3 introduces the case studies utilised in this work, namely, the S.R.Alcaraz et al. 1 3 heat diffusion and image denoising problems; Sect.4 discusses the implementations of the benchmarks; The experimental environment is detailed in Sect.5; Results are shown and discussed in Sect.6; Finally, conclusions and future work are outlined in Sect.7. 2 Related work Heterogeneous computing offers a model that allows further progress in achieving greater computing power, energy efficiency and optimised code development to address specific tasks. However, not all problems can be effectively leveraged using heterogeneous computing schemes. Recommended prerequisites include the possibility to decompose them into independent or loosely coupled tasks and the capacity to avoid data dependencies, ensuring that each device can process its amount of data without conflicts. Regarding the applications that can be improved using heterogeneous platforms, Lukarski and Neytcheva [11] studied the particular case of iterative methods, determining that the codes to solve this type of problem can achieve better performance using parallel and heterogeneous approaches. Numerous studies in the literature related to this topic further reinforce this claim. Among them, Venkatasubramanian etal. [12] present multi-CPU and multi-GPU implementations of Jacobi’s iterative method to solve the 2D Poisson equation. Benner etal. [13, 14] accelerated Newton’s iteration for the matrix sign function using hybrid CPU-GPU platforms. Agulleiro etal. [15] proposed a hybrid scheme using multicore CPUs and NVIDIA GPUs to improve the iterative tomographic reconstruction. More recently, Halbiniak etal. [16] presented a port to CPU+GPU for the solidification modelling using OpenCL. As we explore heat diffusion and image denoising as case studies in this work, it is noteworthy to highlight previous studies that utilised heterogeneous computing to achieve improved performance in addressing these problems. Belhaus etal. [17] conducted a comparative study of parallel heat equation implementations on CPU and NVIDIA GPUs. Sánchez etal. [18] presented an algorithm for removing impulsive noise in images. They used OpenMP (Open Multi-Processing) [19] and CUDA to develop solutions for multi-CPU, multi-GPU and a combination of both. Also, Sarjanoja et al. [20] proposed an implementation of the BM3D (Block-Matching and Three-Dimensional) filtering denoising algorithm for CPU and GPU platforms using OpenCL and CUDA. Concerning the use of Intel OneAPI in heterogeneous architectures, the literature contains works such as Constantinescu etal.’s [21] implementation of Markov decision processes on low-power CPU+GPU SoCs. Moreover, they compared OpenCL + TBB (Threading Building Blocks) and Intel OneAPI + TBB implementations in terms of ease of programming and performance. Yong etal. [22] migrated an ultrasound imaging application from CUDA to Intel oneAPI, tuning it for GPU, FPGA and CPU. Lupescu and Ţăpuş [23] developed a hybrid hashtable for an SoC (System-On-a-Chip) composed of a CPU and an iGPU. Marinelli et Raja [24] ported a GPU-based hash join algorithm from CUDA to DPC++ and tested it on dGPU, iGPU and multicore CPUs. They also developed OneJoin [25], an edit similarity 1 3 Assessing Intel OneAPI capabilities andcloud‑performance… join between architectures for DNA data storage. Nozal and Bosque [26] assess the performance and energy efficiency of popular HPC benchmarks using dynamic and static workloads on CPU+GPU heterogeneous systems. Najmeh etal. [27] provided HosNa, a benchmark suite written in DPC++ for heterogeneous architectures. In addition, they analysed the performance of CPUs and FPGAs—both manufactured by Intel—using algorithms from the suite. The study by Kashino etal. [28] explores the use of GPU+FPGA for complex physical simulations. Groth etal. [29] presented a hybrid and accelerated implementation of various hash table operations for dGPUs and APUs (Accelerated Processing Units), consisting of a combination of a CPU and an iGPU. Specifically, they used Intel OneAPI to develop the search operation in the case of the APU. Finally, Li etal. [30] propose a portable framework for largescale graph processing based on the heterogeneous CPU+GPU scheme. 3 Case studies To study the performance implications of the proposed heterogeneous computing schemes, two diffusion problems have been selected as case studies. This type of problem aims to provide easily understandable examples suitable to be solved using a heterogeneous approach. 3.1 Heat diffusion As an example of isotropic diffusion, we chose the two-dimensional heat equation: The equation states that the rate of change of temperature u with respect to time at a certain point is proportional to the Laplacian of the temperature at that point, which is the sum of the second-order partial derivatives of the temperature with respect to both x and y. In other words, the equation describes how the temperature at a point changes over time due to the temperature differences between its neighbours and the material’s thermal properties in the region, as governed by the diffusion coefficient D. As stated in(1), this problem is an isotropic diffusion because it assumes that heat diffuses uniformly in all directions, regardless of the orientation of the coordinates axes. This is caused by the scalar differential Laplacian operator, which is invariant under rotations of the coordinate axes. The selected initial conditions for our case study are where x and y are the coordinates for the x and y directions, respectively, and the temperature at a point (x,y) on the grid at time t is denoted by ux , y , t . Function(2) is a natural choice for the initial temperature distribution as it satisfies the following properties: firstly, it is a smooth function with no sharp edges or (1) 𝜕u 𝜕t −DΔu= 0. (2) u(x,y,0)=sin(x)sin(y), S.R.Alcaraz et al. 1 3 corners, so the temperature distribution is continuous and differentiable everywhere in the domain; and secondly, the temperature at the boundary is zero, and this function fulfils that condition, since Therefore, the problem can be summarised as follows: 3.1.1 Discretisation To solve the heat equation using finite differences, it is necessary to apply a discretisation of the spatial domain and the time interval. Given the domain Ω=(0, 𝜋)×(0, 𝜋) , it can be discretised in a structured mesh of N×N nodes spaced by 𝛿x and 𝛿y in the x and y axis, respectively such that 𝛿 x=𝛿y= 𝜋 N . Regarding time discretisation, we use a fixed time step, 𝛿t . The problem is solved using an Explicit Euler method, defined as where u is the temperature. In this case, we use the well-known approach 5-point scheme [31], shown in Fig.1. Therefore, it is needed to divide the region of interest into a grid of points with uniform spacing in both x and y directions, being 𝛿x and 𝛿y the spacing between two adjacent grid points in the x and y directions. (3) u(0, y,t)=u(x, 0, t)=u(𝜋,y,t)=u(x,𝜋,t)=0. (4) ⎧ ⎪ ⎨ ⎪ ⎩ 𝜕u 𝜕t−DΔu=0, u(x,y,0)=sin(x)sin(y), u(0, y,t)=u(x, 0, t)=u(𝜋,y,t)=u(x,𝜋,t)= 0. (5) u (x,y,t+𝛿t)=u(x,y,t)+𝛿t 𝜕u(x,y,t) 𝜕t , Fig. 1 Laplacian 5 point scheme 1 3 Assessing Intel OneAPI capabilities andcloud‑performance… The following equations show the approximation of second-order derivatives with respect to x and y using the central difference approximation, respectively: Finally, substituting these approximations into the original heat equation, we obtain which relates the temperature at a point (i,j) at time t+𝛿t to the temperature at the same point and its four neighbours at time t. This equation can be used to iteratively update the temperature at each point on the grid at each time step. This process is known as explicit time integration, and it requires that the time step 𝛿t is small enough to ensure numerical stability. Since a square domain is considered, it is reasonable to use a discretisation where 𝛿x =𝛿y =𝛿 . Therefore, the computations would be 3.2 Image denoising The second problem that we have chosen is image denoising as it can be solved using diffusion methods. To do this, we have taken as reference the solution proposed by Tang et al. [32] to perform image denoising by applying an anisotropic diffusion method. For the sake of simplicity, we have focussed exclusively on the anisotropic flow for brightness that they proposed. Hence, to recreate this study, any grey-scale image with the same dimensions tested in these experiments is appropriate. 3.2.1 Anisotropic diffusion forbrightness A digital image is, in its fundamental form, a representation of a 2-dimensional scene as a finite array of pixels. Each pixel corresponds to a position in the scene and is assigned a value that represents the intensity or colour of the light that was recorded at that position. Therefore, a grey-scale image has only one channel, which represents the brightness or intensity of the image at each pixel. However, colour images typically have three channels, representing the red, green, and blue (RGB) components of the image at each pixel position. A digital image can be defined as a function f∶ ℝ 2 →ℝ c and we can express it as a matrix, where every pair of coordinates M(x,y) is mapped to an c-dimensional vector of real numbers, representing the colour or intensity of the image at that position. (6) 𝜕 2u 𝜕x2≈ u i+1,j,t −2u i,j,t +u i−1,j,t 𝛿2 x ,𝜕2u 𝜕y2≈ u i,j+1,t −2u i,j,t +u i,j−1,t 𝛿2 y . (7) u i,j,t+𝛿t =ui,j,t+𝛿t ( ui+1,j,t−2ui,j,t+ui−1,j,t 𝛿2 x + ui,j+1,t−2ui,j,t+ui,j−1,t 𝛿2 y ), (8) u i,j,t+𝛿t =ui,j,t+𝛿t (u i+1,j,t +u i,j+1,t −4u i,j,t +u i−1,j,t +u i,j−1,t 𝛿2 ). S.R.Alcaraz et al. 1 3 Regarding the diffusion problem, the work of Tang etal. [32] mentioned before proposed two approaches to deal with grey-scale images or brightness. The first is the isotropic flow based on the Laplace equation: The second, which is the diffusion flow that we apply in this work, is the anisotropic flow, which is included in the following equation by adding the term 1∕( 1 +‖∇M‖) : It is worth noting that solving equation(10) involves a larger number of arithmetic calculations than equation (9), which would replicate the problem explained in Sect.3.1. Consequently, we consider two applications with different characteristics in terms of arithmetic intensity, which should allow us to explore the singularities of the studied devices. 3.2.2 Discretisation The spatial discretisation of images is often considered trivial since the resolution of the image is already defined by its pixel grid: This means that we can easily represent each point in the image as a discrete value, and the discretisation process is reduced to determining the appropriate mathematical operations to be applied to these values. In our case, we have chosen to use an Explicit Euler method with a fixed time step to discretise the problem. By selecting a fixed time step, we can ensure that the solution remains stable and accurate throughout the simulation. 4 Implementation details In this section, important details about the implementation of both case studies are depicted. Section4.1 provides information about the tool used for developing both solutions, along with technical details regarding the proposed implementation. Then, the procedure for introducing dynamic load balancing is explained in Sect.4.2. 4.1 Heterogeneous approach: Intel OneAPI The high-level, multi-platform programming tool used to develop this work is Intel OneAPI. It is a unified software toolkit that enables the development of applications across diverse architectures. It includes compilers, libraries, and tools for developing (9) 𝜕M 𝜕t (x,y,t)=Mxx(x,y,t)+Myy(x,y,t)=ΔM(x,y,t) . (10) 𝜕 M 𝜕 t (x,y,t)= � MxxM2 y−2MxMyMxy +MyyM2 x �1 3 1 +‖∇M‖. (11) 𝛿x =𝛿y =1. 1 3 Assessing Intel OneAPI capabilities andcloud‑performance… software that can run on CPUs, GPUs, FPGAs, and other accelerators. Furthermore, OneAPI supports multiple programming languages, such as C++, Fortran, and DPC++. The latter is based on C++17 and includes a modern, heterogeneous programming model based on the SYCL standard. Intel OneAPI also includes extensive tools for profiling, debugging, and optimising code for different architectures such as Intel Advisor [33] or Intel VTune [34]. Concerning memory management, Intel OneAPI provides two ways to deal with it: USM (Unified Shared Memory) and buffers. USM provides a pointer-based approach, with a familiar syntax to C++, to allocate memory on the host (malloc_host), on the device (malloc_device), or both (malloc_shared). The shared allocations are useful in the case of programs where the host and the device frequently access the data since data are accessible on both. The movements of data are hidden as they are performed by runtime mechanisms and lower-level drivers. In the case of device allocations, they occur in the device-attached memory and the movement of data is explicit. This means that the developer must use copy operations to move data between the host and the device. In addition to USM, DPC++ also provides buffers, which represent a region of memory on the device. According to the SYCL 2020 Specification [35], allocations and deallocations of buffers are the responsibility of the SYCL runtime. The program related to this study was developed using the native tools provided by Intel OneAPI mentioned above. This is, using DPC++ mechanisms to manage memory, synchronise devices, and express parallelism. In relation to the latter, SYCL handler class’ parallel_for [36, 37] member function has been used for the GPU kernel. It provides an interface to define and invoke a SYCL kernel to execute in parallel over a specific range of values. We use a two-parameter version, with the first parameter designating the range and the second being a lambda function indicating the code that will be executed on the device. In this study, the workload distribution is performed by rows since the input data are always a matrix or an image, which is essentially the same. Thus, in our code the range is expressed with sycl::range(size, cols) where the size variable would denote the number of rows assigned to be processed on the device in that set of iterations and the cols variable is just the total number of columns of the input matrix or image. Section4.2 provides a detailed explanation of how workload is distributed among devices. Similarly, the FPGA kernel follows the same strategy. Despite knowing that a specific implementation may be needed to achieve maximum performance, this study examines whether it is actually possible to generate valid kernels for different devices without modifying any lines of code using Intel OneAPI. Finally, regarding the code executed by the host, it was parallelised using OpenMP. 4.2 Dynamic load balance In the context of problems that can be optimised using heterogeneous schemes, one of the main priorities is to perform an appropriate workload distribution among the devices involved. The aim is to maximise the benefits derived from the use of heterogeneous computing by taking advantage of the specific capabilities of each device. S.R.Alcaraz et al. 1 3 that could be extracted from the FPGA reports provided by the OneAPI tool—as some low-level design details are treated as a black box towards the programmer for simplicity—the data width of the DDR (Double Data Rate) memory is 512 bits. In addition to the data alignment issues, the performance drops in case of not using powers of two since the compiler always increases the size of the memory ports by using the nearest bigger power of 2 and masking out the extra bytes [42]. This fact degrades performance by wasting bandwidth. In summary, sizes that are powers of 2 provide much more competitive execution times because the time lost in host-device communications over the iterations is significantly lower than that spent with the other matrix sizes. 6.2.2 Image denoising The results obtained with the FPGA in relation to the image denoising problem are similar to the previous ones since the CPU also provides better execution times. Nevertheless, as illustrated in Fig. 7, two notable changes affect the performance of the program. Firstly, the gap between the CPU and the FPGA is significantly less pronounced than the one observed in the context of the heat diffusion problem. Secondly, the performance of the FPGA is relatively lower than that of the CPU, but it provides more consistent results when varying the sizes of the input matrix in relation to the previous study case. This fact makes performing the dynamic workload distribution between the CPU and the FPGA easier. These two reasons explain the possibility of enhancing the results of the devices individually by utilising the Fig. 6 CPU+FPGA: normalised execution times for heat diffusion 1 3 Assessing Intel OneAPI capabilities andcloud‑performance… heterogeneous scheme. However, this is only feasible in experiments where the matrix size is a power of two, as commented before. In other instances, the time required for memory transfers becomes excessively large; therefore, the communication cost does not compensate for workload balancing between the host and the device. 6.3 Comparison This section presents a set of graphical resources that facilitate the comparison of execution times among the different devices and mixed schemes explored in this study. As depicted in Fig. 8a, in the case of the heat diffusion problem, the results obtained with CPU Xeon Gold using OpenMP with the maximum number of threads available—12 threads—were the best for cases where N≥2000 . The next device to provide the best execution times is the alternative CPU model—Xeon E-2176—with 12 threads, closely followed by the CPU+iGPU scheme and the UHD P630 iGPU. Although the two CPUs evaluated have the same number of threads, as mentioned in Sect.5, the Xeon Gold has a larger LLC size. In the case of this problem, the cache size plays an important role in speeding up data access, so the Xeon Gold performance benefits from this. Finally, the FPGA implementations using a non-specific kernel for this type of device yield the worst results. In any case, the CPU+FPGA scheme proves to be better than the FPGA working alone for all sizes. Fig. 7 CPU+FPGA: normalised execution times for image denoising S.R.Alcaraz et al. 1 3 Regarding the image denoising problem, the heterogeneous CPU+iGPU scheme provided the lowest execution times, as shown in Fig. 8b. This scheme is followed by the execution of the iGPU UHD P630 and then by the two CPU models. This problem involves more arithmetic operations, thus favouring the CPU with a higher operating frequency, the Xeon E-2176G. Once again, the worst results are obtained from the non-specific FPGA implementations. However, in this case, the CPU+FPGA scheme is not always better than the isolated execution of the Arria 10. For cases where the input matrix size is not a power of 2, it is notable that the communications between the host and the device penalise the performance significantly. 7 Conclusions This work presents a performance-oriented study of a heterogeneous application which solves two well-known diffusion problems: heat diffusion and image denoising. With this aim, two heterogeneous schemes have been explored: CPU+iGPU and CPU+FPGA, using dynamic workload balancing strategies to improve the final performance and resource utilisation. This way, the workload is distributed between the CPU and the device in runtime. The program has been developed using the highlevel tool Intel OneAPI, and the experiments were conducted on Intel DevCloud. The results confirm the benefits of using a heterogeneous approach to solve problems which are susceptible to being parallelised following this paradigm. On the other hand, it was possible to identify the potential problems in scenarios where the architecture of any of the devices is not appropriately exploited. Moreover, all the experiments demonstrated the importance of dynamic load balancing in heterogeneous schemes due to possible performance irregularities suffered by one of Fig. 8 Comparison of execution times for all the schemes used in the study. Lower is better 1 3 Assessing Intel OneAPI capabilities andcloud‑performance… the devices implicated. In any case, for the image denoising problem, we improved the execution times obtained with the fastest device—in this case, the iGPU—in ≈30% using the CPU+iGPU scheme. Regarding the FPGA experiments, despite not developing a specific kernel for this device, the dynamic load balancing reduces the execution times of the isolated FPGA execution significantly. The best advantages obtained were 92.39% and 86.75% for heat diffusion and image denoising, respectively, in both cases where N=4096 . It is important to note that this advantage was only possible in cases where the memory movements between the host and the device did not imply a burden. Regarding Intel OneAPI, it provides an option to develop heterogeneous solutions for CPUs, GPUs, and FPGAs using a unique and high-level language, DPC++. The emergence of high-level tools such as this one makes it easier to program crossdevice code, facilitating and democratising access to more niche devices such as FPGAs. Nevertheless, in our experience, developing specific code for each platform considering their architectures and features is necessary to achieve optimal performance. In future work, we plan to develop a specific FPGA solution to leverage programmable logic’s advantages in these devices fully. This will enable fairer performance comparisons with the GPU-based scheme for the two study cases explored. Acknowledgements This work has received financial support from the Consellería de Cultura, Educación e Ordenación Universitaria (accreditation ED431C 2022/16 and accreditation ED431G-2019/04) and the European Regional Development Fund (ERDF), which acknowledges the CiTIUS-Research Center in Intelligent Technologies of the University of Santiago de Compostela as a Research Center of the Galician University System. This work was also supported by the Ministry of Economy and Competitiveness, Government of Spain (Grants Nos.28700 PID2019-104834GB-I00 and PID2022-141623NB-I00). Funding Open Access funding provided thanks to the CRUE-CSIC agreement with Springer Nature. Data Availability Statement Specific data are not required to reproduce this study. The work itself explains which input data are used and how to generate it. Declarations Conflict of interest The authors have no conflicts of interest that may have affected the content of this work. Ethics approval Not applicable. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/ licenses/by/4.0/. S.R.Alcaraz et al. 1 3 References 1. Nickolls J (2007) GPU parallel computing architecture and CUDA programming model. In: Proceedings of IEEE Hot chips 19 symposium (HCS), pp 1–12 https:// doi. org/ 10. 1109/ HOTCH IPS. 2007. 74824 91 2. Stone JE, Gohara D, Shi G (2010) OpenCL: a parallel programming standard for heterogeneous computing systems. Computi Sci Eng 12(3):66–73. https:// doi. org/ 10. 1109/ MCSE. 2010. 69 3. Betkaoui B, Thomas DB, Luk W (2010) Comparing performance and energy efficiency of FPGAs and GPUs for high productivity computing. In: International Conference on Field-Programmable Technology, Beijing, China, pp 94–101. https:// doi. org/ 10. 1109/ FPT. 2010. 56817 61 4. Cong J, Fang Z, Lo M, Wang H, Xu J, Zhang S (2018) Understanding performance differences of FPGAs and GPUs. In: IEEE 26th annual international symposium on field-programmable custom computing machines (FCCM), Boulder, pp 93–96. https:// doi. org/ 10. 1109/ FCCM. 2018. 00023 5. Che S, Boyer M, Meng J, Tarjan D, Sheaffer JW, Lee SH, Skadron K (2009) Rodinia: a benchmark suite for heterogeneous computing. In: Proceedings of the IEEE international symposium on workload characterization (IISWC), pp 44–54 .https:// doi. org/ 10. 1109/ IISWC. 2009. 53067 97 6. Vivado High-Level Synthesis (2024) https:// www. xilinx. com/ produ cts/ designtools/ vivado/ integ ration/ esldesign. html. Accessed 12 Jan 7. Koch D, Hannig F, Ziener D (eds) (2016) FPGAs for Software Programmers. Springer. https:// doi. org/ 10. 1007/ 978-331926408-0 8. Intel OneAPI. https:// softw are. intel. com/ conte nt/ www/ us/ en/ devel op/ tools/ oneapi. html. Accessed 12 Jan 2024 9. Reinders J, Ashbaugh B, Brodman J, Kinsner M, Pennycook J, Tian X (2021) Data parallel C++: mastering DPC++ for programming of heterogeneous systems using C++ and SYCL. Apress Berkeley. https:// doi. org/ 10. 1007/ 978-148425574-2 10. SYCL: Khronos Open Standard for C++ heterogeneous parallel programming. https:// www. khron os. org/ api/ sycl. Accessed 12 Jan 2024 11. Lukarski D, Neytcheva M (2014) On the impact of the heterogeneous multicore and many-core platforms on iterative solution methods and preconditioning techniques. Wiley, pp 11–32. Chap. 2. https:// doi. org/ 10. 1002/ 97811 18711 897. ch2 12. Venkatasubramanian S, Vuduc RW (2009) Tuned and wildly asynchronous stencil kernels for hybrid CPU/GPU systems. In: Proceedings of the 23rd International Conference on Supercomputing (ICS). Association for Computing Machinery, New York, pp 244–255. https:// doi. org/ 10. 1145/ 15422 75. 15423 12 13. Benner P, Ezzatti P, Quintana-Orti ES, Remon A (2009) Using hybrid CPU-GPU platforms to accelerate the computation of the matrix sign function. Euro-Par – Parallel Processing Workshops, pp. 132–139. Springer, Berlin, Heidelberg . https:// doi. org/ 10. 1007/ 978-364214122-5_ 17 14. Benner P, Ezzatti P, Kressner D, Quintana-Ortí ES, Remón A (2011) A mixed-precision algorithm for the solution of Lyapunov equations on hybrid CPU-GPU platforms. Parallel Comput 37(8):439– 450. https:// doi. org/ 10. 1016/j. parco. 2010. 12. 002 15. Agulleiro JI, Vázquez F, Garzón EM, Fernández JJ (2012) Dynamic load scheduling on CPU-GPU for iterative tomographic reconstruction. In: IEEE 10th international symposium on parallel and distributed processing with applications, pp 603–608. https:// doi. org/ 10. 1109/ ISPA. 2012. 90 16. Halbiniak K, Szustak L, Olas T, Wyrzykowski R, Gepner P (2021) Exploration of OpenCL heterogeneous programming for porting solidification modeling to CPU-GPU platforms. Concurr Comput Pract Exp 33(4):6011. https:// doi. org/ 10. 1002/ cpe. 6011 17. Belhaous S, Chokri S, Baroud S, Mestari M (2021) Comparative study of the execution time of parallel heat equation on CPU and GPU. J Commun Softw Syst 17(4):350–357. https:// doi. org/ 10. 24138/ jcomss20210133 18. Sánchez MG, Vidal V, Bataller J (2012) Peer group and fuzzy metric to remove noise in images using heterogeneous computing. In: Euro-Par 2011: parallel processing workshops. Springer, Berlin, Heidelberg, pp 502–510. https:// doi. org/ 10. 1007/ 978-364229737-3_ 55 19. Dagum L, Menon R (1998) OpenMP: an industry standard API for shared-memory programming. IEEE Comput Sci Eng 5(1):46–55. https:// doi. org/ 10. 1109/ 99. 660313 20. Sarjanoja S, Boutellier J, Hannuksela J (2015) BM3D image denoising using heterogeneous computing platforms. In: 2015 Conference on Design and Architectures for Signal and Image Processing (DASIP), pp. 1–8. https:// doi. org/ 10. 1109/ DASIP. 2015. 73672 57 1 3 Assessing Intel OneAPI capabilities andcloud‑performance… 21. Constantinescu D, Navarro A, Corbera F, Fernández-Madrigal J-A, Asenjo R (2021) Efficiency and productivity for decision making on low-power heterogeneous CPU+GPU SoCs. J Supercomput. https:// doi. org/ 10. 1007/ s1122702003257-3 22. Yong W, Yongfa Z, Scott W, Wang Y, Qing X, Chen W (2021) Developing medical ultrasound imaging application across GPU, FPGA, and CPU using OneAPI. In: IWOCL’21. Association for Computing Machinery, New York. https:// doi. org/ 10. 1145/ 34566 69. 34566 80 23. Lupescu G, Ţăpuş N (2021) Design of hashtable for heterogeneous architectures. In: 2021 23rd International Conference on Control Systems and Computer Science (CSCS), pp 172–177. https:// doi. org/ 10. 1109/ CSCS5 2396. 2021. 00035 24. Marinelli E, Appuswamy R (2021) XJoin: portable, parallel hash join across diverse XPU architectures with oneAPI. In: ACM (ed) DAMON 2021, 17th International Workshop on Data Management on New Hardware, Held with ACM SIGMOD/PODS, 21 June 2021, China (Virtual Event). https:// doi. org/ 10. 1145/ 34659 98. 34660 12 25. Marinelli E, Appuswamy R (2021) OneJoin: Cross-architecture, scalable edit similarity join for DNA data storage using oneAPI. In: ACM (ed) ADMS 2021, 12th International Workshop on Accelerating Analytics and Data Management Systems Using Modern Processor and Storage Architectures, in Conjunction with VLDB 2021, 16 August 2021, Copenhagen, Denmark, Copenhagen 26. Nozal R, Bosque JL (2021) Straightforward Heterogeneous Computing with the oneAPI Coexecutor Runtime. Electronics 10(19). https:// doi. org/ 10. 3390/ elect ronic s1019 2386 27. Bavarsad NN, Makrani HM, Sayadi H, Landis L, Rafatirad S, Homayoun H (2021). HosNa: a DPC++ benchmark suite for heterogeneous architectures. In: 2021 IEEE 39th International Conference on Computer Design (ICCD), pp 509–516. https:// doi. org/ 10. 1109/ ICCD5 3106. 2021. 00084 28. Kashino R, Kobayashi R, Fujita N, Boku, T (2022). Multi-hetero acceleration by GPU and FPGA for astrophysics simulation on OneAPI environment. In: International Conference on High Performance Computing in Asia-Pacific Region. HPCAsia2022. Association for Computing Machinery, New York, pp 84–93. https:// doi. org/ 10. 1145/ 34928 05. 34928 17 29. Groth T, Groppe S, Pionteck T, Valdiek F, Koppehel M (2023) Hybrid CPU/GPU/APU accelerated query, insert, update and erase operations in hash tables with string keys. Knowl Inf Syst 65:1–19. https:// doi. org/ 10. 1007/ s1011502301891-w 30. Li S, Zhu J, Han J, Peng Y, Wang Z, Gong X, Wang G, Zhang J, Wang X (2023) OneGraph: a crossarchitecture framework for large-scale graph computing on GPUs based on oneAPI. CCF Trans High Perform Comput. https:// doi. org/ 10. 1007/ s4251402300172-w 31. LeVeque R (2007) Finite difference methods for ordinary and partial differential equations: steadystate and time-dependent problems (classics in applied mathematics classics in applied mathematics). Society for Industrial and Applied Mathematics, New York 32. Tang B, Sapiro G, Caselles V (2000) Diffusion of general data on non-flat manifolds via harmonic maps theory: the direction diffusion case. Int J Comput Vis 36(2):149–161. https:// doi. org/ 10. 1023/A: 10081 52115 986 33. Intel Advisor. https:// softw are. intel. com/ conte nt/ www/ us/ en/ devel op/ tools/ oneapi/ compo nents/ advis or. html. Accessed 12 Jan 2024 34. Intel V-Tune. https:// softw are. intel. com/ conte nt/ www/ us/ en/ devel op/ tools/ oneapi/ compo nents/ vtuneprofi ler. html. Accessed 12 Jan 2024 35. SYCL 2020 Specification. https:// www. khron os. org/ regis try/ SYCL/ specs/ sycl2020/ html/ sycl2020. html. Accessed 12 Jan 2024 36. Kronos Group 1.2.1 Specification. https:// regis try. khron os. org/ SYCL/ specs/ sycl-1. 2.1. pdf. Accessed 12 Jan 2024 37. Intel oneAPI GPU Optimization Guide. https:// www. intel. com/ conte nt/ www/ us/ en/ docs/ oneapi/ optim izati onguidegpu/ 2023-1/ overv iew. html. Accessed 8 Jan 2024 38. Laso R, Cabaleiro JC, Rivera F, Muñiz FMC, Alvarez-Dios J (2021) IHP: a dynamic heterogeneous parallel scheme for iterative or time-step methods-image denoising as case study. J Supercomput 77. https:// doi. org/ 10. 1007/ s1122702003260-8 39. Intel Xeon E-2176G Processor. https:// ark. intel. com/ conte nt/ www/ xl/ es/ ark/ produ cts/ 134860/ intelxeone2176gproce ssor12mcacheupto-470ghz. html. Accessed 12 Jan 2024 40. Intel Xeon Gold 6128 Processor. https:// ark. intel. com/ conte nt/ www/ us/ en/ ark/ produ cts/ 120482/ intelxeongold6128proce ssor1925mcache-340ghz. html. Accessed 12 Jan 2024 41. Intel FPGA Arria 10. https:// www. intel. la/ conte nt/ www/ xl/ es/ produ cts/ detai ls/ fpga/ arria/ 10. html. Accessed 12 Jan 2024 S.R.Alcaraz et al. 1 3 42. Zohouri HR, Matsuoka S (2019) The memory controller wall: Benchmarking the Intel FPGA SDK for OpenCL memory interface. In: IEEE/ACM international workshop on heterogeneous high-performance reconfigurable computing (H2RC), pp 11–18 https:// doi. org/ 10. 1109/ H2RC4 9586. 2019. 00007 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.