Adding Support of External Metrics to the AdaptivePerf Profiler
Abstract
AdaptivePerf (1) is an open-source, low-overhead code profiler based on the Linux 'perf' tool (2), designed for comprehensive profiling of on-CPU and off-CPU activities. It generates interactive reports, including thread/process trees and flame graphs. In the beginning, AdaptivePerf supported sampling-based "perf" metrics only. The primary objective of this work was to extend AdaptivePerf's functionality by adding support for any metrics produced by external commands given by the user, such as power consumption and GPU memory usage, which enables more detailed analysis per thread/process, e.g. sliced flame graphs corresponding to the peak of a given user metric.
Full text
CERN openlab Report // 2023 ABSTRACT AdaptivePerf (1) is an open-source, low-overhead code profiler based on the Linux 'perf' tool (2), designed for comprehensive profiling of on-CPU and off-CPU activities. It generates interactive reports, including thread/process trees and flame graphs. In the beginning, AdaptivePerf supported sampling-based "perf" metrics only. The primary objective of this work was to extend AdaptivePerf's functionality by adding support for any metrics produced by external commands given by the user, such as power consumption and GPU memory usage, which enables more detailed analysis per thread/process, e.g. sliced flame graphs corresponding to the peak of a given user metric. 2
CERN openlab Report // 2023 TABLE OF CONTENTS 1. INTRODUCTION........................................................................................................................ 5 Context..................................................................................................................................... 5 2. ADAPTIVEPERF........................................................................................................................5 Structure overview....................................................................................................................5 Framework................................................................................................................................6 Code structure..........................................................................................................................7 3. SCOPE OF WORK.................................................................................................................... 8 Requirements........................................................................................................................... 9 1. Reducing Profiling Overhead..........................................................................................9 2. Synchronizing Data from Different Profilers....................................................................9 3. Finding the Right Tools for Measuring Power Usage / CPU Temperature / GPU Temperature....................................................................................................................... 9 4. Displaying Data as Time Plot Along Flame Graphs......................................................10 4. METHODOLOGY AND IMPLEMENTATION........................................................................... 10 1. Familiarization with AdaptivePerf Codebase................................................................ 10 2. Development of a Profiler-Derived Class..................................................................... 10 3. Integration of the New Profiler Class:........................................................................... 11 4. Saving External Metric Samples...................................................................................12 5. Implementation of Time-Ordered Flame Graph Slicing................................................ 12 6. Slicing Custom Time-Ordered Flame Graphs...............................................................12 5. RESULTS AND CONCLUSION...............................................................................................13 3
CERN openlab Report // 2023 1. INTRODUCTION AdaptivePerf is a profiler1that builds on the Linux 'perf' tool. While many profilers are designed to analyze a single type of processing unit, AdaptivePerf is envisioned in such a way to support multiple architectures, including CPUs, GPUs, and FPGAs. This makes it particularly useful in heterogeneous computing environments. Context AdaptivePerf is a project initiated at CERN as part of the broader SYCLOPS EU project (3), which focuses on open hardware acceleration utilizing SYCL (4) and customized RISC-V architectures. The SYCLOPS project aims to create improved solutions for AI and data mining across large and diverse datasets. The project focuses on developing platforms, and tools specifically designed to accelerate AI applications by integrating knowledge from a wide range of fields, including computer architecture, programming languages, systems and runtimes, Big Data, High-Performance Computing, autonomous systems, High-Energy Physics, and precision oncology. SYCLOPS is a collaborative effort involving several partners across Europe, including Codeplay/Intel, EURECOM, ACCELOM, the University of Heidelberg, CERN, Codasip, INESC-ID, and HIRO Microdatacenters. For CERN, the primary responsibilities under the SYCLOPS project include implementing support for heterogeneous architectures in ROOT through SYCL, as well as profiling, benchmarking, and integration testing. AdaptivePerf is a crucial component of these tasks which will be utilized not only for ROOT but also for other SYCLOPS use cases, such as autonomous systems and genomic analysis, and within CERN itself. 2. ADAPTIVEPERF Structure overview AdaptivePerf is composed of three primary components, as illustrated in Figure 1: 1. AdaptivePerf (Core Profiler): This is the central component of the profiler, running on the machine where the target program is executed. Its primary function is to collect runtime data about the program. Within this core component, AdaptivePerf invokes various profiling tools like Linux perf, in parallel with the program being profiled. This 1A profiler is a performance analysis tool used to measure the behavior of software programs during execution. Profilers collect various types of data, such as CPU usage, memory usage, function calls, and execution time, to help developers understand how their code performs and where bottlenecks may occur. The insights gained from profiling are crucial for optimizing code, improving efficiency, and ensuring that software runs as effectively as possible, especially in complex and resource-intensive applications. 4
CERN openlab Report // 2023 allows it to extract stack traces and other performance metrics, which are subsequently utilized by AdaptivePerfHTML to present the information in an intuitive manner. 2. adaptiveperf-server: The role of this component is to receive and parse the data collected by the core AdaptivePerf component. By default, adaptiveperf-server is executed on the same machine as the profiling program, however, it can also be deployed on a different machine. This feature gives users an option to move compute-intensive result processing to another computer in case the profiled one cannot reliably handle both profiling and processing. 3. AdaptivePerfHTML: This component is responsible for reading and interpreting the data stored by the adaptiveperf-server. It renders the information in a user-friendly format, utilizing various graphical elements such as thread trees, off-CPU/on-CPU time per thread, and navigable flamegraphs2. Figure 1. Abstracted structure of AdaptivePerf Framework The core and server parts of AdaptivePerf are mostly written in C++ which is suited for developing a performance profiling tool, where efficiency, low overhead, and tight system integration are really important. Additionally, the "perf" Python API is used for reading and parsing data produced by "perf" in real-time. For AdaptivePerfHTML, JavaScript and Flask are used. Flask is a web framework for Python. It is designed to be simple and flexible, providing the essential tools for creating web servers and handling requests. Flask was chosen for its simplicity, modularity, and ease of use. However, for 2A flame graph is a visual tool used to represent the performance of a software program by showing the call stack and the time used by the individual functions over time. It displays functions as horizontal bars, with the width of each bar indicating the amount of time spent in that function. The graph's hierarchical structure helps identify "hot spots" where the program consumes the most resources. 5
CERN openlab Report // 2023 this project, C++ will also be incorporated through WebAssembly (WASM)3, to make the AdaptivePerf data processing faster. Code structure Figure 2. AdaptivePerf and adaptiveperf-server dependency tree 3WebAssembly (WASM) is a low-level, binary format designed to run code efficiently in web browsers. It enables high-performance execution of code written in languages like C, C++, and Rust by compiling them to a format that can be executed in the browser alongside JavaScript. WASM is known for its speed and portability, allowing complex applications, such as games or heavy computational tasks, to run in the browser with near-native performance. 6
CERN openlab Report // 2023 As shown in Figure 2, both AdaptivePerf and adaptiveperf-server have their own entry points, where the input commands are parsed. To facilitate ease of use and debugging, a CLI library (5) is employed for this purpose. AdaptivePerf and adaptiveperf-server are designed with parallelization to ensure minimal profiling overhead. On the core side, a separate profiler instance is created for each type of metric, with each instance running processes in parallel with the program being profiled. Correspondingly, on the server side, one or more subclient threads are initiated for each profiler instance to receive the metric samples. These samples are transmitted either through a pipe if the server is on the same machine as the profiling process, or via TCP if the server is hosted on a separate machine. Each subclient processes the received samples to build a flame graph tree structure. As illustrated in Figure 2, both AdaptivePerf and adaptiveperf-server utilize a JSON library (6) for C++ to serialize and transmit these samples between the core and the server. Once all samples are processed by the subclients, the main client gathers the data from each subclient to create a comprehensive JSON structure that stores the flame graphs corresponding to different metrics. For each metric, two types of flame graphs are generated: non-time-ordered and time-ordered. 3. SCOPE OF WORK The scope of my work involved implementing profiling of external metrics in AdaptivePerf to enable real-time measuring of various signals from processing units. This required ensuring compatibility between calling user-provided external commands and the existing Linux perf-based framework, as well as maintaining the performance and accuracy of the profiling data. Key challenges included dealing with the diverse nature of external programs, ensuring seamless data aggregation within AdaptivePerf, and displaying data from them alongside flame graphs intuitively. 7
CERN openlab Report // 2023 Figure 3. Visualization of continuous metrics alongside time-ordered flame-graphs Requirements 1. Reducing Profiling Overhead Profiling overhead refers to the additional time and resources consumed by the profiling tool itself while measuring a program's performance. The goal is to minimize this overhead to ensure that the profiler does not significantly alter the behavior of the program being analyzed. This is crucial for obtaining accurate performance data, as excessive overhead can skew the results and lead to misleading conclusions about the program's performance. Low overhead is also important for a good user experience, as profiling a program at runtime should add almost no extra time compared to running the program normally. 2. Synchronizing Data from Different Profilers When using multiple programs to monitor various aspects of a system (such as CPU performance, power consumption, and temperature), it is essential to synchronize the data collected by these tools. Synchronization ensures that the data from different sources aligns correctly in time, allowing for accurate analysis and correlation of different metrics. This involves implementing mechanisms to timestamp the data accurately with a common time source, and aligning the data streams so that they can be analyzed together in a coherent manner. 3. Finding the Right Tools for Measuring External Metrics Identifying and integrating the appropriate tools to measure relevant metrics such as power usage etc. is critical for comprehensive profiling. This involves researching and selecting tools or APIs that are compatible with the target system and capable of providing accurate, real-time measurements. The selected tools must also be compatible with AdaptivePerf's architecture. This may require testing different options, evaluating their accuracy, overhead, and ease of integration, and ensuring they provide the necessary data for analysis. 8
CERN openlab Report // 2023 4. Displaying Data as Time Plot Along Flame Graphs One of the goals is to enhance the visualization of profiling data by displaying it as time plots alongside flame graphs. Flame graphs provide a static view of where the program spends most of its time, but by adding time plots, users can also see how certain metrics (like CPU temperature or power consumption) change over time in relation to the program’s execution. This combined visualization allows for more dynamic and insightful analysis, making it easier to correlate performance bottlenecks or spikes in resource usage with specific moments in the program's execution. Implementing this requires designing a user interface that can smoothly integrate both types of visualizations, ensuring they are easy to navigate and interpret. An example of how this can be visualized is shown in Figure 3. 4. METHODOLOGY AND IMPLEMENTATION To integrate external metrics into AdaptivePerf, a modular approach was implemented to ensure that each profiler could be added without disrupting the existing framework. The project development was managed using Git, with AdaptivePerf being hosted on GitHub. Additionally, to enhance code readability, comprehensive documentation (7) was created using Doxygen, a tool that generates documentation directly from source code comments. The project implementation was structured around several key tasks, each contributing to the development and enhancement of AdaptivePerf. The following steps represent the current work made for implementing profiling of external metrics in AdaptivePerf: 1. Familiarization with AdaptivePerf Codebase The first task involved getting acquainted with the AdaptivePerf and AdaptivePerfHTML codebases. This included setting up the environment by compiling and installing AdaptivePerf on a profiling node and setting up AdaptivePerfHTML as a Python package. Initial profiling tests were conducted using basic commands to ensure proper setup and to explore various profiling options. 2. Development of a Profiler-Derived Class A new class was created. It is derived from the Profiler base class and designed to capture numeric metrics from user-specified commands. This class executes the metric-reading command at a user-defined frequency, parses the output to extract the relevant metric, and transmits this data to the adaptiveperf-server as a JSON array. Fork and exec system calls were utilized for command execution, ensuring safety and efficiency. 9