Full text
AUTOMATIC SERVING-FOCUSED BENCHMARKING OF MODELS August 2025 AUTHOR(S): Tomasz Wojnar CERN, IT-CD-PI SUPERVISOR(S): Raulian-Ionut Chiorescu
CERN openlab Report PROJECT SPECIFICATION Objective: To create an automated system for benchmarking machine learning models, focusing on their performance in production serving environments. Key Tasks: •Automate Performance Evaluation: Benchmark any registered model. •Measure User-Side Performance: Monitor latency and resource consumption (CPU/RAM) to ensure acceptable thresholds. •Identify Hotspots: Use profiling tools to pinpoint performance bottlenecks within models (e.g., via heatmaps). •Benchmark Diverse Runtimes: Evaluate performance across various serving platforms (e.g., Triton, TFServing, TorchServe, MLflow ModelServer, ONNX). Outcomes: This project will ensure quality of service for production models and enhance the debugging experience for performance issues. AUTOMATIC SERVING-FOCUSED BENCHMARKING OF MODELS 1
CERN openlab Report ABSTRACT The increasing use of machine learning models in production environments requires a systematic and standardized approach to evaluating their performance. Due to significant variations in architecture, size, precision and underlying framework, performance prediction and debugging are complex tasks. This project introduces the ML Inference Benchmark Tool: an automated system designed to benchmark registered models from the user’s perspective. The tool provides a thorough analysis of important performance indicators, such as latency percentiles, throughput and resource consumption (CPU, GPU and RAM). Supporting multiple serving runtimes and frameworks, including PyTorch, TensorFlow, and ONNX, it uses deep profiling capabilities to identify computational hotspots at the operator level. Integrating with Kubeflow pipelines and using TensorBoard for rich visualization automates the evaluation process, ensures quality of service for production-grade models, and significantly improves the user experience when diagnosing performance-related issues. AUTOMATIC SERVING-FOCUSED BENCHMARKING OF MODELS 2
CERN openlab Report TABLE OF CONTENTS 1 INTRODUCTION 4 1.1 Machine Learning in Next Generation Triggers . . . . . . . . . . . . . . . . . . . 4 1.2 The Need for Standardized Model Benchmarking . . . . . . . . . . . . . . . . . 4 1.3 Project Goals and Objectives . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2 BENCHMARKING METHODOLOGY 5 2.1 Supported Backends and Frameworks . . . . . . . . . . . . . . . . . . . . . . . . 5 2.2 CorePerformanceMetrics .............................. 5 2.3 Operator-LevelProfiling ............................... 6 2.4 Data Format and Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 3 SYSTEM ARCHITECTURE 8 3.1 CoreComponents................................... 8 3.2 ExternalIntegrations................................. 9 4 TENSORBOARD VISUALIZATIONS 9 5 CONCLUSION AND FEATURE WORK 10 5.1 SummaryofAchievements.............................. 10 5.2 FutureEnhancements................................. 10 AUTOMATIC SERVING-FOCUSED BENCHMARKING OF MODELS 3
CERN openlab Report 1 INTRODUCTION The deployment of machine learning models in production systems has become standard practice across industries. However, moving a model from a research environment to a live productiongrade service introduces a new set of challenges centered on performance, reliability, and resource efficiency. A model that is highly accurate but fails to meet latency requirements or consumes excessive resources is impractical for real-world applications. In this report, we introduce a comprehensive tool designed to automate the performance benchmarking of machine learning (ML) models and address these critical challenges. 1.1 Machine Learning in Next Generation Triggers In high-energy physics experiments such as those conducted at CERN’s Large Hadron Collider (LHC), vast amounts of data are generated at an enormous rate - billions of collisions per second. It is not possible to store and process all these raw data, so real-time trigger systems are used to filter and select only a tiny fraction of ’interesting’ events for further analysis. The upcoming High-Luminosity LHC (HL-LHC) will present even greater challenges, with significantly higher data rates and more complex events [5]. Machine learning is therefore becoming essential for next-generation trigger systems [4,6,1]. Machine learning is pivotal for these next-generation trigger systems for the following reasons: •Ultra-low-latency decisions: ML models, which are often deployed on specialized hardware such as FPGAs and GPUs, enable microsecond-level decisions that are critical for the initial trigger stages [3]. •Enhanced event selection: ML algorithms provide superior precision in classifying events and identifying complex physics signatures, thereby improving the signal-to-noise ratio [3]. •Anomaly detection: A critical application is the use of ML for anomaly detection, which involves identifying unexpected patterns and is essential for the discovery of new physics beyond the Standard Model [2]. The Next Generation Triggers project is a five-year initiative that was launched in January 2024. It explicitly focuses on leveraging advanced AI and computing to address these challenges and maximize the scientific potential of the HL-LHC [1]. Due to strict latency and throughput requirements, it is essential to rigorously and automatically benchmark these ML models to ensure quality of service and facilitate groundbreaking scientific discoveries at the HL-LHC. 1.2 The Need for Standardized Model Benchmarking Deploying ML models in production requires more than just accuracy; it demands guaranteed performance. Ad hoc testing is insufficient for ensuring quality of service. Key challenges include the diversity of ML frameworks, identifying performance bottlenecks, analysing complex metrics beyond simple averages and ensuring reproducible results. A dedicated and standardised benchmarking tool is essential to address these challenges. AUTOMATIC SERVING-FOCUSED BENCHMARKING OF MODELS 4
CERN openlab Report 1.3 Project Goals and Objectives The primary goal of this project was to develop an automated system for evaluating the performance of any registered model, with a focus on metrics relevant to production serving. The key objectives were: •Develop a multi-framework tool capable of benchmarking models from PyTorch, TensorFlow, and ONNX Runtime. •Implement comprehensive metric collection for latency, throughput, and system resource utilization. •Integrate deep profiling tools to generate operator-level performance data and identify computational hotspots. •Provide rich visualizations of results using TensorBoard for intuitive analysis. •Automate the benchmarking process by creating reusable Kubeflow pipeline components. 2 BENCHMARKING METHODOLOGY An effective evaluation of the performance of machine learning models in a production context demands a systematic and robust methodology. The ML Inference Benchmark Tool takes a comprehensive approach, providing reliable, actionable insights. This includes an initial warmup phase to ensure system stability and proper resource allocation, followed by the precise measurement of key performance indicators and detailed operator-level profiling. The tool also offers flexible configuration options. 2.1 Supported Backends and Frameworks The tool supports benchmarking across the most popular machine learning frameworks, ensuring broad applicability: •PyTorch (.pt, .pth), Input (PyTorch tensors) •ONNX Runtime (.onnx), Input (NumPy arrays) •TensorFlow (.keras, SavedModel), Input (TensorFlow tensors) This broad support enables direct comparisons to be made between the performance of models across different frameworks and optimised runtimes. This provides valuable insights for deployment decisions: 2.2 Core Performance Metrics Following the warm-up phase, a comprehensive suite of metrics is collected to understand the model’s behaviour under realistic conditions. These metrics provide a comprehensive overview of performance, which is essential for optimising models for production environments: AUTOMATIC SERVING-FOCUSED BENCHMARKING OF MODELS 5
CERN openlab Report •Latency This quantifies the time taken for a single inference request to be processed from start to finish. Critical for real-time applications where quick responses are important, such as autonomous systems or real-time data filtering in trigger systems. •Throughput Measures the total number of inference requests or data samples that the system can process in a given time period (e.g. per second). Essential for understanding the model’s processing capacity and scalability. High throughput indicates the efficient utilisation of hardware, which is vital for handling large volumes of data or concurrent requests in high-demand scenarios. •Resource Utilisation Tracks the consumption of system resources (e.g. CPU, GPU and memory) during the inference process. Crucial for cost efficiency and identifying potential hardware bottlenecks. Monitoring resource usage helps to select the right infrastructure, optimise deployment configurations and ensure that the model runs efficiently without overor under-provisioning resources. 2.3 Operator-Level Profiling Beyond external metrics such as latency and throughput, it is important to understand the internal execution dynamics of a model in order to optimise it effectively. In machine learning, profiling involves instrumenting the model’s execution to gather detailed timing and resource consumption data at a granular level, often down to individual operations or layers. This process provides a detailed view of where computational time is being spent, thereby pinpointing bottlenecks that might not be evident from high-level metrics alone. By identifying these ’hotspots’, developers can focus their optimisation efforts on the areas that will have the greatest impact, leading to significant performance gains. To pinpoint performance bottlenecks effectively, the tool uses operation-level profiling tools provided by backends and frameworks. This process captures execution times for individual operations within the model graph. The granular data obtained from operator profiling enables developers to: •Identify specific computationally intensive operations. •Evaluate the performance impact of different model architectures or optimisation techniques on individual operators. •Direct targeted optimisation efforts, such as kernel fusion, precision reduction or the development of custom operators. AUTOMATIC SERVING-FOCUSED BENCHMARKING OF MODELS 6
CERN openlab Report 2.4 Data Format and Configuration The input data for our benchmarking follows the KServe standard for inference requests, ensuring a consistent and widely compatible format. This standardisation simplifies integration with various serving runtimes and enables clear communication of input tensor properties. A KServe JSON payload for a single input tensor typically looks like this: { "inputs": [ { "name":"input_tensor", "shape": [-1,3,224,224], "datatype":"FP32", "data": [[[[0.1,0.2, ...]]]] } ] } Listing 1: Example KServe Inference Request Payload. In addition to data formatting, Hydra manages the tool’s execution parameters. This powerful configuration framework enables flexible overriding of parameters such as batch size, the device to be used (CPU or CUDA) and iteration counts for different measurement phases, either directly from the command line or via structured configuration files. This flexibility is essential for adapting benchmarks to different models, hardware and performance objectives. Below is an example of a config.yaml file, which sets default parameters for a benchmark run: logging_level: INFO profiler: _target_: src.profilers.tensorflow.TensorFlowProfiler model_path: data/example_model.keras payload_path: data/example_payload.json config: warmup_iters: 10 op_iters: 10 latency_iters: 100 throughput_seconds: 5 monitor_intervals: 0.1 batch_size: null data_logger: _target_: src.tensorboard_logger.TensorboardLogger logdir: logs/tf input_adapter: _target_: src.input_adapters.tensorflow.TensorFlowAdapter device: cpu Listing 2: Example Configuration File. AUTOMATIC SERVING-FOCUSED BENCHMARKING OF MODELS 7
CERN openlab Report 3 SYSTEM ARCHITECTURE The ML Inference Benchmark Tool has a modular, extensible architecture designed to support diverse ML frameworks, comprehensive metric collection, and easy integration into MLOps workflows. Its core design principles focus on flexibility, reusability, and clarity, enabling easy expansion to new frameworks or profiling techniques. The sequence Fig. 1illustrates the typical flow of interactions when running a benchmark. Figure 1: UML Sequence Diagram of the ML Inference Benchmark Tool’s Execution Flow. 3.1 Core Components The src/ directory structure reflects the logical separation of concerns within the tool: •Profilers (src/profilers): This is the heart of the benchmarking system. –base.py: Defines an abstract base profiler, establishing a common interface for all framework-specific profilers. This ensures consistency in how benchmarks are initiated, metrics are collected, and results are processed, regardless of the underlying ML framework. –torch.py, onnx.py, tensorflow.py: These modules implement the concrete profiling logic for PyTorch, ONNX Runtime, and TensorFlow respectively. They handle the specific nuances of loading models, preparing inputs, executing inference, and utilizing framework-native profiling tools (e.g., PyTorch profiler, TensorFlow profiler) for operator-level insights. •Input Adapters (src/input_adapters/): Responsible for converting incoming benchmark data, which adheres to the KServe JSON format, into the specific tensor or array formats required by each ML framework (e.g., PyTorch Tensors, NumPy arrays, TensorFlow Tensors). This layer provides a standardised external API while accommodating the requirements of the internal frameworks. •Hydra Config (src/config/): Leverages Hydra for robust and flexible configuration management. Users can easily define and override benchmark parameters (e.g. batch size, device and iteration counts) via command-line arguments or structured configuration files. This promotes reproducibility and ease of experimentation. AUTOMATIC SERVING-FOCUSED BENCHMARKING OF MODELS 8