scieee AI-readable full text Open interactive document viewer

Burst Computing: Quick, Sudden, Massively Parallel Processing on Serverless Resources

Barcelona-Pons, Daniel; Arjona, Aitor; Garcia Lopez, Pedro; Molina-Giménez, Enrique; Klymonchuk, Stepan

Abstract

This is the official published version of the paper presented at USENIX ATC 2025. The original publication is available at https://www.usenix.org/conference/atc25/presentation/barcelona-pons.

Full text

This paper is included in the Proceedings of the 2025 USENIX Annual Technical Conference. July 7–9, 2025 • Boston, MA, USA ISBN 978-1-939133-48-9 Open access to the Proceedings of the 2025 USENIX Annual Technical Conference is sponsored by Burst Computing: Quick, Sudden, Massively Parallel Processing on Serverless Resources Daniel Barcelona-Pons, Universitat Rovira i Virgili and Barcelona Supercomputing Center; Aitor Arjona, Pedro García-López, Enrique Molina-Giménez, and Stepan Klymonchuk, Universitat Rovira i Virgili https://www.usenix.org/conference/atc25/presentation/barcelona-pons Burst Computing: Quick, Sudden, Massively Parallel Processing on Serverless Resources Daniel Barcelona-Pons†‡ Aitor Arjona†Pedro García-López† Enrique Molina-Giménez†Stepan Klymonchuk† †Universitat Rovira i Virgili ‡Barcelona Supercomputing Center Abstract We present burst computing, a novel serverless solution tailored for burst-parallel jobs. Unlike Function-as-a-Service (FaaS), burst computing establishes job-level isolation using a novel group invocation primitive to launch large groups of workers with guaranteed simultaneity. Resource allocation is optimized by packing workers into fewer containers, which accelerates their initialization and enables locality. Locality significantly reduces remote communication compared to FaaS and, combined with simultaneity, it allows workers to communicate synchronously with message passing and group collectives. Consequently, applications unfeasible in FaaS are now possible. We implement burst computing atop OpenWhisk and provide a communication middleware that seamlessly leverages locality with zero-copy messaging. Evaluation shows reduced job invocation and communication latency for a 2 × speed-up in TeraSort and a 98.5% reduction in remote communication in PageRank (13 × speed-up) compared to standard FaaS. 1 Introduction The cloud offers compute infrastructure on demand, but provisioning, adjusting, and managing these resources for largescale data processing applications is an arduous task, especially for non-experts. Furthermore, when the load is unpredictable, dynamic, with varying volumes of data, user-driven, and sometimes interactive, finding the right scale to avoid misprovisioning [25,39,54] becomes very complex. Function-as-a-Service (FaaS) has gained traction as a solution to the resource provisioning problem as it offers rapid, on-demand, no-ops scaling and a pay-as-you-go billing model at very fine granularity (MB per ms). Users do not need to set up a cluster, but the service simply accepts function invocations and fully manages the rest. Moreover, its resource burstability has set FaaS aside from traditional engines like Spark or Dask, allowing to start thousands of short-lived functions in seconds instead of minutes (see Table 1). Several research works [ 15 , 16 , 23 , 44 ] have used FaaS for a myriad of dataand compute-intensive tasks. Table 1: Time to provision cloud compute resources on different services and technologies. Technology Total vCPUs NodesaStart-up time EMR Spark 96 6 296 s 24 431 s Dataproc 96 6 95 s 24 113 s Dask 128 8 184 s 64 253 s Ray 128 8 187 s 64 229 s Knative (Kubernetes) 960 960 54 s OpenWhisk 960 960 21 s AWS Lambda (2 GiB) 960 960 6 s Burst Computing 960 960 1.7 s a AWS EMR Spark and GCP Dataproc use m5 and E2-standard VM families, respectively. Dask and Ray are deployed on user-managed m6i family EC2 VMs. Knative, OpenWhisk, and Burst deployed on 20 c7i.12xlalrge VMs but start 960 functions/workers. This has brought a new concept in cloud computing that refers to the ability to quickly respond to sudden, parallel workloads without provisioning a cluster in advance. Fouladi et al . [16] talk about a “burstable supercomputer-on-demand” and a “burst-parallel swarm of thousands of cloud functions, all working on the same job.” However, literature admits that the current FaaS model is too narrow and precluding for massively parallel data processing programs (MPP) [24]. In this work, we highlight that the key issue of FaaS hindering burst-parallel jobs is its lack of group awareness. Indeed, FaaS users need multiple independent service calls to spawn a fleet of workers, which become strongly isolated from each other. We note that such fine-grained isolation is damaging and unnecessary for collaborative jobs, and thus propose to raise the multi-tenant boundaries to the job level. We present burst computing, a new cloud computing model to deal with quick, sudden, massively parallel workloads, which we call bursts. To this end, we offer a group invocation primitive to handle the whole job as a unit. To the best of our knowledge, we are the first to implement this feature USENIX Association 2025 USENIX Annual Technical Conference 39 in a FaaS platform, clearly differing from all other research efforts that suffer the burden of handling and orchestrating individual function invocations. A group invocation allows to optimize resource allocation, ensure worker parallelism, and perform packing: running multiple workers co-located in the same environment. In addition to speed up worker start-up latency, this enables worker locality and simultaneity, which can be exploited to improve code and data loading, and to aid powerful worker-to-worker communication patterns (e.g., broadcast, all-to-all) that seamlessly leverage shared memory channels with zero-copy mechanisms. From the user side, a burst spawns a fleet of workers that communicate with message-passing, a simple but very powerful abstraction that creates a novel serverless substrate versatile to many applications beyond what is feasible in FaaS. This gives extensive control of the job to advanced users and allows to design compute engines or frameworks on top (e.g., DAGbased) to simplify development, manage the execution, and handle failures. Massively parallel computations are prime burst applications, especially when sudden, unpredictable, and user-driven in nature. Batch-like jobs are also good candidates when run interactively. Some examples are data processing, analytics, and machine learning workloads like exploratory model tuning, SQL, k -means, and large-scale sorting. Bursts may be stateless (e.g., grid search or Monte Carlo simulations) or stateful (e.g., table joins and aggregations). We make the following contributions: • We present burst computing, a novel cloud service model for short, sudden, massively parallel jobs (bursts). We believe that no cloud vendor or research effort has created the necessary substrate to support them. • Burst computing evolves FaaS with a key novel group invocation primitive (a flare) that raises multi-tenant isolation from a single function invocation to the whole job. In consequence, the system launches massive process groups faster, with guaranteed parallelism, and packs workers together to exploit locality. • We implement a burst computing platform by extending OpenWhisk, a state-of-the-art FaaS system. Our implementation includes a specialized Rust worker runtime and a burst communication middleware that seamlessly leverage worker locality with collective code/data loading and zero-copy messaging. • Under evaluation on several burst-parallel workloads against FaaS, burst computing improves job invocation latency (up to 11.5 × faster), worker simultaneity (up to 26.5 × lower median absolute deviation), and group communication (up to 98% in a broadcast), for a speedup of 13×in PageRank and 2×in TeraSort. 2 Motivation: in search of burstability Many works are leveraging serverless services for massive data processing [ 3 , 7 , 9 , 15 , 16 , 23 , 40 , 60 ] despite current (FaaS) hindrances [ 5 , 18 , 24 ] due to the resource burstability 0 2 4 6 8 10 Time (s) 0.0 0.5 1.0 CDF 100 x 256 MiB 100 x 10 GiB 1000 x 256 MiB 1000 x 10 GiB 1 Figure 1: Start-up time (cold start) of 100 and 1000 FaaS functions in AWS Lambda for two memory sizes. of this model [ 17 , 38 ]. Applications benefit from quick, ondemand, no-ops resources at very fine granularity, and pay precisely for what they need, when they need it. This has brought what we call bursts, massively parallel processing (MPP) workloads that appear suddenly and process large, variable volumes of data in a very short time (under 1 or 2 minutes). Such applications have dynamic resource needs that cannot be predicted easily, thus serverless burstability becomes essential [ 49 , 50 , 52 , 56 ]. Consider, e.g., an interactive, scientist-driven workflow in a Jupyter Notebook, where the user dynamically explores large datasets and modifies parameters that significantly impact the workload size. Although they may resemble batch jobs, they are characterized by their sudden, sporadic occurrence, highly dynamic and unpredictable data volumes, and the expectation of low-latency execution. Representative use cases include interactive model tuning via grid search, exploratory data analysis with SQL queries and algorithms such as logistic regression, and data preparation operations like filtering and sorting. The MilliSort and MilliQuery benchmarks [ 31 ] exemplify such short-running workloads. Additional scenarios include real-time data stream or video feed processing, where both data volume and analytical complexity may fluctuate dramatically over time. Current data processing solutions such as Spark, Dask, Flink, or Ray fail to support bursts. A long-lived deployment of these engines is impractical, as it would easily become misprovisioned. They are not offered as a service by any cloud either, which could palliate the issue by multiplexing jobs from multiple tenants, and thus forces per-user deployments that are too slow to set up [ 38 ], even on cloud-managed offerings (e.g., Amazon EMR). Table 1shows that starting one of these technologies is intolerable for critical sporadic or dynamically sized applications. In contrast, FaaS services provide a large-scale compute substrate much faster. Fig. 1shows that AWS Lambda may spawn a fleet of 1000 functions in 6 s ; a much more appropriate time range for bursts. 1 FaaS is also more attractive than Container-as-a-Service (CaaS) or managed Kubernetes ser1 Note that small functions ( 256 MiB ) incur higher invocation latency than large ones ( 10 GiB ). Also found on other providers (e.g., GCP), it is likely due to the overhead of scheduling finer-grained resources. 40 2025 USENIX Annual Technical Conference USENIX Association Client Controller Instance 0 Instance 3 Instance 1 Instance 4 Worker 0 Worker 1 Worker 3 Worker 4 Invoke (x6) A1 A2 FaaSPlatform Client Controller Pack 0 Pack 1 Flare B1 B2 BurstPlatform Instance 2 Worker 2 Instance 5 Worker 5 Worker 0 Worker 1 Worker 3 Worker 4 Worker 2 Worker 5 B3 g=3 Remote indirect Shared memory Worker communication: A3 External Comms. Server External Comms. Server Figure 2: Running a data processing job of 6 workers in FaaS and burst computing with granularity (g) 3. vices due to simpler abstractions [ 26 ] and quicker resource allocation. For example, Knative, a Kubernetes-based FaaS-like implementation, is noticeably slower in spawning workers than dedicated FaaS platforms (Table 1). 2.1 FaaS is holding us back A review of the literature will show us that running bursts atop FaaS brings many challenges [ 18 , 24 , 38 ]. We highlight three friction points: (F1) worker isolation, (F2) job fragmentation with complex orchestration, and (F3) huge data movement. To illustrate them, Fig. 2follows the execution of a parallel job on a FaaS platform. It shows a parallel job with 6 workers. The job could be embarrassingly parallel (stateless) such as a data filtering, or require the workers to coordinate at some point (stateful) such as a table join or, more intensively due to its iterative nature, PageRank. F1 appears because multi-tenant isolation is at the level of a function invocation. FaaS spawns function instances independently, one at a time, requiring multiple HTTP requests A1 to obtain the 6 workers. Besides the added latency of several requests, this is an issue for parallel jobs because the platform is not aware of these workers being collaborators, and thus cannot guarantee their parallelism. This creates delays or skews between workers that potentially harm job execution. Take, for instance, Fig. 1, where the last function starts up to 6 s after the first one. 2 Even more, the platform populates identical environments (instances) for each invocation A2 , 2 Further evaluation (not shown in the plot) reveals that this dispersity may increase to 44 s in GCP, or 20 s in an OpenWhisk deployment. which stresses the system with code, dependency, and data 3 loading that creates memory duplication [41,53]. F2 occurs when workers need to coordinate. For instance, TeraSort à la MapReduce includes a data shuffle amidst the job, and PageRank iteratively globally aggregates a vector. Workers cannot communicate effectively because they may not exist at the same time (F1). Instead, workers read and write intermediate data asynchronously through an external storage solution. This pattern (depicted in Fig. 3) creates job fragmentation (function stages) and complicates its orchestration, especially in iterative algorithms like PageRank (unfeasible with this approach [ 6 ]). First, it increases data movement and requires worker recreation at each stage (adding code and data loading overhead). Second, it needs an active workflow orchestration process to monitor the state of workers and oversee the overall job progress. 4 Centralized solutions add a mostlyidle driver component while decentralized task scheduling adds a layer of complexity to applications [ 9 , 29 , 32 ]. Neither can solve the underlying problem of worker isolation. These issues are emphasized by F3. With so many tiny isolated workers, most communication patterns (e.g., a shuffle) require numerous remote connections A3 . In data processing workloads, this may result in very large data transfers, precluded by the (FaaS) lack of direct communication [10,35]. 3 Burst Computing Burst computing is a novel paradigm for running bursts in the cloud. It overcomes the above frictions with two key principles that evolve FaaS: group awareness and locality exploitation. Fig. 2shows how this changes job invocation. FaaS hinders worker collaboration because multi-tenant isolation is at the level of a single function (F1). Because a job belongs to a single tenant, it makes sense to raise isolation to the job level and handle all its workers as a group. To this end, burst computing provides a group invocation primitive, which we call flare B1 , to instantly launch massive process groups with guaranteed parallelism. To the best of our knowledge, we are the first to implement this kind of primitive in a serverless system. Flares bring group awareness to the service, which is key to perform worker packing B2 , i.e., running multiple workers of a job in the same isolated environment. Packing establishes worker locality and enables several optimizations discussed below. In a flare, all workers have guaranteed parallelism and access to the job context (e.g., the burst size, IDs, or locality), which allows them to communicate synchronously in patterns unfeasible in FaaS, such as worker-to-worker message passing and collectives, that simplify job orchestration and avoid F2. This difference is depicted in Fig. 3.F3 is addressed because communication B3 can seamlessly exploit locality and use shared memory mechanisms between workers in the same pack, which reduces remote transfers. 3For instance, hyperparameter tuning uses the same data in all workers. 4 This can be painful since FaaS does not provide monitoring mechanisms. USENIX Association 2025 USENIX Annual Technical Conference 41 Client/Orch. 𝜆1 𝜆2 Code/data get/send Compute Client/Trigger FaaS Burst Communication ···𝜆3 𝜆4 𝜆1 𝜆2 Wait Service delay Flare Figure 3: Timeline comparison of a parallel job with FaaS and burst computing. 3.1 Worker packing and communication Worker packing To run a flare, the burst platform allocates n workers into m packs; we say that n is the burst size. The number of workers per pack is the burst’s granularity ( g= n/m ). Thus, Fig. 2shows a burst of size n=6 where, by setting g=3 , the platform only spawns 2 packs, each with 3 workers. The higher g , the lower m , reducing the number of environment creations, which is a critical part of function invocation time in FaaS. Then, worker code and dependencies are loaded only once per pack and shared by all co-located workers. This further helps with initialization time (especially when dependencies are large) and optimizes resource usage (e.g., avoiding memory duplication [ 41 ]). A similar reasoning applies for data loading: workers processing the same data (like in hyperparameter tuning) download it just once per pack and utilize their aggregated resources to speed up the transfer (i.e., with parallel downloads). Choosing gis a trade-off between ease of system management and locality maximization. To illustrate that, we identify three strategies 5 for worker packing: (i) heterogeneous, where workers are placed in containers as big as possible in the underlying system machines; (ii) homogeneous, where workers are placed in fixed-size containers; and (iii) mixed, where workers are put in fixed-size packs, but if multiple packs fall onto the same machine, they are merged into a single container. The first approach maximizes locality, but it can become a resource scheduling problem, as it is prone to fragmentation. The homogeneous packing mitigates that issue, but it restricts worker locality. The third strategy is the compromise that allows a fast and flexible management while still maximizing locality (see §5). Given this complexity, we argue that the responsibility for setting the granularity should lie with the platform rather than the user, enabling better control 5 Strategies must consider how many resources we assign to each worker. For simplicity, this paper considers only vCPUs and applies 1 vCPU per worker, but the strategies work for any such assignment. over resource scheduling and providing a more streamlined and user-friendly service. Worker communication Burst applications are elastically distributed and collaborative. They are coded as a single function run by all workers that accepts any worker multiplicity transparently. Then, because workers are guaranteed to be parallel, they may coordinate synchronously by sending messages and with common communication patterns. To simplify this, burst computing includes a worker-toworker, message-passing communication middleware readily available to workers. The middleware seamlessly identifies messages between workers placed in the same pack for local communication (zero-copy). Only messages between packs are transferred remotely, and the middleware optimizes these connections (e.g., a broadcast only sends one message per pack). Remote delivery may be implemented with several technologies. Our contributions are independent of this choice because burst computing reduces any remote communication through packing. In this work, we follow the usual approach in FaaS and only consider indirect solutions using an external communication server B3 . 4 Design and implementation We put the above ideas into a prototype burst computing platform and communication middleware. Here we provide the design details and implementation. Fig. 4shows an overview of the main components and their interactions. The burst platform extends the design of a FaaS platform to implement group invocation and worker packing. Built atop Apache OpenWhisk, our platform shares its components with important modifications (see §4.4). The controller manages user interaction with the platform, it handles inbound HTTP requests to deploy and invoke bursts, oversees system resources, and performs worker packing. A database stores the burst definitions and configuration, as well as the results and execution metadata. Computational resources in the platform are provided by the invokers, a set of machines with capacity for burst packs. Packs are run in containers that isolate a custom runtime environment to run workers. Our burst communication middleware (BCM) has two main components: the core communication library and the remote backends. The library exposes message-based communication to workers, and it is extensible with backends to use different remote message delivery solutions. 4.1 Life cycle overview Fig. 4depicts the life cycle of the system. To deploy a new burst definition, the user first sends 1 a deploy HTTP request. The controller receives it and registers 2 the new definition in the database. Later, when the user desires to trigger the execution of the burst, they send 3 a flare HTTP request with specific parameters. The controller handles the invocation and decides worker allocation 4 based on the current state of the invoker machines. The affected invokers 42 2025 USENIX Annual Technical Conference USENIX Association Burst platform deploy() Controller Invokers Remote communication backend 1 3flare() Pack i Pack (another burst) BCM Burst communication middleware Shared memory Workers BCM 6 - Burst definitions - Results - Metadata 2 5 4 Figure 4: Burst computing platform overview. receive the task to spawn the required runtime environments (packs) with space for as many workers as needed. When the environments boot, their host invoker tells them which burst definition and parameters to load 5 from the database. Then, each pack spawns its workers internally, which will execute the user-defined function ( work in Table 2) in parallel. Workers may use the BCM to coordinate and share data. This seamlessly uses shared memory or remote connections to communicate workers 6 in the same or a different pack, respectively. Additionally, workers may read or write data to external storage systems (e.g., object storage) or produce a result that is stored back to the database, where it may be retrieved later by users through another HTTP request. 4.2 Developing and running bursts User experience is key for burst computing. As a serverless service, all resource management remains hidden. Users interact with the service through a simple interface that allows to define bursts with resource-agnostic code and to schedule their execution. This is similar to FaaS services that allow users to upload their function definitions and then set up triggers or invoke them as needed. The interface and abstractions are summarized in Table 2. Deployment Similar to functions in FaaS, developers package and upload their burst definitions (code) to the cloud, giving them a name and configuration. The configuration includes runtime parameters and worker characteristics (such as language and memory size). Invocation Burst definitions are triggered for execution like functions in FaaS: an event or HTTP request notifies the intent to execute a burst with specific input parameters. We call each burst invocation a burst flare (Table 2). The main difference with FaaS is that a flare will spawn a group of parallel workers (instead of a single function instance). The service ensures that all workers run simultaneously and applies packing. In our prototype, the burst size is explicit on the size of the inputParams array. Hence, users have direct control over it. We believe this to be important because parallelism is strictly application-specific and depends on data volume (e.g., ETL Table 2: Burst computing abstractions and API. Interface Functions Burst deploy(defName,package,conf ) Service upload and deploy a burst definition flare(defName, [inputParams]) invokes a burst Burst abstract work(inputParams,burstContext) Function function to run on each worker Burst workerID →unique ID of this worker within the flare Context burstSize →number of total workers in the flare packID →unique ID of the current pack packSize →number of workers in the current pack numPacks →number of packs within the flare belongToPack(workerID)→packID returns the pack ID to which a worker belongs to isPackLeader() →bool returns true if this worker is its pack’s leader Comm. send(data,dest)→none Primitives recv(source)→data broadcast(data,root)→data allToAll([data])→[data] reduce(data,f(data,data)→data)→data tasks), data content (e.g., dimensionality or sparsity), or algorithm configuration (e.g., the number of clusters in k -means). Smart burst sizing is left for future work, i.e., the platform may automatically calculate the number of workers based on application and data information. Coding Burst definitions are coded as a single function that is run by each worker in the burst ( work in Table 2). This function must be programmed elastically so that it accepts and runs correctly for any burst size. The code is also agnostic to the packing performed by the service. To that end, the work function receives a burst context object through which each worker may obtain information about the worker distribution within the particular flare. For example, a worker can query its own unique ID, the burst size, granularity, or which workers belong to each pack (Burst Context in Table 2). With this information (provided by the platform invoker), the code can implement logic to apply locality optimizations at the pack and burst levels (see an example in §5.4.1). This context object also gives access to the BCM. Communication interface The BCM offers simple yet powerful worker-to-worker communication through message passing similar to MPI. The abstractions are elastic (adapt to the burst size) and available through the burst context. Burst computing programs make use of two basic primitives to connect workers: send and receive. These primitives enable point-topoint communication between workers and are designed to send arbitrary volumes of data efficiently within the burst. To facilitate common communication patterns in parallel jobs, bursts may also use group collectives. As listed in Table 2, our prototype implements broadcast, all-to-all, and reduce. Primitives and collectives are locality-aware, although the USENIX Association 2025 USENIX Annual Technical Conference 43 fn work(params: Input, burst: &BurstContext) -> Output { let num_nodes =params.num_nodes; let mut page_ranks =vec![1.0 /num_nodes; num_nodes]; let mut sum =vec![0.0; num_nodes]; let adjacency_matrix =get_adjacency_matrix(¶ms); while err <ERROR_THRESHOLD { page_ranks =burst.broadcast(page_ranks, ROOT_WORKER); for (node, links) in graph { for link in links { sum[*link] += page_ranks[*node] /out_links(*node); } } let reduced_ranks =burst.reduce(sum, |vec1, vec2|{ vec1.zip(vec2).map(|(a, b)|a+b).collect() }); if burst.worker_id == ROOT_WORKER { err =calculate_error(&page_ranks, &reduced_ranks); page_ranks =reduced_ranks; } err =burst.broadcast(err, ROOT_WORKER); reset_sums(&mut sum); } Output { page_ranks } } Figure 5: Simplified source code of the PageRank work function for burst computing. The accesses to the burst context to obtain the worker ID or communicate are highlighted. programs remain agnostic to it, i.e., co-located workers (same pack) communicate on shared memory and only remote workers hit the network. 4.3 Application example Fig. 5shows an example in Rust code (simplified) of the work function that implements the PageRank application. The algorithm consists of an iterative process in which each worker holds a portion of the adjacency graph (relating links between web pages). In each iteration, the new global ranks are computed in parallel, aggregated, and reduced in a tree structure, then broadcasted from the root worker to the rest of them. The algorithm runs until it converges past a threshold or reaches a limit of iterations. Similarly to the MPI computing model, all workers execute the same code but perform different logic based on the worker ID (the rank in MPI). The example highlights the worker accesses to the BurstContext object to perform collectives and obtain information about the current flare. For example, it is used to perform a collective broadcast to share the updated ranks vector, and later a reduce to aggregate the partial ranks computed among the workers. It also shows how a worker checks its ID when it needs to calculate the convergence, since this is only done by the root worker after collecting the aggregated vector in the reduce. 4.4 Burst platform implementation The prototype implementation is built on top of the popular Apache OpenWhisk platform (v1.0.0). We used OpenWhisk as the basis because it is a well-known, open-source, production-tested FaaS implementation and provides higher burstability than other platforms like Knative (Table 1). Our changes amount to approximately 2K SLOC . They affect the main components of the platform, including the controller, the invoker, and the runtime environment. The controller now supports two new HTTP endpoints for bursts: deploy and flare . It also implements the logic to handle them (§4.1). This includes the packing strategy in the three flavors (§3): heterogeneous, homogeneous, and mixed. Granularity can be configured. In any case, the controller calculates the number and size of the packs based on the specific burst size and the resources available in the invokers. Invokers run a new monitoring logic that can be adjusted to report their load to the controller based on CPU instead of RAM. Our prototype is set to assign 1 vCPU per worker because bursts tend to be compute-intensive jobs and we do not consider parallelism within a worker, 6 but other configurations are possible. Invokers also implement new logic to support the creation and execution of packs, spawning Docker containers of the appropriate size for each burst (by specifying resource limits) and telling each container/runtime the number of workers to run, plus their IDs and context. Containers are currently not reused across bursts. For the runtime, we adapted the official OpenWhisk Rust environment, but it is possible to support others. The new logic allows to spawn multiple workers within it as requested by its host invoker. In particular, the Rust runtime spawns one thread per worker to provide parallelism. Finally, the runtime also includes our BCM built-in. 4.5 BCM implementation The burst communication middleware (BCM) is coded in Rust in about 5K SLOC . It is readily available for our custom Rust runtime and we are working on a binding for Python. 7 It enables the transmission of intra-pack (zero-copy) and interpack (via remote backend) messages. The BCM is instantiated by the runtime (once per pack) and made available to workers as a parameter (in the work function as shown in Table 2). For local communication, BCM uses in-memory queues to send and receive data between workers in the same pack. In the Rust runtime, workers are threads and reside in the same memory space, so shared memory mechanisms are not necessary (e.g., shm_open or mmap ). Instead, workers just pass memory pointers between them. Thanks to Rust’s memory safety guarantees, access to shared data is thread-safe. Rust also provides a reference-counting mechanism for immutable data, so shared data is released when it is no longer used at runtime. For example, the root worker in a broadcast sends a read-only memory pointer to its local workers, and they safely access the message concurrently. To modify the data, one may use mechanisms such as copy-on-write. For remote communication, each pack has a shared connection pool to the remote backend, which allows each worker within the pack to send and receive messages concurrently, with the goal of maximizing the container’s bandwidth. This 6The burst size (number of workers) determines total job parallelism. 7 Other languages may be supported through bindings (Java, C++, Go. . . ). 44 2025 USENIX Annual Technical Conference USENIX Association is especially useful in primitives like all-to-all, where all workers must open channels with all the others. For large messages, the data is split into smaller chunks that are sent and received concurrently. This maximizes network utilization and allows readers to start receiving data from the first chunk, instead of waiting for the full message to be available at the backend. The BCM is extensible, allowing the implementation of more remote backends. Currently, we support Redis, DragonflyDB, RabbitMQ, and S3. The backend interface differentiates between sending direct messages (one-to-one) and broadcast messages (one-to-many). The reason is that direct messages are read only once, while broadcast messages multiple times, so we want to optimize this particular case. For instance, in RabbitMQ, one-to-one messages use direct brokers, while one-to-many use fan-out brokers. To ensure that no messages are lost (at-least-once delivery semantics), the BCM relies first on the backend delivery guarantees (e.g., RabbitMQ uses durable queues to avoid dropping messages). Additionally, the BCM keeps a count of direct messages sent between each pair of workers, and for each collective operation. The middleware handles duplicate and/or out-of-order messages. For that, messages include a header with the source and destination worker, collective type, counter, and, if chunked, the number of chunks and chunk number. Messages with a counter lower than the expected value are ignored and assumed as already processed. Those with a counter greater than expected are cached locally until needed. For chunked messages received out-of-order, a memory region is reserved for the total payload and chunks are written to their respective offset as they come in. 5 Evaluation Our evaluation aims to assess burst computing against current FaaS on the three friction points described in §2.1. Importantly, we show and analyze the effects of worker packing and locality. All experiments run on Amazon Web Services (AWS) in the us-east-1 region. 5.1 Burst group invocation Group invocation is the key element against friction F1. Here we evaluate how job-level isolation improves worker readiness time (invocation latency), ensures their simultaneity, and provides locality for collaborative code and data loading. Setup: The burst platform runs on an Amazon EKS cluster, with the control plane on a t4i.xlarge VM (4 vCPUs and 16 GB RAM), and the invokers on up to 20 c7i.12xlarge VMs (48 vCPUs and 96 GB RAM). This gives us space to accommodate up to 960 workers with 1 vCPU each. Impact on burst invocation latency First, we use the homogeneous packing policy to evaluate how assigning different granularity ( g ) affects burst invocation latency. The exploration is depicted in Fig. 6for two bursts of sizes 48 (left) and 960 (right). 8 It is quickly apparent that as g increases (up to 8 We conducted experiments with similar results for burst sizes in-between. FaaS 2 4 6 12 24 48 0 5 10 15 20 FaaS 2 4 6 12 24 48 0 5 10 15 20 Start-up time (s) Granularity Figure 6: Worker start-up latency distribution within one job for burst computing with different packing granularity and FaaS (equivalent to g=1 ). Left and right show, respectively, burst sizes of 48 and 960. Last worker init Simultaneity Worker life-time 0 250 500 750 1000 0 5 10 15 20 25 0 250 500 750 1000 0 5 10 15 20 25 Time (s) # activation Figure 7: Simultaneity (number of workers running at an instant) in FaaS (left) and Burst with g=48 (right). Each bar represents the life-time of a worker. 48 in both cases), the start-up time decreases, and generally becomes more consistent across workers for all burst sizes. For instance, the latency of having all workers ready in a burst of size 960 reduces by 11.5 × from g=1 (FaaS) to g=48 . We found that container creation dominates invocation latency, hence higher g performs best. This proves that creating the biggest possible containers, and thus the less amount of them (heterogeneous packing), achieves the best start-up latency, since it creates a single container per invoker per flare. By extension, the mixed packing strategy exhibits the same results, but allows the system to manage resources more effectively in small portions to facilitate allocation and avoid resource fragmentation. To assess the impact of granularity, the rest of the evaluation uses homogeneous packing. Impact on worker simultaneity We run a burst with size 960 on FaaS against burst computing with g=48 . For demonstration purposes, each worker performs a 5-second sleep and we plot their execution timeline in Fig. 7. The plot shows that burst computing achieves faster resource allocation and quicker readiness of workers. This ensures worker parallelism. Analyzing dispersity of worker start-up time (also in Fig. 6), the FaaS execution evinces a range of 18.8 s between the start of the first worker and that of the last one, with a median absolute deviation (MAD) of 2.65 s . In contrast, the range USENIX Association 2025 USENIX Annual Technical Conference 45 FaaS 2 4 6 12 24 48 100GB 50GB 0 5s 10s 15s 20s Download size Download time Granularity Figure 8: A burst of 96 workers loading the same 1 GiB object from S3 with different granularity. with g=48 is just 0.44 s (MAD is 0.1 s ). Compared, the range is 43 × lower in burst computing, with MAD showing 26.5 × lower dispersity than FaaS. Dispersity in worker start-up latency precludes FaaS to achieve full parallelism (all workers running simultaneously from start to finish), while burst guarantees it. Impact on data loading Burst computing mitigates the FaaS problem of loading the same data on all functions (§2.1), e.g., in a grid search. We can leverage worker access to locality information to optimize this problem and download the data only once per pack, trivially reducing data ingestion. Specifically, each worker in a pack retrieves a part of the data based on calculations from pack information in the burst context (Table 2). Then they recreate the full data in a local shared memory region. This allows to parallelize the download and complete the process faster than choosing a leader to perform it (also possible through the isPackLeader function in the context, which returns true for the worker with lowest ID within its pack). We evaluate this approach on multiple g and present it in Fig. 8. Burst optimizations achieve a download time speed-up of 32.6×with g=48 compared to FaaS. Takeaway Flares eliminate friction F1 through faster worker group initialization (11.5 × ) and ensured simultaneity (43 × less dispersed workers) that enables locality with packing. In turn, locality may accelerate data download in applications (32.6×), tackling friction F3. 5.2 Burst inter-pack communication Before we evaluate the effects of the BCM on frictions F2 and F3, we want to ensure that an indirect communication model is feasible and to find a backend that sustains the load of bursts at scale. For this, we measure the throughput of several indirect communication backends. Specifically, we test Redis, DragonflyDB (a Redis-compatible multi-threaded alternative), RabbitMQ, and S3. Redis and DragonflyDB evaluate two flavors: using lists or streams. Message chunk size The BCM chunks messages into several blocks to optimize network utilization and allow parallel read/write. The optimal chunk size is a trade-off between latency to first byte and operation overhead, and it varies 64 KiB 1 MiB 64 MiB 128 MiB 256 MiB Chunk Size 0 100 200 300 400 Throughput (MiB/s) RabbitMQ Redis List DragonflyDB List Redis Stream DragonflyDB Stream S3 1 (a) Throughput between two remote workers sending a 1 GiB payload chunked in different sizes. 8 16 32 96 192 384 Burst Size 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Aggregated Throughput (GiB/s) 1 (b) Aggregate throughput of two remote packs, A and B, of varying size ( g= burst size /2 ), where each worker from pack A sends a 256 MiB payload to another worker from remote pack B. Figure 9: Throughput experiments for the different BCM backends. Median values with standard deviation (10 runs). for each communication backend. To find the optimal configuration, we measure the throughput of sending a 1 GiB message between two remote workers. The workers run on two c7i.large machines (4 vCPUs, 8 GB) and we deploy a c7i.16xlarge (64 vCPUs, 128 GB) for the intermediate server. Fig. 9a plots the results. RabbitMQ offers a constant throughput for larger chunk sizes, but does not allow payloads larger than 128 MiB due to AMQP protocol limitations. Redis and DragonflyDB work best at 1 MiB , the latter being slightly superior. S3 offers the lowest throughput because object stores are not designed for small files ( 1 MiB or less exceeds the allowed service request rate limits). Maximum throughput To understand how the different backends scale under parallel load, we measure the aggregated throughput between several pairs of workers communicating simultaneously. In this experiment, we launch a group of workers (burst size from 8 to 384) split into two remote groups. Each worker in a group A sends a fixed message ( 256 MiB ) to a worker in the other, remote group B. As the burst size increases, so does the total data volume sent. Each backend uses the optimal chunk size assessed in the micro-benchmark above. Workers run on two VMs scaled to the burst size (from c7i.xlarge for 8 workers to 46 2025 USENIX Annual Technical Conference USENIX Association and Keith Winstein. 2017. Encoding, Fast and Slow: Low-Latency Video Processing Using Thousands of Tiny Threads. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). USENIX Association, Boston, MA, 363–376. https://www.usenix.org/conference/nsdi17/ technical-sessions/presentation/fouladi [17] Pedro Garcia Lopez, Aleksander Slominski, Bernard Metzler, Michael Berhendt, and Simon Shillaker. 2024. Serverless End Game: Disaggregation enabling Transparency. In Proceedings of the 2nd Workshop on SErverless Systems, Applications and MEthodologies (Athens, Greece) (SESAME ’24). Association for Computing Machinery, New York, NY, USA, 9–14. https://doi.org/10. 1145/3642977.3652094 [18] Joseph M. Hellerstein, Jose Faleiro, Joseph E. Gonzalez, Johann Schleier-Smith, Vikram Sreekanti, Alexey Tumanov, and Chenggang Wu. 2018. Serverless Computing: One Step Forward, Two Steps Back. https: //doi.org/10.48550/ARXIV.1812.03651 [19] Bo Huang, Shengsheng Huang, Jinquan Dai, Jie Huang, and Tao Xie. 2010. The HiBench benchmark suite: Characterization of the MapReduce-based data analysis. In 2010 IEEE 26th International Conference on Data Engineering Workshops (ICDEW 2010). IEEE Computer Society, Los Alamitos, CA, USA, 41–51. https: //doi.org/10.1109/ICDEW.2010.5452747 [20] Aman Jain, Ata F. Baarzi, George Kesidis, Bhuvan Urgaonkar, Nader Alfares, and Mahmut Kandemir. 2020. SplitServe: Efficiently Splitting Apache Spark Jobs Across FaaS and IaaS. In Proceedings of the 21st International Middleware Conference (Delft, Netherlands) (Middleware ’20). Association for Computing Machinery, New York, NY, USA, 236–250. https://doi.org/10. 1145/3423211.3425695 [21] Zhipeng Jia and Emmett Witchel. 2021. Nightcore: efficient and scalable serverless computing for latencysensitive, interactive microservices. In Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’21). Association for Computing Machinery, New York, NY, USA, 152–166. https: //doi.org/10.1145/3445814.3446701 [22] Chao Jin, Zili Zhang, Xingyu Xiang, Songyun Zou, Gang Huang, Xuanzhe Liu, and Xin Jin. 2023. Ditto: Efficient Serverless Analytics with Elastic Parallelism. In Proceedings of the ACM SIGCOMM 2023 Conference (New York, NY, USA) (ACM SIGCOMM ’23). Association for Computing Machinery, New York, NY, USA, 406–419. https://doi.org/10.1145/3603269.3604816 [23] Eric Jonas, Qifan Pu, Shivaram Venkataraman, Ion Stoica, and Benjamin Recht. 2017. Occupy the cloud: distributed computing for the 99%. In Proceedings of the 2017 Symposium on Cloud Computing (SoCC ’17). Association for Computing Machinery, New York, NY, USA, 445–451. https://doi.org/10.1145/3127479. 3128601 [24] Eric Jonas, Johann Schleier-Smith, Vikram Sreekanti, Chia-Che Tsai, Anurag Khandelwal, Qifan Pu, Vaishaal Shankar, Joao Menezes Carreira, Karl Krauth, Neeraja Yadwadkar, Joseph Gonzalez, Raluca Ada Popa, Ion Stoica, and David A. Patterson. 2019. Cloud Programming Simplified: A Berkeley View on Serverless Computing. Technical Report UCB/EECS-20193. EECS Department, University of California, Berkeley. https://www2.eecs.berkeley.edu/Pubs/TechRpts/ 2019/EECS-2019-3.pdf [25] Kostis Kaffes, Neeraja J. Yadwadkar, and Christos Kozyrakis. 2019. Centralized Core-granular Scheduling for Serverless Functions. In Proceedings of the ACM Symposium on Cloud Computing (SoCC ’19). Association for Computing Machinery, New York, NY, USA, 158–164. https://doi.org/10.1145/3357223.3362709 [26] Nima Kaviani, Dmitriy Kalinin, and Michael Maximilien. 2019. Towards Serverless as Commodity: a case of Knative. In Proceedings of the 5th International Workshop on Serverless Computing (Davis, CA, USA) (WOSC ’19). Association for Computing Machinery, New York, NY, USA, 13–18. https://doi.org/10.1145/ 3366623.3368135 [27] Jeongchul Kim and Kyungyong Lee. 2019. FunctionBench: A Suite of Workloads for Serverless Cloud Function Service. In 2019 IEEE 12th International Conference on Cloud Computing (CLOUD). IEEE Computer Society, Los Alamitos, CA, USA, 502–504. https: //doi.org/10.1109/CLOUD.2019.00091 [28] Ana Klimovic, Yawen Wang, Patrick Stuedi, Animesh Trivedi, Jonas Pfefferle, and Christos Kozyrakis. 2018. Pocket: Elastic Ephemeral Storage for Serverless Analytics. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, Carlsbad, CA, 427–444. https://www. usenix.org/conference/osdi18/presentation/klimovic [29] Swaroop Kotni, Ajay Nayak, Vinod Ganapathy, and Arkaprava Basu. 2021. Faastlane: Accelerating Function-as-a-Service Workflows. In 2021 USENIX Annual Technical Conference (USENIX ATC 21). USENIX Association, Boston, MA, 805–820. https://www. usenix.org/conference/atc21/presentation/kotni [30] Collin Lee and John Ousterhout. 2019. Granular Computing. In Proceedings of the Workshop on Hot Topics in Operating Systems (Bertinoro, Italy) (HotOS ’19). Association for Computing Machinery, New York, NY, USA, 149–154. https://doi.org/10.1145/3317550.3321447 [31] Yilong Li, Seo Jin Park, and John Ousterhout. 2021. MilliSort and MilliQuery: Large-Scale Data-Intensive Computing in Milliseconds. In 18th USENIX Symposium on Networked Systems Design and ImplementaUSENIX Association 2025 USENIX Annual Technical Conference 53 tion (NSDI 21). USENIX Association, Boston, MA, 593–611. https://www.usenix.org/conference/nsdi21/ presentation/li-yilong [32] Yiming Li, Laiping Zhao, Yanan Yang, and Wenyu Qu. 2023. Rethinking Deployment for Serverless Functions: A Performance-First Perspective. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Denver, CO, USA) (SC ’23). Association for Computing Machinery, New York, NY, USA, Article 67, 14 pages. https://doi.org/10.1145/3581784.3613211 [33] Bryan Liston. 2016. Ad Hoc Big Data Processing Made Simple with Serverless MapReduce. Retrieved May 02, 2024 from https: //aws.amazon.com/blogs/compute/ad-hoc-big-dataprocessing-made-simple-with-serverless-mapreduce/ [34] David H. Liu, Amit Levy, Shadi Noghabi, and Sebastian Burckhardt. 2023. Doing More with Less: Orchestrating Serverless Applications without an Orchestrator. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). USENIX Association, Boston, MA, 1505–1519. https://www.usenix.org/ conference/nsdi23/presentation/liu-david [35] Fangming Lu, Xingda Wei, Zhuobin Huang, Rong Chen, Minyu Wu, and Haibo Chen. 2024. Serialization/Deserialization-free State Transfer in Serverless Workflows. In Proceedings of the Nineteenth European Conference on Computer Systems (Athens, Greece) (EuroSys ’24). Association for Computing Machinery, New York, NY, USA, 132–147. https://doi.org/10.1145/3627703.3629568 [36] Ashraf Mahgoub, Edgardo Barsallo Yi, Karthick Shankar, Sameh Elnikety, Somali Chaterji, and Saurabh Bagchi. 2022. ORION and the Three Rights: Sizing, Bundling, and Prewarming for Serverless DAGs. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). USENIX Association, Carlsbad, CA, 303–320. https://www.usenix.org/conference/ osdi22/presentation/mahgoub [37] Shruti Mohanty, Vivek M. Bhasi, Myungjun Son, Mahmut Taylan Kandemir, and Chita Das. 2024. FAAStloop: Optimizing Loop-Based Applications for Serverless Computing. In Proceedings of the 2024 ACM Symposium on Cloud Computing (Redmond, WA, USA) (SoCC ’24). Association for Computing Machinery, New York, NY, USA, 943–960. https://doi.org/10.1145/3698038. 3698560 [38] Ingo Müller, Rodrigo Bruno, Ana Klimovic, John Wilkes, Eric Sedlar, and Gustavo Alonso. 2020. Serverless Clusters: The Missing Piece for Interactive Batch Applications? (April 2020), 3 pages. https://doi.org/10. 3929/ethz-b-000405616 Presented at 10th Workshop on Systems for Post-Moore Architectures (SPMA 2020). [39] Gerard París, Pedro García-López, and Marc SánchezArtigas. 2020. Serverless Elastic Exploration of Unbalanced Algorithms. In 2020 IEEE 13th International Conference on Cloud Computing (CLOUD). IEEE Computer Society, Los Alamitos, CA, USA, 149–157. https: //doi.org/10.1109/CLOUD49709.2020.00033 ISSN: 2159-6190. [40] Qifan Pu, Shivaram Venkataraman, and Ion Stoica. 2019. Shuffling, Fast and Slow: Scalable Analytics on Serverless Infrastructure. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). USENIX Association, Boston, MA, 193–206. https: //www.usenix.org/conference/nsdi19/presentation/pu [41] Wei Qiu, Marcin Copik, Yun Wang, Alexandru Calotoiu, and Torsten Hoefler. 2023. User-guided Page Merging for Memory Deduplication in Serverless Systems. In 2023 IEEE International Conference on Big Data (BigData). IEEE Computer Society, Los Alamitos, CA, USA, 159–169. https://doi.org/10.1109/BigData59044.2023. 10386487 [42] Zhenyuan Ruan, Shihang Li, Kaiyan Fan, Seo Jin Park, Marcos K. Aguilera, Adam Belay, and Malte Schwarzkopf. 2025. Quicksand: Harnessing Stranded Datacenter Resources with Granular Computing. In 22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). USENIX Association, Philadelphia, PA, 147–165. https://www.usenix. org/conference/nsdi25/presentation/ruan [43] Josep Sampe, Marc Sanchez-Artigas, Gil Vernik, Ido Yehekzel, and Pedro Garcia-Lopez. 2023. Outsourcing Data Processing Jobs With Lithops. IEEE Transactions on Cloud Computing 11, 01 (Jan. 2023), 1026–1037. https://doi.org/10.1109/TCC.2021.3129000 [44] Josep Sampé, Gil Vernik, Marc Sánchez-Artigas, and Pedro García-López. 2018. Serverless Data Analytics in the IBM Cloud. In Proceedings of the 19th International Middleware Conference Industry (Rennes, France) (Middleware ’18). Association for Computing Machinery, New York, NY, USA, 1–8. https://doi.org/ 10.1145/3284028.3284029 [45] Josep Sampé, Marc Sánchez-Artigas, Pedro GarcíaLópez, and Gerard París. 2017. Data-driven serverless functions for object storage. In Proceedings of the 18th ACM/IFIP/USENIX Middleware Conference (Middleware ’17). Association for Computing Machinery, New York, NY, USA, 121–133. https://doi.org/10.1145/ 3135974.3135980 [46] Marc Sánchez-Artigas, Germán T. Eizaguirre, Gil Vernik, Lachlan Stuart, and Pedro García-López. 2020. Primula: a Practical Shuffle/Sort Operator for Serverless Computing. In Proceedings of the 21st International Middleware Conference Industrial Track (Delft, Netherlands) (Middleware ’20). Association for Computing Machinery, New York, NY, USA, 31–37. https: //doi.org/10.1145/3429357.3430522 54 2025 USENIX Annual Technical Conference USENIX Association [47] Trever Schirmer, Joel Scheuner, Tobias Pfandzelter, and David Bermbach. 2022. Fusionize: Improving Serverless Application Performance through Feedback-Driven Function Fusion. In 2022 IEEE International Conference on Cloud Engineering (IC2E). IEEE Computer Society, Los Alamitos, CA, USA, 85–95. https://doi. org/10.1109/IC2E55432.2022.00017 [48] Carlos Segarra, Simon Shillaker, Guo Li, Eleftheria Mappoura, Rodrigo Bruno, Lluís Vilanova, and Peter Pietzuch. 2025. GRANNY: Granular Management of Compute-Intensive Applications in the Cloud. In 22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). USENIX Association, Philadelphia, PA, 205–218. https://www.usenix. org/conference/nsdi25/presentation/segarra [49] Vaishaal Shankar, Karl Krauth, Kailas Vodrahalli, Qifan Pu, Benjamin Recht, Ion Stoica, Jonathan RaganKelley, Eric Jonas, and Shivaram Venkataraman. 2020. Serverless linear algebra. In Proceedings of the 11th ACM Symposium on Cloud Computing (Virtual Event, USA) (SoCC ’20). Association for Computing Machinery, New York, NY, USA, 281–295. https://doi.org/10. 1145/3419111.3421287 [50] Simon Shillaker and Peter Pietzuch. 2020. Faasm: Lightweight Isolation for Efficient Stateful Serverless Computing. In 2020 USENIX Annual Technical Conference (USENIX ATC 20). USENIX Association, Boston, MA, 419–433. https://www.usenix.org/conference/ atc20/presentation/shillaker [51] Won Wook Song, Taegeon Um, Sameh Elnikety, Myeongjae Jeon, and Byung-Gon Chun. 2023. Sponge: Fast Reactive Scaling for Stream Processing with Serverless Frameworks. In 2023 USENIX Annual Technical Conference (USENIX ATC 23). USENIX Association, Boston, MA, 301–314. https://www.usenix.org/ conference/atc23/presentation/song [52] Vikram Sreekanti, Chenggang Wu, Xiayue Charles Lin, Johann Schleier-Smith, Joseph E. Gonzalez, Joseph M. Hellerstein, and Alexey Tumanov. 2020. Cloudburst: stateful functions-as-a-service. Proceedings of the VLDB Endowment 13, 12 (July 2020), 2438–2452. https://doi.org/10.14778/3407790.3407836 [53] Jovan Stojkovic, Tianyin Xu, Hubertus Franke, and Josep Torrellas. 2023. MXFaaS: Resource Sharing in Serverless Environments for Parallelism and Efficiency. In Proceedings of the 50th Annual International Symposium on Computer Architecture (Orlando, FL, USA) (ISCA ’23). Association for Computing Machinery, New York, NY, USA, Article 34, 15 pages. https://doi.org/10.1145/3579371.3589069 [54] Ali Tariq, Austin Pahl, Sharat Nimmagadda, Eric Rozner, and Siddharth Lanka. 2020. Sequoia: enabling qualityof-service in serverless computing. In Proceedings of the 11th ACM Symposium on Cloud Computing (SoCC ’20). Association for Computing Machinery, New York, NY, USA, 311–327. https://doi.org/10.1145/3419111. 3421306 [55] Shelby Thomas, Lixiang Ao, Geoffrey M. Voelker, and George Porter. 2020. Particle: ephemeral endpoints for serverless networking. In Proceedings of the 11th ACM Symposium on Cloud Computing (SoCC ’20). Association for Computing Machinery, New York, NY, USA, 16–29. https://doi.org/10.1145/3419111.3421275 [56] Bishoy Wadie, Lachlan Stuart, Christopher M. Rath, Bernhard Drotleff, Sergii Mamedov, and Theodore Alexandrov. 2024. METASPACE-ML: Context-specific metabolite annotation for imaging mass spectrometry using machine learning. Nature Communications 15, 9110 (2024), 16 pages. https://doi.org/10.1038/s41467024-52213-9 [57] Michael Wawrzoniak, Gianluca Moro, Rodrigo Bruno, Ana Klimovic, and Gustavo Alonso. 2024. Off-the-shelf Data Analytics on Serverless. In Proceedings of the 14th Conference on Innovative Data Systems Research, CIDR 2024. CIDR, Chaminade, USA, 10 pages. https: //doi.org/10.3929/ethz-b-000646171 [58] Michal Wawrzoniak, Ingo Müller, Gustavo Alonso, and Rodrigo Bruno. 2021. Boxer: Data Analytics on Network-enabled Serverless Platforms. In 11th Conference on Innovative Data Systems Research, CIDR 2021. CIDR, Chaminade, USA, 8 pages. https://doi. org/10.3929/ethz-b-000456492 [59] Xingda Wei, Fangming Lu, Tianxia Wang, Jinyu Gu, Yuhan Yang, Rong Chen, and Haibo Chen. 2023. No Provisioned Concurrency: Fast RDMA-codesigned Remote Fork for Serverless Computing. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). USENIX Association, Boston, MA, 497–517. https://www.usenix.org/conference/osdi23/ presentation/wei-rdma [60] Sebastian Werner and Stefan Tai. 2024. A reference architecture for serverless big data processing. Future Generation Computer Systems 155 (2024), 179–192. https://doi.org/10.1016/j.future.2024.01.029 [61] Zhaorui Wu, Yuhui Deng, Yi Zhou, Jie Li, Shujie Pang, and Xiao Qin. 2024. FaaSBatch: Boosting Serverless Efficiency With In-Container Parallelism and Resource Multiplexing . IEEE Trans. Comput. 73, 04 (April 2024), 1071–1085. https://doi.org/10.1109/TC.2024.3352834 [62] Minchen Yu, Tingjia Cao, Wei Wang, and Ruichuan Chen. 2023. Following the Data, Not the Function: Rethinking Function Orchestration in Serverless Computing. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). USENIX Association, Boston, MA, 1489–1504. https://www.usenix. org/conference/nsdi23/presentation/yu [63] Hong Zhang, Yupeng Tang, Anurag Khandelwal, Jingrong Chen, and Ion Stoica. 2021. Caerus: NIMUSENIX Association 2025 USENIX Annual Technical Conference 55 BLE Task Scheduling for Serverless Analytics. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21). USENIX Association, Boston, MA, 653–669. https://www.usenix.org/ conference/nsdi21/presentation/zhang-hong 56 2025 USENIX Annual Technical Conference USENIX Association