scieee AI-readable full text Open interactive document viewer

Run-Time Adaptation of Complex Event Forecasting

Alevizos, Elias; Giatrakos, Nikos; Artikis, Alexander

Abstract

Complex Event Forecasting (CEF) is a process whereby complex events of interest are forecast over a stream of simple events. CEF facilitates proactive measures by anticipating the occurrence of complex events. This proactive property, makes CEF a crucial task in many domains; for instance, in maritime situational awareness, forecasting the arrival of vessels at ports allows for better resource management, and higher operational efficiency. However, our world’s dynamic and evolving conditions necessitate the use of adaptive methods. For example, for safety reasons, maritime vessels may adapt their routes to avoid powerful swell waves; in fraud analytics, fraudsters evolve their tactics to avoid detection etc. CEF systems typically rely on probabilistic models, trained on historical data. This renders such CEF systems inherently susceptible to data evolutions that can invalidate their underlying models. To address this problem, we propose RTCEF, a novel framework for Run-Time Adaptation of CEF, based on a distributed, service-oriented architecture. We evaluate RTCEF on two use-cases and our reproducible results show that our proposed approach has significant benefits in terms of forecasting performance without sacrificing efficiency.

Full text

Run-Time Adaptation of Complex Event Forecasting Manolis Pitsikalis NCSR Demokritos Athens, Greece [email protected] Elias Alevizos NCSR Demokritos Athens, Greece The American College of Greece Athens, Greece [email protected] Nikos Giatrakos Technical University of Crete Chania, Greece [email protected] Alexander Artikis NCSR Demokritos Athens, Greece University of Piraeus Piraeus, Greece [email protected] Abstract Complex Event Forecasting (CEF) is a process whereby complex events of interest are forecast over a stream of simple events. CEF facilitates proactive measures by anticipating the occurrence of complex events. This proactive property, makes CEF a crucial task in many domains; for instance, in maritime situational awareness, forecasting the arrival of vessels at ports allows for better resource management, and higher operational efficiency. However, our world’s dynamic and evolving conditions necessitate the use of adaptive methods. For example, for safety reasons, maritime vessels may adapt their routes to avoid powerful swell waves; in fraud analytics, fraudsters evolve their tactics to avoid detection etc. CEF systems typically rely on probabilistic models, trained on historical data. This renders such CEF systems inherently susceptible to data evolutions that can invalidate their underlying models. To address this problem, we propose RTCEF , a novel framework for Run-Time Adaptation of CEF, based on a distributed, service-oriented architecture. We evaluate RTCEF on two use-cases and our reproducible results show that our proposed approach has significant benefits in terms of forecasting performance without sacrificing efficiency. CCS Concepts •Computer systems organization → Real-time systems; •Theory of computation → Formal languages and automata theory;•Mathematics of computing → Bayesian computation. This work is licensed under a Creative Commons Attribution 4.0 International License. DEBS ’25, Gothenburg, Sweden ©2025 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-1332-3/25/06 https://doi.org/10.1145/3701717.3730539 Keywords complex event forecasting, run-time adaptation, optimisation ACM Reference Format: Manolis Pitsikalis, Elias Alevizos, Nikos Giatrakos, and Alexander Artikis. 2025. Run-Time Adaptation of Complex Event Forecasting. In The 19th ACM International Conference on Distributed and Eventbased Systems (DEBS ’25), June 10–13, 2025, Gothenburg, Sweden. ACM, New York, NY, USA, 12 pages. https://doi.org/10.1145/3701717. 3730539 1 Introduction Complex Event Forecasting (CEF) is akin to Complex Event Recognition (CER) [ 17 , 22 ], but with a forward-looking perspective. Both tasks operate on a stream of simple events, while their output consists of Complex Events (CEs). For example, in a maritime situational awareness [ 4 ], the stream of simple events would contain positional messages of vessels, while the output stream would contain maritime CEs such as (illegal) fishing activities. The difference between CER and CEF is that, in the former, elements of the output stream refer to CE detections, while in CEF, elements of the output stream refer to the probability of a CE happening in the future. Consequently, CER enables reactive responses upon CE detections, while CEF supports proactive measures by anticipating future CEs. This proactive property renders CEF systems highly desirable. CER and CEF applications span diverse domains, such as maritime situational awareness [ 2 , 28 ] whereby CEs such as fishing are detected or forecast over a stream of maritime data; credit card fraud management [ 2 , 32 ] whereby frauds are detected or forecast over a stream of transaction data; and so on. CEF operates over constantly evolving conditions. Take for example the problem of maritime route optimisation. Vessels may follow a different route depending on the swell wave conditions, i.e., waves that can significantly affect navigation and vessel stability [ 11 ]. Another example is financial fraud 9 DEBS ’25, June 10–13, 2025, Gothenburg, Sweden Pitsikalis et al. detection—fraudsters constantly adapt their tactics to avoid getting caught [ 3 ]. Moreover, CEF systems rely on probabilistic models trained on historical data [ 1 , 25 , 26 ]. This renders CEF systems inherently susceptible to evolutions in the input that can invalidate their underlying models—recall the previous example relating maritime routes with weather. Additionally, as with the majority of trainable models, CEF models have hyperparameters that require fine tuning for optimal performance. Wayeb [ 2 ], a state-of-the-art CEF engine, is no exception to the above. To address the above challenges we propose RTCEF , an open-source framework for Run-Time Adaptation of CEF over constantly evolving data streams. RTCEF adopts a distributed architecture comprising targeted services to effectively (a) enable run-time update of CEF models with little to no downtime, (b) ensure that transition between models does not cause loss of forecasts. In other words, RTCEF supports continuous adaptation to dynamic changes in the input stream with little to no effect on efficiency. Furthermore, RTCEF provides a trend-based policy which acts as a decision making mechanism to distinguish whether hyperparameter optimisation or CEF model retraining, without changing hyperparameters, is the best way to maintain accurate forecasts. In addition to RTCEF , we also present offCEF , a baseline framework for CEF hyperparameter optimisation, that in contrast to RTCEF , optimises CEF in an offline manner. Our contributions are: • we introduce RTCEF , an open-source 1 framework addressing the challenges and requirements of run-time CEF over evolving data streams through a distributed architecture which enables CEF to seamlessly run on par with training or optimisation tasks, ensuring no disruptions; • we formally prove that RTCEF achieves lossless runtime adaptation, i.e., no forecast is lost upon hyperparameter optimisation or retraining decisions; • we extensively evaluate RTCEF and offCEF on two realworld critical use-cases from the maritime and financial domains and our reproducible results validate that RTCEF , compared to offCEF , can significantly improve forecasting performance with little to no lag upon run-time changes. 2 Background CEF is a task that allows forecasting CEs of interest, such as fishing activities or vessel rendezvous, over an input stream of simple events; e.g., timestamped position messages of maritime vessels. Forecasts involve the occurrence of a CE in the future accompanied by a degree of certainty [ 2 ]. This behaviour is usually derived from stochastic models that project into the future evolutions of the input that can cause a detection of a CE. For the task of CEF, we utilise Wayeb, 1https://zenodo.org/records/15229227 Table 1: An example stream for a single vessel composed of five events. Each event has a vessel identifier, a value for that vessel’s speed and a timestamp. vessel ID 78986 78986 78986 78986 78986 ... speed 5 3 9 14 11 ... timestamp 1 2 3 4 5 ... a CEF engine introduced in [ 2 ], which employs symbolic automata as its computational model. The user submits a query/pattern to Wayeb which is then compiled into a symbolic, streaming automaton. This automaton may be used to perform event recognition, i.e., to detect instances of pattern satisfaction upon a stream of input events. Whenever the automaton reaches a final state, a complex event is reported as having occurred. In order to perform forecasting, Wayeb constructs a probabilistic model of the compiled automaton, by using part(s) of a stream for training. The model allows us to infer, at any given moment, the possible paths that the automaton may follow in the future. By searching among the possible future paths, we can estimate when the automaton is expected to reach a final state and thus report a CE. The output of Wayeb thus consists of two streams: a) one reporting the detected events, and b) one reporting the forecasts of events expected to occur in the future. Wayeb has clear, compositional semantics for the patterns expressed in its language and can support most of the common operators [ 17 ]. Wayeb’s patterns are expressed as Symbolic Regular Expressions ( SRE s), where terminal expressions are Boolean expressions, i.e., logical formulae that use the standard Boolean connectives of conjunction ‘ ∧ ’, disjunction ‘ ∨ ’ and negation ‘ ¬ ’ on predicates [ 2 ]. Wayeb SRE s are defined using the grammar below: 𝑅::=𝑅1+𝑅2(union) | 𝑅1·𝑅2(concatenation) |𝑅∗ 1(Kleene-star)| !𝑅1(complement) |𝜓(Boolean expression) 𝑅1, 𝑅2 are regular expressions, and 𝜓 is a Boolean expression. The semantics of the above operators are detailed in [ 2 ]. Evaluation of SRE s on a stream of events requires first their compilation into symbolic automata. Transitions in symbolic automata are labeled with Boolean expressions. For a symbolic automaton to move to another state, it first applies the Boolean expressions of its current state’s outgoing transitions to the element last read from the stream. If an expression is satisfied, then the corresponding transition is triggered and the automaton moves to that transition’s target state. For example, in maritime situational awareness, a domain expert could use Wayeb’s language to specify a pattern 𝑅 : =(𝑠𝑝𝑒𝑒𝑑 > 10 )·(𝑠𝑝𝑒𝑒𝑑 > 10 ) for identifying 10 Run-Time Adaptation of Complex Event Forecasting DEBS ’25, June 10–13, 2025, Gothenburg, Sweden 0 start 1 2 ¬(speed >10) speed >10 ¬(speed >10) speed >10 Figure 1: Streaming symbolic automaton created from the expression 𝑅:=(𝑠𝑝𝑒𝑒𝑑 >10)·(𝑠𝑝𝑒𝑒𝑑 >10). speed violations in specific areas where the maximum allowed speed is 10 𝑘𝑛𝑜𝑡𝑠 . This pattern is satisfied when there are two consecutive events where a vessel’s speed exceeds the threshold. The compiled automaton corresponding to 𝑅 is illustrated in Figure 1. For an input stream consisting of the events in Table 1, the automaton would run as follows. For the first three input events, the automaton remains in state 0. After the fourth event, it moves to state 1and after the fifth event it reaches its final state, state 2, triggering also a CE detection for 𝑅at timestamp =5. To perform CEF, Wayeb needs a probabilistic description for a symbolic automaton derived from a SRE . For this purpose, Wayeb employs Prediction Suffix Trees (PSTs) [ 30 , 31 ]– a form of Variable-order Markov Models. Variable-order Markov Models, compared to fixed-order Markov models, capture longer-term dependencies as in practice they allow for higher order ( 𝑚 ) values than the latter. Each node in a PST contains a “context” and a distribution that indicates the probability of encountering a symbol, conditioned on the context. Figure 4 (top left) shows an example of a PST. Each “symbol” of a PST corresponds to a predicate of the automaton for which we want to build a probabilistic model. For example, the predicate (𝑠𝑝𝑒𝑒𝑑> 10 ) may be such a “symbol” for the pattern 𝑅 . The same predicate, but negated i.e., ¬(𝑠𝑝𝑒𝑒𝑑> 10 ) , may be another such “symbol”. Learning a PST from data is an incremental process that adds new nodes corresponding to symbols only when necessary [ 2 , 30 , 32 ]. The learning process involves two key hyperparameters. First, the pMin ∈ [ 0 , 1 ] hyper-parameter which corresponds to a threshold determining which symbols are deemed to be “too rare” to be taken under consideration by the learning algorithm (symbols with a probability of appearance less than pMin are discarded). Second, the 𝛾 hyperparameter is a symbol distribution smoothing parameter. With the resulting PST, for every state 𝑞 of an automaton and the last 𝑚 (order of the PST) symbols of the input stream, we can calculate the waiting-time distribution ( 𝑊𝑞 ), that is, the probability of reaching a final state in 𝑛 transitions from a state 𝑞 . Recall that a CE is detected whenever an automaton reaches a final state. Figure 4 (middle and bottom left) shows an example of an automaton and the waiting-time distributions learnt from a training dataset. Wayeb then performs CEF as follows. Given the current state 𝑞 of an automaton, using 𝑊𝑞 , we compute the probability of reaching a final state ( 𝑝𝐶𝐸 ) within the next 𝑛 transitions (or, equivalently, input events). If 𝑝𝐶𝐸 exceeds a confidence threshold 𝜃fc ∈ [ 0 , 1 ] , Wayeb emits a “positive” forecast (denoting that the CE is expected to occur), otherwise a “negative forecast” (no CE is expected) is emitted. A forecast for a CE is characterised as a True Positive ( TP ) if a positive forecast (i.e., the CE will occur in the future) was emitted and the CE indeed occurred or, respectively, as a False Positive ( FP ) if the CE did not occur. A forecast for a CE is characterised as a True Negative ( TN ) if a negative forecast is emitted (i.e., the CE will not occur in the future) and the CE does not occur or, respectively, as a False Negative ( FN ) if the CE does occur. Note that a forecast cannot be evaluated as TP , FP , TN or FN upon its emission. It can be evaluated as such after the next 𝑛 input events have arrived, at which point we can know whether the forecast event did occur or not. Given that Wayeb performs both CEF and CER, forecasts are evaluated on-the-fly. Using these classifications of forecasts, the performance of CEF may be quantified through Matthew’s Correlation Coefficient ( MCC ), defined as follows: 𝑀𝐶𝐶 =√︁Precision ×Recall ×Specificity ×NPV −√FDR ×FNR ×FPR ×FOMR (1) where NPV =TN TN+FN , Specificity =TN TN+FP , FDR = 1 − Precision , FNR = 1 −Recall , FPR = 1 −Specificity and FOMR = 1 −NPV .Precision and Recall are defined as usual. Therefore, MCC ∈ [− 1 , 1 ] estimates the agreement, in which case MCC = 1, (or disagreement, resp. MCC =− 1) between the emitted forecasts and observations. In contrast to F1Score, which takes into account only positive instances, MCC takes into account both positive and negative instances. Since Wayeb produces both positive and negative forecasts, MCC is a fitting choice. Given the above, the hyperparameters required for training Wayeb models, i.e., PSTs, are the following. The maximum order 𝑚 of the PST, along with the symbol retaining probability threshold pMin , the symbol distribution smoothing parameter 𝛾 and the confidence threshold 𝜃fc . The naive way to train a Wayeb PST is to manually fix the values of these hyperparameters and then select a training dataset from which a PST may be extracted. This process can be performed offline and Wayeb may then employ the learnt PST for online event forecasting. As we explain below, this is not the proper way to go. 3 Challenges of CEF Performing CEF over constantly evolving data streams exhibits several significant challenges. Challenge 1. CEF hyperparameter optimisation entails complicated trade-offs. 11 DEBS ’25, June 10–13, 2025, Gothenburg, Sweden Pitsikalis et al. 5 10 15 20 25 0 0.2 0.4 0.6 0.8 1 Weeks (Maritime) MCC Wayeb 20 40 60 80 0 0.2 0.4 0.6 0.8 1 Weeks (Finance) Wayeb Figure 2: MCC scores of Wayeb for forecasting a CE related to the arrival of vessels at a port over maritime positional data (left), and forecasting the occurrence of financial frauds over transactional data (right). In Wayeb, although setting the maximum order mgenerally improves accuracy, it leads to longer training times. Similarly, finding the optimal value for 𝜃fc , i.e., the probability threshold for emitting a forecast, is a crucial step as overly low or high 𝜃fc values can cause many false positives or false negatives, respectively. Furthermore, pMin , the threshold determining which symbols are “too rare” to be included during training, can also affect accuracy. High pMin values can produce simpler models but may discard useful symbols. On the other hand, excessively low values of pMin can degrade the accuracy of the forecasts due to overfitting of the PSTs to insignificant symbols. The symbol distribution smoothing parameter 𝛾 behaves in the same manner. Manually fixing the above combination of parameters would lead to sub-optimal results. Moreover, exhaustive hyperparameter space exploration is of high computational complexity making it prohibitive for run-time settings where Wayeb’s hyperparameters need to be tuned multiple times to adjust to data evolutions that invalidate the deployed PSTs. To address these issues,we propose RTCEF , a framework for the run-time adaptation of CEF. RTCEF , presented in the following section, employs Bayesian optimisation to efficiently explore only a small fraction of the parameter space and learn the optimal combination of hyperparameter values for the entire parameter space. Furthermore, in contrast to traditional Bayesian optimisation setups, we do not start each optimisation run from scratch, instead we leverage knowledge from previous runs by constantly refreshing a sample set with new samples. Challenge 2. Run-time CEF optimisation has inherently increased complexity, while CEF applications typically involve processing Big streaming Data with volatile statistical properties that can severely affect CEF performance. Run-time optimisation of Wayeb is a complicated and challenging task because it involves (a) training a Variableorder Markov Model, that is, a PST (b) using it for estimating waiting-time distributions and (c) subsequently performing probabilistic, automaton-based pattern matching (Figure 4). Consequently, employing an analytical formula to model the performance of Wayeb for a given hyperparameter set and input is impossible without training and testing Wayeb. Although an initial optimal hyperparameter set can be found for some historical dataset, in the run-time settings environmental changes might happen that can invalidate the deployed Wayeb’s PST. See for example Figure 2, which shows the MCC score of Wayeb on forecasting two CE in a maritime situational awareness setting and a financial fraud detection setting. In both cases data evolutions in the input stream cause significant fluctuations and drops in CEF scores. Consequently, there is a need for continuous adaptation over evolving data streams. RTCEF addresses this challenge with run-time PST retraining or hyperparameter optimisation. Challenge 3. Time-critical applications employing CEF require undisrupted production of forecasts. In critical applications, such as maritime situational awareness or credit card fraud management, updating the currently deployed PST with a new version should not stall the production of CE forecasts, as such delays would halt the proactive decision-making mechanisms of stakeholders. Consequently, updating the deployed PST with newly revised versions should happen in negligible time ensuring no disruptions in CEF and no loss of forecasts. Finally, although hyperparameter optimisation can result in high performing models, it does not come without a cost. Hyperparameter optimisation is, resource-wise, an expensive procedure which should only happen when necessary. The RTCEF framework addresses the above challenges using a novel distributed, service-oriented architecture. 4 Run-Time CEF Adaptation We start by presenting offCEF , a baseline framework for hyperparameter optimisation of CEF under the stationarity assumption, i.e., assuming that there are no evolutions in the input that might invalidate the CEF model. Subsequently, we present RTCEF , which addresses all challenges of run-time CEF mentioned in Section 3. 4.1 CEF Under the Stationarity Assumption Under the stationarity assumption, a single PST, produced through training on some historical, static dataset, will suffice for future input. Consequently, in this setting, we may use a framework for offline hyperparameter optimisation, hereafter offCEF . The aim of offCEF is the identification of an optimal configuration 𝑐𝑜𝑝𝑡 that yields the best performance for Wayeb, quantified by the MCC score (see Equation (1) ). A configuration 𝑐is defined as follows: 𝑐=[𝑚,𝜃fc,pMin,𝛾] where 𝑚 , 𝜃fc , pMin and 𝛾 are Wayeb’s hyperparameters (see Section 2) with their domain empirically set as: 𝑚∈ [1,5]𝜃fc ∈ [0.0,1.0] pMin ∈ [0.0001,0.01]𝛾∈ [0.0001,0.01] 12 Run-Time Adaptation of Complex Event Forecasting DEBS ’25, June 10–13, 2025, Gothenburg, Sweden Performance Metric x (a) Prior knowledge. Dinit x Performance Metric (b) 𝐷𝑖𝑛𝑖𝑡 samples (red lines). x a(x) value next micro-benchmark (c) Sampling via 𝑎(𝑥). Completed microbenchmarks x Performance Metric (d) BO conclusion. Figure 3: Bayesian Optimisation Operation. Model Factory (Offline) Controller (offline) Wayeb Server • Learn a prediction suffix tree • Estimate waiting time distributions TRAIN TEST C Score BO (GPR) Model C next micro-benchmark Acq. func. value Acquisition function BO optimiser Historical data Saved models ¬convergence ? new micro-bench.:deploy opt. model Hyperparameters [m, θfc, pMin, γ] MCC score Report Perform CEF under the stationarity assumption • Construct forecasts ε, (0.6, 0.4) a, (0.7, 0.3) b, (0.5, 0.5) aa, (0.75, 0.25) ba, (0.1, 0.9) 0 st a r t 1 2 3 4 a b b b a a a b ba 1 2 3 4 5 6 7 8 9 10 11 12 Number of future events 0 0.2 0.4 0.6 0.8 1 Completion Probability state:0 interval:5,12 state:1 state:2 state:3 Figure 4: Architecture of offCEF. Given the infinite parameter combinations, exhaustive search is computationally prohibitive. Furthermore, due to Wayeb’s complexity, performance for a given parameter set cannot be known beforehand. Consequently, to find the optimal configuration 𝑐 we employ Bayesian optimisation (BO) [ 6 , 14 ] i.e., a stochastic method for optimising expensive-to-evaluate objective functions that are complex or cannot be described by analytic formulae. In our work, the objective function is defined as 𝑓(𝑐)=MCC𝑐 , where MCC𝑐 denotes the MCC score of Wayeb given configuration 𝑐. The goal of BO is to find the vector of Wayeb’s hyperparameters that maximises CEF performance, using a minimal set of Wayeb training-test runs, termed ‘micro-benchmarks’, as training samples. Unlike other optimisation methods [ 34 ] BO does not require a high number of micro-benchmarks or an analytical formula [ 6 , 14 , 32 ]. BO employs a probabilistic model—called surrogate model—to approximate the unknown objective function, in our case CEF performance quantified by MCC , and iteratively refines this model. We employ a Gaussian Process Regressor (GPR) as the surrogate model. Initial beliefs about the objective function must be formulated before observing any data. In BO, priors are often specified for the mean and covariance functions of the Gaussian Process model. For example, a prior belief might suggest that the function is smooth and lies within a certain range of values. Priors are represented as: 𝑓(𝑐) ∼ GP(𝜇0(𝑐), 𝑘0(𝑐,𝑐′)) where 𝜇0(𝑐) and 𝑘0(𝑐, 𝑐′) are the prior mean and covariance (kernel) functions, respectively. Every time we observe a new micro-benchmark and collect CEF performance metrics by training and testing Wayeb given a configuration 𝑐 , we acquire a new training sample (𝑐, MCC𝑐) , to fit on the GPR, thereby updating our posterior belief in light of new evidence. The posterior distribution represents our updated knowledge about Wayeb’s performance and after observing 𝑛 new training samples, denoted by Data, the posterior is given by: 𝑓(𝑐) | Data ∼GP(𝜇𝑛(𝑐),𝑘𝑛(𝑐,𝑐′)) 𝜇𝑛(𝑐) and 𝑘𝑛(𝑐, 𝑐′) being the posterior mean and covariance functions updated through Bayesian inference [6, 14]. For selecting training samples, we start by randomly picking points from the input parameter domain, and then execute the respective micro-benchmarks and observe Wayeb’s MCC scores. We call this initial set of configurations 𝑐 , paired with MCCc scores, 𝐷𝑖𝑛𝑖𝑡 . Subsequently, using Bayesian inference, the first posteriors are calculated and the expected result is illustrated by comparing the prior in Figure 3a against the posterior in Figure 3b. After 𝐷𝑖𝑛𝑖𝑡 , the next micro-benchmarks are selected using an acquisition function 𝑎(𝑐) . The acquisition function guides the selection of the next evaluation point by quantifying the utility of sampling a particular point 𝑥 in the input space i.e., the domain of Wayeb’s configurations. 𝑎(𝑐) balances exploration and exploitation. Exploration involves sampling 𝑐 configurations in the input space that are not yet wellexplored or that have high uncertainty associated with them, while exploitation involves sampling 𝑐 points that are likely to yield the best objective function values exploiting the current knowledge. For instance, in the plot of Figure 3c the acquisition function chooses the point in the input domain with the highest uncertainty. Different acquisition functions introduce stochasticity in the BO process by incorporating uncertainty estimates from the probabilistic model. BO concludes either when a micro-benchmark budget is depleted or when the value of 𝑓(𝑐) converges. Figure 3d illustrates a GPR with minimal uncertainty around its mean values, after the microbenchmark budget has been depleted. 13 DEBS ’25, June 10–13, 2025, Gothenburg, Sweden Pitsikalis et al. Figure 4 illustrates the architecture of offCEF , comprising a Model Factory alongside a Controller. The Model Factory includes a Wayeb Server that utilises historical training and validation datasets to construct and evaluate PSTs. The Controller, includes the BO optimiser which is executed offline on a historical dataset. The Controller initialises BO by providing a set of configurations i.e., 𝑐 vectors to the Model Factory, which, respectively, conducts the prescribed micro-benchmarks, saves temporarily the candidate PSTs, and sends reports to the Controller. The Controller will use these reports for updating the GPR surrogate model of BO. offCEF deploys the PST that is expected to maximise MCC based on the hyperparameter vector 𝑐𝑜𝑝𝑡 calculated by BO. On the other hand, offCEF suffers from several disadvantages: (i) it drives its decisions by attributing equal importance to cumulative performance metric statistics, while in a streaming setup we often need to take into consideration only a sliding window of recent measurements and defy obsolete ones; (ii) it cannot optimise CEF hyperparameters at run-time which is a crucial limitation, since fluctuations in the input’s statistical properties in streaming settings is the norm rather than an infrequent situation; (iii) it cannot distinguish whether the hyperparameters for training PSTs should be adjusted through BO or if it is only the Wayeb’s PST that should be retrained, without changing hyperparameters. RTCEF, presented below, addresses these issues. 4.2 CEF Over Evolving Data Streams We propose RTCEF, which is built with three major goals in mind. First, it updates at run-time PSTs according to input data evolutions; second, it performs CEF without disruptions, i.e., PST updating does not cause delays on CEF; and third, it does not overuse resources for producing new PSTs. The architecture of RTCEF consists of five main services, acting as Kafka producers and consumers, running synergistically to ensure undisrupted CEF and dynamic PST retraining or hyperparameter optimisation. Figure 5 illustrates these services and the communication links between them. Synchronisation of the various services is denoted by dotted arrows in Figure 5. Below with details the processing of each service comprising our framework. Observer. In order to determine whether the MCC score of Wayeb has deteriorated, the quality of its forecasts must be monitored. This task is handled by the Observer service (right of Figure 5) which consumes MCC scores from the ‘Reports’ topic and, produces ‘retrain’ or ‘optimise’ instructions as indicated in Algorithm 1. Essentially, a retrain instruction requests a new PST for Wayeb without changing training hyperparameters. An optimisation instruction, requests a new PST, produced through hyperparameter optimisation. Note Collector Datasets Output stream Input stream Wayeb Models Reports Model Factory Commands Scores Controller Observer Instructs Data collection Optimisation and re-training Complex event forecasting Metrics monitoring Figure 5: Architecture of RTCEF . Cylinders and rounded rectangles denote topics and services respectively. For simplicity, we omit synchronisation topics; instead we use gray arrows. Algorithm 1 Observer service Require: 𝑘, guard_n,𝑚𝑎𝑥_𝑠𝑙𝑜𝑝𝑒, min_score 1: scores ← [] 2: guard ← −1 3: while True do 4: score𝑖←consume(Reports) 5: scores.update(score𝑖, k) 6: pit_cond ←score𝑖<min_score 7: slope_cond ←False 8: if guard ≥0then guard ←guard −1 9: if len(|𝑠𝑐𝑜𝑟𝑒𝑠|)>2then 10: (𝑎𝑖,𝑏𝑖) ← fit_trend(scores) 11: slope_cond ←𝑎𝑖<max_slope 12: if (slope_cond and guard ≥0)or pit_cond then 13: send(“instructions”, “optimise”) 14: guard ←guard_n⊲New guard period 15: else if slope_cond then 16: send(“instructions”, “retrain”) 17: guard ←guard_n⊲New guard period that hyperparameter optimisation will provide the best possible hyperparameters, but can be costly procedure, whereas retraining on an updated dataset is a cheaper process. We describe Algorithm 1 following its illustrative execution example for maritime situational awareness presented in Figure 6. The Observer continuously consumes MCC scores from Wayeb and retains the 𝑘 most recent MCC scores to evaluate the performance trend. In the example of Figure 6, Wayeb begins with a PST, referred to as PST 𝑤0 , created using configuration 𝑐𝑤0 . The Observer records the MCC Score at 𝑤0 , however at this point no decision is made since fewer than 𝑘= 3scores have been collected. Once the Observer has at least 𝑘 scores, it computes the first degree polynomial 𝑧𝑖(𝑥)=𝑎𝑖𝑥+𝑏𝑖 (a trend line) so that 𝑎𝑖 and 𝑏𝑖 minimise 14 Run-Time Adaptation of Complex Event Forecasting DEBS ’25, June 10–13, 2025, Gothenburg, Sweden 𝑤0𝑤1𝑤2𝑤3𝑤4𝑤5𝑤6 0.4 0.6 0.8 𝛼𝑤2≃ −0.04𝛼𝑤4≃ −0.05 guard guard Weeks MCC Wayeb rt opt trend trend Figure 6: Execution example of the Observer. ‘opt’ and ‘rt’ stand for ‘optimisation’ and ‘retraining’ respectively. Dashed lines correspond to the trend lines associated with the Observer’s instructions at 𝑤2 and 𝑤4 , and black lines correspond to guard periods. the squared error 𝐸=Í𝑗=𝑘 𝑗=0𝑧𝑖(𝑥𝑗) −𝑦𝑗2 for 𝑥𝑗=𝑗 and 𝑦𝑗=score𝑖−𝑘+𝑗 , where 𝑖 is an increasing integer denoting the ID of the current score (lines 9, 10). If the slope ( 𝑎𝑖 ) of 𝑧𝑖(𝑥) is negative, indicating decrease in performance, and less than a max_slope ∈R− parameter (line 11) then a ‘retrain’ instruction is produced (lines 15-17). In the example, by week 𝑤2 , the Observer has MCC scores for 𝑤0,𝑤1,𝑤2 . Using these points, the Observer computes the trend line 𝑧𝑤2 with a slope 𝛼𝑤2=− 0 . 04 which is steeper than max_slope =− 0 . 02. To remedy this behaviour, the Observer issues a retrain instruction. As a result, a new PST, referred to as PST 𝑤2 , is created using the same configuration as PST 𝑤0 , i.e., 𝑐𝑤0 . This occurs because retraining updates the PST without modifying Wayeb’s hyperparameters. Intuitively, forecasting performance deterioration, demonstrated by 𝑎𝑖<max_slope , shortly after a new PST deployment, indicates that the new PST failed and hyperparameter optimisation should thus be performed. To this end, we place each newly deployed PST in a guard period (lines 14,17). A guard period starts after a PST is deployed, and ends after guard_n performance reports. If the performance of a PST under a guard period deteriorates ( 𝑎𝑖<max_slope ) then a hyperparameter optimisation instruction is produced (lines 12,13). If on the other hand, 𝑎𝑖<max_slope is satisfied after guard_n reports, then a ‘retrain’ instruction is produced for which a new “guard” period begins. In the example of Figure 6, a guard period begins at week 𝑤2 and will last for guard_n= 4reports, i.e., until 𝑤5 . While PST 𝑤2 shows improvement at 𝑤3 , at 𝑤4 performance drops again. The performance drop is also confirmed by the slope of -0.05 computed by the Observer using the MCC scores from 𝑤2 to 𝑤4 . Since the slope 𝛼𝑤4 is again below max_slope , but this time a guard period is active, the Observer issues a hyperparameter optimisation instruction instead of retraining. Consequently, a new PST 𝑤4 is produced through hyperparameter optimisation, resulting in an updated configuration 𝑐𝑤4 , and a new guard period starting at 𝑤4 . Finally, to avoid pitfalls whereby the score drops suddenly very low, we employ an additional condition: if the score of a report is lower than a threshold min_score (line 6) then the Observer asks directly for ‘optimisation’ and omits a ‘retrain’ instruction. Wayeb. The CEF part of RTCEF (top of Figure 5) contains Wayeb. In addition to reading timestamped simple events from the input stream and producing an output stream of CE forecasts, Wayeb produces a stream CEF forecasting performance reports, equally distanced by reporting_distance , and continuously monitors the ‘Models’ topic, which contains updated PSTs. When a new PST is made available in the Models topic, Wayeb replaces its PST with the latest available version. Recall that, to produce a CE forecast, Wayeb will utilise both the automaton corresponding to the symbolic regular expression defining a CE and the PST (see Section 2). The automaton retains information about the current state ( 𝑞 ) and the next states that can lead to an accepting run, while the PST is used for producing the next symbol probabilities and therefore the waiting-time distribution for state 𝑞 ( 𝑊𝑞 ). Below, we show Wayeb PST update is “lossless” i.e., upon PST replacement, any run can continue from its current state and produce forecasts using the new PST. Proposition 1. Given a stream 𝑆={𝜎0, 𝜎1, ..., 𝜎𝑘} where 𝜎𝑖 are symbols, a SRE 𝑅 , its corresponding automaton 𝐴𝑅 , and a PST 𝑇 with order 𝑚∈ [𝑚𝑙,𝑚𝑢] , the replacement of 𝑇 with a 𝑇′ of order 𝑚′∈ [𝑚𝑙,𝑚𝑢] is lossless at any position 𝑖 of the stream 𝑆 if the last 𝑚𝑢 symbols from the 𝜎𝑖 are available. Proof. We prove Proposition 1 by contradiction. Given 𝑆 , 𝑅 , 𝐴𝑅 and 𝑇 with order 𝑚∈ [𝑚𝑙,𝑚𝑢] , assume that 𝑇 is replaced with a PST 𝑇′ with order 𝑚′ at a position 𝑖 . Now, assume that there exists a run that cannot continue from its current state 𝑞 . This is not possible as the SRE 𝑅 remains the same and therefore the automaton 𝐴𝑅 is also the same, consequently the run can continue from 𝑞 . Next, we assume that there exists a run with current state 𝑞 for which the next symbol probabilities and the corresponding waitingtime distribution cannot be computed at position 𝑖 under 𝑇′ . Recall, that for an automaton run at position 𝑗 and a current state 𝑞 , the next symbol probability (and 𝑊𝑞 ) can be computed using the PST 𝑇 and 𝑆[𝑗−𝑚+1,𝑗 ] , denoting the subset of 𝑆 containing the symbols {𝜎𝑗−𝑚+1, ..., 𝜎𝑗} . Since, 𝑆[𝑖−𝑚𝑢+1,𝑖] is available, again the next symbol probabilities as well as the waiting-time distribution for any state 𝑞 can be computed with 𝑇′as 𝑆[𝑖−𝑚′+1,𝑖]⊆𝑆[𝑖−𝑚𝑢+1,𝑖].□ Proposition 1 states that updating a PST 𝑇 with a new 𝑇′ can be lossless if: first, the order 𝑚 of any new PST lies within the same range [𝑚𝑙,𝑚𝑢] ; and, second, the last 𝑚 symbols up to the moment of the replacement are available. RTCEF ensures both conditions are satisfied. Consider, for example, 15 DEBS ’25, June 10–13, 2025, Gothenburg, Sweden Pitsikalis et al. 𝜖 (0.6,0.4) 𝑎 (0.7,0.3) 𝑎𝑎 (0.9,0.1) 𝑏𝑎 (0.6,0.4) 𝑏 (0.5,0,5) (a) PST 𝑇with 𝑚=2. 𝜖 (0.6,0.4) 𝑎 (0.6,0.4) 𝑎𝑎 (0.9,0.1) 𝑏𝑎 (0.7,0.3) 𝑎𝑏𝑎 (0.8,0.2) 𝑏𝑏𝑎 (0.2,0.8) 𝑏 (0.5,0,5) (b) PST 𝑇′with 𝑚′=3. Figure 7: PST 𝑇 with 𝑚= 2and updated PST 𝑇′ with 𝑚= 3, for 𝑅 : =(𝑠𝑝𝑒𝑒𝑑 > 10 )·(𝑠𝑝𝑒𝑒𝑑 > 10 ) and the symbols 𝑎=‘(𝑠𝑝𝑒𝑒𝑑 >10)’ and 𝑏=‘¬(𝑠𝑝𝑒𝑒𝑑 >10)’. the regular expression 𝑅 : =(𝑠𝑝𝑒𝑒𝑑 > 10 ) · (𝑠𝑝𝑒𝑒𝑑 > 10 ) , presented earlier in Section 2 as well as its corresponding automaton illustrated Figure 1. A PST 𝑇 with order 𝑚= 2, along with a revised PST 𝑇′ with order 𝑚′= 3, and 𝑚′<𝑚𝑢 , for 𝑅 are illustrated in Figures 7a and 7b respectively. Here, ‘ 𝑎 ’ corresponds to the symbol ‘ (speed >10) ’, while ‘ 𝑏 ’ corresponds to the symbol ‘ ¬(speed >10) ’. Assume that CEF begins with PST 𝑇 and processes an input stream 𝑆={𝑏, 𝑎, 𝑏, 𝑎} . After consuming 𝑆 , the current state of the automaton of Figure 1 is 1, while the probability of completion in one step, i.e., receiving another 𝑎 , is 0 . 6(see the 𝑏𝑎 node in 7a). At this point, if we choose to replace 𝑇 with 𝑇′ , for a lossless transition we need at most the last 3symbols from 𝑆 —the order of 𝑇′ is 3. Since these symbols are available and the current state is 1, the new probability of completion in one step after consuming 𝑆 , and by considering 𝑇′ , is 0 . 8(see the 𝑎𝑏𝑎 node in Figure 7b). Notably, PST replacement, can be executed in linear time with respect to the number of automata runs and in practice happens in negligible time. Collector. Training datasets evolve over time. Therefore, the data collection part of RTCEF (left of Figure 5) includes the Collector service, a data processing module organising and storing subsets of the input stream that may be used for retraining or hyperparameter optimisation. The Collector service consumes the input stream (see ‘Data collection’ in Figure 5), in parallel to Wayeb, and stores subsets of it in time buckets of fixed bucket_size . The Collector gathers data in a sliding window manner, emitting a new dataset version, containing dt_size buckets, to the ‘Datasets’ topic as soon as the last bucket in the range is full. Old buckets that no longer serve a purpose for training, are deleted for space economy. Controller. The Controller service, based on the instructions of the Observer, initialises hyperparameter optimisation procedures, during which it also serves as the Bayesian optimiser, or retraining procedures, where it supplies Wayeb configurations. When optimisation is required, the Controller initiates the following three phases. Initialisation phase: The Controller sets up the Bayesian optimiser. Similar to [ 10 ], we reuse micro-benchmarks from previous runs. Using the retain_fraction ∈ [ 0 , 1 ] parameter, we uniformly keep ⌊retain_fraction ∗all_samples⌋ observations from the last BO run, where all_samples is the total number of micro-benchmarks. This speeds up optimisation while preserving useful information from previous runs. Step phase: The Controller issues ‘train & test’ commands along with the hyperparameters suggested by the acquisition function—in our case the acquisition function is a combination of lower confidence bound, expected improvement, and probability of improvement. After each ‘train & test’ command, the Controller awaits the corresponding performance report i.e., the value of the MCC(c) objective function. Upon receiving the performance report, the optimiser is updated with the new sample and the hyperparameters for the next step are suggested. The step phase ends when all microbenchmarks are completed or if convergence is achieved. Finalisation phase: Once optimisation concludes and the best hyperparameters are acquired, the Controller sends a finalisation message containing the ID of the best PST. Additionally, the Controller updates the previously best hyperparameters with the newly acquired ones, ensuring availability of the latter for subsequent ‘retrain’ instructions. Model Factory. Similar to offCEF , the primary function of the Model Factory service is to train, test and send up-to-date PSTs to Wayeb. To do this, it will assemble and use the latest dataset version produced by the Collector. Upon receiving a ‘train’ command, the Model Factory trains a PST on the latest dataset and shares this new PST version with Wayeb. For PST production through hyperparameter optimisation, upon receiving an ‘initialisation’ message, the Model Factory ‘locks’ the most recent assembled dataset so that the same dataset is used throughout the optimisation procedure. Next, during the ‘step’ phase, the Model Factory trains, saves and tests candidate PSTs on the locked dataset and reports MCC scores to the Controller. Finally, when the BO ‘finalisation’ message is received, the Model Factory sends the best performing PST to Wayeb. It is only at this point, that Wayeb will stop momentarily for PST replacement. 5 Experimental Evaluation We evaluate our framework on two real-world use-cases. First, in maritime situational awareness, maritime CEs of interest are forecast over real vessel position streams. Second, in credit card fraud management, fraudulent activity is forecast over synthetic transaction data. We describe our experimental setup and then we present our findings. 5.1 Experimental Setup 5.1.1 Datasets & patterns. We present the datasets we employ and the patterns we use for forecasting CEs of interest. 16 Run-Time Adaptation of Complex Event Forecasting DEBS ’25, June 10–13, 2025, Gothenburg, Sweden Maritime situational awareness. We use a real-world, publicly available, maritime dataset containing 18M spatiotemporal positional AIS (Automatic Identification System) messages transmitted between October 1st 2016 and 31st March 2016 (6 months), from 5K vessels sailing in the Atlantic Ocean around the port of Brest, France [ 29 ]. AIS allows the transmission of information such as the current speed, heading and coordinates of vessels, as well as, ancillary static information such as destination and ship type. We evaluate RTCEF on a maritime pattern, which expresses the arrival of a vessel at the main port of Brest [ 32 ]. This pattern is derived after discussions with domain experts from a large, European maritime service provider [27, 28]: 𝑅port :=(¬InPort(Brest))∗· (¬InPort(Brest)) · (¬InPort(Brest)) · (InPort(Brest)) (2) InPort(Brest) is true when a vessel is within 5 km from the port of Brest. Recall that ‘ ¬ ’, ‘ ∗ ’ and ‘ · ’ correspond to negation, iteration (Kleene star) and sequence respectively (see Section 2). Consequently, 𝑅port is satisfied if a sequence of at least three events occur. At least two require the vessel to be away from the port—thus limiting false positives from noisy entrances—, while the last denotes that the vessel has entered the port. This CE is important for port management and logistics reasons. We also perform experiments for a CE named 𝑅fish defined as follows: 𝑅fish :=(¬InArea(Fishing))∗· (¬InArea(Fishing)) · (¬InArea(Fishing)) · (InArea(Fishing) ∧ ¬SpeedRange(Fishing))∗· (InArea(Fishing) ∧ SpeedRange(Fishing)) (3) InArea(Fishing) is true when a vessel is within a fishing area, while speedRange is a predicate satisfied when the vessel has fishing speed [ 28 ]. Therefore, 𝑅fish is satisfied when initially a vessel is outside a fishing area, then the vessel enters the fishing area and; at some point while it is within the fishing area, it has fishing speed. Monitoring (illegal) fishing is important for environmental and sustainability reasons. To cross validate our approach, we create 6 datasets MD𝑖 , 𝑖∈ [0,5]by shifting the starting month in a cyclic manner: MD𝑖=  𝑗=5 𝑗=0month(𝑗+𝑖)mod 6 where ∥ denotes the operation of concatenating two datasets, and monthk corresponds to month 𝑘 of the original dataset. Credit card fraud management. We use a synthetic dataset provided by Feedzai 2 containing 1M credit card transaction events taking place over a period of 82 weeks. Each event contains, among others, the card ID, the amount and time of the transaction. We evaluate RTCEF on a pattern representing a fraudulent behaviour as a sequence of consecutive increasing transactions. Again, this pattern was determined 2https://feedzai.com after discussions with domain experts from a large European credit card management service provider [3]: 𝑅cards :=(amDiff >0)·(amDiff >0)·(amDiff >0)· (amDiff >0)·(amDiff >0)·(amDiff >0)· (amDiff >0) (4) We enrich events of the input stream with an additional attribute amDiff which is equal to the difference between the previous transaction and the current one. Therefore, 𝑅cards is satisfied when 8 consecutive transactions happen with increasing amounts. In order to simulate evolving fraud, we modify the financial dataset by changing randomly every 4 to 8 weeks the range of the highly correlated feature amountDiff . Similar to the maritime dataset, for validating our results, we create 21 datasets FD𝑖,𝑖 ∈ [ 0 , 20 ] by shifting 4 weeks the start of each dataset in a cyclic manner. 5.1.2 Initialisation. We perform offline hyperparameter optimisation with offCEF on the first four weeks of each dataset FD/MD𝑖 and use the resulting PST, hyperparameters and micro-benchmark samples for initialising RTCEF . To showcase the benefit of RTCEF , we also perform CEF with static PSTs yielded by offCEF (see Section 4.1): i.e., for each FD𝑖 / MD𝑖 we adopt the stationarity assumption and perform CEF using the corresponding initial PST of each dataset. In what follows, the experiments that utilise the run-time adaptation framework are labelled with ‘ RTCEF ’ while experiments that are performed only with offline optimised static PSTs models are labelled with ‘ offCEF ’. Both offCEF and RTCEF ingest the input stream with the maximum speed of Kafka. offCEF and RTCEF are implemented in Python 3.9.18, while the Kafka version was 3.5.2. Messages are formatted in JSON, and serialised/deserialised using Apache AVRO format. For BO, we use the scikit-optimize library 0.9.0. The experiments are conducted on a server running Debian 12 with an AMD EPYC 7543 32-Core Processor and 400G of RAM. Each service of RTCEF runs on its own dedicated core. Our framework is open-source and our experiments are reproducible. 5.2 Experimental Results Figures 8 and 10 show MCC over time for 𝑅port and 𝑅cards (see Definitions (2) and (4) respectively), along with the score improvements when using RTCEF as opposed to offCEF for the maritime MD0/2/3/5 and financial FD0/1/8/17 datasets respectively. Results concerning MD0 and FD0 —the datasets in their original order—show that offCEF demonstrates, in both cases, poor performance and significant fluctuations in MCC scores over time. RTCEF , on the other hand, improves scores and reduces fluctuations in both cases. Maritime situational awareness. For MD0 —the maritime dataset in its original order— RTCEF drammatically improves 17