scieee AI-readable full text Open interactive document viewer

Analysis of Resource Working Behaviour and Working Patterns Utilizing High-Level Batch Aggregation

Aliakberova, Liliia

Abstract

This thesis examines the analysis of resource behavior and working patterns using high-level batch aggregation in a European company (referred to as Company C), which specializes in sterilizing medical equipment. The research addresses the question: What patterns in employee work habits and task prioritization can be identified within Company C’s sterilization process, and how can these patterns define the features provided for a digital twin that enhances operational efficiency and compliance? To address this, we developed a methodology that first identifies batching within the Event Knowledge Graph of Company C. These batches are then aggregated into high-level events to reflect iterative patterns. We further analyze resource behavior and working patterns based on this high-level event log. The findings reveal how processing is influenced by workload and highlight common working patterns across weekdays and shifts. These insights are intended to support the development of a digital twin that enhances operational efficiency and compliance.

Full text

Eindhoven University of Technology MASTER Analysis of Resource Working Behaviour and Working Patterns Utilizing High-Level Batch Aggregation Aliakberova, Liliia Award date: 2024 Link to publication Disclaimer This document contains a student thesis (bachelor's or master's), as authored by a student at Eindhoven University of Technology. Student theses are made available in the TU/e repository upon obtaining the required degree. The grade received is not published on the document as presented in the repository. The required complexity or quality of research of student theses may vary by program, and the required minimum study period may vary in duration. General rights Copyright and moral rights for the publications made accessible in the public portal are retained by the authors and/or other copyright owners and it is a condition of accessing publications that users recognise and abide by the legal requirements associated with these rights. • Users may download and print one copy of any publication from the public portal for the purpose of private study or research. • You may not further distribute the material or use it for any profit-making activity or commercial gain Take down policy If you believe that this document breaches copyright please contact us providing details, and we will remove access to the work immediately and investigate your claim. Download date: 17. Nov. 2025 Department of Mathematics and Computer Science Process Analytics Big Data Management and Analytics (BDMA) Analysis of Resource Working Behaviour and Working Patterns Utilizing High-Level Batch Aggregation Master Thesis Liliia Aliakberova Supervisors: Dirk Fahland 25-08-2024 Abstract This thesis examines the analysis of resource behavior and working patterns using high-level batch aggregation in a European company (referred to as Company C), which specializes in sterilizing medical equipment. The research addresses the question: What patterns in employee work habits and task prioritization can be identified within Company C’s sterilization process, and how can these patterns define the features provided for a digital twin that enhances operational efficiency and compliance? To address this, we developed a methodology that first identifies batching within the Event Knowledge Graph of Company C. These batches are then aggregated into high-level events to reflect iterative patterns. We further analyze resource behavior and working patterns based on this high-level event log. The findings reveal how processing is influenced by workload and highlight common working patterns across weekdays and shifts. These insights are intended to support the development of a digital twin that enhances operational efficiency and compliance. Keywords: Resource Behavior, Working Patterns, Batching, High-Level Batch Aggregation, Event Knowledge Graph (EKG), Process Mining Preface As I reflect on this journey, my first thoughts are with my parents. I would like to express my deepest gratitude to my mother and father. They have been my heroes, showing me what success, hard work, and persistence look like. Their unwavering support has been always there for me despite the distance. Their love and encouragement have guided me every step of the way. Without them, none of this would have been possible. My partner, Rick, has been by my side through all the ups and downs. His belief in me, especially when I doubted myself, has given me strength. Your presence in my life has been invaluable, Rick, and I am endlessly grateful for it. I also want to extend my thanks to Rick’s family for their kindness and support throughout this journey. To my friends, both those I have known for years and the ones I have met in the BDMA program—thank you for your support and companionship. Your friendship has made this journey so much richer and more enjoyable. I am deeply grateful to the BDMA consortium for the unique opportunity to study at three distinguished universities. This experience has allowed me to explore Europe, immerse myself in different cultures, and learn new languages. For me, the BDMA program was not just about earning a degree; it was a transformative adventure that expanded my horizons in countless ways. I also want to express my sincere appreciation to the professors for sharing their knowledge and guiding me during my master degree. A special thanks to Professor Dirk Fahland for his guidance, assistance, and for providing both constructive and critical feedback that has been essential in shaping this work. Lastly, I want to thank myself for not giving up. 2 Contents 1 Introduction 6 1.1 ContextandTopic ........................ 6 1.2 StateoftheArt.......................... 7 1.3 ResearchQuestions........................ 9 1.4 ResearchMethod ......................... 10 1.5 Findings.............................. 11 2 Background 14 2.1 AUTO-TWIN Project Description . . . . . . . . . . . . . . . . 14 2.2 Business Understanding . . . . . . . . . . . . . . . . . . . . . 15 2.2.1 Introduction to Company C . . . . . . . . . . . . . . . 15 2.2.2 Sterilization Center Arrangement . . . . . . . . . . . . 15 2.2.3 The Sterilization Process . . . . . . . . . . . . . . . . . 17 2.3 Initial Data Understanding . . . . . . . . . . . . . . . . . . . . 18 2.3.1 Initial Data Understanding: Data Issues . . . . . . . . 20 2.3.2 Initial Data Understanding: Resource Perspective . . . 21 2.4 Problem statements from the business perspective . . . . . . . 23 3 Discussion on Preliminaries and Related Work 24 3.1 ProcessMining .......................... 24 3.2 Graphs............................... 26 3.2.1 Labeled Property Graphs (LPGs) . . . . . . . . . . . . 26 3.2.2 Neo4j and the Cypher Query Language . . . . . . . . . 27 3.3 Multi-dimensional Process Mining . . . . . . . . . . . . . . . . 28 3.4 Resource-Oriented Process Mining . . . . . . . . . . . . . . . . 30 3.4.1 Batching.......................... 30 3.4.2 Human Resource Working Patterns . . . . . . . . . . . 33 3.5 Analytical Techniques in Process and Data-Driven Research: Correlation and Sequence Analysis . . . . . . . . . . . . . . . 33 3.5.1 Normalization and Correlation Analysis . . . . . . . . . 34 3.5.2 Frequent Sequence Analysis . . . . . . . . . . . . . . . 34 3 4 Problem Exposition 36 4.1 Detailed Research Questions . . . . . . . . . . . . . . . . . . 36 4.2 DetailedMethod ......................... 37 5 From EKG Implementation to Extended Resource Exploration and Task Instances Analysis 42 5.1 EKG Implementation . . . . . . . . . . . . . . . . . . . . . . . 43 5.1.1 Preprocessing for EKG Implementation . . . . . . . . . 43 5.1.2 EKG Implementation Statistics . . . . . . . . . . . . . 43 5.2 Extended Resource Exploration Utilizing EKG . . . . . . . . . 44 5.3 Applying Task Instance Framework over Company C’s EKG . 47 5.3.1 Limitations of the Approach . . . . . . . . . . . . . . . 47 5.3.2 Need for Batch Detection Revision . . . . . . . . . . . 48 6 Batching 49 6.1 Batching in the Event Knowledge Graph: Proposing Options forCompanyC .......................... 49 6.2 Batching over Resource . . . . . . . . . . . . . . . . . . . . . . 53 6.2.1 Mathematical Definition of Batching over Resource . . 55 6.2.2 Batching over Resource: Implementation Steps . . . . . 56 6.2.3 Batching over Resource: Implementation Results and Exploration ........................ 59 6.3 Batching over Activity . . . . . . . . . . . . . . . . . . . . . . 65 6.3.1 Mathematical Definition of Batching over Activity . . . 67 6.3.2 Batching over Activity: Implementation Steps . . . . . 67 6.3.3 Batching over Activity: Implementation Results and Exploration ........................ 72 7 High-Level Batching 80 7.1 Motivation and Requirements for High-Level Batch Aggregation 80 7.1.1 Batching Patterns that Lead to High-Level Aggregation 83 7.1.2 Requirements for High-level Batching Aggregation . . . 85 7.2 High-Level Batching: Definitions and Implementation Steps . 86 7.2.1 Mathematical Definition of High-Level Batches . . . . 86 7.2.2 Implementation of High-Level Batching Aggregation . . 87 7.2.3 High-Level Batching: Implementation Results and Representation in the EKG . . . . . . . . . . . . . . . . . . 89 8 High-Level Resource Behavioral Analysis 94 8.1 Forming High-Level Event Log . . . . . . . . . . . . . . . . . . 94 4 8.2 Resource Working Behavior Based on High-Level Batching andWorkload........................... 96 8.2.1 Defining the Workload . . . . . . . . . . . . . . . . . . 97 8.2.2 Detailed Approach for Analyzing Resource Working Behavior.......................... 97 8.2.3 Results Evaluation of Analyzing Resource Working Behavior ...........................102 8.2.4 General Overview of Resource Working Behavior . . . 107 8.3 Frequent Working Patterns Using High-Level Aggregation . . 108 8.3.1 Detailed Approach for Analyzing Frequent Working Patterns..........................108 8.3.2 Results Evaluation of Frequent Working Patterns Analysis.............................109 9 Conclusion 114 9.1 Limitations ............................115 9.2 FutureWork............................116 APPENDICES 121 A Frequent Working Patterns Analysis Results 125 5 Chapter 1 Introduction This chapter outlines the overall scope of the master thesis, which was conducted as part of the Auto-Twin project focused on developing advanced digital twin technologies. Within this project, we concentrate on the process of a European company (referred to as Company C) that specializes in sterilizing medical equipment. This chapter introduces the overall scope of the master thesis. Section 1.1 presents the main topics and provides context for the study. In Section 1.2, we review relevant literature and methodologies in the field. The central research question, which investigates identifiable patterns in employee work habits and task prioritization within Company C’s sterilization process, is discussed in Section 1.3, along with subquestions that further explore how these patterns can inform the development of an effective digital twin to enhance operational efficiency and compliance. The methodology used to investigate these questions, including data exploration, the development of batching techniques, high-level batching aggregation and quantitative analysis, is detailed in Section 1.4. Finally, Section 1.5 presents the findings and their interpretations. 1.1 Context and Topic In healthcare, the sterilization of medical devices is a process that ensures instruments are safe for use in medical procedures. This process relies not only on advanced equipment and strict protocols but also on the actions and decisions of the people involved. The effectiveness of the sterilization process depends significantly on how human resources, such as technicians and operators, interact with equipment, navigate through workstations, and coordinate with one another. Understanding this resource behavior, which includes how employees prioritize tasks and manage workloads, is important 6 for improving the efficiency and reliability of the sterilization process. Recognizing the important role of human resources in these operations, the AUTO-TWIN project focuses on advancing the development and application of Digital Twins that represent digital models that replicate real-world processes. By accurately representing both the physical workflow and the behavior of the people involved, Digital Twins offer a method for optimizing business operations. One of the key areas of focus within the AUTO-TWIN project is the sterilization process at Company C, a provider of sterilization services for medical devices. In this context, the project aims to develop a Digital Twin that not only reflects the physical aspects of the sterilization process but also captures and analyzes the behavior of human resources, which is necessary for improving efficiency, reliability, and compliance. Company C operates sterilization centers where human resources must work together with advanced equipment to ensure that medical instruments are properly sterilized. However, the complexity of these operations means that having advanced equipment and protocols alone is not sufficient. The success of the sterilization process also depends on human resources. Thus, Company C aims to better understand how its human resources work within the sterilization process in order to further develop a Digital Twin that accurately reflects these operations. This research directly contributes to Company C’s objectives by developing a framework that aggregates operational data into high-level features based on resource timelines. These features are designed to capture and analyze how human resources at Company C interact with the sterilization process over time, with a focus on identifying specific behaviors and patterns. By providing these insights, the research supports the creation of a Digital Twin. By accomplishing this, the research not only contributes to Company C but also addresses a broader gap in understanding how human resources interact with complex processes. 1.2 State of the Art Analyzing resource behavior is a complex task, especially in environments where multiple entities interact, like equipment, personnel and stations. This complexity is particularly significant when trying to understand how employees prioritize tasks and perform their work. Process mining is a field that helps address these challenges by analyzing event logs from information systems to discover, monitor, and improve real processes [28]. By examining this event data, process mining provides insights into how processes function, how resources are allocated, and how different components of a process 7 Chapter 2 Background This chapter outlines the background context for this master’s thesis. It starts with an overview of the AUTO-TWIN project in the Section 2.1, establishing the basis for the research. Following this, we describe Company C, its functions and operations. In Section 2.2, we explain the sterilization process at Company C, which is the central subject of this thesis. Section 2.3 offers a preliminary overview of the data utilized for analysis. Lastly, Section 2.4 addresses the problem statements from a business perspective, clarifying the importance of this thesis. Together, these sections provide the necessary background and context to understand the behavior of Company C’s resources in the sterilization process. 2.1 AUTO-TWIN Project Description The AUTO-TWIN project [1] tackles important technological and economic challenges. Its main goal is to improve the development and use of Digital Twins [24], making business operations more efficient, reliable, and sustainable within circular economies [25]. The project introduces several advanced methods and technologies: •Digital Twins: AUTO-TWIN uses an automated, process-aware approach to create Digital Twins [24]. These Digital Twins accurately replicate real-world operations digitally to support reliable business processes [1]. •IDS-based Common Data Space: The project uses an International Data Space (IDS). A data space is a virtual environment that offers a standardized framework for data exchange, utilizing common protocols 14 and formats [21]. This improves data interoperability and resource use in a circular-economy ecosystem [25]. •Smart Green Gateways: A Green Gateway is a strategic decisionmaking tool that helps maximize the value and efficiency of products and materials [10]. AUTO-TWIN integrates advanced hardware into the digital thread to create smart Green Gateways that make efficient and environmentally friendly data-driven decisions using Digital Twins. The AUTO-TWIN project includes three specific use cases, one of which is the Sterilization Process in the Health Sector at Company C. This thesis focuses on this use case. The effectiveness of the sterilization process at Company C largely depends on human resources allocation and behavior, which are key factors in maintaining an efficient workflow. Therefore, this thesis focuses on understanding human resource behavior and working patterns in this process. This understanding is essential for developing a Digital Twin with AUTOTWIN project that accurately reflects the real-world process and improves operational performance and sustainability, aligning with the project’s goals. 2.2 Business Understanding 2.2.1 Introduction to Company C Company C has established itself as a company responsible for distribution and administration of surgical instruments and medical equipment. With many years of experience, the company has adapted to meet the needs of the healthcare industry. It offers a range of products and services aimed at maintaining high standards of medical care. One important contribution of Company C is the development of advanced sterilization centers. 2.2.2 Sterilization Center Arrangement This thesis examines an sterilization process at the advanced sterilization center of Company C inside of a hospital. Figure 2.1 shows the center’s design, which supports an efficient workflow with dedicated rooms and areas. Each area has specific machines and tools for different stages of the sterilization process. The sterilization process commences with acceptance of medical devices in the sterilization center through the hallway. The first area they encounter is the dirty area, where the devices, assembled into medical kits, unpacked 15 Figure 2.1: Facility Arrangement of Company C’s Sterilization Center and pre-washed. This area is essential for removing any contaminants before the kits move to more specialized cleaning processes. Metal kits are placed into automated washing machines that are designed for high-throughput and efficient cleaning, while plastic kits are manually cleaned to prevent damage due to their material properties. After the initial cleaning phase, both types of kits proceed to the medical device storage and entrance cleaning area for a detailed inspection. This area serves as a critical checkpoint where kits are examined for any remaining contaminants or damage. Metal kits are then packed in preparation for sterilization, while plastic kits are unpacked to ensure they are ready for the next phase. Once the devices pass inspection, they move to the clean area. This area is designed for further cleaning processes to ensure the devices are free from any contaminants. Following this, the devices proceed to the sterilization area, where they undergo the sterilization process. The sterilization area is equipped with multiple sterilizers to handle a high volume of devices efficiently. 16 After sterilization, the devices are transferred to the storage area. This zone is used to keep the sterilized devices in a controlled environment until they are needed for use. Throughout the process, scanners are used at various stations to scan the kits as they go through each step of the sterilization process, ensuring proper tracking and documentation. The facility also includes a sterilization corridor that connects different areas of the center, ensuring a smooth transition of devices from one zone to another. Additionally, there is a hatch used for passing items between areas while maintaining the cleanliness and sterility of the environment. Each part of the sterilization center is specifically designed to support a streamlined and effective sterilization process, ensuring that medical devices are handled safely and efficiently from entry to storage. 2.2.3 The Sterilization Process The next important step is to examine the sterilization process in detail. The process presents is a multi-stage operation designed to ensure the highest level of cleanliness and sterility for medical devices, including both plastic and metal kits. The Figure 2.2 represents the C company sterilization process. Figure 2.2: Company C Sterilization Process Diagram The process begins with a Check-in phase (P1) for both plastic and metal kits, where the kits are logged and prepared for the subsequent stages. For metal kits, the process continues with unpacking and pre-washing (P2), followed by machine washing (P3) using racks for batch processing. This method ensures that multiple kits can be processed simultaneously, increasing efficiency. After machine washing, metal kits undergo inspection and 17 packing (P4), where they are thoroughly checked and repacked for sterilization. Plastic kits, on the other hand, follow a slightly different path due to their material properties. They undergo pre-washing (P5), manual washing (P6), and unpacking and inspection (P7). Manual washing is particularly important for plastic kits to prevent any potential damage from automated machines. During the inspection phase, both types of kits are checked for any malfunctions or remaining contaminants. Once the initial cleaning and inspection stages are completed, both types of kits converge for the sterilization stage (P8) using racks for batch processing. This stage is critical as it involves the actual sterilization process, where the kits are exposed to high temperatures and pressure to be fully sterilized. After sterilization, the process diverges again. Metal kits move to the collection stage (P9), where they are prepared for storage. This involves final inspections and packaging to ensure they remain sterile until use. Plastic kits proceed to the packing and collection stage (P10), where they are similarly inspected and packed. Finally, all kits go through the Checkout stage (P11), marking the completion of the sterilization process. This stage involves final logging. In addition, the process involves several human and physical resources. There are 20 unique human resources involved in the sterilization process at Company C, each playing an important role in ensuring the smooth operation of the workflow. They are responsible for scanning barcodes and ensuring compliance with the First In, First Out (FIFO) principle, which ensures that kits are sterilized in the order they arrive. This practice maintains the order and traceability of the kits throughout the entire process. Overall, the sterilization center at Company C exemplifies a well-organized and efficient system designed to meet the highest standards of sterility and safety in the healthcare industry. 2.3 Initial Data Understanding Another important aspect of the project context is the initial understanding of the data. Company C provided the historical process data that spans from January 1st, 2022, to March 31st, 2022, and was provided directly from their internal system. We received a sample of the complete dataset consisting of six original files in CSV format. Notably, provided dataset included the data for metal kits. Each record in the dataset corresponds to a specific event (task) performed on a kit at a particular sterilization station by a designated operator. 18 Below is a detailed list showing the original data column names, their English translations, and explanations for each entity: •Fecha de seguimiento (Monitoring Date): The date when the kit was monitored during the sterilization process. •Hora de seguimiento (Tracking Time): The specific time when the kit was tracked at a particular process stage. •Usuario (User): The operator responsible for handling the kit at a specific process stage. •Nombre punto de control (Checkpoint Name): The name of the activity (station) where the kit is being processed. •Es preliminar (Parameter of Preliminary): Indicates whether the process is in a preliminary stage or has priority. •Es fisico (Parameter of Being Physical): Indicates whether the process involves a physical handling stage (e.g., manual washing). •Es denegado (Parameter of Denial): Indicates whether the kit was denied at a certain checkpoint due to issues. •Cant. (Quantity): The number of medical devices in the kit processed at a particular stage. •Tipo de objeto (Type of a Kit): The type or category of the kit being processed (e.g., metal or plastic). •Codigo (Kit Code): A unique code assigned to each kit for identification. •N/S (Kit Sequence Number): The sequence number of the kit, used to track its order in the process. •Nombre/Descripcion (Description of a Kit): A detailed description of the kit, including its contents and purpose. •Tipo de produccion (Production Type): The type of production process the kit is associated with. •Nombre produccion (Production Name): The name of the specific production process. 19 •Informacion adicional (Additional Information): Any additional information relevant to the kit or process stage. In AUTO-TWIN project, these six initial data files were preprocessed into nine files. Each of the nine files contains data for a specific type of work performed at its corresponding Nombre punto de control (Checkpoint Name). Table 2.1 below provides information about the dataset files, the amount of events at each activity (station) and their associated process stages, as introduced in Section 2.2.3. Table 2.1: Number of Events by Process Activity Name and Associated Process Stages Activity Name Number of Events Process Stage Entrada Material Sucio 13,537 P1 Cargado en carro L+D 19,692 P2 Carga L+D iniciada 24,738 P3 Carga L+D liberada 23,617 P3 Montaje 24,171 P4 Producci´on montada 37,931 P4 Composici´on de cargas 37,813 P8 Carga de esterilizador liberada 37,199 P9 Comisionado 22,958 P11 Total 241,565 - 2.3.1 Initial Data Understanding: Data Issues It is important to mention that we detected several issues with the data that we consider for further processing. These issues include: 1) Duplicated Events: The analysis revealed a significant presence of duplicates within the provided event tables with entities. Table 2.2 illustrates the frequency of duplicates linked to each input file. Notably, these duplicates constitute approximately 10% of the entire dataset. It is important to note that we consider duplicates as errors, and we will not include them while loading the data into EKG presented in the Section 5.1.1 2) Kits Disappearing After a Specific Sterilization Step: Initially, we detected an anomaly where a significant portion of kit codes (based 20 Table 2.2: Number of Duplicates per File Activity Number of Duplicates Entrada Material Sucio 512 Cargado en carro L+D 4,349 Carga L+D iniciada 9,617 Carga L+D liberada 9,319 Montaje 1,270 Producci´on montada 0 Composici´on de cargas 10 Carga de esterilizador liberada 21738 Comisionado 71 Total 22,637 on the C´odigo attribute from the initial data) appeared only during a specific sterilization activity before disappearing, affecting 202 out of 1183 kits. Here we analysed codigo values that have up to 4 unique activities. However, after introducing the KitID in the analysis in the Section 5.1.1, this issue was largely resolved. Only a small portion of these kits that exhibited this disappearing behavior were not considered when loading the data into the Event Knowledge Graph (EKG). The construction of EKG is presented in the Section 5.1 2.3.2 Initial Data Understanding: Resource Perspective From the Resource perspective we mentioned earlier that the process involves 20 unique human resources involved in the sterilization process at Company C. Resources are responsible for scanning barcodes of kits on sterilization stations and ensuring compliance with FIFO principle. FIFO in this case is related to kits that must be sterilized in the order they arrive. We provide additional insights regarding the resources below. Distribution of Work We decided to verify the amount of work per resource from the data to better understand the workload distribution across various resources. The bar chart on the Figure 2.3 visualizes this distribution, with the x-axis representing the resources and the y-axis indicating the number of rows (events) associated with each resource. The analysis reveals that the resource labeled ”ER” handles the most significant volume of tasks, with a total of 22,032 rows. This is followed 21 Figure 2.3: Distribution of Events (rows) for Each Resource by ”MCE” and ”SM,” which manage 20,914 and 18,845 rows respectively. These three resources clearly carry the most of the workload, highlighting their important roles within the process. Resources such as ”PN,” ”EH,” and ”PG,” with task counts ranging from 16,000 to 17,500 rows, represent a middle tier in terms of workload. These resources appear to have a more balanced distribution of tasks, suggesting that they may operate in a collaborative manner within the process. At the lower end of the spectrum, resources like ”SP,” ”MAA,” and particularly ”DF” and ”MGP” have significantly fewer tasks, with ”MGP” managing only 10 rows. This sharp contrast could indicate that these resources are involved in a limited number of activities or underrepresented in the provided portion of the data. In summary, the chart shows a high disparity in workload distribution among the resources. Batch Processing Pattern In the initial examination of the data from a resource perspective, we identified a distinct pattern known as batching processing. This pattern occurs when a resource, represented by the ”Usarion” (User) attribute, performs several similar tasks consecutively. In the dataset, this was evident as sequences of rows where both the ”Usarion” (User) and ”Nombre punto de control” (Checkpoint Name) attributes—indicating the resource and the specific activity they were engaged in—remained consistent across multiple entries. Despite the consistency in user and activity, the kit attributes in these rows varied, signifying that the resource was processing 22 different kits in the same activity. This observation of batching behavior is important because it reveals how resources manage their tasks by grouping similar activities together. Such behavior can have significant implications for understanding working behaviour, resource allocation, and the overall workflow within the sterilization process. Recognizing and analyzing these batching patterns can help us analysis resource workflows. 2.4 Problem statements from the business perspective In the previous section 2.3, we discussed the FIFO principle in Company C, where resources follow FIFO for kits. FIFO from kit perspective in Comapny C means that the first kits to enter a process are also the first to exit. This principle regulate the flow of medical kits through various stages of sterilization, ensuring that older kits are processed before newer ones. This helps maintain the integrity and orderliness of the sterilization process, reducing the risk of outdated or expired kits being used. However, FIFO is implemented only on a task-specific basis and does not apply to the management of human resources. Although kits are handled sequentially, the organization of employee tasks does not follow FIFO. For example, workers might adhere to FIFO in specific tasks like unpacking or pre-washing, but this system does not necessarily influence how they move between different tasks. This inconsistency presents challenges in understanding employee task prioritization and movement, which is crucial for operational insights at Company C.This creates a complex environment that the company currently lacks insight into. For the digital twin simulation [19], which seeks to accurately replicate real-world conditions, it is important to model the decision-making process of workers. Therefore, the simulation model must understand the criteria by which workers decide to remain at a station or switch to another. Understanding and integrating this working behavior and patterns is key to accurately mirror real-world scenarios and enhance the model’s effectiveness. 23 (DF KIT, DF RESOURCE, DF RUN) distinguished by three different colors: pink, blue, and orange, corresponding to the associated entities Kit, Resource, and Run, respectively. In this representation, the Run entity serves as a case. Event nodes are depicted in a light brown color. In this specific EKG, correlation relationships (CORR) connect six Resource entity nodes (PG, EH, VS, ER, PN) displayed in light blue with event nodes, and event nodes connect with two entity nodes in pink labeled as both Run and Kit. The upper entity node corresponds to the Kit HNB-NC.012-1.0 with Run HNB-NC.012-1.0-2, while the lower entity node corresponds to the Kit HUBU-TR.002-1.0 and Run HUBU-TR.002-1.0-9. Creating Event Knowledge Graphs (EKGs) involves importing event table records as nodes, inferring entities and correlations from event attributes, and ordering events by time for each entity to add Directly-Follows relations that show the sequence of events. This process can be automated using a metadata file to describe entities and their relationships [27]. In the AUTO-TWIN project for Company C, this automated approach, based on Object-Centric Event Data Models in Event Knowledge Graphs, is used to transform legacy data into OCED (Object-Centric Event Data) compliant graphs. This ensures that the data structure meets the standards required for comprehensive process mining analysis. By leveraging this methodology, the AUTO-TWIN project integrates complex, multi-entity event data into a unified framework, supporting analysis and the development of CROMA’s digital twin. 3.4 Resource-Oriented Process Mining Understanding and modeling human behavior within process mining is a challenging task due to the complexity and variability of human actions [28]. This complexity is particularly evident in resource-oriented process mining, where the goal is to analyze how resources, particularly human actors, interact with and influence processes. In this thesis, we attempt to address these challenges by leveraging EKG within multi-dimensional process mining domain. 3.4.1 Batching In Section 2.3, we identified the presence of a batching processing pattern in the dataset, where resources performed similar tasks consecutively. This observation is consistent with the broader concept of batching in resource management, where tasks are grouped together to enhance efficiency. Batching is a well-established concept, and it has been discussed extensively in 30 Figure 3.1: Event Knowledge Graph 31 the literature. For example, van der Aalst introduced the Piled Execution pattern in 2005, which closely aligns with the idea of batching [23]. Recognizing and understanding these patterns is crucial for the thesis, as it allows us to analyze how tasks are systematically organized and executed by human resources within the sterilization process. Further, studies such as those by Wen et al., which focus on mining batch processing workflow models from event logs [32], and by Martin et al., which define batch processing and explore its identification in event logs [18], have significantly contributed to the understanding of batch processes. Moreover, while Klijn, Mannhardt, and Fahland introduced a method for aggregating and analyzing tasks using EKGs to reveal patterns in resource behavior [15, 16]. This approach we use in our thesis from batch processing perspective and for revealing on the base of it resource working patterns. The paper initially defines Task Instances, further in groups it into clusters to create high-level events and subsequently analyse. After applying the framework in this thesis analysis shows that the task instance analysis approach, while useful, does not fully capture the workflow of resources in the context of the sterilization process at Company C. The method focuses on tasks within a single case, but our study requires a broader perspective to understand how resources work across multiple cases throughout the day. Consequently, we propose a method that conceptualizes batching differently, focusing on resource behavior rather than solely on task instances. Despite these differences, we do utilize some aspects of their methodology, particularly for defining aggregated relationships such as Node aggregation and Directly-Follows Aggregation. Definition 10. (Node aggregation): The first basic aggregation query Aggnodes(a, X′, ℓ, ℓ′)on an EKG G= (X, Y, Λ,#) proposed in [16] aggregates nodes X′⊆Xby property ainto concept ℓas follows: 1. query all values V={x.a |x∈X′}, 2. for each value v∈Vadd a new node xv∈ℓto Gwith label ℓand set xv.id =v,xv.type =a, 3. for each x∈X′, add new relationship y∈ℓ′with label ℓ′from xto xv, −→ y= (x, xv). Definition 11. (Directly-Follows Aggregation:) The query Aggdf(t, ℓ, ℓ′) aggregates (or lifts) directly-follows (df) relationships between Event nodes for a particular entity type tto nodes in ℓ′along the ℓ′relationships as follows: 32 1. For any two nodes x,x′in ℓ, query the set dft x,x′of all df-edges (e, e′)t∈ df where events e, e′∈Event are related to x, x′via y, y′∈ℓ′,−→ y= (x, e),−→ y′= (x′, e′). 2. If dft x,x′=∅, create a new df-relationship y∗∈df,−→ y∗= (x, x′), set y∗.type =tand set y∗.count =|dft x,x′|. The variant Agg= df (t,ℓ,ℓ′)of the above query that requires x=x′was proposed in [15]. Our approach builds on these studies and their definitions by applying batching within Event Knowledge Graph and introducing time constraints making it a novel approach. This allows us to study how tasks are batched over time, providing a clearer view of resource allocation in the sterilization process. 3.4.2 Human Resource Working Patterns Another key area of focus in resource-oriented process mining is the analysis of human behavior patterns. Goel et al. discuss various patterns in the work. Two patterns highlighted in this paper are particularly relevant to our research: collaboration and utilization [11]. The collaboration pattern involves analyzing frequent sequence patterns to identify tasks often performed collaboratively by teams. Additionally, the utilization pattern, which focuses on workload-based task assignment, is central to the thesis resource analysis. Importantly, these patterns might be additionally influenced by working shifts and weekdays, suggesting that resource behavior could vary significantly depending on the time of week and shift schedules [33, 17]. In the thesis study, we analyse collaboration pattern to examine how resources work together, identifying frequent sequences that indicate collaborative efforts withing shifts and weekdays. Along with that we study utilization pattern to understand how workload influences task assignment and whether processing efficiency is dependent on workload levels. 3.5 Analytical Techniques in Process and DataDriven Research: Correlation and Sequence Analysis This section covers correlation analysis, normalization, and frequent sequence analysis because we use these methods in the behavioral analysis in this 33 thesis. These techniques are important for this thesis research on resource behavior, and we refer to them to their key studies. 3.5.1 Normalization and Correlation Analysis Normalization helps us compare different variables by placing them on a common scale. In this research, we use Percentage of Total Normalization, which scales each value as a proportion of the total sum of all values. The formula for this normalization is: x′ i=xi Px where xiis the original value, and Pxis the total sum of all values. In our approach, normalization applied buffer values and the high-level event log to ensure that the data is on a consistent scale. This step is important before performing correlation analysis, as if we can do correct comparisons between different time bins and activities. Once the data is normalized, we apply correlation analysis to explore the relationships between various process attributes. Specifically, we investigate how the workload at specific stations relates to resource performance. This relationship is measured using the Pearson correlation coefficient: r=P(Xi−X)(Yi−Y) qP(Xi−X)2P(Yi−Y)2 where Xiand Yiare the data points, and Xand Yare their means. Similarly, van der Aalst et al. [29] demonstrated the value of correlation analysis by uncovering previously unrecognized dependencies between different types of events in business process models. This technique is particularly useful for understanding how changes in one part of a process can affect outcomes in another. For instance in this thesis research, changing workload on processing of kits. 3.5.2 Frequent Sequence Analysis Frequent sequence analysis identifies recurring patterns in datasets, often used in process mining to discover common workflows and predict events. Introduced by Agrawal and Srikant [3], and expanded by van der Aalst [2], this method reveals key process paths. In our study, we use the FP-Growth algorithm [12], which mines frequent sequences of high-level batch activities among resources in a particular day of the week and a shift. We formally define frequent sequence and support: 34 Definition 12. (Frequent Sequence): A sequence s=⟨a1, a2, . . . , ak⟩is frequent if it meets a predefined support threshold. Definition 13. (Support): The support of sis the fraction of sequences in the dataset containing s. Using FP-Growth, we identify common high-level batch activity sequences in the high-level event log. 35 Chapter 4 Problem Exposition In the previous chapter 2, we described the sterilization process and its business implications within Company C, and also provided a theoretical scientific base in the Chapter 3. Building on this foundation, we define the specific research questions for this thesis in Section 4.1. We then aim to identify working behavior and patterns in employee work using the research methodology described in Section 4.2. Our main objective is to develop a framework that aggregates operational data into high-level features based on the timeline of a resource. This framework aims to capture user actions, providing a detailed perspective on resource activities over time. These high-level features, centered on specific employee behaviors and patterns, are important for constructing a digital twin designed to provide predictions and improve operational efficiency. 4.1 Detailed Research Questions Given the primary goal, we define the following research question and associated subquestions: Main research question: What are the identifiable patterns in employee work habits and task prioritization within Company C’s sterilization process, and how can these patterns define the features provided for a digital twin that enhances operational efficiency and compliance? Research subquestions: 1. Can the task instances implementation methodology from Klijn, Mannhardt, and Fahland work [15, 16] be effectively applied to present batching processing in the sterilization process and help in identifying and analyzing the working patterns of resources within Company C? 36 2. What additional batching detection techniques can be used to aggregate resource workflows, aiding in the identification and analysis of resource behavior in the sterilization process? 3. Are current batching techniques adequate for capturing the complex work behaviors of resources, or is there a need for more advanced aggregation methods to accurately identify and analyze work patterns? 4. How do employees prioritize tasks within the sterilization process, what factors influence their prioritization decisions? 5. Identify working patterns that could serve as indicators for features that explain work for further modeling in a digital twin? 4.2 Detailed Method In this section we discuss the reseach mothodoly that was used in this thesis research. Figure 4.1: Thesis Research Methodology Outline 37 We provide an outline of this methodology in Figure 4.1 to generalize the approach into major steps. This figure captures the key phases of our research, each of which contributes to the overall analysis. We employed the following methodology to understand the behavior and working patterns of resources in the sterilization process. 1. Data Understanding and Business Understanding The first step involved gaining a comprehensive understanding of both the data and the business processes involved of the company C. (a) Review Documentation: We examined existing documentation related to Company C’s sterilization process to gather contextual knowledge and identify key variables and metrics. This review provided insights into the operational procedures. Additionally, we communicated with Company C to understand the business problem and their perspective on the goal for this thesis. (b) Exploratory Data Analysis (EDA): We conducted an exploratory data analysis to understand the structure, content, and quality of the available data. During EDA, we identified duplicated data and kits that disappears from processing after a particular sterilization stage provided in the Section 2.3. 2. Data Cleaning and Preparation In the data cleaning and preparation step, the focus was on ensuring the data was clean, and data issues in the Section 2.3.2 are resolved. This included removing duplicated data in the Section 5.1.1. Additionally we introduced KitId attribute that helps to resolve issue with kits that are presented only on particular steps. 3. Graph Implementation, Statistics and Analysis Using the cleaned data, we proceeded with graph implementation and resource behavior analysis in the Section 5.1. (a) Running Graph Implementation: We ran graph the provided by the AUTO-TWIN project with adaptations. This step allowed us to visualize the process and understand the interactions and flow within the sterilization process. (b) Graph Implementation Statistics: We provided implementation statistics regarding the graph nodes, entities and edge in the Section 5.1.2 38 (c) Extended Resource Analysis Utilizing EKG: Using insights from the graph analysis, we examined resource shifts and amount of people working on weekends and weekdays in the Section 5.2. (d) Performing Task Instance Analysis: We tried to capture detected batching behaviour by using Task Instance approach introduced by Klijn, Mannhardt, and Fahland [15, 16] and reveal patterns in resource behavior in the Section 5.3. 4. Implementation and Verification of Batching: Due to insufficient results of implemented Task Instance Analysis, we proposed and implemented two batch modeling techniques in EKG in the Section 6.1: batching over resource in the Section 6.2 and batching over process activity 6.3, based on the earlier identified batching pattern. This involved modeling the batches within the sterilization process to gain detailed insights into how resources were utilized and managed. (a) Batching Techniques Analysis: We examined the two batching modelings techniques used in the sterilization process to understand their characteristics and implications. Batching over resource detects batches handled by a single user, focusing on how individual employees manage multiple kits. On the other hand, batching over station identifies batches processed with the same process activity, which can involve multiple users working simultaneously on a particular process activity. We implemented these batching models, analyzed their performance, compared the results, and selected the best approach, which was batching over station. (b) Summarizing Resource Behavior Using Batching Techniques: Batching over activity technique helped to explain and summarize resource behavior differently than on the level of individual events. However, we discovered complex workflows where there are presented overlapping batches that occur simultaneously, with resources alternating between multiple batches executed in parallel. The presence of overlapping batches motivated to propose a high-level batch aggregation method to capture these behaviors. 5. Implementation and Verification of High-Level Batching: After the initial batching, we applied a high-level batching technique to further aggregate the data presented in the Chapter 7. The high-level aggregation considered regular, iterative, and complex iterative resource 39 Figure 5.2: Distribution of Latest Event Timestamps of Resources as shifts. These shifts are likely designed to ensure continuous coverage and maintain optimal productivity throughout the day, with an average interval of 2 to 2.5 hours between each shift. In our analysis, we observe that on weekends, there are primarily 2 shifts, while on weekdays, the number of shifts ranges between 2 and 4. Based on the Figure 5.1 and Figure 5.2 these shifts are structured as follows: •00:00 to 09:30: The early morning shift typically handles initial operations, setting the pace for the day’s activities. •09:30 to 11:30: The mid-morning shift often experiences an increase in activity as it overlaps with peak operational hours. •11:30 to 14:00: The late morning to early afternoon shift continues managing the workload, ensuring continuity and addressing any backlogs from earlier in the day. •14:00 to 23:59: The afternoon to late evening shift covers the remainder of the day, ensuring that operations are sustained until the end of the business day. This analysis of shifts, based on actual start and end timestamps of event nodes correlated to resources, helps us understand how resources are allocated across different parts of the day. The information gathered is also be valuable for development of Company C digital twin and the further in depth high-level resource working patterns analysis. 46 5.3 Applying Task Instance Framework over Company C’s EKG We further proceed to the Task Instances Analysis. In Chapter 3.4.2, we discussed the intention to apply the Task Instance Framework to aggregate EKG to analyse batching processing from resource perspective. This approach is based on the methodology introduced by Klijn, Mannhardt, and Fahland in their study on aggregating event knowledge graphs for task analysis [15, 16]. We aimed to apply this framework on Company C’s sterilization process, using the EKG to gain insights into batching activities and resource behavior within the process. 5.3.1 Limitations of the Approach We encountered challenges in directly applying the Task Instance Framework to the EKG of Company C’s sterilization process. The original framework, while effective in certain contexts, did not fully align with the complexities and unique characteristics of the data we were working with. To address these challenges, we adapted the framework for EKG of Company C. However, we can notice significant limitation for our data. The primary limitation of applying the Task Instance Framework to the sterilization process at Company C lies in its inherent focus on task instances within a single process case, particularly a kit case. This framework is designed to emphasize the process flow related to individual kits, which can be useful in understanding how resources interact with specific cases. However, this approach tends to narrow the analysis to the kit process itself, rather than providing a view of the broader working patterns and behaviors of resources throughout the day. Our implementation revealed that applying the Task Instance Framework to Company C’s process did not provide the expected insights. A key issue was the limited number of process stages in the sterilization process, which led to the aggregation of several stages into a single high-level event. This single event, treated the same as normal events in the graph, did not capture the detailed behaviors of resources. This limitation was particularly evident for resources with lower levels of activity, like ”DF” and ”MGP”, whose contributions were not sufficiently represented, making it difficult to discern their specific behaviors and interactions within the process. These challenges underscore a critical need for an alternative approach that shifts the focus from process cases to a more comprehensive analysis of resource behavior across various tasks and activities throughout the workday. 47 5.3.2 Need for Batch Detection Revision Given the limitations of the Task Instance Framework, our analysis pointed towards the importance of identifying and understanding batching behaviors within the resource workflows. Batching, where resources perform several similar tasks in sequence. In the next chapter, we propose two batching techniques that can fit better for the data. 48 Chapter 6 Batching The goal of this thesis is to understand the working patterns and behavior of resources within Company C’s sterilization process. The Event Knowledge Graph and Implemented Task Instances Analysis Framework do not provide sufficient insight into these patterns and behaviors. Since, we identified in the previous Section 5.3 that resources often work in batches. This is a challenge, thus we believe that representing and summarizing resource work through batching is an improved approach compared to the existing representation. To address this issue, we need to establish batching modeling techniques. In this thesis, we propose two techniques: batching over resource and batching over station. These techniques will help us summarize the resource workflow within the sterilization process. Batching in the thesis context refers to the grouping of event nodes that are correlated to certain entity nodes, such as the same resource or activity, and occur within a defined time window. We begin by examining the general concept of batching in Company C’s EKG and proposing different modeling techniques in Section 6.1. We first formalize the Batching over Resource technique in the Section 6.2. Following this, in Section 6.3, we formalize the second technique, Batching over Activity. Finally, in Section 6.3.3 we discuss and compare these approaches, and identify the most suitable one for Comapny C. 6.1 Batching in the Event Knowledge Graph: Proposing Options for Company C In order to develop a correct batching detection approach for Company C, it is required to address the limitations identified in the methodology presented in Section 5.3 and establish a method for conceptualizing batching. The task instance analysis based on the work by Klijn, E.L., Mannhardt, 49 F., and Fahland, D. [16] specialises on the task instances creation for resources within one particular process case, specifically for the Company C this is a kit case. This means the analysis focuses more on the kit process flow rather than capturing the real behavior (workflow) of the resources within a day. Consequently, it fails to reveal the patterns where resources perform events in groups. In contrast, we need an approach that focuses on resource behavior, which is the main difference from Eva Kleijn’s work. To overcome these limitations, we need to discuss what conceptualize a batch processing. Batching is a process in which similar tasks or items are grouped together and processed as a single unit [31]. From the Company C’ EKG perspective the idea of a batch involves presence of Event nodes with similar node attributes and correlated to the common entity. For batching in the Company C’ EKG we can consider Resource, Activity and Kit Entities because they are main components of the process that interact frequently and impact the workflow. However, it is important to evaluate each entity: •Resource: We can conceptualize batching by examining how individual employees handle same type of work specifically multiple kits over time on the same activity. This approach can provide insights into resource utilization, their generalized workflow and task prioritization. •Activity: By considering an activity as an entity, we can identify batches of events performed at the same process stage. This method allows us to detect multiple events occurring simultaneously or in close succession at the same process stage (station) regardless the resource. •Kit: Considering kits alone does not constitute batching since each kit is processed individually. Thus, we consider defining batching using resources and activity entities. In this context, we define a batch as a group of events that occur within a specific time window, based on the average arrival time of kits at the initial station ”Entrada de Material Sucio” where kits start the process. We analyzed the arrival time and detected that the average arrival time is 5 minutes. A batch considers a singular event and multiple events. Option 1: Batching over Resource This option groups events by the resource performing the activities, based on a sequence of event nodes along the directly-follows edges of the Resource (DF RESOURCE). The logic for batching in this approach is defined by three primary criteria: 50 1. Execution by the Same Resource: Events must be executed by the same resource to be considered part of the same batch. In the EKG concept, this means that we group into one batch event nodes that have a correlation relation to the same resource node. This ensures that the grouping reflects the specific tasks being performed by a single resource. 2. Performance of the Same Activity: Events must involve the same activity to be included in the same batch. This criterion ensures that the events are related to the same type of work or operation. In the EKG concept, all event nodes joined into a batch must have an equivalent activity attribute. 3. Temporal Proximity: Grouped events must occur in consecutive order within a 5-minute window based on their event timestamps. This ensures that the events are temporally close to each other, reflecting a continuous workflow by the same resource. In the EKG concept, the difference between the timestamp attributes of subsequent events along the directly-follows edges of the Resource should be within 5 minutes. Figure 6.1: Example of Batching over Resource in the Event Knowledge Graph 51 Provided Figure 6.1 figure demonstrates the Batching by Resource option within the EKG. It showcases a series of green event nodes, all labeled ”Carga L+D iniciada,” which indicates that each of these events involves the same activity type. These nodes are connected through directly-follow relationships ’DF RESOURCE’, reflecting the sequence of operations performed by the same resource ’SP’. For the resource we have a blue node and pink nodes for kits are processed. This visualization captures a specific instance from the date 12.01.2022 at 14:36, illustrating continuous operational flow of tasks that are temporally close within a 5-minute window. Each node’s transition from one to the next without significant time lapses fulfills the criteria for batching based on execution by the same resource, performance of the same activity, and temporal proximity. This setup shows how we can group event nodes by resource in the EKG for Batching over Resource approach. Option 2: Batching over Activity Another option is to group events based on the Activity Entity. The criteria for batching include: 1. Equivalent Activity: Events must involve the same activity to be included in the same batch. This ensures that the grouping is meaningful and reflects the specific tasks being performed at the station. From the EKG perspective, all event nodes joined into a batch must have an equivalent activity attribute, since we additionally store the Name attributes of Activity Nodes as an attribute of Event nodes. 2. Temporal Proximity: Grouped events must be ordered by their timestamps attributes, with the time difference between consecutive events being less than 5 minutes. This criterion ensures that the events are temporally close to each other, reflecting a continuous workflow. To illustrate Batching over Activity in the EKG we provide following Figure 6.2. This figure displays several green event nodes, each labeled ”Cargado en carro L+D”. These nodes are linked by observed edges to an Activity node with the same label, indicating that similar tasks are performed at this stage of the process. The visualization shows these event nodes with timestamps from 10:48:00 to 10:49:00 on 2022-03-15, highlighting the parallel execution of the same activity by several resources. Each event node is connected to blue resource 52 nodes and pink kit nodes through correlation relationships, demonstrating the integration of resources and kits a within this operational process stage. This arrangement demonstrates parallel task execution, serving as a clear example of how we can groups similar event nodes in the EKG. Figure 6.2: Example of Batching over Activity in the Event Knowledge Graph In the following sections, we formalize each of the approaches, providing detailed implementation steps. 6.2 Batching over Resource In the previous section, we introduced the batching over resource approach. The method focuses on groups of events based on the resources performing them, their temporal proximity, and activity. It is important to note that a batch can consist of a single event or multiple events satisfying the criteria above. 53 In the current section we formalized this method by presenting its mathematical definition and implementation steps. Additionally, we provide provide an example of batching on the Figure 6.3. Figure 6.3: Example of Batching over Resource In addition to the visual representation of the Batching over Resource approach we provide BatchInstance nodes attributes description in the Table 6.3 Based on the Figure 6.3 and Table 6.3 we can observe that green nodes represent event nodes, pink nodes represent nodes of kits, the blue node represents the resource node, and the light brown nodes represent the BatchInstance nodes themselves. The event nodes are connected with DF RESOURCE edges, highlighting the sequence of activities performed by the resource. Additionally, we see that batches are connected with DF BATCH RESOURCE and DF BATCH KIT edges, indicating the relationships between different batches performed by the same resource over kits. This example illustrates how events are grouped into batches and demonstrates the sequence and relationships between different activities and resources. Specifically, we see how the resource ”MCE” performs activities such as ”Montaje” and ”Producci´on montada” within the 5-minute window, processing different kits and demonstrating the workflow continuity for this resource. 54 Table 6.1: BatchInstances Attributes for Batching over Resource Approach Visualization BatchInstance 1 Activity Montaje Batch number 14 Earliest timestamp ”2022-01-01T10:01:00Z” Latest timestamp ”2022-01-01T10:03:00Z” Events number 2 Kits [HNB-GIN.014-33.0,HNB-GIN.015-9.0] Kits number 2 Runs [HNB-GIN.014-33.0-0,HNB-GIN.015-9.0-0] Resource sys id MCE BatchInstance 2 Activity Producci´on montada Batch number 15 Earliest timestamp ”2022-01-01T10:05:00Z” Latest timestamp ”2022-01-01T10:05:00Z” Events number 1 Kits [HNB-GIN.015-9.0] Kits number 1 Runs [HNB-GIN.015-9.0-0] Resource sys id MCE 6.2.1 Mathematical Definition of Batching over Resource Definition 14. (Batch over Resource) The mathematical formula for defining batches over resource based on the criteria mentioned in the Section 6.1 is as follows: Batchresource(ei, ej) =          1if resource(ei) = resource(ej) and activity(ei) = activity(ej) and |timestamp(ei)−timestamp(ej)|<5minutes 0otherwise Where: - eiand ejare consecutive events. - resource(e)represents the resource performing the event. - activity(e)represents the activity of the event. - timestamp(e)is the time the event occurred. 55 Figure 6.6: Iterative BatchInstance Activity Sequences Frequency mented batches. The bar chart displays the frequency of various activity sequences, highlighting that certain sequences occur repeatedly within a short period. For instance, sequences such as ”Montaje ->Producci´on montada -> Montaje” and ”Producci´on montada ->Montaje ->Producci´on montada” appear with high frequency. This pattern indicates that resources often return to their initial activity after completing another, within a 5-minute window. This not only complicates the tracking of resource activities but also undermines the potential benefits of batching by failing to capture the continuity of resource usage. Concurrent Resource Utilization Another significant limitation of the current approach is its inability to accurately represent scenarios where multiple resources work together on the same activity within the same time window. An example of concurrent resource utilization presented on the Figure 6.7. The Figure 6.7 presents a scenario where multiple resources, represented 62 Figure 6.7: An example of concurrent resource utilization by blue nodes (EH, PN, and SP), simultaneously working on the activity ”Montaje,” depicted by the gray node. Each resource is connected to a green event node, which in turn is linked to a light brown BatchInstance node and pink kit node, indicating that while all resources are involved in the same activity on 2022-03-31 at 14:59 and each operates within its own distinct batch. Analysis shows that several resources often perform events on the same activity within a 5-minute timeframe. The existing batching logic fails to account for this collaborative aspect, instead creating separate batches for each resource even when they are effectively working together. To address this, we utilized a query within the EKG to identify BatchInstance nodes that are connected to different Resource nodes yet share an equivalent activity attribute. The query specifically searched for BatchInstance nodes whose activity attributes matched and timestamp attributes occur within a 5-minute window of each other. The visualization on the 63 Figure 6.8 illustrates the results of the query. Figure 6.8: Distribution of Batches where Resources Worked Together within 5-Minute Time Window In this bar chart 6.8, the x-axis represents different activities, and the y-axis shows the number of batchInstances where multiple resources worked together within a 5-minute time window. The chart indicates that activities such as ”Montaje” and ”Producci´on montada” have a significantly higher number of concurrent resource batches, with counts of 8746 and 8676, respectively. This suggests that these activities are more likely to involve collaborative efforts among multiple resources. Other activities, such as ”Cargado en carro L+D” and ”Composici´on de cargas,” also show instances of concurrent resource utilization, though to a lesser extent. The presence of multiple resources working simultaneously on the same activity suggests that the batchInstances created under the current logic may not represent the actual operational scenarios accurately. This limitation is important because it affects the understanding of resource interdependencies and the overall workflow. By not accounting for repetitive resource patterns and the collaborative aspect. These limitations 64 highlight the need for a more refined batching approach that can accurately represent resource workflows. In the next section, we will introduce a new batching approach that addresses some of these limitations and better fits the operational requirements. This new approach aims to improve the resource collaboration aspect. 6.3 Batching over Activity In the previous section 6.2, we explained the batching over resource approach. In the current section we formalize another batching detection technique Bathing over Activity by presenting its mathematical definition and implementation steps. This method focuses on grouping events based on type of activity performed on a process stage (station) and their temporal proximity. Additionally, we provide provide an example of Batching over Activity on the Figure 6.9. In addition to the visual representation of the Batching over Activity approach we provide BatchInstance node attributes description in the Table 6.3 Table 6.3: BatchInstance Attributes for Batching over Activity Approach Visualization BatchInstance 1 Batch number 13974 Activity Montaje Earliest timestamp 2022-03-29T13:18:00Z Latest timestamp 2022-03-29T13:29:00Z Events number 8 Kits [HNB-TR.010-4.0, HUBU-TR.003-5.0, EXTQUI.HARM AZUL, EXT-QUI.012, HUBU-TR.0141.0, DEP-QUI.TRA.STRYK-37.0, HNB-URO.0015.0, HNB-TR.010-3.0] Kits number 8 Runs [HNB-TR.010-4.0-43, HUBU-TR.003-5.0-20, HNBURO.001-5.0-22, HNB-TR.010-3.0-40] Users [MCE, PN, VS, EH, SP] Users number 5 Based on the Figure 6.9 and Table 6.3 we can observe the parallel work of several resources on the same process stage, performing identical types of 65 Figure 6.9: Example of Batching over Activity tasks. In the visualization, green nodes labeled ”Montaje” represent individual event nodes, each signifying an occurrence of the same activity. These nodes are connected with CORR relationships to pink nodes representing kits and blue nodes denoting resources, highlighting the involvement of various kits and resources in the Montaje operations. The DF Sensor edges connecting these event nodes indicate that the kits underwent scanning during the ”Montaje” process, while the CORR edges show the specific kits and 66 resources involved in each event. This BatchInstance captures a group of events at the Montaje station, spanning an 11-minute window from the earliest event timestamp at 202203-29 13:18:00 to the latest at 2022-03-29 13:29:00. The batch comprises 8 events, involves 8 kits, includes 4 runs, and was facilitated by 5 users, demonstrating a batch operation. 6.3.1 Mathematical Definition of Batching over Activity Definition 15. (Batch over Activity) The mathematical formula for defining batches over activity based on the criteria mentioned in the Section 6.1 is as follows: Batchstation(ei, ej) =      1if activity(ei) = activity(ej) and |timestamp(ei)−timestamp(ej)|<5minutes 0otherwise Where: - eiand ejare consecutive events. - activity(e)represents the activity associated with the event. - timestamp(e)is the time at which the event occurred. This formula helps in determining whether two events belong to the same batch based on their activities and the time difference between their occurrences. 6.3.2 Batching over Activity: Implementation Steps The batching over activity process involves several steps: 1. Assigning Batch Attributes: Similarly to the batching over resource approach we need to associate each event node in the EKG to a specific batch. Thus, we introduce a ”batch” attribute for each event node, assigned numeric value that identifies a specific batch of event group. The assignation performed in the following order: (a) Sequential Evaluation: Event nodes are sorted and assessed based on their timestamp attribute to maintain chronological order, alongside the activity attribute, ensuring that each event is associated with one type of work. 67 (b) Grouping Criteria: Each event node is compared with the preceding one to determine if they occur within a 5-minute window, as marked by the timestamp. (c) Batch Attribution: Events that align under the grouping criteria are assigned the same batch number. If the criteria are not met, a new batch is started. This batch number is then recorded as the ’batch’ attribute for each event node. This step insures classifying the event nodes by batch. 2. Creating BatchInstance Nodes and Connecting to Event Nodes: Once the batch attributes are assigned to event nodes, we proceed with creation of BatchInstance nodes, that present these batches in the EKG. Each BatchInstance node represents the events that share the same resource and activity within the 5-minute window. A BatchInstance node for this type of batching includes attributes: (a) Activity – represents the activity (station on which batching occurs) of all associated events of a batch. (b) Earliest timestamp – a timestamp of the first event in the batch, marking the start of the batch. (c) Latest timestamp – a timestamp of the last event in the batch, marking the end of the batch. (d) Events number – the number of events included in the batch, providing a quantitative measure of the batch size. (e) Kits – a list of kits included in the batch, detailing the specific items involved. (f) Kits number – the number of kits included in the batch, indicating the diversity of items processed. (g) Runs – a list of runs associated with kits in the batch, reflecting the operational flow. (h) Users – a list of users who performed the events in the batch, capturing the human resources involved. (i) Users number – the number of users who performed the events in the batch, highlighting the level of collaboration. Example of created BatchInstance node for batching over activity shown on the Figure 6.9 where the node colored in ligh-brown. 68 Following the logic from batching over activity resource, we proceed with connection BatchInstance nodes to Event nodes with correlation relationships (CORR) based on the shared batch number attribute. This edge in EKG ensures that all events included in a batch are associated with the corresponding BatchInstance. Events within the same batch meet the criteria established in the earlier section 6.1. On the Figure 6.9 we can observe green edge CORR that represent established correlated relationship between green Event nodes and a light-brown BatchInstance node. 3. Establishing Relationships: Finally, we create the other required correlation relationships (CORR) between BatchInstance nodes and Resource, and Kit nodes and add directly-follow relationships between BatchInstance nodes: (a) Correlated Relationships: Each BatchInstance node is linked to the corresponding Resource and Kit nodes to accurately represent their interactions within a batch. Specifically, BatchInstance nodes are connected to Resource nodes based on the shared resource’s system ID. It is important to mention that in the batching over activity approach provide a possibly to connect one BatchInstance with several Resource nodes as if we capture simultaneous work of resources on one process stage in a particular time frame. Additionally we establish correlation relationships between BatchInstance nodes and Kit nodes by associating through shared kit identifier. Similarly, we can have multiple correlated relationships from one BatchInstance node to multiple kits. This ensures that all resources and kits involved in the batch are properly represented. In the provided Figure 6.9, we can observe this type of relationship as a green edge named CORR between pink nodes representing Kit nodes, while the blue node representing the Resource node and light-brown BatchInstance nodes. (b) Directly-Follow Relationships: This process includes creating two types of directly-follow relationships DF BATCH KIT and DF BATCH RESOURCE. These relationships are constructed following the methodology of directly-follows aggregation outlined by Klijn, Mannhardt, and Fahland [16] mentioned in the Section 3.4.1 with additional extension. By establishing these directly-follow relationships, we create a detailed view of the workflow within the sterilization process of kits and resource on the batching aggregation level. Details on changes for each of the edge presented 69 below: i. DF BATCH KIT: The relationship has been enhanced by adding attributes kitId and runId to each edge. This adjustment is crucial for identifying which edge corresponds to each kit, especially in scenarios where multiple incoming and outgoing edges exist. By specifying these attributes, we clarify the connection for each kit, ensuring that its path is followed. ii. DF BATCH RESOURCE: The enhancements of this type of relationship focus on refining how we capture the complex nature of resource tasks. These enhancements include the addition of specific attributes to the relationship edges: A. SysId: Including the resource’s identifier to specify which resource performed the activities. B. Count: Counting how many times this edge (relationship) occurs, which helps in understanding the frequency and iterative manner of the work performed by the resource. C. Order: Determining the order of these relationships to capture the sequence of operations on a particular day. This helps in visualizing the daily workflow of the resource. D. Outgoing order: Capturing the sequence of outgoing edges from a source BatchInstance node. This is useful for tracking the multiple transitions a resource might make from a single batch. By specifying these attributes, we clarify the connection for each resource and capture complex workflow, ensuring that a path is followed. We provide additional Figure to illustrate implemented edges 6.10. These relationships for Resources are represented as blue edges DF BATCH RESOURCE named with associated resource system identifier connecting two light-brown BatchInstance nodes, illustrating how batches are sequentially linkedfollowing resource path. Additionally, orange edges DF BATCH KIT for capturing flow of kits between BacthInstances. With this final step, we established the required relationships within the Event Knowledge Graph, allowing us to view a fully aggregated subgraph that illustrates batching over activity. 70 Figure 6.10: Illustration of DF BATCH RESOURCE edge for ”AV” resource In summary, the implementation steps of Batching over Activity detection methodology in this subsection provide an approach for grouping event nodes into BatchInstance nodes and enriching EKG. Initially, it groups event nodes into BatchInstance nodes. It then establishes relationships between the batches and their corresponding events, resources and kits and after within these batches. By linking these components, we create a subgraph within the EKG, enabling visualization of the batching processes, as demonstrated on Figure 6.9 and Figure 6.10. 71 Figure 6.16: Patterns of Resources Returning to Previous BatchInstances the y-axis indicates the frequency with which each sequence occurs. Each sequence is a pattern where a resource performs an initial batch, switches to a different batch, and then returns to the original batch back. The most frequent sequence, ”Producci´on montada ->Montaje ->Producci´on montada” occurs 2,176 times, highlighting that this pattern of returning to the initial batch after a brief switch to another is common. The second most frequent pattern, ”Montaje - >Producci´on montada - >Montaje,” appears 1,481 times, further emphasizing the iterative nature of these activities within the workflow. As we move to less frequent sequences, such as ”Carga L+D iniciada - > Cargado en carro L+D - >Carga L+D iniciada” and others, the frequency drops significantly. Despite being infrequent, these rare sequences—such as those appearing only once or a few times—demonstrate that even uncommon iterative behaviors exist. This suggests that there are instances of variability in how tasks are revisited. This visualization underscores the recurrence of certain workflows and 78 helps identify where resources are frequently returning to previous tasks, which can be important for understanding resource behaviour. Assessing Proposed Batching over Activity Batching over activity in comparison to batching over resource offers several advantages: 1. Reduced BatchInstance and Edges Space: Batching over activity reduces the number of BatchInstance nodes and edges, creating a more streamlined data structure. This makes data storage and retrieval more efficient for further analysis. 2. Improves Batch Size: This approach results in larger, more cohesive batches, capturing more events per batch compared to the batching over resource method. 3. Solved the Issue of Concurrent Resource Utilization: Batching over activity accurately represents scenarios where multiple resources work simultaneously on the same activity. This resolves the limitations of the batching over resource method, ensuring that collaborative work is properly reflected. 4. Enhanced Resource Tracking from Activity Perspective: By focusing on specific activities rather than resource involvement across tasks, this method allows for more precise tracking of resource allocation and performance. 5. Improved Data Clarity: The method’s structured grouping of events reduces data clutter, making it easier to identify patterns and trends, and supports better decision-making. However, while batching over activity provides better batching detection logic, it is still not sufficient for the analysis of resource workflow and working patterns. The current batching approach still falls in fully capturing resource workflows, particularly in addressing the issue of resources returning to previous BatchInstances. This indicates that the current level of data aggregation is insufficient. To overcome this limitation, it is necessary to perform a higher level of aggregation. In the following chapter, we will define and provide the implementation of the high-level aggregation that aims to solve the limitation of Batching over activity and offer a better framework for resource analysis. 79 Chapter 7 High-Level Batching Building on the limitations of the batching over activity approach discussed in Section 6.3, it becomes evident that while this method effectively models collaboration within a batch, it also introduces challenges related to iterative behavior. When using batching over activity, resources switch between tasks, which can cause batches to overlap temporarily. This overlap makes it challenging to track and model what an individual is doing over a period of time accurately. To resolve this issue, we propose high-level aggregation of batches for each resource and a day. The primary goal of high-level aggregation is to address the iterative behavior limitations of the batching over activity approach and to offer a more effective framework for resource analysis by creating a high-level event log. To this end, this chapter first explores the motivation for high-level aggregation in Section 7.1. We then proceed with defining high-level batching and outlining the implementation steps and results of it in the Section 7.2. 7.1 Motivation and Requirements for HighLevel Batch Aggregation The motivation for implementing high-level batching arises from the limitations observed in the batching over activity approach in the Section 6.3. While the activity-based batching method offers a clear and structured grouping of events based on activities and temporal proximity, it falls short in capturing the iterative and repetitive nature of resource workflows. That cause a temporal overlap between resource’s batches. Additionally, it also fails to capture the unique resource input into BatchInstances with several resources. 80 To illustrate these challenges more concretely, consider the following example on the Figure 7.1. Figure 7.1: Example of Overlapping Batches and Iterative Patterns The graph in Figure 7.1 provides an illustrative example of overlapping batches for resource ”BM”, showcasing a segment of BM’s workflow on 02-082022 between 20:59 and 21:32. In this graph, the light brown nodes represent BatchInstances, we numbered them to provide a clear view on the resource’s workflow. Numbering of nodes based determined by the order attribute of the blue DF BATCH RESOURCE edges that are specific to resource ”BM” (edges named BM). This example highlights an iterative pattern between BatchInstance nodes 2 and 3, where resource BM repeatedly switches between BatchInstances of a particular type of work. The detailed information for these nodes is presented in Table 7.1. Resource ”BM” performs in an iterative pattern twice between these two nodes (2 and 3), as indicated by the count attributes of the associated DF BATCH RESOURCE edges. Further, BM exits this iterative pattern by transitioning from node 2 to node 4, demonstrating a shift in workflow. 81 Table 7.1: BatchInstances Attributes for Resource BM’s Iterative Pattern BatchInstance 2 Activity Montaje Batch number 12466 Earliest timestamp ”2022-02-08T21:04:00Z” Latest timestamp ”2022-02-08T21:12:00Z” Events number 4 Kits [EQP-QUI.CP.OPT-1.0, EQP-QUI.OPT-17.0, HNB-CG.003-5.0, HNB-CPL.007-1.0] Kits number 4 Runs [EQP-QUI.CP.OPT-1.0-1, EQP-QUI.OPT-17.0-1, HNB-CG.003-5.0-7, HNB-CPL.007-1.0-0] Users [BM, ER] Users number 2 BatchInstance 3 Activity Producci´on montada Batch number 15300 Earliest timestamp ”2022-02-08T21:06:00Z” Latest timestamp ”2022-02-08T21:12:00Z” Events number 4 Kits [EQP-QUI.OPT-17.0, EQP-QUI.CP.OPT-1.0, HNB-CG.003-5.0, HUBU-ORL.021-3.0] Kits number 4 Runs [EQP-QUI.OPT-17.0-1, EQP-QUI.CP.OPT-1.0-1, HNB-CG.003-5.0-7, HUBU-ORL.021-3.0-2] Users [BM, CLE] Users number 2 Notably, the BatchInstance nodes involved in this iterative pattern also include contributions from other resources, such as ER and CLE. To address these challenges, we propose implementing high-level aggregation. By combining BatchInstances into high-level batches. High-level aggregation also aims to create a sequence of high-level batches for each resource, organized separately within a day, resulting in a unified high-level log that captures the overall workflow of resources within Company C’s sterilization process. 82 7.1.1 Batching Patterns that Lead to High-Level Aggregation The first step in understanding how to correctly perform high-level aggregation is to identify the patterns that lead to this aggregation. These patterns provide a structured approach to how events and activities are grouped within BatchInstances, revealing the iterative behaviour of resource workflows. We define three primary patterns that form the basis for high-level batching: 1. Regular Pattern: A regular pattern describes the behavior of a resource that performs events belonging to one BatchInstance and then moves on to perform events of a subsequent BatchInstance. This pattern reflects a straightforward and linear progression of tasks, providing a clear view of sequential activities handled by a resource. In Figure 7.2, we can observe the regular pattern for resource AV. The resource works sequentially between the activities ”Composici´on de cargas” and ”Producci´on montada,” moving from one BatchInstance to the next without returning to the initial BatchInstance. Figure 7.2: Regular Pattern 2. Iterative Pattern: An iterative pattern describes the behavior of a resource that performs events belonging to one BatchInstance, then moves on to perform events of a subsequent BatchInstance, and finally returns to perform events of the initial BatchInstance. This pattern captures the repetitive nature of some workflows, where a resource frequently revisits previous tasks or BatchInstances. 83 Figure 7.3 demonstrates the iterative pattern for resource MMF. The resource iterates between the activities ”Montaje” and ”Carga L+D liberada,” performing events in one BatchInstance, then in another, and eventually returning to the initial BatchInstance. Figure 7.3: Iterative Pattern 3. Complex Iterative Pattern: A complex iterative pattern describes the behavior of a resource that performs an iterative pattern, then temporarily shifts to handle a different BatchInstance, and subsequently resumes the original iterative pattern. This pattern highlights the complex and multi-faceted nature of resource activities, showing how resources may juggle multiple tasks and revisit them as necessary. Complex iterative pattern considered within a depth of three nodes. Extending this depth further could risk incorporating over 30% of the entire process, given that the sterilization workflow consists of nine distinct stages. In Figure 7.4, we observe resource AV managing multiple BatchInstances by following the DF BATCH RESOURCE edges labeled ”AV.”. The resource shows complex iterative behavior between the activities ”Montaje” and ”Producci´on montada,” temporarily shifting to handle another BatchInstance ”Composicion de cargas” and then resuming the original iterative pattern. This illustrates the resource’s capability to manage complex workflows involving multiple tasks and revisits. 84 Figure 7.4: Complex Iterative Pattern Identifying these patterns is important for further high-level batch aggregation implementation, as they provide method for grouping BatchInstances according to resource workflows. 7.1.2 Requirements for High-level Batching Aggregation To implement high-level batching, two requirements must be met: 1. Eliminating Iterative Behaviors and Temporal Overlaps: Highlevel aggregation must be capable of capturing the patterns defined in the previous Subsection 7.1.1: regular, iterative, and complex iterative 85 patterns considering the temporal overlaps caused by resources switching between batches (BatchInstances). This capability ensures that the aggregation reflects the actual workflow dynamics and provides a basis high-level log creation and further resource analysis 2. Isolating Unique Contributions: This involvement of multiple resources within the same BatchInstance nodes needs to be considered in the construction of a high-level batch structure. To accurately represent each resource’s workflow per day, it is important to isolate the unique contributions of a resource within these BatchInstances before aggregating them into a high-level batch. These requirements are considered for implementation. We present implementation steps in the next section 7.2 7.2 High-Level Batching: Definitions and Implementation Steps In this section, we define high-level batches mathematically and outline the detailed implementation methodology. We also present the implementation results of high-level batching and how they are represented within the Event Knowledge Graph (EKG). 7.2.1 Mathematical Definition of High-Level Batches The mathematical definitions for high-level batches are as follows: Definition 16. (Singular High-Level Batch) A Singular High-Level Batch corresponds to a regular pattern and is defined as: HLBsingular(r, d) = {bi} where rrepresents a particular resource, dis a specific date, and biis a single BatchInstance correlated to the resource on that date. Definition 17. (Joint High-Level Batch) A Joint High-Level Batch corresponds to an iterative pattern involving two BatchInstances and is defined as: HLBjoint(r, d) = {bi, bj} where rrepresents a particular resource, dis a specific date, and bi, bjare two BatchInstances correlated to the resource on that date. 86 This indicates a workflow where the resource revisits a previous task. Definition 18. (Complex Joint High-Level Batch) A Complex Joint High-Level Batch corresponds to a complex iterative pattern involving three BatchInstances and is defined as: HLBcomplex joint(r, d) = {bi, bj, bk} where rrepresents a particular resource, dis a specific date, and bi, bj, bkare three BatchInstances correlated to the resource on that date. This reflects a workflow where the resource revisits tasks, with potential temporary shifts to another task before returning to the original ones. 7.2.2 Implementation of High-Level Batching Aggregation The high-level batching process involves several steps: The primary challenge addressed in this step is associating each BatchInstance node for a resource per day in the Event Knowledge Graph (EKG) with a specific high-level batch. This process involves creating high-level batches and the necessary edges to represent relationships between these batches. 1. Creating High-Level Batch Nodes: (a) Fetching Resource Paths per Day: Begin by retrieving all paths of BatchInstances for a resource in one day. These paths are analyzed based on the order attribute of the DF BATCH RESOURCE edges, which indicates the sequence of BatchInstances. (b) Singular High-Level Batch Creation: For linear paths where a resource transitions from one BatchInstance to another without iterations and demonstrate a regular pattern, created a singular high-level batch node for each of BatchInstance. (c) Joint High-Level Batch Node Creation: If an iterative pattern is identified, where a resource returns to a previous BatchInstance after an intermediate task, merged two BatchInstance nodes into a joint high-level batch. (d) Complex Joint High-Level Batch Creation: For complex patterns where a resource temporarily shifts to handle additional BatchInstances before returning to the original iterative pattern, created a complex joint high-level batch. This high-level batch aggregates three BatchInstances. Each HighLevelBatch node includes the following attributes: 87 Chapter 8 High-Level Resource Behavioral Analysis We have established high-level batches in the Chapter 7 that represent the resource flow, and based on these, we aim to achieve main goal of the thesis by analyzing resource working behavior and patterns. High-level batches provide a comprehensive view of how resources function across the entire workflow, capturing both their aggregated activities and transitions between tasks. The primary method is a detailed analysis of resource behavior and patterns using a series of high-level events constructed from defined high-level batches. These high-level events reflect the daily activities of each resource, enabling a structured examination of their behavior throughout the workday. Thus, this log will allow to examine how resources function over time and identify patterns within their workflows. We begin by forming the high-level event log through high-level batch aggregation, detailed in Section 8.1. Next, in Section 8.2, we analyze resource working behavior using the high-level event log from a workload perspective. Finally, in Section 8.3, we explore working patterns within the high-level event log, using frequent pattern mining to identify common working patterns among resources. 8.1 Forming High-Level Event Log The primary challenge addressed in this section is the formation of a highlevel event log, which is a crucial step before analyzing resource working behavior and patterns. The high-level event log provides a structured representation of resource activities, capturing the aggregated workflow of re94 sources within Company C’s sterilization process. In this thesis research, the high-level event log (HL) is constructed by building a sequence of high-level batches, or high-level traces, for each resource rand each day d. Specifically, for every resource ron a given day d, we construct a high-level trace γ(r, d) as follows: γ(r, d) = ⟨h1, h2, . . . , hn⟩ where each hirepresents a high-level event (or batch) within the trace. These high-level events are derived from the high-level batch aggregation process discussed in Chapter 7. The construction high-level event log represents the collection of each trace γ(r, d) in the data. The process of high-level event log construction involves the following key steps: 1. Defining High-Level Events The first step in the methodology is defining high-level events. As established in the previous chapter, we perform high-level batch aggregation to group related BatchInstances into cohesive high-level batches. Each high-level batch is treated as a distinct high-level event. A highlevel event is characterized by the following attributes: (a) The specific activities it includes. (b) The resource responsible for these activities. (c) The time frame during which these activities occur. Formally, we can define a high-level event hle i as follows: Definition 19. (High-Level Event) hlei=⟨{a1, a2, . . . , an}, r, [tstart, tend]⟩ where: •{a1, a2, . . . , an}represents the set of activities included in the highlevel event. •rrepresents the resource responsible for the high-level event. •[tstart, tend]represents the time frame during which the activities occurred. This formal definition ensures that each high-level event is uniquely identified by its activities, the resource involved, and the time period over which the event spans. 95 2. Extracting Data from the Event Knowledge Graph (EKG) With the high-level events defined, the next step involves extracting the necessary data from the Event Knowledge Graph (EKG). We utilize Neo4j’s Cypher query language to retrieve relevant data, including: (a) Timestamps, activities, resource IDs of the high-level events. (b) Additional attributes such as batch numbers included, the number of kits processed, number of low level events, and whether the high-level event involved collaboration with other resources. This extracted data is then processed into a structured format, ensuring that all dates, timestamps, and other attributes are correctly formatted and ready for further processing. 3. Creating the Complete High-Level Event Log The final step is constructing the high-level event log. The high-level events are sequenced chronologically based on their date and earliest timestamp. This ordered sequence reflects the resources’ workflow over a day, providing a clear timeline of activities. By following these steps, we create a high-level event log that captures resource workflows within the sterilization process. This log is now ready for further analysis, allowing to explore resource behavior and identify patterns. 8.2 Resource Working Behavior Based on HighLevel Batching and Workload In this subsection, we focus to a detailed examination of resource working behavior within Company C’s sterilization process. Our hypothesis, derived from multiple iterations of analysis over sterilization process, batching, and high-level batching, posits that resource behavior is significantly influenced by workload in the sterilization center. Essentially, the hypothesis suggests that resources are more likely to be directed towards areas where work is accumulating. Through this analysis, we seek to answer a key question: How do resources respond to increasing workloads at different stations? To validate this hypothesis and gain understanding of resource working behavior, we first define the workload and perform an in-depth analysis that will identify the correlation between workload and the corresponding actions taken by resources on the high-level aggregation. 96 Understanding how resources work are important for achieving the main objective of this thesis provide high-level resource features for development of Auto-Twin digital twin for company C that accurately mirrors real-world operations. 8.2.1 Defining the Workload Before explaining the approach used for analyzing resource behavior, it is important to define the concept of workload in the context of Company C’s sterilization process. Workload, in this setting, refers to the amount of work accumulated at various activities (process stages), specifically the number of kits waiting to be processed on process stages within a specific time frame. The formal definition of the workload, along with the detailed method for calculating it and analyzing its correlation with high-level resource activities, will be outlined in the following Subsection 8.2.2. 8.2.2 Detailed Approach for Analyzing Resource Working Behavior To conduct an analysis of resource working behavior, we established an approach that utilizes the high-level event log, as defined in Section 8.1, alongside the calculated workload, represented by the number of kits awaiting for processing. This approach involves time-based analysis, enabling us to track the dynamics between workload and the corresponding resource behavior. We utilize a 30-minute time bin in our analysis because it provides an effective balance between capturing detailed insights into resource behavior and keeping the data manageable. This interval is sufficient to observe significant shifts in workload and resource allocation, while still being short enough to accurately reflect changes in the process. We now proceed to the detailed steps of the approach, which are outlined as follows: 1. Calculation of the Workload Utilizing Buffer and Start Kit Values: The first step in our approach involved determining the workload by calculating the buffer and start values. To perform this buffer and start calculation, we utilized the event log provided by Company C. This log, after undergoing a thorough cleaning process, was loaded into the EKG, details of which provided in the Chapter 5. The initial event log allows to track the flow of kits through the sterilization process, enabling monitoring the buffers at each activity (station). By analyzing these flows and transitions, we were able to 97 accurately measure the buffers at each station within 30-minutes time bins, thereby quantifying the workload. The following steps will detail the specific procedures used to calculate these buffers. (a) Initial Event Log Preprocessing: Grouping Events and Sorting within Traces: The first step in the preprocessing involved organizing the initial event log data (L) by grouping low-level events according to their runId. The runId serves as the case identifier for kit traces within the event log, with each runId representing a unique trace corresponding to a specific kit as it progresses through the sterilization process. Definition 20. (Trace) Each trace σis defined as a sequence of events: σ=⟨e1, e2, . . . , en⟩, where (runIdx, activityx, timestampx)∈ex In this context, Ldenotes the set of all traces, Athe set of all activities, and Ethe set of all events. Grouping by runId ensures that we capture the entire sequence of events for each kit. Subsequently, these events were sorted by timestamp to establish a chronological sequence within each trace. This sorting is required for tracking the flow of each kit over time through the sterilization process and for the subsequent calculation of buffers. Time Binning: Following the grouping and sorting of events, the preprocessed event log was divided into 30-minute time bins. (b) Processed Kits Evaluation: We continued with an assessment of kits processed on each process activity within each time bin. Definition 21. (Processed Kits) For each activity Ain a time bin Ti, the value is equal to the number of kits for which there were performed events (e) with activity (activity(e)) following the event log traces (σ): ProcessedKits(a, Ti) = X σ∈L |{e|e∈σ∧activity (e) = a ∧timestamp(e)∈Ti}| ∀a∈A, Ti∈T 98 This calculation is necessary to track the number of kits processed at each station, as it allows us to definitively know what was processed at each activity and with what timestamp based on the initial event log. This step serves as a prerequisite for determining the buffer at each activity. (c) Initializing Starting Traces Kits and Entering Buffer Kits Start Values To ensure the kits tracking that initiate their processing in each time bin, we calculated the start values. Definition 22. (Start Value)The start value represents the number of kits that begin their trace with a specific activity within a given time bin. For each trace (σ), if the first event (e1) in the trace corresponds to a particular activity (a) and falls within the time bin (Ti), the start value for that activity and time bin is incremented. This is formally defined as: Start(a, Ti) = X σ∈L |{e1|e1=σ[1] ∧activity (e1) = a ∧timestamp(e1)∈Ti}| ∀a∈A, Ti∈T This calculation captured the initiation of processes, providing insight into how and where kits start their process flows. Entered Buffer Values In addition to start values, we calculated entered buffer values to identify kits that entered the buffer during each time bin and awaited processing in subsequent events. Definition 23. (Entered Buffer Kit Value)The entered buffer kit value for a specific activity and time bin reflects the accumulation of kits from previous events in the trace. For an event (ej) in a trace (σ), we look at the preceding event (ej−1) to determine its time bin (Ti)and increment the entered buffer value accordingly. The entered buffer value is formally defined as: EnteredBuffer(a, Ti) = X σ∈L |{ej|1< j ≤ |σ| ∧ ej−1=σ[j−1] ∧ej=σ[j]∧activity(ej) = a ∧timestamp(ej−1)∈Ti}| ∀a∈A, Ti∈T 99 Entered buffers were required to evaluate when kit was placed to wait for processing. Addressing the First Event in Traces It is important to note that the first event of a kit’s trace is not included in the EnteredBuffer values for a specific activity, as it lacks a preceding event in the trace and we do not know when kit was placed for awaiting for sterilization on its’ initial trace event. However, the first event is still accounted for in the ProcessedKits for that activity. This distinction could lead to inaccuracies in subsequent buffer calculations. To mitigate this, we introduced the Start value, which captures events that initiate a trace, ensuring tracking of kit flows from the beginning of their processing. (d) Buffers Calculation The buffer values for each activity aand time bin Tiwere calculated to account for kits remaining in the buffer from previous time bins. This calculation ensure that we tracked the flow of kits through the sterilization process, accounting for those that were not processed in earlier time bins: Definition 24. (Buffer) We define buffer as follows: Buffer(a, Ti) =          EnteredBuffer(a, Ti)if i= 1 EnteredBuffer(a, Ti) + Start(a, Ti−1) +Buffer(a, Ti−1)−ProcessedKits(a, Ti−1)otherwise We only updated buffer values for i > 1 as the first bin value was set correctly during the initialization phase. This step is important for understanding the carryover of unprocessed kits, which can affect subsequent time bins’ workload. (e) Workload Calculation The workload for each activity aand time bin Tiwas calculated by summing the buffer values with the start values for the same activity and time bin. This calculation provided a comprehensive measure of the total number of kits awaiting processing within the sterilization process. Definition 25. (Workload) We define workload as follows: Workload(a, Ti) = Buffer(a, Ti) + Start(a, Ti) 100 For the activity Entrada Material Sucio, we only have start values and no buffer values, as it represents the initial stage of the process where kits first enter the sterilization. 2. Prepartion for Correlation Analysis: (a) Normalization of Workload Values To perform the correlation analysis, it was required first to normalize workload values for each time bin to ensure comparability. The workload, which represent the accumulated kits waiting for processing, were computed based on the previous step. Normalization was then applied to the workload values to ensure that the data was on a consistent scale for performing correlation analysis . (b) Preparation of the High-Level Event Log: The high-level event log (HL), which was established in Section 8.1, was further processed to prepare it for correlation analysis. The steps included: i. Aggregating the data to represent the number of kits processed for each activity within each high-level batch. ii. Time-binning the aggregated data, similar to the approach used for the workload values calculation. However, for the high-level event log we considered earliest timestamp attribute to be able to associate high-level event with a particular time bin. iii. Normalizing the data across activities to remove any scale effects and ensure that the correlation analysis could capture the relationships between workload and resource behavior. 3. Performing Correlation Analysis Between workload Values and Prepared High-Level Event Log: Once the workload values and high-level event log were prepared and normalized, we conducted a correlation analysis to examine the relationship between the workload and the resource workflow (captured in the high-level event log). This analysis aimed to identify how variations in workload influenced performing work over kits on the process activities and transitions observed in the high-level event log. 4. Results Evaluation: The results of the correlation analysis were evaluated to determine the significance and patterns of the relationships between the workload values and high-level events. These results are 101 detailed in the next subsection, where we present specific findings and their implications for understanding resource behavior. In conclusion, the detailed approach outlined in this subsection provides a methodology for analyzing resource working behavior by utilizing the highlevel event log and calculated workload. 8.2.3 Results Evaluation of Analyzing Resource Working Behavior After explaining the approach used to analyze resource behavior, we now present the results. The main goal of this analysis is to test our hypothesis that the resource work in Company C’s sterilization process is influenced by the number of kits waiting to be processed at different stations. We aim to understand how the accumulation of kits impacts resource allocation. To do this, we conducted a correlation analysis using data from consistent time bins, focusing on the relationship between workload and the number of kits processed in high-level batches at different activities. The correlation matrix in Figure 8.1 was created using these consistent time bins to maintain accuracy in our analysis. The color gradient shows the strength and direction of the correlations, with dark red indicating strong positive correlations and dark blue indicating strong negative correlations. This correlation matrix helps us understand how workload on specific activities influenced the work performed on those activities. By analyzing the relationships between the workload values (named as Workload activity name) and the processed kits values from the high-level event log (named as kits activity name), we can see which workload levels at certain activities impacted the actual work done. To evaluate the hypothesis that the behavior of resources within Company C’s sterilization process is influenced by the accumulation of kits at various stations awaiting sterilization, we must analyze the correlation matrix with a focus on understanding how workload (buffer and start values) at each station affects the number of kits processed. This analysis allows us to determine whether the workload indeed influences resource behavior and processing efficiency across different stages of the sterilization process. Entrada de Material Sucio (Initial Process Stage): The ”Entrada de Material Sucio” is the first stage in the process, and as such, there are no workload values available at this point since kits have not yet accumulated before processing begins. This is why we only observe the start value in the correlation matrix. The correlation between ”Start Entrada Material Sucio” 102 Figure 8.1: Correlation Matrix for Correlation Analysis Between Workload Values and Processed Kits Based on High-Level Batching Aggregation and ”kits Entrada Material Sucio” is high at 0.70, indicating that the initial inflow of kits directly influences the workload at this stage. This suggests a demand-driven process, where the number of incoming kits closely aligns with the number of kits processed, as expected. Additionally, there is a moderate positive correlation of 0.34 between ”Start Entrada Material Sucio” and ”Workload Carga L+D iniciada”. This suggests that the start of the initial process stage not only affects the immediate workload of processing kits but also has some influence on the subsequent stages, particularly where the kits are being prepared for the next steps (”Carga L+D iniciada”). This indicates that the initial kit inflow could partially predict the workload in subsequent processing stages, demonstrating the interconnected nature of the workload across different stages of the process. 103 a tendency for resources to perform repetitive tasks within a day or shift. (b) The pattern ”Carga L+D liberada (Separate)” appeared frequently across multiple shifts, particularly in Shifts 2 and 4, indicating its critical role in the workflow. Similarly, ”Producci´on montada, Montaje (Together)” emerged as a common collaborative pattern in Shift 1, particularly on Monday, Tuesday, and Wednesday. 2. Shift 1: (a) On Monday, the primary patterns were ”Carga L+D iniciada, Cargado en carro L+D (Together)” with a 19.64% occurrence, ”Montaje (Together)” at 16.07%, and ”Producci´on montada, Montaje (Together)” at 26.79%. These patterns suggest that Shift 1 on Monday is characterized by collaborative work involving multiple resources. (b) Tuesday displayed similar trends, with ”Montaje (Together)” occurring in 36.21% of traces, highlighting the collaborative nature of Shift 1 during this day. The pattern ”Carga de esterilizador liberada (Separate)” was also significant, with 34.48% occurrence, indicating a focus on releasing sterilizers. (c) On Wednesday, the patterns ”Montaje (Together)” and ”Producci´on montada, Montaje (Together)” continued to dominate, with 30.16% and 28.57% occurrence, respectively. This suggests that collaborative activities are consistently important in Shift 1 throughout the week. (d) Thursday and Friday also reflected a high occurrence of collaborative patterns, with ”Producci´on montada, Montaje (Together)” reaching 29.09% on Thursday and 38.00% on Friday, emphasizing the teamwork involved in these tasks. 3. Shift 2: (a) On Monday, the pattern ”Carga L+D liberada (Separate)” was prevalent in 29.41% of traces, reflecting its importance in this shift. Other frequent patterns included ”Carga de esterilizador liberada (Separate)” at 23.53% and ”Composici´on de cargas (Separate)” at 17.65%. (b) Tuesday and Wednesday continued this trend with ”Carga L+D liberada (Separate)” appearing in 31.25% and 41.18% of traces, 110 respectively. This indicates that Shift 2 consistently involves crucial individual tasks related to the handling and releasing of loads. (c) On Thursday, the same pattern was observed with ”Carga L+D liberada (Separate)” appearing in 30.00% of the traces, reinforcing its significance during Shift 2 across multiple days. 4. Shift 3: (a) Shift 3 exhibited lower pattern percentages, indicating a more diverse range of activities with less repetition. For example, on Tuesday, ”Comisionado (Separate)” occurred in 17.65% of traces, and on Friday, ”Producci´on montada, Montaje, Composici´on de cargas (Together)” was present in 17.39% of traces. This suggests that Shift 3 may involve a wider variety of tasks, leading to less frequent repetition of specific patterns. 5. Shift 4: (a) On Monday, the pattern ”Entrada Material Sucio (Separate)” was present in 32.00% of the traces, indicating its critical role during this shift. This pattern continued to be significant on Wednesday with a 41.67% occurrence and on Thursday with 35.56%. (b) Friday also reflected the importance of ”Entrada Material Sucio (Separate)” with a 32.61% occurrence. The consistent presence of this pattern during Shift 4 across multiple days highlights its importance in the workflow, particularly in handling entering the sterilization kits. 6. Saturday and Sunday: (a) Saturday Shift 1 showed a 29% occurrence of the pattern ”Carga L+D iniciada, Cargado en carro L+D (Separate)”, indicating the routine nature of this task on weekends. Shift 4 on Saturday had a higher occurrence of ”Carga L+D iniciada, Cargado en carro L+D (Separate)” at 54% performed by one resource. (b) On Sunday, Shift 1 patterns included ”Carga de esterilizador liberada (Separate)” and ”Composici´on de cargas (Separate)” each at 31%. Shift 4 showed the same pattern with a 50% occurrence, indicating that material handling and load initiation are crucial on Sundays and performed by one resource. 111 Insights Regarding Collaboration within Shifts •Shift 1: This shift often includes combined activities like ”Producci´on montada, Montaje,” reflecting a more collaborative approach. The frequent occurrence of such patterns highlights the teamwork involved in completing these high-level events, particularly on weekdays. •Shifts 2 and 4: These shifts are characterized by a higher prevalence of separate activities, particularly ”Carga L+D liberada” and ”Entrada Material Sucio”. These shifts involve significant individual high-level events handling, reflecting the specialized nature of the tasks performed. •Shift 3: Shows a mix of both separate and combined high-level events sequences, with the smallest values of pattern percentages, indicating high diversity in high-level working behavior. This suggests that Shift 3 may involve more varied tasks, with resources handling different types of activities, leading to less repetition of specific patterns. General Overview of Work Patterns and Expectations The analysis of work patterns across different shifts and weekdays reveals distinct operational behaviors that align with the specific demands of each shift. Shift 1 is predominantly characterized by collaborative activities, especially on weekdays such as Monday, Tuesday, and Wednesday. In approximately 30% to 40% of the traces, resources are engaged in joint tasks, such as ”Producci´on montada, Montaje (Together).” This high level of collaboration underscores the importance of teamwork during Shift 1. In contrast, Shifts 2 and 4are more focused on individual task handling. Patterns such as ”Carga L+D liberada (Separate)” and ”Entrada Material Sucio (Separate)” frequently appear, accounting for about 20% to 35% of the traces. These shifts play a crucial role in maintaining the steady flow of operations by consistently performing routine tasks that require a high degree of autonomy and expertise. The repetition of these patterns across multiple days highlights their importance in the overall workflow. Shift 3 experiences the most diverse work patterns, with no single sequence dominating more than 20% of the traces. This diversity suggests that resources in Shift 3 are likely to encounter a broader range of tasks, leading to a dynamic and varied work environment. Overall, the findings from this analysis strongly support the thesis of High-Level Resource Behavioral Analysis. The results demonstrate that the 112 nature of work in each shift is closely aligned with the specific operational needs of those periods. Shift 1 is characterized by collaborative efforts, while Shifts 2 and 4 are focused on specialized, individual tasks. Shift 3 introduces flexibility into the workflow by accommodating a wide range of activities. These insights not only improve the understanding of how work is distributed across different shifts but also provide an overview of the frequent patterns within shifts and weekdays. This is important for the development of a digital twin of the sterilization process in Company C, as it offers a solid foundation for accurately modeling resource behavior in a dynamic operational environment. 113 Chapter 9 Conclusion This thesis presents an analysis of resource behavior and working patterns within the sterilization process at Company C, utilizing process mining techniques to uncover deep insights into resource operational patterns and task prioritization. This thesis began with an initial exploration of data, highlighting batching processes and shift-based work patterns within Company C’s sterilization process. The research initially aimed to apply the Task Instance Framework, using Company C’ EKG as part of the AUTO-Twin project, for detailed behavioral analysis. However, this framework was found incompatible with the specific needs of the dataset, leading to the development of alternative batch detection techniques. Two methodologies, Batching over Resource and Batching over Activity, were introduced. Batching over Activity was chosen for its effectiveness in defining collaborative work and providing a more streamlined analytical framework. Despite its benefits, this approach alone did not address the repetitive work patterns among resources, leading to the adoption of highlevel aggregation techniques. These techniques categorized work patterns as regular, iterative, or complex iterative, facilitating structured analysis of resource behavior. The research utilized a high-level aggregation log, formed based on highlevel batches of the EKG , to examine the impact of workload—defined as kits awaiting processing—on the processed kits on the sterilization stages comprehensively. For this research we provide the code presented in the GitHub repository that reflects the implementation of this thesis steps. It was found that workload not only directly influences the processing of kits but also impacts subsequent stages, illustrating the interconnected nature of the process workflow. Additionally, frequent sequence mining was GitHub Repository: https://github.com/LiliiaAliakberova/ThesisResearchHighLevelBatchAggregation employed to identify common work patterns across different weekdays, and shifts, uncovering distinct collaborative behaviors and operational nuances within shift structures. The findings from this study provide valuable insights into the dynamics of resource allocation and task execution in complex operational settings, laying the groundwork for the development of effective digital twins. In summary, this thesis not only explores the complexities of analyzing resource behavior within a sterilization context but also provides the way for future advancements in process mining and providing high-level resource features for digital twin implementation. 9.1 Limitations While the study achieved its objectives, several limitations should be acknowledged. One of the main limitations is the reliance on correlation analysis to infer the relationships between workloads and processed kits. Although correlations provide valuable insights, they do not establish causality, and further studies are required to explore these relationships more deeply, possibly through causal analysis or experimental design. Another limitation lies in the scope of high-level aggregation. The current approach limited the depth to three nodes to avoid overgeneralization. However, this restriction may have prevented the capture of more complex iterative behaviors that could be critical for a complete understanding of the process. Future research should consider ways to extend the depth of aggregation without compromising the clarity and precision of the analysis. The approach to batch detection in this thesis primarily focused on specific activities. While this provided detailed insights into those stages, it did not fully account for how batches are formed across multiple activities, particularly those exhibiting iterative behavior. Expanding the detection method to consider multiple activities and refining time constraints could yield a more comprehensive understanding of batching in the process. Lastly, the analysis of resource behavior focused on identifying patterns within individual and group activities, but it did not fully explore how resources collaborate or how their actions during shifts impact overall process efficiency. For the digital twin to more accurately replicate real-world conditions, future research should investigate collaboration patterns among resources to understand how the actions of one resource influence others. Additionally, studying resource behavior during shifts, including the overlaps between shifts, could provide valuable insights into resource allocation and transitions, further enhancing the accuracy and effectiveness of the sim115 ulation model. 9.2 Future Work This thesis has focused on understanding human resource behavior and working patterns in the sterilization process at Company C, with the aim of providing high-level features for the development of a digital twin. While significant progress has been made, there are several areas where further research could enhance the findings, help to solve limitations mentioned in the previous Section 9.1 and contribute to a more accurate and effective digital twin model. Below are some potential directions for future work. Batch Detection Techniques: 1. Exploring Different Batching Detection Methods: Future research could investigate various methods for detecting batches by analyzing the arrival times of kits at different stations. By adapting time constraints for batching to each activity individually, the detection process could become more precise and better aligned with the specific characteristics of each stage. 2. Defining Batching Across Multiple Activities: Another important area for research is defining batching not only within individual activities but also across multiple activities, particularly those that frequently exhibit iterative behavior. This would offer a more comprehensive understanding of how batches and resources with these batches move through the entire sterilization process. High-Level Batching Aggregation 1. Extending the Depth of High-Level Aggregation: The current analysis focused on complex iterative behaviors within a depth of three nodes. Extending this depth could potentially involve more than 30% of the process, considering that the sterilization workflow comprises nine distinct stages. Future research should explore how to extend the depth of high-level aggregation while ensuring it does not exceed 3040% of the total process, to maintain analytical clarity and precision. 2. Developing New Aggregation Methods: Future work could explore creating more compact methods for high-level aggregation that effectively capture complex behaviors while maintaining a clear focus 116 in the analysis. The approach could involve fully abstracting from time to focus solely on the sequence of resource activities. Instead of creating separate high-level nodes for each BatchInstance patterns where a resource performs a batch activity at different times, we could represent each unique batching activity with a single high-level node. The workflow would then be depicted through nodes, with edges capturing the sequence and return of the resource to specific activities. This approach would reduce the space needed for analysis and allow us to concentrate on the sequence of activities, providing a broader view of resource behavior in the process. Resource Working Behavior Analysis 1. Identifying Collaboration Patterns Among Resources: Future research could explore how groups of resources frequently work together. By analyzing the behavior of each resource individually, it could become clearer how the actions of one resource affect others, potentially revealing collaboration patterns that enhance efficiency. 2. Analyzing Resource Behavior During Shifts: Further study could focus on how resources behave within shifts, including the overlaps between shifts. This could offer insights into resource allocation and transitions between shifts. Understanding resource behavior is crucial for developing a digital twin simulation that accurately reflects real-world conditions. The simulation model needs to capture how workers decide whether to stay at a station or move to another. Integrating these decision-making patterns into the digital twin is key to accurately mirroring real-world scenarios and improving the model’s effectiveness. These future research directions could deepen our understanding of batch detection, high-level aggregation, and resource behavior into the sterilization process, providing valuable insights for the development of a more effective digital twin for Company C. 117 Bibliography [1] AUTO-TWIN Project. https://www.auto-twin-project.eu/. Accessed on on 1th July 2024. [2] Wil Aalst. Process Mining: Discovery, Conformance and Enhancement of Business Processes, volume 136. Springer, 01 2011. [3] R. Agrawal and R. Srikant. Mining sequential patterns. In Proceedings of the Eleventh International Conference on Data Engineering, pages 3–14, 1995. [4] Reinhard Diestel. Graph theory. Springer (print edition); Reinhard Diestel (eBooks), 2024. [5] Stefan Esser and Dirk Fahland. Multi-dimensional event data in graph databases. Journal of Data Semantics, 10:109–141, 2021. Received 29 May 2020, Revised 12 November 2020, Accepted 06 March 2021, Published 27 May 2021, Issue Date June 2021. [6] Dirk Fahland. Extracting and pre-processing event logs. CoRR, abs/2211.04338, 2022. [7] Dirk Fahland. Multi-dimensional process analysis. In Claudio Di Ciccio, Remco Dijkman, Adela del R´ıo Ortega, and Stefanie Rinderle-Ma, editors, Business Process Management, pages 27–33, Cham, 2022. Springer International Publishing. [8] Dirk Fahland. Process Mining over Multiple Behavioral Dimensions with Event Knowledge Graphs, pages 274–319. Springer International Publishing, Cham, 2022. [9] Nadime Francis, Alastair Green, Paolo Guagliardo, Leonid Libkin, Tobias Lindaaker, Victor Marsault, Stefan Plantikow, Mats Rydberg, Petra Selmer, and Andr´es Taylor. Cypher: An evolving query language for property graphs. Proceedings of the 2018 International Conference on Management of Data, 2018. 118 [10] Nicla Frigerio, Baris Tan, and Andrea Matta. Green gateways: a concept for decisions in circular-oriented economies. Procedia CIRP, 122:617– 622, 01 2024. [11] Kanika Goel, Tobias Fehrer, Maximilian R¨oglinger, and Moe T. Wynn. Not here, but there: Human resource allocation patterns. In Chiara Di Francescomarino, Andrea Burattin, Christian Janiesch, and Shazia Sadiq, editors, Business Process Management, pages 377–394, Cham, 2023. Springer Nature Switzerland. [12] Jiawei Han, Jian Pei, and Yiwen Yin. Mining frequent patterns without candidate generation. volume 29, pages 1–12, 06 2000. [13] Kongzhang Hao, Zhengyi Yang, Longbin Lai, Zhengmin Lai, Xin Jin, and Xuemin Lin. Patmat: A distributed pattern matching engine with cypher. Proceedings of the 28th ACM International Conference on Information and Knowledge Management, 2019. [14] Olaf Hartig. Foundations to query labeled property graphs using sparql. In SEM4TRA-AMAR@SEMANTiCS, 2019. [15] Eva Klijn, Felix Mannhardt, and Dirk Fahland. Classifying and Detecting Task Executions and Routines in Processes Using Event Graphs, pages 212–229. CEUR-WS.org, 08 2021. [16] Eva L. Klijn, Felix Mannhardt, and Dirk Fahland. Aggregating event knowledge graphs for task analysis. In Marco Montali, Arik Senderovich, and Matthias Weidlich, editors, Process Mining Workshops, pages 493– 505, Cham, 2023. Springer Nature Switzerland. [17] Orlenys L´opez-Pintado, Marlon Dumas, and Jonas Berx. Discovery, simulation, and optimization of business processes with differentiated resources. Information Systems, 120:102289, 2024. [18] Niels Martin, Marijke Swennen, Benoˆıt Depaire, Mieke Jans, An Caris, and Koen Vanhoof. Batch processing: definition and event log identification. 12 2015. [19] Gerardo Santill´an Mart´ınez, Seppo Sierla, Tommi Karhela, and Valeriy Vyatkin. Automatic generation of a simulation-based digital twin of an industrial process plant. In IECON 2018 - 44th Annual Conference of the IEEE Industrial Electronics Society, pages 3084–3089, 2018. [20] Neo4j. Cypher Manual. Accessed on on 16th August 2024. 119 Montaje, Producci´on montada (Together) 9 17 52.94% Monday Shift 3 Montaje, Producci´on montada (Together) 14 51 27.45% Monday Shift 4 Carga L+D liberada (Separate) 11 50 22.00% Composici´on de cargas (Separate) 8 50 16.00% Entrada Material Sucio (Separate) 16 50 32.00% Tuesday Shift 1 Carga L+D iniciada, Cargado en carro L+D (Together) 9 58 15.52% Carga L+D liberada (Separate) 14 58 24.14% Carga de esterilizador liberada (Separate) 20 58 34.48% Montaje (Together) 21 58 36.21% Producci´on montada, Montaje (Together) 20 58 34.48% Tuesday Shift 2 Carga L+D liberada (Separate) 5 16 31.25% Cargado en carro L+D (Separate) 3 16 18.75% Comisionado (Separate) 3 16 18.75% Montaje, Producci´on montada (Together) 3 16 18.75% Producci´on montada, Montaje, Carga L+D liberada (Together) 3 16 18.75% Tuesday Shift 3 Comisionado (Separate) 9 51 17.65% Montaje, Producci´on montada (Together) 10 51 19.61% Tuesday Shift 4 Carga L+D liberada (Separate) 12 48 25.00% Composici´on de cargas (Separate) 8 48 16.67% Entrada Material Sucio (Separate) 16 48 33.33% Montaje (Together) 10 48 20.83% Producci´on montada, Montaje (Together) 13 48 27.08% Wednesday Shift 1 Carga L+D iniciada (Separate) 13 63 20.63% Carga L+D liberada (Separate) 11 63 17.46% 126 Carga de esterilizador liberada (Separate) 16 63 25.40% Composici´on de cargas (Separate) 12 63 19.05% Montaje (Together) 19 63 30.16% Producci´on montada (Together) 10 63 15.87% Producci´on montada, Montaje (Together) 18 63 28.57% Wednesday Shift 2 Carga L+D iniciada (Separate) 4 17 23.53% Carga L+D liberada (Separate) 7 17 41.18% Carga L+D liberada (Separate) → Comisionado (Separate) 3 17 17.65% Carga de esterilizador liberada (Separate) 3 17 17.65% Cargado en carro L+D (Separate) 3 17 17.65% Cargado en carro L+D (Separate) → Carga L+D iniciada (Separate) 3 17 17.65% Comisionado (Separate) 4 17 23.53% Composici´on de cargas (Separate) 4 17 23.53% Montaje, Producci´on montada (Together) 5 17 29.41% Wednesday Shift 3 Montaje, Producci´on montada (Together) 9 54 16.67% Wednesday Shift 4 Carga L+D iniciada (Separate) 9 48 18.75% Carga L+D liberada (Separate) 15 48 31.25% Carga L+D liberada (Separate) →Entrada Material Sucio (Separate) 9 48 18.75% Composici´on de cargas (Separate) 8 48 16.67% Entrada Material Sucio (Separate) 20 48 41.67% Producci´on montada, Montaje (Together) 11 48 22.92% Thursday Shift 1 Carga L+D iniciada, Cargado en carro L+D (Together) 10 55 18.18% Carga L+D liberada (Separate) 16 55 29.09% Carga de esterilizador liberada (Separate) 22 55 40.00% 127 Carga de esterilizador liberada (Separate) →Carga L+D liberada (Separate) 9 55 16.36% Comisionado (Separate) 9 55 16.36% Montaje (Together) 9 55 16.36% Producci´on montada, Montaje (Together) 16 55 29.09% Thursday Shift 2 Carga L+D iniciada (Separate) 3 20 15.00% Carga L+D liberada (Separate) 6 20 30.00% Carga de esterilizador liberada (Separate) 3 20 15.00% Comisionado (Separate) 3 20 15.00% Composici´on de cargas (Separate) 6 20 30.00% Composici´on de cargas (Separate) → Carga L+D liberada (Separate) 4 20 20.00% Montaje, Producci´on montada (Together) 7 20 35.00% Producci´on montada, Montaje, Carga L+D liberada (Together) 4 20 20.00% Thursday Shift 3 Montaje, Producci´on montada (Together) 8 48 16.67% Thursday Shift 4 Carga L+D iniciada (Separate) 7 45 15.56% Carga L+D liberada (Separate) 9 45 20.00% Composici´on de cargas (Separate) 7 45 15.56% Entrada Material Sucio (Separate) 16 45 35.56% Producci´on montada, Montaje (Together) 9 45 20.00% Friday Shift 1 Carga L+D iniciada (Separate) 9 50 18.00% Carga L+D iniciada, Cargado en carro L+D (Together) 8 50 16.00% Carga L+D liberada (Separate) 16 50 32.00% Carga de esterilizador liberada (Separate) 21 50 42.00% Comisionado (Separate) 9 50 18.00% Composici´on de cargas (Separate) 11 50 22.00% Montaje (Together) 12 50 24.00% 128 Producci´on montada, Montaje (Together) 19 50 38.00% Friday Shift 2 Carga L+D iniciada (Separate) 5 18 27.78% Carga L+D liberada (Separate) 6 18 33.33% Carga de esterilizador liberada (Separate) 3 18 16.67% Cargado en carro L+D (Separate) 3 18 16.67% Cargado en carro L+D (Separate) → Carga L+D iniciada (Separate) 3 18 16.67% Composici´on de cargas (Separate) 6 18 33.33% Composici´on de cargas (Separate) → Carga L+D liberada (Separate) 3 18 16.67% Montaje, Producci´on montada (Together) 6 18 33.33% Friday Shift 3 Montaje, Producci´on montada (Together) 10 46 21.74% Friday Shift 4 Carga L+D liberada (Separate) 13 46 28.26% Composici´on de cargas (Separate) 10 46 21.74% Entrada Material Sucio (Separate) 15 46 32.61% Saturday Shift 1 Carga L+D iniciada, Cargado en carro L+D (Separate) 4 14 28.57% Composici´on de cargas (Separate) 3 14 21.43% Saturday Shift 4 Carga L+D iniciada, Cargado en carro L+D (Separate) 7 13 53.85% Sunday Shift 1 Carga L+D iniciada, Cargado en carro L+D (Separate) 4 13 30.77% Carga de esterilizador liberada (Separate) 5 13 38.46% Composici´on de cargas (Separate) 6 13 46.15% Composici´on de cargas (Separate) → Carga de esterilizador liberada (Separate) 4 13 30.77% Sunday Shift 3 129 Carga L+D iniciada, Cargado en carro L+D (Separate) 1 1 100.00% Sunday Shift 4 Carga L+D iniciada, Cargado en carro L+D (Separate) 6 12 50.00% Carga de esterilizador liberada (Separate) 2 12 16.67% Carga de esterilizador liberada (Separate) →Carga L+D iniciada, Cargado en carro L+D (Separate) 2 12 16.67% Producci´on montada, Montaje, Carga L+D liberada (Separate) 5 12 41.67% 130