Full text
Eindhoven University of Technology MASTER Designing an Architecture for a Common Data Environment Supporting Digital Twins Pop, Catalin Award date: 2023 Link to publication Disclaimer This document contains a student thesis (bachelor's or master's), as authored by a student at Eindhoven University of Technology. Student theses are made available in the TU/e repository upon obtaining the required degree. The grade received is not published on the document as presented in the repository. The required complexity or quality of research of student theses may vary by program, and the required minimum study period may vary in duration. General rights Copyright and moral rights for the publications made accessible in the public portal are retained by the authors and/or other copyright owners and it is a condition of accessing publications that users recognise and abide by the legal requirements associated with these rights. • Users may download and print one copy of any publication from the public portal for the purpose of private study or research. • You may not further distribute the material or use it for any profit-making activity or commercial gain Take down policy If you believe that this document breaches copyright please contact us providing details, and we will remove access to the work immediately and investigate your claim. Download date: 17. Nov. 2025
Eindhoven University of Technology Master Thesis Designing an Architecture for a Common Data Environment Supporting Digital Twins Author: C˘at˘alin Pop (1738569) [email protected] University Supervisor: prof. dr. Michel R.V. Chaudron Capgemini Engineering Supervisor: Marc Hamilton 18th September 2023
Abstract Digital Twins (DT) have become more and more popular in recent years in both the academic and industry contexts, as more sectors have started to modernise. Due to the complexity of these systems and the significant amount of data needed to develop and maintain a digital twin, oftentimes errors arise from a lack of communication between different stakeholders, or a lack of consistency between the data needed to create a DT. These issues can be solved with the introduction of a Common Data Environment (CDE), which is a digital system that serves as a repository where all the information of a project is centralised. A CDE provides a way for all the members who work on a Digital Twin to manage and share project-related data, such as design files, sensor data, and other relevant information. The CDE can also provide the data directly to the DT system itself. This opens the door for new functionalities such as the automatic generation of a new version of a DT, based on changes done to the data in the CDE. The goal of this project is to design an architecture for a CDE that supports the development and maintenance of a DT. In order to develop a generic architecture that can be applied to any DT system, the starting point is using a practical use case and creating a CDE architecture for the already existing DT of the soccer robots of Tech United Eindhoven. A prototype CDE web application was developed for the system to confirm the architecture works in practice. The CDE uses a graph database for data storage and leverages the properties of linked data to create relationships between the different elements of the DT. The end result is a generic architecture accompanied by a generic ontology, which can be used in the context of any DT with minimal changes. i
Contents 1 Introduction 1 1.1 ProblemContext ............................. 1 1.2 ResearchQuestions............................ 1 1.3 Definitions................................. 2 1.3.1 DigitalTwin............................ 2 1.3.2 Building Information Modeling . . . . . . . . . . . . . . . . . 5 1.3.3 Common Data Environment . . . . . . . . . . . . . . . . . . . 6 1.4 Theoretical Background: Linked Data, RDF, and Ontologies . . . . . 9 2 Related Work 11 2.1 CDEandBIM............................... 11 2.2 CDEandDT ............................... 13 2.3 Literature Review Conclusions . . . . . . . . . . . . . . . . . . . . . . 15 3 Practical Case Background 16 3.1 SoccerRobots............................... 16 3.1.1 Tech United Eindhoven and Robot Soccer . . . . . . . . . . . 16 3.1.2 TURTLE Robots Physical Elements . . . . . . . . . . . . . . 17 3.1.3 Software.............................. 18 3.2 Soccer Robots Digital Twin . . . . . . . . . . . . . . . . . . . . . . . 19 3.2.1 2DSimulator ........................... 20 3.2.2 Unity Digital Twin . . . . . . . . . . . . . . . . . . . . . . . . 20 4 Proposed Architecture 23 4.1 CDEElements .............................. 23 4.2 Ontology.................................. 24 4.3 Physical Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 4.3.1 DataLayer ............................ 30 4.3.2 Data Coordination Layer . . . . . . . . . . . . . . . . . . . . . 31 4.3.3 Specialized Tools Layer . . . . . . . . . . . . . . . . . . . . . . 31 4.3.4 Tool Coordination Layer . . . . . . . . . . . . . . . . . . . . . 32 4.3.5 Visualisation Layer . . . . . . . . . . . . . . . . . . . . . . . . 32 4.4 Conceptual Architecture . . . . . . . . . . . . . . . . . . . . . . . . . 32 ii
5 Prototype Application 36 5.1 Features.................................. 36 5.1.1 Data Visualisation . . . . . . . . . . . . . . . . . . . . . . . . 37 5.1.2 Filtering by Environment and Object . . . . . . . . . . . . . . 37 5.1.3 External File Download . . . . . . . . . . . . . . . . . . . . . 38 5.1.4 Metadata Update in Graph Database . . . . . . . . . . . . . . 38 5.1.5 Consistency Checks . . . . . . . . . . . . . . . . . . . . . . . . 39 5.2 Technologies................................ 39 5.2.1 DataLayer ............................ 39 5.2.2 Data Coordination Layer . . . . . . . . . . . . . . . . . . . . . 40 6 Discussions 42 6.1 Iterative Research Process . . . . . . . . . . . . . . . . . . . . . . . . 42 6.1.1 FirstIteration........................... 42 6.1.2 Second Iteration . . . . . . . . . . . . . . . . . . . . . . . . . 42 6.1.3 ThirdIteration .......................... 42 6.1.4 Fourth Iteration . . . . . . . . . . . . . . . . . . . . . . . . . . 43 6.2 Encountered Problems . . . . . . . . . . . . . . . . . . . . . . . . . . 43 6.3 CDE-Centered Design . . . . . . . . . . . . . . . . . . . . . . . . . . 43 7 Conclusions 48 7.1 Summary ................................. 48 7.2 ResearchQuestions............................ 50 7.3 FutureWork................................ 52 7.3.1 Implementation of Data Management Services . . . . . . . . . 52 7.3.2 Expanding the Current System . . . . . . . . . . . . . . . . . 52 7.3.3 Security and Access Control . . . . . . . . . . . . . . . . . . . 52 A Appendix 57 iii
1 Introduction This section will serve to provide an introduction to the whole project, by detailing the problem context, defining the research questions, and explaining the most important terms. 1.1 Problem Context As more sectors begin to modernise and start using digital data (e.g. sensor data) for the design, construction, and maintenance of their systems, there is a need for effectively storing and managing large amounts of data for the creation of Digital Twins. DTs can bring multiple benefits, such as reduced development time, reduced costs, better decision-making, and easier maintenance. One of the main issues that the development of a DT faces comes from the complexity of such a system and the large amounts of data that is needed to create it. Generally, there are multiple people who work on such a project, coming from different environments, which can often cause issues in communication. This can lead to errors in the design of the assets, data inconsistencies and more. A solution for this issue is using a Common Data Environment, a digital repository that can centralise data from different environments and external data sources. The CDE can create the connections between the different kinds of available data and link it to the elements of the DT. The research available for such an approach for DTs in practice is limited, so creating an architecture for a CDE would be beneficial since it could be reused in the future for multiple systems. There are multiple aspects that need to be considered, such as how to centralise the data, how to create connections between the elements, how to link data to its corresponding environment, and how to best leverage the capabilities of a CDE to help the development and maintenance of a DT. 1.2 Research Questions The main goal of this project is to design a generic architecture for Common Data Environments supporting Digital Twins, and from this goal, the main research question can be identified: ”What does a generic design for Common Data Environments in the context of Digital Twins look like?”. Multiple other questions and subquestions have been identified that will help guide this research toward the desired result: 1
•RQ1: How is it possible to link the data from multiple sources in the CDE to the right elements of the DT? –What is the available information in the CDE to build the DT from the data? –How to identify the different parts needed to build the DT? •RQ2: How can the data in the CDE be synchronized in order to achieve consistency between different sources? •RQ3: How can the CDE support the construction of a Digital Twin? •RQ4: How to construct an automated pipeline from the CDE to the DT? –Creating automatic updates to the DT based on the changes to the data in the CDE. RQ1 and RQ2 will be answered by examining the available standards and creating an ontology that will help with linking the data in such a way that the CDE can provide functionalities such as consistency checks between elements. RQ3 and RQ4 will be answered through a literature review that examines the current approaches and results, and through building a prototype CDE application for a practical use case, which will provide more insight into the practical side of this approach. In the end, the results achieved from all the steps taken will result in a generic architecture that will answer the main question. Each question has been examined during the different iterations of the development process of this project. More information about the progress made in each iteration can be found in section 6. 1.3 Definitions 1.3.1 Digital Twin The Digital Twin (DT) is still an emerging technology, which means that not all details about it are set in stone, with different authors having different interpretations and definitions of what a Digital Twin is. A popular definition that seems to be used in a significant part of the research done on DTs is the one from Grieves and Vickers, which states that ”the Digital Twin is a set of virtual information constructs that fully describes a potential or actual physical manufactured product from the micro atomic level to the macro geometrical level. At its optimum, any information that could be obtained from inspecting a physical manufactured product can be obtained from its Digital Twin”.[1] 2
An easier way to understand what a DT is, is to look at its uses. Currently, a DT is needed for multiple reasons in the life cycle of a product, such as: •reducing the costs when creating a prototype and executing tests on it. •performing tests under extreme conditions that are hard or impossible to replicate in a real setting. •monitoring a physical asset in real-time and making predictions on its future state. •centralising data from a physical asset and performing real-time analysis on it. As it can be seen from the above requirements, a Digital Twin can end up being a very complex system that encompasses multiple technologies. At its core, a Digital Twin is made out of three main components: •the physical world: this is represented by the actual physical asset, which has sensors and actuators that collect data. •the digital world: the virtual representation that tries to emulate the physical asset as much as possible. •the connection that allows exchanging data between the other two components, which can be one-way or two-way, depending on the level of evolution of the DT. A simple representation of these three components can be seen in Figure 1. Figure 1: Main Components of a Digital Twin, from [2] 3
Depending on the available connections between the physical world and the digital world, a system can be categorized either as a Digital Twin or a Digital Shadow. If the connection between the two components is bidirectional, the system is called a Digital Twin, while if the connection goes only from the physical world to the digital world, the system is called a Digital Shadow.[3] However, during the evolution of the Digital Twin, multiple other components have been attributed to it, such as Machine Learning or Big Data. A. Sharma et al. [4] have performed a literature study that analyses what components different authors consider to be part of the Digital Twin, with the results being summarized in Figure 2. By looking at the table, it is clear that there is no consensus on the extended components of a DT, but some technologies are present in most of the analyzed papers, such as transfer of information, machine learning, IoT, or time-continuous data. 4
2 Related Work In this section, a literature review has been performed to see and compare the findings of other research projects on similar topics. While most available literature doesn’t deal directly with the combination of the CDE and DT, there are some papers on the use of CDEs in BIM, with the results also being applicable to the context of the Digital Twin. 2.1 CDE and BIM There is more research available on the topic of the Common Data Environment being used in BIM systems, since these technologies have also been used in practice more in comparison to Digital Twins. Even though this part of the literature review does not help directly in answering the question of this report, it gives certain pointers towards the right answer, as BIM systems and Digital Twins have some common characteristics and uses. This section will present the available research on these two topics. Possibly the most relevant research available from a CDE perspective is the work of Werbrouck J. et al [12], since it discusses a graph-based Common Data Environment in the context of Building Information Modeling. They argue that a linked data-based CDE ”would need to combine the possibilities offered by CDEs, while allowing to establish links between RDF data and documents, on a fine-grained data level. This way, project sources such as imagery, point clouds, geometry etc. can be linked with RDF graphs about topology, products and properties”. The authors also analyze Solid, a graph-based ecosystem for sharing data in the context of social media, and try to apply it to the BIM environment. In [13], Preidel, C. et al. recognize the major challenge of maintaining consistency during the collaboration between stakeholders of a construction project. The authors go in detail and describe possible consistency issues that can be encountered in a BIM project, like the merging of domain-specific models with the coordination model, which is similar to a merge conflict in a Version Control System. The authors propose that a CDE could implement a kind of locking mechanism on information containers in order to keep the consistency. A framework is introduced, which allows a smooth integration of CDE access into standard BIM authoring tools. It is concluded that the CDE reduces data redundancy and secures the availability of up-to-date data. The Common Data Environment provides higher reusability of information and simplifies the aggregation of model information. In the conclusion, the authors also discuss Project DRUMBEAT, which focuses on the conversion of BIM models from the IFC format to a web-compatible format using linked data. 11
However, this is just a mention of a possible graph-based approach to a CDE and it is not further discussed. One advantage of this approach is mentioned, which is that the building parts from different models can be linked to each other or to external information systems and data sources on the web. Important research for this review was done by ¨ Ozkan, S. et al. [14], where the focus was on the identification of CDE functions during the construction phase of BIM projects. The authors identify these functions by performing several interviews with subject matter experts who have experience in working with Common Data Environments. The identified functions were categorized by different management domains, including ”Design Management”, ”Document Management”, ”Knowledge Management” and so on. The results are presented in the form of a table which can be seen in Figure 6. 12
Figure 6: CDE functions categorized by managerial attributes, from [14] 2.2 CDE and DT As explained before, there is not a lot of research done on Common Data Environments being applied to a Digital Twin. This is because DT is a newer and more complex technology in comparison to BIM systems. However, there has been some 13
recent research on the idea of applying these two technologies together, which is presented in this subsection. Yan J. et al. [15] have recognized the similarities between BIM and DT, and observed the limitation that Digital Twins face with the difficulty of building a shared data environment for collaborators of a project during its life cycle. They have analyzed other research on the topic of DT, which stated that as the complexity of data sources increases (which is the case with all the data used in Digital Twins coming from IoT devices), data management issues start to arise. The paper focuses on the topic of extending Digital Twin technology from a building level to the city level, by reviewing current data management for DTs in other publications, concluding challenges that arise from this data management, and introducing a ”city-level Common Data Environment that addresses the issues identified”, such as interoperability problems and functional and organisational issues. While this paper is mostly theoretical and focuses on a specific type of CDE for cities, not a generic CDE that can be applicable to any DT, it still marks how a CDE generation can help with building and maintaining a Digital Twin. Another paper by Dang N. et al. [16] focuses on creating a Digital Twin for bridge maintenance, arguing that a 3D DT approach could be more efficient than the current way of doing the inspection manually. They argue that representing bridge elements virtually that could hold attributes about them can help the maintenance process and predict future issues that may arise, also with data collected from damage records and logs. The paper briefly mentions a Common Data Environment that stores data about the physical features of the bridge, and also archives data for each element of the bridge. This is in conformity with the design of the CDE, with information containers having different versions and revisions. In this case, the Common Data Environment’s properties can work very well in collaboration with all the data available through the DT, making it easier to compare elements and their properties during a time period, which helps with maintenance and predictions. Meˇza S. et al. [17] have investigated the possibility of building a Digital Twin of a road constructed using secondary raw materials (SRMs) to analyze the long-term durability of such materials and if they can be safely used in construction. They start out with a BIM system and include the integration of sensor data to take the first step in building a Digital Twin. The authors identified the technical requirements of the CDE to be able to support the Digital Twin. Most of them were related to BIM, such as the capability to import IFC model files, and support for BIM metadata, but one interesting technical requirement was the ability to export sensor data in the open IFC file format, with this data being crucial for the DT. 14
2.3 Literature Review Conclusions Through this literature review, multiple general uses of a CDE have been identified, most of which can be applied to DT systems. One of the examined papers specifically dealt with a graph-based CDE, which tried to present the benefits of linking the data through semantic web technologies. The review has helped provide some general insight on how to answer one of the research questions (RQ3) from a theoretical point of view. 15
3 Practical Case Background In this section, the Tech United soccer robots and their associated Digital Twin will be examined, in order to provide the necessary background knowledge needed to understand the context of the selected practical case, and to clearly explain the elements present in the different environments, which will then be used for the prototype application. In subsection 3.1, the context of the real robots will be explained, together with their physical parts, and the needed software. In subsection 3.2, the existing Digital Twin will be analysed, together with its elements and environments. All of these elements and environments will play a role in the development of the ontology presented in subsection 4.2, and the prototype application presented in section 5. 3.1 Soccer Robots 3.1.1 Tech United Eindhoven and Robot Soccer Tech United Eindhoven is a robotics team made up of students (current and former), PhD students and employees of Eindhoven University of Technology. The team initially started working on robots that would compete in the Middle Size League of RoboCup, but they have since then expanded into other fields as well, such as the creation of service robots [18]. RoboCup is an organization founded in 1997, that organizes tournaments for autonomous robots that play soccer, with one of their goals being the push of state-of-the-art technology in robotics. These tournaments are separated into different leagues, based on criteria such as the robots’ shape, size, and capabilities [19]. One of these leagues is the Simulation League, where the robots are not physical entities, only simulations, so the focus is mostly on the software side. Another league where the focus is mostly on the software is the Standard Platform League, where all teams have access to the same hardware, which evens out the field in that department. There is also the Humanoid League, where the robots are shaped similarly to humans, giving a more realistic visual to the game. Finally, there are the Small Size League and the Middle Size League, where teams create their own robots that have to respect certain criteria, such as size and weight. The soccer robot team at Tech United competes in the Middle Size League. The rules of the Middle Size League are similar to the ruleset employed by real-life soccer organizations, such as FIFA or UEFA. A lot of the rules are directly inspired by the FIFA ”laws of the game” ruleset, such as the size of the ball being the same and the field of play being rectangular (with a different size compared to the real game). In the Middle Size League, teams are composed of no more than 5 players, 16
with one of them being the goalkeeper. Similarly to the human sport, fouls can be awarded in the case of damage inflicted by a robot to another robot on the other team, which can result in a free kick or a penalty kick. On top of the ”classic” soccer rules, there is another set of rules which refer to the capabilities of the robots. For example, robots are not allowed to have an overhead camera that grants vision to all of the field, and each robot must do its own processing, so no central server is allowed (however, communication between the robots is permitted). 3.1.2 TURTLE Robots Physical Elements The soccer robots of Tech United are called TURTLEs. Through several iterations over the years, they are now in the fifth generation, and their current representation can be observed in Figure 7. Figure 7: Physical Soccer Robots of Tech United, from [20] There are multiple subsystems that contribute to the functionality of the robots. Some of the notable physical parts are the two cameras used for vision (one ”OmniVision” camera on top of the robot, and a Kinect device), the Omni-wheels for motion (a set of special wheels which have smaller wheels on them, which give an increased level of agility to the robots), and the CPU and GPU, which provide the computing power for the robots to communicate with each other, share information and decide on a strategy. All of these physical elements are controlled by specific software subsystems, that will be detailed in subsubsection 3.1.3. Another important element of the physical robot is its design. Tech United uses 3D CAD models for the design of the robot, with some parts of the robot being 17
directly created from these 3D models. One of these models can be seen in Figure 8. It’s important to note that while these models are mostly accurate, they do not fully depict the real-life robots, which can have some differences, depending on how upto-date the models are. This is something that could be fixed by the introduction of consistency checks in a CDE. Figure 8: 3D CAD Model of the Outfield Robot 3.1.3 Software The soccer robots have their own software subsystems which allow them to perform their tasks during games and communicate with each other. There are four subsystems in total: Strategy, Motion, Vision, and WorldModel. The Strategy module is responsible for deciding what strategy must be used during play, based on the current state of the world, and each robot gets a role assigned based on the chosen strategy. The Motion subsystem is different between the different types of robots (e.g. 3-wheeled robot, 4-wheeled robot, goalkeeper), but its goal is to get the information from the Strategy and translate it into commands for the movement of the robots. The Vision subsystem is responsible for fetching and analysing the images from the cameras, which gives the position of the robot in the field, the position of the robots from the opposing team, the position of the ball, and more. Finally, the WorldModel module creates a model based on the information gathered from all 18
the robots. Each robot has its limited view of the world, but by combining all the available data from multiple robots through communication, it is possible to get a single view of the field and the position of the opposing team robots. Each of these subsystems has its own dataflow graph made in Simulink, which specifies the input/output and how the data flows through the system. It is important to note that there are different versions of each subsystem, one for the real robots, and one for the simulator, with the simulator models being slightly simplified compared to the real ones. The communication between the robots is done through an adapted version of a Real-Time Database (RtDB), which was developed by another soccer robot team and has been used by multiple other teams. [21] The team has another service called Turtle Remote Control (TRC), which is used to control the robots in both the physical and simulator environments. This service gives users the possibility to start and stop the robots, set up free kicks and penalties, and assign roles to the robots. Figure 9: Soccer Robots Software Components 3.2 Soccer Robots Digital Twin A DT prototype for the soccer robots was developed by a TU/e student as part of their master thesis, but before this explicit DT was created, Tech United already had some simulation tools that helped with developing and testing new strategies in a virtual environment, namely a 2D Simulator and a 3D Simulator. For the purpose of this project, the already existing 2D Simulator will be considered as part of the Digital Twin environment, since it is directly connected to the DT, while the 3D Simulator will not be considered, since at the time of developing this project, the team was only using the 2D Simulator for developing their strategy, with the 3D 19
Simulator not being actively used and not bringing any functionality that the DT prototype doesn’t have itself. 3.2.1 2D Simulator The main tool that Tech United uses to develop and test new functionalities and strategies is the 2D Simulator. It offers a top-down view of the pitch, with the robots being represented by 2D sprites. Even though the simulator presents a somewhat simplified version of a real-life scenario, the main advantage it offers is the ease of setting up a game and testing different plays and strategies. The simulator offers the possibility of setting up games with up to three robots in each team. After the game starts, the robots start playing by themselves, while the user can control the environment to test different scenarios, through features such as manually moving the ball in a certain spot and adding obstacles. This makes it easy to test changes by simply restarting the environment after changes in the code are made. A running instance of the simulator can be seen in Figure 10. Figure 10: Running Instance of Tech United’s 2D Simulator 3.2.2 Unity Digital Twin The Digital Twin prototype was created as part of a master’s student thesis, which focused on reusing artefacts for the creation of Digital Twins.[22] As a consequence, a lot of the elements present in the DT are based on already existing artefacts, which creates a relationship between the original elements, and the DT elements (e.g. the 20
Figure 16: Resulting Ontology for Soccer Robots Figure 16 shows the final version of the resulting ontology. It is made out of a combination of elements from the NEN 2660 standard at the base, some elements from the Autonomous Robots ontology, and a few others that are specific for the soccer robots domain (e.g. Ball and Pitch). Starting with the Physical category, that is split into three categories: Process, which was taken from [24], and InformationObject and PhysicalObject, which are imported from the NEN 2660 standard. The Process class is a superclass for RobotBehavior, which is also split into four subclasses that match with the subsystems of the soccer robots software. The InformationObject class represents the objects that describe the PhysicalObject instances in different environments (e.g. Blender model files, CAD files, Matlab models files), and they are the objects that are displayed in the CDE, having additional data properties (e.g. source of the file, last update date). There is also a new object property called ”dependsOn” which has an InformationObject as both the domain and range. This property is used to establish a dependency between files, which helps with maintaining consistency between them in the CDE, together with the ”lastUpdate” data property. The InformationObject class from the NEN 2660 standard can be seen as equivalent to the ContentBearingObject class in the SUMO ontology [25] present in [24] in this case. On the other hand, the PhysicalObject class represents the actual objects in the domain, and every instance of its subclasses needs to be connected to an InformationObject, which is done through the ”isDescribedBy” object property from the NEN 2660 standard. The subclasses of PhysicalObject are the Robot, Pitch and Ball classes, which correspond to the elements identified in subsection 4.1. The Robot class also has subclasses to rep27
resent the different types of soccer robots available, and it has the object property ”hasBehaviour” that connects a robot to the RobotBehaviour class. Moving on to the Abstract category, it has two subclasses, the Environment and Proposition. The Proposition class is taken from the Autonomous Robots ontology, together with its subclasses RobotArchitecture and ElementRA and their properties. These can be used in the future to represent the physical elements of the robot, however, they are not currently used in the CDE application, since the robot itself is considered a single entity. The Environment class is a new addition that is not present in the two base ontologies. It was added by considering the context of Digital Twins and it can be generalised for every such system, no matter the domain. Since all Digital Twins normally have a physical entity and a virtual entity, there are always multiple environments present in such a system (in this case, at least three, since there are two simulators). It is also important to consider the Running Environments of the entities, which can create new environments for both the physical and virtual entities. Considering all these aspects, it was decided to add such classes and connect their instances to InformationObject and PhysicalObject instances through the ”partOf” relationship. This allows for a separation of objects based on environment, which is an important component for data coordination. It can be observed that a big part of the ontology can be reused for any CDE and DT system, with the exception of the domain-specific classes. This means that the developed ontology is generic enough, and the core concepts such as the PhysicalObject, InformationObject, Environment, and their properties can easily be reused, and they will prove to be the most useful in the prototype CDE. 4.3 Physical Architecture In order to develop the generalized conceptual architecture, the start point is to represent the physical architecture for the prototype application presented in section 5. The resulting physical architecture that can be seen in Figure 17 was achieved after analysing the practical use case, identifying the necessary elements present in the Digital Twin, creating the ontology and going through the iterations of the development process for the prototype application. All the elements of this diagram and their connections will be explained here in more detail from an architectural point of view, while the implementation details and technologies used for the CDE are explained in section 5. 28
Figure 17: Physical Diagram for Prototype Application The physical architecture diagram was developed with the concept of the layered 29
architecture style in mind. The first step was to identify the main parts of a digital twin based on the literature study done at the beginning of the project. That resulted in five main layers: the Data Layer, the Data Coordination Layer, the Specialized Tools Layer, the Tool Coordination layer and the Visualisation Layer. Since a CDE’s main concern is data and its structuring, the first two layers have been separated into the Common Data Environment package, while the other three layers have been grouped in the Digital Twin package. This creates a clear separation of concerns, marking that the CDE is concerned with storing and orchestrating the data itself, while the DT is responsible for using that data for services and features. For the soccer robots case, it was already mentioned that the Digital Twin has two different simulation environments, the 2D simulator and the Unity simulator. This is represented in the DT package by the two interconnected subsystems spanning over the three main layers. There are dependencies from elements in the Digital Twin going to external data sources (in this case Gitlab repositories). This is the case for all the elements in the DT, however, only two examples (the behaviour models and robot models, from two different repositories) are shown in this diagram in order to not overload the diagram with multiple relationships that show the same thing. These dependency relationships are present since the DT was not initially designed with a CDE in mind, but ideally, all the data connections would go through the CDE directly. 4.3.1 Data Layer The data layer lies at the base of the Common Data Environment and it is responsible for the storage of the linked data in a knowledge graph. This graph is split up into two categories, the TBox and ABox. The TBox is responsible for the definition of classes and their rules (which come from the ontology that was previously defined), while the ABox is represented by the instances of these classes, which hold the data itself. [26] In practice, the TBox and ABox are usually stored together in the same graph, and the same thing happens here since all the data is stored in a GraphDB graph database instance. In the diagram, the TBox shows a simplified version of the ontology presented in Figure 16, with the instances of the classes and their relationships shown in the ABox. The dependency arrow from the graph database to the Digital Twin package establishes a connection between the CDE and DT, and it represents that the data stored in the database, together with the relationships, depict the structure of the files in the DT. This allows a linking between the data that can then be used to construct 30
the Digital Twin, since there are defined relationships between the elements and files. 4.3.2 Data Coordination Layer The data coordination layer is the subsystem responsible for facilitating the overall management, synchronization and exchange of the data from the data layer. For this implementation, the data coordination layer is represented by the prototype common data environment application that was developed. It has a general clientserver architecture, with the connections between the user interface and the server being established by HTTP request methods. The client is responsible for providing a user interface that allows the user to visualize and interact with the graph database through different views that are made out of components. The server side of the application has two main components, the routers and the services. The routers are responsible for handling the requests received from the client and sending back the responses from the services. The services are responsible for interacting with the data itself. One of the services is concerned with interacting with the graph database, while the other one is concerned with external data sources (in this case, multiple Gitlab repositories where project files are stored). The GraphDB service connects to the data layer through the GraphDB REST API, being able to get and update data from the graph database through SPARQL queries, while the Gitlab service interacts with the external data sources through the Gitlab REST API, being able to fetch data and metadata about project files, which can then be introduced in the graph database through the GraphDB service, or to directly download files to a local computer. 4.3.3 Specialized Tools Layer Moving on to the Digital Twin package, the base layer is the Specialized Tools layer. This layer encompasses the main tools of the DT, which from a design point of view are represented by the Simulink behaviour models in the case of the 2D simulator, and by the main Unity scripts that create the services in the Unity environment, which were shown in Figure 11. These tools are mostly dependent on the specific system and its domain, but they can be generalized based on the most common tools that are usually present in Digital Twins, which were examined during the literature study done in section 1 and section 2. 31
4.3.4 Tool Coordination Layer The Tool Coordination layer is built on top of the Specialized Tools layer and it is responsible for the communication between the different services and for combining different tools in the Specialized Tools layer to create new functionalities. In the soccer robots case, the TurtleCommManager script presented in Figure 11 belongs to this layer, as it is responsible for providing and synchronizing the data to the other services. The Real-Time Database is also part of this layer even though it is outside of the two subsystems. While the RtDB is also responsible for data storage for the robot data, its main use is establishing the communication between the robots in a real-time environment. The RtDB is also responsible for creating the connection between the two simulation environments: the 2D simulator establishes a connection with the RtDB which stores the simulation data, which is then fetched by the TurtleCommManager in the Unity environment through the TurtlePlugin Communication file. In this case, the RtDB is used both for communication purposes, and for creating new functionalities in the Unity environment based on the Behaviour Models in the 2D simulator, and for these reasons, the RtDB was placed in this layer, and not the Data or Data Coordination layers, even though it also acts as a data storage. 4.3.5 Visualisation Layer In general, visualisation can be seen as just another functionality for a system, which means the Visualisation layer could be merged with the Specialized Tools layer and be a part of it. While a lot of DTs have several services and tools in common, visualisation is a central part of any Digital Twin, which makes the separation of this functionality from other services warranted. The Unity environment has an integrated visualisation tool, which uses the files in the Materials directory. This directory contains all the models used in the visualisation, in this case, the Blender models for the robot, the stadium and the ball. 4.4 Conceptual Architecture After achieving a satisfactory version of the physical architecture presented in subsection 4.3, the next step is to try to generalise each element in order to come up with a conceptual architecture, while maintaining the layer structure used in the physical architecture. Starting at the bottom of the CDE package with the Data Layer, the GraphDB database is replaced by a generic linked data storage, with another element being 32
the Ontology, which is then imported as a linked data file in the linked data storage, either manually or automatically, depending on the system. In the Data Coordination layer, the same structure as the one from the physical architecture was kept, using the same client-server architecture for the CDE application. The services are more generalized, with a data retrieval service for the external data sources, another service to connect to the linked data storage in the data layer, and other services that can be added depending on the needed functionalities (e.g. security and access control). The linked data service connects to the storage through an API provided by the chosen linked data storage, while the client that provides the user interface communicates with the server through HTTP requests. In the Digital Twin package, the Specialized Tools layer presents a few of the generic services that are usually available in DTs, as they were researched during the literature study. These tools are connected through a dependency relationship to the Service Coordination module in the Tool Coordination Layer. The service coordination, which is realized through the TurtleCommManager in the physical architecture, synchronizes the other tools and can facilitate the creation of new functionalities or data by combining them. Other elements present in this layer are the Data Communication (which is realized by the Real-Time Database in the physical architecture) and a Brokerage tool. The Visualisation layer has a visualisation tool, which is used to display both the models and additional data information. In the physical architecture, the data visualisation is realized by the additional data panels, which show basic data about the robots during a game, and that’s why it has a dependency to the Data Communication element. Another part of the visualisation layer is represented by the 3D models, which are then used by the Model Visualisation element to display them in the chosen 3D environment which is characteristic of DTs. Another thing that can be noticed that is similar to the physical architecture is the dependencies between the DT files and the external data sources, which have been generalized here as Model Files and the DT Project Files (e.g. scripts, communication files). 33
Figure 18: Conceptual Architecture for CDE and DT As explained before, the elements from the physical architecture have a clear mapping to elements in this conceptual architecture, and the conceptual architecture also presents elements that are generally present in DTs. However, since this conceptual architecture was derived from the physical architecture, which is based on an already existing DT system which had a prototype CDE added to it, it is possible that some elements or relationships might have been missed. To combat this, an additional ar34
chitecture (based on this one) is discussed in subsection 6.3, with the main difference between the two being that this architecture considers the addition of a CDE at a later stage, while the other one considers a DT system that is built using the CDE, which adds new elements and changes some relationships. 35
5 Prototype Application In order to test the developed ontology and architecture of the system, a prototype CDE web application was developed, which takes advantage of the features provided by the graph database. A main view of the application can be seen in Figure 19. The application offers features such as visualising the data in the graph database in a natural manner, filtering the data using SPARQL queries, and accessing information from external data sources based on each object’s source. This application was developed to test and confirm the design in section 4, and while it does not yet have all the properties that are specific for a Common Data Environment, it provides a good start which can later be expanded on. Figure 19: Main View of the CDE Prototype Application 5.1 Features All the features in the application make use of the properties of the graph database through SPARQL queries. While the current features could be implemented in normal applications that make use of a relational database, the goal is to demonstrate that the current ontology works in the context of a CDE, and that it can be generalised and be used in a similar application from a completely different domain, without needing to make changes to it. 36
6.1.4 Fourth Iteration The final iteration focused on reviewing all the findings available up to that point and starting to work on the final version of the report, identifying the elements in the architecture that can be generalized and coming up with the final generic architecture of the system. During this iteration, there was also some development done to the application for the last features. 6.2 Encountered Problems Problems were encountered during each iteration of the project. For the first iteration, the issue started from finding a relative lack of practical research done on CDEs and Digital Twins. To supplement the information, research in BIM systems was also considered through the parallels that can be drawn between these systems and DTs. This provided a context on the available research on CDEs and the lack of standardization available in multiple areas, which has demonstrated the necessity for developing a generic architecture and approach to CDEs in the context of digital twins. In the second and third iterations, the main problem was finding out how to link the data from the different sources in a generic way that can be reproduced in other systems. While the NEN 2660 standard provided some help for this issue, there were still concepts missing that could help in the context of DTs. This led to the extension of the developed ontology by adding a new abstract class for the different environments, linking the elements to the different environments, and adding a dependency relationship that sets the authority between elements. During the last iteration, the problem was settling on a definitive generic architecture. This problem was caused by the practical use case that already had a DT developed before the addition of the CDE. An additional scenario was considered, where the DT would have been developed with a CDE in mind, which brings the possibility of having new functionalities. 6.3 CDE-Centered Design The conceptual architecture diagram presented in subsection 4.4 shows the addition of a CDE to an already existing DT system, which helps with centralizing the elements from the different data sources into a central application, which also provides synchronization between the different environments, and introduces new relation43
ships between the elements of the system. However, to achieve a desirable generic architecture in the end, it would be helpful to consider the case where the DT system is designed with a CDE from the beginning. This changes the relationships between the DT system and the external data sources, and it can also provide the CDE with new functionalities and automation possibilities. Such a design is presented in Figure 22. 44
Figure 22: Extended Architecture for a CDE-Centered Design Most of the elements are the same, including the layers, the elements and most of their connections. The main difference comes from the addition of the Data Management Services package between the CDE and the DT. This package strengthens the 45
connection between the CDE and DT package and allows additional functionalities to be introduced, such as the automatic generation of a DT based on the data in the CDE. Through the connections that are formed between the elements and the different environments in the linked data storage, an additional data management service that connects to the CDE services could get the necessary information that is needed for the creation of a DT system. As an example, in the soccer robots use case, the graph database is aware of all the elements in the domain (e.g. robot, field, ball), the environment used for the simulator, and the files that represent those elements in that environment. That is all the needed information to recreate a DT system from an engineering design point of view. In addition to data management, there could also be an automation tool used to automatically generate a new version of the DT system based on the changes to the data in the CDE. Going back to the soccer robots example, if the Blender model for the robot that is used in the Unity environment has a major update, an automation service together with the data management service would be able to be notified about the change based on the connection to the CDE services and data, and from there, create a new version of the DT system with the updated Blender model. Another possibility is if, for example, the CAD model for the robots is updated, but the Blender model is not. This would be reflected in the CDE through the automatic consistency check. If the data management service would be aware of such a scenario, it could use a tool that automatically converts a CAD model to a Blender model, in order to keep both models up-to-date. There are already existing tools that can do that, such as PiXYZ Software, which could be integrated into the application.[31] This would keep the DT system as updated as possible while reducing the manual work needed to be done to update the system by team members. This type of service would fully display the capabilities and advantages of using a CDE for the development and maintenance of a digital twin. Unfortunately, there was not enough time to implement this idea during this project, so it’s hard to give a clear overview of how this would work in practice and what the interfaces between the Data Management Services and the CDE and DT would look like in a more detailed manner. Automatic DT generation is a pretty new and unexplored topic, based on the available research, but there are new projects that have this as a main focus, such as Auto-Twin.[32] This project tries to use a data-driven method that is based on a process mining approach in order to automatically generate a DT system. What is interesting about this project is their mention of an ”International Data Space (IDS)-based common data space”, which uses linked data and seems to be similar to the common data environment used in this project. Unfortunately, at this point in time, the project is 46
still in its early stages, and there is little to no information publicly available about it, however, it’s possible that such an automation service could be integrated into a system that uses the presented architecture here in the future. Some other changes that can be observed in this extended architecture are the relationships between certain elements. The dependency between the DT and the linked data storage works in the opposite direction in this scenario. Before, the data in the storage would be dependent on the structure of the DT, since the system already existed and was designed before the addition of the CDE, but in the case where a CDE is used from the start, the DT components are based on the available data in the linked data storage. A similar thing can be observed with the lack of dependencies from the DT to the external data sources. In this scenario, all the information can go directly through the CDE: the CDE gets the files and data from the external sources, and then the Data Management Services creates the DT based on the available data, which strengthens the connection between the CDE and DT while reducing outside dependencies. 47
7 Conclusions This chapter is dedicated to explaining the conclusions that were reached during the duration of the project and the possibilities of future work that can be done on this topic. subsection 7.1 presents a summary of the whole report. Next, subsection 7.2 analyses the research questions one more time and shows how they were answered, and finally, subsection 7.3 takes a look at the work and research that can be done to expand this project in the future. 7.1 Summary During this thesis project, multiple steps were taken in order to develop a generic architecture for Common Data Environments in the context of Digital Twins. Multiple possible architectures were considered, with the evolution of the end result happening over the different iterations of the project. Starting with the literature review and researching the different available technologies, this step helped with getting familiar with the problems of this topic, observing what other work has been done in this area, the advantage and drawbacks of different approaches, and seeing what issues were not yet addressed in the available research, that could be explored in this thesis. After getting familiar with the topics listed above, the next step was to analyse the soccer robots use case. This process needed to be as meticulous as possible, because without a clear understanding of all the elements and the different environments that make up the project, there was no way to coordinate all the data in the CDE in a useful manner. There was constant communication with both the developer of the Digital Twin and members of Tech United to understand all the details about the project and how the communication between the systems works, which has proved to be very useful when designing the data storage. Once the main elements of the use case were identified, the development of the ontology was performed. This was a process that was influenced by multiple factors, not just the analysis of the use case. Other contributing factors to the ontology were the available standards and ontologies that had to be followed, namely NEN 2660 and the Autonomous Robots Ontology, but also the development process of the CDE application, which showed how additional data properties that are not present in a standard can help in practice, or the need to keep the ontology as generic as possible, so it could be adapted to other systems as well. All these aspects have made the development of the ontology a continuous process, which led to constant evolution, until all the theoretical and practical necessities were met. 48
The development of the architecture itself was a process that spanned across all iterations of the project, starting from a basic layered architecture of a DT, expanding it by introducing the Common Data Environment, and analysing how this addition changes and splits the identified layers and their connections. The soccer robots use case was used in order to develop a physical architecture for the system with the introduction of the CDE prototype application. Using a concrete practical use case proved to be very useful since it confirmed that such an architecture can work in practice, and seeing how the different elements fit into each layer has given a clear purpose for each layer and clarified how to establish relationships and dependencies between them. The development of the prototype application has also shown how to leverage the capabilities of a graph database and the use of SPARQL queries to access and update data. As explained before, this has also helped shape the design of the final ontology. Once the physical architecture was created, the next stage was to develop the conceptual architecture that was derived from it. The literature study helped with this, since the familiarity with CDEs and DTs have allowed the generalization of most elements, while following the structure set in the physical architecture. The discussions chapter also considers another architecture with additional services and functionalities, where a Digital Twin is designed with a CDE in mind from the beginning. Unfortunately, this can only be researched at a theoretical level during this project, since the available use case presented an already existing DT whose design couldn’t be modified in order to fully fit the CDE, but it is still an interesting situation to consider that could be further analysed. There were multiple constraints and issues encountered during this project that did not allow for a complete analysis of the system and its resulting architecture. The time constraints of a master’s thesis did not fully allow the exploration of all possibilities for such a large project. Another aspect is the use case itself: while it was useful to have a prototype to test the application on, the DT system was actually a Digital Shadow, so it was not possible to explore all the connections between the physical system and the virtual system. Finally, since the DT already existed at the beginning of the project, some aspects of the benefits of a CDE in developing a DT might have been missed. Regardless of these issues, a lot of progress has been made and a generic architecture and ontology were developed, which can be reused for any system in the future. The prototype application that was developed and its functionalities prove the usefulness of the introduction of a CDE in a DT system through the communication that it offers to different team members, the centralization of data for a large project, the automatic consistency checks between elements in different 49
environments, and the possibility to automate the generation of a DT system in the future. 7.2 Research Questions RQ1: How is it possible to link the data from multiple sources in the CDE to the right elements of the DT? The first step is to manually identify the elements needed to create the DT and how they connect to elements represented in the physical world, which requires knowledge about the domain and the system itself. Following the architecture that was created can also help identify all the needed elements of the DT system by analysing what data and files each layer needs. The data from the external data sources is centralised in the CDE through the services provided by the CDE which connect to the external data sources to retrieve the information, and then to the graph database to store them. Through the ontology that was developed, the data is able to be linked based on the environment it belongs to and the object they represent in the DT. Another benefit provided by the created ontology is the ”dependsOn” relationship that creates an authority relationship between the different elements, which also helps strengthen the links between the elements. The CDE leverages the capabilities of linked data and graph databases to create these links between items. RQ2: How can the data in the CDE be synchronized in order to achieve consistency between different sources? Similarly to the first question, consistency between files is achieved through the CDE application services and the way the ontology was developed. The objects stored in the graph database have data properties such as the last update date, which give a clear view on the status of each object. Combining these data properties with the ”dependsOn” relationship that is established between objects can provide automatic consistency checks in the CDE application for the users. In addition to this, the periodic update of metadata in the CDE assures that the information is consistent with the external sources. Once the users can see which files from which environments are outdated, it is much easier to create consistency between the different sources. In addition to this, these functionalities can be leveraged in the future by the automation processes to automatically update outdated files. As an example from the soccer robots use case, if the CAD model used for the physical robots has authority 50
over the Blender model and the Blender model is outdated compared to the CAD model, an automation service could get this information and automatically create a new Blender model based on the updated CAD model. RQ3: How can the CDE support the construction of a Digital Twin? This is a broad question that can be researched on its own during a project, so it is quite hard to find a definitive answer for it in the short time that can be allocated to it during a thesis. Considering the wide range of possible answers, it was decided to gather ideas from all the steps performed during the project to answer it. Starting with the literature review, where it could be observed what benefits the CDE brought to other systems, and continuing with the practical experience gained from the development of the prototype application and the discussions about possible features that could be introduced in similar systems. There are generic advantages of CDEs that can also be applied in the context of digital twins, such as data centralization, data security and access control, data versioning and an overall better communication between the stakeholders of a project. If the context of DTs is specifically considered, some conclusions can be drawn about the benefits of a CDE. A graph-based CDE makes it possible to link data from different environments and different external sources. This can provide additional functionalities that can be observed in the functionalities of the prototype application, such as automatic consistency checks between the files and filtering based on different criteria. Furthermore, linking the different available files to their corresponding object and environment can help with the automation process of generating DT systems. If such an automation tool were to be implemented, it would be easy to query the data in the CDE to see what files are needed for each object and which environment. This is exemplified in the soccer robots use case by the different simulation environments, since there are both a 2D simulator and a 3D simulator available as part of the digital twin. Because of the connections made between the files and the respective environments, it is easy to see which file would be needed to generate an updated version of each part of the DT. RQ4: How to construct an automated pipeline from the CDE to the DT? This question was addressed in subsection 6.3. There was not enough time left during the project to try and create such an automated pipeline and to implement the data management services needed for the automatic generation of a DT. However, as it 51
was discussed in the answers for the previous questions, the linking of the data in the CDE for the simulation environment can prove to be very useful in the creation of an automated pipeline. By leveraging the created relationships between the different available files in a system, it should be possible for the data management service to be aware of which files are needed to create a new version of the DT, and their update status. Such an approach would also allow different automation tools to be integrated into the application. Finally, with the knowledge gained from answering all the questions, the main question about how to create a generic architecture was answered. 7.3 Future Work There are multiple ideas that can be employed in order to expand the current project that were not implemented yet, due to time constraints and the available use cases. 7.3.1 Implementation of Data Management Services An important addition would be to implement the data management services that were discussed in subsection 6.3. This would further confirm the extended architecture and ontology through a practical implementation, while better defining the interfaces between these services and the CDE and DT. The addition of the automatic generation tools would create a new functionality in the application that would further automate the process of maintaining a DT system. This would also help give a clearer answer to the last research question (RQ4). 7.3.2 Expanding the Current System Another way to further test the resulting architecture is to expand the existing system. Currently, the focus is mostly on the engineering environments, but more environments could be added to create new connections to test the scalability of the design. A similar thing could be done for the objects themselves, for example, separating the robots into multiple parts and connecting them in the graph database, which would offer more granularity and create new connections. 7.3.3 Security and Access Control From the point of view of the developed web application, security and access control are two aspects that could become a bigger focus for future work. They were not considered for this project as they were not the focus of the thesis, however, most 52
2sparql = SPARQLWrapper ( url ) 3sparql . setMethod ( POST ) 4sparql . setQuery ( query ) 5sparql . setReturnFormat ( JSON ) 6 7return sparql . query (). convert () 8 9def execute_query_post ( query : str ): 10 ’’’ 11 execute query on GraphDB repository 12 ’’’ 13 url = config [ ’GRAPHDB_URL ’] + "/ statements " 14 print (" url is ") 15 print ( url) 16 query_result = graphdb_update ( url =url , query = query ) 17 18 return query_result Listing 6: GraphDB service methods for executing SPARQL updates with the use of SPARQLWRAPPER 59