Full text
id178717 A KNOWLEDGE-GRAPH DATA LAYER FOR SOFTWARE PROJECT DASHBOARDS ADRIÀ ESPEJO CUBERO Thesis supervisor: XAVIERFRANCHGUTIÉRREZ(DepartmentofServiceandInformationSystem Engineering) Thesis co-supervisor: CARLESFARRETOST(DepartmentofServiceandInformationSystem Engineering) Degree:MasterDegreeinInformaticsEngineering Thesis report Facultat d'Informàtica de Barcelona (FIB) Universitat Politècnica de Catalunya (UPC) - BarcelonaTech 29/06/2023 DRAFT
A Knowledge-Graph Data Layer for Software Project Dashboards page 1 Resumen Los cuadros de mando (dashboards) de proyectos de software nos facilitan una visualización avanzada y técnicas de análisis (como predicciones, simulaciones, análisis de ciertas situaciones hipotéticas y de alertas). Estos operan mediante la recolección de datos de diferentes fuentes (por ejemplo, repositorios de código, herramientas de gestión de proyectos y controles de calidad), conviertiendo estos datos en métricas y combinándolas para formar indicadores. Resulta crucial el hecho de dar soporte a las capacidades de este tipo de software con una capa de datos versátil, eficiente y robusta. El propósito de este TFM es el de diseñar e implementar esta capa de datos a partir del modelo de datos Knowledge Graphs, representando los principales conceptos que los cuadros de mando mantienen y las relaciones entre ellos. Esta nueva capa de datos tendrá que ser integrada con el dashboard de Q-Rapids (acrónimo de Quality-aware Rapid Software Development), un proyecto de software implementado por el Grupo de Investigación de GESSI UPC, de manera no disruptiva. El TFM será evaluado haciendo uso de una personalización del dashboard Q-Rapids en el dominio del aprendizaje.
Report page 2 Resum Els quadres de comandament(dashboards) de projectes de software ens permeten una visualització avançada i tècniques d’anàlisi (com prediccions, simulacions, anàlisi de certes situacions hipotètiques i d’alertes). Aquests operen mitjançant la recol·lecció de dades de diferents fonts (per exemple repositoris de codi, eines de gestió de projectes i control de qualitat), convertint aquestes dades en mètriques i combinant-les per formar indicadors. Esdevé crucial el fet de donar suport a les capacitats d’aquest tipus de software amb una capa de dades versàtil, eficient i robusta. El propòsit d’aquest TFM es el de dissenyar i implementar aquesta capa de dades a partir de Knowledge Graphs, representant els principals conceptes operats pel quadre de comandament i les relacions entre ells. Aquesta nova capa de dades haurà de ser integrada amb el dashboard de Q-Rapids (acrònim de Quality-aware Rapid Software Development), un projecte de software implementat pel Grup d’Investigació de GESSI UPC, de manera no disruptiva. El TFM serà avaluat fent servir una personalització del dashboard Q-Rapids en el domini de l’aprenentatge.
A Knowledge-Graph Data Layer for Software Project Dashboards page 3 Abstract Software project dashboards provide advanced visualization capabilities and analysis techniques (such as predictions, simulations, what-if analysis, and alerts). They operate by gathering data from several sources (such as code repositories, project management tools, and quality checkers), translating this data to metrics, and next combining metrics to assemble indicators. It becomes crucial to support software project dashboards with a versatile, efficient, and powerful data layer. The purpose of this Final Master Thesis is to design and implement this data layer with a Knowledge Graph data model, thus representing the main concepts managed by the dashboard and the relationships between them. This new data layer shall be integrated into the Q-Rapids (an acronym for Quality-aware Rapid Software Development) dashboard, a particular software project dashboard implemented by the GESSI UPC Research Group, in a nondisruptive manner. The TFM will be evaluated using a customization of the Q-Rapids dashboard in the teaching domain.
A Knowledge-Graph Data Layer for Software Project Dashboards page 5 Contents 1 Introduction, Motivation, and Objectives 9 1.1 Introduction ................................... 9 1.2 Motivation .................................... 10 1.3 Tasks ....................................... 10 2 Background 13 2.1 Q-Rapids and the Learning Dashboard .................... 13 2.1.1 Context .................................. 13 2.1.2 System Overview ............................ 14 2.2 Knowledge Graphs ............................... 17 3 State of the Art 21 4 Time planning 25 5 Viability, economical analysis, and comparisons 29 5.1 Human Resources costs ............................ 29 5.2 Material costs .................................. 30 5.3 Total costs .................................... 32 6 Sustainability and social commitment 33 6.1 Economical .................................... 33 6.2 Social ....................................... 33 6.3 Environmental .................................. 34 7 Specification and Solution design 35 7.1 Current System Architecture ......................... 35 7.2 Domain Entities ................................. 40 7.3 Knowledge Graph Model ........................... 49 7.4 Data Layer Architecture ............................ 57 7.5 API Design .................................... 59
Report page 6 8 Development 63 8.1 Tools and Technologies ............................. 63 8.2 Implementation Details ............................. 64 8.3 Program Entity Implementation ........................ 69 8.4 Quality Model update Implementation .................... 73 8.5 QR-Connect Implementation ......................... 75 8.6 QR-Eval implementation ............................ 77 8.7 How to introduce a new DataSource ..................... 79 9 Evaluation 85 9.1 Experimental setup ............................... 85 9.2 Process and results ............................... 86 9.2.1 Learning Dashboard initialization and Quality Model Creation . 86 9.2.2 QR-Connect module .......................... 95 9.2.3 QR-Eval module ............................ 97 10 Conclusions and Future Work 99 11 Bibliography 101
A Knowledge-Graph Data Layer for Software Project Dashboards page 7 List of Figures 1 Quality Model representation ......................... 16 2 Quality Model steps .............................. 16 3 Gantt chart done with teamgantt.com .................... 27 4 Learning Dashboard containerized components deployed ......... 35 5 Architecture overview of the Learning Dashboard ............. 36 6 Metrics Configuration properties file for Acceptance Criteria metric . . . 37 7 Metrics Configuration query file for Acceptance Criteria metric ..... 37 8 Kibana interface for querying ElasticSearch data .............. 39 9 Project Configuration: general information ................. 41 10 Project Configuration: students ........................ 41 11 Strategic Indicator View section ........................ 42 12 Detailed view of strategic indicators ..................... 42 13 View of Strategic Indicators Configuration .................. 43 14 View of Quality Factors ............................. 44 15 Detailed view of Quality factors ........................ 44 16 View of Quality Factors Configuration .................... 45 17 View of metrics ................................. 45 18 View of Metrics Configuration ......................... 46 19 View of Products Configuration ........................ 46 20 View of Iteration Configuration ........................ 47 21 View of Profile Configuration ......................... 48 22 View of User Profile Configuration ...................... 49 23 View of Category Configuration ....................... 49 24 Ontology ..................................... 53 25 UML Class Diagram .............................. 54 26 Data Layer Architecture Diagram ....................... 58 27 Back end classes ................................. 65 28 Resources directory ............................... 66
Report page 14 is then interpreted as Strategic Indicators that decision-makers can use to plan the next steps in the development. This project began in 2016, it is coordinated by UPC and has received funding from the European Union’s Horizon 2020 research and innovation program under a specific grant agreement. Plus, its consortium consists of seven organizations from five different countries: UPC, University of Oulu, Fraunhofer IESE, Bittium, Softeam, Itti, and Nokia. 2.1.2 System Overview The Learning Dashboard was developed with the students of PES and ASW in mind and it is oriented to the learning environment and allows working with Key Performance Indicators (KPI) which are defined by a quality model composed of five levels: data sources, raw data, metrics, factors, and strategic indicators. All of this information about the quality model can be visualized through a web application. The Learning Dashboard architecture is built with three components as its core: •Connectors: Connectors are programs that communicate with each data source in different ways (like consuming their respective API REST or by reading local files) in order to provide and store the desired data in ElasticSearch, where we analyze all the information that is potentially important for computing metrics. A connector has to be implemented every time a data source is added as it is needed to define what data is retrieved for each API call toward the endpoints that provide the information. There are multiple connectors implemented for multiple sources whose nature ranges from code repository hosting services such as GitHub or GitLab to Project Management Tools like Taiga or Jira. A connector example could be the one needed for retrieving data from a specific GitHub code repository. One of the already implemented methods for this connector is to retrieve info on the issues and the commits which after doing so will
A Knowledge-Graph Data Layer for Software Project Dashboards page 15 be stored as unprocessed data on the ElasticSearch running instance where we store all the raw data that comes from the outside of the Learning Dashboard. •Quality Model: There are five different levels that form the Learning Dashboard, enumerated from the lowest to the highest level of abstraction: 1. Data Sources: Services that provide the data to compute the metrics. This data is retrieved by consuming the endpoints that are public in the scope of each project with the use of the right credentials. This data is managed by the connector for selecting the appropriate information within it. 2. Unprocessed data: this level refers to the data retrieved and managed by the connector which then is stored in an ElasticSearch instance. This data cannot be decomposed into smaller units and is used for computing the metrics for the next level. 3. Metrics: metrics become numerical values that are computed by a formula built by parameters that can be filled with unprocessed data. 4. Quality Factors: quality factors are related to metrics that can express a crucial value altogether and can be computed with a pondered average among them. 5. Strategic Indicators: strategic indicators are the highest abstraction level for representing a quality-related aspect that an entity considers important for its decision-making processes. Similar to metrics, strategic indicators are defined with a pondered average among quality factors that are tied to them.
Report page 16 Figure 1: Quality Model representation •Learning dashboard visualization: with the learning dashboard we are able to create different projects for different courses at UPC, within these projects we can define Strategic Indicators and Quality Factors and what conforms to them and how are computed. Nevertheless, metrics are defined with configuration files and are just queried to the ElasticSearch database. Strategic Indicators, Quality Factors, and Metrics values can be seen at all times. We can see how much each factor contributes to the overall entity (Metrics to the Quality Factor and the latter to Strategic Indicators) with radar charts and the current value with a graph that is segmented into different colors that represent customized states for that entity. Furthermore, we can define alerts that notify users that respond to certain entity values by defining thresholds on their creation. Figure 2: Quality Model steps
A Knowledge-Graph Data Layer for Software Project Dashboards page 17 2.2 Knowledge Graphs Knowledge graphs utilize a graph-based data model to organize, integrate, and derive insights from diverse sources of data at scale. In contrast to relational models or other NoSQL alternatives, graph-based knowledge representations offer several advantages. Graphs provide a visual and intuitive representation where nodes and edges capture intricate relationships among different elements within a domain. Unlike rigid schemas in relational databases, graphs allow data to evolve more flexibly without the need for upfront schema definition. Moreover, graph query languages incorporate navigational operators that facilitate traversing arbitrary-length pathways, in addition to standard relational operators like joins, unions, and projections. One widely used graph data model is Directed Edge-labeled Graphs (del. graphs). This model consists of nodes and directed labeled edges connecting them. Nodes represent entities, while edges represent binary relations between these entities. This flexible data modeling approach allows for seamless integration of new data sources, unlike SQL alternatives that rely on predefined schemas or NoSQL models such as documentoriented databases with fixed hierarchies. In the context of knowledge representation, RDF (Resource Description Framework) serves as a standardized data model. RDF enables the creation of statements about resources using subject-predicate-object triples. These triples form the building blocks of RDF graphs, where the subject represents the resource being described, the predicate denotes a specific property or relationship, and the object represents the value or target of the predicate. RDF, with its graph-based structure, facilitates the representation of complex relationships and the integration of data from various sources. By leveraging knowledge graphs and RDF, organizations can unlock the potential of their data by establishing a comprehensive and interconnected knowledge base. This enables advanced data exploration, information retrieval, and reasoning capabilities, driving insights and decision-making in domains ranging from healthcare and finance to e-commerce and social networks.
Report page 18 RDF provides various strong points related to other data models: •expressiveness: Data structure, taxonomies and vocabularies, all types of metadata, reference and master data, and other types of information can be fluently represented using the Semantic Web stack standards RDF(S) and OWL. •Performance: It is simple to model provenance and other structured metadata thanks to the RDF* extension. Every specification has been carefully considered and tested to ensure the effective administration of graphs containing billions of facts and characteristics. •Interoperability: Data serialization, access (SPARQL Protocol for end-points), and management (SPARQL Graph Store) have a variety of specifications. Globally unique identities make publication and data integration easier. •Standardization: To ensure that the needs of various actors are met, all of the aforementioned are standardized through the W3C community process. In this project, we will represent RDF triplets or statements like in Table 1. Every RDF statement is made of Subjects (Robert or PersonResource in general), Predicates (’has a’), and Objects (pencil). Subjects are resources or entities and they are represented by an URI or a blank node and we will represent them visually with appending the name of the entity plus Resource. The subject is usually who or what the statement is about. Properties or predicates represent the relationship between the subject and the object, They are represented also by an URI. Properties are also described by URIs. Finally, objects represent the value of a property in an RDF statement. It can be a literal value like a string, a number, a boolean, etc., or another resource identified by a URI. In this case we identified ’pencil’ as a literal but, if needed by the domain we are working with, we can create a entity class Material and describe characteristics about this material like its name (this would be a literal ’pencil’), its weight (a float) a ’handedTo’ object property that informs the user to who this material has been handed. The possibilities are inifnite and all depends on your final application.
A Knowledge-Graph Data Layer for Software Project Dashboards page 19 Subject Predicate Object Robert writesWith ’pencil’ Robert isFriendWith Laura Table 1: RDF statement schema Ontologies are formal frameworks or specifications that define a domain’s shared understanding. They define the concepts, entities, connections, and rules within a certain domain or subject area to provide a systematic and ordered representation of knowledge. Ontologies are critical in capturing and conveying the semantics of information in the context of knowledge representation and the Semantic Web. They serve as a common lexicon for successful communication and interoperability among various systems, applications, and individuals. Ontologies are often made up of three major components: •Concepts and Classes: Ontologies define concepts or classes that represent the entities or types of objects within a domain. These classes describe the characteristics, attributes, and behaviors associated with the entities they represent. Concepts can be arranged in a hierarchy, with more general classes at the top and more specific subclasses below, forming a taxonomy or ontology structure. •Relationships and Properties: Ontologies define relationships or properties that describe the associations between entities or concepts. These relationships capture the various ways in which entities interact, connect, or depend on each other. Relationships can be simple, like "is-a" or "part-of," or they can be more complex, representing specific domain-specific associations. •Axioms and Constraints: Ontologies can also include axioms and constraints that express logical rules and constraints within the domain. These axioms provide additional knowledge and enable reasoning about the information contained in
Report page 20 the ontology. They can be used to infer new knowledge, validate data consistency, or enforce domain-specific rules. The formal specification languages used to create ontologies include RDF Schema (RDFS), Web Ontology Language (OWL), and others. These languages provide a syntax and semantics to define the structure, relationships, and rules within an ontology.
A Knowledge-Graph Data Layer for Software Project Dashboards page 21 3 State of the Art Knowledge Graphs are widely used in contemporary Knowledge Representation Learning. Its primary application involves generating information encodings for neural networks, as demonstrated in the work by Jiacheng Xu et al. on "Knowledge Graph Representation with Jointly Structural and Textual Encoding" [7]. Knowledge Acquisition, another important area, focuses on constructing knowledge graphs from unstructured text and other structured or semi-structured sources. Xu Han et al. explored this field in their paper on "Neural Knowledge Acquisition via Mutual Attention Between Knowledge Graph and Text" [8]. Additionally, temporal knowledge graphs have been investigated, as evidenced by Chenjin Xu et al.’s research on "Temporal Knowledge Graph Completion Based on Time Series Gaussian Embedding" [9]. Numerous research papers have addressed problems using knowledge graphs. Notably, the work of Antonio Messina et al. on "BioGrakn: A Knowledge Graph-Based Semantic Database for Biomedical Sciences" [10] is noteworthy. This study integrates information from various data sources and aggregates it into a knowledge graph containing biological data. One of the most related papers we found was Virtual knowledge graphs (VKGs): An Overview of Systems and use cases [11] by Xiao et al. VKGs also known as Ontologybased Data Access, offer a flexible approach to data integration and access. Instead of using rigid relational tables, VKGs utilize virtual and graph-based structures, enriched with domain knowledge. The VKG approach combines three key ideas: Data virtualization: VKGs provide a conceptual representation of the domain, known as a global schema, which is familiar to end-users. Integration views are created as suitable views over the data sources, defining the information content of the schema. These views are virtual and not materialized, enabling efficient querying without extensive storage or time-consuming data accessibility. Design and maintenance are simplified
Report page 22 since views can be instantly tested and modified. Graph-based data structure: VKGs represent data as graphs, where nodes represent domain objects and data values, and edges encode properties of objects. This graph structure offers greater flexibility than traditional relational tables, allowing easy integration by merging graphs and handling situations where nodes represent the same real-world entity with different identifiers. The graph representation allows for straightforward integration while preserving object identity. Enrichment with domain knowledge: VKGs are enriched with domain knowledge, such as concept hierarchies, property information, and mandatory properties. In summary, VKGs provide a data integration and access paradigm that utilizes virtualization, graph-based data structures, and the enrichment of domain knowledge. An important paper that guides the development of this project is Converting Relational to graph databases [12] by R. de Virgilio et al. This paper proposes a methodology for converting relational databases to graph databases by leveraging the schema and constraints of the source system. Our project draws inspiration from this paper to convert the data layer of software project dashboards from a relational to a graph database, also non-relational (document-oriented, ElasticSearch in this case). This conversion enables us to utilize the connectivity and scalability of graph models, providing an effective and efficient solution for data storage and query answering in software project dashboards. In our project, we aim to design and implement a data layer for software project dashboards using the Knowledge Graph data model. This data layer will serve as a versatile, efficient, and powerful foundation for advanced visualization capabilities and analysis techniques within the dashboard, such as predictions, simulations, what-if analysis, and alerts. Building upon the concepts presented in the Virtual Knowledge Graphs paradigm and a base knowledge for the latter mentioned paper we will leverage the principles of relational data translation, graph-based data structures, and enrichment with domain
A Knowledge-Graph Data Layer for Software Project Dashboards page 23 knowledge. By adopting these ideas, we can create a data layer that offers a conceptual representation of the software project domain, virtual integration views over multiple data sources, a flexible graph-based structure to capture relationships and properties. Our solution is based on a subset of the Learning Dashboard project, developed by the GESSI UPC Research Group, and we will implement a data layer implemented with Knowledge Graphs that will further grow in the future and be documented for expansion on the different modules. By leveraging the power of the Knowledge Graph data model and drawing inspiration from both last papers mentioned.
Report page 30 Step Period Days Hours RA hours Dev. hours 1 6/2-17/2 12 76.56 76.56 - 2 18/2-23/2 6 38.28 - 38.28 3 24/2-10/3 15 79.75 39.875 39.875 4 6/3-20/3 15 79.75 - 79.75 5 21/3-17/5 58 179.171 89.58 89.58 6 11/4-17/5 37 70.71 - 70.71 7 3/5-24/5 22 46.25 23.125 23.125 8 10/4-18/6 70 278.061 278.061 - 9 19/6-26/6 8 51.04 51.04 - Total - - 900 559 341 Table 2: Hour management between roles After having computed the hours for each role, in Table 3we see the total cost of the Human Resources involved in our project. With the average salary per year of both roles in Barcelona area [13][14] we can deduce the hourly salary and with it the total cost that would suppose each of the profiles which makes a total of 17,174 euros. Role Hours salary/year(€) salary/hour(€) Total cost(€) Research Analyst 559 43,000 20,67 11,554.53 Software Developer 341 34,278 16,48 5,619.68 Total 900 - - 17,174.21 Table 3: HR Economic value 5.2 Material costs For the material resources, a computer with the following components is used during the entirety of the project development, 900 hours: •Intel(R) Core(TM) i7-10700K CPU 3.80 GHz
A Knowledge-Graph Data Layer for Software Project Dashboards page 31 •Corsair 16.0 GB RAM •Samsung M.2 1TB SSD •NVIDIA GeForce RTX 3080 12GB VRAM Plus, other office material, including: •Amazon’s Magnetic Board 40" •MSI MPG27CQ2 (Monitor 1) •MSI MAG274QRF-QD (Monitor 2) For the economic value of the material, we will need to get the Amortized Cost (AC) of each item. The formula for computing it is AC = D/D * Number of Days. Number of Days is 141 in our case. For knowing the rest, we would require the following: •The Base Value (BV) of the item. This is the item cost that we paid at the beginning. •The Residual Value (RV) which is the cost of our item at the end of its Useful Life (UL), we interpret it to be a 10% of the BV for this exercise. •The Depreciation per Day (D/D), that is, the value that our item loses each day. The formula that defines it is D/D = (BV - RV) / UL. In Figure 4we defined each value mentioned for each item, and the total Amortized Cost is 239.7 euros.
Report page 32 Item BV(€) RV(€) UL(d) D/D(€) AC(€) PC 2000 200 1460 1.23 173.43 Monitor 1 500 50 1825 0.24 33.84 Monitor 2 400 40 1643 0.21 29.61 Magnetic board 100 10 3650 0.02 2.82 Total - - - - 239.7 Table 4: Material Economic value 5.3 Total costs Therefore, the sum of both Human Resources costs and Material costs would be equal to 17,413.91 euros. Item Total value(€) HR 17,174.21 Material 239.7 Total 17,413.91 Table 5: Total Economic estimation
A Knowledge-Graph Data Layer for Software Project Dashboards page 33 6 Sustainability and social commitment Sustainability and social commitment’s importance has grown over the years and its analysis is currently taken for granted in the project development and research sectors. At the end of the day, people involved in projects like these, even if they are small-scale, have to realize that poor management and understanding of the mentioned can have a more critical environmental, social, and economic impact if scaled up, so it is best to realize how to make sure these impacts are positive and sustainable over the long term. We will go over the numerous ways that our project is dedicated to sustainability in this section. 6.1 Economical In order to understand the economic sustainability of the project, Table 5is important for interpreting the inequality income > costs as a positive economical state would indicate a good index of economic sustainability. Nevertheless, the income variable cannot be interpreted in a typical way as the Learning Dashboard is a tool that won’t be sold as a product but more as an open-source project. Plus, the total costs associated with the Learning Dashboard component for Q-Rapids are paid by Horizon 2020 research program. All in all, the economic situation can deviate due to time and poor planning. Nonetheless, this can not happen easily as the situation is studied and ready to be handled at all times. Furthermore, the costs associated with this project can be reduced by using less office material, a lower-end computer (as the specifications are too high for this project), and more experienced human resources for lowering the study and preparation time for some domains that had to be researched. 6.2 Social The further development of the Learning Dashboard will help different teams to adapt Rapid practices to their projects. Nokia or Softeam are the most interested in using Q-
Report page 34 Rapids in their projects as they are participants in this research project. Related to the specific part that has been developed in this project, this will be used for comparing different data models for the domain of the Learning Dashboard and being able to define which one is the most suited for accomplishing it. Knowledge Graphs uses a data model that can change easily; therefore we developed the first steps of a solution that can be adapted to different metric domains. Personally, I learned many things about Knowledge Graphs. It is a data model that I learned in ECSDI, an optional subject for the software branch in my bachelor’s degree. But I never had the chance to experience more with it. Plus, doing a project from the ground helps to improve the way you are dealing with future programming problems in the future. 6.3 Environmental For this project a computer is used with the specs mentioned in section 5, thanks to Outervision’s [15] power supply calculator, my PC consumes roughly 500W which represents 450kWh for the project ((500W∗900h)/1000 = 450kWh). To this, I have to add the use of numerous sheets of paper which I used to model the domain when not writing in the magnetic board. Regarding the code, as it is one of the requirements for this project, was written in a modular way that can reuse domain logic for the service layers and only substitute the repository layers for implementing the operations based on the database that is being used.
A Knowledge-Graph Data Layer for Software Project Dashboards page 35 7 Specification and Solution design In this section, we discuss the proposed solution for this project. First, we will explain the current Learning Dashboard system architecture, naming all its components and showing the operations that we cover in detail. After, we show the knowledge graph model that our data layer will use to carry out the operations. Following with the data layer architecture as a whole and our API design approach for the solution. 7.1 Current System Architecture Thanks to previous work done in Specification and design of a dashboard for monitoring the learning process in software projects developed by teams of students master thesis project[16] we replicated the production environment for the Learning Dashboard in our local machine using Docker[17] containers. In Figure 4we can see the Learning Dashboard environment deployed in the local machine. Figure 4: Learning Dashboard containerized components deployed We can observe the overall architecture of the Learning Dashboard that we cover in this
Report page 36 project in Figure 5. Figure 5: Architecture overview of the Learning Dashboard This diagram includes the following components: •Remote DataSources: sources of the data that feed the metrics in the Learning Dashboard. These sources are hosted remotely and are meant to be retrieved with an API among other methods. Examples of data sources can be from Git-hosted remote repository services like GitHub or GitLab to project management tools like Taiga or Jira. Within the scope of each data source, there are numerous objects that can be queried for our metrics computing: commits and issues or user stories and tasks respectively. •QR-Connect: a component that contains the retrieving logic of the raw data from each data source. There exist different ways of handling the data retrieval depending on each data source but commonly this process is done thanks to a REST
A Knowledge-Graph Data Layer for Software Project Dashboards page 37 API implemented by the development team for each web application. After this process, we select the appropriate subset of data that we need for our metric computing and we store all the records in an ElasticSearch data sink following a certain index grouped by data source and project. •QR-Eval: Learning Dashboard component whose task is to compute the metrics specified by configuration files that from now on we will call properties and query files. These files must be declared for defining the properties of an instance of a metric in the Learning Dashboard and how we can compute it. In figures 6and 7 we can see the shape of the property and query files respectively. Figure 6: Metrics Configuration properties file for Acceptance Criteria metric Figure 7: Metrics Configuration query file for Acceptance Criteria metric
Report page 38 For the properties file we define the metric attributes like the name, description, the weight that is going to be used for computing the pondered sum between all metrics that compose a quality factor and the quality factor that it belongs to. The index where the query is called on; indices are grouped by project, data source, and data source object, ex: ASW11, Taiga, Userstories. The remaining properties are parameters assigned based on the query result on the index in the query file. In the query file figure, we query those user stories that do not have a closed milestone with acceptance criteria and we aggregate them into a field called acceptance _criteria _user stories which is referenced by UserStoriesWithAC declared property where the value is stored in. The other parameter references how many user stories with closed milestones are in the index. In the end, these parameters are used for the metric formula and just in case, we have a default value to set in case of error. •ElasticSearch data sink: NoSQL, document-oriented database that stores all the data source objects retrieved by the instantiated connectors and gathers them over indices. Thanks to Kibana we can visualize our data easier with a clear interface. We query data based on Query DSL query language, which is based on JSON and it is ideal for managing big chunks of documents. In Figure 4there are two containers related to this component, ingest_elasticsearch which is the ElasticSearch instance deployed in port 9200, and ingest_Kibana which is a client for visualizing the ElasticSearch data in the previous component. In Figure 8we observe a simple query for getting all records for index github_pes_j12a.commits, which is to group all commits for project pes j12a and data source github.
A Knowledge-Graph Data Layer for Software Project Dashboards page 39 Figure 8: Kibana interface for querying ElasticSearch data •Learning Dashboard application: Core of the Learning Dashboard application, where all the business logic surrounding the quality model is managed. This program interacts with the ElasticSearch data sink and PostgreSQL. The former is used mainly for querying computed metrics while the latter stores all the other entity information of the Learning Dashboard. This is the component that the final user uses for checking the current evaluation of the quality model. The main application runs in a Tomcat web server and Figure 4can be seen as qrapids_tomcat deployed container. •PostgreSQL database: data repository where the information regarding all entities of the Learning Dashboard are stored such as alerts that informs us if a quality model object has surpassed a certain threshold, profiles, users, iteration information, etc. In summary, it contains all the data necessary to manage the Learning Dashboard component, excluding data handling which is done by other components using ElasticSearch as the support repository. Again, as seen in Figure 4,
Report page 46 Figure 18: View of Metrics Configuration •Products: Object that groups the different projects of interest, one example of a product can be all the different projects for the ASW subject. In the following view in Figure 19 we observe the insertion of a Product, its base information, and the Projects that it is composed of. Figure 19: View of Products Configuration •Iterations: Entity that defines a period of time-related to certain projects. Itera-
A Knowledge-Graph Data Layer for Software Project Dashboards page 47 tions are periods during which a development team works on a set of predefined tasks or user stories. In Figure 20 we see the insertion view of this entity in the Learning Dashboard database, In here we complete certain information like the beginning and the end dates and what projects this Iteration applied to. Figure 20: View of Iteration Configuration •Profiles: Profiles inform us about certain configuration parameters related to a User like chart visualization parameters, permissions, and so on. In Figure 21 we see the creation of a Profile, where we input in Step 1 the profile information, in Step 2 profile permissions like which level of detail the User can see the Quality Model objects and also which Strategic Indicators and Projects are allowed to see; ending with Step 3 where we set visualization options of different elements of the Quality Model hierarchy or the whole Quality Model. These visualizations can be seen at the right corner of the screen in a number of Figures like Figure 17.
Report page 48 Figure 21: View of Profile Configuration •Students: Those who are assigned to 1 or more projects, they can be set in the view of Figure 9. For each student, it is configured that they can set their Taiga and GitHub usernames at the moment. •Users: It defines all the account information in our program.
A Knowledge-Graph Data Layer for Software Project Dashboards page 49 Figure 22: View of User Profile Configuration •Categories: Type that creates an evaluation template for Quality Model objects. For example, in figure 14 we observe different kinds of evaluations for each metric expressed by gauges in this case, and in Figure 23 the creation of one of these Configuration entities, which are made of the Category name and an array of thresholds mapped to a color and a type which states what this value range is supposed to mean in a specific context. One example is the Fulfillment of Tasks metric, composed of 3 thresholds, 1 assigned to red, orange, and green which at the same time gives information about the value connotation (it is desirable the gauge pointed to green as we want that everyone does their tasks). Figure 23: View of Category Configuration 7.3 Knowledge Graph Model In order to translate the entity relational data model, and the QR-Connect and QR-Eval processes which the ElasticSearch instance is a data repository for both processes, into a Knowledge Graph data model we followed the steps documented in Neo4J’s site [18] and W3School’s [19]. This transformation process also has to take into consideration
Report page 50 how the data is going to be queried and inserted and try to keep related entities as close as possible to not create performance issues. The ontology creation has followed an iterative approach, instead of a cascading process. Ontologies aspire to extend or change at some time. We applied changes to the ontology even at later stages during the data layer development of the Learning Dashboard and it has been changed according to new requirements that were not taken into account at the moment or by the director’s criteria. The repeatable steps to create an ontology are the following: 1. Study the domain entities: With an established baseline of all the candidate entities. 2. Transform primitive attributes to data properties: Each document field in the ElasticSearch and every column in any table that is based on primitive values like Strings, Booleans, or Floats can be transformed to data properties inside our Knowledge Graph. 3. Transform entity references to object properties: Data models usually contain references to other entities with different document shapes or other tables, these references between entities with Knowledge Graphs are translated into RDF triples with their subjects referencing a Node instead of being a literal value. 4. Create extra nodes: There are situations when it is appropriate to create additional nodes: •N-Ary relations [19]: N-ary relations involve two or more entities where a simple binary association is not enough to represent them. In this project, instances of these are Memberships (Student - DataSource), WeightedQFItems (SIItem - QFItem), and WeightedMetricItems (QFItem - MetricItems). As said, for these situations where an associative relation is taking place, an easy way to handle this is by inserting a new entity between both nodes. •Avoiding repeating information: By not cluttering the database with the
A Knowledge-Graph Data Layer for Software Project Dashboards page 51 same information every time we insert certain items we decided to store some repeating information in "general" objects that can be reused across different projects. Examples are QF and QFItem and SI and SIItem, where it is unnecessary to repeat basic information for different Quality Factors and Strategic Indicators over all the projects that will follow a similar quality model structure. 5. Analyze: Check query building when building the Data Layer and think of different ways to organize the entities, at the same time make sure to create relationships for critical queries if needed. 6. Iterate: Go to step 1 until getting a complete and optimal Knowledge Graph for our purpose. This ontology has been created with a possibility of extension in the future and it has evolved up to this final version. We will enumerate all the classes, stating all their object and data properties and answering why we define these this way. The majority of the domain entities can be easily mapped, even though there are some exceptions. In figure 24 we can identify all the nodes and relations that our ontology is composed of. Plus, there is a UML representation in figure 25 where more information can be seen such as the multiplicities of each relation, and a more visual approach for stating the data properties for each class.
Report page 52 The main classes for the ontology are the following: •Product: An ontology entity that maps to the Product entity in the current Learning Dashboard. A Product contains self-descriptive information and has an object property hasProject with a Project range (accepts only the Project entity for this relation). •Project: A core entity connected to various information. Each project has one or more DataSource instances where its data source information is stored, Students who participate in the projects, and SIItems, which establish a connection between a Project and the different Quality Models’ roots. •Iteration: Represents an Iteration domain entity. An Iteration instance can be connected to a Project. This relation allows retrieval of all the Iterations associated with a Project. •SI: In our knowledge graph, SI instances are used to reuse information across different instances of strategic indicators. We aim to avoid duplicating the same information about a Strategic Indicator repeatedly and instead promote easy reuse across multiple projects. Each SI may be connected to one or more QF nodes. •SIItem: SIItem is the class that represents a specific instance of a Strategic Indicator. Whenever a strategic indicator is needed for a Project, we create an instance of this class with a relationship to an existing source SI. SIItem stores key values for Quality Model management, such as the current value and a specific threshold. Additionally, each SIItem is connected to a Category, which determines how its numerical value is interpreted. SIItems are also related to at least one WeightedQFItem. •QF: Similar to the SI ontology entity, QF represents general and reusable information for Quality Factors. •QFItem: Similar to SIItem, QFItem refers to different instances of a specific Quality Factor. QFItems are connected to source QF nodes to retrieve static information
A Knowledge-Graph Data Layer for Software Project Dashboards page 53 Figure 24: Ontology
Report page 54 Figure 25: UML Class Diagram
A Knowledge-Graph Data Layer for Software Project Dashboards page 55 about a Quality Factor. A relation is established to define the value interpretation using a Category instance. QFItems are related to MetricItems through a WeightedMetricItem relation. •MetricItem: A class representing a particular Metric. Metrics are not generalized like QF and SI, as they are not reusable across projects. A MetricItem is connected to a Category, similar to its parent entities in the Quality Model. To update a metric value, a relation to a DataSource node is established to determine the source from which the base information for computing the metric should be read. Additionally, a MetricItem holds dynamic data properties such as its value, as well as static information that defines how the metric is computed, including its associated method name and an optional target data property (e.g., a username for methods like "retrieve the number of tasks created by [username]"). •WeightedQFItem: This class represents an associative relation between SIItem and QFItem. Since a SIItem may have one or more different QFItems, it is necessary to specify the weight that a QFItem carries for a SIItem. •WeightedMetricItem: Similar to WeightedQFItem, this class represents the associative relation between QFItem and MetricItem. •Category: Each Quality Model hierarchy object can be associated with an interpretation of its performance based on its value. The Category class provides this information and can be reused across multiple instances within the hierarchy. Category nodes consist of one or more CategoryItems. •CategoryItem: CategoryItems are items directly associated with Categories. Every Category must have a set of CategoryItems that define different thresholds for value interpretations. For example, a Category called "Anonymous Commits" may have three CategoryItems indicating three value range interpretations. These CategoryItems contain data properties such as the name, color, and upper threshold to visually represent these values.
Report page 62 •SIItem : – GET By Project: We want to see each SIItem for a Project as reflected in Figure 11. If we look at the UML diagram in Figure 25, there are some entity nodes like WeightedMetricItem, WeightedQFItem, and Membership that we did not mention and are expressed as associative relations, which means that they are highly dependent on both the entities that are partaking in that same relation. It has been decided that information regarding such entities is embedded directly into QFItem, SIItem, and Student GET responses. For integrating the behavior of QREval and QRConnect components, we created two separate modules that replicate the most important features, as our objective is to gather all this logic and information into one data layer powered by a Knowledge Graph, that handles the processes comprising the data retrieval of the different DataSources up to managing to compute all the metrics, update the Quality Model accordingly on demand and querying for information regarding metrics foremost. Furthermore, DataSource-specific objects can not be retrievable at this moment as in the current Learning Dashboard we are not able to see them either and there is no justification enough to do so, save from monitoring.
A Knowledge-Graph Data Layer for Software Project Dashboards page 63 8 Development We go deeper into how we implemented the components that are explained in the last section, naming the technologies used and finishing with a brief subsection on Challenges and Solutions. Our version of the Learning Dashboard can be found in a GitHub repository in the following URL: https://github.com/calandula/LearningDashboard. In the root of this repository we can find a README.md file that explains how to install the project in any local machine, run it and see the results, also we attached an auxiliary Postman collection file for testing it. 8.1 Tools and Technologies The following tools have been used for the development of this project are: •Draw.io[21]: a web-based tool for building flowcharts, diagrams, and more visual representations. Figure 26 or 24 are examples. •Overleaf[22]: Online collaborative LaTeX editor. This report is written and reviewed on this platform. •IntelliJ Community Edition[23]: one of the numerous IDEs from JetBrains, this one is more focused on developing Java applications. This code editor was chosen for the development of the web service. •Apache Jena[24]: Java framework for building semantic web and linked data applications. It provides APIs for working with RDF data. This tool was used to build the data layer of the main application of this thesis. •Github[25]: a web application for version control and collaboration for hosting and managing Git repositories. The repository hosting the application code of the thesis is on GitHub. •Trello[26]: Trello is a web-based application for keeping track of the progress of projects by using cards and managing them in the scope of multiple categories.
Report page 64 •Docker[17]: Docker is an open-source platform for containerization, enabling the creation and deployment of applications in lightweight, isolated containers. It provides a way to package an application with its dependencies into a standardized unit. Docker was used for replicating the production environment in a local machine by deploying all the needed services of the Learning Dashboard. •Postman[27]: the main tool for testing the correct behavior of the application. Postman is used for HTTP requests and examining responses. •Protégé[28]: open-source ontology editor and knowledge management system. Provides a user-friendly interface for editing them. This tool was used for building our ontology in Figure 24. 8.2 Implementation Details After designing the ontology of our data to store the information, we develop the main web application for doing operations to the data at hand. For doing this task we use IntelliJ IDE. We use Java as a programming language, as the Apache Jena toolset is made of different APIs for interacting with Knowledge Graphs - oriented applications. This set of libraries is crucial for our project as it is the way of interacting with the triplestore dataset in a high-level way. The project, called Learning Dashboard like the original, follows the architecture described in section 7.4 which creates a division of responsibilities between different layers, there is the service layer which is where the business logic is implemented, and the repository layer which can be easily interchangeable with another other repository implementations for different databases like PostgreSQL or MongoDB. In Figure 27 we can see the Learning Dashboard project structure. This application is a Maven project and uses Apache Spring which is a powerful framework for building enterprise-grade Java applications. Spring has a set of libraries that helps to promote good software practices. Spring has support for easily configuring a REST API following the three-layer architecture described in section 7.4, granting the ability to create
A Knowledge-Graph Data Layer for Software Project Dashboards page 65 modular code with different layers of responsibility and use dependency injection for managing object dependencies and promoting loose coupling between the different services. Figure 27: Back end classes
Report page 66 Figure 28: Resources directory A description of each of these modules is found in the images: •config: This module includes a global configuration class for the whole application. In JenaConfig we define 3 beans which can be seen in Figure, which in the context of Spring are objects that can be injected into all the services. Information like the data set that we use for storing all of our triple stores, the namespace for our ontology, and prefixes for doing SPARQL queries are initialized following the body of these beans, which in essence are functions that declare how the parameter should be initialized at the beginning of each execution.
A Knowledge-Graph Data Layer for Software Project Dashboards page 67 Figure 29: Resources directory •controller: Where all the controllers in the Learning Dashboard reside, in this module there are all mentioned entities’ controllers and the ones that grant the user to extract data (QR-Connect) and to compute it for updating a metric (QREval). •data source_model: Where the classes for all the DataSource objects are implemented, if there was a new object for the existing DataSources or for new ones it will be stored in this module. •dtos: DTOs stands for Data Transfer Objects which are instances of classes that encapsulate information for passing to different services. We have a unique DTO
Report page 68 for an entity when inserting and updating it. We can see DTOs in this context as data that the user will pass to the program and will also receive. There could be more DTOs for different operations but we kept it simple and bet for only one, a the moment. •repository: This directory groups all the repositories for the different entities. •service: This module gathers all the service layers where the business logic is implemented. •utils: In the utils module we see some functions that are needed over the program and their logic is repeated numerous times, and class implementations of Serializable for different fields of DTOs that their input is composed of an object that refers to an associative class. •MainApp.java: Main class and entry point of the Spring Application, we run this file in order to execute the program. •resources: Directory where items like the TDB database and the ontology file that was exported for Protégé are stored.
A Knowledge-Graph Data Layer for Software Project Dashboards page 69 8.3 Program Entity Implementation Figure 30: General Sequence Diagram In this subsection, we present how every query, update, delete, or insertion is made in our project by using the Category entity as our example. With the sequence diagram shown in Figure 30 we have a reference on how the application flow goes between classes in the majority of operations. The request starts with the controller, where, if needed, a DTO is passed (Figure 31). The function of the controller mapped to the URL, as observed in Figure 32, that user sends an API REST call is the one that is going to be executed. In the mentioned figure, the URL is baseURL/api/categories; this route maps to getAllCategories, the controller only handles the input request and output responses so it directly delegates the responsibility to the service.
Report page 70 Figure 31: Category DTO Figure 32: Category Controller CategoryService handles the business logic, which in this case is not that difficult as
A Knowledge-Graph Data Layer for Software Project Dashboards page 71 we only have to return all the nodes in the Knowledge Graph that are instances of class Category. For implementing this query, which may vary depending on how we implemented our data layer and which data model we chose at the end, we call the appropriate repository class, in this case, we are doing operations over a TDB database which works on RDF triples, therefore all the implementation would be in the CategoryRepository class, regarding the interaction between the repository where we store our information. This class can be seen in Figure 33. Figure 33: Category Service CategoryRepository’s task, as we see in Figure 34, is to interact with Apache Jena API for querying data in our TDB data set, represented by an injected variable defined in the Configuration class of the program. If we recall, a TDB data set is made by an array of RDF Statements, that is composed of Subject, Predicate, and Object. Thanks to the ontology we loaded into our project when we execute it, the sentences related to object class, relationships, and properties are already instantiated. When we save a Category or any other object, we must assign their class in order for us to execute successful queries over the data set, assigning a Resource to a class by the type RDF property. So, we can do queries over Statements that their Subject has a Property of type and its Object is name of the class, for us then to retrieve the Subject Resource which will be a Category instance and we can access to all its connected objects and data properties, again by having the resource, putting a fixed property and obtaining the Object.
Report page 78 Figure 36: QR-Connect sequence diagram
A Knowledge-Graph Data Layer for Software Project Dashboards page 79 8.7 How to introduce a new DataSource One of the main characteristics of our program is that it can be flexible and different DataSource entities can be added over time. For this, we detail how to introduce a new DataSource entity and what changes and implementations must be done. Here are the steps to do so: 1. Create individual classes for each DataSource object we would want to retrieve in the datasource_model module. It is recommended to use libraries that make this process easier like Lombok Annotations, where getters, setters, and full plus no arguments constructors are implemented on the fly. Figure 38: Commit class example 2. Create a new Repository for encapsulating all the new DataSource logic. In this file, we will instantiate as static variables: •Supported objects (ex. Commits, Issues for GithubEntitiesRepository or Tasks and Userstories for TaigaEntitiesRepository) for having them easily read by our IDE and be easily traceable.
Report page 80 Figure 37: QR-Eval sequence diagram
A Knowledge-Graph Data Layer for Software Project Dashboards page 81 •Supported methods (ex. Count total tasks, average number of tasks, percentage of issues created by the student, ...) •base information for accessing the source, like a PATH or a base URL for reaching different entities that will be used across the program. Figure 39: Static Repository class variables 3. Create supportsObject and supportsMethod methods which will return true if the passed object name or method name is among the ones supported by this data source. 4. Implement retrieveData method that will be called by QR-Connect Service, where given an objectName and a DataSource id it branches on different retrieve and save implementations. One would need to implement two functions for each entity, one for gathering all the retrieval logic for the desired object, where we call the client (or read a CSV file, for example), and parse the object to a model instance. The other two functions are save functions and will follow each of the retrieval methods. This save functions contain the algorithm for saving the model object into our repository, as infinite possibilities of objects and relations can be made between entities, we found it appropriate for the user to be the one who implements this.
Report page 82 Figure 40: Retrieve data method overview 5. Implement computeMetric method. This method gets called by QR-Eval where a given DataSource id, a supported method, and an optional parameter are provided. Based on the given method, we will branch to different user-implemented algorithms and return the value given by the computation.
A Knowledge-Graph Data Layer for Software Project Dashboards page 83 Figure 41: Compute metric method overview
A Knowledge-Graph Data Layer for Software Project Dashboards page 85 9 Evaluation In this section, we explain all the preparations done in order to validate our solution comparing it with the current Learning Dashboard. After that, we simulate all the different operations of interest in our software. 9.1 Experimental setup Our implementation process will be validated by simulating similar interactions of our Learning Dashboard with respect to the one deployed in production. The processes that we are validating are the following: 1. Correct creation of the Quality Model and the needed entities to instantiate it. 2. Correct behavior of the implemented QR-Connect module. 3. Correct behavior of the implemented QR-Eval module and its Quality Model update. We are doing this by using the Postman collection that can be found in the thesis project repository [6]. For every post operation, we review the data integrity and correctness. This will be shown here with a Subject, Predicate, Object table to represent how the triples are being stored in our TDB data set.
Report page 86 Figure 42: Postman Overview 9.2 Process and results 9.2.1 Learning Dashboard initialization and Quality Model Creation We are going to simulate the different operations in out software and try to initialize all the entities in our data. Furthermore, we demonstrate the main implemented features of the Learning Dashboard. In order to initialize all entities in the Learning Dashboard we have to do it in this order: 1. In order to initialize the different categories in the program, first we need to ini-
A Knowledge-Graph Data Layer for Software Project Dashboards page 87 tialize all the CategoryItem instances that are going to compose a Category. In the following table we show 3 examples of category items and then a category that associates to the previous ones. Subject Predicate Object CategoryItemResource ’Type’ ’low’ CategoryItemResource ’Color’ ’#FF0000’ CategoryItemResource ’UpperThreshold’ 0.3 CategoryItemResource2 ’Type’ ’medium’ CategoryItemResource2 ’Color’ ’#FFFF00’ CategoryItemResource2 ’UpperThreshold’ 0.6 CategoryItemResource3 ’Type’ ’high’ CategoryItemResource3 ’Color’ ’#00FF00’ CategoryItemResource3 ’UpperThreshold’ 1.0 Table 7: CategoryItem sample instances Thanks to the semantic of Knowledge Graphs, we can associate a Category by an object property, each CategoryItem can be referenced by infinite categories (or the number that the ontology defines), so instead of adding a row in a relational model with redundant data, in here we just need to establish a connection between resources. In the ontology there can be restrictions like each Category has to have at least one relation with a CategoryItem and no more than three, as there is a point where having more seems unclear for the user interacting with the front-end which will be calling our API. In the table, we reference different CategoryItemResources for making it clear how a normal case would unravel, nevertheless in the following tables we will not show all the information of all the Objects a Subject is connected to for brevity.
Report page 94 Subject Predicate Object ProfileResource ’Name’ ’Profile test’ ProfileResource ’Description’ ’Profile for testing’ ProfileResource ’qualityLevel’ ’All’ ProfileResource ’detailedStrategicIndicatorsView’ ’Radar’ ProfileResource ’detailedFactorsView’ ’Radar’ ProfileResource ’metricsView’ ’Gauge’ ProfileResource ’qualityModelsView’ ’Graph’ ProfileResource ’allowedProjects’ ProjectResource ProfileResource ’allowedStrategicIndicators’ SIResource Table 19: Profile sample instance Subject Predicate Object IterationResource ’Name’ ’TFM iteration’ IterationResource ’Subject’ ’TFM project’ IterationResource ’From’ ’2023-06-22’ IterationResource ’To’ ’2023-07-01’ IterationResource ’associatedProject’ ProjectResource Table 20: Iteration sample instance Subject Predicate Object UserResource ’Username’ ’bobtheprofessor’ UserResource ’Email’ ’[email protected]’ UserResource ’Admin’ false UserResource ’securityQuestion’ ’Best data model?’ UserResource ’Answer’ ’KGs’ UserResource ’Password’ ’bobLD’ Table 21: User sample instance
A Knowledge-Graph Data Layer for Software Project Dashboards page 95 9.2.2 QR-Connect module QR-Connect is used to retrieve data from different data sources. QR-Connect is executed as an API REST method and its process has been explained in section 8. Here is an example of how a GitHubDataSource instance would wind up after having called QR-Connect for retrieving all of the commits and issues (imagine we have two of each one). As we see, the have an object property directly referencing each one of the objects to be easily retrievable and referenced for when we used the information to compute metrics: Subject Predicate Object GitHubDataSource ’repository’ ’repo-tfm’ GitHubDataSource ’owner’ ’owner-tfm’ GitHubDataSource ’accessToken’ ’gh_abcdefghijklmnop12345’ GitHubDataSource ’hasCommit’ CommitResource GitHubDataSource ’hasCommit’ CommitResource2 GitHubDataSource ’hasIssue’ IssueResource GitHubDataSource ’hasIssue’ IssueResource2 Table 22: GitHubDataSource with added Commit relationships sample instance The following two tables are Issue and Commit instance samples. In the original project, every time a data source object is fetched, more attributes are mapped to the data repository which they will be stored to. Nevertheless, we decided to get just a few of them, which are the most representative for each object. If, in the future, there is need to add more fields of the original object response, one will have to change the Commit class, for example, change the ontology and do the mapping correctly without worrying about inconsistencies in each entity "schema".
Report page 96 Subject Predicate Object IssueResource ’Title’ ’Insertion error’ IssueResource ’State’ ’active’ IssueResource ’Body’ ’Can’t insert Profile correctly’ IssueResource ’isLocked’ true IssueResource ’createdBy’ MembershipResource IssueResource ’assignedTo’ MembershipResource4 Table 23: Issue sample instance Subject Predicate Object CommitResource ’message’ ’bubble sort added in utils’ CommitResource ’verified’ true CommitResource ’commentCount’ 4 CommitResource ’createdBy’ MembershipResource Table 24: Commit sample instance We represent the same as before but with a TaigaDataSource and Task and UserStory instances. Subject Predicate Object TaigaDataSource ’project’ ’project-tfm’ TaigaDataSource ’accessToken’ ’t_abcdefghijklmnop12345’ TaigaDataSource ’hasTask’ TaskResource TaigaDataSource ’hasTask’ TaskResource2 TaigaDataSource ’hasUserStory’ UserStoryResource TaigaDataSource ’hasUserStory’ UserStoryResource2 Table 25: TaigaDataSource sample instance
A Knowledge-Graph Data Layer for Software Project Dashboards page 97 Subject Predicate Object TaskResource ’Subject’ ’Create landing page’ TaskResource ’isClosed’ true TaskResource ’isBlocked’ false TaskResource ’assignedTo’ MembershipResource Table 26: Task sample instance Subject Predicate Object UserStoryResource ’Subject’ ’Integration 13’ UserStoryResource ’Description’ ’Integrate FlyingStates airline’ UserStoryResource ’isBlocked’ true UserStoryResource ’isClosed’ false UserStoryResource ’createdBy’ MembershipResource Table 27: UserStory sample instance 9.2.3 QR-Eval module QR-Eval module is tested after doing all the process aforementioned in this section. For reference, read section 8. If you try to update a metric value without a defined method, it will throw an exception. If you execute a method that references Commit instances of a specific GitHubDataSource and the operation is incompatible for different reasons, it updates the MetricItem referenced in the API REST call to 0. In this and the successful cases, a value will be assigned to a MetricItem and that will be accompanied with an update of the whole Quality Model hierarchy that this metric acts in. We can see an example of this in section 8.
A Knowledge-Graph Data Layer for Software Project Dashboards page 99 10 Conclusions and Future Work The final section of this Master Thesis Project Report is a rundown of the key ideas that summarize the project and some future changes that may be considered in the future for its continuation. •Utilizing Knowledge Graphs as a web application database is not common, as we have seen that there are not that many tools to integrate this data model into this kind of project, as we have seen there is a lack of object mappers like ORM or ODM makes writing queries and interacting with domain objects a more low-level approach. •Knowledge Graph data model, as it is a schemaless kind, makes the insertion of new objects more flexible but it tends to make developers deal with schema issues more explicitly. Plus, relational databases or document-oriented like MongoDB have mechanisms for indexing certain fields in tables and collections which does not exist anything yet for Knowledge Graphs. •We integrated the base described functionalities of the Learning Dashboard run by relational and document-oriented data layers to a full back-end web application with a data layer using only Knowledge Graphs. With this implementation, we proposed a first approach to tackle the Learning Dashboard domain by using this technology. •We have demonstrated how to extend the application for different data sources. This process can be shortened or generalized in some way in the future for easier integration. •We propose for future research to make the methods available for the different objects for each DataSource more easily declared, as right now for each new method we have to append a new implementation for it in the repository class where this is handled.
Report page 100 •different data sources may not get the information with an API REST, we propose to further extend the different options by which someone can insert data into our database like reading from local files or communicating with remote file repositories, plus handling the different ways in which each data can come, for this project we tackled a REST API response which uses JSON but other formats like XML or CSV can be considered. •It could be interesting to implement the rest of the Learning Dashboard domain related to predictions, alerts, etc. In this thesis, we implemented a baseline that can be extended as a baseline. •Several Apache Jena queries can be improved by avoiding long paths between two sets of nodes by adding more direct relationships. This is a way of improving the performance of the data model. In the future, there can be more queries with added relations between existing entities. •ODIN project[29] is a system that supports the incremental pay-as-you-go integration of data sources into dataspaces and provides user-friendly querying mechanisms for the resulting dataspaces. Since in the Learning Dashboard, we integrate different data sources of heterogenous responses for each of their requested objects, we propose ODIN for overcoming this heterogeneity by extracting the schema from each source and then representing it using a more expressive data model. Source graphs can be generated from data sources in various formats such as JSON, XML, CSV, and relational. Nevertheless, to correct the alignments, user feedback is needed. In the end, a provenance graph is created by applying the accepted alignments.
A Knowledge-Graph Data Layer for Software Project Dashboards page 101 11 Bibliography [1] Gessi. https://gessi.upc.edu/en. [2] Q-rapids project repository. https://github.com/q-rapids. [3] Q-rapids dashboard repository. https://github.com/q-rapids/ qrapids-dashboard. [4] Qr-connect repository. https://github.com/q-rapids/qrapids-connect. [5] Qr-eval repository. https://github.com/q-rapids/qrapids-eval. [6] Learning dashboard repository. https://github.com/q-rapids/ learning-dashboard. [7] Jiacheng Xu, Kan Chen, Xipeng Qiu, and Xuanjing Huang. Knowledge graph representation with jointly structural and textual encoding. arXiv preprint arXiv:1611.08661, 2016. [8] Xu Han, Zhiyuan Liu, and Maosong Sun. Neural knowledge acquisition via mutual attention between knowledge graph and text. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018. [9] Chenjin Xu, Mojtaba Nayyeri, Fouad Alkhoury, Hamed Yazdi, and Jens Lehmann. Temporal knowledge graph completion based on time series gaussian embedding. In The Semantic Web–ISWC 2020: 19th International Semantic Web Conference, Athens, Greece, November 2–6, 2020, Proceedings, Part I 19, pages 654–671. Springer, 2020. [10] Arthur Freire, Mirko Perkusich, Renata Saraiva, Hyggo Almeida, and Angelo Perkusich. A bayesian networks-based approach to assess and improve the teamwork quality of agile teams. Information and Software Technology, 100:119–132, 2018.
Report page 102 [11] Guohui Xiao, Linfang Ding, Benjamin Cogrel, and Diego Calvanese. Virtual knowledge graphs: An overview of systems and use cases. Data Intelligence, 1(3):201–223, 2019. [12] Roberto De Virgilio, Antonio Maccioni, and Riccardo Torlone. Converting relational to graph databases. In First International Workshop on Graph Data Management Experiences and Systems, pages 1–6, 2013. [13] Average salary for a research analyst. https://www.glassdoor.com/Salaries/ barcelona-research-analyst-salary-SRCH_IL.0,9_IM1015_KO10,26.htm. [14] Average salary for a software developer. https://www.glassdoor.com/Salaries/ barcelona-software-engineer-salary-SRCH_IL.0,9_IM1015_KO10,27.htm. [15] Outervision power supply calculator. https://outervision.com/ power-supply-calculator. [16] Alejandra Volkova. Specification and design of a dashboard for monitoring the learning process in software projects developed by teams of students. 2022. [17] Docker, a container-based approach to deploying. https://www.docker.com/. [18] Neo4j’s model: Relational to graph. https://neo4j.com/developer/ relational-to-graph-modeling/. [19] About n-ary relations for knowledge graphs, by w3school. https://www.w3.org/ TR/swbp-n-aryRelations/. [20] Klaus Renzel and Wolfgang Keller. Three layer architecture. 1997. [21] Draw.io. https://app.diagrams.net/. [22] Overleaf. https://www.overleaf.com/. [23] Intellij. https://www.jetbrains.com/es-es/idea/. [24] Apache jena. https://jena.apache.org/.
A Knowledge-Graph Data Layer for Software Project Dashboards page 103 [25] Github. https://github.com/. [26] Trello. https://trello.com/. [27] Postman, a client for sending api requests. https://www.postman.com/. [28] Protégé, an ontology visualization software. https://protege.stanford.edu/. [29] Odin (on-demand data integration). https://github.com/dtim-upc/ODIN.