Full text
“A Model-Driven Engineering Approach for the Uniquely Identity Reconciliation of Heterogeneous Data Sources” TESIS DOCTORAL Autor D. José González Enríquez Directores Doctora D.ª María José Escalona Cuaresma Doctor D. Francisco José Domínguez Mayo Sevilla, 22 de mayo de 2017
“A Model-Driven Engineering Approach for the Uniquely Identity Reconciliation of Heterogeneous Data Sources” DOCTORAL THESIS Author D. José González Enríquez Supervisors Ph.D. Dª. María José Escalona Cuaresma Ph.D. D. Francisco José Domínguez Mayo Seville, 22th May, 2017
To my parents, Antonio y Mercedes, for their tireless support, help and affection.
“My biggest motivation? Just to keep challenging myself.” – Richard Charles Nicholas Branson “La vida sin esfuerzo es una vida mediocre” – Papa Francisco
ACKNOWLEDGEMENTS fter several years of hard work, it is time to write one of the parts that I consider to be the most important part of this doctoral thesis, the acknowledgements, because without the effort, support, advice and encouragement of all the people who I am willing to name, the desire to do this work, which today becomes a reality, would have not been possible. First, thank my family, my parents Antonio and Mercedes, my brothers Antonio and Jesus and my nephews Marta and Alejandro and my sister in law Ana. Thank you for always being supporting me, giving me affection and making the distance not to be a problem for feeling you very close every day. Thanks to my traveling companion on the road to life, Maria. The name of the solution proposed in this work, since it could not have turned out any other way, refers to yours and it is because you are, have been and will be, one of the most important persons of my life. Thank you for understanding my work, for suffering me and waiting for me all the time that I have spent away from home, for being the first shoulder where to find comfort and the reason to smile every day. You are the faithful reflection of effort and constancy. Today I have achieved one of my goals and I am sure that soon, you will do with your own as well, but the most important thing is that we will always achieve what we want together. To my adoptive family, Paqui and Paco, Pablo and the Palacios grandfather. Thank you for making me feel one more of the family, for the advice and encouragement to never give up. To all my friends especially Andrew. Thank you for always accompanying me in good and bad times and for the moments of talk that made me disconnect from the routine of daily work. We still have many projects to do together, thank you for becoming in one more brother. To my directors of this work of doctoral thesis. María José, thank you very much for having trusted me from the beginning and having initiated myself into this wonderful world of research. You have taught me to put a good face on the bad news and not to limit my dreams. Franci, apart from being my tutor, you have become a great friend for me. Thank you very much for your dedication, for your time, for the advice, for the moments of talk, for the discussions, for accompanying me the first time that I presented an article in an international congress, in short, for being always been for everything. Without your work, this doctoral thesis would have been only one more idea. To all my colleagues of the “Ingeniería Web y Testing Temprano” research group for suffering me daily. Thank you very much Julian for all the conversations you had during my stays abroad, on topics related to this work and those which are not, you are an example of overcoming, do not be bored to raise awareness. To Juanmi and Antonio for having always been willing to manage anything while I have been making my research stays and all other companions for your support and conversation moments. To Virginia and Leticia for enduring and suffering me and giving me their affection every moment. A
FIGURES Figure I-1. Entity Reconciliation example (McCallum et al., 2000). ............................... 3 Figure II-1. Method for the Systematic Mapping Study .................................................. 9 Figure II-2. Primary studies Selection Process ............................................................... 17 Figure II-3. Data Synthesis Graphics ............................................................................. 21 Figure II-4. Total of papers by Year and Category ........................................................ 22 Figure II-5. Total of papers by Digital Library .............................................................. 22 Figure III-1. Solution developed in this Doctoral Thesis ............................................... 29 Figure III-2. MaRIA Framework. ................................................................................... 30 Figure III-3. Syntax and Semantic Relation of the Model (Metzger, 2008) .................. 32 Figure III-4. Transformations between Models. ............................................................. 33 Figure IV-1. NDT-Q Framework ................................................................................... 38 Figure IV-2. MaRIA Framework ................................................................................... 40 Figure IV-3. NDT-Q Framework + MaRIA Framework ............................................... 40 Figure IV-4. Extended software development process group ........................................ 42 Figure IV-5. Extended requirement process ................................................................... 43 Figure IV-6. Analyze Data Source Activity ................................................................... 43 Figure IV-7. Extended analysis process ......................................................................... 45 Figure IV-8. Define Entity Reconciliation Problem ....................................................... 45 Figure IV-9. Extended testing process group ................................................................. 47 Figure IV-10. Extended Run the tests process ............................................................... 48 Figure V-1. Framework MaRIA Metamodel .................................................................. 51 Figure V-2. Global view of MaRIA Metamodel ............................................................ 52 Figure V-3. Data Source Metamodel .............................................................................. 53 Figure V-4. Virtual Graph Metamodel ........................................................................... 57 Figure V-5. Transformation Metamodel ........................................................................ 61 Figure V-6. Testing Metamodel ..................................................................................... 69 Figure VI-1. Framework MaRIA Metamodel ................................................................ 75 Figure VI-2. Text Template Architecture (Cook et al., 2007) ........................................ 77 Figure VI-3. Declaration of a Directive.......................................................................... 78 Figure VI-4. Declaration of a Template directive .......................................................... 79 Figure VI-5. Declaration of an Output directive ............................................................ 79 Figure VI-6. Declaration of an Assembly directive ....................................................... 79 Figure VI-7. Declaration of an Import directive............................................................. 79 Figure VI-8. Declaration of an Include directive ........................................................... 79 Figure VI-9. Standard Control Block ............................................................................. 80 Figure VI-10. Expression Control Block ........................................................................ 80 Figure VI-11. Class Feature Control Block .................................................................... 80 Figure VI-12. Write Method ........................................................................................... 81 Figure VI-13. Error Method ........................................................................................... 81 Figure VI-14. Structural Rules Transformation Expression ........................................... 86 Figure VI-15. Structural Rules Transformation Expression ........................................... 87 Figure VI-16. Load Rules Transformation Expression .................................................. 87 Figure VI-17. Load Rules Transformation Expression .................................................. 88 Figure VII-1. MaRIA Framework – MaRIA Tool ......................................................... 90 Figure VII-2. DSL-Tools Meta-Metamodel (Bézivin et al., 2005) ................................ 91 Figure VII-3. Eclipse Modeling Meta-Metamodel (Kolovos, 2016) ............................. 92 Figure VII-4. Graphical Edition Eclipse Modeling Tools (Eclipse, 2016b) .................. 93 Figure VII-5. Graphical Edition DSL Tools (MSDN, 2016) ......................................... 93
Figure VII-6. M2T Transformations .............................................................................. 94 Figure VII-7. Root Element (DomainClass + Diagram) ................................................ 96 Figure VII-8. Root Element (DomainClass + Diagram) ................................................ 96 Figure VII-9. Wrapper Metaclass ................................................................................... 97 Figure VII-10. DataSourceEntity Metaclass .................................................................. 97 Figure VII-11. DataSourceAttribute Metaclass .............................................................. 97 Figure VII-12. EntityVertex Metaclass .......................................................................... 98 Figure VII-13. Attribute Metaclass ................................................................................ 98 Figure VII-14. Transformation Metaclass ...................................................................... 99 Figure VII-15. Transformation Operations Metaclasses ................................................ 99 Figure VII-16. Transformation Metaclass ...................................................................... 99 Figure VII-17. Resolution Operations Metaclasses ...................................................... 100 Figure VII-18. Types of Used Shapes .......................................................................... 100 Figure VII-19. DataSource Shape Decorador............................................................... 101 Figure VII-20. Diagram Element Map Example .......................................................... 101 Figure VII-21. Decorator Maps .................................................................................... 101 Figure VII-22. Decorator Maps .................................................................................... 102 Figure VII-23. Data Source Element Definition .......................................................... 103 Figure VII-24. Wrapper Element Definition ................................................................ 103 Figure VII-25. DataSourceEntity Element Definition .................................................. 104 Figure VII-26. DataSourceAttribute Element Definition ............................................. 104 Figure VII-27. EntityVertex Element Definition ......................................................... 104 Figure VII-28. Attribute Element Definition ................................................................ 105 Figure VII-29. Transformation Element Definition ..................................................... 105 Figure VII-30. Transformation Types Element Definition .......................................... 106 Figure VII-31. Main Resolution Element Definition ................................................... 106 Figure VII-32. Resolution Operation Element Definition ............................................ 106 Figure VII-33. Toolbox ................................................................................................ 107 Figure VII-34. Example of Modeling ........................................................................... 108 Figure VIII-1. MOSAICO Architecture (MOSAICO, 2016) ....................................... 113 Figure VIII-2. Draft Example of Modeling .................................................................. 115 Figure VIII-3. MaRIA Tool .......................................................................................... 116 Figure VIII-4. Real Entity Reconciliation Problem modeled with MaRIA ................. 117 Figure VIII-5. ADAGIO Architecture .......................................................................... 119 Figure IX-1. MaRIA Framework. ................................................................................ 126 Figure IX-2. MaRIA Tool ............................................................................................ 127 Figure A-1. Needed Software for Installation of Visual Studio Enterprise 2015......... 144 Figure A-2. Installation dialog of VS2015 Enterprise .................................................. 144 Figure A-3. Installation process of VS2015 Enterprise ................................................ 145 Figure A-4. Installation dialog of Enterprise SDK 2015 and Modeling SDK 2015 .... 145 Figure A-5. Installation dialog of Enterprise SDK 2015 and Modeling SDK 2015 .... 146 Figure A-6. New Project ............................................................................................... 146 Figure A-7. Domain Specific Language Designer ....................................................... 147 Figure A-8. Minimal Language .................................................................................... 147 Figure A-9. Minimal Language Template .................................................................... 148 Figure A-10. Minimal Language Template .................................................................. 148 Figure A-11. ToolBox .................................................................................................. 149 Figure A-12. DslDefinition.dsl Element ...................................................................... 150 Figure A-13. DSL Definition Diagram......................................................................... 151 Figure A-14. DSL Solutions Explorer .......................................................................... 151
Figure A-15. DSL Elements Explorer .......................................................................... 152 Figure A-16. Properties Window.................................................................................. 153 Figure A-17. DSL Details Tab ..................................................................................... 153 Figure A-18. DSL Details Tab ..................................................................................... 154 Figure A-19. Compilation Meny .................................................................................. 154 Figure A-20. Debug Menu ........................................................................................... 155 Figure B-1. Instance of Support Tool ........................................................................... 157 Figure B-2. Example of Model Design with the Support Tool .................................... 160
CHAPTER I INTRODUCTION
1 CHAPTER I. INTRODUCTION nformation is a phenomenon that provides meaning or sense to things. In general terms, information is an organized set of processed data, which constitutes a message about a concrete entity or phenomenon. The data are perceived, integrated and generate the necessary information to produce the knowledge that is the one that finally allows to make decisions to carry out the daily actions. Information also processes and generates human knowledge. When a concrete problem needs to be solved or a decision must be taken, people usually use different sources of information and build what is generally called knowledge or organized information that allows problem solving or decision making. The main topic of this doctoral thesis is related to the information management. Specifically, this first chapter aims to describes the context of the work developed in this doctoral thesis. To this end, the first section presents a brief introduction to focus work. Then, the second and third sections, describe what the structure of this document and brief conclusions of the chapter respectively. 1. INTRODUCTION Currently, the information management is critical in many aspects of our lives. However, the incorporation of information and communications technology (ICT) in everyday life causes people to experience an overshooting of information, also known by the term “infoxication”. This term refers to the difficulty that someone has to understand a problem and make decisions about it because of an excess of information presence (Yang et al., 2003). In the first era of ICT, the main problem that researchers had was how to find information and how to store and manage it efficiently. Currently, due to the presence of a lot of systems that store and generate information, the biggest problem that researchers have is how to extract knowledge of this information based on the needs of each in an effective way (Enríquez et al., 2015). The complexity of information has increased not only by the high capacity of information production, but these huge amounts are distributed in multiple databases, not just one, and these databases are different, they do not have the same structure, someone’s repeat information and it is not always possible to make a faithful comparison between them, in conclusion, it can be said that these data sources are heterogeneous, it means, although the store information related to the same topic, they share neither structure nor content. Then, it would be very useful to integrate all information related to the same subject into a single data source. It does not means creating a new one unifying all the existing ones, it means that it is necessary to make a system that respond efficiently to queries that are made, despite their differences providing a consolidate information from all the data sources queried and it is very important to make it quickly. In this context, the I
2 problem of reconciling entities in heterogeneous data sources mentioned before (which is the start point of this doctoral thesis) takes a very important value. Entity reconciliation (also called entity resolution or ER) is a fundamental problem in data integration and it is no new. It refers to combining data from different sources for a unified vision or, in other words, identifying entities from the digital world that refers to the same real-world entity. It is an uncertain process because the decision to allocate a set of records with the same entity, cannot be taken with certainty, unless these records are identical in all their attributes or they have a common key (Getoor and Machanavajjhala, 2012; Wang et al., 2013). This problem can be applied to many kinds of scenarios. Entity reconciliation is a well-known problem and it has been investigated since the birth of relational databases (Whang and Garcia-Molina, 2014). If to everything mentioned, a very trending topic nowadays such as the Big Data is added, this problem receives a much more significant attention due to the new challenges that it arises. Although this problem is not new, the management of heterogeneous or big and heterogeneous data sources presents new challenges. (Enríquez et al., 2015; Gal, 2014). In the paper written by Getoor and Machanavajjhala, (2013), the authors exposed some of the main challenges of entity reconciliation in this environment such as: data heterogeneity, it is becoming more common that data are unstructured, unclean or incomplete and also there are diverse data types; data more linked, where it is expressed the necessity of inferring relationships; multi-relational data, dealing with the structure of entities; and building multi-domain systems, trying to customize methods that span across domains. In this sense this doctoral thesis tries to address the most of these problems providing a solution that covers these aspects. In the literature, it is possible to find a wide variety of approaches to try to solve the problem of reconciliation of entities. Taking into account the classification presented by Gal (2014), entity reconciliation problems can be briefly classified into: deterministic rule-based, probabilistic-based methods, learning based techniques and graph-based techniques. Deterministic rule-based: Lee et al., (2013), proposed an approach to coreference resolution that combines the global information and precise features of machine-learning models with deterministic, rule-based systems. Galhardas et al., (2001), presented a language, an execution model and algorithms that enable users to express data cleaning specifications declaratively using for their demonstration an example of a set of bibliographic references. Bhattacharya and Getoor, (2005), proposed a probabilistic model for collective entity resolution for relational domains where references are connected to each other. Thus, they presented an algorithm for collective entity resolution which is unsupervised and takes entity relations into account. Probabilistic methods: Winkler, (2002), presented methods for Record Linkage and Bayesian Networks. Verykios et al., (2003), presented a Bayesian decision model for cost optimal record matching using the ratio of the prior odds of a match along with appropriate values of thresholds to partition the decision space into three decision areas. This model is an improved version of the one proposed by Fellegi and Sunter, (1969). Learning-based techniques: Sarawagi and Bhamidipaty, (2002) presented a system that use a method of interactively discovering challenging training pair using active learning. Cohen and Richman, (2002), described techniques for
3 clustering and matching identifier names that are both scalable and adaptive, in the sense that they can be trained to obtain better performance in a concrete domain. Graph-based techniques: Ioannou et al., (2010), described a framework for entity linkage with uncertainty where the possible linkages are stored alongside the data with their belief value. They use a probabilistic query answering technique to take the probabilistic linkage into consideration. Wang et al., (2013), focused on the construction of effective reference table by relying on co-occurring relationship between tokens to identify suitable entity names. The firstly model data set as graph, and then cluster the vertices in the graph. They also mine synonyms and get the expansive reference table. Wang et al., (2016), modeled the entity resolution problem as the partition of the vertices in a weighted graph into cohesive subgraphs. They propose an approximate algorithm with approximation ratio bound is proposed and a heuristic algorithm for performing entity resolution on a large data set efficiently. Let’s get a simple example of entity reconciliation found in the literature (McCallum et al., 2000) based on a coauthor network from bibliographic data used for InfoVis 2004. At the left side of the Figure I-1, it is possible to see the complete network, but it is possible to note that, there are elements that represents the same author but they are not named equally. In this sense and after applying an entity reconciliation process, it is possible to see how the final result (right side of Figure I-1), presents a much clearer vision of the complete coauthor network. In this process the references to concrete entities have been canonized and linked. Figure I-1. Entity Reconciliation example (McCallum et al., 2000).
4 2. STRUCTURE OF THE DOCTORAL THESIS After this introduction chapter, it will be presented the different chapters that compose this doctoral thesis. All these aspects are developed along them are structured as follows: Chapter II offers a systematic study of the state of the art about which are the related works of this thesis. In it, techniques, methods or tools that allow to solve the entity reconciliation problem are analyzed. This general vision will allow to analyze how the entity reconciliation problem is being and have been solved. The general view presented in the Chapter II helps to understand the actual situation and to set the basis so that during the Chapter III, the definition of the problem to be addressed could be performed. Chapter III provides, so, the definition of the problem, the challenges and objectives to meet, the working environment and a description of the approach to the problem dealt with in this thesis. This chapter also defines in detail the influences that have driven the realization of this work. Raised the problem and achieve the objectives, the next three chapters describe proposed Model-driven entity ReconcilIAtion (MaRIA) Framework, deepening on each of the elements on which the framework is based on. In this sense, Chapter IV presents the MaRIA process, which is a set of activities that should be added to any software developing methodology to carry out an entity reconciliation process. This set of processes will be focused in the requirement, analysis and testing phases. The chapter, includes a concrete example of its use in a web software methodology. Chapter V presents the metamodels used to allow the user of the framework to model entity reconciliation problems and Chapter VI presents how the Early Testing has been included to the solution. In this chapter, the theoretical transformations that will allow to automatically generate the business rules that will make the model testable are presented. These business rules will derive in the test requirements of the application that consume data that generated after the entity reconciliation process. In order to materialize and automate the metamodels, constraints and transformations proposed in Chapters V and VI, a support tool has been developed and it is presented in Chapter VII. Next, Chapter VIII presents a case study taken as validation scenario for this doctoral thesis. The proposal presented has been applied to this scenario using the support tool presented in Chapter VII. Finally, chapter IX closes the memory of this doctoral thesis presenting a set of conclusion, the contributions that this research contributes to the scientific community, presenting the future work that has been proposed as well as the new complementary research lines that have been created from this work. This doctoral thesis is completed with five more sections: references and four annexes. References section represents the description of all the literature references that the student has consulted during the development of this doctoral thesis.
5 Annex A describes an introduction to DSL-Tools (IDE which the support tool of this approach has been developed), explaining the software requirements needed for the installation, a short installation manual and a brief description of the IDE. Annex B describes the user manual of the support tool of this proposal presented in Chapter VII. Annex C presents the glossary of terms that will make easy to understand all the relevant concepts of this doctoral thesis. Finally, Annex D summarizes the recognitions and research activity of the PhD student achieved during the development of this work. 3. CONCLUSIONS The Doctoral Thesis presented in this work is motivated by a problem that has been identified within organizations engaged in the business of software: the need to establish systematic and automated mechanisms that allow to perform the entity reconciliation in heterogeneous data sources. All this to help organizations to improve the management and quality of data they work with. The introduction chapter, in addition to frame and contextualize the work of the thesis, provides a definition of entity reconciliation, as understood throughout this work. Furthermore, it globally shows what the structure of presentation is and the research results which have resulted in the conduct of this doctoral thesis. At this point, it is convenient to mention that this research work has been carried out within the “Ingeniería Web y Testing Temprano” (IWT2) research group of the “Escuela Técnica Superior de Ingeniería Informática (ETSII)” of the University of Seville. This research group is referenced in the Andalusian research plan as PAIDI TIC021. Moreover, and more specifically, this thesis has been developed within a strategic framework of IWT2 group that aims to explore and investigate how to combine satisfactorily the Model-Driven Engineering (MDE) paradigm with the entity reconciliation problem. Finally, this Doctoral Thesis has been sponsored by Fujitsu Laboratories of Europe (FLE). Fujitsu Laboratories Limited has had an active presence in Europe since 1990, when they set up their first facility at Stockley Park near Heathrow. In subsequent years, Fujitsu Laboratories opened branches in France and Germany, undertaking advanced research in the fields of telecommunications, information technology and computational science. In 2015, FLE established a new Data Analytics Research Center in Madrid, Spain, reflecting their commitment to “think global, act local” - applying their R&D expertise to address specific regional challenges.
12 In the systematic search of phase two, once the relevant keywords have been found and some pilot testing was carried out, a Python script was developed for making the combination between all of them. In this context, two category files were created: one for the ER problem and another one for the technologies, tools, frameworks and concepts. Having these two files, the script was programed by taking one of the keywords of the ER problem file and combining it with all the keywords of the second file. Besides, search queries were generated concretely for each database selected to conduct the systematic search. All kind of papers have been included such as: journal papers and presentations at conferences, congresses, tutorials and workshops. A very large number of queries have been executed and for each database they have been customized depending on: the query syntax of the relevant database, the possibilities that the database offers to make filters, year of publication and specific topics. Table II-4 shows some examples of the queries that have been executed. Database Keywords Query 1 Scopus 2010 (TITLE("Data fusion") AND KEY("Entity Matching")) OR (TITLE("Data fusion") AND KEY("Entity Identity")) OR (TITLE("Data fusion") AND KEY("Entity Name System")) OR (TITLE("Data fusion") AND KEY("Entity Recognition")) OR (TITLE("Data fusion") AND KEY("Entity Parsing")) OR (TITLE("Data fusion") AND KEY("Entity Linking")) OR (TITLE("Data fusion") AND KEY("Entity disambiguation")) OR (TITLE("Data fusion") AND KEY("Entity Resolution")) OR (TITLE("Data fusion") AND KEY("Entity Reconciliation")) OR (TITLE("Data fusion") AND KEY("Identity Matching")) OR (TITLE("Data fusion") AND KEY("Identity Management")) OR (TITLE("Data fusion") AND KEY("Identity Resolution")) OR (TITLE("Data fusion") AND KEY("Identity Attributes")) OR (TITLE("Data fusion") AND KEY("Identity Search")) OR (TITLE("Data fusion") AND KEY("Identity Linking")) OR (TITLE("Data fusion") AND KEY("Duplicate Detection")) OR (TITLE("Data fusion") AND KEY("Deduplication")) OR (TITLE("Data fusion") AND KEY("Record Linkage")) OR (TITLE("Data fusion") AND KEY("Object Identification")) OR (TITLE("Data fusion") AND KEY("Reference Matching")) OR (TITLE("Data fusion") AND KEY("Co-Reference Detection")) OR (TITLE("Data fusion") AND KEY("Non-identical Duplicates")) OR (TITLE("Data fusion") AND KEY("Redundancy elimination")) OR (TITLE("Data fusion") AND KEY("Object Matching ")) OR (TITLE("Data fusion") AND KEY("Fuzzy Matching")) OR (TITLE("Data fusion") AND KEY("Similarity join processing")) OR (TITLE("Data fusion") AND KEY("Duplication Detection")) OR (TITLE("Data fusion") AND KEY("Reference Reconciliation")) OR (TITLE("Data fusion") AND KEY("Co-Reference Resolution")) OR (TITLE("Data fusion") AND KEY("Relational Blocking")) Query 2 ACM 2010 ("Title":"large data" AND "Title":"entity matching") OR+("Title":"large data" AND "Title":"entity identity") OR+("Title":"large data" AND "Title":"entity name system") OR+("Title":"large data" AND "Title":"entity recognition") OR+("Title":"large data" AND "Title":"entity parsing") OR+("Title":"large data" AND "Title":"entity linking") OR+("Title":"large data" AND "Title":"entity disambiguation") OR+("Title":"large data" AND "Title":"entity resolution") OR+("Title":"large data" AND "Title":"entity reconciliation") OR+("Title":"large data" AND "Title":"identity matching") OR+("Title":"large data" AND "Title":"identity management") OR+("Title":"large data" AND "Title":"identity resolution") OR+("Title":"large data" AND "Title":"identity attributes") OR+("Title":"large data" AND "Title":"identity search") OR+("Title":"large data" AND "Title":"identity linking") OR+("Title":"large data" AND "Title":"duplicate detection") OR+("Title":"large data" AND "Title":"deduplication") OR+("Title":"large data" AND "Title":"record linkage") OR+("Title":"large data" AND "Title":"object identification") OR+("Title":"large data" AND "Title":"reference matching") OR+("Title":"large data" AND "Title":"co-reference detection") OR+("Title":"large data" AND "Title":"non-identical duplicates") OR+("Title":"large data" AND "Title":"redundancy elimination") OR+("Title":"large data" AND "Title":"object matching ") OR+("Title":"large data" AND "Title":"fuzzy
13 matching") OR+("Title":"large data" AND "Title":"similarity join processing") OR+("Title":"large data" AND "Title":"duplication detection") OR+("Title":"large data" AND "Title":"reference reconciliation") OR+("Title":"large data" AND "Title":"co-reference resolution") OR+("Title":"large data" AND "Title":"relational blocking") Query 3 IEEE 2010 ("Document Title":"Data fusion" AND "Document Title":"Entity Matching") OR ("Document Title":"Data fusion" AND "Document Title":"Entity Identity") OR ("Document Title":"Data fusion" AND "Document Title":"Entity Name System") OR ("Document Title":"Data fusion" AND "Document Title":"Entity Recognition") OR ("Document Title":"Data fusion" AND "Document Title":"Entity Parsing") OR ("Document Title":"Data fusion" AND "Document Title":"Entity Linking") OR ("Document Title":"Data fusion" AND "Document Title":"Entity disambiguation") OR ("Document Title":"Data fusion" AND "Document Title":"Entity Resolution") OR ("Document Title":"Data fusion" AND "Document Title":"Entity Reconciliation") OR ("Document Title":"Data fusion" AND "Document Title":"Identity Matching") OR ("Document Title":"Data fusion" AND "Document Title":"Identity Management") OR ("Document Title":"Data fusion" AND "Document Title":"Identity Resolution") OR ("Document Title":"Data fusion" AND "Document Title":"Identity Attributes") OR ("Document Title":"Data fusion" AND "Document Title":"Identity Search") OR ("Document Title":"Data fusion" AND "Document Title":"Identity Linking") OR ("Document Title":"Data fusion" AND "Document Title":"Duplicate Detection") OR ("Document Title":"Data fusion" AND "Document Title":"Deduplication") OR ("Document Title":"Data fusion" AND "Document Title":"Record Linkage") OR ("Document Title":"Data fusion" AND "Document Title":"Object Identification") OR ("Document Title":"Data fusion" AND "Document Title":"Reference Matching") OR ("Document Title":"Data fusion" AND "Document Title":"Co-Reference Detection") OR ("Document Title":"Data fusion" AND "Document Title":"Non-identical Duplicates") OR ("Document Title":"Data fusion" AND "Document Title":"Redundancy elimination") OR ("Document Title":"Data fusion" AND "Document Title":"Object Matching ") OR ("Document Title":"Data fusion" AND "Document Title":"Fuzzy Matching") OR ("Document Title":"Data fusion" AND "Document Title":"Similarity join processing") OR ("Document Title":"Data fusion" AND "Document Title":"Duplication Detection") OR ("Document Title":"Data fusion" AND "Document Title":"Reference Reconciliation") OR ("Document Title":"Data fusion" AND "Document Title":"CoReference Resolution") OR ("Document Title":"Data fusion" AND "Document Title":"Relational Blocking") Query 4 Web Of Knowle dge 2010 (Data fusion AND Entity Matching) OR (Data fusion AND Entity Identity) OR (Data fusion AND Entity Name System) OR (Data fusion AND Entity Recognition) OR (Data fusion AND Entity Parsing) OR (Data fusion AND Entity Linking) OR (Data fusion AND Entity disambiguation) OR (Data fusion AND Entity Resolution) OR (Data fusion AND Entity Reconciliation) OR (Data fusion AND Identity Matching) OR (Data fusion AND Identity Management) OR (Data fusion AND Identity Resolution) OR (Data fusion AND Identity Attributes) OR (Data fusion AND Identity Search) OR (Data fusion AND Identity Linking) OR (Data fusion AND Duplicate Detection) OR (Data fusion AND Deduplication) OR (Data fusion AND Record Linkage) OR (Data fusion AND Object Identification) OR (Data fusion AND Reference Matching) OR (Data fusion AND Co-Reference Detection) OR (Data fusion AND Non-identical Duplicates) OR (Data fusion AND Redundancy elimination) OR (Data fusion AND Object Matching) OR (Data fusion AND Fuzzy Matching) OR (Data fusion AND Similarity join processing) OR (Data fusion AND Duplication Detection) OR (Data fusion AND Reference Reconciliation) OR (Data fusion AND Co-Reference Resolution) OR (Data fusion AND Relational Blocking) Table II-4. Example of queries Once the queries for each database were created, a new specific Python script was designed for each one. Besides, Selenium, a software testing tool for Web-based applications (Selenium, 2017) was used. The process of searching a paper in each database was replicated and it was automated for getting the results based on the queries created before using this Python script and Selenium. Finally, another Python script was developed for removing the duplicate records found out during the process of search.
14 Because the process of analysis of the results obtained was quite long over time, the searching process previously described was repeated twice. In the last phase of manual search, papers recommended by experts in the ER problem were looked for. These papers were very important because they were very close to the topic and we could discard them because of the problem of bias. The Web version of the application Mendeley was used for managing this amount of data. It is a reference manager tool that helps to handle papers. Mendeley is integrated into the Web browser, allowing adding directly the articles from the digital libraries to a personal document database, avoiding the duplicated ones and saving them (when possible) in PDF format (Mendeley Support Team, 2011). STUDY SELECTION, INCLUSION AND EXCLUSION CRITERIA AND QUALITY INSURANCE This SMS includes papers written in English that refer to ER problems and technologies, tools or frameworks that try to solve this problem, published from 2010 up to January 2017 in indexed journals, such as Journal Citation Reports (JCR) and prestigious conferences, congresses or workshops categorized in the CORE ranking (CORE Conference Ranking). It excludes discussion or opinion papers or those that are only available in PowerPoint or abstract formats, duplicates (always considering the most completed one) and those whose main contribution is not referred to ER problems and technologies, tools or frameworks that try to solve it or just scarcely mention it. The first filter for selecting primary studies was based on the title and abstract of the paper. If it is not relevant to the study, it is automatically excluded. After this process, the inclusion/exclusion criteria were applied when reading the abstracts of the found items. Once read, if there was still any doubt with the inclusion/exclusion criteria, the paper was completely read. The PhD student conducted the selection of the studies and his supervisors, the 30% of the articles to corroborate if the inclusion/exclusion criteria were applied correctly. He/she would consult the other mates in case of doubts or discrepancies. DATA COLLECTION AND ANALYSIS First, a quantitative synthesis considering the number and/or percentage of items in each category was made, illustrating them with tables and graphics, to thereby give an answer to each research question, matching each question with category. Moreover, an interpretation of retrieved results and some suggestions deduced from the synthesis are presented. In addition, it was analyzed: (i) the number of publications per year to detect and justify trends and (ii) the number of publications by publication type to detect the journals in which more has been.
15 2.2.CONDUCTING Once the protocol was agreed, the proper study started. There were two main sections during the process of carrying out this SMS: (i) detect and select primary studies and data extraction, and (ii) apply the inclusion and exclusion criteria for selecting the primary studies that will be used for the work, showing the finally selected ones and the data synthesis phase, where a statistical study was conducted. It showed the main conclusions that obtained after running the previous phase. DETECT AND SELECT PRIMARY STUDIES AND DATA EXTRACTION Papers published between 2010 and 2017 were found using the search strategy defined in the protocol. Because of the limitations that certain search sources offered (for example, not allowing the use of complex search strings), it was necessary to design specific strings for each source and manipulate the outcome of searches to get the same results that may have been obtained using the original search string. The search was made on the title and abstract of the papers except for those databases that did not allow this. In that case, the search had to be performed in the full text. Each search source stored search strings, metadata of found items (title, author, year of publication, etc.) and abstracts of the papers. After reading their abstracts and excluding those irrelevant to the ER problem, 276 papers out of 2,255 were eliminated for being duplicates. Then, according to the range of years that we have chosen, 434 papers that were written before 2010 were also eliminated. Consequently, the inclusion/exclusion criteria were applied to the 1,545 remaining items, and 1,024 papers that were not classified into the computer science or information systems category were eliminated. The last filter was applied to the heterogeneous data area and 382 papers were discarded, remaining 139 candidates. From them, 72 papers were supposed to be duplicated, remaining 67 papers. Then, 6 old versions were also eliminated and finally, 61 primary studies were analyzed in depth reading the full text. As it was classified in the protocol, primary studies were identified and selected by the first author of the article, while the second author chose randomly 30% to corroborate the correct choice. The doubts that arose during the selection of items were resolved among all the PhD student and his supervisors. Table II-5 shows the 61 primary studies selected. Title Reference Entity resolution for distributed probabilistic data (Ayat et al., 2013) Incremental entity resolution on rules and data (Whang and Garcia-Molina, 2014) Efficient entity resolution based on subgraph cohesion (Wang et al., 2015) Domain-specific entity extraction from noisy, unstructured data using ontology-guided search (Bratus et al., 2011) Entity resolution for probabilistic data (Ayat et al., 2014) Entity resolution based EM for integrating heterogeneous distributed probabilistic data (Dharavath and Kumar, 2015) Pay-As-You-Go Entity Resolution (Whang et al., 2013) Interaction between Record Matching and Data Repairing (Fan et al., 2011) Conflict Resolution with Data Currency and Consistency (Fan et al., 2014) Information Fusion for Entity Matching in Unstructured Data (Ali and Cristianini, 2010)
16 Dynamic Sorted Neighborhood Indexing for Real-Time Entity Resolution (Ramadan et al., 2014) Disambiguation of named entities in cultural heritage texts using linked data sets (Brando et al., 2015) Adaptive Connection Strength Models for Relationship-Based Entity Resolution (Nuray-turan et al., 2013) Context-based Entity Description Rule for Entity Resolution (Li et al., 2011) Efficient and Effective Duplicate Detection in Hierarchical Data (Leitao et al., 2013) Entity Disambiguation in Anonymized Graphs Using Graph Kernels (Hermansson et al., 2013) HIL: A High-Level Scripting Language for Entity Integration (Hernández and Koutrika, 2013) A Clustering-Based Framework to Control Block Sizes for Entity Resolution (Fisher et al., 2015) A Probabilistic Model for Linking Named Entities in Web Text with Heterogeneous Information Networks (Shen et al., 2014) A Scalable Machine-Learning Approach for Semi-Structured Named Entity Recognition (Irmak and Kraft, 2010) BEAR: Block Elimination Approach for Random Walk with Restart on Large Graphs (Shin et al., 2015) Beyond 100 million entities: large-scale blocking-based resolution for heterogeneous data (Papadakis et al., 2012) Efficient Entity Resolution for Large Heterogeneous Information Spaces (Papadakis et al., 2011a) Efficient SPectrAl Neighborhood blocking for entity resolution (Shu et al., 2011) Entity Linking on Graph Data (Yu, 2014) Entity Matching across Heterogeneous Sources (Yang et al., 2015) Entity type recognition for heterogeneous semantic graphs (Sleeman and Finin, 2013) A load-balanced mapreduce algorithm for blocking-based entityresolution with multiple keys (Hsueh et al., 2014) Web-based Graphical Querying of Databases through an Ontology: the WONDER System (Calvanese et al., 2010) Large-Scale entity resolution for semantic web data integration (Costa, 2016) Populating Entity Name Systems for Big Data Integration (Kejriwal, 2014) Domain-adapted named-entity linker using Linked Data (Frontini et al., 2015) Learning-based entity resolution with MapReduce. (Kolb et al., 2011) ERGP: A combined entity resolution approach with genetic programming (C. Sun et al., 2014) Learning an accurate entity resolution model from crowdsourced labels (Wang et al., 2014) Entity resolution for high velocity streams using semantic measures (Priya et al., 2015) A confidence-based entity resolution approach with incomplete information (Gu et al., 2014) A framework for entity resolution with efficient blocking (Shu et al., 2012) Entity matching: A case study in the medical domain (Carvalho et al., 2015) An Identification Ontology for Entity Matching (Bortoli et al., 2014) DS-Dedupe: A scalable, low network overhead data routing algorithm for inline cluster deduplication system (Z. Sun et al., 2014) A fast entity resolution method based on wave of records (Liu et al., 2011) Cleaning Framework for Big Data - Object Identification and Linkage (Liu et al., 2015) To compare or not to compare: making entity resolution more efficient (Papadakis et al., 2011b) Entity Resolution for High Velocity Streams Using Semantic Measures (Priya et al., 2015) An Ensemble Blocking Scheme for Entity Resolution of Large and Sparse Datasets (Balaji et al., 2016) Unsupervised Entity Resolution on Multi-type Graphs (Zhu et al., 2016) Entity Matching Across Multiple Heterogeneous Data Sources (Kong et al., 2016) Efficient Entity Resolution on Heterogeneous Records (Lin et al., 2016) Linked Data Entity Resolution System Enhanced by Configuration Learning Algorithm (Nguyen and Ichise, 2016) Linking Heterogeneous Data in the Semantic Web Using Scalable and Domain-Independent Candidate Selection (Song et al., 2016) Using Memetic Algorithm for Instance Coreference Resolution (Xue and Wang, 2016) Rule-Based Method for Entity Resolution (Li et al., 2015) Entity resolution in disjoint graphs: an application on genealogical data (Rahmani et al., 2016)
17 Parallel Meta-blocking for Scaling Entity Resolution over Big Heterogeneous Data (Efthymiou et al., 2016a) Minoan ER: Progressive Entity Resolution in the Web of Data (Efthymiou et al., 2016b) Entity resolution in disjoint graphs: an application on genealogical data (Rahmani et al., 2016) Semantic-Aware Blocking for Entity Resolution (Q. Wang et al., 2016) Online entity resolution using an Oracle (Firmani et al., 2016) Entity Resolution-Based Jaccard Similarity Coefficient for Heterogeneous Distributed Databases (Dharavath and Singh, 2016) A Blocking Scheme for Entity Resolution in the Semantic Web (de Assis Costa and de Oliveira, 2016) Table II-5. Selected primary studies Figure II-2. Primary studies Selection Process DATA SYNTHESIS In this section, the information contained in the data extraction form is displayed to answer the research questions formulated previously. In addition to the quantitative data shown through tables and graphics, an interpretation of the results is also presented. RQ1. As reflected in Table II-6, a classification that divides the publications in those based on algorithms as solution and those data structured-based is presented. There is a total of 68 publications (seven more than the number of primary studies because there are some papers that mention both categories). Studying data, the percentage of publications is very similar in both categories, having 48.53% the structural ones, the 50% the algorithmic ones and just 1.47% represents the others.
18 Type Number Percent Structural 33 48.53% Algorithmic 34 50% Others 1 1.47% Table II-6. RQ1 - Data Synthesis RQ2. As shown in Table II-7, it is interesting to note that most research works have been validated with a theoretical approach (defining a theoretical approach as that which has validated its proposals with any dataset). They represent 95.08%. Two papers (3.28%) do not present any validation and just one (1.64%) presents a validation based on a real-industry scenario. Validation Number Percent Not Validated - Theoretical Approach 2 3.28% Validated - Theoretical Approach 58 95.08% Validated - Approach in Industry 1 1.64% Table II-7. RQ2 - Data Synthesis (validation) Table II-8 presents the datasets used by the authors for validating their approaches. Most of the proposals have been validated with real datasets (76.74%), followed by those which have used both real and synthetic datasets (22.95%) and finally those which have used synthetic datasets (14.75%). It is important to note that there are 77 papers because there are two papers that do not present any validation and 9 of them include the two types of datasets. Dataset Number Percent Real 54 76.74% Synthetic 14 22.95% Real + Synthetic 9 14.75% Table II-8. RQ2 - Data Synthesis (datasets)
19 RQ3. As shown in Table II-9, most research efforts have been focused on graphbased works (26.23%), followed by those based on Clustering/Blocking (22.95%). It also highlights rule-based works (14.75%), and those based on algorithms (16.39%) and probabilistic methods (11.48%). There are two categories based on programming languages and ontologies that represent 4.92%, as well as the learning category that represents 3.28%. Finally, there are three categories based on hints, sorted neighborhood and patterns that represent 1.64%. Method, Technique, Tools Number Percent Rule Based 9 14.75% Probabilistic Method 7 11.48% Learning Based 3 4.92% Graph Based 16 26.23% Programming Languages 2 3.28% Clustering / Blocking Based 14 22.95% Ontology 3 4.92% Patterns 1 1.64% Sorted Neighborhood 1 1.64% Algorithms 10 16.39% Hints 1 1.64% Table II-9. RQ3 - Data Synthesis RQ4. As shown in Table II-10, all the primary studies are based on the operation phase in contrast to the phase of design that only takes 4.92% (three studies present both design and operation phases). More than half of the selected studies (62.30%) apply their experiments to heterogeneous data sources. The lowest result is on multi-applications, where one study that represents 1.64% is mentioned. Finally, automation and multi-domain objectives are poorly represented with only 3.28% and 8.20%, respectively, and just 12 studies that represent 19.67%, mention the multi-relational objective. Objective Number Percent UML Design 3 4.92% Operation 61 100% ER Challenges Automation 2 3.28% Multi-Relational 12 19.67% Multi-Domain 5 8.20% Multi-Applications 1 1.64% Type of Dataset Heterogeneous 38 62.30% Non-heterogeneous 23 37.70% Table II-10. RQ4 - Data Synthesis Once the research questions have been answered and after an in-depth study of the retrieved data, some other conclusions are presented.
20 At it is observed in Table II-11, we can conclude that the topic that we are analyzing in this SMS is arising a lot of interest. From 2010 up to date, the numbers of papers related to ER have been increasing (omitting year 2012 where just one paper less than in 2010 was published). The growth curve between 2012 and 2014 is quite large, almost tripling the number of publications. The number of publications in 2015 remains constant with respect to those published in 2014 and finally, it increases in one more publication having 15 in total. Moreover, Table II-12 summarizes the evolution of the publications based on its category and the year of publication. It shows a clearest trend in this area of research is focused on graph-based methods, techniques or tools followed by those clustering/blocking-based. Those learning-based are the most scattered, finding only one publication in the beginning of the search period and another one in the end. Those sorted neighborhoods and pattern-based but in this case, they are placed at end of the search period. Table II-13 shows the result that has been retrieved from the different digital libraries. In this case, the ACM Digital Library is on top of the selected primary studies with 36.69% followed by Scopus with 25.90%, Web of Knowledge with 20.86%, and finally IEEE Xplore Digital Library with 16.55%. It is important to remark that the amount of papers of this table is higher because the duplications among databases were not eliminated. Year Number Percent 2010 3 4.92% 2011 7 11.48% 2012 2 3.28% 2013 6 9.84% 2014 14 22.95% 2015 14 22.95% 2016 15 24.59% Table II-11. RQ3 - Data Synthesis Year Rules Probabilistic Learning Graph Programming Clustering Blocking Ontology Patterns Sorted Neigh. Algorithm Hints 2010 0 0 1 2 0 0 0 0 0 0 0 2011 2 0 0 0 0 3 1 0 0 1 0 2012 0 0 0 0 0 2 0 0 0 0 0 2013 0 2 0 2 1 0 0 0 0 0 1 2014 3 3 1 1 1 2 1 0 0 3 0 2015 4 1 0 4 0 3 0 1 1 2 0 2016 0 1 1 7 0 4 1 0 0 4 0 Totals 9 7 3 16 2 14 3 1 1 10 1 Table II-12. RQ3 - Data Synthesis Library Total Percent ACM 51 36.69% IEEE 23 16.55% SCOPUS 36 25.90% WOK 29 20.86% Table II-13. RQ3 - Data Synthesis
21 Figure II-3. Data Synthesis Graphics 0 10 20 30 40 Algorithmics Structure Others Papers by Classification 0% 20% 40% 60% 80% 100% Not Validated - Theorical Approach Validated - Theorical Approach Validated - Approach in Industry Validation of Proposals 0% 10% 20% 30% 40% 50% 60% Real World Dataset Synthetic Dataset Real + Synthetic Dataset Validation Datasets 0 10 20 30 40 50 60 70 Design Operation Automation Multi-Relational Multi-Domain Multi-Aplications Heterogenity Total figures by Objectives 0 2 4 6 8 10 12 14 16 Rule-based Probabilistic Method Learning-based Graph-based Prograaming Languages Clustering/Block ing-based Ontology Patterns Sorted Neighborhood Algorithm Hints Totals depending on Classification
28 5. We do not find any approach that addresses the issue of how to test if the expected solution was the one obtained after the entity reconciliation process. With all this, therefore, it is made palpable the need to establish within software organizations, effective and automated mechanisms that enable the effective management of data produced after an entity reconciliation process in heterogeneous data sources. This doctoral thesis focus its work in this sense. However, the definition, implementation and deployment of this process is an extensive work and it is out of the scope of this thesis. In our work, we propose to focus our results in the first phase of the life cycle: requirements and analysis, including also acceptance testing. Thus, the main objective of this Doctoral Thesis can be defined as: “To propose a suitable environment to support the entity reconciliation in the requirements and analysis phases. This environment allows the development team to prepare their future system to guarantee a suitable entity reconciliation with, besides could be systematically tested.” 2. OBJECTIVES OF THE DOCTORAL THESIS Once the problem and the context of the problem to be solve have been raised, it's time to describe in concrete terms the objectives to be achieved with this thesis. Those are: 1. Perform a study of the state of the art of the different existing solutions for the entity reconciliation of heterogeneous data sources, checking if they are being used in real environments. This objective has been achieved with the study presented in the previous chapter. 2. Define and develop a Framework for designing the entity reconciliation models by a systematic way for the requirements and analysis phases. As it is introduced in the next sections, the model based paradigm could offer a suitable environment to support this idea. For this purpose, this objective has been divided in three sub objectives: 2.1. Define a set of activities, represented as a process which can be added to any software development methodology to carry out the activities related to the entity reconciliation in the requirements and analysis phase of a software development life cycle. This objective is solved in the Chapter IV where the proposal is described. 2.2. Define a metamodel that allows us to represent an abstract view of our model-based solution. it is formally defined in the Chapter V of the present doctoral thesis. 2.3. Define a set of derivation mechanisms that allow to stablish the base for automate the testing of the solutions where the framework proposed in this doctoral thesis has been used. Considering that the process will be applied in the early stages of the development, it is possible to say that this proposal applies Early Testing. This objective will be covered in the Chapter VI where a description of how it has been achieved is performed.
29 3. Provide a support tool for the framework. This objective is covered in Chapter VII. The support will allow to a software engineer to define the analysis model of an entity reconciliation problem between different and heterogeneous data sources. The tool will be represented as a Domain Specific Language (DSL). 4. Evaluate the results obtained of the application of the proposal in a real-world case study. Figure III-1 shows the architecture proposed for achieving all the objectives described above. At the left, the “Software Engineer” interacts with the system modeling an entity reconciliation solution for a concrete entity reconciliation problem. This model is transformed automatically in “Business Rules” thanks to the model to text “Transformation Rules” defined between the “Metamodel” and the “Business Rules Metamodel”. In this sense, this business rules are the result of the integration and application of Early Testing in the solution. Figure III-1. Solution developed in this Doctoral Thesis In this moment, once the motivation of the thesis and its objectives have been presented, the next step will be to present a brief introduction of the Framework proposed in this doctoral thesis. Metamodel Model Transformation Rules Business Rules Metamodel Automatic Transformation Business Rules Software Engineer Models Instantiation Instantiation UML UML Text Text Testing Tools Modelling Tools
30 3. PRESENTATION OF THE MARIA (MODEL-DRIVEN ENTITY RECONCILIATION) FRAMEWORK The main objective of the proposed framework is giving support to the final user to model entity reconciliation problems. For this purpose, the Framework has been developed under the umbrella of the Model-Driven Engineering paradigm that will allows to achieve this objective by a systematic and easy way. Figure III-2, shows a global view of the MaRIA Framework. Thanks to the proposed solution, the software engineers will be able to model their presumable entity reconciliation models for solving any type of entity reconciliation problem, understanding presumable, the capacity of test the final solutions for checking the coverage level of the new generated dataset without more efforts than the design of the problem thanks to the integration of Early Testing. MaRIA Framework has been designed so that it can be integrated in any methodology of software development, whether they have classic or agile life cycles. Considering that any software development methodology is composed of several phases, the scope of this Doctoral Thesis has been set up to cover the requirements and analysis phases. Also, the testing phase has been considered adding Early Testing to allow the models to be systematically tested. MaRIA Framework is composed of three fundamental pillars: “MaRIA process”, “Model-Driven Approach” and “MaRIA Tool”. “MaRIA process” defines the set of activities that must be added and performed in any software development methodology to develop a solution to an entity reconciliation problem. “Model-Driven Approach” is defined by the metamodel that allows the user that use this framework to model the entity reconciliation problem and derivation mechanisms that to allow the models defined to be systematically tested. “MaRIA Tool” is the support tool developed to give support to the MaRIA Framework. Figure III-2. MaRIA Framework. Support Tool MaRIA Tool Framework MaRIA MaRIA Process Model-Driven Approach
31 In the next three chapters, this brief introduction to the MaRIA Framework will be developed in detail. 4. CONCEPTUAL AND TECHNOLOGICAL INFLUENCES When defining how to raise the development of this Doctoral Thesis, there were some aspects, technologies and previous works that influenced the way in which it has developed. This section introduces the learned lessons of the systematic mapping study performed and developed in Chapter II. Section 3 of this Chapter provides a global view of the MaRIA Framework. As mentioned before, one of the pillars of the Framework is the “Model-Driven Approach” composed of a metamodel and a set of derivation mechanisms. In this sense, the Model-Driven Engineering paradigm has also been considered for the development of this Doctoral Thesis. Finally, this Model-Driven Approach is also based in the Virtual Graphs technology to represent and model the structure of data to be reconciled. Then, Virtual Graph is also covered by this section. All these influences are exposed in detail below. LEARNED LESSONS OF THE SYSTEMATIC MAPPING STUDY At the beginning of the thesis work, to know which methods techniques or tools were being used for ER was the reason of performing a systematic mapping study. One of the first objectives was to find a method, technique or tool that works with heterogeneous databases and highly scalable allowing to work with any application or domain depending on the necessities of the final user. During the long process of execution of the systematic mapping study and with the help of the suggestions received from reviewers it was not found any method, technique or tool that covered the goals of this work. There were none better than another, since some ones had advantages or disadvantages that others not depending on the area in which they were applied or some ones covered a set of functionalities that others do not. MODEL-DRIVEN ENGINEERING In the research group in which this Doctoral Thesis has been developed, one of the main lines of research that exists is the Model-Driven Engineering (MDE) paradigm. It allows to focus on the concepts and their relationships and to be free of their concrete representation, representing them abstractly. Chapter II showed that the research topics in ER are not focused in the level of automation, multi-applications, multi-domain and design of the problem, the most of techniques, tools or methods found are focused in improving an older solution or compare the results of one proposal with the new one proposed. In this sense, the Model-Driven Engineering (MDE) paradigm provides the level of abstraction and automation that requires this proposal and therefore, it will be the technological support in which the solution of this work will be developed. Also, MDE provides a lot of advantages such as:
32 less error-prone, increased quality, it is more cost-effective or facilitates the reuse of parts of the system in new projects among others. In the scientific activity, abstraction has been and is widely used, and often referred to it as the activity of modeling. If a model is defined as a partial or simplified reality representation, that allow the final user to address a complex task for a specific purpose, it can be considered a model as the result of abstraction (García-Borgoñón, 2015). The complexity of software development has been growing up drastically. In this sense, developers noted in models, an alternative for addressing this complexity. MDE emerged to address the complexity of software systems in order to express the concepts of the problem domain in an effective way (Schmidt, 2006). Thus, in the early stages of development, models are more abstract than in the final stages where the models are much closer to implementation. It means, abstract models are transformed into concrete ones the aim of producing software. Studying this process, Brambilla et al., (2012) defined the two fundamental pillars of the MDE paradigm for creating software automatically: models and transformations. Models must be defined according to the rules of a concrete Modelling Language (ML). This language defines the syntax and semantic of the model. Figure III-3 (Metzger, 2008) shows graphically these relations between syntax and semantic. The ML syntax is composed of a concrete and an abstract syntax. The abstract one defines the language structure and how the different elements can be combined, regardless of its representation. The semantic one, that provides the static and dynamic part, poses restrictions and establishes the meaning of the elements of the language and different ways to combine them. In this moment, it appears the concept of metamodel. A metamodel can be defined as a special type of model that specifies a ML. The metamodel defines the structure and constraints for a family of models (Mellor et al., 2004). Transformations are the mechanisms that allow to derive models from other existing ones. A transformation between models represent a relation between two abstract syntaxes and it is defined by a set of relations between the elements of the metamodels (Thiry and Thirion, 2009). There are two types of transformations: horizontal (the derived model and the original one have the same abstraction level) and verticals (the derived model has a lower abstraction level than the original one). Figure III-3. Syntax and Semantic Relation of the Model (Metzger, 2008)
33 A very interesting concept found in the MDE literature is proposed by Bézivin, (2005) where “Everything is a model”. In this sense transformations, themselves, are also considered as models. Generally, Figure III-4 shows the transformations between models from a MDE perspective. A transformation models program takes as input a model according to an origin metamodel and produces as output a model according to the target metamodel. The transformation program, should be considered as a model itself. One of the advantages of MDE is its support for automation, as the models can be automatically transformed from the early stages of development to the final stages. Therefore, MDE allows automating the tasks involved in a software development, such us the testing tasks. These are, briefly, the main concepts that work in model-based or MDE environments and which are going to produce the theoretical fundaments for building the solution. Metamodel (MOF) Transformation Language Definition of Transformations Origin Metamodel Target Metamodel Origin Model Target Model conformsTo conformsTo conformsTo conformsTo conformsTo conformsToconformsTo Execution of Transformations Figure III-4. Transformations between Models. VIRTUAL GRAPHS Graph technology is a natural solution to treat problems related to data management and especially for the relationships between entities. In the context of entity reconciliation problems, they do not only lie in comparing the content of the entities, but it also is related to compare the context of the entities for trying to discover entities that do not seem to be the same. This makes not all the data structures are good to represent data and considering that graph structures are the most versatile, it has been selected for representing the reconciled data. The wide variety of existing algorithms, for example: Dijkstra, A*, Kruskal, etc. offer great flexibility in different situations. Theoretically, graphs can be displayed in two ways: explicit and implicit.
34 An explicit graph is a collection of items (vertexes and edges) that can be stored in memory, which means that each vertex and each edge of the graph can be completely stored in memory. The problem of using graphs technology is that if they are used implicitly, they occupy a lot of memory resources and in current times where Big Data claims a lot space, makes the task of information processing a handicap. Thus, it was decided to use the virtual version of graph technology for representing data to be reconciled. With this technology, it is possible to build structures on the fly. This will let building different solutions to address many scenarios within a business logic where the predefined data model cannot meet the extensibility or availability of the required data sources. Also, the processing time will be significantly reduced. A virtual (or implicit) graph is a graph that cannot be completely stored in memory for various reasons, such as size or hardware limitations (Mondal and Deshpande, 2012). It is possible to add the label virtual to a set of elements when it is not needed or it is not wanted to have all the elements of the set on memory and there are only available some properties of it. For instance, the virtual set of the even integers that compose the range [a,b). It is easily possible to implement the operations of the type of the set without the must of building a data structure that contains all the elements of the set in memory. A virtual graph is something similar. When the necessities of the problem do not require to store all the vertex and edges of the graph in memory (caused by its size or by the characteristics of the problem), but it is possible to calculate all their neighbors and edges that connect them by an algorithm, we have a virtual graph. A virtual graph has the same types of an explicit graph, such as: directed, non-directed, simple, multigraph, pseudo graph among others. All the vertex and edges of the graph must be unique. For building them, it is available a factory of vertex and another one of edges. In a virtual graph, it is not important adding or removing vertex or edges, knowing the full set of edges or vertex and iterate over them. By the other hand, it is very interesting to decide if a vertex or an edge belongs to a graph or not, knowing if there are an edge between two given vertexes or obtaining all the edges that exists between these two vertexes. 5. CONCLUSIONS In this chapter, the relevant aspects that determine the problem to be solved has been presented. Next, the objectives of the doctoral thesis have been presented. These objectives can be summarized in: study of the state of the art of the different existing solutions for the entity reconciliation of heterogeneous data sources, develop a Framework for designing the entity reconciliation models, provide a support tool for the framework and evaluate the results obtained of the application of the proposal in a real-world case study. This chapter continues presented the MaRIA (Model-driven entity ReconcilIAtion) Framework which proposes this doctoral thesis. This Framework is composed of three
35 main pillars: MaRIA Process (that can be integrated in any software development methodology for solving entity reconciliation problems) MaRIA Metamodel (that will allow to model the entity reconciliation solutions) and Derivation Mechanisms (that will allow to test the designed models). In the following chapters, the Framework will be described depth.
36 CHAPTER IV MARIA (MODEL-DRIVEN ENTITY RECONCILIATION) PROCESS
37 CHAPTER IV. MARIA (MODEL-DRIVEN ENTITY RECONCILIATION) PROCESS hapter III has defined a detailed approach of the solution offered by this Doctoral Thesis, defining the relevant aspects that determine the problem to be solved, the objectives to be achieved in this work, a high level formal definition of the proposed MaRIA Framework (covering requirements, analysis and testing phase of any software development methodology) and the conceptual and technological influences present in this work. This Chapter presents the MaRIA process, one of the three pillars that define the MaRIA Framework. The name of the proposal, MaRIA (Model-driven entity ReconcilIAtion), is because of the close relationship with the Model-Driven Engineering paradigm and with the Entity Reconciliation process. 1. INTRODUCTION In this Doctoral Thesis, it has been developed a Framework that is composed of three main pillars: the MaRIA process, a Model-Driven Approach and the MaRIA tool. This chapter aims to explain in detail the MaRIA process. MaRIA process is defined by a set of activities to be incorporated into any development methodology of any organization to allow preparer the system that is being developed to address an entity reconciliation problem. Concretely, this proposal indicates what are the activities and the artifacts necessary to be able to carry out this type of development. At this point, it is very important to note that the scope of this Doctoral Thesis is to cover the requirements, analysis and testing processes and also, to keep in mind that the goal of the PhD student is to continue this work extended it to all the different processes that complete the software development process of a software system. 2. INCLUDING MARIA PROCESS IN A SOFTWARE DEVELOPMENT METHODOLOGY One of the big research areas of the research group where this Doctoral Thesis has been developed is the Web engineering. Then a Web engineering methodology has been the chosen to include the MaRIA Process to show a real use case application. During the last years, the Web engineering community has proposed several different methodologies for Modeling Web applications with different concepts and definitions such as UWE (UML-based Web Engineering) (Koch et al., 2008), WebML (The Web Modeling Language) (Ceri et al., 2000), OOH4RIA (Meliá et al., 2008), RUX-Method (Rossi et al., 1995) or Navigational Development Techniques (NDT) (Escalona and Aragón, 2008) methodology among others (Domínguez-Mayo et al., 2014). C
44 Analyze Data Access is the activity where the data analysts must study and analyze the data access way of the data sources that the data consumers have defined as problematics. In this sense, they must analyze if each data source refers web service, ODBC any other type of database connection or access. Analyze Records Number is the activity where the data analysts must study and analyze the number of records of each of the data source that the data consumers have defined as problematics (hundred, thousands, millions, etc.). This activity will provide some knowledge and help to choose the technology in which the system will be implemented in the design phase. Generate Data Sources Report it the activity where the data analysts must create the data sources report with all the information studied and analyzed in previous steps. 3.1.2. ANALYSIS PROCESS Analysis process of NDT methodology is composed of five main activities: define services, perform the analysis class model, perform navigational model, perform the set of prototypes and generate analysis document. Define services. This activity aims to define the services that must be created in the software system to be developed. Perform the analysis class model. The conceptual model represents the static structure of the system and it lets to model how the information structure that the system manages will be. It is represented by two elements: conceptual diagram classes and data dictionary. Perform navigational model. The navigational model represents how the user will be able to navigate through conceptual information, which elements will appear in that navigation and how they will adapt to the user interacting with the system and finally, the relationships that appear between said elements of navigation. Perform the set of prototypes. This activity aims to generate a set of prototypes that facilitate the task of validating the models developed in previous activities. Generate analysis document. This final activity describes the process of creating the documentation that collect all what have been created in previous activities. In this context, what MaRIA process proposes is to add three new activities called “Review Data Sources Report”, “Define Entity Reconciliation Problem” and “Generate Analysis Model” in the analysis process group of NDT methodology (Figure IV-7).
45 Define Services Perform the Analysis Class Model Perform Navigational Model Analyze Entity Reconciliation Problem Perform the Set of Prototypes Generate Analysis Document Analysis Model Analysis Document Define Analysis Model Review Data Sources Report Data Sources Report Figure IV-7. Extended analysis process Review Data Sources Report activity is based on the on the study, analysis and review by the software engineer of the strategy document generated in the requirements process group (this document is received as input for performing this activity). This action will provide the software engineer with the knowledge of the data sources needed to model the entity reconciliation problem in the next activity. Define Entity Reconciliation Problem is the activity where the software engineer must model the entity reconciliation problem. This activity will be carried out using the support tool developed in this doctoral thesis, the MaRIA tool, described in Chapter VII. This activity is composed of seven main steps (Figure IV-8). The four first ones, define wrappers, define data sources, define entities, define attributes, can be performed in parallel. Once defined, the software engineer hast to define the connectors, the data structure and finally, define the transformations. Figure IV-8. Define Entity Reconciliation Problem
46 Define Wrappers is the activity where the software engineer must model the wrappers that will allow the transfer of information from the data sources that the data consumers have defined as problematics into the entities that will be defined later. In this sense, the software engineer must use the wrapper element of the MaRIA tool and model in the diagram one element for each data source. Define Data Sources is the activity where the software engineer must model the data sources that the data consumers have defined as problematics. In this sense, the software engineer must use the data source element of the MaRIA tool and model in the diagram one element for each data source. Define Entities is the activity where the software engineer must model the entities where the information coming from the data sources will be stored in. In this sense, the software engineer must use the data entity element of the MaRIA tool and model in the diagram one element for each data source. Define Attributes is the activity where the software engineer must model the attributes that compose the entities. In this sense, the software engineer must use the data source attribute element of the MaRIA tool and model as many attributes as each entity needs. Define Connectors is the activity where the software engineer must model the connectors between the elements that have been defined in the three previous steps. The connectors will relate the wrappers with the data sources, the data sources with the entities and the entities with their attributes. In this sense, the software engineer must use the different types of connectors element that MaRIA tool offers depending on the elements needed to be related. Define Data Structure is the activity where the software engineer must model the data structure where the data reconciled of the final solution will be stored. For performing this activity, the user must use the entity, attribute and connector elements of the MaRIA tool. In this activity, the software engineer must create the entities, related between them if necessary and the attributes that describe each entity. For each data source, it must be created a structure and in addition to this, it must be created another one that will store the final solution. Define Data Transformations is the activity where the software engineer must model the data transformations between the different attributes already created or between the data structures. In this sense, the software engineer must use the transformation operations that the MaRIA tool offers and use it for relating attributes or data structures depending on the necessities of the problem. Finally, the main goal of the Generate Analysis Model activity is to generate the model that have been defined thanks to the achievement of the previous activities. 3.2. TESTING PROCESS GROUP The Testing process group of NDT methodology is composed of four main set of processes: tests of the project, organize the testing phase, manage the testing phase and run the tests (Figure IV-9).
47 Tests of the project. This generic process groups all the processes related to the Testing in the development of a software system. All of these processes are based on ISO/IEC/IEEE 29119 (ISO/IEC/IEEE, 2013). Organize the testing phase. This process formally defines a set of activities for carrying out the specification, maintenance and continuous improvement of the strategic information of the tests within the organization. The information that this specification must contemplate are the objectives and the global scope of the tests within the organization. This specification should also include organizational testing practices and provide a framework for the continuous review and improvement of these policies within the organization, as well as considering aspects such as: what type of evidence will be carried out in the organization, who will be the responsible of them, how and with what techniques they will be executed or what tools will be used among others Manage the testing phase. The test management process can be performed at the software project level for different stages of life cycle testing, it can be represented as: unit testing in development, acceptance testing in implementation or early testing of requirements among others. Run the tests. This phase of the test cycle will consist of those activities that will allow the execution of the test plan and the results of the tests. Once the test plan and control mechanisms have been defined in the test management process, the execution of the test management process will begin. Figure IV-9. Extended testing process group In this context, what MaRIA process proposes is to add two new activities called “Execute Analysis Model” and “Apply Specific Criteria” and the generation of a new document called “Business Rules Document” in the testing process group of NDT methodology (Figure IV-10).
48 Design and Implement Tests Generate B. Rules Document Apply Specific Criteria Execute Tests Test Plan Document Test Plan Document Business Rules Document Report Incidences on Tests Incidence Analysis Model Figure IV-10. Extended Run the tests process Execute Analysis Model activity receives as input the “Analysis Model” generated in the requirements process group. This model contains the definition of the entity reconciliation problem to solve in the software system that it is being developed. This model also contains a set of derivations defined for the transformations that the information must suffer during the process of reconciliation. The execution of this model will generate the “Business Rules Document”. Business Rules Document is the document generated after performing the execute analysis model activity. This document will contain all the business rules that will define the test requirements of the software system that it is being developed after the application of a concrete criteria. Apply Specific Criteria activity receives as input the business rules document. To generate the test cases of the software system that it is being developed, it is necessary to apply a specific criterion to the business rules that the business rules document contains. One example of criteria may be Modified Condition/Decision Coverage (MCDC) (Chilenski, 2001). This coverage criterion has demonstrated its utility in previous work, such as (Tuya et al., 2010) (for testing SQL queries) and (Blanco et al., 2012a) (for testing the user-database interaction). 4. CONCLUSIONS This Chapter has presented the MaRIA Process, one of the three pillars that define the MaRIA Framework. The name of the proposal, MaRIA (Model-driven entity ReconcilIAtion), is because of the close relationship with the Model-Driven Engineering paradigm and with the Entity Reconciliation process.
49 MaRIA process is defined by a set of activities to be incorporated into any development methodology of any organization to allow preparer the system that is being developed to address an entity reconciliation problem. Concretely, this proposal indicates what are the activities and the artifacts necessary to be able to carry out this type of development. This Doctoral Thesis covers the requirements, analysis and testing processes and keeping in mind that, the goal of the PhD student is to continue this work extended it to all the different processes that complete the software development process of a software system. Due to the great baggage of the development of projects of the research group where this Doctoral Thesis has been developed, as well as the high experience and acceptance rate that this methodology has demonstrated in several projects, Navigational Development Techniques (NDT) methodology was selected to integrate MaRIA process in a real context. Finally, considering the scope of this work (the requirements, analysis and testing processes), NDT was extended in its software development process (requirements and analysis processes group) and in its testing process group (run the tests process group), creating a total of 15 new activities and 3 new documents. It is very important to note that, although to illustrate the application of MaRIA Framework in a real software development methodology, it has been used NDT, it may be applied to any software methodology such as SCRUM (Schwaber and Beedle, 2001) or BDD (Lazǎr et al., 2010) among others.
50 CHAPTER V DEFINITION OF METAMODELS
51 CHAPTER V. DEFINITION OF METAMODELS hapter IV has defined a detailed view of the MaRIA Process that is part of the MaRIA Framework. This chapter, describes the languages and concepts needed for modelling the Reconciliation system to covering the requirement, analysis and testing set of processes of the Model-Driven Approach that MaRIA Framework proposes. This approach is composed of two main pillars: the MaRIA metamodel, described in this chapter and a set of derivation mechanisms, described in Chapter VI. The proposed metamodels are formally defined and formally represented by UML class diagrams. The formal definition of the languages indicated in the previous paragraph is taken as a basis in later chapters to propose derivations that will let to be entity reconciliation problem modeled with the MaRIA metamodel systematically tested. Finally, this chapter synthesizes a set of conclusions. 1. MODEL-DRIVEN APPROACH MaRIA metamodel is one of the two pillars that conforms Model-Drvien Approach of the MaRIA Framework. Figure V-1 shows the position that this chapter is focusing inside the Framework. Figure V-1. Framework MaRIA Metamodel The metamodel for the designing of Entity Reconciliation solutions is showed in Figure V-2. It is the final version of the proposal presented in (Enríquez et al., 2015). As it is possible to note, the global view of the metamodel is composed of these four main blocks: virtual graph metamodel, data source metamodel, transformation metamodel and the testing metamodel, all related between them. Support Tool MaRIA Tool Framework MaRIA MaRIA Process Model-Driven Approach Metamodel Derivations C
52 Figure V-2. Global view of MaRIA Metamodel Data Sources Metamodel: allows representing the information of the data sources to reconcile and the way of accessing to them. These data sources can be a structured or unstructured database, a web service, a warehouse or another information generator. Virtual Graph Metamodel: allows the user to design the conceptual data model that represents the reconciled solution to achieve, according to the ER problem domain, as a virtual graph. Transformations Metamodel: represents the different transformations that the data of the sources must undergo to carry out the entity reconciliation and to be consistent with the reconciled solution model. The description of this model is out of the scope of this work. Testing Metamodel: allows representing the testing objectives for the entity reconciliation application in the early stages of the development. The test models can be focused on different level, as unit testing or integration testing. During the process of metamodels design, it has been included the attributes required for meeting with the standard ISO/IEC TR 24774 (OMG Group, 2010). To this metamodel, it will be applied a set of derivations that will generate automatically the test requirements from the model specified by the user. These transformations will be applied to the relations that have been defined in the model. They will be transformations from a MaRIA metamodel to text (business rules). In this sense, the source and the target of these transformations are: Source: MaRIA Metamodel. Target: Set of business rules derived from the MaRIA Metamodel in the form of business rules.2 Virtual Graph Metamodel Transformations Metamodel Data Source Metamodel Testing Metamodel
53 2. DATA SOURCE METAMODEL This section describes the Data Source Metamodel and all its components. Figure V-3 shows the metaclases that compose this metamodel and all their relationships. Also, with a low transparency level, it is shown the relationships with metaclases that belong to other blocks of the metamodel. As mentioned before, this block of the global metamodel allows representing the information of the data sources to reconcile and the way of accessing to them. These data sources can be a structured or unstructured database, a web service, a warehouse or another information generator. Figure V-3. Data Source Metamodel This metamodel is composed of five main metaclasses: «Wrapper», «DataSource», «DataSourceEntity», «DataSourceEntityLink» and «DataSourceAttribute». Now, it will be defined each metaclass in detail, for that, it will be exposed their descriptions, generalizations, attributes, operations, associations and constraints. 2.1.«DATASOURCE» METACLASS Description: This metaclass allows to represent a data source. The final model that the user must define will have as many instances of this metaclass as number of data sources that want to be reconciled.
60 Attributes: This metaclass is defined by two attributes: - destination: String [1] Instance of the destination «EntityVertex» metaclass. - source: String [1] Instance of the source «EntityVertex» metaclass. Operations: N/A Associations: This metaclass is defined by three associations: - has: Graph [1] This relation represents that an instance of the «AssociationEdge» metaclass may only be part of an instance of the «EntityVertex» metaclass. - has: EntityVertex [2] This relation represents that an instance of the «AssociationEdge» metaclass must be part related with two instances of the «EntityVertex» metaclass, being one the source and the other one, the destination. - has: ResolutionContext [1..*] This relation represents the dependency of this block of the metamodel with the testing metamodel, being able to generate business rules trough derivations later. Constraints: N/A 3.5.«VIRTUALGRAPH» METACLASS Description: This metaclass allows to represent a Virtual Graph data structure. This metaclass is modeled as an abstract class that implements the «Graph» metaclass, it will throw an exception in those methods that are not available for a virtual graph and such as: add edges, add vertexes or remove edges among others. Thus, the instantiation of this metaclass produces a virtual graph that will store the entities (and their relationships) that have been reconciled. Generalization: «Graph» metaclass Attributes: N/A Operations: N/A Associations: N/A Constraints: N/A
61 4. TRANSFORMATIONS METAMODEL This section describes the Transformation Metamodel and all its components. Figure V5 shows the metaclases that compose this metamodel and all their relationships. Also, with a low transparency level, it is shown the relationships with metaclases that belong to other blocks of the metamodel. As mentioned before, this block of the global metamodel allows the user to represent the different transformations that the data of the sources must undergo to carry out the entity reconciliation and to be consistent with the instantiation of the «VirtualGraph» metaclass. Figure V-5. Transformation Metamodel This metamodel is composed of sixteen metaclasses: «Context», «Pattern», «Rule», «TransformationContext», «TransformationRule», «ResolutionContext», «ResolutionRule», «ResolutionPatter», «ResolutionClausule», «AggregationClausule», «TransformationPattern», «TransformationClausule», «Load», «FilterClausule», «ConversionClausule» and «SurrogateClausule». Now, it will be defined each metaclass in detail, for that, it will be exposed their descriptions, generalizations, attributes, operations, associations and constraints. Some of the selected operations have been chosen taking the criteria presented by Trujillo & Luján-Mora, (2003).
62 4.1.«CONTEXT» METACLASS Description: This metaclass allows to represent the connections between the types of entities of the «DataSource» metaclass instances and the types of entities of the «VirtualGraph» metaclass instances. These connections impose conditions to be fulfilled to project the entities of the data sources to the entities of the reconciled solution. Generalization: N/A Attributes: N/A Operations: N/A Associations: This metaclass is defined by five associations: - has: Rule [1..*] This relation represents that an instance of the «Context» metaclass may be part of one or more instances of the «Rule» metaclass. - has: Pattern [1] This relation represents that an instance of the «Context» metaclass may only be part of an instance of the «Pattern» metaclass. - has: DataSourceEntity [1..*] This relation represents the dependency of this block of the metamodel with the data source metamodel, being able to generate business rules trough derivations later. - has: EntityVertex [1..*] This relation represents the dependency of this block of the metamodel with the virtual graph metamodel, being able to generate business rules trough derivations later. - has: Attribute [1..*] This relation represents the dependency of this block of the metamodel with the virtual graph metamodel, being able to generate business rules trough derivations later. Constraints: N/A 4.2.«RESOLUTIONCONTEXT» METACLASS Description: This metaclass allows to represent the connections between the types of entities of the «DataSource» metaclass instances and the types of entities of the «VirtualGraph» metaclass instances for the resolution process. Generalization: «Context» metaclass Attributes: N/A Operations: N/A Associations: N/A
63 Constraints: N/A 4.3.«TRANSFORMATIONCONTEXT» METACLASS Description: This metaclass allows to represent the connections between the types of entities of the «DataSource» metaclass instances and the types of entities of the «VirtualGraph» metaclass instances for the transformation process. Generalization: «Context» metaclass Attributes: N/A Operations: N/A Associations: N/A Constraints: N/A 4.4.«PATTERN» METACLASS Description: The instantiation of this metaclass allows to impose the conditions that lead the actions of the Entity Reconciliation process. Generalization: N/A Attributes: N/A Operations: N/A Associations: This metaclass is defined by four associations: - has: Rule [1] This relation represents that an instance of the «Pattern» metaclass may only be part of an instance of the «Rule» metaclass. - has: Context [1] This relation represents that an instance of the «Pattern» metaclass may only be part of an instance of the «Context» metaclass. - has: Attribute [0..*] This relation represents the dependency of this block of the metamodel with the data source metamodel, being able to generate business rules trough derivations later. Constraints: N/A
64 4.5.«RESOLUTIONPATTERN» METACLASS Description: The instantiation of this metaclass allows to impose the conditions that lead the actions of the Entity Reconciliation process in the resolution step. Generalization: «Pattern» metaclass Attributes: N/A Operations: N/A Associations: N/A Constraints: N/A 4.6.«RESOLUTIONCLAUSULE» METACLASS Description: The instantiation of this metaclass allows to impose the conditions that lead the actions of the Entity Reconciliation process in the resolution step, stablishing the level of priority between data sources. Generalization: «ResolutionPattern» metaclass Attributes: This metaclass is defined by three attributes: - name: String [1] Short and concise description with which users can uniquely identify this operation. - description: String [1] Detailed description about the process that this operation must follow. - priority: int [1] This attribute defines the level of priority between data sources. Operations: N/A Associations: N/A Constraints: N/A 4.7.«AGGREGATIONCLAUSULE» METACLASS Description: The instantiation of this metaclass allows to impose the conditions that lead the actions of the Entity Reconciliation process in the resolution step, making a concatenation of all the attributes related. Generalization: «ResolutionPattern» metaclass
65 Attributes: This metaclass is defined by two attributes: - name: String [1] Short and concise description with which users can uniquely identify this operation. - description: String [1] Detailed description about the process that this operation must follow. Operations: N/A Associations: N/A Constraints: N/A 4.8.«TRANSFORMATIONPATTERN» METACLASS Description: The instantiation of this metaclass allows to impose the conditions that lead the actions of the Entity Reconciliation process in the transformation step. Generalization: «Pattern» metaclass Attributes: N/A Operations: N/A Associations: N/A Constraints: N/A 4.9.«TRANSFORMCLAUSULE» METACLASS Description: The instantiation of this metaclass allows to impose the conditions that lead the actions of the Entity Reconciliation process in the resolution step, making a transformation of a source attribute to another one. Generalization: «TransformationPattern» metaclass Attributes: This metaclass is defined by two attributes: - name: String [1] Short and concise description with which users can uniquely identify this operation. - description: String [1] Detailed description about the process that this operation must follow. Operations: N/A Associations: N/A
66 4.10. «LOADCLAUSULE» METACLASS Description: The instantiation of this metaclass allows to impose the conditions that lead the actions of the Entity Reconciliation process in the resolution step, transferring the content of an attribute to another one. Generalization: «TransformationPattern» metaclass Attributes: This metaclass is defined by two attributes: - name: String [1] Short and concise description with which users can uniquely identify this operation. - description: String [1] Detailed description about the process that this operation must follow. Operations: N/A Associations: N/A 4.11. «FILTERCLAUSULE» METACLASS Description: The instantiation of this metaclass allows to impose the conditions that lead the actions of the Entity Reconciliation process in the resolution step, making a filter of the content of one or more attributes. Generalization: «TransformationPattern» metaclass Attributes: This metaclass is defined by two attributes: - name: String [1] Short and concise description with which users can uniquely identify this operation. - description: String [1] Detailed description about the process that this operation must follow. Operations: N/A Associations: N/A 4.12. «CONVERSIONCLAUSULE» METACLASS Description: The instantiation of this metaclass allows to impose the conditions that lead the actions of the Entity Reconciliation process in the resolution step, making a conversion of a source attribute two more than one attributes.
67 Generalization: «TransformationPattern» metaclass Attributes: This metaclass is defined by two attributes: - name: String [1] Short and concise description with which users can uniquely identify this operation. - description: String [1] Detailed description about the process that this operation must follow. Operations: N/A Associations: N/A 4.13. «SURROGATECLAUSULE» METACLASS Description: The instantiation of this metaclass allows to impose the conditions that lead the actions of the Entity Reconciliation process in the resolution step, generating an unique surrogate key. Generalization: «TransformationPattern» metaclass Attributes: This metaclass is defined by two attributes: - name: String [1] Short and concise description with which users can uniquely identify this operation. - description: String [1] Detailed description about the process that this operation must follow. Operations: N/A Associations: N/A 4.14. «RULE» METACLASS Description: The instantiation of this metaclass represents the elements of the rules that constitute this metamodel. It is divided in three mail types: transformation, resolution rules and integration rules (the last ones, defined in the section 5 of this chapter). Generalization: «Context» metaclass Attributes: N/A Operations: N/A
68 Associations: N/A Constraints: N/A 4.15. «TRANSFORMATIONRULE» METACLASS Description: The instantiation of this metaclass represents the elements of the rules that constitute this metamodel in the transformation step. Generalization: «Rule» metaclass Attributes: N/A Operations: N/A Associations: N/A Constraints: N/A 4.16. «RESOLUTIONRULE» METACLASS Description: The instantiation of this metaclass represents the elements of the rules that constitute this metamodel in the resolution step. Generalization: «Rule» metaclass Attributes: N/A Operations: N/A Associations: N/A Constraints: N/A 5. TESTING METAMODEL This section describes the Virtual Metamodel and all its components. Figure V-6 shows the metaclases that compose this metamodel and all their relationships. Also, with a low transparency level, it is shown the relationships with metaclases that belong to other blocks of the metamodel. As mentioned before, this block of the global metamodel allows the user to represent the testing objectives for the entity reconciliation application in the early stages of the development. The test models can be focused on different level, as unit testing or integration testing.
69 Figure V-6. Testing Metamodel This metamodel is composed of six main metaclasses: «IntegrationContext», «IntegrationView», «IntegrationPattern», «Structural», «Load» and «IntegrationRule». Now, it will be defined each metaclass in detail, for that, it will be exposed their descriptions, generalizations, attributes, operations, associations and constraints. It is important to note that «IntegrationContext», «IntegrationPattern» and «IntegrationRule» herigate from the «Context», «Pattern» and «Rule» metaclasses of Transformation Metamodel, although for getting a better clarity of the vision of the metamodel, they have not been added to the figure V-6. 5.1.«INTEGRATIONCONTEXT» METACLASS Description: This metaclass allows to represent the connections between the types of entities of the «DataSource» metaclass instances and the types of entities of the «VirtualGraph» metaclass instances. These connections impose conditions to be fulfilled to project the entities of the data sources to the entities of the reconciled solution. Generalization: «Context» metaclass Attributes: N/A Operations: N/A Associations: This metaclass is defined by five associations: - has: IntegrationRule [1..*] This relation represents that an instance of the «IntegrationContext» metaclass may be part of one or more instances of the «IntegrationRule» metaclass.
76 The proposal presented in this doctoral thesis proposes a set of Model to Text (M2T) transformations to automatically, generate the test requirements of the application that will consume the data that has been generated after the reconciliation process. There are great variety of languages that allow the automatic generation of code from models. Some of the most extended are: Epsilon (Frankel, 2003), a model management platform that provides transformation languages for model-to-model, model-to-text, update-in-place, migration and model merging transformations. MOF Model to Text Transformation Language (Mof2Text or MOFM2T) (Omg, 2008), defined by the Object Management Group (OMG) as a standard for expressing M2T transformations. Java Emitter Templates (JET) (Eclipse, 2016a) used for generating text from Eclipse Modeling Framework (EMF) based models. MOFScript (Oldevik et al., 2005) is a direct result of the OMG RFP for a M2T transformation language. It works with any Meta-Object Facility (MOF) based model and it is very influenced by QVT (Omg, 2008). Basically, it is an imperative language which supports all primitive types and abstract data types. However, this doctoral thesis has bet for the use of Text Template Transformation Toolkit (T4) (Vasudevan and Tratt, 2011) because it is the language defined by Microsoft for M2T transformations in DSL-Tools (tool which has been used for the development of the support tool of the MaRIA methodology). DSL-Tools offers a generator component called “TextTemplating” that allow the generation of code using templates called “T4 templates”. Next sections, will describe the architecture of templates, the typical execution process of a transformation and the main elements of a T4 template following the indications of the manual for domainspecific development with Visual Studio DSL Tools presented by Cook et al., (2007). 1.1.ARCHITECTURE OF TEMPLATES T4 templates are written in C# o Visual Basic language and they have a set of expressions that are exclusives for this technology. The modeling and visualization software of Visual Studio provide a set of tools that allow the generation of code of any type of files taking as a base the content defined in the templates. The transformation of the templates is a process that is performed in two steps the first one, the engine generates a temporal class called “transformation class” that contains the code generated by the control blocks and the directives. The second one take place when the engine is going to compile and execute the transformation class previously generated for creating the output file. The most important components of the tool for generating T4 templates (marked with blue text in Figure VI-2) of the modeling and visualization software of Visual Studio are: Item Directive Processor. It represents the classes that manage text template directives. Host. It is the interface between the engine and the user environment. Visual Studio is a host of the text transformation process.
77 Motor. It controls the process of transforming text templates and core of the template system. It oversees processing the templates and creating an output. Figure VI-2. Text Template Architecture (Cook et al., 2007) Once presented the most important elements of the process of transformation, next section will describe the typical execution sequence of the text templates. 1.2.EXECUTION SEQUENCE OF TEXT TEMPLATES Taking as reference the architecture of Figure VI-2, the execution sequence of a text templates usually follows these steps: The host reads in a template file from disk. The host instantiates a template engine. The host passes the template text to the engine along with a callback. The engine parses the template, finding standard and custom directives and control blocks. The engine asks the host to load the directive processors for any custom directives it has found. The engine produces in memory the code for the skeleton of a Transformation class, derived ultimately from TextTransformation. This template has specified the derived ModelingTextTransformation. The engine gives directive processors a chance to contribute code to the Transformation class for both class members and to run in the body of the Initialize() method.
78 The engine adds the content of class feature blocks as members to the Transformation class, thus appearing to allow methods and properties to be “added to the template.” The engine adds boiler plate inside Write statements and the contents of other standard control blocks to the TransformText() method. The engine compiles the Transformation class into a temporary .NET assembly. The engine asks the host to provide an AppDomain in which to run the compiled code. The engine instantiates the Transformation class in the new AppDomain and calls the Initialize() and TransformText() methods on it via .NET remoting. The Initialize() method uses code contributed by directive processors to load data specified by the custom directives into the new AppDomain. The TransformText() method writes boiler plate text to its output string, interspersed with control code from regular control blocks and the values of expression control blocks. The output string is returned via the engine to the host, which commits the generated output to disk. 1.3.ELEMENTS OF A T4 TEMPLATE The principal elements of a T4 Template may be divided in: directives, text and control blocks and utility methods. Next subsections present a brief description of these elements: DIRECTIVES Directives provide instructions to the text template creation engine on how the transformation code and output file should be generated. The directives must be the first elements in a template. There are five types of directives that need to be defined: template, output, assembly, import and include. <#@ DirectiveName [AttributeName = "AttributeValue"] ... #> Figure VI-3. Declaration of a Directive Template: the most usual is to start the templates with this directive since it is the one that specifies how to process the template. It contains five main properties to define: language, compiler options, culture, debug and line pragmas. o Language: specifies the language to be used as source code in the templates. By default, it is C#. o Compiler options: are compile options for customizing compiler behavior. o Culture: is an attribute used to know the cultural reference of a file when it is generated. o Debug: if this attribute is set as true value it allows the debugging of the code and if it is false does not allow it. o LinePragmas: allows the compiler to display in debug mode the lines of errors either the template or generated code.
79 <#@ template [language="VB"] [hostspecific="true|TrueFromBase"] [debug="true"] [inherits="templateBaseClass"] [culture="code"] [compilerOptions="options"][visibility="internal"] [linePragmas="false"] #> Figure VI-4. Declaration of a Template directive Output: this directive defines the file extension to be generated as the output of the template execution. It also allows to change the encoding of the output file. It contains two main properties: extension and encoding o Extension: represents the extension of the generated file, by default, is represented as “.cs” but may be any type of needed file (“.json”, “.java” or “.html” among others). o Encoding: represents the encoding that the output file will have when it is generated (“utf-8” or “us-ascii” among others) <#@ output extension="fileNameExtension" [encoding="encoding"] #> Figure VI-5. Declaration of an Output directive Assembly: represents a reference to an assembly code so that the template can use the types. <#@ assembly name="[assembly strong name|assembly file name]" #> Figure VI-6. Declaration of an Assembly directive Import: represents the equivalent translation in C# to “using” or “import” in Java language. <#@ import namespace="System.IO" #> Figure VI-7. Declaration of an Import directive Include: allows access from the template that contains the include directive to the templates referenced by this directive. It is used to be able to reuse code between templates. <#@ include file="filePath" [once="true"] #> Figure VI-8. Declaration of an Include directive
80 TEXT BLOCKS It allows inserting text in the output file. It does not require to have a certain format or make use of functions, what is written will be inserted as plain text in the output of the template. CONTROL BLOCKS The control blocks allow writing the template and being able to change the application context of the instructions. This will be very helpfully to create any type of template. There are three type of control blocks. There are differentiated by the opening brackets as well as by the functionality they allow to perform. Those are: standard control block, expression control block and functions control block. Standard control block (<# .... #>): contain the instructions and the blocks can be opened and closed in the middle of sentences (if or for structures among others). <# if (test) { #> // Do something <# } #> Figure VI-9. Standard Control Block Expression control block (<# = #>): is used for code that returns a string that we want to be in the output file. <# string imports = “using System.Collections.Generic”; imports += “\nusing System”; <#= imports #> Figure VI-10. Expression Control Block Class Feature control block (<#+ ... #>): is used to define auxiliary functions or functions that will be reuse within a template. There can only be one block of this type in each template. Everything in this block will be static. <#+ public void GenerateEmptyClass(string name) { #> public partial class <#= name #> { // Some class content } <#+ } #> Figure VI-11. Class Feature Control Block
81 UTILITY METHODS Utility methods allow to have a minimum of tools for writing T4 templates. These methods will always be accessible from the templates and will not have to be imported. They can be briefly classified in: write, bleeding and warning and error methods. Write Methods: they are Write() y WriteLine() methdos that allow write text inside the code. <# int i = 10; while (i-- > 0) { writeLine((i.toString())); } #> Figure VI-12. Write Method Bleeding Methods: these methods are used to format the output generated by the template. It can also be done with the "\t" sentence. o CurrentIndent: Shows the bleeding that is currently being carried. o IndentLenghts: list of bleedings that have been added. o PushIndents: adds a bleeding. o PopIndents: removes a bleeding. o ClearIndents: Cleans the stack of bleedings. Warning and Error Methods: are methods that allow to display errors in the Visual Studio bug list if there were any. <# try { string str = null; writeLine(str.Length.toString()); } catch (Exception e) { Error(e.Message); } #> Figure VI-13. Error Method Once a clear vision of how the transformations must defined and all the components of the DSL-Tools for this purpose have been presented, the next section will present how the transformations have been defined for the proposal of this doctoral thesis.
82 2. INTEGRATION RULES DEFINITION The integration rules, which are statements that define or constraint the business structure or the business behavior (Hay and Healy, 2000), have been used in other approaches focused on testing database applications, such as (Blanco et al., 2012) and (Willmor et al., 2006). On the other hand, as the integration rules are based on the system specification, they could also be used to generate some implementation of the ER application. These integration rules are specially focused on the subsequent derivation of test coverage items that guide the creation of the test data sources and the test reconciled solution. As stated in Chapter IV, integration rules may be defined by two ways: load and structural rules. In this sense and to delimit the scope of this Doctoral Thesis, it will be only covered the integration testing of the model, what means, the load transformations of data and the data structure where the solution will be stored. This concept has been taken from the black box testing method (Beizer, 1995). The following subsections aims to present the patterns that allow expressing the integration context, the integration context view, as well as the integration patterns of each type of integration rule (load and structural). 2.1.INTEGRATION CONTEXTS AND INTEGRATION CONTEXT VIEWS To describe the integration context and the integration context views of an integration rule, it is necessary to define the concept path that is used in their construction. A path (P) is a set of one or more types of entities (instances of the metaclasses DataSourceEntity and EntityVertex) and/or types of relationships (instances of the metaclasses AssociateEdge and DataSourceEntityLink) R1, R2, …, Rn, where each pair (Ri, Ri+1) is directly connected via some attributes in the predicate qi,i+1(): Path P is R1 [q1,2()] R2 [q2,3()] … [qn-1,n()] Rn Each qi,i+1() can contain arithmetic and logical expressions and functions, which involve attributes of R1, R2, …, Ri+1. The definition of the concept path suggests the redefinition of the integration context, the integration context view, as well as the context entities, context relationships and context attributes in terms of this concept, as explained below: An integration context (IC) is a set of one or more paths P1, P2, …, Pm that define the connections between the data source models and the reconciled solution model that are involved in a test condition: Integration context IC is P1, P2, …, Pm
83 If an integration context is formed by only one path, it can be defined directly by: Integration context IC is R1 [q1,2()] R2 [q2,3()] … [qn-1,n()] Rn An integration context view or view, for short, (VIC) of an integration context IC is a subset Rj, Rj+1, Rj+2, …, Rk of a path P of IC, where each pair (Ri, Ri+1) (i=j..k-1) is directly connected via the predicate defined in P: Integration context view VIC is Rj [] Rj+1 [] … [] Rk of IC.P A context entity is a type of entity R of a path P of an integration context IC denoted by IC.R. If R is not unique in IC it is denoted by IC.P.R, where P is a path of IC that contains R. A context entity of a view VIC of an integration context IC is denoted by VIC.R. A context relationship is a type of relationship R of a path P of an integration context IC denoted by IC.R. If R is not unique in IC it is denoted by IC.P.R, where P is a path of IC that contains R. A context relationship of a view VIC of an integration context IC is denoted by VIC.R. A context attribute is an attribute A of a context entity or a context relationship of an integration context IC denoted by IC.A. If A is not unique in IC it is denoted by IC.P.R.A or IC.R.A, where P is a path of IC and R is a context entity of P that contains A. A context attribute of a view VIC of an integration context IC is denoted by VIC.A or VIC.R.A. 2.2.SPECIFICATION OF STRUCTURAL RULES A structural rule establishes the projection from a context entity IC.R (or VIC.R) that belongs to a data source model to one or several context entities and context relationships IC.Si (or VIC.Si) that belong to the reconciled solution model. It also establishes one or several conditions on the context attributes IC.Si.Aj (or VIC.Si.Aj) that constrain their values when the new entities and relationships are created into the current reconciled solution. The projection imposed by the structural rule must be fulfilled by each instance of IC.R (or VIC.R) that belongs to the unreconciled context domain of IC (or the unreconciled view domain of VIC). The integration pattern of a structural rule is described below, using the EBNF notation (Horrocks et al., 2004). The integration pattern of a structural rule is defined as: structural_rule = “Each unreconciled” (IC.R | VIC.R) “generates” gen_cond {“and” gen_cond}; gen_cond = “exactly one” (IC.Si | VIC.Si) “with” att_cond; att_cond = (IC.Si.Aj | VIC.Si.Aj) “=” pj {“and” (IC.Si.Aj | VIC.Si.Aj) “=” pj}; where each pj is a predicate over context attributes of IC.R (or VIC.R) and/or IC.Si (or VIC.Si).
84 2.3.SPECIFICATION OF LOAD RULES A load rule imposes one on several conditions that constrain the value of a context attribute IC.S.A that belongs to the reconciled solution model, according to one or several context attributes IC.Ri.Bj that belong to the data source models. The conditions must be fulfilled by each tuple of the reconciled context domain of IC. The load rules are classified according to two dimensions. The first dimension indicates whether a load rule establishes preconditions that must be fulfilled before constraining the value of a context attribute (conditional rules), or it does not establish any precondition (non-conditional rules). The second dimension indicates the types of conditions that constrain the value of the context attributes according to one or several predicates (IS, OR, AND, XOR rules). These predicates can be either arithmetical or logical expressions or functions over context attributes of the integration context IC, as well as constants or context attributes of IC. The evaluation of the predicates returns a value that fits the type of the context attribute constrained or a null value, which indicates that the predicate was not able to reach a concrete value. The following definitions describe the patterns of each category, using the EBNF notation. A conditional rule is a load rule whose integration pattern is defined as: conditional_rule = “If” p “then” rule_pattern; where p is a predicate over context attributes IC.S.A and/or IC.Ri.Bj whose evaluation returns a boolean value. This predicate defines the preconditions to be fulfilled before constraining the value of the context attribute IC.S.A by means of rule_pattern (described next). An IS rule is a load rule that constrains the value of a context attribute IC.S.A, such that it must be equal to the evaluation of a predicate p. The integration pattern is defined as: IS_rule = “Each” IC.S.A “is” p; An AND rule is a load rule that constrains the value of a context attribute IC.S.A, such that it must be formed by the union of the evaluations of the predicates pi that do not return a null value. The integration pattern is defined as: AND_rule = “It is obligatory that” IC.S.A “is composed of ” pi {“and” pi}; An OR rule is a load rule that constrains the value of a context attribute IC.S.A, such that it can be formed by the evaluation of one or several predicates pi that do not return a null value. The integration pattern is defined as: OR_rule = “It is permitted that” IC.S.A “is composed of ” pi {“or” pi };
85 An XOR rule is a load rule that constrains the value of a context attribute IC.S.A, such that it must be equal to the evaluation of only one predicate pi. Each predicate pi has a different priority ni that indicates the order in which they are evaluated. IC.S.A takes the value of the first predicate pi that does not return a null value. The integration pattern is defined as: XOR_rule = prioritization “Each” IC.S.A “is only ” pi {“or” pi }; prioritization = pi “has priority” ni { pi “has priority” ni } The integration patterns of non-conditional IS, AND, OR and XOR rules are directly described by combining conditional rules with IS, AND, OR and XOR rules defined before respectively. After defining the test conditions as a set of the integration rules, the test coverage items can be derived by means of applying logic criteria, as stated one of the activities defined in MaRIA process of early testing phase over the conditions imposed by these integration rules. 3. DERIVATION MECHANISMS This section details the main derivation mechanisms that aims to transform a model created following the MaRIA metamodel in a set of integration rules defined in section 2 of present Chapter. These transformations take as input the elements identified in the model such as: entities, attributes, data sources or transformations connections among others and generates a business rule conforms to the described previously. As stated in integrated rules definition section, derivation mechanisms defined will be divided in two main categories: structural rules and load rules.
92 Figure VII-3. Eclipse Modeling Meta-Metamodel (Kolovos, 2016) 1.2.METAMODELING The characteristics that each one of the tools offers described in Table VII-1. Modeling Characteristics DSL-Tools Serialization of XML files with .dsl extension Visual diagram of the metamodel in the .dsl.diagram file Multiple inheritance between metaclases Metaclasses with meta-attributes y meta-associations Meta-associations, with roles, multiplicities, navigability and types (assignment or composition) Eclipse Modeling Tools Serialization of XML files with .ecore extension Visual diagram of the metamodel in the .ecore.diagram file Multiple inheritance between metaclases Metaclasses with meta-attributes y meta-associations Meta-associations, with roles, multiplicities, navigability and types (assignment or composition) Table VII-1. Characteristics of DSL and Eclipse Modeling Tools
93 1.3.VISUAL EDITION COMPONENTS Both tools offer a graphical edition interface. In the case of Eclipse (Figure VII-4) the user must derive the graphical definition model, generate graphics and adjust the definition. In the case of DSL-Tools (Figure VII-5), all the changes may be performed from the properties panel. Figure VII-4. Graphical Edition Eclipse Modeling Tools (Eclipse, 2016b) Figure VII-5. Graphical Edition DSL Tools (MSDN, 2016) Similarly, for the mapping of components. On one hand, the visual editor of Eclipse we must derive the mapping model, generate the mapping model and adjust them and on the other and, in Visual Studio all the elements to customize the mappings are placed in the DSL Details window. 1.4.M2T TRANSFORMATIONS DSL Tools uses the “Text template transformation toolkit” and Eclipse uses the “MOFScript” (Figure VII-6). Both environments use similar processes for this kind of transformations. Roughly, the steps that user must take are: Creation of the model. Creation of the file for performing the transformation. Processing of the modeled file. Obtaining the result.
94 Figure VII-6. M2T Transformations 1.5.VISUAL EDITOR DEPLOYMENT Visual Studio allows to generate a VSIX plugin that will allow to use the DSL from an instance of Visual Studio. In Eclipse the user has two options: (i) generate a plugin or (ii) generate a desktop application. In conclusion, both tools seem to be very similar and offer practically the same characteristics. However, for the development of the support tool of this doctoral thesis, it has been selected the Microsoft DSL-Tools for two main reasons: easy to use, it uses a well-known language (C#) by the doctoral student for generating the templates and transformations. Finally, an introduction to Microsoft DSL-Tools and a small guide of how to install it, how to start a new project and the components offered for the creation of a new domainspecific language, is presented in Appendix A. 2. DEFINING THE CONCRETE SYNTAX OF METAMODELS Models and metamodels alone do not explicitly require the use of any particular notation for their representation. The concrete syntax of a metamodel specifies how to visually represent the models through diagrams. It may be represented by two ways: textual or graphical (Fondement and Baar, 2007). As mentioned before, this proposal has been based on the graphical way, supported by the DSL-Tools. In this sense, the concrete syntax will consist of a set of templates, where each template specifies the visual representation of each class of a metamodel. Table VII-1 represents all the elements that have been defined in the concrete syntax that represent the different instances that a metaclass of the metamodel defined in Chapter V may has.
95 Image Tool Reprsentation Data Source Wrapper Data Source Link Entity Attribute Has Link Transformation Link Transform Filter Load Conversion Resolution Link Resolution Aggregation Table VII-1. Concrete Syntax Data Source: represents each data source to be reconciled. Wrapper: represents the way of extracting information from one data source. Data Source Link: represent a connection between the Data Source and the Wrapper. Entity: represents each the entity of the model. Attribute: represents the attributes that define the entity. Has Link: represents the connection between entities. Transformation Link: represents the transformation that data must undergo to carry out the entity reconciliation problem. It is composed by four type of operations: o Transform: an attribute is transformed into another one. o Filter: it applies a filter between some attributes. o Load: it takes a data from a source and moves it to a target. o Conversion: it applies an operation to different attributes. Resolution Link: represents the way in which the reconciliation will be carried out. It is composed by two types: o Resolution: it represents the preferred attribute depending on a priority. o Aggregation: it represents the concatenation of all the values of the attributes involved in the operation. For creating the metamodel presented in Chapter V, using the UML language, to the offered by Visual Studio with DSL Tools, the first step that the doctoral student performed was generating a “Minimal Language” project (see Annex A).
96 DSL Tools contain a set of tools for creating DSLs. Due to the operation of the DSL tools it was necessary to define a “DomainClass” to which must associated a “Diagram” type visual interface (see Figure VII-7). Since all the metaclasses of the metamodel contain an embedded relation with this root element with the same characteristics, it will we removed from the figures for two reasons: (i) improve the quality of the figure and (ii) Do not repeat elements already described. Figure VII-7. Root Element (DomainClass + Diagram) 2.1.DATA SOURCE MODEL This section describes how the data source block of the metamodel presented in Chapter V has been modeled in DSL-Tools. The “DataSource” (Figure VII-8) metaclass has been modeled as a “NamedDomainClass” type So it contains “Name” property by default and it will be unique in all the definition of the metamodel. Also, it has been defined the “URL” property to know the directory path of the data source from which the information wants to be extracted. There is a “Embedding Relationship” type meta-association between the “TreeModel” and the “DataSource” metaclass that indicates that one “TreeModel” may contain zero or more “DataSource”. Also, it is needed another meta-association, in this time of “Reference Relantionship” type that associates the “DataSource” and the “Wrapper” metaclass because it is a reference from one metaclass to the other one. Figure VII-8. Root Element (DomainClass + Diagram)
97 The “Wrapper” (Figure VII-9) metaclass has been defined as a “DomainClass” class because it does not have the necessity of any property for creating the abstraction shape between the “DataSourceEntity” and “DataSource” metaclasses. It is possible to see that it is related with the “DataSourceEntity” metaclass through a “Reference Relationship” meta-association type. Figure VII-9. Wrapper Metaclass The “DataSourceEntity” (Figure VII-10) metaclass is defined as a “NamedDomainClass” type metaclass, it means that, once it has been defined a name, it will not be a other one with the same description in all the diagram. This metaclass has a reference relationship with the “DataSourceAttribute” metaclass. Figure VII-10. DataSourceEntity Metaclass “DataSourceAttribute” (Figure VII-11) metaclass has been defined as a “DomainClass” because the attributes may be duplicated in the diagram. It must be modeled like that because it is necessary for using the same attribute (for example: “descripton”) in more than one instance of the “DataSourceEntity” metaclass. This metaclass contains the “Name” and “Type” string properties, one for specifying the name of the attribute and the other one for specifying the type of the attribute. Figure VII-11. DataSourceAttribute Metaclass
98 2.2.VIRTUAL GRAPH MODEL The “EntityVertex” metaclass (Figure VII-12) is defined as a “DomainClass” metaclass. In this sense, in the final model it could exist more than one instance of the “EntityVertex” metaclass. This metaclass has two properties: “Name” specifying the name of the attributewith a string type “Result” that will stablish if a concrete instance of the “EntityVertex” that has been created as result of a single or a set or operations. Figure VII-12. EntityVertex Metaclass The “Attribute” metaclass (Figure VII-13) follows the same pattern than the “DataSourceAttribute” but with the difference that instead of dealing with transformations it does so with transformations. Otherwise the behavior is similar. Figure VII-13. Attribute Metaclass 2.3.TRANSFORMATION MODEL The “Transformation” metaclass (Figure VII-14) is defined as a “DomainClass” metaclass and it is the metaclass from which all specific transformations inherit, each transformation has a determined behavior so it will behave differently (Figure VII-15). It is possible to see that the multiplicities represent that whenever there is a transformation must have a relation with the attribute, but for an attribute it is not always mandatory to have an associated transformation.
99 Figure VII-14. Transformation Metaclass Figure VII-15. Transformation Operations Metaclasses The “ResolutionPattern” metaclass (Figure VII-16) is defined as a “DomainClass” metaclass and it has a similar representation corresponding to the “Transformation” metaclass, but, in this case, the operations will be carried out between attributes. (Figure VII-15). From this metaclass all specific resolution transformations inherit (Figure VII17), each operation has a determined behavior so it will behave differently. Figure VII-16. Transformation Metaclass
100 Figure VII-17. Resolution Operations Metaclasses 3. EDITOR DESIGN The editor is defined by two types of elements, on one hand, the appearance of the tools that are going to be generated in the tool box and on the hand, the appearance of this tool in the diagram. Firstly, it is necessary to define the appearance that the metaclasses and the metaassociations will have when the execution of the DSL is running. This operation can be easily done adding “Shape” element types for the metaclasses and “Connector” element types for the meta-associations. There are some types of shapes or forms for the metaclasses. In this implementation are used two of them: “Geometry Shape” and “Image Shape” (Figure VII-18). Geometry Shape: shows the metaclass as a form that may be rectangular, oval or rounded. Image shape: it is used when a metaclass behaves like an image but it is necessary that this metaclass still has the logical load that a metaclass has. Figure VII-18. Types of Used Shapes To all the forms can be added son decorators that let change the behavior when the domain classes are in execution. There are three types of decorators (Figure VII-19): Expansion and Collapse Decorator: it will allow that once running the metaclasses instances have a “+” and “–” symbol which will show or hide the properties of the class. Icons Decorator: it will allow add an icon in a concrete position of the instances of the metaclasses in execution time. Text Decorator: it will allow to add the visualization of the properties of the instances of the metaclasses in execution time.
101 Figure VII-19. DataSource Shape Decorador Once created the forms for the metaclasses and the meta-associations, for stablishing the relationship between the metaclasses and meta-associations and the elements of the diagram it will be used the “Diagram Element Map” (Figure VII-20) element. Once selected this tool, for relate a shape to a metaclass the only necessary thing to do is click on the metaclass and click on the shape that wants to be related to. Making this step, both metaclass and shape will be related. Figure VII-20. Diagram Element Map Example Once the relationship has been created the next step is configuring it. For that, clicking on the relationship, it is possible to win the DSL Details windows all the information related with the mapping divided in two tabs: “General” shows all the information autocompleted according to the relationship created and “Decorator Maps” (Figure VII-21) where it is necessary to select the decorator that wants to be used. The same process must be applied to the connectors. Figure VII-21. Decorator Maps
108 Figure VII-34. Example of Modeling 5. CONCLUSIONS During Chapters IV, V and VI, it has been defined by a formal way all the theoretical framework in which the proposal specified in this doctoral thesis is based on. However, to ensure the feasibility and applicability of this theoretical framework within practical environments and a production context, it is convenient and necessary to develop a CASE tool that supports it to make easier the maintenance and improve the quality of the results. This has been the purpose of this chapter and to that end, different aspects have been addressed. At first, the section 2 of this Chapter defines how the concrete syntax of the metamodel proposed in chapter V has been defined. For this purpose and after a comparative study between Eclipse Modeling Tools of Eclipse and DSL Tools of Microsoft, it was decided to use the one provided by Microsoft. Once defined the profile in DSL Tools, section 3 presented the editor design phase. There, it was described how to configure the forms and shapes of the elements (that represent instances of metaclasses) and relationships (that represent instances of metaassociations).
109 Section 4 present how the transformations have been defined to automatically generate the business rules that represent the test requirements of the application to develop and finally, section 5 has present how the final DSL created looks like and an example of model defined with this support tool.
110
111 CHAPTER VIII VALIDATION
112 CHAPTER VIII. VALIDATION his doctoral thesis has presented the MaRIA Framewok and its three main pillars: the MaRIA Process, a set of activities to be added in any software development methodology that allows prepare the software system to be developed, to guarantee a suitable entity reconciliation and having the possibility to be systematically tested, the Model-Driven Approach, composed of the MaRIA Metamodel and a set of Derivations that allows the software engineer to model an test the entity reconciliation problem and finally, the MaRIA Tool, a domain-specific language supported by Microsoft Visual Studio that let some software engineer model entity reconciliation problems in a simple way. Last chapter has presented the more important characteristics of each of the elements that compose the MaRIA Tool. Now it is time to validate the tool. This chapter presents the results of a real proof of concept, based on a problem related to the data management of the cultural heritage of the fixed monuments of the region of Andalusia (Spain). 1. TWO REAL-WORLD CASE STUDIES This section aims to develop to real-world case studies based on different projects that are being carried out in the research group where this Doctoral Thesis has been developed: DIPHDA and ADAGIO. 1.1. DIPHDA One of the main objectives of the “Instituo Andaluz de Patrimonio Histórico (IAPH, Andalusian institute of historical patrimony) in the management of cultural heritage information is concerned with the design and implementation of an effective and operational information system capable of integrating all information resulting from research, documentation, conservation, protection, dissemination, etc. Of the cultural heritage, so that the information reaches the person who needs it at the right time for decision making (Instituto Andaluz de Patrimonio Histórico, 2010). In this context, the IAPH, at the beginning of the 90s, the “Sistema de Información del Patrimonio Histórico de Andalucía” (SIPHA, Andalusian Historical Heritage Information System) was launched, which included, among others, the following advances: Creation of standardized, integrated and computerized standards on the different patrimonial entities. Creation of a standardized documentary language, the Andalusian Historical Heritage Thesaurus, an international pioneer for its multidisciplinary. Incorporation of Geographic Information Systems (GIS). T
113 Photographic and/or audiovisual documentation of patrimony or entities included in the system. Transfer of information through Information Services. Online consultation of the different databases. In recent years, SIPHA sytem has been intetrated with the “Sistema para la Gestión Integral del Patrimonio Cultural” (MOSAICO, System for the Integral Management of Cultural Heritage), project of the Ministry of Culture of the Andalusia region (MOSAICO, 2016). MOSAICO (Figure VIII-1) is a horizontal and global system that aims: (i) offer the technological resources and tools for the management of historical patrimony, (ii) offer a global information system that will store information about all the cultural patrimony and (iii), bring the public and government closer and more specifically, the information related to patrimony. This system was developed by the IAPH to meet their own objectives such as: managing cultural heritage information, protecting cultural heritage information of Andalusia, preserving the cultural heritage of Andalusia, disseminating the values of cultural assets or bringing government to citizen (Ponce et al., 2010). Figure VIII-1. MOSAICO Architecture (MOSAICO, 2016) Help to the Management Information Datasets Queries and Reports Cataloging Authorizations Execution and Follow-up (Others) Commissions and Advisory Bodies Visits and Inspections Studies, Projects and Inventories Knowledge Area Cultural Heritage Management Core of the System Support System Management Interfaces Geographic Information System Actors Registry Object Documentary and Graphic Management Expedient Standard Terminology
114 There are lots of monuments and several data sources where the information is stored so for the IAPH, keep in control all information published about patrimony in the worldwide suppose a very difficult task. In addition, the size and complexity of these data sources make complicated the management of these systems due to the large amount of information stored on them (for example, just MOISAICO, stores terabytes of information). Then, it is necessary to reconcile the existing information about monuments from all data sources. Furthermore, the process in which the information of historical patrimony mentioned is managed, is carried out in a very rudimentary way. When a campaign is done at any point, either externally or internally, it is performed offline, it means, the team in charge of this process perform the campaign and once the delivery finishes, the administrators make the validation of the patrimony found one by one. For example, if there is a reservoir that was studied in the 80’s and now in 2017, a revision is requested the new process requires the following steps The IAPH exports the data that is already in MOSAICO system and it is given to the team in charge of the revision of the monument. The team load, at their discretion, the new things that have been found. Once this process is finished, the team give the report to the IAPH and they perform the integration in the system. In this context, the variabiliy is very high and also, the integrity of data is questioned. Trying to offer a solution to these problems, the project “DIPHDA” (Dynamic Integration for Patrimonial Heritage Data in Andalucía) is being developed in collaboration with the Fujitsu Laboratories of Europe (FLE). The objective of DHIPDA is to achieve significantly improved accuracy and data management efficiency, based on reconciliation logic applied to open data information, as opposed to simple string matching reconciliation. This solution will can integrate management different systems. For this case the “MOSAICO”, Wikipedia and Yelp systems were used. The information that DIPHDA manages is retrieved from the process of reconciliation done in one of its functionalities where the user must define the data structure where the results of the reconciliation process will be stored. This functionality is covered with the domain-specific language developed as support tool of this doctoral thesis presented in Chapter VII. In this context, DIPHDA project provides a DSL (MaRIA) for designing the concrete entity reconciliation problem to be addressed. The software engineer must design the data structure with all the necessary attributes and operations for carrying out the entity reconciliation process. Figure VIII-2, illustrates a brief example of how the instantiation should be made by the software engineer.
115 Figure VIII-2. Draft Example of Modeling To materialize this example using the MaRIA Tool, it was decided to make a proof of concept designing the entity reconciliation problem between two heterogeneous data sources: the first one with the information stored in mosaic and the second one, with the information stored in Yelp. TRANSFORM ATIONS OUTPUT INPUT M OSAICO id: Integer reference: String province: String location: String name: String M osaico Database W rapper M osaico DBPEDIA Monument URL: String Name: String Abstract: String Category: String Art Style: String Building Type: String Location: String Period: String Religion: String Latitude: Double Longitude: Double DBPedia Database Wrapper DBPedia YELP Name: String Contact Phone: Integer Rating: Number Display Address: String City: String Categories: String Yelp Database W rapper Yelp Virtual Graph Solution M onument - Attribute1: String - Attribute2: String - Attribute3: String ... City - Name: String Province - Name: String belongsTo o: String d: String belongsTo o: String d: String Virtual Graph M osaico Virtual Graph DBpedia Virtual Graph Yelp
116 Figure VIII-3. MaRIA Tool Following the example showed in Figure VIII-2 and considering how is the visual aspect of the MaRIA Tool illustrated in Figure VIII-3, the result of the real model can be observed in Figure VIII-4. For achieving to this model and following the MaRIA Process described in Chapter IV, the first step that that software engineer had to design was the data sources from which the information was going to be extracted. As it is possible to see, there are the two different heterogeneous data sources mentioned before: “Mosaico” and “Yelp”. Next, the software engineer had to define the entities where the information was going to be stored during the process of the entity reconciliation. There are two entities, one for each data source. In addition, the user had to define the attributes that defined each data source. In the case of “Mosaico” these were: “id”, “reference”, “province”, “location”, “name” and “buildingType”. In the case of “Yelp” these were: “name”, “contactPhone”, “rating”, “displayAddress”, “city” and “categories”. Once these elements were defined, the user modeled the “wrappers”, one for each data source, thus making it possible the information transfer from a data source to a defined entity.
117 Figure VIII-4. Real Entity Reconciliation Problem modeled with MaRIA The next step that the user had to perform was the designing of the structures where the solution was going to be stored. In this sense, it was defined three equals graph structures, two that were connected to each data source and another one that was going to be the final structure where the data generated from the reconciliation process would be stored. As it is possible to see in Figure VIII-4, the graphs structure was composed of three vertexes and two relationships between them. One of the vertexes (“CityMosaico”, “CityYelp” and “City”), stores information about the city where the monument is located. Other of the vertexes (“ProvinceMosaico”, “ProvinceYelp” and “Province”), stores information about the province where the monument is located. These two vertexes have a “name” property and are related between them with a “isProvinceOf” relationship, indicating that one city belongs to a province. The final (“MonumentMosaico”, “MonumentYelp” and “Monument”) vertex represents the monument itself. They have four attributes that define the monument: “name”, “description”, “rating” and “contact”.
124 In addition, as result of this Doctoral Thesis, within the framework of the IWT2 group, an end of degree works has been performed in the field of the development of a system that let a software engineer to design and analyze an entity reconciliation problem in heterogeneous data sources using the MDE paradigm. Finally, this Doctoral Thesis has been partially sponsored by Fujitsu Laboratories of Europe (FLE). Also, it has been carried out a project that takes this proposal as a basis for its development in collaboration with this organization. 2. CONTRIBUTIONS This second section recapitulates the main contributions of the present Doctoral Thesis to the scientific community, referring to the objectives that were raised at the beginning of the research, proving that there is at least one correspondence for each objective and concluding that all the objectives have been cutlery. In the development of the research carried out in this Doctoral Thesis, it was fundamental to know the current situation regarding the existing solutions to solve entity reconciliation problems in heterogeneous data sources. To know the current situation, a Systematic Mapping Study (SMS) was carried out to first, trying to (i) understand the state-of-the-art of the problem and (ii), identifying any gaps in current research. This SMS was published in (Enríquez, J.G. et al., 2017) A characterization scheme (Table IX-1) was created to achieve these goals. This characterization was divided in three groups: UML, ER challenges and type of datasets. Following Unified Modeling Language (UML) specification (Group, 2017) that classifies diagrams in two categories: (i) structure-based diagrams, which show the static structure of the system and its parts on different abstraction and implementation levels and how they are related to each other, (ii) and behavior-based diagrams, which show the dynamic behavior of the objects in a system extrapolating it to our problem, it was decided to categorize the proposals found in two big groups: design and operation. Challenges for ER proposed in Getoor and Machanavajjhala, (2013), multirelational, dealing with structure of entities, multi-domain, dealing with customizable methods that span across domains and multi-applications, dealing with systems that serve diverse application with different accuracy requirements, level of automation of the proposal. Finally, the types of dataset that were used for the validation of the proposals, understanding them as heterogeneous or non-heterogeneous.
125 Objective UML Design Operation ER Challenges Automation Multi-Relational Multi-Domain Multi-Applications Type of Dataset Heterogeneous Non-heterogeneous Table IX-1. Characterization scheme The analysis of the results showed that the heterogeneity of the datasets is acceptable knowing that more than a half proposals use heterogeneous data sources to test them. Most of the research work has been focused on the operation phase of the reconciliation and not in the design phase. Finally, the efforts made to automate the process of reconciliation, and consider the multi-relational, multi-domain and multiapplications challenges, have been very limited. This covers the objective set first in Chapter III, which was “Perform a study of the state of the art of the different existing solutions for the entity reconciliation of heterogeneous data sources, checking if they are being used in real environments”. 2.1.MARIA FRAMEWORK The main contribution made with the development of the present Doctoral Thesis is the Framework that allows software engineers to design, analyze and test entity reconciliation problems into any software development methodology. It is worth highlighting the great influence that the MDE paradigm had on the development of this reference framework, which has largely guided the solutions adopted. Figure IX-1, shows a global view of the MaRIA Framework. Thanks to the proposed solution, the software engineers will be able to model their presumable entity reconciliation models for solving any type of entity reconciliation problem, understanding presumable, as the capacity of test the final solutions for checking the coverage level of the new generated dataset without more efforts than the design of the problem, all this, thanks to the integration of Early Testing. MaRIA Framework has been designed so that it can be integrated in any methodology of software development, whether they have classic or agile life cycles. Considering that any software development methodology is composed of several phases, the scope of this Doctoral Thesis has been set up to cover the requirements and analysis phases. Also, the testing phase has been considered adding Early Testing to allow the models to be systematically tested.
126 MaRIA Framework is composed of three fundamental pillars: “MaRIA process”, “Model-Driven Approach” and “MaRIA Tool”. “MaRIA process” defines the set of activities that must be added and performed in any software development methodology to develop a solution to an entity reconciliation problem. “Model-Driven Approach” is defined by the metamodel that allows the software engineer that use this framework to model the entity reconciliation problem and derivation mechanisms that to allow the models defined to be systematically tested. “MaRIA Tool” is the support tool developed to give support to the MaRIA Framework. Figure IX-1. MaRIA Framework. This covers the objective set second in Chapter III, which was “Define and develop a Framework for designing the entity reconciliation models by a systematic way for the requirements and analysis phases”. 2.2. MARIA TOOL One of the main pretensions since the beginning of the development of this Doctoral Thesis was that MaRIA Framework could be put into practice to be able to use it in real scenarios. To achieve this objective, the need to develop a support tool that implements the MaRIA Framework encompassing the 3 fundamental pillars that comprise it was raised. The resulting tool, named MaRIA Tool (Figure IX-2), was developed based on a Domain Specific Language that implements the metamodel defined in the Model-Driven pillar of MaRIA Framework. This covers the objective set third in Chapter III, which was “Provide a support tool for the framework”. This objective is covered in Chapter VII. The support will allow to a software engineer to define the analysis model of an entity reconciliation problem between different and heterogeneous data sources. The tool will be represented as a Domain Specific Language (DSL)”. Support Tool MaRIA Tool Framework MaRIA MaRIA Process Model-Driven Approach
127 Figure IX-2. MaRIA Tool To evaluate MaRIA tool, two real-world use cases have been taken as reference: (i) DIPHDA project, in collaboration with the Government of Andalusia and Fujitsu Laboratories of Europe that aimed to reconcile historical patrimony information related to the monuments of the region, and (ii) ADAGIO, a CDTI project in collaboration with Servinform company that aimed to develop an information system or platform that allows: the aggregation, consolidation and standardization of data from different semantic fields, contextualized by environmental and manning data. This covers the last objective set forth in Chapter III, which was “Evaluate the results obtained of the application of the proposal in a real-world case study”. Finally, two more contributions have been achieved: the real validation and the transfer of the knowledge. As mentioned before, this Doctoral Thesis has been applied in two real world case studies (DIPHDA and ADAGIO) transferring the knowledge generated during the development of this work to the companies. 3. FUTURE WORK AND NEW RESEARCH LINES After recapitulating the main contributions made during the development of this Doctoral Thesis, it is important to emphasize that its results lead to the opening of new research lines that will allow us to advance in the selected topic. Considering the objectives defined in Chapter III for this Doctoral Thesis, future work encompasses several avenues. Table IX-2 describes the relation between the objectives proposed for this Doctoral Thesis and the solution proposed each one and the future work and new research lines that emerges from this relation.
128 Objective Solution Future Work and New Research Lines Perform a study of the state of the art of the different existing solutions for the entity reconciliation of heterogeneous data sources, checking if they are being used in real environments Entity Reconciliation in Big Data Sources: a Systematic Mapping Study (Enríquez, J.G. et al., 2017) New Research Lines Increase level of automation in solution for solving entity reconciliation problems in heterogeneous data sources Multi-applications problems, where different applications with different requirements need to be served with the results of the reconciliation process Future Work Include new searches to increase the domains included in the developed SMS Continue developing this king of studies to keep it updated and do not let it become obsolete Define and develop a Framework for designing the entity reconciliation models by a systematic way for the requirements and analysis phases MaRIA Framework New Research Lines Apply MaRIA Framework in other software development methodology such as SCRUM (Szalvay, 2004) or BDD (Lazǎr et al., 2010) among others. Future Work Extend NDT-Methodology to cover, besides the requirements, analysis and testing processes, all the set of processes that define it Extend Testing set of processes of NDTMethodology to: o Cover the unit testing of the transformations applied over the data to carry out the entity reconciliation o Generate automatically the insert queries that cover test requirements Provide a support tool for the Framework MaRIA Tool New Research Lines Extend the tool to cover the new set of processes that gives support to the extended MaRIA Framework Future Work Encompass more transformation operations Offer a big level of abstraction Evaluate the results obtained of the application of the proposal in a real-world case study DIPHDA and ADAGIO projects Future Work Work closely to the companies to keep improving this approach and solve new problems of different domains Table IX-2. Future work and new research lines
129 4. CONCLUSIONS Previous chapters of this PhD thesis hold up the work where we present and define a theoretical and practical Framework. They also describe a work based on a real need in organizations (Chapter I), that later has turned into a specific problem (Chapter III) derived from the results and conclusions obtained after studying the state-of-the-art (Chapter II). Once the context has been specified, the remaining PhD thesis introduces a Framework composed three main pillars: (i) the MaRIA Process (Chapter IV), (ii) the Model-Driven approach (Chapter V) to support the entity reconciliation modeling preparing the system to be developed to be systematically tested defining a set of transformation mechanisms (Chapter VI) and (iii), a support tool to cover the previous two pillars called MaRIA Tool (Chapter VII). To apply the theoretical framework to real environments. In this sense this approach has been put in practice in two real world case studies presented in Chapter VIII. This Doctoral Thesis propose a suitable environment to support the entity reconciliation in the requirements and analysis phases. This environment allows the development team to prepare their future system to guarantee a suitable entity reconciliation with, besides could be systematically tested.
130 REFERENCES
131 REFERENCES Ali, O., Cristianini, N., 2010. Information fusion for entity matching in unstructured data. IFIP Adv. Inf. Commun. Technol. 339 AICT, 162–169. doi:10.1007/978-3-64216239-8_23 Axelos, 2016. PRINCE2 Certification | AXELOS [WWW Document]. PRINCE2 Qualif. URL https://www.axelos.com/qualifications/prince2-qualifications Ayat, N., Akbarinia, R., Afsarmanesh, H., Valduriez, P., 2014. Entity resolution for probabilistic data. Inf. Sci. (Ny). 277, 492–511. doi:10.1016/j.ins.2014.02.135 Ayat, N., Akbarinia, R., Afsarmanesh, H., Valduriez, P., 2013. Entity resolution for distributed probabilistic data. Distrib. Parallel Databases 31, 509–542. Balaji, J., Javed, F., Kejriwal, M., Min, C., Sander, S., Ozturk, O., 2016. An Ensemble Blocking Scheme for Entity Resolution of Large and Sparse Datasets. Beheshti, S.-M.-R., Benatallah, B., Venugopal, S., Ryu, S.H., Motahari-Nezhad, H.R., Wang, W., 2016. A systematic review and comparative analysis of cross-document coreference resolution methods and tools. Computing 1–37. doi:10.1007/s00607016-0490-0 Beizer, B., 1995. Black-box testing: techniques for functional testing of software and systems, Software Testing Verification and Reliability. Bézivin, J., 2005. On the unification power of models. Softw. Syst. Model. 4, 171–188. doi:10.1007/s10270-005-0079-0 Bézivin, J., Hillairet, G., Jouault, F., Kurtev, I., Piers, W., 2005. Bridging the MS/DSL Tools and the Eclipse Modeling Framework. Meta San Diego,. Bhattacharya, I., Getoor, L., 2005. A latent dirichlet allocation model for entity resolution. Proc. 2005 SIAM Int. Conf. Data Min. 47–58. Blanco, R., Tuya, J., Seco, R. V., 2012a. Test adequacy evaluation for the user-database interaction: A specification-based approach, in: Proceedings - IEEE 5th International Conference on Software Testing, Verification and Validation, ICST 2012. pp. 71–80. doi:10.1109/ICST.2012.87 Blanco, R., Tuya, J., Seco, R. V., 2012b. Test adequacy evaluation for the user-database interaction: A specification-based approach, in: Proceedings - IEEE 5th International Conference on Software Testing, Verification and Validation, ICST 2012. pp. 71–80. doi:10.1109/ICST.2012.87 Bortoli, S., Bouquet, P., Bazzanella, B., 2014. An Identification Ontology for Entity Matching. Move to Meaningful Internet Syst. Otm 2014 Work. 8842, 587–596. Brambilla, M., Cabot, J., Wimmer, M., 2012. Model-Driven Software Engineering in Practice, Synthesis Lectures on Software Engineering.
132 doi:10.2200/S00441ED1V01Y201208SWE001 Brando, C.., Frontini, F.. b, Ganascia, J.-G.., 2015. Disambiguation of named entities in cultural heritage texts using linked data sets. Commun. Comput. Inf. Sci. 539, 505– 514. doi:10.1007/978-3-319-23201-0_51 Bratus, S., Rumshisky, A., Khrabrov, A., Magar, R., Thompson, P., 2011. Domainspecific entity extraction from noisy, unstructured data using ontology-guided search. Int. J. Doc. Anal. Recognit. 14, 201–211. doi:10.1007/s10032-011-0149-5 Brizan, D.G., Tansel, A.U., 2006. A Survey of Entity Resolution and Record Linkage Methodologies. Commun. IIMA 6, 41–50. Calvanese, D., Keet, C.M., Nutt, W., Rodr’\iguez-Muro, M., Stefanoni, G., 2010. Webbased graphical querying of databases through an ontology: the Wonder system. Proc. 2010 ACM Symp. Appl. Comput. 1388–1395. doi:10.1145/1774088.1774384 Cartlidge, A., Hanna, A., Rudd, C., Ivor, M., Stuart, R., 2007. An introductory overview of ITIL V3, The UK Chapter of the itSMF. doi:10.1080/13642818708208530 Carvalho, L.F.M., Laender, A.H.F., Meira Jr, W., 2015. Entity Matching: A Case Study in the Medical Domain, in: Alberto Mendelzon International Workshop on Foundations of Data Management. p. 57. Ceri, S., Fraternali, P., Bongio, A., 2000. Web modeling language (WebML): a modeling language for designing Web sites. Comput. Networks 33, 137–157. doi:10.1016/S1389-1286(00)00040-2 Certification Europe, 2014. ISO 27001 Information Security Certification. ISO 27001 Inf. Secur. Chilenski, J.J., 2001. An investigation of three forms of the modified condition decision coverage (MCDC) criterion. Security. Cohen, W.W., Richman, J., 2002. Learning to match and cluster large high-dimensional data sets for data integration. Proc. eighth ACM SIGKDD Int. Conf. Knowl. Discov. data Min. 475–480. doi:10.1145/775107.775116 Cook, S., Jones, G., Kent, S., Wills, A.C., 2007. Domain Specific Development with Visual Studio DSL Tools, Library. Costa, G., Cuzzocrea, A., Manco, G., Ortale, R., 2011. Data de-duplication: A review. Stud. Comput. Intell. 375, 385–412. doi:10.1007/978-3-642-22913-8_18 Costa, G.D.A., 2016. Large-Scale Entity Resolution for Semantic Web Data Integration LARGE-SCALE ENTITY RESOLUTION FOR SEMANTIC. de Assis Costa, G., de Oliveira, J.M.P., 2016. A Blocking Scheme for Entity Resolution in the Semantic Web, in: Advanced Information Networking and Applications (AINA), 2016 IEEE 30th International Conference on. pp. 1138–1145. Dharavath, R., Kumar, C., 2015. The Journal of Systems and Software Entity resolution
133 based EM for integrating heterogeneous distributed probabilistic data. J. Syst. Softw. 107, 93–109. doi:10.1016/j.jss.2015.05.035 Dharavath, R., Singh, A.K., 2016. Entity Resolution-Based Jaccard Similarity Coefficient for Heterogeneous Distributed Databases, in: Proceedings of the Second International Conference on Computer and Communication Technologies. pp. 497– 507. Domínguez-Mayo, F.J., Escalona, M.J., Mejías, M., Ross, M., Staples, G., 2014. Towards a homogeneous characterization of the model-driven web development methodologies. Dorneles, C.F., Gonçalves, R., dos Santos Mello, R., 2011. Approximate data instance matching: A survey. Knowl. Inf. Syst. 27, 1–21. doi:10.1007/s10115-010-0285-0 Eclipse, 2016a. Eclipse Modeling Framework and Java Emitter Templates [WWW Document]. http//www.eclipse.org/emf, Last Accesed December 201. URL http://www.eclipse.org/emf Eclipse, 2016b. Eclipse Graphical Modeling Framework. Efthymiou, V., Efthymiou, V., Papadakis, G., Papastefanatos, G., Stefanidis, K., Palpanas, T., 2016a. Parallel Meta-blocking for Scaling Entity Resolution over Big Heterogeneous Data. Inf. Syst. 65, 137–157. doi:10.1016/j.is.2016.12.001 Efthymiou, V., Stefanidis, K., Vassilis, C., 2016b. Minoan ER : Progressive Entity Resolution in the Web of Data, in: 19th International Conference on Extending Database TechNology, EDBT 2016. pp. 670–671. doi:10.5441/002/edbt.2016.79 Enríquez, J.G., Domínguez-Mayo, F.J., Escalona, M.J., García-García, J.A., Lee, V., Masatomo, G., 2015. Entity Identity Reconciliation based Big Data Federation-A MDE approach, in: International Conference on Information Systems Development (ISD2015). Escalona, M.J.., Gutiérrez, J.J.., Ortega, J.A.., Ramos, I.., Aragón, G.., 2008. {NDT} & Metrica V3 an approach for public organizations based on model driven engineering, in: {WEBIST} 2008 - 4th International Conference on Web Information Systems and Technologies, Proceedings. pp. 224–227. Escalona, M.J., Aragón, G., 2008. NDT. A model-driven approach for web requirements. IEEE Trans. Softw. Eng. 34, 377–394. doi:10.1109/TSE.2008.27 Fan, W., Geerts, F., Tang, N., Yu, W., 2014. Conflict Resolution with Data Currency and Consistency. J. Data Inf. Qual. 5, 6:1--6:37. doi:10.1145/2631923 Fan, W., Li, J., Ma, S., Tang, N., Yu, W., 2011. Interaction between record matching and data repairing, Proceedings of the 2011 international conference on Management of data - SIGMOD ’11. doi:10.1145/1989323.1989373 Fellegi, I.P., Sunter, A.B., 1969. A Theory for Record Linkage. Source J. Am. Stat. Assoc. 64, 1183–1210. doi:10.1080/01621459.1969.10501049