Heterogeneous datasets: a tale of integration and exploration
Abstract
In this presentation, I detail my PhD and post-doctoral work in the data integration and exploration fields.
Full text
Heterogeneous datasets A tale of integration and exploration Nelly Barret Postdoctoral researcher Data Science group Dipartimento di Elettronica, Informazione e Bioingegneria Politecnico di Milano January 24, 2025 Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 1 / 88
Short bio My CS background: Bachelor @ Univ. Lyon Master, AI track @ Univ. Lyon PhD @ Inria Saclay and Ecole Polytechnique Post-doc @ Politecnico di Milano (Italie) My thesis was about user-oriented exploration of semi-structured data. My post-doc is about enabling federating analyses of health data. Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 2 / 88
Outline 1Motivation: data integration and exploration problems 2PhD: exploring unknown semi-structured datasets 3Post-doc: healthcare analytics across hospitals 4Systems developed 5Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 3 / 88
Motivation: data integration and exploration problems Outline 1Motivation: data integration and exploration problems 2PhD: exploring unknown semi-structured datasets 3Post-doc: healthcare analytics across hospitals 4Systems developed 5Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 4 / 88
Motivation: data integration and exploration problems Data integration and exploration Different settings, different needs Structured data models: Tables Relational databases Semi-structured data models: XML documents JSON documents RDF graphs Property graphs Unstructured data models: Text Images Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 5 / 88
Motivation: data integration and exploration problems Data integration and exploration Different settings, different needs Various domains: Health Journalism Transports, ... Sensitivity levels: Enforce privacy rules EU GDPR rules Several actors/users: Different skills Time/money constraints Dataset integration and exploration is hard: large, complex, irregular Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 6 / 88
Motivation: data integration and exploration problems Data integration and exploration Different settings, different needs Various domains: Health Journalism Transports, ... Sensitivity levels: Enforce privacy rules EU GDPR rules Several actors/users: Different skills Time/money constraints Dataset integration and exploration is hard: large, complex, irregular Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 6 / 88
PhD: exploring unknown semi-structured datasets Outline 1Motivation: data integration and exploration problems 2PhD: exploring unknown semi-structured datasets 3Post-doc: healthcare analytics across hospitals 4Systems developed 5Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 7 / 88
PhD: exploring unknown semi-structured datasets What does the dataset describe? Real-world objects and relationships between them Traditional setting: Entity-Relationship models [RG03] Need to compute them from the dataset! Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 8 / 88
PhD: exploring unknown semi-structured datasets Thesis problem and research contributions Thesis problem statement How to facilitate user exploration of unknown heterogeneous semi-structured datasets? Abstra: semi-structured data overviews [BMU22,BMU24] Automatically compute lightweight Entity-Relationship diagrams Ideal for first-sight dataset discovery PathWays: interesting Named Entity connections [BGLM23b,BGLM23a,BGLM25] Compute and rank entity paths in and across datasets Ideal for exploring connections within and across datasets Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 10 / 88
PhD: exploring unknown semi-structured datasets Related work Related work Data summarization Structural Quotient [GGM20,KC10,MS99] (the one we adopt to build G) Non-quotient [GW97] Pattern mining [ZLVK16] Statistical [HS12] Hybrid [RGSB17] Schema inference XML [CGS11] JSON [BCGS19] RDF [GLSW22] PG [LBH21] Data summarization and schema inference are tied to one data model Schemas are often not suited to NTUs Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 11 / 88
PhD: exploring unknown semi-structured datasets Related work A JSON schema from social network data using [BCGS19] Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 12 / 88
PhD: exploring unknown semi-structured datasets Related work What does the dataset describe? Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 13 / 88
PhD: exploring unknown semi-structured datasets Related work The Abstra approach 1Integrate all data sources in a graph (ConnectionLens) [ABC+22] 2Summarize the graph 3Among summary nodes, identify entities and their attributes 4In the summary, identify relationships between the entities 5Propose a simple category to each entity (best-effort) Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 14 / 88
PhD: exploring unknown semi-structured datasets Background Background: from heterogeneous data to data graphs ConnectionLens [ABC+22]: 1Ingests any dataset into a directed graph Generic, flexible, fine granularity 2Extracts Named Entities (NEs) from all text nodes date , email address , People , Place , Organization , ... Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 15 / 88
PhD: exploring unknown semi-structured datasets Background Background: from heterogeneous data to data graphs ConnectionLens [ABC+22]: 1Ingests any dataset into a directed graph Generic, flexible, fine granularity 2Extracts Named Entities (NEs) from all text nodes date , email address , People , Place , Organization , ... Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 15 / 88
PhD: exploring unknown semi-structured datasets Data graph summarization Data graph summarization We need a compact representation of large data graphs Challenges: Heterogeneous graphs originate from different data models Node and/or edge labels may be empty We aim for a quotient graph summary: Based on equivalence between nodes of the original graph We prefer small summaries (number of nodes) Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 16 / 88
PhD: exploring unknown semi-structured datasets Data graph summarization Data graph summarization We need a compact representation of large data graphs Challenges: Heterogeneous graphs originate from different data models Node and/or edge labels may be empty We aim for a quotient graph summary: Based on equivalence between nodes of the original graph We prefer small summaries (number of nodes) Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 16 / 88
PhD: exploring unknown semi-structured datasets Data graph summarization Data graph summarization We need a compact representation of large data graphs Challenges: Heterogeneous graphs originate from different data models Node and/or edge labels may be empty We aim for a quotient graph summary: Based on equivalence between nodes of the original graph We prefer small summaries (number of nodes) Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 16 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Identifying entities in the collection graph G paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Which collections represent entities in the E-R diagram? Which collections represent entity attributes? Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 22 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Identifying entities in the collection graph G paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Which collections represent entities in the E-R diagram? Which collections represent entity attributes? Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 22 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Requirements and algorithm We need an algorithm to identify entity roots and attributes for the E-R diagram For complex, potentially cyclic, collection graphs Greedy selection of few entities in G 1Assign a score to each collection node 2While less than Emax entity roots, or data coverage <covmin 1Elect the next highest-scored eligible collection node as an entity root 2Compute its boundary (set of attributes) 3Update the collection graph to reflect the selection of an entity 4Recompute the scores Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 23 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Requirements and algorithm We need an algorithm to identify entity roots and attributes for the E-R diagram For complex, potentially cyclic, collection graphs Greedy selection of few entities in G 1Assign a score to each collection node 2While less than Emax entity roots, or data coverage <covmin 1Elect the next highest-scored eligible collection node as an entity root 2Compute its boundary (set of attributes) 3Update the collection graph to reflect the selection of an entity 4Recompute the scores Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 23 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships How to score a collection node? Reflect the weight of this node and its structure in the dataset 1wdesck,wleafk: # descendants, leaf descendants, at depth k ×Not clear how to pick k Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 24 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships How to score a collection node? Reflect the weight of this node and its structure in the dataset 1wdesck,wleafk: # descendants, leaf descendants, at depth k ×Not clear how to pick k Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 24 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships How to score a collection node? Reflect the weight of this node and its structure in the dataset 1wdesck,wleafk: # descendants, leaf descendants, at depth k 2Directed Acyclic Graph (DAG) rooted in each node: wDAG Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 25 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Data weight Own weight ow of a leaf node: its in-degree Data weight dw of a leaf collection node: the sum of its nodes’ ow paper author “database” ow = 2 “L´ea” ow = 1 “Data lake” ow = 1 conf name hasWritten writtenBy keyword keyword topic author hasWritten paper writtenBy keyword name #val dw = 1 #val dw = 3 topic conf Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 26 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Data weight DAG propagation Leaf collection dw is propagated back to all ancestors which are not in a cycle Edge transfer factor: |nodes in Cthaving a parent in Cs| |Ct| author hasWritten paper writtenBy keyword name #val dw = 1 #val dw = 3 topic conf 1 1 1 1 1 1 1 1 0.5 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 27 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Data weight DAG propagation Leaf collection dw is propagated back to all ancestors which are not in a cycle Edge transfer factor: |nodes in Cthaving a parent in Cs| |Ct| author dw = 1 hasWritten dw = 0 paper dw = 3 writtenBy dw = 0 keyword dw = 3 name dw = 1 #val dw = 1 #val dw = 3 topic dw = 1.5 conf dw = 1.5 1 1 1 1 1 1 1 1 0.5 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 28 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships The data-weighted PageRank score paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 The reverse collection graph GRwith dw-tuned PR edge weights Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 34 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships The data-weighted PageRank score paper .178 abstract .011 #val .006 year .011 #val .006 title .011 #val .006 wB .107 hW .158 pIn .063 inv .056 author .179 conf .067 name .011 #val .006 date .011 #val .006 email .011 #val .006 affiliation .027 university .024 city .011 #val .006 campus .011 #val .006 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 35 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships The data-weighted PageRank score paper .178 abstract .011 #val .006 year .011 #val .006 title .011 #val .006 wB .107 hW .158 pIn .063 inv .056 author .179 conf .067 name .011 #val .006 date .011 #val .006 email .011 #val .006 affiliation .027 university .024 city .011 #val .006 campus .011 #val .006 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 36 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships The data-weighted PageRank score paper .178 abstract .011 #val .006 year .011 #val .006 title .011 #val .006 wB .107 hW .158 pIn .063 inv .056 author .179 conf .067 name .011 #val .006 date .011 #val .006 email .011 #val .006 affiliation .027 university .024 city .011 #val .006 campus .011 #val .006 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Propagates scores across the collection graph Works on cyclic collection graphs The score reflects the topology and where the data is A collection node distributes its weight Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 37 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships How to compute an entity boundary? Collections in Grepresenting attributes of this entity “Those that contribute to the entity’s weight” The boundary may go far (for deep-structure entities) Easy to define for wdesck,wleafk,wDAG . Example for wdesc2 paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Does not apply for PageRank-based scores Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 38 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships How to compute an entity boundary? Collections in Grepresenting attributes of this entity “Those that contribute to the entity’s weight” The boundary may go far (for deep-structure entities) Easy to define for wdesck,wleafk,wDAG . Example for wdesc2 paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Does not apply for PageRank-based scores Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 38 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships How to compute an entity boundary? Collections in Grepresenting attributes of this entity “Those that contribute to the entity’s weight” The boundary may go far (for deep-structure entities) Easy to define for wdesck,wleafk,wDAG . Example for wdesc2 paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Does not apply for PageRank-based scores Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 38 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Data-acyclic flooding boundary bounddfl−ac Idea: the collection nodes Reachable from the entity root Mainly part of this entity Edge transfer factor ≥fmin At-most-one: each Csnode has at most one child in Ct The path between the entity root and this collection node is not data cyclic If the path in Ghas no in-cycle edges Or, the Gpath has in-cycle edges, but they are not in the data Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 39 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Data-acyclic flooding boundary bounddfl−ac Idea: the collection nodes Reachable from the entity root Mainly part of this entity Edge transfer factor ≥fmin At-most-one: each Csnode has at most one child in Ct The path between the entity root and this collection node is not data cyclic If the path in Ghas no in-cycle edges Or, the Gpath has in-cycle edges, but they are not in the data Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 39 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Data-acyclic flooding boundary bounddfl−ac Idea: the collection nodes Reachable from the entity root Mainly part of this entity Edge transfer factor ≥fmin At-most-one: each Csnode has at most one child in Ct The path between the entity root and this collection node is not data cyclic If the path in Ghas no in-cycle edges Or, the Gpath has in-cycle edges, but they are not in the data Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 39 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Selected entities and their boundaries paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 44 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Finding relationships between entities Relationship: a path from an entity to another paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 paper →wB →author paper →pIn →conf author →hW →paper conf →inv →author Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 45 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Entity classification Assign a semantic category to each entity Input: an entity E, categories K, semantic properties P K: Person, ScientificPaper, Event, Website, Mountain, ... P:{label:"address", domain:[Pers., Org.], range:[Place]}, ... Output: a category for E Algorithm: Compare: The common name of all nodes in the entity root (if it exists) with k∈ K (conf, paper, author) Its attribute names with p∈ P (affiliation, email, ...) Its entity profiles with p.range ∈ P (,,, ...) Each good match votes for one or few categories Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 46 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Entity classification Assign a semantic category to each entity Input: an entity E, categories K, semantic properties P K: Person, ScientificPaper, Event, Website, Mountain, ... P:{label:"address", domain:[Pers., Org.], range:[Place]}, ... Output: a category for E Algorithm: Compare: The common name of all nodes in the entity root (if it exists) with k∈ K (conf, paper, author) Its attribute names with p∈ P (affiliation, email, ...) Its entity profiles with p.range ∈ P (,,, ...) Each good match votes for one or few categories Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 46 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Entity classification Name Similar to Votes for paper ResearchPublication (0.85) ResearchPublication News (0.63) News paper abstract #val year #val title #val 1 1 1 1 1 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 47 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Entity classification Attribute Similar to Votes for abstract abstract (1.0) ResearchPublication summary (0.92) Book preface (0.47) title title (1.0) ResearchPublication honorific title (0.87) Movie Person year year publication (0.85 + )Event Book ResearchPublication, ... paper abstract #val year #val title #val 1 1 1 1 1 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 48 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Entity classification Attribute Similar to Votes for abstract abstract (1.0) ResearchPublication summary (0.92) Book preface (0.47) title title (1.0) ResearchPublication honorific title (0.87) Movie Person year year publication (0.85 + )Event Book ResearchPublication, ... paper abstract #val year #val title #val 1 1 1 1 1 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 49 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Entity classification paper nodes classified as ResearchPublication author nodes classified as Researcher conference nodes classified as Event paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 50 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Abstra output: a lightweight Entity-Relationship diagram 318 person (Person) • person@id (100 %) • phone (49 %) • creditcard (49 %) • homepage (47 %) • address (46 %) • province (52 %) • zipcode (100 %) • country (100 %) • city (100 %) • street (100 %) • emailaddress (100 %) • name (100 %) 150 open_auction (Product) • privacy (56 %) • interval (100 %) • end (100 %) • start (100 %) • type (100 %) • current (100 %) • reserve (51 %) • initial (100 %) • open_auction@id (100 %) • quantity (100 %) watches.watch@open_auction 12 category (Thing) • category@id (100 %) • description (100 %) • text (73 %) • parlist (27 %) • listitem (291 %) • text (87 %) • name (100 %) profile.interest@category seller@person annotation.author@person bidder.personref@person 270 item (schema:how_to_item) • mailbox (64 %) • mail (101 %) • date (100 %) • to (100 %) • from (100 %) • text (100 %) • item@featured (9 %) • item@id (100 %) • shipping (94 %) • description (100 %) • text (73 %) • parlist (27 %) • listitem (291 %) • text (87 %) • payment (94 %) • name (100 %) • quantity (100 %) • location (100 %) itemref@item incategory@category 120 closed_auction (Product) • price (100 %) • type (100 %) • date (100 %) • quantity (100 %) seller@person buyer@person annotation.author@person itemref@item Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 51 / 88
PhD: exploring unknown semi-structured datasets Identifying entities and relationships Abstra output: a lightweight Entity-Relationship diagram Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 52 / 88
PhD: exploring unknown semi-structured datasets Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 57 / 88
PhD: exploring unknown semi-structured datasets Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Abstra selects frequent, coherent and semantically central entities Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 58 / 88
PhD: exploring unknown semi-structured datasets Experimental evaluation Experimental evaluation: scalability Our abstraction method scales up linearly in the data size Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 59 / 88
PhD: exploring unknown semi-structured datasets Extending abstractions to Property Graphs Extending abstractions to Property Graphs Property Graphs (PGs) are graphs whose nodes and edges may carry named attributes Model under standardization [ABD+23,ABD+21] Numerous industrial PG databases (Neo4J, Oracle) Widely used (the Offshore leaks database) For interoperability, we derive a PG schema from any (semi)structured dataset following PG-Schema [ABD+23] Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 60 / 88
PhD: exploring unknown semi-structured datasets Extending abstractions to Property Graphs Deriving a PG schema from an abstraction We need to accommodate to nested attributes: FLAT: wrap the nested attribute in a JSON object CUT: unfold each nested attribute in a PG node 1For each Abstra entity E: 1Create a PG node type for E 2For each attribute a: 1If ais not nested: add ato the PG node 2If ais nested and nesting is FLAT: wrap a 3If ais nested and nesting is CUT: unfold a 2For each Abstra relationship R: 1Create a PG edge type with corresponding PG nodes 3If all Gnodes and edges are in E-R: PG graph type is STRICT, else LOOSE Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 61 / 88
PhD: exploring unknown semi-structured datasets Extending abstractions to Property Graphs Extending abstractions to Property Graphs CREATE GRAPH TYPE myGraphType STRICT { (paperType: Paper { title string, OPTIONAL year integer, OPTIONAL abstract string, ... }) (authorType: Author {name string, email string, ...}), (confType: Conference {name string, year integer, ...}), (:authorType)-[edgeAuthorPaper: HasWritten]->(:paperType), (:paperType)-[edgePaperAuthor: WrittenBy]->(:authorType), (:paperType)-[edgePaperConf: PublishedIn]->(:confType), } Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 62 / 88
Post-doc: healthcare analytics across hospitals Outline 1Motivation: data integration and exploration problems 2PhD: exploring unknown semi-structured datasets 3Post-doc: healthcare analytics across hospitals 4Systems developed 5Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 63 / 88
Post-doc: healthcare analytics across hospitals What does multi-source healthcare data has to reveal? Very low cooperation/normalization between medical centers Few patient data for rare diseases Traditional setting: warehouses [DM88] Need to provide decentralized and federated analyses! Leverage experts’ knowledge + make it as automatic as possible Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 64 / 88
Post-doc: healthcare analytics across hospitals What does multi-source healthcare data has to reveal? Very low cooperation/normalization between medical centers Few patient data for rare diseases Traditional setting: warehouses [DM88] Need to provide decentralized and federated analyses! Leverage experts’ knowledge + make it as automatic as possible Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 64 / 88
Post-doc: healthcare analytics across hospitals What does multi-source healthcare data has to reveal? Very low cooperation/normalization between medical centers Few patient data for rare diseases Traditional setting: warehouses [DM88] Need to provide decentralized and federated analyses! Leverage experts’ knowledge + make it as automatic as possible Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 64 / 88
Post-doc: healthcare analytics across hospitals Related work Related work Data plateforms: EHDEN [PdGdK+23] for tabular data Also: OHDSI [HDS+15], UMG-MeDIC [PSS+23], etc Conceptual models: OMOP [SRR+10] for observational data, also FHIR [fhi] ETL pipelines: D-ETL [OKK+17], also EHDEN’s ETL, OHDSI’s ETL Data platforms are often tied to a single data type (tables, etc.) Conceptual models often design one kind of data (observational, etc.) ETLs provide limited interoperability and require time from experts Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 66 / 88
Post-doc: healthcare analytics across hospitals The I-ETL approach The I-ETL approach 1Analyze datasets and extract their metadata 2Create an interoperable database in each medical center 3Assess interoperability along the pipeline 4Allow federated analyses of data across centers Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 67 / 88
Post-doc: healthcare analytics across hospitals Interoperability Interoperability as in FAIR principles FAIR principles are guidelines for good data management [WDA+16]: FFindable: search for (indexed) resources based on identifiers AAccessible: access data with standard protocols, even after data dies IInteroperable: integrate and refer to datasets following FAIR principles RReusable: reuse datasets in other settings using provenance, etc. Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 68 / 88
Post-doc: healthcare analytics across hospitals Step 1: metadata creation From datasets analysis to metadata Metadata: each dataset can be described by a set of features Name, definition, type, unit, values, ... Specified by medical experts Tabular data of clinical measurements and phenotypic information Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 69 / 88
Post-doc: healthcare analytics across hospitals Step 1: metadata creation From datasets analysis to metadata Metadata: each dataset can be described by a set of features Name, definition, type, unit, values, ... Specified by medical experts The metadata obtained from the tabular data What if “ethnicity” is referred to as “race” in another dataset? What if datasets refer to “Homme”/“Femme” vs. “Male”/“Female”? Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 70 / 88
Post-doc: healthcare analytics across hospitals Step 1: metadata creation From datasets analysis to metadata Metadata: each dataset can be described by a set of features Name, definition, type, unit, values, ... Specified by medical experts The metadata obtained from the tabular data What if “ethnicity” is referred to as “race” in another dataset? What if datasets refer to “Homme”/“Femme” vs. “Male”/“Female”? Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 70 / 88
Post-doc: healthcare analytics across hospitals Step 1: metadata creation The metadata model We aim for a conceptual model for expressive and interoperable metadata Name: the feature name Vocabulary: a vocabulary name Code: the code of the term in the selected vocabulary Kind: phenotypic, clinical, genomic, ... DataType:string,integer,numeric,boolean,category, ... Unit: to interpret values when the data type is numeric; Categories: list of discrete values for categorical features Visibility:public,anonymized,private Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 71 / 88
Post-doc: healthcare analytics across hospitals Step 2: mappings to vocabularies Associate metadata to vocabularies Vocabularies are dictionaries of concepts/values uniquely identified SNOMED CT [SPSW01], LOINC [HRM+98], OMIM [HSA+05], ... We associate each feature and categorical value to an existing vocabulary code →more interoperability Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 72 / 88
Post-doc: healthcare analytics across hospitals Step 3: generic data model The healthcare data model We need a general, extensible healthcare data model Challenges: Use-cases bring very different kinds of data Experts’ metadata needs to be represented We aim for a conceptual model: Based on the notions of features and records Will be populated automatically by an ETL Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 73 / 88
Post-doc: healthcare analytics across hospitals Step 3: generic data model General, extensible healthcare conceptual data model How to automatically populate this data model with hospitals data? Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 74 / 88
Post-doc: healthcare analytics across hospitals Step 5: interoperability assessment Anchor our metrics in FAIR principles 1(Meta)data use a [...] broadly applicable language for knowledge representation Our data model can be implemented within any type of database Metadata can be easily specified using a tabular file 2(Meta)data use FAIR-first vocabularies Associate metadata variables and categories to vocabulary resources Use of widely used vocabularies in healthcare domain 3(Meta)data include qualified references to other data and metadata To db instances: references to a patient, a hospital, and a Feature To the data: from which dataset the value comes Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 80 / 88
Post-doc: healthcare analytics across hospitals I-ETL at work I-ETL at work in the Better project 7 clinical centers across Europe I-ETL is under deployment at each center →7 interoperable databases Working on: Designing a catalogue to: List available datasets and their associated metadata Explore datasets and their aggregated data With visualizations and queries Designing a decentralized federated learning plateform To run federated AI algorithms Secured because no data leaves centers, only aggregates Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 81 / 88
Post-doc: healthcare analytics across hospitals I-ETL at work I-ETL at work in the Better project 7 clinical centers across Europe I-ETL is under deployment at each center →7 interoperable databases Working on: Designing a catalogue to: List available datasets and their associated metadata Explore datasets and their aggregated data With visualizations and queries Designing a decentralized federated learning plateform To run federated AI algorithms Secured because no data leaves centers, only aggregates Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 81 / 88
Post-doc: healthcare analytics across hospitals Better plateform Better plateform: decentralized federated learning Based on the Personal Health Train: stations (centers), trains (queries), central station (results aggregation) →no data leaves centers = privacy preservation Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 82 / 88
Systems developed Outline 1Motivation: data integration and exploration problems 2PhD: exploring unknown semi-structured datasets 3Post-doc: healthcare analytics across hospitals 4Systems developed 5Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 83 / 88
Systems developed Systems developed (1/2) Abstra for data abstraction: 65 Java core classes, 10K LOC Published in EDBT 2024 [BMU24] Demonstrated at CIKM and BDA 2022 [BMU22] PathWays for NE-to-NE paths: 18 Java core classes, 4K LOC Published in ADBIS 2023 [BGLM23a], Info. Sys [BGLM25] Demonstrated at ESWC and BDA 2023 [BGLM23b] ConnectionStudio for NTU data exploration: Web interface by CEDAR engineers Published in CoopIS 2023 [BEG+23] Demonstrated at BDA and SEAGRAPH 2024 [BEMM24,BBE+24] Also to journalists at DataJournos (40) and CFI (60) Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 84 / 88
Systems developed Systems developed (1/2) Abstra for data abstraction: 65 Java core classes, 10K LOC Published in EDBT 2024 [BMU24] Demonstrated at CIKM and BDA 2022 [BMU22] PathWays for NE-to-NE paths: 18 Java core classes, 4K LOC Published in ADBIS 2023 [BGLM23a], Info. Sys [BGLM25] Demonstrated at ESWC and BDA 2023 [BGLM23b] ConnectionStudio for NTU data exploration: Web interface by CEDAR engineers Published in CoopIS 2023 [BEG+23] Demonstrated at BDA and SEAGRAPH 2024 [BEMM24,BBE+24] Also to journalists at DataJournos (40) and CFI (60) Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 84 / 88
Systems developed Systems developed (2/2) I-ETL for interoperable healthcare databases: 31 Python core classes, 8K LOC (restricted access) Reviewed at BMC Med. Info. & Decision Making [BBBP25] Under deployment in the 7 medical centers of the project Data catalogue and decentralized platform: Under development by an IT company With collaboration of Better technical partners Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 85 / 88
Conclusion Outline 1Motivation: data integration and exploration problems 2PhD: exploring unknown semi-structured datasets 3Post-doc: healthcare analytics across hospitals 4Systems developed 5Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 86 / 88
Conclusion Takeaways and next steps (1/2) In my PhD, we introduced: 1A unified view over heterogeneous semi-structured data models 2Abstra: a dataset abstraction system for semi-structured data 3PathWays: an entity-focused exploration system 4ConnectionStudio: a comprehensive data lake exploration tool Next steps: Migrate data graphs into PG graphs reusing [BEMM24] Enrich extracted NEs with RDF knowledge bases Propose an end-to-end data processing/exploration pipeline Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 87 / 88
References IV Benoˆıt Groz, Aur´elien Lemay, Slawek Staworko, and Piotr Wieczorek. Inference of shape graphs for graph databases. In ICDT, volume 220, 2022. Roy Goldman and Jennifer Widom. DataGuides: enabling query formulation and optimization in semistructured databases. In VLDB, 1997. George Hripcsak, Jon D Duke, Nigam H Shah, Christian G Reich, Vojtech Huser, Martijn J Schuemie, Marc A Suchard, Rae Woong Park, Ian Chi Kei Wong, Peter R Rijnbeek, et al. Observational health data sciences and informatics (OHDSI): opportunities for observational researchers. In MEDINFO 2015: eHealth-enabled Health, pages 574–578. IOS Press, 2015. Stanley M Huff, Roberto A Rocha, Clement J McDonald, Georges JE De Moor, Tom Fiers, W Dean Bidgood Jr, Arden W Forrey, William G Francis, Wayne R Tracy, Dennis Leavelle, et al. Development of the logical observation identifier names and codes (LOINC) vocabulary. Journal of the American Medical Informatics Association, 5(3):276–292, 1998. Katja Hose and Ralf Schenkel. Towards benefit-based RDF source selection for SPARQL queries. In Proceedings of the 4th International Workshop on Semantic Web Information Management, pages 1–8, 2012. Ada Hamosh, Alan F Scott, Joanna S Amberger, Carol A Bocchini, and Victor A McKusick. Online mendelian inheritance in man (omim), a knowledgebase of human genes and genetic disorders. Nucleic acids research, 33(suppl 1):D514–D517, 2005. Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 4 / 27
References V Shahan Khatchadourian and Mariano P Consens. ExpLOD: summary-based exploration of interlinking and RDF usage in the Linked Open Data Cloud. In Extended semantic web conference, pages 272–287. Springer, 2010. Hanˆa Lbath, Angela Bonifati, and Russ Harmer. Schema inference for property graphs. In EDBT, 2021. Tova Milo and Dan Suciu. Index structures for path expressions. In International Conference on Database Theory, pages 277–295. Springer, 1999. Toan C Ong, Michael G Kahn, Bethany M Kwan, Traci Yamashita, Elias Brandt, Patrick Hosokawa, Chris Uhrich, and Lisa M Schilling. Dynamic-ETL: a hybrid approach for health data extraction, transformation and loading. BMC medical informatics and decision making, 17:1–12, 2017. Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. The PageRank citation ranking: Bringing order to the web. Technical report, Stanford InfoLab, 1999. Daniel Puttmann, Rowdy de Groot, Nicolette de Keizer, Ronald Cornet, et al. Assessing the FAIRness of databases on the EHDEN portal: A case study on two Dutch ICU databases. International Journal of Medical Informatics, 176:105104, 2023. Felipe Pezoa, Juan L Reutter, Fernando Suarez, Mart´ın Ugarte, and Domagoj Vrgoˇc. Foundations of JSON schema. In Proceedings of the 25th international conference on World Wide Web, pages 263–273, 2016. Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 5 / 27
References VI Marcel Parciak, Markus Suhr, Christian Schmidt, Caroline B¨onisch, Benjamin L¨ohnhardt, Dorothea Keszty¨us, and Tibor Keszty¨us. Fairness through automation: development of an automated medical data integration infrastructure for fair health data in a maximum care university hospital. BMC Medical Informatics and Decision Making, 23(1):94, 2023. Raghu Ramakhrishnan and Johannes Gehrke. Database Management Systems (3rd edition). McGraw-Hill, 2003. Matteo Riondato, David Garc´ıa-Soriano, and Francesco Bonchi. Graph summarization with quality guarantees. Data mining and knowledge discovery, 31:314–349, 2017. Michael Q Stearns, Colin Price, Kent A Spackman, and Amy Y Wang. SNOMED clinical terms: overview of the development process and project status. In Proceedings of the AMIA Symposium, page 662. American Medical Informatics Association, 2001. Paul E Stang, Patrick B Ryan, Judith A Racoosin, J Marc Overhage, Abraham G Hartzema, Christian Reich, Emily Welebob, Thomas Scarnecchia, and Janet Woodcock. Advancing the science for active surveillance: rationale and design for the observational medical outcomes partnership. Annals of internal medicine, 153(9):600–606, 2010. Resource Description Framework (RDF). https://www.w3.org/RDF/. Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 6 / 27
References VII The XML data model. https://www.w3.org/XML/Datamodel.html. W3C XML Document Type Specification. https://www.w3.org/TR/REC-xml/#dt-doctype, 2008. W3C XML Schema Definition Language (XSD). https://www.w3.org/TR/xmlschema11-1/, 2012. Mark D Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Appleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan-Willem Boiten, Luiz Bonino da Silva Santos, Philip E Bourne, et al. The FAIR guiding principles for scientific data management and stewardship. Scientific data, 3(1):1–9, 2016. Mussab Zneika, Claudio Lucchese, Dan Vodislav, and Dimitris Kotzinos. Summarizing linked data RDF graphs using approximate graph pattern mining. In 19th International Conference on Extending Database Technology, 2016. Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 7 / 27
The relational data model According to [RG03]: Arelational schema is a set of relations Each relation has a name and set of named attributes with their domain Aprimary key is a subset of attributes to uniquely identify a tuple Aforeign key is a reference to a primary key Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 8 / 27
The XML data model According to the W3C [W3Cb], a tree of: A (single) document node Element nodes with non-labels, possibly with named attributes Text nodes, carrying values, are children of element nodes Possibility do define a DTD [W3C08] or an XSD [W3C12] Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 9 / 27
The JSON data model According to [PRS+16], a tree where a node can be: Amap (label, one or more a key-value elements) An array (label, zero or more child nodes) Avalue (a string) Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 10 / 27
The RDF data model According to the W3C [W3Ca], an RDG graph contains triples <s,p,o> where: sand pare resource identifiers (URIs) ocan be a resource identifier or a literal (string) Also: blank nodes for anonymous resources (internal ID) Add semantic information with ontologies (incl. RDFS, OWL) Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 11 / 27
The property graph data model Anode is a structured record with: 0..n labels (types) 0..n properties (key-values) Records with the same type set may have different properties Arelationship is a directed labeled edge; possibly have attributes Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 12 / 27
Data summarization techniques Build structured and concise summaries out of datasets; many approaches for semi-structured Structural approaches: groups of equivalent nodes; different notions of node similarity Quotient summaries: groups based on an equivalence relation Non-quotient summaries: other means (dataguides, etc) Pattern mining approaches: discovery of patterns Statistical approaches: counts over data (classes, properties, value types, etc) Hybrid approaches: combine above methods Schema inference techniques: build a schema s.t. the data conforms to it Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 13 / 27
A comprehensive data exploration tool for NTUs Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 20 / 27
A comprehensive data exploration tool for NTUs Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 21 / 27
The EHDEN platform [PdGdK+23,BVD+21] “European Health Data and Evidence Network” Consortium of 15 partners across 10 countries Mapped their data to the OMOP data model Semi-automatic mapping The system proposes mappings Experts have to select/correct them Produced 98 databases Asses quality through a DQ (Data Quality Dashboard) Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 22 / 27
The OHDSI platform [HDS+15] “Observational Health Data Sciences and Informatics” (said Odyssey, child from OMOP) International collaboration for open-source data analytics on healthcare networks Build tools for data exploration and evidence generation Achilles: interactive reports and statistics Hermes: vocabulary browsing and related searches Heracles: build cohorts to assess clinical features on poopulations Homer: risk identification by exploring many clinical dimensions Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 23 / 27
UMG-MeDIC [PSS+23] Medical data integration center; relies on Medical Informatics Initiative (MI-I) funds and HiGHmed consortium Create a technical and legal framework for cross-site secondary use of routine healthcare data Aim for high compliance with FAIR Principles but data integration workflows are complex and inefficient when done manually Operates on a continuous flow of data (6= individual datasets) Periodic integration of new data A central relational database with anonymized data Combine individual pre-processing tasks into workflows Require that each task is documented with “meta-data” Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 24 / 27
The OMOP data model [SRR+10] Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 25 / 27
D-ETL [OKK+17] “Dynamic-ETL”: semi-automatic ETL to map source and target data models Creation of an ETL specification document (vocabularies, data schema, definitions, conventions) Data extraction from initial sources and validation D-ETL rules writing (T1./ T2on T1.a=T2.b) Conversion of rules to SQL statements Testing rules on data; iterate if not satisfying Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 26 / 27
Creating new codes with post-coordination Some healthcare concepts do not have a specific code SNOMED-CT introduces post-coordination as a compositional grammar A post-coordinated code = a sequence of existing codes with operators Nelly Barret (DEIB@PoliMi) Data integration and exploration January 24, 2025 27 / 27