Integrating and exploring heterogeneous datasets
Abstract
In this presentation, I summarize the work that I have done in the fields of heterogeneous data integration and exploration. It goes overy my internships and my PhD works.
Full text
Integrating and exploring heterogeneous datasets Nelly Barret Postdoctoral researcher Data Science group, DEIB, Politecnico di Milano April 19, 2024 Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 1 / 89
Outline 1Motivation: data integration and exploration problems 2Predihood: predicting neighbourhoods’ environment 3GeoAlign: spatial entity matching for Points of Interest 4Abstra: first-sight overview of a dataset 5Pathways: efficiently finding interesting paths 6Systems developed 7Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 2 / 89
Motivation: data integration and exploration problems Outline 1Motivation: data integration and exploration problems 2Predihood: predicting neighbourhoods’ environment 3GeoAlign: spatial entity matching for Points of Interest 4Abstra: first-sight overview of a dataset 5Pathways: efficiently finding interesting paths 6Systems developed 7Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 3 / 89
Motivation: data integration and exploration problems Data exploration and integration Data exploration and integration Structured data models: Relational databases Tables Semi-structured data models: XML documents JSON documents RDF graphs Property graphs Dataset exploration and integration is hard: large, complex, irregular Today’s menu: focus on cartographic and semi-structured data Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 4 / 89
Motivation: data integration and exploration problems Data exploration and integration Data exploration and integration Structured data models: Relational databases Tables Semi-structured data models: XML documents JSON documents RDF graphs Property graphs Dataset exploration and integration is hard: large, complex, irregular Today’s menu: focus on cartographic and semi-structured data Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 4 / 89
Predihood: predicting neighbourhoods’ environment Outline 1Motivation: data integration and exploration problems 2Predihood: predicting neighbourhoods’ environment 3GeoAlign: spatial entity matching for Points of Interest 4Abstra: first-sight overview of a dataset 5Pathways: efficiently finding interesting paths 6Systems developed 7Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 5 / 89
Predihood: predicting neighbourhoods’ environment Motivation: heterogeneous data is everywhere Name: Jane Doe Job: French investigative journalist Sex: F Birth city: Paris Residence city: Lyon Wishes: Learn Lyon neighbourhoods [BDF+21] Visit Lyon’s monuments [BDFM19] Explore new datasets for her investigations [BMU24] Reveal undeclared conflicts of interests [BGLM23a] Skills: Excel: ? ? ?? Word: ? ? ?? Rel. databases: ?Semi-struct. data: N/A Aggregate city-level data Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 6 / 89
Predihood: predicting neighbourhoods’ environment Neighbourhood environment prediction Neighbourhood environment prediction INSEE (French National Institute of Statistics) IRIS: small geo unit of 5K inhabitants (50K IRIS in FR) For each IRIS: 600 quantitative features →No high-level description of neighbourhoods’ characteristics →Too many features for prediction Research contribution Predict automatically the environment of a any French neighbourhood, based on cartographic and city-level data Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 7 / 89
Predihood: predicting neighbourhoods’ environment Neighbourhood environment prediction Neighbourhood environment prediction INSEE (French National Institute of Statistics) IRIS: small geo unit of 5K inhabitants (50K IRIS in FR) For each IRIS: 600 quantitative features →No high-level description of neighbourhoods’ characteristics →Too many features for prediction Research contribution Predict automatically the environment of a any French neighbourhood, based on cartographic and city-level data Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 7 / 89
GeoAlign: spatial entity matching for Points of Interest From cartographic entities to POIS Adaptive formula for geographic entity matching Given two entities e1,e2, the adaptive formula relies on: The similar degree of e1and e2attributes 13 measures: geo, text, type, ... The weight/importance of e1and e2attributes f(e1,e2) = Pn i=1 weighti∗simi(attributei)> θ weight sim. measure attribute Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 14 / 89
GeoAlign: spatial entity matching for Points of Interest From cartographic entities to POIS GeoAlign at work Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 15 / 89
Abstra: first-sight overview of a dataset Outline 1Motivation: data integration and exploration problems 2Predihood: predicting neighbourhoods’ environment 3GeoAlign: spatial entity matching for Points of Interest 4Abstra: first-sight overview of a dataset 5Pathways: efficiently finding interesting paths 6Systems developed 7Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 16 / 89
Abstra: first-sight overview of a dataset Motivation: heterogeneous data is everywhere Name: Jane Doe Job: French investigative journalist Sex: F Birth city: Paris Residence city: Lyon Wishes: Learn Lyon neighbourhoods [BDF+21] Visit Lyon’s monuments [BDFM19] Explore new datasets for her investigations [BMU24] Reveal undeclared conflicts of interests [BGLM23a] Skills: Excel: ? ? ?? Word: ? ? ?? Rel. databases: ?Semi-struct. data: N/A Simple descriptions Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 17 / 89
Abstra: first-sight overview of a dataset What does the dataset describe? Real-world objects and relationships between them Entity-Relationship models [RG03] Need to compute them from the dataset! Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 18 / 89
Abstra: first-sight overview of a dataset What does the dataset describe? Real-world objects and relationships between them Entity-Relationship models [RG03] Need to compute them from the dataset! Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 18 / 89
Abstra: first-sight overview of a dataset What does the dataset describe? Real-world objects and relationships between them Entity-Relationship models [RG03] Need to compute them from the dataset! Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 18 / 89
Abstra: first-sight overview of a dataset What does the dataset describe? Real-world objects and relationships between them Entity-Relationship models [RG03] Need to compute them from the dataset! What about semi-structured data models (nesting)? Keep it simple and of controllable size Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 19 / 89
Abstra: first-sight overview of a dataset What does the dataset describe? Real-world objects and relationships between them Entity-Relationship models [RG03] Need to compute them from the dataset! What about semi-structured data models (nesting)? Keep it simple and of controllable size Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 19 / 89
Abstra: first-sight overview of a dataset Research contribution: data abstraction Abstra: Lightweight Entity-Relationship diagrams [BMU22,BMU24] Automatically and efficiently from semi-structured data Compact yet meaningful data overviews Ideal for first-sight dataset discovery Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 20 / 89
Abstra: first-sight overview of a dataset Data graph summarization Quotient summarization across data models Each data model has its own syntax: XML JSON RDF PG Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 24 / 89
Abstra: first-sight overview of a dataset Data graph summarization Summarization based on same-kind nodes We identify node kinds in each model based on the respective best practices for data design: XML: elements with the same label (or type) JSON: nodes on the same path from the root RDF [GGM20]: depending on node type(s) or, if absent, incoming and outgoing properties PG: adaptation of the above [GGM20] Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 25 / 89
Abstra: first-sight overview of a dataset Data graph summarization The summary (collection graph) G Collection node for each equivalence class paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 26 / 89
Abstra: first-sight overview of a dataset Data graph summarization The summary (collection graph) G Collection node for each equivalence class Collection edge Cs→Ctif a data edge exists paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 27 / 89
Abstra: first-sight overview of a dataset Data graph summarization The summary (collection graph) G Collection node for each equivalence class Collection edge Cs→Ctif a data edge exists Entity profile for each leaf collection node: reflects NEs in the leaves paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 28 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Identifying entities in the collection graph G paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Which collections represent entities in the E-R diagram? Which collections represent entity attributes? Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 29 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Identifying entities in the collection graph G paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Which collections represent entities in the E-R diagram? Which collections represent entity attributes? Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 29 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Identifying entities in the collection graph G paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Which collections represent entities in the E-R diagram? Which collections represent entity attributes? Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 29 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Requirements and algorithm We need an algorithm to identify entity roots and attributes for the E-R diagram For complex, potentially cyclic, collection graphs Greedy selection of few entities in G 1Assign a score to each collection node 2While less than Emax entity roots, or data coverage <covmin 1Elect the next highest-scored eligible collection node as an entity root 2Compute its boundary , i.e., attribute set 3Update the collection graph to reflect the selection of an entity 4Recompute the scores Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 30 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Requirements and algorithm We need an algorithm to identify entity roots and attributes for the E-R diagram For complex, potentially cyclic, collection graphs Greedy selection of few entities in G 1Assign a score to each collection node 2While less than Emax entity roots, or data coverage <covmin 1Elect the next highest-scored eligible collection node as an entity root 2Compute its boundary , i.e., attribute set 3Update the collection graph to reflect the selection of an entity 4Recompute the scores Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 30 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships How to score a collection node? 1wdesck,wleafk: # descendants, leaf descendants, at depth k 2Directed Acyclic Graph (DAG) rooted in each node: wDAG 3wPageRank : PageRank algorithm on G Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 36 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships PageRank score of a collection graph node paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val The collection graph G Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 37 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships PageRank score of a collection graph node paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val The reverse collection graph GR Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 38 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships PageRank score of a collection graph node paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 11 1 1 1 0.5 1 1 1 1 1 0.5 1 1 0.5 0.5 1 1 1 1 1 1 1 1 1 The reverse collection graph GRwith PR edge weights Collections distribute their score based solely on their connectivity Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 39 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships PageRank score of a collection graph node paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 11 1 1 1 0.5 1 1 1 1 1 0.5 1 1 0.5 0.5 1 1 1 1 1 1 1 1 1 The reverse collection graph GRwith PR edge weights Collections distribute their score based solely on their connectivity Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 39 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships How to score a collection node? 1wdesck,wleafk: # descendants, leaf descendants, at depth k 2wDAG :dw bottom-up propagation on G(outside cycles) 3wPageRank : PageRank algorithm on G 4wdwPageRank : PageRank algorithm on Gwith dw-tuned PR edge weights XReflects both the topology and where actual data is Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 40 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships The data-weighted PageRank score paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 The reverse collection graph GRwith dw-tuned PR edge weights Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 41 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships The data-weighted PageRank score paper .178 abstract .011 #val .006 year .011 #val .006 title .011 #val .006 wB .107 hW .158 pIn .063 inv .056 author .179 conf .067 name .011 #val .006 date .011 #val .006 email .011 #val .006 affiliation .027 university .024 city .011 #val .006 campus .011 #val .006 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 42 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships The data-weighted PageRank score paper .178 abstract .011 #val .006 year .011 #val .006 title .011 #val .006 wB .107 hW .158 pIn .063 inv .056 author .179 conf .067 name .011 #val .006 date .011 #val .006 email .011 #val .006 affiliation .027 university .024 city .011 #val .006 campus .011 #val .006 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 43 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships The data-weighted PageRank score paper .178 abstract .011 #val .006 year .011 #val .006 title .011 #val .006 wB .107 hW .158 pIn .063 inv .056 author .179 conf .067 name .011 #val .006 date .011 #val .006 email .011 #val .006 affiliation .027 university .024 city .011 #val .006 campus .011 #val .006 1 1 11 1 1 1 0.66 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Propagates scores across the collection graph Works on cyclic collection graphs The score reflects the topology and where the data is A collection node distributes its weight Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 44 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships How to update the collection graph after selecting an entity? Reflect the allocation of data nodes and edges to one entity 1updateboolean Collection nodes and edges in the boundary of the entity Very efficient Sufficient for wdesck,wleafk,wDAG 2updateexact Graph nodes and edges Much more costly Required for wPageRank ,wdwPageRank Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 49 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships How to update the collection graph after selecting an entity? Reflect the allocation of data nodes and edges to one entity 1updateboolean Collection nodes and edges in the boundary of the entity Very efficient Sufficient for wdesck,wleafk,wDAG 2updateexact Graph nodes and edges Much more costly Required for wPageRank ,wdwPageRank Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 49 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships How to update the collection graph after selecting an entity? Reflect the allocation of data nodes and edges to one entity 1updateboolean Collection nodes and edges in the boundary of the entity Very efficient Sufficient for wdesck,wleafk,wDAG 2updateexact Graph nodes and edges Much more costly Required for wPageRank ,wdwPageRank Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 49 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Exact graph update Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 50 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Exact graph update paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 51 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Selected entities and their boundaries paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 52 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Finding relationships between entities Relationship: a path from an entity to another paper abstract #val year #val title #val wB hW pIn inv author conf name #val date #val email #val affiliation university city #val campus #val 1 1 1 1 1 1 1 1 1 1 1 1 1 0.33 1 1 0.4 0.6 1 1 1 1 1 1 1 1 1 paper →wB →author paper →pIn →conf author →hW →paper conf →inv →author Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 53 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Entity classification Assign a semantic category to each entity Input: an entity E, categories K, semantic properties P K: Person, ScientificPaper, Event, Website, Mountain, ... P:{label:"address", domain:[Pers., Org.], range:[Place]}, ... Output: a category for E Algorithm: Compare: The common name of all nodes in the entity root (if it exists) with k∈ K (conf, paper, author) Its attribute names with p∈ P (affiliation, email, ...) Its entity profiles with p.range ∈ P (,,, ...) Each good match votes for one or few categories Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 54 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Entity classification Assign a semantic category to each entity Input: an entity E, categories K, semantic properties P K: Person, ScientificPaper, Event, Website, Mountain, ... P:{label:"address", domain:[Pers., Org.], range:[Place]}, ... Output: a category for E Algorithm: Compare: The common name of all nodes in the entity root (if it exists) with k∈ K (conf, paper, author) Its attribute names with p∈ P (affiliation, email, ...) Its entity profiles with p.range ∈ P (,,, ...) Each good match votes for one or few categories Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 54 / 89
Abstra: first-sight overview of a dataset Identifying entities and relationships Entity classification Name Similar to Votes for paper ResearchPublication (0.85) ResearchPublication News (0.63) News paper abstract #val year #val title #val 1 1 1 1 1 1 Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 55 / 89
Abstra: first-sight overview of a dataset Experimental evaluation Experimental evaluation On main semi-structured data models: 8 JSON, 7 RDF, 5 XML, 3 PG 10 synthetic, 13 real-world 5M to 14M nodes Collection graphs: 26 to 4.8K collections 14/23 have cycles Graphs stored in PostgreSQL, algorithms in SQL and Java We evaluate: 1Entity selection quality 2Scalability Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 61 / 89
Abstra: first-sight overview of a dataset Experimental evaluation Experimental evaluation On main semi-structured data models: 8 JSON, 7 RDF, 5 XML, 3 PG 10 synthetic, 13 real-world 5M to 14M nodes Collection graphs: 26 to 4.8K collections 14/23 have cycles Graphs stored in PostgreSQL, algorithms in SQL and Java We evaluate: 1Entity selection quality 2Scalability Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 61 / 89
Abstra: first-sight overview of a dataset Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 62 / 89
Abstra: first-sight overview of a dataset Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 63 / 89
Abstra: first-sight overview of a dataset Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 64 / 89
Abstra: first-sight overview of a dataset Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 65 / 89
Abstra: first-sight overview of a dataset Experimental evaluation Entity selection quality with (wdwPageRank ,boundfl−ac ) Dataset name |C| |ME| |MR| cov ME dmax |MEi| Mondial 168 5 8 0.85 City Province Country Organization River 3 3 4 4 4 3,152 1,455 231 168 135 PubMed 26 1 0 1.0 PubMedArticle 5 957 XMark1 136 5 10 0.91 Person Item Open Auction Closed Auction Category 4 7 8 8 2 25,500 21,750 12,000 9,750 1,000 XMark4 136 5 10 0.90 Person Item Open Auction Closed Auction Category 4 7 8 8 2 102,000 87,000 48,000 39,000 4,000 Wikimedia 59 2 0 1.0 Page Namespace 4 3 54,750 32 Abstra selects frequent, coherent and semantically central entities Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 66 / 89
Abstra: first-sight overview of a dataset Experimental evaluation Experimental evaluation: scalability Our abstraction method scales up linearly in the data size Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 67 / 89
Abstra: first-sight overview of a dataset Related work Related work Data summarization Structural Quotient [GGM20,KC10,MS99] (the one we adopt to build G) Non-quotient [GW97] Pattern mining [ZLVK16] Statistical [HS12] Hybrid [RGSB17] Schema inference XML [CGS11] JSON [BCGS19] RDF [GLSW22] PG [LBH21] Data summarization and schema inference are tied to one data model Schemas are often not suited to NTUs Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 68 / 89
Abstra: first-sight overview of a dataset Related work A JSON schema from social network data using [BCGS19] Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 69 / 89
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths connecting Person NEs () to Organization NEs () ←#val ←Name ←Author →Affiliation →#val → ←#val ←Name ←Author ←Authors ←Article →Journal →#val → ←#val ←COI ←Article →Journal →#val →←#val → Which paths are most interesting and deserve to be evaluated? Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 74 / 89
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths connecting Person NEs () to Organization NEs () ←#val ←Name ←Author →Affiliation →#val → ←#val ←Name ←Author ←Authors ←Article →Journal →#val → ←#val ←COI ←Article →Journal →#val →←#val → Which paths are most interesting and deserve to be evaluated? Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 74 / 89
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths connecting Person NEs () to Organization NEs () ←#val ←Name ←Author →Affiliation →#val → ←#val ←Name ←Author ←Authors ←Article →Journal →#val → ←#val ←COI ←Article →Journal →#val →←#val → Which paths are most interesting and deserve to be evaluated? Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 74 / 89
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths connecting Person NEs () to Organization NEs () ←#val ←Name ←Author →Affiliation →#val → ←#val ←Name ←Author ←Authors ←Article →Journal →#val → ←#val ←COI ←Article →Journal →#val →←#val → Which paths are most interesting and deserve to be evaluated? Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 74 / 89
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths are unreliable: we face entity extraction errors E.g., “John Hopkins | {z } person University Hospital” False positives, or wrong entity type attribution, e.g., “THC |{z} org. ” Some paths are structurally weak: we face information dilution E.g., a paper has 50 authors Path interestingness : based on edge reliability and edge force Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 75 / 89
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths are unreliable: we face entity extraction errors E.g., “John Hopkins | {z } person University Hospital” False positives, or wrong entity type attribution, e.g., “THC |{z} org. ” Some paths are structurally weak: we face information dilution E.g., a paper has 50 authors Path interestingness : based on edge reliability and edge force Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 75 / 89
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? Some paths are unreliable: we face entity extraction errors E.g., “John Hopkins | {z } person University Hospital” False positives, or wrong entity type attribution, e.g., “THC |{z} org. ” Some paths are structurally weak: we face information dilution E.g., a paper has 50 authors Path interestingness : based on edge reliability and edge force Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 75 / 89
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? 1Reliability r(Ci99K )of an extraction collection edge The ratio of NEs having the type , and extracted from Ci Path reliability: minimum extraction edge reliability 2Force f(Ci→Cj)of a structural collection edge The inverse of the maximal source node out-degree among data edges represented by Ci→Cj Path force: product of edge forces 3Rank paths on their reliability, then their force 4Take a top-kor those having r≥θ Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 76 / 89
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? 1Reliability r(Ci99K )of an extraction collection edge The ratio of NEs having the type , and extracted from Ci Path reliability: minimum extraction edge reliability 2Force f(Ci→Cj)of a structural collection edge The inverse of the maximal source node out-degree among data edges represented by Ci→Cj Path force: product of edge forces 3Rank paths on their reliability, then their force 4Take a top-kor those having r≥θ Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 76 / 89
Pathways: efficiently finding interesting paths NE-to-NE path enumeration What makes a NE-to-NE path interesting? 1Reliability r(Ci99K )of an extraction collection edge The ratio of NEs having the type , and extracted from Ci Path reliability: minimum extraction edge reliability 2Force f(Ci→Cj)of a structural collection edge The inverse of the maximal source node out-degree among data edges represented by Ci→Cj Path force: product of edge forces 3Rank paths on their reliability, then their force 4Take a top-kor those having r≥θ Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 76 / 89
Pathways: efficiently finding interesting paths Experimental evaluation Experimental evaluation: path interestingness (τ1, τ2) min prel max prel p20 rel |P| |P0|R=|P0| |P| PubMed (Person, Organization) 0.0150 0.9142 0.0409 52 20 38.45% (Person, Location) 0.0150 0.9107 0.0150 30 20 66.66% (Location, Organization) 0.0150 0.9107 0.0232 34 20 58.82% (Person, Person) 0.0150 0.9774 0.0150 24 20 83.33% (Organization, Organization) 0.0150 0.4158 0.0232 31 20 64.51% (Location, Location) 0.0150 0.0954 0.0150 20 20 100.00% Nasa (Person, Organization) 0.0014 0.0645 0.0178 191 20 10.47% (Person, Location) 0.0014 0.0645 0.0077 142 20 14.08% (Location, Organization) 0.0014 0.1016 0.0077 115 20 17.39% (Person, Person) 0.0014 0.0232 0.0077 110 20 18.18% (Organization, Organization) 0.0014 0.0581 0.0077 92 20 21.73% (Location, Location) 0.0014 0.3790 0.0077 67 20 29.85% Yelp (Location, Organization) 0.0002 0.9997 0.0002 8 8 100.00% (Location, Location) 0.0002 1.0000 0.0002 11 11 100.00% Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 81 / 89
Pathways: efficiently finding interesting paths Experimental evaluation Experimental evaluation: path interestingness (τ1, τ2) min prel max prel p20 rel |P| |P0|R=|P0| |P| PubMed (Person, Organization) 0.0150 0.9142 0.0409 52 20 38.45% (Person, Location) 0.0150 0.9107 0.0150 30 20 66.66% (Location, Organization) 0.0150 0.9107 0.0232 34 20 58.82% (Person, Person) 0.0150 0.9774 0.0150 24 20 83.33% (Organization, Organization) 0.0150 0.4158 0.0232 31 20 64.51% (Location, Location) 0.0150 0.0954 0.0150 20 20 100.00% Nasa (Person, Organization) 0.0014 0.0645 0.0178 191 20 10.47% (Person, Location) 0.0014 0.0645 0.0077 142 20 14.08% (Location, Organization) 0.0014 0.1016 0.0077 115 20 17.39% (Person, Person) 0.0014 0.0232 0.0077 110 20 18.18% (Organization, Organization) 0.0014 0.0581 0.0077 92 20 21.73% (Location, Location) 0.0014 0.3790 0.0077 67 20 29.85% Yelp (Location, Organization) 0.0002 0.9997 0.0002 8 8 100.00% (Location, Location) 0.0002 1.0000 0.0002 11 11 100.00% Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 82 / 89
Pathways: efficiently finding interesting paths Experimental evaluation Experimental evaluation: path interestingness (τ1, τ2) min prel max prel p20 rel |P| |P0|R=|P0| |P| PubMed (Person, Organization) 0.0150 0.9142 0.0409 52 20 38.45% (Person, Location) 0.0150 0.9107 0.0150 30 20 66.66% (Location, Organization) 0.0150 0.9107 0.0232 34 20 58.82% (Person, Person) 0.0150 0.9774 0.0150 24 20 83.33% (Organization, Organization) 0.0150 0.4158 0.0232 31 20 64.51% (Location, Location) 0.0150 0.0954 0.0150 20 20 100.00% Nasa (Person, Organization) 0.0014 0.0645 0.0178 191 20 10.47% (Person, Location) 0.0014 0.0645 0.0077 142 20 14.08% (Location, Organization) 0.0014 0.1016 0.0077 115 20 17.39% (Person, Person) 0.0014 0.0232 0.0077 110 20 18.18% (Organization, Organization) 0.0014 0.0581 0.0077 92 20 21.73% (Location, Location) 0.0014 0.3790 0.0077 67 20 29.85% Yelp (Location, Organization) 0.0002 0.9997 0.0002 8 8 100.00% (Location, Location) 0.0002 1.0000 0.0002 11 11 100.00% Both reliability and force downgrade meaningless paths (NE errors or structurally weak) Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 83 / 89
Pathways: efficiently finding interesting paths Related work Related work Structured querying SQL, SPARQL, GQL [DFG+22] Assisted struct. querying Interactive queries [DAB16] Guided query writing [ERAAL18,KKBS10] NL2SQL [KSHL20] Keyword-based search Unidirectional [ABC+02,LOF+08] Bi-directional [ABC+22] Path search in struct. queries SPARQL extensions: [ASMH18,AMSH18, AMM23] For PGs: [DFG+22] Pathways users need no knowledge of the graph structure or values Less intimidating for NTUs Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 84 / 89
Systems developed Outline 1Motivation: data integration and exploration problems 2Predihood: predicting neighbourhoods’ environment 3GeoAlign: spatial entity matching for Points of Interest 4Abstra: first-sight overview of a dataset 5Pathways: efficiently finding interesting paths 6Systems developed 7Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 85 / 89
Systems developed Systems developed Predihood GeoAlign City environment prediction Entity matching for POIs 17 Python core classes 41 PHP core classes DATA 2020 [BDF+21] SIGSPATIAL 2019 [BDFM19] Abstra PathWays Abstractions as E-R diagrams Interesting NE-to-NE paths 65 Java core classes 18 Java core classes CIKM 2022 [BMU22] ESWC 2023 [BGLM23b] Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 86 / 89
Conclusion Outline 1Motivation: data integration and exploration problems 2Predihood: predicting neighbourhoods’ environment 3GeoAlign: spatial entity matching for Points of Interest 4Abstra: first-sight overview of a dataset 5Pathways: efficiently finding interesting paths 6Systems developed 7Conclusion Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 87 / 89
Conclusion Lessons learned Data integration and exploration are difficult: Lack of schema or schema heterogeneity Data quality: wrong, null, missing values, ... Large amounts of data Bring out insights and knowledge from raw data From the user point of view: 1User-friendly interfaces 2No technical detail 3High-level representation Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 88 / 89
Conclusion Thanks Nelly BARRET ¥[email protected] §https://nelly-barret.github.io/ Data Science group DEIB, Politecnico di Milano Milano Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 89 / 89
References I B Aditya, Gaurav Bhalotia, Soumen Chakrabarti, Arvind Hulgeri, Charuta Nakhe, S Sudarshanxe, et al. BANKS: browsing and keyword searching in relational databases. In VLDB’02: Proceedings of the 28th International Conference on Very Large Databases, pages 1083–1086. Elsevier, 2002. Angelos Anadiotis, Oana Balalau, Catarina Conceicao, et al. Graph integration of structured, semistructured and unstructured data for data journalism. Inf. Systems, 104, 2022. Angelos Christos Anadiotis, Ioana Manolescu, and Madhulika Mohanty. Integrating connection search in graph queries. In ICDE, April 2023. Christian Aebeloe, Gabriela Montoya, Vinay Setty, and Katja Hose. Discovering diversified paths in knowledge bases. Proc. VLDB Endow., 11(12):2002–2005, 2018. Code available at: http://qweb.cs.aau.dk/jedi/. Christian Aebeloe, Vinay Setty, Gabriela Montoya, and Katja Hose. Top-k diversification for path queries in knowledge graphs. In Marieke van Erp, Medha Atre, Vanessa L´opez, Kavitha Srinivas, and Carolina Fortuna, editors, Proceedings of the ISWC 2018 Posters & Demonstrations, Industry and Blue Sky Ideas Tracks co-located with 17th International Semantic Web Conference (ISWC 2018), Monterey, USA, October 8th - to - 12th, 2018, volume 2180 of CEUR Workshop Proceedings. CEUR-WS.org, 2018. Mohamed Amine Baazizi, Dario Colazzo, Giorgio Ghelli, and Carlo Sartiani. Parametric schema inference for massive JSON datasets. VLDB J., 28(4), 2019. Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 1 / 14
RDF quotient graph summarization [GGM20] Source clique: set of outgoing properties co-occuring together on at least one node Target clique: set of incoming properties co-occuring together on at least one node Properties “a”, “b”, “d” are in the same source clique Properties “a” and “e” are in the same target clique (c) Pawel Guzewic Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 8 / 14
Strong summary [GGM20] Strong S summary: Two nodes are S equivalent iff they have both the same source and target cliques Source and target cliques for each node Strong summary (c) Pawel Guzewic Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 9 / 14
Typed-strong summary [GGM20] Typed-strong TS summary: Two typed nodes are TS equivalent iff they have the same type set Two untyped nodes are TS equivalent iff they have both the same source and target cliques Source and target cliques for each node + an RDF type Typed-strong summary (c) Pawel Guzewic Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 10 / 14
Disagreement between Flair and ChatGPT False Flair positives: Flair identifies “Av. Peter Henry Rolfs | {z } person 36570-900 Vicosa” Flair mislead by capitalization: Flair identifies “Claudin-7b | {z } person ” (but not ChatGPT) Different token allocation: “University of Alabama | {z } org. ”, “Birmingham | {z } loc. ” “University of Alabama, Birmingham | {z } loc. ” Missed non-English spelling/names: ChatGPT finds “Antonio Gonz´alez | {z } person ” ChatGPT finds “Yoshida, Sakyo-ku, Kyoto 606-8501, Japan | {z } loc. ” Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 11 / 14
A comprehensive data exploration tool for NTUs Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 12 / 14
A comprehensive data exploration tool for NTUs Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 13 / 14
Experimental evaluation: Flair VS ChatGPT NE extractors Flair and ChatGPT mostly agree ChatGPT extraction has better quality Nelly Barret (DEIB@PoliMi) Data integration and exploration April 19, 2024 14 / 14