scieee AI-readable full text Open interactive document viewer

Multimodal data without borders: integration and exploration to the rescue

Barret, Nelly

Abstract

The unprecedented creation, use and share of data around the world has led to new applications and economic opportunities. This data is often large, heterogeneous at a schema and model level, and more or less structured. To bring more order, the World Wide Web consortium recommends sharing data as RDF graphs, which has been mostly adopted in the Open Data initiative, but many other formats are used in practice. Moreover, data is now scattered across places and owners (enterprise silos, open data, Big Tech clouds, etc.) and this adds to the complexity of managing, joining, processing and using various data together. Finally, end users, such as domain experts or decision makers, need tools to generate tangible results they can rely on and share. In this presentation, I will present my PhD and post-doctoral work on this domain. I will first show how to compute structured summaries from any semi-structured dataset. Next, I will show how to homogenize healthcare silos in order to run various federated learning tasks on the underlying data.

Full text

1 Multimodal data without borders: integration and exploration to the rescue Nelly Barret Assistant Professor at INSA Lyon and LIRIS October 27th, 2025 2/56 PhD @ Inria Saclay and Ecole Polytechnique • User-oriented exploration of semi-structured data Post-doc @ Politecnico di Milano (Italy) • FL analyses over multimodal healthcare data Assistant Professor @ INSA Lyon and LIRIS • Data management and Federated Learning SHORT BIO INTRODUCTION 2021-2024 2024-2025 Sept. 2025 3/56 Motivation 4/56 MULTIMODAL DATA IS EVERYWHERE Varied sources, actors, and standards → multimodality is everywhere! INTRODUCTION 5/56 MULTIMODAL DATA IS EVERYWHERE Varied sources, actors, and standards → multimodality is everywhere! INTRODUCTION How to help domain experts make sense of various multimodal data? Multimodal data usage is hard: different formats, schemas, granularities, … 6/56 AGENDA 1. Motivation: multimodal data is everywhere 2. PhD: exploring unknown semi-structured datasets 3. Post-doc: healthcare analytics across hospitals 4. Conclusions INTRODUCTION 7/56 PhD: exploring unknown semi-structured datasets 8/56 WHAT DOES THE DATASET DESCRIBE? •Real-world objects and relationships between them •Traditional setting: Entity-Relationship models [RG03] •Need to compute them from the dataset! •What about semi-structured data models (nesting)? •Keep it simple and of controllable size PhD: EXPLORING SEMI-STRUCTURED DATASETS 9/56 PHD PROBLEM STATEMENT AND CONTRIBUTIONS How to facilitate user exploration of unknown heterogeneous semi-structured datasets? 1. Abstra: semi-structured data overviews [BMU22, BMU24] •Automatically compute lightweight Entity-Relationship diagrams •Ideal for first-sight dataset discovery 2. PathWays: interesting Named Entity connections [BGLM23b, BGLM23a, BGLM24] •Compute and rank entity paths in and across datasets •Ideal for exploring connections within and across sources PhD: EXPLORING SEMI-STRUCTURED DATASETS 16/56 QUOTIENT SUMMARIZATION ACROSS DATA MODELS Each data model has its own syntax: PhD: EXPLORING SEMI-STRUCTURED DATASETS XML RDF JSON PG 17/56 SUMMARIZATION BASED ON SAME-KIND NODES We identify node kinds in each model following best practices for data design: •XML:elements with the same label (or type) • JSON:nodes on the same path from the root •RDF [GGM20]: depending on node type(s) or, if absent, incoming and outgoing properties •PG:adaptation of the above [GGM20] We obtain apartition over the graph: aset of equivalence classes PhD: EXPLORING SEMI-STRUCTURED DATASETS 18/56 THE SUMMARY (COLLECTION GRAPH) 𝒢 •Collection node: one for each equivalence class •Collection edge: Cs→Ctif adata edge exists •Entity profile for each leaf collection node:reflects NEs in the leaves PhD: EXPLORING SEMI-STRUCTURED DATASETS 19/56 IDENTIFYING ENTITIES IN THE COLLECTION GRAPH 𝒢 PhD: EXPLORING SEMI-STRUCTURED DATASETS Which collections represent entities in the E-R diagram? Which collections represent entity attributes? 20/56 REQUIREMENTS AND ALGORITHM We need an algorithm to identify entity roots and attributes for the E-R diagram §For complex, potentially cyclic, collection graphs Greedy selection of few entities in 𝒢 : 1. Assign ascore to each collection node 2. While less than 𝐸!"# entity roots or data coverage <𝑐𝑜𝑣!$%: a. Elect the next highest-scored eligible collection node as an entity root b. Compute its boundary (set of attributes) c. Update the collection graph to reflect the selection of an entity d. Recompute the scores PhD: EXPLORING SEMI-STRUCTURED DATASETS 21/56 HOW TO SCORE A COLLECTION NODE? Objective: reflect the weight of this node and its structure in the dataset •w&'()!and 𝑤*'"+!: #descendants, #leaf descendants, at depth k Not clear how to pick k •𝑊 ,-.: Directed Acyclic Graph (DAG) rooted in each node Does not work on cyclic graphs •𝑊 /0:PageRank algorithm on 𝒢 The score is solely based on the position of the node in the graph •𝑊 &12/0:PageRank algorithm on 𝒢with dw-tuned PR edge weights Reflects both the topology and where actual data is PhD: EXPLORING SEMI-STRUCTURED DATASETS 22/56 THE DATA-WEIGHTED PAGERANK SCORE PhD: EXPLORING SEMI-STRUCTURED DATASETS The original collection graph 𝒢 23/56 THE DATA-WEIGHTED PAGERANK SCORE PhD: EXPLORING SEMI-STRUCTURED DATASETS The reverse collection graph 𝒢0 24/56 THE DATA-WEIGHTED PAGERANK SCORE PhD: EXPLORING SEMI-STRUCTURED DATASETS The reverse collection graph 𝒢0 with dw-tuned PR edge weights 25/56 DATA-WEIGHT PAGERANK SCORES IN 𝒢 PhD: EXPLORING SEMI-STRUCTURED DATASETS 𝑊 &12/0 score computation on 𝒢0 with dw-tuned PR edge weights 32/56 FINDING RELATIONSHIPS BETWEEN ENTITIES PhD: EXPLORING SEMI-STRUCTURED DATASETS Paths between entity roots: •paper → wB → author •paper → pIn → conf •author → hW → paper •conf → inv → author Remaining task: classify each entity into a semantic category 33/56 ABSTRA OUTPUT: A LIGHTWEIGHT E-R DIAGRAM 318 person (Person) •person@id (100 %) •phone (49 %) •creditcard (49 %) •homepage (47 %) •address (46 %) •province (52 %) •zipcode (100 %) •country (100 %) •city (100 %) •street (100 %) •emailaddress (100 %) •name (100 %) 150 open_auction (Product) •privacy (56 %) •interval (100 %) •end (100 %) •start (100 %) •type (100 %) •current (100 %) •reserve (51 %) •initial (100 %) •open_auction@id (100 %) •quantity (100 %) watches.watch@open_auction 12 category (Thing) •category@id (100 %) •description (100 %) •text (73 %) •parlist (27 %) •listitem (291 %) •text (87 %) •name (100 %) profile.interest@category seller@person annotation.author@person bidder.personref@person 270 item (schema:how_to_item) •mailbox (64 %) •mail (101 %) •date (100 %) •to (100 %) •from (100 %) •text (100 %) •item@featured (9 %) •item@id (100 %) •shipping (94 %) •description (100 %) •text (73 %) •parlist (27 %) •listitem (291 %) •text (87 %) •payment (94 %) •name (100 %) •quantity (100 %) •location (100 %) itemref@item incategory@category 120 closed_auction (Product) •price (100 %) •type (100 %) •date (100 %) •quantity (100 %) seller@person buyer@person annotation.author@person itemref@item PhD: EXPLORING SEMI-STRUCTURED DATASETS 34/56 ABSTRA OUTPUT: A LIGHTWEIGHT E-R DIAGRAM PhD: EXPLORING SEMI-STRUCTURED DATASETS 35/56 QUICK OVERVIEW OF THE EXPERIMENTAL EVALUATION Main semi-structured data models: 8 JSON, 7 RDF, 5 XML, 3 PG 10 synthetic, 13 real-world / 5M to 14M nodes Collection graphs: 26 to 4.8K collections, 14/23 have cycles PhD: EXPLORING SEMI-STRUCTURED DATASETS Our abstraction method scales up linearly in the data size Abstra selects frequent, coherent and semantically central entities 36/56 Post-doc: healthcare analytics across hospitals 37/56 WHAT DOES HEALTHCARE DATA HAVE TO REVEAL? •Few cooperation/normalization between centers •Very few patient data for rare diseases •Traditional approach: consolidated warehouse [KR13] •Need to provide decentralized and federated solutions •Need to leverage expert knowledge + automatization POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS 38/56 POST-DOC PROBLEM STATEMENT AND CONTRIBUTIONS How to enable federated analyses of healthcare data across institutions and national borders? 1. I-ETL: an ETL to build interoperable healthcare databases [BBBP25] §Provides a general and extensible metadata and data model §Assesses interoperability along the ETL pipeline §Produces an interoperable warehouse at each medical center 2. A general catalogue for exploring healthcare silos [BBCP+25] §Facilitates easy exploration and federated learning algorithms design §Automatically profiles underlying silos POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS 39/56 EXISTING WORKS POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS Data Platforms EDHEN [ PdGdK+23] OHDSI [ HDS+15] UMG -MeDIC [PSS+23] … • Lack of generality → IT and experts’ knowledge and effort • Use-case tailored approaches → from scratch for every project • Existing models hardly fit on existing data Conceptual models OMOP [ SRR+10] FHIR … Data Catalogues EDHEN: pre - defined statistics OHDSI: medicine - specific interactions … 40/56 THE BETTER APPROACH 1. Analyze datasets and extract their metadata 2. Create an interoperable database in each medical center 3. Assess interoperability along the pipeline 4. Allow federated analyses of data across centers through a catalogue POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS 41/56 INTEROPERABILITY AS PART OF FAIR PRINCIPLES FAIR principles are guidelines for good data management [WDA+16]: •Findable: search for (indexed) resources based on identifiers •Accessible: access data with standard protocols, even after data dies •Interoperable: integrate and refer to datasets following FAIR principles •Reusable: reuse datasets in other settings using provenance, etc. POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS 48/56 MODELING AND PROFILING HEALTHCARE SILOS POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS Silo profiling [BBP+25] 49/56 OUR CATALOGUE DATA MODEL Profiled entities: •Dataset: a file •Feature: any variable captured by a dataset Other entities: •Station (hospital) •Resource (hospital setup) POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS 50/56 MODELING AND PROFILING HEALTHCARE SILOS POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS Silo profiling [BBP+25] 51/56 PROFILING SILOS FOR INDIVIDUAL CATALOGUES 11/09/2025 Healthcare silos unlocked - Nelly Barret et al. 1. For each hospital: §Get station information §List its computational resources 2. For each dataset, compute its DatasetProfile with: §Statistics §Information from experts 3. For each feature: §Compute its FeatureProfile: metadata, statistics, aggregates §Extract its FeatureDomain from metadata POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS 52/56 MODELING AND PROFILING HEALTHCARE SILOS POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS Silo profiling [BBP+25] 53/56 BUILDING THE GLOBAL CATALOGUE Combine individual catalogues of each silo: Union of all the individual catalogues (one per hospital) §Easy thanks to our general catalogue model §Allows to implement visualisations for all centers at a time When hospitals update their data: 1. Recompute the underlying catalogue 2. Replace the old with the new individual catalogue POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS 54/56 I-ETL AND THE CATALOGUE AT WORK 3 use-cases on genetic rare diseases: 1. Understand paediatric intelligence disability 2. Correlate eye dystrophies and gene mutations 3. Understand autistic patients' self harm 3 databases built with synthetic data with I-ETL The catalogue interface is up and helps experts for designing FL algorithms Trains are currently developed for UC1 POST-DOC: HEALTHCARE ANALYTICS ACROSS HOSPITALS 55/56 PHD AND POST-DOC LESSONS LEARNT Existing works: lack of generality, approach per format or use-case Abstracting current approaches/models is promising: ØReuse, use case-agnostic pipelines, various settings Balance generalization vs. domain knowledge Hard to collect and translate experts' knowledge / wishes CONCLUSION 56/56 IF YOU ARE INTERESTED Abstra: team.inria.fr/cedar/projects/abstra PathWays: team.inria.fr/cedar/projects/pathways → ConnectionStudio: connectionstudio.inria.fr Better:https://www.better-health-project.eu/ CONCLUSION Suivez les actualités de l’école www.insa-lyon.fr Rejoignez-nous sur les réseaux sociaux