scieee AI-readable full text Open interactive document viewer

rapid-triples: Adaptive Forms for Semi-automatic Knowledge Collection in RDF

Scrocca, Mario; Carenini, Alessio; Carriero, Valentina Anita; Celino, Irene

Abstract

Slides for the presentation of the paper "rapid-triples: Adaptive Forms for Semi-automatic Knowledge Collection in RDF" accepted for publication at the 1st Workshop on Bridging Hybrid (Artificial) Intelligence and the Semantic Web (HAIBridge 2025) co-located with ISWC 2025. Authors: Mario Scrocca, Alessio Carenini, Valentina Anita Carriero and Irene CelinoAbstract: To reduce inaccuracies or a lack of context, AI applications may heavily benefit from structured knowledge modelled relying on reference ontologies. We present rapid-triples, a customizable and dynamic interface that facilitates human-in-the-loop collection of structured knowledge in RDF format. The system allows domain experts to manually input knowledge or semi-automatically enhance an extraction by an automated system from existing data sources. By utilizing a common schema mapped to a target ontology, rapid-triples ensures that the knowledge collected is semantically interoperable and machine-readable. This tool supports various use cases, including expert-guided knowledge creation, validating and refining outputs from automated extractors, and generating high-quality training data for AI systems. Finally, we discuss the adoption of rapid-triples to support a use case in the industrial domain through the collection and exploitation of procedures as knowledge graphs. Paper: https://ceur-ws.org/Vol-4093/Paper4hai.pdf

Full text

182 211 131 103 183 0 47 60 182 61 177 166 127 211 203 183 233 229 235 249 248 46 133 124 4 77 92 103 183 0 248 203 14 #1 #2 #3 #4 #5 8 134 53 103 183 0 150 189 71 182 211 131 210 228 178 #1 #2 #3 #4 #5 Primary 5 3 14 2 5 3 1 4 2 GRAPHS Sequence of use PERKS IDENTITY 46 133 124 224 83 72 00 00 00 Text Title 1 Title 2 TEXT Secondary Sequence of use BACKGROUND RAG 46 133 124 4 77 92 8 134 53 Mario Scrocca, Alessio Carenini, Valentina Carriero, and Irene Celino Cefriel –Italy HAIBridge 2025 1st Workshop on Bridging Hybrid (Artificial) Intelligence and the Semantic Web, co-located with ISWC 2025, Nara, Japan Adaptive Forms for Semi-automatic Knowledge Collection in RDF rapid-triples 182 211 131 103 183 0 47 60 182 61 177 166 127 211 203 183 233 229 235 249 248 46 133 124 4 77 92 103 183 0 248 203 14 #1 #2 #3 #4 #5 8 134 53 103 183 0 150 189 71 182 211 131 210 228 278 #1 #2 #3 #4 #5 Primary 5 3 14 2 5 3 1 4 2 GRAPHS Sequence of use PERKS IDENTITY 46 133 124 224 83 72 00 00 00 Text Title 1 Title 2 TEXT Secondary Sequence of use BACKGROUND RAG 46 133 124 4 77 92 8 134 53 R A P I D - T R I P L E S ( H A I B R I D G E ’ 2 5 ) Problem Addressed ‣To reduce inaccuracies or a lack of context, AI applications heavily benefit from structured knowledge ‣Collecting the knowledge and building a Knowledge Graph according to reference ontologies is a challenge requiring domain experts' involvement ‣We propose the rapid-triples tool to support expert-guided knowledge creation and AIenabled knowledge completion from existing unstructured documents 182 211 131 103 183 0 47 60 182 61 177 166 127 211 203 183 233 229 235 249 248 46 133 124 4 77 92 103 183 0 248 203 14 #1 #2 #3 #4 #5 8 134 53 103 183 0 150 189 71 182 211 131 210 228 278 #1 #2 #3 #4 #5 Primary 5 3 14 2 5 3 1 4 2 GRAPHS Sequence of use PERKS IDENTITY 46 133 124 224 83 72 00 00 00 Text Title 1 Title 2 TEXT Secondary Sequence of use BACKGROUND RAG 46 133 124 4 77 92 8 134 53 R A P I D - T R I P L E S ( H A I B R I D G E ’ 2 5 ) Workflows for Knowledge Collection •W1 Tacit Knowledge Collection: relevant knowledge in the minds of domain experts and not yet documented. The user must be guided in articulating and formalising it according to the target ontology. •W2 Knowledge Completion from Unstructured Sources: initial knowledge is automatically extracted from unstructured content using AI-based tools. A human-in-theloop (HITL) process ensures that the resulting knowledge is accurate, complete and semantically consistent with the target ontology. •W3 Enhance Automatic Knowledge Extraction Systems: structured knowledge, validated and reviewed by users, is used as training or contextual data to enhance the accuracy of automatic knowledge extraction solutions. 182 211 131 103 183 0 47 60 182 61 177 166 127 211 203 183 233 229 235 249 248 46 133 124 4 77 92 103 183 0 248 203 14 #1 #2 #3 #4 #5 8 134 53 103 183 0 150 189 71 182 211 131 210 228 278 #1 #2 #3 #4 #5 Primary 5 3 14 2 5 3 1 4 2 GRAPHS Sequence of use PERKS IDENTITY 46 133 124 224 83 72 00 00 00 Text Title 1 Title 2 TEXT Secondary Sequence of use BACKGROUND RAG 46 133 124 4 77 92 8 134 53 R A P I D - T R I P L E S ( H A I B R I D G E ’ 2 5 ) Rula, Anisa, et al. "Annotation and Extraction of Industrial Procedural Knowledge from Textual Documents." Proceedings of the 12th Knowledge Capture Conference 2023. 2023. Automatic Knowledge Collection in RDF 182 211 131 103 183 0 47 60 182 61 177 166 127 211 203 183 233 229 235 249 248 46 133 124 4 77 92 103 183 0 248 203 14 #1 #2 #3 #4 #5 8 134 53 103 183 0 150 189 71 182 211 131 210 228 278 #1 #2 #3 #4 #5 Primary 5 3 14 2 5 3 1 4 2 GRAPHS Sequence of use PERKS IDENTITY 46 133 124 224 83 72 00 00 00 Text Title 1 Title 2 TEXT Secondary Sequence of use BACKGROUND RAG 46 133 124 4 77 92 8 134 53 R A P I D - T R I P L E S ( H A I B R I D G E ’ 2 5 ) Manual Knowledge Collection in RDF ‣Automatic annotation from documents fails to understand the knowledge implicitly defined by the document structure (e.g., reference to sub-procedures in other sections) ‣Automatic annotation should be fine-tuned for specific document structures, but often documents follow heterogeneous templates ‣OntoPawls (https://github.com/cefriel/ontopawls) offers a tool to annotate PDFs according to a given ontology ‣Manual effort required by users is high Rula, Anisa, et al. "Annotation and Extraction of Industrial Procedural Knowledge from Textual Documents." Proceedings of the 12th Knowledge Capture Conference 2023. 2023. 182 211 131 103 183 0 47 60 182 61 177 166 127 211 203 183 233 229 235 249 248 46 133 124 4 77 92 103 183 0 248 203 14 #1 #2 #3 #4 #5 8 134 53 103 183 0 150 189 71 182 211 131 210 228 278 #1 #2 #3 #4 #5 Primary 5 3 14 2 5 3 1 4 2 GRAPHS Sequence of use PERKS IDENTITY 46 133 124 224 83 72 00 00 00 Text Title 1 Title 2 TEXT Secondary Sequence of use BACKGROUND RAG 46 133 124 4 77 92 8 134 53 R A P I D - T R I P L E S ( H A I B R I D G E ’ 2 5 ) Semi-automatic Knowledge Collection in RDF •W1 Tacit Knowledge Collection •W2 Knowledge Completion from Unstructured Sources •W3 Enhance Automatic Knowledge Extraction Systems 182 211 131 103 183 0 47 60 182 61 177 166 127 211 203 183 233 229 235 249 248 46 133 124 4 77 92 103 183 0 248 203 14 #1 #2 #3 #4 #5 8 134 53 103 183 0 150 189 71 182 211 131 210 228 278 #1 #2 #3 #4 #5 Primary 5 3 14 2 5 3 1 4 2 GRAPHS Sequence of use PERKS IDENTITY 46 133 124 224 83 72 00 00 00 Text Title 1 Title 2 TEXT Secondary Sequence of use BACKGROUND RAG 46 133 124 4 77 92 8 134 53 R A P I D - T R I P L E S ( H A I B R I D G E ’ 2 5 ) Semi-automatic Knowledge Collection in RDF •W1 Tacit Knowledge Collection •W2 Knowledge Completion from Unstructured Sources •W3 Enhance Automatic Knowledge Extraction Systems 182 211 131 103 183 0 47 60 182 61 177 166 127 211 203 183 233 229 235 249 248 46 133 124 4 77 92 103 183 0 248 203 14 #1 #2 #3 #4 #5 8 134 53 103 183 0 150 189 71 182 211 131 210 228 178 #1 #2 #3 #4 #5 Primary 5 3 14 2 5 3 1 4 2 GRAPHS Sequence of use PERKS IDENTITY 46 133 124 224 83 72 00 00 00 Text Title 1 Title 2 TEXT Secondary Sequence of use BACKGROUND RAG 46 133 124 4 77 92 8 134 53 8 Why rapid-triples? ‣To guide the users in understanding which information is necessary (e.g., tacit knowledge) and how to model it ‣To hide the complexity of the RDF representation from the user ‣To enable usage of the tool seamlessly with other components and support the proposed workflows for semi-automatic knowledge collection R A P I D - T R I P L E S ( H A I B R I D G E ’ 2 5 ) 182 211 131 103 183 0 47 60 182 61 177 166 127 211 203 183 233 229 235 249 248 46 133 124 4 77 92 103 183 0 248 203 14 #1 #2 #3 #4 #5 8 134 53 103 183 0 150 189 71 182 211 131 210 228 278 #1 #2 #3 #4 #5 Primary 5 3 14 2 5 3 1 4 2 GRAPHS Sequence of use PERKS IDENTITY 46 133 124 224 83 72 00 00 00 Text Title 1 Title 2 TEXT Secondary Sequence of use BACKGROUND RAG 46 133 124 4 77 92 8 134 53 ‣The objective is to enable decoupled development and seamless integration of different components often not able to deal directly with the RDF representation ‣JSON-based format introduced relying on the semantics of the target ontology. Provides a common “framing” over the RDF graph representation. ‣A JSON Schema defines valid documents and enables automatic validation ‣A single KG construction process can be implemented relying on the JSON Schema (for simple cases a JSONLD context is sufficient) R A P I D - T R I P L E S ( H A I B R I D G E ’ 2 5 ) Intermediate Exchange Format 182 211 131 103 183 0 47 60 182 61 177 166 127 211 203 183 233 229 235 249 248 46 133 124 4 77 92 103 183 0 248 203 14 #1 #2 #3 #4 #5 8 134 53 103 183 0 150 189 71 182 211 131 210 228 178 #1 #2 #3 #4 #5 Primary 5 3 14 2 5 3 1 4 2 GRAPHS Sequence of use PERKS IDENTITY 46 133 124 224 83 72 00 00 00 Text Title 1 Title 2 TEXT Secondary Sequence of use BACKGROUND RAG 46 133 124 4 77 92 8 134 53 rapid-triples Preliminary Evaluation R A P I D - T R I P L E S ( H A I B R I D G E ’ 2 5 ) We run a preliminary qualitative evaluation of the rapid-triples tool with Beko users: ‣The users were able to generate a valid RDF representation of LOTO procedures using PKO for different machines in the factory ‣The domain experts appreciated the manual collection through the adaptive form, as it effectively guided them in documenting additional tacit knowledge while providing a better user experience with respect to paper or tabular-based approaches ‣Safety managers highlighted the higher quality of the generated procedures, and LOTO operators valued the possibility of having access to more details during the procedure execution. www.perks-project.eu www.perks-project.eu THANK YOU! This project has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement No 210906180 182 211 131 103 183 0 47 60 182 61 177 166 127 211 203 183 233 229 235 249 248 46 133 124 4 77 92 103 183 0 248 203 14 #1 #2 #3 #4 #5 8 134 53 103 183 0 150 189 71 182 211 131 210 228 178 #1 #2 #3 #4 #5 Primary 5 3 14 2 5 3 1 4 2 GRAPHS Sequence of use PERKS IDENTITY 46 133 124 224 83 72 00 00 00 Text Title 1 Title 2 TEXT Secondary Sequence of use BACKGROUND RAG 46 133 124 4 77 92 8 134 53 Mario Scrocca [email protected] rapid-triples •Adaptive form-based interface customizable via JSON Schema •Decoupling JSON output and lifting process to RDF supports the implementation of semi-automatic knowledge collection workflows •Use case considering procedural knowledge collection •Client implementation available on GitHub Future work: Investigate more complex workflows; perform broader and quantitative user evaluation; integrate declarative mapping rules and generative AI to further reduce manual intervention pko-rapid-triples •Check the demo with PKO •Access the rapid-triples template repository •Check how to customize the template