Knowledge Representation and Ontologies in the Era of Large Language Models Tarcisio Mendes de Farias Knowledge Representation Unit
[email protected]
SIB in brief A national network of about 900 scientists A non-profit and independent organization 190 employees in 4 locations The Swiss Node of ELIXIR, the European life science infrastructure
Knowledge Representation in AI
To begin with, a definition…
A brief history of bioinformatics databases… Gene Expression Orthology Protein Interaction FAIR data principles Findable, Accessible, Interoperable, Reusable
Importance of identifiers
What is a Large Language Model? A (deep learning) model trained to predict the next word in a sentence How is it trained? •Self-supervised learning on HUGE amounts of text •“Fill in the blanks” learning Source: https://amitness.com 7
Generalization ●A (deep learning) model trained to predict the next word in a sentence ○Token = words / amino-acids / genes / …. ●What is a Language? ○A vocabulary + ○Sequences of tokens that represent information in that language token sequence 8
Let’s ask an LLM about TIME FOR COFFEE!
How do we extract information from a Knowledge Graph? What are the human genes involved in lung cancer with an ortholog expressed in the mouse? SELECT ?gene ?orthologous_protein2 WHERE { SELECT * { SERVICE <http://sparql.uniprot.org/sparql> { SELECT ?protein1 WHERE { ?protein1 a up:Protein; up:organism/up:scientificName 'Homo sapiens' ; up:annotation ?annotation . ?annotation rdfs:comment ?annotation_text. ?annotation a up:Disease_Annotation . FILTER CONTAINS (?annotation_text, ”lung cancer") } } SERVICE <https://sparql.omabrowser.org/sparql/> { SELECT ?orthologous_protein2 ?protein1 ?gene WHERE { ?protein_OMA a orth:Protein . ?orthologous_protein2 a orth:Protein . ?cluster a orth:OrthologsCluster . ?cluster orth:hasHomologousMember ?node1 . ?cluster orth:hasHomologousMember ?node2 . ?node2 orth:hasHomologousMember* […….] FILTER(?node1 != ?node2) } } SERVICE <https://bgee.org/sparql/> { ?gene genex:isExpressedIn ?anatEntity . ?anatEntity rdfs:label ‘lung' . ?gene orth:organism ?org . ?org obo:RO_0002162 taxon:10090 .}
LLMs as interfaces for scientific knowledge graph exploration: a fine-tuning approach
Fine-tuning LLMs for SPARQL generation •Joint work with the Data Knowledge Organisation Unit •Dr. Norio Kobayashi, Dr. Julio Rangel Reyes •RIKEN, Japan •Challenge: Very little training data •Databases often propose a few examples of queries online •Some queries are hard to connect back to the question •Solution: Augment and semantically enrich existing dataset of examples automatically! In which taxa is the insulin protein present? SWAT4HCLS 2024, full paper available at https://ceur-ws.org/Vol-3890/paper-4.pdf
Fine-tuning an Open LLM for SPARQL generation
Evaluation Experiments’ setup Hugging Face SFTTrainer with 2000 steps Nvidia A100 40GB GPUs OpenLLaMA_7b_v2 (7 billion parameters) Temperature parameter equal to zero Knowledge base (KB): Bgee (~7 billion triples) We rely on four different metrics designed for the evaluation of machine translation output : BLEU, SP-BLEU, METEOR and ROUGE.
Four main experiments and preliminary results Zero-shot evaluation against: Wikidata and Bgee question-SPARQL query sets All metrics including F1-score: either equal or approximately equal to zero. 3rd experiment 2nd experiment
Discussions Systematically augmenting a representative question-to-SPARQL query set over a scientific KG significantly contributes to improving the performance of the OpenLLaMA model for the SPARQL query generation task. Rewriting the SPARQL query to provide more context through comments and meaningful variable names considerably improves OpenLLaMA. Knowledge transfer might deteriorate the LLM performance for the SPARQL query generation over a domain-specific knowledge base.
LLMs as interfaces for scientific knowledge graph exploration: a prompt-tuning with RAG approach
Problem We need a minimal structured way to describe the SPARQL endpoints’ contents to facilitate LLM-based SPARQL query generation. For large and complex knowledge graphs finding the right context is not trivial. Writing SPARQL queries is hard and time-consuming. LLMs are great at it, but they need context!
Generic and reusable methodology that works for most endpoints. •Automatically generate a description of used classes and predicates e.g., Protein isEncodedBy Gene •Provide example queries with their corresponding question in a standard format Make them accessible via the SPARQL endpoint The system will be able to retrieve, index, and use them automatically Properly describing an endpoint with some metadata LLM-based SPARQL Query Generation from Natural Language over Federated Knowledge Graphs, V. Emonet et al, ISWC 2024