scieee AI-readable full text Open interactive document viewer

Ontology Coverage Analysis through Language Models

Abad-Navarro, Francisco; Martínez Costa, Catalina; Fernández Breis, Jesualdo Tomás

Abstract

Ontologies play a crucial role in supporting data and knowledge management in industry by promoting data standardization, interoperability, and facilitating knowledge sharing. However, the growing number of ontologies available in the repositories has made it challenging for developers to select the appropriate ontology for ensuring data interoperability in specific domains. To address this, we propose a method based on artificial intelligence to evaluate how well an ontology covers a particular domain. The results demonstrated that our method effectively identified the most appropriate ontology for each domain. The incorporation of language models in the method enabled it to overcome the limitations of traditional approaches, which often depend on exact string matches. Our method has proven to be an effective tool for assessing how well ontologies cover specific domains, thereby supporting the identification and selection of the most suitable ontologies for intelligent engineering applications.

Full text

Contents lists available at ScienceDirect Engineering Applications of Artificial Intelligence journal homepage: www.elsevier.com/locate/engappai Ontology Coverage Analysis through Language Models Francisco Abad-Navarro , Catalina Martínez-Costa , Jesualdo Tomás Fernández-Breis ∗ Departamento de Informática y Sistemas, Universidad de Murcia, CEIR Campus Mare Nostrum, IMIB-Arrixaca, 30100, Murcia, Spain A R T I C L E I N F O Dataset link:https://github.com/fanavarro/oca lm, https://doi.org/10.5281/zenodo.11220686 Keywords: Ontologies Language Models Domain coverage Knowledge representation FastText Bidirectional Encoder Representations from Transformers A B S T R A C T Ontologies play a crucial role in supporting data and knowledge management in industry by promoting data standardization, interoperability, and facilitating knowledge sharing. However, the growing number of ontologies available in the repositories has made it challenging for developers to select the appropriate ontology for ensuring data interoperability in specific domains. To address this, we propose a method based on artificial intelligence to evaluate how well an ontology covers a particular domain. The method begins by using a text corpus, which represents the domain knowledge. It first identifies noun phrases within the text, which are then matched with the classes in the ontology under evaluation. The alignment process uses a scoring function that combines a Levenshtein-based similarity metric together with FastText and Bidirectional Encoder Representations from Transformers (BERT). These models capture the contextual meaning of both the noun phrases and the ontology classes. The method was tested across four health engineering subdomains— genetics, food, medicine, and law—by selecting domain-specific ontologies and text corpora for each. The results demonstrated that our method effectively identified the most appropriate ontology for each domain. The incorporation of language models in the method enabled it to overcome the limitations of traditional approaches, which often depend on exact string matches. Our method has proven to be an effective tool for assessing how well ontologies cover specific domains, thereby supporting the identification and selection of the most suitable ontologies for intelligent engineering applications. 1. Introduction Ontologies are one of the pillars of Knowledge Graphs, providing the formal meaning of the entities and properties used for the description of domains’ knowledge (Elnagar et al., 2020; Hogan et al., 2021). An ontology is a formal, explicit specification of a shared conceptualization (Studer et al., 1998). In other words, ontologies define concepts of a particular domain by using a formal language that facilitates the data understanding by automated agents, and these definitions must be agreed and shared by the community. In this context, the ontology community promotes the open access to the knowledge and the collaborative development of ontologies, focusing on reusing what has already been done (Katsumi and Grüninger, 2016). Here, the W3C consortium recommended the Web Ontology Language version 2 (OWL2) as a standard language for defining ontologies, being adopted by most ontology developers. These facts are in line with the Findable, Accessible, Interoperable, Reusable (FAIR) principles, which put specific emphasis on enhancing the ability of machines to automatically find and use the data, in addition to supporting its reuse by individuals (Wilkinson et al., 2016). As a consequence, the use of ontologies and knowledge graphs has increased for a variety of purposes in health engineering, such ∗Corresponding author. E-mail addresses: [email protected] (F. Abad-Navarro), [email protected] (C. Martínez-Costa), [email protected] (J.T. Fernández-Breis). as data harmonization framework for secondary data reuse (AbadNavarro and Martínez-Costa, 2024), which is the analysis of existing data collected by others (Donnellan and Lucas, 2013); or event detection and classification in biomedicine (Alrefaie et al., 2024). In addition to this, practical applications can also be found in bridge health monitoring (Ndinga Okina et al., 2023) or software engineering (Bhushan et al., 2024). They have also been applied regularly for data integration (Mulero-Hernández et al., 2024; Gutiérrez et al., 2024). The development of intelligent applications in engineering fields requires a careful selection of the ontologies to be included. This is a difficult task given the number of ontologies available. For example, at the time of writing, the developers of health applications have more than one thousand semantic resources available in BioPortal (Whetzel et al., 2011). Recently, the repository of ontologies for industry (IndustryPortal) (Amdouni et al., 2023) was launched, and it already contains more than one hundred ontologies. Typically, ontology selection is based on how an ontology covers the domain in which it will be used. Here, we define domain coverage as the degree to which the entities and definitions included in the ontology https://doi.org/10.1016/j.engappai.2025.112671 Received 18 September 2024; Received in revised form 16 April 2025; Accepted 6 October 2025 Engineering Applications of Articial Intelligence 162 (2025) 112671 Available online 10 October 2025 0952-1976/© 2025 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY-NC license ( http://creativecommons.org/licenses/bync/4.0/ ). F. Abad-Navarro et al. represent the corresponding domain, the most appropriate ontology being the one that contains the concepts, relationships, and definitions necessary to represent the data of the given domain. However, the perfect ontology for a given purpose rarely exists, so it is necessary to identify suboptimal ontologies that describe part of the given domain in order to extend them by adding new ontology entities. The research question of our work is whether an embedding-based approach can help to measure ontology coverage in a domain given by text corpora. To this end, we propose a method, called OCALM (Ontology Coverage Analysis through Language Models), to measure the coverage of an ontology within a specific domain. This approach leverages a corpus of natural language documents, as text serves as the primary knowledge source in engineering domains. OCALM searches for the matches between the noun phrases detected in the text and the ontology classes. Our hypothesis is that ontologies and text corpora covering the same domain will have the highest scores for their (noun phrase, ontology class) matches. A noun phrase consists of a noun or pronoun, which is called the head, and any dependent words before or after the head (Peters, 2013). Each noun phrase is matched with an ontology class by maximizing a score function that takes into account their lexical similarity and their semantic similarity. On the one hand, lexical similarity accounts for the similarity of the words composing the noun phrase and the annotations of the ontology classes. Annotations are information attached to ontology entities, usually in natural language text, to improve their readability. On the other hand, the semantic similarity takes into account the nearest neighbors of the noun phrase and the class in their respective contexts in order to use them to obtain their similarity in a general language model. We tested OCALM across four domains, namely food, genetics, law and medicine, for which the method has been proven effective. Therefore, we believe that OCALM is a valuable tool for guiding the selection of the most appropriate ontology for a given domain. The main contributions of this work are: •A novel method for supporting the decision of which ontology to use based on the assessment of domain coverage •A text to ontology alignment method, which relies on a score function that considers lexical and semantic similarity and allows non exact matches •The first use of language models for computing semantic similarities with the purpose of coverage assessment The article is organized as follows: Section 2 describes the state of the art in domain coverage measurement; Section 3 explains the method proposed and the experimental approach; Section 4 depicts the results obtained for the experiment, including an ablation study and a comparison with the state of the art; Section 5 discusses our findings and proposes future work; and finally, the conclusions are presented in Section 6. 2. State of the art In ontology engineering, the relevant characteristics of ontologies have traditionally been measured by quantitative metrics, which provide objective measurements that help to profile ontology characteristics. The survey published by Wilson et al. (2021) provides a bibliography review on this aspect. Methods for evaluating the behavior of a set of metrics over a set of ontologies have also been proposed (Bernabé-Díaz et al., 2022). One of the perspectives from which an ontology can be evaluated is estimating the degree in which it provides a good coverage of certain domain knowledge. Domain coverage is defined by Wilson et al. (2023) as ‘the degree to which an ontology covers the axioms which have been specified (i.e., requirement specifications, standard ontologies, standard corpus) with respect to the domain knowledge that the ontology was developed to represent’. Traditionally, domain experts were responsible for assessing the domain coverage of an ontology through expert judgment. Nevertheless, the development of automated methods for assessing it would reduce the effort of domain experts and permit a systematic evaluation of the domain coverage of ontologies. In Zhu et al. (2017), the metric ‘vocabulary coverage’ is used to measure the content of ontologies. This metric is calculated by comparing the ontology being evaluated with another one that acts as gold standard. This is in line with the precision metric proposed in McDaniel et al. (2018), which also compares ontologies with a gold standard dictionary. Another example is the OQuaRE framework (Duque-Ramos et al., 2011), which proposed different metrics linked to the characteristics proposed in its quality model, such as the number of properties per class to measure the functional adequacy, and the HURON framework (Abad-Navarro et al., 2023), which proposed a set of metrics, such as names per class or descriptions per class, to capture the readability of ontologies. Most of the solutions proposed to automatically measure ontology coverage are based on recognizing ontology entities in free text. For instance, BioPortal (Whetzel et al., 2011) and IndustryPortal (Amdouni et al., 2023) are built over OntoPortal (Yang, 2009), which provides a recommendation system (Martínez-Romero et al., 2017) to select the most suitable ontology in their repository given a free text. This system takes into account ontology coverage by identifying ontology entities in the input text through a Named Entity Recognition (NER) (Jonquet et al., 2009), which uses the ontologies stored in the system as a dictionary. Nonetheless, most of the NER systems rely on exact matching between free text and ontology class annotations, which could lead to a poor recall. In addition to dictionary-based approaches, the development of artificial intelligence techniques applied to natural language, such as embedding algorithms (Mikolov et al., 2013a; Bojanowski et al., 2017a; Joulin et al., 2017; Kenton and Toutanova, 2019) or Large Language Models (LLM) (Achiam et al., 2023; Touvron et al., 2023; Jiang et al., 2023; Qiu et al., 2023), is leading to new paradigms in ontology engineering. In this line, there are several works that use LLMs for ontology learning and generation (Babaei Giglou et al., 2023; Saeedizade and Blomqvist, 2024; Behr et al., 2023; Val-Calvo et al., 2025), however, its application to ontology evaluation needs to be further investigated. For example, in Tsaneva et al. (2024) GPT-4 was used for verifying ontology restrictions. Moving into the domain coverage, Zaitoun et al. (2023) proposed a combination of NER and LLM methods trained over a corpus of documents, which were considered as an authoritative source of truth, and whose results are compared to the input ontology. In that work, the coverage is mainly based on the NER, but it is complemented by other metrics supported by the LLM, such as child or parent similarity. The present work deepens the use of LLM methods to create a metric to compute the domain coverage of a given ontology, where the domain is given as a natural language text corpus. In contrast to traditional NER systems, which are typically based on exact matches, the use of LLM techniques can enhance the measurement of domain coverage in ontologies by considering not only exact matches but also related concepts between free text and ontology entities. 3. Materials and methods This section explains the OCALM methodology for measuring the extent to which an ontology covers a domain. The input to OCALM is a corpus of natural language text, written in English, representing the domain knowledge; and the OWL ontology we want to evaluate. OCALM matches each noun phrase identified in the natural language text to an ontology class by using its annotations. A noun phrase is a word or group of words containing a noun and functioning in a sentence as subject, object, or prepositional object, such as ‘Sport drinks’ or ‘coronary artery diseases’; whereas a class annotation is information attached to an ontology class that is primarily intended Engineering Applications of Articial Intelligence 162 (2025) 112671 2 F. Abad-Navarro et al. Table 1 Example of noun phrases detected in the text ‘Diabetes mellitus is another endocrine system disease that affects many people in the United States’ and its normalized form. Noun phrase Normalized noun phrase Diabetes mellitus Diabetes mellitus Another endocrine system disease Endocrine system disease Many people People The United States United States to be informative for humans, including names, synonyms, or descriptions, among others. This matching is based on maximizing a score function that takes into account the lexical and the semantic similarity between the noun phrase and the ontology class. OCALM returns the best ontology class match for each noun phrase, together with the score achieved. Fig. 1 depicts the overview of the method, which consists of the following steps: 1. Normalization of the natural language input text. 2. Generation of a vector space model (𝑀𝑡) from the normalized natural language input text. 3. Normalization of the input ontology. 4. Generation of a vector space model (𝑀𝑜) from the normalized ontology. 5. Application of the score function to obtain the best ontology class match for each noun phrase identified in the normalized natural language input text. The following sections describe each step of the method in detail. 3.1. Text normalization Text normalization consists of replacing each noun phrase in the text by its canonical form. The purpose of this step is to identify different lexical forms of the same concept in order to transform them into the same lexical form. This simplifies further steps by avoiding dealing with different variations of noun phrases that are actually referring to the same concept. The first step of text normalization is to identify the noun phrases that appear in the text. This is achieved by applying a standard pipeline for Natural Language Processing (NLP) consisting of a tokenizer, a part of speech tagger, a dependency parser, a lemmatizer, and a name entity recognizer. In this work, we have used the NLP pipeline offered by SpaCy (Montani et al., 2023), an open source NLP library for Python, and the model ‘en_core_web_sm’, which contains the definition of the pipeline and its components. This model was released under an MIT license, and was trained using texts written in English found on the Web (Explosion, 2023). After applying this pipeline to natural language text, the noun phrases are annotated, thus becoming directly available. Once the noun phrases are identified within the input text, they are transformed into their canonical forms by applying the following rules: •Remove language dependent stop words, such as ‘another’, ‘otherwise’, ‘what’, etc. •Remove special characters, such as quotes or backslashes. •Replace carriage returns by blank spaces. •Replace plural nouns by its lemma. •Convert to lower case. Noun phrases are stored in a list together with their canonical form. Those referring to proper nouns or single numbers are removed from the list. Then, noun phrases in the list are replaced by their canonical forms in the input text, conforming the normalized text. Finally, the output of the text normalization step is the normalized text together with the list of the identified noun phrases, which will be used in further steps. For example, the input text ‘Diabetes mellitus is another endocrine system disease that affects many people in the United States’ results in the list of noun phrases depicted in Table 1. The normalized text is the result of replacing the noun phrases by their normalized form in the original text, resulting in ‘diabetes mellitus is endocrine system disease that affects people in united states’. 3.2. Vector space model creation for the natural language text corpus The vector space model creation for the natural language text takes as input the normalized text and the set of noun phrases identified in the text normalization step (see Section 3.1). The method generates a vector space model 𝑀𝑡 and returns an embedding for each noun phrase. For this purpose, we applied the fastText model (Bojanowski et al., 2017b) to the normalized natural language text with the following parameters: •Vector size = 100 (number of dimensions of the word vectors). •Window = 5 (the maximum distance between the current and predicted word within a sentence). •Min count = 1 (the model ignores all words with total frequency lower than this). •Negative = 5 (how many ‘noise words’ should be drawn). •Iter = 10 (number of epochs over the corpus). •Seed = 1 (seed for random number generator, for reproducibility). The fastText model (Bojanowski et al., 2017b) is derived from the skipgram model with negative sampling introduced in Mikolov et al. (2013b). The goal of the skipgram model is to learn a vectorial representation for each word 𝑤 belonging to a vocabulary of size 𝑊 by predicting words appearing in the context of a given word. For this, the model uses a scoring function to give a score to pairs of (word, context), where the context is a set of words surrounding the given word. Finally, the skipgram model computes the score function as the scalar product between word and context vectors. Nonetheless, this model provides a distinct vector representation for each word, ignoring the internal structure of words. Here, fastText proposed to represent each word as a bag of n-grams that also includes the word itself and to modify the score function to take into account the internal structure of words. In particular, given a dictionary of n-grams of size 𝐺 and a word 𝑤, ⊂{1,…, 𝐺} is the set of n-grams appearing in 𝑤; then, fastText associates a vector representation 𝐳𝑔 to each n-gram 𝑔. Finally, fastText represents a word by the sum of the vector representation of its n-grams, obtaining the scoring function defined in Eq. (1). 𝑠(𝑤, 𝑐) = ∑ 𝑔∈𝑤 𝐳𝑔𝐯𝑐(1) Then, we used the fastText model to assign an embedding to each word in the normalized text. It takes into account not only the words appearing in the text but also the N-grams conforming these words. This facilitates obtaining vectors for multi-word phrases, which is needed to detect the embeddings associated with our noun phrases. Finally, we filter the vector space to keep only the noun phrases detected in the text normalization step. Therefore, the output of this step is a vector space model 𝑀𝑡 containing an embedding for each noun phrase of the normalized text. 3.3. Ontology normalization The ontology normalization step takes an OWL ontology as input and returns its normalized version. This is similar to the text normalization, shown in Section 3.1, as it also consists of identifying and normalizing noun phrases. Here, as the input is an OWL ontology, normalization is applied to the text of each ontology class annotation, which include labels, synonyms, descriptions, or comments, among others, for each ontology class. Engineering Applications of Articial Intelligence 162 (2025) 112671 3 F. Abad-Navarro et al. Fig. 1. Overview of the main steps of OCALM. Both the natural language text corpus and the ontology are normalized in order to generate a vector space model for each one. The score function uses both vector spaces together with a general purpose language model, and the lexical forms of the noun phrases from the text and the annotations of ontology classes to get similarity. The score function is used to get the best ontology class match for each noun phrase detected in the text. Table 2 Example of noun phrases detected in the class GO:0019012 (virion) from Gene Ontology and their normalized form. Noun phrase Normalized noun phrase Wikipedia Wikipedia Virus Virus GO:0019012 Go:0019012 Complete virus particle Complete virus particle The complete fully infectious extracellular virus particle Complete fully infectious extracellular virus particle Noun phrases from ontology class annotations are obtained and normalized according to the already described rules (see Section 3.1). Similarly to the text normalization step, the identified noun phrases are also stored in a list. Then, the ontology is normalized by replacing the detected noun phrases in the class annotations by their normalized form. Finally, the output of this step is the normalized ontology together with the noun phrases list. For example, the class GO:0019012 (virion), from gene ontology, contained the annotations depicted in Fig. 2(a). Table 2 shows the noun phrases detected in the annotations of the class, together with their normalized forms. Then, the class is normalized by replacing the original noun phrases by their normalized ones, resulting in the class showed in Fig. 2(b). 3.4. Vector space model creation for the ontology The vector space model creation for the ontology takes the normalized ontology and the noun phrases identified in Section 3.3 as Engineering Applications of Articial Intelligence 162 (2025) 112671 4 F. Abad-Navarro et al. Fig. 2. Example of the normalization process of the class GO:0019012 (virion). input and returns a vector space model 𝑀𝑜 with a set of embeddings representing each noun phrase. First, the normalized ontology is used as input for the OWL2Vec* embedding algorithm (Chen et al., 2021), which generates embeddings of OWL ontologies. OWL2Vec* generates text documents from the input ontology to apply a text embedding algorithm to them. OWL2Vec* originally applies the Word2Vec algorithm (Mikolov et al., 2013a) to the generated text documents to obtain the ontology embeddings; however, we modified this behavior to use fastText instead, as we used it in the creation of the vector model space for the natural language text, as described in Section 3.2. We used the same fastText parameters as in Section 3.2, whereas the parameters used for OWL2Vec* are shown in Table 3. Since our method is applicable to any ontology, the parameters were selected to take into account most of the information stored in the ontology while preserving good performance. In particular, with the selected parameters, OWL2Vec* first transforms the input OWL ontology into an RDF graph, which is called the ‘projected ontology’. This projection process takes into account all the relationships stated in the ontology and simplifies some OWL constructs preventing the appearance of blank nodes derived from complex class expressions. The projected ontology also includes inverse relationships for rdf:type or rdfs:subClassOf to enable bidirectional walks between entities linked through these relationships. An OWL reasoner could be used to infer new RDF triples from the stated ones; however, this reasoning was disabled for performance reasons. Then, OWL2Vec* performs random walks to traverse the RDF graph containing the projected ontology, where the length of the random walks was set to 3 steps for performance reasons. During these walks, OWL2Vec* generates three text documents by taking the information of the ontology entities that are traversed. On the one hand, the structure document contains sentences resulting from printing the IRIs of the entities being traversed, aiming at capturing both the graph structure and the logical constructors of the ontology. On the other hand, the lexical document contains sentences resulting from printing the annotations (e.g. labels and synonyms) corresponding to the entities. Finally, the combined document contains sentences that mix entity IRIs and annotations with the aim of capturing the correlation between entity IRIs and the lexical information. Then, these three documents are merged and used as input for the fastText algorithm, previously commented in Section 3.2, to generate a vector space for the tokens contained in them. After OWL2Vec* has generated the vector space for the normalized ontology, we keep only the embeddings that refer to the noun phrases identified in the ontology normalization step (see Section 3.3). Thus, the outcome of this phase is a vector space model 𝑀𝑜 containing the noun phrases identified in the normalized ontology. 3.5. Score function A score function is used to measure the degree of relation between a noun phrase identified in the natural language text and an ontology class. The score function returns a matching score for two input strings, one for the noun phrase, and one for the ontology class. Thus, to obtain the matching score between them, the noun phrase is compared to the annotations of the ontology class, which may include labels or synonyms, selecting the one that provides the highest score. In this way, each noun phrase from the natural language text is compared to each ontology class, assigning the class that maximizes the score function to the noun phrase. The score function is the result of a weighted mean between the lexical (lexSim) and the semantic similarity (semSim) of the compared input strings, as described in Eq. (2), where 𝑎 and 𝑏 are the input strings, extracted from the natural language text and from the ontology, respectively; 𝛼 and 𝛽 are the weights given to the lexical and the semantic similarity, respectively; and 𝑀𝑡 and 𝑀𝑜 are the vector space models generated from the natural language text and the ontology, respectively (see Sections 3.2 and 3.4), whereas 𝑀𝑔 is a general vector space model, used for computing the semantic similarity. The details of these similarities are described in the next sections. 𝑠𝑐𝑜𝑟𝑒(𝑎, 𝑏) = 𝛼⋅𝑙𝑒𝑥𝑆𝑖𝑚(𝑎, 𝑏) + 𝛽⋅𝑠𝑒𝑚𝑆𝑖𝑚(𝑎, 𝑏, 𝑀𝑡, 𝑀𝑜, 𝑀𝑔) 𝛼+𝛽(2) 3.5.1. Lexical similarity The lexical similarity compares two text strings, 𝑎 and 𝑏, by considering only their lexical forms. For this, we compare the tokens of both strings, generating pairs of tokens maximizing their Levenshtein similarity (Lcvenshtcin, 1966). Changes in the order or the number of tokens between both penalize the score. Next, we explain how the lexical similarity is computed with examples. An overview of how the lexical similarity is calculated is shown in Fig. 3. First, strings 𝑎 and 𝑏 are divided into tokens by splitting them by using the blank space character as separator. Then, Levenshtein similarity between each token pair between 𝑎 and 𝑏 is calculated obtaining a matrix. For example, the comparison between ‘melitus diabetis’ and ‘diabetes mellitus’ results in the matrix depicted in Table 4. Table 5 shows another example between the strings ‘diabetes type I’ and ‘diabetes mellitus’. The second step is to match the tokens of 𝑎 to the tokens of 𝑏, obtaining token pairs that maximize the Lenveshtein similarity. This is a linear sum assigning problem, where tokens from the first string are Engineering Applications of Articial Intelligence 162 (2025) 112671 5 F. Abad-Navarro et al. Table 3 OWL2VEC parameters used for building the ontology vector space. Parameter Parameter explanation Value Ontology projection Use or not use the projected ontology Yes Projection only taxonomy Projection of only the taxonomy of the ontology without other relationships No Multiple labels Using or not multiple labels/synonyms for the literal/mixed sentences Yes Avoid owl construct Skip OWL constructs like rdfs:subclassof in the document No Axiom reasoner Reasoner to use for inferring axioms None Walker Algorithm to traverse the ontology Random Walk depth Number of hops when traversing the ontology 3 URI Doc Create a document with entities IRIs when traversing the ontology Yes Lit Doc Create a document with entities annotations when traversing the ontology Yes Mix Doc Create a document mixing entities IRIs and annotations when traversing the ontology Yes Mix Type The type for generating the mixture document - all or random All Fig. 3. Overview of the lexical similarity computation. The input strings to be compared are tokenized, and a Levenshtein similarity matrix is computed by generating the Levenshtein similarity of each token pair between the input strings. Tokens pairs that maximize the Levenshtein similarity between the input strings are obtained. This similarity is used, in addition to an order factor that penalizes changes in the order of the tokens between the input strings, to compute the lexical similarity. Table 4 Example of Levenshtein similarity matrix for the strings ‘melitus diabetis’ and ‘diabetes mellitus’. Diabetes Mellitus Melitus 0.40 0.93 Diabetis 0.88 0.38 Table 5 Example of Levenshtein similarity matrix for the strings ‘diabetes type I’ and ‘diabetes mellitus’. Diabetes Mellitus Diabetes 1 0.38 Type 0.17 0.33 I 0 0 assigned to the tokens of the second string maximizing the Levenshtein similarity, giving a token pair assignment matrix, where 1 means that the concerning tokens are selected as pair. Next, Table 6 shows the token assignment matrix obtained for the strings ‘melitus diabetis’ and ‘diabetes mellitus’, whose Leveshtein similarity matrix is described in Table 4, that gives the following token pairs and the corresponding maximized Levenshtein similarity: •Levenshtein similarity (melitus, mellitus) = 0.93 •Levenshtein similarity (diabetis, diabetes) = 0.88 •Maximized Levenshtein similarity = 0.93 + 0.88 = 1.81 For its part, the strings ‘diabetes type I’ and ‘diabetes mellitus’, whose Levenshtein similarity matrix is shown in Table 5, result in the following token pairs with the corresponding maximized Levenshtein Engineering Applications of Articial Intelligence 162 (2025) 112671 6 F. Abad-Navarro et al. Table 6 Example of token pair assignment matrix for the strings ‘melitus diabetis’ and ‘diabetes mellitus’. Diabetes Mellitus Melitus 0 1 Diabetis 1 0 Table 7 Example of token pair assignment matrix for the strings ‘diabetes type I’ and ‘diabetes mellitus’. Diabetes Mellitus Diabetes 1 0 Type 0 1 similarity, derived from the token assignment matrix depicted in Table 7. Note that here, the token ‘I’ remains unassigned because the compared strings have different lengths: •Levenshtein similarity (diabetes, diabetes) = 1 •Levenshtein similarity (type, mellitus) = 0.33 •Maximized Levenshtein similarity = 1 + 0.33 = 1.33 The third step is to calculate the order factor, which penalizes the scores of those strings that have a different order in their tokens. This factor is calculated according to Eq. (3). It returns a number ranged from 0 (maximum penalization) to 1 (no penalization). 𝑜𝑟𝑑𝑒𝑟𝐹 𝑎𝑐𝑡𝑜𝑟(𝑎, 𝑏) = 1 − 𝑛𝑢𝑚𝑏𝑒𝑟𝑂𝑓 𝑈𝑛𝑠𝑜𝑟𝑡𝑒𝑑𝑇 𝑜𝑘𝑒𝑛𝑠(𝑎, 𝑏) 𝑚𝑎𝑥(𝑙𝑒𝑛𝑔𝑡ℎ(𝑎), 𝑙𝑒𝑛𝑔𝑡ℎ(𝑏)) ⋅𝐾(3) where 𝐾 is a configurable parameter ranged from 0 to 1 denoting the weight of the order penalization, and 𝑛𝑢𝑚𝑏𝑒𝑟𝑂𝑓 𝑈𝑛𝑠𝑜𝑟𝑡𝑒𝑑𝑇 𝑜𝑘𝑒𝑛𝑠 is the number of order changes of consecutive tokens of the string 𝑎 with respect to the string 𝑏. The value of 𝑛𝑢𝑚𝑏𝑒𝑟𝑂𝑓𝑈𝑛𝑠𝑜𝑟𝑡𝑒𝑑𝑇 𝑜𝑘𝑒𝑛𝑠 is calculated from the token pair assignment matrix by counting how many consecutive tokens do not have a diagonal of value 1. For example, the token pair assignment matrix for the strings ‘melitus diabetis’ and ‘diabetes mellitus’ (Table 6) has only two consecutive tokens for each input string, so, we check that these two tokens do not form a diagonal of ones, thus having a value of 1 for the 𝑛𝑢𝑚𝑏𝑒𝑟𝑂𝑓 𝑈𝑛𝑠𝑜𝑟𝑡𝑒𝑑𝑇 𝑜𝑘𝑒𝑛𝑠 variable. Contrariwise, the token pair assignment matrix for the strings ‘diabetes type I’ and ‘diabetes mellitus’ (Table 7) has two consecutive tokens forming a diagonal matrix of ones, thus having a value of 0 for the 𝑛𝑢𝑚𝑏𝑒𝑟𝑂𝑓 𝑈𝑛𝑠𝑜𝑟𝑡𝑒𝑑𝑇 𝑜𝑘𝑒𝑛𝑠 variable. Thus, the order factor value for the strings ‘melitus diabetis’ and ‘diabetes mellitus’ is 1 − 1 𝑚𝑎𝑥(2,2) ⋅𝐾= 1 − 0.5𝐾, where the maximum penalization is given if 𝐾= 1 and no penalization is given if 𝑘= 0. On the other hand, the order factor value for the strings ‘diabetes type I’ and ‘diabetes mellitus’ is 1− 0 𝑚𝑎𝑥(3,2) ⋅𝐾= 1. In this case, the order of the tokens match, thus giving a value of 1 (no penalization) for the order factor with independence of the 𝐾 parameter. Finally, the lexical similarity is given by Eq. (4), where 𝑚𝑎𝑥𝐿𝑒𝑣𝑒𝑛 𝑠ℎ𝑡𝑒𝑖𝑛(𝑎, 𝑏) is the maximized Levenshtein similarity found for the strings 𝑎 and 𝑏. Therefore, the lexical similarity between ‘melitus diabetis’ and ‘diabetes mellitus’ is given by 1.81 𝑚𝑎𝑥(2,2) ⋅(1−0.5𝐾) = 0.905⋅(1−0.5𝐾). Note that, in this case, the lexical similarity depends on 𝐾, where higher values of 𝐾 impact negatively on the lexical similarity. On the other hand, the lexical similarity of ‘diabetes type I’ and ‘diabetes mellitus’ is given by 1.33 𝑚𝑎𝑥(3,2) ⋅1 = 0.443. 𝑙𝑒𝑥𝑆𝑖𝑚(𝑎, 𝑏) = 𝑚𝑎𝑥𝐿𝑒𝑣𝑒𝑛𝑠ℎ𝑡𝑒𝑖𝑛(𝑎, 𝑏) 𝑚𝑎𝑥(𝑙𝑒𝑛𝑔𝑡ℎ(𝑎), 𝑙𝑒𝑛𝑔𝑡ℎ(𝑏)) ⋅𝑜𝑟𝑑𝑒𝑟𝐹 𝑎𝑐𝑡𝑜𝑟(𝑎, 𝑏)(4) 3.5.2. Semantic similarity This method measures the similarity between noun phrases from the corpus and the annotations associated with the classes of the ontology. The input are the two strings to compare: string 𝑎, extracted from the natural language text, and string 𝑏, extracted from the ontology, along with their respective vector space models 𝑀𝑡 and 𝑀𝑜. The models 𝑀𝑡 and 𝑀𝑜, created as described in Sections 3.2 and 3.4, provide the semantic context of the input strings 𝑎 and 𝑏. Additionally, a general semantic context 𝑀𝑔 is used to compute the semantic similarity, which is given by Phrase-BERT (Wang et al., 2021a), a general BERT model (Kenton and Toutanova, 2019) focused on phrase embeddings. This general model serves for integrating the specific models 𝑀𝑡 and 𝑀𝑜, created for the natural language corpus and the ontology. The semantic similarity method returns a score in the range 0 (lowest similarity) to 1 (highest similarity). Calculating the semantic similarity requires to execute the following steps. First, we obtain the 10 nearest noun phrases (𝑛𝑝) for each input string by using the corresponding vector space model. For this purpose we apply the cosine distance. In particular, the neighbors of the string 𝑎, 𝑁𝑎, are calculated by using the model 𝑀𝑡, whereas the neighbors of 𝑏, 𝑁𝑏 are extracted from 𝑀𝑜. Fig. 4 depicts an example of the neighbors of ‘diabetes mellitus’ and ‘type 2 diabetes mellitus’, taken from the natural language text and from the ontology vector space models, respectively. For readability reasons, only the 5 nearest neighbors are shown. Second, the 10 nearest neighbors of each input string are extracted from the general Phrase-BERT model 𝑀𝑔, obtaining its vectors in this vector space. This results in two sets of vectors, 𝑉𝑔𝑎 and 𝑉𝑔𝑏, encoding the semantic representation of the neighbors of 𝑎 and 𝑏 in this general, multipurpose vector space model. This allows to obtain a representation of the input strings in a common vector space by using their nearest neighbors as context, thus avoiding polysemy related problems. In addition, this BERT model is specialized for phrases, which is consistent with our phrase-focused method. Fig. 5 shows the neighbors of ‘diabetes mellitus’ and ‘type 2 diabetes mellitus’ represented in the Phrase-BERT model. Third, each set of vectors, 𝑉𝑔𝑎 and 𝑉𝑔𝑏, previously generated, is summarized into a single point in the Phrase-BERT general vector space by applying the mean of all points in the set. As a result, we have two average points, 𝑎𝑝𝑎 (see Eq. (5)) and 𝑎𝑝𝑏 (see Eq. (6)), representing the input strings in the Phrase-BERT vector space model in terms of their neighbors, which were previously extracted from their particular vector space models. Fig. 6 shows these average points for ‘diabetes mellitus’ and ‘type 2 diabetes mellitus’, together with their neighbors. 𝑎𝑝𝑎=∑10 𝑖=1 𝑣𝑖 10 |𝑣𝑖𝜖𝑉𝑔𝑎 (5) 𝑎𝑝𝑏=∑10 𝑖=1 𝑣𝑖 10 |𝑣𝑖𝜖𝑉𝑔𝑏 (6) Finally, the semantic similarity, 𝑠𝑒𝑚𝑆𝑖𝑚, between the input strings 𝑎 and 𝑏 is given by the cosine similarity (𝑐𝑜𝑠𝑆𝑖𝑚) of the computed mean points, 𝑎𝑝𝑎 and 𝑎𝑝𝑏, as depicted in Eq. (7). Although the cosine similarity ranges from −1 to 1, where −1 is a perfect negative similarity, 1 is a perfect positive similarity, and 0 is no similarity, we can assume that only positive values are considered since negative similarities are not considered in this work. 𝑠𝑒𝑚𝑆𝑖𝑚(𝑎, 𝑏) = 𝑐𝑜𝑠𝑆𝑖𝑚(𝑎𝑝𝑎, 𝑎𝑝𝑏) = 𝑎𝑝𝑎⋅𝑎𝑝𝑎 |𝑎𝑝𝑎||𝑎𝑝𝑏|(7) 3.6. Coverage score The score function described in Section 3.5 is used to find the best ontology class match for each noun phrase detected in the natural language text. To achieve this, for each noun phrase 𝑎𝑖 appearing in the natural language text, we selected the noun phrase appearing in the ontology class annotations (restricted to labels and synonyms) 𝑏𝑗 maximizing the score function defined in Eq. (2). More formally, the pairs (𝑎𝑖, 𝑏𝑗) are found according to Eq. (8). 𝑝𝑎𝑖𝑟(𝑎𝑖, 𝑏𝑗) ∶ 𝑠𝑐𝑜𝑟𝑒(𝑎𝑖, 𝑏𝑗) = 𝑚𝑎𝑥(𝑠𝑐𝑜𝑟𝑒(𝑎𝑖, 𝑏𝑘)) (8) Engineering Applications of Articial Intelligence 162 (2025) 112671 7 F. Abad-Navarro et al. Fig. 4. Example of the nearest neighbors of ‘diabetes mellitus’ and ‘type 2 diabetes mellitus’, taken from the natural language text and the ontology class annotations, respectively. Only the 5 nearest neighbors are shown due to readability reasons. Fig. 5. Example of the nearest neighbors of ‘diabetes mellitus’ (blue dots) and ‘type 2 diabetes mellitus’ (red dots), taken from the natural language and the ontology vector models, respectively, represented in the Phrase-BERT general model. Only the 5 nearest neighbors are shown due to readability reasons. The pair (𝑎𝑖, 𝑏𝑗) means that 𝑏𝑗 is the best ontology class annotation for the noun phrase 𝑎𝑖. Furthermore, since the method tracks the source class of the selected ontology class annotations, we can easily translate the pairs (𝑎𝑖, 𝑏𝑗) into pairs of type (𝑎𝑖, 𝑐𝑗), where 𝑐𝑗 is the ontology class that contains the annotation 𝑏𝑗. Thus, the mean of the scores achieved by each pair (𝑎𝑖, 𝑐𝑗) is used as a global score for the coverage of the ontology in the domain given by the natural language text. 3.7. Implementation OCALM has been implemented in a Python 3.8 script, available at GitHub.1 It mainly uses the following libraries: •SpaCy (Montani et al., 2023), for natural language processing and noun phrase identification. •ScyPy (Virtanen et al., 2020), for solving linear sum assignment problems. 1https://github.com/fanavarro/ocalm. Fig. 6. Example of the average points computed in the general Phrase-BERT vector model from the neighbors of ‘diabetes mellitus’ (blue dots) and ‘type 2 diabetes mellitus’ (red dots) extracted from the natural language text and the ontology vector space models, respectively. Only the 5 nearest neighbors are shown due to readability reasons. •Gensim (Rehurek and Sojka, 2011), which provides the fastText implementation. •OWL2Vec* (Chen et al., 2021), for ontology embeddings. The original code OWL2Vec code was modified to use fastText instead of Word2Vec; this modification is included in our GitHub repository.1 •Owlready2 (Lamy, 2017), for managing ontologies. •SentenceTransformers (Reimers and Gurevych, 2019), which allows BERT models to be queried. 4. Results 4.1. Experimental setting We have applied OCALM to four different domains; namely, food, genetics, legal and medical domain. For each domain, we selected a domain ontology and a natural language text corpus covering the corresponding domain. Tables 8and 9 describe the natural language text corpora and the ontologies selected, respectively. We extracted Engineering Applications of Articial Intelligence 162 (2025) 112671 8 F. Abad-Navarro et al. Table 8 Description of the corpora used in the experiment. Domain Corpus Word count Description Food Amazon fine food reviewsa10,099 Reviews of fine foods from amazon. Genetic BioC-BioGRIDb (Islamaj Doğan et al., 2017) 18,610 Full text articles annotated for curation of protein–protein and genetic interactions. Legal Legal case reportsc (Galgani, 2010) 5801 Australian legal cases from the Federal Court of Australia (FCA). Medical i2b2 clinical recordsd13,885 Clinical records used as training datasets in the 2009 medication challenge at i2b2 workshop (Uzuner et al., 2010). a https://www.kaggle.com/snap/amazon-fine-food-reviews/version/2. b https://bioc.sourceforge.net/BioC-BioGRID.html. c https://archive.ics.uci.edu/ml/datasets/Legal+Case+Reports. d https://portal.dbmi.hms.harvard.edu/projects/download_dataset/?file_uuid=3e6f6a8e-7b22-4d7a-8ddd-3cc5e1ab8c08 login required. Table 9 Description of the ontologies used in the experiment. Domain Ontology Number of classes Class annotations Description Food FoodOn 29,906 203,221 Farm-to-fork ontology about food, that describes foods commonly known in cultures from around the world. Genetics GeneOntology 44,085 365,293 Provides a framework and set of concepts for describing the functions of gene products from all organisms. Legal LKIF 206 421 Part of the European project for Standardized Transparent Representations in order to Extend Legal Accessibility. Medical SNOMED CT 364,712 974,828 Clinical healthcare terminology, enabling consistent representation of clinical content in clinical information systems. subsets from the original natural language text corpora due to their size; the column ‘Word count’ indicates the size of the corpus used in the experiment. These files are available at Zenodo (Abad-Navarro, 2024), except for the SNOMED CT ontology and the clinical text corpus, due to licensing restrictions. We ran an experiment for every possible pair (ontology, text corpus) as input for our method. The configuration of the experiment was the following: •The lexical and the semantic similarity had the same weight in the score function; in particular 𝛼=𝛽= 1. •The parameter 𝐾 was set to 0.25, indicating the weight of the order penalization when computing the lexical similarity (see Section 3.5.1). The low value for 𝐾 is intended to prioritize exact string matches over matches where the tokens are in a different order; for example, the match (‘diabetes mellitus’, ‘diabetes mellitus’) will have a higher lexical similarity than (‘mellitus diabetes’, ‘diabetes mellitus’). Furthermore, the low 𝐾 value does not highly penalize matches like ‘slice of bread’ and ‘bread slice’, which will still have a high similarity. Finally, the distributions of the scores for each (noun phrase, ontology class) match provided by the method were compared to find the ontology that better represents each natural language text corpus. 4.2. Findings The output of the experiment described in Section 4.1 consisted of a tabular file for each (ontology, text corpus) pair, available at Zenodo (Abad-Navarro, 2024). Each tabular file contains a list of the noun phrases detected in the corresponding text, together with the best ontology class match found by our method for that noun phrase, and the corresponding score. These files have been summarized in Fig. 7, where the ontologies are compared between them for each text corpus by using the distribution of the scores of the (noun phrase, ontology class) pairs found by OCALM. The figure also shows the p-values returned by the Wilcoxon Rank Sum Test (Wilcoxon, 1947) of each comparison. The Wilcoxon Rank Sum Test is a non-parametric statistical test used to compare the medians of two independent populations. The figure shows that SNOMED CT obtained high scores for all of the natural language text corpora, which means that this ontology covers well the considered domains, namely, food, genetics, legal and medical domains. In addition to this, if we obviate SNOMED CT in the non-medical texts, OCALM found that the best scores are reached by the ontologies in the same domain as the corresponding natural language text. Thus, the highest scores were reached by the pairs (FoodOn, food text), (GeneOntology, genes text), (LKIF, legal text) and (SNOMED CT, medical text). Nevertheless, it is noteworthy that, for the legal text, the LKIF ontology did not stand out against the others, especially FoodOn, whose comparison did not reach statistical significance, assuming 𝛼= 0.05. Additionally, Fig. 8 shows a comparison between the text corpora for each ontology. In other words, this figure depicts which text corpus fits better with each ontology. In this case, each text corpus had the highest score when the corresponding ontology had the same domain, with no exception. Thus, the food text was better represented by FoodOn; the genetic text was better represented by GeneOntology; the legal text was better represented by LKIF; and the medical text was better represented by SNOMED CT. Finally, Table 10 shows the mean of the noun phrase-ontology class matches scores, which serves as coverage score of each ontology for each domain, together with the standard deviation (SD). 4.3. Comparison with the OntoPortal recommender The ontologies and text corpora used as input for OCALM, described in Section 4.1, were also used to test the OntoPortal Recommender system (Martínez-Romero et al., 2017). In particular, a virtual machine was deployed with OntoPortal virtual appliance version 3.2.2. This virtual appliance deploys the OntoPortal web platform together with its services, including the Annotator and the Recommender. Then, Engineering Applications of Articial Intelligence 162 (2025) 112671 9 F. Abad-Navarro et al. Tsaneva, S., Vasic, S., Sabou, M., 2024. LLM-driven ontology evaluation: Verifying ontology restrictions with ChatGPT. In: Data Quality Meets Machine Learning and Knowledge Graphs. In: CEUR Workshop Proceedings, RWTH, URL https://ceurws.org/Vol-3747/dqmlkg_paper3.pdf. Uzuner, Ö., Solti, I., Cadag, E., 2010. Extracting medication information from clinical text. J. Am. Med. Inform. Assoc. 17 (5), 514–518. http://dx.doi.org/10. 1136/jamia.2010.003947, arXiv:https://academic.oup.com/jamia/article-pdf/17/5/ 514/9713258/17-5-514.pdf. Val-Calvo, M., Egaña Aranguren, M., Mulero-Hernández, J., Almagro-Hernández, G., Deshmukh, P., Bernabé-Díaz, J.A., Espinoza-Arias, P., Sánchez-Fernández, J.L., Mueller, J., Fernández-Breis, J.T., 2025. OntoGenix: Leveraging large language models for enhanced ontology engineering from datasets. Inf. Process. Manage. 62 (3), 104042. http://dx.doi.org/10.1016/j.ipm.2024.104042, URL https://www. sciencedirect.com/science/article/pii/S0306457324004011. Virtanen, P., Gommers, R., Oliphant, T.E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S.J., Brett, M., Wilson, J., Millman, K.J., Mayorov, N., Nelson, A.R.J., Jones, E., Kern, R., Larson, E., Carey, C.J., Polat, İ., Feng, Y., Moore, E.W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E.A., Harris, C.R., Archibald, A.M., Ribeiro, A.H., Pedregosa, F., van Mulbregt, P., SciPy 1.0 Contributors, 2020. SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nature Methods 17, 261–272. http://dx.doi.org/10.1038/s41592-019-0686-2. Wang, S., Thompson, L., Iyyer, M., 2021a. Phrase-BERT: Improved phrase embeddings from BERT with an application to corpus exploration. In: Moens, M.-F., Huang, X., Specia, L., Yih, S.W.-t. (Eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, pp. 10837–10851. http://dx.doi.org/10.18653/v1/2021.emnlp-main.846, URL https:// aclanthology.org/2021.emnlp-main.846. Wang, J., Yi, X., Guo, R., Jin, H., Xu, P., Li, S., Wang, X., Guo, X., Li, C., Xu, X., et al., 2021b. Milvus: A Purpose-Built vector data management system. In: Proceedings of the 2021 International Conference on Management of Data. pp. 2614–2627. Whetzel, P.L., Noy, N.F., Shah, N.H., Alexander, P.R., Nyulas, C., Tudorache, T., Musen, M.A., 2011. BioPortal: enhanced functionality via new Web services from the National Center for Biomedical Ontology to access and use ontologies in software applications. Nucleic Acids Res. 39 (suppl_2), W541–W545. Wilcoxon, F., 1947. Probability tables for individual comparisons by ranking methods. Biometrics 3 (3), 119–122. Wilkinson, M.D., Dumontier, M., Aalbersberg, I.J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., da Silva Santos, L.B., Bourne, P.E., et al., 2016. The FAIR guiding principles for scientific data management and stewardship. Sci. Data 3 (1), 1–9. Wilson, R., Goonetillake, J.S., Indika, W., Ginige, A., 2021. Analysis of ontology quality dimensions, criteria and metrics. In: International Conference on Computational Science and its Applications. Springer, pp. 320–337. Wilson, R., Goonetillake, J., Indika, W., Ginige, A., 2023. A conceptual model for ontology quality assessment. Semant. Web 14, 1051–1097. http://dx.doi.org/10. 3233/SW-233393, 6. Yang, S.-Y., 2009. OntoPortal: An ontology-supported portal architecture with linguistically enhanced and focused crawler technologies. Expert Syst. Appl. 36 (6), 10148–10157. Zaitoun, A., Sagi, T., Hose, K., 2023. Automated ontology evaluation: Evaluating coverage and correctness using a Domain Corpus. In: Companion Proceedings of the ACM Web Conference 2023. pp. 1127–1137. Zhu, H., Liu, D., Bayley, I., Aldea, A., Yang, Y., Chen, Y., 2017. Quality model and metrics of ontology for semantic descriptions of web services. Tsinghua Sci. Technol. 22 (3), 254–272. Engineering Applications of Articial Intelligence 162 (2025) 112671 16