scieee AI-readable full text Open interactive document viewer

A Comprehensive Review of Knowledge Graph Integration in Large Language Models for Trust: Challenges and Future Directions

Touameur, Ouissem; Harrag, Fouzi

Abstract

Large Language Models (LLMs) have transformed the landscape of natural language processing, enabling significant application advancements. However, LLMs’ inherent complexity and opacity raise critical concerns regarding their trustworthiness, including bias, misinformation, and lack of interpretability. This paper presents a comprehensive state-of-the-art review on the integration of knowledge graphs (KGs) to enhance trust in LLMs. We explore how KGs can serve as structured frameworks for contextualizing the knowledge embedded in LLMs, providing a means to verify and validate their outputs against reliable sources of information. The review highlights various methodologies that leverage KGs to address trust-related challenges in LLMs, including mechanisms for improving factual consistency, reducing biases, and enhancing interpretability. Additionally, we analyze recent advancements and empirical studies that demonstrate the efficacy of knowledge graph integration in fostering transparency and reliability in LLM applications. By synthesizing the current landscape of research, this paper aims to identify future directions and key challenges in developing trust-aware systems that utilize the synergistic potential of LLMs and knowledge graphs.

Full text

A Comprehensive Review of Knowledge Graph Integration in Large Language Models for Trust: Challenges and Future Directions Touameur Ouissem1and Harrag Fouzi2 1Department of Computer Science, University of Farhat Abbas Setif, [email protected] 2Department of Computer Science, University of Farhat Abbas Setif, [email protected] Abstract Large Language Models (LLMs) have transformed the landscape of natural language processing, enabling significant application advancements. However, LLMs’ inherent complexity and opacity raise critical concerns regarding their trustworthiness, including bias, misinformation, and lack of interpretability. This paper presents a comprehensive state-of-the-art review on the integration of knowledge graphs (KGs) to enhance trust in LLMs. We explore how KGs can serve as structured frameworks for contextualizing the knowledge embedded in LLMs, providing a means to verify and validate their outputs against reliable sources of information. The review highlights various methodologies that leverage KGs to address trust-related challenges in LLMs, including mechanisms for improving factual consistency, reducing biases, and enhancing interpretability. Additionally, we analyze recent advancements and empirical studies that demonstrate the efficacy of knowledge graph integration in fostering transparency and reliability in LLM applications. By synthesizing the current landscape of research, this paper aims to identify future directions and key challenges in developing trust-aware systems that utilize the synergistic potential of LLMs and knowledge graphs. Keywords: LLMs, Trustworthiness, Knowledge graph, NLP, reliability. 1 Introduction The rapid advancement of Large Language Models (LLMs), such as OpenAI’s GPT-3 and Google’s BERT, has revolutionized the field of natural language processing (NLP) by enabling sophisticated text generation, comprehension, and interaction capabilities [3,6]. Despite their impressive performance, significant concerns regarding the trustworthiness of these models have emerged. Issues such as inherent biases, the propensity to generate misleading or factually incorrect information, and the lack of interpretability challenge their deployment in sensitive applications [2,9]. Trust is essential in AI systems, particularly in applications involving critical decision-making, such as healthcare, finance, and law. Users need to rely on the outputs of these models, necessitating a framework that can ensure reliability and transparency. Knowledge graphs (KGs), which represent knowledge in a structured format through entities and their relationships, offer a promising solution for enhancing the trustworthiness of LLMs. KGs can provide contextual information, facilitate the verification of facts, and support the interpretation of model outputs, thereby addressing some of the primary trust-related concerns associated with LLMs [16]. Recent research has explored various ways to integrate KGs with LLMs, highlighting their potential to mitigate issues such as misinformation and bias by providing a robust knowledge base against which LLM outputs can be validated [22]. For instance, KGs can serve as sources of grounding information, enabling LLMs to generate more contextually accurate responses [25]. Furthermore, KGs can enhance model interpretability by elucidating the reasoning behind generated outputs and demonstrating the relationships between concepts [8]. This paper aims to provide a comprehensive review of the current state of research on trust in LLMs, focusing specifically on the integration of knowledge graphs. We will examine the methodologies employed, their effectiveness in enhancing trust, and the challenges that remain in this evolving field. The paper is structured as follows: first, we explore the concept of trust in LLMs, followed by an overview of approaches using knowledge graphs within LLMs. Next, we review related work, present a discussion on key insights, and outline challenges and future directions. Finally, we conclude with a summary of the findings. 139 2 Understanding Trust in LLMs As large language models (LLMs) such as GPT, BERT, and T5 become increasingly embedded in applications affecting daily life, fostering user trust in these models has become crucial. Trust in LLMs rests on several key factors—transparency, reliability, accountability, and interpretability. Each of these elements contributes uniquely to how users perceive, rely on, and interact with LLMs. 2.1 Transparency Transparency refers to the model’s ability to reveal its internal workings, data sources, and decisionmaking processes. In LLMs, transparency encompasses both the interpretability of the model architecture and the visibility of the training data. Transparent LLMs allow users and developers to understand the source of the information, the decisions the model makes, and its reasoning process. For example, the work by Piktus et al. [17] on explaining machine learning classifiers provides insight into how interpretability methods can make black-box models more transparent by offering understandable explanations for predictions. Efforts toward transparency often involve open-sourcing datasets and providing documentation, such as Google’s Model Cards [14], which standardize model reporting to include information about intended use cases, biases, and limitations. 2.2 Reliability Reliability reflects the consistency and accuracy of an LLM’s performance across various tasks and contexts. Users are more likely to trust models that yield accurate and reproducible results, especially when deployed in high-stakes areas such as healthcare or legal domains. For instance, studies on adversarial robustness by Jia & Liang [21] illustrate that LLMs may be vulnerable to small changes in input data, which can lead to unreliable or erroneous outputs. Addressing these issues, some approaches focus on adversarial training and robustness testing to ensure that LLMs perform reliably under diverse conditions [23]. Reliable models thus bolster user trust by providing stable and dependable outcomes. 2.3 Accountability Accountability in LLMs relates to the capacity for tracing and attributing the model’s outputs to its developers or the data it was trained on. Given the wide-reaching influence of LLM-generated content, accountability becomes essential for maintaining public trust and mitigating risks associated with biased or harmful outputs. Studies on responsible AI, such as the work by Binns [20], emphasize the need for accountability mechanisms to ensure that LLMs can be held answerable for their outputs, particularly when they affect individuals or communities. Additionally, frameworks like ”explainable artificial intelligence” (XAI) aim to create models that produce both understandable and accountable outputs, contributing to responsible and trustable AI [4]. 2.4 Interpretability Interpretability is crucial for understanding and trusting LLMs, as it determines how easily a human can make sense of a model’s predictions. In natural language processing, interpretability approaches such as attention visualization [7] provide insights into which parts of the input data an LLM considers most relevant for making predictions. Techniques like Local Interpretable Model-agnostic Explanations (LIME) [27] further support interpretability by explaining individual predictions and model behavior. Interpretable models empower users to verify the model’s reasoning and assess its appropriateness in specific contexts, thus enhancing trust. These four factors—transparency, reliability, accountability, and interpretability—are interdependent and collectively contribute to building and maintaining user trust in LLMs. Ongoing research in each area aims to advance these dimensions of trustworthiness, bringing us closer to AI systems that users feel confident in adopting and integrating into decision-making processes. 140 3 Trust in LLMs with Knowledge Graph 3.1 Definition of Knowledge Graphs Knowledge Graphs (KGs) are structured representations of real-world entities, their attributes, and the relationships between them. Typically organized in a graph format, KGs enable machines to model and reason about domain-specific knowledge by connecting data points through edges representing meaningful relations [5]. They consist of nodes (entities) and edges (relationships) that create a network of interconnected information, thereby facilitating structured queries and efficient knowledge retrieval [18]. KGs are increasingly used with large language models (LLMs) to provide a structured and interpretable backbone for enhancing reasoning and factual correctness. 3.2 Enhancing Trust with Knowledge Graphs Integrating KGs with LLMs can significantly improve transparency and interpretability by providing factual grounding and context for LLM responses. KGs serve as external sources of structured information that can back model predictions, enhancing user trust. By linking facts to entities in a KG, LLMs can provide evidence for their responses, making the output more transparent and understandable for users [26]. For instance, grounding an LLM’s responses with Wikipedia-derived KGs allows for entity resolution, where references in the model’s output can directly link back to a KG entity, thus verifying the source of the information [11]. This grounding ensures that users can trace the origin of certain statements, thus increasing the system’s transparency. 3.3 Examples and Techniques To enhance trust in large language models (LLMs) using knowledge graphs (KGs), several techniques and examples can be employed: 1. KG-Enhanced LLM Interpretability: By integrating KGs into the inference process of LLMs, researchers can improve the interpretability of the model’s outputs. For instance, when an LLM generates a response, the relevant facts from the KG can be highlighted to show the basis for the generated information. This transparency helps users understand how the model arrived at its conclusions, thereby increasing trust. 2. Factual Verification: Techniques such as LLM-facteval can be used to automatically generate probing questions from KGs. These questions can then be used to evaluate the factual knowledge stored in LLMs. By systematically assessing the accuracy of the information provided by LLMs against a trusted KG, users can gain confidence in the model’s reliability. 3. Grounding Responses in KGs: Approaches like KagNet and QA-GNN ground the results generated by LLMs at each reasoning step using KGs. This means that the reasoning process is made explicit by linking the generated outputs to specific entities and relationships in the KG. Such grounding provides a clear rationale for the model’s responses, enhancing user trust. 4. Knowledge Graph-Based Probing: Tools like BioLAMA and MedLAMA utilize domainspecific KGs to probe LLMs for factual knowledge in specialized fields, such as medicine. By evaluating the model’s performance against a trusted medical KG, these techniques can help ensure that the LLM provides accurate and reliable information in critical applications. 5. Instruction-Tuning with KGs: Integrating KGs into the training objectives of LLMs can help improve their factual accuracy. For example, instruction-tuning methods can be employed where LLMs are trained to generate responses that align with the structured knowledge in KGs. This technique can help reduce the occurrence of hallucinations and improve the overall trustworthiness of the model [16]. 4 Related Works The integration of Large Language Models (LLMs) with Knowledge Graphs (KGs) has become a significant area of research, focusing on enhancing the capabilities of LLMs in terms of factual accuracy, 141 reasoning, and trustworthiness. This section reviews notable works in this domain and proposes a taxonomy based on their contributions. We propose a taxonomy to categorize integration approaches of LLMs and KGs, comprising three main dimensions: Type of Integration, Focus Area, and Application Domain. This structured framework helps in understanding the diverse methodologies within LLM and KG integration. 4.1 Performance Evaluation of LLMs Yang et al. [24] discuss the integration of knowledge graphs (KGs) with large language models (LLMs) to improve their ability to recall and utilize factual knowledge. The authors highlight that while LLMs like ChatGPT exhibit impressive conversational abilities, they struggle with generating knowledge-grounded content due to limitations in factual recall. To address this, the paper reviews existing methods for enhancing pre-trained language models (PLMs) with KGs and proposes the development of knowledge graph-enhanced large language models (KGLLMs). Hou et al. [10] examine the limitations of LLMs in biomedical contexts, where accuracy is crucial. The methodology involved conducting experiments where ChatGPT answered questions from the ”Alternative Medicine” sub-category of Yahoo! Answers, while BKGs were queried for relevant knowledge records. Additionally, a prediction scenario was created to evaluate the models’ abilities to suggest potential drug and dietary supplement repurposing candidates for Alzheimer’s Disease (AD). The results indicated that while ChatGPT (especially GPT-4) outperformed earlier versions and provided existing information effectively, BKGs demonstrated higher reliability and accuracy. ChatGPT struggled with novel discoveries and reasoning, particularly in establishing structured links between entities. 4.2 Trustworthiness and Credibility in LLM Outputs Dr. Carlo Lipizzi [13] presents a novel approach to evaluating the trustworthiness of Large Language Models (LLMs) by integrating knowledge graphs, RDF triplets, and a human-in-the-loop system. It emphasizes the importance of accurately representing domain-specific knowledge through knowledge graphs and involves subject matter experts (SMEs) to validate this representation and assess the compatibility of LLM outputs with established knowledge. By focusing on quantitative measures of trustworthiness, the proposed system aims to address the growing concerns regarding the reliability of LLMs, particularly in critical applications. The paper highlights the innovative nature of this approach while acknowledging challenges such as subjectivity, scalability, and the complexity of knowledge representation. Zhang et al. [28] investigate how KGs can enhance the factual recall of LLMs, addressing challenges in integrating various data representations to improve output accuracy. 4.3 Type of Integration The integration strategies can be categorized based on their timing and methodology: •Before-Training Enhancement: Zafar et al. [26] present a novel architecture that integrates large language models (LLMs), knowledge graphs (KGs), and role-based access control (RBAC) to enhance the capabilities of conversational AI systems. It highlights the importance of combining the linguistic proficiency of LLMs with the structured knowledge representation of KGs to address challenges such as explainability, data privacy, and contextual accuracy. The architecture aims to foster user trust and ensure the ethical use of AI technologies. Additionally, the paper introduces LLMXplorer, a comprehensive tool for evaluating various LLMs, which contributes to transparency and informed decision-making in the deployment of conversational AI. •Post-Training Enhancement: Li et al. [12] introduce XTRUST, the first comprehensive benchmark designed to evaluate the multilingual trustworthiness of large language models (LLMs). The authors highlight the remarkable capabilities of LLMs in various natural language processing (NLP) tasks and emphasize the growing concern regarding their trustworthiness, especially in sensitive fields such as healthcare and finance. XTRUST encompasses a wide range of topics, including illegal activities, hallucination, out-of-distribution robustness, mental and physical health, toxicity, fairness, misinformation, privacy, and machine ethics, across ten different languages. The paper presents an empirical evaluation of five widely used LLMs, revealing that many struggle with lowresource languages like Arabic and Russian, indicating significant room for improvement in their multilingual trustworthiness. 142 Alghamdi et al. [1] develop AraTrust, a trustworthiness benchmark for Arabic LLMs. The paper introduces AraTrust, the first comprehensive benchmark designed to assess the trustworthiness of Large Language Models (LLMs) specifically for the Arabic language. It comprises 516 humanwritten multiple-choice questions that cover eight critical categories of trustworthiness: truthfulness, ethics, physical health, mental health, unfairness, illegal activities, privacy, and offensive language. The authors highlight the inadequacies of existing English-centric benchmarks, which fail to address the unique cultural and contextual factors relevant to Arabic users. By providing a culturally aligned and automated assessment framework, AraTrust aims to enhance the safety and reliability of Arabic LLMs and promote further research in this area. The findings indicate that while proprietary models like GPT-4 perform well, many open-source models struggle to meet the benchmark’s standards, underscoring the need for improved trustworthiness in Arabic LLMs. •Real-time Enhancement: Large language models (LLMs) have achieved impressive results across various natural language processing tasks. However, once deployed, LLMs interact with users who possess personalized factual knowledge, which is reflected in their interactions. To enhance the user experience, it is crucial to implement real-time model personalization that allows LLMs to adapt user-specific knowledge based on feedback received during these interactions. Current methods primarily rely on back-propagation to fine-tune model parameters, leading to significant computational and memory overhead. Additionally, these methods often lack interpretability, which can negatively affect model performance as users accumulate personalized knowledge over time. To tackle these challenges, we introduce Knowledge Graph Tuning (KGT), a novel approach that utilizes knowledge graphs (KGs) to personalize LLMs. KGT extracts personalized factual knowledge triples from user queries and feedback, optimizing the KGs without altering the LLM parameters. This method enhances computational and memory efficiency by circumventing backpropagation while ensuring interpretability by making KG adjustments understandable to humans. Experiments with state-of-the-art LLMs, including GPT-2, Llama2, and Llama3, demonstrate that KGT significantly enhances personalization performance while reducing latency and GPU memory usage. In conclusion, KGT presents a promising solution for effective, efficient, and interpretable real-time LLM personalization during user interactions [19]. 4.4 Focus Area The focus of these studies can also be categorized based on their objectives: •Factual Recall: Hou et al. [10] reveal strengths and weaknesses in information retrieval for LLMs. •Trustworthiness Assessment: Lipizzi [13] emphasizes the importance of validation mechanisms for LLM outputs. •Model Personalization: Sun et al. [19] introduces KGT to enhance real-time model personalization using KGs. 4.5 Application Domain The studies vary in their application domains: •General Knowledge: Zafar et al. [26] and Kommineni et al. explore integration methodologies to enhance conversational AI systems. •Biomedical Research: Hou et al. [10] address factual recall challenges in biomedical contexts. •Domain-Specific Applications: Li et al. [12] and Alghamdi et al. [1] focus on trustworthiness in healthcare and Arabic language contexts. Future research should develop efficient integration methodologies that enhance factual grounding while considering computational efficiency. 143 Table 1: Summary of Related Works on LLMs and KGs Integration Reference Type of Integration Focus Area Key Contributions Yang et al. [24] BeforeTraining Factual Recall KGenhanced LLMs for improved factual recall Hou et al. [10] PostTraining Factual Recall Comparison of ChatGPT and BKGs in biomedical contexts Zafar et al. [26] BeforeTraining TrustworthinessLLMs and KGs integration for conversational AI Li et al. [12] PostTraining TrustworthinessXTRUST benchmark for multilingual LLMs Alghamdi et al. [1] PostTraining TrustworthinessAraTrust benchmark for Arabic LLMs Lipizzi et al. [13] Real-time TrustworthinessHuman-inthe-loop for trust assessment in LLMs Sun et al. [19] Real-time Model Personalization KGT for personalized LLMs using KGs 5 Discussion The integration of Large Language Models (LLMs) with Knowledge Graphs (KGs) represents a transformative approach to enhancing the capabilities of AI systems, particularly in domains requiring high accuracy and trust. The reviewed methods showcase significant advancements in the factual recall and reasoning abilities of LLMs, highlighting how structured knowledge representations can mitigate the inherent limitations of these models. For instance, studies by Yang et al. [24] and Hou et al. [10] demonstrate that KGs not only improve information retrieval but also facilitate more contextually relevant outputs, particularly in specialized fields such as biomedical research. Zafar et al. [26] emphasize the comprehensive integration of LLMs and KGs, which enhances explainability and data privacy through robust access control measures, while also allowing for iterative learning and practical applications in real-world scenarios. However, the architecture’s generalizability remains limited as it has predominantly been tested within specific contexts, like media and journalism, raising questions about its performance across diverse industries. Similarly, Li et al. [12] present the XTRUST benchmark, addressing the critical gap in multilingual LLM evaluations and underscoring the importance of trustworthiness in sensitive domains such as healthcare and finance. Yet, their evaluation of only five widely used models and a limited scope of languages may restrict the findings’ applicability. Alghamdi et al. [1] highlights the structured representation of knowledge in KGs, enhancing contextual understanding and facilitating reasoning, but they also point out challenges related to data quality, 144 scalability, and user acceptance, which can undermine trust. Furthermore, approaches focusing on trustworthiness, such as those proposed by Lipizzi et al. [13], underscore the necessity of integrating human expertise and domain-specific knowledge to assess and enhance the reliability of LLMs. Despite these strengths, challenges persist, including knowledge noise within KGs, which can lead to inaccuracies in model outputs, as noted by Yang et al. This highlights the urgent need for robust filtering mechanisms and dynamic updating processes to ensure that KGs remain relevant and accurate. Moreover, the computational complexity associated with integrating KGs with LLM architectures raises concerns about scalability and real-time application, necessitating more efficient methodologies. Addressing the subjectivity involved in trust assessments is equally critical, as inconsistent evaluations may hinder the applicability of these frameworks across diverse contexts. Therefore, future research should prioritize the development of standardized metrics for trust evaluation, optimized integration algorithms, and cross-domain applications of KG-LLM frameworks. By tackling these limitations, researchers can unlock the full potential of integrating KGs with LLMs, paving the way for more reliable, transparent, and contextually aware AI systems. 6 Challenges and future directions 6.1 Challenges While integrating Knowledge Graphs (KGs) with Large Language Models (LLMs) offers potential benefits, several significant challenges persist. Knowledge noise within KGs can lead to inaccuracies in LLM outputs, necessitating robust filtering mechanisms [24], while the complexity of integrating KGs with LLM architectures can introduce substantial computational demands [13]. Accurately capturing the intricacies of a domain in a KG is challenging, as nuanced relationships and exceptions may result in oversimplifications or inaccuracies. Moreover, KGs are often incomplete, leading to gaps in the knowledge accessible to LLMs, which can produce misleading outputs. The dynamic nature of knowledge necessitates continuous updating of KGs to avoid outdated conclusions, a resource-intensive process. Integration complexity arises when aligning structured data from KGs with unstructured data processed by LLMs, requiring sophisticated methods for effective utilization. Additionally, scalability issues become apparent as the size and complexity of KGs increase, complicating maintenance and validation processes. Subjectivity in knowledge selection can lead to inconsistencies in representation, and biases in the underlying data can undermine the trustworthiness of LLM outputs. Variability in knowledge quality further exacerbates trust issues, as inaccurate or outdated information can result in erroneous conclusions. Interpretability challenges emerge when attempting to understand how KGs influence LLM decision-making, compounded by limitations in human validation due to the availability of subject matter experts. Furthermore, the computational overhead of querying and analyzing KGs may impact the efficiency of trust assessments, and user skepticism regarding the reliability of KGs can hinder acceptance, particularly when users are unfamiliar with the methodologies employed in their construction. Finally, interoperability issues can complicate the integration of KGs built using different standards and formats, posing challenges for comprehensive trust assessments across various domains. In summary, while the integration of KGs with LLMs holds promise for enhancing trustworthiness, addressing these multifaceted challenges is critical for achieving effective and ethical outcomes.[15,16,1] 6.2 Future Directions Future directions for using Large Language Models (LLMs) in conjunction with Knowledge Graphs (KGs) to enhance trustworthiness can focus on several key areas. Developing methods for creating and maintaining dynamic knowledge graphs that can automatically update in response to new information, research findings, or changes in domain knowledge is essential. This could involve leveraging real-time data sources and machine-learning techniques to ensure that the knowledge graph remains current and relevant. Additionally, improving the integration of LLM outputs with knowledge graphs through advanced natural language processing techniques could enhance the alignment of LLM-generated content with the structured data in KGs, enabling more accurate assessments of trustworthiness based on contextual relevance. Furthermore, exploring automated or semi-automated validation processes for knowledge graphs, potentially using machine learning algorithms to identify inconsistencies or gaps in the knowledge representation, could reduce reliance on human evaluators and enhance scalability. Encouraging collaboration between domain experts, data scientists, and AI researchers is vital to creating more robust knowledge graphs that accurately reflect the complexities of various fields. This interdisciplinary 145 approach can help ensure that the knowledge represented is comprehensive and trustworthy. Developing user-centric metrics for trust that take into account individual user needs, preferences, and contexts can also enhance the trustworthiness of LLM outputs. Focusing on enhancing the explainability of LLM outputs about the knowledge graph by providing users with clear explanations of how LLM responses are derived from the knowledge graph can build trust through transparency. Moreover, creating crossdomain knowledge graphs that can integrate information from multiple fields allows LLMs to provide more comprehensive and contextually aware responses. This could enhance the trustworthiness of outputs in interdisciplinary applications. Addressing ethical considerations related to trust in LLMs and KGs, including the identification and mitigation of biases in both the knowledge representation and the model outputs, is crucial for broader acceptance. Conducting extensive real-world testing of LLMs combined with knowledge graphs in various applications, such as healthcare, finance, and education, is necessary. Gathering empirical data on their performance and trustworthiness can inform further improvements and refinements. Finally, encouraging community contributions to knowledge graphs, allowing users to add, edit, and validate information, can enhance the richness and accuracy of the knowledge represented, fostering a sense of ownership and trust among users. By pursuing these future directions, the integration of LLMs and knowledge graphs can lead to more reliable, trustworthy, and user-friendly systems that effectively support decision-making across various domains. 7 Conclusion In this paper, we explored the integration of Large Language Models (LLMs) with Knowledge Graphs (KGs) to enhance the trustworthiness of information generated in various domains. We established that while LLMs demonstrate impressive capabilities in natural language processing tasks, their outputs can be limited by biases and inaccuracies inherent in the training data. By combining LLMs with KGs, we can leverage the structured, semantically rich information contained within knowledge graphs to improve the reliability and contextual relevance of LLM-generated content. Our investigation highlighted several promising future directions, including the dynamic updating of knowledge graphs, the enhancement of natural language processing techniques for better integration, and the development of user-centric trust metrics. Additionally, we emphasized the importance of interdisciplinary collaboration and ethical considerations in the deployment of these integrated systems. Through extensive empirical testing and community engagement, we aim to create more robust and trustworthy systems that effectively support decision-making across various fields. The findings of this study pave the way for future research aimed at bridging the gap between LLMs and KGs, ultimately fostering trust and improving the quality of information accessible to users. References [1] Emad A. Alghamdi, Reem I. Masoud, Deema Alnuhait, Afnan Alomairi, Ahmed Ashraf, and Mohamed Zaytoon. Aratrust: An evaluation of trustworthiness for llms in arabic. In Proceedings of the Arabic Language and AI Conference, 2023. [2] Reuben Binns. Fairness in machine learning: Lessons from political philosophy. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pages 149–159, 2018. [3] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, P. Dhariwal, and D. Amodei. Language models are few-shot learners. In Advances in Neural Information Processing Systems, 2020. [4] Erik Cambria, Soujanya Poria, Alexander Gelbukh, and Awais Hussain. Xai meets llms: A survey of the relation between explainable ai and large language models. arXiv preprint arXiv:2407.15248, 2024. [5] Xiaojun Chen, Shengbin Jia, and Yang Xiang. A review: Knowledge reasoning over knowledge graph. Expert Systems with Applications, 141:112948, 2020. [6] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2018. 146 [7] Kaito Fujiwara, Katsuya Nakamura, Jun Matsui, Tatsuya Matsumoto, and Masahiro Hara. Measuring the interpretability and explainability of model decisions of five large language models. arXiv preprint, 2024. [8] Jorge Garnica, Ana Vega, and Juan Guti´errez. Knowledge graphs for explainable artificial intelligence: A survey. In Proceedings of the 2022 International Joint Conference on Neural Networks (IJCNN), pages 1–8, 2022. [9] Kurt Holstein, Jennifer Wortman Vaughan, Hal Daum´e III, and Mike Dudik. Improving fairness in machine learning systems: What do industry practitioners need? In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 2019. [10] Yu Hou, Jeremy Yeung, Hua Xu, Chang Su, Fei Wang, and Rui Zhang. From answers to insights: Unveiling the strengths and limitations of chatgpt and biomedical knowledge graphs. In Proceedings of the Annual Conference on Artificial Intelligence in Medicine, 2023. [11] Tuan Manh Lai. Knowledge Acquisition for Natural Language Understanding. PhD thesis, University of Illinois at Urbana-Champaign, 2023. [12] Yahan Li, Yi Wang, Yi Chang, and Yuan Wu. Xtrust: On the multilingual trustworthiness of large language models. In Proceedings of the International Conference on Multilingual NLP, 2023. [13] Carlo Lipizzi. Tell me the truth: A system to measure the trustworthiness of large language models. In Proceedings of the International Conference on Trustworthy AI, 2023. [14] Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pages 220–229, 2019. [15] Jeff Z. Pan et al. Large language models and knowledge graphs: Opportunities and challenges. arXiv preprint arXiv:2308.06374, 2023. [16] Shirui Pan, Zhiwei Liu, Wei Zhuang, Rui Yang, Lei Zhang, Jialiang Li, and Haifeng Wang. Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engineering, 2024. [17] Aleksandra Piktus, Yi Wang, Vladimir Karpukhin, Ravi Pasunuru, Thibaut Mialon, Nisan Stiennon, Wen-tau Yih, Laleh Koc, and Matthew Efron. The roots search tool: Data transparency for llms. arXiv preprint arXiv:2302.14035, 2023. [18] Ridho Reinanda, Edgar Meij, and Maarten de Rijke. Knowledge graphs: An information retrieval perspective. Foundations and Trends®in Information Retrieval, 14(4):289–444, 2020. [19] Jingwei Sun, Zhixu Du, and Yiran Chen. Knowledge graph tuning: Real-time large language model personalization based on human feedback. arXiv preprint arXiv:2405.19686, 2024. [20] Zhen Tan, Yi Zhang, Xia Shen, Zhen Wang, and Lichao Li. Tuning-free accountable intervention for llm deployment–a metacognitive approach. arXiv preprint arXiv:2403.05636, 2024. [21] Weixuan Wang, Qi Liu, Huilin Xiong, Hang Liu, Liang Hu, Xinyu Zhang, and Yixin Gu. Assessing the reliability of large language model knowledge. arXiv preprint arXiv:2310.09820, 2023. [22] Yujie Wang, Xiaogang Zhang, and Wei Liu. Integrating knowledge graphs into language models: A comprehensive review. Journal of Artificial Intelligence Research, 72:119–143, 2023. [23] Sophie Xhonneux, Pierre Legrand, Wouter De Pauw, Xue Ma, John Sutherland, and Kara Kockelman. Efficient adversarial training in llms with continuous attacks. arXiv preprint arXiv:2405.15589, 2024. [24] Linyao Yang, Hongyang Chen, Zhao Li, Xiao Ding, and Xindong Wu. Give us the facts: Enhancing large language models with knowledge graphs for fact-aware language modeling. In IEEE Transactions on Knowledge and Data Engineering, 2023. [25] Shuo Yu, Zhen Wang, Xiang Zhang, Zhiyuan Liu, and Maosong Sun. Deep learning meets knowledge graphs: A comprehensive survey. arXiv preprint arXiv:2205.02573, 2022. 147