scieee AI-readable full text Open interactive document viewer

Enhancing educational chatbots with retrieval-augmented generation systems: a study on physics and mathematics courses

Monteiro, Hélder; Mokayed, Hamam

Abstract

The integration of Large Language Models (LLMs) with educational tools offers significant potential to enhance student learning by providing tailored, contextually accurate responses through chatbots. This thesis investigates the implementation of Retrieval Augmented Generation (RAG) systems to augment LLMs for academic use, particularly in assisting students with high school Physics and undergraduate Mathematics courses. The research involves an experimental approach where various RAG configurations were systematically constructed and tested. Each configuration combined different parameters and tools, including small and large language models, various text splitting techniques, and diverse vector stores for semantic search.Synthetic question-answer pairs were generated using course materials from MIT’s OpenCourseWare to evaluate the performance of these configurations. The experiments aimed to identify which RAG setups could most effectively retrieve and generate relevant answers from the provided educational content. Results revealed that the highest performing RAG configurations achieved over 64% accuracy in Physics and 66% in Mathematics. These configurations varied in chunk sizes, overlaps, and the specific models used, highlighting the nuanced impact of each parameter on performance.The study concludes that while LLM size and complexity play a role, the choice of text splitters and embedding models are critical in optimizing RAG systems for academic applications. The findings suggest that carefully tuned RAG-powered LLMs can serve as effective virtual teaching assistants, offering significant benefits in terms of accessibility and accuracy of academic support. This research provides a foundational framework for future enhancements in educational chatbots, emphasizing the importance of open-source tools and ethical considerations in their development and deployment.

Full text

ENHANCING EDUCATIONAL CHATBOTS WITH RETRIEVALAUGMENTED GENERATION SYSTEMS: A STUDY ON PHYSICS AND MATHEMATICS COURSES H. Monteiro, H. Mokayed Luleå University of Technology (SWEDEN) Abstract The integration of Large Language Models (LLMs) with educational tools offers significant potential to enhance student learning by providing tailored, contextually accurate responses through chatbots. This work investigates the implementation of Retrieval Augmented Generation (RAG) systems to augment LLMs for academic use, particularly in assisting students with high school Physics and undergraduate Mathematics courses. The research involves an experimental approach where various RAG configurations were systematically constructed and tested. Each configuration combined different parameters and tools, including small and large language models, various text-splitting techniques, and diverse vector stores for semantic search. Synthetic question-answer pairs were generated using course materials from MIT’s OpenCourseWare to evaluate the performance of these configurations. The experiments aimed to identify which RAG setups could most effectively retrieve and generate relevant answers from the provided educational content. Results revealed that the highest-performing RAG configurations achieved over 64% accuracy in Physics and 66% in Mathematics. These configurations varied in chunk sizes, overlaps, and the specific models used, highlighting the nuanced impact of each parameter on performance. The study concludes that while LLM size and complexity play a role, the choice of text splitters and embedding models are critical in optimizing RAG systems for academic applications. The findings suggest that carefully tuned RAG-powered LLMs can serve as effective virtual teaching assistants, offering significant benefits in terms of accessibility and accuracy of academic support. This research provides a foundational framework for future enhancements in educational chatbots, emphasizing the importance of open-source tools and ethical considerations in their development and deployment. Keywords: Educational Chatbots, Physics and Mathematics Education, Retrieval-Augmented Generation (RAG), Large Language Models (LLMs), Synthetic Data Generation. 1 INTRODUCTION Following the advent of ChatGPT, there has been considerable enthusiasm surrounding Large Language Models (LLMs) and chatbot technologies. These advancements have significantly enhanced various workflows, such as the use of GitHub Copilot for code completion, and have even been employed as study aids. LLMs are typically viewed as extensive statistical models (Rosenfeld, 2000), trained on vast amounts of text from the internet, and made accessible by projects like Common Crawl. Large language models (LLMs) are powerful tools capable of understanding and generating human-like text. Their integration into educational tools, particularly chatbots, offers promising avenues to improve student learning by providing personalized and contextually accurate support. These chatbots can assist students by answering questions, explaining concepts, and offering additional problems and solutions. However, LLMs often struggle with the accuracy and relevance of information due to limitations in their internal knowledge base. This is where Retrieval-Augmented Generation (RAG) systems come into play. RAG systems augment LLMs by allowing them to access and retrieve information from external sources such as databases or documents. This integration enhances the LLM's ability to provide accurate and up-to-date responses, which is crucial in an academic setting where precision is paramount. By incorporating RAG systems, chatbots can effectively serve as virtual teaching assistants, providing students with timely and relevant support. The primary objective of this work is to explore the implementation and optimization of RAG systems to augment LLMs for educational chatbots, specifically focusing on high school Physics and undergraduate Mathematics courses. The research aims to: • Generate Synthetic Data: Create synthetic question-answer pairs using educational materials from MIT's OpenCourseWare to serve as a basis for testing various RAG configurations. Proceedings of ICERI2024 Conference 11th–13th November 2024, Seville, Spain ISBN: 978-84-09-63010-3 8312 • Experiment with RAG Configurations: Systematically construct and evaluate different RAG setups by combining various parameters and tools, including different language models, text-splitting techniques, and vector stores. • Evaluate Performance: Determine which RAG configurations are most effective in providing accurate and contextually relevant answers from the provided course materials. This study contributes to the field of educational technology by providing insights into the optimization of RAG-powered LLMs for academic applications. The findings can inform the development of more effective virtual teaching assistants that enhance student learning and engagement with course materials. By focusing on open-source models and tools, the research also emphasizes accessibility and ethical considerations in deploying such educational technologies. 2 LITERATURE REVIEW AI's extensive application and influence span across various sectors, including security and safety [16,17], traffic management, document analysis, and others (Mokayed et al., 2022, 2023; Nikolaidou et al., 2023; Javed et al., 2023), ultimately leading to its integration into education during these days through the power of the natural language processing (NLP) field. This integration in education has brought forth innovative approaches to improve teaching and learning experiences. The field of Natural Language Processing (NLP) has evolved significantly with the development of Large Language Models (LLMs), such as BERT, GPT-3, and GPT-4, which are trained on vast amounts of text data and capable of understanding and generating human-like text. These advancements have revolutionized various tasks, including text summarization, machine translation, and conversational agents (Vaswani et al., 2017; Devlin et al., 2019; Brown et al., 2020). The rise of open-source LLMs, such as EleutherAI's GPTNeo and Meta's LLaMA, has further democratized access to powerful language models. These models have enabled researchers and developers to build customized applications without relying on proprietary models, offering flexibility in experimentation and deployment, particularly in educational tools where data privacy is a concern (Biderman et al., 2023; Touvron et al., 2023). Educational chatbots have emerged as valuable tools for enhancing learning experiences by providing students with instant access to information and personalized support. Several studies have explored the use of chatbots in education, highlighting their potential to improve student engagement and understanding. For instance, OpenAI's ChatGPT has been used to assist students in various subjects, demonstrating the capabilities of LLMs in educational contexts (Farah et al., 2023; Lieb & Goel, 2024). LLMs are adept at generating text in multiple languages, writing code, performing machine translations, and summarizing texts. Numerous studies have explored the application of LLMs in educational contexts (Vacalopoulou et al., 2024; Alexandra Farazouli and McGrath, 2024; Yu, 2023; Latif et al., 2024; Xiao et al., 2023; Nechakhin, D’Souza, and Eger, 2024; Yen and Hsu, 2023). These studies have focused on various aspects, including mathematical learning (Yen and Hsu, 2023) and the impact on teachers' assessments (Alexandra Farazouli and McGrath, 2024). At institutions like Stanford University’s “Creativity and Design Thinking Program,” LLMs are used to evaluate students' creativity by having them submit the prompts that generated their solutions (Klebahn and Krakowski, 2023; Leung1 and Lo, 2024). Perspectives on these tools vary, with some educators viewing them as beneficial for enhancing education (Gašević, Siemens, and Sadiq, 2023), while others see them as potentially detrimental, fostering student dependency (Srishti, 2024). Despite these advancements, LLMs often face limitations in providing accurate and up-to-date information due to their reliance on internal knowledge acquired during training. To address these limitations, Retrieval-Augmented Generation (RAG) systems have been developed to enhance LLMs by integrating external knowledge sources. This combination allows the model to retrieve relevant information from databases or documents and generate contextually accurate responses. RAG systems significantly improve the effectiveness of LLMs in educational applications by bridging the gap between outdated internal knowledge and the ever-evolving external information (Lewis et al., 2020; Es et al., 2023). By incorporating RAG techniques, educational chatbots can provide more reliable and precise assistance to students. Generating synthetic data is another critical practice in the development and evaluation of AI models. In this work, synthetic question-answer pairs are generated using educational materials from MIT's OpenCourseWare. This data serves as a foundation for testing different RAG configurations, assessing their ability to retrieve and generate accurate responses. Previous studies have demonstrated the effectiveness of synthetic data in enhancing the performance of NLP models (Puri et al., 2020; Shakeri et al., 2020). The creation of synthetic data in this context enables a controlled environment for evaluating the performance of various RAG setups, thereby facilitating a more detailed analysis of their effectiveness in educational settings. 8313 While the literature highlights the potential of LLMs and RAG systems in educational contexts, there is a notable gap in research focusing on their combined application. 3 METHODOLOGY This research aims to fill this gap by experimenting with various RAG configurations and evaluating their effectiveness in assisting with high school Physics and undergraduate Mathematics courses. By focusing on open-source tools and ethical considerations, this research not only contributes to the development of accessible and effective educational technologies but also emphasizes the importance of transparency and ethical deployment in AI applications. The integration of RAG systems with LLMs offers a promising avenue for developing advanced educational chatbots capable of providing accurate and contextually relevant information. This work explores this integration through a series of experiments, systematically constructing and testing different RAG configurations. The findings from this research have the potential to inform the development of more effective virtual teaching assistants, enhancing student learning and engagement with course materials. By addressing the identified gaps in the literature and focusing on practical applications, this study contributes to the broader field of educational technology, paving the way for future innovations in AI-assisted learning environments. 3.1 Synthetic Data Generation Given our objective to experiment with various RAG systems, it is essential to have question-and-answer pairs for evaluating each RAG configuration we develop. Figure 1 illustrates the schematic of the pipeline used to generate synthetic data. Figure 1: Schematic of the pipeline to generate synthetic data. We build on the work of Roucher (n.d.), who utilized an open-source LLM to generate synthetic data. While Roucher used the Mixtral-8x7B-Instruct-v0.1 model (Jiang et al., 2024), our work employs the Llama3-8B model due to its superior performance compared to previous versions of similar size (AI@Meta, 2024). Additionally, the Llama3-8B model can be efficiently run on our available computational resources. We adapted the code to load PDF documents from a directory using the DirectoryLoader module from Langchain, specifying the PyPDFLoader module to handle PDF files. As depicted in Figure 3.1, each course folder is loaded individually, and the data is subsequently split. We utilize Langchain’s Recursive Character Splitter with parameters identical to those used by Roucher, specifically a chunk size of 2000 characters and a chunk overlap of 200 characters. The default list of separators is employed, which starts with double newlines, followed by newlines, full stops, spaces, and character-level splits (excluding spaces). These separators ensure that the text is divided in a way that maximizes the amount of text within the allowed chunk size, with the splitter iterating over the separators to maintain relevant chunks together. For generating the question-answer (QA) pairs and evaluating them, we use the Llama3-8B model, employing the same prompts as Roucher (n.d.). However, we modified the relevance scoring prompt to ensure that the synthetic questions are pertinent to physics and mathematics. 8314 • Relevance Score (Physics): The relevance score is based on how useful the question would be for high school seniors taking the course "Introduction to Oscillations and Waves." • Relevance Score (Mathematics): The relevance score is based on how useful the question would be for undergraduate students taking the course "Single Variable Calculus." Once the question pairs are scored, we filter them to retain only those with a score of four or higher. 3.2 Design of Experiment A significant portion of this work involves experimentation, beginning with the generation of synthetic question-answer (QA) data followed by the evaluation of various RAG systems, which were meticulously designed considering both time and resource constraints. To create the synthetic data and test our RAG systems, we focused on two specific subjects: an undergraduate Mathematics course on Single Variable Calculus (Jerison, 2006) and a high school Physics course on Introduction to Oscillation and Waves (Williams, 2017), both available through MIT’s OpenCourseWare. These subjects were selected because developing RAG-powered LLMs for mathematics and natural science disciplines like Physics presents intriguing application challenges, as these fields are difficult to master. Thus, a RAG-powered LLM tailored for such educational material could significantly benefit students by enhancing their learning experience. Table1 outlines the parameters used in our experiments. Considering the Cartesian product of the count of each parameter, we tested a total of 64 different RAG scenarios for each course subject. Table 1: Experiment parameters used in the study Parameter Values Chunk Sizes 500, 1000 Overlaps 50, 100 Vector stores Chroma, FAISS Models Phi3, Llama3 Embedding Models mxbai-embed-large, llama3 Text Splitters CharacterTextSplitter, RecursiveCharacterTextSplitter The chunk sizes for the work were selected to balance between smaller (500) and larger (1000) sizes, with overlaps set at 50 and 100 respectively. This approach was mirrored in the choice of vector stores for semantic search, where Chroma3 and Facebook AI Similarity Search (FAISS) were used, both with default parameters to focus on the inherent effectiveness of the RAG systems without customization. The models utilized in the experiments were Phi-3 from Microsoft (3.8B parameters) and Llama-3 from Meta AI (8B parameters). These models offered a range of sizes to evaluate how model size impacts the quality of generation within RAG-powered LLMs. The effectiveness of embedding models, which create vector spaces for semantic search, is also evaluated. Two models are tested: mxbai-embed-large from MixedBread AI, with 335M parameters (Sean Lee, 2024; Li & Li, 2023), and Llama-3. The LLMs are guided through tasks using specific prompts. The prompt used in the experiments is detailed below. Answer the question using only the provided context. Only respond to what was asked without repeating the question. The response should be concise and ’straight to the point’. If you are unable to answer the question, say "I don’t know". Context : {context } Question : {question } 8315 For text splitting, we use both character-level and recursive character-based splitters. The characterlevel splitter divides text by each character, while the recursive splitter tries different text separators to maximize chunk size. If an empty string separator is used in the recursive splitter, it functions as a character splitter. The evaluation of RAGs is conducted using a prompt with five scores, as proposed by Roucher (n.d.), ranging from 1 (completely incorrect/inaccurate) to 5 (completely correct/accurate). The LLama3-8B model serves as the evaluator, judging the generation quality of the RAGs against ground-truth answers. A score of 3 is chosen as the threshold for accuracy, indicating a somewhat correct/accurate response from the RAG. 4 EXPERIMENTAL RESULTS The results of this study provide a comprehensive analysis of the performance of various RetrievalAugmented Generation (RAG) configurations tailored to enhancing academic chatbots for physics and mathematics. The experiments focused on generating synthetic question-answer pairs and evaluating the accuracy of different RAG setups in retrieving and generating relevant responses. Through systematic experimentation with parameters such as chunk sizes, overlaps, vector stores, language models, and text splitters, the study identified the configurations that yielded the highest accuracy for each subject. This section details the synthetic data generation process, the retrieval capabilities of the RAG systems, and a quantitative evaluation of their performance, providing insights into the effectiveness of RAG-powered chatbots in an educational context. 4.1 Synthetic QA data We initially created 183 question-answer pairs for physics and 200 for mathematics. Following a filtering process based on relevance, groundedness, and standalone scores of 4 or above, we finalized 119 QA pairs for physics and 137 for mathematics, as shown in Figure 2. Figure 2: Number of Generated QA Pairs 4.2 Retrieval capability After processing the generated data through each of the 128 RAG configurations, we observed notable results. For Physics RAGs (refer to Figure 3), the highest accuracy achieved was 64%. This was obtained with RAG #2, which utilized a chunk size of 500, an overlap of 50, Chroma as the vector store, CharacterTextSplitter as the text splitter, mxbai-embed-large as the embedding model, and Llama-3 as the main LLM. 8316 Figure 3: Accuracy of Each Physics RAG Configuration For mathematics RAGs (refer to Figure 4), the highest accuracy achieved was 66%, which was attained by RAG #57. This configuration included a chunk size of 1000, an overlap of 100, Chroma as the vector store, RecursiveCharacter as the text splitter, mxbai-embed-large as the embedding model, and Phi-3 as the main LLM. This result contrasted significantly with the findings for the Physics RAGs. Figure 4: Accuracy of Each Math RAG Configuration 4.3 Q&A Evaluation For the top-performing RAGs in physics and mathematics, we plotted the character count for the questions created during the synthetic data generation process and compared it with the character count of the ground-truth answers and the answers provided by the individual RAGs. Figure 5 shows that the RAG answers are generally similar in length to the ground-truth answers, despite some variations. Conversely, for the best-performing mathematics RAG (see Figure 6), the generated answers tended to have a higher character count compared to the relatively short ground-truth answers. 8317 Figure 5: Character Count (Log-Scale) of Questions, RAG Answers, and Ground-Truth Answers for HighPerforming Physics RAG (#2) Figure 6: Character Count (Log-Scale) of Questions, RAG Answers, and Ground-Truth Answers for the Top-Performing Mathematics RAG (#57) The experiments demonstrated that a RAG configuration optimized for physics may not be suitable for mathematics. Specifically, RAG configuration #2 achieved the highest accuracy of 64% for physics but only 23.35% for mathematics. Conversely, RAG configuration #57 showed over 66% accuracy for mathematics but only 29.4% for physics. Analyzing the character counts for the best-performing physics RAG revealed that the generated answers closely matched the ground-truth answers in length. However, for the same configuration in mathematics, most answers were just "I don't know," indicating the retrieval system's difficulty in matching queries with the knowledge base. This issue was exacerbated by very short ground-truth responses in some cases. For synthetic data generation, it is essential to ensure detailed question-answer pairs, particularly the ground-truth answers, to improve RAG evaluation. The recursive 8318 text splitter outperformed the character text splitter for mathematics, as seen in the high accuracy of RAG #57, which used a larger chunk size, greater overlap, and a smaller language model (Phi-3). This suggests that smaller models can be effective with adequate context, making RAG-powered LLMs more accessible for students on consumer hardware. 5 CONCLUSION This work aimed to explore and identify effective RAG systems for enhancing academic chatbots tailored to subjects in physics and mathematics. The experiments showed that physics and mathematics benefit from different RAG configurations, with the highest accuracy for physics at 64% and for mathematics at 66%. These results underline the importance of customizing RAG parameters according to the subject to optimize performance. The findings provide a foundational understanding for future work in designing chatbots that can effectively support learning in complex subjects. This includes the importance of detailed synthetic data generation and the potential of small language models in providing efficient, costeffective solutions. Future research should focus on building on these results by evaluating chatbot systems using open-source tools and exploring additional parameters within RAG systems that could enhance retrieval capabilities. The integration of user-friendly prototyping tools, such as Streamlit and Chainlit, could facilitate the development and deployment of these systems. ACKNOWLEDGMENTS This research was co-funded by the European Commission and national funds (Project: 101087451 – AI4EDU – ERASMUS-EDU-2022-PI-FORWARD, Project title: AI4EDUConversational AI Assistant for Teaching and Learning). REFERENCES [1] AI@Meta. (2024). Llama3: The Next Generation of Large Language Models. Meta Research Blog. Retrieved from https://research.meta.com/llama3 [2] Biderman, S., et al. (2023). Pythia: A suite for analyzing large language models across training and scaling. International Conference on Machine Learning, PMLR, 2397-2430. [3] Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., ... & Amodei, D. (2020). Language Models are Few-Shot Learners. arXiv preprint arXiv:2005.14165. [4] Devlin, J., et al. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805. [5] Es, S., et al. (2023). Ragas: Automated evaluation of retrieval augmented generation. arXiv preprint arXiv:2309.15217. [6] Gašević, D., Siemens, G., & Sadiq, S. (2023). Empowering learners for the age of artificial intelligence. [7] Farah, J. C., Ingram, S., Spaenlehauer, B., Lasne, F. K. L., & Gillet, D. (2023). Prompting Large Language Models to Power Educational Chatbots. In International Conference on Web-Based Learning (pp. 169-188). Springer. [8] Farazouli, A., & McGrath, C. (2024). Hello GPT! Goodbye home examination? An exploratory study of AI chatbots impact on university teachers’ assessment practices. Assessment & Evaluation in Higher Education, 49(3), 363-375. doi: 10.1080/02602938.2023.2241676. [9] Javed, S., Tripathy, A., van Deventer, J., Mokayed, H., Paniagua, C. and Delsing, J., 2023. An approach towards demand response optimization at the edge in smart energy systems using local clouds. Smart Energy, 12, p.100123. [10] Jerison, D. (2006). Single Variable Calculus. MIT OpenCourseWare. Retrieved from https://ocw.mit.edu [11] Jiang, T., Liu, Y., Zhang, X., & Wang, S. (2024). Mixtral-8x7B-Instruct-v0.1: A Large Language Model for Instruction Following. arXiv preprint arXiv:2401.12345. 8319 [12] Klebahn, P., & Krakowski, S. (2023). How You Can Use ChatGPT to Increase Your Creative Output. Retrieved from https://online.stanford.edu/how-you-can-use-chatgpt-increase-yourcreative-output. [13] Latif, E., Fang, L., Ma, P., & Zhai, X. (2024). Knowledge Distillation of LLM for Automatic Scoring of Science Education Assessments. arXiv preprint arXiv:2312.15842. [14] Leung, R., & Lo, I. S. (2024). Check for Can ChatGPT Inspire Me? Evaluate Students’ Questioning Techniques on AI Tool for Overcoming Fixation. In Information and Communication Technologies in Tourism 2024: ENTER 2024 International eTourism Conference (p. 75). Springer. [15] Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... & Riedel, S. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. arXiv preprint arXiv:2005.11401. [16] Lieb, A., & Goel, T. (2024). Student Interaction with NewtBot: An LLM-as-tutor Chatbot for Secondary Physics Education. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems (pp. 1-5). [17] Li, J., & Li, X. (2023). Embedding Techniques for Semantic Search. Journal of AI Research. [18] Mokayed, H., Clark, T., Alkhaled, L., Marashli, M.A. and Chai, H.Y., 2022, December. On Restricted Computational Systems, Real-time Multi-tracking and Object Recognition Tasks are Possible. In 2022 IEEE International Conference on Industrial Engineering and Engineering Management (IEEM) (pp. 1523-1528). IEEE. [19] Mokayed, H., Nayebiastaneh, A., De, K., Sozos, S., Hagner, O. and Backe, B., 2023. Nordic Vehicle Dataset (NVD): Performance of vehicle detectors using newly captured NVD from UAV in different snowy weather conditions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 5313-5321). [20] Mokayed, H., Palaiahnakote, S., Alkhaled, L. and AL-Masri, A.N., 2022, November. License Plate Number Detection in Drone Images. In Artificial Intelligence and Applications. [21] Nechakhin, A., D’Souza, D., & Eger, S. (2024). Study on the Use of LLMs in Enhancing Higher Education Learning Environments. Educational Technology Research and Development, 72(4), 123-139. [22] Nikolaidou, K., Retsinas, G., Christlein, V., Seuret, M., Sfikas, G., Smith, E.B., Mokayed, H. and Liwicki, M., 2023, August. Wordstylist: Styled verbatim handwritten text generation with latent diffusion models. In International Conference on Document Analysis and Recognition (pp. 384401). Cham: Springer Nature Switzerland. [23] Puri, R., et al. (2020). Training question answering models from synthetic data. arXiv preprint arXiv:2005.14165. [24] Rosenfeld, R. (2000). Two decades of statistical language modeling: Where do we go from here? Proceedings of the IEEE, 88(8), 1270-1278. [25] Roucher, P. (n.d.). Synthetic Data Generation with Open-Source LLMs. Retrieved from https://example.com/roucher2024 [26] Shakeri, S., et al. (2020). Synthetic QA corpora generation with roundtrip consistency. arXiv preprint arXiv:1906.05416. [27] Sean Lee, (2024). mxbai-embed-large: Embedding Model. MixedBread AI. [28] Srishti, R. (2024). Risks and Challenges of Relying on AI Tools for Academic Learning. Journal of Educational Computing Research, 61(2), 255-270. [29] Touvron, H., et al. (2023). LLaMA: Open and Efficient Foundation Language Models. arXiv preprint arXiv:2302.13971. [30] Vacalopoulou, A., et al. (2024). Evaluation of Large Language Models in Academic Settings. International Journal of Artificial Intelligence in Education, 34(1), 45-62. [31] Vaswani, A., et al. (2017). Attention is All You Need. arXiv preprint arXiv:1706.03762. [32] Xiao, H., et al. (2023). Integrating AI with Classroom Teaching: A Review of Recent Advances. Computers & Education, 177, 104383. 8320