scieee AI-readable full text Open interactive document viewer

NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware

Neogi, Madhumita; Dasgupta, Tirtharaj; Mukhopadhyay, Parthasarathi

Abstract

In this rapidly changing world of digital libraries, the typical keywords-based information retrieval systems like Online Public Access Catalogs (OPACs) sometimes fail to fulfill the increasingly complex, multilingual, and natural language-based queries of the users. This study unveils an innovative retrieval framework that consists of a Large Language Model (LLM)-powered frontend and an open-source library discovery platform in the backend and integrates both of them through the Model Context Protocol (MCP). In this framework, VuFind is used as the backend library discovery system, which can accumulate metadata from different data sources like Koha ILS, Greenstone, Omeka Classic, and DSpace and integrate it with Claude Desktop, which is an LLMbased natural language assistant, via a locally deployed MCP server. This integration improves user experience and search accuracy through the implementation of context-aware, multilingual, and natural language information retrieval. The performance of this integrating framework was analyzed by using a set of retrieval scenarios and demonstrated its capabilities to perform the semantic queries and deliver more relevant search results. Although some limitations were identified in structured queries such as call number retrieval. The conclusion of this study highlights how this MCP-based framework has the potential to transform the library discovery services from keywordbased interfaces toward more user-friendly, NLP-based conversational systems. Enhancements in security, scalability, and integration with other open-source LLMs and library discovery services, as well as proprietary ones, are the future suggestions of this study.

Full text

NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware Madhumita Neogi Research Scholar, Department of Library and Information Science,University of Kalyani, Kalyani, Nadia, West Bengal, 741235 email id: [email protected] Tirtharaj Dasgupta Student, Department of Library and Information Science,University of Kalyani, Kalyani, Nadia, West Bengal, 741235 email id: [email protected] Parthasarathi Mukhopadhyay Professor, Department of Library and Information Science,University of Kalyani, Kalyani, Nadia, West Bengal, 741235 email id: [email protected] Corresponding author: Madhumita Neogi, email id: [email protected] Abstract In this rapidly changing world of digital libraries, the typical keywords-based information retrieval systems like Online Public Access Catalogs (OPACs) sometimes fail to fulfill the increasingly complex, multilingual, and natural language-based queries of the users. This study unveils an innovative retrieval framework that consists of a Large Language Model (LLM)-powered frontend and an open-source library discovery platform in the backend and integrates both of them through the Model Context Protocol (MCP). In this framework, VuFind is used as the backend library discovery system, which can accumulate metadata from different data sources like Koha ILS, Greenstone, Omeka Classic, and DSpace and integrate it with Claude Desktop, which is an LLMbased natural language assistant, via a locally deployed MCP server. This integration improves user experience and search accuracy through the implementation of context-aware, multilingual, and natural language information retrieval. The performance of this integrating framework was analyzed by using a set of retrieval scenarios and demonstrated its capabilities to perform the semantic queries and deliver more relevant search results. Although some limitations were identified in structured queries such as call number retrieval. The conclusion of this study highlights how this MCP-based framework has the potential to transform the library discovery services from keywordbased interfaces toward more user-friendly, NLP-based conversational systems. Enhancements in security, scalability, and integration with other open-source LLMs and library discovery services, as well as proprietary ones, are the future suggestions of this study. Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. Keywords: Claude Desktop; Information Retrieval (IR); Large Language Models (LLMs); Library Discovery; Model Context Protocol (MCP); Multilingual Search; Natural Language Processing (NLP); VuFind. 1. Introduction: Information Retrieval (IR) is defined as the process of obtaining material, usually documents, that satisfies an information need from within large collections, typically stored on computers (Manning et al., 2008). Despite significant technological advancements in library services, users continue to face barriers to effective information access. Although library cataloguing or discovery systems have evolved from traditional card catalogues to Online Public Access Catalogues (OPACs), they often fall short in meeting users' increasingly complex and sophisticated information needs. A primary challenge in current IR systems is their reliance on keyword-based retrieval, which often leads to excessively high recall - retrieving a vast number of documents, many of which may be irrelevant to the user's intent (Robertson & Zaragoza, 2009). For example, querying a common term may return thousands or even millions of results, overwhelming users and hindering effective information seeking. This issue is compounded by polysemy, where words with multiple meanings (e.g., "bank" in river bank and finance bank) result in the retrieval of contextually inappropriate documents, thereby lowering precision (Manning et al., 2008). Furthermore, vocabulary mismatches between user queries and the indexing language - such as controlled vocabularies or subject headings - continue to obstruct relevant discovery (Bates, 1989). These factors collectively make it difficult for users to identify useful materials in large-scale library systems. Another significant limitation is the lack of semantic understanding in traditional IR systems, which often treat words as isolated entities without considering contextual meanings and relationships. This limitation hampers the processing of complex queries, especially those employing Boolean operators and advanced search functionalities, leading to user frustration (Brants, 2004). The vast volume of content in digital libraries further contributes to information overload, complicating the retrieval of relevant materials. Basic relevance ranking algorithms that depend primarily on keyword frequency may overlook important but less commonly used terms. Modern users increasingly prefer natural language queries, which typical IR systems struggle to handle effectively. Furthermore, cross-language information retrieval presents challenges, as translating user queries into different languages can introduce ambiguities and inaccuracies (Grefenstette, Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. 1998). Most systems also lack personalization features, delivering identical search results to all users regardless of their unique needs or query context (Teevan et al., 2005). Natural Language Processing (NLP) offers promising solutions to these challenges. By employing techniques such as word embeddings and semantic analysis, NLP enables retrieval systems to understand word meanings and relationships, improving recall by recognizing synonyms and semantically related terms, while enhancing precision through context-based recognition of polysemous phrases (Brants, 2004). NLP significantly improves query formulation and comprehension by analyzing natural language questions, extracting key concepts, incorporating relevant synonyms, and even determining user interests, ultimately delivering more relevant search results. Beyond simple keywords, NLP identifies entities, underlying concepts, and themes in documents, enabling more precise indexing and retrieval even without exact keyword matches. It enhances relevance ranking by assessing semantic similarity between queries and documents, considering contextual relevance rather than mere keyword overlap. NLP facilitates cross-language information retrieval through machine translation techniques and enables personalization by tailoring results based on user profiles and search histories (Teevan et al., 2005). NLP-powered question-answering systems provide direct answers from library collections, offering faster routes to information, while summarization capabilities help users quickly determine content relevance. The integration of NLP with library IR systems promises a more efficient, effective, and user-friendly information discovery experience, ultimately improving utilization of vast library resources and increasing user satisfaction. 2. Background of the study: Model Context Protocol is an open-source protocol that was introduced by Anthropic (Anthropic also developed large language models like Claude). The primary objective of MCP is to standardize how other open-source applications offer context to LLMs. MCP can be considered as the USB-C port of artificial intelligence (AI) implementations. USB-C enables a uniform method to connect your devices to a variety of accessories and equipment; similarly, MCP also offers a defined method for connecting AI models to various software and databases. According to (MSV, 2024), MCP is a significant advancement in the functioning of AI agents. Agents can now undertake helpful, multistep actions like gathering data, summarizing papers, or storing content to a file in addition to answering queries. Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. The VuFind-MCP project, available at h ttps://github.com/MCP-for-VuFind , is a notable opensource initiative aimed at integrating the Model Context Protocol (MCP) with VuFind, an opensource library discovery interface. VuFind enables users to search and browse library holdings through a unified interface, aggregating content from catalogues, repositories, and other systems. With MCP, large language models (LLMs) such as Claude can interact directly with VuFind’s APIs, allowing them to retrieve library records, interpret natural language queries, and provide more intelligent and contextual responses (Anthropic, 2024; McNulty, 2025). MCP functions as a standard interface, akin to a USB-C port, simplifying connections between AI systems and digital services (Anthropic, 2024). Its implementation in library systems like VuFind allows AI agents to perform complex tasks such as document summarization, contextual search, and content storage. The VuFind-MCP integration also serves as a proof of concept for extending MCP to other library and archival platforms such as Koha, Omeka, and DSpace. Although we have not integrated DSpace with our VuFind instance here, integration is feasible due to DSpace’s support for OAI-PMH and RESTful APIs, which VuFind can be configured to harvest or query. As noted in recent developments, such integrations mark a significant step toward enabling AI-driven services across the information retrieval landscape, particularly in library environments (Wikipedia, 2025; McNulty, 2025). 3. Related Literature:- 3.1 Literature Review: Model Context Protocol MCP is all about context. The transformer architecture, a cornerstone of contemporary natural language processing, has revolutionized how models handle sequential data. However, one of its primary limitations is the fixed-length context window, which restricts its capacity to model long-range dependencies. This literature review explores key developments in addressing this limitation, focusing on memory-augmented transformers, session-level memory handling, and scalable context extension techniques. Recent advances in large language models (LLMs) have shifted attention toward improving context persistence and contextual awareness over extended interactions. The Model Context Protocol (MCP) emerges as a proposed standard for structuring contextual memory across sessions, enabling models to reference prior user interactions seamlessly. This literature review explores foundational and current research related to contextual memory, grounding, and the challenges MCP aims to address in NLP. One of the pivotal challenges in natural Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. language processing is achieving persistent memory across sessions, a problem well-addressed in many foundational works. In 2019, one such research work introduced LAMA, a probing method to assess how much factual knowledge is stored in LLMs. This work laid the groundwork for understanding knowledge retention without persistent memory mechanisms (Petroni et al., 2019). A group of researchers proposed Compressive Transformers, which store compressed summaries of past activations in a separate memory stream, allowing models to retain historical context beyond the standard attention window (Rae et al., 2020). Building upon this, a research team introduced RETRO, a retrieval-augmented model that outperforms traditional transformers by dynamically retrieving context from a massive database. RETRO’s design mimics what MCP proposes conceptually - maintaining and accessing past information to improve responses, but it does so within a single session scope (Borgeaud et al., 2022). Khandelwal et al. in 2020 highlighted that LLMs heavily rely on nearby tokens local context, showing limitations in leveraging long-range dependencies. This reinforces the necessity for protocols like MCP to scaffold session-spanning memory (Khandelwal et al., 2020). In contrast, a group of researchers tackled this by proposing Transformer-XL, a model capable of learning dependencies beyond a fixed-length context. This method preserves states across segments, providing a blueprint for session-spanning memory similar to MCP’s goal (Dai et al., 2019). Moreover, memory integration has been explored through external memory mechanisms like Wu et al. in 2022 introduced Memorizing Transformers, equipping models with key-value memories updated during inference. This framework supports dynamic knowledge retrieval and enhances performance on tasks with rare or previously unseen entities (Wu et al., 2022). Efforts such as GPT Cache aim to accelerate inference in LLM applications by caching and reusing responses in vector databases, improving latency and cost-efficiency while maintaining contextual relevance (Bang, 2023). Another key contribution comes from Liu et al. 2023 who investigated long-context understanding using LLMs like Claude and GPT-4. Their analysis identified that even state-of-theart models degrade with increasing context length, suggesting that without structured protocols like MCP, performance suffers over prolonged interactions (Liu et al., 2023). Similarly, other researchers explored grounded generations through search engines, introducing WebGPT, which uses external information to anchor model outputs. MCP could benefit from such grounding mechanisms to improve the accuracy and relevance of recalled context (Krishnan, 2025). A study introduced a framework where LLMs interact with external structured persistent memory banks. This aligns with MCP’s goals by defining clear memory read/write operations, crucial for context Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. persistence (Packer et al., 2023). Recent innovations like LongNet (Ding et al., 2023) scale context length to over 1 million tokens through dilated attention, facilitating linear computational scaling. This approach effectively balances memory efficiency and expressive capacity in LLMs. In a study, researchers highlight (Ehtesham et al., 2025) four agent communication protocols - MCP, ACP, A2A, and ANP - and their specific uses in decentralized agent networks, tool integration, multimodal messaging, and task delegation. They also recommended a phased adoption technique for these protocols, emphasizing how they can enhance security and flexibility in LLM-powered systems. However, the study does not provide any clear assurance of the protocols' long-term performance. In another study researchers (Singh et al., 2025) have explained the concept of MCP and how it supports unifying the context for large language models. This study demonstrated how Model Context Protocol (MCP) can resolve particular security rules and asynchronous interfaces, which are the two present barriers to AI integration. It additionally includes information about the MCP’s potential applications in healthcare and finance but also notes that there hasn't been any practical testing done. One study presents a thorough analysis of MCP from the perspective of communications systems, also unraveling its architecture, outlining its lifecycle and transport semantics, and comparing its uses across numerous domains (Ray, 2025). The security context of MCP is a critical area of study, for example, one study discusses the issues related to the technological implementation strategies and enterprise-level mitigation frameworks for the particular security vulnerabilities of MCP. It also identifies the gaps between the adaptation of AI system and transforming the theoretical issues related to the security into the practical field of study (Narajala & Habler, 2025). On the other hand, (Radosevich & Halloran, 2025) disclose fundamental MCP security flaws, illustrating the various ways in which superior LLMs can be compelled to access systems, and providing MCPSafetyScanner, a safety auditing tool to assess and mitigate these threats. Several research studies explored the possibilities and potential uses of MCP., for example, introduction of MCP Bridge, which is a low-cost RESTful proxy that connecting numerous Model Context Protocol (MCP) servers through a standardize API (Ahmadi et al., 2025). This technique fixes the issues related to MCP implementations, which frequently requires to perform local execution and may not be workable in resource-constrained settings such as edge computing and portable devices. One study suggests a framework for context management strategies, extensible coordination trends and an integrated framework (Krishnan, 2025) Model Context Protocol (MCP) has been suggested as an approach to the drawbacks of conventional data exchange protocols such as TCP/IP (Patil & Lokhande, 2025). This research Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. reveals that whereas current methods provide dependable data transport, they frequently fail to communicating contextual metadata, which results in inefficiencies in dynamic settings. MCP enhances decision-making in distributed systems by integrating contextual data, including location, time, and device condition, into communications. This study highlights how MCP can improve real-time adaptability but does not discuss challenges of the applications of MCP in depth. Lastly, papers like (Hou et al., 2025) and (Jiao et al., 2025) explore all the aspect of MCP, which includes its acceptance, implementation in future research trends. While (Hou et al., 2025) analyze the security and privacy issues related to MCP and offer mitigation strategies, (Jiao et al., 2025) introduce SafeMate, a Model Context Protocol-based multimodal agent designed for disaster preparedness, revealing MCP's potential in practical applications. These studies collectively underscore the fragmented but converging paths toward persistent context in LLMs. MCP represents a vital abstraction layer that could standardize how context is stored, retrieved, and utilized across sessions and applications. The provided abstracts delve into the realm of the Model Context Protocol (MCP), a standardized framework designed to enhance the integration and interoperability of large language models (LLMs) with external tools and systems. A significant theme across these studies is the exploration of MCP’s potential to revolutionize data communication and system interoperability, particularly in the context of artificial intelligence (AI) and its applications. 3.2 Literature Matrix The literature matrix presents a summary of key studies related to Model Context Protocol (MCP). It shows the goals, methods, main findings, and limitations of each work. This helps to understand the development, use, and challenges of MCP across different areas. The matrix also highlights how various techniques have been applied to improve MCP’s performance, security, and usefulness. Table 1: Literature matrix for key papers covered in literature review Study Goal Method Key Findings Limitations (Ahmadi et al., 2025) Create a crossplatform MCP Bridge. Implementation of RESTful proxy makes MCP possible in contexts with limited resources. Inadequate indicators of performance (Bang, 2023) Lower LLM API latency and costs Semantic caching (GPTCache) Cache hit dependency 2–10× Faster cached responses Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. Study Goal Method Key Findings Limitations (Borgeaud et al., 2021) Parameter reduction through retrieval Retrieval-enhanced Transformer (RETRO) GPT-3 matching with fewer parameters Reliance on the quality of the retrieval corpus (Dai et al., 2019) Improve the Transformers context Recurrence at the segment level and enhanced positional encoding Faster analysis and handling of longer dependencies Auto-regressive tasks only (Ding et al., 2023) Increase the length of the sequence to billions of tokens. Distributed training and focused attention Extended lengths are handled by linear complexity. Has to be verified on a variety of tasks. (Ehtesham et al., 2025) Explore the communication protocols used by AI. Comparison between MCP, ACP, A2A, ANP MCP is excellent at integrating tools, ANP is good at decentralized discovery, ACP is good at communicating, and A2A is good at task delegation. Absence of empirical validation (Hou et al., 2025) Examine the security lifecycle of MCP. Synopsis of literature Security threats during the development, use, and update stages of MCP Lack of testing for mitigation analysis (Jiao et al., 2025) Using MCP when responding to emergencies. A retrieval system based on FAISS Improves decisionmaking in emergency situations Limited testing in reality (Krishnan, 2025) Boost the coordination of several agents Protocol for Model Context (MCP) Standardizes the sharing of context Standardizes the sharing of context (Liu et al., 2023) Assess the use of long-context Analysis of positions in QA/kvretrieval Information about midcontext is degraded. Restricted to particular tasks Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. Figure 4: Script for MCP server and the configuration files Claude Desktop currently allows user to configure settings as a developer, as mentioned earlier. The configuration JSON file can be accessed from the developer settings. For implementing the MCP server for VuFind, the configuration file needs to be edited as shown in Table 2. The configuration file contains paths to the server script and the config.ini file. After configuration changes, the Claude Desktop must be restarted. Further functioning of the MCP server requires the server script to be running, and tools to be turned on for working. This successfully connects and allows proper functioning of the MCP server. Table 2: Claude Desktop configuration JSON file Name of the configuration file Configurations claude_desktop_config.json { "mcpServers": { "Vufind": { "command": "python", "args": [ "C:/Users/tirth/my-mcp/MCP-for-VuFind/server.py", "C:/Users/tirth/my-mcp/MCP-for-VuFind/config.ini" ] } } } Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. 6. Results: After proper configuration of the MCP server for VuFind with Claude Desktop, the retrieval results upon searching through the LLM as well as VuFind are observed. Claude Desktop allows users to retrieve information in natural language. Figure 5 demonstrates searching a book by title on VuFind, and Figure 6 shows usage of Claude Desktop for searching the same book in natural language. Figure 5: Retrieval of book by title using VuFind Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. Figure 6: Natural language retrieval of book by title using Claude Desktop An interesting leap in the age of library retrieval can be achieved, when the LLM is asked in different languages, and it successfully retrieves accurate results by analyzing the query. Multilingual retrieval breaks the language barrier which users may face while searching for library resources. Figure 7 displays the retrieval using Bengali language. Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. Figure 7: Bengali retrieval of books written by William Shakespeare using Claude Desktop It is observed that sorting of resources by relevance is different in both VuFind and Claude Desktop. As evident from figure 8 and 9, the sorting of resources varies in both cases. This opens up a question about how the LLM sorts the retrieved resources. Figure 8: Sorting of resources by relevance in VuFind Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. Figure 9: Sorting of resources by relevance in Claude Figure 10 and 11 shows the contradiction in retrieval by call number, where the LLM fails to retrieve, while VuFind successfully retrieves the resources. Figure 10: Retrieval of book by call number on VuFind Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. Figure 11: Failure in retrieval of book by call number using Claude Desktop 7. Conclusion: Model Context Protocol (MCP) offers an open standard for developing servers that enable Large Language Models (LLMs) to interact with external systems through defined tools. In this study, an MCP server was developed for VuFind, a library discovery system, allowing natural language queries through the Claude Desktop interface. This integration may enhance the user experience by enabling semantic search and response generation in natural language, thereby facilitating more intuitive access to library resources. Claude Sonnet 4, an LLM developed by Anthropic, was used as the natural language processing interface. While it is accessible for free through the Claude Desktop application, its usage is subject to limitations such as token caps and rate restrictions. These constraints may impact its applicability in large-scale or continuous information retrieval scenarios where unrestricted and high-volume access is required. Additionally, this study did not evaluate other capable LLMs - such as Llama, Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. DeepSeek, or Mistral - which offer open-source, locally deployable alternatives that could eliminate such usage limits. Since OpenAI and other major players began adopting MCP in early 2025, development activity around MCP has significantly increased. Not only have MCP servers been proposed for a variety of use cases, but MCP clients and bridges - such as those connecting with Ollama for running local LLMs - are also emerging. These developments point toward the feasibility of creating dedicated, domain-specific chatbots for library services using natural language interfaces, without depending solely on proprietary systems. The use of open-source tools and frameworks in this research demonstrates the potential for building flexible and extensible systems for NLP-based information retrieval in libraries. Moving beyond structured search interfaces like traditional OPACs, natural language-based search systems are better aligned with user behavior and expectations. This approach improves user satisfaction and supports more effective information seeking. As MCP-based tools continue to evolve, they may help establish a broader ecosystem for library-centric conversational agents, advancing the mission of libraries to meet users' information needs through innovative and user-friendly technologies. References: Ahmadi, A., Sharif, S. S., & Banad, Y. (2025, April 11). MCP Bridge: A Lightweight, LLM-Agnostic RESTful Proxy for Model Context Protocol Servers. https://www.semanticscholar.org/paper/MCP-Bridge%3A-A-Lightweight%2C-LLMAgnostic-RESTful-for-Ahmadi-Sharif/1caf450f88031258484dcb616daba512fdf78774 Bang, F. (2023). GPTCache: An Open-Source Semantic Cache for LLM Applications Enabling Faster Answers and Cost Savings. In L. Tan, D. Milajevs, G. Chauhan, J. Gwinnup, & E. Rippeth (Eds.), Proceedings of the 3rd Workshop for Natural Language Processing Open Source Software (NLP-OSS 2023) (pp. 212–218). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.nlposs-1.24 Bates, M. J. (1989). The design of browsing and berrypicking techniques for the online search interface. Online Review, 13(5), 407–424. https://doi.org/10.1108/eb024320 Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Driessche, G. van den, Lespiau, J.-B., Damoc, B., Clark, A., Casas, D. de L., Guy, A., Menick, J., Ring, R., Hennigan, T., Huang, S., Maggiore, L., Jones, C., Cassirer, A., … Sifre, L. (2022). Improving language models by retrieving from trillions of tokens (No. arXiv:2112.04426). arXiv. https://doi.org/10.48550/arXiv.2112.04426 Brants, T. (2004). Natural language processing in information retrieval. In Computational Linguistics in the Netherlands 2003: Selected Papers from the Fourteenth CLIN Meeting (pp. 1–13). Centre for Language and Speech Technology. https://clinjournal.org/CLIN_proceedings/XIV/brants.pdf Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., & Salakhutdinov, R. (2019). Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context (No. arXiv:1901.02860). arXiv. https://doi.org/10.48550/arXiv.1901.02860 Ding, J., Ma, S., Dong, L., Zhang, X., Huang, S., Wang, W., Zheng, N., & Wei, F. (2023). LongNet: Scaling Transformers to 1,000,000,000 Tokens (No. arXiv:2307.02486). arXiv. https://doi.org/10.48550/arXiv.2307.02486 Ehtesham, A., Singh, A., Gupta, G. K., & Kumar, S. (2025). A survey of agent interoperability protocols: Model Context Protocol (MCP), Agent Communication Protocol (ACP), Agentto-Agent Protocol (A2A), and Agent Network Protocol (ANP) (No. arXiv:2505.02279). arXiv. https://doi.org/10.48550/arXiv.2505.02279 Feldweg, H. (1999). Implementation and Evaluation of a German HMM for POS Disambiguation. In S. Armstrong, K. Church, P. Isabelle, S. Manzi, E. Tzoukermann, & D. Yarowsky (Eds.), Natural Language Processing Using Very Large Corpora (Vol. 11, pp. 1–12). Springer Netherlands. https://doi.org/10.1007/978-94-017-2390-9_1 Grefenstette, G. (1998). Problems and techniques of cross-language information retrieval. In Proceedings of the First International Conference on Language Resources & Evaluation (pp. 1455–1462). European Language Resources Association. https://aclanthology.org/www.mt-archive.info/LREC-1998-Grefenstette-2.pdf Hou, X., Zhao, Y., Wang, S., & Wang, H. (2025). Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions (Version 2). arXiv. https://doi.org/10.48550/ARXIV.2503.23278 jaohbib. (2025). Jaohbib/MCP-for-VuFind [Python]. https://github.com/jaohbib/MCP-for-VuFind (Original work published 2025) Jiao, J., Park, J., Xu, Y., & Atkinson, L. (2025). SafeMate: A Model Context Protocol-Based Multimodal Agent for Emergency Preparedness (No. arXiv:2505.02306). arXiv. https://doi.org/10.48550/arXiv.2505.02306 Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., & Lewis, M. (2020). Generalization through Memorization: Nearest Neighbor Language Models (No. arXiv:1911.00172). arXiv. https://doi.org/10.48550/arXiv.1911.00172 Krishnan, N. (2025). Advancing Multi-Agent Systems Through Model Context Protocol: Architecture, Implementation, and Applications (No. arXiv:2504.21030). arXiv. https://doi.org/10.48550/arXiv.2504.21030 Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2021). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (No. arXiv:2005.11401). arXiv. https://doi.org/10.48550/arXiv.2005.11401 Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2023). Lost in the Middle: How Language Models Use Long Contexts (No. arXiv:2307.03172). arXiv. https://doi.org/10.48550/arXiv.2307.03172 Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval (1st ed.). Cambridge University Press. https://doi.org/10.1017/CBO9780511809071 Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. McNulty, N. (2025, March 9). The Complete Guide to Model Context Protocol. Medium. https://medium.com/@niall.mcnulty/the-complete-guide-to-model-context-protocol148dca58f148 MSV, J. (2024, November 30). Why Anthropic’s Model Context Protocol Is A Big Step In The Evolution Of AI Agents. Forbes. https://www.forbes.com/sites/janakirammsv/2024/11/30/why-anthropics-model-contextprotocol-is-a-big-step-in-the-evolution-of-ai-agents/ Narajala, V. S., & Habler, I. (2025). Enterprise-Grade Security for the Model Context Protocol (MCP): Frameworks and Mitigation Strategies (No. arXiv:2504.08623). arXiv. https://doi.org/10.48550/arXiv.2504.08623 Packer, C., Wooders, S., Lin, K., Fang, V., Patil, S. G., Stoica, I., & Gonzalez, J. E. (2024). MemGPT: Towards LLMs as Operating Systems (No. arXiv:2310.08560). arXiv. https://doi.org/10.48550/arXiv.2310.08560 Patil, M., & Lokhande, V. (2025). Model Context Protocol (MCP): Enabling Scalable AI Data Integration. International Journal For Multidisciplinary Research, 7(2), 43583. https://doi.org/10.36948/ijfmr.2025.v07i02.43583 Patil, P. (2025). Inside the MCP Protocol: Revolutionizing data communication and system interoperability. World Journal of Advanced Research and Reviews, 26(1), 3055–3071. https://doi.org/10.30574/wjarr.2025.26.1.1401 Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., & Riedel, S. (2019). Language Models as Knowledge Bases? (No. arXiv:1909.01066). arXiv. https://doi.org/10.48550/arXiv.1909.01066 Radosevich, B., & Halloran, J. (2025). MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits (No. arXiv:2504.03767). arXiv. https://doi.org/10.48550/arXiv.2504.03767 Rae, J. W., Potapenko, A., Jayakumar, S. M., & Lillicrap, T. P. (2019). Compressive Transformers for Long-Range Sequence Modelling (No. arXiv:1911.05507). arXiv. https://doi.org/10.48550/arXiv.1911.05507 Ray, P. P. (2025). A Survey on Model Context Protocol: Architecture, State-of-the-art, Challenges and Future Directions. https://doi.org/10.36227/techrxiv.174495492.22752319/v1 Robertson, S., & Zaragoza, H. (2009). The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends® in Information Retrieval, 3(4), 333–389. https://doi.org/10.1561/1500000019 Singh, A., Ehtesham, A., & Kumar, S. (2025). A Survey of the Model Context Protocol (MCP): Standardizing Context to Enhance Large Language Models (LLMs). https://doi.org/10.20944/preprints202504.0245.v1 Teevan, J., Dumais, S. T., & Horvitz, E. (2005). Personalizing search via automated analysis of interests and activities. Proceedings of the 28th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 449–456. https://doi.org/10.1145/1076034.1076111 Voorhees, E. (1999). Natural Language Processing and Information Retrieval. https://tsapps.nist.gov/publication/get_pdf.cfm?pub_id=151418 Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19. Wikipedia contributors. (2025, May 22). Model Context Protocol. In Wikipedia, The Free Encyclopedia. https://en.wikipedia.org/wiki/Model_Context_Protocol Wu, Y., Rabe, M. N., Hutchins, D., & Szegedy, C. (2022). Memorizing Transformers (No. arXiv:2203.08913). arXiv. https://doi.org/10.48550/arXiv.2203.08913 Postprint Neogi, M., Dasgupta, T., & Mukhopadhyay, P. (2025). NLP-based Library Retrieval: Integrating VuFind with LLM through MCP Middleware. Indian Journal of Information Library & Society, 38(1–2), 1–19.