scieee AI-readable full text Open interactive document viewer

Legal Aspects of AI Training and Retrieval Augmented Generation

Beer, Leopold; Johannes, Paul; Koulani, Huda

Abstract

Appeared in: Open Search Symposium 2025, 8-10 October 2025, CSC IT Center for Science, Helsinki, Finland.

Full text

LEGAL ASPECTS OF AI TRAINING AND RETRIEVAL AUGMENTED GENERATION* L. Beer†, Open Search Foundation e.V., Munich, Germany P. C.Johannes§, H.Koulani¶, ITeG, University of Kassel, Germany Abstract This paper explores legal aspects of the development of the Open Web Index (OWI), a publicly funded European initiative designed as an alternative to proprietary web indexes. It examines how the OWI supports the training of Large Language Models (LLMs) and enhances Retrieval Augmented Generation (RAG) systems. The discussion covers the OWI’s architecture, including its distributed crawling and indexing methods, which allow for the collection of vast amounts of web data. By making high-quality, accessible data available, this open infrastructure could benefit smaller companies and research institutions that might otherwise struggle to compete with larger players. The paper delves into the regulatory landscape within the European Union, particularly in relation to the AI-Act and copyright law. It considers the legal challenges surrounding the OWI’s use in LLM training and RAG, emphasizing the importance of data quality, legal compliance, and public trust. The conclusion highlights key areas for future research, including the need to clarify frameworks for rights of use, consent and processing authorisation for data to addresses legal uncertainties. INTRODUCTION Artificial intelligence (AI) has become widely adopted in a variety of areas at lightning speed and has become an integral element of science, business and society. RAG, i.e. the combination of search and a generative component, is no longer a term that only experts understand. In fact, there are numerous RAG systems on the market that are used by millions of people in the EU and elsewhere every day[1]. At the same time, the European Union (EU) has adopted a large number of regulations and directives in the area of digital governance that impose numerous obligations on the developers and users of such applications. Further legislation is also planned for the future to ensure fairness in the digital space and guarantee the competitiveness of European companies. In his report for the EU Commission on the future of European competitiveness, Mario Draghi brings up the problem of extensive regulation: “innovative companies that want to scale up in Europe are hindered at every stage by inconsistent and restrictive regulations”[2]. Deregulation is often a part of the demands by economic interest groups as the many legal requirements are difficult to keep track of, especially for small and medium-sized companies, and therefore might hinder innovation. The PriDI (Privacy enhancing digital infrastructures) research project, has set itself the goal of making the complex digital legislation at national and EU level easy to understand for developers of digital applications[3]. To this end, the consortium of the University of Kassel and the Open Search Foundation is analysing user-related and legal requirements for the training of LLM’s and RAG systems if these applications are based on data from the so-called OWI. Such an index is currently being developed by the EUfunded research project OpenWebSearch.EU[4]. This paper briefly describes the creation and curation of the OWI. It then provides key information on the processes of LLM training and the operation of RAG systems. The authors then shed light on the legal framework within the EU in which the developers of such applications operate. A particular focus is put on the obligations arising from the regulation of AI and the intellectual property rights of the owners of website content. In addition, the paper analyses requirements that arise from the user's perspective with regard to trust and acceptance of AI systems that are based on the OWI. It concludes with an outlook on further research work. THE OPEN WEB INDEX Similar to the Google or Bing indexes, the Open Web Index is being created by systematically crawling the web, analysing the crawled content and storing it with metadata in a database[5]. The OWI is intended to strengthen the EU's digital sovereignty by reducing dependence on the search engine monopolists through a sustainable, freely accessible web index. The researchers have set up a distributed crawling, indexing and hosting architecture for the OWI. This consists of the combination of a frontier crawler, that basically charts the web along embedded links and collects URLs, and distributed worker crawlers, that later on fetch the websites and store the content in so called web archive (WARC) files. Later on, the “raw” web data is further processed, cleaned, filtered, enriched with metadata, classified according to ______________________________________________ * Based on research of project Privacy-enhancing digital infrastructures (PriDI), funded by the German Federal Ministry of Education and Research (BMBF). The responsibility for the content of this publication lies with the authors. Parts of tthe findings presented in this paper have already been submitted for publication to CPDP.AI 2025. † [email protected]g §[email protected] ¶ [email protected] https://doi.org/10.5281/zenodo.17229684 language and web genre and stored as web index charts following the common index file format (CIFF) and additional metadata sets. The system is designed in a way that it federates storage and computing capacities across several high performance computing centres across Europe and can be dynamically extended with additional computing centres being added to the federation. To access the index, providers of LLMs and RAG systems or other scientific users of the OWI can authenticate themselves via a public system and can access and retrieve parts of the index via a command line tool. Currently the web data is made available under a research licence, but the research team is also working to grant access to the system for commercial purposes. The public accessibility of the index is intended to strengthen freedom in internet searches and to form a basis for innovations in science and economy. The researchers have now crawled around 2.23 billion URLs in 185 different languages. The Open Web Index currently has a volume of around 14 TB and is already available to interested developers for initial tests. However, Google’s index with a volume of around 100.000 TB is much larger as it includes also thumbnails and other data, whereas the OWI currently includes text data only[6]. ACTORS When assessing the Open Web Index and its use cases for LLM training and RAG from a legal and user acceptance perspective, a distinction must be made between different stakeholders. Firstly, there are the data subjects. This role describes persons or companies whose personal data and intellectual property are stored in the OWI and are used or may be accessed by tools based on the OWI. The index itself is developed and maintained by the OWI developer. The OWI developers have joined together to form an independent legal entity, the operator consortium. The data retrievers or data consumers of the information contained in the OWI are referred to as application developers. They are persons or systems that request the retrieval of web data from the index in order to create and develop various tools and models based on the retrieved data. They can be individuals, organisations, companies, public institutions or startups that use OWI's open data to develop their own applications and services. Finally, end users also come into contact with the index. They are the natural or legal persons who use the tools and systems developed by the application developers. USE CASES OF THE OWI The Open Web Index can be used in various ways, e.g. as a basis for search engines (see [7] and [8] for more details). This paper focuses on the use of the index’s data to train AI-Systems, i.e. the training of LLM and the development of RAG systems. Training of LLMs For the training of LLMs, comprehensive and highquality pre-filtered web data are essential. A common example of such an LLM is Mistral Large 2 by the French company Mistral AI, on which the company’s chatbot is based. These models are trained with large amounts of text data, which mainly come from online sources. With the help of machine learning using neural networks and deep learning methods, LLMs learn to recognise statistical relationships between words and sentences in order to understand and generate texts. The OWI contributes to the development of new LLMs by providing smaller companies with a sufficient amount of training data at low cost. By offering an open and transparent alternative to proprietary datasets, the OWI enables start-ups and research institutions to train their own models without relying on a few dominant players in the field. This not only fosters innovation and diversity in AI development but also promotes fair competition. As a result, smaller companies can develop high-quality products that can compete with the leading language models currently available on the market. Retrieval Augmented Generation RAG is an advanced approach in AI that enhances text generation models by integrating an information retrieval component. This method combines the generative capabilities of language models with the precision of retrieving relevant data from external knowledge sources such as databases, documents, or the web. The first key component of RAG is the language model, which is trained on vast amounts of text data to understand language and generate coherent responses. This model serves as the foundation for answering queries and can be trained with OWI data. The second component, retrieval, dynamically accesses a knowledge base to fetch relevant information in real time. The OWI can serve as such knowledge base. By combining these two elements, RAG enables AI systems to produce responses that are not only contextually accurate but also based on the most current available data. This approach is particularly valuable in areas where precise and up-to-date information is crucial, such as customer support, medical consultation, legal research, and knowledge management. By overcoming the limitations of static training data, RAG ensures that AIdriven solutions remain relevant and reliable, even in rapidly evolving fields. LEGAL ASPECTS Relevant legislation The complexity of the Open Web Index raises many questions as to how the index and its applications fall https://doi.org/10.5281/zenodo.17229684 under Union and member state law. European law on data and online services has undergone major changes in recent years. Where initially mainly the General Data Protection Regulation (GDPR, Regulation (EU) 2016/679) laid down detailed rules on the handling of personal data, now a network of more or less specialized, directly applicable legal acts has emerged. The Digital Services Act (DSA, Regulation (EU) 2022/2065), the Data Governance Act (DGA, Regulation (EU) 2022/868), the Data Act (DA, Regulation (EU) 2022/868), the Digital Markets Act (DMA, Regulation (EU) 2022/1925) and the AI Act (AIA, Regulation (EU) 2024/1689) form a legal framework for digital services and business models[9]. These regulations are directly and uniformly applicable in all member states. At the same time, there is a harmonized copyright law framework in the European Union. It is primarily governed by a combination of EU Directives, international treaties, and national laws of member states. While the EU aims to harmonize copyright laws across its member states, variations still exist at the national level. This article focuses on legal challenges faced by the OWI developer and the LLM and RAG application developers across the regulatory domains AI regulation and copyright. The initial assessment of the use case highlights the complexities and the need for ongoing studies. Regulation of AI-Systems The AIA aims to regulate systems and practices in the field of artificial intelligence. It was adopted in order to create a robust and flexible legal framework that makes the use of AI and automated decisionmaking systems trustworthy and secure. The AIA introduces a uniform framework for AI systems based on a risk-based approach, see Recital26 AIA. AI systems are in Article3 No.1 defined as amachinebased system that is designed to operate with varying levels of autonomy and that may exhibit adaptiveness after deployment, and that, for explicit or implicit objectives, infers, from the input it receives, how to generate outputs such as predictions, content, recommendations, or decisions that can influence physical or virtual environments. The higher the risk, the more substantial are the obligations put on operators(see Article3 No.8 AIA for definition) of AI systems. AI systems with unacceptable risks, e.g. systems that allow “social scoring” by governments or companies, are considered a clear threat to people's fundamental rights and are therefore banned pursuant to Article5 AIA. To address their specific transparency risk, AI systems like chatbots must clearly inform users that they are interacting with a machine, while certain AI-generated content must be labelled as such, see Article50 AIA. Only a few AI systems with limited risk face no obligation under the AIA. High-risk AI systemsaccording to Article6 AIA and Annex I and II AIA on the other hand, such as AI-based medical software or AI systems used for recruitment, must comply with strict requirements, including riskmitigation systems, high-quality of data sets, clear user information, human oversight. Direct applicability of AIA The developer of the OWI would have to examine to what extent the provisions of the AIA directly apply to the technologies used to facilitate the OWI. It’s plausible that the algorithms utilized by the OWI to assist and coordinate web crawling will be categorized as minimal risk, as they primarily focus on internal operations. However, to conclusively establish this, a thorough risk assessment is required. Following the risk-based approach, AI systems “that may have a significant adverse impact on the health, safety and fundamental rights of persons” (Recital46 AIA) are classified as high-risk AI systems in Article6 AIA, whereby a distinction is made between high-risk AI systems in connection with product regulation (para.1) and stand-alone high-risk AI systems (para.2). AI systems that are safety components of products (Article3 No.14 AIA) or are themselves products covered by the harmonization legislation in AnnexII (e.g. machinery, toys, elevators, radio equipment, cableways, medical devices, motor vehicles and aircraft) are deemed high-risk systems. The OWI would probably not be classified as a high risk systems pursuant to Article6 para.1 AIA, since it is not used as a safety component covered by Annex I of the AIA. Article6 of the AIA also designates high-risk AI systems as those enumerated in Annex III. This includes AI systems in biometrics, critical infrastructures, education, employment, basic services, law enforcement, migration, asylum and border control, as well as the administration of justice and democratic processes. AI used in indexing, as well as determining the exclusion or inclusion of certain website content, generally have tangible external impacts on the index usage by third parties. However, it remains unlikely that these operations, or the entirety of the index, might be classified under one of the sectors specified in AnnexIII, with the exception of the "critical infrastructures" listed as No.2. Indirect applicability Furthermore, it must be asked if the AIA contains regulations that influence the use of OWI data for specific AI applications. For example, certain AI systems are banned under Article 5 AIA. The OWI developer might therefore seek to prohibit the use of its index for the training of such AI systems with unacceptable risks. This could be achieved by means of the index licence or terms of conditions for using the OWI as a service. The OWI operator would have certain leeway, since the prohibition clause does not prevent scientific research into the use of AI system merely capable of prohibited practices. Furthermore, according to Article2 para.6 AIA, the regulation does not apply to AI systems or AI models, including their outputs, which are developed https://doi.org/10.5281/zenodo.17229684 and put into operation for the sole purpose of scientific research and development, see Article3 No.11 AIA. This is intended to promote innovation and protect scientific freedom, see Recital25 AIA. In accordance with Article13 Charter of Fundamental Rights of the European Union, scientific research includes activities with the aim of “gaining new knowledge in a methodical, systematic and verifiable manner”. This includes basic research and applied research in the public (e.g. universities) and private (e.g. industrial research) sectors. Development includes the application and implementation of the knowledge gained through research. Still, the exception is to be interpreted narrowly in terms of wording and well as meaning and purpose. Another example would be, that pursuant to Article10 para.2-5 AIA in conjunction with Article10 para.1 AIA high-risk AI systems must be developed with training, validation and test data sets that meet the certain quality criteria. Article10 para.3 of the AIA stipulates that training, validation and test data sets must be relevant, sufficiently representative and, as far as possible, error-free and complete with regard to the intended purpose. Among other things, it is questionable whether legally erroneous data (e.g. data obtained in violation of data protection or copyright law) or data anonymized or pseudonymized for data protection reasons (e.g. due to added noise) can still be considered error-free and complete[10]. The data records must also have the appropriate statistical characteristics, if necessary also with regard to the persons or groups of persons for whom the high-risk AI system is to be used as intended. OWI data provides a massive and diverse source of information, including text and links. This diversity is crucial for training AI models. In order to be usable under the quality criteria for high-risk AI systems, the OWI developer should use specific techniques to ensure data quality, e.g. data cleaning, augmentation, balancing or annotation. At the same time it could try to create its datasets in a way, subsequent application developers could use or build on to ensure data quality for their specific use case. Copyright Within the European Union, intellectual property rights are mainly determined by European law, but are implemented mostly at national level. For copyright law, which is particularly relevant in the context of LLM training and RAG, the EU legislator has adopted provisions in the Copyright Directive (2001/29/EC) and the Directive on Copyright in the Digital Single Market (2019/790), which have been implemented into national law by the member states. In the following, the legislation in Germany (mainly the German Act on Copyright and Related Rights – UrhG) is taken as an example. The UrhG defines the extent to which copyrightprotected content may be indexed in the OWI and used by applications based on the index. The OWI contains a large amount of data, most of which are protected works under Section2 UrhG. These works are reproduced regularly as part of their inclusion in the index. However, the right of reproduction defined in Section16 UrhG is, in principle, granted to the author of the work and not to the OWI or application developers in accordance with Section15 para.1 No.1 UrhG. Copyright-inducing, at least temporary, reproductions cannot be avoided when creating the index and training LLMs with the index data. However, the Copyright Act contains various exceptions that can justify acts of reproduction. In 2021, the German legislator created the exception rule of Section44b UrhG for general text and data mining in implementing Article4 of the Directive No.2019/790. This is designed to make it possible to analyse large amounts of digital information[11]. According to the legal academia[12-20], the training of AI models can usually be justified by Section44b UrhG and case law also shows a slight tendency in this direction[21]. However, any reservations of the creator pursuant to Section44b para.3 UrhG must be taken into account. If the rights holder opposes to the use of their website content for LLM training or RAG development, the application and OWI developers must adhere to the content owner’s reservations. In the case of acts of reproduction created by crawlers during the indexing of web content, Section44a UrhG also comes into question. In this, the legislator provides for an exception for acts of reproduction that are only of a temporary nature and part of a technical process, have no independent economic significance and serve a purpose of Section44a UrhG. USER ACCEPTANCE AND TRUST While legal compliance, data protection, and privacy are highly valued in Europe, user behaviour suggests that these factors often take a backseat when choosing digital services. In practice, other aspects tend to have a stronger influence on users to adapt new technologies. Many research studies focused on investigating concepts and aspects of user perception and their intention to use and trust technologies[22-26]. Therefore, incorporating key elements of user acceptance into the development of OWI and related tools is essential. To apply user acceptance and trust principles effectively to OWI, our study combines insights from Trust-TAM[26] and UTAUT2[24]. The latter identifies seven key factors—such as performance and effort expectancy, social influence, and facilitating conditions—all of which we aim to adapt to the OWI framework. In terms of establishing trust with developers as actors in these use cases, knowledge-based trust could be conceived introducing familiarity as an antecedent of this trust. Supporting the developers’ familiarity with the structure and functionalities of OWI promote their confidence and https://doi.org/10.5281/zenodo.17229684 trust in the interaction with OWI. Particularly, this familiarity could be implemented in terms of providing web data for LLM training in a format that is standard in these contexts and thus reducing cognitive load required to acquire web data from OWI for the discussed matters. This adaptation is still in its early stages and represents a work in progress, with the mentioned ideas serving as an initial foundation for further development. FUTURE WORK Regardless of the previously conducted analysis of the legal framework, some issues in the context of the use of the OWI for LLM training and the development of RAG are still open. In future works, the interaction, including contractual relationships, between the developer of the OWI and the developer of LLMs and RAG would need to be clarified. Any legal loopholes or unintended consequences de lege lata should be addressed by further developing the law, either on the Union level, or where possible, on the national level. In this context, the focus should be on creating simple and concise provisions that are easy for developers to implement. The PriDI project will also focus on this in its future work and specify the legal requirements outlined above in the form of requirement and design patterns. The OWI can also be used to train other web databased AI applications, such as knowledge representation and reasoning (KRR) systems. In contrast to LLMs, which use statistics to produce texts, KRR systems represent information in a way that a computer can understand it and solve complex problems like a human. The aim is to create intelligent machines that learn from human knowledge and act in the same way. KRR systems are used, for example, in quality management to monitor product quality or to prevent fraud in the insurance industry. Future publications will need to consider whether the above legal requirements also apply to KRR systems. CONCLUSION The OWI data can be used for the training of LLMs and RAG development, since its data would be comprehensive and high-quality, pre-filtered web data. While the OWI operator conceivably would not fall under the AIA, the LLM and RAG developer most likely would. Depending on its specific use case, the LLM or RAG system could even be classified as a high-risk AI system. Either way, the quality of the provided data as well as the legality of its content and its provision are paramount. Even if the AIA would not be applicable for the specific use case (e.g. because of Article 2 para. 6 or 12 AIA), data protection law as well as copyright law most certainly still would. In regards to data protection law, the OWI developer would have to make sure that it is allowed to share or publish the personal data it has collected. The LLM and RAG developer would have to make sure, that it is allowed to process the personal data on the basis of one of the authorisations in Article6 GDPR. For example: The AI or RAG developer most likely would be allowed to process publicly available personal data of persons linked to a business under Article6 para.1 subpara.1 lit.f GDPR. Likewise, in regards to copyright, the LLM and RAG developer would still have to make sure, it could use the provided data. Ideally it could rely on a legal limitation to the copyright, like Section44b UrhG. REFERENCES [1] According to DemandSage, https://www.demandsage.com/perplexity-aistatistics/ the RAG-systems Perplexity AI has about 2 million daily active users. [2] European Commission, https://commission.europa.eu/topics/eucompetitiveness/draghi-report_en, p.6. [3] For ore information on the research project see its website https://pridi-projekt.de/home-en/. [4] For more information on the OWS.EU project see its website https://openwebsearch.eu/. [5] More details on the creation and operation of the OWI can be found in G.Hendriksen et al., “The Open Web Index. Crawling and Indexing the Web for Public Use”, in Advances in Information Retrieval: 46th European Conf. on Information Retrieval. ECIR 2024, Glasgow. UK. March 2024. pp.130-143. [6] Google, https://www.google.com/intl/en_us/search/ho wsearchworks/how-search-works/organizinginformation/. [7] Several other applications are listed in D.Nowakowski, N.Zimmermann and L.Kerner, “Market potential assessment of OpenWebSearch.eu - Exploring the economic and societal impact of an Open Web Index”, Mücke/Roth, Germany, Rep., 2024, https://openwebsearch.eu/wp-content/uploads /2024/09/MarketAssessmentOfOWI-ReportV1.pdf [8] P.C.Johannes, L.Beer and H.Koulani, “Legal challenges of using the OWI for search engines”, presented at OSSYM 2025 - 7th Int. Open Search Symposium, Helsinki, Finland, Oct. 2025, paper #####, this conference, submitted. [9] Also see C.Geminn and P.C.Johannes, Handbuch europäisches Datenrecht. Baden-Baden, Germany: Nomos, in preparation. [10] I.Vogel et al, “Natural Language Processing (NLP) und der Datenschutz - Chancen und Risiken für den Schutz der Privatheit”, in Informatik 2022. Lecture Notes in Informatics (LNI), p./659. doi: 10.18420/inf2022_24 [11] German Bundestag https://dserver.bundestag.de/btd/19/274/192 7426.pdf [12] E.g. D.Bomhard and J.Siglmüller, “AI Act - das https://doi.org/10.5281/zenodo.17229684 Trilogergebnis”, Recht Digital, vol. 5, no. 2, p.50, 2024. [13] M.Dregelies, “KI-Training unter dem AI Act”, Gewerblicher Rechtsschutz und Urheberrecht, vol. 126, no. 20, p.1484, 2024. [14] R.Heine, “Generative KI: Nutzungsrechte und Nutzungsvorbehalt”, Gewerblicher Rechtssschutz und Urheberrecht in der Praxis, vol. 16, no. 4, p.88, 2024. [15] F.Hofmann, “Retten Schranken Geschäftsmodelle generativer KI-Systeme?”, Zeitschrift für Urheberund Medienrecht, vol. 68, no. 3, p.166, 2024. [16] L.Kaede, “Training generativer KI-Modelle ist (auch) Textund Data-Mining. Anwendbarkeit der TDMSchranke des §44b UrhG”, Künstliche Intelligenz und Recht, vol. 1, no. 5, p.162, 2024. [17] N.Maamar, “Urheberrechtliche Fragen beim Einsatz von generativen KI-Systemen”, Zeitschrift für Urheberund Medienrecht, vol. 67, no. 7, p.481, 2023. [18] K.Wagner, “Generative KI: Eine “Blackbox” urheberrechtlicher Haftungsrisiken?. Balanceakt zwischen Innovationsförderung und effektivem Rechtsschutz für Werke Dritter”, Zeitschrift für IT-Recht und Recht der Digitalisierung, vol. 27, no. 4, p.298, 2024. [19] Contrary opinion: T.W.Dornis and S.Stober, Urheberrecht und Training generativer KI-Modelle: Technologische und juristische Grundlagen. BadenBaden, Germany: Nomos 2024. doi: 10.5771/9783748949558 [20] Contrary opinion: T.W.Dornis, ”Generatives KITraining und Textund Data-Mining. Eine funktionale Unterscheidung”, Künstliche Intelligenz und Recht, vol. 1, no. 5, p.156, 2024. [21] LG Hamburg, Judgement of 27 September 2024, Reference 310 O 227/23, https://www.itm.nrw/wpcontent/uploads/2024/09/2495651-en.pdf [22] F.D.Davis, “User acceptance of information technology: system characteristics, user perceptions and behavioral impacts” International Journal of ManMachine Studies, vol. 38, no. 3, pp.475–487, Mar. 1993. doi: 10.1006/imms.1993.1022 [23] V. Venkatesh, M. Morris, G. Davis, and F. D. Davis, “User acceptance of information Technology: toward a unified view,” MIS Quarterly, vol. 27, no. 3, p. 425, Jan. 2003. doi: 10.2307/30036540 [24] V.Venkatesh, J.Thong, and X.Xu, “Consumer Acceptance and use of Information technology: Extending the unified theory of acceptance and use of technology” MIS Quarterly, vol. 36, no. 1, p.157, Jan. 2012. doi: 10.2307/41410412 [25] M.Söllner, A.Hoffmann, and J.M.Leimeister, “Why different trust relationships matter for information systems users” European Journal of Information Systems, vol. 25, no. 3, pp.274–287, Dec. 2015. doi: 10.1057/ejis.2015.17 [26] D.Gefen, E.Karahanna, and D.Straub, “Trust and TAM in online shopping: an integrated model” MIS Quarterly, vol. 27, no. 1, p.51, Jan. 2003. doi: 10.2307/30036519 https://doi.org/10.5281/zenodo.17229684