Small-scale Domain-specific Web Crawling for Complementing Established LLM Data Sources
Abstract
Poster for OSC25 presenting the domain-specific web crawling infastructure of the Leipzig Corpora Collection (Wortschatz Leipzig) and its relevance to complementing open-access data sources for LLMs and other machine learning models.
Full text
OSCAR 22.01 + 23.01 German LCC Crawl 2022 German LCC Crawl 2023 German Source URL Overlaps 4.04% 4.57% 20.33% 572,167,630 643,637,710 172,075,879 Total Number of Source URLs LCC DE 2022 LCC DE 2023 OSCAR DE 22.01+23.01 2 - https://oscar-project.org We compare the LCC's annual news and web crawls of the .de top-level domain with the German subset of the OSCAR² corpus, a popular resource for machine learning and artificial intelligence applications, for the same crawling period. Using identical content extraction and sentence-level deduplication procedures on both resources and considering their levels of content density, we establish solid estimates of their degree of complementarity. The German LCC crawl of 2023 (396 billion tokens in total) yields at least 35.9 billion tokens of cleaned text data deduplicated at document-level, with around 5% overlap with the German OSCAR subset of 2022 and 2023 and around 20% overlap with the LCC data of the previous year’s crawling round. Overlaps were calculated by comparing the URL source lists of the three datasets with each other, as a heuristic for general content overlap. This demonstrates that: –Small-scale crawling infrastructures that ensure complete control over source selection and thematic focus can still provide a significant gain in LLM training material compared to established resources. –Established training datasets based on crawling are still far from achieving a 100% coverage of the Web. –Regular re-crawls (e.g., on an annual basis) continue to result in significant growth in new data and potential training material. Dataset Comparisons Processing Impacts on Token Counts (in Millions) Raw After cleaning After deduplication 0 50,000 100,000 150,000 200,000 250,000 300,000 350,000 400,000 450,000 LCC DE 2022 LCC DE 2023 OSCAR-DE –Training of a German foundation model at the Center for Scalable Data Analytics and ArtificialIntelligence (ScaDS.AI). –Constrained Retrieval-Augmented LLMs (CORAL): Research project examining LLMs under legal, technical, resource-based and data-based constraints. –Multilingual, day-to-day data to monitor European news and for digital press reviews. Data Usage for Large Language Models Since June 2021, we annually collect over 130 billion tokens of cleaned documents (or ~27-35 billion tokens after sentence-wise deduplication), which can be freely used as current, diverse, high-quality and sourceable training data for LLMs and other forms of text and data mining. Use-cases include: Crawling Cleaning Enrichment DATA SOURCES https://wortschatz-leipzig.de For more than 30 years, the Leipzig Corpora Collection (LCC), or Wortschatz Leipzig, provides digital corpora and corpus-based dictionaries. Currently, data for more than 250 languages can be accessed and downloaded. For many of those languages, the project provides the largest freely available text resources on the web. This is possible due to the LCC's own crawling infrastructure and processing pipeline, both of which were developed and grew with the project. These processes were built to enable crawling with limited hardware and the creation of smallerand large-scale datasets, depending on availability of sources and resources. This infrastructure allows for the targeted creation of domain-specific corpora, many of them the only ones of their kind. The LCC data processing pipeline can be divided into three major steps: 1) Crawling and scraping of web page text, using the Internet Archive's open-source crawler Heritrix¹. 2) Cleaning of the raw HTML data, i.e. text extraction, language separation, removal of non-text and other artefacts, sentence and token splitting, etc. 3) Various data enrichments are added in subsequent post-processing steps: different annotations, like part-of-speech tags, and language statistics such as word frequencies and co-occurrences data. WORTSCHATZ LEIPZIG 1 - https://github.com/internetarchive/heritrix3 LCC Web Crawling Small-scale Domain-specific Web Crawling for Complementing Established LLM Data Sources The NFDI consortium Text+ is funded by the German Research Foundation (DFG) – Project number 460033370 Universität Leipzig / InfAI / ScaDS.AI Christopher Schröder Institut für Angewandte Informatik (InfAI) Frank Binder Sächsische Akademie der Wissenschaften zu Leipzig Thomas Eckart, Felix Helfer, Erik Körner Participants