scieee AI-readable full text Open interactive document viewer

Towards the Extraction of Location References and Topics from Semi-Structured Textual Data from the Open Web Index using Open-Source Large Language Models

Gadziomski, Patryk; Rittlinger, Vanessa; Pfeffer, Magnus; Voigt, Stefan

Abstract

Appeared in: Open Search Symposium 2025, 8-10 October 2025, CSC IT Center for Science, Helsinki, Finland.

Full text

TOWARDS THE EXTRACTION OF LOCATION REFERENCES AND TOPICS FROM SEMI-STRUCTURED TEXTUAL DATA FROM THE OPEN WEB INDEX USING OPEN-SOURCE LARGE LANGUAGE MODELS P. Gadziomski*1, V. Rittlinger†2, M. Pfeffer‡1, S. Voigt§2 1Media University (HdM), Stuttgart, Germany 2German Aerospace Center (DLR), Earth Observation Center, Oberpfaffenhofen, Germany Abstract With the steadily growing relevance of the Web as information source and the associated increase in web content, the systematic extraction of structured information from it is becoming increasingly important. Every day, thousands of social media posts and news articles are published, containing not only thematic but also geo-spatial information such as location names and addresses. These data are relevant for numerous applications and research fields—including open-data projects like OpenStreetMap, the optimization of search engine indices, or their use for example in crisis management. The automated extraction of addresses from web sources—particularly from imprint pages—could efficiently capture legal information, such as compliance with the General Data Protection Regulation (GDPR) or improve deep learning models for Named Entity Recognition (NER). Despite the high relevance of this task, existing methods for address extraction from text have so far yielded only limited results due to inconsistent formatting of addresses, the ambiguity of words, and the embedding of addresses in unstructured texts. Since rulebased methods for address extraction achieve only limited quality, the use of Large Language Models (LLMs) is proposed as a promising alternative to specifically extract addresses from imprint pages. Since the release of GPT-3, LLMs have enabled significant advancements in various fields, particularly in automated text processing. Information extraction, as a subfield of Natural Language Processing (NLP), is gaining increasing relevance due to LLMs availability and functionality and becomes an active research topic. Given the continuous evolution of these models, this trend is expected to persist. This applies both to the technical developments of LLMs and to advancements in prompting methods, and refinement of model output through targeted inputs. The data used in this study originates from the "Legal" datasets of the Open Web Index of the OpenWebSearch.EU projects, providing substantial amounts of imprint data. The data is restricted to German language. The extracted dataset is annotated using LLMs, followed by manual correction. The result is the creation of an annotated dataset with manually collected “gold standard” of geo-location samples for comparison and quality assessment. The hit rate of the LLM is documented to establish a well-founded basis for further work. The geo-localization results are documented to compare with different model outputs and the different applied prompting techniques. Following the extraction, an evaluation is conducted to determine at which level the models can extract relevant geo-information. Addresses consist of country, postal code, city, street name, and house number. A specific score is assigned to each of these elements. This metric is designed to assess the effectiveness of address extraction using LLMs. Additionally, it is examined whether the German language yields better results for German addresses or, if in general, the English language enables better extraction. To make optimal use of spatial data, the websites in the dataset are classified thematically e.g. by company type or a more diverse classification. For this classification task, the LLMs are provided with pre-existing thematic categories. The websites are classified according to the plain text and the URL of the website. The thematically classified data enables targeted evaluation of the geocoded data points through subsequent visualization and analysis. The data obtained includes all addresses found in the imprint as well as the classification of the website. The output data from models with many parameters is expected to be complete and will likely surpass previous rule-based and data-driven approaches. However, there remains a possibility that the models, regardless of their parameter size, may not be able to perform address extraction and classification at sufficient quality. Extracting the spatial context in form of coordinates from the addresses enables a largescale geographic analysis of imprint entries. The thematic context can indicate the type of institution on the site. Based on the achieved quality and hit rate, the computational power required by each model is analyzed to determine to optimize for the required computational resources. Therefore, the evaluation does not only capture the outputs but also records the number of generated tokens, the resulting costs, and the processing time. LLMs are expected to achieve a significantly higher hit rate in address extraction than conventional methods through targeted prompting. The knowledge gained from this study can contribute to the improvement of data-driven geo-spatial text data analysis and can be used in areas such as geo-spatial search engines and many types of open data projects. It should be noted that this study is work in progress, and the results presented reflect first analysis results. _____________________ * [email protected] † [email protected] ‡ [email protected] § stefan.voi[email protected] https://doi.org/10.5281/zenodo.17238018