scieee AI-readable full text Open interactive document viewer

DH Benelux Journal, Volume 7: Breaking Silos, Connecting Data: Advancing Integration and Collaboration in Digital Humanities

Bunout, Estelle; Gheldof, Tom; Smits, Thomas; Dillen, Wout; Fantoli, Margherita; Neugarten, Julia; Paccosi, Teresa; van Erp, Marieke

Abstract

Seventh volume of the DH Benelux Journal: DH Benelux Journal – Journal of Digital Humanities in the Benelux The DH Benelux Journal publishes every year a selection of the papers accepted for the DH Benelux Conference. The papers in this issue were presented in the 2024 Edition of the DH Benelux Conference, organized at KU Leuven (Leuven, Belgium), with the theme: Breaking Silos, Connecting Data: Advancing Integration and Collaboration in Digital Humanities. Guest Editors: Estelle Bunout, Tom Gheldof, Thomas Smits Permanent Editors: Wout Dillen , Margherita Fantoli, Julia Neugarten, Teresa Paccosi, Marieke van Erp

Full text

DH Benelux Journal Volume 7: Breaking Silos, Connecting Data: Advancing Integration and Collaboration in Digital Humanities Fall 2025 General Editors Wout Dillen, Margherita Fantoli, Julia Neugarten, Teresa Paccosi, and Marieke van Erp Guest Editors Estelle Bunout, Tom Gheldof, Thomas Smits DH Benelux Journal 7. Breaking Silos, Connecting Data: Advancing Integration and Collaboration in Digital Humanities. ISSN: 2666-6952 cb 2025 This journal, including all its contents, is licensed under a Creative Commons Attribution 4.0 International license, and made available in Open Access at https://journal.dhbenelux.org. The authors of the individual contributions, who are identified as such, retain the copyright over their original work. For more information on the CC BY 4.0 license, please refer to: https://creativecommons.org/licenses/by/4.0/deed.en. This journal was typeset in L A T E X by Margherita Fantoli using varianTeX — a reusable template for journals in the Humanities, developed by Wout Dillen. varianTeX is open source, available on GitHub, and deposited in the Zenodo Open Science Repository. DOI: 10.5281/zenodo.3484651. DH Benelux Journal 7 Breaking Silos, Connecting Data: Advancing Integration and Collaboration in Digital Humanities General Editors Wout Dillen, Margherita Fantoli, Julia Neugarten, Teresa Paccosi, and Marieke van Erp Guest Editors Estelle Bunout, Tom Gheldof, Thomas Smits Contents Editors’Preface............................ i Introduction: Breaking Silos, Connecting Data: Advancing Integration and Collaboration in Digital Humanities . . . . . . . . . iii Essays Reader Response to the Unconventional Narrative Tense in Different Genres Katja Tereshko, Marijn Koolen, Eva Viviani, and Peter Boot . . . . . . . . . . . . . . . . . 1 Persons in Context: Potential and Pitfalls for the CLARIAH Infrastructure Sytze Van Herck, Ivo Zandhuis, Richard L. Zijdeman, and Rick J. Mourits . . . . . . . . 17 Multilingual Automated Subject Indexing: a Comparative Study of LLMs vs Alternative Approaches in the Context of the EHRI Project Maria Dermentzi, Mike Bryant, Fabio Rovigo, and Herminio Garc´ıa-Gonz´alez . . . . . . . . 33 1 2CONTENTS Can Stones be Made (Artificially) Intelligent? A Pilot Study for AI-based Repeated Pattern Detection of Architectural Decoration in Roman Asia Minor Julie Verlinden, Thorsten Mahieu, and Toon Goedem´e.................... 57 Breaking Down Barriers: The Transition of ODIS from a Relational to a Triple Store Database Bert Aernouts, Mehmet Celik, and Joris Colla 69 Linking Het Amsterdams Stadsjournaal: A Case Study in Emerging Linked Open Data (LOD) Approaches to Audio-Visual Heritage Meg Weijers, and Christian Olesen . . . . . . 87 Herbaria Heritage: Visualizing Colonial Bias in Natural History Collections Sakura Morales Furuta, Ana Reviejo Salamanca, Michaela Todd, Lise Stork, and AndreasWeber ..................103 Small-Scale Testing on Generative AI and Post-OCR Correction in Historical Datasets Florentina Armaselu . . . . . . . . . . . . . . 123 CONTENTS 3 Uncovering Interwar Comics: The Challenge of Labeling Graphics in Belgian Weekly Magazines Bas Vercruysse, Julie M. Birkholz, Benoˆıt Crucifix, Erwin Dejasse, S´ebastien Hermans, and Krishna Kumar Thirukokaranam Chandrasekar ....................137 Documenting Jewish Mobility in Antiquity: Connecting Archaeological, Inscriptional, and Historical Data from around the Mediterranean Stefan Dingemans, Tijmen C. Baarda, Berit Janssen, Arjan Mossel, and Leonard V. Rutgers.......................161 Accessing the Republic. Entity Extraction from the Resolutions of the Dutch States-General Marijn Koolen, Esger Renkema, Nienke Groskamp, Frank Smit, Jirsi Reinders, Ger Dijkstra, Ronald Sluijter, Rik Hoekstra, and Joris Oddens . . 175 STUDIUM.AI: Datafying and Connecting the ‘Webs of Knowledge’ around the Premodern University of Leuven (1425-1797) Yann Ryan, Margherita Fantoli, Yanne Broux, and Violet Soen . . . . . . . . . . . . . . . . 207 The Wartime Propaganda Puzzle Mari Wigham, Rana Klein, Vincent Kuitenbrouwer, Marjet Brolsma, and Roeland Ordelman .....................229 cb 2025. This work is licensed under a Creative Commons “Attribution 4.0 International” license. Editors’ Preface Seven years after its creation, the DH Benelux Journal continues to gather contributions based on accepted papers from the DH Benelux conference. The conference is organized in turn by universities in the Netherlands, Belgium, and Luxembourg, but it welcomes a broad international community that joins each year. The current issue presents papers from the 2024 edition. Each accepted presentation at the conference can be submitted as a full paper to the journal, undergoing a completely new round of peer review by at least two reviewers. This year, 13 papers were accepted, covering a broad range of research topics and authored by scholars from multiple institutions. The 2024 edition—the 11th DH Benelux conference—was hosted by KU Leuven from June 4th to 7th in the charming setting of a 17th-century Franciscan College, the Irish College of Leuven, which today serves as an international residential centre for education, training, and networking. The conference theme, chosen by program chairs Thomas Smith, Estelle Bunout, Tom Gheldof, and Margherita Fantoli — "Breaking Silos, Connecting Data: Advancing Integration and Collaboration in Digital Humanities"—certainly contributed to the warm atmosphere of community building, and the sharing of best practices, data, and experiences that characterized these four sunny days. The articles in this issue contribute to this overarching theme from different perspectives, as the guest editors explain in their preface. We thank them for proposing such a stimulating program and for curating the publication process of the papers. We would also like to thank the local organizers who made the DH Benelux 2024 conference possible, ensuring a pleasant stay for all participants and providing space for all selected contributions, despite the steadily rising number of submissions. As guest editors, we hope you enjoy this volume and that you will join us at future editions of the DH Benelux conference. November 28, 2025 Amsterdam, Borås, Nijmegen and Leuven Wout Dillen Margherita Fantoli Julia Neugarten Teresa Paccosi Marieke van Erp i otherwise they would run the risk to break their wheels on the Roman ruts (Hilton, 2023). Although this narrative represents only one version of reality—failing to capture all the nuances of what is happening and often referred to as the myth of the origin of the railway track width standard—it serves as a compelling illustration of the power and historical stability of norms and standards in shaping human activity and organization. Moreover, substantial evidence shows that people tend to repeat habitual actions unconsciously. Over a century ago, psychologist William James famously described humans as "bundles of habit" (James, 1890), and this observation remains relevant in the modern world. Crucially, the specific context in which habitual behavior occurs is a decisive factor in triggering and sustaining these actions, as it provides the cues that prompt repetition and reinforce routine patterns (Neal et al., 2012; Ouellette and Wood, 1998). We propose that literary genre can function as a context in literature, shaping readers’ expectations about both content and style, including the narrative tense. For instance, when we think of fantasy literature, we often have a general sense of its themes and narrative style, as is true for many other genres. In this article, we will explore to what extent different genres are conventional in terms of themes and style. To examine style, we will focus on the primary narrative tense, as this stylistic element is fundamental to any literary text. Additionally, in a previous study (Tereshko (2024)), we demonstrated that narratives written in the present tense within the literary fiction genre tend to elicit more reader engagement, as evidenced by a higher volume of online reviews. In the present study, we focus not only on literature in the present tense, but on books written in a tense that is not conventional for a particular genre. For example, historical novels are most often written in the past tense, therefore such narrative in the present tense would be considered a deviation from what is conventional within the genre, and vice versa in the case of non-fiction, where such a deviation would be the choice of past tense narration. The aim of this study is to reveal how an average reader responds to literary texts in different genres which deviate from the genre conventions in terms of the choice of the main narrative tense. 1.1 Concept of normality and fiction If we are so inclined to choose what is familiar, where are the opportunities for progress and creativity? It is conceivable that human existence is constituted by the tension between two opposing forces – the desire for development, which implies renewal, and inertia. Neither of these is bad in itself, but the consequence of inert behaviour is the notion of normality. However, this term is debated in the humanities. Normality is a concept originating from the Latin word "norma," which means rate, scale, or figuratively, a rule. Norma (standard) is in a general sense a criterion for comparing and evaluating. Normality is defined as an agreement with the standard (Havigerová, 2012). These definitions are often debated by researchers in the field of social psychology, where the views on the operation of the concept of norm vary from a complete rejection of the concept (Gibbs, 1965) to the reduction of it to an empirical definition. In the latter case, a norm may be defined from an evaluative point of view - whether a person’s actions are judged to coincide with a standard of behaviour – or from a reactivist point of view – whether actions are encouraged or caught up. There is also a statistical approach to normality in science, in which case normality is defined as that which occurs frequently. While in social psychology this definition of norm has faded into the background, in linguistics or digital humanities it is now as 2 relevant as ever. For example, in the field of cognitive linguistics, the term ‘conventional metaphor’ has become firmly established, i.e. a metaphor that occurs in texts and speech so often that it may no longer even be recognised as a metaphor. Another example is the assimilation of the principles of conventionality/norm in children’s language acquisition (Diesendruck, 2005). While the notion of norm, due to its frequent use in various contexts, including unscientific contexts, has acquired a lot of connotations, the notion of conventionality is more neutral, so we will use it in this article. From the linguistic point of view, conventionality is an assumption or knowledge of interlocutors that for certain concepts there are forms of expression that are used in a certain linguistic community, these forms can be words, idioms, language constructions or their combination (Clark, 2010). These conventions are used even without participants being aware of them. Within a single language, conventions inherent in different types of discourse, for example, can stand out. In the areas between social sciences and linguistics, there is also room for discussion of conventions. For example, literature, as an art form, is on the one hand subject to its own laws – conventional, but on the other hand, as any art form creative by definition. For more on views on the conventionality of literature and the understanding of literary text, see Olsen (2000). Regardless of whether the pendulum swings in the direction of conventions or new forms, to some extent there are always rules in literature. In the same article, Olsen (2000) cites as an example "a HIGHLY conventionalised genre of literature, such as the sonnet" (author’s italics). If the sonnet is a highly conventionalised genre, we can assume that there are also less conventionalised genres. Nevertheless, the very division into genres proves the existence of rules by which they are distinguished. Distinguishing genres in literature is not always an easy task, and this task often has no unambiguous solution. For example, in Pratt (2016), she distinguishes linguistic expressions that occur in descriptive passages of fiction and non-fiction (historical) texts. Most often these expressions combine, undoubtedly repeating the formal structure, the meaning to be conveyed and the form of the utterance. But it is not the only way of distinguishing different genres; for example, we can use tense as a marker of genre, but even as a marker of fiction in general. In her study of the use of tense in the medieval French literature, Fleischman (1990) points out that the use of tense in narrative can be seen as abnormal compared to the use of tense in ordinary language. We suggest that tense can also be a marker of different genres. Hyland (2002) moving towards the practical application of the concept of genre in education, defines it as follows: "Genres are rhetorical actions that we draw on to respond to received repeated situations; we recognise certain language/meaning choices as representing effective ways of getting things done in particular contexts" (p. 116). This definition again notes the relationship between the content of a text and its form. The genre thus can be understood very broadly, but he also notes that speaking about the literary genres we have to stick to literary genre theory, as also mentioned in Frow (2014), where the purpose of genres is mentioned, namely the use of discrimination and taxonomy to organise texts into recognisable classes. Genres have not only a practical meaning, they are also historically defined by social entities. In the same chapter about the literary genre theory, Frow (2014) cites the book ‘Genre in Discourse’ by Tzvetan Todorov, where he writes that genres are “only the classes of texts that have been historically perceived as such [...]. In a given society the recurrence of certain discursive properties is institutionalised, and individual texts are produced 3 and perceived in relation to the norm constituted by that codification. A genre, whether literary or not, is nothing other than the codification of discursive properties’ (Todorov, 1990). Given this, the genre creates a ‘horizon of expectation’ (Frow, 2014) for the readers within the society and historical period. That is why it is important to take the conventions of the genre in consideration while analysing reader response on literary works in different genres. In this article we will rely on the existing genre system used by publishers in the Netherlands (NUR-codes, discussed in details in section 3.1), which reflects the modern perspective of genres and is broadly used in the country. Taking into account the fuzzy genre boundaries and the results of our previous research showing that readers tend to leave reviews of literary fiction in the present tense more often, this article will address the following questions: 1. Is it possible to distinguish more and less conventional genres and what could such a distinction be based on? 2. How do readers react to books that do not obey the canons of genre in terms of their choice of the main narrative tense? 2 Data and method Thanks to the efforts of the National Library in The Hague and permission from many publishers, our project, Impact and Fiction at the Huygens Institute, got access to 18.885 books published in Dutch language, including translations into Dutch, for our analysis. The only selection criteria for this corpus were the language (Dutch) and the availability of the book in digital format, as it was not feasible to digitize print versions of the books. The collection includes all the books meeting these criteria, and contains books in a span from 1956 to 2022, the majority having been published between 1996 and 2018. To create a database with all these texts they were parsed by Trankit Van Nguyen et al. (2021), the meta-information about the books was added, such as ISBN, the year of publication and NUR-code(s) to detect the specific genres. After dropping the books for which we did not know the year of publication, 17736 ISBN’s were left in the database. Based on the NUR-codes, these books are divided in 11 genres (see Table 1). The same was done for the Reviews database, containing 634,614 reviews from different digital platforms such as Goodreads, Hebban.nl, Amazon, Bol.com and others. This database is based on the Online Dutch Book Reviews (ODBR) corpus (Boot, 2017), but includes also later additions. This corpus does not cover the whole corpus of books, but we have at least one review for a majority of the books. There are also books with uncredibly high number of reviews, for example "Fifty Shades of Grey" has 828 reviews, "The Kite Runner" and "A Thousand of Splendid Suns" - 817 and 612 reviews respectively, and "The Girl on the Train" - 630, however the books with more than 100 reviews make up less than 1 per cent of our database; 9% of the books got only one review; most of the books, which is 27%, have more than 1 but less than 10 reviews; another 15% have more than 10 and less than 50 reviews; and only 2% have more than 50 and less than 100 reviews. For this article we applied statistical analysis, such as permutation test, regression analysis and other statistical tests to the text and reviews database. To the literary texts we have also applied a topic-modelling method, to test whether there is a correlation between the topics of the book and its genre. For our topic modeling of the novels, 4 Table 1: Number of books in each genre of our book collection. Genre Number of books Literary fiction 5408 Non-fiction 4947 Other fiction 1931 Suspense 1626 Regional fiction 859 Literary thriller 800 Children fiction 518 Young adult 485 Romantic fiction 483 Fantasy fiction 352 Historical fiction 327 we used the Top2Vec model (Angelov, 2020), treating entire books as individual documents. Based on previous research (Sobchuk and Še l , a, 2023; Uglanova and Gius, 2020; Van Zundert et al., 2022), we focused on content words – nouns, verbs, adjectives, and adverbs – while removing any person names identified by the Trankit NER tagger, since we consider them as not relevant for topics and they can affect the modelling by grouping books or series of books around specific characters. However, we left the location names for the modelling, since they can be an important part of the setting, especially if the story takes part in real world. After filtering our corpus of books reduced from 1,92 billion to 826 million tokens. We then applied a frequency filter, removing terms that appeared in less than 1% or more than 50% of the documents. This left us with 190 million tokens, which is about 23% of all content words and just under 10% of the total tokens. To test for the differences in the response to different genres we used the Dutch Reading Impact Model (DRIM), a rule-based model that works at the sentence level, using 275 rules to analyze impact in four categories: Affect, Aesthetic, Narrative, and Reflection. We excluded Reflection from our analysis due to unvalidated rules (Boot and Koolen, 2020). Aesthetic and Narrative are sub-categories of Affect, with relevant expressions requiring both an impact word or phrase and a book aspect term. In our analysis, we focus on Narrative and general Affect, identifying over 2 million impact expressions across reviews. To clarify distinctions, we treat generic Affect separately from Narrative and Aesthetic impacts, with 667,672 expressions identified for Aesthetic impact, 690,184 for Narrative, and 731,720 for generic Affect. 3 Conventions and genre 3.1 Genre and topic If we accept as axiomatic the fact that the genres are defined in a certain historical period by social institutions and serving as a horizon of expectation for readers of this period are conventional, then it is necessary to define what their conventionality consists of. That is, it is necessary to determine on what principles the genre division takes place. Turning to classical sources, such as Aristotle’s Poetics, we can identify three principles of genre division: medium, theme and style. Aristotle writes of the 5 Figure 1: Use of themes in a selection of genres kinds of poetry as ’modus imitatio’ and says that they all differ from each other ‘either by a difference of kind in their means, or by differences in the objects, or in the manner of their imitations’ (McKeon et al., 2009). In the context of our study, medium is of the least value as a genre feature, since the literary work most often appears as a written (less often audio) text. Of greater interest are thematic and stylistic features, which can be the basis for distinguishing genres. In this study we use the division into genres according to NUR (Nederlandstalige Uniforme Rubrieksindeling) – Dutch Uniform Subject Classification, which is based on an earlier version, NUGI (Nederlandse Uniforme Genre Indeling) – Dutch Uniform Genre Classification. Both systems use the subject division as a basis (van Rees, 2006). To check if the genres are connected with the themes of the books, we have conducted the topic analysis on the whole text database. Afterwards, the computationally selected topics were manually tagged to merge them into more recognizable themes. Figure 1 shows how these themes are represented in a selection of genres. To make these radar-plots the first 5 represented topics in each book were considered as relevant. The figures clearly show that there are thematic preferences within each genre. However, there are more robust genres such as for instance Fantasy fiction, Suspense or Historical fiction in comparison with looser genres in terms of topic preferences such as Literary fiction, Literary thriller or Children fiction. This observation suggests the conventionality of genre may be a question of degree. 3.2 Genre and tense Having seen that genre division is related mainly to themes, we can ask what stylistic differences are associated with genre division. Considering the volume of the material under study and the specificity of computer-based methods of analysis, we focused on the use of tenses as a stylistic device. In the previous study (Tereshko (2024)), we managed to identify a pattern of reader reactions to literary works in the present tense. The conclusion of that study was that readers more often tend to write an online review to the books written mainly in present tense. However, now we have a much larger number of texts at our disposal, among which more genres are represented. For this study, we determined the narrative tense of each book by utilizing grammatical parsing with Trankit. Following this, we developed a set of rules to assign a tense tag to each clause (see Table 2). In complex sentences, which may contain multiple clauses, these clauses can differ in tense, especially in Dutch. By calculating the proportions of the simple present, present perfect, simple past, and past perfect tenses, we identified the primary narrative tense for each book, regardless of genre. In Dutch, simple tenses are formed using only the main verb, while perfect tenses require an auxiliary verb combined with the perfect participle of the main verb. 6 Table 2: Examples of present and perfect tenses in Dutch Tense Example in Dutch Translation of the example Present Simple Ik werk. I work. Past Simple Ik werkte. I worked. Present Perfect Ik heb gewerkt. I have worked. Past Perfect Ik had gewerkt. I had worked. Figure 2: Distribution of books according to present ratio (left - simple present, right - simple and perfect present We have not looked for the future tense because, firstly, it is not used as the main narrative tense and, secondly, in Dutch the present tense can carry the meaning of the future tense. In the previous study, we included the present perfect in the grammatical present (based on grammatical parsing performed by Trankit), since we counted the main verb in the sentence. However, the present perfect refers to past events which are represented as given facts, which creates a break in the semantic and grammatical meaning and leads to confusion when determining the main narrative tense by counting only conjugated verbs. By counting clauses and assigning the tense to a clause rather than to the main verb in the sentence, we were able to measure the use of tenses more accurately. However, the use of perfect tenses in books seems to be insignificant for the main narrative tense, which is also logical, because this aspect cannot be used for the whole narrative and is represented in the text in small amounts. To avoid the debate over the use of the perfect, we assigned narrative tenses and divided the texts into categories – present, past, or mix - based on the proportions of the simple present and simple past tenses. Figure 2 shows the distribution of books in our database based on the percentage of present tense usage (x-axis). In the graph we can see that many books contain few clauses in the present tense (the wave on the left), so these are mostly written in the past tense, but the right side of the graph also shows that quite a large number of books use the present tense in more than half of the clauses, which makes us believe that in this case present tense might be the main narrative tense. The left graph shows this distribution based on the simple tenses only, the proportion in this case is calculated as follows: the number of sentences in the simple past/present divided by the number of all sentences in the book, including sentences in the perfect. The right graph shows the same, but including the clauses in the perfect tenses, and the graph is slightly 7 Figure 3: Distribution of present ratios by genre (KDE-plot) shifted to the right due to the larger number of counted sentences, but it does not change its shape, which means that, as mentioned above, sentences in the perfect do not significantly affect the narrative tense. Given the notion of genre conventions, we might assume that the choice of narrative tense may be specific to each individual genre, which turns out to be true. Looking at Figure 3 , we see that although most genres gravitate towards the past tense, while the present may be considered an exception to the rule, non-fiction clearly shows the opposite: these books are much more often written in the present tense than in the past tense. It is also interesting to note that authors in the genre of children’s and teenage literature do not have strong preferences in the choice of narrative tense, and the same is true also for regional novels: books in these genres are almost equally found in both present and past tense. On the other hand, fantasy and detectives are genres where the choice of past tense is clearly preferred. Thus, we can speak of a greater or lesser conventionalisation of the genre in terms of the choice of the main narrative tense. To illustrate this scale we assigned categories to each book in our database: if the percentage of present tense clauses in a book exceeds 60%, it is categorised as ’present’, if the same is true for past tense sentences it is ’past’, otherwise the book is defined as ’mix’. Then we grouped the database by genre to see how the books of each category are distributed within different genres. For this we counted the proportion of each category, considering the number of books within the genres (bar-plot on the right side of Figure 4). Then we looked at the difference between the proportions of categories, 8 Figure 4: Genres scaled by their robustness; the colour coding indicate the dominant category within the genre (left); Proportions of texts of each category within different genres (right) namely, the difference between proportionally the most represented category and the less represented one. The bigger this difference, the more robust is the genre in terms of the choice of the main narrative tense. On the left side of Figure 4 we can for example see that in Fantasy fiction the past tense is highly preferred (coded blue), while for Regional fiction the difference is just a few percent (there are slightly more books in the present tense – coded red). Thus, Figure 4 shows on the one hand that each genre is more or less conventionalised and on the other hand highlights the most dominant tense used within particular genre. If there is a tense which is used in most of the cases within a genre, there must also be the least represented tense, which is also occasionally used. The most used tense can thus be seen as an unmarked and the less used one – marked, as Fleischman (1990) explains: In the unmarked context of ordinary (nonnarrative) language the PRESENT is generally regarded as the unmarked tense while the PAST is marked <...> in a narrative context the PRESENT – or any tenseaspect category other than the PAST – is marked with respect to one or more of a set of properties that together define the unmarked tense of narration, the PAST. This passage supports our observations and is valid for most of the genres, except of non-fiction, other fiction, children literature and regional novel. In section 5 we’ll spotlight the reactions of readers on the books which deviate from the conventions of the genre, such as a non-fiction book written in the past tense, or a literary thriller expressed in the present tense. 3.3 Topic and tense As we have seen, literary genres are often distinguished by their narrative focus and seem to follow certain stylistic conventions, particularly in terms of narrative tense. This raises an interesting question: is there a connection between the topic of a text and its predominant tense? To investigate that we can use the topic analysis described before but split the genres 9 Figure 5: Topics of books in marked and unmarked tense in different genres into subgroups of books within the genre which are written in a marked tense and those where the unmarked (including the ‘mix’ category, except for other fiction) tense is used. In some genres the number of books in the unmarked tense would be very low; for instance, there are just a few historical books written in the present tense. To be sure that the number of books in the subgroup does not affect the results, we repeated the same procedure of plotting topics 1000 times with samples of 10 books from each subgroup – unmarked and marked. The results remained stable and comparable to the results that we got for the whole, which are represented in Figure 5 for selected genres. On Figure 5 for each genre, the radar plot on the left shows the themes used in books written in an ‘unmarked’, convenient, tense for this genre, while on the right shows the themes of the books in ‘marked’ group, where the choice of tense deviates from the conventions of the genre. The numbers given indicate the number of books in each category respectively. For some genres the difference between the used themes is not so big, such as in the case of romance or suspense, these genres seem to be quite theme oriented. In non-fiction the choice of tense is also clearly seen on the radar plots: the non-fiction books written in the past tense (a marked tense for non-fiction) tend to be 10 Figure 6: Average number of reviews for books written in the marked and unmarked tense more likely about historical events and geographical settings. Fantasy fiction written in the present tense (marked for fantasy) is more diverse in terms of themes used there. On the other side, young adult literature in present tense (marked for young adult books) seems to be more focused on feelings and behaviour in comparison with the young adult books written in the past tense, where for example the themes of supernatural, fantasy and science fiction play a great role. Literary fiction in its turn appears very diverse in terms of the choice of narrative themes, so it is difficult to draw conclusions about the correlation between the choice of the main narrative time and the themes raised within this genre. 4 Reader response to deviations in tense use within the genre In previous section we have seen that the genre, being defined mostly by theme, often has also a preferred tense. In this and the next section we test our hypothesis that the narrative tense is reflected in the online book reviews. To do so, we have labelled the books written in a tense, which is the least represented within a genre as ‘marked’ and other books within each genre as ‘unmarked’. After that we have merged the data about the online book reviews and the data about the books. For all genres together the ‘marked’ category of books got 28316 reviews for 2832 books, and ‘unmarked’ category – 107476 reviews for 14904 books. Figure 6 shows the average number of reviews per category, and it is clear, that there are more reviews for books written in the marked tense. The left side of Figure 6 shows the average number of reviews for the two categories for all genres, while on the right side we can see the average number of reviews for each genre. The significance of the difference in number of reviews for the two categories follows from applied permutation test with observed difference in average number of reviews between marked and unmarked categories 2.787 and p-value 0.0. To ensure the robustness of our analysis, we identified and removed outliers from the dataset based on the interquartile range (IQR) method. Specifically, for each combination of genre and category, we calculated the IQR and defined outliers as 11 resources to enhance our understanding of the past. The method of modeling the observations of personal characteristics is very important. Reconstructions of persons and family relations must follow standards to support combining or comparing datasets over time and space. Standards also ensure the quality and accountability of the data (Meroño Peñuela et al., 2020; Woltjer et al., 2024). The importance of traceability is underlined by the long tradition in humanities research of making references to the sources on which the result is based. These references are traditionally organised in textual form as footnotes, end notes, or table captions. Such sources may include not only primary source material, but also secondary literature. There are important reasons to facilitate ‘provenance trails’ from archival source data to final research output. For one, references facilitate the reproducibility of the study and as a result increase its trustworthiness. References also reduce the time required to reproduce results, enhancing opportunities to evaluate research results. Finally, references explicate assumptions made in the data wrangling or research interpretation phase. This allows work to be revisited in the future, when more knowledge becomes available on assumptions made today. Yet, in this day and age of Digital Humanities the field has hardly made progress in enhancing the degree of traceability. Reconstruction of historical persons requires data to flow from archival institutions to genealogists and researchers. Archival institutions scan, transcribe and disclose person observations from historical records. These observations are shared online for genealogists and researchers, who then create life course reconstructions based on these observations using manual methods or automated algorithms, of which the LINKS project has been the largest and most successful example in the Netherlands (Mandemakers et al., 2023). However, software made by different researchers and institutes is often not reusable, and the rationale behind reconstructions, also known as provenance, is often not retraceable. Moreover, in communicating the life course reconstructions, all links to underlying sources are aggregated into a handful of footnotes on which archival materials were used. Having more detailed references to underlying sources, which we will call ’provenance trails’, is feasible. Many of the larger research platforms stem from the 1990s and early 2000s, when FAIR was not yet an accepted philosophy, and especially memory or the use of the internet to link resources was limited. Yet, the same can be said for platforms such as WieWasWie.nl and OpenArchieven.nl, which both show how aggregator platforms can give credit to the original source in different ways, by directly pointing to that source online. The fact that this state-of-the-art state is not yet embraced by research is in our view problematic for at least three reasons. The first is that the actual reproduction of research is still a very cumbersome process, requiring many manual, none traceable, actions. Second, as with research, archives are increasingly dependent on the explication of relevance to science and society, often expressed in usage indicators for funding. The occasional acknowledgment of archives in research papers is not retraceable for those archives and cannot be used to show their relevance to research. The third is a possibly waning relationship between GLAM institutions and academic research. FARO, the Flemish support institute for cultural heritage, has already called on Flemish heritage institutions to democratise their research function and emphasises the need to step away from measuring expertise based on academic and scientific merit releasing their exclusive authority as heritage institutions (Van Oost and Vander Stichele, 2024, p. 21). In the Netherlands, there is a long-standing relationship between exchange of his18 torical person records via the archives, centralised by the Center for Family History (CBG), and person and family reconstruction via research embodied in the HSNDB at the International Institute of Social History (IISH). Research-made reconstitutions are returned to the CBG’s website as suggestions to help persons find their ancestors. 1 The CBG has recently, with the research and heritage community, released a Linked Data Vocabulary for the description of persons observations and person reconstructions in the heritage (archival) setting and the research setting. This vocabulary is called “Persons in Context” (PiCo) and will phase out an existing XML data model to support the ever-increasing exchange and use of historical person data, as meticulously described in Woltjer et al. (2024). The Netwerk Digitaal Erfgoed (NDE) urges Dutch GLAM institutions to provide their collections as linked data, archives are thus stimulated to use PiCo to describe person observations in their collection (Gaakeer et al. (2021); NDE and ministerie van OC&W (2024)). Linked Data “refers to a set of best practices for publishing structured data on the Web”. 2 Linked Data follows existing and widely adopted standards using vocabularies such as PiCo to model the data. In this paper, we want to show how the PiCo vocabulary could be used to enhance the link between research and archival sources. 3 We will do so within the Common Lab Research Infrastructure for the Arts and Humanities (CLARIAH) infrastructure framework, and discuss the implications of Persons in Context for the CLARIAH person reconstruction and evaluation tools, burgerLinker and C2RC. 4 We also introduce the notion of ’provenance trails’, a ‘bread crumb trail’ from research result to primary source as a means to capture the way that person reconstructions are built. By doing so, we underline the value of PiCo for (1) reusability of data and tooling, (2) enhancing the interoperability of tools and data in a person reconstruction pipeline and (3) the retraceability of research results to its archival resources. The importance of these concepts extends far beyond the presented pipeline and will be pivotal for sustainable Digital Heritage and Digital Humanities. 2 Generating Person Reconstructions We first evaluate the software that generates person reconstructions from person observations in the indexed civil registry for the HSNDB project LINKS (Mandemakers et al., 2023). A Person Observation refers to a single instance of a person observed in a single source. LINKS aims to reconstruct all persons who were born, married, or died within the Netherlands within the 19 th and 20 th centuries by matching and harmonising person observations from civil registry records. A combination of several person observations that presumably refer to the same person is called a person reconstruction. The initiators of LINKS chose to base their person reconstructions on the civil registry, as historically willingness to register was high, checks on the registration practice stringent, and double storage meant that the source survived bombardments, fires, floods, and other mishaps (Vulsma, 1988). Although the civil registry contains a rather complete recollection of all vital events within the borders of the Netherlands, it can be tricky to identify to which person each event belongs. The civil registry chronologically lists all births, marriages, and deaths that occurred in a municipality. Persons are only followed passively, and are 1See https://WieWasWie.nl 2’Linked Data’, W3CWiki (24 September 2023), https://www.w3.org/wiki/LinkedData. 3We would like to thank Pieter Woltjer for his insightful feedback and comments. 4See https://www.clariah.nl. 19 not observed outside of events, so while we get names of index persons and relatives, we cannot know whether and when a previous or next observation occurs (Alter et al., 2009; Gill, 1997; Van den Berg et al., 2021). As a result, any person reconstruction based on person observations in civil records is always to some extent open to interpretation, as we cannot ascertain that all events associated with a person have been added to a reconstruction. Hence provenance data and a transparent pipeline are key. Figure 1: Civil Registries schema for burgerLinker (Raad et al., 2020) (CC0) Data models help to create reusable software and introduce standards on data provenance. However, the current pipeline (Van Herck and Mourits, 2023) was developed before the creation of PiCo. The Civil Registries schema in figure 1 was purpose-built resulting from ad hoc needs for matching person observations with the custom-made tool burgerLinker. 5 burgerLinker was designed to solve computational problems in comparing millions of name combinations. burgerLinker generates person reconstructions regardless of minor differences in name spelling using the Levenshtein distance to measure similarity between names on two different records (Raad et al., 2020). In order to retrieve unique matches, multiple names on civil certificates are matched, for example ego, father, and mother. The matching strategy in figure 2 demonstrates how person names can be compared for seven types of links, six of which are included in burgerLinker.6 Because person observations are converted and checked before person reconstructions can be generated, this schema increases the overhead. In addition, burgerLinker cannot be generalised. Any variation in other sources or similar sources from other countries will likely cause errors. Moreover, comparing output with other matching algorithms is cumbersome, as each package describes provenance slightly differently in their metadata. If we were to adopt the PiCo standard, dates will be standardised to YMD ISO 8601, names are edited according to the rules of the Person Name Vocabulary (PNV) and Schema.org, certificate types are listed in PiCo terminology, relations are listed in Schema.org, and roles according to PiCo terminology (Woltjer et al., 2024). However, 5 Joe Raad, ’CLARIAH/burgerLinker’, Java (2021; repr., CLARIAH, 30 November 2023), https: //github.com/CLARIAH/burgerLinker 6Within D-M is currently not included. 20 Figure 2: burgerLinker matching strategy (Van Herck and Mourits, 2023). Image reproduced with permission from the copyright owner. other parties involved in historical person reconstruction have generally used their own schemas and developed their own algorithms to match historical person records, which has hampered the interoperability and reusability of software (Mourits et al., 2023). The PiCo vocabulary will supersede the Civil Registries vocabulary specifically created for burgerLinker. When creating burgerLinker, no existing ontology fully covered what we needed, therefore we created the Civil Registries vocabulary. The Civil Registries vocabulary was kept simple, because creating a proper schema would be another project entirely. burgerLinker’s architecture was designed to incorporate an ontology with as little effort as possible. The adoption of the ’new’ PiCo vocabulary, will therefore come with minimal adaption cost as the community has anticipated the arrival of a more fully fletched ontology to describe person observations. We note three conceptual differences between PiCo and the Civil Registries vocabulary that are important to implement anew in burgerLinker. First, unlike the Civil Registries vocabulary, PiCo discerns the date of the registration of the event and the actual event. While a marriage is usually registered on the same day, there can be a time difference in registration of death and birth. This seemingly small distinction in time is of crucial importance when registrations of baptisms are used to estimate birth dates and in research focusing on infant mortality. Second, PiCo’s elaborate modelling of family relationships and roles supersedes the Civil Registries schema. Third, while the Civil Registries vocabulary was designed with vital event registers in mind, PiCo was designed to be compatible with a wide variety of historical documents, making PiCo more versatile and interoperable. To illustrate some of the more detailed differences between the PiCo and Civil Registries vocabularies, we created Figure 3, showing the properties relating an event to a person; picot:newborn refers to the role of a newborn on a birth certificate, while picom:deceased has become a property of a person observation, schema:spouse refers to both bride and groom, and schema:parent replaces mother and father. The Civil Registries schema implies gender within the language used to describe the role, whereas 21 Figure 3: burgerLinker matching strategy remodeled according to PiCo information on a person’s gender in PiCo is explicated via the property schema:gender. To illustrate the difference in handling names between the Civil Registries Schema and PiCo, figure 4 provides an example of a marriage certificate. 7 While the Civil Registries schema uses the Schema ontology for names, the PiCo model extends Schema with the Person Name Vocabulary (PNV) to include peculiarities of Dutch names. 8 Similarly PiCo allows for other regional and language specific additions to accomodate the expression of names. burgerLinker already dropped the prefix of Dutch surnames for matching, as prefixes are generally seen as separate from the surnames in the source material and can introduce spelling variations in the full sdo:familyName. This split between prefixes and surnames becomes easier when PiCo is implemented. Specifically, burgerLinker can use PiCo’s distinction between pnv:surnamePrefix and pnv:baseSurname for family names. These are all minor tweaks to burgerLinker, but would make the program much easier and less labour-intensive to use. Provenance is central to the PiCo data model, whereas the Civil Registries schema only indirectly links to the source through a civ:registrationID. We created figure 5 to illustrate the PiCo model and to show the link between person observations and archive components in PiCo. 9 This figure illustrates the relation between a node representing the PersonObservation on the left (in blue), with all its potential properties. Moreover, PersonObservation is linked to an ArchiveComponent (in yellow) on the right side of the figure with a property from the PROV-O ontology, specifically designed to capture provenance. This relation is called the ’hadPrimarySource’ property. One of the properties describing the ArchiveComponent is the URL that points to the web 7 CBG, ’PiCo/examples/various-sources/huwelijksakte.ttl’, 14 December 2024 https://github.com/CBG-Centrum-voor-familiegeschiedenis/PiCo/blob/ a01283ffc1469285c2dcb280e32842b616187791/examples/various-sources/huwelijksakte. ttl#L77-L88 ; CBG, ’PiCo’, 20 september 2024, https://github.com/CBG-Centrum-voorfamiliegeschiedenis/PiCo/blob/main/PiCo%20Specificatie.md. 8 Schema.org, 9 January 2024, https://schema.org/ . Lodewijk Petram, Elvin Dechesne, and Gijsbert Kruithof, ‘Person Name Vocabulary’, Person Name Vocabulary, 1 July 2019, https://www. lodewijkpetram.nl/vocab/pnv/doc/. 9Visual created in Grafo (https://gra.fo). 22 Figure 4: PiCo example for the observation of a person on a marriage certificate page provided by an archival institution where the source is described more precisely and potentially could be studied through a scan of the original paper record. (Woltjer et al., 2024). Figure 5: PiCo Person Observation including PNV linked to Archive Component. The increased use of standardized value lists, such as the HSNDB’s ages, occupations, and place names (Mandemakers et al., 2020; Mourits et al., 2024), allow for secondary quality checks on the data. For example, to test whether a matched individual has a similar place or region of birth in its birth and marriage certificate (see e.g. (Van Toor et al., 2022)). Thanks to these various standardisations, we can build interoperable tooling to create PersonReconstructions and keep the links to the observation and the sources on which they are based, provided by, and stored on websites of the managing GLAM institutions. Again, the PROV-O ontology provides us with a property, called wasDerivedFrom, 23 to relate the established PersonReconstruction to the PersonObservations. Because of all these links, a PersonReconstruction can be traced back to its original source, constituting a provenance trail containing all the steps from the research result to the sources it was based upon. In this way we are describing the knowledge production process, which enables research replication in the future.(Stapel and Zandhuis, 2024) 3 Identifying Person Reconstructions The second challenge is generating a unique identifier for every person. When civil registries were introduced in the late nineteenth century, governments did not assign identifiers. The civil registry never followed persons over time and ’only’ reports on the births, marriage, and deaths that occurred that year within a municipality. However, when the civil registry for a province or the entire country is combined, the person observations can be used to reconstruct life courses and families. In order to keep track of these person reconstructions, we suggest to create a unique identifier for each person reconstruction linking to all observations of that individual. 10 Such a identifier helps to keep track of provenance on changes in sets of observations that ’reconstruct’ a person’s life course, including edits or dissolution of such a reconstruction. Generating person identifiers on a national scale over two centuries is no small feat, especially across borders and sources. It requires a balancing act between a stringent protocol stable between different versions of the data and a system tailored to continuous changes. To compare person reconstructions between versions, identifiers cannot be generated randomly. This also means that the name of a person reconstruction cannot rely on the combination of matched certificates, as this would not allow reconstructions to be checked for merging and splitting. PiCo enables us to model person reconstructions with a unique identifier and provenance to the underlying observations. However, tracking all mutations would generate an unfathomable change log. We suggest using an authority list because some historical person records are more useful than others to determine which unique persons lived in the past. A person reconstruction would receive an identifier based on the highest authority certificate. For the civil registry a logical authority list would be first birth, then death, and finally marriage certificates. This system allows for person identifiers that can be traced whilst allowing for changes between versions. For example, if a person reconstruction contains two sisters named Anna and Anna Maria, it should be split in two separate person reconstructions. One reconstruction with the identifier of the birth certificate of Anna to identify the first person reconstruction and another with an identifier from the remaining certificate with the highest authority to identify Anna Maria. An authority list can also be a set of thoroughly checked person reconstructions. In that case, an alternative to the previous logical historical identifier would be to assign serial identifiers. When mistakes in links are discovered the serial identifier would need to be changed and could possibly result in multiple identifiers for person reconstructions with updated life courses. Historical person identifiers can be maintained at any level, but ideally at the national level. In the Netherlands, the Center for Family History (CBG) is aiming to become the central network partner that identifies historical persons and assigns them historical person identifiers. In Belgium a logical approach to identifiers would be to transpose the current logic of national person identification numbers (rijksregisternummers) to the 10 Al Idrissou, ‘C2RC’, Python, 26 December 2023, https://github.com/CLARIAH/C2RC 24 past, whereby the date of birth (YY.MM.DD) is followed by a serial number of three digits according to the order of registration of a person with even numbers for female persons and uneven numbers for male persons. The final two digits are calculated by dividing the birth date and serial number by 97 as a checksum. Persons born in or after the year 2000 are assigned an additional digit (2) at the start of their national person identification number. With each century the order of registration begins again, which causes problem for collections of historical person data that span multiple centuries. Therefore, the identifier should be expanded to YYYY.MM.DD-XXX.XX.11 4 Evaluating and Validating Person Reconstructions (C2RC) Thirdly, we investigate the role of PiCo in the evaluation and validation of the - possibly automatically generated - reconstructions. The Civil Registry Reconstitutions Cleaner (C2RC), developed in the context of CLARIAH similar to burgerLinker, was designed to systematically evaluate record linkage, such as output from burgerLinker. Reconstructing life courses from person observations is a challenge. Information on records is often incomplete or slightly inaccurate. Moreover, explicitly stating why some matches are considered sound, while others are not, supports reproducibility. Therefore C2RC provides an adaptable set of hard rules and soft rules. Hard rules are matches deemed impossible, such as a reconstructed life course containing multiple birth records. Soft rules are more domain-specific and for example specify a person’s oldest feasible age or the maximum number of children a person might have. In addition to evaluation, C2RC allows users to improve upon incongruencies in the linkage based on the hard and soft rules. In line with the concept of provenance trails, C2RC monitors edits based on these violations. To track these improvements C2RC uses the VOID+ vocabulary (Idrissou et al., 2022) as shown in Figure 6.12 The VoID+ vocabulary documents the Creation, Manipulation and Evaluation of Links for Reuse and Reproducibility and indicates which sources and entities are covered by the dataset, the algorithms and criteria used to generate the data, whether links are validated or not, and if for instance person observations are clustered(Idrissou et al., 2022, p. i). The VoID+ vocabulary reuses the existing vocabularies VoID and PROV-O to create provenance trails(Idrissou et al., 2022, p. ii). Figure 6 illustrates the example from the previous section of a person reconstruction that contains the two siblings Anna and Anna Maria. The initial person reconstruction is now identified as the original cluster that is split into two subclusters. For each of these new subclusters VoID+ tracks what cluster this new person reconstruction was derived from and which person observations it is composed of. VoID+ allows us to keep track of how data is generated, which validations were applied and how person reconstructions can be split or merged. However, this does not resolve issues in identifying person reconstructions between different matching efforts, as validations and changes to the dataset are lost with each iteration. Furthermore, we could end up with an unfathomable change log that makes the dataset harder to work with despite increasing the chances of proper reproducibility. While PiCo does not implement VoID+ as a way to establish provenance trails, it enables including the PROV vocabulary to model the prov:Activity and thus keep track of the research 11 See https://www.ibz.rrn.fgov.be/fileadmin/user_upload/nl/rr/toegang/bestand-rr. pdf 12 ’VoIDPlus’, 1 August 2022. https://github.com/VoIDPlus-owl/EKAW2022/blob/main/voidPlus.owl. 25 Figure 6: C2RC export of the manual validation of a person reconstruction. Screenshot from a user session in C2RC. process of generating a Person Reconstruction as we have illustrated in figure 7. Figure 7: PiCo Person Reconstruction. As PiCo is expressed using RDF, an alternative and better approach would be to use the standardised Shapes Constraint Language (SHACL) for validation purposes. A shape graph consists of triples describing what patterns can be expected, for example, that a birth year should be expressed as a date in YYYY.MM.DD and the date must fall within a certain time range. This shape graph is compared to a data graph to show incongruencies, similar to the ‘hard’ and ‘soft’ rules in C2RC. SHACL will provide an important contribution as it can be used to see whether data was put in the right data format, and reused in similar tools across the ecosystem. We have included an example of a translation from C2RC’s rule to SHACL in Figure 8. Ideally, there 26 should be different SHACL shapes for ‘hard’ rules and ‘soft’ rules, standardising current evaluation purposes. Internationally providing these SHACL rules holds huge potential as most evaluation is now done within countries on country-specific linkage methods. Figure 8: SHACL example for PiCo. Image reproduced with permission from the copyright owner. When the indexed civil registry data is remodelled into PiCo, the Civil Registry Reconstitutions Cleaner (C2RC) needs an extensive update. C2RC relies on the graph structure to evaluate person reconstructions against a set of rules, changes to the graph structure would require changes to C2RC. In addition to the benefits of SHACL for record linkage evaluation, SHACL would also make C2RC interoperable with other software components. 5 Conclusion This paper describes potential benefits of applying a standardised data model to an existing pipeline. First we discussed how PiCo (Woltjer et al., 2024) can impact and expand burgerLinker, the software that generates person reconstructions. The second challenge is identifying person reconstructions by generating a unique and persistent identifier. Finally, we explained how a remodel of the indexed civil registry enables a standardised SHACL implementation of the Civil Registry Reconstitutions Cleaner (C2RC), a tool to evaluate and validate person reconstructions. We conclude that PiCo makes each component of the pipeline reusable by improving the interoperability of datasets and algorithms for historic archival sources containing person observations. The model is only the beginning of expanding across time, space, and source types. PiCo as a shared data standard allows us to further integrate tools into an ecosystem. Anyone can select only part of the ecosystem for their project as long as they adhere to the standard. The tools that are already developed based on the standard can be further integrated into other pipelines. An additional advantage in saving both input and output, is that by using provenance trails, provenance is saved independent from 27 1 Introduction Since its launch in 2010, one of the principal goals of the European Holocaust Research Infrastructure (EHRI) has been to catalyse transnational Holocaust research by making information about dispersed archival materials more interconnected, coherent, and accessible. For researchers into the Holocaust, the landscape of archival sources is a highly complex one, for reasons including the deliberate destruction of evidence, dispersal of materials, and the migration of survivors (Blanke et al., 2014). This has resulted in material being fragmented, duplicated, and separately described by a multitude of institutions with different mandates, cataloguing practices, and approaches to curation. The EHRI Portal is an attempt to integrate archival metadata from these dispersed sources into a single framework, creating a virtual observatory within which fragmented collections can be connected and contextualised (Blanke et al., 2017).1 Aggregating and integrating archival descriptions from hundreds of institutions around the world poses considerable institutional, social, and technical challenges. As previously observed (Erez et al., 2020; García-González and Bryant, 2023; Rodriguez et al., 2016), in the context of Holocaust-related archives, different institutions take a wide range of approaches to describing their materials, even putting aside those differences imposed by factors such as their specific collection management systems and institutional environments. The use of index terms — controlled vocabularies of subject headings, people, organisations, and places — is a cornerstone technique in enhancing the discovery and retrieval of archival material. Yet across the hundreds of institutions holding Holocaust-related material, or specialising in the field, there is negligible commonality in the application of index terms or the use of common controlled vocabularies. In most cases where index terms are used, an in-house vocabulary is the typical approach. While a small minority of institutions do use general-purpose subject heading thesauri such as the Library of Congress Subject Headings (LCSH) 2 , that does not necessarily mean that they interpret and apply the index terms to their material in consistent or interoperable ways (Erez et al., 2020). This lack of a common vocabulary for Holocaust-related material was one of the problems EHRI set out to address in its first phase (2010-2014) with the creation of the EHRI Terms vocabulary, a hierarchically organised, multilingual set of subject headings. While the starting point for EHRI Terms was a controlled vocabulary in use at Yad Vashem (Erez et al., 2020), 3 it has since undergone considerable expansion and restructuring in both the terms and available languages, and at the time of writing consists of 913 terms translated in 12 languages. 4 The vocabulary is used by EHRI’s own subject-matter experts when creating new descriptions for as-yet uncatalogued material, and as an integration and interlinking point when ingesting into the EHRI Portal material described by EHRI’s many partner institutions. It is also available in the Resource Description Framework (RDF) in Simple Knowledge Organization System (SKOS) format, along with two other controlled vocabularies focusing on World War II ghettos and camps.5 While EHRI has built a robust data integration infrastructure, improving the discoverability and interconnectedness of multilingual collection metadata, sourced from 1https://portal.ehri-project.eu/. 2https://www.loc.gov/aba/publications/FreeLCSH/LCSH44-Main-intro.pdf 3https://www.yadvashem.org/ 4 English, Hebrew, Italian, Dutch, Russian, Ukrainian, Czech, Hungarian, French, Polish, SerboCroatian, and German. As of 24 January 2025. 5https://portal.ehri-project.eu/vocabularies/ 34 multiple institutions, remains key objective for the project. The harmonisation of subject headings, using the EHRI Terms vocabulary as an integration point, is one particular area where there is room for improvement, not only because much of the aggregated metadata does not currently align with it but because many archival collections lack index terms entirely. At present, 25% of collection-level descriptions have no index terms at all. 6 Of those that do have index terms, only 30% have subject terms aligned with the EHRI Terms vocabulary.7 Past work has described how vocabularies in use by a selection of partner institutions were systematically co-referenced to EHRI Terms (Erez et al., 2020). This paper, however, investigates whether archival descriptions in the EHRI Portal could be made more discoverable and interlinked at scale through Automated Subject Indexing (ASI), i.e., the use of machine-based methods, such as computational linguistics and statistics, to perform the subject indexing steps typically performed by human indexers (Golub, 2021). The findings of this investigation are applicable beyond EHRI and its partners, as they could help other institutions address similar challenges in their specific contexts. 8 Experimenting with ASI is also part of a larger effort to integrate automated metadata generation techniques into EHRI’s workflows. Other techniques we have been exploring include multilingual Named Entity Recognition (Dermentzi and Scheithauer, 2024) and Automatic Speech Recognition (Wynne, 2023). Such efforts fit into the broader context of cultural heritage institutions viewing their collections as data that can be used to experiment with and develop innovative digital tools as described in Candela et al. (2023) and Candela et al. (2024). The rest of this paper is structured as follows: in Section 2 we provide an overview of ASI and the related work in this area; in Section 3 we describe in more detail the EHRI Terms vocabulary and a corpus of documents derived from metadata in the EHRI Portal; in Section 4 we present our experiments, the different tools under evaluation, the quantitative and qualitative methodology and the metrics used; in Section 5 we present the results obtained from our evaluation, followed in Section 6 by a discussion of various issues identified in this work. Finally, in Section 7 we draw the future lines of work based on the obtained results and in Section 8 we present our conclusions. 2 Automated Subject Indexing Approaches to ASI vary based on the purpose of the application and the field from which it originates (Golub, 2021). The two prevailing approaches are the statistical/associative and lexical techniques respectively (Suominen and Koskenniemi, 2022; Toepfer and Seifert, 2020). 9 The statistical associative approach involves Multi-label Text Classification (MTC) methods, where a supervised Machine Learning (ML) model is trained on the already indexed texts of a collection. The model learns weights that are used to predict the correct set of terms given the text of a document. Lexical approaches, on the other hand, employ string-matching to match terms in the controlled vocabulary with words in the text of the record’s description using similarity measures (Golub, 2021). Recent research has also suggested fusion approaches that 6Accessed 29th Jan 2024: https://portal.ehri-project.eu/api/datasets/E136TY2zwL 7Accessed 29th Jan 2024: https://portal.ehri-project.eu/api/datasets/CtaJXtPuZA 8 This is especially the case given EHRI’s use of archival standards, e.g., General International Standard Archival Description (ISAD(G)), and the fact that the tools used by the authors are open-source, meaning that others can reuse and adapt them for their use cases. 9 Lexical approaches are also known as string-matching or rule-based approaches (Golub, 2021). 35 combine statistical and lexical methods using ensemble techniques (Suominen and Koskenniemi, 2022; Toepfer and Seifert, 2020). More recently, the emergence of Large Language Models (LLMs) has enabled a novel approach: transfer learning using zero-shot or few-shot classification (Chow et al., 2024). A general-purpose pre-trained LLM can be used to predict suitable labels from a list of candidate terms (derived from a controlled vocabulary) without needing fine-tuning on domain-specific data (Zhang et al., 2023, p. 3). Presently, the zeroshot classification approach (where the model is made to return predictions without having been provided with any examples of its task) is commonly used as a data augmentation method to overcome class imbalance or the lack of training examples (Møller et al., 2023; Van Nooten and Daelemans, 2023). Zhang et al. (2023) and Chow et al. (2024) have experimented with subject indexing using few-shot learning (meaning the instructions provided to the model included a few examples of the task along with the list of the candidate subject terms). While EHRI has not so far deployed ASI, it has used lexical techniques in its Entity Matching Tool (EMT) for co-referencing index terms from third-party vocabularies to EHRI’s terms (or to GeoNames in the case of place names). 10 If the third-party vocabulary used to index an archival description uses terms that are morphologically very similar to the EHRI Terms, and in a language included within the EHRI Terms vocabulary, this method can return accurate results. This tool, however, was intended for assisting keyword-to-keyword alignment or aligning a set of discrete textual references, and not for generating keywords from the running text of an archival description. It does not consider the semantics of specific terms, and cannot assist with terms in languages other than those included in the target vocabulary. 3 Dataset This section characterises our dataset — derived from archival descriptions within the EHRI Portal — and the EHRI Terms vocabulary that comprises our label set. As an aggregator of archival metadata, EHRI integrates descriptions of Holocaust-related archival material from a wide range of collection holders around the world. These range from small local archives and museums with very limited digital capabilities, to large and comparatively well-resourced institutions dedicated to memorialising the Holocaust. Among the latter include prodigious aggregators of physical archival material, such as Yad Vashem and the United States Holocaust Memorial Museum (USHMM) 11 , who have amassed extensive collections of material sourced from other institutions and described using their own in-house standards and styles. A more in-depth explanation on how data integration takes place within EHRI’s current phase can be consulted in (Vanden Daelen et al., 2024). As noted in the introduction, diversity and complex provenance are notable characteristics of the archival landscape relating to the Holocaust. One result of this is that some important collections do not yet have any metadata available in digital form. Since 2010 EHRI has sought to mitigate this by providing its own English-language descriptions of such collections in the EHRI Portal. Thus, the Portal contains not just metadata sourced from third parties, but original descriptions authored by subject matter experts following EHRI’s ISAD(G)-based descriptive standard and indexed against EHRI’s vocabularies, including EHRI Terms. 10 https://emt.ehri-project.eu/ 11 https://www.ushmm.org/ 36 EHRI Terms 12 is structured as a Directional Acyclic Graph (DAG) with 15 top-level terms (“top concepts” in SKOS parlance 13 ). The hierarchical structure has a maximum depth of 10 from broadest to narrowest, but the median depth is shallower at just 3 terms. The vocabulary is, additionally, not strictly hierarchical; many terms have multiple “broader” concepts, for example, the term “Healthcare” has three broader terms (“State Functions”, “Aid, Welfare, Rescue”, and “Daily Life”). Overall, 291 of 913 subjects fit within more than one higher-level category. To construct a training and test set from structured metadata in the EHRI Portal we were restricted to the subset of archival descriptions which had already been assigned subject headings, either manually (by EHRI’s cataloguers) or via the co-referencing process described in Erez et al. (2020). This resulted in 48,854 individual archival units at all levels of description (from collection to item level). We then transformed each description into a single text by concatenating the fields considered most likely to capture the topical content of the material, as opposed to administrative details. The fields selected were ISAD(G) elements 3.1.2. Title; 3.2.2. Administrative / Biographical History; 3.2.3. Archival History; and 3.3.1. Scope and Content. We acknowledge that there is inevitably some imprecision in the selection of these fields, in terms of relevancy. This dataset contained archival descriptions in 17 different languages. Polish and Finnish were represented in the dataset by only one example each. To retain these two descriptions in the dataset before splitting it into training and test sets would mean that descriptions in Polish and Finnish could only ever be present either in the training or in the test set. To avoid training on only one example of these two languages, or testing on them without having seen any in the same language during training, we decided to remove these descriptions from our dataset, leaving us with 48,852 entries. After de-duplicating our dataset to remove descriptions of the same archival units in different languages, we were left with 36,801 archival units. In common with many multi-label classification tasks, particularly those featuring a large number of labels, class imbalance was a significant feature of our dataset. After some exploratory data analysis, we noticed that the most populous term, “Photographs”, was represented by 4,819 positive examples, while some of the least popular terms, like “Jewish religious leaders and functionaries”, were represented by only one example. Given the hierarchical nature of the label set, it is unsurprising that certain broader terms, especially terms pertaining to the genre of the material, such as “Photographs” or “Letters” were more frequently used than highly specific ones (e.g., “Activists in justice against criminals”). In fact, 215 terms had fewer than five positive examples in the dataset. To ensure better label distribution across training and test datasets, we removed terms with fewer than five examples. As a result, some descriptions matched to only one subject term lost their positive label and were also removed. The final dataset included 36,759 archival descriptions. Of the 913 subject terms in the EHRI vocabulary, only 554 met our threshold of at least five examples per term. As in related work (Skenderi et al., 2021), to obtain a training and a test dataset we used the iterative multi-label stratification method proposed by Sechidis et al. (2011) as implemented in the Python package iterative-stratification (Bradberry, 2021). We opted for this method of stratification to increase the chances that all labels (as well as all languages) – no matter how rare in the dataset – would still have positive 12 https://portal.ehri-project.eu/vocabularies/ehri_terms 13 https://www.w3.org/TR/skos-reference/#L2446 37 examples in both the training and test set. This was crucial to ensure that the evaluation of our model on the test set would be as complete as possible, covering the entire label space. It would also reduce the probability that certain multi-label classification metrics would be impossible to calculate due to the lack of positive examples as explained in Sechidis et al. (2011). Specifically, we performed an initial multi-label stratified shuffle split of the dataset to obtain one split, where the test set size would include 30 per cent of the dataset (11,027 archival descriptions) and the train set would include the remaining 70 per cent of the dataset, i.e. 25,732 archival descriptions. The test split obtained through this process was further stratified/split once more in the same manner to obtain a third, considerably smaller, test set for evaluation purposes, which included 167 archival descriptions out of the 11,027 descriptions obtained through the first split. Since we wanted to perform both quantitative and qualitative evaluations to compare the different models used in our experiments, it was essential to have a manageable yet representative final test set that would be feasible for human evaluators to process in only a few working hours. Moreover, since we wanted to experiment with at least one LLM for zero-shot classification, obtaining predictions for a very large number of archival descriptions using 554 candidate labels would have been very costly and time-consuming to run on locally available hardware. While the larger test set obtained through the second splitting process (comprising 10,860 archival descriptions) was used to estimate model performance while training and fine-tuning new models, the smaller test set (167 archival descriptions), henceforth called the “evaluation” set to avoid confusion, was reserved for the final quantitative and qualitative evaluation and comparison of all the models in our experiment and was not used during training or fine-tuning. However, due to the very small size of the evaluation set and despite following the iterative stratification method for obtaining it, only the 134 most frequent labels had positive examples within it, while the remaining 420 did not. Similarly, although the test set used during training included examples in all 15 languages present in the training set, the evaluation set included examples only in the nine most frequent languages of our dataset, 14 represented proportionally based on their distribution within our dataset. While performing these splits, a random state seed was used to make our experiment reproducible. The dataset has been made openly available following best practices on publishing cultural heritage datasets (Alkemade et al., 2023; Candela et al., 2023).15 4 Experiments This section describes how we selected the tools we put to the test: a transformer-based zero-shot classifier, a transformer-based model for fine-tuning, and other classical and non-LLM approaches. Key considerations in selecting the transformer-based models were support for multilingual text in languages relevant to EHRI’s material, being open source and easily accessible, having well-documented training data, and giving us the ability to quantitatively evaluate their performance using our in-house compute infrastructure. 14 English, Hebrew, Dutch, Czech, German, Multiple (used for descriptions containing text in more languages than one), French, Russian, Italian. 15 https://doi.org/10.5281/zenodo.14697253 38 4.1 Zero-shot Classification Approach The first approach, zero-shot classification, is the labelling of input data without the use of any prior context, e.g. labelled examples. In the context of LLMs it relies on the model having sufficient training data to have encountered related examples in its pre-training phase and to be able to make suitable inferences when given new data to classify. The purpose of this approach was to help us answer the question of whether using an LLM out of the box for classification purposes could be sufficiently reliable to be used as part of an automated or semi-ASI process or, alternatively, whether it could be used to augment our labelled dataset in order to reduce our class imbalance problem. While at the time of writing the most advanced generally available LLM is OpenAI’s GPT4 (OpenAI, 2023), we decided to omit this from the comparison at this stage due to the fact that it is not open source and due to the lack of transparency surrounding its development and training data. For practical reasons, we were also unable to include some newer LLMs capable of being self-hosted, such as TII’s Falcon 16 and Meta’s Llama217, both of which have exceedingly demanding inference requirements. The model we selected for the zero-shot approach was mDeBERTa-v3-base-mnli-xnli (Laurer et al., 2023) available through Hugging Face (HF) which uses mDeBERTa v3 (He et al., 2023) as base model. mDeBERTa v3 is a multilingual model pre-trained on the CC100 multilingual dataset covering 100 languages (including all of the languages in our dataset) (He et al., 2023). mDeBERTa-v3-base-mnli-xnli is a fine-tuned version of mDeBERTa v3 for Natural Language Inference (NLI), trained using the XNLI (Conneau et al., 2018) and MNLI (Williams et al., 2018) datasets. These datasets include texts in 15 languages, but some of the languages that are present in our EHRI MTC dataset were not included in the datasets used to fine-tune mDeBERTa v3. However, since multilingual models for NLI are known for their cross-lingual transfer learning capabilities (Conneau et al., 2020; Zhang et al., 2023), the fine-tuned model can still classify texts in the rest of the languages covered in the dataset that was used to train the base model despite not having seen any examples of them during fine-tuning. 4.2 Fine-tuning Approach The second approach was to fine-tune a pre-trained LLM on EHRI’s data, thus giving us some insight into whether the already-indexed archival descriptions can be harnessed to train models capable of performing effective classification of the remaining material. Although there is no shortage of pre-trained multilingual Transformer models to fine-tune, we decided to select BERT-base, Multilingual Cased (Devlin et al., 2019). In addition to being multilingual and open-source, we picked this model over more state-of-the-art (SOTA) alternatives for other important reasons. First, compared to more recent models it is much lighter and does not require significant resources to fine-tune. Second, it was trained on Wikipedia data rather than a CommonCrawl dataset, eliminating the chance of the model having already seen data in our training or test set (derived from either the EHRI Portal itself or EHRI-affiliated institutions). Third, as the original transformer implementation, it has gained a lot of popularity and is still widely used for tasks like text classification. It is also often used as a baseline with which to compare more recent models, helping us determine whether spending 16 https://huggingface.co/tiiuae/falcon-180B 17 https://ai.meta.com/llama/ 39 more resources to fine-tune SOTA models would be something worth pursuing. As suggested by the authors in the instructions accompanying the GitHub repository of the BERT family of models (Devlin et al., 2023), 18 the “cased” version of the model is suggested for use in multilingual settings, especially with non-Latin alphabets, since it does less normalisation of the input. According to the authors, the data that was used to train the model comprised the entire Wikipedia dump for each of the top 100 languages with the largest Wikipedia instances. 19 This list of languages includes all of those present in our train, test and evaluation sets. The transformers library (Wolf et al., 2020) developed by HF, as well as tutorials developed by the HF community (Rogge, 2021), provided an accessible way to proceed with fine-tuning Transformer models for multi-label classification. Underneath the hood, fine-tuning the pre-trained base BERT model on our EHRI dataset for MTC adds an extra linear layer on top of BERT’s final layer, which is trained to output scores that indicate how likely it is for each label to be the correct label for each text in the input. To fine-tune BERT, we used a batch size of 32 and trained for 25 epochs, selecting the model with the best micro-averaged F1 score. We set the learning rate to 2e-5. To avoid overfitting on the training set, we set the hidden dropout probability hyperparameter to 0.2 when instantiating the model and the weight decay hyperparameter to 0.01 when specifying the training arguments. The best micro-averaged F1 score on the large test set (≈0.56) was achieved by the model obtained in the final epoch.20 4.3 Other Approaches As a framework for quantitative evaluation we chose to employ Annif (Suominen, 2019; Suominen and Koskenniemi, 2022; Suominen et al., 2022, 2023), a tool for multilingual subject indexing already used by a number of cultural heritage institutions. Annif has been designed to enable the use of different multi-label classification backends tailored for specific types of material. It includes implementations of classical indexing algorithms like TF-IDF, as well as state-of-the-art approaches such as fastText (Joulin et al., 2017), and Parabel (Prabhu et al., 2018). Annif is written in Python and can be extended with additional backends by implementing specific subclasses. These new backends can then leverage Annif’s evaluation tools for generating metrics such as precision, recall, and F1 score (per-document, as well as using macro or micro averaging) along with ranked metrics such as non-discounted cumulative gain (NDCG). In addition to the LLM-based approaches detailed in Section 4.1 and 4.2, we also evaluated a selection of Annif backends in order to provide a performance comparison. 21 These included: TF-IDF A baseline method capable of fast, resource-efficient labelling. MLLM Maui-like Lexical Matching, inspired by Medelyan (2009). fastText A fastText classifier. 18 https://github.com/google-research/bert/blob/master/multilingual.md 19 See https://github.com/google-research/bert/blob/master/multilingual.md#datasource-and-sampling 20 EHRI’s fine-tuned BERT-based model is available on Hugging Face ( https://huggingface.co/ mdermentzi/finetuned-bert-base-multilingual-cased-ehri-terms ) and Zenodo ( https://doi. org/10.5281/zenodo.14680007) 21 More details about these backends can be found in the Annif wiki page: https://github.com/ NatLibFi/Annif/wiki 40 Omikuji A Parabel implementation as an example of a state-of-the-art non-LLM method. NN Ensemble A neural-network-based method of dynamically combining the outputs of several other discrete classification tools, in this case TF-IDF, MLLM, fastText, and Omikuji. While it is capable of being configured for use in a multilingual setting, Annif does not currently have explicit support for heterogeneous corpora, e.g. those containing documents in a variety of languages. Since testing on our predominantly English evaluation dataset showed minimal impact on performance, however, we did not seek to address this limitation at this stage, noting that it could be worthwhile to explore as future work. We created custom Annif backends for our fine-tuned BERT model and our zero-shot model of choice run via the HF transformers library, allowing us to deploy them and evaluate their performance in a consistent manner alongside Annif’s native backends. For the purpose of this experiment we only implemented the inference/suggestion functionality, however as future work the fine-tuning process could also be reimplemented to fit within Annif’s backend framework, including training and hyper-parameter optimisation.22 4.4 Evaluation Metrics We decided to evaluate the results of our experiments both quantitatively and qualitatively. Properly evaluating subject terms and consequently subject indexing tools is a very challenging task for a plethora of reasons (Golub et al., 2016). First and foremost, assigning and evaluating subject terms is a very subjective process and a term that may be considered pertinent by some may not be perceived as such by others. In fact, several inter-indexer consistency studies have proven that consistency levels between different indexers are typically very low (Tonta, 1991; Vaughan and Rafferty, 2011). Moreover, depending on the target audience of an institution, narrower and more specific subject terms may be preferred over broader ones, while other institutions may prioritise indexing records with as many terms as potentially useful rather than only a few specific ones (Golub et al., 2016). Ultimately, people browsing a catalogue to search for information on a specific topic will perceive a subject term to be correctly assigned only if it leads them to records relevant to their specific needs. Ideally, the evaluation of ASI tools should be based on how successful they are at helping users find answers to their questions. However, performing this kind of evaluation is extremely difficult and due to the subjective nature of the task may still lead to inconclusive results. Usually, already indexed records are used as gold standard datasets based on which different tools can be evaluated quantitatively (Ahmed et al., 2023; Golub, 2021; Golub et al., 2016; Suominen et al., 2022). However, since subject indexing is a subjective task that is influenced by many factors (level of indexer’s expertise, institutional priorities, length of the vocabulary, etc.) (Golub, 2021; Golub et al., 2016), it is very hard to develop a truly comprehensive gold standard dataset on any real-life collection indexed with a sizeable enough controlled vocabulary. While we acknowledge this limitation, for the purposes of the quantitative evaluation herein we assume that the terms 22 A branch containing modified Annif code is available on Github ( https://github.com/EHRI/ Annif/tree/ehri_masi) and Zenodo (https://doi.org/10.5281/zenodo.14697325). 41 assigned to EHRI-ingested records through the methods described in the previous sections are complete and correct (Golub et al., 2016). As explained by Suominen et al. (2022), treating parts of already indexed collections as gold standard collections allows for quick and easy experimentation and evaluation of different algorithms. However, the authors agree with Suominen et al. (2022) that this method is best seen as a “ballpark estimate of quality” and other methods of evaluation should also be considered. Similar to previous work (Ahmed et al., 2023; Asula et al., 2021; Chou and Chu, 2022; Gárdos et al., 2023; Skenderi et al., 2021; Toepfer and Seifert, 2020), to quantitatively evaluate the different multi-label classification methods we computed precision, recall and F1 scores. We computed these scores in three different ways: micro-averaged (taking into account total true positives, false negatives and false positives), weighted macro-averaged (computing the metrics per label weighted by each label’s number of true instances to account for class imbalance) and averaged per document. All of these scores range from zero to one, with one being the best possible score. We have used as our primary metric the F1 score as provided by the Annif framework (Suominen et al., 2022) which considers the top five suggestions returned by the algorithm, averaged across all text samples. Since labels display some degree of imbalance due to the hierarchical nature of the vocabulary and preponderance of broader terms, other metrics display a degree of divergence from this figure and have also been given in Table 1. The F1 micro average, which considers true positive, false positive, and false negative values summed across all text samples, gives proportionally higher weight to performance on common labels. The F1 weighted macro score, on the other hand, averages individual scores across all labels and gives a better indication of performance on rare labels. Since labels are unordered we have not included NDCG or other ranked measures. Since only 134 of the 554 subject terms we are considering in this experiment had positive examples in the evaluation set, our evaluation of these models only covers around 24 per cent of the labels that the models were trained (or–in the case of zero-shot classification–asked) to predict. We acknowledge that this is an important limitation but at this stage of our research, this evaluation is regarded more like a first soundness check before proceeding with a more thorough evaluation. Nevertheless, thanks to the stratification method that we followed, the evaluation set still contained sufficient labels with positive examples which helped us understand the overall strengths and weaknesses of the different models under evaluation. 5 Results 5.1 Quantitative evaluation The results of the quantitative evaluation, in terms of F1 document average, weighted subject average (weighted macro average), and micro average are shown in Table 1 and Figure 1. Considering the F1 document average and F1 subject average scores, Annif’s Neural Network ensemble classifier gave us the best results with 0.5884 and 0.5568 respectively. However, the fine-tuned BERT model achieved the highest micro average F1 score with 0.5916, and beat all other non-ensemble methods across all metrics, including Omikuji/Parabel, the best-performing non-ensemble classifier. This might be explained by the fact that during training the fine-tuned BERT model was optimised for better micro-averaged F1 scores. 42 The quantitative evaluation scores achieved by our zero-shot LLM model, mDeBERTA v3, were generally poor compared with other methods, with a document average F1 of just 0.1099, narrowly ahead of MLLM but behind other classification methods. TF-IDF MLLM fastText Omikuji/Parabel NN Ensemble mDeBERTa BERT-EHRI Classifier 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Score Classifier F1 scores F1 (doc avg) F1 (weighted subj avg) F1 (micro avg) Figure 1: F1 scores for 7 classifiers under test F1 (doc avg) F1 (weighted subj avg) F1 (micro avg) TF-IDF 0.2373 0.3640 0.2533 MLLM 0.0849 0.1762 0.1284 fastText 0.3606 0.4331 0.3768 Omikuji/Parabel 0.3771 0.4776 0.3975 NN Ensemble 0.5884 0.5568 0.5256 mDeBERTa 0.1099 0.2055 0.1126 BERT-EHRI 0.5556 0.4844 0.5916 Table 1: F1 document average, weighted subject average, and micro average scores for different models. The results in bold represent the best metrics per type of measure. The two horizontal grey lines delimit the three categories of tools under evaluation, namely the first group is composed of the Annif originally supported multi-label classification backends, the second group comprises the zero-shot LLM and the final group contains the fine-tuned LLM. 43 evaluation. Quite consistently it output labels that were considered good by all three human judges, along with a relatively large quantity of labels considered poor by one judge. As an assistive tool, therefore, it could be considered more useful than the trained models, which — consistent with the label distribution in the dataset they were trained on — tend towards broader, less specific terms. Potentially offsetting this, however, are the computational requirements of the zero-shot model, which was several orders of magnitude slower than alternative methods on the same hardware. As highlighted in Section 6, based on our experiments the ASI tools described in this paper can realistically be considered only as part of a semi-ASI workflow. The finetunedmultilingualBERTmodelcanreturn goodlabelsets, butthe NNEnsemble model shows very similar performance while being considerably less resource-intensive to run. Given its tendency to result in more specific terms, zero-shot classification using LLMs could be potentially useful as part of data augmentation efforts to improve the representation of more specific labels, resulting in a richer and more balanced dataset that could, in turn, contribute to training more accurate domain-specific models. Funding This work has been carried out in the context of the EHRI-3 project funded by the European Commission under the call H2020-INFRAIA-2018-2020, with grant agreement ID 871111 and DOI 10.3030/871111. References Mustak Ahmed, Mondrita Mukhopadhyay, and Parthasarathi Mukhopadhyay. Automated Knowledge Organization AI ML based Subject Indexing System for Libraries. DESIDOC Journal of Library & Information Technology, 43:45–54, March 2023. doi: 10.14429/djlit.43.01.18619. Henk Alkemade, Steven Claeyssens, Giovanni Colavizza, Nuno Freire, Jörg Lehmann, Clemens Neudeker, Giulia Osti, Daniel van Strien, et al. Datasheets for digital cultural heritage datasets. Journal of open humanities data, 9(17):1–11, 2023. doi: 10.5334/johd.124. URL https://doi.org/10.5334/johd.124. Marit Asula, Jane Makke, Linda Freienthal, Hele-Andra Kuulmets, and Raul Sirel. Kratt: Developing an Automatic Subject Indexing Tool for the National Library of Estonia. Cataloging & Classification Quarterly, 59(8):775–793, November 2021. ISSN 0163-9374. doi: 10.1080/01639374.2021.1998283. URL https://doi.org/10.1080/ 01639374.2021.1998283. Tobias Blanke, Veerle Vanden Daelen, Michal Frankl, Conny Kristel, Kepa Rodriguez, and Reto Speck. The Past and the Future of Holocaust Research: From Disparate Sources to an Integrated European Holocaust Research Infrastructure. In Andrea Rapp, Norbert Lossau, and Heike Neurot, editors, Evolution der Informationsinfrastruktur, pages 157–177. Verlag Werner Hülsbusch, Glückstadt, May 2014. ISBN 978-3-86488-043-8. Tobias Blanke, Michael Bryant, Michal Frankl, Conny Kristel, Reto Speck, Veerle Vanden Daelen, and René Van Horik. The European Holocaust Research Infrastructure Portal. Journal on Computing and Cultural Heritage, 10(1):1:1–1:18, January 2017. 50 ISSN 1556-4673. doi: 10.1145/3004457. URL https://dl.acm.org/doi/10.1145/ 3004457. Trent J. Bradberry. iterative-stratification, 2021. URL https://github.com/trentb/iterative-stratification. original-date: 2018-02-04T00:32:10Z. Gustavo Candela, Nele Gabriëls, Sally Chambers, Milena Dobreva, Sarah Ames, Meghan Ferriter, Neil Fitzgerald, Victor Harbo, Katrine Hofmann, Olga Holownia, et al. A checklist to publish collections as data in GLAM institutions. Global Knowledge, Memory and Communication, 2023. doi: 10.1108/GKMC-06-2023-0195. URL https://doi.org/10.1108/GKMC-06-2023-0195. Gustavo Candela, Vicky Dritsou, Sally Chambers, Agiatis Benardou, Alba Irollo, and Toma Tasovac. Introduction to Collections as Data. DARIAH-Campus, September 2024. URL https://campus.dariah.eu/en/resource/posts/introduction-tocollections-as-data. Publisher: DARIAH-Campus. Charlene Chou and Tony Chu. An Analysis of BERT (NLP) for Assisted Subject Indexing for Project Gutenberg. Cataloging & Classification Quarterly, 60(8):807– 835, November 2022. ISSN 0163-9374. doi: 10.1080/01639374.2022.2138666. URL https://doi.org/10.1080/01639374.2022.2138666. Eric H. C. Chow, T. J. Kao, and Xiaoli Li. An Experiment with the Use of ChatGPT for LCSH Subject Assignment on Electronic Theses and Dissertations. Cataloging & Classification Quarterly, 62(5):574–588, July 2024. ISSN 0163-9374. doi: 10.1080/ 01639374.2024.2394516. URL https://doi.org/10.1080/01639374.2024.2394516 . Publisher: Routledge _eprint: https://doi.org/10.1080/01639374.2024.2394516. Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. XNLI: Evaluating Cross-lingual Sentence Representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2018. Place: Brussels, Belgium. Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised Cross-lingual Representation Learning at Scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.747. URL https://www.aclweb.org/anthology/2020. acl-main.747. Maria Dermentzi and Hugo Scheithauer. Repurposing Holocaust-Related Digital Scholarly Editions to Develop Multilingual Domain-Specific Named Entity Recognition Tools. In Isuri Anuradha, Martin Wynne, Francesca Frontini, and Alistair Plum, editors, Proceedings of the First Workshop on Holocaust Testimonies as Language Resources (HTRes) @ LREC-COLING 2024, pages 18–28, Torino, Italia, May 2024. ELRA and ICCL. URL https://aclanthology.org/2024.htres-1.3. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pretraining of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 51 pages 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N191423. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT, October 2023. URL https://github.com/google-research/bert . original-date: 2018-1025T22:57:34Z. Sigal Arie Erez, Tobias Blanke, Mike Bryant, Kepa Rodriguez, Reto Speck, and Veerle Vanden Daelen. Record linking in the EHRI portal. Records Management Journal, 30(3):363–378, 2020. ISSN 09565698. doi: 10.1108/RMJ-08-2019-0045. URL https: //www.proquest.com/docview/2466656464/abstract/620A0D6747804A5CPQ/1. Herminio García-González and Mike Bryant. The Holocaust Archival Material Knowledge Graph. In Terry R. Payne, Valentina Presutti, Guilin Qi, María Poveda-Villalón, Giorgos Stoilos, Laura Hollink, Zoi Kaoudi, Gong Cheng, and Juanzi Li, editors, The Semantic Web – ISWC 2023, Lecture Notes in Computer Science, pages 362– 379, Cham, 2023. Springer Nature Switzerland. ISBN 978-3-031-47243-5. doi: 10.1007/978-3-031-47243-5_20. Koraljka Golub. Automated Subject Indexing: An Overview. Cataloging & Classification Quarterly, 59(8):702–719, November 2021. ISSN 0163-9374. doi: 10.1080/01639374. 2021.2012311. URL https://doi.org/10.1080/01639374.2021.2012311. Koraljka Golub, Dagobert Soergel, George Buchanan, Douglas Tudhope, Marianne Lykke, and Debra Hiom. A framework for evaluating automatic indexing or classification in the context of retrieval. Journal of the Association for Information Science and Technology, 67(1):3–16, 2016. ISSN 2330-1643. doi: 10.1002/asi.23600. URL https://onlinelibrary.wiley.com/doi/abs/10.1002/asi.23600 . Number: 1 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/asi.23600. Judit Gárdos, Julia Egyed-Gergely, Anna Horváth, Balázs Pataki, Roza Vajda, and András Micsik. Identification of social scientifically relevant topics in an interview repository: a natural language processing experiment. Journal of Documentation, 80(2):354–377, January 2023. ISSN 0022-0418. doi: 10.1108/JD-12-2022-0269. URL https://doi.org/10.1108/JD-12-2022-0269 . Publisher: Emerald Publishing Limited. Pengcheng He, Jianfeng Gao, and Weizhu Chen. DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing, March 2023. URL http://arxiv.org/abs/2111.09543. arXiv:2111.09543 [cs]. Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. Bag of Tricks for Efficient Text Classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 427–431, Valencia, Spain, April 2017. Association for Computational Linguistics. URL https://aclanthology.org/E17-2068. Moritz Laurer, Wouter van Atteveldt, Andreu Casas, and Kasper Welbers. Less Annotating, More Classifying: Addressing the Data Scarcity Issue of Supervised Machine Learning with Deep Transfer Learning and BERT-NLI. Political Analysis, pages 1–17, June 2023. ISSN 1047-1987, 1476-4989. doi: 10.1017/pan.2023.20. URL https://www.cambridge.org/core/journals/political-analysis/article/ 52 less-annotating-more-classifying-addressing-the-data-scarcity-issueof-supervised-machine-learning-with-deep-transfer-learning-andbertnli/05BB05555241762889825B080E097C27. Olena Medelyan. Human-competitive automatic topic indexing. Thesis, The University of Waikato, 2009. URL https://researchcommons.waikato.ac.nz/handle/10289/ 3513. Accepted: 2009-12-18T00:41:54Z. Anders Giovanni Møller, Jacob Aarup Dalsgaard, Arianna Pera, and Luca Maria Aiello. Is a prompt and a few samples all you need? Using GPT-4 for data augmentation in low-resource classification tasks, April 2023. URL https://arxiv.org/abs/2304. 13861v1. OpenAI. GPT-4 Technical Report, March 2023. URL http://arxiv.org/abs/2303. 08774. arXiv:2303.08774 [cs]. Yashoteja Prabhu, Anil Kag, Shrutendra Harsola, Rahul Agrawal, and Manik Varma. Parabel: Partitioned Label Trees for Extreme Classification with Application to Dynamic Search Advertising. In Proceedings of the 2018 World Wide Web Conference on World Wide Web - WWW ’18, pages 993–1002, Lyon, France, 2018. ACM Press. ISBN 978-1-4503-5639-8. doi: 10.1145/3178876.3185998. URL http://dl.acm.org/ citation.cfm?doid=3178876.3185998. Kepa J Rodriguez, Vladimir Alexiev, Laura Brazzo, Charles Riondet, Yael Gherman, and Reto Speck. EHRI-2 - D.11.2 Road Map Domain Vocabularies. Deliverable GA no. 654164, 2016. URL https://www.ehriproject.eu/sites/default/files/downloads/ehri_downloads/D%2011.2% 20Road%20map%20domain%20vocabularies.pdf. Issue: GA no. 654164. Niels Rogge. Fine-tuning BERT (and friends) for multi-label text classification, November 2021. URL https://github.com/NielsRogge/TransformersTutorials/blob/master/BERT/Fine_tuning_BERT_(and_friends)_for_multi_ label_text_classification.ipynb. Konstantinos Sechidis, Grigorios Tsoumakas, and Ioannis Vlahavas. On the Stratification of Multi-label Data. In Dimitrios Gunopulos, Thomas Hofmann, Donato Malerba, and Michalis Vazirgiannis, editors, Machine Learning and Knowledge Discovery in Databases, Lecture Notes in Computer Science, pages 145–158, Berlin, Heidelberg, 2011. Springer. ISBN 978-3-642-23808-6. doi: 10.1007/978-3-642-23808-6_10. Erjon Skenderi, Jukka Huhtamäki, andKostasStefanidis. Multi-KeywordClassification: A Case Study in Finnish Social Sciences Data Archive. Information, 12(12):491, December 2021. ISSN 2078-2489. doi: 10.3390/info12120491. URL https://www. mdpi.com/2078-2489/12/12/491. Osma Suominen. Annif: DIY automated subject indexing using multiple algorithms. LIBER Quarterly: The Journal of the Association of European Research Libraries, 29(1):1–25, July 2019. ISSN 2213-056X. doi: 10.18352/lq.10285. URL https://liberquarterly. eu/article/view/10732. Osma Suominen and Ilkka Koskenniemi. Annif Analyzer Shootout: Comparing text lemmatization methods for automated subject indexing. The Code4Lib Journal, (54), August 2022. ISSN 1940-5758. URL https://journal.code4lib.org/articles/ 16719. 53 Osma Suominen, Juho Inkinen, and Mona Lehtinen. Annif and Finto AI: Developing and Implementing Automated Subject Indexing. JLIS.it, 13(1):265–282, January 2022. ISSN 2038-1026. doi: 10.4403/jlis.it-12740. URL https://www.jlis.it/index. php/jlis/article/view/437. Osma Suominen, Juho Inkinen, Tuomo Virolainen, Moritz Fürneisen, Bruno P. Kinoshita, Sara Veldhoen, Mats Sjöberg, Philipp Zumstein, Robin Neatherway, and Mona Lehtinen. Annif, August 2023. URL https://github.com/NatLibFi/Annif . Martin Toepfer and Christin Seifert. Fusion architectures for automatic subject indexing under concept drift. International Journal on Digital Libraries, 21(2):169–189, June 2020. ISSN 1432-1300. doi: 10.1007/s00799-018-0240-3. URL https://doi.org/10. 1007/s00799-018-0240-3. Yaşar Tonta. A Study of Indexing Consistency Between Library of Congress and British Library Catalogers, 1991. URL http://eprints.rclis.org/9464/ . Issue: 2 Number: 2 Pages: 177-185 Volume: 35. Jens Van Nooten and Walter Daelemans. Improving Dutch Vaccine Hesitancy Monitoring via Multi-Label Data Augmentation with GPT-3.5. In Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment, & Social Media Analysis, pages 251–270, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.wassa-1.23. URL https://aclanthology.org/ 2023.wassa-1.23. Veerle Vanden Daelen, Herminio García González, Dorien Styven, Jarno Maertens, Michala Lônčíková, Angel Chorapchiev, Eliot Nidam, Eli Furman, Johannes Meerwald, Florine Miez, Rachel Pistol, Mike Byrant, Martin Posch, Éva Kovács, Fabio Rovigo, László Csősz, Mantas Šikšnianas, Giorgos Antoniou, Laura Brazzo, Rebecca Dillmeier, and Michał Czajka. EHRI-3 Deliverable 9.6: Overview of Data Integration, 2024. URL https://www.ehri-project.eu/wp-content/uploads/2024/12/D9.6Overview-of-data-integration.pdf. Hughes Alan Vaughan and Pauline Rafferty. Inter-indexer consistency in graphic materials indexing at the National Library of Wales. Journal of Documentation, 67(1): 9–32, January 2011. ISSN 0022-0418. doi: 10.1108/00220411111105434. URL https: //doi.org/10.1108/00220411111105434 . Publisher: Emerald Group Publishing Limited. Adina Williams, Nikita Nangia, and Samuel Bowman. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics, 2018. URL http://aclweb.org/anthology/N181101. Place: New Orleans, Louisiana. Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. HuggingFace’s Transformers: State-of-the-art Natural Language Processing, July 2020. URL http://arxiv.org/abs/1910.03771. arXiv:1910.03771 [cs]. 54 Martin Wynne. Using Holocaust Testimonies as Research Data, May 2023. URL https: //www.clarin.ac.uk/article/using-holocaust-testimonies-research-data. Shiwei Zhang, Mingfang Wu, and Xiuzhen Zhang. Utilising a Large Language Model to Annotate Subject Metadata: A Case Study in an Australian National Research Data Catalogue, October 2023. URL http://arxiv.org/abs/2310.11318 . arXiv:2310.11318 [cs]. 55 cb 2025. This work is licensed under a Creative Commons “Attribution 4.0 International” license. Can Stones be Made (Artificially) Intelligent? A Pilot Study for AI-based Repeated Pattern Detection of Architectural Decoration in Roman Asia Minor Julie Verlinden1,2, Thorsten Mahieu2, and Toon Goedemé2 1 Ancient History Research Group, Department of History, Ghent University, Belgium 2PSI-EAVISE Research Group, Department of Electrical Engineering, KU Leuven, Belgium The disciplines of History, Archaeology and Computer Vision have each stimulated the development and use of diverse digital tools within their own domain. But what if we dare to ask a research question positioned right at the crossroad of these fields, necessitating the computer-based analysis of people’s past via both inscriptions and monuments? In this paper we return to the dawn of our era (1st century BCE - 3rd CE) in the historical region of Asia Minor (roughly modern-day Turkey) to put this to the test. Our starting point? The many legacy photographs of so called Bauornamented buildings. With the help of pattern recognition techniques, we attempted to design a (proof-of-concept) AI-algorithm capable of detecting and segmenting the various complex designs in pictures of such embellished building blocks. In doing so, we focused on mitigating the bottlenecks of the traditional time-consuming manual comparative approaches, allowing to unlock this data’s true historical potential1. Keywords: Computer Vision, Repeated Pattern Detection, AI, Architectural Decoration, Roman Asia Minor 1 Introduction Our starting point is a simple, yet significant observation about the Roman empire: archaeologists and ancient historians encounter a remarkable degree of similarity in the material culture and architecture of its many studied provincial cities (Van Oyen and Pitts, 2017). Focusing on buildings, we observe that, despite differences of detail, 1 This research was funded by the Research Foundation Flanders (FWO), grant number 11PDX24N. 57 constructions look remarkably similar, also in Asia Minor. What processes explain this similarity in an empire where the central government was bureaucratically minimalist (Garnsey and Saller, 2015, 1987) and local communities enjoyed a lot of autonomy (Zuiderhoek, 2009)? Furthermore, the socio-cultural effects of imperial politics and policies were never top-down, nor straightforward, and could take many forms (Mattingly, 2011). A pivotal dataset, omnipresent but so far overlooked in any attempt to unravel this commonality 2 , forms our dataset of choice: that of architectural decoration. Also known as Bauornamentik, the canon of geometrical and vegetal designs found on a building has no structural function, yet is present on nearly every monument with some grandeur in the empire. And, it is prone to fashions, changing over time and space, yet displaying the above-mentioned similarities (Ginouvès and Martin, 1985; Vandeput, 1997). Especially in Asia Minor, many monuments are also linked to surviving inscriptions. These provide information on who commissions or pays for these buildings; who sets trends or copies them; who plays which role in the game of social status. These texts, in other words, connect the buildings and their decorations to the people behind the stones: the elite protagonists who built their lives in these ornamented cities. In brief, we investigate if the similarities identified in the ornaments match with other (un)known forms of connectivity (institutional, cultural, social, religious) between the same cities and their elites, known from written sources3. The opportunity for computer vision techniques here pertains to the quantity and granularity of detail that needs to be captured. At this stage, our data compendium contains just under 2,000 images of decorated blocks from ca. 300 monuments on 40 archaeological sites. Since every building block typically exhibits multiple different designs, mapping and comparing these manually is practically impossible, both in terms of time consumption and because the sheer number of research components exceeds the human capacity of analysis and comparison. Hence, AI-based automatic pattern recognition (step 1) and comparison (step 2) techniques are indispensable, placing architectural decoration studies on an exciting and innovative new footing. Indeed our proposed workflow consists of a two-step approach: 1) isolate the separate repetitive decorative patterns as identified in literature from the legacy photographs, 4 , yielding a digital library of pattern snippets per different canonized design class, 2) sub-cluster the latter per ornament class to unveil, measure and quantify the mentioned (dis)similarity via an embedding learning or metric learning approach. This consists of 2 Different scholars have attempted to answer this question, e.g. based on the concepts of globalization (Pitts and Versluys, 2015) or on social and cultural revolutions (Wallace-Hadrill, 2008). To date, however, the main merit of the field of architectural decoration studies (typically part of the broader domain of Bauforschung) lies in its contribution to establishing chronological frameworks for archaeological buildings, sites and regions e.g. Gliwitzky (2010). As a result, this archaeological dataset’s potential to answer broader socio-historical questions has so far remained underexplored. 3 Roman Imperial Asia Minor provides an ideal and relevant laboratory to perform such stylistic, architectural, urban and historical exercise, considering the comparatively high degree of urbanization and monumentalization of the vast peninsula (Willet, 2020), the number of extensively excavated ancient cities in the region (Gates, 2011), their well-studied architectural (Yegül and Favro, 2019)(esp. ch.10) and exceptionally rich epigraphical record. Moreover, there is the simultaneous academic recognition of both ancient historians and archaeologists of the significance of these topics (Koller et al., 2021; Lohner-Urban and Quatember, 2020; Marek, 2016) 4 For a comprehensive overview of the nomenclature of the main decorative designs, see eg.Vandeput (1997) in English, Cavalier (2005) in French, Rumscheid (1994) in German. This Ph.D. project also works on the disambiguation of the existing nomenclature as apt terminology in English is for some patterns still problematic. This will be presented in the database which will be made publicly available. 58 converting a photo/snippet into a descriptor vector (the embedding) that describes the image content. Distances between those vectors are then trained to correspond to the visual (dis)similarity which can vary per region and/or time window. This allows to compare a set of snippets of the same family/class of patterns (e.g. Ionic kyma decoration) and divide them into subsets of (more) homogeneous groups (e.g. eggand-dart, egg-and-tongue, intermediate forms) according to the needed granularity (De Feyter et al., 2023). In this paper, we focus on the first step of our workflow, the application of artificial intelligence (AI) for the automated recognition of the architectural decoration canon in Roman Asia Minor. Our core task is thus automating the identification and classification of repetitive decorative patterns found in Roman architecture, thereby enabling a more efficient exploration and deeper understanding of cultural transmission and stylistic evolution across the Roman Empire. 2 Proposed AI pipeline In short, we have developed an AI-driven system that can recognize, segment, and classify these architectural patterns from photographic data. In this paper, we explore state-of-the-art computer vision techniques that would render this task totally automatic, without any of the manual labor traditionally involved in analyzing these motifs. Indeed, the very fact that these patterns consist of a unit that consistently repeats multiple times on the image makes a clue to detect and isolate each pattern without supervision. A limited but representative dataset of 50 images of already published Roman blocks from across Asia Minor was compiled, serving as the basis for training, testing and refining our algorithm. Since this dataset purely served the purpose of establishing our novel pipeline, we did not impose any further chronological or spatial selection of the Roman building blocks depicted. Preprocessing techniques were applied to improve the accuracy. First, Contrast Limited Adaptive Histogram Equalization (CLAHE) was used (Reza, 2004), where histogram equalization is applied to small parts of the image (’tiles’) and combined via bilinear interpolation. This improved the contrast between shadow and sunlight without disrupting major contrasts. In addition, gamma correction further refined the contrast. For the actual repreating pattern detection, we explored and compared two types of algorithms, leading to the new fusion of them presented in this paper: semantic pattern discovery to segment each pattern from the image and convolutional neural networks (CNNs) activation analysis for detecting the recurring pattern units. In the text below, we will first tackle each separately, to then demonstrate what our novel synthesis provides as assets and improvements. 2.1 Repeating pattern segmentation After having scanned the existing repeating pattern detection algorithms to identify the most apt one for the task here at hand (Das and Mandal, 2018; Mahmudnejad et al., 2023; Qu et al., 2020; Sakai et al., 2024), a first algorithm that showed potential was "Unsupervised Semantic Discovery through visual patterns detection" by Pelosin et al. (2021). It is the most optimized state-of-the-art open source algorithm that can detect multiple patterns by deploying the Canny algorithm, DAISY and SLIC techniques. The algorithm works independently of a fixed geometric scheme and proved more performant than GRASP (Liu and Liu, 2013). This makes it possible to detect a wider 59 Sugata Das and Sekhar Mandal. Figure spotting in indian heritage image. Journal of Cultural Heritage, 32:133–143, 2018. Floris De Feyter, Kristof Van Beeck, and Toon Goedemé. Deep visual recognition for the real world. Phd dissertation, 2023. P. Garnsey and R. Saller. The Roman Empire: Economy, Society and Culture. University of California Press, 2015. P. Garnsey and R.P. Saller. The Roman Empire: Economy, Society and Culture. University of California Press, 1987. Charles Gates. Ancient Cities: The Archaeology of Urban Life in the Ancient Near East and Egypt, Greece, and Rome. Classical studies and archaeology. Routledge, 2011. René Ginouvès and Roland Martin. Dictionnaire méthodique de l’architecture grecque et romaine. 1 Matériaux, techniques de construction, techniques et formes du décor. Collection de l’École française de Rome 84,1. Ecole française de Rome, Rome, 1985. Christian Gliwitzky. Späte Blüte in Side und Perge: Die pamphylische Bauornamentik des 3. Jahrhunderts n. Chr. Peter Lang Verlag, Lausanne, Switzerland, 2010. doi: 10.3726/978-3-0351-0128-7. Kai Jes, Richard Posamentir, and Michael Wörrle. Der tempel des zeus in aizanoi und seine datierung. In Klaus Rheidt, editor, Aizanoi und Anatolien: neue Entdeckungen zur Geschichte und Archäologie im Hochland des westlichen Kleinasien, pages 59–87. 2010. Gamze Kaymak. Die Cumanın Camii in Antalya: ihre Baugeschichte und ihre byzantinischen Ursprünge: Bauaufnahme, Bauforschung, Denkmalpflege = Antalya Cumanın Camii: mimari tarih ve Bizans kökeni: rölöve, yapı analizi, anıt koruma ve bakımı. Number 9 in Adalya Supplementary Series. Suna-Inan Kıraç Akdeniz Medeniyetleri Araştirma Enstitüsü, 2009. Karin Koller, Ursula Quatember, and Elisabeth Trinkl, editors. Stein auf Stein. Number 9 in Keryx. Unipress Verlag, 2021. Louis Lettry, Michal Perdoch, Kenneth Vanhoey, and Luc Van Gool. Repeated pattern detection using cnn activations. In 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 47–55. IEEE, 2017. Jingchen Liu and Yanxi Liu. Grasp recurring patterns from a single view. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2003–2010, 2013. Ute Lohner-Urban and Ursula Quatember. Zwischen Bruch und Kontinuitat: Architektur in Kleinasien am Ubergang vom Hellenismus zur römischen Kaiserzeit. Byzas 25. Ege Yayınları, Istanbul, 2020. A. Mahmudnejad, E. Andaroodi, and M. Saadatseresht. Advanced clustering of architectural geometric ornaments using small scale machine learning, case study of ilkhanid geometric patterns. ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, X-4/W1-2022:417–422, 2023. doi: 10.5194/isprs-annalsX-4-W1-2022-417-2023. Christian Marek. In the Land of a Thousand Gods: A History of Asia Minor in the Ancient World. Princeton University Press, 2016. 66 D.J. Mattingly. Imperialism, Power, and Identity: Experiencing the Roman Empire, volume 1 of Miriam S. Balmuth Lectures in Ancient History and Archaeology. Princeton University Press, 2011. Francesco Pelosin, Andrea Gasparetto, Andrea Albarelli, and Andrea Torsello. Unsupervised semantic discovery through visual patterns detection. In Structural, Syntactic, and Statistical Pattern Recognition: Joint IAPR International Workshops, S+ SSPR 2020, Padua, Italy, January 21–22, 2021, Proceedings, pages 272–281. Springer, 2021. M. Pitts and M.J. Versluys. Globalisation and the Roman World: World History, Connectivity and Material Culture. Cambridge University Press, 2015. Jeroen Poblome, Peter Talloen, Inge Uytterhoeven, and Julie Verlinden. Spolia, stenen met een verhaal. Peeters; Leuven, 2022. Hong Qu, Yanghong Zhou, K. P. Chau, and P. Y. Mok. Repeated pattern extraction with knowledge-based attention and semantic embeddings. Proceedings of the 14th IADIS International Conference Computer Graphics, Visualization, Computer Vision and Image Processing 2020, CGVCVIP 2020 and Proceedings of the 5th IADIS International Conference Big Data Analytics, Data Mining and Computational Intelligence 2020, BigDaCI 2020 and Proceedings of the 9th IADIS International Conference Theory and Practice in Modern Computing 2020, TPMC 2020 - Part of the 14th Multi Conference on Computer Science and Information Systems, MCCSIS 2020, pages 99–106. IADIS Press, 2020. Ali M Reza. Realization of the contrast limited adaptive histogram equalization (clahe) for real-time image enhancement. Journal of VLSI signal processing systems for signal, image and video technology, 38:35–44, 2004. Frank. Rumscheid. Untersuchungen zur kleinasiatischen Bauornamentik des Hellenismus. Number Bd. 14 in Beiträge zur Erschliessung hellenistischer und kaiserzeitlicher Skulptur und Architektur. P. von Zabern, 1994. Masato Sakai, Akihisa Sakurai, Siyuan Lu, Jorge Olano, Conrad M. Albrecht, Hendrik F. Hamann, and Marcus Freitag. Ai-accelerated nazca survey nearly doubles the number of known figurative geoglyphs and sheds light on their purpose. Proceedings of the National Academy of Sciences, 121(40):e2407652121, 2024. A. Van Oyen and M. Pitts. Materialising Roman Histories. University of Cambridge Museum of Classical Archaeology Monographs. Oxbow Books, 2017. Lutgarde Vandeput. The architectural decoration in Roman Asia Minor: Sagalassos : a case study. Studies in Eastern Mediterranean archaeology 1. Brepols, Turnhout, 1997. A. Wallace-Hadrill. Rome’s Cultural Revolution. Cambridge University Press, 2008. Rinse Willet. The geography of urbanism in Roman Asia Minor. Equinox Publishing Ltd., Sheffield, South Yorkshire, 2020. F.K. Yegül and D.G. Favro. Roman Architecture and Urbanism: From the Origins to Late Antiquity. Cambridge University Press, 2019. Arjan Zuiderhoek. The Politics of Munificence in the Roman Empire: Citizens, Elites and Benefactors in Asia Minor. Cambridge University Press, 2009. 67 cb 2025. This work is licensed under a Creative Commons “Attribution 4.0 International” license. Breaking Down Barriers: The Transition of ODIS from a Relational to a Triple Store Database Bert Aernouts1, Mehmet Celik2, and Joris Colla1 1KADOC-KU Leuven/ODIS 2LIBIS-KU Leuven In this paper, we discuss the ongoing transformation of the ODIS database – a key digital infrastructure for historical research on civil society in Flanders and Brussels – from a relational to a triple store architecture. Originally developed in the early 2000s and overhauled between 2009 and 2013, ODIS is now undergoing a comprehensive renewal to align with linked data principles and the FAIR data guidelines. The transition involves the implementation of a Virtuoso triple store, a redesigned RDF-based data model, and a new user interface. Central to this effort is the development of SOLIS (Smart Ontology Layer for Interoperable Systems), a tool that bridges domain expertise and technical implementation by generating SHACL-based APIs from spreadsheet-defined ontologies. The paper discusses the rationale behind adopting a triple store, the challenges of semantic modelling, multilingual data handling, and the integration of external authorities such as OpenStreetMap. It also highlights the benefits of semantic richness, interoperability, and enhanced data visualisation through network and geographical tools. This case study offers insights and best practices for digital humanities projects seeking to modernise their data infrastructures and foster interdisciplinary, data-driven research. Keywords: civil society, contextual disclosure of heritage, digital humanities, heritage and research database, linked open data 1 Introduction As research trends in the humanities change over the years, tools evolve to follow and sometimes even shape these new approaches. ODIS, a contextual relational database on the history of civil society that is used by a growing number of research and heritage organisations in Flanders and Brussels, was developed between 2000 and 2003. From 2009 to 2013, the database underwent a major overhaul to better meet the needs of its users. With the broader adoption of linked data in cultural heritage and humanities research, and with researchers demanding more advanced tools for data visualisation 69 and analysis, the database is currently undergoing another renewal process, propelled by an infrastructure grant from the Research Foundation Flanders (FWO). At the core of the ‘ODIS renewal project 2022–2027’ is the development of a new triple store database and a new user interface. This stems from a desire to create new ways to query, analyse, and visualise the ODIS datasets, as well as to develop an interoperability and discovery layer. This paper focuses on the backend transition from a relational to a triple store database, discussing the decisions, methods, and workflows implemented to enable the desired functionalities for the new user interface. We first provide some background information on ODIS, its history, and the current architecture of the database. We then discuss the approaches to data(re)modelling, as well as the tools used and created. Our aim is to offer an account of our efforts and journey through this transition. We seek to share our experiences with the digital humanities community, advocate for a linked data approach, and present some best practices.1 2 ODIS: Past and Present ODIS is a research-driven relational and contextual database on the history of civil society (ODIS, n.d.b). It was developed between 2000 and 2003, thanks to a grant from the Research Foundation Flanders (FWO – Max Wildiers Fund) (Heyrman and Weber, 2007). Four major cultural archives in Flanders, all seeking a system to store, interconnect, and disclose contextual metadata related to their heritage collections, formed the basis of the project: ADVN - archive for national movements, Amsab-Institute of Social History (Amsab-ISG), KADOC-KU Leuven, and Liberas (the former Liberal Archives). Other heritage and research organisations and projects in Flanders and Brussels soon joined the partnership, making ODIS one of the most widely used historical research instruments in Flanders. 2 As ODIS does not receive structural funding, its management and maintenance are covered by the annual operational usage fees paid by the partners. In 2006, the management of the database was entrusted to a non-profit association under Belgian law. 3 The universities of Antwerp (UA), Brussels (VUB), Ghent (UGent), and Leuven (KU Leuven) are represented in its General Assembly. KADOC, the interfaculty Documentation and Research Centre on Religion, Culture and Society at KU Leuven, is responsible for the day-to-day management of ODIS, while LIBIS-KU Leuven is the main technical service provider. From 2009 to 2014, a research infrastructure grant from the Hercules Foundation of the Flemish authorities enabled the development of a new database (De Maeyer, 2015). The data structure of ODIS was expanded with new modules for the description of buildings, families, and events. Bilingualism (Dutch–English) became a key feature of the database, significantly increasing its usability, though also adding complexity to the data model. A mix of language-dependent and language-independent fields makes up the descriptions, 1 We would like to thank Peter Heyrman (KADOC-KU Leuven), Stephan Pauls (LIBIS-KU Leuven), Rachele Ricceri (LIBIS-KU Leuven), Winand Van Meerbeek (KADOC-KU Leuven), and Roxanne Wyns (LIBIS-KU Leuven) for their valuable input. 2 A complete overview of the ODIS partners, with links to their websites, is available on the ODIS website (ODIS, n.d.a). 3 The bylaws of the non-profit organisation (vzw) ‘Onderzoekssteunpunt en Databank Intermediaire Structuren in Vlaanderen (19de-20ste eeuw)’ were published in the Belgian Official Gazette (Belgisch Staatsblad), 13 November 2006, 884.703.544. Their most recent update was published on 8 January 2024 (noa, 2024). 70 resulting in shared data between the two language versions of a record, but also discrepancies. Translation projects aimed at closing this information gap have already yielded significant progress, though the issue remains on the agenda. Stimulating cross-fertilisation between heritage and research has always been a major goal of ODIS. By offering researchers and heritage experts a shared platform, the database aims to foster interdisciplinary research into the history and heritage of civil society. First and foremost, it provides its partners with a stable, reliable, and user-friendly environment to centralise and store contextual information on civil society and its documentary heritage. This enables them to safeguard, supplement, update, and correlate their dispersed datasets and repertories, keeping them accessible in the long term and ensuring the reproducibility of research. Secondly, the system is used to validate data and publish them on the World Wide Web, via the ODIS public catalogue (OPAC) and specific OPACs on the websites of partners and projects. Finally, ODIS enables its partners to analyse their datasets by means of advanced search tools (Angelaki et al., 2019; Colla and Heyrman, 2019; Heyrman, 2025). The broad scope of the ODIS partnership is reflected in the comprehensive content of the database. Many thematic fields are addressed at the international, national, regional, and local levels: politics, social movements and organisations, culture, art and architecture, religion, national movements, education, care provisions, migration, and more. In May 2021, the ODIS datasets on religion, culture, and society, managed by KADOC, were recognised as a core research facility of KU Leuven (KADOC-KU Leuven, 2024; KU Leuven, 2024). As of early 2024, ODIS contained 308,144 records on organisations, persons, families, buildings, events, archives, and publications (mainly periodicals). Nearly 48% of them are available for consultation under a Creative Commons licence (BY-NC-SA 4.0), facilitating their reuse in other initiatives and projects. In 2023, the ODIS OPAC was visited 91,642 times by 51,517 unique visitors, resulting in 604,982 page views. Visitors use the database in several ways: 1) as an encyclopaedia; 2) as a heuristic instrument; 3) as an authority database; and 4) as a digital humanities research tool.4 The current ODIS data model consists of eight interlinked modules (see Figure 1). The structure of the modules is based on international standards wherever possible: ISAAR(CPF) for organisations, persons, and families; the DOCOMOMO Guidelines for buildings; ISAD(G) for archives; ISBD for publications; and ISDIAH for repositories (Docomomo International, n.d.; International Council on Archives, 2024a,b,c; International Federation of Library Associations, 2011). All modules share a parallel structure. Each includes ‘authority entries’ for identification data (e.g. names, titles, dates and places of birth and death). Furthermore, data inputters can use both free text fields and repeatable fields/field groups. In the latter category, systematic data input is carried out based on validated thesauri and lists of choices. Finally, ODIS features relational field groups. These ensure the interconnection between the ODIS modules and enable the creation of clear and unambiguous links between records. By interconnecting records, the ODIS data model makes it possible, for example, to map networks and discover unexpected connections. With some research effort, the contextual information clusters in the database provide a multifaceted view of civil society in Flanders/Belgium, its key players and documentary heritage. Since only some authority entries are mandatory, authors and partners can determine the scope and depth of their datasets in alignment with future research objectives. They also 4 More information on the content and use of ODIS can be found in the ODIS annual reports, available on the ODIS website (in Dutch) (ODIS, n.d.a). 71 Figure 1: Entity-relationship diagram of the ODIS database. independently determine whether they would like to publish their series in the OPAC. Striving for versatility and cooperation, ODIS seeks to address and even anticipate the technical needs of its diverse partners and users. These users are a balanced mix of academic researchers, heritage experts, individuals involved in local historical research, genealogists, journalists, and other information seekers. As several features of the database have become somewhat outdated and no longer provide the expected functionalities, nearly ten years after the so-called ‘ODIS Hercules Project 2009–2014’, another comprehensive renewal of ODIS became necessary. Supported by a research infrastructure grant from FWO and benefiting from a technical collaboration with LIBIS-KU Leuven, the database has embarked on a transformative journey since 2022. 5 The ‘ODIS renewal project 2022–2027’ aims to update the mission and goals of ODIS, foster data input and data quality management, strengthen participation, and implement a revamped communication strategy. However, its core ambition is the technical renewal of the database. Although the project resources did not allow for extensive user research, ODIS had a good understanding of the technical needs of its backend and frontend users through a user survey conducted in 2020, and through several technical workshops at the annual ODIS community day. A key element is the development of new ways to query, analyse, and visualise the ODIS datasets, such as network and geographical visualisation. This will enable ODIS to provide its users with more hands-on research tools. Secondly, we aim to develop an interoperability and discovery layer, offering more and better connections with other catalogues, research instruments, platforms, and linked open data resources. This serves as the cornerstone of a sustainable open access policy, based on the so-called FAIR principles (Findable, Accessible, Interoperable, Reusable) (GO FAIR initiative, 2022). Endpoints (JSON:API and SPARQL endpoint) will enable end users – both ODIS partners and external users – to retrieve data from ODIS for reuse in other databases or within research tools. Rights and responsibilities of the data creators and users will be defined in a data management and sharing plan, which is currently being developed. The endpoints will also facilitate the semantic enrichment of the ODIS content through interconnections with complementary datasets. To achieve these project goals, the development of a new database and user interface is underway. For the database, it was decided to switch from a relational Oracle database to a triple store database developed in Virtuoso. In the following sections of this contribution, we will discuss the challenges involved in the backend renewal of 5 Research Foundation Flanders, I011122N. The supervisors of the renewal project are Prof. Dr. Kim Christiaens (KADOC-KU Leuven), Dr. Peter Heyrman (KADOC-KU Leuven), and Jo Rademakers (LIBIS-KU Leuven). 72 Figure 2: Left: the current ODIS search interface. Right: the new ODIS search interface. ODIS in more detail. 3 A Database in Transition: ODIS Yet to Come In the current phase of the project, a new frontend has been envisioned to maintain a strong connection with users’ evolving needs. To implement this renewed vision within the existing infrastructure, a new backend had to be designed and developed. To better understand the choices made for the backend, allow us to outline a vision of ODIS Yet to Come. Thewebsite willundergoa facelift, featuring anewlogo, newbrandcolours, updated graphics, and more (see Figure 2). Users will be able to query the OPAC more easily through updated simple search functionalities, a more user-friendly advanced search, the ability to filter the result set using facets, and the option to represent the result list either in a table or in a more detailed manner. All of this should help accommodate researchers of all levels. Navigation will be more intuitive within the website and between the various records, allowing for seamless movement back and forth within the OPAC. We also aim to enable a ‘re-presentation’ of the rich data already stored within ODIS by offering new ways to explore them both visually and analytically. Not only will the records have a new layout, but there will also be an integrated geographical tool to plot results on a map, helping to contextualise the data spatially. To conclude our vision of ODIS Yet to Come, we must mention one final visualisation tool: the network graph. Inspired by, for instance, the Gemeinsame Normdatei (GND), managed by the German National Library, we will integrate a network representation within our records (GND-Explorer., n.d.). Users will be able to open the network from a record and further explore the multitude of relationships interconnecting records, hopefully discovering network clusters at a higher level. In short, the current OPAC will be updated to meet the latest standards in search functionalities, along with some 73 useful extras. Tackling the renewal of the entire infrastructure is no easy task, and neither is describing the process. To meet all the needs of the envisaged frontend, the entire backend architecture must be stable and versatile enough to support it. This can be broken down into a few components: a new database environment, a new data model to structure that environment, and a new API layer to facilitate the exchange of data between the database and the OPAC, as well as between the database and external authorities. The first task we faced was selecting a new database platform that would meet our needs. Our decision was partially influenced by the type of data model we intended to use. The cultural heritage sector has increasingly shifted towards a linked data approach, exemplified by the International Council on Archives (ICA) launching its new description standard, Records in Contexts (RiC), 6 and the Flemish authorities with their Open Standards for Linking Organisations (OSLO) (International Council on Archives, 2024d; Vlaanderen, n.d.). 7 Given that our frontend applications strongly suggest a linked data framework, this direction appeared to be the most logical choice. Separate from the frontend requirements, the opportunities for the research data in ODIS to be disseminated more widely and at the same time being enriched by external authorities are too important to be overlooked. ODIS data were already interlinked by the many relationships between the different modules, such as between organisations and persons. These links, however, were not always as easy to perceive on a larger scale. This is where an RDF (Resource Description Framework) approach could allow the data to shine. Although RDF models can be stored in relational databases – typically yielding high-quality results (Hernández et al., 2016) – we decided to embrace the triple store approach. 8 While many are familiar with the concept of relational databases, it is helpful to briefly describe what a triple store entails. A triple store is a database optimised for storing and querying RDF data. Conceptually, it functions as a key-value store, where the key is split into a subject (identifier) and predicate (property name). Unlike relational databases, triple stores lack predefined table structures or constraints, instead relying on standards like RDF and SHACL for data organisation and validation. While triple stores offer unparalleled flexibility, their layered approach can make them difficult to adopt. Developers often perceive linked data technologies as overly complex, with limited immediate value. This situation underscores the importance of having an API that takes care of validation and database communication. Several challenges are connected to the adoption of a triple store database, often cited by developers as reasons to refrain from embracing this novel approach. Firstly, it is well known that linked data has a steep learning curve, involving effort and time investment that are often unavailable within the constraints of a specific project. Secondly, some inherent challenges of adopting a triple store system include higher complexity when integrating with other platforms, such as frontends. Specific queries 6 The International Council on Archives (ICA) launched its first drafts of a conceptual model and ontology for Records in Contexts in 2016 and most recently published version 1.0 in 2023. 7 The Flemish authorities have been rolling out linked open data standards across several fields and have been in the process of working on a standard for cultural heritage since 2020. Their aim is to help facilitate data exchange within the heritage sector in Flanders, for example, between archives and museums, but also to connect with international players. 8 Ravat et al. (2020) have shown that for querying vast datasets, relational databases still have the upper hand. However, we believe that a triple store approach is appropriate for our case since, despite being large in comparison to other historical research infrastructures, our dataset (308,144 records as of 1 January 2024) is still relatively small compared to those in most performance tests. 74 must be constructed to make relevant information human-readable. In terms of maintenance, we still observe more complex debugging, as well as performance issues for larger triple stores. Moreover, as ecosystems are not yet as expanded as RDBMSs, triple stores come with a tooling complexity that can be perceived as too high at first. Yet, we believe that the path to linked data can and should be pursued using triple stores, and that these challenges are outweighed by numerous long-term advantages. Firstly, data models are more flexible and adaptable to the needs of evolving projects, accommodating the needs of different users. A new record structure or graph can thus be easily created using construct queries. Additionally, a significant added value of a triple store approach includes semantic richness and interoperability. These two reasons support our choices. Semantic richness is guaranteed by the possibility of referring to different schemas (e.g. schema.org), and to maintain consistent values. Agent entities are more easily integrable into the data model. Interoperability is a major asset of the new data model built using triple stores, as it is achieved both by smoothly exporting information to other triple stores and by acquiring data from external systems. The apparent tooling complexity can be overcome by implementing an appropriate API layer to facilitate data exchange. Finally, the relative performance issues that can be experienced when running large triple stores are mitigated by using Elasticsearch. We were not the first humanities database infrastructure in Flanders to transition from a relational to a triple store database. Archiefbank, the database of Archiefpunt, made a similar transition between 2020 and 2022 (Archiefpunt, n.d.; FARO Vlaams steunpunt voor cultureel erfgoed vzw., 2020). 9 Like our colleagues at Archiefpunt, we opted for a Virtuoso Triple Store. Virtuoso is open-source software that provides all the functional necessities for our project: it is a well-maintained product, offers highperformanceSPARQL querying forefficient handlingof complexqueries andoptimised RDF graph traversal, supports federated SPARQL queries enabling remote RDF data source queries, is scalable for large datasets through distributed deployment, and includes reasoning and inference support. Additionally, it facilitates data integration across multiple formats, such as XML, CSV, and other databases, while ensuring ACID-compliant transactions. 10 From a financial perspective, it also offers excellent support at a relatively low cost. It is important, however, not to become dependent on one database platform. The way we have built our data model and surrounding infrastructure (as discussed in the next section) ensures that we can easily migrate between platforms if and when we choose to do so. The only requirement for the platforms is that they have an HTTP SPARQL endpoint. 4 Working out Solutions, Challenges Along the Way The successful implementation of data-driven applications in the digital humanities hinges on the seamless integration of domain expertise with technological proficiency. While scholars possess invaluable knowledge of historical sources, cultural contexts, and research methodologies, they often lack the technical expertise to translate their 9 Archiefpunt vzw is a cultural heritage organisation that aims to raise awareness of private archives in Flanders and Brussels. Its Archiefbank database makes thousands of archival descriptions available to the public. In their ‘DATA project’, spanning 2020 to 2022, the infrastructure (backend and frontend) underwent a total facelift to better serve its stakeholders. 10 ACID (Atomicity, Consistency, Isolation, and Durability) is a set of rules expected from database platforms. 75 to the new one, the existing structure had to be converted to the new triples format. Mappings were made to facilitate an easy transition, which served as a useful test to identify any omissions or logic errors in the new model. Once the conversion scripts were completed, the newly converted data were ingested in the triple store via a JSON:API. This functioned as a preliminary test to ensure that the necessary constraints and logic were being enforced. The process resulted in several iterations of the model, as well as updates to the SHACL and API. This highlighted the advantages of this approach, allowing for easy adjustments to the model in the spreadsheet and seamlessly triggering subsequent changes to the SHACL and API. 5 Concluding Remarks In this paper, we shared insights from the ongoing technical renewal project for the contextual database ODIS. The journey to the new ODIS environment is an adventurous one; we are currently halfway through and on track to complete the project on time. We have established a new triple store database in Virtuoso, reworked the data model, and set up a user-friendly workflow for easy modelling and API generation with SOLIS. Furthermore, we have successfully converted the data and rigorously tested the setup. The shift to a triple store approach enables us to break down barriers not only between ODIS and other data repositories but also between the separate modules within the database. Additionally, it has enabled us to align the ODIS data model with Records in Contexts. Alongside redesigning the database structure, the public interface of ODIS will undergo a revamp. The technical renewal programme aims to strengthen the flexibility and multifunctional use of ODIS, fostering its broad applicability in individual and collaborative scientific research, all within a multidisciplinary and international-comparative perspective. The system will help ODIS users compose and query consistent data packages. It will also enable them to visualise and analyse ODIS datasets in novel ways, supporting, for example, a more in-depth understanding of associational life, the prosopographical analysis of civil society organisations, the visualisation of biographical and associational itineraries, the spatial and territorial impressions of civil society, and more. In addition to opening new perspectives for data-driven research in the social sciences and humanities, we are confident that the investment programme will also encourage more research units to join the ODIS partnership and incorporate their own data series, further strengthening the thematic scope and research potential of the database. We hope that our journey will inspire similar projects and demonstrate that it is possible to navigate the challenging middle ground between business and technical teams. We also hope to have offered convincing reasons for developers to take the leap into the RDF domain and overcome the obstacles. Our approaches can serve as a foundation for others to further develop their own workflows and streamline their transition processes. References Onderzoekssteunpunt en Databank Intermediaire Structuren in Vlaanderen (19de20ste eeuw). Belgisch Staatsblad, 2024. URL https://www.ejustice.just.fgov.be/ tsv_pdf/2024/01/08/24004971.pdf. 82 Bert Aernouts and Mehmet Celik. ODIS Ontologie, 2025. URL https://data.odis. be/_doc. Georgia Angelaki, Karolina Badzmierowska, David Brown, Vera Chiquet, Joris Colla, Judith Finlay-McAlester, Klaudia Grabowska, Vanessa Hannesschläger, Natalie Harrower, Freja Howat-Maxted, Maria Ilvanidou, Wojciech Kordyzon, Magdalena Król, Antonio Gabriel Losada Gómez, Maciej Maryl, Sanita Reinsone, Natalia Suslova, Mark Sweetnam, Kamil Śliwowski, and Marcin Werla. How to Facilitate Cooperation between Humanities Researchers and Cultural Heritage Institutions. Guidelines. Technical report, Digital Humanities Centre at the Institute of Literary Research of the Polish Academy of Sciences, Warsaw, Poland, March 2019. URL https://zenodo.org/records/2587481. Archiefpunt. Home - Archiefpunt, n.d. URL https://archiefpunt.be/. Mehmet Celik. GitHub - mehmetc/solis: Turn a SHACL file into an API, April 2025. URL https://github.com/mehmetc/solis. original-date: 2021-09-17T12:17:43Z. Joris Colla and Peter Heyrman. Het middenveld en zijn erfgoed in context: het veelzijdig gebruik van de onlinedatabank ODIS (2000-2017). In D Leyder and N Istasse, editors, Ten dienste van de gebruiker. Over de zoekinstrumenten ontwikkeld door archivarissen en bibliothecarissen. Akten van het colloquium van 8 december 2017, volume 105 of Archiefen Bibliotheekwezen in België, pages 99–111. Archiefen Bibliotheekwezen in België, 2019. Jan De Maeyer. ODIS: Database Intermediary Structures Flanders (AKUL043). Final Report, June 2015. Technical report, 2015. URL https://www.odis.be/hercules/ docs/FinalReport_AKUL043_ODIS.pdf. Docomomo International. Docomomo International – Architecture Archive, n.d. URL https://docomomo.com/. FARO Vlaams steunpunt voor cultureel erfgoed vzw. Digitale Architectuur Traject Archiefpunt (DATA). Verbetering van gebruiksvriendelijkheid en interconnectiviteit in Linked Open Data tijd, fase 1 | FARO Vlaams steunpunt voor cultureel erfgoed vzw, 2020. URL https://faro.be/project/digitale-architectuur-trajectarchiefpunt-data-verbetering-van-gebruiksvriendelijkheid-en. GND-Explorer. Homepage, n.d. URL https://explore.gnd.network/. GO FAIR initiative. FAIR Principles - GO FAIR, 2022. URL https://www.go-fair. org/fair-principles/. Daniel Hernández, Aidan Hogan, Cristian Riveros, Carlos Rojas, Enzo Zerega, Fabian Flöck, Elena Simperl, Marta Sabou, Alasdair Gray, Paul Groth, Freddy Lecue, Yolanda Gil, and Markus Krötzsch. Querying Wikidata: Comparing SPARQL, Relational and Graph Databases. In The Semantic Web–ISWC 2016: 15th International Semantic Web Conference, Kobe, Japan, October 17–21, 2016, Proceedings, Part II 15, volume 9982 of Lecture Notes in Computer Science, pages 88–103. Springer International Publishing AG, Switzerland, 2016. ISBN 0302-9743. doi: 10.1007/978-3-319-46547-0_10. Peter Heyrman. Online tools and instruments supporting the study of transnational elites: the functionalities of the contextual database ODIS. In Andrea Ciampani 83 and Thomas Kroll, editors, Transnational Encounters?, European Elites, International Associations and National States (1882-1914). De Gruyter Oldenbourg, 2025. ISBN 978-3-11-168059-0. Peter Heyrman and Donald Weber. Intermediaire structuren in Vlaanderen: ODIS. In Marijke Hoflack, editor, Het Max Wildiersfonds geëvalueerd: archief en wetenschappelijk onderzoek., Archiefkunde : verhandelingen aansluitend bij Bibliotheeken archiefgids 9, pages 73–80. VVBAD Vlaamse vereniging voor bibliotheek-, archiefen documentatiewezen, Antwerpen, 2007. ISBN 978-90-72679-32-1. International Council on Archives. ISAAR (CPF): International Standard Archival Authority Record for Corporate Bodies, Persons and Families, 2nd Edition, 2024a. URL https://www.ica.org/resource/isaar-cpf-international-standardarchival-authority-record-for-corporate-bodies-persons-and-families2nd-edition/. International Council on Archives. ISAD(G): General International Standard Archival Description - Second edition, 2024b. URL https://www.ica.org/resource/isadggeneral-international-standard-archival-description-second-edition/. International Council on Archives. ISDIAH: International Standard for Describing Institutions with Archival Holdings, 2024c. URL https: //www.ica.org/resource/isdiah-international-standard-for-describinginstitutions-with-archival-holdings/. International Council on Archives. Records in Contexts - Ontology, 2024d. URL https://www.ica.org/resource/records-in-contexts-ontology/. International Federation of Library Associations. International Standard Bibliographic Description Consolidated Edition 2011, 2011. URL https://www.ifla.org/wp-content/uploads/2019/05/assets/cataloguing/ isbd/isbd-cons_20110321.pdf. Yiyao Jin. The Network Visualization Approach for the ODIS Database. PhD thesis, KU Leuven. Faculteit Wetenschappen, Leuven, 2024. JSON:API. JSON:API — A specification for building APIs in JSON, n.d. URL https: //jsonapi.org/. KADOC-KU Leuven. KADOC - ODIS: KU Leuven Core Facility on Religion, Culture and Society, 2024. URL https://kadoc.kuleuven.be/english/3_research/31_ ourresearch/ODIS. KU Leuven. KU Leuven Core Facilities, 2024. URL https://research.kuleuven.be/ en/research-areas/core-facilities. ODIS. About ODIS, n.d.a. URL https://www.odis.be/hercules/_en_overODIS. php. ODIS. Home, n.d.b. URL https://www.odis.be. OpenStreetMap. OpenStreetMap, n.d. URL https://www.openstreetmap.org/. 84 Franck Ravat, Jiefu Song, Olivier Teste, and Cassia Trojahn. Efficient querying of multidimensional RDF data with aggregates: Comparing NoSQL, RDF and relational data stores. International Journal of Information Management, 54:102089, October 2020. ISSN 0268-4012. doi: 10.1016/j.ijinfomgt.2020.102089. URL https: //www.sciencedirect.com/science/article/pii/S0268401219306097. Vlaanderen. OSLO, n.d. URL https://www.vlaanderen.be/digitaal-vlaanderen/ onze-diensten-en-platformen/oslo. W3C. Shapes Constraint Language (SHACL), July 2017. URL https://www.w3.org/ TR/shacl/. Guohui Xiao, Diego Calvanese, Roman Kontchakov, Domenico Lembo, Antonella Poggi, Riccardo Rosati, and Michael Zakharyaschev. Ontology-Based Data Access: A Survey. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, pages 5511–5519, Stockholm, Sweden, July 2018. International Joint Conferences on Artificial Intelligence Organization. ISBN 978-0-9992411-2-7. doi: 10.24963/ijcai.2018/777. URL https://www.ijcai.org/proceedings/2018/777. 85 cb 2025. This work is licensed under a Creative Commons “Attribution 4.0 International” license. Linking Het Amsterdams Stadsjournaal: A Case Study in Emerging Linked Open Data (LOD) Approaches to Audio-Visual Heritage Meg Weijers1and Christian Olesen2 1Eye Filmmuseum 2University of Amsterdam Taking the case of Het Amsterdams Stadsjournaal (ASJ) as a starting point, this article aims to advance emerging Linked Open Data (LOD) approaches to AV heritage. The archiving of ASJ materials,a political film collective that produced politically engaged documentaries between 1974 and 1984, raises several challenges and questions concerning archival access to AV heritage and standardisation because the materials preserved are located at three different institutions with diverging missions and user groups: respectively Eye Filmmuseum (Eye), the Netherlands Institute for Sound and Vision (NISV), and the Stadsarchief Amsterdam (SAA). Addressing this circumstance, the article evaluates the potential of Linked Open Data as a way to connect and enhance data on ASJ to make the preserved materials more widely accessible. This entails critically discussing the current access given to ASJ films in the CLARIAH Media Suite in light of the infrastracture’s Wikidata features, focusing on how these functionalities may accommodate the user requirements identified at the institutions preserving ASJ materials. By emphasizing the specific challenges of AV archives, the article also seeks to broaden the scope of current digital humanities discussions concerning LOD as these have hitherto had a tendency to leave out collections preserved by AV heritage institutions. Keywords: audiovisual archiving, linkedopen data, documentary cinema, Wikidata 1 Introduction Focusing on the example of Het Amsterdams Stadsjournaal, this article aims to advance discussions around emerging Linked Open Data (LOD) approaches to AV heritage. 1 1 The research for this article was initiated in the context of Meg Weijers’ MA thesis, and further developed in the context of the CLARIAH Media Suite Learn initiative. 87 Het Amsterdams Stadsjournaal, henceforth ASJ, was a political film collective that produced critical documentary productions about the misconducts and malpractices in capitalist society between 1974 and 1984 (Heijs, 1984). The archiving of the ASJ materials, and their associated datasets, raises several challenges and questions concerning archival access to AV heritage and standardisation because the materials are located at three different institutions with diverging missions, user groups and definitions of film works: Eye Filmmuseum (Eye), the Netherlands Institute for Sound and Vision (NISV), and Stadsarchief Amsterdam (SAA). Addressing this challenge, and elaborating on the reflection developed in the context of the DH Benelux 2024 Conference theme “Breaking Silos, Connecting Data”, this article discusses the research done on evaluating the potential of Linked Open Data as a way to connect and enhance data on ASJ to make it more widely accessible. Moreover, this article discusses and critically evaluates the access currently given to a part of ASJ’s films in the CLARIAH Media Suite and the environment’s recently developed LOD and Wikidata features, focusing on how these functionalities may accommodate the user requirements identified at the institutions. By pointing out specific challenges we deal with within AV archives, we aim to shed new light on discussions concerning LOD and digital humanities. 1.1 The Shared Custodianship of ASJ One of the challenges raised by the ways in which ASJ materials are currently accessible is the shared custodianship over the collection. Currently, the films are preserved and documented by three different institutions, each operating with a distinct perspective and definition of the films, namely as respectively art (Eye), historical artefact (NISV), and as part of a specific urban history (SAA). Due to this shared custodianship, the records have been provided with different information and metadata that lacks standardisation, a topic we shall return to later in the article in more detail. This is problematic: the archives use different databases and have ingested them for different reasons. As a result, some films have been provided with elaborate metadata whereas other films are only provided with a minimal amount, as you can see in Figures 1 and 2, screenshots taken from Eye’s database CE, and Figure 3, a screenshot taken from the NISV’s database DAAN. Because of an extensive research project by curators and volunteers of Eye on the ASJ collection, CE has an elaborate description of the film in its database, including keywords and metadata on the physical aspects of the material such as sound, color, and image ratio. In contrast, the metadata on the same film in DAAN lacks some metadata such as a description of the film but does have technical metadata and keyframes. It should be noted that this incoherence is not a unique phenomenon within audiovisual heritage practices. Even within an archive, metadata can differ greatly between films from the same collection. At Eye for instance, metadata is continuously added when films are taken out of the vaults for research or viewings. However, a lack of metadata makes it harder to depart from the collection of an archive when doing research because it means you need a considerable amount of tacit knowledge to understand the information in a database.2 88 Figure 1: Screenshot of record of Amsterdams Stadsjournaal 8 in CE Figure 2: Screenshot of record of Amsterdams Stadsjournaal 8 in CE 2 By tacit knowledge, we mean the knowledge possessed by individuals working at institutions that is not made explicit to others, for example user groups and coworkers. For a better understanding of the role of tacit knowledge within audiovisual archives, see Share That Knowledge!: A Road Map for Sharing Knowledge across Generations of Audiovisual Archivists (van Dalen and Šičarov, 2023). 89 Figure 3: Screenshot of record of Amsterdams Stadsjournaal 8 in DAAN This need for adequate and coherent metadata is even more crucial for a collection like ASJ, dispersed over multiple archives. To be able to approach them as a collection, you need to be able to link the individual films to each other. In short, the shared custodianship requires a certain coherence in metadata without changing the databases of the three archives. Here, metadata should enhance the collection’s access to a broad public. Moreover, the information needs to reflect these three perspectives on the nature of film and the different user groups of the institutes. As interest in unlocking the materials grew, an alternative way had to be found to enable access to them as a whole, while being preserved at three different institutions. The research presented in this article first approaches this challenge conceptually and practically, while experimenting with potential solutions, focusing on the following main question: • What are the potential affordances of metadata and Linked Open Data (LOD) in overcoming the challenge of shared custodianship and cataloguing contextual information in the archiving of the Amsterdams Stadsjournaal? The overarching aim of this research is to contribute to two academic discussions in the AV heritage field, namely: 1. the archiving of films that are produced outside the mainstream film circuit from the perspective of archival access, 2. the emergence of Linked Open Data in the AV heritage field. As these are relatively unknown topics in the AV heritage field, this research can be seen as an experiment with this technology, how it interacts with metadata in catalogues dealing with filmographic metadata, and to examine its potential and pitfalls within film archival practices, thus advancing discussions around LOD in connection to film and AV heritage, which still tend to be underrepresented. In this research, we particularly focus on the two LOD–technologies Wikibase and Wikidata. By taking this focus, the article furthers research around film heritage and LOD, while also building on current efforts by considering non-theatrical films, a type of film that has not yet been a main focal point for LOD in film and AV archives. Moreover, Wikibase and Wikidata are already being used by one of the custodian archives (NISV) and incorporated in the annotation functions of the CLARIAH Media Suite. Hence, Wikibase and Wikidata were logical platforms to build this research on. 2 Theoretical Framework Before diving deeper into the case of ASJ, we want to sketch the theoretical context and reference frame our endeavor is situated in. With the help of some examples, we 90 discuss archival access and linked open data in light of AV heritage to understand the link between the two and challenges that are particular to our field. In the past decades, archival access has been considered an increasingly important archival paradigm in the (AV) heritage sector. From the 1980s onwards, with computerization and digitization, archival access has increasingly become a concern for AV archives. In "Access – the reformulation of an archival paradigm", Angelika Menne-Haritz (Menne-Haritz, 2001) defines access as a professional strategy that is user-oriented. In contrast to the precedent, custodial paradigm, in which records are considered as evidence of which archivists are impartial (but not really impartial) custodians (Cook, 2013), the access paradigm is “neutral towards the content but passionate concerning openness and availability of information potentials” (MenneHaritz, 2001, p. 63). Within this paradigm, the role of the archivist is to enable the user to find their way through the archive and evaluate the relevance of the found records themselves. In other words, archival access is about providing tools rather than giving meaning to the records. Therefore, Menne-Haritz argues that archival descriptions and presentations should be user-driven, as each researcher has their own specific research question and is the only one to determine what is relevant for answering that question. For Prelinger (2009), one of the more prominent voices within this debate, archival access helps AV archivists validate our work and ensure against institutional irrelevance. As he states about the current state of archives concerning access: “Unlike public libraries, which have longestablished traditions of access, we lack a strategy that might help move us toward greater openness” (p. 168). This lack of strategy is related to a certain hesitation by many film archivists to open up their archives because of risks concerning copyright, piracy, and loss of control. In “Archives and Access in the 21st Century”, (Prelinger, 2007) explains how, traditionally, archivists have privileged preservation over access for several reasons that are more or less tied to "copyright maximalism" (p. 115). However, he argues that access is more than a binary choice between publicly usable, fully available records or only private, in-house viewings. Between these options, there is a great range of possible uses that circumvent the risks of copyright infringement. An example of this is the solutions of the CLARIAH Media Suite for works of which film archives only partially have the rights, such as the Peter Rubin Collection of the Eye Filmmuseum. Peter Rubin was an American filmmaker and VJ active in the Amsterdam club scene and German techno collectives Mayday and Love Parade (Tzialli, 2017). His VJ works contain footage of television items filmed on VHS alternating the original work. Whereas Eye is allowed to put Rubin’s work in the Media Suite, the segments with television items are at risk of copyright infringement. These parts can be blocked in the Media Suite, but users are still able to see where they were being used in the work in question. This allows users to see the ‘copyright safe’ segments in the right context. In terms of archival access, the example above shows the particular challenges in AV heritage such as copyright. The emergence of digitization and digital tools have further complicated these challenges as they have changed users’ perception of access. For many users, archival access means searching by theme or topic and accessing materials on a streaming platform, creating high expectations which often cannot be met because of insufficient finding aids (Heftberger, 2022). Cataloging is key here, as Adelheid Heftberger writes in “Access is Not a One-Way Street: The Relation Between Access to Collections and Cataloguing” (Heftberger, 2022). In this article, she explains how cataloging practices “define the kind of knowledge and concepts we make accessible 91 need a Wikipedia page and which do not. Another way in which Wikidata helps users navigate through the information is the multilingualism of the technology. As explained on their introductory page, Wikidata is multilingual which means that “Data entered in any language is immediately available in all other languages” (Wikidata, b). This enables users for whom Dutch is not their native language to access the metadata as well. More specific to the films of ASJ, this affordance can also be used to open up the archive to certain user groups such as the target groups of the films. For instance, Gebroken Tijd and Surinamers in Nederland targeted Turkish and Surinamese immigrants. This affordance helps to make the collection accessible for those more fluent in Turkish and Sranantongo (or other languages spoken in Suriname) than Dutch. Despite the helpful affordances described above, some affordances of Wikidata need more critical examination. One of the affordances is the free and openly collaborative nature of the database, which opens up some questions concerning ethics and the challenge of shared custodianship. As mentioned on Wikidata itself, their data is “published under the Creative Commons Public Domain Dedication 1.0, allowing the reuse of the data in many different scenarios” (Wikidata, b). This needs careful consideration of the information put on the Wikidata as reusing data could also mean recontextualization and loss of control. Therefore, the archives need to determine whether there is any data too sensitive to publish under creative commons. Next to ethical issues like these, one may also argue that co-editors outside the archives are not desirable in the case of ASJ. Because it is a collaborative data repository, solely using Wikidata means that all the contextual information could be edited by users outside the three archives. This could mean that the data will become unevenly described, further complicating the need for coherency between the films and thus the challenge of the shared custodianship. Therefore, the archives need to determine which data would fit in Wikidata and which data should only be created and edited by three archives in the Wikibase. 4.2 Wikibase The section above shows that Wikidata has several affordances that can help with contextualizing and linking the films. However, as previously argued, it might not be desirable to contextualize and link the films by solely using Wikidata as the archives might want to retain some control over the information stored on this platform. Here, Wikibase could be used to create an individual platform to adapt it to the unique characteristics of the ASJ collection, circumventing the limitations of Wikidata while using the technology at the same time. This does not only help to archive the collection coherently, as argued in the previous section, but it also enables the archives to create collection-specific statements. For instance, Surinamers in Nederland as an item in the Wikibase would be followed by statements about the film. Some statements would be very generic, such as the length and spoken language. However, as the target groups, aims, and cooperating organizations were also considered relevant to be added as metadata, these statements could potentially be quite confusing if target groups were used as properties in a different context in other Wikidata items. Therefore, metadata that is very specific to ASJ needs to get a property in an individual platform to be as unambiguous as possible. In this sense, it shows similarities to the conclusions of Duchesne’s and Heftbergers article on cataloguing practices. Another limitation in Wikidata that can be circumvented by creating an individual platform, is the descriptions of the items. As can be seen in Figure 4 below, items are 98 Figure 4: Screenshot of item Sprited Away and its discription minimally described. Wikidata itself has limited item descriptions, as they explain on their help page: “For those who are familiar with Wikipedia, it may be tempting to initially think of items as like the Wikidata version of Wikipedia articles. While items and articles are both pages for storing information about different concepts or topics of human knowledge, it’s important to keep in mind that Wikidata is not just a database of Wikipedia content” (Wikidata, a). In other words, the item’s description is just to distinguish the item from items with the same name. However, for the items referring to the films of ASJ, it would be helpful to state the film’s content more elaborately. Other contextual information, such as the societal debates and perspectives, can be put as statements. Here, Wikibase can be used to adapt the platform to the specific needs of the ASJ collection. In this respect, it would make sense to set up a dedicated Wikibase insofar as it offers more autonomy. Yet, on the other hand, it should also be highlighted that while such a solution may appear more accommodating when it comes to contextualizing the collections and defining ontologies, it also comes with significant technical overheads that the institutions involved may not be able to dedicate to a subcollection. 5 Conclusion The user requirement study showed which contextual information was relevant specifically for understanding ASJ as a film collective and for understanding ASJ as a form of art, a historical artefact, and part of urban history. Moreover, the findings suggested a need for providing this information on a meta-level, noting down only distinctive events, developments, or concepts without further elaboration. However, this could be problematic for users who do not have the required knowledge or research skills. To solve this problem, the interviewees suggested an external place for this kind of information, helping this specific target group further without overloading the databases. The findings from the affordances analysis implied that metadata and LOD-based technologies such as Wikidata and Wikibase can be used to capture contextual information of the individual ASJ films to adequately catalog the collection to such an extent that all users can understand it regardless of their background. Hence, their potential is to open up the collection for a broad group of users. Concerning the challenges of the ASJ collection, metadata should be used to capture the information relevant to all three institutes whereas Wikidata and Wikibase can be used to store more specific information and to link the individual films to each other, helping users 99 understand the collection as a whole. Additionally, it can offer users starting points for further research by providing links to other web pages containing contextual information. Therefore, LOD has two potential roles: to connect the films and to offer users specific tools to investigate topics, concepts, and persons further if wanted. The latter is currently facilitated in the CLARIAH Media Suite. Here, the results show the biggest potential of using LOD in audiovisual archives in light of current practices: as a way to contextualise political films without having to change an entire database. References Terugkeer-plan Surinamers verworpen. De Volkskrant, 1976. URL https://www.delpher.nl/nl/kranten/view?query=Terugkeer-plan+ Surinamers+verworpen+&coll=ddd&identifier=ABCDDD:010880581:mpeg21: a0176&resultsidentifier=ABCDDD:010880581:mpeg21:a0176&rowid=1. Kathy Baxter and Catherine Courage. Understanding Your Users: A Practical Guide to User Requirements Methods, Tools, and Techniques. Morgan Kaufmann Publishers, 2005. Terry Cook. Evidence, memory, identity, and community: four shifting archival paradigms. Arch Sci, (13), 2013. doi: 10.1007/s10502-012-9180-7. FIAF. Linked open data for film archives. Website, 2019. URL https://www.fiafnet. org/pages/E-Resources/LoD-Task-Force-Workshop-2019.html. Adelheid Heftberger. Access is not a One-way Street: The Relation Between Access to Collections and Cataloguing. Journal of Film Preservation, 107, 2022. Adelheid Heftberger and Paul Duchesne. Linked open data for film archives. Website, 2020. URL https://www.fiafnet.org/pages/E-resources/CataloguingPractices-Linked-Open-Data.html. Jan Heijs. 10 Jaar Amsterdams Stadsjournaal: 1974-1984. Uitgeverij NADA, 1984. Angelika Menne-Haritz. Access – the reformulation of an archival paradigm. Archival Science, 1, 2001. Christian Gosvig Olesen. Visualizing Film History: Film Archives and Digital Scholarship. Indiana University Press, 2025. Rick Prelinger. Archives and Access in the 21st Century. Cinema Journal, 46(3), 2007. Rick Prelinger. Points of Origin: Discovering Ourselves Through Archival Access. The Moving Image, 9(2), 2009. Roger Smither. Formats and Standards: A Film Archive Perspective on Exchanging Computerized Data. American Archivist, 50(3), 1987. URL https://americanarchivist.kglmeridian.com/view/journals/aarc/50/3/article-p324.xml. Eleni Tzialli. On handling the peter rubin collection. Website, 2017. URL https://www.eyefilm.nl/en/magazine/on-handling-the-peter-rubincollection/326365. 100 Janneke van Dalen and Nadja Šičarov. Share That Knowledge! A Road Map for Sharing Knowledge across Generations of Audiovisual Archivists. FIAF, 2023. URL https://www.fiafnet.org/images/tinyUpload/2023/11/sharethat-knowledge-digital_book-spreads.pdf. Seth van Hoogland and Ruben Verborgh. Linked Data for Libraries, Archives and Museums. Facet Publishing, 2014. Nanne van Noord, Christian Olesen, Roeland Ordelman, and Julia Noordegraaf. Automatic Annotations and Enrichments for Audiovisual Archives. ARTIDIGH, 1, 2021. Wikidata. Help: Items. Website, a. URL https://www.wikidata.org/wiki/Help: Items. Wikidata. Wikidata:introduction. Website, b. URL https://www.wikidata.org/ wiki/Wikidata:Introduction. Wikidata. Joop den Uyl. Website, c. URL https://www.wikidata.org/wiki/Q318320 . 101 cb 2025. This work is licensed under a Creative Commons “Attribution 4.0 International” license. Herbaria Heritage: Visualizing Colonial Bias in Natural History Collections Sakura Morales Furuta1∗, Ana Reviejo Salamanca2*, Michaela Todd1*, Lise Stork3, and Andreas Weber4 1Leiden University, Leiden, The Netherlands 2Vrije Universiteit, Amsterdam, The Netherlands 3University of Amsterdam, Amsterdam, The Netherlands 4University of Twente, Enschede, The Netherlands As a result of the colonial entanglements of many natural history collections, much of the world’s biodiversity heritage is housed in Europe. Increasingly, natural history institutions have begun to address their colonial past. However, in this regard, computational methods for analyzing large collections tend to consist of static visualizations of collection provenance. Moreover, in regions where cutting-edge visualization technologies are only scarcely available - often these are the same areas from which most European plant collections originate - there is a lack of simple, accessible solutions that can generate meaningful results. Thus, we argue that accessible, simple, yet interactive visualizations of collection provenance allow users to understand colonial bias in natural history collections better. Our solution allows users to focus on content gaps and highlights historical patterns and trends in collection data. Using a dataset containing metadata of five million entries from the Naturalis Biodiversity Center botanical collection as a use case, we created an interactive visualization with Microsoft Power BI. The visualization showcases the origins and movements of botanical specimens from former Dutch colonies to the Netherlands on an interactive map and timeline. This addresses not only a gap in historical research on the colonial legacy of Dutch botanical collections but also a gap in digital humanities research regarding simple and easy-to-use computational techniques for distant reading of natural heritage data. The particular use case demonstrates only a fraction of the research possibilities that this tool enables. Our interactive visualization increases the accessibility of the available scientific data, and contributes to a better understanding of the relationship between cultural history and natural history, highlighting ∗Sakura Morales Furuta, Ana Reviejo Salamanca and Michaela Todd share the first authorship of this paper. 103 the greater importance of easy-to-use, interactive and accessible visualizations of biodiversity collection histories. Ultimately, our project suggests a way forward for natural history museums not only in the Netherlands to reinterpret the colonial past of their collections. Keywords: visualization, colonial bias, natural history collections, digitization 1 Introduction While Natural History museums and herbaria in the Global North house much of the world’s biodiversity heritage, they have tended to disassociate themselves from the cultural histories inherent to their collections (Borren and Drieënhuizen, 2022; Das and Lowe, 2018; Drew et al., 2017; Kaiser et al., 2023). This lack of recognition of the histories behind collections becomes even clearer when situated alongside other museum disciplines that have been grappling with this problem for years (Haarnack, 2021). However, acknowledging the inherent colonial entanglements of natural history collections is the first step in writing more nuanced histories of these collections. This entails a greater analytical focus on the histories of acquisition and appropriation in the field (Dubald and Madruga, 2022; Hünniger, 2021) and on the subsequent movement of objects to Europe. Such a focus ensures the political context of collections is not hidden under the guise of biodiversity science and rationality, as is often the case today (Driver et al., 2021; Nadim et al., 2024). Additionally, this type of recognition becomes vital for understanding the implications of such colonial entanglements for global biodiversity and conservation efforts (Johnson et al., 2023; Park et al., 2023). The work done to uncover these hidden histories has tended to concern researchers working primarily with textual or object-focused approaches, lacking computational methods and visualizations (e.g. Alcantara-Rodriguez et al. (2022); Drieënhuizen and Sysling (2021); Jacobs and Koch (2021); Müller (2023); Wingerden (2020)). Some studies by Mohammed et al. (2022) and Park et al. (2023) utilized heat maps, world maps, etc. to illustrate botanical specimen movements across the globe and categorize trends. However, these existing visualizations, such as herbaria snapshots, lack interactivity, leaving a rich trove of information untapped. This lack of interactivity means users cannot currently engage with the different temporalities that natural historical collections and their digitized counterparts entail (Hopman, 2024). Moreover, in localities where advanced and cutting-edge visualization technologies are limited – often the same regions where many European plant collections originate – there is a shortage of simple, accessible tools capable of generating meaningful results (Buschke et al., 2023). Our tool is therefore not intended to be cutting-edge from a technological point of view. On the contrary, we seek to provide a simple, accessible and user-friendly solution that can be used as widely as possible. Given the increasing digitization of natural history collections and the availability of large datasets related to these collections (Beltrame et al., 2024; Hedrick et al., 2020; Pickering, 2024; Stork et al., 2019, 2021), there is great potential to exploit the data in simple,meaningful and importantly, interactive ways. Thus, we propose interactive visualizations as a computational technique for distant reading of natural heritage data. Focusing on the provenance of botanical specimens – specifically those stored in Naturalis Biodiversity Center (Leiden, The Netherlands) – we create an interactive and multi-layered visualization to trace historical movements of species from Indonesia 104 and Papua New Guinea 1 using Microsoft Power BI (see Figure 2). Our dashboard and accompanying use case are insightful for the processing, design and visualizations of other collection data. Since Microsoft BI is already used in the domain to visualize digitization progress and the composition of collections, we foresee a fast uptake of our solution (Breugelmans and Trekels, 2023; DiSSCo, 2023). We are of course aware that there is a deep disciplinary and institutional divide between the natural sciences and cultural, colonial and historical studies. While natural science oriented researchers consider digitized natural history collections merely as abstract and decontextualized datapoints, historians read them as entangled cultural objects with a long history of collection, appropriation, circulation and datafication. While some have suggested that hiring more diverse staff at natural history museums is the way forward to achieve a more holistic understanding of such collections (Das and Lowe, 2018), we think that an intensified training of historians in critical and reflective digital humanities, for instance with a focus on interactive visualization techniques, should play an important role to fill this epistemological and methodological gap (Díez Díaz et al., 2025). A smart use of interactive and easy to use visualization tools such as Power BI can help to reconnect datasets to their cultural and historical roots. Taken together, we hope that our paper serves as inspiration for (biodiversity) scientists, historians, data providers and a general public - not only in the Western world - to actively continue and engage in conversations about new ways of access and visualization of collection datasets in the field of natural history. To outline our research, we first describe the dataset (Section 2), before detailing our use case and the historical, practical relevance of the case. Secondly, we compare existing computational methods to further highlight the gap we aim to address in the related work (Section 3). With this background, we describe our method (Section 4), before discussing the obtained results from a qualitative perspective (Section 5). In the final section (Section 6), we reflect upon our methodology, suggesting follow-up work. Ultimately, this article aims to bridge the gap between the historical and digital, exploring how data visualization can be used as a tool to uncover the histories hidden behind scientific collections. 2 Dataset Our paper employs the “Naturalis Biodiversity Center (NL) - Botany” dataset, obtained through the GBIF website (Creuwels, 2015) 2 . GBIF, or the Global Biodiversity Information Facility (GBIF), is an international network and data portal that aggregates data from natural historical institutions all over the world (Feng et al., 2022). The physical plant specimens that this dataset represents are stored in the depots of Museum Naturalis in Leiden. This freely available tabular dataset comprises more than five million entries registering botanical specimens found in the Naturalis natural history collection. It has 51 columns with metadata about collected specimens, from which we use: country of origin (referred to as country in the dataset), recordedBy, continent, family and typeStatus. These five categories have been implemented as filters 1 We are aware of the complex colonial history of Papua New Guinea. In-depth historical research would be necessary to find out which - and even more important when - specific botanical specimens were registered under the locality of Papua New Guinea. 2 The dataset can be downloaded through the following link: https://www.gbif.org/dataset/ 15f819bd-6612-4447-854b-14d12ee1022d , however, its download is not necessary in order to run the Power BI file. 105 Figure 1: General Interface, including two pie charts, map, and timeline. within the interactive interface, as well as through a world map to plot specimen occurrences. Looking more closely at the data from Figure 1, the Netherlands accounts for 650,464 (22.21%) of Naturalis’ specimens, followed by Indonesia at 568,517 (19.42%) and Papua New Guinea at 152,914 (5.22%). Taking the Orchidaceae family as an example (see Figure 3), the visualization reveals that Indonesia leads with 19,695 occurrences (37.38%), followed by ‘Unknown’, with 13,519 specimens (25.66%), and finally Papua New Guinea with 8,437 occurrences (16.01%). Significant spikes in registered specimens occurred in 1909 (907 from 51 countries), 1936 (1,328 from 34 countries), and 1986 (1,791 from 47 countries). Orchids are interesting, since in the nineteenth and early twentieth centuries, orchids from colonial areas were an important economic luxury good and played an important role in scientific debates within the botanical sciences. Not only for individual botanists, but also for scientific institutions such as herbaria or natural history museums, it was important to acquire and publish about new specimens collected in colonized areas (Endersby, 2016, pp.105-128). Next to orchid hunters, colonial botanical gardens such as the one in Buitenzorg (now Bogor, Indonesia) or gardens in British India played key role in establishing global networks of plant exchange (Baber, 2016; Drayton, 2000; Goss, 2011; Weber and Wille, 2018). 3 Related Work 3.1 Colonial Roots of Natural History Collections Although a contentious issue, an important body of literature exists exploring various topics which relate to the broader issue of addressing the colonial roots of natural history collections (Coote et al., 2017; Curry et al., 2018). There is consensus surrounding the finding that it is not uncommon for natural history museums to present their collections in a way that blurs “its entanglement with the human world” (Drieënhuizen and Sysling, 2021, p.305). For instance, scholars have found that Eurocentric perspectives 106 Figure 2: Indonesia and Papua New Guinea Interface, including a pie chart, filters, a donut chart and a timeline. Figure 3: General Interface with the Orchidaceae family selected, shifting the visualizations to be filtered by this family. 107 taceae (clove) and Meliaceae (cotton fruits, commonly consumed in South East Asian regions) were collected during colonial times. According to our data, the Naturalis collection entails ca. 20.000 Moracea specimens from Indonesia with a collection peak in 1961. Legumes, peas and other edible plants were also accumulated, particularly between the nineteenth and twentieth centuries, matching the timeline showcased on the interface (Wasino et al., 2017, pp. 265-267). Observing the bottom of Figure 4, Meliaceae, most commonly known as neem, was widely traded and forms an important part of the museum’s botanical collection. Its consumption took many forms, particularly for "pressed seed oil, mainly used as insecticide, but also for cosmetic, medicinal and agricultural uses" (Sujarwo et al., 2016, p.186). Data with respect to other species, such as Labiatae (mint), and Compositae (daisies) also show interesting spikes which would require further contextualization. Taken together, the in-depth examination of data related to specific plant families illustrate how the collection’s composition was also shaped by economic motives geared towards extracting agricultural resources from areas which fell under Dutch colonial rule. 6 Discussion and Conclusion This project thus contributes to filling a gap in historical research regarding the colonial legacy of Dutch botanical collections, highlighting the intertwined nature of scientific endeavors and history. Furthermore, it addresses a gap in visualization-oriented digital humanities research by providing a concrete example of how easy-to-use computational techniques can help in understanding the cultural histories of scientific collections in an interactive way, and can thereby help decolonization of natural history collections. Specifically, we have demonstrated how simple and interactive visualizations can help researchers to develop a better understanding of a collection’s colonial bias and how this bias has evolved historically (Nelson and Ellis, 2019). Our paper presents an opportunity for natural history museums to make their data more accessible through a user-friendly and visual tool that is easy to deploy and does not consume much computational resources. Other Dutch museums such as the Wereldmuseum have addressed colonial Dutch history by decolonizing their collections and exhibits. 5 Their work has shown how addressing colonial history can lead to making collections and exhibitions more inclusive and reflective of diverse cultural histories. Enabling museum visitors and users of scientific collection-related datasets to study colonial provenance acknowledges the colonial biases of natural history collections. Tools such as Power BI play a crucial role in reconnecting datasets to their cultural and historical roots by enabling dynamic visualizations that reveal patterns, biases, and historical trajectories otherwise hidden in raw data. By integrating spatial, temporal, and categorical filters, these tools allow scientists to trace the provenance of specimens, historians to contextualize collection practices within colonial frameworks, and a broader multidisciplinary audience to engage with data through an intuitive interface. The ability to filter by collector, location, time period, inter alia, highlights how scientific knowledge production was shaped by historical power dynamics, making the tool particularly useful, i.e. for museum professionals and researchers investigating provenance. At the same time, its accessibility, which does not require specialized coding skills, expands its reach to the general public, fostering greater engagement with the ethical dimensions of natural history collections. By bridging these disciplinary and public divides, Power BI 5See also here: ../why-exhibition-about-our-colonial-inheritance, accessed 10 October 2024. 114 serves as a vital instrument for uncovering, analyzing, and sharing the layered histories embedded within scientific datasets. Thus, our interactive tool links the Netherlands’ colonial past with Naturalis’ scientific plant collection through multi-layered visualizations. Additionally, it helps researchers to uncover important stories of how the collections have been brought together, contributing to a more nuanced interpretation of their collections’ evolution over a longer period. The initial colonial context behind botanical acquisitions set the trajectory for future Dutch botanical endeavors in the various countries, even after they gained independence. Our dashboards - and our focus on the two former Dutch colonies - allows users to employ a historical lens to reveal dynamic relationships between natural history collection patterns and broader socio-political contexts. Since the interactive visualizations allow natural history collections to be studied in terms of their movement (via the map tool) over time (via the timeline) and allow for detailed analysis (via the metadata filters), it can be used to investigate a range of questions related to the cultural histories of scientific collections. The flexibility it allows the user, its ability to make data more accessible, its potential use in flagging missing data and the clear ability to reveal information that would otherwise be hidden shows that this interactive visualization can best contribute to understanding the cultural histories of scientific collections in a way that static visualizations otherwise would not. As depicted, the visualization tool enhances the collections’ accessibility through visual reading, with the potential to reveal insights otherwise hidden and uncover data-related concerns. The tool thus aids biologists and those interested in tracing the provenance, geographical spread of natural objects in collections and the ethics of obtaining these specimens in the context of colonial power asymmetries. However, certain limitations remain. Firstly, this study focuses solely on the botany subsection of Naturalis, leaving the broader applicability of our approach to the entire natural history collection untested, particularly in terms of merging multiple datasets and assessing Power BI’s scalability. Due to the scope of this research, we were unable to evaluate the generalization of our experiment by incorporating additional large datasets, an avenue we intend to explore in future work. Expanding this approach to include datasets from different institutions would provide a more thorough assessment of Power BI’s ability to handle larger and more diverse data sources. While the platform is well-suited for structured data visualization, its performance may be challenged by increased data volume and variability, potentially necessitating extensive preprocessing or alternative tools for optimal scalability. Future research should examine how integrating datasets from various natural history museums, each with distinct formats and metadata standards, influences the visualization process while ensuring that insights into colonial biases remain clear and accessible. Secondly, since the Naturalis botany dataset originates from the GBIF platform, adjustments may be required to accommodate datasets from other institutions, as formatting conventions are unlikely to align seamlessly. Lastly, leveraging Power BI for such visualizations may demand significant data preprocessing, which could present a barrier for users with limited data management skills. Nonetheless, this visualization tool provides a starting point for tracing the histories of any museum collection, provided they have the data. The recordedBy metadata, for instance - through which explorations and records made by botanists can be tracked and mapped - is an interesting starting point to understand the specific acquisition histories of certain specimens. For example, Carl Ludwig Blume - whose records show that 95,77% of the specimens he recorded belong to Indonesia (see Figure 5) 115 Figure 5: General Interface with the Recorded By filter set to Blume CL, showing that the majority of specimens he recorded came from Indonesia. - is a relevant finding, indicating the Dutch botanical connection to their colonies. Blume was the founder of ‘s Rijksherbarium (National Herbarium), an institutional predecessor of what is now museum Naturalis. The sheer amount of orchids taken from Indonesia by a single researcher is also an interesting finding in itself, opening a Pandora’s box of other insightful questions. Thus, the scope of such a tool can undoubtedly be extended to address other questions of contemporary relevance. While in the nineteenth century, orchids were primarily seen as scientific study objects and luxury products from the East and West Indies, they are now a globally traded horticultural commodity whose presence in almost every European household has severe environmental consequences (Soode et al., 2015). 7 Data Access The Power BI file used for this paper can be downloaded through Zenodo: https: //doi.org/10.5281/zenodo.13958314. References Godwin Ubong Akpan, Isah Mohammed Bello, Kebba Touray, Reuben Ngofa, Daniel Rasheed Oyaole, Sylvester Maleghemi, Marie Babona, Chanda Chikwanda, Alain Poy, Franck Mboussou, et al. Leveraging polio geographic information system platforms in the african region for mitigating covid-19 contact tracing and surveillance challenges. JMIR mHealth and uHealth, 10(3):e22544, 2022. Mireia Alcantara-Rodriguez, Tinde Van Andel, and Mariana Françozo. Nature portrayed in images in Dutch Brazil: Tracing the sources of the plant woodcuts in the Historia Naturalis Brasiliae (1648). bioRxiv, pages 2022–10, 2022. 116 Tinde van Andel. Open the treasure room and decolonize the museum. Universiteit Leiden, 2017. Jack Ashby and Rebecca Machin. Legacies of colonial violence in natural history collections. Journal of Natural Science Collections, 8:44–54, 2021. Zaheer Baber. The plants of empire: Botanic gardens, colonial power and botanical knowledge. Journal of Contemporary Asia, 46(4):659–679, 2016. Vijay Barve and Javier Otegui. bdvis: visualizing biodiversity data in r. Bioinformatics, 32(19):3049–3050, 2016. Tiziana N Beltrame, Elena Canadelli, and Luca Tonetti. The natures of digital practices: People, objects, and data mobilities in natural history collections. Nuncius, 39(3): 741–758, 2024. Elizabeth H Boakes, Philip JK McGowan, Richard A Fuller, Ding Chang-qing, Natalie E Clark, Kim O’Connor, and Georgina M Mace. Distorted views of biodiversity: spatial and temporal bias in species occurrence data. PLoS biology, 8(6):e1000385, 2010. Peter Boomgaard. The making and unmaking of tropical science: Dutch research on indonesia, 1600-2000. Bijdragen tot de taal-, land-en volkenkunde/Journal of the Humanities and Social Sciences of Southeast Asia, 162(2):191–217, 2008. Marieke Borren and Caroline Drieënhuizen. Dossier: The coloniality of natural history collections. Locus: Tijdschrift voor Cultuurwetenschappen, 2022. Lissa Breugelmans and Maarten Trekels. Implementation experience report for the developing latimer core standard: The dissco flanders use-case. Biodiversity Information Science and Standards, 7:e113766, 2023. Falko T Buschke, Claudia Capitani, El Hadji Sow, Yvonne Khaemba, Beth A Kaplin, Andrew Skowno, David Chiawo, Tim Hirsch, Elizabeth R Ellwood, Hayley Clements, et al. Make global biodiversity information useful to national decision-makers. Nature Ecology & Evolution, 7(12):1953–1956, 2023. Anne Coote, Alison Haynes, Jude Philp, and Simon Ville. When commerce, science, and leisure collaborated: the nineteenth-century global trade boom in natural history collections. Journal of Global History, 12(3):319–339, 2017. Jeroen Creuwels. Naturalis Biodiversity Center (nl) - Botany Leiden. Occurrence dataset. Naturalis Biodiversity Center, 2015. Helen Anne Curry, Nicholas Jardine, James Andrew Secord, and Emma C Spary. Worlds of natural history. Cambridge University Press, 2018. Subhadra Das and Miranda Lowe. Nature read in black and white: decolonial approaches to interpreting natural history collections. Journal of Natural Science Collections, 6(4):4–14, 2018. Verónica Díez Díaz, Sara Akhlaq, Katja Kaiser, Ina Heumann, and Daniela Schwarz. Digitization as a research methodology in colonial natural history collections. Nature Reviews Biodiversity, pages 1–2, 2025. 117 DiSSCo. Nationaal natuurhistorisch collectieoverzicht. https://app.powerbi.com/view?r=eyJrIjoiNDQ1YmQzMzMtNmY5YS00MDQzLWI5M2Y tNmRhOTM2MTg2NTU0IiwidCI6IjhjZDI0OTg0LTBhYTMtNGZjNS1iMD liLTRkNmVjZmFhNThmYiIsImMiOjl9, accessed 19 October 2024, 2023. Richard Drayton. Nature’s government: Science, imperial britain, and the ‘improvement’of the world, 2000. Joshua A Drew, Corrie S Moreau, and Melanie LJ Stiassny. Digitization of museum collections holds the potential to enhance researcher diversity. Nature ecology & evolution, 1(12):1789–1790, 2017. Caroline Drieënhuizen and Fenneke Sysling. Java man and the politics of natural history: An object biography. Bijdragen tot de taal-, land-en volkenkunde/Journal of the Humanities and Social Sciences of Southeast Asia, 177(2-3):290–311, 2021. Felix Driver, Mark Nesbitt, and Caroline Cornish. Mobile museums: Collections in circulation. UCL Press, 2021. Déborah Dubald and Catarina Madruga. Introduction: Situated nature: Field collecting and local knowledge in the nineteenth century. Journal for the History of Knowledge, 3 (1):1–11, 2022. Jim Endersby. Orchid: A cultural history. University of Chicago Press, 2016. Xiao Feng, Brian J Enquist, Daniel S Park, Brad Boyle, David D Breshears, Rachael V Gallagher, Aaron Lien, Erica A Newman, Joseph R Burger, Brian S Maitner, et al. A review of the heterogeneous landscape of biodiversity databases: Opportunities and challenges for a synthesized biodiversity knowledge base. Global Ecology and Biogeography, 31(7):1242–1260, 2022. M.R. Fernando. Coffee Cultivation in Java, 1830–1917, pages 157–172. 06 2003. ISBN 9780521818513. doi: 10.1017/CBO9780511512193.017. Nico M Franz and Beckett W Sterner. To increase trust, change the social design behind aggregated biodiversity data. Database, 2018:bax100, 2018. Emilio García-Roselló, Jacinto González-Dacosta, and Jorge M Lobo. The biased distribution of existing information on biodiversity hinders its use in conservation, and we need an integrative approach to act urgently. Biological Conservation, 283: 110118, 2023. Andrew Goss. Decent colonialism? Pure science and colonial ideology in the Netherlands East Indies, 1910–1929. Journal of Southeast Asian Studies, 40(1):187–214, 2009. Andrew Goss. The floracrats: State-sponsored science and the failure of the enlightenment in Indonesia. Univ of Wisconsin Press, 2011. Andrew Goss. Reinventing the kebun raya in the new republic: Scientific research at the Bogor Botanical Gardens in the age of decolonization. Studium, 11(3), 2018. Andrew Goss. Decolonizing botany: Indonesia, unesco, and the making of a global science. Journal of the History of Biology, 56(3):495–523, 2023. 118 Carl Haarnack. Met andere ogen: Over de slavernijtentoonstelling in het Rijksmuseum. Jaarboek De Achttiende Eeuw, 53(1):193–199, 2021. Brandon P Hedrick, J Mason Heberling, Emily K Meineke, Kathryn G Turner, Christopher J Grassa, Daniel S Park, Jonathan Kennedy, Julia A Clarke, Joseph A Cook, David C Blackburn, et al. Digitization and the future of natural history collections. BioScience, 70(3):243–251, 2020. P. Hiepko. The collections of the Botanical Museum Berlin-Dahlem (b) and their history. Englera, 7:219–252, 1987. Roos Hopman. Snails, time, data: On the politics of mass-digitization and the possibility of data drift. Big Data & Society, 11(3):20539517241267760, 2024. Alice C Hughes, James B Dorey, Silas Bossert, Huijie Qiao, and Michael C Orr. Big data, big problems? how to circumvent problems in biodiversity mapping and ensure meaningful results. Ecography, page e07115, 2024. Dominik Hünniger. Visible labour? Productive forces and imaginaries of participation in european insect studies, ca. 1680–1810. Berichte zur Wissenschaftsgeschichte, 44(2): 180–210, 2021. Sharif Islam, Andreas Weber, and Erzsébet Tóth-Czifra. From green deal to cultural heritage: Fair digital objects and european common data spaces. Research Ideas and Outcomes, 8:e93815, 2022. Hans J Jacobs and André Koch. On the discovery and scientific description of the Emerald Tree Monitor, Varanus prasinus (Schlegel, 1839). Bibliotheca herpetologica, 15(7):61–76, 2021. Hans J Jacobs and Glenn M Shea. Eight weeks in lobo bay. the natuurkundige commissie on new guinea in 1828. i. scincus and centroplites (scincidae). Bibliotheca herpetologica, 16(6):48–81, 2022. Kirk R Johnson, Ian FP Owens, and Global Collection Group. A global approach for natural history museum collections. Science, 379(6638):1192–1194, 2023. Katja Kaiser, Ina Heumann, T Nadim, H Keysar, M Petersen, M Korun, and F Berger. Promises of mass digitisation and the colonial realities of natural history collections. Journal of Natural Science Collections, 11:13–25, 2023. Kelsey Leonard. Decolonizing botanical gardens. Qualitative Research Journal, 24(5): 536–554, 2024. Carla Maldonado, Carlos I Molina, Alexander Zizka, Claes Persson, Charlotte M Taylor, Joaquina Albán, Eder Chilquillo, Nina Rønsted, and Alexandre Antonelli. Estimating species diversity and distribution in the era of b ig d ata: to what extent can we trust public databases? Global Ecology and Biogeography, 24(8):973–984, 2015. Maarten Manse. Kennis is macht: de veelzijdige expedities van botanicus Pieter Willem Korthals (1807–1892). Studium, 6(1), 2013. Asla Medeiros e Sá, Franklin Alves de Oliveira, Bruno Schneider, Karina Rodriguez Echavarria, and Cristiana Silveira Serejo. Visually overviewing biodiversity open data digital collections. In Proceedings of the Symposium on Open Data and Knowledge for a Post-Pandemic Era ODAK22, UK, pages 1–8. BCS Learning & Development, 2022. 119 Yamikani Mgusha, Deliwe Bernadette Nkhoma, Msandeni Chiume, Beatrice Gundo, Rodwell Gundo, Farah Shair, Tim Hull-Bailey, Monica Lakhanpaul, Fabianna Lorencatto, Michelle Heys, et al. Admissions to a low-resource neonatal unit in malawi using a mobile app and dashboard: a 1-year digital perinatal outcome audit. Frontiers in Digital Health, 3:761128, 2021. Eulàlia Gassó Miracle. Coenraad Jacob Temminck and the emergence of systematics (1800– 1850). Brill, 2021. Ryan S Mohammed, Grace Turner, Kelly Fowler, Michael Pateman, Maria A NievesColón, Lanya Fanovich, Siobhan B Cooke, Liliana M Dávalos, Scott M Fitzpatrick, Christina M Giovas, et al. Colonial legacies influence biodiversity lessons: how past trade routes and power dynamics shape present-day scientific research and professional opportunities for caribbean scientists. The American Naturalist, 200(1): 140–155, 2022. Johannes Müller. Naming the world: Pieter Bleeker’s travels and the challenges of archipelagic biodiversity. Animals in Dutch travel writing, ed. by R.A.M. Honings and E. Beek, pages 79–98, 2023. Tahani Nadim, Mareike Vennen, Ina Heumann, and Filippo Bertoni. Logistical natures: Trade, traffics, and transformations in natural history collecting. Historical Studies in the Natural Sciences, 54(2):125–134, 2024. Gil Nelson and Shari Ellis. The history and impact of digitization and digital data mobilization on biodiversity research. Philosophical Transactions of the Royal Society B, 374(1763):20170391, 2019. Daniel S Park, Xiao Feng, Shinobu Akiyama, Marlina Ardiyani, Neida Avendaño, Zoltan Barina, Blandine Bärtschi, Manuel Belgrano, Julio Betancur, Roxali Bijmoer, et al. The colonial legacy of herbaria. Nature Human Behaviour, 7(7):1059–1068, 2023. Victoria Pickering. Mobilising historical botanical data as research. Nuncius, 39(3): 759–774, 2024. Aníbal Quijano. Coloniality and modernity/rationality. Cultural studies, 21(2-3): 168–178, 2007. Willem Jan van der Schoor. Zuivere en toegepaste wetenschap in de tropen: biologisch onderzoek aan particuliere proefstations in Nederlands-Indië 1870-1940. Ph. D. Thesis, 2012. Eveli Soode, Paul Lampert, Gabriele Weber-Blaschke, and Klaus Richter. Carbon footprints of the horticultural products strawberries, asparagus, roses and orchids in germany. Journal of Cleaner Production, 87:168–179, 2015. Lise Stork, Andreas Weber, Eulàlia Gassó Miracle, Fons Verbeek, Aske Plaat, Jaap van den Herik, and Katherine Wolstencroft. Semantic annotation of natural history collections. Journal of Web Semantics, 59:100462, 2019. Lise Stork, Andreas Weber, Jaap van den Herik, Aske Plaat, Fons Verbeek, and Katherine Wolstencroft. Large-scale zero-shot learning in the wild: Classifying zoological illustrations. Ecological informatics, 62:101222, 2021. 120 Wawan Sujarwo, Ary P Keim, Giulia Caneva, Chiara Toniolo, and Marcello Nicoletti. Ethnobotanical uses of neem (azadirachta indica a. juss.; meliaceae) leaves in Bali (Indonesia) and the Indian subcontinent in relation with historical background and phytochemical properties. Journal of Ethnopharmacology, 189:186–193, 2016. F. Sysling, H. Vellinga, S. Kompier, and T. Niederer. Who did all the work. https://whodidallthework.nl, accessed 19 October 2024, 2024. Marieke Sophia van de Loosdrecht, Nicholaas M Pinas, Jerry R Tjoe Awie, Frank FM Becker, Harro Maat, Robin van Velzen, Tinde van Andel, and M Eric Schranz. Maroon rice genomic diversity reflects 350 years of colonial history. bioRxiv, pages 2024–05, 2024. Lorella Viola. Networks of migrants’ narratives: a post-authentic approach to heritage visualisation. ACM Journal on Computing and Cultural Heritage, 16(1):1–21, 2023. M Wasino, Hartatik Endah Sri, et al. Peasant economy under the plantation capitalism (the study of peasant agricultural plantation in java at the nineteenth to the beginning of twentieth century). Man in India International Journal of Anthropology. Vol. 97 No. 5, 2017., 2017. Andreas Weber. A garden as a niche: Botany and imperial politics in the early nineteenth century Dutch empire. Studium, 11(3):178–190, 2018. Andreas Weber and Esther Turnhout. A langur from sumatra: Digital futures, material presents and colonial pasts. Nuncius, 39(3):775–788, 2024. Andreas Weber and Robert-Jan Wille. Laborious transformations: Plants and politics at the bogor botanical gardens. Studium, 11(3):169–177, 2018. Robert-Jan Wille. Mannen van de microscoop: De laboratoriumbiologie op veldtocht in Nederland en Indië, 1840-1910. Vantilt, 2019. Florian Windhager, Paolo Federico, Günther Schreder, Katrin Glinka, Marian Dörk, Silvia Miksch, and Eva Mayr. Visualization of cultural heritage collection data: State of the art and future challenges. IEEE transactions on visualization and computer graphics, 25(6):2311–2330, 2018. Pieter van Wingerden. Science on the edge of empire: E.A. Forsten (1811–1843) and the Natural History Committee (1820–1850) in the Netherlands Indies. Centaurus, 62(4):797–821, 2020. Pieter van Wingerden. The Natuurkundige Commissie in the Netherlands Indies (1820-1850). Brill, 2025. 121 cb 2025. This work is licensed under a Creative Commons “Attribution 4.0 International” license.cb 2025. This work is licensed under a Creative Commons “Attribution 4.0 International” license. Small-Scale Testing on Generative AI and Post-OCR Correction in Historical Datasets Florentina Armaselu1 1Luxembourg Centre for Contemporary and Digital History (C2DH), University of Luxembourg 1. Introduction Recent developments in large language models (LLMs) and generative AI (GenAI) chatbots such as Chat-GPT, Google Bard (now Gemini), and YouChat (Brown et al., 2020; Chaka, 2023; Manyika and Hsiao, 2023) have fostered new types of interaction that can lower the barrier in human-machine communication through conversation in natural languages. We assume that such chatbots may be able to act as conversational assistants in tasks that otherwise require more complex processing to improve the results produced by simpler or earlier, less accurate techniques. This article proposes a set of small-scale tests with GenAI chatbots on post-OCR correction in historical datasets. It illustrates, through examples of responses obtained from GenAI agents integrated into post-OCR correction and assessment tasks, what types of challenges should be addressed in this context when working with historical datasets. Previous studies have shown that OCR errors in input data can have a non-negligible impact on downstream language processing, such as sentence segmentation, named entity recognition (NER), topic modelling, word embedding, sentiment analysis and information retrieval (Nguyen et al., 2022; Strien et al., 2020). Therefore, various methods have been envisaged to tackle this problem, including post-OCR detection and correction of errors. Recent surveys classified these methods into manual and (semi-)automatic , isolated-word and context-dependent types, and emphasised the trend in the development of neural networkand context-based approaches (Nguyen et al., 2022). Other enquires focused on the use of machine learning techniques to automatically estimate text quality and select candidates for OCR rerun within cultural institutions that deal with lower quality historical data (Schneider and Maurer, 2022). On the other hand, studies on post-OCR error detection and correction investigated the use of large language models in this type of tasks. For example, pre-trained language models from the GPT-2 family were tested on datasets of English books and texts from post-OCR competitions by combining multiple OCR versions of an object and choosing the best-scored option, with the goal of reducing the number of errors (Gupta et al., 2021). Tests with other models, such as Llama 2, and a corpus of 19th century British newspapers containing aligned excerpts of machine and manual transcriptions, were also performed to assess the ability of prompt-based approaches to detect and correct OCR errors (Thomas et al., 2024). The performance of generative LLMs in post-OCR 123 Figure 2: Generic pipeline for small-scale testing of GenAI post-OCR correction. illustrates a generic pipeline that considers both the preservation and modernisation branches, taking into account the various types of OCR errors due to misinterpretation of letters, specific or older fonts, print quality, etc., and the categories of differences between the original and corrected text resulting from the particularities of the generative AI technology. One of the typical emergent abilities of large language models, which are not present in smaller pre-trained models, is in-context learning. Compared to fine-tuning that implies adaptation of a model to a specific task by training it on a tailored dataset and updating its weights, or to retrieval-augmented generation (RAG) that combines information retrieval and domain-relevant knowledge bases, in-context learning involves the use of zero-, oneor few-shot learning including in the prompts task descriptions with or without examples, and no update of the model parameters (Brown et al., 2020; Li, 2023; Zhao et al., 2024). While various fine-tuning and RAG techniques are available (Parthasarathy et al., 2024), in-context learning may represent an approach that is potentially interesting to explore for tasks such as LLM-based post-OCR correction. Recent multilingual benchmarks have shown both promise and challenges in applying LLMs in zeroand few-shot settings to post-OCR correction of historical datasets (Boros et al., 2024; Kanerva et al., 2025), some of them similar to observations presented in this paper (e.g., tendency of the models to overcorrect and paraphrase). The decision on which of these categories of adaptation techniques to choose also depends on the complexity of the project and the budget allocated to it. Although the utility of large-scale benchmarking is undeniable, small-scale testing, including data preparation, processing, and evaluation, and iterative prompt tuning, can prove its usefulness as a faster and cheaper way of identifying categories of problems from the outset, suggesting directions for adaptation, further investigation, or assessment aligned with existing standards, that may inform the design of subsequent larger-scale phases or applications. 4. Conclusion and Future Work The article proposes a small-scale investigation on the use of GenAI agents in post-OCR correction workflows for historical datasets. Although preliminary results show a certain potential for this type of technology in solving tasks in this category, more tests are necessary to assess the capacity to respond to specially conceived prompts for historical text processing. In particular, it was shown that the agents have a tendency to replace historical forms with more modern ones, to reformulate whole phrases, 130 or to change word order and punctuation. Specific instructions should be devised to prevent these forms of modification. However, building modernised layers for documents with old spelling may be considered a possibly interesting application in tasks such as the transformation of older texts to be read by modern users or computer programs. Another potential challenge resides in a certain degree of instability observed in computing character and word error rates (CER, WER). This indicates that comparing the GenAI values with independently computed results using programming languages such as Python or R should be envisaged. The study was limited to interactions with three agents that involved the use of online platforms, while the integration of this type of technology into larger-scale pipelines would require more code-oriented or batch processing solutions that need to be further examined. Future work can include additional testing with other GenAI agents, open-source large language models in a programming environment, datasets in other languages, and further enquiry about the application of these technologies in post-OCR correction for historical and digital humanities research. Although limited in scope, it may be argued that small-scale testing of the kind presented in this paper can serve as a pilot phase in the pipeline, that allows for early identification of problems, synthesis of observations, assessment, and articulation of hypotheses for further testing or potential solutions, which can inform the design of subsequent larger-scale phases in the workflow or applications. References N. Abadie, E. Carlinet, J. Chazalon, and B. Duménieu. A Benchmark of Named Entity Recognition Approaches in Historical Documents Application to 19th Century French Directories, volume 13237 of Lecture Notes in Computer Science, page 445–460. Springer International Publishing, Cham, 2022. ISBN 978-3-031-06554-5. doi: 10.1007/978-3031-06555-2_30. URL https://link.springer.com/10.1007/978-3-031-065552_30. Florentina Armaselu, Barbara McGillivray, Chaya Liebeskind, Paola Marongiu, Giedr ˙ e Val ¯ unait ˙ e Oleškevičien ˙ e, Elena-Simona Apostol, and Ciprian-Octavian Truică. Multilingual word embedding and linguistic linked open data for tracing semantic change. Rasprave Instituta za hrvatski jezik i jezikoslovlje, 50(2):219–257, 2024. ISSN 18490379, 13316745. doi: 10.31724/rihjj.50.2.1. Kyunga Bang. Exploring Generative Large Language Models for Post-OCR Enhancement of Historical Texts. Master’s thesis, Graduate School of Sungkyunkwan University, April 2024. URL https://www.researchgate.net/publication/ 382113790_Exploring_Generative_Large_Language_Models_for_PostOCR_Enhancement_of_Historical_Texts. Emanuela Boros, Maud Ehrmann, Matteo Romanello, Sven Najem-Meyer, and Frédéric Kaplan. Post-correction of historical text transcripts with large language models: An exploratory study. In Proceedings of LaTeCH-CLfL 2024, page 133–159. Association for Computational Linguistics, March 2024. URL https://aclanthology.org/2024. latechclfl-1.14/. Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, 131 Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. (arXiv:2005.14165), July 2020. URL http://arxiv.org/abs/2005.14165. arXiv:2005.14165 [cs]. Chaka Chaka. Generative AI chatbots - ChatGPT versus YouChat versus Chatsonic: Use cases of selected areas of applied English language studies. International Journal of Learning, Teaching and Educational Research, 22(6):1–19, June 2023. ISSN 16942493, 16942116. doi: 10.26803/ijlter.22.6.1. Guillaume Chiron, Antoine Doucet, Mickael Coustaty, and Jean-Philippe Moreux. ICDAR2017 competition on post-OCR text correction. In 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), page 1423–1428, Kyoto, November 2017. IEEE. ISBN 978-1-5386-3586-5. doi: 10.1109/ICDAR.2017.232. URL http://ieeexplore.ieee.org/document/8270163/. Krissy Davis. The best AI chatbots: ChatGPT and other alternatives, May 2023. URL https://www.wearedevelopers.com/en/magazine/236/best-aichatbots-chatgpt-and-other-alternatives. Harsh Gupta, Luciano Del Corro, Samuel Broscheit, Johannes Hoffart, and Eliot Brenner. Unsupervised multi-view post-OCR error correction with language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, page 8647–8652, Online and Punta Cana, Dominican Republic, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.680. URL https://aclanthology.org/2021.emnlp-main.680. Jenna Kanerva, Cassandra Ledins, Siiri Käpyaho, and Filip Ginter. OCR error postcorrection with LLMs in historical documents: No free lunches. (arXiv:2502.01205), February 2025. doi: 10.48550/arXiv.2502.01205. URL http://arxiv.org/abs/2502. 01205. arXiv:2502.01205 [cs]. Romain Karpinski, Devashish Lohani, and Abdel Belaid. Metrics for complete evaluation of OCR performance. In The 22nd Int’l Conf on Image Processing, Computer Vision, & Pattern Recognition, Las Vegas, United States, July 2018. URL https://inria.hal.science/hal-01981731v1. Yinheng Li. A practical survey on zero-shot prompt design for in-context learning. In Proceedings of the Conference Recent Advances in Natural Language Processing - Large Language Models for Natural Language Processings, page 641–647, 2023. doi: 10.26615/978-954-452-092-2_069. URL http://arxiv.org/abs/2309.13205 . arXiv:2309.13205 [cs]. James Manyika and Sissie Hsiao. An overview of Bard: an early experiment with generative AI. 2023. URL https://ai.google/static/documents/google-about-bard.pdf. Andrew Cameron Morris, Viktoria Maier, and Phil Green. From WER and RIL to MER and WIL: improved evaluation measures for connected speech recognition. In Interspeech 2004, page 2765–2768. ISCA, October 2004. doi: 10.21437/Interspeech. 2004-668. URL https://www.isca-archive.org/interspeech_2004/morris04_ interspeech.html. 132 Thi Tuyet Hai Nguyen, Adam Jatowt, Mickael Coustaty, and Antoine Doucet. Survey of post-OCR processing approaches. ACM Computing Surveys, 54(6):1–37, July 2022. ISSN 0360-0300, 1557-7341. doi: 10.1145/3453476. Venkatesh Balavadhani Parthasarathy, Ahtsham Zafar, Aafaq Khan, and Arsalan Shahid. The ultimate guide to fine-tuning LLMs from basics to breakthroughs: An exhaustive review of technologies, research, best practices, applied research challenges and opportunities. (arXiv:2408.13296), October 2024. doi: 10.48550/ arXiv.2408.13296. URL http://arxiv.org/abs/2408.13296 . arXiv:2408.13296 [cs]. Christophe Rigaud, Antoine Doucet, Mickael Coustaty, and Jean-Philippe Moreux. ICDAR 2019 competition on post-OCR text correction. In 2019 International Conference on Document Analysis and Recognition (ICDAR), page 1588–1593, Sydney, Australia, September 2019. IEEE. ISBN 978-1-7281-3014-9. doi: 10.1109/ICDAR.2019.00255. URL https://ieeexplore.ieee.org/document/8978127/. F. Rosset. L’art de conduire et regler les pendules et les montres. chez la Veuve de J. B. Kleber, Imprimeur de Sa Majesté, Luxembourg, 1789. URL https://viewer. eluxemburgensia.lu/ark:70795/dqgfr3/pages/17/articles/DTL612. Pit Schneider and Yves Maurer. Rerunning OCR: A machine learning approach to quality assessment and enhancement prediction. Journal of Data Mining & Digital Humanities, November 2022. ISSN 2416-5999. doi: 10.46298/jdmdh.8561. URL http://arxiv.org/abs/2110.01661. arXiv:2110.01661 [cs]. Daniel van Strien, Kaspar Beelen, Mariona Ardanuy, Kasra Hosseini, Barbara McGillivray, and Giovanni Colavizza. Assessing the impact of OCR quality on downstream NLP tasks. In Proceedings of the 12th International Conference on Agents and Artificial Intelligence, page 484–496, Valletta, Malta, 2020. SCITEPRESS - Science and Technology Publications. ISBN 978-989-758-395-7. doi: 10.5220/ 0009169004840496. URL http://www.scitepress.org/DigitalLibrary/Link. aspx?doi=10.5220/0009169004840496. Alan Thomas, Robert Gaizauskas, and Haiping Lu. Leveraging LLMs for postOCR correction of historical newspapers. In LT4HALA 2024@LREC-COLING2024, page 116–121. ELRA Language Resources Association, 2024. URL https: //aclanthology.org/2024.lt4hala-1.14.pdf. Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. A survey of large language models. (arXiv:2303.18223), October 2024. doi: 10.48550/arXiv.2303.18223. URL http://arxiv.org/abs/2303.18223. arXiv:2303.18223 [cs]. 133 A. Appendix Table 4: Prompt and response excerpts, reminding the language (Google Bard, 18.12.2023). User GenAI agent (Bard) Please try to identify the words that may be due to OCR errors in the following list and assign a confidence level in percentage to each proposed correction. The words are in French and the corrections should also be in French. ... Sure, here are the words in the list that may be due to OCR errors, along with proposed corrections and confidence levels: ... confération confédération 100% ... Table 5: Prompt and response excerpts, “supposant” (YouChat, 16.12.2023). User GenAI agent (YouChat) Would your decision be different if you consider that usually elongated ’s’ in older texts is producing OCR errors, such as considering ’f’ instead of ’s’? Considering the possibility of an OCR error substituting "f" for elongated "s" is indeed a relevant factor to consider. Given this information, the potential correct form for the word "fuppofant" might indeed be "supposant" if we consider the substitution of "s" with "f" due to OCR errors in older texts. Taking this into account, the corrected potential form for "fuppofant" could be "supposant" instead of "fantastic." ... Table 6: Prompt and response excerpts, "roue de chaussée" (ChatGPT-4, 15.12.2023). User GenAI agent (ChatGPT-4) Can you please reanalyze the word "chaufée" given the context "roue de chaufée" and the fact that the text refers to clock making in the 18th century? In the context of clockmaking during the 18th century, the term "chaufée" in "roue de chaufée" could be a historical or technical term related to the mechanism of a clock. Considering the OCR misinterpretation of long ’s’ as ’f’, and given the clockmaking context, the term could be "roue de chaussée", with "chaussée" potentially referring to a part or a process in the clock’s mechanism. However, without specific knowledge of 18th-century clockmaking terminology, it’s difficult to ascertain the exact term. "Roue de chauffée" could refer to a heated part or a wheel that is treated in a particular way, but this seems less likely in the context of mechanical clockwork where terms like "chaussée" do not clearly fit. Considering this, a more likely correction would be "roue de chaussée" if "chaussée" is a term used in the context of clock mechanics at the time. However, if "chauffée" is not due to an OCR error and is actually a correct transcription of the original text, it might require specialized historical knowledge to interpret accurately. Based on common French usage and the potential for an OCR error, "roue de chaussée" could be a plausible correction with a medium to high confidence level, but without specific expertise in historical clockmaking terminology, I cannot provide a definitive correction. 134 Table 7: Prompt and response excerpts, "afrotement", "engrenne" (ChatGPT-4, 15.12.2023) User GenAI agent (ChatGPT-4) Do you have any information about the word "afrotement ", in the text "tige entre afrotement dur", and its possible usage in French in the 18th century? Same question for "engrenne ", in "ce pignon engrenne". It is possible that these are historical usages that should be kept in the transcription of the text. The goal is to correct just the errors that may be due to OCR misinterpretation. In the 18th century, French language and spelling were less standardized than they are today, and many words could have various spellings or be spelled according to phonetics. With this in mind: •Afrotement: The correct modern French word is "frottement," which means "friction." It is plausible that "afrotement" is a historical spelling or a typographical error in the original text. If "afrotement" is not an OCR misinterpretation, it could be an archaic spelling of "frottement" that was used at the time. This would be especially true in the context of clock mechanisms where friction is an important concept. •Engrenne: The modern French word "engrène" comes from the verb "engrener," which means "to mesh" or "to engage," as gears do in a clock. "Engrenne" could be an old conjugation or spelling of the verb. If "engrenne" is correctly transcribed, it may not be an OCR error but rather a historical usage that should be preserved in the text. ... 135 cb 2025. This work is licensed under a Creative Commons “Attribution 4.0 International” license. Uncovering Interwar Comics: The Challenge of Labeling Graphics in Belgian Weekly Magazines Bas Vercruysse1, Julie M. Birkholz1,2, Benoît Crucifix2,3, Erwin Dejasse2,4, Sébastien Hermans2, and Krishna Kumar Thirukokaranam Chandrasekar5 1Ghent Centre for Digital Humanities, Ghent University 2KBR - Royal Library of Belgium 3Cultural Studies Research Group, KU Leuven, 4Centre de recherches en Etudes littéraires, philologiques et textuelles, Université Libre de Bruxelles 5Royal Museums of Art and History - KMKG Comics in Belgian weekly magazines remain understudied due to their fragmentary presence and the challenges of retrieval within large periodical collections. While comics are crucial artifacts of visual culture, their marginal status in general-interest magazines and their multimodal nature hinder both their discoverability and scholarly analysis. This article addresses the need for systematic identification of comics in mass-digitized archives by developing a semi-automatic detection workflow grounded in computer vision and object detection techniques. Drawing on a representative corpus of six general-interest illustrated magazines published in Belgium between 1934 and 1940, the study applies a distant viewing approach using YOLO models to detect and label comics alongside related visual categories such as cartoons, photographs, and advertisements. The methodology combines expert-driven annotation, collaborative labeling with weighted majority voting, and model evaluation to fine-tune visual classification. Our results highlight the conceptual and technical complexities of defining comics computationally, the importance of contextual metadata, and the benefits of hierarchical and iterative model training. The proposed workflow not only enhances access to hidden visual materials in historical magazines but also contributes to broader discussions on the remediation of cultural heritage collections and the role of computational methods in comics studies. Keywords: Historical Comics, Object Detection, Computer Vision, Periodical Studies, Distant Viewing, Visual Culture 137 1 Introduction Comics are important objects of study for understanding visual culture. The emergence of broad-audience illustrated magazines in the late 19th century arguably transformed visual culture and thus the vehicle for disseminating comics. However, only recently have magazines and periodical media been a focus of comic studies, favoring more permanently available and visible formats to piece together comic histories (Brinker et al., 2022; Letourneux and Vérilhac, 2025) 1 . Comics make up only a small portion of the varied and heterogeneous contents of such mass-market magazines, complicating their retrieval. The core issues for understanding the place and function of comics within periodicals, even with increasingly available digital collections, remain those of identification, accessibility, and archival visibility. The new opportunities and challenges raised by mass-digitization programs for cultural and media histories are well-known namely: if online repositories have radically altered the means of accessing and exploring archival sources, they nevertheless echo (in a different tone) the same issues of selection and erasure that bear on analogue archives (Bunout et al., 2023). The increased availability of digitized collections, as Patrick Leary famously argued in a 2005 piece on the impact of online archives on periodicals research, also produces and maintains an “offline penumbra [...] that increasingly remote and unvisited shadowland into which even quite important texts fall if they cannot yet be explored, or perhaps even identified, by any electronic means” (Leary, 2005). As Pierre-Carl Langlais suggests, “the main stake today is to develop ways of making visible what has been hidden or excluded from the historical indexing systems” (Langlais, 2021) 2 . With the amount of archival data exponentially increasing, the issue of the digital divide has gradually shifted from the ‘what’ to the ‘how’: how to facilitate and enhance accessibility and findability within such massive corpora? How to navigate online collections beyond textual characteristics (e.g. keyword search)? Against this background, identifying comics in magazines brings up important issues: not only because comics make up an often overlooked and neglected corpus, but also because their identification and indexing within digitized collections is particularly complex due to their multimodal and intermedial features. Manually browsing these analogue magazines would be too time-consuming and would reduce the research to a small corpus better suited for close-reading. Working with the digitized corpus, a text-based keyword search also falls short: many comics feature no text captions or titles, and text in comics is more often than not handwritten, making Optical Character Recognition (OCR) of little use for that specific purpose. In this context, it becomes essential to turn to image-based computational models as a means to identify comics in these magazines. This research reflects on an attempt and proposes a method for semi-automatic detection of comics with the support of computer vision workflows. Specifically we investigate this method on a corpus of digitized facsimile images of mass-market illustrated weekly magazines published in Belgium during the interwar period, as part of the ARTPRESSE project at Royal Library of Belgium (KBR) (Lemmers et al., 2023). The ARTPRESSE project centered on the representation of fine arts in general-audience 1 The focus in comics histories has moreover mostly lied on comics magazines and periodicals that feature a significant portion of comics material (children’s magazines, but also comics magazines for adults, for instance); rather than on the publication of comics within general-audience and multi-subject magazines. 2All translations are ours, unless otherwise indicated. 138 magazines, and it has made an extensive corpus of general-interest magazines digitally available. The place and function of comics in general-audience weeklies is an important blind spot in histories of comics in Belgium, and which remain extremely patchy in their account of the interwar period (Paques, 2011). The interwar indeed marks a fluctuating period where various models of comics and graphic narratives coexist across different formats (such as broadsheet newspapers, children’s magazines, woodcut novels), with the 1930s as a crucial decade for the “breakthrough” of comics in the press (Lefèvre, 2009). This rich corpus offers many more possible avenues of inquiry in terms of historical research and further allows for new computational approaches to grasp with its volume. The underlying research for this article thus participates in a further remediation of these digitized magazines following in the approach to “collections as data” (Padilla et al., 2023) and ultimately, in the stage of the project, aims to produce a dataset based on a single genre within these periodicals. 3 This article thus is a mid-stage reflection on the preliminary results and the identification of a workflow to identify comics within a selected corpus of mass-market magazines. This is explained in detail from the corpus, the proposed computer vision workflow used to detect, label and extract comics from the ARTPRESSE corpus, and thereby considers the conceptual and technical challenges for image classification raised by graphic material in such magazines. 2 Theoretical Framework 2.1 Computational Approaches: From Text to Object detection Traditionally, computer vision methods for historical sources have largely focused on the extraction of text (Bhatt et al., 2021; Paaß and Konya, 2011) using both Optical Character Recognition (OCR) for printed sources and Handwritten Text Recognition (HTR) for handwritten sources. These advancements made it possible for maps, letters, books, and more to be transcribed at continuously larger scales (Nockels et al., 2024). OCR was quickly embedded within most display interfaces and access platforms for cultural heritage materials, making text-based search and data mining the most common operations. In the last decade, the introduction of deep-learning algorithms has nevertheless allowed for the development of new approaches that more fully account for images and illustrations in large digitized archives. Melvin Wevers and Thomas Smits, among others, have argued for a “visual digital turn”, suggesting that convolutional neural networks (CNNs) “open up a new world of intuitive and serendipitous exploration of the visual side of digital archives.” (Wevers and Smits, 2020) Computational approaches to comics reflect those trends in digital humanities. OCR technology has contributed to direct efforts towards transcribing text balloons (Baskaran and Devi, 2023; Guo et al., 2006). Some recent and notable research includes Soykan, Yuret and Sezgin who showed that training newer OCR models, combined with segmentation techniques and pre/post-processing methods, significantly improves the accuracy of text detection within text balloons (Soykan et al., 2022). Hiroe and Hotta used OCR to find exclamation points in text balloons with the idea of enabling the detection of scene changes and providing a rough under3 The present project thus follows in the footsteps and draws inspiration from previous research aiming to automatically retrieve information such as literary feuilletons from KBR’s newspaper collections (Ali et al., 2024) 139 Figure 7: Combining the strengths of the different dates in the Compare Tool. Yellow: Folder dates, pink: manual dates, blue: automatic dates date information was precise to the day, highly accurate, but only available for 6% of the collection. To get the most out of this date information, we decided to publish it all. Researchers can then switch between the dates depending on their needs. To find as much material as possible, they can use the folder date, whereas to pinpoint found pages more accurately, they can use the manual or automatic dates. The user can even utilise the Compare Tool to combine the strengths of both. With the graph shown in Figure 7, the researcher can use the folder date line to see that there was a dip in the middle of November, with far fewer hits for ’Europa’ (Europe) than in the latter half of the month. Perhaps more interesting for researchers, though incomplete, are the automatic and manual date lines, which suggest dates for further investigation, such as 9th and 26th November, perhaps caused by the publicity surrounding Hitler’s speech on the 18th anniversary of the Beer Hall Putsch on the 8th of November, or the renewal of the Anti-Comintern pact on the 25th, which was presented as a milestone in the history of European unity. Obviously, given the incomplete coverage of the automatic and manual dates, this functionality must be used with care. Such hypotheses require further research into the propaganda content, for example with close reading. In general, the characteristics of the BNO collection have a large impact on how a researcher can work with it. The poor OCR quality influences search results, as search terms are less likely to be successfully found in transcripts with poor OCR quality. Furthermore, the incomplete availability of dates influences both what can be found when searching by date, and what is visualised in the Compare tool, as only items with dates can be plotted on the Compare tool graph. To allow the researcher to work with the collection and the Media Suite tools appropriately, it is essential that they are properly informed about the limitations of both. For this reason, we documented the BNO collection, with particular emphasis on the processing that was applied to generate dates (Figure 8). The documentation was produced by a team of data engineers and propaganda researchers to ensure that it included relevant information presented in such a way that researchers could understand the implications. This 242 Figure 8: Documentation page for the BNO collection documentation is included in the Media Suite and can be reached via a link in the collection(Figure 9)15. The tools available in the Media Suite are also documented, and this information is available via icons in the tool itself and the Help section of the Media Suite. In this way, the researcher is supported in both data and tool criticism, equipping them to work with the BNO collection in an informed manner. 15 https://mediasuitedata.clariah.nl/dataset/wartime-radio-bulletin-transcripts Figure 9: Collection Selector showing the BNO collection and its documentation link 243 Figure 10: An example of an audio recording with an ASR transcript in the Media Suite 4.2 Matching of radio audio with transcripts and newspapers Publishing the BNO transcript collection in the Media Suite gives unprecedented possibilities to explore the collection. But to truly start putting the propaganda puzzle together, we want to link the collections. Then a researcher could find a radio transcript discussing a given topic, listen to the audio, and look at what the newspapers were saying about that topic on the same day. The radio transcripts can also be used to fill some of the holes in the audio collection, and vice versa. There are far more radio transcripts (to give an indication, the digitised BNO collection alone comprises 68,986 pages) than there are audio recordings (2262). At the same time, there are cases where an audio recording is present while the transcript is missing, such as the news bulletin of 8am on 25th July 1943. In order to link transcripts to the relevant audio recording, we proposed the following steps: •Generate speech transcripts of the audio recordings • For each audio recording, find the transcripts within a small date range of the broadcast date of the audio (to allow some margin for errors) • Match the words in the speech transcript with the words in the candidate transcripts • With a sufficient number of matching words, link the transcript to the audio recording For the first step, we used Kaldi NL16 to generate speech recognition transcripts of the audio recordings. For 27% of the collection, no transcript could be produced. For another 16%, the produced transcripts contained fewer than 50 words, which could indicate poor quality. The remaining 57% of the collection has transcripts that are longer than 50 words. All transcripts were added to the radio audio collection in the Media Suite, and can be viewed there (Figure 10). However, a review of the transcripts quickly showed that the quality is very poor in many cases, so much so that it is often impossible to make sense of it (Figure 10). 16 https://github.com/opensource-spraakherkenning-nl/Kaldi_NL 244 Figure 11: An example of an audio recording with a poor quality ASR transcript The poor performance is most likely due to the poor audio quality of the recordings. Given this poor quality, combined with the poor coverage of the date information in the BNO, it was decided to abandon attempts to link the audio with the BNO transcripts within this project. Despite the fact that the items within the collections could not be explicitly linked, their publication together in the Media Suite still enables quantitative analysis of combined collections, as we indicated in the section on ’the radio audio collection’. To demonstrate this, we proceeded to compare semantic patterns in the BNO transcripts with the wartime newspaper collection in the Media Suite. 4.3 Quantitative analysis of combined collections Via the visualisations provided in the Compare tool, the Media Suite supports some basic quantitative analysis of the combined collections. In the example in Figure 12, the search results for the term ’Europa’ (Europe) are shown per month for the Nazi party newspapers (pink line) and the BNO transcripts (blue for folder date, yellow for the automatic page date). It can be seen that the two lines for the BNO transcript follow the same pattern, showing that despite the incompleteness of the automatic date information it is still usable to show trends. All three graphs show noticeable peaks in July 1941 and February 1943. The first peak reflects the increasing propaganda value of the ’Europe’ construct for the Nazis after the start of the attack on the Soviet Union on 22nd June. The Nazis framed the fight against the ’godless Bolshevism’ and the ’Asian’ Soviet Russia as a common European fight or crusade to preserve the ’Christian European civilisation’. The second peak shows the urgency of ’Europe’ in the propaganda offensive of the Nazis after the loss of the Battle of Stalingrad, in which Goebbels emphasised that only total war could neutralise the acute Bolshevik danger and that the fight against the Soviet Union was a matter of life and death, a struggle for the continued existence of Germany and Europe (Brolsma (2022)). As explained in Section 3.3, the Compare Tool offers an option to provide relative graphs, compensating for the imbalance between collection sizes and therefore making it easier to obtain a meaningful comparison. Figure 13 shows the query results of the keyword term ’Europa’ (Europe) relative to their own category. This graph confirms, 245 Figure 12: Comparing search results over time for the search term ’Europa’. Pink: Nazi party newspapers, yellow: BNO with automatic page dates, blue: BNO with folder dates Figure 13: Comparing search results over time for the search term ’Europa’, relative to the category/- collection size. Pink: Nazi party newspapers, yellow: BNO with automatic page dates, blue: BNO with folder dates and indeed foregrounds, the urgency of Europe as a propaganda concept in the context of Nazi Germany’s invasion of the Soviet Union. However, compared to Figure 12 it suggests that the peak of February 1943 was less significant, as Europe continued to be frequently used as a propaganda frame, also in the last two years of the war. These visualisations invite researchers to reflect on the continuities and discontinuities of propaganda narratives throughout the Second World War. This sort of data visualisation is therefore useful both to identify important moments for a closer investigation of the media content, and to map out the similarities and differences between various categories of wartime media, such as between pro-/anti-Nazi media and/or between newspapers and radio broadcasts. In other words, semantic patterns in quantitative visualisations can guide researchers to interesting ’media moments’ which merit further investigation. The texts produced by OCR and ASR assist researchers in searching collections. However, they are also a rich resource for quantitative analysis. We wrote a Jupyter Notebook, using the Media Suite APIs, that extracted the words occurring just before or just after a given search term. 17 This provides a tool for researchers to investigate in what context a word is used. 17 Stop words such as ’de’, ’het’, ’van’, ’in’, etc were excluded from this analysis 246 Table 1: The top ten most frequently occurring words before the word ’Europa’, per (sub)collection 1 2 3 4 5 Anti-nazi west oost oorlog rijk geheel Pro-nazi nieuwe geheel strijdt nieuw west BNO nieuwe geheel oost west landen Radio audio heel oost west landen we 6 7 8 9 10 Anti-nazi landen bevrijd heel vasteland nieuwe Pro-nazi oost landen oorlog midden vasteland BNO oorlog nieuw vrijheid vasteland heel Radio audio zuid binnen Nederland vanuit rest This notebook was used to conduct an analysis of the use of the word ’Europa’ (Europe) in the antiand pro-Nazi newspapers, the BNO collection (via OCR) and the radio audio collection (via ASR). The results are shown in Table 1 Interesting is that the most frequently occurring word in both the pro-Nazi and BNO collections is ’nieuwe’ (new). The variant ’nieuw’ also appears high in the ranking for these collections. These results show that the construct ’Europe’ was very important for the occupying authorities to promote a national-socialist vision of the future in the nazified printed press and via Radio Hilversum. Of course, the poor - and variable - quality of the ASR and OCR means that this analysis is at best incomplete, and may also be inaccurate. For example, if a frequently occurring word is poorly recognised, then it will be counted much less often, causing it to appear to occur less frequently that it actually does. However, the fact that the results in the BNO and pro-Nazi categories are so similar seems to indicate that the quality is sufficient to discover important Nazi media frames via digital research methods, such as the frame of a ’new Europe’. The results in the ASR of the (both proand anti-Nazi) radio seem to consist mainly of geographical indications. Due to the poor quality of the texts, such analyses cannot provide reliable conclusions, but can be used to discover possible interesting trends for further qualitative investigation. 5 Conclusion During the Second World War in the Netherlands, the Berichtendienst Nederlandsche Omroep (BNO) made an important contribution to the propaganda of the Nazi occupying regime through their news broadcasts. The publication of the digitised BNO transcripts in the CLARIAH Media Suite enables researchers to more effectively search through the transcripts. We did not achieve the goal of linking these transcripts to the audio recordings also stored in the Media Suite. Yet despite the absence of explicit links, the publication of collections of newspapers, radio transcripts and radio audio recordings in the same online environment means that researchers can analyse the BNO collection in conjunction with other categories of war media (categories per political-ideological signature and media type) for the first time. The availability of the OCR and ASR transcripts make it possible for them to combine qualitative analysis of the media content with various quantitative research techniques. 247 For both quantitative and qualitative research of the BNO transcripts, date information is essential. Making the transcripts searchable by date enables historians and media scientists not only to effectively analyse the reporting around various events - as demonstrated in the examples above - but also to make optimum use of the digital tools that the Media Suite offers. For example, they can compare semantic patterns in various media categories in the Compare Tool to identify important ’media moments’ or study the use of certain frames over time. With the algorithm we developed, we succeeded in extracting dates for part of the collection. The biggest obstacle to the automatic recognition of date information is the poor quality of the OCR. A priority in the future is therefore to improve the OCR, so that more pages can be reliably dated. This could be done by using new OCR tools, and potentially by re-scanning difficult pages. This should be discussed with the company that produced the original scans, as they may be able to suggest improvements. Alternatively, a more labour-intensive approach would be to manually annotate more pages with the date. As was demonstrated in this project, both approaches can be combined. The large variation in date formats used in the collection also presented a challenge. An interesting option for the future could be to finetune or fully train a NER model on part of the BNO collection. This would ensure that the model is trained on dates in the formats and context such as they typically appear in the BNO collection (including the typical OCR errors) and hence can recognise them with more accuracy. This would require manual effort to identify all the variations and tag these to use as training material. In addition to a better date recognition, improving radio ASR transcript quality could also help to make linking the radio broadcasts with the radio transcripts feasible. New speech recognition models are now available, which offer the potential for improvement. These should be tested on the radio broadcasts to see if it is possible to achieve better ASR quality. To increase the value of the collection of wartime media in the CLARIAH Media Suite for future media-historic research into propaganda in Dutch newspapers and radio, a further expansion of this wartime media collection is desirable. This can be achieved by digitizing and publishing (parts of) the sizable radio archive of the years 1940-1945, such as the Radio Oranje transcripts and the reports of the listening service of the Dutch government in London exile. Such an addition would allow for research that will deepen our understanding of the dynamics in radio propaganda during the Second World War, and more specifically of the transnational war of words between Nazi-controlled transmissions from Hilversum and pro-Allied broadcasts from London. In the course of the Media War Matching project, we learned a number of valuable lessons about the publication of such collections. The first was that quality starts at the source, and we would recommend that any future projects intending to publish a collection first conduct a thorough investigation of the data to identify quality issues up front. For large collections this is a challenge, as manual checking of everything is impossible, but checking samples alone can be misleading, as was the case in this project. The second lesson concerns the importance of data and tool criticism when dealing with such collections. The limitations of both data and tools must be clearly communicated with researchers to allow proper use of the collection. The third lesson is that cooperation between data engineers and historical experts is essential both to achieve this clear communication and to publish the collection (including development of tools and processing of data) in such a way that it is optimally usable by researchers. 248 This ongoing dialogue between experts in the various Digital Humanities disciplines brings us to our final and pivotal point. While it is important to keep in mind the ideal situation of digitised collections, each with complete metadata and linked to related collections, it is possible to achieve a great deal while falling far short of that ideal. In this particular project, we didn’t put together the complete puzzle of the three propaganda collections, because of both computational challenges and the wartime circumstances that affected the production and preservation of these sources. However, what we did achieve has already proven to be of great value to researchers, showing how quantitative visualisations of media collections can serve as an entry point for further qualitative research. We would therefore encourage others to embark on the challenge of publishing such collections, to discover the limitations of the data and tools, and to use these to assist transparent and usable publication of the data, rather than letting them prevent it. 6 Acknowledgement This work was enabled by the CLARIAHPLUS project funded by NWO (Grant 184.034.023), and the Mondriaan Fonds (75 jaar vrijheid).18 References Marjet Brolsma. Propagandaslag om europa: Wisselwerkingen tussen de nederlandse genazificeerde en antinazistische pers na operatie barbarossa. Tijdschrift voor Geschiedenis, 135, 2022. doi: https://doi.org/10.5117/TvG2022.2/3.003.BROL. Mark Connelly, Jo Fox, Stefan Goebel, and Ulf Schmidt. Prologue. power and persuasion’: Propaganda into the twenty-first century. Propaganda and Conflict. War, Media and Shaping the Twentieth Century, 2019. Vincent Kuitenbrouwer. The traces of a media war: Archives of dutch broadcasts from london during the second world war. Journal of Media History, 25, 2022. doi: https://doi.org/10.18146/tmg.821. Vincent Kuitenbrouwer and Marjet Brolsma. Audio on paper: The merits and pitfalls of the dutch digital media archive for studying transnational entanglements during the second world war. Journal of European Television History & Culture, 12, 2023. doi: https://doi.org/10.18146/view.306. Vincent Kuitenbrouwer and Huub Wijfjes. Media war. Journal of History, 135, 2022. doi: https://doi.org/10.5117/TvG2022.2/3.001.KUIT. Onne Sinke. Onderling strijdend voor de goede zaak: Radio oranje en de brandaris. Journal of Media History, 8:97–109, 2005. doi: https://doi.org/10.18146/tmg.540. Onno Sinke. Verzet vanuit de verte. De behoedzame koers van Radio Oranje. Augustus, 2009. Hans van den Heuvel and Gerard Mulder. Het vrije woord. De illegale pers in Nederland. SDU Uitgevers, 1990. 18 https://www.mondriaanfonds.nl/subsidie-aanvragen/regelingen/open-oproep-75jaar-vrijheid/ 249 Dick Verkijk. Radio Hilversum 1940-1945. De omroep in de oorlog. Uitgeverij de Arbeiderspers, 1974. René Vos. Niet voor publicatie.De legale Nederlandse pers tijdens de Tweede Wereldoorlog. Uitgeverij Sijthoff, 1988. Ivo van de Wijdeven. De macht van het verleden. Geschiedenis als politiek wapen. Unieboek Het Spectrum, 2022. Lydia E. Winkel. De ondergrondse pers 1940-1945. Martinus Nijhoff, 1954. Mariëtte Wolf and Frank van Vree. De krant. Een cultuurgeschiedenis, chapter Oorlog herstel en vernieuwing 1940-1950, pages 205–218. Boom, 2019. 250