Full text
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 1 Ir0000014 – Itserr D8.1.3 - Whitepaper on the Results Using uBIQUity on Bible and Qur'anic Commentaries (ONFIELD) Document reference: ITSERR-WP8-D8.1.3-ONFIELD Version number: 01.00 Status: FINAL Last revision date: 30/07/2025 by: Anna Mambelli, Fabrizio D’Avenia, Sara Abram, Marcello Costa, Chiara Palillo, Fabio Tutrone Verification date: DD/MM/YYYY by: Board Approval date: DD/MM/YYYY by: MUR Subject: IR0000014 – ITSERR Whitepaper on the results using uBIQUity on Bible and Qur'anic commentaries Filename: ITSERR_WP8_D8.1.3_01.00_FINAL.docx This document is available in the ITSERR WP Management document repository at: ITSERR-WP8\_Deliverables\D8.1.3_ONFIELD\
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 2 Change history Version Number Date Status Summary of main or important changes 00.01 30/05/2025 WORKING Working version 00.02 27/06/2025 DRAFT Complete revision of the document after ITSERR Board review 01.00 30/07/2025 FINAL Finalisation Distribution List Name Company Role ITSERR Board Members All ITSERR partners
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 3
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 4 Table of Contents 1. Document Overview 5 2. Objectives 6 3. Scope 7 4. State of the Art 7 4.1. The Golden Age of Biblical Studies and Computational Technologies 8 4.2. Semantic Analysis of the Greek Bible 10 4.3. Different Approaches to Textual Semantics 11 4.4. Beyond Canon: Multiplicity and Transmission in the Prophetic Tradition and Islamic Exegesis 15 4.5. From Access to Analysis: Digital Resources and Methodological Gaps in Islamic Intertextual Studies 16 5. The New Tool for Intertextual Reference Retrieval 17 5.1. Greek Search Engine Evaluation 17 5.2. Latin Search Engine 21 5.3. Arabic Search Engine 25 6. Methodology 27 7. Findings 28 7.1 Greek and Latin 28 7.2. Arabic 35 8. Conclusion 42 9. Appendix 43 9.1. uBIQUity - Latin and Greek Tool Test (WP8): Feedback Questionnaire 43 9.2. uBIQUity - Arabic Tool Test (WP8): Feedback Questionnaire 84
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 5 1. Document Overview The Ministero dell’Università e della Ricerca (MUR) has concluded a Grant Agreement (GA) with the ITSERR consortium (ITSERR)1 for the provision of IT services to support the existing national infrastructure and bring it to a higher level of maturity, in terms of involvement of technology and ability to increase the innovation, quality and variety of the knowledge produced by the community of Religious Studies. Within the structure of the ITSERR project, the Work Package 8 (WP8)–uBIQUity, which incorporates the “BI” of the Bible(s) and the “QU” of the Qurʾān in its title, aims at investigating the sacred texts of Christianity and Islam in different environments and historical periods through two huge corpora: Greek and Latin Christian commentaries (broadly understood as exegetical works) on the Bible(s) written from the Patristic age until the Late Byzantine period, and classical commentaries on the Qurʾān written in Arabic (tafāsīr) from the rise of Islam until the 15th century. These works are unique sources for the study of knowledge, readings, and hermeneutics of the sacred texts through the centuries. The intertextual references, conscious or unconscious, that the ancient commentaries contain work as invisible “places of memory”, making sacred texts “ubiquitous” (hence the title of the project). By interweaving the research methods of the Humanities and state-of-the-art Computer Science research, uBIQUity is developing a novel research tool that can identify references to the Bible(s), the Qurʾān, and the ʾaḥādīṯ (the Sayings of the Prophet) in ancient Christian and Islamic exegetical works with a higher degree of accuracy than pre-existing resources. The WP8 team members are: ● Sara ABRAM, University of Palermo (UNIPA): Arabic texts. ● Marcello COSTA, UNIPA: visual communication design. ● Fabrizio D’AVENIA, WP8 Leader, UNIPA. ● Cinzia FERRARA, UNIPA: visual communication design. ● Anna MAMBELLI, WP8 Scientific Coordinator and Product Owner, University of Modena and Reggio Emilia (UNIMORE): Greek and Latin texts. ● Chiara PALILLO, UNIPA: visual communication design. ● Fabio TUTRONE, UNIPA: Latin texts. Regarding research in the field of Computer Science, for the Latin section, the collaboration with the UNIMORE team composed of Davide CAFFAGNI, Federico COCCHI, Marcella 1 ITSERR is a consortium composed of National Research Council (CNR), University of Modena and Reggio Emilia (UNIMORE), University of Naples L’Orientale (UNIOR), University of Palermo (UNIPA), and University of Torino (UNITO).
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 6 CORNIA, Rita CUCCHIARA, and also with Cesare CONCORDIA and Carlo MEGHINI of the National Research Council (CNR, Pisa) is fundamental. As for the Arabic section of WP8, Giovanni PUCCETTI of the CNR is working on the IT side. Concerning research and work on Ancient Greek, WP8 has decided to collaborate with the team of another project with similar aims, which has been incubated in the framework of RESILIENCE RI, namely the PRIN (Italian Research Project of National Interest, n. 20229E83B3) Resilient Septuagint. An Initial Exploration of the Semantics of Killing and Healing in the Septuagint and its Reception in Patristic and Late Antique Sources (3rd cent. BCE-5th cent. CE). This PRIN project involves the Alma Mater Studiorum University of Bologna (UNIBO, PI Research Unit), the University of Bari (UNIBA), and the University of Catania (UNICT). The Resilient Septuagint team members are: Luca ARCARI (University of Naples Federico II), Laura BIGONI (UNIBO), Laura CARNEVALE (UNIBA Research Unit Chief), Davide DAINESE (Principal Investigator [PI], UNIBO), Isabella PIGNOCCO (UNICT-UNIBO), Arianna ROTONDO (UNICT Research Unit Chief and deputy PI), Giorgia SAMPÒ (DDC University of Southern Denmark-UNIBO), Gianluca SCATIGNO (UNIBO), and Marco ZANELLA (University of Padua-UNIBO). This collaboration between WP8 and Resilient Septuagint has included and will continue to include the development of a shared methodological framework and interoperable datasets for the study of Greek biblical texts and their heritage. This document is the third deliverable from the WP8 of ITSERR (D8.1.3) and illustrates the scientific advances enabled by the uBIQUity methodology and software, drawing on the experience of specialists/testers from the scholarly community. It not only describes uBIQUity’s performance in theory (as previously detailed in D8.1.2) but also presents the results of its “on-field” application. More specifically, the whitepaper first devotes considerable attention to the state of the art at the intersection of biblical and Islamic studies and digital technologies. The sections titled “The New Tool for Intertextual Reference Retrieval”, “Methodology”, and “Findings” then provide significant and detailed information on the tool and the results of uBIQUity’s field operations, showing both the outcomes of user testing and the effectiveness of the algorithms. In doing so, the document highlights the role of the uBIQUity platform not only as a prototype but also as a potential instrument for generating new knowledge and opening new avenues of research. 2. Objectives Building on the project’s characteristically fruitful combination of philological methods, computer science, computational linguistics, and design, this whitepaper outlines the current outcomes of the uBIQUity platform as an innovative tool that offers (Digital) Humanities researchers advanced LLM(Large Language Model)-based functionalities and a new way of approaching the knowledge produced through intertextual analysis. The textual corpora digitized and/or revised and semantically enriched in the earliest stages of the project (see
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 7 D8.1.1), together with the User Experience (UX)/Data Visualization Design perspective (as detailed in D8.1.2), provided a solid basis for testing a tool that aims to change the experience of readers of early and medieval Christian and Islamic literatures. Since its very beginning, the uBIQUity platform has prioritized a qualitative—rather than quantitative— approach to text comparison, which is reflected in its use of a node-based user interface (UI) designed to manage, display, and compare complex data and metadata across different levels (see D8.2 for further details on the tool’s interface and user interface testing). In this crucial phase, respected experts in early and medieval Christian and Islamic textual traditions have been invited to use and comment on the uBIQUity platform, with the main purpose of assessing its performativity, user-friendliness, scientific potential, and accuracy. These tests have ultimately allowed the WP8 team to reassess the platform’s sensitivity to different cognitive learning styles, improving its flexibility and enhancing its potential for innovation. 3. Scope The project’s interdisciplinary essence—reflected in the integration of the humanities and digital technologies and in its openness to the multilingual nature of religious traditions— has remained a key factor during the testing phase of the uBIQUity search engine, which the WP8 IT team developed on the basis of the methodology, digital contents, and tools produced in the earlier stages of the project (see the deliverable D8.1.1). The team of experts in philology, computer science, computational linguistics, and design has played an integral role in the development of the uBIQUity platform. At the present stage, their different skills and domains of expertise have been further enriched by the contributions of additional scholars and specialists, selected to provide on-field feedback and to offer a wide range of stimulating questions and comments. Across all roles within the project, there was a shared understanding that evaluating the functionalities of such a complex and multi-layered product as a new semantic search engine inherently requires precise assessment informed by diverse disciplinary perspectives and viewpoints. At the same time, all team members consistently worked to consolidate the various forms of feedback, knowledge backgrounds, and skill sets into a coherent framework, with the aim of providing the Religious Studies scientific community with an agile and innovative digital tool. 4. State of the Art For scholars of ancient and medieval texts, religious traditions, and intertextuality, computational tools for text reuse detection have become indispensable, allowing for the systematic analysis of vast interconnected corpora that would be extremely difficult to examine manually. During the testing phase of the uBIQUity platform as a shared repository for Greek, Latin, and Arabic texts, the WP8 research team maintained a constant focus on existing text reuse tools, recognizing them as representing the current state of the art in the field. The primary goal of this awareness was to expand search capabilities in response to
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 8 the evolving needs of the scientific community, paving the way for novel approaches to intertextual research through the use of Artificial Intelligence (AI) and LLMs, and thereby overcoming the known limitations of the tools currently available. 4.1. The Golden Age of Biblical Studies and Computational Technologies Although it has not yet made a major impact in the AI field, biblical studies represent an area where digital transformations developed early and in a well-structured manner. It is precisely in continuity with these experiences—and with the already established infrastructure for the digital processing of sacred texts—that our project aims to position itself, proposing a possible extension toward integration with AI tools and methods. By biblical studies, we do not refer solely to the contributions of theologians or historians of Christianity.2 Rather, we are referring to the most significant results that have emerged from decades of scholarly exchange, research, and collaboration—activities that have, directly or indirectly, revolved around the Society of Biblical Literature since the 1980s. The initial developments are associated with John R. Abercrombie and date back to the early 1980s, within the broader context of digitizing Greek literary corpora for the Thesaurus Linguae Graecae. The digitization of Rahlfs’ 1935 edition of the Septuagint (LXX) highlighted the need to align the Greek text with the Hebrew/Aramaic of the Masoretic Text.3 This led to the CATSS project (Computer Assisted Tools for Septuagint/Scriptural Study, 1986), which aimed to create a database preserving every single variant reading of the First Testament in Greek, in addition to the existing alignment, and to provide a morphological analysis of all words in both the LXX and the Masoretic Text.4 From a technological perspective, this meant offering texts that were searchable (using a David Packard IBYCUS personal computer), and morphologically analyzed.5 From a practical standpoint, the implementation of a straightforward binary algorithm, operating within a word table and a hierarchical tree structure for desinences, proved to be decisive.6 2 Regarding the digital turn, it is worth noting that theologians have provided an authoritative review of the ethical dilemmas posed by certain issues, particularly those related to digital platforms. This is exemplified by C. Clivaz, “The Bible in the Digital Age: Multimodal Scriptures in Communities,” in T. Hutchings/C. Clivaz (eds.), Digital Humanities and Christianity. An Introduction, Berlin/Boston: De Gruyter, 2021, 21–46. This essay also includes a mapping of academic efforts to digitize the scriptural corpora of the Jewish and Christian traditions. 3 See E. Tov, The Greek and Hebrew Bible. Collected Essays on the Septuagint, Leiden/Boston/Köln: Brill, 1999, 31. 4 See http://ccat.sas.upenn.edu/rak/catss.html; E. Tov, A Computerized Database for Septuagint Studies: The Parallel Aligned Text of the Greek and Hebrew Bible, Stellenbosch: JNSL, 1986. 5 In that case, this was made possible thanks to a program David Packard designed in 1973 for IBM Mainframe 390/91 at Nasa’s Goddard Flight Center in Washington, DC, as well as the supervision of a team of experts. 6 See D.W. Packard, “Computer-Assisted Morphological Analysis of Ancient Greek”, in A. Zampolli/N. Calzolari (eds.), Computational and Mathematical Linguistics: Proceedings of the International Conference on Computational Linguistics, Firenze: Olschki, 1980, 343–355, on p. 348.
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 9 In more general terms, the late 1980s and 1990s were particularly fortunate for computational linguists interested in the sacred scriptures of the Jewish and Christian traditions.7 In 1998, CATSS was integrated into Accordance (https://www.accordancebible.com/), the second of the so-called biblical software created to handle ancient Greek,8 after BibleWorks (https://www.bibleworks.com/), developed in 1992 to support preaching and pastoral activities.9 Today, the digital landscape of biblical studies offers a wide array of source collections that, from a technological standpoint, focus more on data representation (and knowledge) than on data processing. These platforms manage the copyrights of various dictionaries and editions for a fee, including Olive Tree (https://www.olivetree.com/), BibleWorks, Accordance, Logos (https://www.logos.com/), Brill Dead Sea Scrolls Concordance (which uses and refines Martin G. Abegg’s database—through texts edited by the Oxford series Discoveries in the Judean Desert—and now offers it integrated into Brill’s Electronic Library: https://brill.com/display/package/9789004310391?srsltid=AfmBOoozmj_Yt77uLHPNEx3 M_hBzwnYAsa_glzInrn3pyYj5VAZC609N), or are open source, such as the STEP Bible (https://www.stepbible.org/). 7 On the Second Testament see: M.E. Davison, “New Testament Greek Word Order”, Literary and Linguistic Computing (LCC) 4 (1989), 19–28; H.H. Greenwood, “St Paul Revisited—A Computational Result”, LCC 7 (1992), 43–47, Id., “St Paul Revisited—Word Clusters in Multidimensional Space”, LCC 8 (1993) 211–219; Id., “Common Word Frequencies and Authorship in Luke’s Gospel and Acts”, LCC 10 (1995), 183–187; G. Ledger, “An Exploration of Differences in the Pauline Epistles using Multivariate Statistical Analysis”, LCC 10 (1995), 85-97; D.L. Mealand, “Correspondence Analysis of Luke”, LLC 10 (1995) 171-182; Id., “Measuring Genre Differences in Mark with Correspondence Analysis”, LLC 12 (1997), 227–245; A.J.M. Linmans, “Correspondence Analysis of the Synoptic Gospels”, LLC 13 (1998), 1–13; G.K. Barr, “A Computer Model for the Pauline Epistles”, LLC 16 (2001), 233–250; Id., “Interpolations, Pseudographs, and the New Testament Epistles”, LLC 17 (2002), 439–455; Id., “Two Styles in the New Testament Epistles”, LLC 18 (2003), 235–248; A. Wilson, “Developing Conceptual Glossaries for the Latin Vulgate Bible”, LLC 17 (2002), 413–426. On the First Testament it is worth mentioning the third section of the Historical Dictionary of the Hebrew Language (see R. Merkin/Z. Busharia/E.Meir, “The Historical Dictionary of the Hebrew Language”, LLC 4 [1989], 271–273) and the application of the famous TUSTEP (Tübingen System von Textverarbeitungsprogrammen) to the Book of Daniel (W. Bader [ed.], „Und die Wahrheit wurde hinweggefegt“: Daniel 8 linguistisch interpretiert, Tübingen: Francke, 1994). 8 Accordance achieved this by merging Bibles from Roy Brown’s biblical software, originally ThePerfectWord, later MacBible, and the Dead Sea Scrolls word database created by another Mac-user, Martin G. Abegg. See Tov, The Greek and Hebrew Bible, 43. 9 Collections of digitized biblical versions and biblical commentaries began to circulate in 1980.
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 16 authenticity. Rather, they offer valuable clues for reconstructing the relationship between the prophetic legacy and Qurʾānic interpretation, as well as for tracing the processes through which meaning was constructed and reshaped over time. From this perspective, the transmission of the prophetic and exegetical heritage—not only through the reuse of aḥādīth, but also through the intertextual reworking of tafsīr traditions themselves—emerges not as a fixed canon, but as a dynamic continuum: historically mobile, yet consistently grounded in specific exegetical contexts. The plurality of versions and uses is not a flaw to be corrected, but rather an essential dimension to be investigated. Acknowledging this multiplicity does not undermine the integrity of the tradition; instead, it opens new avenues for exploring how meaning was preserved, adapted, and transformed over centuries of Islamic scholarship. 4.5. From Access to Analysis: Digital Resources and Methodological Gaps in Islamic Intertextual Studies When it comes to the Qurʾān and other foundational texts of the Islamic tradition—such as ḥadīth collections and major works within the tafsīr corpus from the 8th to the 15th century— digital libraries and Arabic-language websites (e.g., Shamela and IslamWeb; see D8.1.1) offer an impressive range of accessible resources. From a technological perspective, however, these platforms primarily focus on the digitization and structured presentation of sources rather than on computational processing or semantic enrichment. Many tools, such as Altafsir.com (https://www.altafsir.com/) for the tafāsīr or Sunnah.com for the aḥādīth (https://sunnah.com/), provide fully searchable texts organized according to their internal divisions (chapters and verses for the Qurʾān; books and chapters for ḥadīth collections). While these platforms facilitate navigation and citation retrieval, they do not yet offer mechanisms for semantic search or intertextual analysis. Most existing tools rely on literal string-matching or keyword-based search functions. While effective for retrieving exact phrases, they do not support the identification of paraphrastic reuse, conceptual similarity, or stylistic transformation—features that are central to understanding how the Qurʾān and the prophetic tradition have been quoted, adapted, and interpreted over time. In this regard, the limitations observed in Arabic studies mirror those in Greek and Latin: N-gram-based or literal-matching algorithms are highly sensitive to morphological or syntactic variation, and even minimal reformulations can prevent detection. This produces a reductive model of intertextuality that fails to capture the full complexity of historical textual transmission and exegetical practices. Moreover, the Arabic textual tradition still lacks certain tools that have long been available in other fields, such as a unified, TLG-style environment where texts can be queried, compared, and annotated in semantically rich ways. Widely used platforms like Sunnah.com, despite their accessibility, only allow for surface-level retrieval and do not incorporate mechanisms for identifying intertextual relationships beyond exact matches. This technological gap becomes even more pronounced when considering the lack of advanced tools for detecting and analyzing text reuse and intertextuality in Islamic religious
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 17 texts. No existing resource currently allows for the efficient analysis of intertextual references using methods such as automatic lemmatization of inflected forms, in-depth semantic analysis, or N-gram-based concordance. While tools such as dictionaries, translations, and morphological analyzers are available, they are often scattered across different platforms, rarely updated, and lack integration. This fragmentation severely limits scholars’ ability to trace meaningful connections between tafāsīr and other genres, particularly the collections of aḥādīth. Moreover, no single platform currently brings together linguistic support, semantic enrichment, and computational approaches for detecting both literal and non-literal reuse. An integrated infrastructure combining lemmatization, morphological annotation, and semantic similarity metrics would represent a major advance for Islamic Studies as a whole. It would enable researchers to explore intertextual references with greater precision and to reconsider the methodological frameworks through which Islamic sources are analyzed. The capacity to visualize and interact with a large quantity of data through a user-friendly platform interface would improve accessibility and inclusivity in the field of Religious Studies and support the emergence of new perspectives and analytical approaches. This scope has guided the WP8 team in developing a rigorous methodology for identifying and analyzing intertextual references, thereby laying the groundwork for a semantic search engine. uBIQUity represents a first valuable step toward addressing this gap by exploring the potential of combining morpho-syntactic data and vector-based similarity measures to support the study of textual reformulations in Islamic exegetical traditions. 5. The New Tool for Intertextual Reference Retrieval 5.1. Greek Search Engine Evaluation To facilitate intertextual research within uBIQUity, a search engine dedicated to the Septuagint—the oldest known Greek translation of the Hebrew Bible—has been developed. In this system, the atomic unit of retrieval is the biblical verse, each enriched with historically attested textual variants. The search engine enables both syntactic and semantic search, supporting flexible exploration of textual parallels. The syntactic component employs language-agnostic string-matching techniques, including common word overlaps, trigrams, and shingles. For semantic search, SPhilBERTa, a transformer-based language model trained on Greek philosophical and theological texts,21 was used to generate vector embeddings for each verse and query fragment. These embeddings enable semantic similarity matching beyond literal word overlap. 21 See F. Riemenschneider/A. Frank, “Graecia capta ferum victorem cepit. Detecting Latin Allusions to Ancient Greek Literature”, in Proceedings of the Ancient Language Processing Workshop, Shoumen: Incoma, 2023, 30–38.
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 18 To evaluate the system’s performance, test data were provided from humanist WP8 researchers, who identified approximatively 80 intertextual references between selected works of Clement of Alexandria and passages from the Septuagint (for further details on these selected exegetical works, see D8.1.1). For each reference, the corresponding fragment from Clement was used as a query to assess whether the system could retrieve the relevant biblical passage(s) within the top 50 results, ranked by pertinence, a metric commonly known as Success@K or Recall@K. Three search configurations were tested: syntactic-only search (based on surface string similarity), embedding-only search (based on semantic vector similarity), and hybrid search (combining syntactic and semantic scoring). In addition, each configuration was evaluated in two modes: ● Full Search, conducted over the entire Septuagint corpus. ● Aided Search, restricted to a subset of biblical books potentially relevant to Clement of Alexandria (Genesis, Exodus, Leviticus, Numbers, Deuteronomy, Psalms, and Isaiah). The results for the Full Search were as follows: ● Syntactic Search: 51 out of 78 references found (65.4%). ● Embedding Search: 38 out of 78 references found (48.7%). ● Hybrid Search: 55 out of 78 references found (70.5%). The Aided Search mode yielded slightly improved results but followed the same overall trend. Extended Evaluation of Search Configurations To further assess and refine the retrieval system, a series of controlled experiments were conducted on the full Septuagint corpus. These tests examined the impact of different retrieval strategies and data configurations on two standard metrics: Recall@10 (R@10), which measures how often the relevant verse appears in the top 10 results, and Mean Reciprocal Rank (MRR), which captures how highly the correct result is ranked on average. The following search strategies were tested: 1. Token-Based Search: matching based on individual word tokens (normalized forms). 2. Fully Language-Agnostic Search: a composite approach combining token overlap, trigrams, and word-level shingles. 3. Semantic Search: vector-based similarity using SPhilBERTa embeddings.
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 19 4. Hybrid Search: a combined scoring approach blending syntactic and semantic similarity through a linear combination. Each configuration was executed in two modes: ● With Apparatus: incorporating textual variants and alternate readings from the critical apparatus of the reference edition. ● Without Apparatus: using only the main text established by the modern editor(s), without variant expansions. The following table summarizes the evaluation results: Without Apparatus With Apparatus Search Type R@10 MRR R@10 MRR Token-Based 0.42 0.32 0.47 0.37 Language-Agnostic 0.40 0.32 0.56 0.42 Semantic (Embeddings) 0.33 0.26 0.33 0.21 Hybrid 0.42 0.33 0.69 0.49 Table 1. Different search strategies and their respective evaluation results. To illustrate the performance patterns, the following plots show R@K for increasing values of K across configurations: Figure 1. Token-based search type (with apparatus).
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 20 Figure 2. Language-agnostic search type (with apparatus). Figure 3. Semantic (embeddings) search type (with apparatus).
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 21 Figure 4. Hybrid search type (with apparatus). These results reveal several key insights: ● Incorporating the critical apparatus consistently improves performance across all methods. ● The semantics approach alone yields the worst performance, possibly due to lack of high-quality models for ancient languages. ● Hybrid search with apparatus achieves the best overall performance, validating the benefit of combining surface-level and semantic cues. These findings indicate that combining syntactic and semantic methods significantly improves retrieval performance in detecting intertextual references, thereby underscoring the value of multimodal search strategies in Digital Humanities research. The Greek semantic search engine has been integrated into uBIQUity as a dedicated node type, enabling flexible, research-driven textual comparisons. 5.2. Latin Search Engine To support intertextual analysis within uBIQUity for Latin texts, a dedicated search engine was developed with a focus on retrieving semantically related biblical passages referenced in Latin patristic literature. The primary corpus consists of two major Latin Bible versions, Jerome’s Vulgate (W_VULG) and the Vetus Latina (S_VL), alongside annotated intertextual references from Augustine’s De Genesi ad litteram libri duodecim. The engine is built around a sentence-level semantic search model where each biblical verse is treated as a standalone retrievable unit. Unlike traditional keyword-based systems, this
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 22 search engine leverages dense vector embeddings trained to capture semantic similarity, thereby enabling the identification of allusive and stylistically varied intertextual connections. A key innovation lies in the use of LLM-generated synthetic data to address the low-resource nature of Latin. Specifically, a state-of-the-art Latin embedding model (i.e., Latin BERT or LaBERTa in our experiments) was fine-tuned using synthetic triplets (source, positive, negative) produced by GPT-4. Positive samples retained the meaning of the source passage while varying in form, and negative samples exhibited superficial similarity but lacked true semantic correspondence. This contrastive training approach shaped a highly contextualized embedding space for Latin texts. The prompting strategy used to generate these samples is illustrated in Fig. 5. Figure 5. Example of the prompt template employed for generating positive and negative samples. Experiments were conducted on a training set comprising 35k verses from the W_VULG corpus and 23k from the S_VL corpus, including 22k overlapping passages. During training, positive pairs were formed using either GPT-4-generated positives or aligned verses from the two Bible versions. Negative samples were sourced exclusively from the LLMgenerated outputs. To enhance variability, one of the two positive options was selected at random for each training instance. Representative synthetic examples are provided in Fig. 6.
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 23 Figure 6. Synthetic positive and negative passages generated by GPT-4, starting from a source passage from W_VULG. Evaluation was performed on a test set of 376 intertextual references manually annotated by the WP8 humanists from Augustine’s commentary De Genesi ad litteram, which are mapped to verses in both W_VULG and S_VL. Tables 2 and 3 show the retrieval results, measured in terms of the average number of queries for which the target passage is retrieved within the first k passages, with k={1,2,3,5,10,20} (i.e. Recall@k). The proposed fine-tuning pipeline is compared against standard fine-tuning, where the models were trained without positive and negative passages generated by GPT-4 but only leveraging the correspondences between W_VULG and S_VL to build positive pairs for contrastive learning. Zero-shot results of the models are also provided as a reference. Table 2. Results on the W_VULG corpus comparing the proposed fine-tuning strategy using LLMgenerated synthetic data against standard fine-tuning (without generated samples) and the original (non-finetuned) model.
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 24 Table 3. Results on the S_VL corpus comparing the proposed fine-tuning strategy using LLM-generated synthetic data against standard fine-tuning (without generated samples) and the original (non-fine-tuned) model. Across all configurations, models fine-tuned with the synthetic data pipeline consistently outperformed those using only standard fine-tuning, with particularly strong gains observed in Recall@1. These improvements were consistent across both biblical corpora and embedding architectures, despite synthetic data being generated solely from the W_VULG corpus. This cross-corpus generalization underscores the robustness and transferability of the synthetic data approach for retrieval in low-resource settings. Further validation was conducted on the most challenging subset of the test set composed of references with minimal lexical overlap (corresponding to similarity scores between 0.0 and 0.25). Results presented in Table 4 confirm that the proposed approach significantly improves retrieval effectiveness in these difficult cases. For instance, Recall@1 for Latin BERT with token averaging on W_VULG increased from 3.9% to 11.8%, representing a relative gain of over 200%. In the S_VL corpus, the same configuration achieved 8.9% Recall@1 where standard fine-tuning failed to retrieve any correct passage (0.0%). These results highlight the particular utility of synthetic data augmentation for identifying subtle, non-obvious intertextual relationships.
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 25 Table 4. Performance comparison on the “hard” subsets of W_VULG and S_VL corpora (i.e., queries with the lowest similarity to the corresponding biblical verse). 5.3. Arabic Search Engine For the Arabic component of the tool, which involves semantic indexing of the corpus, the WP8 team parsed the texts with an emphasis on preserving their original structure wherever possible. This included maintaining distinctions between headers, paragraphs, and verse lines, which are integral to the organization and meaning of classical Arabic texts. To prepare the texts for embedding, we implemented a chunking strategy that imposed both minimum (50) and maximum (250) length limits for string segments, balancing computational efficiency with semantic coherence. Rather than splitting arbitrarily, we aimed to preserve logical units of meaning—such as paragraph boundaries or complete verses—within each chunk. This approach ensured that the embeddings captured contextually rich and meaningful representations of the text, while respecting the stylistic and structural features of the original corpus. To maximize the usefulness and interpretability of each embedded chunk, we augmented every chunk with comprehensive metadata drawn from the corpus and file structure. This includes attributes such as author name, work title, estimated date, source collection, and original file path. When available, we also incorporated classification information, such as literary genre or subject, derived from the OpenITI metadata files (https://openiti.org/). This enriched metadata enables both users and systems to filter, organize, and contextualize search results, going beyond the semantic content of the text alone. In addition, we appended the text of the immediately preceding and following chunks to each main chunk as auxiliary fields. While these surrounding texts were not used in the semantic embedding itself, they provide browsing and interpretive context—helping users quickly understand how and where a chunk fits within the broader structure of the work.
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 32 FREE QUERIES (LATIN)
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 33 FREE QUERIES (GREEK) RETRIEVED RESULTS AND INTERFACE
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 34
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 35 7.2. Arabic The evaluation of the Arabic prototype tool by three domain experts yielded a range of constructive insights, revealing both the scholarly potential of the tool and several key areas requiring refinement. Despite the small size of the test group, the responses were detailed and substantive, offering diverse perspectives on the tool’s semantic capabilities, interface design, and research utility within the fields of Arabic and Islamic studies. All three evaluators acknowledged the value of the tool’s semantic search capabilities. The tool was consistently praised for retrieving not only exact matches but also semantically relevant paraphrases and reuses. Even when no literal overlap was found, the level of semantic sensitivity of the engine often produced contextually meaningful results, offering valuable support for research in Islamic textual traditions and textual reuse across different literary genres. When testing the three preset queries, participants largely rated the results as “mostly Relevant” or “very Relevant.” However, concerns were raised regarding the clarity and completeness of metadata, including missing author information in some cases, inconsistencies in date presentation, and occasional duplication or truncation of excerpts. These issues are already being addressed through ongoing efforts to improve the metadata structure, normalize entries, and enhance bio-bibliographic accuracy across the corpora. Another recurring observation concerned the ordering of results. While some testers expected canonical sources (e.g., Bukhārī, Muslim) to appear first, it is important to note that the current engine intentionally avoids privileging “canonical” over “non-canonical” texts. In future iterations, users will be able to customize ranking criteria—e.g., by semantic similarity, chronology, or source type—without predefining what constitutes “canonical.” This design choice is intended to challenge traditional biases and encourage the discovery of new textual associations. Upcoming updates will also introduce additional filtering
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 36 options, allowing users to prioritize canonical sources if desired, without enforcing such hierarchies by default. In their custom queries, testers explored a wide range of inputs, from well-known ʾaḥādīṯ to key terms from the Islamic religious sphere, as well as bibliographic metadata. While narrative or conceptual queries produced strong results, queries such as book titles (e.g., Rasāʾil Ibn Rushd al-Ṭibbīyyah) or acronyms like “ISBN” in Arabic yielded few or no relevant matches. This is to be expected, as the tool is not currently designed to retrieve modern catalog metadata or structured identifiers. These results highlight the need to clarify the intended scope of the tool and the kinds of queries it is optimized to handle. As for metadata, testers found it to be often overwhelming, unclear, incomplete, or overly technical—an issue stemming from its origins in the OpenITI corpus and other legacy data models. In response, ongoing efforts are focused on refining the metadata presentation by reducing visual clutter, standardizing field order, and distinguishing between essential bibliographic metadata and advanced technical tags. Suggestions, such as adding both Hijrī and Gregorian dates, or enabling corpus-specific filtering (e.g., Shamela), are actively being considered. The tool’s interface was described as generally clear and usable, though still at a prototypical stage. Each participant rated the tool’s intuitiveness as a 3 out of 4. Feedback emphasized the importance of normalizing the layout of filters, improving their naming conventions, and providing structured input fields (e.g., for author, title, source). The interface presented during testing was only provisional, and it has already been completely redesigned to improve the user experience. A guided query builder and expanded search options have been introduced to better accommodate different research workflows. The testers also indicated that training resources and documentation would be especially helpful for new users—an area we are already addressing through the development of comprehensive training materials that cover theoretical, methodological, visual, and technical dimensions. Despite its prototypical status, the tool was regarded as highly promising for advancing Arabic textual studies. All three scholars recognized its capacity to accelerate research, identify meaningful textual connections, and support broader digital humanities workflows. This evaluation thus validated both the current achievements and the roadmap for future improvements of the uBIQUity tool for Arabic texts. Based on the results of the questionnaires, the following charts and diagrams have been created to visually represent the feedback received, highlighting both the backgrounds of the consulted experts and the functionalities of the tool. Meanwhile, the full responses of the three domain experts in the field of Arabic-Islamic studies are provided in the Appendix below.
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 37
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 38 TEAM-SUGGESTED AND FREE QUERIES
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 39
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 40 RETRIEVED RESULTS AND INTERFACE
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 41
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 48 4. numbers 23:14 [rahlfs] καὶ παρέλαβεν αὐτὸν εἰς ἀγροῦ σκοπιὰν ἐπὶ κορυφὴν λελαξευμένου καὶ ᾠκοδόμησεν ἐκεῖ ἑπτὰ βωμοὺς καὶ ἀνεβίβασεν μόσχον καὶ κριὸν ἐπὶ τὸν βωμόν. Scores: 2.363, 0.175 (normalized), 5.165 (standardized) 5. numbers 23:14 [gottingen] καὶ παρέλαβεν αὐτὸν εἰς ἀγροῦ σκοπιὰν ἐπὶ κορυφὴν λελαξευμένου, καὶ ᾠκοδόμησεν ἐκεῖ ἑπτὰ βωμούς, καὶ ἀνεβίβασεν μόσχον καὶ κριὸν ἐπὶ τὸν βωμόν. Scores: 2.363, 0.175 (normalized), 5.165 (standardized) 6. psalms 16:4 [rahlfs] ἔνυξέν με ὡς κέντρον ἵππου ἐπὶ τὴν γρηγόρησιν αὐτοῦ, ὁ σωτὴρ καὶ ἀντιλήπτωρ μου ἐν παντὶ καιρῷ ἔσωσέν με. Scores: 2.305, 0.170 (normalized), 5.002 (standardized) 7.maccabeorum-iii 12:12 [gottingen] Ἰούδας δὲ ὑπολαβὼν ὡς ἀληθῶς ἐν πολλοῖς αὐτοὺς χρησίμους ἐπεχώρησεν εἰρήνην ἄξειν πρὸς αὐτούς· καὶ λαβόντες δεξιὰς εἰς τὰς σκηνὰς ἐχωρίσθησαν. Scores: 2.175, 0.161 (normalized), 4.635 (standardized) Q: Please rate the overall quality of the results of Query 1 (GREEK) A: 3 Q: Please rate the overall quality of the results of Query 2 (GREEK) A: 4 Q: LATIN. Please perform two queries of your own choice and then search for one specific word. Please annotate here your queries together with the first seven results returned by the model for each query: A: QUERY 1: Aug. De Gen. ad litt. 12.14.30 :ea ipsa quae in disco demonstrabantur, tamquam vera animalia 1.wisdom 19:18 [weber] agrestia enim in aquatica convertebantur et quaecumque erant natantia in terram transiebant Scores: 2.199, 0.971 (normalized), 4.783 (standardized) 2.Wis 19:18 [sabatier-vetus-latina] Agrestia enim in aquatica convertebantur, et quaecumque erant natantia, in terram transibant, Scores: 2.182, 0.963 (normalized), 4.731 (standardized)
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 49 3.Ezek 1:14 [sabatier-vetus-latina] Et animalia currebant et revertebantur quasi species bezec. Scores: 2.149, 0.948 (normalized), 4.634 (standardized) 4. romans 13:9 [sabatier-versio-antiqua] Etenim: Non adulterabis: Non occides: Non furaberis: Non concupisces: et si quid est aliud mandatum, in hoc verbo* instauratur: Diliges proximum tuum tanquam teipsum. Scores: 2.134, 0.941 (normalized), 4.589 (standardized) 5.esther 2:9 [weber] quae placuit ei et invenit gratiam in conspectu illius ut adceleraret mundum muliebrem et traderet ei partes suas et septem puellas speciosissimas de domo regis et tam ipsam quam pedisequas eius ornaret atque excoleret Scores: 2.128, 0.939 (normalized), 4.573 (standardized) 6. romans 13:9 [weber] nam non adulterabis non occides non furaberis non concupisces et si quod est aliud mandatum in hoc verbo instauratur diliges proximum tuum tamquam te ipsum Scores: 2.108, 0.929 (normalized), 4.512 (standardized) 7. deuteronomy 4:9 [weber] custodi igitur temet ipsum et animam tuam sollicite ne obliviscaris verborum quae viderunt oculi tui et ne excedant de corde tuo cunctis diebus vitae tuae docebis ea filios ac nepotes tuos Scores: 2.096, 0.924 (normalized), 4.478 (standardized) QUERY 2: Aug. De Gen. ad litt. 4.9.16: donum Spiritus sancti, per quem diffunditur caritas in cordibus nostris. 1.romans 5:5 [weber] spes autem non confundit quia caritas Dei diffusa est in cordibus nostris per Spiritum Sanctum qui datus est nobis Scores: 3.355, 1.000 (normalized), 6.129 (standardized) 2.romans 5:5 [sabatier-versio-antiqua] spes autem non confundit: quia caritas Dei diffusa est in cordibus nostris per Spiritum sanctum, qui datus est nobis. Scores: 3.349, 0.998 (normalized), 6.115 (standardized) 3.corinthias-2 1:22 [weber] et qui signavit nos et dedit pignus Spiritus in cordibus nostris Scores: 3.013, 0.896 (normalized), 5.323 (standardized)
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 50 4.corinthians-2 1:22 [sabatier-versio-antiqua] et qui signavit nos, et dedit pignus Spiritus in cordibus nostris. Scores: 3.011, 0.896 (normalized), 5.318 (standardized) 5.timothy-2 1:14 [sabatier-versio-antiqua] bonum depositum custodi per Spiritum sanctum, qui habitat in nobis. Scores: 2.857, 0.849 (normalized), 4.958 (standardized) 6.timothy-2 1:14 [weber] bonum depositum custodi per Spiritum Sanctum qui habitat in nobis Scores: 2.851, 0.847 (normalized), 4.943 (standardized) 7.ezekiel 48:8 [weber] et super terminum Iuda a plaga orientali usque ad plagam maris erunt primitiae quas separabitis viginti quinque milibus latitudinis et longitudinis sicuti singulae partes a plaga orientali usque ad plagam maris et erit sanctuarium in medio eius Scores: 2.748, 0.816 (normalized), 4.700 (standardized) QUERY 3 (SINGLE WORD): propitiatorium 1. Ex 25:20 [sabatier-vetus-latina] extendentia alas suas, et obumbrantia super propitiatorium, et facies contra se super propitiatorium, Scores: 1.193, 1.000 (normalized), 19.334 (standardized) 2.exodus 39:34 [weber] velum arcam vectes propitiatorium Scores: 1.183, 0.990 (normalized), 19.113 (standardized) 3.Num 14:20 [sabatier-vetus-latina] ... Propitius ero illis. Scores: 1.079, 0.885 (normalized), 16.786 (standardized) 4.Psa 24:11 [sabatier-vetus-latina] Propter nomen tuum Domine et propitiaberis peccato meo, copiosum est enim. Scores: 1.067, 0.873 (normalized), 16.518 (standardized) 5. Sir 35:3 [sabatier-vetus-latina] Et propitiationem litare sacrificii super iniustitias: et deprecatio pro peccatis, recedere ab iniustitia. Scores: 1.062, 0.868 (normalized), 16.412 (standardized)
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 51 6.psalms 24:11 [weber] propter nomen tuum Domine et propitiaberis peccato meo multum est enim Scores: 1.053, 0.859 (normalized), 16.216 (standardized) 7.psalms-iuxtra-hebraicum 24:11 [weber] propter nomen tuum propitiare iniquitati meae quoniam grandis est Scores: 1.053, 0.859 (normalized), 16.215 (standardized) Q: Please rate the overall quality of the results of your first query A: 3.0 Q: Please rate the overall quality of the results of your second query A: 4.0 Q: Please rate the overall quality of the results of your third query A: 3.0 Q: GREEK. Please perform two queries of your own choice and then search for one specific word. Please annotate here your queries together with the first seven results returned by the model for each query: A: QUERY 1. Clem. Paed. 1.5.13.2: ὁ κύριος ἐν τῷ εὐαγγελίῳ μυωπίζει τοὺς γνωρίμους, προσέχειν αὐτῷ παρορμῶν ὡς ἤδη σπεύδων πρὸς τὸν πατέρα. 1. deuteronomy 15:9 [gottingen] πρόσεχε σεαυτῷ, μὴ γένηται ῥῆμα κρυπτὸν ἐν τῇ καρδίᾳ σου, ἀνόμημα, λέγων Ἐγγίζει τὸ ἔτος τὸ ἕβδομον, ἔτος τῆς ἀφέσεως, καὶ πονηρεύσηται ὁ ὀφθαλμός σου τῷ ἀδελφῷ σου τῷ ἐπιδεομένῳ, καὶ οὐ δώσεις αὐτῷ, καὶ βοήσεται κατὰ σοῦ πρὸς κύριον, καὶ ἔσται ἐν σοὶ ἁμαρτία μεγάλη. Variants: [R_LXX] πρόσεχε σεαυτῷ μὴ γένηται ῥῆμα κρυπτὸν ἐν τῇ καρδίᾳ σου, ἀνόμημα, λέγων Ἐγγίζει τὸ ἔτος τὸ ἕβδομον, ἔτος τῆς ἀφέσεως, καὶ πονηρεύσηται ὁ ὀφθαλμός σου τῷ ἀδελφῷ σου τῷ ἐπιδεομένῳ, καὶ οὐ δώσεις αὐτῷ, καὶ βοήσεται κατὰ σοῦ πρὸς κύριον, καὶ ἔσται ἐν σοὶ ἁμαρτία μεγάλη. [G_α] πρόσεχε σεαυτῷ, μὴ γένηται ῥῆμα κρυπτὸν ἐν τῇ καρδίᾳ σου, ἀποστασίας τῷ λέγειν Ἐγγίζει τὸ ἔτος τὸ ἕβδομον, ἔτος τῆς ἀφέσεως, καὶ πονηρεύσηται ὁ ὀφθαλμός σου τῷ ἀδελφῷ σου τῷ ἐπιδεομένῳ, καὶ οὐ δώσεις αὐτῷ, καὶ βοήσεται κατὰ σοῦ πρὸς κύριον, καὶ ἔσται ἐν σοὶ ἁμαρτία μεγάλη. Scores: 3.239, 0.979 (normalized), 5.115 (standardized) 2.kings-1 25:23 [rahlfs] καὶ εἶδεν Αβιγαια τὸν Δαυιδ καὶ ἔσπευσεν καὶ κατεπήδησεν ἀπὸ τῆς ὄνου καὶ ἔπεσεν
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 52 ἐνώπιον Δαυιδ ἐπὶ πρόσωπον αὐτῆς καὶ προσεκύνησεν αὐτῷ ἐπὶ τὴν γῆν Scores: 3.173, 0.959 (normalized), 4.973 (standardized) 3. judges-vaticanus 20:23 [rahlfs] καὶ ἀνέβησαν οἱ υἱοὶ Ισραηλ καὶ ἔκλαυσαν ἐνώπιον κυρίου ἕως ἑσπέρας καὶ ἠρώτησαν ἐν κυρίῳ λέγοντες Εἰ προσθῶμεν ἐγγίσαι εἰς παράταξιν πρὸς υἱοὺς Βενιαμιν ἀδελφοὺς ἡμῶν; καὶ εἶπεν κύριος Ἀνάβητε πρὸς αὐτούς. Scores: 3.159, 0.955 (normalized), 4.942 (standardized) 4.kings-4 6:32 [rahlfs] καὶ Ελισαιε ἐκάθητο ἐν τῷ οἴκῳ αὐτοῦ, καὶ οἱ πρεσβύτεροι ἐκάθηντο μετ' αὐτοῦ. καὶ ἀπέστειλεν ἄνδρα πρὸ προσώπου αὐτοῦ· πρὶν ἐλθεῖν τὸν ἄγγελον πρὸς αὐτὸν καὶ αὐτὸς εἶπεν πρὸς τοὺς πρεσβυτέρους Εἰ οἴδατε ὅτι ἀπέστειλεν ὁ υἱὸς τοῦ φονευτοῦ οὗτος ἀφελεῖν τὴν κεφαλήν μου; ἴδετε ὡς ἂν ἔλθῃ ὁ ἄγγελος, ἀποκλείσατε τὴν θύραν καὶ παραθλίψατε αὐτὸν ἐν τῇ θύρᾳ· οὐχὶ φωνὴ τῶν ποδῶν τοῦ κυρίου αὐτοῦ κατόπισθεν αὐτοῦ; Scores: 3.021, 0.912 (normalized), 4.646 (standardized) 5.kings-1 2:36 [rahlfs] καὶ ἔσται ὁ περισσεύων ἐν οἴκῳ σου ἥξει προσκυνεῖν αὐτῷ ὀβολοῦ ἀργυρίου λέγων Παράρριψόν με ἐπὶ μίαν τῶν ἱερατειῶν σου φαγεῖν ἄρτον. Καὶ τὸ παιδάριον Σαμουηλ ἦν λειτουργῶν τῷ κυρίῳ ἐνώπιον Ηλι τοῦ ἱερέως· καὶ ῥῆμα κυρίου ἦν τίμιον ἐν ταῖς ἡμέραις ἐκείναις, οὐκ ἦν ὅρασις διαστέλλουσα. Scores: 2.925, 0.882 (normalized), 4.441 (standardized) 6. judges-alexandrinus 20:23 [rahlfs] καὶ ἀνέβησαν οἱ υἱοὶ Ισραηλ καὶ ἔκλαυσαν ἐνώπιον κυρίου ἕως ἑσπέρας καὶ ἐπηρώτησαν ἐν κυρίῳ λέγοντες Εἰ προσθῶ προσεγγίσαι εἰς πόλεμον μετὰ Βενιαμιν τοῦ ἀδελφοῦ μου; καὶ εἶπεν κύριος Ἀνάβητε πρὸς αὐτόν. Scores: 2.894, 0.873 (normalized), 4.373 (standardized) 7.daniel-theodotionis 2:9 [rahlfs] ἐὰν οὖν τὸ ἐνύπνιον μὴ ἀναγγείλητέ μοι, οἶδα ὅτι ῥῆμα ψευδὲς καὶ διεφθαρμένον συνέθεσθε εἰπεῖν ἐνώπιόν μου, ἕως οὗ ὁ καιρὸς παρέλθῃ· τὸ ἐνύπνιόν μου εἴπατέ μοι, καὶ γνώσομαι ὅτι τὴν σύγκρισιν αὐτοῦ ἀναγγελεῖτέ μοι. Scores: 2.892, 0.872 (normalized), 4.369 (standardized) QUERY 2. Clem. Paed. 1.8.62.1: καὶ ὁ φοβούμενος κύριον ἐπιστρέφει ἐπὶ καρδίαν αὐτοῦ 1.sirach 21:6 [rahlfs] μισῶν ἐλεγμὸν ἐν ἴχνει ἁμαρτωλοῦ, καὶ ὁ φοβούμενος κύριον ἐπιστρέψει ἐν καρδίᾳ. Scores: 3.957, 1.000 (normalized), 11.414 (standardized)
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 53 2.sirach 34:14 [rahlfs] ὁ φοβούμενος κύριον οὐδὲν εὐλαβηθήσεται καὶ οὐ μὴ δειλιάσῃ, ὅτι αὐτὸς ἐλπὶς αὐτοῦ. Scores: 3.304, 0.834 (normalized), 9.261 (standardized) 3.sirach 31:16 [gottingen] ὁ φοβούμενος κύριον οὐδὲν εὐλαβηθήσεται καὶ οὐ μὴ δειλιάσει, ὅτι αὐτὸς ἐλπὶς αὐτοῦ. Scores: 3.298, 0.832 (normalized), 9.240 (standardized) 4.sirach 21:6 [gottingen] μισῶν ἐλεγμὸν ἐν ἴχνει ἁμαρτωλοῦ, καὶ ὁ φοβούμενος κύριον ἐπιστρέψει ἐν καρδίᾳ. Scores: 3.133, 0.790 (normalized), 8.696 (standardized) 5.sirach 15:1 [rahlfs] Ὁ φοβούμενος κύριον ποιήσει αὐτό, καὶ ὁ ἐγκρατὴς τοῦ νόμου καταλήμψεται αὐτήν· Scores: 3.019, 0.761 (normalized), 8.320 (standardized) 6.sirach 15:1 [gottingen] Ὁ φοβούμενος κύριον ποιήσει αὐτό, καὶ ὁ ἐγκρατὴς τοῦ νόμου καταλήμψεται αὐτήν· Scores: 3.019, 0.761 (normalized), 8.320 (standardized) 7.sirach 6:17 [rahlfs] ὁ φοβούμενος κύριον εὐθυνεῖ φιλίαν αὐτοῦ, ὅτι κατ' αὐτὸν οὕτως καὶ ὁ πλησίον αὐτοῦ. Scores: 2.944, 0.742 (normalized), 8.071 (standardized) QUERY 3 (SINGLE WORD): βδέλυγμα 1. proverbs 15:26 [rahlfs] βδέλυγμα κυρίῳ λογισμὸς ἄδικος, ἁγνῶν δὲ ῥήσεις σεμναί. Scores: 1.650, 1.000 (normalized), 13.416 (standardized) 2. proverbs 27:20a [rahlfs] βδέλυγμα κυρίῳ στηρίζων ὀφθαλμόν, καὶ οἱ ἀπαίδευτοι ἀκρατεῖς γλώσσῃ. Scores: 1.631, 0.986 (normalized), 13.190 (standardized) 3. sirach 15:13 [gottingen] πᾶν βδέλυγμα ἐμίσησεν κύριος, καὶ οὐκ ἔστιν ἀγαπητὸν τοῖς φοβουμένοις αὐτόν. Scores: 1.587, 0.956 (normalized), 12.688 (standardized) 4. sirach 15:13 [rahlfs] πᾶν βδέλυγμα ἐμίσησεν ὁ κύριος, καὶ οὐκ ἔστιν ἀγαπητὸν τοῖς φοβουμένοις αὐτόν. Scores: 1.569, 0.943 (normalized), 12.478 (standardized) 5. genesis 4:10 [gottingen]
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 54 καὶ εἶπεν ὁ θεός Τί ἐποίησας; φωνὴ αἵματος τοῦ ἀδελφοῦ σου βοᾷ πρὸς με ἐκ τῆς γῆς. Variants: [R_LXX] καὶ εἶπεν ὁ θεός Τί ἐποίησας; φωνὴ αἵματος τοῦ ἀδελφοῦ σου βοᾷ πρός με ἐκ τῆς γῆς. Scores: 1.086, 0.604 (normalized), 6.934 (standardized) 6.sirach 13:20 [rahlfs] βδέλυγμα ὑπερηφάνῳ ταπεινότης· οὕτως βδέλυγμα πλουσίῳ πτωχός. Scores: 1.061, 0.587 (normalized), 6.655 (standardized) 7.sirach 13:20 [gottingen] βδέλυγμα ὑπερηφάνῳ ταπεινότης· οὕτως βδέλυγμα πλουσίῳ πτωχός. Scores: 1.061, 0.587 (normalized), 6.655 (standardized) Q: Please rate the overall quality of the results of your first query .1 A: 2.0 Q: Please rate the overall quality of the results of your second query .1 A: 4.0 Q: Please rate the overall quality of the results of your third query .1 A: 4.0 Q: Did the model return semantically relevant results, even if not literal matches? Did it do so among the top 5, 10 or 15 results? A: Yes Q: Did you find the results meaningful in your area of expertise? A: Yes Q: Did you notice any errors or inconsistencies in the suggested results? If yes, please describe them. A: Allusions seem to be more hard to capture than literal or quasi-literal quotations . Moreover, in one case, even if my query included an explicit reference to the Gospel (ἐν τῷ εὐαγγελίῳ), the search engine's results referred to the Old Testament. Q: How well do you think the model captures contextual meaning in Latin and Greek? A: Pretty well
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 55 Q: Do you think the amount of surrounding context provided is sufficient to understand the reuse? A: Yes Q: Did you find the metadata below each result useful and clear? A: I must confess I had difficulties in understanding the meaning of metadata Q: Did you encounter any difficulties interpreting the results? A: No Q: Please describe your general impression of the tool. A: The tool is very useful and promising. It just needs to be improved and refined Q: Would you suggest adding other types of metadata? If yes, which ones? A: I am not a fan of metadata Q: In which ways do you think this model could enrich your research work? A: This model can provide significant textual parallels, creating unexpected connections between famous and less famous texts on the basis of their semantic relationships. Q: What improvements or additions would you suggest to make the tool more practical or useful for your specific research context? A: Clear instructions should be given in order to allow users to adjust research criteria (text, semantic et sim.) to their specific needs. For instance, in some cases semantic connections can be far more important than trigrams - and users should be instructed to disable the latter criterion. Q: Which improvements would you propose to extend the scope of the model and scale it to your colleagues' needs? A: The scope of the model should be extended by uploading additional texts - ultimately including non-religious sources. Q: Do you think specific training should be offered along with the tool? A: Yes Q: Please feel free to add any further observations, comments, or suggestions: A: Fingers crossed for the next steps of this ambitious work!
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 56 Interviewee 2 Q: What digital tools do you currently use to analyze the Greek and Latin sources you work on? A: TLG, TLL, Bibleworks, DBG Q: What are the main features you typically look for when working with digital tools for text analysis and text reuse? A: Searching for single lemmas or phrases; parallel analysis of Hebrew, Greek and Latin texts of the Bible(s); searching for words or expressions in large corpora (Greek and Latin literature; LXX; Latin Bibles); analysis of textual recurrences; study of rewritings and ways of reproposing biblical texts in subsequent ancient literature Q: Are you familiar with text reuse tools (such as TRACER)? A: Scarcely familiar Q: What is your level of familiarity with AI-based technologies in the humanities? A: Beginner Q: Are you familiar with transformer-based models such as BERT? A: Scarcely familiar Q: LATIN. Please perform the two queries suggested below. Please annotate here the first seven results returned by the model for each query: Query 1. Aug. De Gen. ad litt. 7.25: ratio reddenda est in iudicio Dei, recepturo unoquoque secundum ea, quae per corpus gessit, sive bonum sive malum. Query 2. Aug. De Gen. ad litt. 12.2.5: ara, unde carbo assumptus Prophetae labia mundavit. A: The first search has an effectiveness of 60% compared to the reference texts identified. The search tool identifies a series of textual links to biblical texts that revolve around the relationship between people and judgment, with the identification, in 6 out of 10 cases, of particularly relevant passages. The second search proved completely inadequate. I believe that the use of the term carbo misled the tool, which linked it to names of precious stones or, more generally, to natural elements. I did not find any relevant references in the first 10 positions.
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 57 Q: Please rate the overall quality of the results of Query 1 (LATIN) A: 3 Q: Please rate the overall quality of the results of Query 2 (LATIN) A: 1 Q: GREEK. Please perform the two queries suggested below. Please annotate here the first seven results returned by the model for each query: Query 1. Clem. Paed. 1.6.27: ὡς τὸ θέλημα αὐτοῦ ἔργον ἐστὶ καὶ τοῦτο κόσμος ὀνομάζεται Query 2. Clem. Paed. 1.7.56: ἐκύκλωσεν αὐτὸν καὶ ἐπαίδευσεν αὐτὸν καὶ διεφύλαξεν ὡς κόρην ὀφθαλμοῦ A: The first search showed remarkable effectiveness in 5 out of 10 cases, all revolving around the term thelema. The other references identified are quite relevant. The second search proved very effective in 3 out of 10 cases, as it is an almost literal quotation from Deut 32:10. The other references identified do not appear to be relevant. Q: Please rate the overall quality of the results of Query 1 (GREEK) A: 3 Q: Please rate the overall quality of the results of Query 2 (GREEK) A: 2 Q: LATIN. Please perform two queries of your own choice and then search for one specific word. Please annotate here your queries together with the first seven results returned by the model for each query: A: Gen. Litt. 1.5.10 Non enim habet informem vitam Verbum Filius, cui non solum hoc est esse quod vivere, sed etiam hoc est vivere, quod est sapienter ac beate vivere. Gen. Litt. 1.8.14 Ita etiam rebus ex illa inchoatione perfectis atque formatis, vidit Deus quia bonum est inchoatione Q: Please rate the overall quality of the results of your first query A: 2.0
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 64 Q: Please feel free to add any further observations, comments, or suggestions: A: nan Interviewee 4 Q: What digital tools do you currently use to analyze the Greek and Latin sources you work on? A: Diogenes; Phi corpus; TLG Q: What are the main features you typically look for when working with digital tools for text analysis and text reuse? A: allusivity; quotationa Q: Are you familiar with text reuse tools (such as TRACER)? A: No Q: What is your level of familiarity with AI-based technologies in the humanities? A: Intermediate Q: Are you familiar with transformer-based models such as BERT? A: No Q: LATIN. Please perform the two queries suggested below. Please annotate here the first seven results returned by the model for each query: Query 1. Aug. De Gen. ad litt. 7.25: ratio reddenda est in iudicio Dei, recepturo unoquoque secundum ea, quae per corpus gessit, sive bonum sive malum. Query 2. Aug. De Gen. ad litt. 12.2.5: ara, unde carbo assumptus Prophetae labia mundavit. A: Query 1. Eccl 12:14; Tob 14:6; Tob 3:16; 2Mac 3:6; Lev 12:2; 1Mac 1:57; Psa 21:27 Query 2: esdrae-4 8:52; judges 1:11; revelation 21:19; Is 54:11; job 25:4; maccabees-1 6:33; 1Mac 2:23 Q: Please rate the overall quality of the results of Query 1 (LATIN) A: 4
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 65 Q: Please rate the overall quality of the results of Query 2 (LATIN) A: 3 Q: GREEK. Please perform the two queries suggested below. Please annotate here the first seven results returned by the model for each query: Query 1. Clem. Paed. 1.6.27: ὡς τὸ θέλημα αὐτοῦ ἔργον ἐστὶ καὶ τοῦτο κόσμος ὀνομάζεται Query 2. Clem. Paed. 1.7.56: ἐκύκλωσεν αὐτὸν καὶ ἐπαίδευσεν αὐτὸν καὶ διεφύλαξεν ὡς κόρην ὀφθαλμοῦ A: Query 1: daniel-theodotionis 4:35; sirach 35:17; sirach 32:17; psalms 101:21; deuteronomy 15:9; daniel-theodotionis 11:3; esdras-a 9:9 Query 2: deuteronomy 32:10; odes 2:10; deuteronomy 32:10; numbers 23:14; numbers 23:14; psalms 16:4; maccabeorum-iii 12:12 Q: Please rate the overall quality of the results of Query 1 (GREEK) A: 4 Q: Please rate the overall quality of the results of Query 2 (GREEK) A: 4 Q: LATIN. Please perform two queries of your own choice and then search for one specific word. Please annotate here your queries together with the first seven results returned by the model for each query: A: Query 1: Laudate Dominum de terra, dracones et omnes abyssi Psa 148:7; psalmsiuxtra-hebraicum 148:7; Psa 148:1; psalms-iuxtra-hebraicum 148:1; psalms-iuxtrahebraicum 116:1; daniel 3:60; jeremiah 51:48 Query 2: alias capite albo sicut lana, alias inferiore parte corporis sicut aurichalcum: song-of-solomon 7:5; chronicles-2 21:19; Esth 2:7; Sir 34:6; ephesians 6:5; esdrae-2 9:24; leviticus 13:10 Query 3: multiplicamini: Gen 9:7; genesis 1:22; leviticus 26:9; isaiah 8:9; zachariah 10:8; psalms 138:18; Lev 26:9 Q: Please rate the overall quality of the results of your first query A: 3.0 Q: Please rate the overall quality of the results of your second query A: 2.0 Q: Please rate the overall quality of the results of your third query
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 66 A: 4.0 Q: GREEK. Please perform two queries of your own choice and then search for one specific word. Please annotate here your queries together with the first seven results returned by the model for each query: A: nan Q: Please rate the overall quality of the results of your first query .1 A: nan Q: Please rate the overall quality of the results of your second query .1 A: nan Q: Please rate the overall quality of the results of your third query .1 A: nan Q: Did the model return semantically relevant results, even if not literal matches? Did it do so among the top 5, 10 or 15 results? A: Yes, not always in the same way Q: Did you find the results meaningful in your area of expertise? A: Yes Q: Did you notice any errors or inconsistencies in the suggested results? If yes, please describe them. A: Some passages have only a vague resemblance to the biblical source Q: How well do you think the model captures contextual meaning in Latin and Greek? A: Sufficiently well Q: Do you think the amount of surrounding context provided is sufficient to understand the reuse? A: yes Q: Did you find the metadata below each result useful and clear? A: yes, the indication of the specific biblical version is useful, but other metadata should be added
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 67 Q: Did you encounter any difficulties interpreting the results? A: No Q: Please describe your general impression of the tool. A: The tool is useful and sufficiently balanced, but needs to be improved in order to capture further allusive reuses Q: Would you suggest adding other types of metadata? If yes, which ones? A: The date and place of origin of all sources could be indicated Q: In which ways do you think this model could enrich your research work? A: The model could support the investigation of quotations and allusions in a wide range of late antique texts Q: What improvements or additions would you suggest to make the tool more practical or useful for your specific research context? A: More information should be provided regarding the context and the origin of each source Q: Which improvements would you propose to extend the scope of the model and scale it to your colleagues' needs? A: Other texts of different kinds could be included among the sources available on the tool Q: Do you think specific training should be offered along with the tool? A: Yes Q: Please feel free to add any further observations, comments, or suggestions: A: Nothing to add Interviewee 5 Q: What digital tools do you currently use to analyze the Greek and Latin sources you work on? A: TLG, PHI Corpus, Corpus Corporum Q: What are the main features you typically look for when working with digital tools for text analysis and text reuse?
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 68 A: Advanced research also with co-occurrence Q: Are you familiar with text reuse tools (such as TRACER)? A: No, I am not Q: What is your level of familiarity with AI-based technologies in the humanities? A: Beginner Q: Are you familiar with transformer-based models such as BERT? A: No Q: LATIN. Please perform the two queries suggested below. Please annotate here the first seven results returned by the model for each query: Query 1. Aug. De Gen. ad litt. 7.25: ratio reddenda est in iudicio Dei, recepturo unoquoque secundum ea, quae per corpus gessit, sive bonum sive malum. Query 2. Aug. De Gen. ad litt. 12.2.5: ara, unde carbo assumptus Prophetae labia mundavit. A: Query 1: corinthias-2 5:10 [weber]; corinthians-2 5:10 [sabatier-versio-antiqua], chronicles-1 28:8 [weber], Eccl 12:14 [sabatier-vetus-latina], ecclesiastes 12:14 [weber], ezekiel 31:12 [weber], corinthias-2 1:6 [weber] Query 2: judges 1:11 [weber]; revelation 21:19 [weber]; Is 54:11 [sabatier-vetus-latina]; job 25:4 [weber]; maccabees-1 6:33 [weber]; judges 1:11 [weber], esdrae-4 8:52 [weber], esdrae-4-s 8:52 [weber] Q: Please rate the overall quality of the results of Query 1 (LATIN) A: 4 Q: Please rate the overall quality of the results of Query 2 (LATIN) A: 2 Q: GREEK. Please perform the two queries suggested below. Please annotate here the first seven results returned by the model for each query: Query 1. Clem. Paed. 1.6.27: ὡς τὸ θέλημα αὐτοῦ ἔργον ἐστὶ καὶ τοῦτο κόσμος ὀνομάζεται
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 69 Query 2. Clem. Paed. 1.7.56: ἐκύκλωσεν αὐτὸν καὶ ἐπαίδευσεν αὐτὸν καὶ διεφύλαξεν ὡς κόρην ὀφθαλμοῦ A: Query 1: daniel-theodotionis 4:35 [rahlfs]; sirach 35:17 [gottingen], psalms 102:21 [rahlfs],psalms 101:21 [gottingen], daniel-theodotionis 11:3 [rahlfs], deuteronomy 15:9 [gottingen] Query 2: deuteronomy 32:10 [gottingen], odes 2:10 [rahlfs], deuteronomy 32:10 [rahlfs], numbers 23:14 [rahlfs], numbers 23:14 [gottingen], psalms 16:4 [rahlfs], maccabeorum-iii 12:12 [gottingen] Q: Please rate the overall quality of the results of Query 1 (GREEK) A: 4 Q: Please rate the overall quality of the results of Query 2 (GREEK) A: 4 Q: LATIN. Please perform two queries of your own choice and then search for one specific word. Please annotate here your queries together with the first seven results returned by the model for each query: A: Query 1 (Aug. De anima et eius orig. 1.4.4): in hominis faciem sufflavit, eique illo modo animam fecit; Results: Gen 2:7 [sabatier-vetus-latina]; Judith 2:7 [sabatier-vetus-latina]; genesis 2:7 [weber], isaiah 9:7 [weber], matthew 26:29 [weber], genesis 19:28 [weber], Tob 5:2 [sabatier-vetus-latina] Query 2 (Aug. De anima et eius orig. 1.8.9): alioquin gratia iam non est gratia Results: romans 11:6 [weber]; romans 11:6 [sabatier-versio-antiqua], ephesians 2:8 [weber], ephesians 2:8 [sabatier-versio-antiqua], galatians 2:21 [weber], galatians 2:21 [sabatier-versio-antiqua], romans 4:4 [weber] Query 3: perfusum Results: leviticus 8:30 [weber], ezekiel 24:7 [weber], Ex 4:9 [sabatier-vetus-latina], revelation 16:4 [weber], leviticus 8:12 [weber], 1Mac 1:39 [sabatier-vetus-latina], hosea 10:7 [weber] Q: Please rate the overall quality of the results of your first query A: 4.0 Q: Please rate the overall quality of the results of your second query A: 4.0
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 70 Q: Please rate the overall quality of the results of your third query A: 4.0 Q: GREEK. Please perform two queries of your own choice and then search for one specific word. Please annotate here your queries together with the first seven results returned by the model for each query: A: nan Q: Please rate the overall quality of the results of your first query .1 A: nan Q: Please rate the overall quality of the results of your second query .1 A: nan Q: Please rate the overall quality of the results of your third query .1 A: nan Q: Did the model return semantically relevant results, even if not literal matches? Did it do so among the top 5, 10 or 15 results? A: Yes, with the exception of one allusive reference Q: Did you find the results meaningful in your area of expertise? A: Yes Q: Did you notice any errors or inconsistencies in the suggested results? If yes, please describe them. A: The tool is sometimes less efficient in detecting allusive re-uses Q: How well do you think the model captures contextual meaning in Latin and Greek? A: Pretty well Q: Do you think the amount of surrounding context provided is sufficient to understand the reuse? A: Yes, but it would be useful if one could switch to the full context of each quotation Q: Did you find the metadata below each result useful and clear?
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 71 A: The reference to the biblical version is useful. Instead, the metadata following each result are a bit obscure to me Q: Did you encounter any difficulties interpreting the results? A: Not really Q: Please describe your general impression of the tool. A: It is a useful tool, especially in order to retrieve non-literal quotations, allusions, and paraphrases Q: Would you suggest adding other types of metadata? If yes, which ones? A: One could mention the volume and the publication date of each edition of biblical books Q: In which ways do you think this model could enrich your research work? A: By offering parallels and materials for text analysis Q: What improvements or additions would you suggest to make the tool more practical or useful for your specific research context? A: Other metadata could be added; the critical apparatus could be added for all texts Q: Which improvements would you propose to extend the scope of the model and scale it to your colleagues' needs? A: See above Q: Do you think specific training should be offered along with the tool? A: Yes, of course Q: Please feel free to add any further observations, comments, or suggestions: A: Nothing Interviewee 6 Q: What digital tools do you currently use to analyze the Greek and Latin sources you work on? A: TLG, BiblIndex
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 72 Q: What are the main features you typically look for when working with digital tools for text analysis and text reuse? A: Reliable platforms for speed, clarity, and ease of consultation. I prefer those with up-todate and comprehensive source codes. Q: Are you familiar with text reuse tools (such as TRACER)? A: No, I recently started using Tracer (a year ago). Q: What is your level of familiarity with AI-based technologies in the humanities? A: Intermediate Q: Are you familiar with transformer-based models such as BERT? A: No Q: LATIN. Please perform the two queries suggested below. Please annotate here the first seven results returned by the model for each query: Query 1. Aug. De Gen. ad litt. 7.25: ratio reddenda est in iudicio Dei, recepturo unoquoque secundum ea, quae per corpus gessit, sive bonum sive malum. Query 2. Aug. De Gen. ad litt. 12.2.5: ara, unde carbo assumptus Prophetae labia mundavit. A: Query 1: 2Cor 5:21 [nestle-aland-28] τὸν μὴ γνόντα ἁμαρτίαν ὑπὲρ ἡμῶν ἁμαρτίαν ἐποίησεν, ἵνα ἡμεῖς γενώμεθα δικαιοσύνη θεοῦ ἐν αὐτῷ. Scores: 0.806, 1.000 (normalized), 1.054 (standardized) Rom 5:18 [nestle-aland-28] Ἄρα οὖν ὡς δι’ ἑνὸς παραπτώματος εἰς πάντας ἀνθρώπους εἰς κατάκριμα, οὕτως καὶ δι’ ἑνὸς δικαιώματος εἰς πάντας ἀνθρώπους εἰς δικαίωσιν ζωῆς· Scores: 0.803, 0.992 (normalized), 1.038 (standardized) sirach 16:12 [gottingen] κατὰ τὸ πολὺ ἔλεος αὐτοῦ, οὕτως καὶ ὁ ἔλεγχος αὐτοῦ· ἄνδρα κατὰ τὰ ἔργα αὐτοῦ κρινεῖ. Scores: 0.796, 0.977 (normalized), 1.006 (standardized)
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 73 James 2:13 [nestle-aland-28] ἡ γὰρ κρίσις ἀνέλεος τῷ μὴ ποιήσαντι ἔλεος· κατακαυχᾶται ἔλεος κρίσεως. Scores: 0.794, 0.973 (normalized), 0.997 (standardized) Heb 5:1 [nestle-aland-28] Πᾶς γὰρ ἀρχιερεὺς ἐξ ἀνθρώπων λαμβανόμενος ὑπὲρ ἀνθρώπων καθίσταται τὰ πρὸς τὸν θεόν, ἵνα προσφέρῃ δῶρά τε καὶ θυσίας ὑπὲρ ἁμαρτιῶν, Scores: 0.793, 0.970 (normalized), 0.991 (standardized) 1John 3:9 [nestle-aland-28] Πᾶς ὁ γεγεννημένος ἐκ τοῦ θεοῦ ἁμαρτίαν οὐ ποιεῖ, ὅτι σπέρμα αὐτοῦ ἐν αὐτῷ μένει, καὶ οὐ δύναται ἁμαρτάνειν, ὅτι ἐκ τοῦ θεοῦ γεγέννηται. Scores: 0.792, 0.968 (normalized), 0.988 (standardized) sirach 32:17 [rahlfs] ἄνθρωπος ἁμαρτωλὸς ἐκκλινεῖ ἐλεγμὸν καὶ κατὰ τὸ θέλημα αὐτοῦ εὑρήσει σύγκριμα. Scores: 0.792, 0.968 (normalized), 0.987 (standardized) Query 2: kings-3 13:5 [rahlfs] καὶ τὸ θυσιαστήριον ἐρράγη, καὶ ἐξεχύθη ἡ πιότης ἀπὸ τοῦ θυσιαστηρίου κατὰ τὸ τέρας, ὃ ἔδωκεν ὁ ἄνθρωπος τοῦ θεοῦ ἐν λόγῳ κυρίου. Scores: 0.799, 1.000 (normalized), 1.038 (standardized) leviticus 8:11 [gottingen] καὶ ἔρρανεν ἀπʼ αὐτοῦ ἐπὶ τὸ θυσιαστήριον ἑπτάκις, καὶ ἔχρισεν τὸ θυσιαστήριον καὶ ἡγίασεν αὐτό, καὶ πάντα τὰ σκεύη αὐτοῦ καὶ τὸν λουτῆρα καὶ τὴν βάσιν αὐτοῦ, καὶ ἡγίασεν αὐτά· καὶ ἔχρισεν τὴν σκηνὴν καὶ πάντα τὰ ἐν αὐτῇ, καὶ ἡγίασεν αὐτήν. Scores: 0.796, 0.993 (normalized), 1.024 (standardized) leviticus 9:20 [rahlfs] καὶ ἐπέθηκεν τὰ στέατα ἐπὶ τὰ στηθύνια, καὶ ἀνήνεγκαν τὰ στέατα ἐπὶ τὸ θυσιαστήριον. Scores: 0.796, 0.992 (normalized), 1.022 (standardized) leviticus 9:20 [gottingen] καὶ ἐπέθηκεν τὰ στέατα ἐπὶ τὰ στηθύνια, καὶ ἀνήνεγκαν τὰ στέατα ἐπὶ τὸ θυσιαστήριον· Scores: 0.795, 0.991 (normalized), 1.020 (standardized) leviticus 8:11 [rahlfs]
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 80 αὐτοῦ· πλυνεῖ ἐν οἴνῳ τὴν στολὴν αὐτοῦ καὶ ἐν αἵματι σταφυλῆς τὴν περιβολὴν αὐτοῦ· [R_LXX] δεσμεύων πρὸς ἄμπελον τὸν πῶλον αὐτοῦ καὶ τῇ ἕλικι τὸν πῶλον τῆς ὄνου αὐτοῦ· πλυνεῖ ἐν οἴνῳ τὴν στολὴν αὐτοῦ καὶ ἐν αἵματι σταφυλῆς τὴν περιβολὴν αὐτοῦ· Scores: 0.896, 0.645 (normalized), 9.883 (standardized) exodus 3:17 [gottingen] καὶ εἶπα Ἀναβιβάσω ὑμᾶς ἐκ τῆς κακώσεως τῶν Αἰγυπτίων εἰς τὴν γῆν τῶν Χαναναίων καὶ Χετταίων καὶ Εὑαίων καὶ Ἀμορραίων καὶ Φερεζαίων καὶ Γεργεσαίων καὶ Ἰεβουσαίων, εἰς γῆν ῥέουσαν γάλα καὶ μέλι. Variants: [R_LXX] καὶ εἶπον Ἀναβιβάσω ὑμᾶς ἐκ τῆς κακώσεως τῶν Αἰγυπτίων εἰς τὴν γῆν τῶν Χαναναίων καὶ Χετταίων καὶ Αμορραίων καὶ Φερεζαίων καὶ Γεργεσαίων καὶ Ευαίων καὶ Ιεβουσαίων, εἰς γῆν ῥέουσαν γάλα καὶ μέλι. [R_LXX] καὶ εἶπον Ἀναβιβάσω ὑμᾶς ἐκ τῆς κακώσεως τῶν Αἰγυπτίων εἰς τὴν γῆν τῶν Χαναναίων καὶ Χετταίων καὶ Αμορραίων καὶ Φερεζαίων καὶ Γεργεσαίων καὶ Ευαίων καὶ Ιεβουσαίων, εἰς γῆν ῥέουσαν γάλα καὶ μέλι. Scores: 0.869, 0.621 (normalized), 9.366 (standardized) genesis 3:6 [gottingen] καὶ εἶδεν ἡ γυνὴ ὅτι καλὸν τὸ ξύλον εἰς βρῶσιν, καὶ ὅτι ἀρεστὸν τοῖς ὀφθαλμοῖς ἰδεῖν καὶ ὡραῖόν ἐστιν τοῦ κατανοῆσαι, καὶ λαβοῦσα τοῦ καρποῦ αὐτοῦ ἔφαγεν· καὶ ἔδωκεν καὶ τῷ ἀνδρὶ αὐτῆς μετʼ αὐτῆς, καὶ ἔφαγον. Variants: [R_LXX] καὶ εἶδεν ἡ γυνὴ ὅτι καλὸν τὸ ξύλον εἰς βρῶσιν καὶ ὅτι ἀρεστὸν τοῖς ὀφθαλμοῖς ἰδεῖν καὶ ὡραῖόν ἐστιν τοῦ κατανοῆσαι, καὶ λαβοῦσα τοῦ καρποῦ αὐτοῦ ἔφαγεν· καὶ ἔδωκεν καὶ τῷ ἀνδρὶ αὐτῆς μετ᾽ αὐτῆς, καὶ ἔφαγον. Scores: 0.852, 0.606 (normalized), 9.030 (standardized) proverbs 26:26 [rahlfs] εδρίοις. Scores: 0.818, 0.575 (normalized), 8.369 (standardized) Query 3: Nonnus Panopolitanus, Paraphrasis sancti evangelii Joannei, I 1-3 : Ἄχρονος ἦν, ἀκίχητος, ἐν ἀρρήτῳ λόγος ἀρχῇ καὶ λόγος αὐτοφύτοιο θεοῦ φάος, ἐκ φάεος φῶς John 1:1 [nestle-aland-28] Ἐν ἀρχῇ ἦν ὁ λόγος, καὶ ὁ λόγος ἦν πρὸς τὸν θεόν, καὶ θεὸς ἦν ὁ λόγος. Scores: 2.927, 1.000 (normalized), 6.161 (standardized)
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 81 Acts 7:17 [nestle-aland-28] Καθὼς δὲ ἤγγιζεν ὁ χρόνος τῆς ἐπαγγελίας ἧς ὡμολόγησεν ὁ θεὸς τῷ Ἀβραάμ, ηὔξησεν ὁ λαὸς καὶ ἐπληθύνθη ἐν Αἰγύπτῳ Scores: 2.801, 0.956 (normalized), 5.842 (standardized) psalms 105:48 [gottingen] Εὐλογητὸς κύριος ὁ θεὸς Ισραηλ ἀπὸ τοῦ αἰῶνος καὶ ἕως τοῦ αἰῶνος. καὶ ἐρεῖ πᾶς ὁ λαός Γένοιτο, γένοιτο. Scores: 2.770, 0.946 (normalized), 5.764 (standardized) psalms 105:48 [rahlfs] Εὐλογητὸς κύριος ὁ θεὸς Ισραηλ ἀπὸ τοῦ αἰῶνος καὶ ἕως τοῦ αἰῶνος. καὶ ἐρεῖ πᾶς ὁ λαός Γένοιτο γένοιτο. ––– Scores: 2.770, 0.946 (normalized), 5.764 (standardized) 1Tim 6:13 [nestle-aland-28] παραγγέλλω [σοι] ἐνώπιον τοῦ θεοῦ τοῦ ζῳογονοῦντος τὰ πάντα καὶ Χριστοῦ Ἰησοῦ τοῦ μαρτυρήσαντος ἐπὶ Ποντίου Πιλάτου τὴν καλὴν ὁμολογίαν, Scores: 2.708, 0.924 (normalized), 5.607 (standardized) genesis 14:19 [gottingen] καὶ εὐλόγησεν τὸν Ἀβρὰμ καὶ εἶπεν Εὐλογημένος Ἀβρὰμ τῷ θεῷ τῷ ὑψίστῳ, ὃς ἔκτισεν τὸν οὐρανὸν καὶ τὴν γῆν, Variants: [R_LXX] καὶ ηὐλόγησεν τὸν Αβραμ καὶ εἶπεν Εὐλογημένος Αβραμ τῷ θεῷ τῷ ὑψίστῳ, ὃς ἔκτισεν τὸν οὐρανὸν καὶ τὴν γῆν, Scores: 2.698, 0.921 (normalized), 5.582 (standardized) judith 13:17 [gottingen] καὶ ἐξέστη πᾶς ὁ λαὸς σφόδρα καὶ κύψαντες προσεκύνησαν τῷ θεῷ καὶ εἶπαν ὁμοθυμαδόν Εὐλογητὸς εἶ, ὁ θεὸς ἡμῶν ὁ ἐξουδενώσας ἐν τῇ ἡμέρᾳ τῇ σήμερον τοὺς ἐχθροὺς τοῦ λαοῦ σου. Scores: 2.689, 0.918 (normalized), 5.560 (standardized) Q: Please rate the overall quality of the results of your first query .1 A: 2.0 Q: Please rate the overall quality of the results of your second query .1 A: 1.0 Q: Please rate the overall quality of the results of your third query .1
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 82 A: 3.0 Q: Did the model return semantically relevant results, even if not literal matches? Did it do so among the top 5, 10 or 15 results? A: No, after the first 15 results Q: Did you find the results meaningful in your area of expertise? A: Yes, he made me proposals that I hadn't identified and which in some cases are relevant. Q: Did you notice any errors or inconsistencies in the suggested results? If yes, please describe them. A: Yes, when it comes to queries on Nonno, neologisms are a problem; it doesn't identify the exact meaning of more formally searched words (dialectal or archaic variants, neologisms), and it can't identify synonyms among the solutions it suggests. Q: How well do you think the model captures contextual meaning in Latin and Greek? A: It is of little success if the form of the text becomes complex, as in the case of epic poetry. Q: Do you think the amount of surrounding context provided is sufficient to understand the reuse? A: This aspect could be implemented, I don't think it's ever enough. Q: Did you find the metadata below each result useful and clear? A: Yes. Q: Did you encounter any difficulties interpreting the results? A: Sometimes, scores aren't easy to understand. Graphs are very helpful. Q: Please describe your general impression of the tool. A: It's fast and comprehensive, but I don't think it's very effective at identifying synonyms yet. The formal differences between the texts are crucial to the accuracy of the results. Could something be done about this during the machine training phase? Q: Would you suggest adding other types of metadata? If yes, which ones? A: Perhaps stylistic metadata (meter: dactylic hexameters)
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 83 Q: In which ways do you think this model could enrich your research work? A: It can help me find new connections with biblical texts implicit in a particular text such as a poetic paraphrase. Q: What improvements or additions would you suggest to make the tool more practical or useful for your specific research context? A: I would enhance its stylometric capabilities, making it more adaptable to formally diverse texts, such as poetic ones; I would improve its linguistic capabilities (in Greek, the ability to recognize dialectal variants...). Q: Which improvements would you propose to extend the scope of the model and scale it to your colleagues' needs? A: I don't know, perhaps a graphical interface that makes its use and advantages immediately evident. Q: Do you think specific training should be offered along with the tool? A: yes Q: Please feel free to add any further observations, comments, or suggestions: A: nothing
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 84 9.2. uBIQUity - Arabic Tool Test (WP8): Feedback Questionnaire Note on Response Format: To avoid redundancy and keep the appendix concise, each general question is listed only once, followed by the three individual responses labeled as Answer 1, Answer 2, and Answer 3. In contrast, for queries involving practical tool testing—where each participant tested the system with their own input (e.g., a ḥadīth, a tafsīr excerpt, or a keyword)—both the question and each participant’s answer are presented individually. This allows for a clearer understanding of how different testers interacted with the tool. Q: What digital tools do you currently use to analyze the Arabic sources you work on? A1: Qwen VL72B; Google Vision AI; Kraken OCR; Diffusion models; Different LLMs, Transformers (AraBERT) and other hybrid solutions. A2: al-Shamila. A3: I do not use any specialized tools at the moment; my work primarily involves organizing and analyzing data using Excel spreadsheets. Q: What are the main features you typically look for when working with digital tools for text analysis and text reuse? A1: Consistency in output performance, reusability, language independence in multilingual contexts, semantic or other metadata output and their quality. Hence, from a more Humanist perspective, insights into possible new connection between and reuse in different texts over decades and centuries. A2: The ability to search for quotations; support for Arabic script and morphology; export and citation. A3: At this stage, I do not rely on specialized digital tools for text analysis or text reuse. My work is primarily conducted using structured Excel files. However, when considering digital tools for future use, I would prioritize features such as support for right-to-left scripts (especially Arabic), the ability to handle large textual datasets, reliable search and matching functions, and compatibility with metadata extraction workflows. Q: Are you familiar with text reuse tools?
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 85 A1: Been at the Aga Khan University at the CDH center which developed and integrated the passim algorithm into the Kitab Project corpus they have previously built in the OpenITI frame. The passim system used in the text reuse case based on text alignment algorithm such as the Smith-Waterman. At the same time the CDH. A2: No A3: Yes Q: What is your level of familiarity with AI-based technologies in the humanities? A1: Expert A2: Beginner A3: Advanced Q: Are you familiar with transformer-based models such as BERT? A1: Use of BERT-based models fine-tuned for Arabic language such as araBERT for developing topic modeling solutions for cataloguing non-Latin and Islamic studies texts. A2: No A3: Yes Q: We are now providing three preset queries. For each query, please: 1. Examine at least the first 10 results returned by the tool, in the order in which they appear. 2. Rate the overall relevance of the results on a scale from 1 to 4. 3. Provide a brief explanation of your rating. Query 1: ﺮﺴﯿﺗ ﺎﻣ ﮫﻨﻣ اوؤﺮﻗﺎﻓ فﺮﺣأ ﺔﻌﺒﺳ ﻰﻠﻋ لﺰﻧأ نآﺮﻘﻟا نإ (From Ṣaḥīḥ al-Bukhārī, Kitāb al-Khuṣūmāt) How relevant were the results returned by the tool? A1: Mostly relevant
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 86 A2: Mostly relevant A3: Mostly relevant Q: Please explain your rating for Query 1. Did the tool return semantically relevant results, even if they were not literal matches? Did relevant results appear among the top 5, 10, or 15? A1: Results are returned, and not all are only exact matching but also paragraphs with semantic relevance and close to the hadith searched are outputted which is a good result. Semantic similarity provided and ranked in a proper way. Reducing the top k retrieved passages to 15, 10 or even less will output the most similar in decreasing order. A2: I chose ‘mostly relevant’ because the first result wasn’t complete. In result 1 the author’s name is missing even if it appears in the ‘filename’ (Nizam al-din al-Nisaburi). I found the author’s birth date (650 H. /1253 C. E.?), also missing, and other biographical information in the following sources: VIAF (Virtual International Authority File); English Wikipedia; al-Kindi Catalogue. In result 2 we found a lot of precise information, but a little bit confused in second part of section ‘inf’: in this section the tool quotes many sources to demonstrate that the work belongs to the author, but this info is not too relevant for our search. The third result is correct, and we verified the info on Shamila. All the results appear in the correct chronological order. A3: I assigned a rating of 3 (Mostly Relevant). The tool successfully retrieved several results that are either exact matches or semantically equivalent paraphrases of the hadith: “ﺮﺴﯿﺗ ﺎﻣ ﮫﻨﻣ اوؤﺮﻗﺎﻓ فﺮﺣأ ﺔﻌﺒﺳ ﻰﻠﻋ لﺰﻧأ نآﺮﻘﻟا اﺬھ نإ”. These include narrations from major hadith collections such as Ṣaḥīḥ Ibn Ḥibbān, Sunan Abī Dāwūd, and al-Mustakhraj ʿalā ṢaḥīḥMuslim, which appeared within the top 5 results. Query 2: ﻲﻓ ثﺪﺣ يﺬﻟا اﺬھ نإ ﺲﯿﻠﺑإ لﺎﻗ ﺐﮭﺸﻟﺎﺑ اﻮﻤﺟرو ءﺎﻤﺴﻟا ﺖﺳﺮﺣ ﺎﻤﻠﻓ ﻊﻤﺴﻟا قﺮﺘﺴﺗ ﺖﻧﺎﻛ ﻦﺠﻟا نأ ﻚﻟذ ﺐﺒﺳ نﺎﻛو ضرﻷا ﻲﻓ ثﺪﺣ ءﻲﺸﻟ ءﺎﻤﺴﻟا (From Tafsīr al-Thaʿlabī) A1: Mostly relevant A2: Mostly relevant A3: Very relevant Q: Please explain your rating for Query 2. Did the tool return semantically relevant results, even if they were not literal matches? Did relevant results appear among the top 5, 10, or 15?
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 87 A1: Same as the first case, however as expected the reuse of longer sentences with different semantics meaning and the intrinsic polysemy of many of the terms involved cause the similarity score to drop significantly. Nonetheless, the output is fundamentally sound in every case of top k retrieved passages. A2: Result 1: the author’s name is not complete (Abu Hafs Umar ibn Ali ibn Adil alDimashqi al-Hanbali) and the birth date is missing (675 H. / 1276 CE). Result 2 contains a mistake: the last highlighted sentence occurs twice (see al-Shamila) and the last word is incomplete. Result 3 is correct except for the death date in Gregorian calendar (965 instead of 944-45). Result 4 repeats result 3. The chronological order of the sources is reversed. A3: Although the tool did not return any exact textual matches for the query, it retrieved several passages that were highly semantically relevant. Relevant results appeared consistently within the top 5 and top 10. Query 3: لﺎّﺟﺪﻟا A1: Mostly relevant A2: Slightly relevant A3: Very relevant Q: Please explain your rating for Query 3. Did the tool return semantically relevant results? Did relevant results appear among the top 5, 10, or 15? A1: A single word of course will find many matches and reuse case on several text. Here semantic similarity is hard to define since different passages with different contexts happen to have the same word inside and as imaginable the similarity is like the previous drop staying around 0.68 to 0.58. Results are relevant and related to the single word sentence even though contexts are different. A2: The results are semantically relevant, except the result 4 where the query does not appear at all. Anyway, the criterion used in the choice of sources is unclear. They seem to have been chosen at random. For ex., the first quotation is taken from the collection of hadith by Abu Dawud, which is not the most relevant, as the query is also present in Bukhari and Muslim. Another example: in the second result Razi’s tafsir has been chosen, but why not Tabari, Ibn Kathir or Zamakhshari? The result 4 pick up the word in a noncanonical source (al-Tabarani).
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 88 A3: The tool returned strong results on the topic of al-Dajjāl, including canonical hadith and narrative or eschatological discussions. However, not all top 10 results are equally focused or relevant. Relevant results appeared among the top 5 (especially Results 1, 2, and 5). Results (6–9) were highly relevant. Interviewee 1 Q: You may now test the tool using three queries of your choice. Feel free to select a passage, an expression, and a keyword from your research (e.g., a ḥadīth, a tafsīr excerpt, and a specific term), and evaluate how well the tool retrieves relevant results. Please write here your first query and comment on the first results returned by the tool. A1: ﺮﺒﻛﻷا دﺎﮭﺠﻟا ﻰﻟإ ﺮﻐﺻﻷا دﺎﮭﺠﻟا ﻦﻣ ﺎﻨﻌﺟر Score of similarity is rather low even though the semantic output is correct and relevant in most cases including from exact matching to more articulate semantic contexts. I think the tool in this sense is functional to providing same passages through different textual traditions and outlining eventual reuse cases. Note also that this hadith is considered by most studies as not Sahih. Q: How relevant were the results returned by the tool for your first query? A1: Mostly relevant Q: Please write your second query and comment on the first results returned by the tool. A1: ﻢﻠﻌﻟا ﺾﺒﻘﯾ ﻻ ﷲ نإ":لﺎﻗ ﻢﻠﺳو ﮫﯿﻠﻋ ﷲ ﻰﻠﺻ ﷲ لﻮﺳر نأ ﺎﻤﮭﻨﻋ ﷲ ﻲﺿر صﺎﻌﻟا ﻦﺑ وﺮﻤﻋ ﻦﺑ ﷲ ﺪﺒﻋ ﺚﯾﺪﺣ اﻮﻠﺌﺴﻓ ، ً ﻻﺎﮭﺟ ﺎًﺳوؤر سﺎﻨﻟا ﺬﺨﺗا ﻢﻟﺎﻋ ﻖﺒﯾ ﻢﻟ اذإ ﻰﺘﺣ ،ءﺎﻤﻠﻌﻟا ﺾﺒﻘﺑ ﻢﻠﻌﻟا ﺾﺒﻘﯾ ﻦﻜﻟو ،دﺎﺒﻌﻟا ﻦﻣ ﮫﻋﺰﺘﻨﯾ ﺎًﻋاﺰﺘﻧا اﻮﻠﺿأو اﻮﻠﻀﻓ ،ﻢﻠﻋ ﺮﯿﻐﺑ اﻮﺘﻓﺄﻓ" Similarity score, results and top k ranking seems to work properly on this query, pretty much as done in the others. Q: How relevant were the results returned by the tool for your second query? A1: Very relevant Q: Please write your third query (a single keyword), and comment on the first results returned by the tool.
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 89 A1: ةرﺎﻣأ Results were relevant, in this case most of them related to Arabic dictionaries (general or qur’anic ones). Basically, the tool behaved as did in the single keyword of the third query in the first section. Q: How relevant were the results returned by the tool for your third query? A1: Mostly relevant Interviewee 2 Q: You may now test the tool using three queries of your choice. Feel free to select a passage, an expression, and a keyword from your research (e.g., a ḥadīth, a tafsīr excerpt, and a specific term), and evaluate how well the tool retrieves relevant results. Please write here your first query and comment on the first results returned by the tool. A2: ْ ﺮِﻄْﻓَأَو ﺎًﻣْ ﻮَﯾ ْﻢُﺻ :َلﺎَﻗ ﻰﱠﺘَﺣ َلاَز ﺎَﻤَﻓ ،َﻚِﻟَذ ْﻦِﻣ َﺮَﺜْﻛَأ ُﻖﯿِطُأ :َلﺎَﻗ ،ٍمﺎﱠﯾَأ َﺔَﺛَ ﻼَﺛ ِ ﺮْﮭﱠﺸﻟا َﻦِﻣ ْﻢُﺻ ِّﻞُﻛ ﻲِﻓ َنآ ْ ﺮُﻘْﻟا ِأَﺮْﻗا :َلﺎَﻘَﻓ ،ﺎًﻣْ ﻮَﯾ ٍث َ ﻼَﺛ ﻲِﻓ :َلﺎَﻗ ﻰﱠﺘَﺣ َلاَز ﺎَﻤَﻓ ،َﺮَﺜْﻛَأ ُﻖﯿِطُأ ﻲِّﻧِإ :َلﺎَﻗ ، ٍ ﺮْﮭَﺷ We choose a hadith from the Sahih of Bukhari. The tool didn’t recognize the original source (Bukhari) and the first two results are variants taken from the same source: Abu Dawud, which is onother canonical source. The other results are selected from noncanonical sources (Ibn Habban, Ibn Farra’ al-Baghawi, al-Maturidi). Also in this case, the criterion for selecting the sources is not clear. Q: How relevant were the results returned by the tool for your first query? A2: Slightly relevant Q: Please write your second query and comment on the first results returned by the tool. A2: ﻢﮭﻨﺌﺒﻨﺘﻟ ﮫﯿﻟإ ﺎﻨﯿﺣوأو) :ﮫﻟﻮﻗ ﻲﻓ ،ﺪھﺎﺠﻣ ﻦﻋ ،ﺢﯿﺠﻧ ﻲﺑأ ﻦﺑا ﻦﻋ ،ءﺎﻗرو ﻦﻋ ،ﷲ ﺪﺒﻋ ﺎﻨﺛﺪﺣ ،لﺎﻗ قﺎﺤﺳإ ﺎﻨﺛﺪﺣ ،لﺎﻗ ﻚﻟﺬﺑ نوﺮﻌﺸﯾ ﻻ ﻢھو ،اﻮﻌﻨﺻ ﺎﻤﺑ ﻢﮭﺌﺒﻨﯿﺳ ْنأ ّﺐﺠﻟا ﻲﻓ ﻮھو ﻒﺳﻮﯾ ﻰﻟإ ﻰﺣوأ :لﺎﻗ (نوﺮﻌﺸﯾ ﻻ ﻢھو اﺬھ ﻢھﺮﻣﺄﺑ ﻲﺣﻮﻟا The second query was selected from Ṭabarī’s tafsīr (Q 12:15). The tool does not recognize the source from which the passage was taken (Ṭabarī) and does not select other tafsirs
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01 96 A1: Arabic Sources as corpora is not clear what contains and from where it is taken. Is there any overlap with the other corpus? Are they complementary? What kind of sources? tafasir, ahadith collections, jarh wa ta’dil, biographies, tarajim, tabaqat, fiqh, lexicography and ʿilm al-lughah? A2: No answer A3: No answer (End of Document)
Ir0000014 – Itserr Status: FINAL ITSERR-WP8-D8.1.3-ONFIELD Version: 00.01