scieee AI-readable full text Open interactive document viewer

Quantifying the Environmental Footprint of Curating Datasets with LLMs

Lang, Sarah; Pitawanik, Wishyut; Belouin, Pascal; Sevink, Emma; Olszynko-Gryn, Jesse; Freeborn, Alfred; Benson, Etienne

Abstract

This study evaluates the environmental trade-offs of using large language models to curate cross-collection oral-history datasets in the Commoning Oral Histories of Knowledge (CORAL) project. Manual screening of 2,606 interviews was benchmarked against a workflow that tested four instruction-tuned LLMs and two prompt designs. Environmental impact was approximated using token-based inputs to EcoLogits, although implementing such assessments remains non-trivial. Ultimately, we conclude that the environmental impact of our project's use case could be considered moderate compared to common academic activities such as traveling to conferences. However, such impacts should be monitored closely, as they may vary significantly across different research setups and are likely to scale with larger datasets and broader adoption of LLMs in the field. Finally, the paper urges sufficiency-oriented practices and transparent carbon reporting in Computational Humanities research.

Full text

Working Paper Quantifying the Environmental Footprint of Curating Datasets with LLMs Sarah Lang1, Wishyut Pitawanik1, Pascal Belouin1, Emma Sevink2, Jesse Olszynko-Gryn1, Alfred Freeborn1, Etienne Benson1 1Max Planck Institute for the History of Science, Berlin, Germany 2Freie Universität Berlin, Berlin, Germany Abstract This study evaluates the environmental trade-offs of using large language models to curate cross-collection oral-history datasets in the Commoning Oral Histories of Knowledge (CORAL) project. Manual screening of 2,606 interviews was benchmarked against a workflow that tested four instruction-tuned LLMs and two prompt designs. Environmental impact was approximated using token-based inputs to EcoLogits, although implementing such assessments remains non-trivial. Ultimately, we conclude that the environmental impact of our project’s use case could be considered moderate compared to common academic activities such as traveling to conferences. However, such impacts should be monitored closely, as they may vary significantly across different research setups and are likely to scale with larger datasets and broader adoption of LLMs in the field. Finally, the paper urges sufficiency-oriented practices and transparent carbon reporting in Computational Humanities research. Keywords: environmental footprint, environmental sciences, oral history, Large Language Models, digital humanities, critical digital humanities Sarah Lang, Wishyut Pitawanik, Pascal Belouin, Emma Sevink, Jesse Olszynko-Gryn, Alfred Freeborn and Etienne Benson. “Quantifying the Environmental Footprint of Curating Datasets with LLMs.” Working Paper, 2025. https://doi.org/10.5281/zenodo.17902822. ©2025 by the authors. Licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). 0 1 Introduction Recent scholarship and activism – both within and beyond the digital humanities – have begun to engage critically with the environmental implications of artificial intelligence (AI), as well as with the complicity of academic research in adopting such technologies [15]. In the digital humanities, a number of initiatives have emerged in response to the ecological footprint of computational tools. Debates around this issue remain ongoing and, at times, polarised. Some argue for the wholesale rejection of AI due to ethical and ecological risks, while others advocate more moderate, exploratory approaches. Many digital humanities practitioners continue to experiment with AI-driven tools to test their potential value. Although discourse on the environmental costs of AI is expanding, one of the primary challenges lies in the difficulty of generating precise and reliable measurements of these impacts. Tools such as EcoLogits, CodeCarbon, CallToChange, and the AI Energy Score seek to address this by visualising energy consumption, thereby making its implications more accessible to researchers. This paper contributes to debates surrounding the environmental impact of the adoption of AI technologies in humanities research by evaluating the effectiveness and environmental costs of using large language models (LLMs) to identify relevant across disciplinary boundaries in relevant oral history collections. We assess the qualitative value and environmental cost of LLM use in our specific scenario, thereby enabling a more informed judgment about their justification in scholarly workflows. A wealth of historical and humanities-related resources is available online, for example through digital libraries, that scholars can use for their research. Yet, while much material has been digitised, researchers often face difficulties navigating these ever-more abundant collections. Scholars are often required to manually collect and enrich data to create datasets that reflect their unique needs, yet much material still remains underused or buried in vast digital corpora. Advancing digital humanities scholarship entails more than ensuring access to and the preservation of digitised materials; it also requires the development of effective entry points and analytical tools that can facilitate meaningful engagement with the existing resources. This is particularly the case for oral history, where researchers frequently encounter challenges such as non-standard metadata, diverse access protocols, and highly variable data formats. These factors complicate discovery, comparison, and analysis, underscoring the necessity of tailored infrastructures and methodologies to support scholars in deriving insights from such heterogeneous collections. A significant challenge in the historical study of science, technology, medicine, and the environment is that, although interviews held across various institutions represent a valuable source base, they are frequently underutilised due to inconsistencies in access policies, metadata standards, and user interfaces. As a result, many potentially significant narratives remain overlooked. The Commoning Oral Histories of Knowledge (CORAL) project responds to this need by supporting the creation of personalized, curated oral history datasets, drawing across existing collections.1 Such datasets tailored to specific research interests are vital not only for advancing computational analysis techniques but also for traditional scholarship engaging with the growing body of digitised sources. CORAL is not merely an aggregator for oral history interviews. Rather, it provides a 1https://www.mpiwg-berlin.mpg.de/research/projects/coral-commoning-oral-histories-knowledge Developed by the Department on Knowledge Systems and Collective Life at the Max Planck Institute for the History of Science (MPIWG), CORAL aims to address this issue by providing a digital platform for cross-institutional discovery and thematic analysis of oral history interviews. https://coral.mpiwg-berlin.mpg.de While the platform does not host or provide direct access to interview recordings or transcripts – users must access these through the holding institutions – it enables the identification of relevant materials and enhances the discoverability of oral history collections. CORAL builds directly on the earlier Commoning Biomedicine (ComBio) project, which was developed by the MPIWG’s independent research group “Practices of Validation in the Biomedical Sciences,” led by Lara Keuck between 2021 and 2024. ComBio aimed to centralise access to disparate oral history resources in the biomedical sciences and currently provides searchable metadata for 1,637 records drawn from 15 different collections. https://combio.mpiwg-berlin.mpg.de. This working paper was written in July 2025 and represents the state of research at the time of writing. 1 workflow to help scholars craft custom collections and curate datasets for further analysis, digital or traditional, tailored to their unique research aims. To make this process more efficient, we have experimented with using LLMs in selected cases. While CORAL is designed to address a wide range of themes, in the case of the Storying the Earth and Environmental Sciences (SEES) subcollection we discuss in this article, we are specifically focusing on environment-related sources. Such topics are often subsumed under a wide range of synonyms and abstract terms, which complicates traditional filtering methods like simple, targeted keyword searches. LLMs may offer a means of improving the efficiency of discovery and retrieval of relevant materials in such contexts. It is important to note that LLMs do not always add value and thus, should only be employed selectively, especially given their environmental impacts. For topics suited to keyword searches, where relevant terms are neither obscured by umbrella concepts nor require complex semantic interpretation, LLM-based approaches may be unnecessary. LLMs might even introduce errors, thus not only failing to add value but actually subtracting it. Our present research thus involves a comparative evaluation of this LLM-supported workflow in terms of its environmental impact and usefulness in our research. By measuring energy consumption and comparing performance against conventional methods, we aim to determine whether such approaches offer a sustainable and effective means of supporting our research practices. This paper exemplifies an approach that balances the use of advanced computational technologies to enhance humanities research with critical reflection on our own practices [12, cf. p. vii–xii, p. 1–7]. As digital humanities scholars have noted, we must ‘turn the macroscope on ourselves [13, p. 73] from time to time to understand how these tools shape our work and whether their use is justified – accepting that, in some cases, LLMs may not offer sufficient benefits to warrant their environmental costs.2 2 Environmental Impact and Sustainability of Large Language Models The Digital Humanities and the Climate Crisis Manifesto, authored by a transnational collective of scholars [3], recognises that the field of digital humanities both participates in and perpetuates the global climate crisis through its practices, emphasising that the digital is inherently material.3Prendergrass and colleagues [29] critically examine digital preservation infrastructures and their environmental impacts, observing that the term sustainability in this context often refers more to workforce continuity and financial viability than to environmental concerns. However, these infrastructures do have an environmental footprint, which can be partially mitigated through technological adaptations such as using greener hosting providers. The paper ultimately calls for environmentally conscious archival practices to be recognised as parts of a professional ethics. Baillot contends that the apparent efficiency gains achieved through AI in the Global North result in corresponding losses of time and resources in the Global South [2]: Our AI usage obscures its asymmetric global distribution of labour and environmental burden, effectively reinforcing existing colonial structures of power. By externalising the material and human costs of digital technologies, AI development contributes to the persistence of exploitative global hierarchies, wherein the benefits accrued in one region are made possible by systemic extractions from another. The training and deployment of LLMs entail considerable energy consumption, a demand that 2In many research contexts within the humanities, smaller, task-specific machine learning models may, in fact, be more appropriate. Such models are not only more stable and easier to fine-tune, but they also consume considerably fewer computational resources. This approach may be more compatible with a commitment to sustainability, not only in ecological terms but also in relation to financial and technical feasibility: Large models necessitate access to highperformance computing infrastructure, which many institutions cannot afford or maintain. A use-what-is-necessary approach corresponds with both ecological sufficiency and institutional resource constraints. 3Additional institutional efforts are underway to promote transparency and responsible computing. The ‘Greening DH’ working group of the German Digital Humanities association Digital Humanities im deutschsprachigen Raum (DHd), for example, compiles best practices and empirical data to foster sustainability in digital scholarship [4]. Baillot also assess the costs of digital access to text [1]. A related area of research is Digital Environmental Humanities [34]. 2 is expected to increase significantly in the absence of regulatory measures [10].4Although some argue that the environmental impact of occasional individual use is negligible – particularly when compared to emissions generated by practices such as international conference travel – this comparison should not be taken as justification for uncritical or excessive use.5Taken together, existing contributions to this debate suggest that machine learning and generative AI are not inherently unsustainable. Sustainability depends on how models are trained and deployed. Nevertheless, many scholars advocate for mandatory carbon disclosure, ideally within a centralised emissions repository, and for comprehensive life cycle assessments of environmental impact.6While energy use dominates discussions of sustainability in computational research, calculating it remains complex [6, 7, 11, 14, 16, 18, 19, 21, 22, 23, 24, 27, 30, 31, 33]. Tools for estimating the energy footprint of custom-trained models are now available (see overview in appendix A), but many widely-used AI platforms, such as ChatGPT, do not offer transparent data on energy usage. Even much more privacy-conscious infrastructure providers like the Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen (GWDG), which support high-performance computing for researchers, still lack integrated mechanisms to track or report energy consumption.7 Beyond energy, other forms of environmental cost remain insufficiently addressed. As Kate Crawford has highlighted, water usage is a significant but often overlooked factor [10, 20]. Many data centres are situated in desert regions where substantial water resources are diverted for cooling purposes – resources that might otherwise serve essential human needs. This raises serious ethical questions about the prioritisation of computational infrastructure over the welfare of local populations. Electronic waste (e-waste) constitutes an additional concern. The continuous advancement of AI systems necessitates ongoing hardware upgrades, leading to the disposal of older equipment. A significant proportion of this e-waste is exported to countries in the Global South, frequently without proper recycling or regulatory oversight. Thus, the environmental impact of 4Kate Crawford’s seminal Atlas of AI underscores the material and ecological costs embedded in AI systems [9]. This concern also highlighted in the now well-known ‘Stochastic Parrots’ paper, particularly in relation to LLMs [5]. 5It is nonetheless worth noting that even environmentalists concur that typical everyday interactions with LLMs by individual users generate only minimal environmental cost during inference: Andy Masley, in a widely circulated 2025 blog post [26], argues that focusing climate concerns on chatbot usage of individuals is misguided, as it distracts from the structural causes of emissions. He calls for evidence-based prioritisation in climate strategy, contending that public and activist attention directed toward climate anxiety around ChatGPT diverts focus from more impactful policy-level and systemic issues. Masley supports his argument with calculations showing that while emissions from LLM usage scale with the number of queries, they remain small compared to everyday electricity and water consumption in the Global North. He advocates redirecting attention toward higher-impact behaviours such as eating meat, air travel, and structural interventions. Even the energy-intensive training phase, he notes, becomes negligible when averaged across billions of uses. In his estimate, individual ChatGPT interactions consume less than 4 Wh. However, this steady demand may drive AI companies to invest substantial resources in model training, which remains far more energy-intensive. 6Patterson et al. argue that machine learning papers requiring significant computational resources should, where possible, make their energy consumption explicit, and that CO2emissions should be a key evaluation metric—covering both training and inference [28]. Strubell et al. similarly contend that comparing models through cost–benefit analyses, including accuracy and environmental impact, would be beneficial [33]. Luccioni and Hernandez-Garcia present a broader survey of emissions across 95 models [22]. Overall, factors such as model architecture, data centre location, and infrastructure choices play a substantial role in determining environmental outcomes, and researchers should pay close attention to these variables, which leads Patterson et al. [27] to argue that widespread adoption of best practices could significantly reduce emissions, given the large differences resulting from design and deployment choices. To assess this, full life cycle assessments (LCA) of AI services would be ideal [21]. Discussing the dilemma of sustainable scaling in AI, Desroches et al. confirm that large models consume far more energy than traditional ones, reinforcing the point that architectural decisions matter [11]. We can also draw on existing work work on making AI more sustainable [8, 17, 24, 35]. 7The Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen (GWDG) functions as a service organisation jointly operated by the University of Göttingen and the Max Planck Society. It serves as a central data and IT service centre, offering infrastructure and support to research institutions and universities. Among its core responsibilities is the operation and maintenance of high-performance computing systems for academic use. 3 high-performance computing extends well beyond the immediate sites of technological production and use. Even areas in Digital Humanities not directly related to AI, such as the research data preservation or the long-term archiving of digital cultural heritage, are a factor in these systemic issues [29]. Researchers must therefore consider the full material footprint of AI technologies [21]. The large-scale operation of LLMs, especially when avoidable, contributes to broader patterns of ecological degradation. However, efficiency in prompting and measured API use can mitigate both financial and environmental costs. Since energy usage, as calculated by measurement tools available today, seems to correlate closely with token volume, optimising inputs and outputs offers a practical way to reduce impact in larger-scale applications. A key concern going forward is how exactly to measure LLM energy consumption, which remains methodologically difficult to capture accurately. Reliable data on energy use per conversation, message, or token is still scarce. Toolkits such as CodeCarbon, EcoLogits, and Callto-Change offer way to track energy consumption at runtime.8These software libraries already include features to track environmental metrics related to inference-time token usage. These tools may become increasingly relevant for institutions such as libraries or digital humanities centres conducting large-scale projects. However, they often exclude training-related costs, which are better captured by aggregate metrics such as the AI Energy Score. It is important to emphasise that most environmental impact estimates for large language model usage, such as those derived using the EcoLogits Calculator [32], are based on publicly available data and generalised assumptions.9 These assumptions typically include average figures for energy consumption per token, model size, and hardware efficiency. As a result, such estimates offer a valuable point of comparison between different models or workflows, but they should not be interpreted as precise measurements of real-world emissions. In practice, actual energy usage and associated carbon emissions can vary considerably. Factors such as GPU performance, hardware architecture, thermal conditions, model optimisation strategies, and load-balancing across distributed systems all play a role in shaping the environmental footprint of an LLM task. Furthermore, the source of electricity powering the compute infrastructure, whether derived from renewable or non-renewable sources, can have a significant impact on carbon output, yet is often not reflected in estimation tools. Thus, while tools like EcoLogits provide a useful and accessible framework for environmental benchmarking, they remain approximations. They are most effective for comparative analysis, helping researchers identify relatively more or less efficient models or configurations. Their outputs should be interpreted with caution and contextualised by an awareness of the broader system-level variables that influence emissions. While there is ample work on measuring the environmental impact of LLMs in Computer Science contexts [6, 7, 11, 14, 16, 18, 19, 21, 22, 23, 24, 27, 30, 31, 33], practical implementations are still rare in the Digital Humanities, with many researchers still not taking advantage of available resources.10 This may be due to a lack of preliminary work to guide the process. With this publication, we aim to address that gap by applying these tools to our research setup and documenting the relevant context and potential difficulties scholars may encounter when using them in their own work. 8We provide an overview with the results of our survey of relevant tools in appendix A. 9The following overview draws on information from available calculator tools and specifically the disclaimer from https://llmemissions.com/, a carbon calculator based on [31]. 10 The CorDeep project, for instance, has introduced an interface feature allowing users to generate an environmental impact statement during inference (of its non-LLM ML application) displayed before users agree to run the analysis, with the primary aim of raising awareness about the environmental costs of deploying machine learning. It supports the principle of sustainability by reducing redundant efforts – such as the need for researchers to retrain models independently – by offering the trained model as a web service: https://cordeep.mpiwg-berlin.mpg.de 4 3 Workflow Development and Model Evaluation in CORAL The initial phase of the CORAL project focused on surveying existing oral history collection platforms to create a comparative overview of available materials. This involved compiling structured data for each collection in a shared spreadsheet accessible to technical staff. Following this groundwork, researchers established inclusion criteria to guide the selection of relevant interviews. For instance, in the Storying the Earth and Environmental Sciences (SEES) subproject, we initially focused on interviews from academic experts on the history of earth and environmental sciences. In this first round of analysis, interviews centered on policy, activism, or experiential environmental knowledge were excluded. The intent was to ensure a coherent and manageable dataset for researchers working within a clearly defined domain. However, in practice, the implementation of these criteria proved complex. As many interviews fell into a grey area between relevance and irrelevance, it would have been challenging to detect relevant interviews solely using simple keyword searches. Once criteria were set, we assessed which interviews from identified collections met these standards. Manual review began with the Voices oral history collection of the US National Oceanic and Atmospheric Administration.11 This process involved examining each sub-collection, reading abstracts, and, where necessary, consulting transcripts or researching the interviewee’s background. Three broad outcomes emerged: 1) fully relevant sub-collections; 2) small, partially relevant sub-collections manageable for full manual review; and 3) large, mixed-content collections that were too labour-intensive to screen manually. This latter category raised questions about scalability and led to the exploration of large language models (LLMs) as a filtering tool which would assess interview relevance and provide brief justifications, with a human expert reviewing all positively identified interviews prior to final inclusion. This approach was intended to reduce the tedium of manually screening large datasets, not to replace human judgement. In this case study, only 16% (411) of 2,606 interviews in the collections were found to meet inclusion criteria, highlighting the scale of irrelevant material and the potential labour savings offered by automated filtering. To that end, the team tested a variety of prompts, instructing the LLM to categorise responses as ‘relevant’ or ’irrelevant’ based on specified criteria, including keyword density and thematic cues. Despite significant effort to optimise prompts through systematic rephrasing, practical testing initially showed limited improvement through adapted prompts. Given the review structure, the project prioritised recall over precision; false positives (0;1) were acceptable as they would still undergo human review, whereas false negatives (1;0) risked the exclusion of relevant material without further verification. A formal evaluation followed, comparing LLM classifications against the manually reviewed NOAA dataset. Analysis of classification discrepancies identified several recurring issues. The LLMs often failed to recognise academic or scientific credentials unless explicitly stated. It also struggled with adjacent disciplines, such as engineering or applied environmental roles, misclassifying them as irrelevant despite their academic dimensions. In some cases, the model offered no rationale for its decisions. Misclassifications also arose from overreliance on surface features, such as titles (e.g., “Dr.”) or environmental terminology, which led to false positives among consultants, journalists, or legal professionals outside the project’s academic focus. Conversely, obliquely phrased academic affiliations often resulted in false negatives. In some instances, the model generated inaccurate academic attributions based on keyword inference rather than verifiable evidence. These outcomes demonstrated the model’s sensitivity to phrasing. Misjudgments frequently stemmed from conflating environmental language with academic expertise, especially in interviews referencing aspirations, informal learning, or non-research roles. The ongoing task is to refine the prompt to improve recall and precision, with attention to maintaining systematic records of prompt versions 11 NOAA (National Oceanic and Atmospheric Administration) Voices Oral History Archives: https://www. climate.gov/maps-data/dataset/noaa-voices-oral-history-archives. 5 and their outcomes to prevent duplication. Crucially, the usefulness of specific prompts to the task at hand is closely tied to the clarity of inclusion criteria, which must be precisely formulated for each collection. Cases falling into the ‘grey zone’ posed challenges both for human reviewers and the LLM.12 4 Prompt Design, Model Selection, and Evaluation Methodology To evaluate the potential of LLMs in classifying oral history interviews by relevance, we employed models offered by the Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen (GWDG).13 All models were accessed via the infrastructure provided by GWDG and no architectural modifications were made. Initially, the following pre-trained instruction-tuned LLMs were tested: llama-3.1-sauerkrautlm-70b-instruct,mistral-large-instruct,meta-llama-3.1-8b-instruct and llama-4-scout-17b-16e-instruct. Each model received a structured prompt consisting of: (1) a short topic description, (2) the extractive summary of the interview transcript, and (3) an instruction to produce a JSON-formatted response containing a binary classification (‘relevant’ or ‘irrelevant’) and a justification for the decision. To improve reproducibility and reduce stochastic variation, all models were run with a temperature setting of 0. Each transcript was evaluated using a single prompt instance. Our study then focused on two LLMs capable of structured output using the Instructor module14,llama-3.1-sauerkrautlm-70b-instruct15 and mistral-large-instruct.16 mistral-large-instruct took significantly longer to process transcripts compared to the llama variant. Prompt Variants Two types of prompts were tested (see appendix B). The basic prompt offered concise instructions asking whether an interview reflected an academic perspective within the environmental sciences. The detailed prompt extended this with examples, a list of subfields (e.g., ecology, climatology, conservation), and clarification on relevant academic profiles.17 If an official summary was available, it was inserted before the key sentences to provide additional context. Dataset and Task The dataset comprised 2,606 oral history interviews, each stored in PDF format. Of these, 2,491 transcripts were successfully processed and evaluated. A subset of 115 interviews was excluded due to missing or unprocessable transcripts. Among the evaluated transcripts, 411 were manually labelled by a researcher as relevant to the defined inclusion criteria. 12 The manual classification of interviews was the most labour-intensive component of the workflow, with significant variation in difficulty depending on the collection’s structure. The breakdown of labour for this phase of CORAL is as follows: •Researching relevant oral history collections (including legal and contact information): approximately 15 hours by the current student assistant, but building on earlier work. •Drafting and testing LLM prompts: approximately 8 hours. •Manual classification of NOAA interviews (n=2,455): approximately 50 hours. •Comparative analysis of LLM vs. human classification: approximately 15 hours. •Workflow development, coordination, and reporting: approximately 10 hours. 13 https://docs.hpc.gwdg.de/services/chat-ai/models/index.html 14 https://python.useinstructor.com/ 15 https://huggingface.co/VAGOsolutions/Llama-3.1-SauerkrautLM-70b-Instruct 16 https://huggingface.co/mistralai/Mistral-Large-Instruct-2407. Other models, including meta-llama-3.1-8b-instruct, were tested but ultimately excluded due to incompatibility with our requirement of generating structured output. 17 Future work could explore further improving the detailed prompt, for instance by explicitly requesting information on academic background and relevance to environmental science, as our results suggest that increased specificity and clarity in prompt wording significantly improve its effectiveness. 6 Large language models (LLMs) were selected over other machine learning classifiers for several reasons. First, LLMs offer strong generalisability to new topics without the need for retraining or model-specific tuning. Second, the LLMs we used do not retain training data internally, thereby reducing privacy concerns. Third, and crucially for this application, LLMs are capable of producing interpretable outputs in the form of natural-language justifications for each classification decision – an affordance typically unavailable in standard machine learning approaches. The task of identifying relevant interviews within oral history collections was framed as a binary classification problem characterised by a highly imbalanced label distribution. Of the 2,606 interviews examined, only 411 (16%) were deemed relevant. The primary objective was to maximise the detection of these relevant cases; thus, the evaluation prioritised recall over other metrics, with the F1 score considered secondary. Preprocessing and Summarisation Interview transcripts were extracted using pypdf18. We initially considered three summarisation strategies for inputs which exceed the 125k token limit: A first-N-token heuristic, where extracted content us truncated to fit the token limit (125k tokens), starting from the beginning of the transcript, keyword frequency (term frequency) and TF-IDFbased sentence extraction.19 For input selection, we finally adopted the first-n-token heuristic method, feeding the first portion of each transcript up to the token limit. This was based on the assumption that interviewees typically introduce their professional background at the beginning. More elaborate summarisation strategies, such as keyword frequency or TF-IDF-based extraction, were explored but ultimately abandoned.20 Only five of 2,491 interviews (0.2%) exceeded the LLM context window, making such methods unnecessary for the vast majority of transcripts. Evaluation Setup and Metrics Each model was prompted once per transcript using temperature set to 0 to minimise randomness from the LLM. Responses were returned in JSON format, including a binary classification decision (relevant or irrelevant) and a justification. If an API or formatting error occurred, the prompt was retried up to five times.21 After five failures, the transcript was excluded from the evaluation. Models were evaluated based on standard metrics like accuracy, precision, recall and F1 score. Since our primary objective was to identify as many relevant interviews as possible, we attempted to optimise our prompts for recall. This reflects the underlying curation strategy, wherein false positives (i.e., irrelevant interviews flagged as relevant) could be manually reviewed, but false negatives risked permanent exclusion from the curated dataset. Ongoing efforts include testing summarisation methods on the few interviews exceeding the context 18 https://pypi.org/project/PyPDF2/ 19 The token limit for both Sauerkraut and Mistral is 128k, but we set it to 125k to maintain a 3k buffer for Instructor and other components. 20 Both extractive techniques involved sentence ranking based on token frequency weights, a method with roots in early automatic summarisation research from the 1950s and 1970s. In our experimental implementation, an extractive summary was generated using a sentence-ranking algorithm based on term frequency, reminiscent of early summarisation techniques [25]. The summarisation procedure involved the following steps: 1. Tokenisation and linguistic annotation using a SpaCy NLP model. 2. Removal of stop words, punctuation, numeric characters, and whitespace tokens. 3. Computation of term frequency for the remaining tokens, normalised by the maximum frequency observed. 4. Ranking of sentence importance by summing token frequencies within each sentence. 5. Iterative selection of top-ranked sentences, in descending order of importance, until the LLM token limit was reached. Sentences were then reordered to match their original sequence. 21 An example of an error message illustrating the types of issues encountered is: “Generated extractive summary for Record ID 5804 (Person_Name) is empty.” or “None.” 7 window and developing a more targeted prompt that further improves recall while maintaining an acceptable F1 score. Each trial combination was evaluated three times for consistency. For each model-prompt combination, we report the mean and standard error of each metric. Future work may also explore ensemble voting strategies or chunk-based processing for longer transcripts, e.g. sliding-window methods. Furthermore, some justifications produced by the models were uninformative, such as generic references to the interview subject without elaboration on their relevance. Additionally, we did not explicitly instruct the models to identify and state the interviewee’s academic background or research contributions, which may have limited the usefulness of the generated justifications for human reviewers. Results and Discussion Our results are summarised in table 1: The detailed prompt consistently improved recall across both models, though at the cost of lower precision. For example, llama-3.1-sauerkrautlm-70b-instruct with the detailed prompt correctly identified 97% of relevant interviews but had only a 51% precision rate, meaning that researchers would need to manually review a large number of false positives. The best balance was achieved by mistral-large-instruct with the detailed prompt, which yielded high recall (96%) and moderate precision (61%).22 Responses generated using detailed prompts were also longer on average, which may have implications for inference-time energy use. Prompts could potentially be shortened, or justifications removed entirely, to reduce token count. However, this would also defeat the purpose of using a technology capable of delivering justifications, a key motivation why we chose to use this technology over others alternatives. Model Prompt Evaluated Recall Precision F1 Score Avg. Tokens llama-70b basic 2459 0.93 0.60 0.73 69.3 llama-70b detailed 2457 0.97 0.51 0.67 78.3 mistral basic 2452 0.79 0.69 0.74 80.2 mistral detailed 2450 0.96 0.61 0.74 101.0 Table 1: Model performance across prompts (more details in appendix C). These values represent the averages of three runs with very little variation between them. 5 Implementing Environmental Impact Assessments in Digital Humanities To estimate the environmental impact of our workflow, we evaluated the available tools discussed in appendix A with regard to their usability in estimating the carbon footprint of our LLM usage scenarios. Ecologits is a promising solution for emissions estimation during inference, with an API and calculator interface.23 However, its coverage of models hosted by the GWDG infrastructure is still limited. CodeCarbon, while detailed, tracks only local machine emissions and is, thus, not applicable to our current workflow. Other tools, such as LLMCarbon [14] and the llmemissions. com calculator based on [31], offer theoretical estimates based on model size and usage but do not track live resource consumption. The environmental and legal implications of deploying large language models are playing an increasingly important role in research policy and institutional decision-making. The Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen (GWDG) serves as a pivotal infrastructure partner for our project in this regard, as it ensures that data remains within German jurisdiction, 22 See also tables 3 and 4 in Appendix C. 23 https://ecologits.ai/latest/ 8 A Survey of Available Environmental Impact Estimation Tools As a variety of approaches and tools have emerged to estimate the energy costs associated with LLMs, we want to provide a brief overview of those we found most relevant: CodeCarbon is a Python package that estimates carbon emissions based on the power consumption of CPU, GPU, and RAM. It operates by either measuring power usage in real time or approximating it based on hardware specifications.31 CodeCarbon then correlates these values with the carbon intensity of the electricity mix from the local grid or a specified cloud provider. When exact carbon intensity data is not available, it relies on estimates derived from national or regional electricity profiles. It tracks emissions from local machines and is thus not applicable to remote-hosted models but would be suitable if LLMs were locally deployed. EcoLogits concentrates on quantifying the environmental footprint of generative AI systems, particularly during the inference phase.32 It enables users to assess the environmental costs of interacting with large language models via API calls, covering major providers such as OpenAI and Anthropic. Grounded in life-cycle assessment methodologies, EcoLogits accounts not only for electricity used during the usage phase but also for environmental impacts associated with hardware production, transfer, and infrastructure in the embodied phase.33 The EcoLogits Calculator [32] offers a visual interface for estimating the footprint of individual interactions with generative AI, building on the underlying EcoLogits Python library.34 It presents metrics across multiple dimensions from electricity consumption, greenhouse gas emissions, abiotic resource depletion to primary energy use. To assist with interpretation, it contextualises results by comparing them to familiar activities such as walking, driving an electric vehicle, streaming video, or taking a transatlantic flight between Paris and New York. CallToChange is another Python library designed to log and estimate carbon emissions associated with API calls to large language models.35 It often requires as little as a single line of code to operate. Its core functionality is analysing LLM calls to compute an estimated CO2-equivalent emission, taking into account factors such as model size, token length, cloud provider or geographic location, and hardware type. While CodeCarbon provides system-level measurements and EcoLogits focuses on providerspecific model footprints, CallToChange targets emissions at the level of user interactions, making it especially useful for companies deploying AI services at scale. Collectively, these tools offer complementary perspectives on the environmental footprint of LLM use: CodeCarbon for system-wide energy profiling, EcoLogits for life-cycle-based analysis of provider services, and CallToChange for granular, interaction-level accounting. The AI Energy Score initiative responds to a gap in sustainability reporting in light of the growing environmental footprint of generative AI technologies, responding to a lack of consistent benchmarks for assessing the energy and material demands of individual models across a range of tasks.36 The AI Energy Score addresses this by proposing a systematic framework for evaluating energy efficiency, extending beyond direct emissions to incorporate broader environmental metrics such as water consumption, depletion of critical materials, and the production of electronic waste.37 31 https://codecarbon.io/ 32 https://ecologits.ai/ 33 While EcoLogits provides an API and GUI for estimating emissions, limitations for the presented project include incompatibility with some GWDG-hosted models. 34 https://huggingface.co/spaces/genai-impact/ecologits-calculator 35 https://calltochange-theta.vercel.app/ 36 https://huggingface.github.io/AIEnergyScore/ 37 At the core of the framework is a rating system that assigns models a score ranging from one to five stars. This score is based on GPU water consumption during inference across specific machine learning tasks. The framework currently encompasses ten standardised tasks, including text generation and summarisation, for which custom datasets have been developed to ensure consistency and reproducibility. Performance and environmental data are compiled into a public 15 B Prompt Variations B.1 Basic prompt The following are sentences extracted from an interview transcript with {person_name}. Based only on these sentences, please determine if the full interview is likely relevant to the topic: ‘Academic perspective from a profession within the environmental sciences’. For the evaluation of relevance especially consider the questions: Does this interview portray an academic perspective from a profession within the environmental sciences? Is this interview held with an academic expert within the field of the environmental sciences research? The official interview summary (if it exists) is added, followed by the extracted sentences from the interview. B.2 Detailed prompt The following are sentences extracted from an interview transcript with {person_name}. Based *only* on these sentences, please determine if the full interview is likely relevant to the topic: ‘Academic perspective from a profession within the environmental sciences’. When evaluating the interview, consider that environmental science contains various sub-disciplines such as, but not limited to: environmentalism, ecology, geography, geology, geoscience, climatology, metereology, oceanography, hydrology, toxicology, biodiversity, agriculture, atmospheric sciences, conservation and more. For the evaluation of relevance especially consider the questions: Does this interview portray an academic perspective from a profession within the environmental sciences? Is this interview held with an academic expert within the field of the environmental sciences research? Note that the interviewee’s educational background does not necessarily have to be in environmental sciences, but their work and insights should be relevant to the field. (*) Please rule out interviews that do not present the perspective of academic experts, even if they are involved in environmental sciences. For example, a conservationist without an academic background should not be considered relevant, while an academic regardless of their field discussing their research or work in conservation or oceanography would be relevant. (**) Please rule out interviews that do not present the perspective of academic experts, or those that focus on environmental movements, activism, policy, or on practical forms of knowing the environment. The official interview summary (if it exists) is added, followed by the extracted sentences. In both cases, if an official summary of the interview existed, we added a line ‘Additional context related to this interview: {official_summary}’ before the inputting the transcript sentences. During the iterative development of the prompting strategy, two specific sentences—marked here as (*) and (**)—were identified as contributing to a significant decline in recall. Though initially intended to clarify the topic and reinforce inclusion criteria, these sentences caused the language model to adopt an overly restrictive filtering behaviour, excluding a substantial number of interviews that were judged relevant by domain experts. Sentence (*) served as a topic clarification, while sentence (**) was derived from scholar-authored instructions provided to the LLM to guide leaderboard, which is updated twice per year. In addition to total electricity usage and associated carbon emissions, the framework reports on water consumption and other forms of resource use. 16 its classification decisions. In practice, both sentences led the model to apply more rigid criteria than intended, thus compromising its effectiveness in identifying relevant material. Given that the overarching goal was to maximise recall while maintaining an acceptable F1 score, these two sentences were removed from the prompt. Commenting them out improved the model’s ability to identify borderline or implicitly relevant interviews, which were often overlooked under the stricter instructions. This underscores the sensitivity of LLM output to prompt phrasing and the importance of empirical testing in prompt design. C Results Tables In this section, we present two tables: the first ‘master table’ is the full data for all runs, whereas the second one with mean +/- standard error represents the aggregated results. C.1 Model Performance Metrics 17 Table 3: Model Performance Metrics Model Prompt Evaluated Errors Relevant Irrelevant Accuracy Precision Recall F1 Avg Tokens (min–max) Sum Tokens Interviews llama-3.1-sauerkrautlm-70b-instruct basic 2459 32 403 2056 0.89 0.60 0.94 0.73 69.37 (17–315) 170583 2491 llama-3.1-sauerkrautlm-70b-instruct basic 2459 32 403 2056 0.88 0.59 0.93 0.73 69.15 (17–303) 170033 2491 llama-3.1-sauerkrautlm-70b-instruct basic 2459 32 403 2056 0.89 0.60 0.93 0.73 69.49 (17–344) 170869 2491 llama-3.1-sauerkrautlm-70b-instruct detailed 2457 34 402 2055 0.84 0.50 0.97 0.66 78.37 (18–387) 192551 2491 llama-3.1-sauerkrautlm-70b-instruct detailed 2459 32 403 2056 0.84 0.50 0.97 0.66 77.97 (18–417) 191720 2491 llama-3.1-sauerkrautlm-70b-instruct detailed 2455 36 402 2053 0.85 0.52 0.97 0.67 78.54 (17–342) 192810 2491 mistral-large-instruct basic 2453 38 399 2054 0.91 0.69 0.80 0.74 80.39 (15–472) 197187 2491 mistral-large-instruct basic 2453 38 399 2054 0.91 0.69 0.80 0.74 80.56 (15–472) 197615 2491 mistral-large-instruct basic 2451 40 398 2053 0.91 0.68 0.79 0.73 79.64 (16–338) 195193 2491 mistral-large-instruct detailed 2451 40 397 2054 0.89 0.61 0.97 0.75 101.65 (18–470) 249154 2491 mistral-large-instruct detailed 2447 44 398 2049 0.89 0.60 0.96 0.74 101.03 (18–595) 247216 2491 mistral-large-instruct detailed 2454 37 401 2053 0.89 0.61 0.94 0.74 100.19 (15–552) 245857 2491 18 Table 4: Metrics with mean +/- standard error Model Prompt Evaluated Errors Relevant Irrelevant Accuracy Precision Recall F1 Avg Tokens (min–max) Sum Tokens Interviews llama-3.1-sauerkrautlm-70b-instruct basic 2459.00 ±0.00 32.00 ±0.00 403.00 ±0.00 2056.00 ±0.00 0.89 ±0.00 0.60 ±0.00 0.93 ±0.00 0.73 ±0.00 69.34 ±0.10 (17.00 ±0.00–320.67 ±12.17) 170495.00 ±245.31 2491.00 ±0.00 llama-3.1-sauerkrautlm-70b-instruct detailed 2457.00 ±1.15 34.00 ±1.15 402.33 ±0.33 2054.67 ±0.88 0.84 ±0.00 0.51 ±0.01 0.97 ±0.00 0.67 ±0.00 78.29 ±0.17 (17.67 ±0.33–382.00 ±21.79) 192360.33 ±328.78 2491.00 ±0.00 mistral-large-instruct basic 2452.33 ±0.67 38.67 ±0.67 398.67 ±0.33 2053.67 ±0.33 0.91 ±0.00 0.69 ±0.00 0.79 ±0.00 0.74 ±0.00 80.19 ±0.28 (15.33 ±0.33–427.33 ±44.67) 196665.00 ±746.30 2491.00 ±0.00 mistral-large-instruct detailed 2450.67 ±2.03 40.33 ±2.03 398.67 ±1.20 2052.00 ±1.53 0.89 ±0.00 0.61 ±0.00 0.96 ±0.01 0.74 ±0.00 100.96 ±0.43 (17.00 ±1.00–539.00 ±36.67) 247409.00 ±956.64 2491.00 ±0.00 19 C.2 Ecological Impacts and Equivalents Table 20 Table 5: Ecological impact numbers and equivalents Model Prompt kWh GWP ADPE PE Equivalents meta-llama/Meta-Llama-3.1-70B-Instruct basic 1.93 – 2.35 1.3 – 1.57 3.16e-06 – 3.2e-06 17.46 – 21.1 Usage 1.93 – 2.35 1.26 – 1.53 1.7e-07 – 2.1e-07 16.9 – 20.54 Embodied – 0.04 – 0.04 2.99e-06 – 3e-06 0.56 – 0.56 walk: 89.08 – 107.68 km run: 59.39 – 71.79 km ev: 11.36 – 13.81 km stream: 20.28 – 24.51 h flight: 0.07 – 0.09 % mistralai/Mistral-Large-Instruct-2407 detailed 4.64 – 5.25 3.12 – 3.51 7.1e-06 – 7.15e-06 41.86 – 47.18 Usage 4.64 – 5.25 3.02 – 3.41 4.1e-07 – 4.6e-07 40.61 – 45.93 Embodied – 0.1 – 0.1 6.69e-06 – 6.69e-06 1.25 – 1.25 walk: 213.57 – 240.73 km run: 142.38 – 160.49 km ev: 27.31 – 30.88 km stream: 48.62 – 54.79 h flight: 0.18 – 0.2 % mistralai/Mistral-Large-Instruct-2407 basic 3.67 – 4.16 2.47 – 2.78 5.62e-06 – 5.66e-06 33.13 – 37.34 Usage 3.67 – 4.16 2.39 – 2.7 3.2e-07 – 3.7e-07 32.14 – 36.35 Embodied – 0.08 – 0.08 5.29e-06 – 5.3e-06 0.99 – 0.99 walk: 169.02 – 190.52 km run: 112.68 – 127.01 km ev: 21.61 – 24.44 km stream: 38.48 – 43.36 h flight: 0.14 – 0.16 % mistralai/Mistral-Large-Instruct-2407 detailed 4.61 – 5.21 3.09 – 3.48 7.04e-06 – 7.1e-06 41.53 – 46.82 Usage 4.61 – 5.21 2.99 – 3.39 4e-07 – 4.6e-07 40.29 – 45.57 Embodied – 0.1 – 0.1 6.64e-06 – 6.64e-06 1.24 – 1.24 walk: 211.91 – 238.86 km run: 141.27 – 159.24 km ev: 27.09 – 30.64 km stream: 48.24 – 54.36 h flight: 0.17 – 0.2 % continued on next page 21 Model Prompt kWh GWP ADPE PE Equivalents meta-llama/Meta-Llama-3.1-70B-Instruct basic 1.93 – 2.34 1.3 – 1.57 3.15e-06 – 3.19e-06 17.4 – 21.04 Usage 1.93 – 2.34 1.25 – 1.52 1.7e-07 – 2.1e-07 16.84 – 20.48 Embodied – 0.04 – 0.04 2.99e-06 – 2.99e-06 0.56 – 0.56 walk: 88.79 – 107.33 km run: 59.2 – 71.55 km ev: 11.33 – 13.77 km stream: 20.22 – 24.43 h flight: 0.07 – 0.09 % meta-llama/Meta-Llama-3.1-70B-Instruct detailed 2.18 – 2.65 1.47 – 1.77 3.57e-06 – 3.61e-06 19.71 – 23.82 Usage 2.18 – 2.65 1.42 – 1.72 1.9e-07 – 2.3e-07 19.07 – 23.19 Embodied – 0.05 – 0.05 3.38e-06 – 3.38e-06 0.63 – 0.63 walk: 100.55 – 121.55 km run: 67.04 – 81.03 km ev: 12.83 – 15.59 km stream: 22.89 – 27.66 h flight: 0.08 – 0.1 % meta-llama/Meta-Llama-3.1-70B-Instruct detailed 2.17 – 2.64 1.46 – 1.77 3.56e-06 – 3.6e-06 19.62 – 23.72 Usage 2.17 – 2.64 1.41 – 1.72 1.9e-07 – 2.3e-07 18.99 – 23.09 Embodied – 0.05 – 0.05 3.37e-06 – 3.37e-06 0.63 – 0.63 walk: 100.12 – 121.02 km run: 66.75 – 80.68 km ev: 12.77 – 15.53 km stream: 22.79 – 27.54 h flight: 0.08 – 0.1 % meta-llama/Meta-Llama-3.1-70B-Instruct basic 1.94 – 2.35 1.3 – 1.57 3.17e-06 – 3.21e-06 17.49 – 21.14 Usage 1.94 – 2.35 1.26 – 1.53 1.7e-07 – 2.1e-07 16.93 – 20.58 Embodied – 0.04 – 0.04 3e-06 – 3e-06 0.56 – 0.56 walk: 89.23 – 107.86 km run: 59.49 – 71.91 km ev: 11.38 – 13.84 km stream: 20.32 – 24.55 h flight: 0.07 – 0.09 % meta-llama/Meta-Llama-3.1-70B-Instruct detailed 2.18 – 2.65 1.47 – 1.78 3.58e-06 – 3.62e-06 19.74 – 23.85 Usage 2.18 – 2.65 1.42 – 1.73 1.9e-07 – 2.3e-07 19.1 – 23.22 continued on next page 22 Model Prompt kWh GWP ADPE PE Equivalents Embodied – 0.05 – 0.05 3.38e-06 – 3.39e-06 0.63 – 0.63 walk: 100.69 – 121.71 km run: 67.13 – 81.14 km ev: 12.84 – 15.61 km stream: 22.92 – 27.7 h flight: 0.08 – 0.1 % mistralai/Mistral-Large-Instruct-2407 basic 3.68 – 4.16 2.47 – 2.79 5.63e-06 – 5.67e-06 33.2 – 37.42 Usage 3.68 – 4.16 2.39 – 2.71 3.2e-07 – 3.7e-07 32.21 – 36.43 Embodied – 0.08 – 0.08 5.31e-06 – 5.31e-06 0.99 – 0.99 walk: 169.39 – 190.93 km run: 112.93 – 127.29 km ev: 21.66 – 24.5 km stream: 38.56 – 43.46 h flight: 0.14 – 0.16 % mistralai/Mistral-Large-Instruct-2407 detailed 4.58 – 5.18 3.08 – 3.47 7e-06 – 7.06e-06 41.31 – 46.56 Usage 4.58 – 5.18 2.98 – 3.37 4e-07 – 4.6e-07 40.07 – 45.32 Embodied – 0.1 – 0.1 6.6e-06 – 6.6e-06 1.24 – 1.24 walk: 210.74 – 237.55 km run: 140.5 – 158.36 km ev: 26.94 – 30.48 km stream: 47.97 – 54.06 h flight: 0.17 – 0.2 % mistralai/Mistral-Large-Instruct-2407 basic 3.64 – 4.11 2.44 – 2.75 5.56e-06 – 5.6e-06 32.79 – 36.96 Usage 3.64 – 4.11 2.36 – 2.67 3.2e-07 – 3.6e-07 31.81 – 35.98 Embodied – 0.08 – 0.08 5.24e-06 – 5.24e-06 0.98 – 0.98 walk: 167.31 – 188.59 km run: 111.54 – 125.73 km ev: 21.39 – 24.2 km stream: 38.09 – 42.92 h flight: 0.14 – 0.16 % 23 D Methodology for Estimating Environmental Impact Following the Ecologits Framework To approximate the environmental impact of our model runs, we reproduced the calculations implemented by the Ecologits Calculator, drawing upon the methodological explanations provided in its documentation.38 In order to conduct these estimates, we developed a custom wrapper that emulated the internal structure of the Ecologits wrapper. This allowed us to input our own model parameters while selecting from the pre-defined models listed in the tool. For the purposes of our estimations, we matched our usage to the closest available configurations. Specifically, we selected MistralAI / mistral large instruct, i.e. mistralai/Mistral_large_instruct-2407. In the case of the Sauerkraut variant of Meta’s LLaMA 3.1 70B model, developed by the German company Vago Solutions, we mapped it to the most comparable option in the tool: Meta’s LLaMA 3.1 70B. Although the Sauerkraut model is distinct from Meta’s original, this approximation was necessary given the available presets in the Ecologits framework. Our analysis relies on various environmental metrics derived from three principal impact indicators provided by Ecologits: Primary Energy (PE), Total Energy (in kWh), and Global Warming Potential (GWP, measured in kgCO2eq). These are used to generate intuitive equivalencies: To convey energy consumption in familiar physical terms (as in the EcoLogits Calculator [32]), we translated Primary Energy values (given in megajoules, MJ) into estimated distances for walking and running.39 After converting MJ to kilojoules (kJ) by a factor of 1,000, we applied energy expenditure rates: •Walking: 196 kJ/km, such that the walking distance is calculated as: Distance (km) = 38 See ‘Methodology’ in: https://huggingface.co/spaces/genai-impact/ecologits-calculator. 39 Our environmental impact estimations are aligned with the methodological assumptions outlined in the ‘Methodology’ tab of the Ecologits Calculator [32]. The following conversion parameters were used in the derivation of equivalent activity metrics: For physical activity equivalents, energy expenditures are based on average values associated with movement at specific speeds. Walking is calculated at 196 kJ/km, corresponding to a speed of 3 km/h, while running is assumed at 294 kJ/km, associated with a pace of 10 km/h. Electric vehicle (EV) distance is based on an average consumption rate of 0.17 kWh per kilometre and used to estimate how far a standard electric vehicle would travel on the same energy budget as that consumed by a given model inference run. Streaming time equivalence is based on the global warming potential (GWP) of the request, with 1 kgCO2eq considered equivalent to 15.6 hours of video streaming. For comparisons with air travel, the calculator estimates that a return flight between Paris and New York City emits 1,770 kgCO2eq per passenger. Assuming an average passenger load of 100 per flight, the emissions per flight are scaled accordingly. Our implementation of the ‘flight percentage’ calculation differs from the version used on the Ecologits Calculator web interface (https://huggingface.co/spaces/genai-impact/ecologits-calculator), which relates the flight emissions to the question: “What if 1% of the planet does this request every day for 1 year?” In this model, the impact of a single request is scaled by the factor 0.01 ×8billion people ×365 days, yielding a hypothetical global-use scenario. We instead based our values on the explanation provided under the ‘Methodology’ tab, specifically under the section on the number of Paris to New York City return flights, where it is stated: We compare the GHG emissions (scaled) of the request and of a return flight Paris ↔ New York City. From impactco2.fr (https://impactco2.fr/outils/comparateur?value=1& comparisons=&equivalent=avion-pny) we consider that a return flight Paris →New York City → Paris for one passenger emits 1,770 kgCO2eq and we consider an overall average load of 100 passengers per flight. We divide the scaled GHG emissions by this value to get the equivalent number of return flights. Our version, however, does not apply the global scaling step. We compare the GHG emissions from a single execution of our workflow with the 1,770 kgCO2eq emitted by a return flight for one passenger. We do not scale this further to estimate the number of flights that would result if the request were executed by 1% of the global population daily for a year. This is because our use case is not intended to be run repeatedly or widely deployed. The process is designed to be executed once (or a small number of times, once sufficient output quality is achieved) in order to support scholars in filtering and analysing data – not as a continually repeated step in an automated pipeline. These parameters, drawn directly from the Ecologits Calculator’s own documentation, inform all derived environmental equivalencies presented in our analysis. However, it must be acknowledged that many assessments of the ecological impact of specific activities, such as estimates of the environmental impact of an intercontinental flight per passenger, are themselves controversial. 24