scieee AI-readable full text Open interactive document viewer

AI and Contemporary Historiography: From Instrumental Use to Methodological Transformation (Preprint v2, November 2025)

Werner, Wiktor

Abstract

Status: Preprint (non-peer-reviewed working paper). This preprint explores how artificial intelligence (AI) and machine learning (ML) are reshaping the methodology of contemporary historiography. It distinguishes three levels of AI integration into historical research: (1) technical reconstruction of sources, (2) factual extraction and verification through NLP and machine learning, and (3) interpretative modeling and hypothesis generation via probabilistic reasoning. By referencing both classic and current works—from Alan Turing and Marvin Minsky to recent neural-network applications such as the Ithaca project—the paper situates AI within a longer history of epistemological transformation. It argues that AI-assisted historiography represents a gradual transition from narrative hermeneutics to data-driven empiricism, analogous to the earlier methodological shifts seen in archaeology and cliometrics. A brief empirical illustration (Werner & Wilczyński 2024) demonstrates the application of classification algorithms to prosopographic data from the PRL Security Service archives. The paper also discusses ethical risks, including the misuse of generative models, and references Marnie Hughes-Warrington’s Artificial Historians (2025) for context. Disclaimer:– This version (v1, November 2025) is a pre-submission working draft shared for scholarly feedback.– It has not undergone peer review and should not be cited as a published article.– A revised version will be submitted to a peer-reviewed journal

Full text

Wiktor Werner (AMU, Poznan, Poland) AI and Contemporary Historiography: From Instrumental Use to Methodological Transformation (Preprint, November 2025) 10.5281/zenodo.17593628 Keywords machine learning, historiography, social network analysis, NLP, digital history, AI, archeology Summary (EN) This preprint explores how artificial intelligence (AI) and machine learning (ML) are reshaping the methodology of contemporary historiography. It distinguishes three levels of AI integration into historical research: (1) technical reconstruction of sources, (2) factual extraction and verification through NLP and machine learning, and (3) interpretative modeling and hypothesis generation via probabilistic reasoning. By referencing both classic and current works—from Alan Turing and Marvin Minsky to recent neural-network applications such as the Ithaca project—the paper situates AI within a longer history of epistemological transformation. It argues that AI-assisted historiography represents a gradual transition from narrative hermeneutics to data-driven empiricism, analogous to the earlier methodological shifts seen in archaeology and cliometrics. A brief empirical illustration (Werner & Wilczyński 2024) demonstrates the application of classification algorithms to prosopographic data from the PRL Security Service archives. The paper also discusses ethical risks, including the misuse of generative models, and references Marnie Hughes-Warrington’s Artificial Historians (2025) for context. Disclaimer: – This version (v1, November 2025) is a pre-submission working draft shared for scholarly feedback. – It has not undergone peer review and should not be cited as a published article. – A revised version will be submitted to a peer-reviewed journal Development of Artificial Intelligence and Machine Learning The initial research on the specifics of machine learning is linked to the work of Alan Turing, who introduced the (then theoretical) concept of the computer as a universal computing machine. He described the operation of the computer as the cooperation of an infinite tape, a read–write head, and a finite set of internal machine states (Turing 1936). In his article "Computing Machinery and Intelligence" (Turing 1950: 433-460), published in 1950, he also outlined the theoretical framework for the potential development of machine intelligence. His reflection involved the reconceptualization of the very idea of "thinking"—as a computational process of data processing using algorithms. Turing also attempted to define an empirical test (trial) of machine intelligence—as the ability to imitate intelligent human actions, particularly communication practices, and developed a logical model of machine thinking called the "Turing machine," which is a formal description of a mechanism capable of performing any (possible) algorithmic procedure. The direction set by Turing was continued in the reflections and research practices of Marvin Minsky. In 1951, he constructed tthe first analog neural network model— SNARC (Stochastic Neural Analog Reinforcement Calculator)—one of the earliest machines simulating learning behavior—built from about 3,000 vacuum tubes, simulating a network of 40 neurons. SNARC could "learn" to navigate through a virtual maze by searching for correct paths. In 1959, together with John McCarthy, Minsky launched the Artificial Intelligence Project at MIT, which was formally reorganized in 1970 as the MIT Artificial Intelligence Laboratory (Minsky, 1959). In his work "Steps Toward Artificial Intelligence," Minsky argued that reasoning can be modeled as a sequence of symbolic operations — suggesting that intelligent behavior, whether human or artificial, could be analyzed in computational terms. (Minsky, 1961). In his book "Computation: Finite and Infinite Machines" (Minsky, 1967), he indicated that intelligence consists of a network of cooperating processes—a preview of his later theory "The Society of Mind" (Minsky, 1986), in which consciousness is an emergent property of complex interactions of many simple cognitive modules, leading him to the concept of neural networks—computational algorithms operating on a nonlinear principle of parallel process cooperation. In the 1960s and 1970s, Minsky, along with Seymour Papert (Minsky, M. & Papert, 1969), designed and studied perceptrons—early neural networks. Single-layer perceptrons were linear models (nonlinearity and parallel process cooperation appeared only in later multi-layer networks). Early neural-network (Minsky & Papert, 1969) were thus limited to linearly separable problems. However, it was neural networks that accelerated the development of machine learning in the 1980s. In the latter half of that decade, a breakthrough method was developed—the backpropagation algorithm (Rumelhart, Hinton, Williams, 1986). It was a mechanism that allowed neural networks to correct errors by sending a signal from the output layer to earlier neurons, essentially "reversing" the error to the point where it occurred, in order to correct the procedures that led to it (Werbos, 1974). This enabled neural networks to recognize nonlinear relationships between data and also independently "learn from their own mistakes" through iterative data analysis (Bishop C. M., 1995). The technique of backpropagation, developed in the 1980s, found application in natural language processing (NLP) during the 1990s and 2000s. Initially, it was used in recurrent neural networks (RNN), which allowed the differentiation of sequence events – such as words in a sentence (Elman, 1990) – and later in long short-term memory models (LSTM), capable of maintaining context even in long sequences like complex sentences or time series, thanks to "memory gates" (Hochreiter & Schmidhuber, 1997). In these models, the network learns to predict subsequent sequence elements, and the prediction error is propagated backwards to adjust weights and better reflect contextual dependencies. This mechanism enables networks to recognize complex structures and regularities, not as logical rules, but as hidden probabilistic patterns. The architecture of ChatGPT, which is based on Transformer algorithms (Vaswani et al., 2017), represents another step in this tradition. Although different from classic RNNs, it still uses backpropagation as a learning method. The difference is that instead of predicting a sequence over time, transformers learn the relationships between all tokens simultaneously (also known as self-attention). Each network layer simultaneously transforms input data non-linearly and backpropagates the error signal until the entire model learns to predict the most probable next word. It is through such learning that models like GPT-3 (2020), GPT-4 (2023), and other contemporary LLM architectures have been developed, neural networks utilizing over 100 billion parameters that learn the structure of language through millions of iterations using the backpropagation mechanism (Vaswani, A. et al., 2017). Artificial Intelligence and Historical Research The mere existence of a particular technology does not imply its uniform adoption across all domains of human activity. However, there are technologies that exert such a strong influence on operations—increasing their efficiency—that ignoring them is impossible or only possible for a very limited time. Such technologies include writing, printing, digitization, and now, artificial intelligence. Even the relatively conservative, almost inertial field of knowledge that is historiography (Werner, Falkowski, 2024) does not reject the cognitive benefits derived from the application of AI technology (although, of course, individual researchers may effectively ignore its existence for a very long time). In practice, this means that research utilizing AI algorithms and the tools that employ them, even if not conducted by all historians, will be carried out by some—with visible beneficial effects. The applications of artificial intelligence methods and tools in historical research are currently noticeable at three levels: technical applications (basic level – 1), establishing new facts (factual level – 2), and establishing relationships between facts (interpretative level – 3). On the basic level, digital tools have begun to be used in the process of preparing source resources for participation in historical research. These include well-known tasks related to reading heavily damaged manuscripts or manuscripts that, due to damage, were not readable at all. A spectacular success was the pioneering use of computer tomography and fluorescence analysis for the virtual reading of charred and partially carbonized scrolls and books without opening them, which would have destroyed the objects (Albertin et al., 2015). This demonstrated the benefits of broad implementation of digital methods in historical research. A subsequent breakthrough was the introduction of AI tools (and not just digital technology), which surpassed other previously used methods in effectiveness—including older digital methods. Noteworthy are the results obtained in the Ancient Lives project (University of Oxford, Zooniverse), where the effectiveness of CNN neural networks was compared with more "traditional" OCR technology (technology that transforms photographs or scanned documents into editable, digital text, usually developed by specific scanner manufacturers) for recognizing Greek letters in low-quality papyri. CNN substantially outperformed baseline OCR on Ancient Lives papyri (see Swindall et al., 2021/2022). An important example of AI success in historical research is the Ithaca project, developed by a team from Google DeepMind and the University of Oxford (Assael et al., 2022), where an AI model was created that learned from a corpus of 78,608 Greek inscriptions. This model can reconstruct missing text fragments with an accuracy of up to 62%, date inscriptions with an average error of about 30 years, and determine geographical origins with 71% accuracy. The Ithaca model is based on a deep neural network architecture of the Transformer type, similar to those used in contemporary language models (e.g., GPT). The learning involved analyzing linguistic and topographic context—missing fragments are predicted based on the probability of symbol sequences in the surrounding context. Technically, Ithaca uses the self-attention mechanism and backpropagation to minimize text prediction errors. Another example is the work of Israeli researchers using recurrent neural networks (RNN) to automatically restore missing fragments of cuneiform tablets from Babylon (Achaemenid period). The model was trained on a corpus of transliterated Akkadian texts, which allowed for accurate prediction of missing characters and words: from 65% to 95% (Fetaya, Lifshitz, Aaron, Gordin, 2020). The text corpus was prepared using functions and modules of the Python programming language (considered a kind of lingua franca for artificial intelligence systems). Benefits arising from the application of digital technology, with particular emphasis on self-learning algorithms (AI and machine learning), are so evident that remaining at the "technical" level no longer seems possible. New technology is thus being introduced into further stages of the research process, meaning it not only participates in developing a source to a state in which it can be read by a historian but also takes part in verifying hypotheses about facts based on non-obvious regularities and connections in data that a historian (a human) would not be able to perceive with a mind lacking digital instrumentation. We are therefore operating at the second, factual level where tools of NLP play a significant role: NER algorithms (Named Entity Recognition) for recognizing names of places, people, and events, as well as stylometric ones for determining authorship of texts and sentiment analysis (emotional charge of texts). Very interesting stylometric studies of the text of the Polish Chronicle by Anonymous ("Gallus") were conducted by Maciej Eder (Eder, 2015), who, using NLP methods, confirms Tomasz Jasiński’s hypothesis about the Italian identity of Anonymous. Significant importance can also be attributed to relatively simple machine learning algorithms (CART – classification and regression tree, decision tree, random forest) belonging to the group of supervised learning algorithms, where the model learns based on labeled training data. For example, in the classification of historical texts, an algorithm may receive a set of texts labeled with categories: "document," "letter," "press article," and then learn to recognize these categories in new sets. Zbigniew Wilczyński and Wiktor Werner, in their studies on the activities of the PRL Security Service (Department VI in Szczecin), applied a single "classification and regression tree" (CART) to recognize the model types of SB collaborators and the relationships between specific parameters of their description (age, profession, education, gender, etc.) and their tendency to make personal reports (Werner, Wilczyński, 2024). For studying all relationships between phenomena (people, institutions, states, cities, etc.), network analysis algorithms (Social Network Analysis) are useful, which use graph theory as their basis. The methodology of SNA and graph theory has already found application in many historical studies covering a wide range of different sources. Anderson Pereira Antunes's research (Antunes, 2021) shows how one can analyze social structures and personal dependencies in the history of science. In his work Social Network Analysis in the History of Sciences, Antunes used the Gephi program and the networkX library (Python) to analyze relationships among participants of 19th-century scientific expeditions (Antunes, 2023). Network nodes represented people, and edges represented scientific contacts, correspondence, and co-authorships. The results were visualized and analyzed using indicators: centrality (betweenness, eigenvector centrality), network density and modularity, cluster analysis (community detection). Katherina Kaska examined the relationships between scribes and manuscripts in Cistercian monasteries Heiligenkreuz, Zwettl, and Baumgartenberg in the 12th century. The aim was to understand the structure of scribes' collaboration and the development of the book production "ecosystem" (Kaska, 2023). Ruedi Epple studied socio-political networks in the canton of Basel-Landschaft (Switzerland) at the end of the 19th century. He showed how local elites, clergy, and activists created networks of influence in the context of initiatives and political voting (Epple, 2022). Louis Bissières examined merchant contact networks and credit connections in North America at the end of the 18th century using the multilayer networks method, which allowed for the presentation of temporal dependencies (Bissières, 2023). In the context of the real results being achieved, it is not surprising to notice the already visible trend of elevating AI tools to an even higher level of research and engaging this technology in the interpretation of facts (e.g., to search for causal relationships among them) and historical modeling. At the third level (interpretative), we encounter models capable of assisting in generating hypotheses about causality, testing counterfactual scenarios, and simulating historical processes. Here, we find: agent-based models (ABM) – simulating social behaviors and decisions (e.g., migrations, conflicts), Bayesian networks – for estimating the probability of events with incomplete data, and Large Language Models (LLM) – for integrating narrative and factual data (e.g., GPT used for classifying temporal relations in archives). An example of such research is the pioneering and currently unique work of Atin Basuchoudhary, James T. Bang, John David, and Tinni Sen titled "Identifying the Complex Causes of Civil War" (Basuchoudhary et al., 2021). The authors use decision tree models to identify the structural factors of civil wars; the model applied there "learns" causal dependencies from extensive historical and economic datasets. In the presented approach, machine learning methods are treated not only as instruments of quantitative analysis but as analogs of cognitive processes inherent in historical research. The mentioned authors developed the Empirically Informed Covariate Selection (EICS) procedure, combining variable selection, causality analysis, and dependency visualization in a Bayesian model. Their research scheme includes: data imputation using the Miss Forest method (a method based on Random Forest, which is a set of decision trees that collectively learn to predict the missing values of each variable based on the other columns in the dataset), a procedure for eliminating insignificant features, Recursive Feature Elimination (RFE), analysis of variable impact (PDP), and causal modeling based on Bayesian statistics (BART). MissForest corresponds to the historical process of reconstructing incomplete sources. The model does not fill these gaps arbitrarily but restores the missing values while maintaining nonlinear relationships between variables (Basuchoudhary et al., 2021, p. 41). It is thus an algorithmic equivalent of reconstructing "possible states of the past" – a form of probabilistic hermeneutics, where absence becomes part of the cognitive process rather than a blocking factor. Recursive Feature Elimination (RFE) constitutes a heuristic process akin to the selection of meanings in historical interpretation (Basuchoudhary et al., 2021, p. 53). In humanistic terms, it is a procedure for distinguishing significant information from random data – a digital version of traditional source interpretation that selects data and limits the risk of overinterpretation. Partial Dependence Plots (PDP) visualize how a change in one variable affects the model outcome while keeping other factors constant (Basuchoudhary et al., 2021, p. 72). Finally, Bayesian Additive Regression Trees (BART) introduce causality modeling based on probability distributions under variable uncertainty. BART consists of many small decision trees combined additively (i.e., their predictions sum up). Each tree is very shallow – responsible for a "local" fragment of the data space. The algorithm draws a set of such trees, trains them on different fragments of the dataset, and combines their results into a probability distribution of the final outcome (Basuchoudhary et al., 2021, p. 83). In epistemological terms, this corresponds to the historian's awareness that knowledge of the past always remains probabilistic and sensitive to randomness. The model does not eliminate uncertainty but makes it an element of description. The set of methods MissForest – RFE – PDP – BART can thus be treated as a proposal for computational hermeneutics, where algorithms do not replace interpretation but mimic its structure: they reconstruct missing links in the message, select the excess of meanings, explore counterfactual relationships, and capture knowledge in terms of probability. In this way, machine learning methods become not only a technique for data analysis but a model of thinking about the past, where historical and computational cognition meet in the area of reflection on the uncertain image of the world obtained from an incomplete dataset. The next stage in the adoption of AI in historiography may involve not their use, but their misuse. It is also impossible to ignore issues related to the engagement of generative algorithms associated with large language models (LLM) for creating various forms of historical narratives (which are not accompanied by an adequate research process!) and iconographic pseudo-sources. It should be remembered that generative algorithms do not differentiate between past and present, as well as truth and falsehood. They learn from large datasets to generate outputs consistent with patterns found in their training material. Where we might see truth or falsehood, machines see consistency or its lack. Generally, therefore, the growing density of AI-generated information environments poses many threats to professional historiography and broader historical culture. As Marnie Hughes-Warrington emphasizes in her book "Artificial Historians" (HughesWarrington et al, 2025), the logic of artificial intelligence (as we know, based on geometric convergences between data vectors) fundamentally differs from the logic of Turing, A.M. (1936) ‘On Computable Numbers, with an Application to the Entscheidungsproblem.’ Proceedings of the London Mathematical Society, 42(1), pp. 230–265. Turing, A.M. (1950) ‘Computing Machinery and Intelligence.’ Mind, 49, pp. 433–460. Vaswani, A. et al. (2017) ‘Attention Is All You Need.’ Advances in Neural Information Processing Systems (NeurIPS). Werbos, P. (1974) Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences. PhD thesis, Harvard University. Werner, W. & Falkowski, T. (2024) ‘Problem naukowości historiografii.’ In: Kuklo, C. & Walczak, W. (eds.) Człowiek twórcą historii. T. 6: Warsztat nowoczesnego humanisty historyka na progu XXI wieku, cz. 2. Białystok: Uniwersytet w Białymstoku, pp. 155– 178. ISBN 978-83-67846-13-4. White, H. (2000) ‘An Old Question Raised Again: Is Historiography Art or Science? (Response to Iggers)’, Rethinking History, 4(3), pp. 391–406. doi: 10.1080/136425200456958a. Wilczyński, Z. & Werner, W. (2024) ‘Portret tajnego współpracownika „bezpieki”. Badania źródeł archiwalnych SB z wykorzystaniem algorytmów uczenia maszynowego oraz analizy prozopograficznej.’ Historia i Polityka, 49(56), pp. 9–25. doi: http://dx.doi.org/10.12775/HiP.2024.019 dr hab. Wiktor Werner, prof. UAM Wydział Historii UAM, Poznań ORCID: 0000-0002-3004-6021 e-mail: [email protected]