scieee AI-readable full text Open interactive document viewer

Processable and Computable Media in TEI: Usecases and Strategies

Roeder, Torsten; Shtohryn, Tomash; Herbst, Yannik

Abstract

Processable and Computable Mediain TEI: Usecases and Strategies Torsten Roeder¹ | 0000-0001-7043-7820 | [email protected] Shtohryn¹ | 0009-0000-4597-603X | [email protected] W. Herbst¹ | 0000-0002-6547-9599 | [email protected] ¹Julius-Maximilians-Universität Würzburg, Germany Beyond Readability? In everyday usage, “text” typically refers to physical or electronic inscriptions—material or digital—that are intended for human reading. Our notion of textual heritage predominantly encompasses classical forms such as literature, correspondence, journals, and newspapers, all seemingly bound to natural language. The TEI Guidelines offer extensive support for encoding such textual complexity, accommodating a wide range of human-created text types. Although such texts are, in principle, amenable to processing by computational rulesets or human interpretation, this kind of processing is not usually the primary focus. This prompts a set of critical questions: Is “text”—and by extension, the TEI—necessarily bound to human natural language? Conversely, are there classes of textual heritage that require mechanical or rule-based processing either before, during, or after human or machine reading? If so, how can these be encoded within the existing TEI framework, and what expansions might be necessary? Furthermore, how can the rulesets for such processing themselves become integral to the preservation of cultural heritage? Examples The questions raised may appear broad, if not esoteric, thus it is appropriate to illustrate them with concrete examples. Consider a multiple-choice examination sheet completed by a student and graded by a teacher. The form itself encodes an implicit ruleset that governs how the examination is to be completed and assessed—thus undergoing multiple stages of processing to produce a graded result. Another example stems from early digital literature: HyperCard-based works require specific computational environments (e.g., pre-OS X Macintosh systems) for proper functionality. The user’s interaction with these works is defined by both the hardware and software, raising the question of whether essential aspects of historical digital interfaces can be preserved through TEI encoding (cf. Ensslin 2007 for a more general discussion on early hypertext fiction). Similarly, early digital periodicals such as “diskmags”—early software magazines distributed on floppy disk—often employed custom character sets and graphical conventions that defy standard extraction and representation methods. Here, knowledge of the original processing mechanisms is essential for reconstructing the intended presentation of text, images, and audio content (cf. Shtohryn 2025). Emails, a contemporary form of correspondence, further complicate notions of text (cf. Beshero-Bondar and Bauman 2024). Beyond metadata such as sender and recipient information, the software environment critically influences the creation, display, and transmission of emails. Variations between client applications and server-side modifications introduce layers of textual genetics and performance dynamics that are relevant for scholarly documentation. Finally, historical program code published in print media (“paperware”) exemplifies the hybrid nature of text and process (cf. Höltgen 2014). Such code can be “read” and executed by both humans and machines. Punchcards, featuring both computer-readable perforations and human-readable printed text, reinforce the idea that digitality often straddles material and electronic forms (cf. Feichtinger 2023). Perspectives These examples collectively demonstrate that distinctions between "digital" and "material" documents are secondary when considering a document's intentional processability. While this poster centers primarily on born-digital heritage, it recognizes the blurred boundaries between digital and material media. Recent TEI developments—such as the introduction of the <post> element by the special interest group for “Computer-Mediated Communication” in 2025 (cf. TEI 2025)—illustrate growing recognition of these evolving forms. Given the increasing volume of born-digital cultural heritage and its vulnerability to hardware decay, software obsolescence, corporate restrictions, and political censorship, it is urgent to develop strategies for preserving processable and computable forms of text. This poster will present the aforementioned examples to initiate a discussion on extending TEI’s capabilities in this area. We aim to establish a new TEI Special Interest Group dedicated to processable and computable media—recognizing that the traditional notion of “text” may not be sufficiently inclusive in this context. Our initial focus will be on utilizing existing guidelines, proposing semantic extensions where necessary, and, where appropriate, developing new attributes or elements. References Beshero-Bondar, Elisa; Bauman, Syd (2024): Can we apply the new CMC chapter to the TEI Listserv Archives? An experiment with TEI for Correspondence and Computer-Mediated Communication, TEI Conference 2024. https://www.conftool.pro/tei2024/[…] Ensslin, Astrid (2007): Canonizing Hypertext. Explorations and Constructions. Feichtinger, Moritz (2023): Annotation, Simulation und Analyse eines historischen Datenbanksystems. In: Burghardt, M. & Weiß, C. (eds.): Lecture Notes in Informatics (LNI), Gesellschaft für Informatik. Höltgen, Stefan (2014): Humanities of the Digital. Philologische Perspektiven auf Source Codes als Beitrag einer computerarchäologischen Knowledge Preservation. In: Bartelmus, M./Nebrig, A. (eds.): Digitale Schriftlichkeit. Programmieren, Prozessieren und Codieren von Schrift, p. 207–229. Shtohryn, Tomash (2025): Semi-automatisierte Erzeugung eines Textkorpus von deutschsprachigen Diskettenmagazinen für das Heimcomputersystem Commodore 64 im TEI-Format, Universität Würzburg. https://doi.org/10.25972/OPUS-40282 TEI (2025): Computer-mediated Communication. In: TEI: Guidelines for Electronic Text Encoding and Interchange. P5 Version 4.9.0, 24th January. https://tei-c.org/release/doc/tei-p5-doc/en/html/CMC.html

Full text

Processable and Computable Media in TEI: Usecases and Strategies University of Würzburg Centre for Philology und Digitality Research Unit DACHS Text Encoding Initiative TEI Special Interest Group Computable Text and Media Torsten Roeder Tomash Shtohryn Yannik Herbst Correspondence Chess FORMS, INTERACTION & RULES form layers input areas complex layouts input types selection modes conditional fields layer with filled-out form pointers to template fields interpreted content transposed format How about a nice game of chess? Software Magazine COVERS, CONTAINERS & ENCODINGS source description for hybrid media descriptors for media and contained files additional descriptors for analytical data inclusion of file system references to full source descriptions transcription of conceptional layer as on screen translated individual and special characters Commodore 64 Programming Manual COMPUTER CODE, ENVIRONMENT & OUTPUT extended descriptors for environments manufacturer for objects from serial production use existing elements for programming languages multiple layers within choice: logical vs. conceptional analytic markup for historical program code use @xml:lang for more than RFC-3066 Atari 800XL