scieee AI-readable full text Open interactive document viewer

Heritage Data Processor (HDP): A Modular Architecture for Processing and Persistently Storing Multimodal Cultural Heritage Data

Ukolov, Dominik

Abstract

This working paper presents preliminary findings from my ongoing doctoral research at Leipzig University and is part of a monograph-based PhD project in musicology. It has not been peer reviewed and substantial parts may later appear in revised and extended form in my doctoral dissertation and any subsequent book publication. The work was also carried out in the context of the EU-funded project “3D Big Data to enrich the European Data Space for Cultural Heritage” (3DBigDataSpace). Please cite this Zenodo record when referring to these results, and note that later versions may contain significant updates or corrections.

Full text

Heritage Data Processor (HDP): A Modular Architecture for Processing and Persistently Storing Multimodal Cultural Heritage Data Working Paper [v1] 2025 Dominik Ukolov Digital Humanities (Image/Object), Friedrich-Schiller-University Jena Research Center DIGITAL ORGANOLOGY, Leipzig University [email protected] Working Paper [v1] submitted on 21 November, 2025. DOI [v1]: 10.5281/zenodo.17643051 Abstract The acceleration of data acquisition technologies within the Digital Humanities has precipitated a proliferation of heterogeneous research assets, creating significant challenges for long-term preservation and repository ingestion. Existing workflows, frequently dependent on ephemeral scripting and manual curation, fail to guarantee the reproducibility and structural integrity required by trusted digital repositories. This paper details the Heritage Data Processor (HDP), a modular architectural framework designed to automate the preservation of multimodal cultural heritage data to the Zenodo infrastructure. Distinguishing itself from monolithic legacy tools, the system implements a persistent project state through a SQLite-based container format, thereby decoupling local data management from immediate network dependencies. The architecture utilizes a component-based pipeline strategy to enforce strict dependency isolation, facilitating the integration of diverse processing tasks ranging from high-throughput 3D rendering to AI-driven metadata extraction. By formalizing the transition from raw operational storage to archival records, the HDP establishes a standardized protocol that enhances compliance with FAIR principles and ensures the functional longevity of digital heritage collections. Keywords Software ·Cultural Heritage ·Data Processing ·Persistence ·Storage Heritage Data Processor (HDP) Working Paper [v1] Contents 1 Introduction 1 1.1 The Challenge of Heterogeneous Heritage Data . . . . . . . . . . . . . . . . . . . . . . . . . . 1 1.2 Evolution from Scripts to Systems: Zenodo Toolbox . . . . . . . . . . . . . . . . . . . . . . . 1 1.3 The Heritage Data Processor (HDP) Solution . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 2 System Architecture and Implementation 2 2.1 Technology Stack and Hybrid Application Structure . . . . . . . . . . . . . . . . . . . . . . . 2 2.1.1 DependencyManagement .................................. 3 2.1.2 SystemDependencies .................................... 4 2.2 Data Persistence: The .hdpc Database . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2.2.1 Schema Versioning and Migration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2.3 ConfigurationManagement ..................................... 5 2.3.1 Dynamic Configuration and Runtime Reloading . . . . . . . . . . . . . . . . . . . . . . 5 2.3.2 Security and Credential Isolation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.4 Deployment and Lifecycle Automation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.4.1 Automated Provisioning and Validation . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.4.2 Runtime Integrity Enforcement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.4.3 Transactional Maintenance and Rollback . . . . . . . . . . . . . . . . . . . . . . . . . . 6 3 Data Organization Paradigms 6 3.1 Project Initialization and Modalities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 3.2 Batch Entity Modes: Conceptualizing the Record . . . . . . . . . . . . . . . . . . . . . . . . . 6 3.2.1 RootMode .......................................... 6 3.2.2 SubdirectoryMode...................................... 7 3.2.3 HybridMode......................................... 7 3.3 BundlingStrategies.......................................... 7 3.3.1 StemMatching........................................ 7 3.3.2 PatternandRegexMatching ................................ 7 3.3.3 PrimarySourceDesignation................................. 7 4 Metadata Engineering and Automation 7 4.1 TheMetadataMappingWizard................................... 8 4.2 Automated Enrichment and LLM Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 4.2.1 Constructed Fields and Templates . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 i Heritage Data Processor (HDP) Working Paper [v1] 4.2.2 AutomappingAlgorithms .................................. 8 4.2.3 LocalLLMIntegration.................................... 9 4.3 Schema Validation and Integrity Enforcement . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 4.3.1 StructuralValidation..................................... 9 4.3.2 SemanticValidation ..................................... 9 4.3.3 Relational Integrity Checks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 4.3.4 LLMOutputVerification .................................. 9 5 The Modular Pipeline Ecosystem 10 5.1 Component Architecture: The Triad Structure . . . . . . . . . . . . . . . . . . . . . . . . . . 10 5.2 The Centralized Registry Middleware . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 5.2.1 Asynchronous Synchronization and Integrity Verification . . . . . . . . . . . . . . . . . 12 5.2.2 UnifiedDiscoveryEndpoint................................. 12 5.3 Component Implementation Case Study: The Fast Renderer . . . . . . . . . . . . . . . . . . . 12 5.3.1 Specification and Interface Congruency . . . . . . . . . . . . . . . . . . . . . . . . . . 12 5.3.2 Metaprogramming in the Logic Layer . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 5.4 Component Implementation Case Study: The LLM Processor . . . . . . . . . . . . . . . . . . 13 5.4.1 InterfacePolymorphism ................................... 13 5.4.2 Logic Layer: Validation and Self-Correction . . . . . . . . . . . . . . . . . . . . . . . . 13 5.4.3 AI Provenance and Paradata . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 5.5 ExecutionandOrchestration .................................... 14 5.5.1 AsynchronousExecution................................... 14 5.5.2 Real-Time Monitoring via SSE . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 5.5.3 EnvironmentIsolation.................................... 14 5.6 Pipeline Construction and Logic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 5.6.1 Step-wise Execution and Serialization . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 5.6.2 Metadata Overrides and Injection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 5.7 High-Performance Computing (HPC) Integration . . . . . . . . . . . . . . . . . . . . . . . . . 15 5.7.1 Remote Orchestration Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 5.7.2 Profile-Based Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 6 Repository Integration and Publication Lifecycle 16 6.1 The Zenodo Workflow State Machine . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 6.1.1 PreparationPhase ...................................... 16 6.1.2 Draft Management and Environments . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 ii Heritage Data Processor (HDP) Working Paper [v1] 6.1.3 VersioningLogic ....................................... 16 6.2 API Interaction and Fault Tolerance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 6.2.1 RateLimiting......................................... 16 6.2.2 Error Recovery: Discard and Restore . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 6.2.3 Mitigating Repository Retrieval Limits . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 6.3 BatchProcessingCapabilities.................................... 17 7 Interfaces and Extensibility 17 7.1 User Stability and Feature Gating . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 7.2 Graphical User Interface Modules . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 7.2.1 The Pipeline Constructor Interface . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 7.2.2 TheComponentRepository................................. 18 7.2.3 Interactive Component Execution and Testing . . . . . . . . . . . . . . . . . . . . . . 18 7.2.4 Automated Installation Wizard . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 7.2.5 The Dashboard and Upload Manager . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 7.3 The Command-Line Interface (CLI) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 7.3.1 HeadlessOperation...................................... 19 7.3.2 DaemonManagement .................................... 19 7.4 RESTAPIDesign .......................................... 20 7.4.1 APIStructureOverview................................... 20 7.4.2 TheCLIRoutesAPI..................................... 20 7.5 WebHDP: The Server-Hosted Browser Interface . . . . . . . . . . . . . . . . . . . . . . . . . . 20 7.5.1 Microservice Architecture and Stack . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 7.5.2 Multi-User Collaboration and Role-Based Access Control (RBAC) . . . . . . . . . . . 21 7.5.3 SecurityandDeployment .................................. 21 8 Evaluation and Case Studies 21 8.1 3D Heritage Workflow Case Study . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 8.2 Performance of Distributed Rendering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 8.2.1 Hybrid Parallelization Strategy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 8.2.2 ThroughputandOverhead ................................. 22 8.3 Deployment and Validation in Research Infrastructures . . . . . . . . . . . . . . . . . . . . . . 22 9 Discussion: Ethics, Standards, and Governance 23 9.1 FAIRComplianceAssessment.................................... 23 9.1.1 Findability (F): PIDs and Rich Metadata . . . . . . . . . . . . . . . . . . . . . . . . . 23 iii Heritage Data Processor (HDP) Working Paper [v1] 9.1.2 Accessibility (A): Protocol Standardization . . . . . . . . . . . . . . . . . . . . . . . . 23 9.1.3 Interoperability (I): Modular Specifications . . . . . . . . . . . . . . . . . . . . . . . . 23 9.1.4 Reusability (R): Provenance and State Persistence . . . . . . . . . . . . . . . . . . . . 24 9.2 ReproducibilityandStandards ................................... 24 9.2.1 The Portable Project State (.hdpc) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 9.2.2 Component Specification (component.yaml) . . . . . . . . . . . . . . . . . . . . . . . . 24 9.3 Repository Suitability and Policy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 9.3.1 Infrastructural Permanence and Data Integrity . . . . . . . . . . . . . . . . . . . . . . 25 9.3.2 Legal Framework and GDPR Alignment . . . . . . . . . . . . . . . . . . . . . . . . . . 25 9.3.3 The Limitations of Closed Access . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 9.3.4 LiabilityandLicensing.................................... 25 9.3.5 The Limits of Passive Persistence: Bit-Level vs. Functional Preservation . . . . . . . . 26 9.3.6 Comparative Analysis: Generalist vs. Domain-Specific Infrastructures . . . . . . . . . 26 9.3.7 API Usage and Operational Stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 9.4 Ethical Governance: Integrating CARE Principles and Indigenous Data Sovereignty . . . . . 27 9.4.1 The Tension Between Open Data and Data Sovereignty . . . . . . . . . . . . . . . . . 27 9.4.2 Mitigating Algorithmic Bias and ”Data of Disregard” . . . . . . . . . . . . . . . . . . 28 9.4.3 Operationalizing Ethics in the Pipeline . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 9.5 Ethical Risks in Digital Dissemination: Trafficking, Sensitivity, and Security . . . . . . . . . . 28 9.5.1 Countering the Illicit Trade and Looting . . . . . . . . . . . . . . . . . . . . . . . . . . 28 9.5.2 Bioarchaeological Ethics and Contextualization . . . . . . . . . . . . . . . . . . . . . . 29 9.5.3 Intellectual Property and Geometric Protection . . . . . . . . . . . . . . . . . . . . . . 29 10 System Availability and Development Status 29 10.1CurrentReleaseState ........................................ 29 10.2 Usage and Stability Disclaimer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29 11 Conclusion 29 11.1FutureDevelopments......................................... 30 11.1.1 State Re-hydration and Large-Collection Resilience . . . . . . . . . . . . . . . . . . . . 30 11.1.2 Active Preservation and Default Format Normalization . . . . . . . . . . . . . . . . . . 30 11.1.3 Automated Sensitivity Detection and Ethical Guardrails . . . . . . . . . . . . . . . . . 31 11.1.4 Embedded Intelligence via Small Language Models . . . . . . . . . . . . . . . . . . . . 31 References 33 iv Heritage Data Processor (HDP) Working Paper [v1] 1 Introduction The rapid digitization of cultural heritage has precipitated a fundamental shift in the scale and complexity of research data within the Digital Humanities (DH). As storage capacities expand and acquisition technologies become ubiquitous, spanning high-resolution photogrammetry, multispectral imaging, and spatialized multidimensional audio, researchers are confronted with an increasingly heterogeneous landscape of digital assets. While the production of such data has accelerated considerably, the workflows required to curate, standardize, and preserve it according to the FAIR principles (Findable, Accessible, Interoperable, Reusable) [1] have often lagged behind. Current practices frequently rely on ad-hoc scripts and manual intervention, which compromise both efficiency and reproducibility. This paper introduces the Heritage Data Processor (HDP), a modular, pipeline-driven architecture designed to bridge the operational gap between local data management and long-term preservation repositories. 1.1 The Challenge of Heterogeneous Heritage Data Modern heritage datasets are characterized by substantial complexity and a lack of uniformity. They encompass a diverse array of modalities, spanning high-resolution photography, three-dimensional models, audio recordings, audiovisual documentation, and structured textual data. This heterogeneity creates significant impediments to the publication and archival process. For instance, a single three-dimensional digital artifact is rarely a solitary file; rather, it is a complex aggregate involving geometry data (such as Wavefront .obj), material libraries (.mtl), and texture maps. These components must be maintained as a coherent conceptual unit to ensure future computational reproducibility and visual fidelity. A critical systemic challenge arises from the disconnect between the raw operational storage of these assets and the rigorous metadata schemata mandated by trusted digital repositories. Researchers frequently encounter difficulties in maintaining the semantic lineage between a digital object and its descriptive metadata, such as provenance and authorship, during the ingest phase. Furthermore, the absence of standardized configurations for validating specific data types frequently results in inconsistent data quality and ingestion errors. In the absence of automated validation and structural integrity checks, the long-term utility and reusability of archived heritage data remain compromised. 1.2 Evolution from Scripts to Systems: Zenodo Toolbox The architectural design of the HDP represents an evolutionary successor to the Zenodo Toolbox [2]. Originally implemented as a Python module, this toolbox was engineered to facilitate large-scale batch processing and bulk ingestion into the Zenodo repository [3]. Although effective for targeted automation tasks, the utility displayed critical architectural constraints inherent to its script-based origins. Primarily, the initial iteration lacked a mechanism for persistent project state management, relying solely on a monolithic SQLite database for record-keeping. In the absence of a dedicated serialization format for projects, metadata and processing configurations remained ephemeral. These configurations persisted only as logs of upload operations, lacking the contextual continuity required for iterative research workflows. Furthermore, the unified software design precipitated significant challenges regarding maintainability. As the toolkit expanded to encompass diverse data processing tasks, the accumulation of conflicting libraries within a singular Python environment resulted in unresolvable dependency conflicts. In this scenario, updates required for specific functionalities frequently compromised the stability of adjacent components. An intermediate, unpublished iteration of the Zenodo Toolbox introduced an interactive Command-Line Interface (CLI) and implemented persistent project files to address state management. However, the fundamental issue of monolithic dependency management remained unresolved. Additionally, a critical usability barrier persisted: despite the provision of multiple Jupyter Notebooks containing simplified examples and workflow templates, the script-based and command-line interfaces remained inaccessible to researchers without programming expertise. This accessibility gap represented a significant impediment to broader 1 Heritage Data Processor (HDP) Working Paper [v1] adoption within the Digital Humanities community, where domain specialists often lack formal training in computational methods. 1.3 The Heritage Data Processor (HDP) Solution The Heritage Data Processor (HDP) addresses legacy limitations through a comprehensive architectural philosophy centered on persistent state management, strict component isolation, and versatile deployment capabilities. The system offers a unified operational framework accessible via a Graphical User Interface (GUI), a CLI, and an Application Programming Interface (API), ensuring functional congruence across all interaction modalities. 1. Persistent Project Management: Central to the HDP workflow is the .hdpc (Heritage Data Processor Container) project file. Unlike its predecessors, HDP utilizes this SQLite-based structure to store not only file inventories and metadata mappings but also the complete processing history and configuration state of a dataset. This persistent architecture facilitates asynchronous workflows, allowing data preparation, validation, and publication to span extended periods without loss of contextual integrity. 2. Encapsulated Modular Architecture: To eliminate dependency conflicts, HDP adopts a rigorously modular design. Processing logic is encapsulated within independent units designated as HDP Components. Each component is defined by a YAML [4] interface specification and executed within isolated environments, utilizing either ‘uv‘-managed virtual environments [5] or Docker containers [6]. This isolation allows disparate components to operate with conflicting dependencies or distinct Python versions while remaining orchestratable by a central pipeline manager. 3. Scalable Deployment and Integration: The HDP architecture is agnostic to its hosting environment, capable of operating as a standalone desktop application, a server-side backend for web services, or within High-Performance Computing (HPC) infrastructures. Furthermore, it provides a robust bridge to Zenodo, supporting both Sandbox and Production environments. The system automates the complex lifecycle of repository records, from initial draft creation to versioned publication, ensuring precise DOI assignment and continuous metadata synchronization. By formalizing these processes into a versatile environment that supports both interactive and script-based operation, HDP standardizes the preservation of complex heritage data, ensuring it remains accessible and interoperable for future research applications. 2 System Architecture and Implementation The HDP is engineered based on core principles of modularity, portability, and robust state management. It is implemented as a hybrid desktop application that decouples the user interface from the core processing logic, a significant departure from the transient, script-based nature of its predecessors. This architecture is designed to satisfy the dual requirements of a responsive, cross-platform user interface and high-performance data processing. 2.1 Technology Stack and Hybrid Application Structure To achieve a strict separation of concerns, the HDP employs a bifurcated structure, separating a frontend built on web-standard technologies from a dedicated Python backend for system-level computation [7]. The frontend is developed using the Electron framework [8], which facilitates cross-platform deployment by combining a Chromium rendering engine [9] with a Node.js runtime [10]. To ensure lightweight performance and maintainability, the interface avoids complex single-page application frameworks, instead relying on 2 Heritage Data Processor (HDP) Working Paper [v1] standard ECMAScript (ES6+) with ES Modules (ESM) for its logical structure [11]. The visual design is implemented with Tailwind CSS [12], a utility-first styling framework. This frontend communicates with the Python backend via a local RESTful API [13]. The backend, built upon the Flask web framework [14], is responsible for the computational workload, including file system scanning, metadata extraction, and pipeline execution. To enhance user interface responsiveness, the frontend directly integrates sql.js, a WebAssembly-based [15] implementation of SQLite [16]. This allows the application to perform direct, low-latency read operations on the local .hdpc project database, minimizing inter-process communication for data-querying tasks and offloading read operations from the Python backend. HDP Technology Stack & Architecture Frontend (Electron + Node.js) Electron Framework Chromium Rendering Engine Node.js Orchestration UI Logic & Styling Standard ECMAScript (ES6+) & ES Modules Tailwind CSS (Utility-First Design) Local Data Access (sql.js - WebAssembly) .hdpc Project Database Dependency Management npm Backend (Python + Flask) Flask Server API Handling Pipeline Execution Data Processing & Computation File Scanning Metadata Extraction HDP Components (Modular Pipelines) System-Level Dependencies libmagic (File Type Identification) Dependency Management uv (High-Performance Python Package Installer) Isolated Virtual Environments (.venv, pyproject.toml) RESTful API Calls Figure 1: Architectural overview of the Heritage Data Processor (HDP). The diagram illustrates the hybrid application structure, decoupling the Electron-based frontend (handling UI and direct data reads via sql.js ) from the Python backend (responsible for processing logic and database write operations). Inter-process communication is mediated by a local RESTful API, with the .hdpc SQLite database serving as the shared persistent storage layer. 2.1.1 Dependency Management Given the complexity of heritage data workflows, rigorous dependency management is essential to prevent environment conflicts and ensure reproducibility. The HDP employs distinct strategies for frontend and backend dependencies. Frontend libraries are managed via npm [17], following standard Node.js conventions. For backend environment management, the system adopts uv, a high-performance Python package installer and environment manager. The use of uv facilitates the rapid creation of isolated virtual environments (.venv) [18] based on the pyproject.toml specification [19], ensuring that the HDP’s dependencies remain isolated from system-level Python installations. This isolation is particularly critical for the modular HDP Components, which may require conflicting library versions. For instance, different processing components might depend on incompatible versions of widely used libraries such as pandas [20] or numpy [21], conflicts that are resolved through encapsulated environments. 3 Heritage Data Processor (HDP) Working Paper [v1] 2.1.2 System Dependencies Beyond language-specific package management, the HDP integrates system-level tools to ensure robust and platform-agnostic data handling. A critical dependency is libmagic [22], the C library underlying the file command on Unix-like systems. This library provides definitive file type identification based on file signatures (byte-level patterns within the file header), rather than relying on potentially misleading file extensions. The HDP accesses this functionality via python-magic [23], a Python wrapper for the underlying system library. This capability is essential for digital preservation workflows, as it ensures that MIME types recorded in metadata accurately reflect the true format of digital objects across different operating systems, including macOS, Linux, and Windows. 2.2 Data Persistence: The .hdpc Database State management within the HDP represents a fundamental departure from the ephemeral, file-based approaches of its predecessors. The system introduces the .hdpc file format, which functions as a portable, self-contained project database. Technically, an .hdpc file is implemented as a SQLite relational database. This design choice provides transactional integrity and supports complex relational queries that flat file formats such as JSON or YAML cannot accommodate. The database schema organizes project data into several core relational tables: •project_info : Stores high-level project metadata, including the project name, short code, and schema version. •source_files : Maintains a comprehensive inventory of all local assets, tracking absolute file paths, computed SHA-256 cryptographic hashes [24], MIME types [25], and processing statuses (e.g., pending, valid, processed). •zenodo_records : Manages the publication lifecycle, linking local file sets to Zenodo draft identifiers, Digital Object Identifiers (DOIs) [26], and metadata snapshots. •record_files_map : A relational join table that handles the many-to-one association between source files and Zenodo records, tracking the upload status of individual assets within a deposit. Beyond portability, the SQLite architecture enables advanced project consolidation capabilities. In massdigitization workflows, operators frequently create distinct projects for individual physical objects to ensure logical isolation during data acquisition. The HDP supports the merging of these disparate .hdpc files into a unified master database. This functionality permits a decentralized creation workflow, where individual datasets can be combined subsequently for bulk processing and unified publication. Critically, this consolidation preserves relational integrity and avoids data duplication, ensuring that provenance chains and processing histories remain intact across merged projects. 2.2.1 Schema Versioning and Migration To ensure longitudinal software compatibility and the preservation of project data integrity, the HDP incorporates a rigorous database migration strategy utilizing Alembic [27]. As the software evolves, for instance by introducing new columns to the zenodo_records table to support additional metadata standards or repository features, the internal structure of the .hdpc container must be updated without compromising existing user data. Upon initializing a project file, the application compares the revision identifier stored within the database against the schema definition of the current software version. If a version discrepancy is detected, the system automatically initiates an idempotent migration routine. Alembic executes the necessary upgrade scripts to modify the table structure transactionally. This process ensures that a project created with an older version of the HDP remains fully operational in the latest release. This approach effectively prevents the ”format 4 Heritage Data Processor (HDP) Working Paper [v1] ROOT MODE model.obj render.png readme.pdf model.obj Record 1 (model.obj) render.png Record 2 (render.png) readme.pdf Record 3 (readme.pdf) Input Directory Each File = 1 Record (Homogenous, Independent Assets) SUBDIR MODE Input Directory Each Subdirectory = 1 Record (Compound, Aggregated Assets) 3DModel_A 3DModel_B mesh.obj mesh.mtl tex.png Record 1 3DModel_A Aggregated mesh.mtl scan.ply Record 2 3DModel_B Aggregated HYBRID MODE Input Directory Mix: Root Files & Subdirectories as Records (Messy, Mixed Datasets) report.pdf data.csv Record 1 (report.pdf) Record 2 (data.csv) Record 3 (Project_X) Record 4 (Project_Y) Project X Project Y mod.obj tex.png code.py Figure 3: Schematic representation of HDP Batch Entity Modes. The diagram illustrates the three structural paradigms used to map local file systems to archival records. Root Mode maps individual files to distinct records (1:1). Subdirectory Mode aggregates entire folders into single compound records (n:1), ensuring that multi-file assets like 3D models remain semantically linked. Hybrid Mode enables the simultaneous processing of both structures to handle mixed datasets. 11 Heritage Data Processor (HDP) Working Paper [v1] command-line arguments. This layer is responsible for input validation, argument parsing, and instantiating the core processor class. A strict congruency is enforced here: every parameter defined in component.yaml must have a corresponding handler in main.py. • The Logic Layer ( processor.py ): This layer contains the computational logic, implemented as a class-based structure. By decoupling the algorithmic core from the CLI interface, the processor remains technology-agnostic and testable, facilitating its reuse in other Python contexts external to the HDP environment. 5.2 The Centralized Registry Middleware To decouple local processing clients from the latency and strict rate-limiting constraints of the Zenodo API, the ecosystem relies on a centralized middleware service, the HDP Component Registry. This service is architected as a lightweight, containerized microservice utilizing the Flask framework. 5.2.1 Asynchronous Synchronization and Integrity Verification Unlike a passive database, the Registry actively maintains synchronization with the hdp-components Zenodo Community 1 through an autonomous background scheduler. Implemented via the APScheduler library, the SchedulerService executes periodic update cycles (configurable via UPDATE_INTERVAL_HOURS , typically 24 hours). Crucially, the synchronization process implements a strict Supply Chain Security protocol to mitigate the risk of arbitrary code execution. During the update cycle, the ZenodoService does not merely download assets; it enforces a cryptographic handshake. The system calculates the SHA-256 checksum of the downloaded component_complete.zip and validates it against a trusted, immutable ledger of signed releases maintained by the community moderator. This ensures that even if a community contributor’s Zenodo account is compromised, malicious code cannot be injected into the ecosystem, as the modified archive would fail the signature verification step and be rejected by the Registry. 5.2.2 Unified Discovery Endpoint The aggregated and validated metadata is served through a high-performance REST endpoint, /hdp/v1/available-components . This endpoint provides HDP clients with a unified, categorized manifest of all valid HDP Components, including their version history, direct download links, and calculated dependency trees. By consolidating discovery and security logic into this middleware, the system ensures that end-users interact with a cached, verified state of the ecosystem. This significantly improves the responsiveness of the pipeline construction interface while maintaining a secure chain of custody. 5.3 Component Implementation Case Study: The Fast Renderer To illustrate the practical application of the three-tier architecture, we examine the Fast Renderer component [38]. This high-throughput 3D visualization tool was designed to evaluate the limits of the HDP’s modular capabilities. The component encapsulates the Blender 3D creation suite [39], exposing its rendering engine through a standardized HDP interface. 5.3.1 Specification and Interface Congruency The component.yaml file acts as the binding contract, defining the component’s interface independently of its underlying technology. It declares a required input input_dir for raw 3D models and an output_dir for the generated imagery. Crucially, it abstracts complex rendering parameters into user-friendly controls. For 1https://zenodo.org/communities/hdp-components/ 12 Heritage Data Processor (HDP) Working Paper [v1] instance, the quality parameter offers presets (”draft”, ”medium”, ”ultra”) which map to specific internal sample rates and resolutions. The interface layer ( main.py ) ensures strict congruency with this specification. Using the argparse library [40], it maps every parameter defined in the YAML, such as --light-energy or --camera-dist , to executable arguments. This layer also handles infrastructure-level logic that lies outside the core processing scope. For the Fast Renderer, this includes a portable GPU detection routine ( detect_available_gpus ) that inspects the host environment for NVIDIA or Apple Metal accelerators to optimize resource allocation. 5.3.2 Metaprogramming in the Logic Layer The core logic layer ( processor.py ) demonstrates the power of the HDP’s encapsulation strategy. Rather than interacting with Blender’s API directly (which would require the HDP to run within Blender’s embedded Python environment), the OptimizedRendererProcessor class employs a metaprogramming approach. The processor utilizes a string template, OPTIMIZED_BATCH_TEMPLATE , to dynamically generate a dedicated Python script at runtime. This script injects the user-defined parameters, such as resolution, lighting angles, and engine selection, directly into the Blender context. The processor then orchestrates the execution via the subprocess module, streaming the standard output in real-time to capture distinct BATCH_LOG signals. This separation ensures that the HDP’s dependency environment remains isolated from the strict requirement sets of external tools like Blender, preventing library conflicts while maintaining full control over the execution lifecycle. 5.4 Component Implementation Case Study: The LLM Processor While the Fast Renderer demonstrates deterministic media processing, the LLM Processor component illustrates how the HDP architecture manages the probabilistic nature of Generative AI. This component is designed to extract structured metadata from unstructured text or tabular data using local LLMs via the Ollama framework. 5.4.1 Interface Polymorphism The interface layer ( main.py ) of this component handles input polymorphism. The component.yaml specification defines two mutually exclusive data inputs: a simple string ( input_text ) and a file path ( table_file_path ) for batch processing CSV or Excel datasets. The CLI wrapper logic parses these arguments to determine the execution mode. In Table Mode, it utilizes the table_column_mapping parameter to dynamically inject row values into prompt placeholders, such as mapping a ”Bio” column to a biography variable. This effectively treats the LLM as a row-by-row transformation engine. 5.4.2 Logic Layer: Validation and Self-Correction The core challenges of integrating LLMs into scientific pipelines are hallucination and structural inconsistency. The LLMProcessor class addresses this through a rigorous ”Generate-Repair-Validate” cycle implemented in the logic layer. Instead of accepting raw model output, the processor enforces a json_schema constraint. The execution flow, managed by the _run_llm method, proceeds as follows: 1. Generation: The system formats a prompt using the PromptStore , which decouples prompt engineering from code, and queries the local Ollama instance. 2. Repair: A heuristic attempt_json_repair utility strips Markdown code fences (e.g., ```json ) and fixes common syntax errors, such as trailing commas, which frequently occur with smaller models. 3. Validation and Fallback: The result is validated against the defined schema. Uniquely, if validation fails, the component can trigger a fallback_prompt_key (e.g., a ”final” attempt with stricter instructions) to auto-correct the output before flagging it as a failure. 13 Heritage Data Processor (HDP) Working Paper [v1] 5.4.3 AI Provenance and Paradata To satisfy transparency requirements and processing provenance, the component implements comprehensive paradata logging. The output JSON does not merely contain the extracted results; it includes a paradata object recording the exact execution_start time, the specific model version used (e.g., ”qwen3:8b”), and crucially, the full text of the system and user prompt templates employed during the run. This ensures that AI-derived metadata remains auditable and reproducible. 5.5 Execution and Orchestration The orchestration of these components is managed by the component_runner_api , which handles the lifecycle of execution, resource allocation, and status monitoring. 5.5.1 Asynchronous Execution When a component is triggered, the system initiates an asynchronous execution process, returning a unique UUID (Universally Unique Identifier, Version 4) [41] to track the instance. This identifier allows the frontend to query the execution state (transitioning through Running , Completed , or Failed statuses) without blocking the main application thread. The runner utility constructs the specific shell command by inspecting the component’s CLI patterns (e.g., detecting support for --output-dir versus --output-file flags) and resolving absolute file paths. 5.5.2 Real-Time Monitoring via SSE To provide immediate feedback to the user, HDP implements Server-Sent Events (SSE). The /components/logs/<execution_id> endpoint opens a unidirectional stream that pushes log messages (categorized as info , warning , or error ) from the Python subprocess directly to the user interface in real-time. This mechanism includes heartbeat signals to maintain connection stability during long-running processes. 5.5.3 Environment Isolation A critical feature of the HDP is its resolution of dependency conflicts. The system leverages uv to manage distinct virtual environments for each component. The runner dynamically resolves the Python executable path within the specific component’s environment (e.g., looking into env/bin on Unix or env/Scripts on Windows). This ensures that a component requiring legacy libraries can coexist in the same pipeline as one utilizing cutting-edge dependencies. 5.6 Pipeline Construction and Logic Pipelines in HDP are defined not merely as ordered lists of tasks, but as complex, serialized structures that handle data flow and metadata transformation. 5.6.1 Step-wise Execution and Serialization Pipelines are defined as YAML files and database representations where each element represents a processing step containing a component_name , parameters , and an inputMapping configuration. The input mapping logic resolves dependencies dynamically; for instance, an input for Step 2 can be explicitly mapped to the output of Step 1 using a sourceType: pipelineFile reference. This allows for the construction of non-linear workflows where intermediate derivatives are passed between isolated components. A typical use case of this principle is executing geolocalization components on photographies, and mapping the resulting coordinates to linked data carriers such as IIIF (International Image Interoperability Framework) manifest files [42] or Europeana Data Model (EDM) XMLs [43]. 14 Heritage Data Processor (HDP) Working Paper [v1] 5.6.2 Metadata Overrides and Injection A powerful capability of the pipeline engine is the programmatic extraction and injection of metadata. Pipeline steps can be configured with an outputMapping , which instructs the system to parse the JSON output of a component and map specific keys to Zenodo metadata fields. For example, a ”Metadata Extractor” component might analyze an FITS [44] image header and output a JSON file containing technical metadata. The pipeline manager can be configured to extract this specific data and inject it into the description or keywords fields of the target Zenodo record. This mechanism effectively overrides the initial, static metadata mapping derived from the spreadsheet, allowing for data-driven, automated enrichment of the final archival record. 5.7 High-Performance Computing (HPC) Integration To accommodate the computational demands of modern heritage data processing, such as large-scale photogrammetry or AI-driven geolocation, the HDP architecture incorporates a specialized hpc-connector module. Although currently implemented and validated within test environments, this middleware is scheduled for public release in subsequent updates. It mediates the interaction between the local HDP instance and remote supercomputing clusters, functionally integrating an HPC node as an ephemeral extension of the local pipeline. 5.7.1 Remote Orchestration Architecture The integration utilizes a lightweight client-side orchestrator ( hpc_orchestrator.py ) that manages the lifecycle of a remote job without requiring a persistent server agent on the cluster. The workflow adheres to a strict execution sequence: 1. Dynamic Allocation: The system connects to a defined SSH gateway (e.g., draco ) and requests resources via the Slurm Workload Manager using salloc . It employs a state tracker to monitor the transition from PENDING to READY , parsing real-time standard output (stdout) streams to identify the allocated node (e.g., gpu01). 2. Data Synchronization: Upon node allocation, the connector utilizes rsync over SSH to transfer input assets from the local project to the remote workspace. This ensures that only necessary data is transferred, thereby optimizing bandwidth usage. 3. Wrapper-Based Execution: To maintain compatibility with the HDP component standard, remote tools are wrapped in a specialized CLI script (e.g., hdp_cli.py ). This wrapper standardizes argument parsing and ensures that the remote application accepts inputs and produces outputs in a format congruent with local pipeline definitions. 4. Result Retrieval and Cleanup: Upon successful execution (detected via configurable success indicators in the log stream), results are synchronized back to the local output directory, and temporary remote files are purged to maintain cluster hygiene. 5.7.2 Profile-Based Configuration The complexity of heterogeneous cluster environments is abstracted through a centralized config.json file. This file defines ”Resource Profiles” (e.g., gpu_large , testgpu_quick ) that map to specific Slurm partition and memory configurations, as well as CUDA Profiles that handle the loading of environment modules and library paths. This decoupling enables the HDP to adapt to disparate HPC infrastructures solely through configuration updates, avoiding alterations to the core codebase. 15 Heritage Data Processor (HDP) Working Paper [v1] 6 Repository Integration and Publication Lifecycle The primary objective of the HDP is the reliable, automated publication of data to persistent repositories. To achieve this, the system implements a deterministic state machine that governs the transition of digital objects from local storage to the Zenodo infrastructure, ensuring transactional integrity across the publication lifecycle. 6.1 The Zenodo Workflow State Machine The HDP workflow is designed as a strict state machine, transitioning records through distinct phases: Preparation,Draft Management, and Publication. 6.1.1 Preparation Phase The lifecycle commences with the transition from raw file processing to internal database preparation. Upon the execution of the metadata mapping, the system validates the extracted metadata against the Zenodo JSON schema. Validated records are stored in the local zenodo_records table with a status of prepared . This internal staging area allows researchers to review and refine metadata locally (applying manual overrides or pipeline-derived enrichments) without interacting with external APIs or consuming network resources. 6.1.2 Draft Management and Environments Once prepared, records transition to the Draft state through interaction with the Zenodo API. The HDP architecture strictly enforces environment separation, utilizing distinct API credentials for the Sandbox (testing) and Production (live) environments. When a draft is instantiated via the create_api_draft endpoint, the system retrieves the remote Deposition ID and pre-reserves the DOI, binding them to the local record. This allows researchers to embed the persistent identifier into their data during runtime, such as linked data XML files, prior to the final upload. In practice, this capability is utilized to self-reference the record’s version within descriptions (e.g., changelogs) and, specifically, to generate and upload XML files compliant with the Europeana Data Model (EDM) [see 43]. 6.1.3 Versioning Logic A critical component of the lifecycle is the automated management of versioning. Zenodo utilizes a Concept DOI to represent all versions of a dataset, while specific Version DOIs identify individual snapshots. HDP automates this complexity through the create_new_version_draft function. When updating a published record, the system queries the Concept ID to identify the latest version and automatically calculates the next semantic version number (defaulting to a patch increment, e.g., v1.0.0 to v1.0.1 , with user overrides for minor or major updates) before creating the new draft. This ensures citation continuity without manual intervention. 6.2 API Interaction and Fault Tolerance Given the constraints of external web services, HDP incorporates robust fault tolerance mechanisms to handle network instability and API limitations. 6.2.1 Rate Limiting To maintain compliance with Zenodo’s usage policies, all outgoing requests are mediated by an internal rate_limiter_zenodo service. This middleware enforces a mandated wait time between requests, preventing the application from triggering ”429 Too Many Requests” errors during bulk operations. 16 Heritage Data Processor (HDP) Working Paper [v1] 6.2.2 Error Recovery: Discard and Restore Draft operations can fail due to various external factors, leaving the local database out of synchronization with the remote repository. HDP implements a ”Discard and Restore” pattern to mitigate this risk via the discard_zenodo_draft endpoint. Before any destructive API call, the system creates a snapshot in the metadata_backups table. If a draft requires discarding, whether due to user request or API failure, the system not only sends the delete command to Zenodo but also invokes the _restore_record_metadata routine. This function rolls back the local record state to its pre-draft configuration, restoring denormalized fields and clearing invalid IDs, effectively reversing the failed operation. 6.2.3 Mitigating Repository Retrieval Limits A significant architectural limitation of the Zenodo API is the retrieval cap of 10,000 records. While heuristic strategies like sorting inversion can extend this to 20,000, it leaves an informational gap for collections exceeding this size, preventing efficient synchronization via API calls alone. HDP circumvents this by treating the local .hdpc database (rather than the remote repository) as the source of truth. By persisting the link between local assets and Zenodo Deposit IDs internally, the system eliminates the need to fetch the full deposit list from the API. This allows HDP to manage collections of arbitrary size (e.g., exceeding 50,000 items) while circumventing retrieval ceilings and pagination bottlenecks. 6.3 Batch Processing Capabilities To support institutional-scale digitization, HDP exposes a batch_actions_api capable of executing lifecycle transitions on large-scale datasets simultaneously. This API supports bulk operations for every stage of the workflow: prepare_metadata , create_api_draft , upload_main_files , and publish . To mitigate the inherent instability of network operations during longrunning tasks, the batch processor implements a robust auto-retry mechanism with exponential backoff. When a transient error occurs (such as a 503 Service Unavailable or a network timeout), the system does not immediately mark the record as failed. Instead, it pauses execution for a calculated interval (base delay of 2 s, increasing exponentially up to 30 s) and retries the operation up to three times. This ensures that temporary API instabilities do not disrupt overnight ingestion workflows. Furthermore, the processor employs a ”Multi-Status” error handling strategy, analogous to HTTP 207. If retries are exhausted and a definitive failure occurs, the system isolates the exception to the specific record. The process continues for all remaining items, and the final response provides a granular report detailing which IDs succeeded and which failed, along with specific error diagnostics. This capability allows operators to address only the problematic entries without the necessity of restarting the entire batch operation. 7 Interfaces and Extensibility While the GUI lowers the barrier to entry for users, the HDP architecture acknowledges that large-scale data engineering requires automation, reproducibility, and headless operation. To satisfy these requirements, the system exposes its functionality through two primary programmatic interfaces: a comprehensive CLI and a structured REST API. 7.1 User Stability and Feature Gating To maintain a robust user experience for non-experts, HDP implements a strict feature-gating middleware, the Feature Stability Gate. Recognizing that experimental features can compromise data integrity or confuse users, this system selectively disables UI elements (such as specific buttons or navigation views like Storage 17 Heritage Data Processor (HDP) Working Paper [v1] Optimization or Pipeline View) that are currently flagged as work-in-progress. Crucially, this blocking logic applies strictly to the GUI to prevent accidental misuse. The functionalities remain fully accessible via the CLI and API, allowing developers and power users to utilize and test alpha features programmatically without exposing general users to instability. 7.2 Graphical User Interface Modules While the backend manages data integrity, the Electron-based frontend translates these abstract architectures into a visual workspace designed for non-technical users. This interface is organized into distinct modules that expose the system’s logic through interactive visual paradigms. 7.2.1 The Pipeline Constructor Interface The Pipeline Constructor serves as a visual programming environment that abstracts the JSON-based pipeline definitions described in subsection 5.6. Unlike linear script editors, this view renders the processing workflow as a sequential series of Step Blocks, each representing a discrete processing unit. • Visual Data Flow: The interface implements a File Card system where input and output assets are represented as manipulatable UI objects. Users configure data flow not by writing paths, but by selecting file cards generated by previous steps using a dynamic input selector. • Metadata Injection Logic: A dedicated Initial Step block allows for the global configuration of the Zenodo draft. Subsequent steps feature an Output Mapping modal where users can visually map specific component outputs (e.g., a JSON log file) to Zenodo metadata fields, with the UI creating visual flags (e.g., ”Overwrites Metadata”) to warn users of data precedence conflicts. • Description Templating: The module includes a Description Constructor step that utilizes a variable-based templating system (e.g., ${output_file.coordinates.x} ), allowing the final description to be dynamically generated from the analytical results of the pipeline. 7.2.2 The Component Repository To manage the modular ecosystem, the GUI includes the Component Repository. Surpassing a static enumeration, this interface acts as a real-time dashboard for the local component library. • Categorized Management: Components are organized into collapsible categories (e.g., ”3D Processing”, ”Metadata Extraction”) with visual status indicators distinguishing between ”Installed”, ”Available”, and ”Invalid” states. • Update Management: The view integrates with the remote registry to perform version checks. An Available Updates modal allows users to compare local versus remote versions and view changelogs rendered directly from the remote repository’s Markdown files before authorizing an update. • Online Discovery: ADownload Components interface queries the central registry for new modules, allowing users to expand their processing capabilities without leaving the application context. 7.2.3 Interactive Component Execution and Testing Recognizing that pipelines are complex to debug, the HDP includes a Component Execution interface that functions as an interactive unit-testing ground. Accessible via the Run button on any installed component, this modal allows users to execute a processor in isolation. • Dynamic Parameterization: The interface dynamically generates form controls based on the component’s YAML specification, rendering file pickers for path types, toggle switches for boolean flags, and dropdowns for enums. 18 Heritage Data Processor (HDP) Working Paper [v1] • Template System: To facilitate repetitive testing, the manager supports a Parameter Template system, allowing users to save and load specific configuration sets (e.g., ”High-Res Photogrammetry Settings”) for rapid reuse. • Real-Time Feedback: Execution logs are streamed in real-time via SSE to a dedicated console window, allowing operators to diagnose failures or verify outputs immediately without inspecting backend log files. 7.2.4 Automated Installation Wizard The installation of complex research software is streamlined through the Component Installation Manager. This wizard abstracts the underlying uv environment provisioning into a guided graphical process. • Pre-Flight Validation: Before installation begins, the manager parses the component’s requirements to perform system checks (e.g., verifying GPU availability or Python versions). • Dependency Resolution: It visualizes the requirement for auxiliary files (such as machine learning model weights or binaries), providing direct download links and validation logic to ensure all dependencies are present before the environment is built. 7.2.5 The Dashboard and Upload Manager The publication lifecycle state machine is visualized through the Upload Manager. This view replaces the abstract concept of database states with a tabbed interface (”Pending Preparation”, ”Drafts”, ”Published”), allowing users to promote records through the preservation lifecycle via action buttons. It includes a real-time Dashboard that aggregates project statistics (such as total file count, record distribution across environments, and metadata completion rates), providing immediate visual feedback on the project’s archival readiness. 7.3 The Command-Line Interface (CLI) The HDP CLI serves as an orchestration tool designed for server-side environments and automated workflows, effectively decoupling the processing logic from the Electron-based frontend. 7.3.1 Headless Operation The CLI exposes the entire preservation lifecycle through a suite of hdp commands, enabling operations on headless servers where a GUI is unavailable by default. •hdp upload : This command automates the ingestion process by scanning a specified input directory, registering files in the .hdpc database, executing metadata preparation logic, and instantiating Zenodo drafts in the target environment (Sandbox or Production). •hdp process : This command triggers pipeline execution on existing records. It supports granular filtering options (such as --since , --until , or --title-pattern ), allowing operators to target specific subsets of data for transformation (e.g., processing only records created in the last 24 hours). •hdp publish : Representing the final step in the lifecycle, this command validates that all file uploads are complete and commits the drafts to the repository, making them publicly accessible via DOI. 7.3.2 Daemon Management To facilitate integration into system services or Continuous Integration/Continuous Deployment (CI/CD) pipelines, the CLI implements an automatic daemon management strategy. When an hdp command is issued, the tool checks for an active instance of the backend server. If none is detected, it automatically spawns the hdp-server process to handle the request and manages its lifecycle, thereby eliminating the need for 19 Heritage Data Processor (HDP) Working Paper [v1] manual server startup. Conversely, advanced users can utilize the --no-auto-server flag to interface with an existing, long-running server instance. 7.4 REST API Design Underlying both the Electron GUI and the CLI is a RESTful API architecture built on the Flask framework. This API is segmented into logical namespaces that reflect the resource-oriented nature of the application. 7.4.1 API Structure Overview The API surface is organized into distinct blueprints to enforce a separation of concerns: •/project : Endpoints under this namespace manage the project state, handling file scanning, database persistence, and metadata preparation routines. •/pipelines : This controller manages the definition, Create/Read/Update/Delete (CRUD) operations, and execution of processing pipelines. It handles the serialization of pipeline steps and the mapping of input/output artifacts between components. •/pipeline_components : Dedicated to the component ecosystem, these endpoints facilitate the discovery, installation, and configuration of modular processors. They also expose health check and storage optimization routines. •/zenodo : These endpoints abstract the complexity of the external Zenodo API, providing simplified methods for record retrieval and file synchronization (often integrated within project routes). 7.4.2 The CLI Routes API A distinctive feature of the HDP architecture is the cli_routes_api ( /cli ), designed specifically to support robust batch scripting. Recognizing that GUI-driven endpoints often include logic for user feedback (e.g., flash messages) or single-item error blocking, the CLI routes are engineered for fault tolerance. For example, the /cli/pipelines/<name>/execute endpoint implements a ”Multi-Status” response pattern (HTTP 207). Instead of halting a batch operation upon a single record failure (a scenario catastrophic in long-running automated tasks), this endpoint isolates exceptions. It processes the entire list of provided record IDs and returns a comprehensive report detailing successful transactions alongside specific error messages for failed items, enabling retry logic without manual intervention. 7.5 WebHDP: The Server-Hosted Browser Interface Complementing the Electron desktop client and the CLI, WebHDP represents the third component of the system’s interface strategy. Currently in active development and successfully validated in internal trials, WebHDP transforms the HDP from a single-user local tool into a centralized, multi-user platform suitable for institutional deployment. 7.5.1 Microservice Architecture and Stack WebHDP is architected as a containerized microservice ecosystem orchestrated via Docker Compose. • Reactive Frontend: The user interface is a Single Page Application (SPA) built with Vue 3 [45] and Vite [46], utilizing Pinia [47] for state management and Tailwind CSS [12] for responsive design. Unlike the desktop version, this frontend is completely decoupled from the file system, interacting solely via HTTP requests. • Orchestration Backend: A dedicated Flask application ( web ) serves as the middleware layer. It manages user authentication using JSON Web Tokens (JWT) [48], session persistence, and database 20 Heritage Data Processor (HDP) Working Paper [v1] Alternatively, domain-specific repositories such as MorphoSource [69] or tDAR [70] offer native, integrated handling of complex metadata and visualization [71]. These platforms support granular active preservation workflows, such as format migration and deep indexing, which Zenodo’s generic schema cannot match [71]. Yet, they often impose higher barriers to entry, including membership costs, whereas Zenodo’s inclusive approach allows for the rapid, cost-free deposit of heterogeneous datasets [58]. Therefore, the HDP’s strategic choice of Zenodo represents a compromise that no longer requires sacrificing visualization: it utilizes Zenodo for robust, citable storage [64] while enabling client-side or external visualization tools to consume the data directly [66]. This model aligns with the decoupled architecture where Zenodo acts as the permanent source for assets that can be integrated into dynamic frontend experiences. 9.3.7 API Usage and Operational Stability While Zenodo offers robust archival capabilities, its operational constraints regarding API request limits and bandwidth usage make it unsuitable as a direct backend for live, high-frequency applications, positioning it instead as a persistent repository for digitized research artifacts. Moreover, Zenodo reserves the right to restrict or remove user access where it considers that use of the service “interferes with its operations” [60]. In practice, this often manifests as the blocking of IP addresses if an API client performs “an excessive amount of automated requests” which impacts service stability [59]. The HDP proactively addresses this operational risk through two specific mechanisms: • Client-Side Rate Limiting: The internal Rate Limiter enforces a mandatory delay between requests, ensuring the client respects the repository’s stability requirements and avoids triggering 429 error blocks [59]. • Pagination Avoidance: By maintaining the file inventory and deposit IDs locally in .hdpc files, HDP minimizes the need to repeatedly query the Zenodo API for full deposit lists. This significantly reduces the request volume compared to standard API clients, aligning with the requirement to avoid interference with operations [60]. 9.4 Ethical Governance: Integrating CARE Principles and Indigenous Data Sovereignty While the HDP architecture emphasizes technical efficiency and adherence to the FAIR principles, the automated processing of cultural heritage data necessitates a rigorous ethical framework. Digitization is not a neutral technical act; it is a process of interpretation that can reinforce historical power imbalances or, conversely, serve as a tool for restitution and empowerment [72]. Consequently, the HDP workflow must be evaluated not only against technical metrics but also against the CARE Principles for Indigenous Data Governance (Collective Benefit, Authority to Control, Responsibility, Ethics). 9.4.1 The Tension Between Open Data and Data Sovereignty A central tension in modern digital preservation exists between the “open by default” ethos of repositories such as Zenodo and the rights of Indigenous peoples to control data about their communities, territories, and knowledge. Carroll et al. [73] argue that while the FAIR principles focus on data attributes (machinereadability and persistence), they fail to address the people and purposes behind the data. The CARE principles address this deficit by prioritizing Authority to Control, ensuring that Indigenous rights and interests are recognized in governance protocols [73]. For HDP users, this distinction is critical. The system’s Root Mode’ or Subdirectory Mode’ scanning could inadvertently aggregate and publish sensitive data (such as sacred rituals or traditional ecological knowledge) that was never intended for public dissemination. Tsosie [74] emphasizes that data sovereignty is a foundational aspect of political self-determination, noting that tribal governments have an inherent right to govern the collection, ownership, and application of their data. This architecture supports the critical distinction 27 Heritage Data Processor (HDP) Working Paper [v1] highlighted by Lovett et al. [75]: it provides the technical mechanisms for Indigenous governance of data, which is a prerequisite for mobilizing that data for governance and internal community decision-making. The uncritical upload of such data to open repositories, even under Restricted Access, may violate this sovereignty. 9.4.2 Mitigating Algorithmic Bias and ”Data of Disregard” The automation provided by HDP also carries the risk of amplifying ”data of disregard,” defined as data that frames Indigenous peoples solely through lenses of disparity or disadvantage, rather than nation-building or cultural resilience [72]. Manžuch [72] highlights that Western metadata schemas often force Indigenous knowledge into ill-fitting categories, stripping away context and spiritual significance. To mitigate this, the HDP’s metadata_mapping module must be utilized not just for schema compliance, but for ethical contextualization: • Authority to Control (CARE): Users should utilize HDP’s exclusion filters to prevent the ingestion of sensitive subdirectories, respecting the provenance and ownership of the data as defined by the community [73]. • Responsibility (CARE): Metadata descriptions generated via the LLMProcessor should be prompted to include attribution statements or “CARE Notices” that reflect the community’s worldview, rather than relying solely on generic archival descriptions [72, 73]. 9.4.3 Operationalizing Ethics in the Pipeline Ultimately, software cannot automate ethics, but it can provide the mechanisms for ethical decision-making. The HDP’s architecture facilitates the decoupling of preservation (managed locally via .hdpc project states) from publication (upload to Zenodo). This aligns with the view that digitization should support participatory archives, where communities are active partners in the selection and presentation of their heritage [72]. By allowing users to curate specific batches for publication while retaining others in local-only storage, HDP enables a workflow where Indigenous Data Sovereignty is respected, ensuring that data functions as a resource for collective benefit rather than extraction [73, 74]. By placing these curation tools directly in the hands of communities, the system fosters the technical capacity building required to transform abstract rights into actionable data practices [75]. 9.5 Ethical Risks in Digital Dissemination: Trafficking, Sensitivity, and Security While HDP streamlines the publication process, the uncritical dissemination of high-resolution heritage data introduces significant risks that must be managed through the system’s configuration features. The transition from local storage to global platforms transforms heritage assets into accessible digital commodities, necessitating a careful evaluation of the nature and method of data sharing. 9.5.1 Countering the Illicit Trade and Looting The automation of uploads must not facilitate the illicit antiquities market. Votey [76] demonstrates that digital platforms have become central nodes in the trafficking of cultural goods, where the ease of sharing images and connecting buyers bypasses traditional regulatory checkpoints. Furthermore, Rouhani [77] warns that “open data” is not inherently benign; the publication of precise geospatial coordinates and high-resolution site maps can be utilized to facilitate looting or unauthorized excavations, particularly in conflict zones or politically unstable regions. To address this, HDP users should utilize the exclusion_filters to strip sensitive geolocation metadata from public uploads or withhold high-resolution survey data from the open Zenodo record, retaining them in the local .hdpc database for restricted access. 28 Heritage Data Processor (HDP) Working Paper [v1] 9.5.2 Bioarchaeological Ethics and Contextualization Specific ethical safeguards are required when processing human remains. Ulguim [78] argues that sharing 3D models of human remains without adequate contextual metadata risks dehumanizing the subject, reducing ancestors to mere “technical showcases” stripped of their osteobiographical narrative. The lack of consent or community consultation in these uploads can constitute an ethical violation. HDP’s metadata enforcement ensures that models are not uploaded as orphan files; the system requires the association of descriptive metadata (such as provenance, age, and ethical statements) before a batch is cleared for publication, preventing the de-contextualization that is prevalent on platforms like Sketchfab [78]. 9.5.3 Intellectual Property and Geometric Protection Finally, the protection of the digital asset itself remains a concern for custodians. Koller [79] highlight that curators are often reluctant to share high-fidelity 3D models due to the risk of intellectual property theft and unauthorized commercial reproduction (e.g., 3D printing of museum artifacts). While HDP currently uploads file-based assets, its modular architecture supports the integration of future components that could implement “remote rendering” or geometric decimation [79]. This capability allows users to share low-resolution proxies for public visualization while securely archiving the high-resolution geometry locally, balancing the mandate for public access with the imperative of data security [77]. 10 System Availability and Development Status The HDP is currently released as open-source software under the GNU General Public License v3.0 (GPLv3) [80]. The source code, binary releases, and technical documentation are available on GitHub3. 10.1 Current Release State The software described in this manuscript represents the complete architectural design and intended functionality of the HDP system. At the time of writing, the publicly available release (version v0.1.0-alpha.5 ) is in an Alpha stage of development. Consequently, readers should be aware that not all features or modules discussed in this paper are fully implemented in the latest public distribution. The current codebase serves primarily as a proof-of-concept for the proposed architecture, with core functionalities operational but advanced features potentially disabled or subject to change. 10.2 Usage and Stability Disclaimer This software is provided for testing, experimentation, and feedback purposes only. As is standard for alpha-stage research software, there is no guarantee of reliability, correctness, or continuous availability. It is not recommended for deployment in production environments or for use in safety-critical systems without extensive local validation. Furthermore, backward compatibility for project files ( .hdpc ) and configuration schemas is intended, but not guaranteed during this phase, as APIs and data models may evolve significantly between versions. We invite the Digital Humanities community to participate in the testing process and to consult the repository for the most up-to-date implementation roadmap. 11 Conclusion The preservation of digital cultural heritage demands a shift from ad-hoc, disparate data management practices to standardized, reproducible engineering systems. The HDP represents a significant architectural advancement in this domain, offering a robust solution to the friction between heterogeneous research data 3https://github.com/Digital-Humanities-Jena/heritage-data-processor 29 Heritage Data Processor (HDP) Working Paper [v1] and rigorous archival standards. By introducing the persistent .hdpc project state, HDP effectively decouples data preparation from immediate upload, allowing for asynchronous curation workflows that were previously difficult to manage with transient scripts. The strict modularity of the Pipeline Component ecosystem addresses the dependency management challenges endemic to complex preservation pipelines, ensuring that processing logic remains portable and reproducible across institutions. Furthermore, the dual-interface design democratizes access to these advanced tools: the Electron-based GUI lowers the technical barrier for users and domain experts, while the comprehensive CLI and REST API provide data engineers with the headless automation capabilities required for institutional-scale ingestion. 11.1 Future Developments With the architectural foundations of the hpc-connector and WebHDP established, future development will broaden its scope to address not only the federation of resources and data lifecycle resilience, but also the integration of active preservation strategies, automated ethical governance protocols, and embedded on-device intelligence. • Federated Research Infrastructure: The convergence of WebHDP and the HPC connector will enable a “Federated Processing” model. Future releases will allow multiple users via the WebHDP interface to submit jobs to a shared, authenticated HPC backend, transitioning beyond the current single-user local client model. • Expanded Component Registry: The Component Registry will be enhanced to support complex, hardware-dependent configurations declared in the component.yaml , such as CUDA versions for GPU-accelerated modules. • Repository Agnosticism: While currently optimized for Zenodo, the modular architecture will be abstracted further to support additional emergent repository standards. This ensures HDP serves as a universal bridge for digital heritage preservation, necessitating further evaluations of available repository infrastructures. 11.1.1 State Re-hydration and Large-Collection Resilience A critical area for optimization addresses the retrieval limitation inherent in the Zenodo REST API, which caps record retrieval at 10,000 items. While the HDP currently mitigates this by treating the local .hdpc database as the source of truth, the loss of this local file poses a significant risk for large-scale collections, effectively compromising synchronization integrity. To resolve this, future iterations will implement a State Re-hydration Utility designed for recovering metadata related to the processed records. Distinct from the standard REST API, this module will leverage the OAI-PMH (Open Archives Initiative Protocol for Metadata Harvesting) interface [81]. By utilizing OAI-PMH’s flow control and resumption tokens, the utility will be capable of iteratively harvesting metadata for collections of arbitrary size to reconstruct the local zenodo_records table. This ensures that the local project state can always be fully reconstituted from the remote archive, guaranteeing system resilience and data continuity regardless of local storage failures. 11.1.2 Active Preservation and Default Format Normalization Addressing the distinction between bit-level preservation (ensured by Zenodo) and functional preservation, future iterations of the HDP will introduce an integrated Active Preservation Layer. While the repository guarantees that the stored bitstream remains unchanged, it does not guard against the functional obsolescence of proprietary file formats. To mitigate this, the pipeline will implement Automated Format Normalization as a default, “opt-out” 30 Heritage Data Processor (HDP) Working Paper [v1] behavior during the ingestion phase. Upon detecting proprietary or risk-prone formats (e.g., camera-specific RAW files, proprietary 3D binaries, or closed document formats), the system will automatically generate a normalized, open-standard derivative (e.g., uncompressed TIFF, PDF/A, or ASCII-based OBJ) to serve as the long-term archival copy. By ensuring that every deposited record contains at least one representation in a documented, non-proprietary standard, the HDP effectively shifts the preservation strategy from passive storage to active risk management, ensuring the data remains interpretable by future scholarship independent of specific software vendors. 11.1.3 Automated Sensitivity Detection and Ethical Guardrails Another key area for future development is the integration of an “Automated Sensitivity Detection” middleware within the HDP pipeline. While the current architecture streamlines ingestion, it relies on the user’s manual discretion to identify sensitive materials. To better align with the CARE Principles [73], future iterations will implement algorithmic safeguards designed to flag potential ethical risks prior to repository upload. Geospatial Redaction and ”Red Lists” To mitigate the risks of looting and site exploitation facilitated by open data, the detector will analyze file metadata (EXIF/XMP) for high-precision geolocation coordinates. By cross-referencing these coordinates against defined risk zones (such as areas of active conflict or known looting hotspots) [76, 77], the system will automatically prompt the user to redact spatial data or restrict the record’s access level. This functionality aims to prevent the misuse of heritage data while maintaining the scholarly value of the non-spatial dataset [77]. Bioarchaeological and Sacred Content Flagging Building on the challenges of digital bioarchaeology, the system will employ metadata keyword scanning and potentially computer vision heuristics to identify human remains or sacred objects. Upon detection, the pipeline will halt the batch processing for those specific records, requiring the operator to explicitly confirm the presence of ethical contextualization, such as osteobiographies or consent statements [78]. This human-in-the-loop (HITL) friction is intentional, designed to prevent the inadvertent “datafication” of ancestors or sensitive indigenous knowledge [72]. Integration with Local Contexts Finally, the HDP will move toward direct integration with the Local Contexts Hub API. This will allow users to apply Traditional Knowledge (TK) and Biocultural (BC) Labels directly within the metadata_mapping wizard. By embedding these persistent digital identifiers into the .hdpc database, the system ensures that Indigenous authority and governance protocols travel with the data, regardless of the repository environment [73, 74]. By embedding these ethical and technical standards directly into the processing architecture, HDP ensures that the future of digital heritage is not only persistent but also culturally responsive and technically reproducible. 11.1.4 Embedded Intelligence via Small Language Models To enhance the system’s cognitive capabilities without compromising the local-first privacy architecture or requiring high-end workstation hardware, future releases will integrate optimized Small Language Models (SLMs) directly into the client application. Specifically, the system will leverage high-efficiency models such as Qwen3-4B-Instruct-2507 [82] to provide real-time, on-device intelligence. This embedded module will be utilized for granular tasks including metadata mapping suggestions, semantic validation of descriptive fields, and contextual UI hinting. Recognizing the variability in user hardware environments, this feature is designed with a strict ”opt-in” performance toggle; users can enable the module for advanced assistance during complex curation tasks or disable it entirely to prioritize system responsiveness on lower-specification machines. 31 Heritage Data Processor (HDP) Working Paper [v1] Acknowledgments The development of the Heritage Data Processor (HDP) was conducted within the scope of the 3DBigDataSpace Project [51]. The conceptual foundations and the predecessor software, ”Zenodo Toolbox,” were developed during the INDUX-R Project: Industrial Research and Innovation for Cultural Heritage [50]. 32 Heritage Data Processor (HDP) Working Paper [v1] References [1] Mark D Wilkinson et al. “The FAIR Guiding Principles for scientific data management and stewardship”. In: Scientific data 3.1 (2016), pp. 1–9. [2] Dominik Ukolov. Zenodo Toolbox. 2024. url: https://github.com/Digital-Humanities-Jena/ zenodo-toolbox (visited on 09/24/2024). [3] European Organization For Nuclear Research and OpenAIRE. Zenodo. en. 2013. doi: 10.25495/7GXKRD71.url:https://www.zenodo.org/ (visited on 11/21/2025). [4] YAML Language Development Team. YAML Ain’t Markup Language (YAML) Version 1.2. Revision 1.2.2 (2021-10-01). Tech. rep. Human-readable data serialization standard commonly used for configuration files. YAML Language Development Team, Oct. 2021. url: https://yaml.org/spec/1.2.2/ (visited on 11/21/2025). [5] Astral. GitHub: uv. Astral, Nov. 2025. url: https : / / github . com / astral - sh / uv (visited on 11/21/2025). [6] Dirk Merkel et al. “Docker: lightweight linux containers for consistent development and deployment”. In: Linux j 239.2 (2014), p. 2. [7] Guido Van Rossum, Fred L. Drake, and Python Development Team. Python. Python Programming Language, Version 3.11.14. Python Software Foundation, 2022. url: https://docs.python.org/3.11/ (visited on 11/21/2025). [8] OpenJS Foundation. Electron. Build Cross-Platform Desktop Apps with JavaScript, HTML, and CSS. OpenJS Foundation. url:https://www.electronjs.org/ (visited on 11/21/2025). [9] The Chromium Project. Blink. Browser Rendering Engine of the Chromium Project. The Chromium Project. url:https://www.chromium.org/blink/ (visited on 11/21/2025). [10] Node.js contributors. Node.js: JavaScript Runtime. OpenJS Foundation. url: https://nodejs.org/ (visited on 11/21/2025). [11] ECMA International. ECMAScript 2015 Language Specification. Standard ECMA-262, 6th Edition. Tech. rep. ECMA International, 2015. url: https://262.ecma-international.org/6.0/ (visited on 11/21/2025). [12] Tailwind Labs. Tailwind CSS. A Utility-First CSS Framework.url: https://tailwindcss.com/ (visited on 11/21/2025). [13] Roy Thomas Fielding. “Architectural Styles and the Design of Network-based Software Architectures”. PhD thesis. University of California, Irvine, 2000. url: https://www.julianbrowne.com/article/100year-architecture/assets/downloads/fielding-dissertation.pdf (visited on 11/21/2025). [14] Pallets Projects. Flask. A Python Micro Web Framework.url: https://flask.palletsprojects.com/ (visited on 11/21/2025). [15] World Wide Web Consortium (W3C). WebAssembly Core Specification. W3C Recommendation. Version 1.0. World Wide Web Consortium (W3C), Dec. 5, 2019. url: https://www.w3.org/TR/wasmcore-1/ (visited on 11/21/2025). [16] SQLite Development Team. SQLite. Self-Contained SQL Database Engine.url: https://sqlite.org/ (visited on 11/21/2025). [17] npm Inc. npm. Node Package Manager. Default package manager and registry for Node.js and JavaScript packages. npm, Inc., 2025. url:https://www.npmjs.com/ (visited on 11/21/2025). [18] Guido Van Rossum and Python Development Team. venv — Creation of Virtual Environments. Python Software Foundation. 2025. url: https://docs.python.org/3.11/library/venv.html (visited on 11/21/2025). 33 Heritage Data Processor (HDP) Working Paper [v1] [19] Python Packaging Authority. pyproject.toml Specification. Standard configuration file format for modern Python packaging, defined following PEP 518, PEP 621, Pep 639, and PEP 794. Python Packaging Authority. Oct. 2025. url: https://packaging.python.org/en/latest/specifications/pyprojecttoml/ (visited on 11/21/2025). [20] The pandas development team. pandas-dev/pandas. Python Data Analysis Library. 2020. doi: 10.5281/ zenodo.3509134.url:https://doi.org/10.5281/zenodo.3509134 (visited on 11/21/2025). [21] Charles R. Harris, K. Jarrod Millman, Stéfan J. van der Walt, et al. “Array Programming with NumPy”. In: Nature 585.7825 (2020). Core array programming library for scientific computing in Python, pp. 357– 362. doi: 10.1038/s41586-020-2649-2 .url: https://www.nature.com/articles/s41586-0202649-2. [22] libmagic. libmagic. File Type Identification Library. 2022. url: https://manpages.ubuntu.com/ manpages/jammy/man3/File::LibMagic.3pm.html (visited on 11/21/2025). [23] Adam Hupp and contributors. GitHub: python-magic. 2022. url: https://github.com/ahupp/pythonmagic (visited on 11/21/2025). [24] National Institute of Standards and Technology. Secure Hash Standard (SHS). Federal Information Processing Standards Publication FIPS 180-4. Specifies the SHA-1, SHA-224, SHA-256, SHA-384, and SHA-512 hash functions. Gaithersburg, MD: National Institute of Standards and Technology, Aug. 2015. url:https://doi.org/10.6028/NIST.FIPS.180-4 (visited on 11/21/2025). [25] Ned Freed, John Klensin, and Tony Hansen. Media Type Specifications and Registration Procedures. Request for Comments 6838. Standardizes media (MIME) type naming and registration procedures. Internet Engineering Task Force, 2013. doi: 10.17487/RFC6838 .url: https://doi.org/10.17487/ RFC6838 (visited on 11/21/2025). [26] International DOI Foundation. DOI Handbook. Primary reference for the Digital Object Identifier (DOI) System. International DOI Foundation. Apr. 2023. doi: 10.1000/182 .url: https://www.doi.org/theidentifier/resources/handbook/ (visited on 11/21/2025). [27] Mike Bayer and Alembic contributors. Alembic. A Lightweight Database Migration Tool for SQLAlchemy. Open-source database schema migration tool for Python applications. 2025. url: https://alembic. sqlalchemy.org/ (visited on 11/21/2025). [28] Scott Chacon and Ben Straub. Pro Git. 2nd ed. Berkeley, CA: Apress, July 25, 2025. doi: 10.1007/9781-4842-0076-6.url:https://git-scm.com/book/en/v2. [29] Annika Jacobsen et al. FAIR principles: interpretations and implementation considerations. 2020. [30] Python Software Foundation. difflib — Helpers for computing deltas. Python Software Foundation. Oct. 2025. url:https://docs.python.org/3.11/library/difflib.html (visited on 11/21/2025). [31] John W Ratcliff, David E Metzener, et al. “Pattern matching: The gestalt approach”. In: Dr. Dobb’s Journal 13.7 (1988), p. 46. [32] ollama. GitHub: ollama. Version 0.13.0. Nov. 2025. url: https://github.com/ollama/ollama (visited on 11/21/2025). [33] vLLM. GitHub: vLLM. Version 0.11.2. Nov. 2025. url: https://github.com/vllm-project/vllm (visited on 11/21/2025). [34] Woosuk Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention. 2023. arXiv: 2309.06180 [cs.LG].url:https://arxiv.org/abs/2309.06180. [35] Zenodo. Zenodo REST API. 2025. url:https://developers.zenodo.org/ (visited on 11/21/2025). [36] The Linux Foundation. Software Package Data Exchange (SPDX) Specification, Version 3.0.1. The Linux Foundation. Apr. 2024. url:https://spdx.dev/specifications/ (visited on 11/21/2025). [37] International Organization for Standardization. Information technology – SPDX specification V2.2.1. Standard. Geneva, CH: International Organization for Standardization, Aug. 2021. 34 Heritage Data Processor (HDP) Working Paper [v1] [38] Dominik Ukolov. HDP Component — Fast Renderer (fast_renderer) (0.1.0). 2025. doi: 10.5281/ zenodo.17622912.url:https://doi.org/10.5281/zenodo.17622912. [39] Blender. Blender 5.0. Oct. 2025. url:https://www.blender.org/ (visited on 11/27/2025). [40] Python Software Foundation. argparse — Parser for command-line options, arguments and subcommands. Python Software Foundation. Oct. 2025. url: https://docs.python.org/3.11/library/ argparse.html (visited on 11/21/2025). [41] Kyzer Davis, Brad Peabody, and Paul J. Leach. Universally Unique IDentifiers (UUIDs). RFC 9562. Internet Engineering Task Force, May 2024. doi: 10.17487/ RFC9562 .url: https:// www . rfc - editor.org/info/rfc9562 (visited on 11/21/2025). [42] Michael Appleby et al. International Image Interoperability Framework (IIIF) Presentation API 3.0. Specification. Version 3.0.0. IIIF Consortium, June 2020. url: https://iiif.io/api/presentation/ 3.0/ (visited on 11/21/2025). [43] Antoine Isaac et al. “Europeana data model primer”. In: (2013). url: https://pro.europeana. eu / files / Europeana _ Professional / Share _ your _ data / Technical _ requirements / EDM _ Documentation/EDM_Primer_130714.pdf. [44] IAU FITS Working Group. Definition of the Flexible Image Transport System (FITS). Standard. Version 4.0. Version 4.0, updated 2018-08-13. International Astronomical Union, Aug. 2018. url: https://fits.gsfc.nasa.gov/fits_standard.html (visited on 11/21/2025). [45] Evan You and The Vue Team. Vue.js: The Progressive JavaScript Framework. Version 3.0. Vue.js, 2025. url:https://vuejs.org (visited on 11/21/2025). [46] Evan You and The Vite Team. Vite: The Build Tool for the Web. 2025. url: https://vite.dev (visited on 11/21/2025). [47] Eduardo Posva. Pinia: The Intuitive Store for Vue.js. 2025. url:https://pinia.vuejs.org (visited on 11/21/2025). [48] Michael Jones, John Bradley, and Nat Sakimura. JSON Web Token (JWT). RFC 7519. Internet Engineering Task Force, May 2015. doi: 10.17487/RFC7519 .url: https://www.rfc-editor.org/ info/rfc7519 (visited on 11/21/2025). [49] Michael Bayer. “SQLAlchemy”. In: The Architecture of Open Source Applications Volume II: Structure, Scale, and a Few More Fearless Hacks. Ed. by Amy Brown and Greg Wilson. aosabook.org, 2012. url: http://aosabook.org/en/v2/sqlalchemy.html (visited on 11/21/2025). [50] INDUX-R Project: Industrial Research and Innovation for Cultural Heritage. https://cordis.europa. eu/project/id/101135556. Project ID: 101135556. 2024. [51] 3DBigDataSpace Project.https://3dbigdataspace.eu/. 2025. (Visited on 11/21/2025). [52] Richard Khulusi, Jakob Kusnick, Josef Focht, and Stefan Jänicke. “musiXplora: Visual Analysis of a Musicological Encyclopedia”. In: Proceedings of the 15th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications (IVAPP). Vol. 3. SCITEPRESS, 2020, pp. 76–87. doi:10.5220/0008977100760087. [53] Dominik Ukolov, Josef Focht, and Research Center DIGITAL ORGANOLOGY. musiXplora: Documentation - Data Retrieval. June 2024. doi:10.5281/zenodo.11582200. [54] Dominik Ukolov. Large-Scale Georeferenced 3D Reconstruction of Pipe Organs from Multi-Source Heritage Data: An End-to-End Pipeline. Version v1. 2025. doi: 10.5281/zenodo.17629438 .url: https://doi.org/10.5281/zenodo.17629438. [55] Sean R Wilkinson et al. “Applying the FAIR principles to computational workflows”. In: Scientific Data 12.1 (2025), p. 328. [56] Gaelle Bequet, Martin Matthiesen, Jessica Parland-von Essen, and René Belsø. “Risks and Trust in Pursuit of a Well-functioning Persistent Identifier Infrastructure for Research”. PhD thesis. Knowledge Exchange, 2021. 35 Heritage Data Processor (HDP) Working Paper [v1] [57] Zenodo. Infrastructure and Organisation. https://about . zenodo.org/infrastructure/ . 2025. (Visited on 11/21/2025). [58] Irene del Rosario Crespo Garrido, María Loureiro García, and Johannes Gutleber. “The Value of an Open Scientific Data and Documentation Platform in a Global Project: The Case of Zenodo”. In: The Economics of Big Science 2.0: Essays by Leading Scientists and Policymakers. Springer Nature Switzerland Cham, 2024, pp. 181–200. [59] Zenodo. FAQ: Policies. https://support.zenodo.org/help/en-gb/13-policies . Includes specific policies on GDPR, DPAs, and Content Removal. Published updates: 05/12/2024. 2024. [60] Zenodo. Terms of Use.https://about.zenodo.org/terms/. DOI: 10.5281/zenodo.3896780. 2020. [61] Vinod Gurav and Sudhir R Nagarkar. “ZENODO: A PLATFORM FOR OPEN ACCESS AND SUSTAINABLE DIGITAL RESEARCH REPOSITORY”. In: TechnoLibrarianship: A Gateway Towards Future Libraries-2025, Shivaji University, Kolhapur, PP-152-159 (2025). [62] Matteo Lombardi et al. “Sustainability of 3D heritage data: life cycle and impact”. In: Archeologia e Calcolatori 34.2 (2023), pp. 339–356. [63] Simon Schiff and Ralf Möller. “Persistent Data, Sustainable Information.” In: CHAI@ KI. 2023, pp. 5–14. [64] Nicola Amico and Achille Felicetti. 3D Data Long-Term Preservation in Cultural Heritage. eArchiving Initiative, European Commission. eArchiving Initiative, 2024. [65] Erik Champion and Hafizur Rahaman. “Survey of 3D Digital Heritage Repositories and Platforms”. In: Virtual Archaeology Review 11.23 (2020), pp. 1–15. [66] Sander Münster et al. “4D Geo Modelling from Different Sources at Large Scale”. In: Proceedings of the 6th workshop on the analySis, Understanding and proMotion of heritAge Contents. 2024, pp. 13–17. [67] Igor Bajena et al. “DFG-3D-Viewer–Development of an infrastructure for digital 3D reconstructions”. In: Digital Humanities 2022 Conference Abstracts The University of Tokyo, Japan 25-29 July 2022. DH2022 Local Organizing Committee. 2022, pp. 117–120. [68] Sander Münster et al. “A digital 4D information system on the world scale: research challenges, approaches, and preliminary results”. In: Applied Sciences 14.5 (2024), p. 1992. [69] Doug M Boyer, Gregg F Gunnell, Seth Kaufman, and Timothy M McGeary. “Morphosource: archiving and sharing 3-D digital specimen data”. In: The Paleontological Society Papers 22 (2016), pp. 157–181. [70] Francis P McManamon, Keith W Kintigh, Leigh Anne Ellison, and Adam Brin. “tDAR: a cultural heritage archive for twenty-first-century public outreach, research, and resource management”. In: Advances in Archaeological Practice 5.3 (2017), pp. 238–249. [71] Juliet L Hardesty et al. “3D Data Repository Features, Best Practices, and Implications for Preservation Models”. In: College & Research Libraries 81.5 (2020), p. 789. [72] Zinaida Manžuch. “Ethical issues in digitization of cultural heritage”. In: Journal of Contemporary Archival Studies 4.2 (2017), p. 4. [73] Stephanie Carroll et al. “The CARE principles for indigenous data governance”. In: Data science journal 19 (2020). [74] Rebecca Tsosie. “Tribal data governance and informational privacy: Constructing indigenous data sovereignty”. In: Mont. L. Rev. 80 (2019), p. 229. [75] Raymond Lovett et al. “Good data practices for Indigenous data sovereignty and governance”. In: Good data 2019 (2019), pp. 26–36. [76] Maxwell Votey. “Illicit Antiquities and the Internet: The Trafficking of Heritage on Digital Platforms”. In: NYUJ Int’l L. & Pol. 54 (2021), p. 659. [77] Bijan Rouhani. “From ruins to records: Digital strategies and dilemmas in cultural heritage protection”. In: J. Art Crime 2025 (2025), pp. 35–51. 36