scieee AI-readable full text Open interactive document viewer

D2.3 Data Centre Services - Operation Management Policy

DEMO, Rudy

Abstract

This D2.3 Data Centre Services – Operation Technical Documentation covers HORTUS, the platform delivering the ITSERR WP2 related services of ITSERRto the Religious Studies research community.

Full text

Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 1 Ir00000014 - Itserr Data Centre Services - Operation Management Policy Document reference: ITSERR-WP2-D2.3-DCS-OperationTechnicalDoc Version number: 01.01 Status: FINAL Last revision date: 24/10/2025 by: UNIPA Verification date: dd/05/2023 by: Board Approval date: DD/MM/YYYY by: MUR Subject: IR00000014 - ITSERR Data Centre Services - Operation Management Policy Filename: ITSERR_WP2_D2.2_DC_OperationManagementPolic y_01.00_FINAL.docx Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 2 This document is available in the ITSERR WP Management document repository at: ITSERR-WP1\05_Deliverables\D1.2_QAP\ Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 3 Change history Version Number Date Status Summary of main or important changes 00.01 20/11/2024 WORKING Working version 00.02 17/10/2025 DRAFT Draft version contains updates from D4Science and Fincons. 01.00 24/10/2022 FINAL Final version 01.01 28/10/2025 FINAL Readiness improvements Distribution List Name Company Role ITSERR Board Members All ITSERR partners Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 4 Table of Contents 1 Document Info .............................................................................................................................. 11 1.1 Scope ....................................................................................................................................... 11 1.2 Applicable documents ............................................................................................................. 11 1.3 Reference documents ............................................................................................................. 11 2 Introduction .................................................................................................................................. 12 2.1 Context of the ITSERR Project ................................................................................................. 12 2.2 The ITSERR HORTUS ................................................................................................................ 12 2.3 Objectives of this document ................................................................................................... 15 2.4 Discriminative Application of the Operational Management Policy ....................................... 18 3 Operational Governance ............................................................................................................... 20 3.1 Organizational Structure ......................................................................................................... 20 3.2 Organizational Structure - Third-party stakeholders .............................................................. 21 3.3 Roles and Responsibilities ....................................................................................................... 21 3.4 Decision-Making Processes ..................................................................................................... 24 3.5 Interface with ITSERR Governance ......................................................................................... 24 3.6 Escalation path ........................................................................................................................ 25 4 Core ITIL Functions ........................................................................................................................ 27 4.1 IT Operations Management .................................................................................................... 27 4.2 Service Desk ............................................................................................................................ 27 4.3 Technical Management ........................................................................................................... 28 4.4 Application Management ........................................................................................................ 28 4.5 Relevance for ITSERR HORTUS Operations ............................................................................. 29 5 Service Operation Model .............................................................................................................. 30 5.1 Service Desk & Support ........................................................................................................... 30 5.2 Monitoring and Control Service (Event Management) ........................................................... 31 5.3 Incident Management ............................................................................................................. 32 5.4 Request Fulfilment .................................................................................................................. 33 5.5 Problem Management ............................................................................................................ 33 5.6 Access Management ............................................................................................................... 34 5.7 Technical Management ........................................................................................................... 34 5.8 Application Management ........................................................................................................ 35 5.9 IT Operations Management .................................................................................................... 35 Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 5 5.10 External Service Integration .................................................................................................. 36 5.11 Services escalation path ........................................................................................................ 37 5.12 Conclusion ............................................................................................................................. 37 6 ITIL-Aligned Operational Management Task Catalogue ............................................................... 38 6.1 Purpose & context ................................................................................................................... 38 6.2 Service Desk Triage (Level 1 Support) ..................................................................................... 38 6.3 Account Management operations .......................................................................................... 41 6.4 Steady-state application maintenance in VREs ....................................................................... 43 6.5 Monthly Service Reporting ...................................................................................................... 46 6.6 RACI Table ............................................................................................................................... 48 7 Service Level Management ........................................................................................................... 50 7.1 Service Description .................................................................................................................. 50 7.2 Service Level Objectives (SLOs) ............................................................................................... 50 7.3 Measurement Criteria ............................................................................................................. 50 7.4 Reporting and Monitoring ....................................................................................................... 51 7.5 Roles and Responsibilities ....................................................................................................... 51 7.6 Service Review and Improvement .......................................................................................... 51 7.7 Remedial Measures ................................................................................................................. 51 7.8 Duration and Termination ....................................................................................................... 52 7.9 Amendments and Changes ..................................................................................................... 52 8 Operational Standards and Procedures ........................................................................................ 53 8.1 Compliance with Standards .................................................................................................... 53 8.2 Version Control and Change Management ............................................................................. 54 8.3 Incident Management Procedures .......................................................................................... 55 8.4 Request Management Procedures .......................................................................................... 55 9 Risk Management ......................................................................................................................... 56 9.1 Risk Identification and Assessment ......................................................................................... 56 9.2 Risk Mitigation Strategies........................................................................................................ 56 9.3 Risk Monitoring and Reporting ............................................................................................... 56 10 Security and Privacy Management ............................................................................................... 58 10.1 Security Protocols ................................................................................................................. 58 10.2 Data Protection and Encryption ............................................................................................ 58 10.3 Compliance with Regulations (GDPR, etc.) ........................................................................... 58 Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 6 10.4 User Privacy and Confidentiality ........................................................................................... 59 11 Documentation Management ...................................................................................................... 60 11.1 Storage, Version Control, and Accessibility .......................................................................... 60 11.2 Regular Review and Updating ............................................................................................... 60 12 Maintenance and Sustainability ................................................................................................... 61 12.1 Maintenance Strategy and Schedule .................................................................................... 61 12.2 Update and Upgrade Cycles .................................................................................................. 61 12.3 Long-Term Sustainability Plan ............................................................................................... 62 13 Remedial Measures ...................................................................................................................... 63 13.1 Procedures for Addressing Service Failures .......................................................................... 63 13.2 Corrective Action Plans ......................................................................................................... 63 14 Duration, Amendments, and Termination of the OMP................................................................ 65 14.1 Policy Duration and Review Schedule ................................................................................... 65 14.2 Amendment Procedures ....................................................................................................... 65 14.3 Termination Conditions and Procedures .............................................................................. 65 15 Appendices ................................................................................................................................... 67 15.1 OMP – Service Catalogue – Interaction Diagrams ................................................................ 68 15.2 OMP – Service Catalogue – Check List .................................................................................. 73 15.3 External Service Integration (ESI) Playbook .......................................................................... 74 15.4 RESILIENCE Risk Management Model of Resilience .............................................................. 78 Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 7 List of Tables Table 1: Acronyms .......................................................................................................................... 10 Table 2: Issues Priority Definition .................................................................................................... 39 Table 3:Service Desk Triage (Level 1 Support) KPIs ......................................................................... 41 Table 4: Account Management KPIs ............................................................................................... 43 Table 5:Steady-state application maintenance KPIs ....................................................................... 46 Table 6: Monthly Service Reporting KPIs ........................................................................................ 47 Table 7: OMS - RACI Table .............................................................................................................. 49 List of Figures Figure 1: HORTUS logo ................................................................................................................... 12 Figure 2: Internal Organisation Structure ........................................................................................ 20 Figure 3: external stakeholders ....................................................................................................... 21 Figure 4: ITIL V3 Service Operation Model ...................................................................................... 30 Glossary and Acronyms Acronym Full Form AAI Authentication and Authorization Infrastructure ACK Acknowledgement – confirmation that a request or ticket has been received. AI Artificial Intelligence API Application Programming Interface – a set of functions and protocols enabling software components to communicate. APISIX Apache APISIX – a high-performance, cloud-native API gateway. ASVS Application Security Verification Standard – an OWASP framework for verifying web application security. AWS Amazon Web Services BE Belgium (country code) – used for regional context in some requirements. CAB Change Advisory Board CAP Capacity (planning) – managing sufficient compute/storage resources to meet service demand. CCP Cloud Computing Platform (D4Science service for compute workloads) CI/CD Continuous Integration / Continuous Deployment CMDB Configuration Management Database – a database that stores information about all components (CIs) used in IT services. CNR Consiglio Nazionale delle Ricerche – Italy’s National Research Council. CSAT Customer Satisfaction – a user rating of service quality. CSI Continual Service Improvement CTO Chief Technical Officer Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 8 DB Database – used generically to describe the data storage layer (e.g., PostgreSQL, MongoDB). DC Data Centre DCS Data Centre Services – the ITSERR work package covering infrastructure services. DPO Data Protection Officer – the role responsible for overseeing data protection and GDPR compliance. ELK Elasticsearch, Logstash, Kibana (centralized logging stack) ERIC European Research Infrastructure Consortium – a legal framework for pan-European research infrastructures. ESFRI European Strategy Forum on Research Infrastructures – the roadmap framework ITSERR belongs to. EU European Union – the political union funding the project. FAIR Findable, Accessible, Interoperable, Reusable – principles for data management and stewardship. gCat D4Science Cataloging Service GDPR General Data Protection Regulation HTTP / HTTPS Hypertext Transfer Protocol / Hypertext Transfer Protocol Secure – protocols for web communications. IAM Identity and Access Management IDS/IPS Intrusion Detection System / Intrusion Prevention System – tools for monitoring and defending networks against malicious activities. INFRADEV Development of New Pan-European Research Infrastructures – an EU programme that funded ITSERR. IP Internet Protocol – the principal communications protocol in the internet protocol suite. ISO International Organization for Standardization – standards body referenced for security frameworks (e.g., ISO 27001, ISO 31000). IT Information Technology – the field encompassing computer systems and networks. ITIL Information Technology Infrastructure Library (IT service management framework) ITSERR Italian Strengthening of the ESFRI RI RESILIENCE (project/Research Infrastructure) ITSM IT Service Management – the practice of designing, delivering and managing IT services. ITSO IT Support Office – the operational entity responsible for daily IT operations, maintenance and monitoring. JSON JavaScript Object Notation – a lightweight data interchange format. JWT JSON Web Token KB Knowledge Base KEDB Known Error Database – a database of known errors and their work-arounds. KPI Key Performance Indicator L1 Level 1 (first-line) support LMS Learning Management System – a software platform for training and education (e.g., Moodle). MAINT Maintenance – shorthand prefix used for maintenance KPIs (e.g., MAINT-01 for uptime vs SLR). ML Machine Learning Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 9 ML/AI Machine Learning / Artificial Intelligence – advanced analytics techniques utilised in some ITSERR services. MON Monitoring/Observability – functions responsible for collecting metrics and logs (ELK, Prometheus/Grafana). MSR Monthly Service Reporting – the process of compiling and reviewing monthly operational performance. MTBF Mean Time Between Failures MTRC Mean Time to Route Correctly – the average time from ticket open to assignment to the correct resolver group. MTTR Mean Time To Restore (or Repair) MUR Ministero dell’Università e della Ricerca – the Italian Ministry for Universities and Research, noted as the approving authority. NIST National Institute of Standards and Technology – U.S. standards body referenced for risk and security frameworks. NLP Natural Language Processing OIDC OpenID Connect (identity protocol) OMP Operational Management Policy OMSP Operation Management Service Provider OMSP Operations Management Service Provider – a provider (often external) that executes operational tasks for ITSERR. Ops Admin Operations Administrator OS Operating System – software that manages computer hardware and software resources (e.g., Linux, Windows). OWASP Open Web Application Security Project P1, P2, P3 Incident and Request classification from P1 (Critical), P2(High) to P3 (Medium/Low). PIR Post-Incident Review PMO Project Management Office PO/PL Product Owner / Project Leader – roles responsible for specific research work packages or projects. RACI Responsible, Accountable, Consulted, Informed (role matrix) RBAC Role-Based Access Control RCA Root Cause Analysis – a method for identifying the underlying cause of incidents or problems. RFC Request for Change RI Research Infrastructure S3 Simple Storage Service (AWS) SBOM Software Bill of Materials SD Service Desk – the single point of contact for all user requests and incidents (Level 1 support). SDLC Software Development Life Cycle – the structured process for planning, creating, testing and deploying software. SDP Service Design Package (or Service Delivery Plan) – documentation describing all aspects of an IT service through each stage of its lifecycle. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 16 Outlines the four foundational ITIL functions—IT Operations Management, Service Desk, Technical Management, and Application Management—and explains their relevance for stable and efficient HORTUS operations. 5 Service Operation Model Details how services are run day-to-day: Service Desk & Support, Monitoring & Event Management, Incident, Request, and Problem Management, Access Control, Technical and Application Management, and integration of external services. 6 ITIL-Aligned Operational Management Task Catalogue Translates the OMP into practical tasks: • L1 Service Desk triage • Onboarding and deployment of new VRE systems • IAM operations (role design, provisioning, de-provisioning, recertification) • Steady-state application maintenance • Monthly service reporting. Each task includes workflows, artefacts, KPIs, and checklists. 7 Service Level Management Defines service descriptions, SLOs, measurement criteria, monitoring/reporting practices, roles, and review/improvement cycles. Establishes how service quality is tracked and improved. 8 Operational Standards and Procedures Covers compliance with standards (GDPR, ISO, ITIL, FAIR), version control, change management, and procedures for incident/request handling. Provides minimum baseline practices for providers. 9 Risk Management Introduces ITSERR’s risk framework: identification, assessment, categorisation, mitigation strategies, and monitoring/reporting through monthly reviews and annual audits. 10 Security and Privacy Management Sets principles for security protocols, data protection and encryption, GDPR compliance, user privacy, and expected roles (e.g., DPO). Highlights federated IAM and secure external provider agreements. 11 Training and Documentation Outlines training plans for end-users and IT staff, documentation management (Zenodo, Git), version control, and periodic reviews to ensure up-to-date resources and skills. 12 Maintenance and Sustainability Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 17 Defines preventive maintenance, update/upgrade cycles, and lifecycle planning. Emphasises strategic partnerships and long-term sustainability of ITSERR services. 13 Remedial Measures Establishes how service failures are managed: incident detection, RCA, corrective action plans, service restoration, and continuous improvement feedback loops. 14 Duration, Amendments, and Termination Explains validity, annual reviews, amendment procedures, and conditions for termination/transition of the OMP to ensure governance continuity. 15 Appendices Provides supporting materials: interaction diagrams, RACI tables, lightweight checklists for tasks, and references to external standards and deliverables. 2.3.3 Long-term Objectives The Operational Management Policy for ITSERR provides a strategic framework designed to ensure the sustainable, secure, and efficient operation of services within the ITSERR Research Infrastructure. It addresses the comprehensive management of operational activities, resource allocation, compliance with relevant standards, risk management, and user satisfaction, ensuring that ITSERR consistently meets the evolving requirements of the religious studies research community. A primary objective of this policy is to establish consistency and standardization across all operational aspects, extending beyond software development. This unified approach guarantees seamless integration, effective collaboration, and optimal resource management across different projects and operational activities within ITSERR. The policy aims to achieve operational excellence through clearly defined roles, responsibilities, and processes, enhancing transparency and accountability. Security and privacy management constitute another fundamental pillar of the policy, reflecting the sensitive nature of data typically encountered in religious studies. The Operational Management Policy mandates advanced security measures, data protection strategies, and stringent compliance with relevant privacy regulations, thereby safeguarding the integrity and confidentiality of user data. This policy places significant emphasis on user-centric service provision. Every service under ITSERR is created and maintained with the end-user’s perspective at the forefront, ensuring intuitive use, comprehensive documentation, prompt support, and high user satisfaction. The policy fosters an environment of collaboration, promoting integration capabilities both within ITSERR’s internal systems and with external platforms, positioning ITSERR as a pivotal hub for collaborative research. Continuous improvement is a critical component of the Operational Management Policy. Through regular monitoring, feedback mechanisms, and performance assessments, ITSERR is committed to ongoing refinement of operational practices and service offerings. This dynamic process ensures that ITSERR remains adaptive to technological advancements, user feedback, and changing operational demands. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 18 Sustainability and maintenance strategies are clearly articulated within the policy, emphasizing the long-term relevance and operational efficiency of ITSERR services. Detailed strategies for regular service updates, preventive maintenance, incident resolution, and, if necessary, comprehensive service overhauls are incorporated to ensure sustained operational performance. In conclusion, the Operational Management Policy embodies ITSERR’s dedication to delivering highquality, secure, and user-focused operational services, underpinning the research infrastructure’s capacity to effectively support the dynamic needs of the religious studies research community. 2.4 Discriminative Application of the Operational Management Policy While ITSERR rigorously manages operational processes for internally managed resources, a discriminative approach is employed for in-kind contributions due to practical limitations regarding detailed evaluation and governance. The compliance of the resources (internal and/or in-kind) to this document is the responsability of the ITSO. 2.4.1 Application to In-kind Contributions (external software and IT services) • Scope of contributions. In-kind contributions include not only data or documents but also external software and IT services provided by research institutions, consortia or community projects such as Moodle, Wikibase, library catalogues or engineered analysis tools. These services may be operated on infrastructure controlled by the contributor, but they are accessed via the ITSERR HORTUS. • Recommended framework. Contributors are expected to treat this Operational Management Policy as a baseline framework. Because ITSERR cannot govern all aspects of externally operated systems, the policy is advisory rather than prescriptive for these contributions. • Due-diligence and vendor assessment. External software or services should come from reputable providers with strong security measures; providers should be vetted through a risk-assessment process that considers their security posture, including vendor security scoring, ongoing monitoring and contractual reviews. Where the contribution involves an API, ensure the provider’s controls (authentication, encryption and key management) align with ITSERR’s risk tolerance. • Inventory and documentation. ITSERR will maintain an inventory of integrated external software and services and monitor them for updates and vulnerabilities (akin to the API inventories recommended for third-party integrations. • Minimum security standards. Contributors must satisfy a minimum set of security requirements before integration: o Anonymisation. All sensitive or personal data must be de-identified or anonymised prior to being shared with ITSERR. o Malware screening and vulnerability assessment. Contributed software must be scanned for viruses or malicious code and undergo vulnerability testing appropriate Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 19 to its function (e.g., static and dynamic analysis as recommended for third-party APIs. o Secure coding and OWASP compliance. Applications should adhere to the OWASP Application Security Verification Standard (ASVS) or equivalent secure-coding guidelines; open-source components must be kept up-to-date and free of known vulnerabilities. o Encryption and access control. Data transmitted to or from external services must be encrypted commensurate with its sensitivity. Access must be restricted to authorised users through federated identity management and RBAC consistent with D4Science. o Key and credential management. Where integration relies on API keys or tokens, contributors should implement regular key rotation and ensure keys are not exposed in code or documentation (as recommended for third-party integrations. • Ongoing responsibilities. Contributors must notify the HORTUS operator of updates or security incidents affecting their service. During de-provisioning, they must ensure that no residual user accounts or access tokens remain active on their systems. The primary data and user accounts remain under the control of the external provider. 2.4.2 Application to In-house Created Resources Once the HORTUS implementation is well under way: • Full compliance. IT resources created and operated by ITSERR partners must fully comply with this Operational Management Policy and the associated security frameworks. This includes implementing advanced security controls proportional to identified operational risks and following established change-management, incident-response and patch-management processes. Use of open-source components in these resources must be tracked and monitored, and a software bill of materials (SBOM) should be maintained to support vulnerability management. • Exception management. Any deviations from the policy require documented justification and prior approval from the RI Security Officer. • Regular security audits. Internal resources are subject to periodic security assessments to verify adherence to high standards and to detect emerging vulnerabilities. Audit results inform continual improvement of security controls. • Data handling. Sensitive or confidential data must not be hosted on the ITSERR platform. Instead, alternate secure storage solutions compliant with relevant regulations (e.g., GDPR) must be used. • Sustainability. In-house services should incorporate lifecycle planning for maintenance, updates and decommissioning to ensure long-term sustainability and alignment with the HORTUS’s evolving operational governance. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 20 3 Operational Governance 3.1 Organizational Structure The organizational structure for managing operations of the IT Services within ITSERR Technical Infrastructure includes dedicated governance bodies and roles, specifically aligned with the overarching ITSERR governance model: - Board Holds overall responsibility for ensuring strategic alignment of IT operational management. - Chief Technical Officer (CTO) Provides executive oversight for compliance, security, resource allocation, and approves significant operational changes or improvements. This role is directly responsible for operational performance, continuity of services, monitoring and reporting, and executing service-level objectives. - IT Support Office (ITSO) [if no Office, then by the CTO] The operational entity dedicated explicitly to IT management, responsible for direct implementation, maintenance, and daily operation of the technical in frastructure. This unit is led by the designated Chief Technical Officer and comprises technical and operational staff from the partner institutions. - Product Owners/Project Leaders (PO/PL) These person provides support specifically related their research project/WP realised during the ITSERR project. These people are the interfaces to various research project actors involved in research/IT service user documentation maintenance and communication related to WP service issues and incidents. Oversees the coordination of IT infrastructure services. Figure 2: Internal Organisation Structure Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 21 3.2 Organizational Structure - Third-party stakeholders These are not member of the ITSERR projects but are stakeholders external to the ITSERR project that play a role in the operational management of the ITSERR HORTUS services. - Operational Maintenance Service Provider (OMSP) The external service provider that will appoint a team to perform all operation management activities described in section 6. - D4Science Infrastructure provider – Support Team (D4S-ST) The D4Science Hybrid Data Infrastructure provides Virtual Research Environments, storage, software hosting and support services. Infrastructure managers ensure that platform services meet Service Level Agreements specified on the D4Science agreed availability and response time targets. - External Service Providers (for corrective/evolutive SW maintenance) (ESP) Providers of third-party applications integrated into HORTUS (e.g., Moodle for the Training Centre). Each external service has a local PoC administrator. 3.3 Roles and Responsibilities The effectiveness of IT service operations within ITSERR depends on clearly defined and communicated roles. A detailed RACI table is provided at section 6.6. Here are the defined roles and responsibilities: Within the ITSERR organisation - Board: o Approves strategic IT service management policies. o Makes strategic decisions regarding substantial investments or changes in IT services. - Chief Technical Officer (CTO): o Directly responsible for operational management of IT services. o Ensures adherence to defined service-level objectives (SLOs), security policies, and operational guidelines. o Manages the IT operations team, coordinates technical maintenance activities, and resolves escalated incidents. o Coordinates the IT Support Office’s activities with other operational units. o Monitors monthly service-level performance, incidents, and user satisfaction. o Facilitates regular operational reviews and recommends improvements to the Board. Figure 3: external stakeholders Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 22 - IT Support Office (ITSO) [if no Office, then by the CTO]: o Provides logistical and administrative support for IT operational processes. o Liaises for IT-related communication towards other governance bodies and thirdparty stakeholders. o Responsible for daily operation, maintenance, and monitoring of IT infrastructure. o Handles technical support, incident management, routine updates, security audits, and compliance checks. o Oversees compliance with operational standards and security frameworks. o Reports directly to the CTO and indirectly to the Board through regular service performance reporting o Coordinates operation services provisioning, documentation, inventory management, and communication related to infrastructure services. - Operations Administrator (ITSO-OA): o Is a member of the ITSO dedicated to the daily operation management of the ITSERR platform. Outside the ITSERR organisation These are not member of the ITSERR projects but are stakeholders external to the ITSERR project that play a role in the operational management of the ITSERR HORTUS services. - Operational Management Service Provider (OMSP) is a team composed of: o Service Manager (interface to CTO) o The OMSP shall appoint a single Service Manager as the primary point of contact for the ITSERR Chief Technical Officer. o The Service Manager is responsible for contractual alignment, escalation handling, monthly reporting, KPI tracking, and ensuring the OMSP fulfills its commitments under the Service Level Requirements (SLR). o The Service Manager participates in regular governance meetings and is empowered to commit resources and initiate corrective actions on behalf of the OMSP. o Service Desk Team (Level 1 Operations) o The OMSP shall staff and operate the Service Desk as the single point of contact (SPOC) for all ITSERR HORTUS operational incidents and service requests. o Responsibilities include ticket logging, triage, first-level diagnosis, and routing to the appropriate resolver group (e.g., IT Support Office, D4Science Support, External Service Providers). o The Service Desk must provide user communication, acknowledgements, and status updates, ensuring compliance with agreed response and resolution times. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 23 o Wherever feasible, the Service Desk shall maximise its capacity to resolve incidents and fulfil requests (eg account management) at first level, applying knowledge bases, known error databases, and standard procedures before escalating to higher support tiers. o Monitoring Team (MON) o The OMSP shall appoint a dedicated Monitoring Team (MON) responsible for continuous oversight of all ITSERR HORTUS services. o The MON team configures and maintains dashboards and alerting systems (e.g., ELK, Prometheus/Grafana) to track performance, availability, capacity, and security posture. o The MON team collaborates with ITSO, D4Science, and external providers to escalate anomalies, enforce proactive maintenance, and provide input to capacity and risk planning. - D4Science Infrastructure provider(s) (D4S-ST) o Operate and maintain the core D4Science infrastructure (compute, storage, networking, middleware). o Provide and manage Virtual Research Environments (VREs) and core services: o StorageHub (secure file storage & preservation). o gCat (metadata cataloging). o SocialService (collaboration & notifications). o CCP (Cloud Computing Platform for ML/AI workloads). o Ticketing/Wiki for issue tracking and documentation. o Guarantee service availability and performance against agreed SLA (See R3). o Implement security measures: GDPR compliance, federated authentication (IAM), RBAC, encryption, incident detection. o Provide monitoring and logging (ELK, Prometheus/Grafana) and report incidents. o Support scalability and elasticity (Docker Swarm/Kubernetes orchestration, load balancing). o Apply preventive maintenance, updates, and backup/recovery procedures. o Liaise with ITSERR HORTUS Operators for: o Incident escalation and resolution. o Infrastructure upgrades and integration with external services. o Ensure long-term sustainability of infrastructure resources and negotiate continuation of free or subsidized services after the initial funding period. - External service/application providers (ESP) o Provide and manage their own application/service (e.g., LMS, bibliographic tools, analytics engines). o Ensure compatibility with ITSERR policies (e.g., data protection, AAI integration, access control). o Provide a responsible contact for configuration, maintenance, and second-line support. o Guarantee minimum security standards before integration: Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 24 o Secure coding practices (e.g., OWASP ASVS). o Malware/vulnerability checks before deployment. o Encryption of sensitive data in transit and at rest. o Implement/maintain federated identity (integrating with D4Science IAM). o Notify ITSO and about updates, changes, or security incidents. o Participate in incident/problem resolution when their service impacts HORTUS workflows. o Provide documentation, versioning, and SLA commitments where applicable. 3.4 Decision-Making Processes Effective IT operational governance relies on structured, transparent decision-making: - Strategic IT Decisions (e.g., major infrastructure investments, new service adoption): Approved by the Board, based on recommendations and strategic analysis presented by the Chief Technical Officer and the ITSO. - Tactical IT Decisions (e.g., routine service-level adjustments, operational improvements): Managed and decided by the Chief Technical Officer, based on ongoing operational assessments, performance monitoring, and user feedback. - Operational IT Decisions (e.g., daily management, incident response, maintenance activities): Handled directly by the IT Support Office with regular reporting and coordination with the CTO. - Incident and Problem Management Decisions: Follow defined escalation processes, clearly documented in the Service Level Requirements (SLR) [R3] and approved by the CTO. Highimpact incidents require immediate notification and potential intervention or approval by the Board. All decisions are documented, monitored, and regularly reviewed to ensure ongoing compliance and operational excellence. 3.5 Interface with ITSERR Governance The operational governance of IT services within the ITSERR Technical Infrastructure maintains clear and structured interfaces with broader ITSERR governance entities: - Board: o Acts as the primary interface for strategic and high-level operational alignment, resource authorization, and overall accountability for the performance of IT infrastructure services. - Chief Technical Officer(CTO): o Serves as the main operational governance lead, regularly interfacing with the board to report IT service performance, compliance status, risks, and recommended operational improvements. - IT Support Office (ITSO): o Provides regular, detailed operational reports to the PMO and escalates strategiclevel issues to the Chief Technical Officer as required. It ensures transparent and Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 25 efficient communication of operational metrics, incidents, and service-level adherence. By clearly establishing these interfaces, ITSERR ensures integrated and robust governance of its IT services, effectively linking operational activities with strategic oversight, compliance, and continuous improvement objectives. 3.6 Escalation path Here is a two-part escalation path derived from this OMP. a) Internal to ITSERR 1. Entry & triage All incidents/requests enter through the Service Desk acting as the single point of contact; the desk logs, acknowledges, classifies (incident vs request), and routes based on impact/urgency and the SLR categories. 2. Operational handling The IT Support Office (ITSO), with its Operations Administrator, owns day-to-day operations: monitoring, incident handling, maintenance, documentation, and user communications; ITSO attempts restoration or workaround and keeps the ticket updated. 3. Escalation to CTO ITSO escalates to the Chief Technical Officer when any of the following occur: a critical or security incident, a breach of agreed SLO, multi-team coordination is required, or policy/major change approvals are needed; the CTO leads tactical decisions and risk/SLO ownership. 4. Escalation to Board If impact is strategic (e.g., major investment/change, high reputational or compliance risk), the CTO elevates to the Board for direction and approvals; the Board provides strategic oversight while the CTO manages execution and reporting. 5. Closure & learning After restoration, ITSO and the OMSP perform post-incident activities (e.g., PIR/RCA as applicable), updates runbooks/KB, and includes items in monthly reporting and continuous improvement. b) Between ITSERR roles and external stakeholders 1. Infrastructure issues → D4Science Support Team When symptoms indicate platform/infrastructure causes (VRE, StorageHub, gCat, SocialService, CCP, IAM), ITSO routes to D4Science Support with evidence from monitoring/logs; D4Science acts under agreed SLAs, performs fixes or workarounds, reports incidents, and hands back to ITSO for communication and closure. 2. External application issues → External Service Provider (ESP) For integrated third-party apps (e.g., Moodle), ITSO engages the ESP local administrator for second-line diagnosis, patching/configuration, and security notifications; ESP aligns with Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 32 Proactive Monitoring and Observability Enhancement: To strengthen the capacity for early detection and prevention of service issues, the ITSERR HORTUS operational framework shall adopt a proactive monitoring approach rather than relying solely on post-incident analysis. In line with industry best practices, Prometheus and Grafana are established as the de facto standard tools for metrics collection, alerting, and visualization across all platform services. To ensure optimal integration and consistency, the ITSERR PMO—in collaboration with the HORTUS Software Provider—will conduct a library and client selection process to identify the most appropriate Prometheus-compatible exporters or agents to be deployed across the service stack. Following this selection, a dedicated unified monitoring dashboard shall be developed and made available to operators, enabling real-time visibility of service health, SLA compliance, and anomaly trends. This initiative will form part of an improvement proposal before the end of ITSERR, supporting continuous service improvement and operational resilience within the ITSERR ecosystem. 5.3 Incident Management Description: Incident management focuses on restoring service operation as quickly as possible when disruptions occur. It uses information from monitoring, user reports and service desk logs to identify incidents and coordinate resolution. Responsibilities: • Classify, prioritize and investigate incidents reported by the service desk or monitoring service. • Provide rapid restoration of services by coordinating with technical and application management teams. • Escalate incidents to D4Science infrastructure providers or external service providers when necessary. • Keep users informed of progress and closure. • Conduct post-resolution reviews to identify improvements. Interfaces: • Service Desk (L1) submits incident records and receives restoration updates. • Monitoring service (MON) supplies event data that triggers incident identification. • Technical management (D4S-ST) and application management (ESP) supply diagnostics and resolution resources. • Problem management (L2 could be D4S-ST and/or ESP) receives incident history for root cause analysis. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 33 5.4 Request Fulfilment Description: Request fulfilment handles user requests that are not incidents (e.g., account creation, role assignment, workspace storage allocation, etc.). It provides efficient and standardized procedures for handling low-risk, routine service requests. Responsibilities: • Define and maintain service request models and approval workflows. • Validate user entitlements based on functional roles. • Coordinate with access management to grant or revoke rights. • Ensure timely fulfilment and communication of request status. Interfaces: • Users submit requests through the service desk or a self-service portal. • HORTUS Operational Admin. review and approve requests. • Service Desk (L1) as access management updates privileges, while technical management (D4S-ST) performs any required infrastructure security changes. • External service providers (e.g., training systems) process enrolment or provisioning actions. 5.5 Problem Management Description: Problem management aims to prevent incidents from recurring by identifying the root cause of incidents and implementing permanent solutions. It also maintains a known error database (KEDB) that can assist incident resolution. Responsibilities: • Analyze trends in incident data to detect recurring issues. • Perform root cause analysis and document known errors. • Work with technical and application management to develop permanent resolutions or work-arounds. • Coordinate changes or releases through change management when solutions require modifications to services. • Maintain documentation of problems, errors and work-arounds to improve service desk efficiency. Interfaces: • SD-L1 receive incident records and logs from incident management and operations logs 11 . • Collaborates with technical management (D4S-ST as L2) and application management ( ESP as L2) to identify root causes and implement fixes. • Notifies ITSO and monitoring (MON) services about known errors and work-arounds. 11 Operations logs and shift schedules document activities and provide a basis for problem management and reporting. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 34 • Proposes changes to change management (ITSO) and service transition processes. 5.6 Access Management Description: Access management ensures that only authorized users can use HORTUS services. It executes policies defined in service strategy and design and enforces the D4Science acceptable use policy. Responsibilities: • Provision, modify and revoke user accounts and roles. • Verify user identities and enforce multi-factor authentication where applicable. • Manage privileges for external services (e.g., Moodle) and ensure de-provisioning on exit. • Audit access logs for compliance and detect unauthorized access. • Coordinate with request fulfilment for approval workflows. Interfaces: • HORTUS users submit access requests via the service desk or self-service portal. • HORTUS Operational Admin. approve role assignments based on project requirements. • Infrastructure providers and external service providers integrate authentication and authorization mechanisms. • Incident management receives notifications of security breaches or access violations. 5.7 Technical Management Description: Technical management provides technical expertise and resources to support the IT infrastructure. It ensures that the underlying hardware, network and platform services are reliable and secure. IT operations management is divided into IT operations activities and facility management 12 . Facility management manages the physical environment such as data centers, security, power and environmental controls 13 . Responsibilities: • Plan, install, configure and maintain servers, storage, networks and middleware. • Provide technical support to incident and problem management. • Perform capacity and performance tuning and plan upgrades in collaboration with D4Science. • Manage backups, restores, print/output and job scheduling tasks 14 . 12 Facility management manages the physical environment of IT operations, including physical infrastructure and consolidation projects. 13 idem 14 Specific IT operations tasks include console management, backup and restore, print/output and job planning. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 35 • Ensure that facility management duties (e.g., power, cooling, physical security) are performed or coordinated with D4Science data centers 15 . Interfaces: • MON supported by D4S-ST monitors and control service consumes infrastructure metrics. • Incident and problem management require technical diagnostics and fixes. • Application management (ESP) relies on infrastructure for deploying and running applications. • D4Science infrastructure providers supply physical resources and infrastructure remote services. 5.8 Application Management Description: Application management is responsible for managing applications throughout their lifecycle. It ensures that the HORTUS’s integrated applications (catalog system, analytics engines, training platforms) perform efficiently and meet user requirements. Responsibilities: • Assist with design, deployment and improvement of applications integrated into HORTUS. • Provide subject-matter expertise to incident and problem management. • Implement patches, upgrades and release packages. • Define and maintain application configurations and dependencies. • Ensure that external service integrations adhere to platform policies. Interfaces: • Incident management and service desk rely on application management (ESP) to resolve application-related incidents. • Request fulfilment coordinates with application management (ESP) for provisioning features. • Technical management (D4S-ST) provides the underlying infrastructure to host applications. • ESP share application updates and collaborate on integration issues. 5.9 IT Operations Management 16 Description: IT operations management executes day-to-day operational activities to maintain stability and consistency. It has a dual role: maintaining the status quo for stability while adapting to changing business requirements. IT operations performs tasks such as console management, backup 15 Facility management manages the physical environment of IT operations, including physical infrastructure and consolidation projects. 16 These services should be further detailed in the future Request of OMP Service so that the IT-OM services are well segregated between the OMP-SP and D4Science. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 36 and restore, print/output, job scheduling and maintenance. Operations activities are executed from an operations bridge 17 . Responsibilities: • Run and monitor scheduled jobs and scripts. • Execute backups and restores; ensure data is stored securely and can be recovered when needed 18 . • Manage print/output services and ensure correct delivery of information 19 . • Perform routine maintenance and performance activities commissioned by technical or application management 20 . • Maintain operations logs and shift schedules to document activities and support problem analysis. Interfaces: • Receives operational procedures from technical management and application management. • Provides incident and problem management with logs and operational data 21 . • Coordinates with facility management for data center operations. • Communicates with the service desk to provide updates on operational status and scheduled maintenance. 5.10 External Service Integration The integration of external service is a process that involves all the project stakeholders. In such a case, we provide in section 15.3 a detailed version of this process. Description: HORTUS integrates systems (e.g., training systems, analysis tools) provided by ESP that are governed by their own administrators. These services inherit core HORTUS policies and must be monitored and managed to ensure they operate consistently with the platform. Responsibilities: • Onboard new external services, ensuring that they comply with D4Science policies and data protection requirements. • Maintain service-level agreements (SLAs) with external providers, if feasible. • Manage federated identity integration and ensure that de-provisioning removes access rights. • Coordinate incident and problem resolution with external administrators. 17 IT Operations activities are executed from an operations bridge, where services and infrastructure are monitored and managed. 18 Specific IT operations tasks include console management, backup and restore, print/output and job planning. 19 idem 20 idem 21 idem Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 37 Interfaces: • HORTUS Operational Admin. and external service administrators manage integration and escalation paths. • Service Desk - Access management synchronizes user identities and privileges across services. • Technical and application management teams integrate APIs and data flows. • Infrastructure providers may host some external services or provide network access. 5.11 Services escalation path A diagram at section 15.1.1 represents this escalation path between these ITIL services. 5.12 Conclusion The service operation model for the ITSERR HORTUS builds on ITIL guidance by establishing clear services with defined responsibilities and interfaces. Implementing a robust service desk improves user satisfaction and provides a single point of contact 22 23 . Monitoring and control through an operations bridge enable early detection of events and support rapid incident response 24 25 . Coordinated incident, request and problem management processes ensure that services remain stable while continuously improving. Technical, application and operations management teams maintain the underlying infrastructure and applications, support integration with external services and work closely with D4Science to meet availability targets. Together, these services will deliver reliable and user-focused operation of HORTUS and allow it to evolve to meet the needs of the Religious Studies community. 22 A service desk provides advantages such as improved customer service, increased accessibility and improved infrastructure management. 23 The principal goal of the service desk is to restore normal service as soon as possible, whether by resolving errors, fulfilling requests or answering questions. 24 IT Operations activities are executed from an operations bridge, where services and infrastructure are monitored and managed. 25 Specific IT operations tasks include console management, backup and restore, print/output and job planning. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 38 6 ITIL-Aligned Operational Management Task Catalogue 6.1 Purpose & context This catalogue translates the Operational Management Policy (OMP) into concrete, ITIL-aligned tasks for day-to-day operations of the ITSERR HORTUS, with specific focus on D4Science-backed Virtual Research Environments (VREs), federated identity, and continuous monitoring. It follows the service operation functions and roles defined in the OMP (Service Desk, Technical Management, Application Management, and IT Operations) and the governance split between the Chief Technical Officer, IT Support Office, D4Science infrastructure providers and the Operational Management Service providers. Foundational capabilities referenced here: • D4Science core services used by ITSERR: VRE provisioning, StorageHub (files/workspace), gCat (catalog/metadata), SocialService (notifications), CCP (compute). • Monitoring & logging stack adopted by ITSERR: ELK for logs; Prometheus/Grafana for metrics and alerting. • Service levels & operations guardrails: infrastructure availability, scheduled downtime notifications, and integrated external service rules (e.g., de-provision within 24h). 6.2 Service Desk Triage (Level 1 Support) 6.2.1 Objective Ensure all requests arriving at the Service Desk (L1 Support) are captured, acknowledged, analyzed, and routed or resolved appropriately, acting as the Single Point of Contact (SPOC) for all HORTUS users. The Service Desk operates during standard service hours: eight (8) hours per day, Monday through Friday, for a total time coverage of forty (40) hours per week. Activities are carried out on working days, excluding Saturdays, Sundays, and national holidays, unless otherwise agreed between the parties. 6.2.2 Scope & triggers • Triggered by any incoming request (incident, service request, information query, defect, or project-related request). • Applies to all HORTUS services, whether provided by D4Science, external service providers, or ITSERR project teams. 6.2.3 End-to-end workflow (ITIL mapping in ==> ITIL:) 1. Reception & acknowledgement ==> ITIL: Request Fulfilment / Incident Logging o Capture the request in the ticketing system; send an acknowledgement to the user confirming receipt and tracking ID. 2. Analysis & classification ==> ITIL: Incident/Request Categorisation Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 39 o Identify whether it is: ▪ Infrastructure issue → redirect to D4Science support team with priority, context, and evidence. ▪ Application defect → log as a bug record; assign to the IT Service/Application Provider responsible under warranty/maintenance. ▪ Information request → if knowledge base or documentation provides the answer, resolve directly; else escalate to subject matter expert. ▪ Project-related request → forward to the relevant Project Leader, including all context provided. ▪ In-scope Operational Mgt task → this task is managed directly by the Service Desk. 3. Action & routing ==> ITIL: Service Desk Resolution or Escalation o Route to the correct support group or resolve at Level 1 if possible. Maintain clear logs of actions taken and response times. 4. Incident and Request Priority Classification (P1–P3) All tickets registered in the ITSERR HORTUS Service Desk are assigned a priority level based on business impact and urgency. This classification ensures proportional response and resolution times across all support tiers (L1–L3) and aligns with the Service Level Objectives (SLOs) defined in the SLR [R3]. Priority Definition Typical Examples P1 – Critical Major outage or severe degradation of a core HORTUS function affecting a majority of users or critical research operations. No viable workaround exists. Full VRE unavailability, IAM outage, data loss risk, major security incident. P2 – High Significant functionality loss or performance degradation affecting a subset of users or key project deliverables, but a temporary workaround is possible. Partial VRE slowdown, failed deployment, IAM role provisioning issue, patch regression. P3 – Standard / Low Minor service impact, cosmetic defect, information request, or planned change with no immediate operational consequence. User request for access, UI issue, feature clarification, documentation update. Table 2: Issues Priority Definition Key principles: • Priority is assigned by L1 Service Desk during triage and may be re-evaluated by ITSO/CTO during incident review. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 40 • Escalation to higher tiers (ITSO, D4S, ESP) must respect these timeframes. • Compliance with P1–P3 timelines is measured monthly through KPIs such as MTTR, MTRC, and % tickets within SLA. • Post-Incident Reviews (PIRs) are mandatory for all P1 and P2 incidents. 6.2.4 Diagram See section §15.1.1. 6.2.5 Artifacts • Ticket log (including classification, actions, routing, SLA priority). • Bug record (with defect description, replication steps, severity). • Knowledge base artefacts (updated when information requests are resolved). 6.2.6 Service Levels 6.2.6.1 KPIs • % of tickets acknowledged within SLA (e.g., <1h for high priority). • % of tickets resolved at L1 without escalation. • Mean time to route correctly. • User satisfaction with Service Desk interactions. 6.2.6.2 Why it matters • SD-KPI-01 — Acknowledgement within SLA: Timely acknowledgements reassure users, establish trust, and kick-off coordinated response early, which reduces duplicate tickets and accelerates recovery for critical VRE services. • SD-KPI-02 — L1 Resolution Rate: A high L1 fix rate shortens cycle time and lowers cost by resolving common issues without escalation, keeping specialist teams free for complex work and improving overall service throughput. • SD-KPI-03 — Mean Time to Route Correctly (MTRC): Fast, accurate routing prevents “ticket ping-pong,” cuts MTTR, and protects availability targets for ITSERR’s VREs by getting the right experts engaged the first time. • SD-KPI-04 — User Satisfaction (CSAT): CSAT validates that operational processes actually meet researcher needs, providing a direct signal for continuous improvement and adoption of the HORTUS’s services. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 41 6.2.6.3 Target Service Levels ID Name Description Metric (formula) Service level (target)26 SD-01 Acknowledgement within SLA Tickets acknowledged on time after reception (# tickets ack’d within SLA ÷ total tickets) × 100 P1: ≤ 30m. → ≥95% P2: ≤ 2h → ≥90% P3: ≤ 6h → ≥90% (monthly avg) SD02 L1 Resolution Rate Share of tickets fully resolved at L1 (no escalation) (# resolved at L1 ÷ total tickets) × 100 ≥50% (monthly) SD03 Mean Time to Route Correctly (MTRC) Average time from ticket open to first assignment to correct resolver group Σ (time to correct routing) ÷ # routed tickets P1: ≤ 1h P2: ≤ 4 h P3: ≤ 8 h (monthly avg) SD04 User Satisfaction (CSAT) User rating of L1 interaction/resolution Average post-ticket score (1–5) or % of scores ≥4 ≥4.5/5 or ≥90% ≥4 (monthly) Table 3:Service Desk Triage (Level 1 Support) KPIs 6.3 Account Management operations 6.3.1 Objective Design, operate, and continuously improve identity, roles, and permissions across HORTUS services (including new systems), ensuring least-privilege access, timely provisioning, and compliant deprovisioning. 6.3.2 Scope & triggers • New service or feature requires new application roles; • Joiner/mover/leaver events; • User support requests (enrolment, role changes). 6.3.3 Operational activities (ITIL mapping in ==> ITIL:) 1. Role design & modelling ==> ITIL: Change Enablement o Map business roles to D4Science IAM groups/claims; define new application roles for the system (admin/maintainer/contributor/viewer); document authorization rules and SoD constraints. 26 See service conditions at §6.2.1. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 48 6.6 RACI Table # Task / Step ITSERR Operation Mgt Service Provider D4Science Third-Party CTO ITSO Oper. Admin. Serv. Desk - L1 MON Serv. Mgr Support Team ESP Project Leader 6.2 Service Desk Triage 6.2.1 Reception & acknowledgement I C I R – A I I I 6.2.2 Analysis & classification I C C R – A C C C 6.2.3 Action & routing (resolve/esc) I C C R – A C C C 6.2.4 Ticket updates & closure comms I C C R – A C C C 6.3 IAM Operations 6.4.1 Role design & modelling (SoD) A R C C – I C C C 6.4.2 Provisioning & entitlement I C C R – A R C I 6.4.3 De-provision & token revocation (≤24h) I C C R – A R C I 6.4.4 User lifecycle support (password/reset) I C C R – A I I I 6.4.5 Periodic reviews & compliance I C C R – A I I I 6.4 Steady-State App Maintenance 6.5.1 Health monitoring & observability I C R C R A C C I 6.5.2 Incident response (restore/comms/PIR) I C R C R A C C I 6.5.3 Patch & release mgt (24h notice, backup, deploy) I C R R C A C C I 6.5.4 Config & documentation (CMDB/runbook/KB) I C R R C A C C I 6.5.5 Capacity / perf / cost (periodic) I C R R C A C I I Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 49 6.5 Monthly Service Reporting 6.6.1 Define/agree template (periodic) A A R I C R C C I 6.6.2 Collect monthly data I C I R C A C C I 6.6.3 Draft & internal review I C C R C A C C I 6.6.4 Submit & pre-review C C C R I A I I I 6.6.5 Review meeting & actions A C I R I R C C I 6.6.6 CSI register update I C C R I A I I I A = Accountable, R = Responsible, C = Consulted, I = Informed, – = Not typically involved. Table 7: OMS - RACI Table Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 50 7 Service Level Management The subsections of this chapter are detailed in SLR [R3]. 7.1 Service Description A clear description of IT services provided within the ITSERR Technical Infrastructure is fundamental for effective service management. Services include a variety of technical and user-oriented offerings, designed to ensure robust and reliable infrastructure supporting the ITSERR research community. Each service is clearly described by outlining its purpose, main functionalities, and essential components, including hardware, software, network infrastructure, and supporting systems. Non-exhaustive examples of Services Provided: - Support Services: Direct user support addressing queries effectively and promptly. - Helpdesk Operations: A primary contact point providing immediate assistance and problem resolution. - Incident Management: Management and resolution of service disruptions to ensure minimal impact on operations. 7.2 Service Level Objectives (SLOs) Service Level Objectives (SLOs) define the measurable standards to which the ITSERR services adhere. These objectives ensure clarity, mutual understanding, and accountability, and they encompass key metrics such as: - Availability: Defined percentage uptime, including clear descriptions of acceptable downtime and maintenance windows. - Performance: Metrics including response times, throughput rates, and data processing speeds. - Reliability: Specification of redundancy measures, failover mechanisms, and regular preventive maintenance practices. - Incident Response Time: Clearly defined expected timeframes to acknowledge incidents, categorized by priority. - Resolution Time: Timeframes for resolving incidents according to their severity, with procedures clearly outlined for escalation when initial resolution attempts are unsuccessful. 7.3 Measurement Criteria Effective monitoring and assessment of service delivery rely on precise and clearly articulated measurement criteria, which include: - Metrics Definition: Clearly identified performance metrics, ensuring each is Specific, Measurable, Achievable, Relevant, and Time-bound (SMART). - Measurement Methods: Defined methods and standardized tools for accurate and consistent data collection and validation. - Data Collection Processes: Structured procedures for routine data gathering, analysis, and storage. Regularly collected data is analysed to assess compliance with defined SLOs and to identify trends and opportunities for service improvement. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 51 7.4 Reporting and Monitoring Regular reporting and proactive monitoring are integral for continuous improvement and transparent communication with stakeholders: - Reporting Frequency: Defined intervals for performance reports (monthly, quarterly, annually), differentiated by metrics as necessary. - Reporting Formats: Clear, accessible formats such as dashboards, detailed reports, and summary presentations ensuring comprehension by all stakeholders. - Monitoring Tools and Techniques: Use of specialized monitoring tools enabling real-time tracking of performance against objectives, proactive identification of issues, and implementation of preventive actions to maintain service standards. 7.5 Roles and Responsibilities Clear allocation of responsibilities ensures effective service delivery and management. Roles are delineated between service providers and users: - ITSERR RI Responsibilities: o Maintaining operational standards, managing incidents, performing preventive maintenance, and continuous performance monitoring. o Providing clear communication regarding service status and incident management processes. - Researchers Responsibilities: o Prompt reporting of incidents, providing necessary information for issue resolution, and adhering to agreed service usage guidelines. - Escalation Procedures: o Structured escalation hierarchy with defined contact points, enabling effective resolution of complex or unresolved incidents, clearly specifying levels of responsibility and communication pathways. 7.6 Service Review and Improvement Service delivery undergoes continuous evaluation to identify and implement improvements: - Review Process: Scheduled reviews of service performance against SLOs involving key stakeholders, ensuring alignment with user expectations and technological advancements. - Feedback Mechanisms: Structured mechanisms to regularly gather and analyse feedback from service users, informing service enhancements and prioritizing improvements. - Continuous Improvement Procedures: Established processes for testing, validating, and integrating improvements into service delivery, maintaining the relevance and effectiveness of ITSERR IT services. 7.7 Remedial Measures Defined actions address failures to meet agreed service objectives, ensuring accountability and continuous operational improvement: Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 52 - Failure to Meet SLOs: Processes for identifying the root causes of performance failures, immediate corrective actions, and detailed documentation of incidents and resolutions. - Penalties and Compensation: Clearly defined terms under which penalties or compensations are applicable, ensuring fairness and accountability. - Corrective Action Plans: Structured corrective plans to prevent recurrence of identified issues, specifying responsibilities, implementation timelines, and review procedures. 7.8 Duration and Termination Clear contractual terms outline the validity period and conditions governing the conclusion or renewal of Service Level Agreements (SLAs): - SLA Duration: Explicit definitions of contract periods, including start and end dates, and conditions for possible extension or renewal. - Conditions for Termination: Specific conditions under which either party may terminate the SLA, including required notice periods and any related costs or penalties. - Renewal Process: Detailed description of processes for SLA renewal, including timelines, responsibilities, and approval requirements. 7.9 Amendments and Changes Established processes ensure clarity, transparency, and stakeholder alignment regarding any amendments to the SLA: - Process for SLA Amendments: Clearly outlined procedures for proposing, reviewing, and approving changes to service agreements. - Documentation of Changes: Detailed record-keeping and documentation of all SLA modifications, maintaining version control and historical records for accountability. - Communication of Changes: Transparent and timely dissemination of information regarding SLA changes to all stakeholders, ensuring clear understanding and continued agreement. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 53 8 Operational Standards and Procedures This section outlines the key operational standards and procedures for managing IT services within the ITSERR infrastructure. It covers compliance requirements, version control and change management processes, as well as guidelines for incident and request management. 8.1 Compliance with Standards ITSERR will base its infrastructure on service providers that adopt industry best practices and relevant legal frameworks to ensure secure, efficient, and user-focused IT operations. The following subsections outline minimum expected standards in security, privacy, software development, and accessibility. Any specific procedures or controls implemented should be aligned with these overarching principles. 8.1.1 Security and Privacy Standards Regulatory and Framework Adherence - ITSERR services must comply with GDPR requirements, as well as any applicable national regulations on data protection and user privacy. - ITIL security management principles are recognized as a long-term objective; relevant practices and processes should be progressively integrated as resources allow. Technical Security Controls - While detailed controls (e.g., encryption in transit/at rest) are not mandated in this policy, ITSERR service providers must demonstrate sufficient safeguards to protect data, consistent with GDPR and recognized good practices. - Potential examples include encryption, secure access controls, and periodic vulnerability assessments; the specifics of these controls should be defined when implementing the Operational Management Policy. Compliance Reporting - Monthly reporting to the Chief Technical Officer is required from all D4S-ST and ESP, highlighting security posture, any detected vulnerabilities, and incidents that occurred or were mitigated. 8.1.2 Software Development and Documentation Standards All these standards have been developed in the Software Development Plan Template [R6]. Software Development Lifecycle (SDLC) - ITSERR follows Agile and Scrum methodologies for software development. All teams involved in developing or maintaining software must plan, implement, and iterate in accordance with these methodologies. Open Science/Open Source Standards Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 54 - ITSERR encourages use of open-source technologies and aims to align with open science principles. Repositories must be managed in a transparent, collaborative manner, supporting knowledge sharing and community-driven development. Reference to Existing Documentation Guidelines - A dedicated Software Development Plan (SDP) template (See [R4]) and associated guidelines exist to standardize coding practices, version control usage, and documentation artifacts. This Operational Management Policy does not introduce additional documentation requirements; rather, it references the approved SDP for details on requirements, design documents, testing, and release notes. 8.1.3 Accessibility and Usability Standards Public Funding Requirements - Because ITSERR is publicly funded, user-facing systems must adhere to commonly accepted usability and accessibility guidelines, such as WCAG 2.1 principles where feasible. - Any interfaces, portals, or websites serving researchers or the public should ensure inclusive access. Future Testing and Guidelines - Although no formal usability testing framework is in place yet, ITSERR acknowledges the importance of user experience (UX) testing. Procedures will be defined to incorporate accessibility and usability reviews into all front-end IT products, projects, and services as they mature. 8.2 Version Control and Change Management Version Control Tools - Git is the mandated version control system for all software projects under ITSERR. Repositories must follow best practices, including clear commit messages, feature branches, and merge reviews. Change Advisory Board (CAB) - While a dedicated CAB is not yet established, this policy recognizes the importance of having a formal approval process for production-critical changes. The Project Management Office, or a sub-group designated by it, may serve in this advisory capacity until a standalone CAB is formed. Release Management - Release management procedures (e.g., designated release windows, notifications to stakeholders, and roll-back plans) are defined in the Software Development Plan [R4] This includes tracking release versions in Git and ensuring that any new release meets necessary testing and documentation requirements. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 55 8.3 Incident Management Procedures Classification of Incidents - ITSERR follows the classification system defined in section 6.2, which categorizes incidents (e.g., Critical(P1), High(P2), Medium/Low(P3)) and sets associated response times. Incident Reporting and Tools - The specific incident management platform is not yet defined. Possible solutions include Jira Service Desk, ServiceNow, or a tailored portal. - Incident reporting should be standardized across ITSERR services, ensuring all users know how to escalate issues. Incident Response - The OMSP handles incident responses, investigations, and resolution activities. ITIL-Informed Practices - ITSERR incident response aligns with ITIL principles to ensure consistent processes for triage, analysis, resolution, and post-incident review. 8.4 Request Management Procedures Differentiation Between Incidents and Requests - Incidents relate to service disruptions or failures (e.g., bugs, outages), whereas Requests (often new features or enhancements) capture user-driven improvements to existing services. User Request Channels - A service on ITSERR portal (e.g. a contact form, or feedback form) will be in place for researchers and other stakeholders to submit requests, directed towards the RI Helpdesk and then forwarded to the relevant body or person. Users should receive clear guidance on how to log feature enhancements or routine service requests. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 56 9 Risk Management ITSERR recognizes that proactive risk management is essential for maintaining service reliability, protecting data, and ensuring compliance with internal policies and external regulations. This section outlines the framework for identifying, assessing, mitigating, and reporting risks related to IT operations. ITSERR will apply the RESILIENCE Risk Management Model as summarised in Section 15.4. 9.1 Risk Identification and Assessment Internal Procedures - ITSERR does not currently adopt a formal industry-standard risk management framework (such as ISO 31000 or NIST SP 800-30). Instead, it employs an internal, generic procedure tailored to the organization’s scale and resource constraints. This procedure is detailed in section 4 of D6.4 Quality Assessment Plan [R7]. Responsibilities - Both the IT Support Office and the Project Management Office (PMO) are jointly responsible for identifying and reviewing new or evolving risks. - Risks may be discovered through routine operations, incident reports, audits, or ad hoc observations by staff and stakeholders. Risk Categorization and Matrix - Identified risks are evaluated based on likelihood and impact, plotted on a risk matrix to determine priority and urgency. - Acceptable risk thresholds are established by the PMO in consultation with the IT Support Office, ensuring alignment with ITSERR’s strategic objectives and available resources. 9.2 Risk Mitigation Strategies Standard Controls - Backups, disaster recovery testing, patch management, and other preventive controls are documented as part of ITSERR’s baseline mitigation strategies. These measures help minimize service downtime, data loss, and security breaches. Budget and Resource Allocation - The Research Infrastructure must forecast budgetary needs to address critical or high-priority risks. Where shortfalls exist, the PMO should be informed of the potential impact on service continuity or data protection, and any decisions to accept residual risks must be formally recorded. 9.3 Risk Monitoring and Reporting Frequency of Reviews and Audits Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 57 - Monthly Risk Reviews: The IT Support Office and PMO conduct regular risk evaluations at least once a month, focusing on newly identified risks, the status of ongoing mitigations, and any changes in impact or likelihood. - Annual Risk Audits: A more in-depth audit is performed yearly, reviewing the effectiveness of existing controls, identifying emerging threats, and ensuring alignment with the Operational Management Policy. Reporting Requirements - Monthly Reports: The Chief Technical Officer consolidates risk-related updates and submits them to the PMO and, if necessary, the Board. These reports summarize new risks, status of mitigations, and potential resource implications. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 64 Lessons learned from these incidents inform process improvements, including possible amendments to standard operating procedures, training, or resource allocations. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 65 14 Duration, Amendments, and Termination of the OMP This section clarifies how long the Operational Management Policy remains valid, under what conditions it may be updated, and the scenarios leading to its termination or replacement. 14.1 Policy Duration and Review Schedule Initial Validity This Operational Management Policy (OMP) is effective upon approval by the CTO and remains in force until superseded by an updated version or formally terminated. Review Intervals The OMP is reviewed annually to confirm its ongoing relevance and compliance with evolving regulations, user needs, and infrastructure changes. Interim reviews may occur if new services are introduced or significant modifications to existing operations are required. Responsible Entities The Chief Technical Officer coordinates the review process in collaboration with the Project Management Office (PMO). Proposed amendments or major updates are presented to the Board for final approval. 14.2 Amendment Procedures Proposing Amendments Any governance body involved may propose changes to the OMP. Proposals must include a rationale, potential impacts, and revised policy text. Amendments to sections covering compliance, security, or major resource allocation require Board’s approval. Documentation and Version Control All updates or amendments are recorded with a new version number, date, and summary of changes. Previous versions are archived for reference, ensuring clear policy evolution tracking. Stakeholder Communication Once approved, amendments are communicated to all operational teams and relevant stakeholders. Mandatory briefing sessions or refresher training may be scheduled if amendments significantly alter roles, responsibilities, or processes. 14.3 Termination Conditions and Procedures Conditions for Termination Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 66 A complete termination of this OMP may be warranted if the ITSERR consortium ceases operations, undergoes major restructuring (e.g., merges with another RI), or if a new overarching policy supersedes it entirely. Notice Period A written notice must be issued at least 60 days prior to termination, specifying the rationale and the effective date. Exceptions to this notice period can apply in emergencies (e.g., regulatory directives, severe security breaches) but require explicit justification from the Board. Transition and Handover If a new policy or operational framework replaces the OMP, the Board and PMO ensure seamless transition, preserving essential processes and documentation. The Chief Technical Officer remains responsible for final archival of the retiring policy and for briefing relevant parties on new or updated governance structures. Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 67 15 Appendices Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 68 15.1 OMP – Service Catalogue – Interaction Diagrams 15.1.1 ITIL Services Diagram Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 69 15.1.2 Service Desk Triage (Level 1 Support) Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 70 15.1.3 Identity & Access Management (IAM) operations for HORTUS Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 71 15.1.4 Steady-state application maintenance in VREs Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 72 15.1.5 Monthly Service Reporting Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 73 15.2 OMP – Service Catalogue – Check List 15.2.1 Service Desk Triage • Ticket logged & ACK sent • Classified (infra/app/info/project/op-task) • Action taken (resolve or route with full context & priority) • Work notes updated; SLA class recorded • Status/resolution sent to requester; ticket closed 15.2.2 New system in a new VRE • Intake complete; success criteria agreed • Architecture & IAM fit reviewed (VRE, IAM, StorageHub, gCat, CCP) • Monitoring/alerts wired (ELK, Prometheus/Grafana) • UAT plan executed and signed • Cutover executed; rollback verified; comms sent • Catalogue entry, runbook, KB, SLOs published 15.2.3 IAM Operations • Role mapped (RBAC/claims); CTO approval when needed • Provisioning completed across D4S/ESP; evidence captured • De-provision & tokens revoked ≤24h; evidence captured • Quarterly recertification and dormant cleanup done • KB updated (password/SOPs) 15.2.4 Steady-state Maintenance • Alert triage; escalation path chosen • Restore/workaround; user comms • PIR completed; runbook/KB updated • Patch: test → notice (≥24h) → backup test → deploy → closure comms • CMDB/config & dashboards updated; monthly capacity review done 15.2.5 Monthly Service Reporting • Template valid; data collected from all sources • Draft reviewed; submit ≥5 days before month-end Ir00000014 - Itserr Status: FINAL Data Centre Services - Operation Management Policy Version: 01.01 Page 80 (End of Document)