scieee AI-readable full text Open interactive document viewer

D2.6 ML Model Certification – v1

Kao, Ching-Yu

Abstract

This deliverable presents the initial design, architecture, and implementation state of the machine learning (ML) model evidence extractors of WP2, which we call AI-SEC. They contribute to the key result KR1-EXTRACT of EMERALD, a framework to continuously extract knowledge from well-trained ML models and prepare suitable evidence based on them. EMERALD follows a knowledge graph-based approach to provide a unified view of the cloud service under certification at different layers of the service, ranging from the infrastructure layer (e.g., virtual resources), to the business layer (e.g., policies and procedures), to the implementation layer (e.g., source code files) and data layer (e.g., increasingly used AI models) in cloud applications. The ML model evidence extractors, developed in Task 2.4 and described in this deliverable, aim at identifying critical security-related features, such as adversarial robustness, privacy, security, and explainable AI. Other related deliverables in WP2, all due at project Month 12 (October 2024), provide functional and technical details on further evidence extractors from different sources, i.e., D2.2 on source code evidence extraction, D2.4 on evidence extraction from policy documents in Task 2.3, and D2.8 on runtime data extraction in Task 2.5. All these details contributed to D2.1 on the overall information model of the certification graph in Task 2.1.

Full text

Deliverable D2.6 ML Model Certification – v1 Editor(s): Ching-Yu Kao Responsible Partner: Fraunhofer Institute for Applied and Integrated Security (FhG AISEC) Status-Version: Final – v1.0 Date: 28.10.2024 Type: OTHER Distribution level: PU D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 2 of 19 www.emerald-he.eu Project Number: 101120688 Project Title: EMERALD Title of Deliverable: D2.6 – ML model certification – v1 Due Date of Delivery to the EC 31.10.2024 Workpackage responsible for the Deliverable: WP2 – Methodology for knowledge extraction Editor(s): Ching-Yu Kao (FhG) Contributor(s): -- Reviewer(s): Marinella Petrocchi CNR Cristina Martínez, Juncal Alonso (TECNALIA) Approved by: All Partners Recommended/mandatory readers: WP1, WP2, WP3, WP4, and WP5 Abstract: This deliverable presents components for evidence extraction from machine learning models that can be integrated with the certification graph. It is the result of work performed in Task 2.4. This document is a first/interim version, the final version on source evidence extractors will be reported in D2.7 Keyword List: Knowledge extraction, machine learning, deep learning, robustness, security, technical evidence Licensing information: This work is licensed under Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0 DEED https://creativecommons.org/licenses/by-sa/4.0/) Disclaimer Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union. The European Union cannot be held responsible for them. D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 3 of 19 www.emerald-he.eu Document Description Version Date Modifications Introduced Modification Reason Modified by v0.1 03.09.2024 First draft version, outline Ching-Yu Kao (FHG AISEC) v0.2 07.10.2024 Added contents to AI-SEC Ching-Yu Kao (FHG AISEC) v0.3 18.10.2024 Finalization Ching-Yu Kao (FHG AISEC) v0.4 26.10.2024 Internal review Marinella Petrocchi (CNR) v0.5 28.10.2024 Modification after QA review Ching-Yu Kao (FHG AISEC) v0.6 28.10.2024 Final Review Cristina Martínez/ Juncal Alonso (TECNALIA) v0.7 29.10.2024 Modifications after final review Ching-Yu Kao (FHG AISEC) v1.0 31.10.2024 Submitted to the European Commission Cristina Martínez/ Juncal Alonso (TECNALIA) D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 4 of 19 www.emerald-he.eu Table of contents Terms and abbreviations ............................................................................................................... 5 Executive Summary ....................................................................................................................... 6 1 Introduction ........................................................................................................................... 7 1.1 About this deliverable ................................................................................................... 7 1.2 Document structure ...................................................................................................... 7 2 ML model evidence extractors in the EMERALD architecture .............................................. 8 3 AI-SEC..................................................................................................................................... 9 3.1 Functional description ................................................................................................... 9 3.2 Technical description .................................................................................................. 10 3.2.1 Prototype architecture.................................................................................... 10 3.2.2 Technical specifications .................................................................................. 11 3.3 Delivery and usage ...................................................................................................... 12 3.3.1 Package information ....................................................................................... 12 3.3.2 Installation ...................................................................................................... 12 3.3.3 Instructions for use ......................................................................................... 13 3.3.4 Example for Running the Tool ......................................................................... 14 3.3.5 Licensing information ..................................................................................... 16 3.3.6 Download ........................................................................................................ 16 3.4 Limitations and future work ........................................................................................ 16 4 Conclusions .......................................................................................................................... 18 5 References ........................................................................................................................... 19 List of tables TABLE 1. REQUIREMENT AI-SEC.01 - EXTRACTION OF SECURITY FEATURES FROM ML MODELS ...................... 9 TABLE 2. OVERVIEW AND DESCRIPTION OF PACKAGE STRUCTURE FOR THE AI-SEC ...................................... 12 TABLE 3. SETUP FOR THE ML MODEL USING MNIST DATASET ................................................................. 15 TABLE 4. SETUP FOR THE ML MODEL USING CIFAR10 DATASET .............................................................. 15 TABLE 5. RESULTS ON MNIST AND CIFAR10 USING CLEVER SCORE, SHAPR SCORE, DATA POISONING AND LIME. .................................................................................................................................... 15 List of figures FIGURE 1. EMERALD COMPONENT OVERVIEW DIAGRAM [6]. THE RED RECTANGLE HIGHLIGHTS THE ML MODEL EVIDENCE EXTRACTION COMPONENTS, WHICH ARE DESCRIBED IN THIS DELIVERABLE. ............................. 8 FIGURE 2. AI-SEC ARCHITECTURE. TO EXTRACT FEATURES FROM ML MODELS, WE NEED TO EVALUATE THE POISONING LEVEL (ATTACK COMPONENT), ROBUSTNESS (DATA PROCESSOR1), PRIVACY LEVEL (DATA PROCESSOR1) AND EXPLANATIONS (DATA PROCESSOR2) ................................................................. 10 D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 5 of 19 www.emerald-he.eu Terms and abbreviations AI Artificial Intelligence AI-SEC AI Security Evidence Collector AMOE Assessment and Management of Organisational Evidence API Application Programming Interface BSI Bundesamt für Sicherheit in der Informationstechnik BSI C4 Artificial Intelligence Cloud Services Compliance Criteria Catalogue CertGraph Certification Graph CIFAR Canadian Institute For Advanced Research Codyze Static Code Analyzer from FHG CSA or EU CSA EU Cybersecurity Act CSP Cloud Service Provider CSV Comma-Separated Values CLEVER Cross Lipschitz Extreme Value for nEtwork Robustness DoA Description of Action EC European Commission eknows Platform for Software Analysis from SCCH GA Grant Agreement to the project KPI Key Performance Indicator MEDINA Predecessor project of EMERALD MIT Massachusetts Institute of Technology MNIST Modified National Institute of Standards and Technology database LIME Local Interpretable Model-agnostic Explanations MEDINA Predecessor project of EMERALD PNG Portable Network Graphics SW Software SHAPr SHapley Additive exPlanations TOM Technical and Organisational Measure TRL Technology Readiness Level WP Work Package D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 6 of 19 www.emerald-he.eu Executive Summary This deliverable presents the initial design, architecture, and implementation state of the machine learning (ML) model evidence extractors of WP2, we call it AI-SEC. They contribute to the key result KR1-EXTRACT of EMERALD, a framework to continuously extract knowledge from well-trained ML models and prepare suitable evidence based on them. EMERALD follows a knowledge graph-based approach to provide a unified view of the cloud service under certification at different layers of the service, ranging from the infrastructure layer (e.g., virtual resources), to the business layer (e.g., policies and procedures), to the implementation layer (e.g., source code files) and data layer (e.g., increasingly used AI models) in cloud applications. The ML model evidence extractors, developed in Task 2.4 and described in this deliverable, aim at identifying critical security-related features, such as adversarial robustness, privacy, security and explainable AI. Other related deliverables in WP2, all due at project Month 12 (October 2024), provide functional and technical details on further evidence extractors from different sources, i.e., D2.2 [1] on source code evidence extraction, D2.4 [2] on evidence extraction from policy documents in Task 2.3 and D2.8 [3] on runtime data extraction in Task 2.5. All these details contributed to D2.1 [4] on the overall information model of the certification graph in Task 2.1. The main part of this deliverable provides functional and technical descriptions of the evidence extractor AI-SEC, including its purpose and scope, the (current and planned) coverage of the EMERALD requirements, and the components’ internal architecture. These descriptions are complemented by information on delivery and usage, as well as on limitations and future work. Finally, the document concludes with a short summary. The ML model evidence extractors described in this deliverable contribute to KR1-EXTRACT by providing next-generation evidence gathering tools and techniques based on a knowledge graph approach. The presented extractors currently have the initial prototypes implemented and ready to be (to some degree) integrated with other components of the EMERALD architecture. Based on the work described in this deliverable, the ML model evidence extractors will be further extended and integrated into the EMERALD framework. This is the first iteration of the deliverable coming from Task 2.4. The second and final version of this deliverable (D2.7 [5] ) with the updated extractors will be delivered in project Month 24 (October 2025). Evidence will be prepared according to the integrated, graph-based model of semantically linked and combined evidence. D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 7 of 19 www.emerald-he.eu 1 Introduction EMERALD aims to provide a next generation set of evidence gathering tools and techniques based on a knowledge graph approach. KR1-EXTRACT supports an improved and unified toolsupported approach to continuously extract knowledge from different layers of a cloud service, e.g., infrastructure, platform, runtime information, policy documents, software, and AI models. The objective of WP2 is to establish a unified view of the cloud service under certification by extracting and enriching knowledge of the different layers of the service and providing suitable evidence for security metrics. A graph-based model, called the certification graph (CertGraph), serves as a common structure that is filled by all evidence extraction tools. 1.1 About this deliverable This deliverable focuses on the design, implementation, and initial evaluation of the tools and techniques that form the backbone of EMERALD's evidence-gathering framework. The deliverable emphasizes the role of AI-SEC in creating evidence from the ML models. 1.2 Document structure The document is structured as follows. In Section 3 we report on the design and implementation of AI-SEC ML model extractor. For the ML model extractor, functional and technical descriptions are provided, including their purpose and scope, the (current and planned) coverage of the EMERALD requirements, the components’ internal architecture, their subcomponents, and details about the programming language, libraries, etc. used. These descriptions are complemented by information on delivery and usage, including package information, installation instructions, user manual, licensing and download information, as well as limitations and future work. D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 8 of 19 www.emerald-he.eu 2 ML model evidence extractors in the EMERALD architecture This section describes how the ML model evidence extractors interact with (selected) EMERALD components on a conceptual level. Figure 1 shows the EMERALD high-level architecture as a component diagram, as described in D1.1 [6]. In EMERALD, a component is defined as “any part of the EMERALD ecosystem that has a specific functionality and can be considered a separate entity with respect to other components” (see D1.3 [7]). The components for collecting evidence about technical and organisational measures, i.e., AMOE, eknows, AI-SEC, Clouditor-Discovery, and Codyze, are represented at the bottom part of Figure 1. The ML model evidence extractor AI-SEC, which obtains technical evidence from the analysis of the ML model of cloud applications, is highlighted using a thick frame. AI-SEC is a newly developed component in EMERALD and analyses AI models for several key evidence regarding robustness against adversarial attacks, explainability, and fairness. Figure 1. EMERALD component overview diagram [6]. The red rectangle highlights the ML model evidence extraction components, which are described in this deliverable. D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 9 of 19 www.emerald-he.eu 3 AI-SEC In alignment with the requirements defined in Section 6.2, "Security & Robustness Objective," of the BSI Criteria Catalogue C5, AI-SEC is designed to meet these critical security criteria to ensure comprehensive protection and compliance. Our solution addresses four key aspects: privacy, adversary resistance, explainability, and data leakage prevention. These elements are fundamental in establishing a robust and secure system capable of withstanding potential threats while maintaining transparency and data integrity. 3.1 Functional description Overall purpose. The prototype provides a comprehensive toolkit for evaluating and improving the security of machine learning models by focusing on adversarial robustness testing, privacy vulnerability assessment, data poisoning attacks, and model interpretability. It effectively assesses vulnerabilities using techniques such as CLEVER score calculation [8], SHAPr leakage analysis [9], backdoor data poisoning [10], and LIME-based explanations [11]. These features will be collected as evidence for the certification graph. Context and scope. The toolkit assumes that users already have pre-trained models available, which can be directly utilized for evaluation. Additionally, it supports standard datasets for robustness and privacy assessments, while also allowing custom dataset imports for tailored evaluations. Motivation. The toolkit aims to streamline the security evaluation of machine learning models by integrating multiple security assessments into a unified system. Innovation. AI-SEC will focus on the following innovations: • Integration of multiple security assessments, robustness, privacy, and interpretability into a single, unified system. • Automation of calculations and result generation through command-line inputs Requirements. The relevant requirements with their respective implementation state (partially / fully / not implemented) and a brief description of how they are / will be implemented are provided in Table 1. Table 1. Requirement AI-SEC.01 - Extraction of security features from ML models Field Description Requirement ID AI-SEC.01 Short title The extractor tool includes defined criteria Description The designed AI-SEC has the features based on BSI AIC4 Status Work in Progress Priority Must Component AI-SEC Source Component, KPI Type Technical Related KR KR5_AIPOC Related KPI KPI 5.1 Validation acceptance criteria Code review: Review code and check if analysis methods work for different ML models. Progress Partially implemented – 35% Milestone MS5: Components V2 (M24) D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 16 of 19 www.emerald-he.eu Data Poisoning original image poisoned image original image poisoned image LIME-based explanations original image model prediction original image model prediction 3.3.5 Licensing information Since LIME and SHAPr use permissive MIT licenses, we choose to license this tool under the MIT License. 3.3.6 Download The currently implemented parts are stored on EMERALD's Gitlab 11 . 3.4 Limitations and future work The current tests using accessible models have shown reasonable results, demonstrating the AISEC potential. However, a significant limitation is that many cloud services do not grant direct access to their models, which poses a challenge for comprehensive evaluation. To address this, a potential solution is to use a proxy model as a substitute for the cloud-based model. While promising, further experimentation is required to evaluate the effectiveness of the proxy model in accurately reflecting the behaviour of the original cloud model. Additionally, optimization of 11 https://git.code.tecnalia.com/emerald/public/components/ai-sec D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 17 of 19 www.emerald-he.eu the tool is necessary to enhance its efficiency and ensure it can run smoothly in various environments. Future work will focus on refining the AI-SEC performance and exploring alternative methods for model evaluation in scenarios where direct access is restricted. These efforts will help improve the adaptability and robustness of AI-SEC across different use cases. Future activities will also cover the integration of the AI-SEC component in the EMERALD CaaS framework as evidence extractor tool and with the EMERALD UI. All these changes will be reported in the subsequent version of this deliverable, namely, D2.7 [5] in project month M24. D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 18 of 19 www.emerald-he.eu 4 Conclusions In this deliverable, as an initial output of Task 2.4, we presented the design, architecture, and current implementation status of the EMERALD model evidence extraction components. These components follow the holistic approach of the EMERALD framework and are aligned with the technical requirements gathered in WP1 (D1.3 [12]). The report outlines the relationship between the presented component, AI-SEC and other parts of the EMERALD framework, detailing the internal structure of the component, its subcomponents, and relevant information about its technical implementation. The component introduced in the report, AI-SEC, supports evidence extraction for machine learning models. At the current stage of the project, this component, based on preliminary work, has a working prototype that can (partially) integrate with other EMERALD components and has been tested using accessible ML models. Future work will involve testing this tool with more complex models. The subsequent and final iteration of this report (D2.7 [5]), which will provide updates on the progress of the component, is planned for project month 24. D2.6 - ML model certification – v1 Version 1.0 – Final. Date: 31.10.2024 © EMERALD Consortium Contract No. GA 101120688 Page 19 of 19 www.emerald-he.eu 5 References [1] EMERALD Consortium, “D2.2 Source Evidence Extractor–v1,” 2024. [2] EMERALD Consortium, “D2.4 AMOE – v1: Evidence extraction from policy documents that can be integrated with the certification graph,” 2024. [3] EMERALD Consortium, “D2.8 Runtime evidence extractor – v1: Evidence extraction from runtime data that can be integrated with the certification graph,” 2024. [4] EMERALD Consortium, “D2.1 Graph Ontology for Evidence Storage: Description of a uniform schema for storing and linking heterogenous data,” 2024. [5] EMERALD Consortium, “D2.7 ML model certification–v2”. [6] EMERALD Consortium, “D1.1 Data modelling and interaction mechanisms - v1,” 2024. [7] EMERALD Consortium, “EMERALD Glossary in D1.3EMERALD solution architecture - v1,” 2024. [8] T.-W. Weng, H. Zhang, P.-Y. Chen, Y. Jinfeng, D. Su, Y. Gao, C.-J. Hsieh and L. Daniel, “Evaluating the robustness of neural networks: An extreme value theory approach,” arXiv preprint, 2018. [9] V. Duddu, S. Szyller and N. Asokan, “Shapr: An efficient and versatile membership privacy risk metric for machine learning,” arXiv preprint, 2021. [10] H. Souri, L. Fowl, R. Chellappa, M. Goldblum and T. Golstein, “Sleeper agent: Scalable hidden trigger backdoors for neural networks trained from scratch,” Advances in Neural Information Processing Systems 35, 2022. [11] M. T. Ribeiro, S. Singh and C. Guestrin, “"Why should I trust you?" Explaining the predictions of any classifier,” Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1135-1144, 2016. [12] EMERALD Consortium, “D1.3 EMERALD solution architecture - v1,” 2024. [13] A. Saha, A. Subramanya and H. Pirsiavash, “Hidden trigger backdoor attacks,” in Proceedings of the AAAI conference on artificial intelligence, , vol. 34, no. 07, pp. 1195711965, 2020.