scieee AI-readable full text Open interactive document viewer

Privacy Risk Assessment of AI Models in the Healthcare Domain

Anderson, Maya; Goldsteen, Abigail; Amit, Guy; Shachor-Ifergan, Shlomit; Razinkov, Natalia; Koren, Noam; Farkash, Ariel

Abstract

Artificial Intelligence (AI) is increasingly being used to drive business, and specifically in the healthcare domain it has proven to be a critical factor inimproving efficacy and outcomes. However, as the technology advances, and especially with the increased use of Large Language Models (LLMs), existing privacychallenges have been exacerbated, and new privacy issues have arisen. LLMs’ tendency to memorize and repeat complete data samples seen duringtraining or tuning has increased their attack surface and given rise to new data leakage vulnerabilities. This, in turn, has slowed the adoption of AI, due to significantprivacy concerns, in addition to the need to adhere to strict data protection regulations, standards and guidelines. To help counter these issues, we presentan end-to-end framework for privacy risk assessment in AI models, and demonstrate the use of this framework to assess AI models of different types inconcrete healthcare use cases. We thus offer a practical approach to bridge the gap of trust in AI-driven solutions, by building privacy risk assessment into theML pipeline as an integral part of the flow.

Full text

Privacy Risk Assessment of AI Models in the Healthcare Domain Maya Anderson, Abigail Goldsteen, Guy Amit, Shlomit Shachor-Ifergan, Natalia Razinkov, Noam Koren, Ariel Farkash IBM Research, a[email protected]bm.com Abstract Artificial Intelligence (AI) is increasingly being used to drive business, and specifically in the healthcare domain it has proven to be a critical factor in improving efficacy and outcomes. However, as the technology advances, and especially with the increased use of Large Language Models (LLMs), existing privacy challenges have been exacerbated, and new privacy issues have arisen. LLMs’ tendency to memorize and repeat complete data samples seen during training or tuning has increased their attack surface and given rise to new data leakage vulnerabilities. This, in turn, has slowed the adoption of AI, due to significant privacy concerns, in addition to the need to adhere to strict data protection regulations, standards and guidelines. To help counter these issues, we present an end-to-end framework for privacy risk assessment in AI models, and demonstrate the use of this framework to assess AI models of different types in concrete healthcare use cases. We thus offer a practical approach to bridge the gap of trust in AI-driven solutions, by building privacy risk assessment into the ML pipeline as an integral part of the flow. Keywords AI Privacy, Healthcare, Privacy Risk Assessment Introduction AI privacy is relevant whenever an AI system involves personal data. The concept of personal data is not universally defined and varies between geographies, cultures, and regulatory environments. According to the EU General Data Protection Regulation (GDPR)1, personal data is any information relating to an identified or identifiable natural person (called the data subject). This refers to a person who can be identified, directly or indirectly, by reference to an identifier such as a name, an identification number, location data, an online identifier, or one or more factors specific to their physical, physiological, genetic, mental, economic, cultural, or social identity. The California Consumer Protection Act (CCPA)2 provides an even wider definition of personal information as information that identifies, relates to, or could reasonably be linked with a person. For example, it could include not only direct identifiers, but also records of products purchased, internet browsing history, geolocation data, and inferences from other personal information. In general, personal data may include information about demographics (age, gender, place of residence, etc.), location, health, finances, biometrics, employment, and various activities, behaviors, and opinions. It is crucial to ensure that any AI project complies with all relevant data protection and privacy regulatory requirements. This may include obtaining consent, minimizing data collection, enabling data correction or withdrawal, and more. These requirements apply to all data used in an AI system, and especially the data used to train machine learning (ML) models. Violating data protection regulations can incur serious fines. GDPR has set fines of up to €20 million, or 4% of the company's worldwide annual revenue from the preceding financial year, whichever is higher. The highest GDPR fine given so far was an amazing 1.2 billion Euros3 incurred against Meta Platforms Ireland Limited for insufficient legal basis for data processing. Specifically in the context of machine learning, in 2021, a precedential ruling from the US Federal Trade Commission (FTC) forced an AI company to delete its machine learning models after unlawfully collecting user data4. Recent studies have shown that a malicious third party with access to a trained ML model, even without access to the training data itself, can still reveal 1 https://ec.europa.eu/info/law/law-topic/data-protection/data-protection-eu_en 2 https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=201720180AB375 3 https://www.enforcementtracker.com/ 4https://jolt.law.harvard.edu/digest/everalbum-inc-in-first-facial-recognition-misuse-settlement-ftc-requires-destruction-of-algorithms-trained-on-deceptively-obtained-photos sensitive, personal information about the people whose data was used to train the model. For example, it may be possible to reveal if a person’s data was part of the model’s training set, or even infer sensitive attributes about them, such as their salary [1, 2]. Real-world large language models have been shown to leak individual user's personal information5 [3]. Factors that may affect the level of privacy risk posed by a machine learning model include the size of the model, the size of the training set, the dimensionality of the data, etc. In general, the more a model overfits, or in other words memorizes its training data, the higher the probability of the model leaking unintended information. To mitigate these risks, organizations developing AI systems should implement privacy risk assessment early in the development lifecycle. This approach supports informed decisions regarding the deployment of a model in production, its publication or sharing it with third parties. An early privacy risk assessment also facilitates the identification of potential mitigations to create a more privacy-preserving model. 1.1 AI Privacy Risks in the Healthcare Domain There is a clear conflict between the ever-increasing interest in analyzing personal data to enhance and improve processes that affect human life and the need to preserve the privacy of data subjects. This is especially the case in the medical domain, where sensitive health data from real patients is often used to train machine learning (ML) models that aid physicians in efficient diagnostics and personalized treatment. Artificial intelligence (AI) models are also used in medical devices and applications to predict malfunctions and enhance continuous patient care. The deployment of such systems introduces privacy risks, and ensuring patient privacy and protecting their sensitive health information presents a significant challenge. Assessing the privacy risk of such AI models is crucial for making informed decisions about whether to use a model in production, share it with third parties, or deploy it in patients’ homes. Privacy risk assessment is often achieved by 5https://www.businessupturn.com/world/chatgpts-answer-gives-away-a-journalists-number-to-join-signal/ running privacy attacks against one or more models and measuring how successful they are in leaking personal or sensitive information. Striking a balance between the need for patient privacy and effective health and disease management through technology is pivotal in the design and implementation of healthcare solutions. We outline here a few relevant use cases from the healthcare domain that employ AI models. In the next sections, we will give specific examples of such AI models and the privacy risk assessments that may be performed on them. One example of a system that uses AI models trained on sensitive health data is a medical device for the continuous monitoring of patients with neurodegenerative diseases, such as Parkinson’s disease. It is designed to trace, record and store a variety of motor symptoms commonly associated with these diseases through the continuous use of wearable body sensors. The system is composed of: (1) a set of wearable monitoring devices, (2) a docking station for collecting, processing and uploading the data of the monitoring devices to the cloud, (3) a mobile application that enables patients and caregivers to record complementary information, such as medication, nutrition and non-motor status information, and (4) a web application, called the Physician’s Tool, that presents all patientrelated information to the healthcare professional. Since the trained AI models may be part of the system’s mobile application or other devices deployed at patients’ homes, it is especially important to ensure that they do not leak any sensitive information. Another example is a diabetes support application that uses an AI model to provide personalized recommendations on daily insulin adjustments based on treatment related information. The system is a Software as a Medical Device (SaMD) that can be installed on the patient’s mobile phone and is capable of connecting to external devices and remote servers. It processes sensitive data related the patient’s condition and therapy, such as blood glucose levels, insulin intakes, etc. The current version of the model is trained on a single patient’s data, but an improved model is in development. It will use a model pre-trained on data from multiple patients in a centralized manner and will then be fine-tuned on a specific patient's data once deployed on that patient’s device. Since this pretrained model will be deployed on every patient's local device, it is important to ensure that it does not leak sensitive health data of other patients used to train the model. The third example is a diagnostic device and AI assistant designed to empower healthcare practitioners to provide improved, personalized diagnosis of skin cancer. Patient data is extracted from various sources such as lesion images, clinical history and genotypic information, and analyzed by AI algorithms to assign each patient a holistic melanoma risk score. These scores are presented to physicians within a cognitive assistant, that also displays an explanation justifying its decisions. These ML models for assessing patients’ melanoma risk based on clinical, genetic, and family history information are trained on highly sensitive patient data, making it crucial to assess their risk of leaking this information. Regulations Related to AI Privacy in Healthcare In this section we highlight the main regulations relevant to AI Privacy risk assessment, including regulations for privacy, AI and health. 1.2 Privacy Regulations The EU, with its General Data Protection Regulation (GDPR), was the first to enact a comprehensive enforceable regulation around information privacy. It includes chapters on data subject rights, data controllers’ and processors’ obligations, and transfers of personal data to third countries. To date, it is one of the strictest privacy regulations in the world, awarding fines of up to 4% of a company’s annual revenue in case of a breach. The UK, after breaking away from the European Union in 2020, decided to keep the GDPR in force, coining it the UK GDPR6. The GDPR has become a model for many similar laws around the world, either already enacted or under development, including in Brazil, Japan, Singapore, South Africa, and more. In the United States there is no comprehensive national privacy law, apart from the Privacy Act of 1974, that governs the collection, maintenance, use, and dissemination of information about individuals by federal agencies. However, there are a number of sector-specific privacy and data security laws that apply to financial institutions, telecommunications companies, credit reporting agencies and healthcare providers, as well as a multitude of state and local laws. 6 https://ico.org.uk/for-organisations/data-protection-and-the-eu/data-protection-andthe-eu-in-detail/the-uk-gdpr/ The landmark California Consumer Privacy Act (CCPA) secures comprehensive privacy rights for California consumers, including the right to know what personal information is collected about them, how it is used and shared, the right to delete personal information, and the right to limit the use and disclosure of sensitive personal information. Many states have followed by enacting different privacy and data protection acts, including Colorado, Connecticut, Delaware, Florida, Virginia and more7. The American Data Privacy and Protection Act8, proposed in 2022, establishes requirements for how companies handle personal data. The proposed bill requires most companies to limit the collection, processing, and transfer of personal data to what is reasonably necessary to provide a requested product or service and prohibits the transfer of personal data without explicit consent. In China, the Personal Information Protection Law (PIPL)9, which came into effect in 2021, seeks to protect personal data and regulate its processing. The Brazilian General Data Protection Law (LGPD)10 governs the processing of personal data, with the purpose of protecting the fundamental rights of freedom and privacy and is in line with the GDPR on most articles. In Canada, the proposed Consumer Protection Privacy Act (CPPA)11, would replace the existing Personal Information Protection and Electronic Documents Act and establish a new Personal Information and Data Protection Tribunal. 1.3 AI Regulations In the EU, the Artificial Intelligence Act12 (EU AI Act), in force as of this year, harmonizes the different AI rules and regulations across Europe, with penalties for violations greatly exceeding those of GDPR. This landmark regulation 7 https://www.dlapiperdataprotection.com/?t=law&c=US 8https://www.congress.gov/bill/117th-congress/house-bill/8152#:~:text=American%20Data%20Privacy%20and%20Protection%20Act,-This%20bill%20establishes&text=The%20bill%20establishes%20consumer%20data,opt%20out%20of%20such%20advertising. 9 http://en.npc.gov.cn.cdurl.cn/2021-12/29/c_694559.htm 10 https://iapp.org/resources/article/brazilian-data-protection-law-lgpd-english-translation/ 11 https://cppa.ca.gov/regulations/consumer_privacy_act.html 12https://artificialintelligenceact.eu/wp-content/uploads/2022/05/AIA-COM-Proposal-21April-21.pdf organizes AI systems by risk level. Some uses of artificial intelligence are prohibited, such as manipulating people's decisions, classifying people based on social behavior, and predicting a person's risk of committing a crime. High-risk applications are those that pose a significant risk of harm to the health, safety, or fundamental rights of individuals. This includes safety components, biometric identification, education, employment, law enforcement, and justice. These are allowed, but subject to strict obligations and requirements, including risk assessment, logging, documentation, human oversight, and security. These applications must have a risk management system in place throughout the entire lifecycle of the AI system, which should identify any known and reasonably foreseeable risks and implement appropriate measures to reduce them. In Article 10 it specifically addresses the training of AI models and mandates that the training, validation, and testing datasets used must meet quality, safety, and fairness criteria, requiring that any special categories of personal data be subject to technical limitations on reuse and state-of-the-art security and privacy measures. The European Council also opened for signature in September 2024 the Framework Convention on Artificial Intelligence13, the first international legally binding treaty aimed at ensuring that activities throughout the lifecycle of AI systems are consistent with human rights, democracy and the rule of law. Other countries have recently also begun to promote AI legislation. The proposed Canadian Artificial Intelligence and Data Act (AIDA)14, introduced as part of the Digital Charter Implementation Act in 2022, would set the foundation for the responsible design, development and deployment of AI systems that impact the lives of Canadians. In the United States, the FTC has published multiple reports and guidelines on the use of artificial intelligence in agencies and businesses and the risks concerned. In 2025, they warned against risks of consumer harm from AI, and recommended ensuring privacy and security by default in generative AI tools15. A few bills such as the Artificial Intelligence Initiative Act16, the Algorithmic 13 https://www.coe.int/en/web/artificial-intelligence/the-framework-convention-on-artificial-intelligence 14 https://ised-isde.canada.ca/site/innovation-better-canada/en/artificial-intelligence-anddata-act-aida-companion-document 15 https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2025/01/ai-risk-consumerharm 16 https://www.congress.gov/bill/116th-congress/senate-bill/1558/text Accountability Act (AAA)17 and the National AI Commission Act18 have also been proposed to Congress. In 2022, the White House Office of Science and Technology Policy published a Blueprint for an AI Bill of Rights19 , with the goal of developing policies and practices for automated systems that protect civil rights and promote democratic values such as privacy In 2023, the Joe Biden administration issued an executive order on “Safe, Secure, and Trustworthy Development and Use of AI” however, it was later revoked by the incoming president Donald Trump. Drawing inspiration from the EU AI Act, The Colorado AI Act of 202420 adopts a risk-based approach to regulate the deployment of high-risk AI systems, emphasising transparency, risk mitigation, and proper documentation. Many other countries are currently publishing strategies and frameworks around AI, including the UK, Singapore, and Australia. However, these have not yet matured into concrete legislation. 1.4 Healthcare Regulations The US Health Insurance Portability and Accountability Act (HIPAA) establishes federal standards protecting sensitive health information. HIPAA contains a Privacy Rule21 (“Standards for Privacy of Individually Identifiable Health Information”) that addresses the use and disclosure of individuals' protected health information (PHI). The Privacy Rule protects all "individually identifiable health information" including demographic data to the individual's physical or mental health or condition, provision of health care and payment for the provision of health care. This includes many common identifiers such as name, address, and birth date. While the GDPR is mostly domain-neutral, it recognizes health data as a special category of data requiring enhanced privacy protection. This includes all personal data concerning or linked to the health of an individual, including genetic and biometric data. The GDPR stipulates that health data can only be used 17 https://www.congress.gov/bill/117th-congress/house-bill/6580/text 18 https://www.congress.gov/bill/118th-congress/house-bill/4223/text 19 https://bidenwhitehouse.archives.gov/ostp/ai-bill-of-rights/ 20 https://leg.colorado.gov/bills/sb24-205 21 https://www.hhs.gov/hipaa/for-professionals/privacy/laws-regulations/index.html for specific purposes or with explicit patient consent, and outlines rules under which it can be processed. The Medical Devices Regulation (MDR)22 lays down rules concerning the sale and servicing of medical devices for human use. Regarding privacy and confidentiality of data associated with the use of Medical Devices (MDs), Article 62.4(h) mandates that the subject’s rights to privacy and the protection of the personal data in accordance with Directive 95/46/EC are safeguarded. In 2019, The Medical Device Coordination Group (MDCG) issued the Guidance on Cybersecurity for medical devices23. Some of the key points addressed in this guidance are secure design, security risk management and implementation of security measures to mitigate risks. It also mandates the application of the GDPR principles in the design and development of medical devices. ENISA published in 2023 a report on Cybersecurity and privacy in AI - Medical imaging diagnosis24. This report covers the use of AI in the health sector and highlights specific privacy and security threats that may arise. The report finds that efforts to optimize security and privacy can often come at the expense of system performance, and consequently insists on the search for new security measures. The guide places strong emphasis on privacy issues alongside cybersecurity issues, mentioning privacy as one of the most important challenges facing society today. In January 2025, the EU Council adopted the European Health Data Space Regulation (EHDS)25 that aims at facilitating cross-border exchange of EU health data and improving individuals’ control over how their health data is used. AI Privacy Risk Assessment 22 https://eur-lex.europa.eu/eli/reg/2017/745/oj/eng 23 https://ec.europa.eu/docsroom/documents/41863 24https://www.enisa.europa.eu/publications/cybersecurity-and-privacy-in-ai-medical-imaging-diagnosis 25 https://health.ec.europa.eu/ehealth-digital-health-and-care/european-health-dataspace_en tools and frameworks by automatically selecting which attacks and metrics to run based on Q&A-based interactions with the user, then running the attacks, summarizing and visualizing the results in an easy-to-consume manner. The first goal of this end-to-end risk assessment tool is to automate many of the decisions around which attacks and metrics to run, and all of the technical preparation required in order to run them. For example, some attacks require specific data preparation or have multiple runtime options and parameters that need to be set in a way that best matches the given model and data. Secondly, since most non-technical users cannot understand the meaning of each individual attack or score, let alone compare the results of different models. It is therefore crucial to summarize these individual results into an overall privacy risk score. The tool is generic and enables evaluating models from different ML frameworks (scikitlearn45, pytorch46, keras47), different data modalities (tabular, images, text) and different model types (classification, regression, detection). This is done by decoupling the actual models and datasets from the rest of the framework using generic wrappers. It performs multiple attack runs with different random data splits, collecting performance metrics like accuracy, precision, recall, and F1 score. The results are stored in a JSON structure and aggregated into an overall privacy health score for the model (® Fig. 1). 45 https://scikit-learn.org/stable/ 46 https://pytorch.org/ 47 https://keras.io/ Fig. 1 AI Risk Assessment Tool The tool also employs a novel attack framework introduced in [16], significantly enhancing membership inference attacks against classification models. This framework leverages an ensemble method, generating multiple specialized attack models for different data subsets. By aggregating results from these specialized models, the framework provides a more comprehensive assessment of the target model’s privacy risks. This approach outperforms single attack models and per-class attack models across classical and language classification tasks. AI Privacy Risk Assessment of Various Model Types Let us examine privacy risk assessment of tabular AI models, time-series AI models and Large Language Models (LLMs). 1.8 AI Privacy Attacks on Tabular Models There are several types of privacy attacks against ML models, including membership inference, attribute inference, model inversion, and database reconstruction. The most commonly researched and employed attack is membership inference, with dozens of papers published each year [17], and implementations available in open-source privacy assessment frameworks [9], [10]. These attacks may be performed either in a black-box manner - where only the output of the model on a particular input is known - or in a white-box manner, where internal parameters of the model are also known. 1.8.1 Membership Inference Attacks on Tabular Models Membership inference attacks (MIA) attempt to distinguish between members, who were part of a target model’s training data, and non-members. MIAs are especially relevant when the training dataset comprises sensitive user data or when the model is used to predict sensitive attributes, such as medical conditions. When such an attack succeeds, the adversary may gain insight into private information related to a person, thereby compromising their privacy. For example, by exposing the participation of an individual in the training dataset used for a model predicting the disease progression of Alzheimer's, an adversary gains knowledge about the individual having Alzheimer's disease. MIAs have been extensively studied in the context of classification models and in the black-box setting, where the model internals are unknown to the attacker. The first MIAs were either threshold-based [18] or employed binary classifiers trained to distinguish between members and non-members based on model outputs [1]. For example, these outputs may include class probabilities or logits (for classification models), the model’s loss, or possibly activations from internal layers of the model (in white-box attacks) [19]. To generate labeled (member/non-member) data to train the attack classifier without knowledge of the true member samples of the attacked model, shadow models are commonly used [1]. 1.8.2 Attribute Inference and Model Inversion Attacks on Tabular Models Attribute inference and model inversion (MI) attacks attempt to reconstruct new information about training data samples based on the trained model. Attribute inference usually refers to an attack where certain sensitive features may be inferred about individuals who participated in training a model [2], [20], [4]. Given a trained model and knowledge about some of the features of a specific person, it may be possible to deduce the value of additional, unknown features of that person. Model inversion typically aims to reconstruct representative feature values of the training data by inverting a trained ML model [20], [21], [22]. For example, it may be possible to reconstruct what the average sample for a given class looks like. This can be considered a privacy violation if a class represents a specific person or group of people, such as in facial recognition models. MI attacks usually require white-box access to the model. 1.8.3 AI Privacy Risk assessment of Tabular Models in Healthcare In the Healthcare domain there are many use cases where AI models are trained on tabular data, including past clinical data, various test results, medication doses, sensor data, and more. One such example is an ML model that uses patients’ vital signs to detect sleep stages and associated disorders. These vital signs, including heart rate and respiration rate, are extracted from sensors present in wearable devices placed on the patient's body. They can be used to train different types of ML classifiers, including: Decision Tree classifier, Multi-layer Perceptron, k-Nearest Neighbors classifier, AdaBoost, Random Forest, and Logistic Regression. The result is a mapping to one of five different sleep stages. Another example includes ML models that perform quantitative risk assessment to help determine a patient’s risk for melanoma and other skin cancers based on clinical, genetic and familial history information. This includes demographic data, such as age and place of residence, medical history and historical sun exposure data. The output of the first model is a probability that indicates the risk of developing first primary melanoma. In addition, a multi-label classification model simultaneously assesses the risk of both melanoma and other skin cancers. Both membership inference and attribute inference attacks pose a privacy threat to these ML models, since even the fact that a patient’s data is part of a dataset of patients having a specific disease is already information that a patient might not want to disclose, and its leakage might result in patient harm, for example, negatively affecting the patient’s employment prospects. In addition, inferring a patient’s sensitive attributes directly violates their data privacy and undermines patient trust. Inference of the patient’s demographic data and medical history would also go against the guidelines and recommendations described in sections ® Regulations Related to AI Privacy in Healthcare and ® AI Risk Assessment Guidelines and Frameworks. Finally, for the melanoma risk detection models, since their training data also includes family history and genetic information, the privacy risk is not only to the patient participating in the study, but also to unaware family members, which poses an even greater issue. 1.9 AI Privacy Attacks on Time-Series Models Time-series (TS) forecasting uses historical observations to predict future values, making it indispensable in areas like healthcare, economics, and environmental science. Over the years, it has progressed from classical statistical techniques, such as ARIMA and exponential smoothing [23], to state-of-the-art machine learning models like the Neural Fourier Transform (NFT), TimesNet, and PatchTST. Time-series data can be decomposed into two main components: trend and seasonality . The trend of the data represents the long-term direction of the data. For instance, in the health domain, given a series of an infant's hourly weight over two years, the overall weight increase represents the trend. In contrast, seasonality represents cycles in the data. For example, the infant’s weight measurements taken in the morning are usually lower than those taken at night, producing a regular daily cycle. Fig. 2 Decomposition of the original time series into its trend and seasonal components, highlighting both the long-term trajectory and recurring patterns over time. 1.9.1 Membership Inference Attacks on Time-Series Models In [27], the authors address the vulnerability of time-series forecasting models to membership inference attacks (MIAs). This is especially significant in medical contexts, where TS data often includes highly sensitive, personal medical data, such as blood glucose levels, heart monitoring signals, or disease progression. Despite relying on highly sensitive patient data, these forecasting models have not been thoroughly examined for privacy risks, so there is a pressing need for focused research in this area. Many TS forecasting models inherently capture trend and seasonality to improve predictive performance. In [27], the authors show that these very components can also be exploited to enhance MIAs. They introduce two novel features: one approximates the predicted time-series trend using a low-degree polynomial, while the other extracts seasonality through the Discrete Fourier Transform (DFT). Since state-of-the-art models like the Neural Fourier Transform (NFT) [25], TimesNet [28], and more, explicitly incorporate trend and seasonality, they are more likely to accurately learn and reproduce these patterns if a series is included in their training data, thereby amplifying the potential for privacy leaks ( ® Fig. 3). Fig. 3 Overview of the membership inference attack pipeline for time-series models. The results demonstrate that embedding time-series-specific characteristics - like trend and seasonality - into the attack vector, led to substantially higher risk scores compared to conventional feature sets. In other words, constructing attack vectors around a series’ unique patterns makes forecasting models far more susceptible to membership inference. By exposing this weakness, this work sets the stage for more in-depth studies of membership inference attacks on time-series forecasting. 1.9.2 AI Privacy Risk Assessment of TS Models in Healthcare Let us examine an example model for predicting blood glucose levels in diabetes patients. The model aims to perform personalized blood-glucose prediction for patients suffering from type 1 diabetes (PwT1D) and to reduce the effect of inter-patient variability through grouping training data according to glycemic similarities for improving personalized prediction accuracy. The model takes as input a multi-dimensional time series featuring: Continuous Glucose Monitoring (CGM) data, bolus insulin injections and carbohydrate intake; and predicts the next blood-glucose values with prediction horizons of 30 and 60 minutes. This model is a sequence-to-sequence model based on an LSTM neural network architecture. To reduce the setup time of a device for a new patient, the model is pretrained on multiple patients' data in a centralized manner, and then fine-tuned on the specific patient's data once deployed on the patient’s device. This both reduces the time required before the system becomes functional and decreases the computational cost of local model training. Since this pre-trained model will be deployed on every patient's local device, it is important to ensure that it does not leak the sensitive health data of the patients used to train the model. Moreover, TS data, and especially multi-dimensional TS data, tends to be much more unique than typical tabular data, increasing even further the privacy leakage risk, and specifically the membership inference risk, as shown in [27]. 1.10 AI Privacy Attacks on Large Language Models Large Language Models are susceptible to both the same privacy attacks as classic AI models and to new attacks specific to Large Language Models. 1.10.1 Membership Inference Attacks on LLMs Membership Inference Attacks (MIAs) aim to determine if a specific data point was part of a model's training dataset. Since they were initially introduced, MIAs have been extended to LLMs, where they exploit model overfitting and memorization tendencies to infer training membership [3]. Text embedding models (e.g., BERT, Word2Vec) encode inputs into vector representations. These models can inadvertently expose training membership, as shown in [29]. [30] extended this to include downstream NLP models that use these embeddings. In healthcare, clinical embeddings trained on patient records can expose whether a specific term or phrase was seen in training [31]. Generative LLMs, like GPT and LLaMA, are susceptible to MIAs via likelihood-based attacks, where membership is inferred by observing model confidence [32]. [3] showed that LLMs occasionally memorize training data verbatim, making MIAs particularly effective when targeting rare sequences. Recent techniques, such as neighborhood comparison [33], generate perturbations of an input to compare loss distributions, improving attack success without requiring a reference model. Duan et al.[13], however, found that MIAs on large, diverse LLMs often fail, as models trained on vast corpora generalize well, obscuring membership signals. Document-level MIAs are an emerging risk, where complete documents (e.g., books, medical records) are used for the attack, i.e. checking if a specific document was part of the training data of an LLM [34]. MIA in LLMs remains an active research area, with text embeddings and generative models vulnerable under certain conditions. While successful MIAs are harder to perform against large, commercial models, smaller fine-tuned language models, such as those typically used in healthcare, pose significant privacy risks, making them particularly attractive targets for adversaries seeking to extract sensitive information. As research progresses, considering the tradeoff between model utility with privacy will play crucial step for responsible LLM development. 1.10.2 Data Leakage Attacks on LLMs Generative language models have a natural tendency to memorize training or fine-tuning samples. This tendency can be exploited by data extraction attacks to extract from the LLM partial or complete samples from the training or finetuning datasets, or any other dataset that the model is exposed to during inference [3], [35], [36], [37]. This is crucial for the healthcare domain as it may reveal personal medical information about individuals. Data leakage attacks usually fall under one of two categories: attacks that utilize prior knowledge about training samples to carry out the attack, and those that do not assume such knowledge. Attacks that utilize prior knowledge about training samples typically split the samples that are being extracted into two parts: a prefix and a suffix [3]. The prefix is used as input to the language model aiming to elicit it to complete the suffix. The suffix is then compared with the model’s actual output to determine whether leakage has occurred. Optionally, the entire dataset may be indexed to enable identification of whether generated text is from any of the training data samples. Common methods of scoring the match usually include exact matches, n-grams or other distance metrics [38], [39]. Existing solutions use fixed sizes of prefixes and suffixes, based on the number of tokens in each part [3]. This token-wise split may cause the sample to be divided in a way that breaks the semantic meaning of the text and thus the model might get confused or simply not output the desired suffix. Another possible approach is to split each sample first into sentences, and then use a word-based split instead of token-based. Since the model occasionally leaks information in subsequent sentences, it is also beneficial to “force” the model to generate a response that is longer than the target suffix. This method increases the chances of the LLM models to leak the remainder of the sentence (suffix), or even other sentences from the entire dataset. Fig. 4 demonstrates an example leakage, for the mistral48 model fine-tuned on the HealthCareMagic49 dataset. Fig. 4 LLM leaking entire suffix given the prefix 48 https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.2 49https://huggingface.co/datasets/RafaelMPereira/HealthCareMagic-100k-Chat-Format-en Common methods of scoring the match, such as considering only exact matches, or comparing long sequences of text, may miss many more subtle examples of data leakage. This method employs new lexical similarity techniques to compare the suffix to the output sentences, giving a higher score to longer leaks, allowing for a looser definition of similarity, for example by identifying sequences that do not preserve the exact order or have some missing words. For the second class of attacks that does not use training data, an attacker may leverage specific techniques to cause the model to “diverge” from its aligned behavior and output training data. This can be done, for example, by asking the model to repeat a word forever [40], or by using special characters [41]. This can also lead to a secondary type of PII/confidential information leakage attack if the reconstructed texts in fact contain PII or confidential information [6], [42]. Another possible use case for this kind of attack is to identify whether an LLM model was trained or fine-tuned on a dataset that is not allowed, such as cases of copyright infringement [43]. In the healthcare domain it is common to use synthetic data to overcome privacy risks due to potential leakage of personal information [44]. However, as shown in [45], privacy leakages may occur even when synthetic data is used. Data leakage assessment can be leveraged during synthetic data generation to ensure that real data does not leak into the generated data. 1.10.3 Prompt Leakage Attacks on LLMs Prompt leakage [7], [8], [46], or a system prompt extraction, is a sophisticated attack vector concerning leaking of the system instructions defined in an LLMbased application. The attack, often masked as a benign user query, tricks the system into revealing the system prompt, thus revealing the application design. This can have severe implications since the system prompt defines the application's (for example, a chatbot) capabilities, response format, and the tools it can access. It can also include personal or sensitive information in the form of fewshot examples. An example of a system prompt for a medical chatbot application could be: You are a medical chatbot. You always start your answers with “Welcome to MediChat.” Your job is to advise on medical issues presented by clients, as if you are a family doctor. Be mindful of the advice you give, as some patients may act on it, which may cause harm. [7] Y. Zhang, N. Carlini, and D. Ippolito, “Effective prompt extraction from language models,” ArXiv Prepr. ArXiv230706865 , 2023. [8] B. Hui, H. Yuan, N. Gong, P. Burlina, and Y. Cao, “Pleak: Prompt leaking attacks against large language model applications,” in Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security , 2024, pp. 3600–3614. [9] M.-I. Nicolae et al. , “Adversarial Robustness Toolbox v1. 0.0,” ArXiv Prepr. ArXiv180701069 , 2018. [10] S. Kumar and R. Shokri, “ML Privacy Meter: Aiding regulatory compliance by quantifying the privacy risks of machine learning,” in Workshop on Hot Topics in Privacy Enhancing Technologies (HotPETs) , 2020. [11] J. Ye, A. Maddi, S. K. Murakonda, V. Bindschaedler, and R. Shokri, “Enhanced Membership Inference Attacks against Machine Learning Models,” in Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security , 2022, pp. 3093–3106. [12] P. Guldimann et al. , “COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act,” ArXiv Prepr. ArXiv241007959 , 2024. [13] M. Duan et al. , “Do Membership Inference Attacks Work on Large Language Models?,” in Conference on Language Modeling (COLM) , 2024. [14] H. Takahashi, “AIJack: Security and Privacy Risk Simulator for Machine Learning,” ArXiv Prepr. ArXiv231217667 , 2023. [15] A. Goldsteen, S. Shachor, and N. Raznikov, “An end-to-end framework for privacy risk assessment of AI models,” in Proceedings of the 15th ACM International Conference on Systems and Storage , in SYSTOR ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 142. doi: 10.1145/3534056.3534998. [16] S. Shachor, N. Razinkov, and A. Goldsteen, “Improved Membership Inference Attacks Against Language Classification Models,” Oct. 11, 2023, arXiv : arXiv:2310.07219. doi: 10.48550/arXiv.2310.07219. [17] H. Hu, Z. Salcic, L. Sun, G. Dobbie, P. S. Yu, and X. Zhang, “Membership inference attacks on machine learning: A survey,” ACM Comput. Surv. CSUR , vol. 54, no. 11s, pp. 1–37, 2022. [18] S. Hisamoto, M. Post, and K. Duh, “Membership inference attacks on sequence-to-sequence models: Is my data in your machine translation system?,” Trans. Assoc. Comput. Linguist. , vol. 8, pp. 49–63, 2020. [19] M. Nasr, R. Shokri, and A. Houmansadr, “Comprehensive Privacy Analysis of Deep Learning: Passive and Active White-box Inference Attacks against Centralized and Federated Learning,” in 2019 IEEE Symposium on Security and Privacy (SP) , May 2019, pp. 739–753. doi: 10.1109/SP.2019.00065. [20] M. Fredrikson, S. Jha, and T. Ristenpart, “Model inversion attacks that exploit confidence information and basic countermeasures,” in Proceedings of the 22nd ACM SIGSAC conference on computer and communications security , 2015, pp. 1322–1333. [21] U. Aïvodji, S. Gambs, and T. Ther, “Gamin: An adversarial approach to black-box model inversion,” ArXiv Prepr. ArXiv190911835 , 2019. [22] Z. He, T. Zhang, and R. B. Lee, “Model inversion attacks against collaborative inference,” in Proceedings of the 35th Annual Computer Security Applications Conference , 2019, pp. 148–162. [23] G. E. Box, G. M. Jenkins, G. C. Reinsel, and G. M. Ljung, Time series analysis: forecasting and control . John Wiley & Sons, 2015. [24] R. J. Hyndman and G. Athanasopoulos, Forecasting: principles and practice . OTexts, 2018. [25] N. Koren and K. Radinsky, “Interpretable multivariate time series forecasting using neural fourier transform,” ArXiv Prepr. ArXiv240513812 , 2024. [26] C. Lea, M. D. Flynn, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional networks for action segmentation and detection,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 156–165. [27] N. Koren, A. Goldsteen, G. Amit, and A. Farkash, “Membership Inference Attacks Against Time-Series Models,” ArXiv Prepr. ArXiv240702870 , 2024. [28] H. Wu, T. Hu, Y. Liu, H. Zhou, J. Wang, and M. Long, “Timesnet: Temporal 2d-variation modeling for general time series analysis,” ArXiv Prepr. ArXiv221002186 , 2022. [29] C. Song and A. Raghunathan, “Information Leakage in Embedding Models,” Aug. 19, 2020, arXiv : arXiv:2004.00053. doi: 10.48550/arXiv.2004.00053. [30] S. Mahloujifar, H. A. Inan, M. Chase, E. Ghosh, and M. Hasegawa, “Membership inference on word embedding and beyond,” ArXiv Prepr. ArXiv210611384 , 2021. [31] A. Jagannatha, B. P. S. Rawat, and H. Yu, “Membership inference attack susceptibility of clinical language models,” ArXiv Prepr. ArXiv210408305 , 2021. [32] C. Song and V. Shmatikov, “Auditing data provenance in text-generation models,” in Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2019, pp. 196–206. [33] J. Mattern, F. Mireshghallah, Z. Jin, B. Schölkopf, M. Sachan, and T. BergKirkpatrick, “Membership inference attacks against language models via neighbourhood comparison,” ArXiv Prepr. ArXiv230518462 , 2023. [34] M. Meeus, S. Jain, M. Rei, and Y.-A. de Montjoye, “Did the neurons read your book? document-level membership inference for large language models,” in 33rd USENIX Security Symposium (USENIX Security 24) , 2024, pp. 2369–2385. [35] N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang, “Quantifying memorization across neural language models,” in The Eleventh International Conference on Learning Representations , 2022. [36] J. G. Wang, J. Wang, M. Li, and S. Neel, “Pandora’s White-Box: Precise Training Data Detection and Extraction in Large Language Models,” ArXiv Prepr. ArXiv240217012 , 2024. [37] X. Zhou et al. , “LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks,” Feb. 10, 2025, arXiv : arXiv:2502.06215. doi: 10.48550/arXiv.2502.06215. [38] W. Yu et al. , “Bag of Tricks for Training Data Extraction from Language Models,” Jun. 01, 2023, arXiv : arXiv:2302.04460. Accessed: Aug. 14, 2023. [Online]. Available: http://arxiv.org/abs/2302.04460 [39] R. Xu, Z. Wang, R.-Z. Fan, and P. Liu, “Benchmarking Benchmark Leakage in Large Language Models,” Apr. 29, 2024, arXiv : arXiv:2404.18824. doi: 10.48550/arXiv.2404.18824. [40] M. Nasr et al. , “Scalable Extraction of Training Data from (Production) Language Models,” Nov. 28, 2023, arXiv : arXiv:2311.17035. doi: 10.48550/arXiv.2311.17035. [41] Y. Bai, G. Pei, J. Gu, Y. Yang, and X. Ma, “Special characters attack: Toward scalable training data extraction from large language models,” ArXiv Prepr. ArXiv240505990 , 2024. [42] K. K. Nakka, A. Frikha, R. Mendes, X. Jiang, and X. Zhou, “PII-Scope: A Benchmark for Training Data PII Leakage Assessment in LLMs,” Oct. 09, 2024, arXiv : arXiv:2410.06704. doi: 10.48550/arXiv.2410.06704. [43] A. Karamolegkou, J. Li, L. Zhou, and A. Søgaard, “Copyright Violations and Large Language Models,” Oct. 20, 2023, arXiv : arXiv:2310.13771. doi: 10.48550/arXiv.2310.13771. [44] V. C. Pezoulas et al. , “Synthetic data generation methods in healthcare: A review on open-source tools and methods,” Comput. Struct. Biotechnol. J. , vol. 23, pp. 2892–2910, Dec. 2024, doi: 10.1016/j.csbj.2024.07.005. [45] M. Meeus, L. Wutschitz, S. Zanella-Béguelin, S. Tople, and R. Shokri, “The Canary’s Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text,” Feb. 19, 2025, arXiv : arXiv:2502.14921. doi: 10.48550/arXiv.2502.14921. [46] T. Sternak, D. Runje, D. Granoša, and C. Wang, “Automating Prompt Leakage Attacks on Large Language Models Using Agentic Approach,” Feb. 18, 2025, arXiv : arXiv:2502.12630. doi: 10.48550/arXiv.2502.12630. [47] S. Zeng et al. , “The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG),” Feb. 23, 2024, arXiv : arXiv:2402.16893. Accessed: Mar. 13, 2024. [Online]. Available: http://arxiv.org/abs/2402.16893 [48] M. Anderson, G. Amit, and A. Goldsteen, “Is my data in your retrieval database? membership inference attacks against retrieval augmented generation,” ArXiv Prepr. ArXiv240520446 , 2024. [49] Y. Li, G. Liu, C. Wang, and Y. Yang, “Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation,” Sep. 26, 2024, arXiv : arXiv:2406.19234. doi: 10.48550/arXiv.2406.19234. [50] A. Naseh, Y. Peng, A. Suri, H. Chaudhari, A. Oprea, and A. Houmansadr, “Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation,” Feb. 01, 2025, arXiv : arXiv:2502.00306. doi: 10.48550/arXiv.2502.00306. [51] Z. Qi, H. Zhang, E. Xing, S. Kakade, and H. Lakkaraju, “Follow my instruction and spill the beans: Scalable data extraction from retrieval-augmented generation systems,” ArXiv Prepr. ArXiv240217840 , 2024. [52] C. Jiang, X. Pan, G. Hong, C. Bao, and M. Yang, “RAG-Thief: Scalable Extraction of Private Data from Retrieval-Augmented Generation Applications with Agent-based Attacks,” Nov. 21, 2024, arXiv : arXiv:2411.14110. doi: 10.48550/arXiv.2411.14110. [53] M. Laymouna, Y. Ma, D. Lessard, T. Schuster, K. Engler, and B. Lebouché, “Roles, users, benefits, and limitations of chatbots in health care: rapid review,” J. Med. Internet Res. , vol. 26, p. e56930, 2024. [54] Z. Zhang, C. Yan, and B. A. Malin, “Membership inference attacks against synthetic health data,” J. Biomed. Inform. , vol. 125, p. 103977, Jan. 2022, doi: 10.1016/j.jbi.2021.103977. [55] A. R. Sarkar, Y.-S. Chuang, N. Mohammed, and X. Jiang, “De-identification is not enough: a comparison between de-identified and synthetic clinical notes,” Sci. Rep. , vol. 14, no. 1, p. 29669, 2024. [56] A. Elmahdy, H. A. Inan, and R. Sim, “Privacy Leakage in Text Classification: A Data Extraction Approach,” Jun. 09, 2022, arXiv : arXiv:2206.04591. doi: 10.48550/arXiv.2206.04591. [57] Lewis, Patrick, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler et al. "Retrieval-augmented generation for knowledge-intensive nlp tasks." Advances in neural information processing systems 33 (2020): 9459-9474. [58] Gao, Yunfan, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. "Retrieval-augmented generation for large language models: A survey." arXiv preprint arXiv:2312.10997 2, no. 1 (2023). [59] Rigaki, M. and Garcia, S., 2023. A survey of privacy attacks in machine learning. ACM Computing Surveys, 56(4), pp.1-34.