Full text
X-Vent: ICU Ventilation with Explainable Model-Based Reinforcement Learning Farhad Safaeia,e,*,1, Milos Nenadovi´ cb,1, Roman Liessnera,e,1, Raphael Theilenc,1, Jakob Wittensteinc,1, Jens Lehmanna,d,1 and Sahar Vahdati a,f aInstitute for Applied Computer Science (InfAI) e. V., Leipzig bInstitute Mihajlo Pupin cCarl Gustav Carus Dresden University Clinic, Technical University of Dresden dAmazon eDeutsche Bahn (DB) fScaDS.AI, Technical University of Dresden ORCID (Farhad Safaei): https://orcid.org/0000-0002-2606-3152, ORCID (Milos Nenadovi´ c): https://orcid.org/0000-0002-7840-7578, ORCID (Roman Liessner): https://orcid.org/0009-0004-2014-6899, ORCID (Raphael Theilen): https://orcid.org/0009-0001-8277-2922, ORCID (Jakob Wittenstein): https://orcid.org/0000-0003-4397-1467, ORCID (Jens Lehmann): https://orcid.org/0000-0001-9108-4278, ORCID (Sahar Vahdati ): https://orcid.org/0000-0002-7171-169X Abstract. This study introduces a Model-Based Deep Reinforcement Learning approach to enhance the effectiveness and transparency of mechanical ventilation treatment in the critical care setting of Intensive Care Units (ICUs). Distinct from conventional model-free methods, our approach benefits from the model-based algorithms’ capability to learn and interrogate dynamics models, enabling better generalization through synthetic data generation and a deeper understanding of the system dynamics. Coupled with Explainable AI (XAI) techniques, we focus on uncovering the underlying mechanisms of patient-ventilator interactions as learned by the AI. Our findings show a significant improvement in treatment efficacy, measured by Fitted Q Evaluation (FQE) metrics, achieved without the need for auxiliary rewards. This advancement not only highlights the potential of model-based reinforcement learning in healthcare but also emphasizes the importance of transparent AI design in healthcare applications. 1 Introduction In the intensive care unit (ICU), invasive mechanical ventilation [12] stands out as one of the most frequently employed life-saving treatments. It’s particularly crucial for patients experiencing acute respiratory failure, providing essential support to lung function. The significance of MV has been especially underscored during the current COVID-19 pandemic. However, while MV remains a crucial medical intervention, its potential risks are increasingly acknowledged. Inappropriate MV settings in ICU patients have been linked to organ damage. Currently, there are different strategies to protect the lungs from injury by the ventilator. Several protective MV settings have shown ∗Corresponding Author. Email: [email protected] 1The first two authors share an equal contribution. The work for Jens Lehmann is done outside of Amazon, and for Farhad Safaei, and Roman Liessner, outside DB. efficacy to reduce ventilator-induced lung injury, including low distending pressures, use of positive end-expiratory pressure (PEEP), low peak flow, low tidal volumes, and limited respiratory frequency. Several combinations of them are possible, but not all of them result in the lowest mechanical energy transferred to the lungs or impact on mean arterial pressure. Moreover, these general rules cannot be utilized to the same effect for different cases. Most concerning is that the lack of individualized ventilator settings could result in lung injury in a significant number of mechanically ventilated patients. Not to mention that competing treatment priorities can impair the urgent decision-making of attending doctors. It is to a great emergency that mechanical ventilation requires precise and customized handling to mitigate associated risks [19]. The integration of Artificial Intelligence (AI) in such critical healthcare settings, is revolutionizing patient care management [26]. Pioneering this domain, the application of Reinforcement Learning (RL), particularly Conservative Q-Learning (CQL) [7], has emerged as a promising method in recent studies for optimizing such complex medical interventions [6, 10]. While model-free RL methods, including CQL, have shown potential, Model-based Deep Reinforcement Learning (MBRL) emerges as a potent alternative, especially under data-constrained conditions [27]. The ability of MBRL to synthesize training data addresses the critical need for robust learning mechanisms in high-stakes settings such as ICU environments. Building upon the foundational work in model-free RL, such as the DeepVent study [6], we introduce X-Vent 2, a novel approach to provide decision-making support in mechanical ventilation through the integration of model-based reinforcement learning with explainable AI (XAI) techniques. This work is done in the context of the IntelliLung project3. 2The code for X-Vent can be found here: https://github.com/NIMI-research/ IntelliLung-X-Vent 3https://intellilung-project.eu/ ECAI 2024 U. Endriss et al. (Eds.) © 2024 The Authors. This article is published online with Open Access by IOS Press and distributed under the terms of the Creative Commons Attribution Non-Commercial License 4.0 (CC BY-NC 4.0). doi:10.3233/FAIA241069 4719
Enhanced Policy Performance: We demonstrate that our MBRL framework, using Conservative Offline Model-Based Optimization (COMBO), achieves superior policy performance as measured by Fitted Q Evaluation (FQE) [9] without the need for additional heuristic rewards. This could indicate a tendency of the works such as DeepVent [6] to align closely with the specific guidance provided by heuristic rewards, potentially at the expense of exploring a broader solution space. Explainability and Trust in RL for critical healthcare: Employing Explainable AI (XAI) techniques, we provide a dual explainability layer: one for the underlying transition model and the other for the policy. This not only enhances the trustworthiness and transparency of AI in critical medical applications but also aids in understanding the complex dynamics of the ventilation-patient system. Healthcare professionals further enriched the work by reviewing the results of the dual explainability layer. This collaborative examination underscores the practical applicability and reliability of our findings, bridging the gap between theoretical AI models and real-world clinical needs. MBRL in Data-Restricted Healthcare Settings: What sets our approach apart is the integration of dynamic system modeling with the robustness of CQL. COMBO’s capability to synthesize realistic training data from limited real-world ICU scenarios addresses a critical gap in existing methods. This enables our model to adapt and optimize ventilation strategies more effectively than previous RL applications. This research advances the application of MBRL in mechanical ventilation. We shall note that our vision is towards building a decision-support system as a human-in-the-loop process, underscoring the principle that the algorithm’s decisions are not intended for direct application without expert human intervention. The remainder of this paper is structured as follows: Section 2 reviews related work in the field, Section 3 discusses preliminaries, Section 4 details our methodology, Section 5 presents the results, Section 6 concludes the paper and outlines future research directions, and Section 7 delves into ethical considerations. 2 Related Work The subsequent text provides a comprehensive overview of research studies that have employed RL within the ICU setting for various applications, notably including the optimization of mechanical ventilation through RL. These studies are categorized into three groups: initial efforts, contemporary research, and studies incorporating explainability. Early efforts in applying RL in ICU. RL is used in [15, 16] for sepsis treatment, and in [5] for finding optimal treatment when delivering fluids, medications directly into a patient’s bloodstream. In [25], a Supervised-Actor-Critic (SAC) method and in [24] an inverse RL method are used to mitigate the problem of traditional RL on maximizing a long-term reward function that may cause a fatal impact on the patient. In these works, the authors focused on managing ventilation and sedative dosing decisions. Several works focus on learning dynamic treatment regimes, such as [14] that uses Qlearning to learn optimal policies. A deep learning approach is used in [11] for personalized dosing in critical care, showing its effectiveness in reducing medication complications by analyzing intensive care patient data, but not MV setting optimization. Recent advances in using RL for Mechanical Ventilation. It is only recently that RL-based methods have started being utilized in optimizing the settings of mechanical ventilation. In [10], the AIVE model (Artificial Intelligence model for Ventilation control during Emergence) is introduced for optimizing ventilation control during emergence from general anesthesia. Utilizing ventilatory and hemodynamic parameters from surgical cases, the model’s performance was tested and compared against clinicians’ policies. In another work, Dueling Double-Deep Q Network and kernel-based RL are used [15] as a mixture-of-experts framework for tailoring treatment to the individual patient of Sepsis disease. DeepVent [6] uses deep RL and suggests solutions for sparse reward and value overestimation. However, all of the above-mentioned works focus on a special emergency situation of mechanical ventilation using model-free RL. Model-based RL can make better use of limited data by learning a model of the environment. Apart from [10] which uses SHAP for exploring feature importance, most of the studies on RL treatments neglect model-based methods and explainability. Contrasting with [10], we employ a Model-Based Reinforcement Learning algorithm that additionally creates a simulation model, which is interpreted using Explainable AI (XAI) techniques and evaluated through feedback from physicians. Explainable AI in ICU Studies. [2] highlights that current algorithms used in ICU studies often operate as black boxes, making it difficult to understand which data features influence policy suggestions. A basic graphical model was used in [22] to illustrate the changes in patient health status and treatments. It uses traditional RL to produce medication recommendations for patients in the ICU. This simple graphical model is not applicable, nor effective for the complex setting of MV optimization. An algorithm based on Gradient Boosted Decision Trees is used in [23] for providing explanations on early sepsis prediction. This work provides Shapley values, which were calculated based on the change in expected risk output when a specific feature was present versus absent. The mentioned attempts only focus with single variable recommendation, despite the acknowledged importance of XAI in ICU settings. By providing a novel simulation model combining human and mechanical ventilation, our contribution is a step forward that opens opportunities for interpretation of recommendations for MV optimization, as a decision support system. 3 Preliminaries 3.1 Offline Reinforcement Learning Reinforcement learning architecture facilitates an agent to learn and optimize actions based on its experiential interactions. A standard RL model is structured around an agent, encompassing a set of states S, a repertoire of actions A, and a reward function R:S×A×S→ R. The primary objective of the agent is to develop a policy π(a|s) that is optimized to maximize the expected cumulative reward [20]. Offline RL involves training a reinforcement learning agent using a pre-collected, fixed dataset. Definition: In Offline RL, the agent learns from a dataset D= {(si,a i,r i,s i)}consisting of state transitions, actions, and rewards, without further environment interaction. 3.2 Conservative Q-Learning (CQL) Conservative Q-Learning (CQL) [7] is a method in offline-RL that addresses the overestimation bias commonly observed in Q-learning, especially in offline settings with out-of-distribution (OOD) actions and function approximation errors. CQL operates on the principle of learning a conservative Q-function, which provides a lower bound to the true value of a policy, mitigating overestimation. F. Safaei et al. / X-Vent: ICU Ventilation with Explainable Model-Based Reinforcement Learning4720
3.3 Model-Free vs. Model-Based RL in Offline Settings Model-Free RL in Offline Settings: In offline Model-Free RL, the agent learns from a fixed dataset of environment interactions, without ongoing access to the environment. It learns a policy π(a|s), value function V(s), or action-value function Q(s, a)based on stored experiences, constraining the learning process to the scope of the dataset. The action-value function update is typically represented as: QMF (s, a)=E[R(s, a)+γmax aQMF (s,a )].(1) where (s, a, R(s, a),s )are elements from the dataset. Model-Based RL in Offline Settings: In contrast, offline ModelBased RL involves learning or utilizing a model of the environment using the same fixed dataset. This model, symbolized as M(s, a), predicts the next state and reward. The agent uses this model to generate synthetic data, supplementing the offline dataset and potentially covering unexplored state-action spaces. The action-value function in this paradigm is updated as: QMB(s, a)=E(s,r)∼D⊕M(s,a)[r+γmax aQMB(s,a )].(2) This method is particularly useful in scenarios where real-world data is limited. The complexity inherent in accurately modeling the diverse and intricate nature of medical data presents a significant challenge. However, the ability of Model-Based RL to function as a virtual environment for simulating various medical situations makes it a compelling choice, especially where explainability is crucial. 3.4 Conservative Offline Model-Based Policy Optimization (COMBO) COMBO [27] addresses limitations in offline model-based reinforcement learning (MBRL) by eliminating the need for explicit uncertainty estimation, which can be challenging with complex datasets. This algorithm, extending CQL, focuses on optimizing a lower bound of policy performance. COMBO augments the offline policy evaluation of CQL by integrating a learned dynamics model to enrich policy improvement. The approach begins with the training of a probabilistic dynamics model, denoted as Tθ(s,r|s,a)= N(μθ(s,a),Σθ(s,a)), on an offline dataset D. This model generates synthetic transitions, supplementing the replay buffer Dmodel. Policy improvement is then conducted under the conservative critic. COMBO leverages a mix of real data from Dand synthetic data from model rollouts, enhancing its applicability in offline datasets with a limited variety of data which is typically the case in critical healthcare settings. 3.5 Explainable AI: SHapley Additive exPlanations Explainable AI, particularly through SHapley Additive exPlanations (SHAP) [1], plays a crucial role in our study, providing transparency and interpretability to Model-Based Reinforcement Learning models. Definition: SHAP is grounded in cooperative game theory and utilizes Shapley values to quantify the contribution of each feature in a model. For a prediction model f, the SHAP value of a feature ifor a given input xis defined as: φi(f,x)= S⊆N\{i} |S|!(|N|−|S|−1)! |N|! ·[fx(S∪{i})−fx(S)]. (3) where Nis the set of all features, Sis a subset of features excluding i, and fx(S)is the model’s prediction using features in S. Application in MBRL: Our Model-Based Reinforcement Learning framework utilizes SHAP values to elucidate the impact of patient parameters on decision-making at two levels: first, by revealing how state features contribute to transition models, thus clarifying the underlying patient-ventilation dynamics; and second, by highlighting the role of features within the policy framework. Such interpretability is essential in healthcare settings, where comprehending the rationale behind AI-driven recommendations is of paramount importance. 4 Methodology 4.1 Dataset Utilization and Processing Strategy Our study’s data preprocessing methodology was directly adopted from the protocol utilized in the DeepVent study, leveraging the MIMIC-III database [3]. This database includes comprehensive records from 61,532 ICU admissions at the Beth Israel Deaconess Medical Center, spanning the years 2001 to 2012. The primary motivation behind adopting DeepVent’s approach was to ensure the comparability of our results with this established benchmark in the field. Following the DeepVent model, we segmented the patient data into structured four-hour intervals. This included a detailed extraction of parameters such as vital signs, laboratory results, demographics, fluid administration, and ventilation settings, with a particular focus on the initial 72-hour period of mechanical ventilation. The data was methodically arranged into state, action, and reward arrays. In dealing with missing data, our study replicated DeepVent’s multi-tiered imputation approach. For cases with less than 30% missing data, the k-nearest-neighbor (KNN) method with k=3was applied [17]. Where missing data ranged from 30% to 95%, we employed a sample-and-hold technique, carrying forward the initial data point until a new value was recorded or a set limit reached. Mean value imputation was used when the initial data point was missing. Consistent with DeepVent, variables with over 95% missing data were excluded from our analysis to maintain data integrity and ensure a reliable comparison with their findings. 4.2 Formulation of the RL Problem In our study, we encounter a scenario where not all patient state features are fully observed, due to unmeasured or inherently unmeasurable factors. This situation ideally suits a Partially Observable Markov Decision Process (POMDP). However, to balance the complexity of POMDP with the practicality of our analysis, we adopt a Markov Decision Process (MDP) framework as an operational assumption. This simplification, aligned with the methodology in DeepVent, allows for a more tractable application of Reinforcement Learning (RL). Our RL episodes are defined from the start of intubation to 72 hours post-intubation. State Representation The state space Sconsists of 38 distinct variables, which include: •Demographic factors: Age, gender, weight, ICU readmission status, Elixhauser comorbidity score. •Vital signs: SOFA score, SIRS criteria, GCS, heart rate, systolic BP, diastolic BP, mean BP, shock index, respiratory rate, body temperature, oxygen saturation. •Laboratory parameters: Levels of potassium, sodium, chloride, glucose, BUN, creatinine, magnesium, carbon dioxide, hemoglobin, white blood cell count, platelet count, PTT, PT, INR, F. Safaei et al. / X-Vent: ICU Ventilation with Explainable Model-Based Reinforcement Learning 4721
arterial pH, partial pressure of carbon dioxide, base excess, bicarbonate. •Fluid management data: Urine output, use of vasopressors, intravenous fluid administration, cumulative fluid balance. Action Definition The ventilator settings that comprise our action space Ainclude: Adjusted tidal volume (Vt) based on ideal weight, Positive End Expiratory Pressure (PEEP), and Fraction of inspired oxygen (FIO2). These settings combine to form a tuple a=(v,o, p), with each element corresponding to a specific range within Vt,F IO2, and PEEP. Reward Mechanism At the core of our model lies the terminal reward function, r(st,a t,s t+1), designed to gauge patient survival outcomes. This function is structured to align with our primary goal of enhancing patient longevity, assigning a score of -1 in the event of a patient’s demise within 90 days post-treatment, and +1 if the patient survives this period. In addition to our main reward metric, we investigated the heuristic reward mechanism introduced by DeepVent, based on the Apache II score, for comparative analysis. The Apache II score [4], a widely recognized metric in ICU settings, aggregates various physiological indicators to assess patient health severity. 4.3 Off-Policy Evaluation In our study, evaluating the efficacy of reinforcement learning policies within a healthcare environment necessitated an Off-Policy Evaluation (OPE) technique, given the impracticality and ethical constraints of real-world patient interactions. This approach aligns with established practices in healthcare-related RL research. We adopted the Fitted Q Evaluation (FQE) method, following the precedent set in [9, 21]. Our implementation of FQE, facilitated by the d3rlpy toolkit [18], processes a dataset of state transitions D={st,a t,s t+1,r t}n t=1 along with a given policy π. The FQE algorithm iterates to compute yt=rt+γQk−1(st+1,π(st+1)) for each data point in D. The goal is to minimize the function Qk=argminf∈Fn i=1 (f(st,a t)−yt)2, where Fencompasses the neural network’s function class, resulting in a network Qπthat estimates state-action pair values under policy π. To assess policy performance, we focus on the average value of initial states, primarily the first four hours of ventilation data. Our FQE evaluation relies solely on the terminal reward, based on Dand the policy πactions. Similar to DeepVent, we analyzed the performance of the physician policy, which generated the episodes in our dataset, by computing its cumulative discounted rewards, thereby offering a comprehensive understanding of policy effectiveness in a real-world medical setting. 4.4 Manual Evaluation The results were carefully analyzed and described through a manual evaluation conducted by three clinicians. Among these experts, two are co-authors of this paper, ensuring that their insights and expertise directly contributed to the evaluation process. Their collective analysis provided a comprehensive assessment of the results, ensuring a thorough understanding from both clinical and AI perspectives. To ensure the evaluation was fair, the clinicians were introduced to the results in a manner that minimized bias. 5 Results 5.1 Action Distribution Analysis An analysis of the histogram distributions for PEEP, FIO2, and Adjusted Tidal Volume was conducted (shown in Figure 1) to compare the decision-making patterns of two reinforcement learning algorithms, DeepVent and our approach (X-Vent), against the recorded actions of physicians. PEEP and FIO2Settings. Both algorithms demonstrated a preference for moderate PEEP (0-5 cmH2O) and FIO2(35-50%) values, closely mirroring the decisions made by physicians. This suggests that the RL models are capable of learning clinically accepted practices for these settings. Adjusted Tidal Volume. A notable difference between the RL algorithms’ decisions was observed in the Adjusted Tidal Volume, with X-Vent favoring higher volumes, similar to the physicians. This pattern shows agreement with clinical practice over conservatism. Comparison of DeepVent and X-Vent. Despite similarities in action distributions, X-Vent exhibited a more conservative action profile across two of its settings, allowing higher ranges only for one. The alignment between DeepVent and X-Vent’s control actions indicates that both models have learned similar strategies from the available data, wherein the divergence in Tidal Volume suggests less conservatism on the side of X-Vent, which may be the result of utilizing a model for greater exploration of the solution space. Figure 1. Comparing the action distributions of DeepVent, X-Vent, and physicians for PEEP, FIO2, and Adjusted Tidal Volume. The results show that X-Vent effectively emulates the latest clinically validated practices, which have been corroborated by physician expertise. This supports the potential of X-Vent to not only replicate but also enhance ventilatory treatment strategies through its learned policies. 5.2 Policy Performance In our study, we evaluated the performance of various methods using Fitted Q Evaluation (FQE) metrics, as summarized in Table 1. Our findings reveal significant differences in the effectiveness of these methods for mechanical ventilation management. The X-Vent method, implemented without the heuristic reward (NO HeuristicRew), achieved the highest performance, recording an FQE score of F. Safaei et al. / X-Vent: ICU Ventilation with Explainable Model-Based Reinforcement Learning4722
Table 1. Performance of different methods evaluated by FQE. POLICY PERFORMANCE(FQE) X-VENT (NO Heuristic-Rew)0.793±0.004 X-VENT (Heuristic-Rew) 0.776±0.006 DEEPVENT (Heuristic-Rew) 0.743±0.005 DEEPVENT (NO Heuristic-Rew) 0.729±0.002 BEHAVIORAL CLONING 0.572±0.002 PHYSICIAN 0.502±0.007 0.793 ±0.004. This is particularly noteworthy as heuristic rewards are often chosen in reinforcement learning to improve performance by incorporating human knowledge. However, the superior performance of X-Vent with sparse rewards suggests that avoiding heuristic rewards can be advantageous, preventing the potential limitation of the solution space. Figure 2. SHAP plot referencing the influence of features on the policy’s choice of PEEP. Conversely, the DeepVent method demonstrated better results with the inclusion of heuristic rewards (Heuristic-Rew), scoring 0.743 ± 0.005, compared to its performance without them (NO HeuristicRew), which was 0.729 ±0.002. This could indicate a tendency of the DeepVent algorithm to align closely with the specific guidance provided by heuristic rewards, potentially at the expense of exploring a broader solution space. Behavioral Cloning, with an FQE score of 0.572 ±0.002, and the dataset representing traditional clinical decision-making, labeled ’Physician’, with a score of 0.502 ±0.007, lagged behind the adFigure 3. SHAP plot referencing the influence of features on the policy’s choice of FIO2. vanced reinforcement learning techniques. The Physician score does not reflect the quality of human decision-making but rather the inherent complexities in clinical decision-making processes, highlighting the potential of sophisticated AI models like X-Vent to support healthcare professionals in critical settings. These results underscore the effectiveness of the model-based XVent algorithm, particularly when employing sparse rewards, in enhancing mechanical ventilation strategies over other AI-based approaches and clinical practices. 5.3 Explanation of the Transition Model To ensure the results generated by the model are both medically meaningful and possess clinical validity, the model underwent a comprehensive evaluation process involving three clinicians. Furthermore, to enhance trustworthiness in the model’s decisions, a comprehensive medical explanation of the XAI is provided below. This step underscores the model’s reliability by integrating clinical expertise into its validation process and offering transparent insights into its decision-making mechanisms through XAI. Overview: Physiologic organ support, including mechanical ventilation, is critical in intensive care medicine for ensuring sufficient gas exchange, which entails proper oxygenation and carbon dioxide elimination [13]. Clinicians adjust various parameters e.g., tidal volume (Vt), positive end-expiratory pressure (PEEP), and inspiratory oxygen fraction (FIO2), each influencing oxygenation and CO2 elimination differently. To optimize therapy, clinicians assess patient states through physiological and laboratory data, set therapeutic F. Safaei et al. / X-Vent: ICU Ventilation with Explainable Model-Based Reinforcement Learning 4723
goals, and select strategies accordingly, employing pattern recognition based on their experience. This involves predicting patient status changes post-therapy to align closer with therapeutic objectives. Explaining the Policy: Figures 2, 3, and 4 showcase parameters that significantly impact the policy, as indicated by their SHAP values. This visualization facilitates a direct comparison between the decision-making processes of clinicians and the algorithm, highlighting the influence of these key parameters. PEEP: Positive End-Expiratory Pressure (PEEP), is the residual pressure in the lungs at the end of expiration. Lower PEEP levels are linked to increased collapsed lung tissue, adversely affecting gas exchange and oxygenation. As depicted in Figure 2, our algorithm predominantly bases PEEP selection on peripheral oxygen saturation (SpO2), a critical measure of blood oxygen content. This mirrors clinical practice, where SpO2is a key parameter for assessing oxygenation. Clinicians typically increase PEEP in response to low SpO2readings, a strategy that our algorithm appears to emulate, suggesting a parallel in decision-making processes between clinicians and the algorithm. Contrary to Figure 2’s implication, respiratory rate (RR) is not as emphasized in clinical practice as the algorithm suggests. However, body weight, positively correlated with PEEP, plays a significant role: higher body weight, associated with more collapsed lung tissue, necessitates increased PEEP for effective lung ventilation. RR, typically linked to CO2elimination rather than oxygenation, becomes crucial in cases of impaired lung function where both oxygenation and CO2elimination are compromised. Consequently, increased RR for CO2elimination often requires concurrent oxygenation improvement, potentially through increased PEEP. This interdependence might account for the algorithm’s emphasis on RR in PEEP determination. FIO2:Similar to PEEP, FIO2is crucial for oxygenation, and as such, SpO2is a key parameter for its selection, aligning with clinical practices as illustrated in Figure 3. However, the involvement of other parameters in FIO2selection is less direct. Among these, the Glasgow-Coma scale (GCS), shock index, and the presence of systemic inflammatory response syndrome (SIRS) are composite values derived from various distinct parameters. Consequently, the pattern recognized by the ML algorithm in choosing optimal FIO2could be influenced by multiple underlying processes. In clinical settings, these variables (GCS, shock index, and SIRS) typically provide broader patient status insights rather than directly informing the need for increased oxygenation support, which is a primary consideration for higher FIO2. Tidal Volume: Tidal volume (Vt) represents the air volume delivered to the lungs with each breath. Combined with respiratory rate (RR), Vtcontributes to the total gas volume available for exchange per minute, known as minute volume (Vm). Unlike PEEP and FIO2, which are linked to oxygenation, Vm, and thus Vtand RR, play a crucial role in CO2elimination. To maintain effective CO2removal, a reduction in Vtmust be balanced by an increased RR. This interplay is depicted in Figure 4, indicating the algorithm’s adoption of a similar pattern for adjusting Vt. Clinicians are tasked with not only assessing the current patient state but also predicting future changes based on past therapeutic decisions. The algorithm in our study can directly influence certain parameters through ventilation settings. For others, like diastolic blood pressure, it predicts future developments based on training data patterns, not solely on physiological processes. As shown in Figure 5, key parameters like systolic and mean blood pressure are crucial for predicting diastolic blood pressure. The formula for mean Figure 4. SHAP plot referencing the influence of features on the policy’s choice of Adjusted Vt. blood pressure, given as Mean blood pressure = 1/3 Systolic blood pressure + 2/3 Diastolic blood pressure, highlights this interrelation. Interestingly, Figure 5 reveals an inverse relationship between diastolic blood pressure and the other blood pressure measurements, a phenomenon not fully explained by physiological norms alone. This suggests the influence of clinical interventions. In ICU settings, clinicians use pharmaceuticals to manage blood pressure, creating scenarios where low systolic or mean pressures lead to increased diastolic pressures post-treatment, and high systolic or mean pressures lead to treatments that lower diastolic pressures. This inverse trend demonstrated by the algorithm underscores the critical role of therapeutic decisions in stabilizing blood pressure. Understanding the influence of age and the Systemic Inflammatory Response Syndrome (SIRS) criteria in our model is more complex. Younger patients often maintain more stable blood pressure and circulation even as their illness progresses, suggesting that age may inversely relate to changes in diastolic blood pressure. Consequently, the algorithm might interpret increased age as a risk factor for future decreases in diastolic blood pressure. On the other hand, the relevance of SIRS criteria in clinical practice has diminished over time, with more sensitive and specific assessment scores now preferred. Despite this shift, the algorithm still assigns significance to SIRS criteria, which, as a composite of diverse clinical parameters, presents a challenge in determining the exact factors influencing this discrepancy. Considering the role of pharmaceutical circulatory support in predicting diastolic blood pressure, as depicted in Figure 6, the relationship between mean blood pressure and predicted diastolic blood F. Safaei et al. / X-Vent: ICU Ventilation with Explainable Model-Based Reinforcement Learning4724
Figure 5. SHAP value distribution for the Diastolic blood pressure, showing interaction effects with other states and actions. Figure 6. SHAP value distribution for mean blood pressure augmented by the total amount of vasopressors administered. pressure is evidently influenced by current treatments. In cases of low doses of vasoactive medication, an increase in mean blood pressure corresponds to a rise in diastolic pressure, aligning with normal physiological patterns. However, during circulatory instability, pharmacological intervention is essential to sustain adequate blood pressure. If this leads to overcompensation and excessively high blood pressure, clinicians are likely to reduce medication dosage. Thus, in scenarios where high doses of vasoactive drugs coincide with elevated mean blood pressure, a subsequent decrease in diastolic pressure is often anticipated following medication adjustment. These observations confirm that our model-based RL approach effectively learns meaningful correlations and dependencies, which can be crucial for creating trust in its clinical recommendations. By capturing these nuances, the transition model not only aids in achieving a high degree of accuracy but also ensures the explainability of the AI-driven decision-making process. 6 Conclusion and Future Work This study presents X-Vent, an Explainable Reinforcement Learning (RL) framework designed to optimize mechanical ventilation strategies in Intensive Care Units (ICUs). Our evaluation demonstrates that X-Vent not only significantly exceeds the performance of traditional physician implemented strategies and Behavioral Cloning, a competing AI approach for offline RL, but also surpasses DeepVent, the leading state-of-the-art offline RL technique, in performance. This achievement is particularly noteworthy as it is accomplished without the need for auxiliary heuristic rewards that are a hallmark of DeepVent’s strategy. Moreover, X-Vent introduces two avenues: it facilitates the exploration of the model’s validity through verification of the transition model, and enables the extraction of insights from the AI policy. We illustrate how to interpret these results using contemporary Explainable AI (XAI) visualization techniques, highlighting the potential for deeper investigation into the model’s decisionmaking processes. Future Work. Our study marks significant progress in applying Model-Based Reinforcement Learning in critical care settings, yet several promising directions for further refinement and validation of our model remain. Key future explorations include advancing learning and representation methodologies to enhance model explainability and interactivity, specifically through sophisticated models and intuitive visualization interfaces. This will deepen our understanding of the complex dynamics between actions and patient features, and foster trust among healthcare practitioners. Moreover, developing methods that surpass Fitted Q Evaluation (FQE) in accuracy and reliability for off-policy evaluation is another crucial focus. Assessing the model’s generalizability across diverse ICU settings and patient populations, potentially through multi-center studies, will also be vital. Expanding the range of patient features and treatment options in our model could further improve its predictive capabilities, leading to more effective patient care strategies. In conclusion, XVent represents a significant step forward in the application of AI in healthcare, particularly in the critical domain of ICU ventilation. By enhancing performance, trust, and understanding, our work not only contributes to the field of AI in healthcare but also opens up new possibilities for future research and practical application. Ethical Considerations In conducting this study, stringent ethical standards were upheld, especially concerning patient data privacy and consent. All patient data, sourced from the MIMIC-III database, was de-identified and utilized in accordance with regulations, ensuring confidentiality and compliance with ethical guidelines. We acknowledge the criticality of AI system deployment in healthcare, particularly in ICU settings, and emphasize that our AI model, X-Vent, is designed with a focus on augmenting, not replacing, clinical decision-making. Acknowledgement Our work is supported by the European Union IntelliLung project with Grant Agreement No. 101057434. F. Safaei et al. / X-Vent: ICU Ventilation with Explainable Model-Based Reinforcement Learning 4725
References [1] D. Fryer, I. Strümke, and H. Nguyen. Shapley values for feature selection: The good, the bad, and the axioms, 2021. [2] O. Gottesman, F. Johansson, J. Meier, J. Dent, D. Lee, S. Srinivasan, L. Zhang, Y. Ding, D. Wihl, X. Peng, et al. Evaluating reinforcement learning algorithms in observational health settings. arXiv preprint arXiv:1805.12298, 2018. [3] A. E. W. Johnson, T. J. Pollard, L. Shen, L.-W. H. Lehman, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. A. Celi, and R. G. Mark. Mimic-iii, a freely accessible critical care database. Scientific Data, 3:160035, 2016. doi: 10.1038/sdata.2016.35. [4] W. A. Knaus, E. A. Draper, D. P. Wagner, and J. E. Zimmerman. Apache ii: a severity of disease classification system. Critical Care Medicine, 13(10):818–829, 1985. [5] M. Komorowski, L. A. Celi, O. Badawi, A. C. Gordon, and A. A. Faisal. The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care. Nature medicine, 24(11):1716–1720, 2018. [6] F. Kondrup, T. Jiralerspong, E. Lau, N. de Lara, J. Shkrob, M. D. Tran, D. Precup, and S. Basu. Towards safe mechanical ventilation treatment using deep offline reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 15696–15702, 2023. [7] A. Kumar, A. Zhou, G. Tucker, and S. Levine. Conservative q-learning for offline reinforcement learning, 2020. [8] P. Langley. Crafting papers on machine learning. In P. Langley, editor, Proceedings of the 17th International Conference on Machine Learning (ICML 2000), pages 1207–1216, Stanford, CA, 2000. Morgan Kaufmann. [9] H. M. Le, C. Voloshin, and Y. Yue. Batch policy learning under constraints, 2019. [10] H. Lee, H.-K. Yoon, J. Kim, J. S. Park, C.-H. Koo, D. Won, and H.-C. Lee. Development and validation of a reinforcement learning model for ventilation control during emergence from general anesthesia. npj Digital Medicine, 6(1):145, 2023. [11] R. Lin, M. D. Stanley, M. M. Ghassemi, and S. Nemati. A deep deterministic policy gradient approach to medication dosing and surveillance in the icu. In 2018 40th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), pages 4927– 4931. IEEE, 2018. [12] A. M. Luks. Ventilatory strategies and supportive care in acute respiratory distress syndrome. Influenza and other respiratory viruses, 7:8–17, 2013. [13] J. C. Marshall, L. Bosco, N. K. Adhikari, B. Connolly, J. V. Diaz, T. Dorman, R. A. Fowler, G. Meyfroidt, S. Nakagawa, P. Pelosi, J.-L. Vincent, K. Vollman, and J. Zimmerman. What is an intensive care unit? a report of the task force of the world federation of societies of intensive and critical care medicine. Journal of Critical Care, 37: 270–276, 2017. ISSN 0883-9441. doi: https://doi.org/10.1016/j.jcrc. 2016.07.015. URL https://www.sciencedirect.com/science/article/pii/ S0883944116302404. [14] A. Peine, A. Hallawa, J. Bickenbach, G. Dartmann, L. B. Fazlic, A. Schmeink, G. Ascheid, C. Thiemermann, A. Schuppert, R. Kindle, et al. Development and validation of a reinforcement learning algorithm to dynamically optimize mechanical ventilation in critical care. NPJ digital medicine, 4(1):32, 2021. [15] X. Peng, Y. Ding, D. Wihl, O. Gottesman, M. Komorowski, H. L. Liwei, A. Ross, A. Faisal, and F. Doshi-Velez. Improving sepsis treatment strategies by combining deep and kernel-based reinforcement learning. In AMIA Annual Symposium Proceedings, volume 2018, page 887. American Medical Informatics Association, 2018. [16] A. Raghu, M. Komorowski, I. Ahmed, L. Celi, P. Szolovits, and M. Ghassemi. Deep reinforcement learning for sepsis treatment. arXiv preprint arXiv:1711.09602, 2017. [17] C. Salgado, C. Azevedo, H. Proença, and S. Vieira. Missing data. in: Secondary analysis of electronic health records. Springer, Cham., 2016. [18] T. Seno and M. Imai. d3rlpy: An offline deep reinforcement learning library, 2022. [19] M. Smith and R. C. Heath Jeffery. Addressing the challenges of artificial intelligence in medicine. Internal Medicine Journal, 50(10): 1278–1281, 2020. doi: https://doi.org/10.1111/imj.15017. URL https: //onlinelibrary.wiley.com/doi/abs/10.1111/imj.15017. [20] R. S. Sutton and A. G. Barto. Reinforcement Learning: An Introduction. The MIT Press, second edition, 2018. URL http://incompleteideas.net/ book/the-book-2nd.html. [21] S. Tang and J. Wiens. Model selection for offline reinforcement learning: Practical considerations for healthcare settings. Proceedings of the 6th Machine Learning for Healthcare Conference, 2021. [22] C. P. Utomo, X. Li, and W. Chen. Treatment recommendation in critical care: A scalable and interpretable approach in partially observable health states. 2018. [23] M. Yang, C. Liu, X. Wang, Y. Li, H. Gao, X. Liu, and J. Li. An explainable artificial intelligence predictor for early detection of sepsis. Critical care medicine, 48(11):e1091–e1096, 2020. [24] C. Yu, J. Liu, and H. Zhao. Inverse reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units. BMC medical informatics and decision making, 19(2):111–120, 2019. [25] C. Yu, G. Ren, and Y. Dong. Supervised-actor-critic reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units. BMC medical informatics and decision making,20 (3):1–8, 2020. [26] C. Yu, J. Liu, S. Nemati, and G. Yin. Reinforcement learning in healthcare: A survey. ACM Computing Surveys (CSUR), 55(1):1–36, 2021. [27] T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn. Combo: Conservative offline model-based policy optimization, 2022. F. Safaei et al. / X-Vent: ICU Ventilation with Explainable Model-Based Reinforcement Learning4726