scieee AI-readable full text Open interactive document viewer

A review of approximate dynamic programming applications within military operations research

Rempel, M.,Cai, J.

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Rempel, M.; Cai, J. Article A review of approximate dynamic programming applications within military operations research Operations Research Perspectives Provided in Cooperation with: Elsevier Suggested Citation: Rempel, M.; Cai, J. (2021) : A review of approximate dynamic programming applications within military operations research, Operations Research Perspectives, ISSN 2214-7160, Elsevier, Amsterdam, Vol. 8, pp. 1-15, https://doi.org/10.1016/j.orp.2021.100204 This Version is available at: https://hdl.handle.net/10419/325707 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/4.0/ Operations Research Perspectives 8 (2021) 100204 Available online 14 October 2021 2214-7160/Crown Copyright © 2021 Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). Contents lists available at ScienceDirect Operations Research Perspectives journal homepage: www.elsevier.com/locate/orp A review of approximate dynamic programming applications within military operations research M. Rempel a,∗, J. Cai b aCentre for Operational Research and Analysis, Defence Research and Development Canada, 101 Colonel By Dr., K1A 0K2, Ottawa, Canada bCanadian Joint Operations Command, 1600 Star Top Road, K1B 3W6, Ottawa, Canada ARTICLE INFO Keywords: Sequential decision problem Markov decision process Approximate dynamic programming Reinforcement learning Military ABSTRACT Sequences of decisions that occur under uncertainty arise in a variety of settings, including transportation, communication networks, finance, defence, etc. The classic approach to find an optimal decision policy for a sequential decision problem is dynamic programming; however its usefulness is limited due to the curse of dimensionality and the curse of modelling, and thus many real-world applications require an alternative approach. Within operations research, over the last 25 years the use of Approximate Dynamic Programming (ADP), known as reinforcement learning in many disciplines, to solve these types of problems has increased in popularity. These efforts have resulted in the successful deployment of ADP-generated decision policies for driver scheduling in the trucking industry, locomotive planning and management, and managing highvalue spare parts in manufacturing. In this article we present the first review of applications of ADP within a defence context, specifically focusing on those which provide decision support to military or civilian leadership. This article’s main contributions are twofold. First, we review 18 decision support applications, spanning the spectrum of force development, generation, and employment, that use an ADP-based strategy and for each highlight how its ADP algorithm was designed, evaluated, and the results achieved. Second, based on the trends and gaps identified we discuss five topics relevant to applying ADP to decision support problems within defence: the classes of problems studied; best practices to evaluate ADP-generated policies; advantages of designing policies that are incremental versus complete overhauls when compared to currently practiced policies; the robustness of policies as scenarios change, such as a shift from high to low intensity conflict; and sequential decision problems not yet studied within defence that may benefit from ADP. 1. Introduction Many decisions are not made in isolation—decisions are made; new information, which was previously uncertain, is observed; given this new information, further decisions are made; more new information arrives; and so on. These types of decisions are aptly described as sequential decision problems,sequential decision making under uncertainty, or multistage decision problems and are characterized by decisions having an impact on future rewards received or costs incurred, the feasibility of future decisions, and in some cases the exogenous events that occur between decisions [1–3]. In essence, ‘‘today’s decisions impact on tomorrow’s and tomorrow’s on the next day’s’’ [2, p. 1], and if the relationship between decisions is not accounted for, then the outcomes achieved may be neither efficient nor effective. It has been known since the 1950s that such sequential decisions may be modelled as a Markov Decision Process (MDP), which consists of five components: a set of candidate actions; rewards that are received ∗Corresponding author. E-mail addresses: [email protected] (M. Rempel), [email protected] (J. Cai). as a result of selecting an action; the epochs at which decisions are made; the state, which is the information required to select an action, determine the rewards, and inform how the system evolves; and transition probabilities that define how the system transitions from one state to the next [4]. Given a MDP, the objective is then to find a decision policy—‘‘a rule (or function) that determines a decision given the information available’’ [3, p. 221], also referred to as a contingency plan, plan, or strategy [2, p. 22]—that makes decisions which result in the system performing optimally with respect to a given criterion. The classic approach to finding an optimal decision policy is to solve Bellman’s optimality equation via Dynamic Programming (DP) [5]. Within a defence context, DP has been applied to determine decision policies for a variety of sequential decision problems, including fleet maintenance and repair [6], scheduling basic training [7], selecting research and development projects [8], stay-or-leave decisions for military personnel [9], and dispatching medical evacuation assets [10]. https://doi.org/10.1016/j.orp.2021.100204 Received 26 May 2021; Received in revised form 26 September 2021; Accepted 30 September 2021 Operations Research Perspectives 8 (2021) 100204 2 M. Rempel and J. Cai Although DP provides an elegant framework to solve sequential decision problems, its limited usefulness in many real-world applications has been long acknowledged. This is due to the curse of dimensionality [5]—‘‘the extraordinarily rapid growth in the difficulty of problems as the number of variables (or dimensions) increases’’ [11]—and the curse of modelling, which is the need for an explicit model of how the system transitions from one state to the next [12]. While today’s computers can solve sequential decision problems with millions of states [13], many problems remain too large to be solved efficiently via classic DP methods. In addition, it is often the case that the transition probabilities between states are simply not known. Sequential decision problems with these characteristics permeate throughout defence, spanning the spectrum of force development, generation, and employment. For example: •within force development, decisions regarding capability investments, which may number in the hundreds, typically occur at fixed times during a business planning cycle and repeat on a yearly basis. Decision makers must consider both the short and long-term impact of the selected investments, as well as the investments not selected, while accounting for uncertainty surrounding future military obligations, changes in coalition and adversaries’ capabilities, defence specific inflation, etc.; •within force generation, decisions regarding how many commissioned and non-commissioned members to recruit in order to meet requirements across the spectrum of military occupations, while respecting the nation’s authorized strength and accounting for various uncertainties including yearly retirements, promotions, attrition, etc.; and •within force employment, decisions regarding which individuals to load onto helicopters during a mass evacuation operation, such as a major maritime disaster, while accounting for uncertainties including changes in weather, the individuals’ health, helicopter breakdowns, etc. As a result of these challenges, in these types of problems it is often not possible to find an optimal decision policy, and alternative approaches that focus on finding a good or near-optimal policy are required. The first approach, based on functional approximation, was suggested by Bellman and Dreyfus [14], with additional approaches being developed throughout the following decades across a variety of fields, including operations research, control theory, and computer science—see Powell [15] for a detailed discussion and list of relevant references. In addition, the field of mathematical programming, in particular stochastic programming, has developed sophisticated algorithms to solve problems with high-dimensional decision and state vectors, often seen in real-world sequential decision problems [16]. Within operations research, these approaches have been developed under a variety of names; in particular, neuro-dynamic programming, adaptive dynamic programming, and Approximate Dynamic Programming (ADP). The popularity of these approaches has grown over the last 25 years, as depicted in Fig. 1, with 2286 articles being published between 1995 and 9 April 2021 and the yearly publication rate growing from a single article to nearly 250 per year. More recently, the term ADP—‘‘a method of making intelligent decisions within a simulation’’ [17, p. 205] where ‘‘the resulting policies are not optimal, so the research challenge is to show that we can obtain [high-quality decision policies] that are robust under different scenarios’’ [18, p. 3]— has become the more commonly used term [3]. (Authors have recently started to use the label reinforcement learning as well, as evidenced by the recently published book entitled Reinforcement learning and Optimal Control [19], and a forthcoming book entitled Reinforcement Learning and Stochastic Optimization: A unified framework for stochastic decisions [20].) Of note, ADP-generated decision policies have been successfully deployed into industry, including policies to schedule drivers in the trucking industry [21–23], plan and manage locomotives [24, 25], and manage high-value spare parts within manufacturing [26]. Fig. 1. Number of ADP related articles published per year between 1995 and 9 April 2021. Source: Web of Science using the search pattern ‘‘approximate dynamic programming’’ OR ‘‘adaptive dynamic programming’’ OR ‘‘neuro-dynamic programming’’ in the title, abstract, and keywords. In this article we present the first review of applications of ADP set within a defence context. In particular, we focus on peer-reviewed literature within the field of military operations research; that is ‘‘[t]he application of quantitative analytic techniques to inform military [or civilian] decision making’’ [27]. This article’s main contributions are twofold. First, we review 18 decision support applications, spanning the spectrum of force development, generation, and employment, that use an ADP-based strategy and for each highlight how its ADP algorithm was designed, evaluated, and the results achieved. Second, based on the trends and gaps identified we discuss five topics relevant to applying ADP to decision support problems in defence: the classes of problems studied; best practices to evaluate ADP-generated policies; advantages of designing policies that are incremental versus complete overhauls when compared to currently practiced policies; the robustness of policies as scenarios change, such as a shift from high to low intensity in a conflict; and we suggest additional sequential decision problems within defence that may benefit from ADP-generated policies. The remainder of this article is organized as follows. Section 2provides relevant background information. Section 3presents the methodology used to conduct this review. Section 4and Section 5provide the main body of the review. Section 4reviews the 18 identified applications of ADP for decision support within defence, and Section 5 presents five topics relevant to applying ADP within a defence context. Finally, concluding remarks are given in Section 6. 2. Background This section provides background information on MDPs, DP, and ADP. While an exhaustive discussion of each topic is beyond the scope of this work, each section provides sufficient background to support the remaining sections of this article. For the interested reader, Puterman [2] provides an in-depth discussion on MDPs and DP, and Powell [3] provides a comprehensive introduction to ADP from an operations research perspective. In addition, Bertsekas and Tsitsiklis [12] discuss ADP from a controls perspective and Sutton and Barto [13] from a computer science viewpoint. 2.1. Markov decision process A MDP is a model of sequential decision making under uncertainty, and consists of five components: decision epochs,states,actions,rewards, and transition probabilities [2, Ch. 2]. These components are briefly described as follows. Operations Research Perspectives 8 (2021) 100204 3 M. Rempel and J. Cai •Decision epoch: A decision epoch is a point at time 𝑡at which a decision is made, where is the set of all decision epochs. This set may be a continuum or discrete. When a continuum, decisions may be made continuously, at random points when events occur, or at times chosen by a decision maker. In this case, the model is labelled as continuous-time. In the discrete case, decisions are made at all decision epochs and the model is labelled as discrete time. In addition, the set may be finite or infinite, with the model correspondingly labelled as finite horizon or infinite horizon respectively. •State: The state of a system 𝑆𝑡∈, where is the set of all possible states, may be defined as the minimally dimensioned function of history that is necessary and sufficient to select an action, determine the rewards or costs associated with the decision, and determine the transition probabilities to the next state [3, p. 179]. The state 𝑆𝑡is also referred to as the pre-decision state. •Actions: An action 𝑎𝑡∈𝑡(𝑆𝑡), where 𝑡(𝑆𝑡)is the set of all possible actions available in state 𝑆𝑡at time 𝑡, is an option selected by the decision maker that changes the state of the system such that it is in a new state at 𝑡+ 1. •Rewards: A reward is received, or a cost is paid, as a result of a decision maker selecting action 𝑎𝑡while in state 𝑆𝑡. This is defined as a real-valued function 𝐶(𝑆𝑡, 𝑎𝑡),∀𝑆𝑡∈, 𝑎𝑡∈𝑡(𝑆𝑡), and is known as the contribution function. •Transition probabilities: A transition probability function is a non-negative function 𝑝(𝑠′|𝑆𝑡, 𝑎𝑡)that denotes the probability that given the state 𝑆𝑡and the decision maker’s selected action 𝑎𝑡, the system is in state 𝑠′at time 𝑡+ 1. It is usually assumed that ∑𝑠′∈𝑝(𝑠′|𝑆𝑡, 𝑎𝑡)=1. Given a MDP, a decision maker’s goal is to find a decision policy, also referred to as a contingency plan, plan, or strategy [2, p. 22], that makes a sequence of decisions which results in the system performing optimally with respect to a given criterion. This is known as a Markov decision problem. For example, if the criterion is to maximize the expected total discounted contribution over a finite horizon, then this may be stated as [2] max 𝜋∈𝛱 E𝜋(𝑇 ∑ 𝑡=0 𝛾𝑡𝐶(𝑆𝑡, 𝐴𝜋 𝑡(𝑆𝑡))),(1) where 𝑇is the final decision epoch, 𝛾is a discount factor, and 𝐴𝜋 𝑡(𝑆𝑡) is a decision policy that determines the action selected for a given state. The objective is then to find the best decision policy—‘‘a rule (or function) that determines a decision given the available information in state 𝑆𝑡’’ [3, p. 221]—from the family of decision policies (𝐴𝜋 𝑡(𝑆𝑡))𝜋∈𝛱. The expectation operator E𝜋implies that the policy affects exogenous events that arise between decision epochs. For example, within the context of a ballistic missile scenario, the decision to launch an interceptor against an incoming missile may influence the adversary’s decision to launch further missiles. If the policy does not affect exogenous events, then the operator is written as E(⋅). It should also be noted that other objective functions may be used, such as maximizing the average contribution [2, p. 332], minimizing a risk measure [28], or a robust objective [29, p. 114]. 2.2. Dynamic programming In 1957, Bellman [5] showed that for a finite horizon MDP an optimal decision policy can be found by recursively computing Bellman’s optimality equation, 𝑉𝑡(𝑆𝑡) = max 𝑎𝑡∈𝑡(𝑆𝑡)(𝐶(𝑆𝑡, 𝑎𝑡) + 𝛾∑ 𝑠′∈ 𝑝(𝑠′|𝑆𝑡, 𝑎𝑡)𝑉𝑡+1(𝑠′)), (2) where the value 𝑉𝑡(𝑆𝑡)of being in state 𝑆𝑡is the value of taking the optimal decision, resulting in an immediate reward from the contribution function and the expected future value 𝑉𝑡+1(𝑆𝑡+1). Eq. (2) is referred as the standard form of Bellman’s equation [3, p. 60]. The resulting optimal decision for each state is then given as 𝐴𝜋 𝑡(𝑆𝑡) = arg max 𝑎𝑡∈𝑡(𝑆𝑡)(𝐶(𝑆𝑡, 𝑎𝑡) + 𝛾∑ 𝑠′∈ 𝑝(𝑠′|𝑆𝑡, 𝑎𝑡)𝑉𝑡+1(𝑠′)).(3) For infinite horizon MDPs, the subscript 𝑡is dropped as these types of problems are often studied in the steady-state, i.e., 𝑉(𝑠) = lim𝑡→∞𝑉𝑡(𝑆𝑡)[3, p. 67]. DP is a mathematical optimization approach to solving discrete time Markov decision problems. Regarding finite horizon MDPs, backwards induction provides an exact solution [2, p. 92], value iteration and policy iteration algorithms are employed to solve infinite horizon problems, and linear programming may be used for either. Thus, DP should be though of as ‘‘a collection of algorithms that can be used to compute optimal policies given a perfect model of the environment as a Markov decision process’’ [13, p. 73]. 2.3. Approximate dynamic programming In some applications, such as many of those studied in the field of operations research, a decision maker selects a vector of decisions 𝑥𝑡∈𝑡, where 𝑡represents a feasible region (defined by a set of constraints that depend on 𝑆𝑡, i.e., 𝑡=(𝑆𝑡)), rather than a single action from a small set of possible actions. Given this, and that the transition probabilities 𝑝(𝑆𝑡+1|𝑆𝑡, 𝑥𝑡)are often not known, Eq. (2) may be restated in expectation form [30, p. 240], 𝑉𝑡(𝑆𝑡) = max 𝑥𝑡∈𝑡(𝐶(𝑆𝑡, 𝑥𝑡) + 𝛾E(𝑉𝑡+1(𝑆𝑡+1)|𝑆𝑡)), (4) where 𝑆𝑡+1 =𝑆𝑀(𝑆𝑡, 𝑥𝑡, 𝑊𝑡+1). The function 𝑆𝑀(⋅)is the transition function that returns the next state 𝑆𝑡+1,𝑊𝑡+1 is exogenous information that arrives after a decision is made at time 𝑡and before 𝑡+1 as depicted in Fig. 2, and the expectation operator replaces the summation over probabilities in Eq. (2) and is over the random variable 𝑊𝑡+1. The decision policy is then given as 𝑋𝜋 𝑡(𝑆𝑡) = arg max 𝑥𝑡∈𝑡(𝐶(𝑆𝑡, 𝑥𝑡) + 𝛾E(𝑉𝑡+1(𝑆𝑡+1)|𝑆𝑡)). (5) In many real-world applications, computing an optimal policy via Eq. (4) is not feasible due to the three curses of dimensionality [3, pp. 3–6]: first, the state space may be large resulting in a prohibitively large number of times Eq. (4) must be computed; second, for each state 𝑆𝑡, the expectation over 𝑊𝑡+1 must be computed; and third, the expectation must be computed for each decision 𝑥𝑡. ADP aims to overcome these limitations through employing the concept of the postdecision state variable [3, p. 129–139], and using an approximation of the value function. The tradeoff is that the resulting policies are often not optimal; rather, the goal of ADP is to seek high-quality policies that are robust under different conditions. Similar to DP, ADP should be thought of as a framework rather than a single algorithm. In the remainder of this subsection, we present an overview of key concepts of the ADP framework within the context of a finite horizon MDP as it makes the presentation more explicit. We also assume the action to be taken at time 𝑡is represented by a vector of decisions 𝑥𝑡∈𝑡, and thus a decision policy is given as 𝑋𝜋 𝑡(𝑆𝑡)as opposed to 𝐴𝜋 𝑡(𝑆𝑡). First, we discuss how ADP overcomes the curses of dimensionality. Next, we summarize four classes of decision policies, i.e., implementations of Eq. (5). We follow this by discussing the algorithmic strategies employed to search for a good policy, where the choice of algorithm depends on the policy class selected. Lastly, we briefly show how ADP-generated policies are evaluated. Operations Research Perspectives 8 (2021) 100204 4 M. Rempel and J. Cai Fig. 2. Example of a sequence of decisions as a MDP with the post-decision state variable. At each decision epoch 𝑡a vector of decisions 𝑥𝑡is selected via a decision policy 𝑋𝜋 𝑡(𝑆𝑡). 2.3.1. Addressing the curses of dimensionality Post-decision state variable. The post-decision state variable 𝑆𝑥 𝑡represents the state of a system after a decision 𝑥𝑡is made, but before new exogenous information 𝑊𝑡+1 arrives as depicted in Fig. 2. Using this concept, the transition function 𝑆𝑀(⋅)may be broken down into two steps, 𝑆𝑥 𝑡=𝑆𝑀,𝑥(𝑆𝑡, 𝑥𝑡), (6) 𝑆𝑡+1 =𝑆𝑀,𝑊 (𝑆𝑥 𝑡, 𝑊𝑡+1), (7) resulting in 𝑆𝑡also being labelled as the pre-decision state variable. It follows that the optimality equation around the post-decision state is given as [3, p. 138]1 𝑉𝑥 𝑡−1(𝑆𝑥 𝑡−1) = E(max 𝑥𝑡∈𝑡(𝐶(𝑆𝑡, 𝑥𝑡) + 𝛾𝑉 𝑥 𝑡(𝑆𝑥 𝑡)|𝑆𝑥 𝑡−1)). (8) The advantage of Eq. (8) over Eq. (4) is that the expectation operator is outside the max operator and does not need to be computed for each decision. In addition, Eq. (8) requires solving a series of deterministic optimization problems—a significant advantage over Eq. (4). Approximating the value function. Eq. (8) assumes 𝑉𝑥 𝑡(𝑆𝑥 𝑡)is a lookup table with a one-to-one relationship between value and state. While this is an effective approach is some instances, it is a limitation in others. In order to overcome this limitation, ADP-based strategies often approximate the value function around the post-decision state variable 𝑉𝑥 𝑡(𝑆𝑥 𝑡) ≈ 𝑉𝑥 𝑡(𝑆𝑥 𝑡). There are a range of methods to approximate functions [33]. Strategies to approximate the value function (or the decision policy) may be grouped into three strategies [3, p. 223]: •Lookup table: A lookup table returns a discrete value for each state. An example of a lookup table is one based on hierarchical aggregation [3, pp. 290–304]; •Parametric representation: A parametric representation is an analytical function, designed by an analyst, that involves a vector of tunable parameters 𝜃. Parametric representations may be characterized as having either a linear or non-linear architecture; however, both are based on basis functions 𝜙𝑓(𝑆𝑥 𝑡), 𝑓 ∈, where is a set of features based on information from the post-decision state variable [3, p. 305]; and •Nonparametric representation: A nonparametric representation is an approach that builds local approximations based on observations. Examples of nonparametric representations are Neural Network (NN)s, kernel regression, 𝑘-nearest neighbour, etc. [3, pp. 316–324]. 2.3.2. Decision policies There are four classes of decision policies [3, Ch. 6]: myopic, and its extension known as myopic Cost Function Approximation (CFA); Policy Function Approximation (PFA); Value Function Approximation (VFA); 1The discount factor 𝛾is not included in this equation in Powell [3, p. 138], however it is included in other sources, e.g., Mes and Rivera [31] and McKenna et al. [32]. and a Direct Lookahead Approximation (DLA).2In addition, hybrid policies may be created by mixing policies from two or more of these categories. Table 1 lists the policy classes, an example policy for each, and a description of the policy class. These policy classes may be grouped into two categories: policy search and lookahead approximations, as depicted in Fig. 3 [34]. The policy search category includes PFAs and myopic CFAs; that is, those policies that are characterized by a low-dimensional vector 𝜃that must be tuned and do not explicitly model the downstream impact of a decision. The lookahead approximation category includes DLAs and VFAs. While these approximations also include tunable parameters, these types of policies explicitly model the impact of today’s decisions on the future. 2.3.3. Algorithms The choice of an algorithmic strategy to search for a good decision policy is dependent on the class of decision policy itself as depicted in Fig. 3. As an in-depth discussion of each algorithmic strategy is beyond the scope of this work, for a deeper appreciation of these algorithms see previously cited works by Powell and the references provided below. Policy function approximation: A PFA is characterized by a vector 𝜃of tunable parameters, which may range from one parameter to thousands, and thus searching for a good decision policy amounts to searching for the best vector 𝜃. This may be tackled by a variety of algorithms, generally described as stochastic search, with individual algorithms being categorized as being either derivative-based or derivative-free [20, Ch. 12]. Derivative-based methods tend to employ a stochastic gradient algorithm where the gradient is either based on a numerical derivative, such as in simultaneous perturbation stochastic approximation [35,36], or an exact derivative; in situations where access to the derivative is not possible, such as when a policy is being evaluated using a complex simulation or obtaining observations are expensive in terms of time or money, derivative-free methods, such as genetic algorithms, may be used [35]. Myopic cost function approximation: A myopic CFA is in essence an optimization model. When using a non-parameterized approximation, the decision function may be solved via an appropriate mathematical programming method. When using a parameterized approximation, the parameter vector 𝜃may be tuned via stochastic search while a solver is used to determine a decision for a given parameter vector. Direct lookahead approximation: When decisions are vector valued a DLA may be modelled as a mathematical program—deterministic in the case of point forecasts, and stochastic when point estimates are replaced by sample realizations. In the latter situation, two-stage or multistage stochastic programming are core approaches to solving direct lookahead policies [16]. Value function approximation: When using this type of approximation, searching for a good policy is akin to tuning the approximation’s parameter vector 𝜃. Two common strategies employed are Approximate Value Iteration (AVI) and Approximate Policy Iteration (API). AVI performs one iteration of policy evaluation and uses this information to conduct policy improvement, looping over these two 2Note that many authors use the term ‘cost function approximation’ or ‘cost-to-go’ to refer to a value function approximation [15, p. 328]. Operations Research Perspectives 8 (2021) 100204 5 M. Rempel and J. Cai Table 1 Summary of decision policies. See [3, Ch. 6] for a complete description. Class Example policy and class description Myopic 𝑋𝜋(𝑆𝑡) = argmax 𝑥𝑡∈𝑡 𝐶(𝑆𝑡, 𝑥𝑡) Makes decisions based on near-term rewards without regard to long-term impacts. An extension of this policy is known as CFA [29, p. 120], which may include a parameter vector 𝜃, changing or adding constraints, or adding a correction term. PFA 𝑋𝜋(𝑆𝑡|𝜃) = 𝜃0+𝜃1𝜙1(𝑆𝑡) Returns a decision 𝑥𝑡directly based on the state 𝑆𝑡without solving an optimization problem. May take the form of a lookup table, e.g., if military airlift is not available, then use contracted airlift; a parametric model; or a nonparametric model. VFA 𝑋𝜋 𝑡(𝑆𝑡|𝜃) = argmax 𝑥𝑡∈𝑡 (𝐶(𝑆𝑡, 𝑥𝑡) + 𝛾∑𝑓∈𝜃𝑓𝜙𝑓(𝑆𝑥 𝑡)) Explicitly accounts for the downstream impact of a decision through approximating its value around the post-decision state variable. Approximation may take the form of a lookup table, parametric, or nonparametric approximation. DLA 𝑋𝜋(𝑆𝑡|𝜃) = argmax 𝑥𝑡𝑡 ,…, 𝑥𝑡,𝑡+𝐻∑𝑡+𝐻 𝑡′=𝑡𝐶( 𝑆𝑡𝑡′, 𝑥𝑡𝑡′) Makes a decision at 𝑡= 0 by optimizing decisions over a finite horizon from 𝑡to 𝑡+𝐻. When the number of possible actions and outcomes over a short time horizon are small, an approach based on tree search, using either complete enumeration, Monte Carlo sampling of outcomes, or employing a roll-out heuristic to evaluate what may happen once reaching a state, is feasible [3, pp. 225–227]. When the problem involves a vector of decisions, a deterministic or stochastic DLA based on mathematical programming may be used (deterministic example shown). In the deterministic case, point forecasts of future exogenous information available at time 𝑡are used to design the policy, such as [29, pp. 121–122]. In the stochastic case, point estimates are replaced by sample realizations. Fig. 3. Decision policies and algorithms. Decision policy categories are based on Powell [34]. steps for a fixed number of iterations 𝑁. In contrast, API performs multiple policy evaluations and uses this data to improve the policy. API involves two nested loops—an outer loop that controls the number of policy improvements 𝑀and an inner loop that controls the number of policy evaluations 𝑁. Both AVI and API simulate the decision making process within policy evaluation, removing the need to loop over all possible states. AVI and API algorithms are widely used within ADP, and generally follow the aforementioned outlines; however, implementations tend to be problem specific—see Powell [3, Ch. 10] for several examples and Bertsekas [37] for an overview of API. Factors that affect an implementation include whether the problem is modelled as having a finite or infinite horizon, the type of value function approximation chosen, and how the approximation’s parameter vector 𝜃is learnt. Regarding the latter, many options exist that produce a good value function approximation (see Geist and Pietquin [38] for an overview), such as Least Squares Temporal Difference (LSTD) [39], Least Squares Policy Evaluation (LSPE) [40], or Support Vector Regression (SVR) [41]. In addition, within the 𝑚th policy improvement step an update equation of the form 𝜃𝑚= (1 − 𝛼𝑚−1)𝜃𝑚−1 +𝛼𝑚−1  𝜃𝑚, (9) is often used to update the estimate 𝜃𝑚for the parameter vector, where  𝜃𝑚is the most recent observation and 𝛼𝑚−1 is a step size. There are many step size options, including constant, generalized harmonic, bias-adjusted Kalman filter, etc. [42]. As with the remainder of the algorithmic options, ‘‘strategies for stepsizes are problem dependent’’ [3, p. 451]. 2.3.4. Evaluating ADP-generated policies When evaluating an ADP-generated policy, a benchmark should be prepared because the algorithm selected to search for the policy or the policy itself may be poorly designed. Suggested benchmarks include: an optimal policy via DP; an optimal deterministic version of the problem; or a myopic policy [30]. In particular, the first option enables an ADPgenerated policy to be compared to one that is optimal, however this is only feasible for relatively simple problems. Comparing to a myopic policy establishes the benefit of incorporating the downstream value in a decision policy. In addition, ‘‘direct comparison of [an ADP] algorithm’s performance to earlier heuristic solutions should be made’’ [43, p. 10], where earlier solutions may include currently practices policies or previously published results. Regardless of the benchmark selected, these comparisons may be done by simulating the implementation of the policies and comparing their objective function values, given in Eq. (1), and ‘‘[computing] the variance of the estimate of the difference to see if it is statistically significant’’ [3, p. 154]. 3. Review methodology To identify relevant articles for this review, Web of Science was searched on 9 April 2021 for articles published from 1995 onwards. Title, abstract, and keywords were searched using the following pattern: •(military OR army OR navy OR ‘‘air force’’) AND (‘‘approximate dynamic programming’’ OR ‘‘adaptive dynamic programming’’ OR ‘‘neuro-dynamic programming’’). Operations Research Perspectives 8 (2021) 100204 6 M. Rempel and J. Cai This search returned 16 results. We manually refined these results to those that focus on military operations research—‘‘[t]he application of quantitative analytic techniques to inform military [or civilian] decision making’’ [27]—and were deemed to be an application-based article. This resulted in a set of 10 articles. In addition, although the journals Military Operations Research and The Journal of Defence Modelling and Simulation are included in the Web of Science databases, the above search pattern did not return results from these journals. This is most likely due to the search pattern requiring (military OR army OR navy OR ‘‘air force’’); thus, title, abstract, and keywords in each of these journals were searched using this pattern. Two articles from each were returned. One article from Military Operations Research, Flint et al. [44], focused on best practices for random number generation and variance reduction within the application of ADP algorithms to defence problems. Hence, it was deemed not to be an application paper and thus not within the scope of this review. Lastly, articles were added via references contained within these 13 articles. In total, 18 application-based articles were identified as relevant and are discussed in the following section. 4. Applications of ADP within military operations research In this section, we present a summary of the 18 application-based articles identified through the aforementioned literature search. Table 2 lists each study reviewed, its application area, and characteristics of the ADP policies and algorithms implemented. The characteristics listed focus on those discussed in Section 2.3, namely: •type of decision policy—Myopic CFA, PFA, VFA, DLA, or hybrid; •value function approximation strategy—lookup table, parametric, or nonparametric; •value function model—hierarchical aggregation, linear architecture, NN, etc.; •algorithmic strategy—stochastic search, mathematical programming, stochastic programming, AVI, API; •approach to updating the value function model parameters 𝜃— temporal difference learning, LSTD, LSPE, SVR, etc.; and •step size 𝛼𝑚—constant, generalized harmonic, polynomial, etc. For some articles listed, not enough information was provided in order to identify how certain characteristics were addressed by the authors. In this case, the characteristic is listed as Not specified. In addition, for some articles listed certain characteristics are not applicable. In this case, the characteristic is listed as N/A. Further details are given below. Studies are organized into three categories—force development, force generation, force employment—and then listed in chronological order. 4.1. Force development Within the force development category, two articles have been published: one that focuses on investment planning [45] and one on force structure analysis [46]. These two articles employ significantly different ADP strategies to solving their respective sequential decision problems, where the difference is rooted in the choice of decision policy. Fisher et al. [45]: Every fiscal year, the Royal Canadian Navy selects a subset of candidate non-strategic capital projects that will receive funding, where each requires a fixed expenditure over a multiyear period and has an associated value to the Navy. Since each fiscal year’s funding does not carry over to the next, historically Navy planners have tried to allocate the entire available budget while striving for high long-term value. Fisher et al. [45], with the aim to maximize the long-term value, designed a knapsack-based myopic CFA parameterized by a single scalar parameter 𝛼that reserves a portion of the budget, between zero and 100%, for future projects. The authors employed a simple enumeration method to assess the impact of the policy’s parameter. The remaining characteristics listed in Table 2 are not relevant to this article, and thus listed as N/A. The results obtained from this policy were compared to that obtained from a single optimization model that considered projects across the entire planning period simultaneously, i.e., the model knows how the future will unfold. Fisher et al. showed that for a 25 year planning period when 𝛼= 0.4(40% of the budget reserved) the myopic CFA was able to achieve 90% of the value found by the single optimization model, which was considerably better than the 73% obtained when 𝛼was set was zero (no budget reserved). The authors concluded that ‘‘planners who try to expend all of their planning space in the current year end up implementing a set of projects that have considerable lower value and higher cost’’ [45, pp. 88–89]. Southerland and Loerch [46]: The authors proposed an ADP algorithm based on AVI to support deliberations within United States Department of the Army’s annual force structure review. This review, known as Total Army Analysis, aims to ‘‘identify changes, within total personnel constraints mandated by Congress, to the existing force structure that maintain or improve the Army’s ability to meet Combatant Command mission requirements while maintaining the ability to respond to future crises’’ [46, p. 25]. The algorithm’s objective included two weighted contributions: first, the force structure’s ability to meet mission requirements over the following year; and second, the ability to meet readiness requirements. The authors employed a VFA based on diffusion wavelet transform [61], where the advantage of using this approach within an ADP algorithm is that ‘‘it generates the best basis functions as part of the approximation process [whereas] regression methods require the basis functions to be pre-specified’’ [62, p. 59]. In addition, the approach to learning the value function model’s parameters is based on a three phase approach: initialization in which states are selected based on random decisions; transition in which a transition occurs from random decisions to VFA-based decisions; and an exploitation phase where states are selected based solely on the VFA. It should be noted that the ADP algorithm’s pseudocode is not listed in this article; see Southerland [63, pp. 52–54] and Balakrishna [62, pp. 57–58] for details. The authors conducted a set of experiments, varying the weights within the objective function and the number of sampled states used within the VFA to model the entire state space. The resulting ADPgenerated policies were then compared to a myopic policy. To do so, 1000 20-year scenarios of mission and readiness requirements for 20 units types were generated. Each policy was then applied within each scenario three times, each time with a different initial force structure. With regards to meeting mission requirements, the ADPgenerated policies showed an average improvement between 1.8% and 15.7% as compared to the myopic policy, whereas the average improvement to meet readiness requirements ranged between 7.6% and 19.9%. It is worth noting that increasing the number of states used in the diffusion wavelet transform approximation did not result in a better ADP-generated policy. 4.2. Force generation One application within the force generation domain has been published [47]. This study, which focused on personnel sustainment, differs in its ADP strategy from those used to date in the force development domain; in particular, the use of a VFA that employs a parametric strategy and as a result the selected approach to update the value function model parameters. However, this approach aligns with those used in the force employment domain. Hoecherl et al. [47]: In this study the authors investigated the problem of personnel recruitment and promotion decisions in the United States Air Force. The authors formulated two ADP algorithms to obtain decision policies. The first algorithm uses an API strategy with a VFA based on a set of unspecified basis functions in a linear Operations Research Perspectives 8 (2021) 100204 7 M. Rempel and J. Cai Table 2 ADP applications within military operations research during the period 1995–2021. Articles are divided by horizontal lines into three groups: force development (top group), force generation (middle group), and force employment (bottom group). Article Application area Policy type Approx. Strategy Value function model Algorithm strategy Model update algorithm Step size Fisher et al. [45] Investment planning CFA Parametric N/A Enumeration, mathematical programming N/A N/A Southerland and Loerch [46] Force structure analysis VFA Non-parametric Diffusion Wavelet Transform AVI Exploration/Exploration Not specified Hoecherl et al. [47] Personnel sustainment VFA Parametric Linear/Separable piecewise linear API/AVI LSTD/Concave Adaptive Value Estimation (CAVE) Decreasing Ross et al. [48] Situational awareness VFA Lookup table Q-factor AVI Q-learning Not specified Bertsekas et al. [49] Missile defence VFA Non-parametric /Parametric NN/Linear API Least squares/TD(𝜆) Decreasing Popken and Cox [50] Air combat DLA N/A N/A Game theory, Multistage stochastic program N/A N/A Sztykgold et al. [51] Battlefield strategy VFA Lookup table N/A AVI TD(𝜆) Not specified Wu et al. [52] Airlift VFA Parametric Linear AVI N/A Not specified Powell et al. [18] Airlift VFA Parametric Piecewise linear concave AVI Not specified Adaptive Ahner and Parson [53] Weapon target assignment VFA Lookup table N/A AVI N/A Not specified Rettke et al. [54] Combat medical evacuation VFA Parametric Linear API LSTD Harmonic Davis et al. [55] Missile defence VFA Parametric Linear API LSTD Harmonic Laan et al. [56] Illegal fishing patrols VFA Lookup table Hierarchical aggregation AVI N/A Harmonic McKenna et al. [32] Inventory routing VFA Parametric Linear API Least squares/LSTD Harmonic Robbins et al. [57] Combat medical evacuation VFA Lookup table Hierarchical aggregation API N/A Harmonic Summers et al. [58] Missile defence VFA Parametric Linear API LSPE/LSTD Harmonic Jenkins et al. [59] Combat medical evacuation VFA Parametric Linear API SVR Polynomial Jenkins et al. [60] Combat medical evacuation VFA Parametric/Nonparametric Linear/NN API LSTD/NN Polynomial architecture, LSTD, and a decreasing step size to update the VFA’s tunable parameters. The authors developed variants of this algorithm, specifically using instrumental variables in the regression equation as they ‘‘have been found to significantly improve regression performance’’ [47, p. 3080] and Latin Hypercube Sampling ‘‘to generate an improved set of post-decision states [and] help ensure uniform sampling across all possible dimensions’’ [47, p. 3080]. The second algorithm uses an AVI strategy, a VFA based on separable, piecewise linear value function approximations, and the CAVE algorithm [64] to update the approximations. The authors compared the ADP-generated policies against the currently employed sustainment line policy in two 50-year scenarios of sustaining personnel: a small problem instance and a larger problem instance, where the latter was of particular interest to the Air Force. The difference between the two scenarios is defined in terms of number of career fields, number of officer grades, number of commissioned years of service, and number of new recruits. Across the two problem instances the effect of changing several scenario and algorithmic parameters were evaluated, including order of the basis function, number of hiring and promotion decisions, and inclusion or exclusion of the aforementioned variants. Overall, the LSTD-generated policy performed best in the smaller problem instance, showing approximately an 8% improvement in the objective function (reduction in cost) as compared to the benchmark policy when using fourth-order basis functions and the two previously mentioned variants. However, this algorithm performed significantly worse than the benchmark policy in the larger problem. The CAVE-generated policy performed better in the larger scenario instance, showing an improvement of 2.8%. Although not published in a peer-reviewed journal or conference, it should be noted that similar problems have been studied by Bradshaw [65] and West [66] as part of their graduate programs at the Air Force Institute for Technology, and Situ [67] at George Mason University. 4.3. Force employment Within the force employment category, 15 articles have been published. Combat medical evacuation is the topic with the most publications [54,57,59,60], with the remaining articles covering a wide variety of topics including airlift [18,52], missile defence [49,55], and inventory routing [32]. All studies in this category developed policies based on VFA, with the exception of Popken and Cox [50] which used a DLA. Within those studies that used a VFA, all three approximation strategies—lookup table, parametric, non-parametric—were used, and six used an algorithm based on AVI with the remaining eight being based on API. Ross et al. [48]: The requirements of a Commander’s situational awareness evolve over time, and thus information fusion strategies must adapt to new information and new requirements. In this application, Bayesian networks were used to infer, based on a variety of data sources, the type of military unit formed by clusters of military vehicles. Under the assumption that available computational resources are insufficient to evaluate all incoming data in the Bayesian network, the ADP algorithm’s aim is to optimize the use of the data to maximize the knowledge about the battlespace in order to inform a military commander. The ADP algorithm is based on Q-learning [13, pp. 131–132], however a detailed algorithm is not presented within this article. Operations Research Perspectives 8 (2021) 100204 8 M. Rempel and J. Cai The algorithm was tested in a scenario with up to five military unit types operating within a 75 km by 90 km area. Simulated training data was used to train the ADP algorithm, where on average at any time there were 14 clusters of military vehicles present and seven false alarm clusters. The results indicated the ADP algorithm outperformed a random controller, increasing the probability of correct classification to 0.53 from 0.31. Bertsekas et al. [49]: This study examined a theatre missile defence problem in which the attacker has a limited inventory of multiple types of missiles, the defender has multiple types of interceptors where the number launched at a given time is limited by the of number launchers, and the defender must allocate specific interceptors to specific missiles. With the aim to maximize the expected value of targets surviving at the end of the battle, the authors describe an API algorithm that used a non-parametric VFA based on a NN with a single hidden layer, as well as an API algorithm that used a VFA with a linear architecture. Several variations to compute the estimate of the weights within both were explored: Monte Carlo simulation and least squares; temporal difference learning (TD(𝜆)); and one labelled as optimistic policy iteration which is similar to an AVI algorithm. A decreasing step size is then used to update the weights. The authors compared their ADP-generated policies within a representative scenario, which included one attack missile type, one interceptor missile type, and three targets to be defended. In this scenario, the authors used four features to approximate a state’s value: missile leakage, surviving target value, number of targets, and number of interceptors. Twenty-three test cases were developed that spanned a wide range of scenario instantiations, including those that overwhelmed the defender or the attacker, each with a different degree of interceptor effectiveness. For each test case, results for the API algorithm using Monte Carlo simulation and a NN, Monte Carlo simulation and a linear architecture, and the optimistic policy iteration algorithm were reported. As each algorithm includes a variety of parameter settings, the authors used ‘‘a mixture of insight and preliminary experimentation to arrive at a combination of settings for each algorithm that would work robustly’’ [49, p. 47]. The articles lists some, but not all, of the parameter settings. Results were compared to both an exact decision policy and a heuristic. Although a detailed comparison of the test cases was not reported, amongst the authors’ conclusions it is worth noting that ‘‘algorithm performance is highly dependent on the tuning of these parameters’’ [49, p. 50] and that no single VFA was superior across the tests conducted. Popken and Cox [50]: In this article the authors examined the modelling of air combat to support air warfare operational planning in an effort to demonstrate ‘‘embedding of optimization algorithms [that account for the complexities and uncertainties in real-world conflicts] into an operational planning cycle operating over a multi-period conflict can be made practical’’ [50, p. 127]. The authors modelled the problem as a stochastic two-player game, implemented as a simulation– optimization, with a hierarchy of decisions, i.e., force allocation decisions down to flight control decisions. Within the force allocation level, at each decision epoch aircraft were assigned to one of five roles: counter air, air defence, target reduction, anti-aircraft fire suppression, or other. These decisions are generated by a DLA via a multistage stochastic program. Target assignments were then based on a greedy heuristic and a stochastic simulation was used to determine the outcomes. In testing, the authors utilized a representational scenario of an air conflict between the United States and North Korea; in particular, 15 base locations for United States forces positioned in and around South Korea, and 201 targets consisting of various infrastructure types in North Korea. ADP-generated policies were generated through varying two algorithmic parameters and one scenario parameter. The two algorithmic parameters were the number of improvement steps in the simulation–optimization algorithm, and horizon weight which balanced the need for forces to survive in the near term versus the long term. The one scenario parameter varied was the quality of the intelligence assessment, which influenced the accuracy of information and thus decisions on how to allocate aircraft. The authors concluded that the optimal tradeoff between time and performance occurred around 100 improvement iterations, a lower horizon weight to not overshoot enemy targets, and poor intelligence quality which contributes to lower horizon weight and balances value against survivability. Sztykgold et al. [51]: This study explored decision-making in a two-sided terrestrial conflict where the goal is to determine a battle strategy for one side to control a target location while maintaining a minimum strength ratio between friendly and adversarial forces. The authors modelled the problem as a stochastic game in a multilayered graph, and designed an algorithm based on temporal difference learning (TD(𝜆)) which employs a variation of generalized policy iteration [13]. For the purposes of this review, we have classified this algorithm as an AVI using VFA based on a lookup table, and the value function model as N/A since the authors do not describe any form of state aggregation. The algorithm’s pseudocode is not provided in the article. In testing, the AVI algorithm’s result was compared with two optimal policies found via enumeration. This was possible due to the small size of the scenario, which included 16 vertices and a strength ratio constraint of two to one for allies to enemies. The authors ran three experiments, each with 25 trails with 1000 iterations of the algorithm and varied 𝜆to affect how learning was performed. The results showed that ‘‘more than 84% of trials return the optimal control, and the value of 𝜆does not change this ratio’’ [51, p .6]. The authors did not provide further details regarding the experiments. Wu et al. [52]: In this article the authors discussed a variety of existing approaches to optimize military airlift, and presented an ADP approach to improve solution quality for the United States Air Mobility Command. With the aim to maximize cargo throughput, the authors describe an ADP policy that uses a linear approximation centred around the post-decision state variable. In addition, the authors describe how the policy may be extended to incorporate expert knowledge. Of note, while the authors describe how the value estimate of the post-decision state variable may be updated via Eq. (9), they neither specify the step size nor provide the pseudocode for the ADP algorithm implemented. However, given the description in the article it is plausible that the algorithm used is similar to the Value iteration using a post-decision state variable described by Powell [3, p. 391]. For this reason, in Table 2 we have chosen to list the algorithm as AVI and the model update algorithm as N/A due to that a specific algorithm, other than Eq. (9), is used to update the value estimate of the post-decision state variable. In testing, the authors utilized a representative scenario of managing six aircraft types to move passengers and cargo between the United States and Saudi Arabia, where the total requirements were four times the total capacity of all the aircraft. To integrate the risk of breakdowns, the scenario assumed a failure probability of 20% on the C-141B freighter with an associated five day repair period, and that all other aircraft could be repaired without delay. The policy generated via the API algorithm was compared to four other policies, which in order of the amount of information considered are: •a rule-based policy that examined one requirement and one aircraft at a time; •a myopic cost-based policy that looked at one requirement and a list of aircraft, knowable now and actionable now; •a myopic cost-based policy that looked at a list of requirements and a list of aircraft, knowable now and actionable now; and •a myopic cost-based policy that looked at a list of requirements and a list of aircraft, knowable now and actionable in the future. The results showed a maximum improvement in the objective function of 100% for the policy generated via the API algorithm as compared to the first policy (rule-based policy) and a minimum improvement of 10% compared to the fourth policy. The policy generated via Operations Research Perspectives 8 (2021) 100204 15 M. Rempel and J. Cai [51] Sztykgold A, Coppin G, Hudry O. Dynamic optimization of the strength ratio during a terrestrial conflict. In: 2007 IEEE international symposium on approximate dynamic programming and reinforcement learning, 2007. p. 241–6. [52] Wu T, Powell W, Whisman A. The optimizing-simulator: An illustration using the military airlift problem. ACM Trans Model Comput Simul 2009;19. [53] Ahner D, Parson C. Weapon tradeoff analysis using dynamic programming for a dynamic weapon target assignment problem within a simulation. In: Proceedings of the 2013 winter simulation conference: Simulation: Making decisions in a complex world. WSC ’13, IEEE Press; 2013, p. 2831–41. [54] Rettke A, Robbins M, Lunday B. Approximate dynamic programming for the dispatch of military medical evacuation assets. European J Oper Res 2016;254:824–39. [55] Davis M, Robbins M, Lunday B. Approximate dynamic programming for missile defence interceptor fire control. European J Oper Res 2017;259:873–86. [56] Laan C, Barros A, Boucherie R, Monsuur H. In: Monsuur H, Jansen J, Marchal F, editors. Security games with restricted strategies: An approximate dynamic programming approach. NL ARMS Netherlands annual review of military studies 2018: Coastal border control: From data and tasks to deployment and law enforcement, The Hague: T.M.C. Asser Press; 2018, p. 171–91. [57] Robbins M, Jenkins P, Bastian N, Lunday B. Approximate dynamic programming for the aeromedical evacuation dispatching problem: Value function approximation utilizing multiple level aggregation. Omega 2020;91:102020. [58] Summers D, Robbins M, Lunday B. An approximate dynamic programming approach for comparing firing policies in a networked air defence environment. Comput Oper Res 2020;117:104890. [59] Jenkins P, Robbins M, Lunday B. Approximate dynamic programming for the military aeromedical evacuation dispatching, preemption-rerouting, and redeployment problem. European J Oper Res 2021a;290:132–43. [60] Jenkins P, Robbins M, Lunday B. Approximate dynamic programming for military medical evacuation dispatching policies. INFORMS J Comput 2021b;33:2–26. [61] Coifman R, Maggioni M. Diffusion wavelets. Appl Comput Harmon Anal 2006;21:53–94. [62] Balakrishna P. Scalable approximate dynamic programming models with applications in air transportation (Ph.D. thesis), Fairfax, Virginia: George Mason University; 2009. [63] Southerland J. Using approximate dynamic programming to adapt a military force mix (Ph.D. thesis), Fairfax, Virginia: George Mason University; 2017. [64] Godfrey G, Powell W. An adaptive dynamic programming algorithm for dynamic fleet management, I: Single period travel times. Transp Sci 2002;36:21–39. [65] Bradshaw A. United States Air Force officer manpower planning problem via approximate dynamic programming (M.Sc. thesis), Ohio: Air Force Institute of Technology, Wright-Patterson Air Force Base; 2016. [66] West K. Approximate dynamic programming for the United States Air Force officer manpower planning program (M.Sc. thesis), Ohio: Air Force Institute of Technology, Wright-Patterson Air Force Base; 2017. [67] Situ J. An approximate dynamic programming appraoch to analyzing military personnel end-strength planning (Ph.D. thesis), Fairfax, Virginia: Systems Engineering and Operations Research, George Mason University; 2018. [68] Powell W, Ruszczyński A, Topaloglu H. Learning algorithms for separable approximations of discrete stochastic optimization problems. Math Oper Res 2004;29:814–36. [69] Salgado E. Using approximate dynamic programming to solve the stochastic demand military inventory routing problem with direct delivery (M.Sc. thesis), Ohio: Air Force Institute of Technology, Wright-Patterson Air Force Base; 2016. [70] Government of Canada. TERMIUM plus. 2021, https://www.btb.termiumplus.gc. ca/tpv2alpha/alpha-eng.html?lang=eng&i=1&index=alt&codom2nd_wet=1. [71] Brown G, Dell R, Newman A. Optimizing military capital planning. Interfaces 2004;34:415–25. [72] Rempel M, Young C. VIPOR: A visual analytics decision support tool for capital investment planning. Scientific Report DRDC-RDDC-2017-R129, Ottawa, Canada: Defence Research and Development Canada; 2017, https://cradpdf.drdc-rddc.gc. ca/PDFS/unc290/p805944_A1b.pdf. [73] Harrison K, Elsayed S, Garanovich I, Weir T, Galister M, Boswell S, et al. Portfolio optimization for defence applications. IEEE Access 2020;8:60152–78. [74] Gallo A. Understanding military doctrinal change during peacetime. (Ph.D. thesis), New York: Graduate School of Arts and Science, Columbia University; 2018. [75] Sacco W, Navin D, Fiedler K, Waddell I.I. R, Long W, Buckman Jr. R. Precise formulation and evidence-based application of resource-constrained triage. Acad Emerg Med 2005;12:759–70. [76] Saran C. AI struggles with data silos and executive misconceptions. 2019, ComputerWeekly. https://www.computerweekly.com/news/252464222/AI-struggleswith-data-silos-and-executive-misconceptions, (Accessed: 19 April 2021). [77] Bakhshi N, Khera A, Bilato A. Navigate data management challenges to enable AI initiatives, Deloitte. 2020, https://www2.deloitte.com/content/dam/Deloitte/ nl/Documents/strategy-analytics-and-ma/deloitte-nl-strategy-analytics-dailsense-whitepaper.pdf, (Accessed 27 April 2021). [78] Scott W. Why data silos are bad for business, Forbes. 2021, https: //www.forbes.com/sites/forbestechcouncil/2018/11/19/why-data-silos-arebad-for-business/?sh=2d0ff8765faf, (Accessed 23 April 2021). [79] Teeple N, Dean R. NORAD modernization. CDA Institute; 2020, https:// cdainstitute.ca/norad-modernization-report-three-jadc2-jado/.(Accessed 23 April 2021). [80] Walker W, Rahman S, Cave J. Adaptive policies, policy analysis, and policy-making. European J Oper Res 2001;128:282–9. [81] Stasko T, Gao H. Developing green fleet management strategies: Repair/retrofit/replacement decisions under environmental regulation. Transp Res A 2012;46:1216–26. [82] Abdul-Malak D, Kharoufeh J. Optimally replacing multiple systems in a shared environment. Probab Engrg Inform Sci 2018;32:179–206. [83] Sadeghpour H, Tavakoli A, Kazemi M, Pooya A. A novel approximate dynamic programming approach for constrained equipment replacement problems: A case study. Adv Prod Eng Manag 2019;14:355–66. [84] Fang J, Zhao L, Fransoo JC, Van Woensel T. Sourcing strategies in supply risk management: An approximate dynamic programming approach. Comput Oper Res 2013;40:1371–82. [85] Geng Y, Klabjan D. Approximate dynamic programming based approaches for green supply chain design. Technical Report, Industrial Engineering and Management Sciences, Northwestern University; 2014. [86] Ghanmi A. A stochastic model for military air-to-ground munitions demand forecasting. In: 2016 3rd international conference on logistics operations management (GOL), 2016. p. 1–8. [87] Nozhati S, Ellingwood BR, Chong EK. Stochastic optimal control methodologies in risk-informed community resilience planning. Struct Saf 2020;84:101920. [88] Karamanis D. Stochastic dynamic programming methods for the portfolio selection problem. (Ph.D. thesis), London: The London School of Economics and Political Science; 2013. [89] MacLeod M, Rempel M, Roi M. Decision support for optimal use of joint training funds in the Canadian Armed Forces. In: Evans G, Biles W, Bae K, editors. Analytics, operations, and strategic decision making in the public sector. Hershey, PA: IGI Global; 2019, p. 255–76. [90] Séguin R. PARSim, a simulation model of the Royal Canadian Air Force (RCAF) pilot occupation: An assessment of the pilot occupation sustainability under high student production and reduced flying rates. In: Vitoriano B, Parlier G, editors. Proceedings of the international conference on operations research and enterprise systems. Lisbon, Portugal: SciTePress; 2015, p. 51–62. [91] Hunter G, Chan J, Rempel M. Assessing the operational impact of infrastructure on Arctic operations. Scientific Report DRDC-RDDC-2021-R024, Ottawa, Canada: Defence Research and Development Canada; 2021, https://cradpdf.drdc-rddc.gc. ca/PDFS/unc356/p812844_A1b.pdf. [92] Shin K, Lee T. Emergency medical service resource allocation in a mass casualty incident by integrating patient prioritization and hospital selection problems. IISE Trans 2020;52:1141–55. [93] Sidoti D, Han X, Zhang L, Avvari GV, Ayala DFM, Mishra M, et al. Context-aware dynamic asset allocation for maritime interdiction operations. IEEE Trans Syst Man Cybern 2020;50:1055–73. [94] Rempel M, Cai J, MacLeod M. Military versus contracted airlift: an approach to finding the right mix. Scientific Letter DRDC-RDDC-2021-L086, Ottawa, Canada: Defence Research and Development Canada; 2021.