First-order sensitivity of the optimal value in a Markov decision model with respect to deviations in the transition probability function
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Kern, Patrick; Simroth, Axel; Zähle, Henryk Article — Published Version First-order sensitivity of the optimal value in a Markov decision model with respect to deviations in the transition probability function Mathematical Methods of Operations Research Provided in Cooperation with: Springer Nature Suggested Citation: Kern, Patrick; Simroth, Axel; Zähle, Henryk (2020) : First-order sensitivity of the optimal value in a Markov decision model with respect to deviations in the transition probability function, Mathematical Methods of Operations Research, ISSN 1432-5217, Springer, Berlin, Heidelberg, Vol. 92, Iss. 1, pp. 165-197, https://doi.org/10.1007/s00186-020-00706-w This Version is available at: https://hdl.handle.net/10419/288279 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
Mathematical Methods of Operations Research (2020) 92:165–197 https://doi.org/10.1007/s00186-020-00706-w ORIGINAL ARTICLE First-order sensitivity of the optimal value in a Markov decision model with respect to deviations in the transition probability function Patrick Kern1·Axel Simroth2·Henryk Zähle1 Received: 23 January 2019 / Revised: 2 September 2019 / Published online: 2 March 2020 © The Author(s) 2020 Abstract Markov decision models (MDM) used in practical applications are most often less complex than the underlying ‘true’ MDM. The reduction of model complexity is performed for several reasons. However, it is obviously of interest to know what kind of model reduction is reasonable (in regard to the optimal value) and what kind is not. In this article we propose a way how to address this question. We introduce a sort of derivative of the optimal value as a function of the transition probabilities, which can be used to measure the (first-order) sensitivity of the optimal value w.r.t. changes in the transition probabilities. ‘Differentiability’ is obtained for a fairly broad class of MDMs, and the ‘derivative’ is specified explicitly. Our theoretical findings are illustrated by means of optimization problems in inventory control and mathematical finance. Keywords Markov decision model ·Model reduction ·Transition probability function ·Optimal value ·Functional differentiability ·Financial optimization Electronic supplementary material The online version of this article (https://doi.org/10.1007/s00186020-00706-w) contains supplementary material, which is available to authorized users. BHenryk Zähle [email protected].de Patrick Kern [email protected].de Axel Simroth [email protected] 1Department of Mathematics, Saarland University, Saarbrücken, Germany 2Fraunhofer Institute for Transportation and Infrastructure Systems, Dresden, Germany 123
166 P. Kern et al. 1 Introduction Already in the 1990th, Müller (1997a) pointed out that the impact of the transition probabilities of a Markov decision process (MDP) on the optimal value of a corresponding Markov decision model (MDM) can not be ignored for practical issues. For instance, in most cases the transition probabilities are unknown and have to be estimated by statistical methods. Moreover in many applications the ‘true’ model is replaced by an approximate version of the ‘true’ model or by a variant which is simplified and thus less complex. The result is that in practical applications the optimal (strategy and thus the optimal) value is most often computed on the basis of transition probabilities that differ from the underlying true transition probabilities. Therefore the sensitivity of the optimal value w.r.t. deviations in the transition probabilities is obviously of interest. Müller (1997a) showed that under some structural assumptions the optimal value in a discrete-time MDM depends continuously on the transition probabilities, and he established bounds for the approximation error. In the course of this the distance between transition probabilities was measured by means of some suitable probability metrics. Even earlier, Kolonko (1983) obtained analogous bounds in a MDM in which the transition probabilities depend on a parameter. Here the distance between transition probabilities was measured by means of the distance between the respective parameters. Error bounds for the expected total reward of discrete-time Markov reward processes were also specified by Van Dijk (1988) and Van Dijk and Puterman (1988). In the latter reference the authors also discussed the case of discrete-time Markov decision processes with countable state and action spaces. In this article, we focus on the situation where the ‘true’ model is replaced by a less complex version (for a simple example, see Subsection 1.4.3 in the supplemental article Kern et al. (2020)). The reduction of model complexity in practical applications is common and performed for several reasons. Apart from computational aspects and the difficulty of considering all relevant factors, one major point is that statistical inference for certain transition probabilities can be costly in terms of both time and money. However, it is obviously of interest to know what kind of model reduction is reasonable and what kind is not. In the following we want to propose a way how to address the latter question. Our original motivation comes from the field of optimal logistics transportation planning, where ongoing projects like SYNCHRO-NET (https://www.synchronet.eu/) aim at stochastic decision models based on transition probabilities estimated from historical route information. Due to the lack of historical data for unlikely events, transition probabilities are often modeled in a simplified way. In fact, events with small probabilities are often ignored in the model. However, the impact of these events on the optimal value (here the minimal expected transportation costs) of the corresponding MDM may nevertheless be significant. The identification of unlikely but potentially cost sensitive events is therefore a major challenge. In logistics planning operations engineers have indeed become increasingly interested in comprehensibly quantifying the sensitivity of the optimal value w.r.t. the incorporation of unlikely events into the model. For background see, for instance, Holfeld and Simroth (2017) and Holfeld et al. (2018). The assessment of rare but risky events takes on greater importance also 123
First-order sensitivity of the optimal value in a MDM 167 in other areas of applications; see, for instance, Komljenovic et al. (2016), Yang et al. (2015) and references cited therein. By an incorporation of an unlikely event into the model we mean, for instance, that under performance of an action aat some time na previously impossible transition from one state xto another state ygets now assigned small but strictly positive probability ε. Mathematically this means that the transition probability Pn((x,a), ·) is replaced by (1−ε)Pn((x,a), •)+εQn((x,a), •)with Qn((x,a), •):= δy[•], where δyis the Dirac measure at y. More generally one could consider a change of the whole transition function (the family of all transition probabilities) Pto (1−ε)P+εQ with ε>0 small. For operations engineers it is here interesting to know how this change affects the optimal value, V0(P). If the effect is minor, then an incorporation can be seen as superfluous, at least from a pragmatic point of view. If on the other hand the effect is significant, then the engineer should consider the option to extend the model and to make an effort to get access to statistical data for the extended model. At this point it is worth mentioning that a change of the transition function from Pto (1−ε)P+εQwith ε>0 small can also have a different interpretation than an incorporation of an (unlikely) new event. It could also be associated with an incorporation of an (unlikely) divergence from the normal transition rules. See Sect. 4.5 for an example. In this article, we will introduce an approach for quantifying the effect of changing the transition function from Pto (1−ε)P+εQ, with ε>0 small, on the optimal value V0(P)of the MDM. In view of (1−ε)P+εQ=P+ε(Q−P), we feel that it is reasonable to quantify the effect by a sort of derivative of the value functional V0at Pevaluated at direction Q−P. To some extent the ‘derivative’ ˙ V0;P(Q−P) specifies the first-order sensitivity of V0(P)w.r.t. a change of Pas above. Take into account that V0(P+ε(Q−P)) −V0(P)≈ε·˙ V0;P(Q−P)for ε>0small.(1) To be able to compare the first-order sensitivity for (infinitely) many different Q,it is favourable to know that the approximation in (1)isuniformin Q∈Kfor preferably large sets Kof transition functions. Moreover, it is not always possible to specify the relevant Qexactly. For that reason it would be also good to have robustness (i.e. some sort of continuity) of ˙ V0;P(Q−P)in Q. These two things induced us to focus on a variant of tangential S-differentiability as introduced by Sebastião e Silva (1956) and Averbukh and Smolyanov (1967) (here Sis a family of sets Kof transition functions). In Section 3 we present a result on ‘S-differentiability’ of V0for the family Sof all relatively compact sets of admissible transition functions and a reasonably broad class of MDMs, where we measure the distance between transition functions by means of metrics based on probability metrics as in Müller (1997a). The ‘derivative’ ˙ V0;P(Q−P)of the optimal value functional V0at Pquantifies the effect of a change from Pto (1−ε)P+εQ, with ε>0 small, assuming that after the change the strategy π(tuple of the underlying decision rules) is chosen such that it optimizes the target value Vπ 0(P)(e.g. expected total costs or rewards) in πunder the new transition function P:= (1−ε)P+εQ. On the other hand, practitioners are also interested in quantifying the impact of a change of Pwhen the optimal strategy (under 123
168 P. Kern et al. P) is kept after the change. Such a quantification would somehow answers the question: How much different does a strategy derived in a simplified MDM perform in a more complex (more realistic) variant of the MDM? Since the ‘derivative’ ˙ Vπ 0;P(Q−P) of the functional Vπ 0under a fixed strategy πturns out to be a building stone for the derivative ˙ V0;P(Q−P)of the optimal value functional V0at P, our elaborations cover both situations anyway. For fixed strategy πwe obtain ‘S-differentiability’ of Vπ 0even for the broader family Sof all bounded sets of admissible transition functions. The ‘derivative’ which we propose to regard as a measure for the first-order sensitivity will formally be introduced in Definition 7. This definition is applicable to quite general finite time horizon MDMs and might look somewhat cumbersome at first glance. However, in the special case of a finite state space and finite action spaces, a situation one faces in many practical applications, the proposed ‘differentiability’ boils down to a rather intuitive concept. This will be explained in Section 1 of the supplemental article Kern et al. (2020) with a minimum of notation and terminology. In Section 1 of the supplemental article Kern et al. (2020) we will also reformulate a backward iteration scheme for the computation of the ‘derivative’ (which can be deduced from our main result, Theorem 1) in the discrete case, and we will discuss an example. In Section 2 we formally introduce quite general MDMs in the fashion of the standard monographs Bäuerle and Rieder (2011), Hernández-Lerma and Lasserre (1996), Hinderer (1970), Puterman (1994). Since it is important to have an elaborate notation in order to formulate our main result, we are very precise in Section 2. As a result, this section is a little longer compared to the respective sections in other articles on MDMs. In Section 3 we carefully introduce our notion of ‘differentiability’ and state our main result concerning the computation of the ‘derivative’ of the value functional. In Section 4 we will apply the results of Section 3 to assess the impact of one or more than one unlikely but substantial shock in the dynamics of an asset on the solution of a terminal wealth problem in a (simple) financial market model free of shocks. This example somehow motivates the general set-up chosen in Sections 2–3. All results of this article are proven in Sections 3–5 of the supplemental article Kern et al. (2020). For the convenience of the reader we recall in Section 6 of the supplemental article Kern et al. (2020) a result on the existence of optimal strategies in general MDMs. Section 7 of the supplemental article Kern et al. (2020) contains an auxiliary topological result. 2 Formal definition of Markov decision model Let Ebe a non-empty set equipped with a σ-algebra E, referred to as state space.Let N∈Nbe a fixed finite time horizon (or planning horizon) in discrete time. For each point of time n=0,...,N−1 and each state x∈E,letAn(x)be a non-empty set. The elements of An(x)will be seen as the admissible actions (or controls) at time n in state x. For each n=0,...,N−1, let An:= x∈E An(x)and Dn:= (x,a)∈E×An:a∈An(x). 123
First-order sensitivity of the optimal value in a MDM 169 The elements of Ancan be seen as the actions that may basically be selected at time nwhereas the elements of Dnare the possible state-action combinations at time n. For our subsequent analysis, we equip Anwith a σ-algebra An, and let Dn:= (E⊗An)∩Dnbe the trace of the product σ-algebra E⊗Anin Dn. Recall that a map Pn:Dn×E→[0,1]is said to be a probability kernel (or Markov kernel) from (Dn,Dn)to (E,E)if Pn(·,B)is a (Dn,B([0,1]))-measurable map for any B∈E, and Pn((x,a), •)∈M1(E)for any (x,a)∈Dn.HereM1(E)is the set of all probability measures on (E,E). 2.1 Markov decision process In this subsection, we will give a formal definition of an E-valued (discrete-time) Markov decision process (MDP) associated with a given initial state, a given transition function and a given strategy. By definition a (Markov decision) transition (probability) function is an N-tuple P=(P0,...,PN−1) whose n-th entry Pnis a probability kernel from (Dn,Dn)to (E,E). In this context Pnwill be referred to as one-step transition (probability) kernel at time n (or from time n to n +1) and the probability measure Pn((x,a), •)is referred to as one-step transition probability at time n (or from time n to n +1) given state x and action a. We denote by Pthe set of all transition functions. We will assume that the actions are performed by a so-called N-stage strategy (or N-stage policy). An (N -stage) strategy is an N-tuple π=(f0,..., fN−1) of decision rules at times n=0,...,N−1, where a decision rule at time n is an (E,An)-measurable map fn:E→Ansatisfying fn(x)∈An(x)for all x∈E.Note that a decision rule at time nis (deterministic and) ‘Markovian’ since it only depends on the current state and is independent of previous states and actions. We denote by Fnthe set of all decision rules at time n, and assume that Fnis non-empty. Hence a strategy is an element of the set F0×···×FN−1, and this set can be seen as the set of all strategies. Moreover, we fix for any n=0,...,N−1someFn⊆Fnwhich can be seen as the set of all admissible decision rules at time n. In particular, the set Π:= F0×···× FN−1can be seen as the set of all admissible strategies. For any transition function P=(Pn)N−1 n=0∈P,strategyπ=(fn)N−1 n=0∈Π, and time point n∈{0,...,N−1}, we can derive from Pna probability kernel Pπ nfrom (E,E)to (E,E)through Pπ n(x,B):= Pn(x,fn(x)), B,x∈E,B∈E.(2) The probability measure Pπ n(x,•)can be seen as the one-step transition probability at time ngiven state xwhen the transitions and actions are governed by Pand π, respectively. 123
170 P. Kern et al. Now, consider the measurable space (Ω, F):= (EN+1,E⊗(N+1)). For any x0∈E,P=(Pn)N−1 n=0∈P, and π∈Πdefine the probability measure Px0,P;π:= δx0⊗Pπ 0⊗···⊗ Pπ N−1(3) on (Ω, F), where x0should be seen as the initial state of the MDP to be constructed. The right-hand side of (3) is the usual product of the probability measure δx0and the kernels Pπ 0,...,Pπ N−1; for details see display (16) in Section 2 of the supplemental article Kern et al. (2020). Moreover let X=(X0,...,XN)be the identity on Ω, i.e. Xn(x0,...,xN):= xn,(x0,...,xN)∈EN+1,n=0,...,N.(4) Note that, for any x0∈E,P=(Pn)N−1 n=0∈P, and π∈Π,themapXcan be regarded as an (EN+1,E⊗(N+1))-valued random variable on the probability space (Ω, F,Px0,P;π)with distribution δx0⊗Pπ 0⊗···⊗ Pπ N−1. It follows from Lemma 1 in the supplemental article Kern et al. (2020) that for any x0,x0,x1,...,xn∈E,P=(Pn)N−1 n=0∈P,π=(fn)N−1 n=0∈Π, and n=1,...,N−1 (i) Px0,P;π[X0∈•]=δx0[•], (ii) Px0,P;π[X1∈•X0=x0]=P0(x0,f0(x0)), •, (iii) Px0,P;π[Xn+1∈•(X0,X1,...,Xn)=(x0,x1,...,xn)] =Pn(xn,fn(xn)), •, (iv) Px0,P;π[Xn+1∈•Xn=xn]=Pn(xn,fn(xn)), •. The formulation of (ii)–(iv) is somewhat sloppy, because in general a (regular version of the) factorized conditional distribution of Xgiven Yunder Px0,P;π(evaluated at afixedsetB∈E)isonlyPx0,P;π Y-a.s. unique. So assertion (iv) in fact means that the probability kernel Pn(( ·,fn(·)), •)provides a (regular version of the) factorized conditional distribution of Xn+1given Xnunder Px0,P;π, and analogously for (ii) and (iii). Note that the factorized conditional distribution in part (ii) is constant w.r.t. x0∈E. Assertions (iii) and (iv) together imply that the temporal evolution of Xnis Markovian. This justifies the following terminology. Definition 1 (MDP) Under law Px0,P;πthe random variable X=(X0, ...,XN)is called (discrete-time) Markov decision process (MDP) associated with initial state x0∈E, transition function P∈P, and strategy π∈Π. 2.2 Markov decision model and value function Maintain the notation and terminology introduced in Sect. 2.1. In this subsection, we will first define a (discrete-time) Markov decision model (MDM) and introduce subsequently the corresponding value function. The latter will be derived from a reward 123
First-order sensitivity of the optimal value in a MDM 171 maximization problem. Fix P∈P, and let for each point of time n=0,...,N−1 rn:Dn−→ R be a (Dn,B(R))-measurable map, referred to as one-stage reward function.Here rn(x,a)specifies the one-stage reward when action ais taken at time nin state x.Let rN:E−→ R be an (E,B(R))-measurable map, referred to as terminal reward function.Thevalue rN(x)specifies the reward of being in state xat terminal time N. Denote by Athe family of all sets An(x),n=0,...,N−1, x∈E, and set r:= (rn)N n=0. Moreover let Xbe defined as in (4) and recall Definition 1. Then we define our MDM as follows. Definition 2 (MDM) The quintuple (X,A,P,Π,r)is called (discrete-time) Markov decision model (MDM) associated with the family of action spaces A, transition function P∈P, set of admissible strategies Π, and reward functions r. In the sequel we will always assume that a MDM (X,A,P,Π,r)satisfies the following Assumption (A). In Sect. 3.1 we will discuss some conditions on the MDM under which Assumption (A) holds. We will use Ex0,P;π n,xnto denote the expectation w.r.t. the factorized conditional distribution Px0,P;π[•Xn=xn].Forn=0, we clearly have Px0,P;π[•X0=x0]=Px0,P;π[•]for every x0∈E; see Lemma 1 in the supplemental article Kern et al. (2020). In what follows we use the convention that the sum over the empty set is zero. Assumption (A) supπ=(fn)N−1 n=0∈ΠEx0,P;π n,xn[N−1 k=n|rk(Xk,fk(Xk))|+|rN(XN)|] < ∞for any xn∈Eand n=0,...,N. Under Assumption (A) we may define in a MDM (X,A,P,Π,r)for any π= (fn)N−1 n=0∈Πand n=0,...,NamapVP;π n:E→Rthrough VP;π n(xn):= Ex0,P;π n,xnN−1 k=n rk(Xk,fk(Xk)) +rN(XN).(5) As a factorized conditional expectation this map is (E,B(R))-measurable (for any π∈Πand n=0,...,N). Note that for n=1,...,Nthe right-hand side of (5) does not depend on x0; see Lemma 2 in the supplemental article Kern et al. (2020). Therefore the map VP;π n(·)need not be equipped with an index x0. The value VP;π n(xn)specifies the expected total reward from time n to N of X under Px0,P;πwhen strategy πis used and Xis in state xnat time n. It is natural to ask for those strategies π∈Πfor which the expected total reward from time 0 to Nis maximal for all initial states x0∈E. This results in the following optimization problem: 123
172 P. Kern et al. VP;π 0(x0)−→ max (in π∈Π)!(6) If a solution πPto the optimization problem (6) (in the sense of Definition 4ahead) exists, then the corresponding maximal expected total reward is given by the so-called value function (at time 0). Definition 3 (Value function)ForaMDM(X,A,P,Π,r)the value function at time n∈{0,...,N}is the map VP n:E→Rdefined by VP n(xn):= sup π∈Π VP;π n(xn). (7) Note that the value function VP nis well defined due to Assumption (A) but not necessarily (E,B(R))-measurable. The measurability holds true, for example, if the sets Fn,...,FN−1are at most countable or if conditions (a)–(c) of Theorem 2 in the supplemental article Kern et al. 2020) are satisfied; see also Remark 1(i) in the supplemental article Kern et al. (2020). Definition 4 (Optimal strategy)InaMDM(X,A,P,Π,r)astrategyπP∈Πis called optimal w.r.t. Pif VP;πP 0(x0)=VP 0(x0)for all x0∈E.(8) In this case VP;πP 0(x0)is called optimal value (function), and we denote by Π(P)the set of all optimal strategies w.r.t. P. Further, for any given δ>0, a strategy πP;δ∈Π is called δ-optimal w.r.t. PinaMDM(X,A,P,Π,r)if VP 0(x0)−δ≤VP;πP;δ 0(x0)for all x0∈E,(9) and we denote by Π(P;δ) the set of all δ-optimal strategies w.r.t. P. Note that condition (8) requires that πP∈Πis an optimal strategy for all possible initial states x0∈E. Though, in some situations it might be sufficient to ensure that πP∈Πis an optimal strategy only for some fixed initial state x0. For a brief discussion of the existence and computation of optimal strategies, see Section 6 of the supplemental article Kern et al. (2020). Remark 1 (i) In practice, the choice of an action can possibly be based on historical observations of states and actions. In particular one could relinquish the Markov property of the decision rules and allow them to depend also on previous states and actions. Then one might hope that the corresponding (deterministic) history-dependent strategies improve the optimal value of a MDM (X,A,P,Π,r). However, it is known that the optimal value of a MDM (X,A,P,Π,r)can not be enhanced by considering history-dependent strategies; see, e.g., Theorem 18.4 in Hinderer (1970) or Theorem 4.5.1 in Puterman (1994). (ii) Instead of considering the reward maximization problem (6) one could as well be interested in minimizing expected total costs over the time horizon N. In this case, 123
First-order sensitivity of the optimal value in a MDM 179 The following lemma, whose proof can be found in Subsection 3.2 of the supplemental article Kern et al. (2020), provides an equivalent characterization of ‘Hadamard differentiability’. Lemma 2 Let M⊆Mψ(E),φbe another gauge function, and V:Pψ→L be any map. Fix P∈Pψ. Then the following two assertions hold. (i) If Vis ‘Hadamard differentiable’ at Pw.r.t. (M,φ) with ‘Hadamard derivative’ ˙ VP, then we have for each triplet (Q,(Qm), (εm)) ∈Pψ×PN ψ×(0,1]Nwith dφ ∞,M(Qm,Q)→0and εm→0that lim m→∞ V(P+εm(Qm−P)) −V(P) εm −˙ VP(Q−P) L=0.(13) (ii) If there exists an (M,φ)-continuous map ˙ VP:PP;± ψ→L such that (13) holds for each triplet (Q,(Qm), (εm)) ∈Pψ×PN ψ×(0,1]Nwith dφ ∞,M(Qm,Q)→0 and εm→0, then Vis ‘Hadamard differentiable’ at Pw.r.t. (M,φ)with ‘Hadamard derivative’ ˙ VP. 3.5 ‘Differentiability’of the value functional Recall that A,Π, and rare fixed, and let VP;π nand VP nbe defined as in (5) and (7), respectively. Moreover let ψbe any gauge function and fix some Pψ⊆Pψbeing closed under mixtures. In view of Lemma 1(with P:= {P}), condition (a) of Theorem 1below ensures that Assumption (A) is satisfied for any P∈Pψ. Then for any xn∈E,π∈Π, and n=0,...,Nwe may define under condition (a) of Theorem 1functionals Vxn;π n:Pψ→Rand Vxn n:Pψ→Rby Vxn;π n(P):= VP;π n(xn)and Vxn n(P):= VP n(xn), (14) respectively. Note that Vxn n(P)specifies the maximal value for the expected total reward in the MDM (given state xnat time n) when the underlying transition function is P. By analogy with the name ‘value function’ we refer to Vxn nas value functional given state xnat time n. Part (ii) of Theorem 1provides (under some assumptions) the ‘Hadamard derivative’ of the value functional Vxn nin the sense of Definition 8. Conditions (b) and (c) of Theorem 1involve the so-called Minkowski (or gauge) functional ρM:Mψ(E)→R≥0(see, e.g., Rudin (1991, p.25)) defined by ρM(h):= inf λ∈R>0:h/λ ∈M,(15) where we use the convention inf ∅:=∞,Mis any subset of Mψ(E), and we set R>0:= (0,∞). We note that Müller (1997a) also used the Minkowski functional to formulate his assumptions. 123
180 P. Kern et al. Example 6 For the sets M(and the corresponding gauge functions ψ) from Examples 1–5we have ρMTV (h)=sp(h),ρMKolm (h)=V(h),ρMBL (h)=hBL,ρMKant (h)= hLip, and ρMH¨ol,α(h)=hH¨ol,α, where as before MTV and MKolm are used to denote the maximal generator of dTV and dKolm, respectively. The latter three equations are trivial, for the former two equations see Müller (1997a, p.880). Recall from Definition 4that for given P∈Pψand δ>0thesetsΠ(P;δ) and Π(P)consist of all δ-optimal strategies w.r.t. Pand of all optimal strategies w.r.t. P, respectively. Generators Mof dMwere introduced subsequent to (10). Theorem 1 (‘Differentiability’ of Vxn;π nand Vxn n)Let M⊆Mψ(E)and Mbe any generator of dM.FixP=(Pn)N−1 n=0∈Pψ, and assume that the following three conditions hold. (a) ψis a bounding function for the MDM (X,A,Q,Π,r)for any Q∈Pψ. (b) supπ∈ΠρM(VP;π n)<∞for any n =1,...,N. (c) ρM(ψ) < ∞. Then the following two assertions hold. (i) For any xn∈E, π=(fn)N−1 n=0∈Π,n=0,...,N, the map Vxn;π n:Pψ→R defined by (14) is ‘Fréchet differentiable’ at Pw.r.t. (M,ψ)with ‘Fréchet derivative’ ˙ Vxn;π n;P:PP;± ψ→Rgiven by ˙ Vxn;π n;P(Q−P) := N−1 k=n+1 k−1 j=nE ···E rk(yk,fk(yk)) Pk−1(yk−1,fk−1(yk−1)), dyk ···(Qj−Pj)(yj,fj(yj)), dyj+1···Pn(xn,fn(xn)), dyn+1 + N−1 j=nE ···E rN(yN)PN−1(yN−1,fN−1(yN−1)), dyN ···(Qj−Pj)(yj,fj(yj)), dyj+1···Pn(xn,fn(xn)), dyn+1.(16) (ii) For any xn∈E and n =0,...,N, the map Vxn n:Pψ→Rdefined by (14)is ‘Hadamard differentiable’ at Pw.r.t. (M,ψ)with ‘Hadamard derivative’ ˙ Vxn n;P: PP;± ψ→Rgiven by ˙ Vxn n;P(Q−P):= lim δ0sup π∈Π(P;δ) ˙ Vxn;π n;P(Q−P). (17) If the set of optimal strategies Π(P)is non-empty, then the ‘Hadamard derivative’ admits the representation ˙ Vxn n;P(Q−P)=sup π∈Π(P) ˙ Vxn;π n;P(Q−P). (18) 123
First-order sensitivity of the optimal value in a MDM 181 The proof of Theorem 1can be found in Section 4 of the supplemental article Kern et al. (2020). Note that the set Π(P;δ) shrinks as δdecreases. Therefore the right-hand side of (17) is well defined. The supremum in (18) ranges over all optimal strategies w.r.t. P. If, for example, the MDM (X,A,P,Π,r)satisfies conditions (a)– (c) of Theorem 2 in the supplemental article Kern et al. (2020), then by part (iii) of this theorem an optimal strategy can be found, i.e. Π(P)is non-empty. The existence of an optimal strategy is also ensured if the sets F0,...,FN−1are finite (a situation one often faces in applications). In the latter case the ‘Hadamard derivative’ ˙ Vxn n;P(Q−P)can easily be determined by computing the finitely many values ˙ Vxn;π n;P(Q−P),π∈Π(P), and taking their maximum. The discrete case will be discussed in more detail in Subsection 1.5 of the supplemental article Kern et al. (2020). If there exists a unique optimal strategy πP∈Πw.r.t. P, then Π(P)is nothing but the singleton {πP}, and in this case the ‘Hadamard derivative’ ˙ Vx0 0;Pof the optimal value (functional) Vx0 0at Pcoincides with ˙ Vx0;πP 0;P. Remark 3 (i) The ‘Fréchet differentiability’ in part (i) of Theorem 1holds even uniformly in π∈Π; see Theorem 1 in the supplemental article Kern et al. (2020)forthe precise meaning. (ii) We do not know if it is possible to replace ‘Hadamard differentiability’ by ‘Fréchet differentiability’ in part (ii) of Theorem 1. The following arguments rather cast doubt on this possibility. The proof of part (ii) is based on the decomposition of the value functional Vxn nin display (26) of the supplemental article Kern et al. (2020) and a suitable chain rule, where this decomposition involves the sup-functional Ψintroduced in display (27) of the supplemental article Kern et al. (2020). However, Corollary 1 in Cox and Nadler (1971) (see also Proposition 4.6.5 in Schirotzek 2007) shows that in normed vector spaces sup-functionals are in general not Fréchet differentiable. This could be an indication that ‘Fréchet differentiable’ of the value functional indeed fails. We can not make a reliable statement in this regard. (iii) Recall that ‘Hadamard (resp. Fréchet) differentiability’ w.r.t. (M,ψ) implies ‘Hadamard (resp. Fréchet) differentiability’ w.r.t. (M,φ)for any gauge function φ≤ ψ. However, for any such φ‘Hadamard (resp. Fréchet) differentiability’ w.r.t. (M,φ) is less meaningful than w.r.t. (M,ψ). Indeed, when using dφ ∞,Mwith φ≤ψinstead of dψ ∞,M,thesetsKfor whose elements the first-order sensitivities can be compared with each other with clear conscience are smaller and the ‘derivative’ is less robust. (iv) In the case where we are interested in minimizing expected total costs in the MDM (X,A,P,Π,r)(see Remark 1(ii)), we obtain under the assumptions (and with the same arguments as in the proof of part (ii)) of Theorem 1that the ‘Hadamard derivative’ of the corresponding value functional is given by (17)(resp.(18)) with “sup” replaced by “inf”. Remark 4 (i) Condition (a) of Theorem 1is in line with the existing literature. In fact, similar conditions as in Definition 5(with P:= { Q}) have been imposed many times before; see, for instance, Bäuerle and Rieder (2011, Definition 2.4.1), Müller (1997a, Definition 2.4), Puterman (1994, p.231 ff), and Wessels (1977). 123
182 P. Kern et al. (ii) In some situations, condition (a) implies condition (b) in Theorem 1.Thisis the case, for instance, in the following four settings (the involved sets Mand metrics were introduced in Examples 1–5). (1) M:= MTV and ψ:≡ 1. (2) M:= MKolm and ψ:≡ 1, as well as for n=1,...,N−1 –RVP;π n+1(y)Pn(( ·,fn(·)), dy),π=(fn)N−1 n=0∈Π, are increasing, –rn(·,fn(·)),π=(fn)N−1 n=0∈Π, and rN(·)are increasing. (3) M:= MBL and ψ:≡ 1, as well as for n=1,...,N−1 –sup π=(fn)N−1 n=0∈Πsupx=ydBL(Pn((x,fn(x)), •), Pn((y,fn(y)), •))/ dE(x,y)<∞, –sup π=(fn)N−1 n=0∈Πrn(·,fn(·))Lip <∞and rNLip <∞. (4) M:= MH¨ol,α and ψ(x):= 1+dE(x,x)αfor some x∈Eand α∈(0,1]. (recall that MH¨ol,α =MKant for α=1), as well as for n=1,...,N−1 –sup π=(fn)N−1 n=0∈Πsupx=ydH¨ol,α(Pn((x,fn(x)), •), Pn((y,fn(y)), •))/ dE(x,y)α<∞, –sup π=(fn)N−1 n=0∈Πrn(·,fn(·))H¨ol,α <∞and rNH¨ol,α <∞ The proof of (a)⇒(b) relies in setting 1) on Lemma 1(with P:= {P}) and in settings 2)–4) on Lemma 1(with P:= {P}) along with Proposition 1 of the supplemental article Kern et al. (2020). The conditions in setting 2) are similar to those in parts (ii)– (iv) of Theorem 2.4.14 in Bäuerle and Rieder (2011), and the conditions in settings 3) and 4) are motivated by the statements in Hinderer (2005, p.11f). (iii) In many situations, condition (c) of Theorem 1holds trivially. This is the case, for instance, if M∈{MTV,MKolm,MBL}and ψ:≡ 1, or if M:= MH¨ol,α and ψ(x):= 1+dE(x,x)αfor some fixed x∈Eand α∈(0,1]. (iv) The conditions (b) and (c) of Theorem 1can also be verified directly in some cases; see, for instance, the proof of Lemma 7 in Subsection 5.3.1 of the supplemental article Kern et al. (2020). In applications it is not necessarily easy to specify the set Π(P)of all optimal strategies w.r.t. P. While in most cases an optimal strategy can be found with little effort (one can use the Bellman equation; see part (i) of Theorem 2 in Section 6 of the supplemental article Kern et al. 2020), it is typically more involved to specify all optimal strategies or to show that the optimal strategy is unique. The following remark may help in some situations; for an application see Sect. 4.4. Remark 5 In some situations it turns out that for every P∈Pψthe solution of the optimization problem (6) does not change if Πis replaced by a subset Π⊆Π(being independent of P). Then in the definition (7) of the value function (at time 0) the set Πcan be replaced by the subset Π, and it follows (under the assumptions of Theorem 1) that in the representation (18) of the ‘Hadamard derivative’ ˙ Vx0 0;Pof Vx0 0at Pthe set Π(P)can be replaced by the set Π(P)of all optimal strategies w.r.t. Pfrom 123
First-order sensitivity of the optimal value in a MDM 183 the subset Π. Of course, in this case it suffices to ensure that conditions (a)–(b) of Theorem 1are satisfied for the subset Πinstead of Π. 3.6 Two alternative representations of ˙ Vxn; n;P In this subsection we present two alternative representations (see (19) and (20)) of the ‘Fréchet derivative’ ˙ Vxn;π n;Pin (16). The representation (19) will be beneficial for the proof of Theorem 1(see Lemma 3 in Subsection 4.1 of the supplemental article Kern et al. 2020) and the representation (20) will be used to derive the ‘Hadamard derivative’ of the optimal value of the terminal wealth problem in (28) below (see the proof of Theorem 3in Subsection 5.3 of the supplemental article Kern et al. 2020). Remark 6 (Representation I) By rearranging the sums in (16), we obtain under the assumptions of Theorem 1that for every fixed P=(Pn)N−1 n=0∈Pψthe ‘Fréchet derivative’ ˙ Vxn;π n;Pof Vxn;π nat Pcan be represented as ˙ Vxn;π n;P(Q−P)= N−1 k=nE ···EE VP;π k+1(yk+1)(Qk−Pk)(yk,fk(yk)), dyk+1 Pk−1(yk−1,fk−1(yk−1)), dyk···Pn(xn,fn(xn)), dyn+1(19) for every xn∈E,Q=(Qn)N−1 n=0∈Pψ,π=(fn)N−1 n=0∈Π, and n=0,...,N. Remark 7 (Representation II) For every fixed P=(Pn)N−1 n=0∈Pψ, and under the assumptions of Theorem 1, the ‘Fréchet derivative’ ˙ Vxn;π n;Pof Vxn;π nat Padmits the representation ˙ Vxn;π n;P(Q−P)=˙ VP,Q;π n(xn)(20) for every xn∈E,Q=(Qn)N−1 n=0∈Pψ,π=(fn)N−1 n=0∈Π, and n=0,...,N, where (˙ VP,Q;π k)N k=0is the solution of the following backward iteration scheme ˙ VP,Q;π N(·):= 0, ˙ VP,Q;π k(·):= E ˙ VP,Q;π k+1(y)Pk(·,fk(·)), dy +E VP;π k+1(y)(Qk−Pk)(·,fk(·)), dy,k=0,...,N−1. (21) Indeed, it is easily seen that ˙ VP,Q;π n(xn)coincides with the right-hand side of (19). Note that it can be verified iteratively by means of condition (a) of Theorem 1and Lemma 1(with P:= { Q}) that ˙ VP,Q;π n(·)∈Mψ(E)for every Q∈Pψ,π∈Π, and n=0,...,N. In particular, this implies that the integrals on the right-hand side 123
184 P. Kern et al. of (21) exist and are finite. Also note that the iteration scheme (21) involves the family (VP;π k)N k=1which itself can be seen as the solution of a backward iteration scheme: VP;π N(·):= rN(·), VP;π k(·):= rk(·,fk(·)) +E VP;π k+1(y)Pk(·,fk(·)), dy,k=1,...,N−1; see Proposition 1 of the supplemental article Kern et al. (2020). 4 Application to a terminal wealth optimization problem in mathematical finance In this section we will apply the theory of Sections 2–3 to a particular optimization problem in mathematical finance. At first, we introduce in Sect. 4.1 the basic financial market model and formulate subsequently the terminal wealth problem as a classical optimization problem in mathematical finance. The market model is in line with standard literature as Bäuerle and Rieder (2011, Chapter 4) or (Föllmer and Schied 2011, Chapter 5). To keep the presentation as clear as possible we restrict ourselves to a simple variant of the market model (only one risky asset). In Sect. 4.2 we will see that the market model can be embedded into the MDM of Sect. 2. It turns out that the existence (and computation) of an optimal (trading) strategy can be obtained by solving iteratively None-stage investment problems; see Sect. 4.3. In Sect. 4.4 we will specify the ‘Hadamard derivative’ of the optimal value functional of the terminal wealth problem, and Sect. 4.5 provides some numerical examples for the ‘Hadamard derivative’. 4.1 Basic financial market model, and the target Consider an N-period financial market consisting of one riskless bond B=(B0, ...,BN)and one risky asset S=(S0,...,SN). Further assume that the value of the bond evolves deterministically according to B0=1,Bn+1=rn+1Bn,n=0,...,N−1 for some fixed constants r1,...,rN∈R≥1, and that the value of the asset evolves stochastically according to S0>0,Sn+1=Rn+1Sn,n=0,...,N−1 for some independent R≥0-valued random variables R1,...,RNon some probability space (Ω, F,P)with (known) distributions m1,...,mN, respectively. Throughout Section 4 we will assume that the financial market satisfies the following Assumption (FM), where α∈(0,1)is fixed and chosen as in (24) below. 123
First-order sensitivity of the optimal value in a MDM 185 In Examples 7and 8we will discuss specific financial market models which satisfy Assumption (FM). Assumption (FM) The following three assertions hold for any n=0,...,N−1. (a) R≥0yαmn+1(dy)<∞. (b) Rn+1>0P-a.s. (c) P[Rn+1= rn+1]=1. Note that for any n=0,...,N−1thevaluern+1(resp. Rn+1) corresponds to the relative price change Bn+1/Bn(resp. Sn+1/Sn) of the bond (resp. asset) between time nand n+1. Let F0be the trivial σ-algebra, and set Fn:= σ(S0,...,Sn)= σ(R1,...,Rn)for any n=1,...,N. Now, an agent invests a given amount of capital x0∈R≥0in the bond and the asset according to some self-financing trading strategy. By trading strategy we mean an (Fn)-adapted R2 ≥0-valued stochastic process ϕ=(ϕ0 n,ϕ n)N−1 n=0, where ϕ0 n(resp. ϕn) specifies the amount of capital that is invested in the bond (resp. asset) during the time interval [n,n+1). Here we require that both ϕ0 nand ϕnare nonnegative for any n, which means that taking loans and short sellings of the asset are excluded. The corresponding portfolio process Xϕ=(Xϕ 0,...,Xϕ N)associated with ϕ=(ϕ0 n,ϕ n)N−1 n=0is given by Xϕ 0:= ϕ0 0+ϕ0and Xϕ n+1:= ϕ0 nrn+1+ϕnRn+1,n=0,...,N−1. A trading strategy ϕ=(ϕ0 n,ϕ n)N−1 n=0is said to be self-financing w.r.t. the initial capital x0if x0=ϕ0 0+ϕ0and Xϕ n=ϕ0 n+ϕnfor all n=1,...,N. It is easily seen that for any self-financing trading strategy ϕ=(ϕ0 n,ϕ n)N−1 n=0w.r.t. x0the corresponding portfolio process admits the representation Xϕ 0=x0and Xϕ n+1=rn+1Xϕ n+ϕn(Rn+1−rn+1)for n=0,...,N−1. (22) Note that Xϕ n−ϕncorresponds to the amount of capital which is invested in the bond between time nand n+1. Also note that it can be verified easily by means of Remark 3.1.6 in Bäuerle and Rieder (2011) that under condition (c) of Assumption (FM) the financial market introduced above is free of arbitrage opportunities. In view of (22), we may and do identify a self-financing trading strategy w.r.t. x0with an (Fn)-adapted R≥0-valued stochastic process ϕ=(ϕn)N−1 n=0satisfying ϕ0∈[0,x0] and ϕn∈[0,Xϕ n]for all n=1,...,N−1. We restrict ourselves to Markovian selffinancing trading strategies ϕ=(ϕn)N−1 n=0w.r.t. x0which means that ϕnonly depends on nand Xϕ n. To put it another way, we assume that for any n=0,...,N−1 there exists some Borel measurable map fn:R≥0→R≥0such that ϕn=fn(Xϕ n). Then, in particular, Xϕis an R≥0-valued (Fn)-Markov process whose one-step transition probability at time n∈{0,...,N−1}given state x∈R≥0and strategy ϕ=(ϕn)N−1 n=0 (resp. π=(fn)N−1 n=0) is given by mn+1◦η−1 n,(x,fn(x)) with ηn,(x,fn(x))(y):= rn+1x+fn(x)(y−rn+1), y∈R≥0.(23) 123
186 P. Kern et al. The agent’s aim is to find a self-financing trading strategy ϕ=(ϕn)N−1 n=0(resp. π=(fn)N−1 n=0) w.r.t. x0for which her expected utility of the discounted terminal wealth is maximized. We assume that the agent is risk averse and that her attitude towards risk is set via the power utility function uα:R≥0→R≥0defined by uα(y):= yα(24) for some fixed α∈(0,1)(as in Assumption (FM)). The coefficient αdetermines the degree of risk aversion of the agent: the smaller the coefficient α, the greater her risk aversion. Hence the agent is interested in those self-financing trading strategies ϕ=(ϕn)N−1 n=0(resp. π=(fn)N−1 n=0) w.r.t. x0for which the expectation of uα(Xϕ N/BN) under Pis maximized. In the following subsections we will assume for notational simplicity that r1,...,rN are fixed and that m1,...,mNare a sort of model parameters. In this case the factor 1/BNin uα(Xϕ N/BN)in display (25) is superfluous; it indeed does not influence the maximization problem or any ‘derivative’ of the optimal value. On the other hand, if also the (Dirac-) distributions of r1,...,rNwould be allowed to be variable, then this factor could matter for the derivative of the optimal value w.r.t. changes in the (deterministic) dynamics of BN. 4.2 Embedding into MDM, and optimal trading strategies The setting introduced in Sect. 4.1 can be embedded into the setting of Sections 2–3 as follows. Let r1,...,rN∈R≥1be a priori fixed constants. Let (E,E):= (R≥0,B(R≥0)) and An(x):= [0,x]for any x∈R≥0and n=0,...,N−1. Then An=R≥0and Dn=D:= {(x,a)∈R2 ≥0:a∈[0,x]}.Let An:= B(R≥0).In particular, Dn=B(R2 ≥0)∩Dand the set Fnof all decision rules at time nconsists of all those Borel measurable functions fn:R≥0→R≥0which satisfy fn(x)∈[0,x] for all x∈R≥0(in particular Fnis independent of n). For any n=0,...,N−1, let the set Fnof all admissible decision rules at time nbe equal to Fn. Let as before Π:= F0×···× FN−1. Moreover let rn:≡ 0 for any n=0,...,N−1, and rN(x):= uα(x/BN), x∈R≥0.(25) Consider the gauge function ψ:R≥0→R≥1defined by ψ(x):= 1+uα(x). (26) Let Pψbe the set of all transition functions P=(Pn)N−1 n=0∈Pconsisting of transition kernels of the shape Pn(x,a), •:= mn+1◦η−1 n,(x,a)[•],(x,a)∈Dn,n=0,...,N−1 (27) 123
First-order sensitivity of the optimal value in a MDM 187 for some mn+1∈Mα 1(R≥0), where Mα 1(R≥0)is the set of all μ∈M1(R≥0)satisfying R≥0uαdμ<∞, and the map ηn,(x,a)is defined as in (23). In particular, Pψ⊆Pψ (with Pψdefined as in Sect. 3.3), and (1−ε)P+εQ∈Pψfor all P,Q∈Pψand ε∈(0,1)(i.e. Pψis closed under mixtures). Moreover it can be verified easily that ψ given by (26) is a bounding function for the MDM (X,A,Q,Π,r)for any Q∈Pψ (see Lemma 7(i) of the supplemental article Kern et al. 2020). Note that Xplays the role of the portfolio process Xϕfrom Sect. 4.1. Also note that for some fixed x0∈R≥0, any self-financing trading strategy ϕ=(ϕn)N−1 n=0w.r.t. x0may be identified with some π=(fn)N−1 n=0∈Πvia ϕn=fn(Xϕ n). Then, for every fixed x0∈R≥0and P∈Pψthe terminal wealth problem introduced in the second to last paragraph of Sect. 4.1 reads as Ex0,P;π[rN(XN)]−→max (in π∈Π)!(28) AstrategyπP∈Πis called an optimal (self-financing) trading strategy w.r.t. P(and x0)if it solves the maximization problem (28). Remark 8 In the setting of Sect. 4.1 we restrict ourselves to Markovian self-financing trading strategies ϕ=(ϕn)N−1 n=0w.r.t. x0which may be identified with some π= (fn)N−1 n=0∈Πvia ϕn=fn(Xϕ n). Of course, one could also assume that the decision rules of a trading strategy πalso depend on past actions and past values of the portfolio process Xϕ. However, as already discussed in Remark 1(i), the corresponding historydependent trading strategies do not lead to an improved optimal value for the terminal wealth problem (28). 4.3 Computation of optimal trading strategies In this subsection we discuss the existence and computation of solutions to the terminal wealth problem (28), maintaining the notation of Sect. 4.2. We will adapt the arguments of Section 4.2 in Bäuerle and Rieder (2011). As before r1,...,rN∈R≥1are fixed constants. Basically the existence of an optimal trading strategy for the terminal wealth problem (28) can be ensured with the help of a suitable analogue of Theorem 4.2.2 in Bäuerle and Rieder (2011). In order to specify the optimal trading strategy explicitly one has to determine the local maximizers in the Bellman equation; see Theorem 2(i) in Section 6 of the supplemental article Kern et al. (2020). However this is not necessarily easy. On the other hand, part (ii) of Theorem 2ahead (a variant of Theorem 4.2.6 in Bäuerle and Rieder 2011) shows that, for our particular choice of the utility function (recall (24)), the optimal investment in the asset at time n∈{0,...,N−1} has a rather simple form insofar as it depends linearly on the wealth. The respective coefficient can be obtained by solving the one-stage optimization problem in (29) ahead. That is, instead of finding the optimal amount of capital (possibly depending on the wealth) to be invested in the asset, it suffices to find the optimal fraction of the wealth (being independent of the wealth itself) to be invested in the asset. 123
188 P. Kern et al. For the formulation of the one-stage optimization problem note that every transition function P∈Pψis generated through (27)bysome(m1,...,mN)∈Mα 1(R≥0)N. For every P∈Pψ,weuse(mP 1,...,mP N)to denote any such set of ‘parameters’. Now, consider for any P∈Pψand n=0,...,N−1 the optimization problem vP;γ n:= R≥0 uα1+γy rn+1 −1mP n+1(dy)−→ max (in γ∈[0,1])!(29) Note that 1 +γ(y/rn+1−1)lies in R≥0for any γ∈[0,1]and y∈R≥0, and that the integral on the left-hand side (exists and) is finite (this follows from displays (34)–(36) in Subsection 5.1 of the supplemental article Kern et al. 2020) and should be seen as the expectation of uα(1+γ(Rn+1/rn+1−1)) under P. The following lemma, whose proof can be found in Subsection 5.1 of the supplemental article Kern et al. (2020), shows in particular that vP n:= sup γ∈[0,1] vP;γ n is the maximal value of the optimization problem (29). Lemma 3 For any P∈Pψand n =0,...,N−1, there exists a unique solution γP n∈[0,1]to the optimization problem (29). Part (i) of the following Theorem 2involves the value function introduced in (7). In the present setting this function has a comparatively simple form: VP n(xn)=sup π∈Π Ex0,P;π n,xn[rN(XN)](30) for any xn∈R≥0,P∈Pψ, and n=0,...,N. Part (ii) involves the subset Πlin of Πwhich consists of all linear trading strategies, i.e. of all π∈Πof the form π=(fγ n)N−1 n=0for some γ=(γn)N−1 n=0∈[0,1]N, where fγ n(x):= γnx,x∈R≥0,n=0,...,N−1.(31) In part (i) and elsewhere we use the convention that the product over the empty set is 1. Theorem 2 (Optimal trading strategy) For any P∈Pψthe following two assertions hold. (i) The value function V P ngiven by (30)admits the representation VP n(xn)=vP nuα(xn/Bn) for any xn∈R≥0and n =0,...,N−1, where vP n:= N−1 k=nvP k. 123
First-order sensitivity of the optimal value in a MDM 195 Δ -1.2 -1 -0.8 -0.6 -0.4 -0.2 0 ν=0.01 ν=0.03 ν=0.035 ν=0.04 Δ -1.2 -1 -0.8 -0.6 -0.4 -0.2 0 ν=0.01 ν=0.03 ν=0.035 ν=0.04 00.10.2 0.3 0.4 0.5 0.6 0.7 0.8 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Δ -1.2 -1 -0.8 -0.6 -0.4 -0.2 0 ν=0.01 ν=0.03 ν=0.035 ν=0.04 Fig. 2 ‘Hadamard derivative’ ˙ Vx0 0;P(QΔ,τ −P)for α=0.5, μP=0.05, and σP=0.2 in dependence of the ‘jump’ Δand the drift νof the bond, showing N=4inthefirst,N=12 in the second, and N=52 in the third column Δ -1.2 -1 -0.8 -0.6 -0.4 -0.2 0 α=0.25 α=0.5 α=0.75 Δ -1.2 -1 -0.8 -0.6 -0.4 -0.2 0 α=0.25 α=0.5 α=0.75 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Δ -1.2 -1 -0.8 -0.6 -0.4 -0.2 0 α=0.25 α=0.5 α=0.75 Fig. 3 ‘Hadamard derivative’ ˙ Vx0 0;P(QΔ,τ −P)for N=12, μP=0.05, and σP=0.2 in dependence of the ‘jump’ Δand risk aversion parameter α, showing ν=0.02 in the first, ν=0.03 in the second, and ν=0.04 in the third column -12 -10 2 -8 0.8 4 -6 -4 0.6 6 -2 Δ 80.4 0 10 0.2 12 0 -12 -10 2 -8 0.8 4 -6 -4 0.6 6 -2 Δ 80.4 0 10 0.2 12 0 -12 -10 2 -8 0.8 4 -6 -4 0.6 6 -2 Δ 80.4 0 10 0.2 12 0 Fig. 4 ‘Hadamard derivative’ ˙ Vx0 0;P(QΔ,τ() −P)for N=12 in dependence on ∈{1,...,N}and Δ∈[0,0.8]showing α=0.25 and ν=0.02 (left), α=0.5andν=0.03 (middle), and α=0.75 and ν=0.04 (right) As appears from Fig. 3, for any Δ∈[0,0.8]the (negative) effect of incorporating a ‘jump’ Δin the dynamics S=(S0,...,SN)of an asset price is the smaller the higher the agent’s risk aversion, no matter what the drift ν∈{0.02,0.03,0.04}of the bond looks like. Take into account that the extent of this effect is influenced via (41)–(43) by the optimal fraction γP BSM to be invested into the asset which in turn depends on the risk aversion parameter α(see (36)). Finally, let us briefly touch on the case where more than one jump may appear. More precisely, instead of QΔ,τ (with τ∈{0,...,N−1}) consider the transition function QΔ,τ() (with 1 ≤≤N,τ() =(τ1,...,τ ),τ1,...,τ ∈{0,...,N−1} pairwise distinct) which is still generated by means of (40) but with the difference that at the different times τ1,...,τ the distribution mPis replaced by δΔ.Justas in the case =1, it turns out that it does not matter at which times τ1,...,τ exactly 123
196 P. Kern et al. these jumps occur. Figure 4shows the value of ˙ Vx0 0;P(QΔ,τ() −P)in dependence on and Δ. It seems that for any fixed Δ∈[0,0.8]the first-order sensitivity increases approximately linearly in . Supplement The supplement Kern et al. (2020) illustrates the setting of Sects. 2–3in the case of finite state space and finite action spaces, and contains the proofs of the results from Sects. 3–4. Moreover, supplemental definitions and results to Sect. 2are given and the existence of optimal strategies in general MDMs is discussed. Finally, a supplemental topological result is shown. Acknowledgements Open Access funding provided by Projekt DEAL. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. References Averbukh VI, Smolyanov OG (1967) The theory of differentiation in linear topological spaces. Russ Math Surv 22:201–258 Bäuerle N, Rieder U (2011) Markov decision processes with applications to finance. Springer, Berlin Bellini F, Klar B, Müller A, Rosazza Gianin E (2014) Generalized quantiles as risk measures. Insur Math Econ 54:41–48 Cox SH Jr, Nadler SB Jr (1971) Supremum norm differentiability. Ann Soc Math Pol 15:127–131 Dall’Aglio G (1956) Sugli estremi di momentidetle funzioni di ripartizione doppia. Ann Sc Norm Super Pisa 10:35–74 Dudley RM (2002) Real analysis and probability. Cambridge University Press, Cambridge Fernholz LT (1983) Von Mises calculus for statistical functionals. Springer, Berlin Föllmer H, Schied A (2011) Stochastic finance. An introduction in discrete time. de Gruyter, Berlin Gill RD (1989) Nonand semi-parametric maximum likelihood estimators and the von mises method—I. Scand J Stat 16:97–128 Hernández-Lerma O, Lasserre JB (1996) Discrete-time Markov control processes: basic optimality criteria. Springer, Berlin Hinderer K (1970) Foundations of non-stationary dynamic programming with discrete time parameter. Lecture notes in economics and mathematical systems 33. Springer, Berlin Hinderer K (2005) Lipschitz continuity of value functions in Markovian decision processes. Math Methods Oper Res 62:3–22 Holfeld D, Simroth A (2017) Learning from the past—risk profiler for intermodal route planning in SYNCHRO-NET. In: International conference on operations research (OR2017), Berlin Holfeld D, Simroth A, Li Y, Manerba D, Tadei R (2018) Risk analysis for synchro-modal freight transportation: the SYNCHRO-NET approach. In: 7th international workshop on freight transportation and logistics (Odysseus 2018), Cagliari Kantorovich LV, Rubinstein GS (1958) On a space of completely additive functions. Vestnik Leningrad University 13:52–59 123
First-order sensitivity of the optimal value in a MDM 197 Kern P, Simroth A, Zähle H (2020) Supplement to “First-order sensitivity of the optimal value in a Markov decision model with respect to deviations in the transition probability function” Kiesel R, Rühlicke R, Stahl G, Zheng J (2016) The Wasserstein metric and robustness in risk management. Risks 4:32 Kolonko M (1983) Bounds for the regret loss in dynamic programming under adaptive control. Z Oper Res 27:17–37 Komljenovic D, Gaha M, Abdul-Nour G, Langheit C, Bourgeois M (2016) Risks of extreme and rare events in asset management. Saf Sci 88:129–145 Krätschmer V, Zähle H (2017) Statistical inference for expectile-based risk measures. Scand J Stat 44:425– 454 Krätschmer V, Schied A, Zähle H (2012) Qualitative and infinitesimal robustness of tail-dependent statistical functionals. J Multivar Anal 103:35–47 Krätschmer V, Schied A, Zähle H (2017) Domains of weak continuity of statistical functionals with a view toward robust statistics. J Multivar Anal 158:1–19 Lemor JP, Gobet E, Warin X (2006) Rate of convergence of an empirical regression method for solving generalized backward stochastic differential equations. Bernoulli 12:889–916 Merton RC (1969) Lifetime portfolio selection under uncertainty: the continuous-time case. Rev Econ Stat 51:247–257 Müller A (1997) How does the value function of a Markov decision process depend on the transition probabilities ? Math Oper Res 22:872–885 Müller A (1997) Integral probability metrics and their generating classes of functions. Adv Appl Probab 29:429–443 Pham H (2009) Continuous-time stochastic control and optimization with financial applications. Springer, Berlin Puterman ML (1994) Markov decision processes: discrete stochastic dynamic programming. Wiley, New York Rachev ST (1991) Probability metrics and the stability of stochastic models. Wiley, New York Römisch W (2004) Delta method, infinite dimensional. Encyclopedia of statistical sciences. Wiley, New York Rudin W (1991) Functional analysis. McGraw-Hill, New York Schirotzek W (2007) Nonsmooth analysis. Springer, Berlin Sebastião e Silva J (1956) Le calcul différentiel et intégral dans les espaces localement convexes, réels ou complexes, Nota I. Rendiconti, Atti della Accademia Nazionale dei Lincei, Serie VIII, Vol VIII, pp. 743–750 Shapiro A (1990) On concepts of directional differentiability. J Optim Theory Appl 66:477–487 Vallender SS (1974) Calculation of the Wasserstein distance between probability distributions on the line. Theory Probab Appl 18:784–786 Van Dijk NM (1988) Perturbation theory for unbounded Markov reward processes with applications to queueing. Adv Appl Probab 20:99–111 Van Dijk NM, Puterman ML (1988) Perturbation theory for Markov reward processes with applications to queueing systems. Adv Appl Probab 20:79–98 Villani C (2003) Topics in optimal transportation, vol 58. American Mathematical Society, Providence Wessels J (1977) Markov programming by successive approximations with respect to weighted supremum norms. J Math Anal Appl 58:326–335 Yang M, Khan F, Lye L, Amyotte P (2015) Risk assessment of rare events. Process Saf Environ Prot 98:102–108 Zolotarev VM (1983) Probability metrics. Theory Probab Appl 28:278–302 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. 123