The folk theorem with imperfect public information in continuous time
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Frei, Christoph; Bernard, Benjamin Article The folk theorem with imperfect public information in continuous time Theoretical Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Frei, Christoph; Bernard, Benjamin (2016) : The folk theorem with imperfect public information in continuous time, Theoretical Economics, ISSN 1555-7561, The Econometric Society, New Haven, CT, Vol. 11, Iss. 2, pp. 411-453, https://doi.org/10.3982/TE1687 This Version is available at: https://hdl.handle.net/10419/150282 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/3.0/
Theoretical Economics 11 (2016), 411–453 1555-7561/20160411 The folk theorem with imperfect public information in continuous time Benjamin Bernard Department of Mathematics and Statistical Sciences, University of Alberta Christoph Frei Department of Mathematics and Statistical Sciences, University of Alberta We prove a folk theorem for multiplayer games in continuous time when players observe a public signal distorted by Brownian noise. The proof is based on a rigorous foundation for such continuous-time multiplayer games. We study in detail the relation between behavior and mixed strategies, and the role of public randomization to move continuously across games within the same model. Keywords. Folk theorem, repeated games, continuous time, imperfect observability. JEL classification. C73. 1. Introduction Folk theorems constitute a set of key results in the theory of discrete-time repeated games. These theorems state that the set of equilibrium payoffs expands to the set of all feasible and individually rational payoffs V∗as players are increasingly patient. A central requirement for the standard folk theorem to hold is that profitable deviations from strategy profiles can be detected. Clearly, this is satisfied if players perfectly observe each other’s actions. When players observe only a public outcome with noise, players’ deviations need to be sufficiently identifiable, that is, deviations of different players can be statistically distinguished. Under suitable identifiability assumptions, Fudenberg et al. (1994) establish the folk theorem in these games of imperfect public information. A continuous-time analogue to these important games of moral hazard has been introduced by Sannikov (2007). He studies a class of continuous-time games with two players, where the public signal is distorted by Brownian noise and players’ actions affect the drift rate of the signal. Rather than focusing on a folk theorem, Sannikov (2007) characterizes the set of pure strategy equilibrium payoffs via a differential equation of Benjamin Bernard: [email protected] Christoph Frei: [email protected] We would like to thank the editor and three anonymous referees for valuable comments helping us to significantly improve the structure and exposition of the paper. It is a pleasure to thank Tilman Klumpp, Semyon Malamud, and Yuliy Sannikov for valuable comments and interesting discussions. Financial support by the Natural Sciences and Engineering Research Council of Canada under Grant RGPIN/402585-2011 is gratefully acknowledged. Copyright ©2016 Benjamin Bernard and Christoph Frei. Licensed under the Creative Commons Attribution-NonCommercial License 3.0. Available at http://econtheory.org. DOI: 10.3982/TE1687
412 Bernard and Frei Theoretical Economics 11 (2016) its boundary ∂Ep(r) for any discount rate r>0. His paper offers seminal insight into continuous-time games, but his techniques rely heavily on using an ordinary differential equation, which is only possible in two-player games. The purpose of this paper is to prove the folk theorem in an extension of these games to any number of players. In the course of this, we have to construct continuous-time perfect public equilibria (PPE) that achieve nearly efficient payoffs. Our construction of equilibrium profiles resembles the iterative procedures of discrete time, in the sense that the equilibrium profiles are constructed using a local optimality condition on random time intervals [0τ1)[τ1τ2) for a suitable increasing sequence of stopping times (τ)≥0with τ→∞. Similarly to Fudenberg et al. (1994), the proof of the folk theorem canbesummarizedinthefollowingtwomainsteps: Step 1. Under suitable identifiability conditions, any smooth closed set W⊆intV∗ is locally self-generating, that is, for any w∈W, there exists a neighborhood Uwand a discount rate rwsuch that any payoff v∈Uw∩Wis attained by the same enforceable strategy profile for any r∈(0rw)with a continuation value that remains in Wfor a time interval [ττ+1)of positive length. Step 2. For a sufficiently small discount rate, any compact locally self-generating set is self-generating, i.e., payoffs can be attained with a continuation value that remains in Wforever. In continuous-time games, the continuation value of a strategy profile is characterized by a stochastic differential equation (SDE) similar as in Sannikov (2007). Step 1 thus requires us to find a solution to an SDE that remains in Wup to a suitable stopping time. The uniformity condition means that these solutions exist on a fixed probability space for the entire neighborhood Uw, that is, locally, these are strong solutions to the SDE. This is important when we concatenate these local solutions to a global solution in Step 2: by compactness of W, we have to deal with only finitely many probability spaces and it is thus easy to define an enlarged probability space that contains the concatenation. However, requiring that these local solutions are strong solutions in the definition of local self-generation creates additional difficulties in proving Step 1, as the conditions for the existence of strong solutions are much more stringent than for weak solutions. Nevertheless, we are able to show that local self-generation holds under suitable identifiability conditions. These conditions are essentially the same as in discrete time, plus an additional condition to ensure that strong solutions exist in a neighborhood of coordinate payoffs, i.e., payoffs that maximize or minimize a player’s payoff on W. Requiring strong solutions to the SDE in Step 1 above entails that the local strategy profiles are constant. The constructed equilibrium profiles are therefore constant on each of the intervals [ττ+1), which is a very desirable feature from both an implementation and an interpretation standpoint. Indeed, the finite variation property of the equilibrium profiles precludes strategies of unbounded oscillation, that is, agents do not switch actions infinitely often in finite time. Note also that agents adapt their strategies only at stopping times (τ)≥0. This leads to the interpretation of a continuously repeated game as a discretely repeated game where the length of the periods is not fixed but random. Indeed, on each of these intervals of random length, equilibrium profiles are constant and the corresponding SDE has a strong solution. This is consistent with
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 413 discrete games, where the sampled random variable at time tis necessarily fixed on the interval [tt +). Our methods indicate that rather than approximating the continuous-time world by discrete-time models with fixed time intervals, one should consider approximations where the length of a period is random. This construction avoids the “chattering phenomenon” that arises in optimal control when passing from discrete to continuous time. In comparison, consider a discrete-time game where players observe cumulative outcomes of a diffusion process at fixed times 2. Because it is not possible to bound the change in the public signal with fixed times as it is with stopping times, strategy profiles in the limit as →0will typically exhibit unbounded oscillation or one has to consider weaker forms of convergence as in Staudigl and Steg (2014). Thus, one can only obtain well implementable solutions in the above sense either by directly working in a continuous-time setting or by defining a suitable sequence of discrete games where the time periods are of random length. Not only is this paper the first to formally model continuously repeated games with any finite number of players, but it is also the first paper to introduce continuous-time strategies in mixed actions in games of imperfect public monitoring and to establish the continuous-time analogue of Kuhn’s theorem (realization equivalence of behavior strategies and mixed strategies). Additional results include a square-root law, establishing that changes in the discount rate r, the drift rate m, or the volatility σof the public signal do not affect the game as long as the informativeness of the signal relative to the players’ discounting σ(σσ)−1m/√rremains constant. This is similar in spirit to a square-root law that is obtained in Faingold and Sannikov (2011) and the news dependence of equilibria in Daley and Green (2012). In contrast to their results, however, we show that equilibrium strategies are transformed with a time change, and that the time-changed equilibria give rise to exactly the same path of the continuation values, at a different speed. Moreover, we show that public randomization can be used to move continuously over games within the same class of games when σ(σσ)−1m/√ris increased. Together, these two results imply that E(r) is monotonic in the discount rate if players have access to a public randomization device. While we study directly the continuous-time situation rather than a discrete-time approximation, we briefly mention some recent literature on the connection between discrete and continuous time in relation to the folk theorem. Fudenberg and Levine (2007) analyze a specific example between one long-lived and one short-lived player, where efficient limit equilibria can be obtained as the length of the time period shrinks to 0if the long-lived player’s actions affect the volatility of the Brownian signal. Sannikov and Skrzypacz (2010) consider games between two long-lived players, where the public signal has both a Brownian and a Poisson component, and players’ actions affect the drift of the Brownian motion and the intensity of the Poisson jumps. Building upon the methods of Fudenberg et al. (1994), they show a folk theorem uniformly for small and highlight the different impacts of Brownian and Poisson signals. Osório (2012) studies a model where players’ actions incur pairwise identifiable jumps in the signal otherwise given by a Brownian motion. As goes to 0, the signal becomes perfectly informative
414 Bernard and Frei Theoretical Economics 11 (2016) and a folk theorem is obtained. This is similar to discrete-time repeated games with perfect monitoring, where r→0and →0have the same effect. The remainder of the paper is organized as follows. We introduce the continuoustime model with strategies in mixed actions in Section 2.Section 3 states and discusses our main results and highlights some of the similarities and differences between discrete and continuous time. The proofs of all main results are contained in Section 4. In Appendix A, we provide a mathematical framework for continuous-time strategies in mixed actions and prove the continuous-time analogue of Kuhn’s theorem. Similarly to discrete time, the role of mixing is to weaken the conditions under which the folk theorem holds. When players are restricted to pure strategies, the same techniques can be used to construct equilibrium profiles, but the conditions need to be strengthened as we explain in the supplementary file available on the journal website, http://econtheory.org/supp/1687/supplement.pdf.Appendix B shows how we use public randomization to prove monotonicity of the equilibrium payoff set in the discount rate. Finally, Appendices Cand Dcontain proofs of auxiliary results. 2. The multiplayer setting 2.1 The model We consider a multiplayer game, where agents i=1n continuously take actions from the finite sets Aiat each moment of time t∈[0∞). Players may be allowed to mix their actions, in which case they continuously choose an element from (Ai),thesetof distributions over Ai. We denote by A=A1×···×Anand (A)=(A1)×···×(An) the spaces of pure and mixed action profiles, respectively. Agents cannot see their opponents’ actions and observe only the outcome of a public signal Y=σZ instead, where the constant volatility matrix σ∈Rd×kis of rank dand Zis a k-dimensional Brownian motion on some probability space ( FP). The arrival of public information is captured by the filtration F=(Ft)t≥0. It can be strictly larger than the filtration generated by Yto allow for public randomization, but this is not needed in our proof of the folk theorems.1 Definition 1. A (public) behavior strategy of player iis an F-progressively measurable stochastic process Ai:×[0∞)→(Ai). Distributions in (Ai)may be degenerate, so that behavior strategies contain pure strategies as special cases. A function m:A→Rddescribes the impact that a chosen pure action profile has on the drift rate of the public signal. The drift rate is extended to any mixed action profile α∈(A)by multilinearity, that is, μ(α) = a∈A m(a)α1(a1)···αn(an) (1) 1In continuously repeated games it is currently unknown whether the PPE payoff set E(r) is monotonic in the discount rate r. Hence, even when the folk result applies, E(r) may not expand monotonically to V∗ without public randomization. Using public randomization, we are able to show monotonicity of E(r) with an appropriate time change; see Theorem 3.
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 415 The choice of a strategy profile affects the future distribution of the public signal through a change of the probability measure. This is the natural analogue to taking expectations under conditional distributions in discrete time; see, for example, Fudenberg et al. (1994, p. 1000). Strategy profile Ainduces a family QA=(QA t)t≥0of probability measures, defined via its density process dQA t dP:=expt 0 μ(As)(σσ)−1dYs−1 2t 0 μ(As)(σσ)−1μ(As)ds(2) under which players observe the game when strategy profile Ais played. Under this family of probability measures, the public signal can be decomposed into Y=μ(As)ds+σZA where ZA:= Z−σ(σσ)−1m(As)dsis a Brownian motion on [0t]under QA tby Girsanov’s theorem. That is, from the players’ perspective, the signal Yconsists of a drift term given by μ(A) and Brownian noise. Remark 1. Because independence is always subject to a certain probability measure, a change of measure might affect the outcome of mixed actions. We assert in Lemma 14 that mixed actions remain unaffected by the change of measure in (2). Anderson (1984) and Simon and Stinchcombe (1989) demonstrate that seemingly simple strategies need not necessarily lead to unique outcomes in continuous time. This is not a problem in our model because actions taken by agents do not immediately generate information. Indeed, since the normal distribution has unbounded support, any realization of the signal is possible after play of any strategy profile. Therefore, these are games of full-support public monitoring. Restricting attention to public strategies also means that the probability space can be identified with the path space of all publicly observable processes. For a realized path ω∈, a pure strategy profile Athus naturally leads to the unique outcome A(ω).2This is analogous to discrete-time repeated games with full-support public monitoring; see Mailath and Samuelson (2006) for a thorough exposition of discrete-time games. Players i=1nreceive an expected flow payoff according to gi:A→Rthat depends on the action profile a−iof player i’s opponents only through m(a).3,4Again, giis extended to mixed action profiles by multilinearity at all times. 2For nondegenerate behavior strategy profiles, the outcome of strategy profile Aadditionally includes the outcome of players’ mixing. See Appendix A for the construction of a unified probability space on which these outcomes live. 3This is the most general payoff structure in continuous-time games of imperfect public information. Compare this to discrete-time repeated games, where players receive an expected instantaneous payoff gi(At)=EQA t[fi(Ai t˜ Y)|Ft],where ˜ Y=Yt+1−Ytis the change in the public signal. In continuous time, the infinitesimal change in the public signal dYthas drift m(At)dtunder QA. 4Because Brownian information can only be used to transfer value linearly, we need to impose an affine payoff structure gi(a) =bi(ai)m(a) −ci(ai)to show that Pareto-efficient payoffs are enforceable. This special payoff structure is used in Theorem 2, but not in our other main results. For pure strategy profiles,
416 Bernard and Frei Theoretical Economics 11 (2016) Definition 2. Let r>0be a common discount rate. Player i’s discounted expected future payoff under a behavior strategy profile Aat time tis given by Wi t(A;r):=∞ t re−r(s−t)EQA s[gi(As)|Ft]ds (3) We will omit the argument rwhen there is no chance of confusion. Because discounted expected future payoffs are normalized, Wt(A) lies in the set of feasible payoffs V:=conv{g(a)|a∈A}at all times t≥0with probability 1. Definition 3. A behavior strategy profile Ais a perfect public equilibrium (PPE) for discount rate rif for every player i=1nand every t≥0,wehave Wi t(A;r)≥Wi t(˜ A;r) a.s. for all public behavior strategy profiles ˜ Awith ˜ A−i=A−ia.e.5We denote the set of payoffs achievable by perfect public equilibria by E(r) :={x∈V|there exists a PPE Awith W0(A;r) =xa.s.} Similarly to discrete time, any player has a public best response to any public strategy profile of his opponents; see Lemma 15. Therefore, in a PPE, any player i=1ncan ensure that his payoff rate dominates his minmax payoff vi=min α−imax ai∈Aigi(aiα−i) at all times by myopically maximizing against his opponents’ strategy profile. Therefore, E(r) ⊆V∗,whereV∗:= {w∈V|wi≥vi∀i}denotes the set of feasible and individually rational payoffs. 2.2 Enforceable strategy profiles and self-generation Definition 4. An action profile αis said to be enforceable if there exist sensitivities β=(β1βn)∈Rn×dto the public signal, such that for every player i,thesumof expected flow payoff gi(α) and promised continuation rate βiμ(α) is maximized in αi. That is, for i=1nand every ai∈Ai, gi(α) +βiμ(α) ≥gi(aiα−i)+βiμ(aiα−i)a.s. (4) this payoff structure is the same as in Sannikov (2007), and one can show that Wt(A) is the Ft-conditional expectation under some probability measure QA ∞of ∞ t re−r(s−t)(bi(Ai s)dYs−ci(Ai s)ds) That is, biis the sensitivity of player i’s payoff to the public signal and ciis a cost-of-effort function. 5Because players maximize their discounted expected future payoff, deviations with time measure 0or probability 0are irrelevant. Therefore, two strategy profiles lead to the same continuation value if they are P⊗Lebesgue-a.e. the same.
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 417 A behavior strategy profile is enforceable if there exists a progressively measurable process (βt)t≥0such that (4) is satisfied a.e.6 It follows from (3) that the continuation game after time tis equivalent to the entire game. In discrete time, this feature gives rise to an iterative procedure over the time periods; in continuous time it leads to an SDE characterizing the infinitesimal change in the continuation value. The following lemma is the analogue of Theorem 1 in Sannikov (2007) adapted to the multiplayer setting with behavior strategies. Lemma 1. For an n-dimensional process Wand a behavior strategy profile A, the following statements are equivalent: (a) The process Wis the discounted expected future payoff under A. (b) The process Wis a bounded semimartingale that satisfies for i=1nthat dWi t=r(W i t−gi(At))dt+rβi t(σ dZt−μ(At)dt)+dMi t(5) for a martingale Mi(strongly) orthogonal to σZ with Mi 0=0and a progressively measurable process βiwith EQA T[T 0|βi t|2dt]<∞for all T≥0. Moreover, a behavior strategy profile Ais a PPE if and only if β=(β1βn)related to W(A)by (5)enforces A. Definition 5. A set W⊆Rnis self-generating for discount rate r>0if for every w∈W there exists a solution (WAβZM) to (5)suchthatβenforces A,W0=wa.s., and Wτ∈Wa.s. for every stopping time τ. Lemma 2. The set E(r) is the largest bounded self-generating set. This result is the equivalent of Theorem 1 in Abreu et al. (1990). It follows from Lemma 1 that any self-generating set is contained in E(r). The idea for the proof of the converse is that a PPE is subgame perfect, hence Wt∈E(r) a.s. since the continuation game is equivalent to the repeated game. For a formal proof one needs to deal with some measurability issues that we discuss in Section 4.1. 2.3 Enforceability and identifiability The folk theorem will follow from Lemma 2 once we find suitable conditions, under which any smooth set W⊆int V∗is self-generating for a sufficiently small discount rate. This means that we need to construct enforceable strategy profiles whose continuation values do not escape W. In this section we will motivate some necessary conditions for that to be possible. By examining (5), we see that on ∂W, both of the following statementshavetobesatisfied: 6It is enough to consider deviations to pure strategies, since any behavior strategy has a realization equivalent mixed strategy by Theorem 4, and a deviation to a mixed strategy can only be profitable if it has at least one profitable pure strategy in its support.
418 Bernard and Frei Theoretical Economics 11 (2016) w Nw g(α) drift rβt(σ dZt−μ(α)dt) ∂W dWt r2tr(βσσβ)dt Figure 1. At every point w∈∂W, the tangential diffusion rβt(σ dZt−μ(α)dt) leads to an outward-pointing drift of order r2tr(βσσβ)dt. For sufficiently small r,thistermisdominatedby the inward-pointing drift r(Wt−g(At)) dt. (i) The drift points inward, that is, N w(g(α) −w) > 0,whereNwis the outer-pointing normal vector at w∈∂W. (ii) The volatility is tangential to W. Indeed, otherwise the continuation value would immediately escape W;seealsoFigure 1. The first condition translates to enforceability of action profiles with extremal payoffs. The second condition is achieved using enforceability on hyperplanes and the related concept of orthogonal enforceability. Definition 6. (i) Let T∈Rn×(n−1)be a matrix whose column vectors T1Tn−1span the hyperplane H⊆Rn.Anactionprofileαis enforceable on the hyperplane Hif there exists amatrixB∈R(n−1)×dsuch that αis enforced by β=TB. (ii) Let Nbe a vector in Rn.Amatrixβ∈Rn×denforces αorthogonal to Nif it enforces αand satisfies Nβ=0. Observe that the two notions of enforceability are equivalent, i.e., αis enforceable on ahyperplaneHif and only if it is enforceable orthogonal to the normal vector Nof H.7 While enforceability on hyperplanes has the nice interpretation of transferring future value among players, it is often easier to work with the related concept of orthogonal enforceability because the normal vector to the smooth hypersurface ∂Wis unique. We distinguish two types of hyperplanes. Definition 7. AhyperplaneHis said to be coordinate if it is orthogonal to a coordinate axis. A hyperplane His regular if it is not coordinate. For an enforceable action profile α, the additional requirement to be enforceable on a coordinate hyperplane means that the corresponding player does not make any 7Indeed, if β=TB,thenNβ=0.Conversely,ifNβ=0, then all column vectors βjlie in H,which means they can be written as linear combinations of the Tj.Thisisequivalenttoβ=TB.
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 425 w g(α) Sw ∂W E[u(Y)] vw g(α) S w ∂W E[u(Y)] Figure 2. The left panel shows the decomposition of the payoff w=δg(α) +(1−δ)E[u(Y)]on the boundary ∂Win discrete time. The continuation payoff u(Y) lies on a hyperplane parallel to the tangent hyperplane Swfor all realizations yof Y. The right panel shows that payoffs close to ware decomposable with respect to the same hyperplane. time, the folk theorem is established by showing that any smooth set W⊆int V∗is selfgenerating for a sufficiently small discount rate rand therefore has to be a subset of the PPE payoff set by Lemma 2. To show that Wis self-generating, we construct equilibrium strategies with continuation values in Win the following steps: Step 1. For any payoff w∈W, there exists an enforceable strategy profile Awwith initial payoff wand continuation value in Wfor a short but positive amount of time τw. The discount rate rwmay depend on the payoff w. Step 2. There exist a neighborhood Uwof wand ˜ r>0such that for any r∈(0˜ r),the time τand the strategy profile Acan be chosen uniformly across Uw. Step 3. By compactness we can concatenate these local solutions to a global solution. In both discrete and continuous time, payoffs in the interior of Ware attainable by a static Nash equilibrium for a sufficiently low discount rate. We will thus focus on payoffs w∈∂Win this motivation. Step 1. In discrete time, Fudenberg et al. (1994) decompose any payoff w∈∂Winto a current-period payoff g(α) outside of Wand a continuation promise in the interior of W. In their decomposition, the continuation promise is parallel to the tangent hyperplane Swat w,thatis,αis enforceable on Sw;seealsoFigure 2. A payoff set Wis called decomposable on tangent hyperplanes if such a decomposition is possible for any w∈∂W, which is sufficient for Wto be self-generating in discrete time. As we mentioned in Section 2.3, in continuous time it is not enough that action profile αis enforceable on Swonly. Instead, αhas to be enforceable on all nearby hyperplanes so that the movement of the continuation value can be continuously adjusted to follow ∂W. For self-generation in continuous time, a payoff set Whas to be uniformly decomposable on tangent hyperplanes,thatis,foranyw∈∂W, there exists an enforceable action profile αwith g(α) strictly separated from Wby Swso that (α Nw)satisfies one of the four conditions of Lemma 5,whereNwis the unique outer-pointing normal vector to ∂Wat w.Thenw→βw αis locally Lipschitz continuous on ∂W, i.e., if an action profile is enforceable on a given hyperplane, then it can be enforced on nearby hyperplanes without changing the volatility significantly. Indeed, since Wis smooth, w→Nw
426 Bernard and Frei Theoretical Economics 11 (2016) is Lipschitz continuous, hence the concatenation with the locally Lipschitz continuous map from Lemma 5 is Lipschitz continuous on a suitable neighborhood Uwof w.Itis for this difference that the continuous-time folk theorem has the additional conditions (iii) and (iv).13 The stopping time τwcan be chosen as the time when the conditions of Lemma 5 are no longer satisfied, i.e., τw=inf{t≥0|Ww t/∈Bε(w)}for a suitable ε>0. Step 2. Because continuation payoffs are strictly separated from the hyperplane Sw in discrete time, it is possible to decompose payoffs vclose enough to won a translate to Swby moving the continuations by a small (and constant) amount; see also Figure 2. Because the variance of the continuation payoff is decreasing in δ, any payoff decomposable for ˜ δis also decomposable for δ∈(˜ δ1)by convexity of W. In continuous time, payoffs v∈Uwclose to wcan be decomposed by the same action profile αand Lipschitz continuous function v→βv αfor sufficiently small ˜ rby uniform decomposability. The main difficulty in this step is to ensure that the stopping times τvare uniformly bounded from below by a strictly positive stopping time, that is, infv∈Uwτv>0a.s. For this step it is important that the SDE (5) locally admits a strong solution so that the probability space and the Brownian motion Zcan be fixed on the entire neighborhood Uw. We show that the SDE is “sufficiently nice” so that solutions Wv to (5) with initial value v∈Uwhave the continuous flow property, that is, v→Wvis continuous for almost every ω∈. This implies that infv∈Uwτv>0a.s. if Uwis sufficiently small; see the proof of Lemma 9 for details. Step 3. The neighborhoods in Step 2 form an open cover of W, hence by compactness, there exists a finite subcover U={U1UN}. It follows that rU:= minU∈U˜ rUis strictly positive. In discrete time, the action profiles found in the one-period decomposition are concatenated at every step to form an enforceable strategy profile with continuations that remain in Wforever. In continuous time, we also take the minimum over the stopping times, i.e., we set τU:=minU∈UτU, which is strictly positive almost surely. For this step it is crucial again that (5) locally admits strong solutions, so that we have to deal with only finitely many 13In some games, slightly weaker conditions for enforceability on coordinate hyperplanes may be sufficient. However, if αdoes not have the unique best response property for player i, then some sort of identifiability conditions has to be satisfied for all deviations ai∈Aiwith gi(aiα−i)=gi(α) to ensure enforceability on hyperplanes infinitesimally close to being coordinate. Indeed, let j=ibe a player for whom αjis not a best response to α−j. Any neighborhood of eicontains vectors N(ε):=εej+1−ε2eifor εarbitrarily close to zero. From N(ε)β(ε) =0it follows that βi(ε) =−ε 1−ε2βj(ε) For aj∈Ajwith gj(ajα−j)−gj(α) =δ>0, enforceability implies βj(ε)(μ(α)−μ(ajα−j)) ≥δand hence βj(ε) is not orthogonal to μ(α) −μ(ajα−j)for any εclose enough to zero. The enforceability condition for player i,imposes 0≤εβj(ε)(μ(aiα−i)−μ(α)) It follows that either the right hand side is identically zero, that is, βj(ε) is orthogonal to μ(aiα−i)−μ(α), or βj(ε)(μ(aiα−i)−μ(α)) takes opposite signs for positive and negative ε, respectively. Both imply that μ(aiα−i)−μ(α) is linearly independent of μ(ajα−j)−μ(α). That is, some sort of identifiability condition has to be satisfied.
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 427 probability spaces and filtrations. These can be enlarged at the beginning of the concatenation to contain all the necessary information. Therefore, τUis indeed a stopping time, i.e., is measurable with respect to the enlarged filtration. Observe that τUdepends on the specific subcover Uthat is chosen. By construction, the stopping time τUis a functional of the public signal such that the continuation value Wvdoes not escape W on [0τU)regardless of the starting value v. Since Wv τU∈W,weusethesameprocedure to find a solution to (5)on[τUτU2)starting at Wv τU,whereτU2−τUis independent of τUand identically distributed as τUby independence and stationarity of the increments of Brownian motion. An iteration of this procedure leads to a sequence of stopping times (τU)≥0and solutions (W AβZ)≥0to (5)on[τUτU+1)such that τU+1−τU are independent and identically distributed (i.i.d.) as τU.14 This implies τU →∞a.s. Indeed, any sequence (τU)≥0of random variables with strictly positive and i.i.d. increments τU+1−τU diverges to ∞a.s. by the strong law of large numbers; see Lemma 11 for details. Therefore, a countable concatenation of the solutions (W AβZ)to (5) yields a solution on [0∞). This shows that Wis self-generating and hence W⊆E(r).Observe that the global solution is only a weak solution to (5) because the underlying Brownian motion was also concatenated at τU, depending on what element of the cover WτU fell into; hence the Brownian motion cannot be fixed a priori. We elaborate on some technical difficulties that arise with weak solutions in Section 4.1. 3.4 Finite-variation property of equilibrium profiles Because the constructed equilibrium profiles are concatenations of locally constant strategy profiles at i.i.d. copies of a positive stopping time τU, the resulting equilibrium profiles exhibit finitely many changes on every finite time interval. This is a very desirable feature for implementation because it seems unrealistic that agents can adapt their strategy profiles arbitrarily often. In this section, we present an example of such a strategy profile and compare it to the techniques used in Sannikov (2007).Consider the two-player partnership example of Section 2 in Sannikov (2007),reproducedinFigure 3 for the sake of exposition. To illustrate that the finite-variation property does not depend on the players’ ability of mixing, we restrict the example to pure strategies. Figure 3 shows a possible cover for a smooth payoff set Win the interior of V∗,such that on each element of the cover, the SDE (5) admits a strong solution. To ensure that the stopping times can be chosen strictly positive uniformly on each element of the cover, payoffs in a band of width εaround the element of the cover need to be decomposable with respect to the same pure action profile. The strategy profile is changed only when the continuation value leaves this band, ensuring that the strategy profile remains constant for a small but positive amount of time.15 14We omit mentioning Mas a part of the solution because M≡0inastrongsolutionto(5). 15The stopping times τU constructed in the proof are functionals of the signal such that the continuation value would not escape the band of width εif it were to start at any point of any neighborhood. For all practical purposes (including this example), one may think of these times as the times ˜τat which the continuation value leaves the band of width εaround UW˜τ−1.Because˜τ≥τU a.s. for any , the concatenation still extends to ∞.
428 Bernard and Frei Theoretical Economics 11 (2016) Figure 3. The matrix of static payoffs (w1w2)is shown to the left, and the right panel shows acoverof∂W(bold black line) into four overlapping sets (solid lines in gray scale), such that payoffs in a band of width εaround the sets (dashed lines in gray scale) can be decomposed with respect to the same pure action profile for discount rate r=01.Thecoverof Wis completed by playing the static Nash equilibrium in the interior of W. Also depicted is ∂Ep(01)(thin black line) constructed with the techniques in Sannikov (2007). In comparison, the construction of equilibrium profiles in Sannikov (2007) works even on the boundary ∂Ep(r), where the constructed strategy profiles are constant up to a finite number of “switching points.” However, due to unbounded variation of Brownian motion, the players will switch between action profiles an infinite number of times during a finite time interval when the continuation value crosses a switching point; see also Figure 4. While our approach of constructing equilibrium profiles is more general in the sense that it is applicable to any finite number of players, this example shows that it can have advantages even in two-player games. If players are not restricted to pure strategies, the realizations of their strategies are drawn continuously. Therefore, players switch actions infinitely often on finite time intervals even for constant (but mixed) strategy profiles. However, mixing is done individually for each player, and because of the multilinearity in (1), the public signal is not affected by the different realizations of a player’s mixed action as long as his strategy profile remains constant. The strategies of a player’s opponents are therefore not affected by the realizations of his mixed strategy; hence a change of actions within a constant strategy profile is a less complicated operation than a change of strategy profile. Moreover, because mand gare extended to mixed action profiles by multilinearity, continuoustime mixing may be interpreted as a division of effort among the pure actions in its support. This is a common formulation in continuous-time games of strategic experimentation; see, for example, Bolton and Harris (1999) and Keller and Rady (2010).
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 429 Figure 4. The left panel shows the simulation of the continuation value of a PPE in a zoom-in of Figure 3. Lines in light gray, dark gray, and black mean that action profiles (11),(01),and (00), respectively, are played; see Figure O.1 in the Supplement for a colored version. When the continuation value leaves the band around the cover of ∂W, the static Nash equilibrium is played until the boundary of Wis reached. The upper right panel shows the corresponding strategy profile. The lower right panel shows a strategy profile constructed with the techniques in Sannikov (2007) with unbounded variation when the continuation value crosses the switching point S(left panel). 4. Proofs of the main results 4.1 Weak definition of E(r) and self-generation In Section 2.1 we introduced our main object of study, the set E(r) of payoffs that are achievable by public perfect equilibria. For the proof of our results it is necessary to elaborate on what it means for a payoff x∈Vto be achieved by some PPE A. As we outlined in Section 3.3, we will construct weak solutions to the SDE (5) achieving x.Thismeans that the public signal, the filtrations, and the whole probability space may depend on x. We arrive at the specification E(r) :=x∈V there exists (FFP)containing an (FP)Brownian motion Zand a PPE Awith W0(A) =xP-a.s. We will also call (FFPZ)a stochastic framework for A. Remark 2. Technically, this weak definition is necessary to ensure existence of the solutions. But even aside from the technical advantages, the weak solution concept is appropriate here. From an interpretation standpoint, the difference between a strong solution and a weak solution to an SDE lies in the causality of the noise. If the noise is defined exogenously and not affected by players’ actions, this corresponds to a strong solution. In games of imperfect information, however, the noise is induced by players’ strategies and thus cannot be fixed at the beginning. This is in line with discrete-time games, where we only care about the distribution of the public signal and not on what probability space the distribution is realized.
430 Bernard and Frei Theoretical Economics 11 (2016) As illustrated in Section 3.3, we will piece together local solutions to obtain a global solution. At any stopping time τ, the value Wτ(A) is a random variable. Because of the weak formulation, the probability space depends on the point x∈E(r), and hence it is not clear what measurability conditions a random variable in E(r) should satisfy. This is clarified by the following lemma, whose rather technical proof is contained in Appendix C. Lemma 7. Let Xbe an F0-measurable random variable in a stochastic framework ( FFPZ). Then the following statements are equivalent: (a) We have X∈E(r) a.s. (b) There exists a PPE Awith W0(A) =Xa.s. This means we can only achieve random variables by a PPE on a fixed probability space. At first glance, this might seem like a major restriction because we are dealing with weak solutions. However, we will need this result only to concatenate locally strong solutions to a global solution, at which point we only have finitely many probability spaces by compactness. These probability spaces can be enlarged at the beginning of the concatenation, after which it remains fixed. From Lemmas 1and 7and we obtain the following stochastic characterization of E(r). Lemma 8. The following statements are equivalent for an F0-measurable random variable Xin a stochastic framework ( FFPZ): (a) We have X∈E(r) a.s. (b) There exists a strategy profile A, a square-integrable progressively measurable process β,amartingaleMorthogonal to σZ, and a bounded semimartingale Wsuch that βenforces A,W0=Xa.s., and A,β,W,Z,andMsatisfy (5). We conclude this section with the proof of self-generation. The argument sheds some first insight into the necessity of weakly defining E(r). Proof of Lemma 2.ByLemma 1, any bounded self-generating set Wis contained in E(r). Since E(r) is bounded, it remains to show that E(r) is self-generating. Take x∈E(r) so that Lemma 1 yields the existence of a stochastic framework (FFPZ), a behavior strategy profile Aenforced by β, a martingale Morthogonal to σZ, and a bounded semimartingale Wsatisfying (5)withW0=xa.s.Wenowfixastoppingtimeτand show that Wτ∈E(r) a.s. To do so, we set ˜ X=Wτ,˜ Ft=Ft+τ,˜ Zt=Zτ+t−Zτ,˜ Mt=Mτ+t−Mτ, ˜ Wt=Wτ+t,˜ βt=βτ+t,and ˜ At=Aτ+t. Because the tilde processes and filtrations satisfy condition (b) in Lemma 8, we obtain that Wτ=˜ X∈E(r) a.s. 4.2 Construction of continuous-time equilibria In Sections 2.3 and 3.3, we motivated the condition of uniform decomposability on tangent hyperplanes for a payoff set Wto be self-generating. In this section we will show
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 431 that this condition is also sufficient. In a first step, we show that any uniformly decomposable set Wis locally self-generating. Definition 9. A set W⊆Rnis called locally self-generating if for every point w∈W, there exist an open neighborhood Uwof w, a stochastic framework (FFPZ),an enforceable strategy profile A, a martingale Morthogonal to σZ,and˜ r>0such that for every discount rate r∈(0˜ r) there exists a stopping time τ>0such that for all v∈Uw∩W,thereexistWvwith Wv 0=va.s. and βvsuch that v→ Wvand v→ βvare Borel measurable and on 0≤t≤τ, the processes Wv,βv,A,M,Zare related by (5), βv t enforces Ata.s., and Wv t∈Wa.s. Lemma 9. Suppose that a smooth set W⊆V∗is uniformly decomposable on tangent hyperplanes. Then Wis locally self-generating. Proof. Suppose first that wis in the interior of W.LetUw=Bε(w),whereε>0is chosen such that the open ball B2ε(w) is contained in W. For a static Nash equilibrium αe, the constant strategy profile A≡αeis enforced by β≡0.Foranyr>0and any v∈Uw,letWvbe a strong solution to dWv t:=r(W v t−g(αe))dt with initial condition Wv 0=va.s. The explicit solution Wv t=v+(ert −1)(v −g(αe)) is measurable in v.Lett0(r) := log(1+ε/(w−g(αe)+ε))/r.Thenforanyv∈Bε(w) it follows that Wv t−v≤(ert −1)ε+w−g(αe)≤ε on [0t0]. Therefore, in the interior of W, we can support any discount rate r>0by choosing Uw=Bε(w) and the deterministic time τ≡t0(r) > 0. For any w∈∂W, denote by Nwthe outward unit normal to ∂Win wand denote by Swthe tangent hyperplane to ∂Win w. By smoothness of W, both of these are unique and continuous in w∈∂W. Fix now a payoff w∈∂W. It will be convenient to work in a coordinate system with origin in wand a basis consisting of an orthonormal basis of Sw and Nw, where we choose the nth coordinate in the direction of Nw. Since ∂Wis a C2submanifold, we can locally parametrize it by a twice differentiable function ϕ.Letˆ x=(x1xn−1)denote the projection onto the first n−1components so that the boundary ∂Wis locally given by (ˆ xϕ(ˆ x)). By assumption, there exists an enforceable action profile αsuch that g(α) is strictly separated from Wby Sw.Letβα be the locally Lipschitz continuous function from Lemma 5, which assigns to any vector x∈Rnamatrixβenforcing αorthogonal to x. Choose an ε>0such that the following statements hold: (i) We have N vNw>0for all v∈B2ε(w) ∩∂W. (ii) For all v∈B2ε(w),∇ϕ(ˆ v)≤p1and |ij ϕ(ˆ v)|≤p2for i j =1nand constants p1p2>0,whereij ϕdenotes the second partial derivative of ϕwith respect to ˆ viand ˆ vj.
432 Bernard and Frei Theoretical Economics 11 (2016) (iii) The map βα(−∇ϕ(ˆ·)1)is Lipschitz continuous on B2ε(w) and |(βασσβ α)ij |≤B for i j =1nand a constant B<∞. (iv) We have c(ε) :=N w(g(α) −w) −2ε(1+p1)−p1g(α) −w>0. The first condition makes sure that a local parametrization ϕexists with bounded gradient as in condition (ii). Since ∂Wis assumed to be C2, the first two derivatives of ϕare continuous, hence locally bounded. In particular, by letting εsmall enough we get the first two conditions to hold. For the third condition, observe that ϕis continuous with bounded derivative by condition (ii), hence is Lipschitz continuous. Since the projection ˆ·is Lipschitz continuous with Lipschitz constant 1and the composition of Lipschitz continuous functions is Lipschitz again, the third condition holds in a small neighborhood of Nw=(−∇ϕ( ˆ w)1)by Lemma 5. Finally, N w(g(α) −w) in condition (iv) is positive by strict separation of g(α) from W.Because∇ϕis continuous and ∇ϕ( ˆ w) =0, p1can be made arbitrarily small by choosing a small ε. This implies that c(ε) > 0for sufficiently small ε. Fix a stochastic framework (FFPZ) and an εsatisfying all of the above conditions. Denote Uw:= Bε(w), fix a discount rate r≤2c(ε)/((n −1)2p2B) =: ˜ r,andlet A≡α.Forallv∈Bε(w),letWvdenote the strong solution to dWv t=r(W v t−g(α))dt+rβα(−∇ϕ( ˆ Wv t)1)(σ dZt−μ(α)dt) on J0τvKwith W0=va.s., where τv:=inf{t>0|Wv t/∈B2ε(w)}. Using that βα◦(−∇ϕ(ˆ·)1) is uniformly bounded and Lipschitz continuous on B2ε(w) by condition (iii), a strong solution to this SDE exists by Theorem 5.2.1 of Øksendal (1998).16 Note that the process βt:=βα(−∇ϕ( ˆ Wv t1)) is progressively measurable on J0τvKas a concatenation of a progressively measurable process with a Borel measurable function. Moreover, it enforces Aand it is bounded on J0τvKby Lemma 5, hence is locally square integrable. Let Dv t:=Wvn t−ϕ( ˆ Wv t)measure the distance from Wtto ∂Win the direction of Nw as shown in Figure 5. By Itô’s formula, dDv t=(−∇ϕ( ˆ Wv t)1)dWv t−1 2 n−1 ij=1 ∂2ϕ( ˆ Wv t) ∂xi∂xj dWviWvjt =r(−∇ϕ( ˆ Wv t)1)(W v t−g(α)) −r 2 n−1 ij=1 ∂2ϕ( ˆ Wv t) ∂xi∂xj (βtσσβ t)ij dt +r(−∇ϕ( ˆ Wv t)1)βα(−∇ϕ( ˆ Wv t)1)(σ dZt−μ(α)dt) ≤rr 2(n −1)2p2B−c(ε)dt 16Uniform boundedness implies the linear growth condition needed for the existence result. The function f(x)=r(x−g(α)) is linear, hence Lipschitz continuous and of linear growth.
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 433 v w Nw g(α) drift Sv ∂W Wv t Dv t Dv t N(ˆ Wv tϕ( ˆ Wv t)) βα Figure 5. The distance of Wv tto ∂Wis measured by Dv tin the direction of Nwfor all vin a neighborhood of w. The continuation values Wv tremain in Wif and only if Dv t≤0. The sensitivity βα of Wvto Zis chosen orthogonal to (−∇ϕ( ˆ Wv t)1), the normal vector to ∂Win the projection (ˆ Wv tϕ( ˆ Wv t)) of Wv tonto ∂W. whereweusedthatxβα(x) =0for all x∈B2ε(w) and that conditions (ii) and (iv) imply (−∇ϕ( ˆ Wv t)0)+Nw(W v t−w+w−g(α)) ≤∇ϕ( ˆ Wv t)2ε+g(α) −w+2ε−N w(g(α) −w) ≤−c(ε) This implies that for any r∈(0˜ r),Dvis absolutely continuous with dDv t/dt≤0on J0τvK,whereτvdepends on r. Since Dv 0≤0for all v∈Uw∩W, it follows that Dv t≤0 on J0τvK. Next, we show that the stopping times τvare uniformly bounded from below by a stopping time τ>0. The idea is that this SDE is sufficiently nice such that the flow v→ Wvis continuous and thus Wvcan be approximated by W¯ vfor ¯ vclose to v.This leads to a cover of Bε(w) with a finite subcover, over which the minimum of stopping times is still positive. Denote Vv t:=e−rt(W v t−v) and derive from the product rule that it is the solution to the SDE dVv t=re−rt(v −g(α))dt+re−rtβα(−∇ϕ( v+ertVv t)1)(σ dZt−μ(α)dt) Fix a time horizon T>0to make ert bounded and Lipschitz continuous. To apply Theorem V.37 of Protter (2005),wewriteVvin its integrated form Vv t=r(1−e−rt)(v −g(α)) +t 0 F(V v)s(σ dZs−μ(α)ds) t ≤T where F(V v)s=re−rsβα(−∇ϕ( v+ersVv s)1). Both the finite variation part and Fare Lipschitz,17 hence Theorem V.37 of Protter (2005) applies and we deduce that the flow 17The result requires that Fis functional Lipschitz, which is satisfied for any operator induced by a Lipschitz function; see Protter (2005, p. 251).
434 Bernard and Frei Theoretical Economics 11 (2016) v→Vv(ω) is continuous for almost all ω.18 Let σv:= inf{t>0|ertVv t>ε/2},whichis strictly positive by continuity of Vv.Foranyv ¯ v∈Bε(w), define the auxiliary process v¯ v t:=ertVv t1J0σ¯ v)) +erσ¯ vVv σ¯ v1Jσ¯ v∞)) .Forfixed¯ v∈Bε(w), the map v→v¯ vis continuous for almost all ω.Because¯ v¯ v≤ε/2,thereexistsanε¯ v>0such that v¯ v<εfor all v∈Bε¯ v(¯ v). This implies Wv−v<εon J0σ¯ v)) for all v∈Bε¯ v(¯ v) ∩Bε(w) and thus τv>σ¯ va.s. Since Bε(w) is compact, there exists a finite subcover Bε¯ v1(¯ v1)Bε¯ vm(¯ vm) of Bε(w) and thus τ:=inf=1m σ¯ vis strictly positive. Since v→Wvis continuous, it is Borel measurable and because βis a continuous functional of W,soisv→βv. Having locally constructed strong solutions to (5), we need to piece them together to construct a global weak solution. The following lemma tells us that this is possible if the corresponding payoff set is compact. Lemma 10. Let W⊆Rnbe a compact locally self-generating set. Then there exists a discount rate ˜ rsuch that W⊆E(r) for any r∈(0˜ r). Proof. The family of open neighborhoods (Uw)w∈Wforms an open cover of W;hence by compactness there exists a finite subcover (Uk)k=1N . By making these sets disjoint, we obtain a finite, Borel measurable cover of W. On each of these (now disjoint) sets Uk we have a stochastic framework (kFkFkPkZk).Let(FFP) be the product space, that is, := 1×···×N,F:= F1⊗···⊗FN,andsimilarlyforFt,t≥0,and define P(B) := P(B1)···P(BN)for B=B1×···×BN∈F. Choose any discount rate r smaller than ˜ r=min(˜ r1˜ rN)>0. Then, dependent on r, for every Ukthere exist a strategy profile Ak, a martingale Mkorthogonal to σZk, and a stopping time τk>0 such that for every v∈Uk∩W,thereexistβv,Wvsatisfying the appropriate conditions. Let now τ(ω1ωN):=min(τ1(ω1)τN(ωN)1),whichispositiveP-a.s. Fix v∈Wand let κbe the index such that v∈Uκ.OnJ0τKwe set Z=Zκ,A=Aκ, M=Mκ,W=Wv,andβ=βv, and we know that they have the desired properties. In particular, Wτ∈W; hence we can concatenate the solution with the processes related to the neighborhood Ukthat Wτfalls into. More precisely, on Jτ ˜τK,where˜τ−τis identically distributed as τ,weset Zt(ω) :=Zκ τ(ω)(ωκ)+ N k=1 Zk t−τ(ω)(ωk)1{Wτ(ω)∈Uk} and similarly for M. It follows from the strong Markov property of each Zkthat Zis an (FP)-Brownian motion. Since τis bounded, the optional stopping theorem implies 18Here, Vv(ω) is to be understood as an element of the space Dnof càdlàg functions from [0∞)to Rn with the topology of uniform convergence on compacts. A compatible metric is given by d(fg) := ∞ k=1 1 2k1∧sup 0≤s≤kf(s)−g(s); see Protter (2005, p. 220).
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 441 Definition 10. A mixed strategy for player iis a probability measure κion the set Pi={γi:C[0∞)d×[0∞)→Ai|γiis progressively measurable}20 Amixed strategy profile κis given by the product measure κ=κ1⊗···⊗κnon the product σ-algebra on P=P1×···×Pn,whereκiis a mixed strategy for player i=1n.Player i’s discounted expected future payoff of a mixed strategy profile is given by Wi t(κ) =P Wi t(γ)dκ(γ) =P1···PnWi t(γ1γn)dκ1(γ1)⊗···⊗dκn(γn) Since mixing over strategies is a more complicated procedure than mixing over actions, it is usually easier to work with behavior strategies than with mixed strategies. In the remainder of this appendix we show in an analogue to Kuhn’s theorem (see Kuhn 1953) that the two notions are equivalent. A mixed strategy profile κand a behavior strategy profile Aare realization equivalent if they lead to the same distribution over outcomes, that is, for any a∈A, A(a) =κ({γ∈P|γ=a})a.e. Theorem 4 (Analogue of Kuhn’s theorem). Every mixed strategy profile is realization equivalent to some behavior strategy profile. Conversely, every behavior strategy profile has a realization equivalent mixed strategy profile. Proof.Letκbe a mixed strategy profile. Fix a player iand define for any ai∈Aiand any (ω t) ∈C[0∞)d×[0∞), Ai t(ai;ω) :=κi{γi∈Pi|γi t(ω) =ai} It can be deduced that Ai(ai)is almost everywhere well defined and a progressively measurable process for all ai∈Ai. Indeed, the sets S(ai)={(γiωt)∈Pi×C[0∞)d×[0∞)|γi t(ω) =ai} are elements of the product σ-algebra of σPi(see footnote 20) and the progressive σalgebra on C[0∞)d×[0∞). Since {γi∈Pi|γi t(ω) =ai}are the (ωt) sections of S(ai), it follows from measurable induction that the mapping (ω t) →κi{γi∈Pi|γi t(ω) =ai} is progressively measurable, which means that Ai(ai)is progressively measurable. Moreover, the processes Ai(ai)are nonnegative and their sum over ai∈Aiis 1. Since κ=κ1⊗···⊗κn, it follows that A(a) := A1(a1)···An(an)indeed defines a realization equivalent behavior strategy profile. 20Formally, κiis defined not on Piitself, but on the σ-algebra σPion Pigenerated by the coordinate maps πt:Pi→(C[0t]d→Ai)given by πt(γi):=γi t;seealsoBillingsley (1986, p. 509).
442 Bernard and Frei Theoretical Economics 11 (2016) Let now Abe a behavior strategy profile and let U1Unbe independent processes with standard uniform marginals as in the proof of Lemma 12 on some probability space (˜ ˜ F˜ P).Foranyplayeri, enumerate Ai={ai 1ai Ki}and define γi t(˜ωω) = Ki j=1 ai1{j−1 =1Ai t(ai ;ω)≤Ui t(˜ω)<j =1Ai t(ai ;ω)} which we consider as a mapping in ˜ωfrom ˜ to the set Piof progressively measurable processes C[0∞)d×[0∞)→Ai. Using the σ-algebra from footnote 20, this mapping becomes measurable; hence we can define a probability measure κion Pias the preimage of γiunder ˜ P,thatis,κi=˜ P◦(γi)−1. Therefore, κ=κ1⊗···⊗κnindeed defines a realization equivalent mixed strategy profile. Appendix B: Time-changed PPEs and monotonicity of E(r) Proof of Lemma 6. For a strategy profile Ain a stochastic framework (FFPZ), we define the time-changed processes ˜ At:=Aλt˜ Zt:= 1 √λZλt˜ Yt:= ˜σ˜ Zt Observe that ˜ Ais progressively measurable with respect to the time-changed filtration ˜ F=(˜ Ft)t≥0,where ˜ Ft:= Fλt,and ˜ Zis an ˜ F-Brownian motion by the scaling property of Brownian motion. The family (˜ Q˜ A t)t≥0of probability measures induced by ˜ Awith respect to ˜ m,˜σ,and ˜ Yis defined as d˜ Q˜ A t dP:=expt 0˜μ( ˜ As)(˜σ˜σ)−1d˜ Ys−1 2t 0˜μ( ˜ As)(˜σ˜σ)−1˜μ( ˜ As)ds Abbreviate ˆ m:= ˜ m/√λand observe that ˜σ(˜σ˜σ)−1ˆ m=σ(σσ)−1mby assumption. With the substitution d˜ s=λdswe arrive at d˜ Q˜ A t dP=expt 0ˆμ(Aλs)(˜σ˜σ)−1˜σdZλs −1 2t 0ˆμ(Aλs)(˜σ˜σ)−1ˆμ(Aλs)λ ds =expλt 0 μ(A˜ s)(σσ)−1σdZ˜ s−1 2λt 0 μ(A˜ s)(σσ)−1μ(A˜ s)d˜ s=dQA λt dP and hence ˜ Q˜ A tcoincides with QA λt on ˜ Ft=Fλt. Observe that the expected flow payoff gi(a) =fi(aim(a)) =fi(aiσ˜σ(˜σ˜σ)−1˜ m(a)/√λ) still depends on a−ionly through ˜ m(a). Substituting d˜ s=λdsagain, we obtain for every t≥0, ˜ Wi t(˜ A;λr ˜ m ˜σ) := ∞ t λre−λr(s−t)E˜ Q˜ A s[gi(˜ As)|˜ Ft]ds (9) =∞ λt re−r(˜ s−λt)EQA ˜ s[gi(A˜ s)|Fλt]d˜ s=Wi λt(A;r m σ) a.s.
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 443 Because all unilateral deviations of (Aλt)t≥0correspond to unilateral deviations of A,it follows from (9) that (Aλt)t≥0is a PPE with respect to (λr ˜ m ˜σ) if and only if Ais a PPE with respect to (rmσ). Proof of Theorem 3. To show (a), we need to show that a PPE Awith respect to (rmσ) can be transformed to a PPE with respect to (rmσ).Let(FFPZ) be the stochastic framework of A,andletβ,Z,andMbe the processes from Lemma 1 that satisfy (5)forAand W=W(A)with respect to μand σ.LetZ⊥be a k-dimensional Brownian motion orthogonal to both Zand M, and denote by ˆ Fthe augmented filtration generated by Fand Z⊥.Becauseis symmetric, we can write =QDQ for an orthogonal matrix Qand a diagonal matrix D=⎛ ⎝ I00 0−Im0 00 ˜ D⎞ ⎠ such that ˜ Dhas entries in (−11).Define ˜ :=QIk−D2Qso that 2+˜ 2=Ik.Then ˆ Z:=Z +˜ Z⊥is a Brownian motion with respect to ˆ F.Set ˆ :=Q00 0(Ik−−m−˜ D2)−1Q and define ˆ Z⊥:= ˆ (Z −ˆ Z). It follows from dˆ Z⊥ˆ Z⊥t=00 0Ik−−mdt and dˆ Z⊥ˆ Zt=0that ˆ Z⊥is a martingale orthogonal to ˆ Z. This gives us the decomposition Z=ˆ Z+˜ ˆ Z⊥and hence dWt=r(Wt−g(At))dt+rβt(σ dZt−μ(At)dt)+dMt =r(Wt−g(At))dt+rβt(σdˆ Zt−μ(At)dt)+rβtσ˜ dˆ Z⊥ t+dMt Since ker(σ) =ker(σ) by assumption, it follows that there exists a matrix ∈Rd×dsuch that σ =σ.21 Therefore, σ ˆ Z=2σZ+σ˜ Z⊥is orthogonal to Mand hence also to ˆ M:=M+rβsσ˜ dˆ Z⊥ s. It follows that Walso fulfills (5)forσ with processes β,ˆ Z,and ˆ M. Since βenforces A,Ais also a PPE in the continuous-time game with volatility σ by Lemma 1. Note, however, that Ais a PPE as an ˆ F-progressively measurable process, that is, when players use the new orthogonal information suitably. We derive from Itô’s formula that d(e−rtWt)=−re−rtg(At)dt+re−rt ˆ βt(σdˆ Zt−μ(At)dt)+e−rt dˆ Mt(10) 21Indeed, σis an isomorphism from ker(σ)⊥to Rdwith inverse σ(σσ)−1and hence the vectors bi:=σ(σσ)−1eifor i=1d form a basis of ker(σ) =ker(σ),whereeidenotes the ith standard basis vector in Rd. Define δi:= σbifor i=1d and set =(δ1δd).Byconstruction,wehave σbi=δi=σbi.Since(b1bd)can be completed to a basis of Rnwith any basis of ker(σ) and ker(σ) =ker(σ), it follows that σ =σ.
444 Bernard and Frei Theoretical Economics 11 (2016) Observe that ker(σ) =ker(σ) implies rank(σ) =d; hence we can define a family (˜ QA t)t≥0of probability measures as in (2)withrespecttomand σ. Integrating (10) from tto Tand taking ˜ QA T-conditional expectations on ˆ Ftyields Wt=T t e−r(s−t)E˜ QA s[g(As)|ˆ Ft]ds+e−r(T−t)E˜ QA T[WT|ˆ Ft] Taking the limit as T→∞yields Wt(A;r m σ) =ˆ Wt(A;r m σ) a.s. since Wis bounded. This concludes the proof of (a). For statement (b), let first ˜σ=σ.Definethek×kmatrix :=σ(σσ)−1σ so that σ=σ. Note that =σ(σσ)−1σ=, i.e., is symmetric. Every vector in the kernel of σis an eigenvector of with eigenvalue 0.Letλbe an eigenvalue of to an eigenvector vthat is not in the kernel of σ.Thenσv =σv=λσv,thatis,λis an eigenvalue of for eigenvector σv. Since this applies to all eigenvalues of outside the kernel of σ, the eigenvalues of lie in [−11]and ker(σ)=ker(σ).Moreover,σ= σ has rank dbecause is invertible and σhas rank d. The statement now follows by applying (a) to . Observe that the change ˜ m=−1mis completely equivalent since σ(σσ)−1m=σ−1(σσ)−1−1m=σ(σσ)−1−1m and hence the induced probability measures coincide at all times. For statement (c), Lemma 6 implies ˜ Wt((Aλt)t≥0λrmσ/√λ) =Wλt(Armσ) a.s. for λ∈(01).The statement follows from (a) for the matrix =diagk(√λ) applied to σ/√λ. Appendix C: Proofs of Lemmas 1,5,and 7 The statement of Lemma 1 is rather intuitive, since (3) can be rewritten as Wi t(A) =rertlim u→∞ EQA uu 0 e−rsgi(As)dsFt−t 0 e−rsgi(As)ds and hence dWi t=rWtdt−rgi(At)dt+d“martingale” by the product rule. However, the limiting probability measure QA ∞is not equivalent to Pon F∞; hence we cannot immediately apply a martingale representation result.22 Proof of Lemma 1.Toshow(a)⇒ (b), observe first that Wi:=Wi(A) is bounded, as it remains in Vat all times. Fix T>0and derive from (3) that wi T:=Wi T−rT 0 (W i t−gi(At))dt (11) =Wi T+rT 0 gi(At)dt−r∞ 0s∧T 0 re−r(s−t)EQA s[gi(As)|Ft]dtds 22The probability measure QA ∞that coincides with QA ton Ftfor every t≥0. It is not obtained as the limit in (2), but its existence is asserted by Proposition I.7.4 of Karatzas and Shreve (1998).
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 445 is a bounded FT-measurable random variable. By Corollary 1 to Theorem IV.37 of Protter (2005), any square-integrable QA T-martingale can be decomposed uniquely into a stochastic integral with respect to σZ −μ(As)dsand a square-integrable martingale Morthogonal to Z. Applying this to EQA T[wi T|Ft],weobtainanF0-measurable ci T, a progressively measurable process (βi tT )0≤t≤Twith EQA T[T 0|βi tT|2dt]<∞,anda QA T-martingale (Mi tT)0≤t≤Torthogonal to Zwith Mi 0T =0such that wi T=ci T+T 0 rβi tT(σ dZt−μ(At)dt)+Mi TT To prove that (b) holds, we need to show that ci T,βi tT ,andMi tT do not depend on T.Let ˜ T≤Tand take in (11) conditional expectations on F˜ Tunder QA Tto deduce that EQA T[wi T|F˜ T]−wi ˜ T=EQA T[wi T|F˜ T]−Wi ˜ T+rT ˜ T EQA t[gi(At)|F˜ T]dt −r∞ ˜ Ts∧T ˜ T re−r(s−t)EQA s[gi(As)|F˜ T]dtds =EQA T[wi T|F˜ T]−Wi ˜ T−∞ T re−r(s−t)EQA s[gi(As)|F˜ T]ds +∞ ˜ T re−r(s−˜ T)EQA s[gi(As)|F˜ T]ds =0 using that EQA T[X|F˜ T]=EQA s[X|F˜ T]for Fs-measurable Xand s≤T. Taking ˜ T=0,this shows that ci T=Wi 0does not depend on T.Italsoimplies wi ˜ T=EQA T[wi T|F˜ T]=Wi 0+˜ T 0 rβi tT(σ dZt−μ(At)dt)+Mi ˜ TT which yields βi ·T =βi ·˜ Ta.e. and Mi ˜ TT =Mi ˜ T˜ Ta.s. by the uniqueness of the orthogonal decomposition. Taking Ft-conditional expectations, we deduce Mi t ˜ T=Mi tT a.s. for every t;henceW(A)satisfies (b). To prove (b) ⇒ (a), we derive from Itô’s formula that d(e−rtWi t)=−re−rtgi(At)dt+re−rtβi t(σ dZt−μ(At)dt)+e−rt dMi t(12) Integrating (12)fromtto Tand taking QA T-conditional expectations on Ftthus yields Wi t=T t re−r(s−t)EQA s[gi(As)|Ft]ds+e−r(T−t)EQA T[Wi T|Ft] Since Wis bounded, the second summand converges to zero a.s. as Ttends to ∞;hence Wi tis indeed the discounted expected future value of A.
446 Bernard and Frei Theoretical Economics 11 (2016) For the last statement, fix a player iand a time t,andlet ˜ Abe a strategy profile with ˜ A−i=A−i.Forβrelated to W(A)by (5), we obtain from (12)foru≥tthat Wi t(A) =e−r(u−t)Wi u(A) −u t re−r(s−t)βi s(σ dZs−μ(As)ds) −gi(As)ds−dMi s As we let u→∞,theterme −r(u−t)Wi u(A) vanishes since Wi(A) is bounded. Since Mis a martingale up to time ualso under Q˜ A u,weobtain Wi t(˜ A) =lim u→∞ EQ˜ A uu t re−r(s−t)gi(˜ As)dsFt =Wi t(A) +lim u→∞ EQ˜ A uu t re−r(s−t)(gi(˜ As)−gi(As))ds +βi s(σ dZs−μ(As)ds)Fta.s. Because wi Tdefined in (11) is bounded, it follows from the construction of βthat the process · tre−r(s−t)βi s(σ dZs−μ(As)ds) is up to any time u∈(t ∞)aboundedmean oscillation (BMO) martingale under the probability measure QA u. This implies by Theorem 3.6 of Kazamaki (1994) that · tre−r(s−t)βi s(σ dZs−μ( ˜ As)ds) is a BMO martingale under Q˜ A u. Together with Fubini’s theorem this implies Wi t(˜ A) −Wi t(A) =∞ t re−r(s−t)EQ˜ A sgi(˜ As)−gi(As)+βi s(μ( ˜ As)−μ(As))|Ftdsa.s. (13) If βenforces A, the above conditional expectation is nonpositive; hence Ais a PPE. To show the converse, assume toward a contradiction that there exists a player iand a set ⊆×[0∞)with P⊗Lebesgue() > 0, such that for some other strategy ˆ Ai, gi(ˆ AiA−i)−gi(A) +βi(μ( ˆ AiA−i)−μ(A)) > 0on Set ˜ Ai:= ˆ Ai1+Ai1c.Becauseβis progressively measurable, we can and do choose and ˆ Ato be progressively measurable as well. In particular, ˜ Aiis a behavior strategy for player i.Forsuchan˜ A,theexpectationin(13) is strictly positive for t=0,a contradiction. Proof of Lemma 5. Statement (i). Let βi α(N) =Bias in (6), which is locally Lipschitz continuous in N. The statement holds by choosing UNsuch that the first two coordinates are bounded away from zero. Statement (ii). The case where Nis not parallel to a coordinate axis is shown in statement (i); hence suppose now that N=e1.Let ˜ β1 ˜ βnbe defined as in the proof of Lemma 4 and set β1 α(x) :=− n i=2 xi x1˜ βiβ i α(x) =˜ βii=2n (14)
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 447 Along the lines of the proof of Lemma 4, it follows that βα(x) enforces αorthogonal to x if x1=0. The statement follows by choosing UNbounded away from {x1=0}. Statement (iii). Suppose N=e1and that α1=a1∈A1is a unique best response to α−1.Letβi α(x) as in (14), except that ˜ βiare replaced by βi.Then,clearly,(4)is fulfilled for players i=2n. Because of the unique best response property, there exists an ε>0such that g1(α1)≥g1(˜ a1α−1)+εfor every ˜ a1∈A1\{a1}.Letusset B=maxi=2n max˜ a1∈A1|βi(μ(˜ a1α−1)−μ(α))|, which is finite as βis fixed. If B=0, then αis a Nash equilibrium and the result holds by statement (iv). Suppose, therefore, that B>0. Then for all xin Ue1:=x∈Rnx−e1≤ ε B(n −1)+ε x1is bounded away from 0, and hence |xi|/x1≤ε/(B(n −1)). It follows that β1 α(x)(μ(˜ a1α−1)−μ(α))= n i=2 |xi| x1βi(μ(a1α−1)−μ(α))≤ε for every ˜ a1∈A1. Together with the unique best response property for player 1,this shows that βα(x) enforces αfor x∈Ue1.Forallx∈Ue1,βα(x)x=0by construction and βαis Lipschitz continuous and bounded since x1is bounded away from 0. Statement (iv). This is clear since βα(x) =0for all x∈Rn. Proof of Lemma 7. To show that (a) ⇒ (b), let X∈E(rm) a.s. Although we may have different probability spaces in (a) for each realization X=x, we can use the fact that the models all share the same path space to construct a regular conditional probability on that space. The path space of a behavior strategy Aand its stochastic framework is given by D:=(A)[0∞)×C[0∞)d,andC[0∞)dis the space of continuous functions [0∞)→Rd. One can show that =V×Dis complete and separable,23 hence by Theorem V.3.19 in Karatzas and Shreve (1998) there exists a regular conditional probability Px(F) :V×F→[01], which means that it has the following properties: (i) For each x∈V,Pxis a probability measure on (F). (ii) For each F∈F, the mapping x→Px(F) is B(V)-measurable. (iii) For each F∈F,Px(F) =P(F|X=x) for ν-a.e. x∈V,whereνis the distribution of X. We know that for each x∈E(rm),thereexistsaPPEAxachieving x.LetAnow be the process defined pointwise by Axon {X=x}. It follows from the properties of a regular conditional probability that Ais a PPE achieving X. Indeed, for any player iand any 23Vis complete and separable as a closed subset of Rn,C[0∞)dis both complete and separable with respect to the uniform metric by Theorem 43.6 of Munkres (2000) and (A)[0∞)is compact by Tychonov’s theorem and hence complete and separable.
448 Bernard and Frei Theoretical Economics 11 (2016) behavior strategy profile ˜ Awith ˜ A−i=A−i, P(W i 0(A) ≥Wi 0(˜ A)) =E(rm) P(W i 0(A) ≥Wi 0(˜ A)|X=x)dν(x) =E(rm) Px(W i 0(Ax)≥Wi 0(˜ A))dν(x) =1 and in the same way P(W0(A) =X)=E(rm) Px(W0(Ax)=x) dν(x) =1. To show the implication (b) ⇒ (a), suppose that X/∈E(r m) on an F0-measurable set with ν() > 0. Since there are only finitely many players, this implies the existence of an F0-measurable set ˜ with ν( ˜ ) > 0such that some player ican improve his strategy to ˜ Axi for x∈˜ . Letting ˆ Ai:=Ai1˜ c(x) +˜ Axi1˜ (x), it follows that P(W i 0(A) ≥Wi 0(ˆ AiA−i)) =V Px(W i 0(A) ≥Wi 0(˜ AxiA−i))dν(x) < 1 contradicting the assumption that Ais a PPE. Appendix D: Auxiliary results In this appendix we provide some auxiliary results related to enforceability and pairwise identifiability. These results are needed in the proofs of the folk theorems as indicated in Figures 6and 7. They differ from the results in Fudenberg et al. (1994) in that we model the change of the signal’s distribution through a change of its drift. Lemma 16. An action profile αis enforceable if for every player ioneofbelowconditions holds: (i) The profile αhas individual full rank for player i. (ii) The action αiis a best response of player ito α−i. The enforceability condition (4)imposesnsystems of linear inequalities, one for each player i.BecauserankMi(α) ≤|Ai|−1, we cannot simply solve the system iby applying the left inverse of Mi(α), but we need to additionally exploit that the linear dependence among the columns of Gi(α) and Mi(α) is the same. Proof of Lemma 16.Fixaplayeriand suppose first that 1. is satisfied for action profile α.Letai∈Aibe an action with αi(ai)>0and enumerate Ai={ai 1ai Ki}such that ai=ai Kiis the last element. For the sake of brevity, denote by Mi j(α) the column of Mi(α) corresponding to action ai j. Because of the linear dependence among the columns of Mi(α),weobtain Mi Ki(α) =− Ki−1 j=1 α(ai j) α(ai Ki)Mi j(α)
Theoretical Economics 11 (2016) Folk theorem with imperfect public information 449 Figure 8. The right panel shows Vand V∗for the stage game with payoffs given in the table to the left. In the proof of the folk theorems, we only need that we can approximate extremal points of V∗and the minmax payoffs, which all lie in V. Condition 1 implies that there is no other linear dependence among the columns of Mi(α) and thus the d×(|Ai|−1)-dimensional submatrix ˜ Mi(α) consisting of the first |Ai|−1columns has full column rank. In particular, ˜ Mi(α) has a left inverse ˜ Mi L(α) and βi=Gi(α)˜ Mi L(α) 0 solves the system for player iwith equality. Indeed, Gi(α)˜ Mi L(α) 0Mi(α) =Gi(α)IKi−1Ki−1 j=1 α(ai j) α(ai Ki)˜ Mi L(α) ˜ Mi j(α) 00 =Gi 1(α)Gi Ki−1(α)− Ki−1 j=1 α(ai j) α(ai Ki)Gi j(α) whereweusedthat ˜ Mi L(α) ˜ Mi j(α) =ej. The claim under condition 1 follows since Gi Ki(α) =− Ki−1 j=1 α(ai j) α(ai Ki)Gi j(α) Under condition 2, βi=0solves the inequalities for player i. The following lemmas are the analogues of Lemmas 6.1–6.3 in Fudenberg et al. (1994) in our setting. Let Vdenote the set of payoffs achievable in mixed actions. While it may be strictly smaller than Vfor some games (see Figure 8), the extremal payoffs always correspond to pure action profiles and hence are contained in V. Lemma 17. Suppose that for every pair of players i,j, there exists a mixed action profile αij having ij-pairwise full rank. Then the set of payoffs Cof action profiles with pairwise full rank for all pairs of players is dense in V. Proof.LetE⊆(A)denote the set of mixed action profiles with pairwise full rank. By Lemma 6.2 of Fudenberg et al. (1994),Eis dense in (A). Since the map g:(A)→V is continuous and surjective, C⊇g(E) is dense in V.
450 Bernard and Frei Theoretical Economics 11 (2016) Lemma 18. Suppose that every pure action profile has individual full rank. Then for any ε>0and any player i, there exists an enforceable action profile αwith best response property for player iand |gi(α) −vi|<ε. Moreover, if every pure action is pairwise identifiable orthebestresponseofplayerito the minmax profile αi −iis unique, then the pair (α−ei) satisfies condition (ii) or (iii) of Lemma 5, respectively. Note that the first part of the statement is identical to Lemma 6.3 of Fudenberg et al. (1994). However, we also need that the resulting action profile leads to a locally uniform decomposition as in Lemma 5. Proof of Lemma 18.Fixaplayeriand let α−i idenote a minmax profile against player i. By assumption, (aia−i)has individual full rank for every ai∈Aiand every a−i∈A−i. Therefore, similar as in the proof of Lemma 6.2 in Fudenberg et al. (1994), one can find a sequence of profiles (α−i (k))k≥1converging to α−i i,suchthat(aiα−i (k))has individual full rank for every ai∈Aiand all k. Let ai kbe a best response for player ito α−i (k). The profiles (ai kα−i (k))are enforceable orthogonal to −eiby Lemmas 16 and 3.Letai∈Aibe an accumulation point of (ai k)k≥1 and choose a subsequence (kν)ν≥1such that ai kν=aifor all ν∈N.Observethataiis also a best response to α−i ibecause gi(aiα−i i)=lim ν→∞gi(aiα−i (kν))≥lim ν→∞gi(˜ aiα−i (kν))=gi(˜ aiα−i i) hence ai∈argmax gi(·α−i i)and gi(aiα−i i)=vi. Therefore, for any ε>0we can find ν large enough such that |gi(aiα−i (kν))−vi|<ε. Under condition 1, (α−i (k))k≥1canbechoseninawaythat(aiα−i (k))has pairwise full rank for every ai∈Aiand all k. Therefore, ((aiα−i (kν))−ei)satisfies the second condition of Lemma 5. Under condition 2, it follows from multilinearity that there exists a νlarge enough such that aiis also a unique best response to α−i (kν). Therefore, condition (iii) of Lemma 5 is fulfilled for the pair ((aiα−i (kν))−ei). Definition 11. An action profile αPareto-dominates aprofile ˜αif gi(α) ≥gi(˜α) for every player iand gj(α) > gj(˜α) for at least one player j. An action profile is Pareto-efficient if it is not Pareto-dominated by any other action profile. Lemma 19. Suppose that gi(a) =bi(ai)m(a) −ci(ai). Then any Pareto-efficient pure action profile is enforceable. Proof. Fix a Pareto-efficient pure action profile a∈A. Because its payoff is on the “upper right” boundary of V, there exists a direction N∈Rnwith Ni>0for every isuch that g(a) =argmaxv∈VNv.Thenβwith row vectors βi:=j=ibj(aj)Nj/Nienforces a.