scieee AI-readable full text Open interactive document viewer

Uncertain identification

Giacomini, Raffaella,Kitagawa, Toru,Volpicella, Alessio

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Giacomini, Raffaella; Kitagawa, Toru; Volpicella, Alessio Article Uncertain identification Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Giacomini, Raffaella; Kitagawa, Toru; Volpicella, Alessio (2022) : Uncertain identification, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 13, Iss. 1, pp. 95-123, https://doi.org/10.3982/QE1671 This Version is available at: https://hdl.handle.net/10419/296270 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/ Quantitative Economics 13 (2022), 95–123 1759-7331/20220095 Uncertain identification Raffaella Giacomini Department of Economics, University College London and Federal Reserve Bank of Chicago Toru Kitagawa Department of Economics, University College London and Department of Economics, Kyoto University Alessio Volpicella School of Economics, University of Surrey Uncertainty about the choice of identifying assumptions is common in causal studies, but is often ignored in empirical practice. This paper considers uncertainty over models that impose different identifying assumptions, which can lead to a mix of pointand set-identified models. We propose performing inference in the presence of such uncertainty by generalizing Bayesian model averaging. The method considers multiple posteriors for the set-identified models and combines them with a single posterior for models that are either point-identified or that impose nondogmatic assumptions. The output is a set of posteriors (postaveraging ambiguous belief ), which can be summarized by reporting the set of posterior means and the associated credible region. We clarify when the prior model probabilities are updated and characterize the asymptotic behavior of the posterior model probabilities. The method provides a formal framework for conducting sensitivity analysis of empirical findings to the choice of identifying assumptions. For example, we find that in a standard monetary model one would need to attach a prior probability greater than 0.28 to the validity of the assumption that prices do not react contemporaneously to a monetary policy shock, in order to obtain a negative response of output to the shock. Keywords. Partial identification, sensitivity analysis, model averaging, Bayesian robustness, ambiguity. JEL classification. C11, C32, C52. Raffaella Giacomini: [email protected] Toru Kitagawa: [email protected] Alessio Volpicella: [email protected] We would like to thank Frank Kleibergen, José Montiel-Olea, Ulrich Müller, John Pepper, James Stock, three anonymous referees, and participants at numerous seminars and conferences for valuable comments and beneficial discussions. Financial support from ERC grants (numbers 536284 for R. Giacomini and 715940 for T. Kitagawa) and the ESRC through the ESRC Centre for Microdata Methods and Practice (CeMMAP) (grant number RES-589-28-0001) is gratefully acknowledged. This paper is based on and develops the first chapter of the Ph.D. thesis, Volpicella (2020). The views expressed are those of the authors and do not necessarily reflect those of the Federal Reserve Bank of Chicago or the Federal Reserve System. ©2022 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE1671 96 Giacomini, Kitagawa, and Volpicella Quantitative Economics 13 (2022) 1. Introduction The choice of identifying assumptions is the crucial step that allows researchers to draw causal inferences using observational data. This is often a controversial choice, and there can be uncertainty about which assumptions to impose from a menu of plausible ones, but this uncertainty and its effects on inference are typically ignored in empirical work. This paper proposes a formal framework for sensitivity analysis via Bayesian model averaging in the presence of uncertain identification, which we characterize as uncertainty over a class of models that impose different sets of identifying assumptions. The class of models can include ones where parameters are set-identified, which occurs when the assumptions are underidentifying or take the form of inequality restrictions. For these models, we advocate adopting the multiple-prior approach of Giacomini and Kitagawa (2021). In our context, the approach has the additional advantage of isolating the component of each model that depends on the identifying restrictions, making it possible, for example, to compare models that only differ in the restrictions they impose. The paper makes both a methodological and a theoretical contribution. The methodological contribution is to extend Bayesian model averaging/selection to allow for models characterized by multiple priors (associated here with set identification). The theoretical contribution is to clarify how the different components of the models affect inference in terms of model averaging/selection in finite samples and asymptotically. There are several examples in economics where empirical researchers face uncertainty about identifying assumptions that lead to pointor set-identification of a common causal parameter of interest. The first is macroeconomic policy analysis based on structural vector autoregressions (SVARs), where assumptions include causal ordering restrictions (Sims (1980)) and long-run neutrality restrictions (Blanchard and Quah (1993)). Subsets of these assumptions deliver set-identified impulse-responses, as do sign restrictions (Canova and Nicolo (2002), Faust (1998), and Uhlig (2005)). The second example is microeconometric causal effect studies with assumptions such as selection on observables (Ashenfelter (1978)andRosenbaum and Rubin (1983)), selection on observables and unobservables (Altonji, Elder, and Taber (2005)), exclusion and monotonicity restrictions in instrumental variables methods (Imbens and Angrist (1994), yielding set-identification of the average treatment effect), and monotone instrument assumptions (Manski and Pepper (2000), also yielding set-identification). The third example is missing data with assumptions such as missing at random, Bayesian imputation (Rubin (1987)), and unknown missing mechanism (Manski (1989), yielding set-identification). Finally, estimation of structural models with multiple equilibria relies on assumptions about the equilibrium selection rule, with different assumptions (or lack thereof) delivering pointor set-identification (e.g., Bajari, Hong, and Ryan (2010), Beresteanu, Molchanov, and Molinari (2011), and Ciliberto and Tamer (2009)). The common practice in empirical work is to report results based on what is deemed the most credible set of identifying assumptions, or sometimes, based on a small number of alternative assumptions, viewed as an informal sensitivity analysis. Our method provides a formal framework for investigating the sensitivity of empirical findings to Quantitative Economics 13 (2022) Uncertain identification 97 specific identifying assumptions and/or for aggregating results based on different identifying assumptions, which can be more practical than reporting separate results when there are many restrictions.1 The idea of model averaging has a long history in econometrics and statistics since the pioneering works of Bates and Granger (1969)andLeamer (1978). The literature has considered Bayesian approaches (see, e.g., Hoeting, Madigan, Raftery, and Volinsky (1999)), frequentist approaches (Hansen (2007,2014), Hjort and Claeskens (2003), Liu and Okui (2013)), and hybrid approaches (Hjort and Claeskens (2003), Kitagawa and Muris (2016), and Magnus, Powell, and Prüfer (2010)), but none of them allows for setidentification/multiple priors in any candidate model. This paper takes a Bayesian perspective. The standard approach to Bayesian model averaging delivers a single posterior that is a mixture of the posteriors of the models, with weights equal to the posterior model probabilities.2This approach could in principle be extended to our context if one could obtain a single posterior for every model, including set-identified ones. Assuming a single prior under set identification is however problematic from a robustness viewpoint as the choice of a single prior can lead to spuriously informative posterior inference for the object of interest (Baumeister and Hamilton (2015)). The severity of the problem is magnified by the fact that the effect of the prior choice persists asymptotically, unlike in the case of point-identified models (Moon and Schorfheide (2012), Poirier (1998), among others). The key innovation of our approach to Bayesian model averaging is that we do not assume availability of a single posterior for the set-identified models. Rather, we allow for multiple priors (an ambiguous belief ) within the set-identified models (as in Giacomini and Kitagawa (2021)), and then combine the corresponding multiple posteriors with single posteriors for models that are either point-identified or that impose nondogmatic identifying assumptions in the form of a Bayesian prior for the structural parameters (as in Baumeister and Hamilton (2015)). The output of the procedure is a set of posteriors (post-averaging ambiguous belief ), that are mixtures of the single posteriors and any element of the set of multiple posteriors, with weights equal to the posterior model probabilities. To summarize and visualize the post-averaging ambiguous belief, one can report the set of posterior quantities (e.g., the mean or median) and the associated credible region (an interval to which any posterior in the class assigns a certain credibility level), which are easy to compute in practice. The method proposed in this paper provides a formal framework for conducting sensitivity analysis of causal inferences to the choice of identifying assumptions. First, one can perform reverse-engineering exercises that compute the minimal prior probability one would need to attach to a set of identifying assumptions in order for the 1For example, the SVAR literature often considers models with a large number of sign restrictions (e.g., Amir-Ahmadi and Uhlig (2015), Korobilis (2020), Furlanetto, Ravazzolo, and Sarferaz (2019), and Matthes and Schwartzman (2019)). 2When a constrained model is a lower dimensional submodel of a large model, performing inference conditional on the constrained model may suffer from the Borel paradox; see, for example, Drèze and Richard (1983). Bayesian model averaging offers a practical way to avoid the Borel paradox in such context. 98 Giacomini, Kitagawa, and Volpicella Quantitative Economics 13 (2022) averaging to obtain a certain conclusion (e.g., that the set of posterior means for the impact response of output to a monetary policy shock is contained in the negative real half-line). This exercise has a similar motivation as the breakdown frontier analysis in Horowitz and Manski (1995)andMasten and Poirier (2020). Second, when a setidentified model nests a point-identified model, our method can be used to assess the posterior sensitivity in the point-identified model with respect to perturbations of the prior in the direction of relaxing some of the point-identifying assumptions. This exercise can be seen as an example of the -contamination sensitivity analysis developed in Huber (1973)andBerger and Berliner (1986). Our approach to sensitivity analysis therefore differs from and complements the approaches proposed by Giacomini, Kitagawa, and Uhlig (2019)andHo (2019), which specify the class of priors as a Kullback–Leibler neighborhood of a benchmark prior. Our method can also be viewed as bridging the gap between pointand setidentification. When focusing solely on a point-identified model, a researcher who is not fully confident about the choice of identifying assumptions may doubt the robustness of the conclusions. On the other hand, discarding some of the point-identifying assumptions and reporting estimates of the identified set may appear “excessively agnostic,” and often results in uninformative conclusions. Our averaging procedure reconciles these two extreme representations of the posterior beliefs by exploiting the prior weights that one can assign to alternative sets of identifying assumptions. This paper contributes to the growing literature on Bayesian inference for partially identified models (Giacomini and Kitagawa (2021), Kline and Tamer (2016), Moon and Schorfheide (2012)). We follow the multiple-prior approach to model the lack of knowledge within the identified set as in Giacomini and Kitagawa (2021). When a set-identified model is the only model considered, the set of posteriors generated by the approach provides posterior inference for the identified set. When there is uncertainty about the identifying assumptions, however, the usual definition of identified set is not available without conditioning on the model. The multiple prior viewpoint has an advantage in this case since the set of posteriors has a well-defined subjective interpretation even in the presence of model uncertainty. The paper makes two main analytical contributions to the literature on Bayesian model selection and averaging. First, we clarify under which conditions the prior model probabilities can be updated by data. We show that the updating occurs if some models are “distinguishable” for some distribution of data and/or the priors for the reducedform parameters differ across models. Second, we investigate the asymptotic properties of the posterior model probabilities and of the averaging method. We show that, when only one model is consistent with the true distribution of the data, our method asymptotically assigns probability one to it. When multiple models are observationally equivalent and “not falsified” at the true data generating process, the posterior model probabilities asymptotically assign nontrivial weights to them. We clarify what part of the prior input determines the asymptotic posterior model probabilities in such case. The consistency property of Bayesian model selection has been well studied in the statistics literature (e.g., Claeskens and Hjort (2008) and references therein), but there is no discussion about the asymptotic behavior of posterior model probabilities when the models differ Quantitative Economics 13 (2022) Uncertain identification 99 in terms of the identifying assumptions but can be observationally equivalent in terms of their reduced form representations. These new results therefore could be of separate interest. The empirical application in this paper considers SVAR analysis with uncertainty over the classes of identifying assumptions typically used in empirical work. The choice of identifying assumptions has often been a source of controversy in this literature, and researchers have differing opinions about their credibility. To our knowledge, little work has been done on multimodel inference in the SVAR literature, and the methods proposed in this paper could therefore prove helpful in reconciling the controversies about the identifying assumptions that are widespread in this literature. As an example, the empirical application documents the high sensitivity of the conclusion in standard monetary SVARs that output decreases after a contractionary monetary policy shock to the choice of identifying assumptions. The remainder of the paper is organized as follows. Section 2illustrates the motivation and the implementation of the method in the context of a simple model. Section 3 presents the formal analysis in a general framework and provides a computational algorithm to implement the procedure. Section 4applies our method to impulse response analysis in monetary SVARs. The Appendix in the Online Supplementary Material (Giacomini, Kitagawa, and Volpicella (2021)) contains proofs and details about computation. 2. Illustrative example We present the key ideas and the implementation of the method in a price-quantity static model, subject to common types of identifying assumptions. The model is Aqt pt=d t s t,A=a11 a12 a21 a22,t=1, ,T, (2.1) where (qt,pt)are price and quantity of a certain good/service in a given market and (d t,s t)is an i.i.d. normally distributed vector of demand and supply shocks with variance-covariance the identity matrix. Ais the structural parameter and the contemporaneous impulse responses are elements of A−1. For example, in the labor market (qt,pt)can be replaced by employment and wages, respectively. The reduced-form model is indexed by , the variance-covariance matrix of (qt,pt), which satisfies =A−1(A−1). Denote its lower triangular Cholesky decomposition with nonnegative diagonal elements by tr =σ11 0 σ21 σ22 with σ11 ≥0andσ22 ≥0, and define the reduced form parameter as φ=(σ11,σ21,σ22 )∈=R+×R×R+.3Let the mapping from the structural parameter to the reduced-form parameter be denoted by φ=g(A). Suppose the object of interest is the response of the first variable to a unit positive shock in the first variable, α≡(1, 1)-element of A−1. Without identifying assumptions, the structural parameter is set-identified since knowledge of the reduced-form parameter φcannot uniquely pin down the structural parameter (φ=g(A)is a many-to-one 3The positive semidefiniteness of does not constrain the value of φother than σ11 ≥0 and σ22 ≥0. 100 Giacomini, Kitagawa, and Volpicella Quantitative Economics 13 (2022) mapping). Imposing assumptions can lead to a set or a point for α, depending on the type and number of assumptions. A Bayesian model is the combination of a likelihood and a prior input. The prior input can be either a single prior or multiple priors. In point-identified models, the prior input is a single prior for the structural parameter A, which implies the prior for the reduced-form parameter φ. In set-identified models, one could either specify a single prior for A(e.g., as a way of imposing nondogmatic identifying assumptions) or consider multiple priors as in Giacomini and Kitagawa (2021). In the latter case, a model is the combination of a likelihood, a single prior for the reduced-form parameter φ(which is revised) and multiple priors for A|φ(which are not revised).4 The division that we introduce in the paper is between single-prior models (which could be pointor set-identified) and multiple-prior models (which are always setidentified). We now illustrate how this interplays with identifying assumptions in two examples. 2.1 Dogmatic identifying assumptions First, consider dogmatic identifying assumptions, which are equality or inequality restrictions on (functions of) the structural parameter that hold with probability one. Scenario 1: Candidate models •Model Mp(point-identified): The demand is inelastic to price, a12 =0. •Model Ms(set-identified): The price elasticity of demand is nonpositive, a12 ≥0, and the price elasticity of supply is nonnegative, a21 ≤0. Model Mprestricts Ato be lower-triangular, as in the classical causal ordering assumptions of Sims (1980)andBernanke (1986). Combined with the sign normalization restrictions requiring the diagonal elements of Ato be nonnegative, the assumption implies that the impulse responses can be identified by A−1=tr. The parameter of interest is α=αMp(φ)≡σ11. Model Msimposes sign restrictions that only set-identify α. The Appendix in the Online Supplementary Material shows that the identified set for αis ISα(φ)≡⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ σ11 cosarctanσ22 σ21 ,σ11,forσ21 >0, 0, σ11 cosarctan−σ21 σ22 ,forσ21 ≤0. (2.2) Note that the identified set is nonempty for any φ. Hence, models Mpand Msare observationally equivalent at any φ∈and neither of them is falsifiable, that is, for any 4See Giacomini and Kitagawa (2021) for a discussion about and motivation for assuming a single prior for φ. An additional advantage of this assumption in the context of model selection is that it allows one to isolate the component of the model that depends on the identifying restrictions. This enables one, for example, to compare models that only differ in the restrictions they impose. Quantitative Economics 13 (2022) Uncertain identification 101 φ∈in both models there exists a structural parameter Athat satisfies the identifying assumptions.5 We start by specifying a prior for φin each model. Given the observational equivalence of the two models, it might be reasonable to specify the same prior: πφ|Mp=πφ|Ms=˜πφ, (2.3) where ˜πφis a proper prior, such as the one induced by a Wishart prior on .Thesame prior for φin observationally equivalent models leads to the same posterior: πφ|Mp,Y=πφ|Ms,Y=˜πφ|Y. (2.4) In model Mp, the posterior for φimplies a unique posterior for α,πα|Mp,Y,viathe mapping α=αMp(φ).InmodelMs, on the other hand, the posterior for φdoes not yield a unique posterior for α, since the mapping in (2.2) is generally set-valued. Following Giacomini and Kitagawa (2021), we formulate the lack of prior knowledge by considering multiple priors (ambiguous belief). Formally, given the single prior πφ|Ms,weformthe class of priors for Aby admitting arbitrary conditional priors for Agiven φ,aslongas they are consistent with the identifying assumptions: A|Ms≡πA|Ms= πA|Ms,φdπφ|Ms:πA|Ms,φAsign ∩g−1(φ)=1, πφ|Ms-a.s., where Asign ={A:a12 ≥0, a21 ≤0, diag(A)≥0}is the set of structural parameters that satisfy the sign restrictions and the sign normalizations and g−1(φ)is the set of observationally equivalent structural parameters given the reduced-form parameter φ. Since the likelihood depends on the structural parameter only through the reducedform parameter, applying Bayes’ rule to each prior in the class only updates the prior for φ, and thus leads to the following class of posteriors for A: A|Ms,Y≡πA|Ms,Y= πA|Ms,φdπφ|Ms,Y:πA|Ms,φAsign ∩g−1(φ)=1, πφ|Ms-a.s.. (2.5) Marginalizing the posteriors in A|Ms,Yto αleads to the class of α-posteriors: α|Ms,Y≡πα|Ms,Y=˜  πα|Ms,φdπφ|Ms,Y:πα|Ms,φ(ISα(φ)) =1, πφ|Ms-a.s.. (2.6) We view this class as a representation of the posterior uncertainty about αin the setidentified model. The class contains any α-posterior that assigns probability one to the identified set, and it represents the lack of belief therein in terms of Knightian uncertainty (ambiguity). This is a key departure from the standard approach to Bayesian 5When σ21 >0, the point-identified αin model Mpis the upper bound of the identified set in model Ms,whereaswhenσ21 <0, the identified set in model Msdoes not contain the point-identified α.Thisis because in model Mpwe have a12 =− σ21 σ11σ22 , which is positive if σ21 <0, meaning that the point-identifying assumptions a12 =0 and σ21 <0 are not compatible with the restriction a21 ≤0. 102 Giacomini, Kitagawa, and Volpicella Quantitative Economics 13 (2022) model averaging, which requires a single posterior for all models, including those where the parameter is set-identified. Suppose that the researcher’s prior uncertainty over the two models can be represented by prior probabilities πMp∈[0, 1]for model Mpand (1−πMp)for model Ms. Our proposal is to combine the single posterior for αin model Mpand the set of posteriors for αin model Msaccording to the posterior model probabilities πMp|Yand πMs|Y (the posterior model probability for model Msdepends only on the single prior for the reduced-form parameter, so it is unique in spite of the multiple priors for the structural parameter). The combination delivers a class of posteriors α|Y,thepost-averaging ambiguous belief : α|Y={πα|Mp,YπMp|Y+πα|Ms,YπMs|Y:πα|Ms,Y∈α|Ms,Y}. (2.7) As we show in Section 4.1, our proposal can be interpreted as applying Bayes’ rule to each prior in a class that has the form of an -contaminated class of priors (Berger and Berliner (1986)). A key result of the paper is to establish conditions under which the prior model probabilities are updated by the data, which we show occurs when the models are “distinguishable” for some reduced-form parameter values and/or they specify different priors for φ(see Lemma 3.1 below). In the current scenario, the two models are indistinguishable, so the prior model probabilities are not updated if they use a common φ-prior. In practice, we recommend reporting as the output of the procedure the postaveraging set of posterior means or quantiles of α|Yand its associated robust credible region with credibility γ∈(0, 1), defined as the shortest interval that receives posterior probability at least γfor every posterior in α|Y. Proposition 3.1 shows that the set of posterior means is the weighted average of the posterior mean in model Mpand the set of posterior means in model Ms: inf πα|Y∈α|Y Eα|Y(α),sup πα|Y∈α|Y Eα|Y(α) =πMp|YEα|Mp,Y(α)+πMs|YEφ|Ms,Yl(φ),Eφ|Ms,Yu(φ), (2.8) where (l(φ),u(φ)) are the lower and upper bounds of the nonempty identified set for α shown in (2.2), a+b[c,d]stands for [a+bc,a+bd],andEφ|Ms,Y(·)denotes the posterior mean with respect to πφ|Ms,Y=˜πφ|Y. Since the set of posterior means can be viewed as an estimator for the identified set in model Ms, our procedure effectively shrinks the estimate of the identified set in the set-identified model toward the point estimate in the point-identified model, with the amount of shrinkage determined by the posterior model probabilities. The robust credible region for αwith credibility γcan be computed as follows. We first draw z1,,zGrandomly from a Bernoulli distribution with mean πMp|Yand then generate g=1, ,Grandom draws of the “mixture identified set” for αaccording to ISmix α(φg)=α(φg),φg∼πφ|Mp,Y=˜πφ|Yif zg=1, l(φg),u(φg),φg∼πφ|Ms,Y=˜πφ|Yif zg=0. (2.9) Quantitative Economics 13 (2022) Uncertain identification 109 (ii) For any measurable subset Hin R,the lower and upper bounds of the posterior probabilities on {α∈H}in the class α|Y(the lower and upper posterior probabilities of α|Y)are inf πα|Y∈α|Y πα|Y(H)= Mp∈Mp πα|Mp,Y(H)πMp|Y + Ms∈Ms πφMs|Y,MsISαφMs|Ms⊂H·πMs|Y, sup πα|Y∈α|Y πα|Y(H)= Mp∈Mp πα|Mp,Y(H)πMp|Y + Ms∈Ms πφMs|Y,MsISαφMs|Ms∩H= ∅·πMs|Y. If ISα(φMs|Ms)is a connected interval at every reduced-form parameter value, then we can view [EφMs|Y,Ms[l(φMs|Ms)],EφMs|Y,Ms[u(φMs|Ms)]] as an estimator of the identified set in model Ms. We can thus interpret the set of post-averaging posterior means as the weighted Minkowski sum of the Bayesian point estimators in the point-identified models and the identified set estimators in the set-identified models. The second claim of the proposition provides an analytical expression for the lower probability of α|Yas a mixture of the containment functionals of the random sets, which in turn can be viewed as the containment functional of the mixture random sets Pr(ISmix α⊂A),whereISmix αis generated according to M∼Multinomial{πM|Y}M∈M, ISmix α={α},α|(Mp,Y)∼πα|Mp,Y,forMp∈Mp, ISαφMs|Ms,φMs|(Ms,Y)∼πφMs|Ms,Y,forMs∈Ms. (3.10) This way of interpreting the lower probability of α|Ysimplifies its computation and justifies the algorithm presented in (2.9). 3.4 Computation To report the set of posteriors based on the analytical expressions in Proposition 3.2, we need to compute (i) the posterior model probabilities (equivalently, the marginal likelihood in each M∈M), (ii) the posterior for αfor each single-prior model, and (iii) the identified set ISα(φMs|Ms)and the posterior for φMsfor each multiple-prior model. Estimation of the single-prior models in (ii) is standard, and we assume some suitable posterior sampling algorithm is applicable to obtain Monte Carlo draws of α∼πα|Mp,Y. For (i), efficient and reliable algorithms to compute the marginal likelihood are available in the literature; for example, see Chib and Jeliazkov (2001), Geweke (1999), and Sims, Waggoner, and Zha (2008). When all the models admit an identical reduced-form, computing the marginal likelihoods is not necessary since the posterior model probabilities depend only on the posterior-prior plausibility ratios OM. 110 Giacomini, Kitagawa, and Volpicella Quantitative Economics 13 (2022) In each multiple-prior model, the posterior-prior plausibility ratio OMscan be computed by plugging in numerical approximations for the prior and posterior probabilities of the non-emptiness of the identified set into (3.4). The denominator of OMsis computed by drawing many φ’s from the prior ˜πφand computing the fraction of draws that yield nonempty identified sets. The numerator of OMsis computed similarly except that the φ’s are drawn from the posterior ˜πφ|Y. Whether checking the nonemptiness of ISα(φ|Ms)is simple or not depends on the application. In the application in Section 4to SVARs with sign restrictions, we consider two ways to check the nonemptiness of ISα(φ|Ms). The first (Algorithm A.1 in the Appendix of the Online Supplementary Material) builds on Algorithm 1 of Giacomini and Kitagawa (2021) and assesses nonemptiness based on the Monte Carlo draws of the impulse responses. The second approach (Algorithm A.2 in the Appendix of the Online Supplementary Material), which is novel in the literature and can be of independent interest, exploits the analytical features of the identifying restrictions in sign restricted SVARs. See the Appendix for the details of these algorithms. Monte Carlo draws of the lower and upper bounds of the identified set in model M∈Mscan be obtained by first drawing φ’s from the posterior ˜πφ|Y, then retaining the draws of φthat yield a nonempty ISα(φ|Ms), and computing the corresponding l(φ|Ms)and u(φ|Ms). Their sample averages approximate Eφ|Ms,Y(l(φ|Ms)) and Eφ|Ms,Y(u(φ|Ms)). Implementation of this procedure requires computability of the lower and upper bounds of the identified set for each φ. In the SVAR application of Section 4,wecomputel(φ|Ms)and u(φ|Ms)by numerical optimization. Utilizing the mixture random set representation shown in (3.10), we can use the following algorithm to approximate the lower posterior probability. Algorithm 3.1. Step 1: Draw a model M∈Mfrom a multinomial distribution with parameters (πM|Y:M∈M). Step 2: If the drawn Mbelongs to Mp, then draw α∼πα|M,Yand set ISmix α={α}(a singleton). If the drawn Mbelongs to Ms,drawφM∼πφ|M,Yand set ISmix α= ISα(φM|M).15 Step 3: Repeat Steps 1 and 2 many (G) times and obtain Gdraws of ISmix α:ISmix α,1 ,, ISmix α,G. Step 4: Let [lmix g,umix g]be the lower and upper bounds of ISmix α,g,g=1, ,G,where lmix g=umix gif ISmix α,gis a singleton (i.e., gth draw of Mbelongs to Mp). Approximate the mean bounds of the post-average posterior class by inf πα|Y∈α|Y Eα|Y(α)=1 G G  g=1 lmix g,sup πα|Y∈α|Y Eα|Y(α)=1 G G  g=1 umix g. (3.11) 15Note that since πφ|M,Yis supported only on the set of φ’s yielding a nonempty identified set, ISα(φ|M) computed subsequently is nonempty. Quantitative Economics 13 (2022) Uncertain identification 111 Approximate the lower probability of the post-averaging posterior class at H⊂ Rby inf πα|Y∈α|Y πα|Y(H)≈1 G G  g=1 1ISmix α,g⊂H. (3.12) The draws of ISmix αobtained in Steps 1–3 in Algorithm 3.1 are also useful for constructing robust credible regions, which are intervals that attain a certain level of credibility uniformly over the posterior class. Applying Proposition 1 of Giacomini and Kitagawa (2021) to the Monte Carlo draws of ISmix α, we can easily approximate the shortest robust credible region for α. 3.5 Asymptotic properties This section analyzes the asymptotic properties of our method. The method is finitesample exact (up to Monte Carlo approximation errors), but the asymptotic analysis can be valuable to understand what aspects of the prior input, if any, remain influential in large samples. In this section, we make the sample size explicit by denoting a size nsample by Yn. We assume that at least one model is correctly specified, so that the data-generating process is given by p(Yn|φtrue ),whereφtrue ∈is the true reduced-form parameter value. We denote the unconstrained maximum likelihood estimator for φby ˆ φ≡ argmaxφ∈p(Yn|φ)and the true probability law of the sampling sequence {Yn:n= 1, 2, }by PY∞|φtrue . We impose the following regularity assumptions. Assumption 3.2. (i) Madmits an identical reduced-form (Definition 3.1)and every M∈Msatisfies either one of the following conditions: (A) Mcontains φtrue in its interior. (B) c Mcontains φtrue in its interior. MA,denoting the set of models satisfying condition (A), is nonempty. (ii) Let ln(φ)≡n−1logp(Yn|φ).There exist an open neighborhood Bof φtrue and n0≥ 1, such that for any {Yn:n=n0,n0+1, },ln(·)is third-time differentiable with the third-order derivatives bounded uniformly on B. (iii) Let Hn(ˆ φ)≡−∂2ln(ˆ φ) ∂φ∂φ .Hn(ˆ φ)is a positive definite matrix and lim inf n→∞ detHn(ˆ φ)>0, with PY∞|φtrue -probability one. 112 Giacomini, Kitagawa, and Volpicella Quantitative Economics 13 (2022) (iv) For any open neighborhood Bof φtrue, lim sup n→∞ sup φ∈\Bln(φ)−ln(φtrue )<0 holds with PY∞|φtrue -probability one. (v) For every M∈M,πφ|Mhas probability density fφ|M(φ)≡dπφ|M dφ (φ)with respect to the Lebesgue measure on Mand fφ|M(φ)is continuously differentiable with a uniformly bounded derivative.For every M∈MA,fφ|M(φtrue )>0. Assumption 3.2(i) implies that none of the models has φtrue on the boundary of its reduced-form parameter space. Assumptions 3.2(iii) and (iv) impose regularity conditions that imply almost sure consistency of ˆ φ. Assumptions 3.2(ii) and (v) allow an application of the Laplace method to approximate the large sample marginal likelihood. Assumptions similar to Assumptions 3.2(ii)–(v) appear in Kass, Tierney, and Kadane (1990) in their validation of the higher-order expansion of the marginal likelihood. The next proposition derives the limits of the posterior model probabilities. Proposition 3.3. (i) Suppose Assumption 3.2 holds.Then πM|Y∞≡lim n→∞ πM|Yn=⎧ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎩ fφ|M(φtrue)·πM  M∈MA fφ|M(φtrue)·πM ,for M∈MA, 0, for M/∈MA. (3.13) with PY∞|φtrue -probability one. (ii) Suppose that Assumption 3.2 holds and a prior for φgiven Mis constructed according to (3.2)with a proper prior ˜πφ.If ˜πφ(M)>0for all M∈M, πM|Y∞=⎧ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎩ ˜πφ(M)−1·πM  M∈MA ˜πφ(M)−1·πM ,for M∈MA, 0, for M/∈MA. (3.14) with PY∞|φtrue -probability one. (iii) Under the assumptions of Lemma 3.1(iii), πM|Y∞=πMholds for every M∈Mfor any sampling sequence {Yn:n=1, 2, }. The proposition clarifies the large sample behavior of the posterior model probabilities when the models admit an identical reduced-form. First, it shows that our procedure asymptotically screens out models whose identifying assumptions are misspecified, M/∈MA, as their posterior probabilities converge to zero. If there is only one model consistent with the data-generating process, asymptotically it has probability one. Second, if MAcontains multiple models, their asymptotic probabilities depend Quantitative Economics 13 (2022) Uncertain identification 113 on the prior model probabilities and on the φ-priors evaluated at φtrue. This implies that the post-averaging posterior is asymptotically sensitive to the choices of φ-priors and prior model probabilities when multiple models are observationally equivalent at φtrue. Third, when the φ-priors are common, the asymptotic model probabilities are proportional to the reciprocal of the prior probability that the data is consistent with the identifying assumptions. Hence, the asymptotic posterior model probabilities are higher for more observationally restrictive models, that is, if M1⊂M2for M1,M2∈MA,we have πM1|Y∞≥πM2|Y∞. This result is in line with the principle of parsimony (Ockham’s razor)— we should prefer a more parsimonious model among those that explain the data equally well.16 3.6 Discussion We discuss how our method relates to the literature on -contaminated class of priors and to a hierarchical Bayesian way to bridge the gap between structural and reducedform models. Our method can be directly linked to performing robust Bayes’ analysis using an - contaminated class of priors (Huber (1973), Berger and Berliner (1986)). Consider the case of one single-posterior model and one multiple-posterior model, M={Mp,Ms} that share the same parameterization of the structural model and where the likelihood for the common structural parameters θdoes not depend on the model. Given (πMp,πMs),πθ|Mp,andθ|Msas in (3.6), consider the set of priors for θconstructed by marginalizing θ,Mof Proposition 3.1 to θ, θ≡{πθ=πθ|MpπMp+πθ|MsπMs:πθ|Ms∈θ|Ms}. (3.15) A general formulation of an -contaminated class of priors is given by  θ≡πθ=(1−)π0 θ+qθ:qθ∈Q, (3.16) where 0 ≤≤1 is a prespecified constant, π0 θis a benchmark prior for θ,andQis a set of priors of θ. Following Berger and Berliner (1986), is interpreted as the amount of contamination, qθcaptures how π0 θdiffers from the most credible prior and Qis the set of possible departures. The prior input of our procedure in (3.15) has the same form as the -contaminated class of priors (3.16)—θis an -contaminated class of priors where the benchmark prior is from the single-prior (point-identified) model π0 θ=πθ|Mp,the amount of contamination is the prior model probability assigned to the set-identified model =πMsand Qcorresponds to the multiple priors for the set-identified model θ|Ms. This clarifies a robust Bayes interpretation of our method: If the point-identified model is a possibly misspecified benchmark, averaging it with the set-identified model with weight πMscan be interpreted as performing sensitivity analysis by contaminating 16For instance, in a SVAR, a model point-identified by equality restrictions is not observationally restrictive, while a model set-identified by sign restrictions can be observationally restrictive. If the φ-priors satisfy (3.2) and the models are observationally equivalent at φtrue, then, relative to the prior model weights, the sign-restricted model receives a larger weight than the point-identified model in large samples. 114 Giacomini, Kitagawa, and Volpicella Quantitative Economics 13 (2022) the prior of the point-identified model by an amount πMsin every possible direction subject to the set-identifying assumptions. Our method could be viewed as a way to bridge the gap between structural and reduced-form models, for example, as an alternative to the hierarchical Bayesian approach of, for example, (Del Negro and Schorfheide (2004)), in which the structural parameters in a DSGE model act as hyperparameters of a prior for SVAR parameters. The two approaches differ in several ways. First, the hierarchical Bayesian approach always leads to a single posterior for the parameter, even if it is not identified in the SVAR model. If the parameter is not identified, this means that the priors have some part that is unrevisable by the data, leading to posterior sensitivity. In contrast, our procedure would classify the DSGE model as a single-prior model and the set-identified SVAR as a multiple-prior model, thus removing sensitivity to the choice of prior. Second, in the hierarchical Bayesian approach the prior confidence assigned to the structural model is the tightness of the prior predicted by the DSGE model, while in our procedure it is the model probability. It is however important to distinguish the notions of confidence in the two approaches, since the former is in terms of Bayesian probabilistic uncertainty while the latter is in terms of Knightian uncertainty. 4. Empirical application We illustrate our method in the context of a conventional monetary SVAR for the federal funds rate it, real output growth gdptand inflation πt,asinAruoba and Schorfheide (2011). The model has three lags (as selected by the HQ information criterion). Following Definition 3 in Giacomini and Kitagawa (2021), we order the variables so that we can verify the conditions guaranteeing convexity of the identified set using their Proposition B.1: A0yt=c+ 3  j=1 Ajyt−j+t,fort=1, ,T, (4.1) where yt=(it,gdpt,πt)and A0=⎛ ⎜ ⎝ a11 a12 a13 a21 a22 a23 a31 a32 a33 ⎞ ⎟ ⎠. (4.2) Assume t=(mp t,d t,s t)are i.i.d. normally distributed with mean zero and variancecovariance the identity matrix I3. The first equation in (4.1) is interpreted as a monetary policy function, while the second and third represent aggregate demand (AD) and aggregate supply (AS), respectively. Thus, mp t,d t,ands tare monetary policy-, aggregate demand-, and aggregate supply shocks, respectively. The data are quarterly observations from 1965:1 to 2005:1 from the FRED2 database. The reduced-form VAR is yt=b+ 3  j=1 Bjyt−j+ut, (4.3) Quantitative Economics 13 (2022) Uncertain identification 115 where b=A−1 0c,Bj=A−1 0Aj,ut=A−1 0t,var(ut)=E(utu t)==A−1 0(A−1 0).Thereduced form parameter is φ=(b,B1,,B4,). The prior belongs to the normal inverse-Wishart family: ∼IW(,d),β|∼N(¯ b,⊗), where β≡vec([b,B1,,B4]).=I3is the location matrix of ,d=4 is a scalar degrees of freedom hyperparameter and =100I10 is the variance-covariance matrix of β.The prior mean ¯ bis consistent with a random walk representation for the observables. In what follows, we perform Algorithm 3.1 with 1000 draws of φ’s from the normal inverseWishart posterior. Following Christiano, Eichenbaum, and Evans (1999), we always impose the sign normalization restrictions so that the diagonal elements of A0are nonnegative. 4.1 Averaging indistinguishable models Suppose we are interested in the cumulative output growth response17 to a unit (contractionary) monetary policy shock mp tat horizon h,IRh gdp,mp, and consider the following two sets of identifying assumptions. •Model 1 (M1, point-identified) Assume that output growth and inflation do not react on impact to the monetary policy shock, so that the (2,1) and (3,1) elements of the matrix of contemporaneous impulse responses IR0=A−1 0are zero. This identification scheme point identifies IRh gdp,mp. •Model 2 (M2, set-identified through zero restrictions) The identification scheme in Model 1 is controversial.18 Thus, in Model 2 we leave inflation unrestricted and the zero restriction is only imposed on the (2,1) element of A−1 0. By Proposition B.1 in Giacomini and Kitagawa (2021), Model 2 delivers a convex identified set for IRh gdp,mp. Panels (a), (b), and (e) of Figure 1focus on the output response at horizon h=3 implied by Model 1, Model 2, and their average for uniform prior model probabilities. In panel (a), the vertical solid lines for Model 1 are the 90% credible region for the pointidentified output response based on a single posterior; in panel (b), the vertical dashed lines for Model 2 are the posterior mean bounds (consistent estimator of the identified set) for the output response and the solid line represents credible regions piled up from the 95% (bottom) to 5% (top) with increasing credibility by 5%. Panel (e) reports the model averaging results. The vertical dashed lines for the averaged model can be viewed as shrinking the identified set estimator from Model 2 toward the point estimator from Model 1. Figure 2reports the results for multiple horizons. 17From now on, any impulse response is cumulative. 18See Kilian (2013) for a discussion. 116 Giacomini, Kitagawa, and Volpicella Quantitative Economics 13 (2022) Figure 1. Density and robust credible region of output impulse responses. Note:Figure1reports output impulse response at horizon h=3. For set-identified models (panel (b), (c) (e), (f), (g), (h)), step lines represent the Robust Credible Region (RCR) at different credibility levels. The vertical dashed lines represent the posterior mean bounds. For point-identified models (panel (a) and (d)), the vertical solid lines display the standard credible region. In such a case, we report its posterior density. Note that, as is common for point-identified small-scale SVARs, Model 1 shows a negative response of output in the short run, whereas the set-identified Model 2 is consistent with both positive and negative effects. This is confirmed by the last row in TaFigure 2. Plots of output impulse responses. Note: For set-identified models (panel (b), (c) (e), (f), (g), (h)), the vertical bars show the posterior mean bounds and the dashed curves connect the upper/lower bounds of posterior robust credible regions with credibility 90%. For point-identified models (panel (a) and (d)), the points plot the (unique) posterior mean and the dashed curve represent the highest posterior density regions with credibility 90%. Quantitative Economics 13 (2022) Uncertain identification 117 ble 2, reporting the lower and upper probability that the post-averaging interval of posterior means of the output response lies in the negative real half-line. Averaging the models still does not rule out a positive output response, as the 90% robust credibility region always contains positive values. Note that, since the models are indistinguishable, the prior model probabilities are not updated by the data. 4.2 Averaging distinguishable models We now consider a case where the prior model probabilities are updated, by adding two popular models: a sign-restricted SVAR and a structural DSGE model: •Model 3 (M3, set-identified through sign restrictions) We consider the following sign restrictions: the inflation response to a contractionary monetary policy shock is nonpositive and the interest rate response is nonnegative at h=0, 1. As in Uhlig (2005), the output response is unrestricted. By Lemma 5.2 in Giacomini and Kitagawa (2021), the identified set in Model 3 is convex. Consider averaging Model 1 and Model 3 with equal prior probabilities. In contrast to the previous example, the prior probabilities can now be updated using equation (3.5) because the models are distinguishable due to the observationally restrictive sign restrictions. The Appendix in the Online Supplementary Material provides two algorithms for approximating the posterior-prior plausibility ratio for the sign-restricted SVARs. We report results based on Algorithm A.1 (Algorithm A.2 produces almost identical results). Panel (f) of Figures 1and 2reports the results of averaging the two models: as in the case of Model 2, Model 3 does not rule out a positive output response (this is also the conclusion of Uhlig (2005), however based on a single-prior approach). Table 1shows that the posterior model probabilities favor Model 3 (with posterior probability 0.55), and the average of the two models does not exclude a positive output response. •Model 4 (M4, DSGE) We consider the Bayesian DSGE model in An and Schorfheide (2007), which is a simplified version of Smets and Wouters (2003)andChristiano, Eichenbaum, and Evans (2005). In order to estimate the model, we rely on the prior specification in An and Schorfheide (2007), Table 2and use output, inflation, and interest rate as observables. We use the Laplace approximation to compute the marginal likelihood. Panel (g) of Figures 1and 2shows the results of averaging Models 3 and 4; note the different scale for Model 4. These models do not admit an identical reduced form, so the (equal) prior probabilities are updated according to equation (3.3). We see that Model 4 implies a negative output response; however, its posterior model probability is only 0.13, and the averaged model is consistent with both a positive and negative output response. Finally, Panel (h) of Figures 1and 2reports the results of averaging all models (with equal prior weights). The posterior model probabilities (Table 1, last column) show evidence supporting the sign-restricted SVAR, while the support for the DSGE model is again weak. As in all previous cases, the averaged model does not rule out a positive output response. 118 Giacomini, Kitagawa, and Volpicella Quantitative Economics 13 (2022) Table 1. Output responses: prior and posterior weights. Averaging M1, M2 Averaging M1, M3 Averaging M3, M4 Averaging M1, M2, M3, M4 Prior w10.50 0.50 /0.25 Prior w20.50 // 0.25 Prior w3/0.50 0.50 0.25 Prior w4//0.50 0.25 O111/1 O21// 1 O3/1.21 1.21 1.21 O4//11 ln ˜ p(Y)−779.61 −779.61 −779.61 −779.61 lnp(Y|M1)−779.61 −779.61 −779.61 −779.61 lnp(Y|M4)//−781.29 −781.29 Posterior w∗ 10.50 0.45 /0.29 Posterior w∗ 20.50 // 0.29 Posterior w∗ 3/0.55 0.87 0.36 Posterior w∗ 4//0.13 0.06 Note:Priorwi,Oi,andposteriorw∗ idenote prior model probability, posterior-prior credibility ratio, and posterior model probability for candidate Model i, respectively; ln ˜ p(Y),lnp(Y|M1),andlnp(Y|M4)represent log marginal likelihood for the common reduced form, for Model 1 and for Model 4, respectively. 4.3 Reverse-engineering prior model probabilities We now conduct the reverse engineering exercise discussed in Section 2, which computes the prior weight one would need to assign to a set of restrictions in order for the posterior mean bounds for the output response to be contained in the negative real halfline. First, consider Model 1 and Model 2. Letting wbe the prior probability of Model 1, the post-averaging interval of posterior means is inf πα|Y∈α|Y Eα|Y(α),sup πα|Y∈α|Y Eα|Y(α) =πM1|YEα|M1,Y(α)+πM2|YEφ|Y,M2lφM2|M2,Eφ|Y,M2uφM2|M2 and the posterior model probabilities are equal to the prior probabilities (since the models are indistinguishable), that is, πM1|Y=wand πM2|Y=1−w. We compute the prior model probability wsuch that the post-averaging interval of posterior means is contained in the negative real half-line for h=3. We find that one would need w>0.28 to support the conclusion. We next consider Model 1 and Model 3 (set-identification through sign restrictions). The only difference is that now the posterior model probabilities are updated and are equal to πM1|Y=O1·w O1·w+O3·(1−w)and πM3|Y=O3·(1−w) O1·w+O3·(1−w). We find that one would need to attach very high prior probability (w>0.83) to the point-identifying restrictions in Model 1 to obtain a negative output response.