Demand estimation with infrequent purchases and small market sizes
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Hortaçsu, Ali; Natan, Olivia R.; Parsley, Hayden; Schwieg, Timothy; Williams, Kevin R. Article Demand estimation with infrequent purchases and small market sizes Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Hortaçsu, Ali; Natan, Olivia R.; Parsley, Hayden; Schwieg, Timothy; Williams, Kevin R. (2023) : Demand estimation with infrequent purchases and small market sizes, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 14, Iss. 4, pp. 1251-1294, https://doi.org/10.3982/QE2147 This Version is available at: https://hdl.handle.net/10419/296353 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Quantitative Economics 14 (2023), 1251–1294 1759-7331/20231251 Demand estimation with infrequent purchases and small market sizes Ali Hortaçsu Department of Economics, University of Chicago and NBER Olivia R. Natan Haas School of Business, University of California Hayden Parsley Department of Economics, University of Texas Timothy Schwieg Booth School of Business, University of Chicago Kevin R. Williams School of Management, Yale University and NBER We propose a demand estimation method that allows for a large number of zerosale observations, rich unobserved heterogeneity, and endogenous prices. We do so by modeling small market sizes through Poisson arrivals. Each of these arriving consumers solves a standard discrete choice problem. We present a Bayesian IV estimation approach that addresses sampling error in product shares and scales well to rich data environments. The data requirements are traditional market-level data as well as a measure of market sizes or consumer arrivals. After presenting simulation studies, we demonstrate the method in an empirical application of air travel demand. Keywords. Discrete choice modeling, demand estimation, zero-sale observations, Bayesian methods, airline markets. JEL classification. C11, C18, L93. Ali Hortaçsu: [email protected] Olivia R. Natan: [email protected] Hayden Parsley: [email protected] Timothy Schwieg: [email protected] Kevin R. Williams: [email protected] This paper was previously titled “Incorporating sales and arrivals information in demand estimation.” The views expressed herein are those of the authors and do not necessarily reflect the views of the National Bureau of Economic Research. We thank the anonymous airline for giving us access to the data used in this study. All authors have no material financial relationships with entities related to this research. We thank three anonymous referees. A replication file is posted (Hortaçsu, Natan, Parsley, Schwieg, and Williams (2023)). ©2023 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE2147
1252 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) 1. Introduction Demand estimation techniques are often used in applied work where researchers have access to aggregate data. Although existing methods perform well in large markets when observed market shares can be assumed to be the same as their model equivalents, digitization has enabled the collection and analysis of high frequency disaggregated sales data. Often, many or most observations in these micro-data contain no purchases. Situations with zero transactions are problematic for existing methods because these approaches require observed market shares to be strictly positive in order to be estimable. This practical estimation challenge raises the concern that these standard demand models are conceptually inappropriate for high-frequency, detailed micro data sets. In this paper, we propose an approach to modeling and estimating discrete choice demand suitable for data environments with sparse sales. Our model combines Poisson arrivals and discrete choice demand models that accommodate random coefficients, flexible latent product characteristics, and endogenous prices. Although alternative recent methodological advances propose solutions to the zeros problem by considering inference under large market sizes, in many cases the number of consumers will never grow in a way that reduces sampling error in product shares. This occurs because purchase opportunities tend to grow too slowly relative to the number of products offered. In contrast, our approach explicitly models small market sizes, which allows for low purchase rates and a significant amount of zero-sale observations. The data requirements are conventional market-level data as well as a measure of observed market sizes, such as information on consumer arrival intensity. These data are becoming increasingly available in economics and marketing and not need pertain only to e-commerce. For example, arrival intensity may involve foot-traffic statistics or a proxy for market sizes, such as the number of consumers who purchased a particular good, for example, milk in the grocery context. We present simulation studies to compare to alternative approaches to handling sparse sales. We show that our approach performs well in situations where alternative methods produce biased demand estimates, including misspecifying the supply-side. Finally, we extend our methodology to discrete random coefficients and apply it to the study of airline markets, where the daily demand for flights is low—product-level zero sales exceed 85%. We compare demand estimates across estimation approaches and use our model to explore the underlying forces that cause cyclical demand for air travel. In Section 2, we consider the workhorse demand model of Berry, Levinsohn, and Pakes (1995), henceforth BLP (1995), which is often used to flexibly estimate substitution patterns across differentiated products. In the BLP (1995) approach, empirical product shares are matched to their model counterparts via a market share inversion that requires all sales quantities (and thus empirical shares) to be strictly positive. In estimation, this inversion is not possible in the presence of zero-sale observations, as it requires taking the log of empirical shares equal to zero. Dropping the zeros is known to create a selection bias in the demand estimates (Berry, Linton, and Pakes (2004)). Although aggregation of the data may be possible, this may smooth over the heterogeneity
Quantitative Economics 14 (2023) Demand estimation 1253 of interest—in our empirical application, we find that aggregation also yields implausible estimates of demand. In our model, market sizes are modeled through Poisson distributions. Under our modeling assumptions, demand is also distributed Poisson. This allows us to rationalize zero-sale observations and account for the sampling distribution of sales driven by small market sizes. Relative to other proposed solutions to the zeros problem, including Quan and Williams (2018), Li (2019), Adam, He, and Zheng (2020), Dubé, Hortaçsu, and Joo (2021), Lima (2021), and Gandhi, Lu, and Shi (2023), we explicitly leverage market-size variation as a source of zero sales. An alternative approach would be to fix the market size and model the multinomial distribution of sales, similar to Conlon and Mortimer (2013). Our approach combines methodologies that allow for sampling error in product shares, price endogeneity, and rich unobserved heterogeneity into a single framework. Market participation is modeled according to flexible Poisson distributions. That is, consumers consider all products or do not participate in the market at all. This contrasts with individual-level models of limited information, where consumers search across subsets of products (e.g., Amano, Rhodes, and Seiler (2022), Abaluck, Compiani, and Zhang (2022)). Our approach relates to Burda, Harding, and Hausman (2012), who consider Poisson demand with rich individual-level heterogeneity, and Vulcano, Van Ryzin, and Ratliff (2012), who suggest using choice set variation in Poisson demand estimation. However, both of these works abstract from endogenous variables.1Similarly, a growing literature in empirical industrial organization (e.g., Buchholz (2021), Williams (2022)) and operations (e.g., Newman, Ferguson, Garrow, and Jacobs (2014), Jain, Rudi, and Wang (2015), Abdallah and Vulcano (2021), Wang (2021)) consider Poisson demand. Relative to these works, we allow for prices to have a flexible correlation structure with latent demand characteristics, that is, prices are endogenous to unobserved product qualities. We develop a Bayesian instrumental variables estimator for the model in Section 3. We build on the methods proposed by Jiang, Manchanda, and Rossi (2009)byaddingan explicit model of market size that accommodates unobserved product shares (Poisson demand), discrete or continuous random coefficients, and a flexible treatment of addressing price endogeneity. By augmenting the data with unobserved shares, we can use the market share inversion of BLP (1995) even though there may exist zero-sale observations and sampling error in product shares. We present two approaches to handling price endogeneity that use limited information pricing equations. The first treatment considers a nonparametric relationship between price and the demand unobservables by applying a Dirichlet process prior. We also present a semi-nonparametric treatment using a mixture normal model in the Appendix. In Section 4, we demonstrate that our estimator can recover accurate and precise parameter values relative to existing methods in several simulation studies. Our estimator provides coverage even if market sizes are very small, for example, when arrival rates are five individuals per market. In this setting, we show that common zero-share solutions introduce considerable bias to the demand estimates and overstate the dispersion 1See also Chen and Kuo (2001) and Lee, Green, and Ryan (2017) for related work that involves random effects without endogenous characteristics.
1254 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) of preferences within and across markets. In addition, we show that our estimator can be robust to some forms of misspecification, including misspecifying the dispersion of consumer arrivals and restricting the flexibility in our treatment of price endogeneity. Lastly, we conduct a set of simulation studies that better represents retail scanner data, where consumers face choice sets with dozens of products. Finally, we provide an adaptation of our estimator to mass point random coefficients (Kamakura and Russell (1989), Berry, Carnall, and Spiller (2006)) and consider an empirical application to the airline industry in Section 5. Our empirical setting emphasizes the data requirements for estimation. We use data provided by a large international airline based in the United States. Our sample contains sales quantities, prices, as well as a measure of consumer arrivals. More precisely, we use search query counts at granular levels to inform market sizes in estimation. The demand for air travel is difficult to estimate because sales are sparse. We find that 85% of observations are zero-sale observations. Moreover, typically just a few consumers arrive per day. Using our method, we estimate mean product elasticities to be −1.2. However, we estimate demand to be far more inelastic by using existing demand approaches. Existing approaches yield price elasticities between −0.13 and 0, with substantial masses of estimates very close to zero. This occurs because imputing (small) shares when sales are zero causes price variation to have no impact on shares due to attenuation bias. Similarly, dropping zeros lead to price elasticity estimates very close to zero, since observations with positive shares feature higher willingness to pay. We find that low arrival rates prevent any product from having consistently well-measured market shares, which causes the estimator of Gandhi, Lu, and Shi (2023) to yield biased demand estimates as well. With our model estimated, we explore preference heterogeneity across markets and decompose the driving forces of cyclicality in demand for air travel. Our analysis shows that periods of low demand feature both low market participation and consumers with lower willingness to pay. We estimate substantial variation in preferences across travel itineraries: passengers are willing to pay $96 more for the most popular week for travel over the least popular week on an identical route. Preferences across departure times reflect observed price differences. However, this variation is inflated significantly using existing approaches, suggesting consumers are willing to pay thousands of dollars to switch flights. Finally, we discuss why demand cyclicality would be amplified, absent by the use of dynamic pricing and frequent capacity adjustments. 2. Model of consumer demand We model aggregate demand using the widely applied random coefficients logit demand model. Consumers, indexed by i, arrive in market t, and make a discrete choice, choosing among market-specific products (j∈Jt)and an outside option (j=0).Marketst may be defined across time or other dimensions, such as space. The data may comprise a panel. In the baseline model, consumers are drawn from a continuum of types. We also consider mass-point random coefficients in Section 5.
Quantitative Economics 14 (2023) Demand estimation 1255 2.1 Utility specification We assume that indirect utilities are linear in product characteristics and are given by ui,j,t=Xj,tβi+ξj,t+εi,j,t,j∈Jt, εi,0,t,j=0, where Xare product characteristics, including price, ξare unobserved (to the econometrician) product characteristics that are potentially correlated with price, and εare independent and identically distributed error terms. We assume these errors are distributed type-1 extreme value. The random coefficients are assumed to be distributed jointly normal across consumers and are independent of characteristics. That is, βi=¯ β+bi, where βi∼N(0, I)is a multivariate, standard normal distribution, and is the Cholesky decomposition of a positive definite matrix. This allows for a general variance pattern between demand parameters and the random coefficients. Some of the parameters may be linear, meaning that there are no associated random coefficients with these characteristics. All consumers solve a straightforward utility maximization problem; consumer i chooses product jif and only if ui,j,t≥ui,j,t,∀j∈J∪{0}. The distributional assumption on the idiosyncratic error term leads to analytical expressions for the individual choice probabilities of consumers. In particular, the probability that consumer ipurchases product jis equal to si,j,t=exp(Xj,tβi+ξj,t) 1+ k∈Jt exp(Xk,tβi+ξk,t) . Integrating over all consumers, we obtain product market shares, which are equal to sj,t=i si,j,tdF(bi). 2.2 Distribution on market sizes We model the distribution on market sizes using Poisson distributions, that is, At∼Poisson(λt), such that λt:=exp(Wt)and Wtis a full-rank matrix. Note that we assume arrivals are measured specific to trather than specific to j,t. That is, all arriving consumers have full information about all products upon arrival. The matrix Wcould contain, for example, time, location, or other market-specific covariates, depending on the application.
1256 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) Our model of market participation is a highly stylized representation of consumer search. Consumers either search among all products or they do not participate in the market. This is because we do not leverage individual-level data; see, for example, Kim, Albuquerque, and Bronnenberg (2010), Honka (2014), and Armona, Lewis, and Zervas (2021). Our model uses aggregate, market-level data. Two assumptions allow us to construct analytic expressions for demand: (1) the realizations of arrivals are conditionally independent of preferences (A⊥ξ,p|X,W),and (2) consumers solve the above utility maximization problems. With these assumptions, conditional on prices, product characteristics, and observed determinants of market size, demand for each product jis distributed according to a conditionally independent Poisson distribution, that is, qj,t∼Poisson(λt·sj,t). Importantly, the demand for product jis independent of the demand for product j, conditional on Wt,Xt,andpt,2which still allows correlation in the observed exogenous determinants of utility and the arrival rates. For example, this may accommodate higher arrival rates on average in markets with higher average demand (e.g., in the case of seasonal demand). Xand Wcould contain the same variables in some applications. 3. Demand estimation We propose a Bayesian estimator (Poisson RC) that can be used to recover a rich set of demand parameters. Our approach differs from frequentist estimators that use marketlevel data in two key ways. First, we directly accommodate small market sizes by providing an alternative to the empirical share inversion step of Berry (1994)andBLP(1995). Our estimator uses data augmentation to directly sample from the distribution of unobserved demand shocks, thus only requiring an inversion of model shares. Second, our estimator scales well to many markets and rich demand covariates, including many fixed effects. Without small market sizes and sampling error, our method closely follows Jiang, Manchanda, and Rossi (2009).3Unlike in settings with large market sizes, with low arrival rates and sparse sales, we are forced to treat market shares as unobserved. That is, we cannot simply average observed sales and equate them to market shares because zero-market shares can be due to zero-consumer arrivals or arriving consumers choosing not to purchase. Observed sales in the data are qj,t, which are not only a function of the product shares, but also the number of people that arrive in each time period. This is important because in periods with low arrivals, the probability of qj,t=0 is quite high, but sj,tis never equal to zero. Thus, when we observe few arrivals, we must account for the sampling variation to be expected in sales quantities. Note that if we did not have endogenous product characteristics, our estimator would closely resemble existing Poisson-logit maximum likelihood estimators (e.g., 2This result follows from the properties of splitting Poisson processes (Gallager (1996)). 3In addition, we extend their approach to mass-point random coefficients and make the pricing equation estimation more flexible using a Dirichlet process prior and mixture of normal distributions.
Quantitative Economics 14 (2023) Demand estimation 1257 Burda, Harding, and Hausman (2012), Vulcano, Van Ryzin, and Ratliff (2012)). These approaches could, in the absence of price endogeneity, accommodate zero-sale observations and provide flexible arrival patterns. However, in the presence of endogenous product attributes, such estimators will fail to produce unbiased estimates of key parameters, including the price coefficient(s). 3.1 Accounting for price endogeneity To account for price endogeneity, we model pricing through a limited information pricing equation with observed exogenous components and an unobserved component. Using a set of instruments Zj,t, we specify pj,t=Z j,tη+υj,t, where υis unobservable to the econometrician. We allow the aggregate demand shocks ξto be correlated with prices through υand follow standard Bayesian frameworks for simultaneity with discrete choice models (Rossi and Allenby (1993), Jiang, Manchanda, and Rossi (2009), Rossi, Allenby, and McCulloch (2012)). We provide additional flexibility by estimating the joint distribution prices and the demand shocks ξnonparametrically using a Dirichlet process prior, instead of assuming it to be a single joint normal distribution, for example, in Rossi, Allenby, and McCulloch (2012). The Dirichlet process provides the researcher with more control over the generality of the distribution when data are sparse. We show that these estimators, while misspecified to many equilibrium pricing models, allow for a sufficiently general approximation of equilibrium behavior between unobserved components of demand and price while still allowing for likelihood-based estimation. For example, they preclude that pjdepends on ξkor conditional on the contents of Z, characteristics and cost shocks for other products. A more general likelihood-based approach, which does not allow for zero shares, is considered in Grieco, Murry, Pinkse, and Sagl (2023). Alternatively, the econometrician can fully specify a supply-side model and explicitly model price endogeneity, while still remaining non-parametric with respect to the demand unobservables. We comment on this in more detail below. 3.2 The Dirichlet process prior for (ξ,υ) We use a Dirichlet process prior to allow for an arbitrary number of distributions to be mixed together to approximate the joint distribution of (ξ,υ). Appendix Bprovides details for the more restrictive, though computationally cheaper, finite mixture of normal distributions that can be used to semi-nonparametrically model the joint distribution of ξand υ. Here, we state properties of the Dirichlet process prior and detail the algorithm used to implement the method within our MCMC algorithm. The Dirichlet process can be viewed as a distribution over distributions. Each residual pair (ξ,υ)is drawn from some distribution, and the Dirichlet process clusters these distributions together. There are two components to the process: ˜α, the tightness parameter, and G0, a prior distribution for each residual pair. Drawing from the process
1258 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) yields a distribution for the residuals (ξ,υ), and we condition on the drawn distribution for each step of the chain in the rest of estimation. Each residual pair has its own distribution, indexed by parameters θn. Many residual pairs will have the same θnand be drawn from the same distribution at each step of the chain. The prior distribution G0describes the distribution of a new cluster. It must be diffuse enough to cover the data, but not so diffuse that the likelihood of drawing from it even at likely parameter values is too low. As with any nonparametric estimator, care must be taken to choose priors such that we can approximate the residuals well. We choose a normal distribution for G0so that G0∼N(μ,). We are interested in the conditional distribution of θ. We follow the Blackwell– MacQueen Pólya Urn representation (Blackwell and MacQueen (1973)) where θn|θ−n∼ ˜αG0+ j=n 1θj ˜α+N−1, which is a mixture distribution between G0and 1θj, a point-mass located at θj.Given other draws from the Dirichlet process, a new draw has a positive probability of drawing the same parameters as previous draws, and ˜α ˜α+N−1probability of a new draw from the distribution G0. By drawing the classifiers from a Dirichlet process, there is always a possibility of introducing a new normal distribution for each data point, but this is disciplined by the choice of the prior distribution, which has a tendency to cluster observations together. We choose the tightness parameter ˜αto be constant and base it on the data.4 The Dirichlet process produces a prior that is very similar to what one might select for a finite mixture of normal model, but with two substantial differences. The number of mixtures is changing with every step of the chain, and there is a positive probability of adding another mixture component. The prior probability of adding another mixture to the model is governed by the hyperparameter ˜α. In practice, this means that for K existing clusters, the prior probability of the nth data point being drawn from the existing clusters is Nk ˜α+N−1,whereNkis the number of data points currently in cluster K.The prior probability of a new cluster is ˜α ˜α+N−1. Conditional on θn, the residuals are distributed bivariate normal, a key fact that we leverage in the rest of estimation to construct several conditional likelihoods. Because of the tendency of the Dirichlet process to cluster distributions together, we index these clusters by κj,t, which is sufficient for θn. All residual pairs with the same κnshare the same distribution. Let each cluster be indexed by k,observationsinthekth cluster are distributed normally with a shared mean and variance. Observations in the kth cluster 4Rossi (2014) provides a more in-depth look at the choice of hyperparameters, including treatments where ˜αis random, determined by data, and a further parameterization of a,ν,v. We omit these for simplicity.
Quantitative Economics 14 (2023) Demand estimation 1265 Table 3. Monte Carlo results for λt=5, J=25, with true values α=−2, =0.2. Estimator M Aggregate Zeros α MAD MAE Bias MSE MAD MAE Bias MSE Poisson-RC – – – 0.31 0.30 0.30 0.10 0.16 0.13 −0.12 0.02 (0.05, 0.50)( −0.18, 0.06) BLP (1995) 80 No Drop 1.84 1.74 1.72 3.15 0.20 0.28 −0.06 0.16 (0.30, 1.91)( −0.20, 1.28) BLP (1995) 80 No Adjust 1.57 1.53 1.52 2.38 0.27 0.31 0.31 0.15 (1.28, 1.69)( 0.08, 0.63) BLP (1995) 80 Yes Drop 1.21 1.47 0.38 5.24 0.20 0.67 0.42 1.75 (−5.72, 1.66)( −0.20, 3.76) BLP (1995) 80 Yes Adjust 1.15 1.32 0.90 5.24 0.20 0.36 −0.00 1.31 (0.14, 1.59)( −0.20, 1.00) BLP (1995) Realized No Drop 1.86 1.85 1.85 3.42 0.20 0.26 −0.06 0.17 (1.66, 2.02)( −0.20, 1.06) BLP (1995) Realized No Adjust 1.75 1.71 1.71 2.94 0.20 0.32 0.15 0.31 (1.49, 1.82)( −0.20, 0.84) BLP (1995) Realized Yes Drop 1.25 1.35 1.20 2.53 0.20 0.66 0.37 4.75 (−0.68, 2.92)( −0.20, 5.65) BLP (1995) Realized Yes Adjust 1.11 1.12 1.08 1.48 0.20 0.39 0.11 1.02 (0.15, 1.93)( −0.20, 2.20) GLS (2023) 80 No – 1.07 1.05 1.05 1.13 0.79 0.77 0.77 0.60 (0.71, 1.36)( 0.56, 0.90) GLS (2023) 80 Yes – 0.79 0.76 0.73 0.72 0.82 0.82 0.82 0.69 (−0.19, 1.41)( 0.60, 1.00) GLS (2023) Realized No – 1.51 1.52 1.52 2.33 0.80 0.80 0.80 0.66 (1.29, 1.75)( 0.53, 1.05) GLS (2023) Realized Yes – 1.51 1.49 1.47 2.79 0.78 0.76 0.76 0.59 (0.18, 3.23)( 0.59, 0.89) Note: Reported in the table are the median absolute deviation (MAD), mean absolute error (MAE), bias, and mean squared error (MSE). In parenthesis, we enclose the 2.5 percentile and the 97.5 percentile of the bias for all 100 simulations. Column M refers to the market size, which is either handled as part of the model (Poisson-RC), used as data (“Realized”), or calibrated to 80. Column Aggregate refers to whether observations are aggregated prior to estimation. In the case of aggregation, sales, covariates, and arrivals are averaged across 10 adjacent observations. Column Zeros refers to how zero sales are handled—either directly accommodated by the estimator, or observations with zero empirical shares are dropped, or zero empirical shares are replaced with their zero-adjusted (Laplace) shares. Simulations that had the absolute value of the estimated parameter larger than 10 times the true parameter value were excluded from calculations. (rows with zeros set to “Drop” in Table 2and Table 3). This is more severe in the case with the smaller market size. Although the drop-zero estimates sometimes perform better in estimating the price coefficient than the adjusted zero methods, this approach fails to capture other parameters accurately, including both the mean and variance of the random coefficient. Across various BLP (1995) specifications, we find that the median absolute deviation for the random coefficient variance is −0.20, which in part is due to constraining the parameter value to be positive in estimation. Removing this constraint
1266 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) Figure 1. Estimation bias in product characteristics. Note: Reported in each figure is the distribution of parameter bias for the first exogenous characteristic (β1) across simulations for different estimators. We report only results the first element of ¯ βsince each element of both Xand ¯ βis generated i.i.d.—results are interchangeable across the exogenous product attributes. The blue solid lines report the distribution of mean parameter bias across simulations using Poisson RC. The orange dashed lines report the distribution of mean parameter bias across simulations and estimation adjustments for BLP (1995). That is, it reports the bias distribution across all BLP (1995) adjustments (aggregation and disaggregation, dropping zeros or adjusting zeros) and across simulations. leads to a much more dramatic bias due to unconstrained estimation, which results in large negative values. Adjusting product shares when they are equal to zero results in bias (rows with zeros set to “Adjust” in Table 2and Table 3), but this adjustment biases results on average less than dropping the zero-share observations. Adjusting zero shares yields parameter estimates, which consistently bias both the mean and variance in the random coefficient on price. The distribution of product shares when sales are zero is centered below the distribution of shares when sales are positive. Dropping the zero shares creates selection since zero shares reflect higher prices or lower demand shocks (Berry, Linton, and Pakes (2004)). Figure 2(a) plots the distribution of the difference between true shares and adjusted shares. Dropping zeros results in a distribution of empirical shares that are lower than the truth. Adjusting zero shares with a small value also understates the true share. Consequently, price sensitivity estimates are attenuated since imputing a tiny share is inaccurate when the zeros occur due to few consumers arriving. Figure 2(b) shows the conditional distribution of product shares when sales are zero. Imputing an arbitrary small value performs poorly because imputed shares are on the large end of common imputed definitions of zero shares, but they are lower than true shares when quantity sold is zero. As a result, observations with small true shares (e.g., when price is higher than usual) will be imputed to have even smaller shares. This leads estimates to understate the price sensitivity of consumers. An alternative solution to minimize the frequency of zero shares is to aggregate the data. We find that applying share adjustments after aggregation results in fewer adjusted shares, but it does not improve the performance of such estimators (rows with aggregate equal to Yes in Table 2and Table 3). Instead, aggregating the data removes variation in prices, shares, and the instruments, which results in additional bias in some parame-
Quantitative Economics 14 (2023) Demand estimation 1267 Figure 2. True shares compared to share adjustments. Note: (a) Distributions of the difference between the “true” model shares that generated the data from the various zero-share adjustments. Adjusted shares refers to taking any observations where the empirical shares would equal zero and replace it with an arbitrary small number ε, which is set by sA t=Mtst+1 Mt+Jt+1. Drop refers to dropping all observations where the empirical shares are equal to zero. (b) The density of the log of the model shares when quantity sold is zero. Plotted in the dashed vertical grey line is the average ε. ters. This approach results in smaller shares on average compared to the disaggregated results. We find that, in general, using observed market size realizations improves the results of BLP (1995) where zero shares are dropped, however, our main results hold: the BLP (1995) estimator with zero adjustments performs poorly when market sizes are small. Our Poisson RC estimation method uses an unknown number of mixture of normal distributions to approximate any joint distribution of (ξ,υ). We find that our approach is able to recover this joint distribution well.8In implementing the Dirichlet process (DP) prior on the correlation between prices and ξ, our approach requires minimal ex ante specification—all components have the same prior mean and variance. In addition, the method requires a single prior parameter governing the variability of DP components. Our Monte Carlo results show that the Poisson RC model can accurately measure consumer preferences and heterogeneity under very small market sizes. Our method accounts for the sampling error to be expected in sales. Alternative solutions to the zeros problem conduct inference only under large market sizes. For example, we find that the approach of Gandhi, Lu, and Shi (2023) produces significant bias in the price coefficient and noisier estimates than the Poisson RC model in these small market settings. We hypothesize that this is due to the lack of “safe products” used in estimation. Safe products are products in which empirical shares are observed with minimal measurement error in sample (the identities of these products need not be known). In our simulations, zero empirical shares are largely driven by a small market size, which causes all empirical shares to be noisy measures of true shares. Note that Gandhi, Lu, and Shi (2023)also 8In the case that the researcher is certain about the number of components to use to approximate this distribution, we provide an extension to that allows for semi-nonparametric estimation of the distribution of residuals in Appendix B.
1268 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) suggest a partial identification strategy. This approach does not require the presence of safe products. Additionally, the partial identification approach of Dubé, Hortaçsu, and Joo (2021) could be feasible if point estimates are not required. We also note that our approach is not as general as typical moment-based estimators, for example, BLP (1995), which only assume E[ξZ]=0. Our pricing equation does not correspond directly to typical models of differentiated products competition, where prices may depend on demand shocks of all products. This is because the pricing equation we leverage does not match directly the supply model generating the data in our Monte Carlos. Petrin and Train (2010) provide a detailed discussion of the limitations of this approach. At the same time, our approach does satisfy exclusion and relevance for our instruments. Moreover, specifying a full supply model that would be more efficient than our approach (e.g., Yang, Chen, and Allenby (2003)). Nonetheless, as we have shown, our flexible pricing equation is able to accurately recover demand fundamentals. In the case in which market sizes are larger, our method still performs well. We report results in a setting with fewer than 1% zero-sale observations in Appendix C. Existing approaches perform well in such settings, and these moment-based estimators may be considerably faster than our approach if zero shares are not a significant problem in the data. 4.2.1 Estimator performance in larger choice sets Many empirical applications feature larger choice sets. We test our estimator in settings where J=45. The rest of the data generating process remains unchanged. Increasing the number of products does slow estimation—in nearly doubling the number of products, estimation takes 3–4 times longer. This is driven primarily because of greater computational burden at each step of the chain (e.g., more share inversions). We expect that using our approach in settings with hundreds of products may be infeasible unless using a highly optimized code. The results of our estimator and competing methods are presented in Table 4.Qualitatively, estimates using our Poisson RC method are similar to those in smaller choice sets. In this setting, however, our estimator results in some small bias in the estimation of the coefficients on exogenous characteristics (Figure 1, panel c). As in our previous simulations, implementing ad hoc fixes for BLP (1995) do not allow us to recover sensible estimates and typically return price parameters, which understate price sensitivity significantly. Similarly, we find that the estimator of Gandhi, Lu, and Shi (2023)resultsin significant bias due to the lack of “safe products.” 4.2.2 Results with misspecified models We also test our estimator to two forms of misspecification. Results are shown in Table 5. We simulate a misspecified arrival process (overdispersed with twice the variance of our λt=25 setup) and a less flexible form of correlation between the pricing error and the demand shock (estimating the correlation structure between all pairs of ξand υusing a single normal distribution). We find that our estimator performs nearly identically in the case of overdispersion. We also find that using a less flexible form of price endogeneity also provides nearly identical estimates on average, however, the tails of the (bias) distribution are more dispersed. We have found that when the conditional expectation of ξgiven υis approximately linear, or when the correlation summarizes the dependence well, a normal approximation performs well. In
Quantitative Economics 14 (2023) Demand estimation 1269 Table 4. Monte Carlo results for λt=25, J=45, with true values α=−2, =0.2. Estimator M Aggregate Zeros α MAD MAE Bias MSE MAD MAE Bias MSE Poisson-RC – – – 0.19 0.18 0.18 0.04 0.14 0.12 −0.09 0.02 (0.00, 0.29)( −0.18, 0.11) BLP (1995) 80 No Drop 1.60 1.60 1.60 2.57 0.20 0.20 −0.20 0.04 (1.55, 1.65)( −0.20, −0.20) BLP (1995) 80 No Adjust 1.32 1.31 1.31 1.72 0.30 0.31 0.31 0.11 (1.12, 1.43)( 0.10, 0.63) BLP (1995) 80 Yes Drop 0.37 0.40 0.32 0.24 0.20 0.24 −0.10 0.12 (−0.24, 0.82)( −0.20, 0.79) BLP (1995) 80 Yes Adjust 0.48 0.51 0.45 0.33 0.20 0.25 −0.09 0.14 (0.04, 0.96)( −0.20, 0.94) BLP (1995) Realized No Drop 1.64 1.62 1.62 2.63 0.20 0.37 0.03 0.54 (1.32, 1.79)( −0.20, 2.06) BLP (1995) Realized No Adjust 1.44 1.41 1.41 2.00 0.20 0.31 0.15 0.16 (1.17, 1.54)( −0.20, 0.94) BLP (1995) Realized Yes Drop 0.45 0.60 0.19 0.65 0.20 0.52 0.24 2.87 (−1.61, 1.42)( −0.20, 2.60) BLP (1995) Realized Yes Adjust 0.52 0.56 0.33 0.49 0.20 0.45 0.18 0.77 (−0.75, 1.32)( −0.20, 2.46) GLS (2023) 80 No – 1.14 1.12 1.12 1.27 0.79 0.78 0.78 0.61 (0.81, 1.32)( 0.55, 0.90) GLS (2023) 80 Yes – 0.41 0.45 0.39 0.31 0.80 0.80 0.80 0.66 (−0.30, 1.11)( 0.57, 1.06) GLS (2023) Realized No – 1.38 1.36 1.36 1.88 0.83 0.83 0.83 0.69 (1.05, 1.56)( 0.70, 0.95) GLS (2023) Realized Yes – 0.70 0.83 0.78 1.01 0.78 0.74 0.73 0.57 (−0.31, 2.14)( 0.04, 0.91) Note: Reported in the table are the median absolute deviation (MAD), mean absolute error (MAE), bias, and mean squared error (MSE). In parenthesis, we enclose the 2.5 percentile and the 97.5 percentile of the bias for all 100 simulations. Column M refers to the market size, which is either handled as part of the model (Poisson-RC), used as data (“Realized”), or calibrated to 80. Column Aggregate refers to whether observations are aggregated prior to estimation. In the case of aggregation, sales, covariates, and arrivals are averaged across 10 adjacent observations. Column Zeros refers to how zero sales are handled—either directly accommodated by the estimator, or observations with zero empirical shares are dropped, or zero empirical shares are replaced with their zero-adjusted (Laplace) shares. Simulations that had the absolute value of the estimated parameter larger than 10 times the true parameter value were excluded from calculations. Table 5. Monte Carlo results for misspecified distributions. A∼NegBinom(25, 0.5)Misspecified Residual α0.11 0.07 (−0.06, 0.25)( −0.06, 0.22) 11 −0.16 −0.04 (−0.18, 0.06)( −0.18, 0.11) Note: Reported in the table are the median, the 2.5 percentile, and the 97.5 percentile of the difference between the point estimate and the true parameter across 100 simulations. The first column simulates data in the same manner as our main Monte Carlo experiments, but using only a different arrival process. This process has an identical mean but has twice the variance of the baseline Poisson. The second column presents results when the model is estimated assuming that the joint distribution between ξand νis normal and not a mixture, despite being generated from a mixture of normal distributions.
1270 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) cases where there is strong dependence but low correlation, such as ξbeing a symmetric function of υ, this simpler specification may be restrictive and lead to biased demand estimates. Symmetry may be unrealistic because it implies that demand shocks are associated with both low and high prices. 5. Empirical application to the airline industry We use our approach to estimate the demand for air travel with data from a large international air carrier based in the United States.9Our primary aim is to show how to adapt our method to a relevant empirical setting. We also show how our method can flexibly estimate preferences over time and compare these estimates to existing approaches. 5.1 Data We use data from the air carrier’s booking system to construct the quantity of tickets sold and the price paid for every flight, each day before departure. For this analysis, we concentrate on nonstop bookings. In addition to prices and quantities, we also extract basic product characteristics, such as the departure time for each flight and the date of departure. The key additional market-level data for estimation is a measure of market sizes. We calculate consumer arrivals using the number of consumers who initiate search requests on the air carrier’s website using consumer clickstream data. In our setting, consumers arrive at the air carrier’s website and their activity within a browsing session is tracked. We then aggregate search activity to the level of origin-destination-search date-departure date. Note that we measure search just on one website, but consumers may shop via online travel agencies, such as Expedia. We cannot directly measure searches made from other sites. However, because we observe all bookings, we account for searches made via the unobserved sites through scaling factors. Each scaling factor is based on the fraction of sales directly through the airline and the average number of passengers per booking. If we know observed arrivals account for 50% of total bookings, assuming consumers who shop elsewhere have the same distribution of preferences, we can scale up estimated arrival rates by two.10 We present a basic summary analysis of the 207 markets in the sample in Table 6. The sample size is 35 million. We find that the average number of nonstop bookings per flight-day before departure is low (0.2). Choice sets are small, with around 2.5 flights offered per market. Measured arrivals are also small—around two consumers search for a given market-departure date-day before departure. Finally, we find that 89% of observations do not contain any sales. In Figure 3, we plot the 30-day moving average of bookings, fares, fraction of sold-out flights, total capacity for all routes in our sample from August 2018 to August 2019. Also, included in the graph is a measure of opportunity cost, which has the interpretation of 9The airline has elected to remain anonymous. 10In-depth summary analysis of the data and how unobserved searches are accounted for can be found in Hortaçsu, Natan, Parsley, Schwieg, and Williams (forthcoming).
Quantitative Economics 14 (2023) Demand estimation 1271 Table 6. Summary statistics for the sample. Variable Mean Std. Dev. Median 5th Pctile 95th Pctile One-Way Fare (p) 173.9 129.7 138.5 68.5 369.0 Num. Fare Changes 9.3 4.2 9.0 3.0 17.0 Purchase Rate (Q)0.2 0.7 0.0 0.0 1.0 Ending Load Factor (%) 82.2 21.4 90.0 36.0 102.0 Size Choice Set (J)2.5 1.6 2.0 1.0 6.0 Arrival Rate (A)1.9 4.8 0.0 0.0 9.0 Zero-Sale Obs. 88.8 8.3 90.9 71.1 98.3 Note: Summary statistics for the 407 markets included in the sample. The sample contains 298,817 unique flight departures and 35,542,018 observations. the marginal cost in this setting. We will use this measure as an instrument for demand. The figure shows that all these variables are positively correlated. For example, all curves peak around the winter holiday season as well as during summer. During peak periods, demand is high. At the same time, on the supply-side, prices, opportunity costs, flight capacity, and the percentage of flights that sell-out are also high. 5.2 Empirical specification We define a market (m) as an origin, destination, and departure date tuple and let the time index (t) denote days until the departure date. That is, we index markets by m,t.We extend the estimator to allow for discrete support random coefficients. Following Berry, Carnall, and Spiller (2006), we assume consumers are one of two discrete types, corresponding to leisure (L) travelers and business (B) travelers. An individual consumer is denoted as iand her consumer type is denoted by ∈{B,L}. The probability that an arriving consumer is a business traveler is equal to γtand varies by days until the departure date. These types need not correspond to the consumer’s purpose for travel; they merely are commonly used names for discrete consumer types. The less-price-sensitive type is typically referred to as business. We assume the indirect utilities are linear in product Figure 3. Demand, prices, opportunity costs, and capacity for markets in data. Note: 30-day moving average of sales, fares, opportunity costs, probability of flights selling out, and total capacity by departure date. Values are normalized by the respective value for the first departure date in our data, 08-01-2018.
1272 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) characteristics and given by ui,j,t,m=Xj,t,mβ−pj,t,mα(i)+ξj,t,m+εi,j,t,m,j∈J(t,m), εi,0,t,m,j=0. As before, we assume that observed product characteristics Xj,t,mare uncorrelated with the unobserved product characteristic ξj,m,t. These exogenous characteristics include departure time, week, and day of week fixed effects. We include week fixed effects in the utilities to flexibly capture seasonal variation in the value of travel. The consumer types differ in their preferences on price, α(i), and we assume that ξj,m,tis correlated with price. Given our assumption on εi,j,t,m, the probability that consumer iwants to purchase product jis equal to si j,t,m=exp(Xj,t,mβ−pj,t,mα(i)+ξj,t,m) 1+ k∈J(t,m) exp(Xk,t,mβ−pk,t,mα(i)+ξk,t,m) . Since consumers are one of two discrete types, we define sL j,t,mas the conditional choice probability for leisure type consumers (and sB j,t,mfor business types). Integrating over consumer types, we have sj,t,m=γtsB j,t,m+(1−γt)sL j,t,m. In this mass-point random coefficient model, we parameterize the change in the composition of consumers as follows. We assume γtis equal to γt=expf(t) 1+expf(t), where f(t)is an orthogonal polynomial basis of degree 5 with respect to days from departure. This parametric assumption allows for a flexible, nonmonotonic relationship between the composition of consumer types and time while producing values bounded between 0 and 1. Depending on the application, this function can be adjusted accordingly. In addition to allowing for discrete random coefficients, we also adjust the likelihood to account for the possibility of binding capacity constraints (sell-out events). In particular, when capacity is binding, we observe a right-censored estimate of the true number of individuals that wished to purchase. That is, for a given capacity Cj,t,m, qj,t,m=min{˜ qj,t,m,Cj,t,m}, ˜ qj,t,m∼Poisson(λt,m·sj,t,m). Note that when the capacity constraint is observed to bind, the likelihood contribution is instead 1 −Fq(qj,t,m|·),whereFqis the cumulative distribution function of the above Poisson.11 11Wedonotmodelchoicevariationwithinanm,tbecause arrival/booking rates are low. See Conlon and Mortimer (2013) for a method that accounts for choice set variation within a market.
Quantitative Economics 14 (2023) Demand estimation 1273 We parameterize arrival rates by a set of multiplicative fixed effects across markets m and time t.Thatis,λm,t=exp(Wtλt+Wmλm),whereWtis a dummy matrix with a column for each day from departure and Wmis a dummy matrix with a column for each departure market m(an origin, destination, departure date tuple). This parameterization approach allows us to capture general increases in market size toward departure across all seasonal markets. In addition, we then have two sources of seasonal variation in participation and preferences built directly into the model, which will enable us to distinguish seasonal market participation from seasonal preferences. We identify seasonal market participation variation directly from variation in arrivals, which allows us to include the same sets of fixed effects for departure markets in both preferences and arrivals. Given these arrivals, we can identify seasonal preferences from variation in quantities. Finally, we instrument for price to address endogeneity. We use the opportunity cost of capacity for a given flight, advance purchase discount indicators, and the number of inbound or outbound bookings from a route’s hub airport as our instruments.12 We leverage the expiry of advance purchase discounts since these changes alter prices in a predetermined fashion, regardless of realizations of demand shocks. The opportunity cost of capacity directly influences price-setting, as residual variation (after our use of fixed effects) is driven by bookings on onward itineraries. The total number of inbound or outbound bookings to a route’s hub airport captures the change in opportunity cost for flights that are driven by demand shocks in other markets. For example, consider a flight from A→B,whereBis a hub, which serves many markets. We construct all onward traffic from Bonward to other destinations Cor D. We assume that unobserved, systematic demand shocks are independent across routes, so shocks to demand for travel from B→Cand B→Dare unrelated to unobserved shocks to demand for the focal route A→B. Pricing decisions across routes are related via capacity: if a positive shock to demand out of hub Bis realized, the opportunity cost to provide service from A→B→Cor A→B→Drises. This increase in opportunity cost for connecting tickets also raises the opportunity cost of capacity on the A→Bleg, which raises the price on A→B. Our instruments are strong: the correlation between price and the opportunity cost of capacity is 0.72, and the correlation between price and the onward connecting traffic measure is 0.31. Pseudo first stage regressions (presented in Table 7) with our selected specification have a R2of 0.83.13 We report results on the first stage in 7and compare our results to alternative sets of instruments below. 12Opportunity costs during a specific period tdepend on past pricing decisions. We include a set of fixed effects in both the exogenous characteristics and in the instrument set, which capture persistent differences in demand across departure dates. Conditional on these fixed effects, we assume that demand shocks are independent over time. For a route with origin Oand destination D,whereDis a hub, the total number of outbound bookings from the route’s hub airport is defined as the DQD,D,whereQD,Dis the total number of bookings in period t, across all flights, for all routes where the origin is the original route’s destination. If the route’s origin is the hub, we calculate the total number of inward bound bookings, which would be OQO,O,whereQO,Ois the total bookings from all routes where the original route’s origin is the destination. 13We denote these pseudo first stage regressions as we present frequentist OLS estimates of the first stage model we adapt in our estimator.
1274 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) Table 7. Pseudo-first stage regressions. (1) (2) (3) (4) (5) Onward Connecting Traffic 0.597 0.591 0.377 0.369 (0.010) (0.011) (0.006) (0.006) Opportunity Cost 11.069 11.308 11.383 (0.059) (0.056) (0.057) Opportunity Cost27.288 6.506 6.483 (0.073) (0.072) (0.073) Opportunity Cost3−4.287 −3.747 −3.732 (0.053) (0.053) (0.053) Opportunity Cost41.545 1.408 1.399 (0.063) (0.062) (0.063) Opportunity Cost50.162 0.122 0.140 (0.073) (0.073) (0.073) APD FEs N Y N N Y DoW FEs Y Y Y Y Y Week FEs Y Y Y Y Y Departure Time FEs Y Y Y Y Y Adjusted R20.300 0.302 0.798 0.825 0.829 F-Stat 389.003 369.252 3358.269 3958.341 3851.889 Note: Pseudo-first stage results for our instruments. Columns 1 and 2 exclude polynomial terms of the opportunity cost measured by the airline’s algorithm. Column 3 excludes the onward connecting traffic term, and both columns 3 and 4 exclude fixed effects for advance purchase discounts. Column 5 includes the full specification used for estimation. Fixed effect indicators denote inclusion of advance purchase discount, day of week, week of year, and departure time fixed effects, respectively. For this application, we specify a time-varying block structure on the pricing equation, and (ξ,υ)have a block-varying joint normal distribution. That is, within a daysfrom-departure block, (ξt,υt)are distributed jointly normal, and this distribution may vary across blocks. Though such a specification may appear more restrictive than the Dirichlet process prior, this specification allows us to tailor our specifications to our empirical context where the pricing equation clearly changes over time due to advance purchase discounts.14 5.3 Estimation procedure We modify our estimator to accommodate discrete unobserved consumer heterogeneity. We consider a two-type model, though this can be extended to more than two points of discrete support. Conditional on sampled parameters and the different likelihood function, many of our estimation steps remain unchanged. The modified algorithm used for estimation is shown in Algorithm 2. Adjusted steps relative to the continuous random coefficients case are highlighted with NEW. 14Our misspecified specifications in Section 4provide an example where a restrictive distribution of this type performs relatively well compared to the fully flexible estimator.
Quantitative Economics 14 (2023) Demand estimation 1281 responses, this suggests that cyclicality would be even higher, as in the movie industry case. 6. Conclusion We propose a method to estimate product-level demand with small market sizes. Our approach allows for many zero-sale observations, endogenous prices, and rich unobserved consumer heterogeneity. We derive a Bayesian IV estimator to recover random coefficients logit demand parameters with Poisson arrivals. We show through simulation studies that this method can outperform typical zero-sale adjustments and provide unbiased estimates, even in very small markets or under a misspecified pricing function. Our approach can be applied to many settings where granular demand estimates are necessary in order to evaluate counterfactuals or address firms’ decision-making. The key data requirements in our approach are traditional market-level outcomes and one additional data column—measures of consumer arrival intensity. These data are becoming increasingly available to researchers, with relevant applications in e-commerce, retailing, and transportation, among others. Appendix A: Estimation routine using Dirichlet process prior A.1 Markov chain Monte Carlo details A.1.1 Sampling arrival parameters To update the parameters describing the arrival rate of consumers, we use arrival and quantity data. We define the likelihood to be the joint probability of observing Atand qj,t, conditional on sj,t. Arrivals are distributed Poisson. Conditional on shares, we split the arrival process (with rate λt) by the shares to obtain the distribution of quantities sold. Each purchase is drawn from a Poisson distribution with rate λt·sj,t. Because data on arrivals may be sparse—perhaps only a single data point per market—we suggest parameterizing the arrival rate with a series of fixed effects whenever possible, λt:=exp(Wt), where Wis a full rank matrix (composed of 0 and 1s if using fixed effects). Other specifications are possible. Arrivals are distributed Poisson, At∼Poisson(λt). Note that purchase quantities also depend on arrivals. Using the properties of the Poisson distribution, we have qj,t∼Poisson(λtsj,t).
1282 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) We note that a conjugate prior choice for w is log-Gamma distribution such that each element exp()∼(k,ζ). Therefore, the posterior distribution of exp(w),forapartition Wdefined by the matrix Wt,isgivenby exp(w)∼ t∈WAt+ j qj,t+k,ζ 1+ζ t∈W (2−s0,t). A.1.2 Sampling shares and utility parameters Updating shares The Dirichlet process allows for complex distributions of (ξ,υ)to be approximated by a series of normal distributions through a component classifier κ.Conditional on this classifier, each pair of residuals (ξ,υ)are distributed bivariate normal. We apply the standard treatment of simultaneity by conditioning on the variance structure of the normal and the respective residuals. The following sections condition upon κ and derive the sampler for a multivariate normal joint distribution of the demand shock and pricing residual. In the final sections, we discuss sampling the classifier and the component means and variances. Conditional on β,,κ,μ,,andυ, the shares are an invertible function of ξ.The conditional distribution of ξis also normal, which implies a distribution of shares. We compute the likelihood of any particular set of share draws by inverting the demand system for these shares. We derive a distribution of shares via a standard change of variables theorem. Since ξis assumed to be correlated with price, we follow the Bayesian framework for simultaneity with discrete choice models (Rossi and Allenby (1993), Jiang, Manchanda, and Rossi (2009), Rossi, Allenby, and McCulloch (2012)). Using a set of exogenous and relevant instruments Zt,d, we assign ξj,t=f−1(sj,t|β,,Xt) υj,t=pj,t−Z j,tη κ=k∼Ni.i.d.(μk,k)such that k=σ2 k,11 ρk ρkσ2 k,22. For notational parsimony, we omit the conditioning statement, but note that each function is implicitly conditioned on the other demand parameters. We refer to the share equation as sj,t,d=f(ξj,t,d). Since fis invertible, the density of sj,t,dis given by fsj,t(x)=fξj,tf−1(x)·| Jξj,t→sj,t|−1. With this notation, Jξj,t→sj,trepresents the Jacobian matrix of model shares with respect to ξand |Jξj,t→sj,t|−1denotes the inverse of the determinant of the Jacobian. Since υand ξare assumed to be jointly normal, knowing υprovides information about the magnitude of the demand shock. This joint normality does not factor into the Jacobian of the shares distribution, because neither sj,tor ξj,tare in the pricing equation and it is assumed it to be a linear system. However, we must use the correct conditional distribution for ξ. Conditioning on both ηand is enough to pin down the the correlation structure between ξand υ,andto“observe”υas well. Drawing on the structure of
Quantitative Economics 14 (2023) Demand estimation 1283 the bivariate normal distribution, we have ξ|υ,κ=k∼Nμk,2 +ρkυ σ2 k,11 ,σ2 k,22 −ρ2 k σ2 k,11, where ξ υ|κ=k∼N(μk,k),k=σ2 k,11 ρk ρkσ2 k,22. One interpretation of this treatment of simultaneity is that price gives information about the realized demand shock ξ, so the conditional distribution of ξis higher or lower depending on the unobserved component υthat influences price. The conditional distribution of shares is then given by J(t) j=1 ⎡ ⎢ ⎢ ⎢ ⎢ ⎢ ⎢ ⎣ φ ⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ f−1(sj,t)−ρkυ σ2 k,11 σ2 k,22 −ρ2 k σ2 k,11 ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ ⎤ ⎥ ⎥ ⎥ ⎥ ⎥ ⎥ ⎦ ·| Jξt→st|−1, where φ(·)is the standard normal density function. Shares directly shape the distribution of sales. The distribution of purchases is a split Poisson process given by qj,t∼Poisson(λtsj,t). Since the Poisson draw is only dependent on the demand parameters through the shares, qj,tis conditionally independent of ξ. Thus, the likelihood of a particular market’s shares is given by the product of the density of ξand the mass function of qj,t,given by (s.,t)= J(t) j=1 ⎡ ⎢ ⎢ ⎢ ⎢ ⎢ ⎢ ⎣ φ ⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ f−1(sj,t)−ρkυ σ2 k,11 σ2 k,22 −ρ2 k σ2 k,11 ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ λtsj,t)qj,texp(−λtsj,t) qj,t! ⎤ ⎥ ⎥ ⎥ ⎥ ⎥ ⎥ ⎦ ·|Jξt→st|−1. The posterior likelihood is constructed by taking the product of the each market’s inversion multiplied by the likelihood contribution of each product’s quantity sold. There is no conjugate prior distribution, so we sample from the posterior using a Metropolis– Hastings step. Our candidate distribution for share draws is a transformation of a normal distribution added to ξ. This allows for easy tuning of the candidate distribution via the variance of the normal. However, as a complication, the candidate distribution is not reversible. That is q(a|b)= q(b|a). As a result, we reweight the Metropolis–Hastings step according to the implied p.d.f. to make the chain reversible.
1284 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) Updating distribution of consumer types, We use a random coefficients demand specification, where demand parameters can be grouped into nonlinear and linear parameters. We order these demand parameters such that the first Lparameters are distributed normally, and the remaining K−Lare constant across consumers. That is, ui,j,t=xj,tβi+ξj,t+i and βi=¯ β1:L+Ui ¯ βL+1:K, where Ui∼N(0, IL),andis the Cholesky decomposition of a variance matrix. This allows for a flexible covariance between the demand parameters with random coefficients, while maintaining linearity in Kparameters. We use the Cholesky decomposition for computational simplicity. We sample from the posterior distribution of nonlinear demand parameters with a Metropolis–Hastings step. The distribution of ξremains unchanged, and we evaluate a candidate in a similar manner as to drawing shares, but without incorporating the likelihood of purchases. The likelihood of a particular is constructed from the implied distribution of the demand shock ξfrom inverting the demand system. The likelihood of the shares, given ,isgivenby sj,t|υj,t,κj,t,∼Nμκj,t,2 +ρκj,t σ2 κj,t,11 υ,σ2 κj,t,22 −ρ2 κj,t σκj,t,11 υj,tJ−1 ξ→s. We use a shorthand distribution here of a distribution times the Jacobian to mean that the p.d.f. of ξ.,t(evaluated at a set of shares s.,t) is the p.d.f. of a normal distribution with those parameters multiplied by the determinant of the Jacobian. However, a candidate cannot be drawn in a trivial manner, as we must sample from the set of Cholesky decompositions of positive-definite (variance) matrices. We employ the parameterization suggested by Jiang, Manchanda, and Rossi (2009), which lets ij =⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ exp(rij ),fori=j, rij ,fori<j, 0otherwise. This enforces a strictly positive diagonal upper-triangular matrix for any candidate draw r. We have given the likelihood of the demand residual, to complete the posterior likelihood of , we must also define a prior distribution over . Following Jiang, Manchanda, and Rossi (2009), we impose normal priors over each r.Jiang, Manchanda, and Rossi (2009) explore the implications of this prior specification: rij ∼N(0, ψ2 ij ).
Quantitative Economics 14 (2023) Demand estimation 1285 The posterior distribution of is given by J(t) j=1$φf−1(sj,t)−μκj,t,2 −ρ σ2 11 υj,t σκj,t,2|1%|Jξj,t→sj,t|−1× i≤j 1 ψij φ(rij ). For alternative utility specifications, the same procedure can be used. Note that only a few nonlinear parameters may be estimated in a single step, as a Metropolis–Hastings step searching in a high dimension traverses the stationary distribution slowly. Updating type-invariant parameters, ¯ βTo sample from the linear demand parameters, we define δj,tsuch that δj,t=xj,t¯ β+ξj,t. Since ξj,thas a normal distribution and we impose a normal prior on ¯ β,wehaveastandard Bayesian linear regression after we account for the influence of the pricing residual, and the different variances in each element of ξj,t. We accomplish this by normalizing each component of equation (A.1.2) by subtracting the expected value of ξj,tand diving both sides by the standard deviation. We then perform a Bayesian linear regression on this collection of normalized equations, as these rescaled errors have unit variance. Let σk,2|1=&σ2 k,22 −ρ2 k σ2 k,11 be the variance of ξconditional on υand . δj,t−μκj,t,2 −ρκj,t σ2 κj,t,11 υ σκj,t,2|1 =1 σκj,t,2|1 xj,t¯ β+Uβ j,t, where Uβ∼N(0, 1). We follow the typical conjugate prior distribution for a linear regression— ¯ β∼ N(¯ β0,V0). The posterior distribution is then a shrinkage estimator of OLS. Let ˆ Xj,t=xj,t σκj,t,2|1 , and ˆ δ= δj,t−μκj,t,2 −ρκj,t σ2 κj,t,11 υ σκj,t,2|1 . Then the posterior distribution of ¯ βis ¯ β∼N(βN,VN),where βN=ˆ Xˆ X+V0−1−1V0−1β0+ˆ Xˆ δ, and VN=V0−1+ˆ Xˆ X−1.
1286 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) A.1.3 Sampling price-endogeneity parameters Updating pricing equation, ηThe pricing equation is given by pj,t=Zj,tη+υj,t. Conditional on shares, ,and ¯ β,ξis known, so we use the conditional distribution of υgiven ξto perform another Bayesian linear regression in the same manner as ¯ β. We impose a Normal prior, subtract the expected value and divide by the conditional variance. Define σκj,t,1|2='σ2 κj,t,11 −ρ2 κj,t σ2 κj,t,22 .Then pj,t−μκj,t,1 −ρκj,t σκj,t,22 ξj,t σκj,t,1|2 =1 σκj,t,1|2 xj,t¯η+Uη j,t. After this normalization, Uη j,tis a standard normal error term. We draw from ηusing a standard Gibbs sampler draw from a linear regression with unit variance, which is the same process as used for ¯ β. Updating component classifier Using the properties of the Dirichlet process, the prior probability of each cluster is weighted by the likelihood of each data point being sampled from the cluster. The posterior distribution of θis θn|θ−n,υj,t,ξj,t˜α∼ q0G0+ i=n qi1θi q0+ i=n qi , where qi=1 ˜α+N−1Pr(υj,t,ξj,t)|θifor i= 0, and q0=˜α ˜α+N−1Pr(υj,t,ξj,t)|θiG0(dθi). This is a mixture distribution with weights q0for a new cluster, and πiqifor existing clusters, where πiis the sum of data points in cluster idivided by the total data points. While qipresents a similar form as a finite mixture model, q0is difficult to calculate. Because we assume G0∼N,q0is the prior predictive distribution, that is, the likelihood of a data point over the distribution of possible normal distributions θimight take on.18 As 18We impose standard conjugate priors for computational ease, so μ|∼N0, aμ−1, and ∼IW (ν,νvI),
Quantitative Economics 14 (2023) Demand estimation 1287 shown in Murphy (2007), this quantity is distributed multivariate t. Applying our priors, the form is given by Pr(x|θi)G0(dθi)∼tν−10, 1 νvI(aμ+1) aμ(ν−1). We can evaluate the p.d.f. of θat each of the residuals to determine the posterior probability of adding a new cluster. It is important to draw a connection between θnand κ, the component classifier. There are at most nunique values of θn, and usually far fewer due to the clustering nature of the Dirichlet process. κis then drawn from a categorical distribution, with weights q0for a new cluster, and Nkqkfor each cluster k. The number of unique values of θn is constantly changing, so the size of κmust be adjusted whenever θchanges in every estimation step. If a new cluster is drawn from the categorical distribution, we must know what distribution to sample. The prior distribution of a new cluster is G0, but since the residual pair belongs to the cluster, we sample from its posterior distribution. This is the same process as sampling means and variance for a finite mixture basis that contains only a single point—a multivariate Bayesian linear regression. We draw from its posterior distribution in the standard way. Let Yk=(υk,ξk), which is the residual pair for the new cluster, the posterior distribution of component variance and mean are k∼IW (ν+1, V+S), μk|k∼N˜μ,1 1+aμ k, with S=Yk−ι¯μ kYk−ι¯μ k,+aμ(˜μk−¯μ)(˜μk−¯μ), ˜μ=(1+aμ)−1(¯ yk+aμ¯μ),and ¯ yk=Y kι, where ιis a corresponding length vector of all ones. To combine all of the above steps, we present the algorithm for updating the component classifier κin Algorithm 3. where prior parameters νdetermines the tightness of the Inverse-Wishart distribution, and aμdetermines the scale of variance of means. We allow prior parameter vto determine the location via mode()=ν ν+2vI.
1288 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) Algorithm 3 Drawing component clusters under a DP prior. 1: for n=1toNdo 2: Compute probability of new cluster, q0, for residual pair n 3: for k=1toKdo 4: Compute Bayes Factor qk. 5: end for 6: Draw classifier κn∼Multinomial(q) 7: if q== 0then 8: Draw cluster mean, μK+1and variance, K+1 9: Update K=K+1 10: end if 11: Check if a cluster has been orphaned. Adjust K 12: end for Updating the component distributions, Kand μkConditional on κj,t, each pair of residuals is known to come from a particular component of the mixture normal. If we only consider the residual pairs drawn from a particular component k,thenitisasifall of the residuals are drawn from the same distribution, and the standard Inverse-Wishart parameterization can be used to draw the variance parameters. We follow that procedure here for each component, with an extra step to allow for each component to have a different mean parameter as well. Since there is no intercept in the demand parameters, there is an extra degree of freedom in this problem that we use to sample as a mean for each component bivariate normal distribution. We sample from this mean using a multivariate regression with only a constant, since each component distribution is normal. Some care must be made since the residuals are not independent, so we use a Bayesian multivariate regression to correctly sample from their joint distribution. To utilize the standard Bayesian machinery for such a regression, we impose standard (normal) priors to exploit conjugate priors. For any component k,thevariancekhas an Inverse Wishart prior IW (ν,V)and the mean μk|khas a normal prior distribution N(¯μ,a−1 μk). Define the vector Yk=(υ,ξk), which is only the collection of residual pairs such that κj,t=k.Wecanwrite Yk=ιμ+U, where U∼N(0, k). The posterior covariance and conditional mean of the components are then k∼IW (ν+nk,V+S) and μk|k∼N˜μ,1 nk+aμ k,
Quantitative Economics 14 (2023) Demand estimation 1289 wherewedefine S=Yk−ι¯μ kYk−ι¯μ k+aμ(˜μk−¯μ)(˜μk−¯μ), ˜μ=(nk+aμ)−1(nk¯ yk+aμ¯μ),and ¯ yk=1 nk Y kι. The vector ιis a corresponding length vector of all ones, and nkis the number of observations in cluster k. This is repeated for each component k. Appendix B: Extension:Finite mixture components For computational speed or researcher preference, one may wish to put some restrictions on the joint distribution of the demand shock and the pricing error. We provide an extension from our more flexible model presented in the body of the paper to allow for a finite number of mixture components. That is, we treat the number of component distributions as fixed, and thus do not need to evaluate whether to add or remove components at each step of the sampler. After updating the unconditional mixture weights for each component, we only need to update the component distribution probabilities for each observation (ξj,t,υj,t)for the fixed set of components. To do so, we take the current candidate draw of πas a prior and evaluate the likelihood that this observation of (ξj,t,υj,t)is drawn from that component distribution given the current candidate mean and variance. Rather than clustering means and variance, we augment the data with a classifier for each observation using ˆπ, the posterior probabilities. Conditional upon the classifier, the residuals are distributed bivariate normal. We then update posterior cluster mean and variances based on the classified observations with a standard Gibbs step. All steps are shown in Algorithm 4. Algorithm 4 Hybrid Gibbs sampler: Finite mixture. 1: for c=1toCdo 2: Update arrivals λ(Gibbs) 3: Update shares s(·)(Metropolis–Hastings) 4: Update linear parameters β(Gibbs) 5: Update nonlinear parameters (Metropolis–Hastings) 6: Update pricing equation η(Gibbs) 7: (NEW) Update unconditional mixture weights π(Gibbs) 8: (NEW) Update component classifier κ(Gibbs) 9: Update mixture component parameters k,μk(Gibbs) 10: end for
1290 Hortaçsu, Natan, Parsley, Schwieg, and Williams Quantitative Economics 14 (2023) B.1 Details of sampling price-endogeneity parameters Updating mixing probabilities We assume a Dirichlet prior on the mixture probabilities, π∼Dirichlet(¯α). Conditional on the classifier κ, we have information about which data points fall into which classifier, and the posterior distribution of πis given by π∼Dirichlet(˜α), ˜αk=nk+¯αk. This gives the unconditional probability that a data point is drawn from classifier k. Updating component classifier This step now skips the need to (a) evaluate new component probabilities and (b) check for orphaned components. Rather than using a classifier κthat is sufficient for all unique values of θn, we augment the data with the classifier at each step of the chain. Each residual can be treated as drawn from a single, unobserved normal distribution, simplifying the computation required when evaluating its distribution. The classification of each data point can be thought of as a multinomial draw with πas the prior probability of each classification. The remaining information can be gathered from the likelihood of each component. We exploit the conjugacy nature of the multinomial distribution and the Dirichlet distribution, so that κj,t|π∼ Mulitnomial(¯πj,t)and ¯πj,t,k=πkφk((υj,t,ξj,t) K i=1 πiφi(υj,t,ξj,t) , where φk(x)is the likelihood of the kth component evaluated at x. This step is computationally expensive, as the number of computations is O(N×K). It requires evaluating the likelihood of each residual at every distribution, this must be evaluated with every draw, as ξ,υ, and the mean and variance of each component change each draw. Through careful application of the prior ω, and priors on the mean and variance of the components, complex distributions can be approximated with relatively few mixtures, which can reduce the computational burden of this procedure. Updating the component distributions, Kand μkThis step proceeds identically to the Dirichlet process case.