Full text
Journal of Econometrics 240 (2024) 105689 Available online 10 February 2024 0304-4076/© 2024 The Author(s). Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). Contents lists available at ScienceDirect Journal of Econometrics journal homepage: www.elsevier.com/locate/jeconom Non-representative sampled networks: Estimation of network structural properties by weighting✩ Chih-Sheng Hsieha, Yu-Chin Hsub,c,d, Stanley I.M. Koe, Jaromír Kováříkf,g,∗, Trevon D. Loganh,i aDepartment of Economics, National Taiwan University, Taipei, Taiwan bInstitute of Economics, Academia Sinica, Taipei, Taiwan cDepartment of Finance, National Central University, Taoyuan City, Taiwan dDepartment of Economics, National Chengchi University, Taipei, Taiwan eGraduate School of Economics and Management, Tohoku University, Japan fDpto. del Análisis Económico, University of the Basque Country UPV-EHU, Bilbao, Spain gFaculty of Arts & Faculty of Economics, University of West Bohemia, Pilsen, Czech Republic hDepartment of Economics, The Ohio State University, 410 Arps Hall, 1945 N. High Street, Columbus, OH, 43210, United States of America iNBER, United States of America ARTICLE INFO Keywords: Networks Weighting (Post-)stratification Non-representativeness Measurement errors ABSTRACT This paper analyzes statistical issues arising from non-representative samples of a network. Sampled network data could systematically bias the network properties and generate non-classical measurement error problems. Apart from the sampling rate and the elicitation procedure, the biases on network structural measures depend non-trivially on which subpopulations of nodes are missing with higher probability. We propose a methodology, adapting weighted estimators to networked contexts, which enables researchers to recover several network-level statistics and reduce the biases in the estimated network effects. The proposed weighted estimators are consistent and asymptotically normally distributed and have good performance in finite samples. Notably, our approach does not require users to assume any network formation model and is straightforward to implement. 1. Motivation There is growing interest in understanding the role of networks in Economics (Vega-Redondo,2007;Jackson,2010). Different ‘‘micro’’ and ‘‘macro’’ features of network architecture shape diffusion, learning, behavior, and other substantive phenomena in a variety of contexts. Due to the increasing availability of large network data sets and increasing computational power, empirical network research is now a dynamic and growing part of this literature. At the same time, empirical network analysis generates ✩We are grateful to Isaiah Andrews, Aureo de Paula, Marco van der Leij, and participants at numerous seminars for comments and suggestions. Hsieh acknowledges financial support from the National Science and Technology Council of Taiwan (NSTC110-2410-H-002-195). Hsieh and Hsu gratefully acknowledge the research support from the Center for Research in Econometric Theory and Applications of National Taiwan University, Taiwan (Grant no. 112L8601). Hsu gratefully acknowledges research support from the National Science and Technology Council of Taiwan (NSTC112-2628-H-001-001), and the Academia Sinica Investigator Award of Academia Sinica, Taiwan (AS-IA-110-H01). Kovářík acknowledges financial support from Ministerio de Economía y Competividad, Spain and Fondo Europeo de Desarrollo Regional (PID2019-106146GB-I00), the Basque Government, Spain (IT1461-22), and the Grant Agency of the Czech Republic (21-22796S). ∗Corresponding author at: Dpto. del Análisis Económico, University of the Basque Country UPV-EHU, Bilbao, Spain. E-mail addresses: [email protected] (C.-S. Hsieh), [email protected] (Y.-C. Hsu), [email protected] (S.I.M. Ko), [email protected] (J. Kovářík), [email protected] (T.D. Logan). https://doi.org/10.1016/j.jeconom.2024.105689 Received 27 July 2022; Received in revised form 7 January 2024; Accepted 10 January 2024
Journal of Econometrics 240 (2024) 105689 2 C.-S. Hsieh et al. new econometric challenges (Fortin and Boucher,2015;De Paula,2017;Jackson et al.,2017). This paper tackles the challenges that arise when network data come from non-representative samples of the population, which is the most commonly encountered scenario in practical applications. The vast majority of empirical network studies analyze sampled data, and the sampling rates are typically low.1Even though the literature across several disciplines has noted that using sampled data may lead to considerable biases and other statistical issues (see below for references), the typical approach is to treat the sampled data ‘‘as if’’ it were complete. Chandrasekhar and Lewis (2016) show formally that, even if the nodes are selected representatively through simple random sampling (SRS, henceforth), the statistics of the sampled networks differ significantly from those of the population network. This disparity results in measurement errors and inconsistency problems when we estimate network effects through regressions. The estimates from sampled networks may suffer from attenuation, expansion, or even sign-switching. As a result, one cannot rely on solutions to classical measurement-error problems to correct these issues, even if the sample is representative. Furthermore, nodes observed in network samples are typically non-representative. First, non-representativeness may be caused by the sampling design itself (Frank,1981;Kolaczyk,2009;Handcock and Gile,2010). For instance, the star subgraph sampling design analyzed in this paper is prone to including nodes with higher connectivity than nodes with a small number of network neighbors. The reason is that star subgraphs encompass not only the initially sampled nodes but also their network neighbors even if the latter were not initially sampled. Having more connections thus increases the probability of a node being included. This is an example of a design that generates samples in which the inclusion probabilities of nodes are endogenous to the underlying population network structure. Non-representativeness may also arise when the inclusion (or missing) probabilities are orthogonal to the population network architecture. For instance, when network samples are collected with specified boundaries such as within schools, or within villages, etc., it is not guaranteed that samples within boundaries are representative of the entire population. Such boundary-induced network samples are equivalent to the induced subgraph sampling design analyzed in this paper. Other common sources of non-representativeness in sampling studies are non-responses or disproportionate stratified sampling. Many studies exploit stratified samples to improve precision and sampling efficiency. Unfortunately, it is difficult and costly to stratify for all relevant characteristics. To intuitively explain the issues arising from sampled networks, we decompose the problem into two sources, scaling and nonrepresentativeness.Scaling refers to observing fewer nodes and edges than there exist in the whole network, independently of the (non-) representativeness of the sample. In contrast, non-representativeness arises when different nodes have unequal probabilities of being included in the sample. If nodes appear in the sample with equal probability, only scaling matters. As an example of the effect of scaling, let us consider the average degree of a network. When the links between the sampled and non-sampled nodes are not observed, the sample average degree is biased downwards by construction. Furthermore, suppose the average degree is correlated with the network’s diffusion properties. As a result, using the sample average degree in a regression analysis leads to an overestimation of the average degree’s impact on diffusion, even when samples are representative. This is an example of the expansion of the estimated effect and thus, non-classical measurement error. However, if nodes appear in the sample with unequal probabilities, whether the observed average degree and the estimates are inflated or attenuated will depend on who is missing. For example, if less connected nodes are missing with higher probability, scaling and non-representativeness can bias the average degree and the estimates in opposite directions, and one cannot easily predict which force will dominate. In contrast to the average degree, the global clustering coefficient and the homophily index can be unbiased in representative samples. In samples in which different types of nodes are missing with different probabilities, homophily will be biased by definition. Since clustering is typically associated with connectivity in social networks (Jackson and Rogers,2007), it is also likely to be mismeasured. The magnitude and direction of the biases in these characteristics and their estimated effects in regressions again depend crucially and non-trivially on who is missing. In this study, we systematically analyze the problems arising from sampled network data elicited via two widely employed sampling methods, and proposes a solution enabling to recover the true structural features of a network (e.g., average degree) and mitigate biases in regressions which study the impact of these network features on either individual or group-level behaviors and outcomes.2We first derive analytically weighted estimators for a set of network structural properties from sampled networks assuming that nodes appear in the sample with unequal probabilities according to their types. Secondly, we study the asymptotic properties of the proposed weighted estimators and evaluate their finite-sample performance numerically. Lastly, the proposed methodology is applied to a widely employed stratified data set on Indian villages (Banerjee et al.,2013).3This data set is suited for our approach because it contains a relatively large number of networks, and we document that the network data have been collected from a non-representative sample of the population under scrutiny. This study shows that relying on the assumption of representativeness to adjust network samples, which is rarely satisfied in real-world applications, can be as biased as using raw network samples without any adjustments. Since the direction and magnitude 1The reasons behind the common use of network samples are that the impracticality of analyzing the entire population and the higher costs associated with network elicitation compared to collecting basic individual characteristics (Aral,2016;Breza et al.,2020). Chandrasekhar and Lewis (2016) report that the median sampling rate in applied work in Economics is 25% and more than 66% of network studies have a sampling rate lower than 51%. Similar rates are found in other fields. 2Our study also improves inferences in network-formation applications studying contextual determinants of the network architecture (i.e., applying network properties as regressands). Since network formation represents a key topic in the network literature (see Jackson,2005 and De Paula,2020 for reviews), it enlarges the applicability of the proposed methodology. However, this study focuses on regressions including network properties as regressors. 3See, e.g., Jackson et al. (2012), Banerjee et al. (2013,2014), Chandrasekhar and Lewis (2016) and De Paula et al. (2018), among others.
Journal of Econometrics 240 (2024) 105689 3 C.-S. Hsieh et al. of the biases depend on who is missing, we demonstrate the necessity of accounting for potentially different missing rates of different types of nodes in applied work. This is particularly important in network data where population and distributional parameters are of primary interest. As the main contribution, we propose weighted estimators for a selected set of network characteristics that are widely used in applications: average degree, global clustering coefficient, epidemic threshold, and homophily index. These network features represent fundamental aspects of network architecture employed in theoretical and empirical research and provide intuitive insights regarding the way social organization shapes individual and group-level phenomena (Jackson et al.,2017).4To that aim, we assume that network members can be divided into a finite number of disjoint types, and that sampling rates differ across these types. Taking explicit account of the differing sampling rates across types, we adapt standard (network-free) Horvitz–Thompson (H–T) estimators to networked contexts and propose (post-) stratification as a viable approach to correct sampling biases caused either by the sampling procedure or due to varying non-response rates among different demographic or socioeconomic categories (or both) in order to improve the precision of sample estimates for objective variables of interest (Smith,1991;Little,1993). The main difference between the standard H–T estimators and our approach is to weight on network objects, such as links, triples, or triangles, rather than on nodes.5We prove that, in sparse networks, the proposed weighted estimators are consistent and asymptotically normally distributed.6 We also provide sufficient conditions so that we can ignore the estimation effects when the regression analysis includes the proposed weighted network measures as covariates. Our numerical analysis shows that our methodology performs well in finite samples and substantially outperforms both the naive (uncorrected) statistics from the raw data and corrections designed for representative samples. Our empirical application shows that the Indian village network data stratified on religion and geography are non-representative in terms of age and gender. We then show that not accounting for unequal missing rates for nodes of different types affects the estimated network effects substantially and one cannot easily predict the direction and magnitude of the biases. Given the differences, applied researchers should carefully consider to what extent their results might be driven by the non-representativeness of their samples. The present paper connects to three pieces of literature. First, our methodology complements emerging econometric literature on imperfectly measured network data and the estimation of network effects. Chandrasekhar and Lewis (2016) show that estimations with network data coming from representative samples suffer from non-classical measurement errors and propose a method to ensure consistent estimates. Their methodology consists of two alternative approaches. First, they provide formal corrections for several network measures. Our approach generalizes this first strategy. As a second approach, they propose a graphical reconstruction technique that delivers consistent estimates in both network-level and individual-level regressions. The procedure is first to estimate a network formation model and then employ the estimated model to interpolate over missing parts of the network. The network reconstruction approach requires a correct model specification and certain assumptions to ensure the consistency of the network effects. However, this second approach does not necessarily recover the structural properties of the population network, which is the primary objective of our study. Most importantly, from the perspective of the present work, both approaches are restricted to the case that the sample is representative. Chandrasekhar and Jackson (2016) propose a network formation model similar in spirit to our methodology in that it is also based on subgraphs in the function of types of nodes. However, none of these approaches can effectively recover the true network formation process from non-representative samples because, when the network-formation model is fitted on non-representative network samples, the estimated parameters in the first stage will likely be biased and potentially inconsistent even if the assumed model is correct. Thirkettle (2019) proposes a network formation model enabling the estimation of bounds on network statistics from partially observed networks. The advantage of our approach, as opposed to the graphical reconstruction techniques, is that our methodology does not rely on any assumed network formation model. Our work complements and expands the above studies by providing the first step toward the statistical treatment of network data coming from non-representative samples of the population, which is the most common type of network data available.7 Second, we contribute to the statistical sampling theory that has developed procedures for recovering the true network structural parameters from samples if the only source of non-representativeness comes from the network sampling design (see Kolaczyk, 2009 for a survey). Our methodology nests these procedures as a special case (e.g., Frank,1981,Kolaczyk,2009,Chandrasekhar and Lewis,2016). Our weighting method shares the same goals with these approaches but differs substantially in the underlying assumptions and applicability. Unlike these approaches which are typically suitable for specific sampling designs, our method can be applied or adapted to various sampling procedures.8Furthermore, our approach remains effective even in cases where nonrepresentativeness is caused by factors unrelated to the sampling design, such as non-response or the presence of hard-to-reach subpopulations. Most importantly, existing approaches assume certain forms of representativeness in the sampling process ex-ante, while our proposed methodology targets both ex-ante and ex-post non-representativeness of the sample. If the sampling rates are set by the researcher before the data collection (as in the approaches discussed above and in standard stratification), they are treated 4Section 6discusses the extension of our approach to other network characteristics. 5Such network objects are referred to as subgraphs, subnetworks, or network motifs in different fields. 6Since virtually all real-life social and economic networks are sparse, our asymptotic results are broadly applicable to empirical research. 7Boucher and Houndetoungan (2020) study the estimation of peer effects when the researchers only observe consistent estimates of aggregate network statistics. Hence, our methodology and their approach naturally complement each other in non-representative samples since our methodology delivers such consistent estimates from non-representative samples. 8For the sake of brevity, we concentrate on two sampling designs commonly used in economics. However, our methodology can be applied or adapted to other sampling designs. See Section 6.
Journal of Econometrics 240 (2024) 105689 4 C.-S. Hsieh et al. as known parameters. If the sampling rates are learned after the data collection, our methodology corresponds to post-stratification by exploiting the non-representativeness of the sample and treats the sampling rates as unknown parameters to be estimated. Last, we contribute to better practices for empirically evaluating the effects of global network features in socio-economic environments. Our study shows that, despite other econometric issues, mismeasured network features with non-representative samples might lead to a serious misunderstanding of network effects. However, our methodology mitigates this issue and provides an additional argument for the employment of sampling in empirical network work. With the increasing use of network data and corresponding empirical techniques, our proposed approach can improve the design of network sampling strategies and the inference we draw from network studies more generally. Moreover, it can serve as a standard robustness check of empirical results. 2. Framework 2.1. Notation A graph or network is defined by 𝐺𝑛= (𝑉 , 𝐸), where 𝑉is the set of vertices (nodes) with 𝑛=|𝑉|denoting the cardinality of 𝑉, and 𝐸is the set of edges (links). The network can be represented by an 𝑛×𝑛adjacency matrix 𝑊𝑛. We focus on unweighted and undirected networks; i.e., 𝑊𝑖𝑗,𝑛 = 1(0) if 𝑖and 𝑗are (not) connected and 𝑊𝑖𝑗,𝑛 =𝑊𝑗𝑖,𝑛 for each 𝑖, 𝑗 ∈𝑉. Following the convention, we exclude self-loops by setting 𝑊𝑖𝑖,𝑛 = 0. We assume that the nodes can be classified into 𝑇disjoint types with a generic type 𝑡∈= {1,2,…, 𝑇 }. One can view this classification as stratification, which can be carried out either before or after sample collection. When conducted after the collection, this process is commonly referred to as post-stratification. We write 𝑡𝑖=𝑡if node 𝑖is of type 𝑡. Then, 𝑡𝑖=𝑡𝑗(𝑡𝑖≠𝑡𝑗) indicates that 𝑖and 𝑗are (not) of the same type. Let 𝑉𝑡be the set of nodes of type 𝑡,𝑛𝑡=|𝑉𝑡|is the size of this set, and ∑𝑇 𝑡=1 𝑛𝑡=𝑛. Rather than the whole network 𝐺𝑛, researchers only observe the sampled network, which is also referred to as a subgraph of 𝐺𝑛. Let 𝑉∗⊆ 𝑉 be the set of sampled nodes of size 𝑚=|𝑉∗|and let 𝜓denote the sampling rate. Analogously, 𝑉∗ 𝑡denotes the set of nodes of type 𝑡in the sample and 𝑚𝑡=|𝑉∗ 𝑡|is the number of sampled nodes of type 𝑡and ∑𝑇 𝑡=1 𝑚𝑡=𝑚. We use 𝜓𝑡to denote type 𝑡’s sampling rate. We assume that 𝜓𝑡> 𝜏 for some 𝜏 > 0for each 𝑡and is independent of 𝑛. Crucially, we assume that, within each type, individual nodes have an equal probability of being selected into the sample. Our framework primarily focuses on non-representative samples, i.e., 𝜓𝑡≠𝜓𝑠for at least one 𝑡, 𝑠 ∈, while also encompassing the representative sample, i.e., 𝜓𝑡=𝜓for all 𝑡∈, as a special case. In the context of (ex-ante) stratification, the true value of 𝜓𝑡is a known quantity specified by the researcher. However, when it comes to post-stratification, the true value of 𝜓𝑡is treated as unknown and can be estimated by 𝜓𝑡=𝑚𝑡 𝑛𝑡, the ratio of the number of nodes of type 𝑡included in the sample to the population number of nodes of type 𝑡. We denote 𝜑𝑖=∑𝑇 𝑡=1 𝜓𝑡𝟏(𝑡𝑖=𝑡)the sampling probability of node 𝑖, conditional on her type, and 𝜑𝑖=∑𝑇 𝑡=1 𝜓𝑡𝟏(𝑡𝑖=𝑡)the corresponding estimator based on 𝜓𝑡. Given sampled nodes, this paper focuses on two designs for eliciting network edges. The first is the induced subgraph, in which the sampled network is denoted by 𝐺I 𝑛= (𝑉∗, 𝐸I). In 𝐺I 𝑛, the set 𝑉∗involves 𝑚sampled nodes, and the set 𝐸I⊆ 𝐸 involves network links among these 𝑚sampled nodes. 𝑊I 𝑛is the 𝑚×𝑚adjacency matrix corresponding to 𝐺I 𝑛. The second is the star subgraph, in which the sampled network is denoted by 𝐺S 𝑛= (𝑉∗, 𝐸S).9In 𝐺S 𝑛, there are 𝑚initially sampled nodes in the set 𝑉∗ 0. However, researchers observe not only the network links among these 𝑚sampled nodes, but also the links of the 𝑚sampled nodes to unsampled nodes in 𝑉. Hence, we use 𝐸Sto denote the set of edges such that at least one node of the corresponding dyad is in 𝑉∗ 0. The set 𝑉∗ 0is enlarged to 𝑉∗by including all the vertices 𝑖∈𝑉⧵𝑉∗ 0that are connected through the observed links to at least one sampled node from 𝑉∗ 0. The size of this enlarged vertex set is denoted by 𝑚′=|𝑉∗|, and the corresponding sampling rate is denoted by 𝜓′. Let 𝑊S 𝑛be the 𝑚′×𝑚′adjacency matrix corresponding to the graph 𝐺S 𝑛. In both the induced and star subgraphs, we assume that edges are reported without errors. We study several network structural properties (measures) and refer to a generic population network measure as 𝛬. Let 𝛬(𝐺𝑛) denote the estimated network measure for 𝛬based on the whole network data, and let 𝛬(𝐺𝑛),𝐺𝑛∈ {𝐺I 𝑛,𝐺S 𝑛}, represent the corresponding estimated network measure based on the sampled network 𝐺𝑛. We call 𝛬(𝐺𝑛)the naive estimator of network property. Additionally, let 𝛬(𝐺𝑛)denote the weighted network measure proposed to mitigate sample biases with respect to the whole network. For example, 𝛬(𝐺𝑛) = 1 𝑛∑𝑖∈𝑉∑𝑗∈𝑉𝑊𝑖𝑗,𝑛 is the average degree of a graph, which we denote 𝑑(𝐺𝑛)below. Hence, 𝑑(𝐺𝑛)is the average degree of the sampled network, and 𝑑(𝐺𝑛)is the proposed weighted estimator to mitigate biases of 𝑑(𝐺𝑛). In applications, researchers may observe multiple networks. We use a generic subscript 𝑟∈= {1,2,…, 𝑅}when a measure refers to network 𝑟. That is, 𝐺𝑟,𝑛𝑟denotes the graph 𝑟, and 𝐺𝑟,𝑛𝑟∈ {𝐺I 𝑟,𝑛𝑟,𝐺S 𝑟,𝑛𝑟}denotes the corresponding sampled network. Therefore, 𝑛𝑟,𝑡 and 𝑚𝑟,𝑡 are the number of nodes of type 𝑡in the whole network 𝑟and its corresponding number in the sample. 2.2. Regression with network measures In addition to the reconstruction of network properties of interest, we also consider regression analysis with network measures. Throughout the analysis, we focus on regressions in which researchers are interested in understanding whether and how the global measures of network properties influence a particular outcome. Formally, 𝑦𝑟=𝛼+𝛽𝛬𝑟+𝑥𝑟𝛾+𝜀𝑟,(1) 9𝐺S 𝑛is referred to as the labeled star subgraph in Kolaczyk (2009) because the unsampled nodes which connect to sampled nodes are identified and labeled.
Journal of Econometrics 240 (2024) 105689 5 C.-S. Hsieh et al. where 𝑦𝑟is the outcome variable of network (or community) 𝑟,𝑥𝑟is the set of network-level controls, and 𝛬𝑟is the population network property of interest for 𝑟th network population. The researchers are interested in estimating the parameters 𝛼,𝛽, and 𝛾. Examples of the applications of (1) in the literature include Alatas et al. (2016) which regress the ability of villagers to aggregate information on a set of network characteristics in Indonesian villages, Banerjee et al. (2013) who model microfinance take-up rate in rural India in function of the average centrality of the initial seeds, Currarini et al. (2009) and Golub and Jackson (2012) who relate homophily with school-level statistics using Add Health data, or Fleming et al. (2007) who model the ability of different regions to generate knowledge depending on the structure of regional research networks. Such regressions are also of interest theoretically. For example, the overall clustering of a network may explain the magnitude and efficiency of risk-sharing within a society (Bloch et al.,2008), and the stability of behavior in a society may be related to the minimal eigenvalue of the adjacency matrix (Bramoullé et al.,2014). The proposed approach also applies to models investigating the influence of a network’s global measure on individual-level outcomes: 𝑦𝑖𝑟 =𝛼+𝛽𝛬𝑟+𝑥𝑖𝑟𝛾+𝜇𝑟+𝜀𝑖𝑟, where 𝑦𝑖𝑟 is the outcome of an individual 𝑖in network 𝑟,𝑥𝑖𝑟 captures individual heterogeneity (that can also include the heterogeneity of 𝑖’s neighborhood), and 𝜇𝑟is a network random effect. For instance, the decision of an individual to adopt a product (e.g., microfinance as in Banerjee et al.,2013), participate in an activity (e.g., recreational activity as in Bramoullé et al.,2009), or behave in a particular way (Centola,2010) can depend on the overall structure of the network. In the same vein, the innovation literature studies how the structure of regional networks shapes the innovative performance of individual innovators (Schilling and Phelps,2007). There also exist theories arguing that the overall structure of a network may determine the behavior at the individual level (see, e.g., Ballester et al.,2006;Bramoullé et al.,2014). With sampled data, researchers observe 𝐺𝑟,𝑛𝑟∈ {𝐺I 𝑟,𝑛𝑟, 𝐺S 𝑟,𝑛𝑟}, and the naive estimator 𝛬(𝐺𝑟,𝑛𝑟)is not a consistent estimator for 𝛬𝑟. Therefore, when researchers estimate 𝑦𝑟=𝛼+𝛽𝛬(𝐺𝑟,𝑛𝑟) + 𝑥𝑟𝛾+𝑢𝑟,(2) it leads to a measurement error in the regressor. The classic measurement error and the resulting attenuation bias are based on several assumptions that are generally not satisfied in the case of network measures.10 Chandrasekhar and Lewis (2016) show analytically and via simulations that the biases are generally not tractable and can lead to expansion or sign switching under representativeness. The issues become even more problematic if the representativeness assumption is violated. On the other hand, when researchers estimate 𝑦𝑟=𝛼+𝛽 𝛬(𝐺𝑟,𝑛𝑟) + 𝑥𝑟𝛾+𝑢𝑟,(3) it leads to consistent estimation of the parameters. In this regard, we later provide sufficient conditions such that we can ignore the estimated effect of 𝛬(𝐺𝑟,𝑛𝑟)in the OLS regression. 3. Weighted estimators for sampled network measures This section proposes weighted estimators for commonly used network measures when sampled data are used. We also address the biases present in both the naive (unweighted) estimators and weighted estimators that solely account for scaling effects. One key assumption made throughout this section is that the network measures (statistics) under consideration are well-defined. For example, when the sampling rate is extremely low, the global clustering coefficient could be zero if no closed triplets are observed in the sampled network. In such scenarios, both the naive estimator and our proposed weighted estimator are null and it is not possible to recover the true value of the coefficient. Hence, we stress that our corrections for network statistics are applicable under sampling rates in which well-defined naive estimators exist. To maintain notational simplicity, we will omit the network index 𝑟 in the subscripts throughout Sections 3.1 and 3.2. We reintroduce it when discussing the asymptotics of regressions with network measures in Section 3.3. 3.1. Average degree The degree is the number of connections of a node, which is a basic measure of a node’s importance or local centrality. The average degree of the graph 𝐺𝑛is simply the average number of network links per node in the network, defined as 𝑑(𝐺𝑛) = 1 𝑛∑𝑖∈𝑉∑𝑗∈𝑉𝑊𝑖𝑗,𝑛. It has been applied as a regressor in numerous empirical studies of different contexts (see, e.g., Branas-Garza et al.,2010;Banerjee et al.,2013;Alatas et al.,2016, among many others). For an induced subgraph, the naive estimator of the average degree is computed as 𝑑(𝐺I 𝑛) = 1 𝑚∑ 𝑖∈𝑉∗∑ 𝑗∈𝑉∗ 𝑊I 𝑖𝑗,𝑛 =1 𝑚∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛𝐷𝑖𝐷𝑗,(4) where 𝐷𝑖is a binary variable that takes the value 1 if 𝑖∈𝑉∗, and 0 otherwise. To correct the biases from both scaling and nonrepresentativeness in (4), we propose the weighted sample average degree by multiplying each observed sample edge 𝑊I 𝑖𝑗,𝑛 with the 10 Although network regressions often face additional challenges such as endogeneity and omitted variable problems, we contend that the sampling issue persists even in the absence of these problems.
Journal of Econometrics 240 (2024) 105689 6 C.-S. Hsieh et al. weight, (𝜑𝑖𝜑𝑗)−1, which is the inverse of the estimated inclusion probability. Thus, the weighted sample average degree is given by 𝑑(𝐺I 𝑛) = 1 𝑛∑ 𝑖∈𝑉∗∑ 𝑗∈𝑉∗ 𝑊I 𝑖𝑗,𝑛(𝜑𝑖𝜑𝑗)−1 =1 𝑛∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛 𝐷𝑖 𝜑𝑖 𝐷𝑗 𝜑𝑗 .(5) As the true value of the inclusion probability (𝜑𝑖𝜑𝑗)is typically unknown and needs to be estimated from the sample, we refer to the weighted estimator in (5) as a post-stratification estimator. However, when the true value of (𝜑𝑖𝜑𝑗)is known and applied in (5), the estimator follows the general principle of the H–T estimator (Horvitz and Thompson,1952). To show why the proposed weighted estimator (5) removes the bias in (4), assume known (𝜑𝑖𝜑𝑗). Then, E(1 𝑛∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛 𝐷𝑖 𝜑𝑖 𝐷𝑗 𝜑𝑗|||||| 𝐺𝑛)=1 𝑛∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛 (E(𝐷𝑖𝐷𝑗|𝐺𝑛) 𝜑𝑖𝜑𝑗)=1 𝑛∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛 (𝜑𝑖𝜑𝑗 𝜑𝑖𝜑𝑗)=1 𝑛∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛. That is, the expected value of our weighted estimator is the true average degree of the population network. The intuition behind (5) is as follows. There are ∑𝑖∈𝑉∑𝑗∈𝑉𝑊𝑖𝑗,𝑛 edges to account for in 𝐺𝑛. However, due to variations in the inclusion probabilities of sample edges, we only observe ∑𝑖∈𝑉∑𝑗∈𝑉𝑊𝑖𝑗,𝑛(𝜑𝑖𝜑𝑗)edges in an induced subgraph in expectation. Even if samples are representative (i.e., 𝜑𝑖=𝜑𝑗=𝜓), as long as 𝜓 < 1, a bias emerges due to scaling. Moreover, as 𝜑𝑖and 𝜑𝑗are not necessarily the same, we have the second source of bias, non-representativeness, and the issues become more complicated. For a star subgraph, the naive (sample) average degree is defined as 𝑑(𝐺S 𝑛) = 1 𝑚′∑ 𝑖∈𝑉∗∑ 𝑖∈𝑉∗ 𝑊S 𝑖𝑗,𝑛 =1 𝑚′∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛(1 − (1 − 𝐷𝑖)(1 − 𝐷𝑗)),(6) where 𝐷𝑖is a binary variable that takes the value 1 if 𝑖∈𝑉∗ 0, and 0 otherwise. To correct the bias, we propose the following weighted sample average degree, 𝑑(𝐺S 𝑛) = 1 𝑛∑ 𝑖∈𝑉∗∑ 𝑗∈𝑉∗ 𝑊S 𝑖𝑗,𝑛 (1 − (1 − 𝜑𝑖)(1 − 𝜑𝑗))−1 =1 𝑛∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛 1 − (1 − 𝐷𝑖)(1 − 𝐷𝑗) 1 − (1 − 𝜑𝑖)(1 − 𝜑𝑗).(7) Once again, assuming that 𝜑𝑖’s are known, we can demonstrate the following: E(1 𝑛∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛 1 − (1 − 𝐷𝑖)(1 − 𝐷𝑗) 1 − (1 − 𝜑𝑖)(1 − 𝜑𝑗)|||||| 𝐺𝑛)=1 𝑛∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛 (E(1 − (1 − 𝐷𝑖)(1 − 𝐷𝑗)|𝐺𝑛) 1 − (1 − 𝜑𝑖)(1 − 𝜑𝑗)) =1 𝑛∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛 (1 − (1 − 𝜑𝑖)(1 − 𝜑𝑗) 1 − (1 − 𝜑𝑖)(1 − 𝜑𝑗))=1 𝑛∑ 𝑖∈𝑉∑ 𝑗∈𝑉 𝑊𝑖𝑗,𝑛. This result justifies why the weighted sample average in (7) mitigates the bias problem. The weighted estimators proposed in (5) and (7) account for two phenomena. Firstly, they account for the differing inclusion probabilities of the links in the function of the types of the involved nodes. Secondly, they respect the correlations in who is connected to whom in the observed part of the network (i.e., they respect the network homophily). If one applies the corrections assuming representativeness of the sample, (5) and (7) will change to 𝑑(𝐺I 𝑛) = 1 𝑛∑ 𝑖∈𝑉∗∑ 𝑗∈𝑉∗ 𝑊I 𝑖𝑗,𝑛(𝜓2)−1 (8) and 𝑑(𝐺S 𝑛) = 1 𝑛∑ 𝑖∈𝑉∗∑ 𝑗∈𝑉∗ 𝑊S 𝑖𝑗,𝑛(1 − (1 − 𝜓)2)−1,(9) respectively. These corrections are exactly the same as shown by Chandrasekhar and Lewis (2016). However, biases would still emerge in (8) and (9) if the sample is not truly representative. Importantly, there is no reason for these biases to be smaller than in the raw (uncorrected) data as their size depends on who is missing. One can perceive the proposed weighted estimators (5) and (7) as a design-based approach. However, the remainder of this subsection characterizes the asymptotic properties of the estimators. To this aim, we envision that the underlying finite-population network (𝐺𝑛)expands progressively toward a hypothetical superpopulation network. This prompts a natural transition to a modelbased approach, targeting the unknown (model) parameters that characterize this hypothetical superpopulation for asymptotic statistical inference.11 Consequently, we advocate a synthesis of design-based and model-based approaches (Binder and Roberts, 2003;Sterba,2009). We expand upon the framework introduced by Bickel et al. (2011) to account for non-representativeness of nodes under the assumption that the network is sparse. Sparse networks refer to networks where the number of observed links is considerably lower than the maximum number of possible links, a common feature of real-life social networks. Formally, sparseness 11 An alternative possibility is asymptotic analysis with finite population sampling (Prášková and Sen,2009;Li and Ding,2017). However, the methodology cannot currently handle sparse networks in finite populations. Consequently, we adopt the approach in Bickel et al. (2011) to investigate the asymptotics of a superpopulation and defer the analysis of a finite population for future research.
Journal of Econometrics 240 (2024) 105689 7 C.-S. Hsieh et al. is defined as the property of an infinite sequence of graphs where the (average) degree is bounded as 𝑛→∞(Bickel and Chen, 2009;Lovász,2012).12 In addition, similar to Bickel et al. (2011), we assume that the adjacency matrix of the whole network 𝑊𝑛is exchangeable.13 As a result, according to the Aldous–Hoover theorem (Aldous,1981;Hoover,1979), the adjacency matrix can be represented by 𝑊𝑖𝑗,𝑛 D =𝑔𝑛(𝜉𝑖, 𝜉𝑗, 𝜖𝑖𝑗 , 𝑡𝑖, 𝑡𝑗),(10) where D =denotes equality in distribution, and 𝑔𝑛is a measurable function symmetric in its first two and last two arguments. In (10), 𝜉𝑖and 𝜖𝑖𝑗 are i.i.d. uniform random variables on [0,1],𝜖𝑖𝑗 =𝜖𝑗𝑖, and {𝑡𝑖}𝑛 𝑖=1 are independent of {𝜉𝑖}𝑛 𝑖=1 and {𝜖𝑖𝑗 }𝑛 𝑖,𝑗=1. Note that this implies 𝑊𝑖𝑗,𝑛 =𝑊𝑗𝑖,𝑛. Since the function 𝑔𝑛(.)in (10) cannot be uniquely identified (Bickel and Chen,2009), it would be advisable to explore an alternative parameterization, ℎ𝑡𝑠,𝑛(𝑢, 𝑣)≡P[𝑊𝑖𝑗,𝑛 = 1|𝜉𝑖=𝑢, 𝜉𝑗=𝑣, 𝑡𝑖=𝑡, 𝑡𝑗=𝑠]for 𝑡, 𝑠 ∈, which refers to the unique canonical ℎ𝑡𝑠,can such that ∫1 0ℎ𝑡𝑠,can(𝑢, 𝑣)𝑑𝑣 is monotone non-decreasing in 𝑢. Also, let 𝑝𝑡= P(𝑡𝑖=𝑡)and assume for all 𝑡∈,𝑝𝑡≥𝜏for some 𝜏 > 0, and is independent of 𝑛. Under these assumptions, we have ℎ𝑡𝑠,𝑛(𝑢, 𝑣) = ℎ𝑠𝑡,𝑛(𝑢, 𝑣)and ℎ𝑛(𝑢, 𝑣) = P[𝑊𝑖𝑗,𝑛 = 1|𝜉𝑖=𝑢, 𝜉𝑗=𝑣] = ∑𝑇 𝑡=1 ∑𝑇 𝑠=1 ℎ𝑡𝑠,𝑛(𝑢, 𝑣)𝑝𝑡𝑝𝑠. Let 𝜌𝑛=∫1 0∫1 0 ℎ𝑛(𝑢, 𝑣)𝑑𝑢 𝑑𝑣 (11) be the probability of an edge in the network (i.e., network density). We can then write 𝑤𝑡𝑠,𝑛(𝑢, 𝑣) = 𝜌−1 𝑛ℎ𝑡𝑠,𝑛(𝑢, 𝑣), which represents the conditional density of (𝜉𝑖, 𝜉𝑗)given that there is an edge between 𝑖and 𝑗. The expression 𝑤𝑡𝑠,𝑛 decouples the network density from the inhomogeneity structure. For the asymptotics, we will assume that 𝑤𝑡𝑠,𝑛(𝑢, 𝑣) = 𝑤𝑡𝑠(𝑢, 𝑣), where 𝑤𝑡𝑠(𝑢, 𝑣)is independent of 𝑛. Let 𝑤𝑡(𝑢, 𝑣) = ∑𝑇 𝑠=1 𝑤𝑡𝑠(𝑢, 𝑣)𝑝𝑠and 𝑤(𝑢, 𝑣) = ∑𝑇 𝑡,𝑠=1 𝑤𝑡𝑠(𝑢, 𝑣)𝑝𝑡𝑝𝑠=∑𝑇 𝑡=1 𝑤𝑡(𝑢, 𝑣)𝑝𝑡. We will control the rate of the expected degree 𝜆𝑛= (𝑛− 1)𝜌𝑛>0as 𝑛→∞.14 The asymptotics of the average degree for the whole network 𝐺𝑛,𝑑(𝐺𝑛), and our proposed weighted estimators 𝑑(𝐺I 𝑛)and 𝑑(𝐺S 𝑛) can be summarized in the following theorem.15 Theorem 1. Suppose that ∫1 0∫1 0𝑤2(𝑢, 𝑣)𝑑𝑣𝑑𝑢 < ∞and lim𝑛→∞𝜆𝑛=𝜆 < ∞. Then, (a) under the population (non-sampled) network, 𝑑(𝐺𝑛)𝑝 →𝜆, √𝑛(𝑑(𝐺𝑛) − 𝜆)𝑑 →(0, 𝜎2 𝑑(𝐺)) for some 𝜎2 𝑑(𝐺)>0; (b) under the induced subgraph, 𝑑(𝐺I 𝑛)𝑝 →𝜆, √𝑛( 𝑑(𝐺I 𝑛) − 𝜆)𝑑 →(0, 𝜎2 𝑑(𝐺I)) for some 𝜎2 𝑑(𝐺I)>0; (c) under the star subgraph, 𝑑(𝐺S 𝑛)𝑝 →𝜆, √𝑛( 𝑑(𝐺S 𝑛) − 𝜆)𝑑 →(0, 𝜎2 𝑑(𝐺S)) for some 𝜎2 𝑑(𝐺S)>0. Theorem 1 establishes that, if the network is sparse, our weighted estimators 𝑑(𝐺I 𝑛)and 𝑑(𝐺S 𝑛)are consistent and asymptotically normally distributed with finite variance.16 This is the case independently of whether the sampling rates are treated as estimators or not. Section 4complements the asymptotic analysis (in Theorem 1 and Supplementary Appendix B) with numerical analysis assessing to what extent the corrections proposed in this section differ from their true values in finite samples. 12 The sparse networks that we consider here are restricted to a particular class of networks with 𝑜(𝑛2)edges, or equivalently 𝑜(𝑛)network degrees, and do not contain dense spots. This may thus preclude some real-life social networks that exhibit the power-law degree distribution (Borgs et al.,2019). 13 To be precise, a network is relatively exchangeable with respect to the type variable 𝑡if [𝑊𝜎𝑡(𝑖)𝜎𝑡(𝑗),𝑛]D = [𝑊𝑖𝑗,𝑛] for all 𝑛and all permutations 𝜎𝑡satisfying [𝑡𝜎𝑡(𝑖)]𝑖∈𝑛= [𝑡𝑖]𝑖∈𝑛(Crane and Towsner,2018). Exchangeability implies a particular dependence structure across the elements of 𝑊𝑖𝑗,𝑛. In particular, 𝑊𝑖𝑗,𝑛 and 𝑊𝑖′𝑗′,𝑛 are dependent if 𝑖=𝑖′or 𝑗=𝑗′. This type of dependence is implied by many statistical and econometric network formation models, such as stochastic blockmodels (Holland et al.,1983), latent position model (Hoff et al.,2002), and other conditional edge independence models (Chandrasekhar,2016). More details of this framework are available in the handbook chapter by Graham (2020). 14 Specifically, we require 𝜌𝑛=𝛩(1∕𝑛), i.e., 𝜌𝑛grows as fast as 1∕𝑛, so that 𝜆𝑛converges to a non-zero constant when 𝑛goes to infinity. 15 The proofs of Theorem 1 and Lemma 2 are relegated to the Supplementary Appendix B. 16 The analytical expressions for the asymptotic variances of 𝑑(𝐺I 𝑛)and 𝑑(𝐺S 𝑛)are complex due to intricate network patterns (Bickel et al.,2011;Bhattacharyya and Bickel,2015;Graham,2020) and we therefore leave their derivations for future research. Nevertheless, Bickel et al. (2011) propose subsampling bootstrap methods for approximation and conjecture–although do not prove–that these methods might work properly in sparse networks; see p. 2291–2292 of their paper.
Journal of Econometrics 240 (2024) 105689 8 C.-S. Hsieh et al. 3.2. Other network measures In addition to the average degree, we also study three other fundamental network measures: the global clustering coefficient, epidemic threshold, and homophily index. We will now provide a brief introduction to these three network measures. Global clustering coefficient. The global clustering coefficient is defined as the ratio between the number of closed triplets 𝑇𝑐(𝐺𝑛) and the number of connected triples 𝑁𝑐(𝐺𝑛)in the network (Watts and Strogatz,1998),17 calculated as 𝑐(𝐺𝑛) = 𝑇𝑐(𝐺𝑛) 𝑁𝑐(𝐺𝑛),(12) where 𝑇𝑐(𝐺𝑛) = 1 2∑ 𝑖∈𝑉∑ 𝑗∈𝑉∑ 𝑘∈𝑉 𝑖≠𝑗≠𝑘 𝑊𝑖𝑗,𝑛𝑊𝑗𝑘,𝑛𝑊𝑘𝑖,𝑛 and 𝑁𝑐(𝐺𝑛) = 1 2∑ 𝑖∈𝑉∑ 𝑗∈𝑉∑ 𝑘∈𝑉 𝑖≠𝑗≠𝑘 𝑊𝑖𝑗,𝑛𝑊𝑗𝑘,𝑛. The global clustering coefficient has traditionally been considered a measure of social capital. For example, it plays an important role in risk-sharing (Bloch et al.,2008), trust building (Karlan et al.,2009), job search (Ruiz-Palazuelos et al.,2023), and enhancing cooperation (Granovetter,1985). Several empirical studies have used the global clustering coefficient as a regressor or a dependent variable (e.g., Fleming et al.,2007;Alatas et al.,2016). Supplementary Appendix A.1 shows that the naive estimators of the global clustering coefficients (12) under the induced and star subgraphs display biases. We propose their corrections, which differ from those for the average degree: rather than edges connecting dyads (pairs of individuals), we adjust ‘‘relationships’’ involving three individuals, taking into account their interconnections as closed triplets or connected triples, and accounting for the associated sampling probabilities. Nevertheless, research has demonstrated that the global clustering coefficient in (12) approaches zero in sparse networks as 𝑛→∞(see Supplementary Appendix B.2 for a formal proof; see also Bhattacharyya and Bickel 2015 and Graham 2020 for further discussion). Therefore, the asymptotic analysis of 𝑐(𝐺𝑛)is uninformative. To overcome this issue, we follow the literature, employing a normalized global clustering coefficient which converges to a non-zero value asymptotically and is robust to network size, network density, and degree heterogeneity. In particular, we focus on the normalized coefficient proposed in Li et al. (2019), calculated as 𝑐𝑛𝑜𝑟𝑚(𝐺𝑛) = 𝑇𝑐(𝐺𝑛) 3(𝑛 3)(𝑛𝑑(𝐺𝑛) 2(𝑛 2))3 (𝑁𝑐(𝐺𝑛) 3(𝑛 3))3=(𝑛− 2)2𝑡𝑟(𝑊3 𝑛)(𝟏′𝑊𝑛𝟏)3 𝑛(𝑛− 1)(𝟏′𝑊2 𝑛𝟏−𝑡𝑟(𝑊2 𝑛))3,(13) where 𝟏is 𝑛-dimensional vector of 1’s. It is straightforward to see that the normalization in (13) balances the exponents regarding the network size 𝑛and the adjacency matrix 𝑊𝑛in the numerator and denominator. Therefore, as 𝑛→∞, the numerator and the denominator will converge at the same rate.18 After rearranging, (13) can be expressed as follows: 𝑐𝑛𝑜𝑟𝑚(𝐺𝑛) = 𝜁𝑛 𝑇𝑐(𝐺𝑛)𝑑(𝐺𝑛)3 (𝑁𝑐(𝐺𝑛) 𝑛)3,(14) with 𝜁𝑛=(𝑛−2)2 4𝑛(𝑛−1) . Supplementary Appendices A.1 and B.2 analyze the biases in the naive estimators of the normalized global clustering coefficient in (14), provide the corresponding corrections, and show that the weighted estimators are consistent and asymptotically normally distributed as 𝑛→∞. Epidemic threshold. There is an increasing interest in understanding the diffusion properties of networks. The epidemic threshold is one way to quantify how easy it is for a disease, information, idea, or behavior to propagate through a network. The applications range from product adoption (Banerjee et al.,2013), spread of information (Alatas et al.,2016) to spread of behaviors (Centola,2010). There is a large variety of epidemic thresholds, depending on the diffusion conditions and network properties (see, e.g., Vega-Redondo,2007, and Jackson,2010). We focus on the following widely used version, based on the mean-field approximation (Pastor-Satorras and Vespignani,2002): 𝛿(𝐺𝑛) = 1 𝑛∑𝑖∈𝑉∑𝑗∈𝑉𝑊𝑖𝑗,𝑛 1 𝑛∑𝑖∈𝑉(∑𝑗∈𝑉𝑊𝑖𝑗,𝑛)2. Supplementary Appendices A.2 and B.3 show that the naive estimators are biased and our proposed weighted estimators are consistent and normally distributed. 17 The number of closed triplets also equals three times the number of triangles. A triangle refers to a complete subnetwork of three individuals, which consists of three closed triplets, one centered on each node. A connected triple is a three-node subnetwork in which at least two edges are present. Hence, every triangle is a connected triple, but the reverse is not necessarily true. 18 In addition to (13), there is an alternative normalized global clustering coefficient proposed in Bhattacharyya and Bickel (2015), which we discuss in further detail in Supplementary Appendix B.2. We focus on (13) in the main text for the sake of brevity.
Journal of Econometrics 240 (2024) 105689 9 C.-S. Hsieh et al. Homophily index. Social and economic networks exhibit a feature called homophily, a tendency to bond with similar individuals. In social and economic networks, who links with whom is typically correlated with characteristics such as gender, age, race, and social and economic status, among others (see McPherson et al.,2001 for a survey). This phenomenon of ‘‘birds of a feather flock together’’ gains particular relevance in our approach because we explicitly consider the types of nodes in the network. Homophily is an important measure of cross-type segregation and affects many economically relevant phenomena such as diffusion or learning and their speeds (Golub and Jackson,2012), labor market outcomes (Calvo-Armengol and Jackson,2004), or individual and firm-level success (McPherson and Smith-Lovin,1987). We adopt the homophily index from Currarini et al. (2009). The index for type 𝑡is defined as 𝐻𝑡(𝐺𝑛) = 𝑑𝑡𝑡(𝐺𝑛) 𝑑𝑡(𝐺𝑛), where 𝑑𝑡𝑡(𝐺𝑛) denotes the average number of friendships that agents of type 𝑡have within the same type and 𝑑𝑡(𝐺𝑛)denotes the average number of friendships that type 𝑡form regardless of others’ types. Supplementary Appendix A.3 contains detailed derivations of the weighted estimator for 𝐻𝑡(𝐺𝑛)under induced and star subgraphs. Supplementary Appendix B.4 again proves that our weighted estimators are consistent and asymptotically normally distributed. 3.3. Asymptotics of regressions with estimated network measures This section discusses the asymptotic properties of OLS regressions in (3), in which our weighted estimators are employed as regressors. Suppose we have 𝑅networks. Let 𝑛𝑟denote the number of nodes in the 𝑟th network for 𝑟= 1,…, 𝑅. Let 𝑛𝑟=𝑎𝑟⋅𝑛and assume that 0< 𝜍𝓁≤𝑎𝑟≤𝜍𝑢<∞for all 𝑟with 𝜍𝓁and 𝜍𝑢being constants that do not depend on 𝑟. That is, we assume that the number of nodes in each network is of the same order. It is well-known that in OLS regressions with covariates being estimated, if the estimating error of the regressor is independent of the regression error 𝜖𝑟and max𝑟=1,…,𝑅{| 𝛬(𝐺𝑟,𝑛𝑟) − 𝛬𝑟|} = 𝑜𝑝(1), the estimation effect can be ignored asymptotically. To be specific, let (𝛼𝑖𝑛, 𝛽𝑖𝑛, 𝛾′ 𝑖𝑛)′= arg min (𝛼,𝛽,𝛾′) 1 𝑅 𝑅 ∑ 𝑟=1(𝑦𝑟−𝛼−𝛽𝛬𝑟−𝑥𝑟𝛾)2, (𝛼, 𝛽, 𝛾′)′= arg min (𝛼,𝛽,𝛾′) 1 𝑅 𝑅 ∑ 𝑟=1 (𝑦𝑟−𝛼−𝛽 𝛬(𝐺𝑟,𝑛𝑟) − 𝑥𝑟𝛾)2,(15) where (𝛼𝑖𝑛, 𝛽𝑖𝑛, 𝛾′ 𝑖𝑛)′denotes the infeasible OLS estimator because the true 𝛬𝑟’s are not observable and (𝛼, 𝛽, 𝛾′)′denotes the OLS estimator when the true 𝛬𝑟’s are replaced with their estimates. If max𝑟=1,…,𝑅{| 𝛬(𝐺𝑟,𝑛𝑟) − 𝛬𝑟|} = 𝑜𝑝(1), then √𝑅((𝛼, 𝛽, 𝛾′)′− ( 𝛼𝑖𝑛, 𝛽𝑖𝑛, 𝛾′ 𝑖𝑛)′)=𝑜𝑝(1), i.e., the infeasible OLS estimator and the OLS estimator based on estimated 𝛬(𝐺𝑟,𝑛𝑟)’s are asymptotically equivalent. In other words, we can treat 𝛬(𝐺𝑟,𝑛𝑟)’s as the true 𝛬𝑟’s in the regression without the need to correct for the estimation effect of 𝛬(𝐺𝑟,𝑛𝑟)’s. The following lemma provides sufficient conditions for max𝑟=1,…,𝑅{| 𝛬(𝐺𝑟,𝑛𝑟) − 𝛬𝑟|} = 𝑜𝑝(1). Lemma 2. Assume that the variance of √𝑛( 𝛬(𝐺𝑟,𝑛𝑟) − 𝛬𝑟)is uniformly bounded above by a finite constant 𝑀, for all 𝑟and 𝑛≥𝑁for some finite large number 𝑁, and 𝑅∕𝑛→0. Then, max𝑟=1,…,𝑅{| 𝛬(𝐺𝑟,𝑛𝑟) − 𝛬𝑟|} = 𝑜𝑝(1). 4. Monte Carlo simulations This section complements the previous one in assessing the performance of our approach in finite samples. In particular, we evaluate numerically the estimation biases in the network measures under study (in Sections Section 4.1), as well as the network effects when using these measures as regressors in regression analysis (in Section 4.2). The evaluation considers various factors such as the sampling design (induced vs. star subgraph), the sampling rate, and whether SRS (representativeness) is assumed when applying the weighted estimators. We quantify the biases present in the naive estimators and the corrections made under the SRS assumption and compare their performances vis-à-vis our post-stratification estimators. For ease of interpretation, we concentrate on the scenarios that mimic our modeling assumptions. In this simulation exercise, we demonstrate the effectiveness of our post-stratification approach by analyzing the network measures discussed in Section 3. These measures include the average degree, global clustering coefficient, normalized global clustering coefficient, epidemic threshold, and homophily index.19 The network data in our simulation study are adopted from the Add Health Wave-I In-school data.20 In particular, we adopt one school as a prototype.21 By adopting the real-life friendship network 19 We include both the standard global clustering coefficient and its normalized variant. Our corrections of the latter are asymptotically well-behaved and we would like to assess its performance in finite samples. However, the (non-normalized) coefficient is widely employed in the literature. Hence, although we know it converges to zero in sparse networks asymptotically (see Section 3.2), we analyze its performance in finite samples. 20 This is a program project designed by J. Richard Udry, Peter S. Bearman, and Kathleen Mullan Harris, and funded by a grant P01-HD31921 from the National Institute of Child Health and Human Development, with cooperative funding from 17 other agencies. Special acknowledgment is due Ronald R. Rindfuss and Barbara Entwisle for assistance in the original design. Persons interested in obtaining data files from Add Health should contact Add Health, Carolina Population Center, 123 W. Franklin Street, Chapel Hill, NC 27516-2524 ([email protected]). No direct support was received from grant P01-HD31921 for this analysis. 21 This adopted school is a public suburban school with 1606 students from grades 9 to 12. The school is located in the southern U.S.
Journal of Econometrics 240 (2024) 105689 16 C.-S. Hsieh et al. Table 1 Population and sample shares of different characteristics and labor market outcomes in the Indian rural village data from Banerjee et al. (2013). Population Sample Diff. (𝑝-value) Age <30 38.71% 30.97% 7.74% (0.000) 30–50 39.60% 54.11% −14.51% (0.000) >50 21.69% 14.92% 6.77% (0.000) Male 50.34% 44.57% 5.77% (0.000) Household size <317.26% 15.49% 1.77% (0.038) 3–8 71.57% 73.48% −1.91% (0.039) >811.17% 11.03% 0.14% (0.879) Labor market outcome employed 62.49% work outside village 21.21% Number of villages 75 75 Observations 48,646 16,995 collected the census information for each household in all villages. Subsequently, they conducted a comprehensive follow-up survey with a subset of each village, wherein they also recorded the networks of relationships among surveyed individuals. As is common in most studies, the survey respondents only represent a sample of each village, and their reported network is an induced subgraph of the whole network. The average sampling rate across villages is 35%. The crucial aspect of the sampling design in Banerjee et al. (2013) is the stratification by religion and geographic sub-location, generating a representative sample with respect to these two variables. This is a common approach in many applications. Despite the stratification based on religion and geography, Table 1 reveals that the data are not representative in terms of age, gender, and–to a lesser extent–household size. Below, we show to what extent the differences between the village population and sample shares of these categories affect the estimation of network effects in regressions discussed in Section 2.2. The data contain several variables regarding the labor market outcomes of the participants, such as their employment status, whether they work outside the village, and their occupation. Since the important role of social networks in labor markets is widely acknowledged (Granovetter,1985;Calvo-Armengol and Jackson,2004;Cingano and Rosolia,2012), we ask how the village employment rate and the fraction of people working outside the village correlate with the global features of the underlying network of relationships within the village.25 Theoretical literature suggests that both connectivity and the global clustering coefficient can have a direct impact on employment prospects (Calvo-Armengol and Jackson,2004;Ruiz-Palazuelos et al.,2023). Additionally, the epidemic threshold can indirectly influence labor outcomes by affecting the flow of labor-market information (Calvo-Armengol and Jackson,2004). Similarly, the degree of segregation can determine which individuals have access to job information and those who do not. Most importantly, for the present study, we ask how the estimated network effects change if we account for non-representativeness of the network sample. We hypothesize that the over-representation of individuals aged 30–50 and the under-representation of men in the sample (as evident in Table 1), who are typically more active participants in labor markets in a country like India, could bias the estimated network effects if this misrepresentation is not taken into account. Table 2 reports the estimated network effects in a series of regressions differing in (i) the dependent variable (employment rate or fraction of working outside the village), (ii) whether raw sample statistics or corrections are used and (iii) different network measures. As for (ii), to separate the effect of scaling from the effect on non-representativeness of the sample, we use the naive estimators (denoted Raw in Table 2), corrections assuming SRS (denoted SRS), and our approach in which we weight on crosscharacteristics (incorporating the information on age, gender, and household size; denoted Cross). Table 1 illustrates the distributions of these three variables, from which we compute the 𝜓𝑡for the 3 × 2 × 3 = 18 types according to the variable Cross. Each row reports the estimated network effect (and the standard error robust to heteroskedasticity in parentheses) from a separate regression of one dependent variable on the corresponding network statistic and village size, mimicking the structure of the regressions in Section 2.2. We also apply the post-stratification weighting on the dependent variables (i.e., employment and working outside villages) at the village level to correct measurement errors.26 Consequently, the columns Cross provide a typical example of standard post-stratification with a reasonable number of stratification groups, where the sampling rates are estimated from the differences between the sample and population shares of auxiliary variables. Since we show that our approach delivers consistent estimates, we believe that applied researchers should report estimates such as those in the columns Cross as their main result while estimating the effect of network measures on outcomes in non-representative samples or, at least, as a robustness check of their main analysis. As for the influence of village networks on labor market outcomes, our findings support existing literature, highlighting the significant role played by the structure of social networks in shaping labor markets. By accounting for the non-representativeness 25 To maintain simplicity and align better with the assumptions of our analysis, we focus on a simpler application compared to Banerjee et al. (2013), who propose a more intricate estimation strategy. 26 We use the network constructed by the union of all relationships reported by survey respondents (e.g., borrowing, lending, seeking advices, going to temple together, visiting home, etc.). We find similar results if we only focus on friendships (see Table C.1 in the Supplementary Appendix).
Journal of Econometrics 240 (2024) 105689 17 C.-S. Hsieh et al. Table 2 Estimated network effects on the labor market outcomes of villagers in rural India villages. Dependent variable (I) Employed (%) (II) Work outside village (%) Raw SRS Cross Raw SRS Cross Average degree 0.0269*** 0.0091** 0.0088** −0.0235*−0.0093** −0.0101* (0.0095) (0.0035) (0.0039) (0.0120) (0.0044) (0.0051) Global clustering 0.4989** 0.4989** 0.4240** −0.6410** −0.6410** −0.4996*** (0.1930) (0.1930) (0.1830) (0.2666) (0.2666) (0.1879) Norm. global clustering −0.0047 −0.0047 −0.0022 0.0038 0.0038 0.0037 (0.0061) (0.0061) (0.0051) (0.0048) (0.0048) (0.0054) Epidemic threshold −1.1530*** −2.3017*** −2.0965** 0.9357** 2.3498** 2.3924** (0.3589) (0.8442) (0.8341) (0.4148) (0.9967) (1.0430) HI-male 0.1939*0.1939*0.1445 −0.1374 −0.1374 −0.0463 (0.1028) (0.1028) (0.0939) (0.1490) (0.1490) (0.1621) HI-middle age 0.2848 0.2848 −0.2386 −0.5150** −0.5150** 0.0010 (0.1956) (0.1956) (0.2119) (0.2033) (0.2033) (0.2428) HI-small household size 0.0930 0.0930 0.1866** −0.2677** −0.2677** −0.0818 (0.0991) (0.0991) (0.0856) (0.0992) (0.0992) (0.0920) Note: Regressions are based on 75 villages. Standard errors robust to heteroskedasticity are reported in parentheses. Each row represents a separate regression with a different network measure, and the village size is included in every regression as a default control. Raw indicates an unweighted sample statistic, SRS signifies the correction based on the representativeness assumption, and Cross denotes the weighting on the Cross characteristic variable. * Stand for significance at 10%. ** Stand for significance at 5%. *** Stand for significance at 1%. of the sample (as indicated by the Cross columns in Table 2), certain features of the social networks have a meaningful impact on average labor outcomes within the village. Moreover, the effects of these various network characteristics largely exhibit consistency with one another. Regarding the main purpose of this exercise, Table 2 shows the sensitivity of the results with respect to (non-)representativeness of network samples. In contrast to Section 4, we do not know the true impact of the different network measures. However, since the data and the performed regressions match the assumptions behind our approach, all the previous analysis suggests that the results using our methodology are consistent, less biased, and more stable than either the naive estimators or corrections assuming representativeness. As a result, the following discussion provides an informal assessment of the disparities among the results obtained from naive estimators, the corrections assuming SRS, and our approach. Table 2 documents that the estimates using raw data or corrections based on the representativeness assumption are mostly expanded compared to the corrections that account for both scaling and the non-representativeness of the network data. However, we also observe instances of attenuation and even sign-switching. There are three cases in which we observe a network effect when employing the naive estimators or corrections under SRS, but this effect does not show up using our weighting approach. In one other case, the network effect is absent with the naive estimators and corrections for scaling, but this effect becomes significant in the Cross column. All these four cases are associated with the impact of homophily. Quantitatively speaking, the effect of the average degree, when based on raw data, is overestimated in Table 2 by over 130% compared to the effect observed through our weighting approach. Likewise, the effect of the global clustering coefficient is overestimated by more than 17%, while the effect of the epidemic threshold is underestimated by at least 45%. Hence, some of these differences are economically significant. The corrections assuming representativeness either alleviate or maintain the biases when compared to the results obtained through our approach. These corrections effectively reduce the biases with respect to the Cross column to below 10% for the average degree and epidemic threshold. However, the biases remain economically significant for network measures that are unbiased in representative samples but generally biased in non-representative samples, such as the clustering coefficients and the homophily indices. In sum, significant differences are present between the estimates obtained using our approach and those from the naive estimators as well as the corrections assuming representativeness. These findings suggest that false positives (or negatives), expansion of network effects, and sign switching might be common phenomena resulting from non-representativeness of network samples. Given that most network data share the underlying properties of this data, the results here imply that applied researchers should consider the effect of weighting on the sign, size, and magnitude of network effects. More importantly, the direction and the magnitude of the biases depend non-trivially on the particular network statistics, the dependent variable under study, and who is missing. Hence, this exercise corroborates that researchers cannot easily predict the direction of the biases and consequently, they should not rely on classical measurement-error solutions, even in the simplest cases analyzed here. 6. Discussion This section discusses potential extensions and limitations of our methodology and provide several recommendations concerning the selection of auxiliary variables for weighting.
Journal of Econometrics 240 (2024) 105689 18 C.-S. Hsieh et al. Alternative Network Sampling Designs. Although this paper focuses on the induced and star subgraphs, the proposed methodology can be adapted to other sampling schemes as long as the researcher knows the strategy employed for the elicitation of the sample and possesses some information about the whole population. We present several examples illustrating how the proposed approach can be applied to different sampling strategies and discuss cases where our methodology cannot be directly applied, or requires modification. As a first example, consider the issue known as the boundary specification problem. Researchers sometimes set a boundary to determine the whole network of interest. Imagine a researcher who collects a network sample from a few classes within a school, excluding individuals from other classes and any connections between the classes under investigation and individuals outside the class. Although the sampled network may provide a comprehensive representation of the analyzed classes, it remains incomplete in capturing the entirety of the true social network within the school. If one would like to study the school network, and individual characteristics are available for the whole school, one can mitigate the boundary specification problem by applying our method directly because setting a boundary is mathematically equivalent to the induced subgraph sampling. As a second example, consider snowball sampling, a sampling procedure commonly applied in Sociology, Marketing, and Epidemiology (see, e.g., Berg,2004;Browne,2005). In snowball sampling, a researcher begins by randomly selecting seed nodes. These seeds serve as the starting point for the first wave, during which the researcher collects information on all the contacts of the initially selected nodes. In subsequent waves, the researcher expands the sample by eliciting the contacts of the nodes identified in the previous wave, and this process continues iteratively. Note that conducting a one-wave snowball sampling is essentially equivalent to the star subgraph sampling approach discussed earlier, thus making our methodology directly applicable. The literature has suggested corrections for one-wave snowball sampling (Frank,1977;Kolaczyk,2009), but these corrections only align with our approach when the initial seeds are representative samples of the population. We argue this is rarely the case even in very carefully and systematically collected data sets. Although the computation becomes increasingly complex as more waves are performed, one can adapt our approach to multiple waves of snowball sampling taking into account the missing frequencies of each type and the information about the within-type and across-type connectivity from the observed part of the network using combinatorial arguments. In fact, our methodology has certain parallelism with Respondent Driven Sampling (Heckathorn,1997), a weighting approach on snowball samples to compensate analytically for the non-randomness of snowball-sampling procedures. In contrast to this approach that corrects for the non-representativeness ex-ante, our approach adjusts for these issues ex-post by mitigating the discrepancy between the sampled and population networks and treating the sampling rates as estimators. Unsurprisingly, our corrections cannot be applied to some alternative sampling designs or should be tailored to the specific sampling strategy employed in the corresponding study. Consider, for example, random selection of links (also known as random edge sampling) where an individual 𝑖is included in the sample if at least one of her edges is sampled. Such sampling is commonplace in communication data, where only random samples of phone calls or e-mails are selected. We do not target this procedure in this study as additional assumptions would be necessary, but see, e.g., Kolaczyk (2009) for a potential direction. Relatedly, our approach assumes that, conditionally on observing a particular sample of nodes and the sampling design, the links are observed perfectly. That is, this study specifically analyzes issues arising from imperfect observation of network members but cannot solve issues arising from mismeasured links (see, e.g., Hardy et al.,2019). A notable example of this issue is the truncated fixed-choice survey design, where respondents are constrained to nominate a certain number of friends (e.g., up to ten friends). Our approach mitigates the biases due to the non-representativeness but not those due to the truncation. However, both issues might be targeted simultaneously by combining our post-stratification weighting with the approach proposed by Griffith (2022), which is specifically designed to mitigate the issues due to the truncation. Similarly, additional applications of our approach might result from combining our approach with methods designed for other purposes. The extension of our approach to these other more specialized sampling procedures is left for future research. Other Network Measures. Due to their theoretical and empirical relevance, this study focuses on four fundamental network measures commonly seen in the empirical literature. Nevertheless, one can adapt the methodology to other measures that solely require the knowledge of nodes’ local information.27 The first set of examples allows for a direct application of our methodology, which includes the assortativity coefficient and the average size of the second-order neighborhood. Assortativity plays a crucial role in the process of diffusion, as it can either impede or facilitate the transmission of diseases, behaviors, and social norms (Newman, 2002;Jackson et al.,2017). The average size of the second-order neighborhood enables us to assess how fast diffusion spreads, and it is important in labor markets (Calvo-Armengol and Jackson,2004). Since the computation of both the assortativity coefficient and the second-order neighborhood only requires the knowledge of an individual’s degree and the degrees of their neighbors, their weighted corrections follow directly from Section 3. Other measures do not follow directly from Section 3, but our approach can still be applied. For instance, Eagle et al. (2010) apply the concept of entropy to capture the diversity of connections of an individual to different types in the network. Since their measure only relies on the neighborhood of each node, the corrected variation of this measure for sampled networks is straightforward. Similarly, cycles of length four have recently received certain attention in sociology (Opsahl,2013) and economics (Ruiz-Palazuelos et al.,2023). One can recover it following our approach using the combinatorial logic. Since these characteristics are extensions of the ideas of homophily and the global clustering coefficient, respectively, we focus on the more common variations and do not propose the corrections of these two in this study. 27 In this paper, local information always refers to the firstand second-order neighborhoods of each node. One can go further and incorporate more distant neighbors probably at the cost of lower precision of the proposed corrections.
Journal of Econometrics 240 (2024) 105689 19 C.-S. Hsieh et al. The proposed methodology cannot recover global network measures computed based on the entire network architecture. This includes spectral properties, average betweenness or eigenvalue centrality, and network distances. However, there is a rich literature proposing approximations, bounds, or ‘‘plug-in’’ estimators computed on the basis of nodes’ local information (e.g., Van Mieghem, 2010;Comellas and Gago,2007). Hence, one can correct these bounds and approximations using our approach either directly or by plugging some of our corrections into more general expressions. Future research shall establish the finite-sample as well as asymptotic properties of such bounds, approximations, and plug-in estimators. The proposed approach cannot recover the network characteristics at the individual node level. Selection of (Auxiliary) Weighting Variables. A natural question arising from the proposed methodology is the choice of the (auxiliary) weighting variables for post-stratification. The evidence points out that different characteristics matter in different contexts and situations. For instance, Morelli et al. (2017) report that positive emotions explain positioning in network reflecting time sharing, while empathy plays a role in intimate networks of the same people describing trust and support. Similarly, firms may form ties differently if searching for providers (or buyers) compared to innovation collaborations. Hence, one has to know the particular application under study to assess which node-level characteristic might provide valuable information about the network and we prefer to refrain from making general recommendations regarding the application of particular variables. For this reason, we would generally encourage applied researchers to first analyze the degree of non-representativeness and then use that information to inform the variables chosen for the correction. Practically speaking, most data sets are limited to a relatively small set of variables that encompass census information. Since our results show that the performance improves with more information and applying variables that provide no information about the network does not affect the performance negatively, we recommend employing all the available information in such cases. In contrast, when many variables are available for weighting, a problem would be to have too few observations in each stratified cell. This can lead to an increase in variance, resulting in reduced efficiency of the weighting estimates for the characteristic being studied. One straightforward solution is to apply the principal component analysis to filter the relevant independent information from a large number of potentially correlated variables and construct the weights using the discretized components. Another solution can be a simple two-step algorithm, outlined in Supplementary Appendix E, that we propose for the selection of the ‘‘right’’ variables. We remain agnostic about the specific approach a researcher would take for a particular project. However, that researchers should be aware of the inferential problem addressed here and the general limits of treating the network as if it were complete or assuming representativeness of the network sample. Given that sensitivity, a variety of weights should be used to discover if the results are sensitive to accounting for non-representativeness. Such analysis should serve as a standard robustness check of empirical network results, giving scholars confidence that the results reflect network effects and are not a figment of the sampling strategy. Appendix A. Supplementary data Supplementary material related to this article can be found online at https://doi.org/10.1016/j.jeconom.2024.105689. References Alatas, Vivi, Banerjee, Abhijit, Chandrasekhar, Arun G., Hanna, Rema, Olken, Benjamin A., 2016. Network structure and the aggregation of information: Theory and evidence from Indonesia. Amer. Econ. Rev. 106 (7), 1663–1704. Aldous, David J., 1981. Representations for partially exchangeable arrays of random variables. J. Multivariate Anal. 11 (4), 581–598. Aral, Sinan, 2016. Networked experiments. In: The Oxford Handbook of the Economics of Networks. Oxford, UK: Oxford University Press, pp. 376–411. Ballester, Coralio, Calvó-Armengol, Antoni, Zenou, Yves, 2006. Who’s who in networks. Wanted: The key player. Econometrica 74 (5), 1403–1417. Banerjee, Abhijit, Chandrasekhar, Arun G., Duflo, Esther, Jackson, Matthew O., 2013. The diffusion of microfinance. Science 341 (6144), 1236498. Banerjee, Abhijit, Chandrasekhar, Arun G., Duflo, Esther, Jackson, Matthew O., 2014. Gossip: Identifying central individuals in a social network. No. w20422 NBER Working paper. Berg, Sven, 2004. Snowball sampling—I. Encycl. Stat. Sci. 12. Bhattacharyya, Sharmodeep, Bickel, Peter J., 2015. Subsampling bootstrap of count features of networks. Ann. Statist. 43 (6). Bickel, Peter J., Chen, Aiyou, 2009. A nonparametric view of network models and Newman–Girvan and other modularities. Proc. Natl. Acad. Sci. 106 (50), 21068–21073. Bickel, Peter J., Chen, Aiyou, Levina, Elizaveta, 2011. The method of moments and degree distributions for network models. Ann. Statist. 39 (5), 2280–2301. Binder, David A., Roberts, Georgia R., 2003. Design-based and model-based methods for estimating model parameters. Anal. Survey Data 29, 33–54. Bloch, Francis, Genicot, Garance, Ray, Debraj, 2008. Informal insurance in social networks. J. Econom. Theory 143 (1), 36–58. Borgs, Christian, Chayes, Jennifer, Cohn, Henry, Zhao, Yufei, 2019. An 𝐿𝑝theory of sparse graph convergence I: Limits, sparse random graph models, and power law distributions. Trans. Amer. Math. Soc. 372 (5), 3019–3062. Boucher, Vincent, Houndetoungan, Aristide, 2020. Estimating peer effects using partial network data. Working paper. Bramoullé, Yann, Djebbari, Habiba, Fortin, Bernard, 2009. Identification of peer effects through social networks. J. Econometrics 150 (1), 41–55. Bramoullé, Yann, Kranton, Rachel, D’amours, Martin, 2014. Strategic interaction and networks. Amer. Econ. Rev. 104 (3), 898–930. Branas-Garza, Pablo, Cobo-Reyes, Ramón, Espinosa, María Paz, Jiménez, Natalia, Kovářík, Jaromír, Ponti, Giovanni, 2010. Altruism and social integration. Games Econom. Behav. 69 (2), 249–257. Breza, Emily, Chandrasekhar, Arun G., McCormick, Tyler H., Pan, Mengjie, 2020. Using aggregated relational data to feasibly identify network structure without network data. Amer. Econ. Rev. 110 (8), 2454–2484. Browne, Kath, 2005. Snowball sampling: using social networks to research non-heterosexual women. Int. J. Soc. Res. Methodol. 8 (1), 47–60. Calvo-Armengol, Antoni, Jackson, Matthew O., 2004. The effects of social networks on employment and inequality. Amer. Econ. Rev. 94 (3), 426–454. Centola, Damon, 2010. The spread of behavior in an online social network experiment. Science 329 (5996), 1194–1197. Chandrasekhar, Arun, 2016. Econometrics of network formation. In: The Oxford Handbook of the Economics of Networks. pp. 303–357. Chandrasekhar, Arun G., Jackson, Matthew O., 2016. A network formation model based on subgraphs, Working paper. Available at SSRN: https://ssrn.com/ abstract=2660381. Chandrasekhar, Arun, Lewis, Randall, 2016. Econometrics of sampled networks, Working paper.
Journal of Econometrics 240 (2024) 105689 20 C.-S. Hsieh et al. Cingano, Federico, Rosolia, Alfonso, 2012. People I know: job search and social networks. J. Labor Econ. 30 (2), 291–332. Comellas, F., Gago, S., 2007. Spectral bounds for the betweenness of a graph. Linear Algebra Appl. 423 (1), 74–80. Crane, Harry, Towsner, Henry, 2018. Relatively exchangeable structures. J. Symbolic Logic 83 (2), 416–442. Currarini, Sergio, Jackson, Matthew O., Pin, Paolo, 2009. An economic model of friendship: Homophily, minorities, and segregation. Econometrica 77 (4), 1003–1045. De Paula, Aureo, 2017. Econometrics of network models. In: Advances in Economics and Econometrics: Eleventh World Congress. In: Econometric Society Monographs, Cambridge University Press, Cambridge, pp. 268–323, De Paula, Áureo, 2020. Econometric models of network formation. Annu. Rev. Econ. 12, 775–799. De Paula, Áureo, Rasul, Imran, Souza, Pedro, 2018. Recovering social networks from panel data: Identification, simulations and an application. Working paper. Eagle, Nathan, Macy, Michael, Claxton, Rob, 2010. Network diversity and economic development. Science 328 (5981), 1029–1031. Fleming, Lee, King, III, Charles, Juda, Adam I., 2007. Small worlds and regional innovation. Organ. Sci. 18 (6), 938–954. Fortin, Bernard, Boucher, Vincent, 2015. Some challenges in the empirics of the effects of networks. In: The Oxford Handbook of the Economics of Networks. Frank, Ove, 1977. Survey sampling in graphs. J. Statist. Plann. Inference 1 (3), 235–264. Frank, Ove, 1981. A survey of statistical methods for graph analysis. Sociol, Methodol, 12, 110–155. Golub, Benjamin, Jackson, Matthew O., 2012. How homophily affects the speed of learning and best-response dynamics. Q. J. Econ. 127 (3), 1287–1338. Graham, Bryan S., 2020. Network data. In: Handbook of Econometrics, vol. 7, Elsevier, pp. 111–218. Granovetter, Mark, 1985. Economic action and social structure: The problem of embeddedness. Am. J. Sociol. 91 (3), 481–510. Griffith, Alan, 2022. Name your friends, but only five? the importance of censoring in peer effects estimates using social network data. J. Labor Econ. 40 (4), 779–805. Handcock, Mark S., Gile, Krista J., 2010. Modeling social networks from sampled data. Annals of Applied Statistics 4 (1), 5. Hardy, Morgan, Heath, Rachel M., Lee, Wesley, McCormick, Tyler H., 2019. Estimating spillovers using imprecisely measured networks. arXiv preprint arXiv:1904.00136. Heckathorn, Douglas D., 1997. Respondent-driven sampling: a new approach to the study of hidden populations. Soc. Problems 44 (2), 174–199. Hoff, Peter D., Raftery, Adrian E., Handcock, Mark S., 2002. Latent space approaches to social network analysis. J. Amer. Statist. Assoc. 97 (460), 1090–1098. Holland, Paul W., Laskey, Kathryn Blackmond, Leinhardt, Samuel, 1983. Stochastic blockmodels: first steps. Soc. Netw. 5 (2), 109–137. Hoover, Douglas N., 1979. Relations on Probability Spaces and Arrays of Random Variables, Preprint. vol. 2, Princeton, NJ, p. 275. Horvitz, Daniel G., Thompson, Donovan J., 1952. A generalization of sampling without replacement from a finite universe. J. Amer. Statist. Assoc. 47 (260), 663–685. Jackson, Matthew O., 2005. A survey of network formation models: stability and efficiency. Group Form. Econ. Netw. Clubs Coalitions 11–49. Jackson, Matthew O., 2010. Social and Economic Networks. Princeton University Press. Jackson, Matthew O., Rodriguez-Barraquer, Tomas, Tan, Xu, 2012. Social capital and social quilts: Network patterns of favor exchange. Amer. Econ. Rev. 102 (5), 1857–1897. Jackson, Matthew O., Rogers, Brian W., 2007. Meeting strangers and friends of friends: How random are social networks? Amer. Econ. Rev. 97 (3), 890–915. Jackson, Matthew O., Rogers, Brian W., Zenou, Yves, 2017. The economic consequences of social-network structure. J. Econ. Lit. 55 (1), 49–95. Karlan, Dean, Mobius, Markus, Rosenblat, Tanya, Szeidl, Adam, 2009. Trust and social collateral. Q. J. Econ. 124 (3), 1307–1361. Kolaczyk, Eric D., 2009. Statistical Analysis of Network Data: Methods and Models. Springer Science & Business Media. Li, Xinran, Ding, Peng, 2017. General forms of finite population central limit theorems with applications to causal inference. J. Amer. Statist. Assoc. 112 (520), 1759–1769. Li, Ting, Yu, Xianshi, Jing, Bing-Yi, 2019. Measuring the clustering strength of a network via the normalized clustering coefficient. arXiv preprint arXiv:1908.00523. Little, Roderick J.A., 1993. Post-stratification: a modeler’s perspective. J. Amer. Statist. Assoc. 88 (423), 1001–1012. Lovász, László, 2012. Large Networks and Graph Limits, vol. 60, American Mathematical Society. McPherson, J. Miller, Smith-Lovin, Lynn, 1987. Homophily in voluntary organizations: Status distance and the composition of face-to-face groups. Am. Sociol. Rev. 370–379. McPherson, Miller, Smith-Lovin, Lynn, Cook, James M., 2001. Birds of a feather: Homophily in social networks. Annu. Rev. Sociol. 27 (1), 415–444. Morelli, Sylvia A., Ong, Desmond C., Makati, Rucha, Jackson, Matthew O., Zaki, Jamil, 2017. Empathy and well-being correlate with centrality in different social networks. Proc. Natl. Acad. Sci. 114 (37), 9843–9847. Newman, Mark E.J., 2002. Assortative mixing in networks. Phys. Rev. Lett. 89 (20), 208701. Opsahl, Tore, 2013. Triadic closure in two-mode networks: Redefining the global and local clustering coefficients. Social Networks 35 (2), 159–167. Pastor-Satorras, Romualdo, Vespignani, Alessandro, 2002. Immunization of complex networks. Phys. Rev. E 65 (3), 036104. Prášková, Zuzana, Sen, Pranab Kumar, 2009. Asymptotics in finite population sampling. Handbook of Statist. 29, 489–522. Ruiz-Palazuelos, Sofía, Espinosa, María Paz, Kovářík, Jaromír, 2023. The weakness of common job contacts. Eur. Econ. Rev. 160, 104594. Schilling, Melissa A., Phelps, Corey C., 2007. Interfirm collaboration networks: The impact of large-scale network structure on firm innovation. Manage. Sci. 53 (7), 1113–1126. Smith, Terence M.F., 1991. Post-stratification. J. Royal Stat. Soc. Series D 40 (3), 315–323. Sterba, Sonya K., 2009. Alternative model-based and design-based frameworks for inference from samples to populations: From polarization to integration. Multivar. Behav. Res. 44 (6), 711–740. Thirkettle, Matthew, 2019. Identification and estimation of network statistics with missing link data. Working paper. Van Mieghem, Piet, 2010. Graph Spectra for Complex Networks. Cambridge University Press. Vega-Redondo, Fernando, 2007. Complex Social Networks, No. 44. Cambridge University Press. Watts, Duncan J., Strogatz, Steven H., 1998. Collective dynamics of ‘‘small-world’’ networks. Nature 393 (6684), 440–442.