scieee AI-readable full text Open interactive document viewer

Ethnic and religious polarization and social conflict

Esteban, Joan; Mayoral, Laura

Abstract

In this paper we examine the link between ethnic and religious polarization and conflict using interpersonal distances for ethnic and religious attitudes obtained from the World Values Survey. We use the Duclos et al (2004) polarization index. We measure conflict by means on an index of social unrest, as well as by the standard conflict onset or incidence based on a threshold number of deaths. Our results show that taking distances into account significantly improves the quality of the fit. Our measure of polarization outperforms the measure used by Montalvo and Reynal-Querol (2005) and the fractionalization index. We also obtain that both ethnic and religious polarization are significant in explaining conflict. The results improve when we use an indicator of social unrest as the dependent variable.

Full text

Ethnic and Religious Polarization and Social Conflict Joan Esteban ∗ Laura Mayoral † January 13, 2011 First version November 2, 2009 Abstract In this paper we examine the link between ethnic and religious polarization and conflict using interpersonal distances for ethnic and religious attitudes obtained from the World Values Survey. We use the Duclos et al (2004) polarization index. We measure conflict by means on an index of social unrest, as well as by the standard conflict onset or incidence based on a threshold number of deaths. Our results show that taking distances into account significantly improves the quality of the fit. Our measure of polarization outperforms the measure used by Montalvo and Reynal-Querol (2005) and the fractionalization index. We also obtain that both ethnic and religious polarization are significant ∗ Institut d’Anàlisi Económica, CSIC, and Barcelona GSE; email: [email protected]. Financial support from the Axa Research Fund, the Government of Catalonia, and the CICYT project n. SEJ2006-00369 is gratefully acknowledged. † Institut d’Anàlisi Económica, CSIC; email: [email protected]. Member of the Barcelona GSE Research Network funded by the Government of Catalonia. Financial support from the CICYT project n. SEJ2006-00369 is gratefully acknowledged. 1 in explaining conflict. The results improve when we use an indicator of social unrest as the dependent variable. JEL-Classification: Key-words: conflict, polarization, fractionalization, ethnicity, religion. Acknowledgements: Ignacio Ortuño-Ortín has gracefully computed for us the linguistic distances between groups based on the ethnologue dataset. We are grateful to Borek Vasicek for a highly competent assistance. 2 1. INTRODUCTION In this paper we examine the link between ethnic and religious polarization and conflict. Income inequality has traditionally been seen as a major potential cause of conflict. Early empirical studies focussed on the personal distribution of income or of landownership. 1 However, as the survey article by Lichbach (1989) concluded, the empirical results obtained were lacking of significance and ambiguous. This apparent lack of connection between economic inequality and social conflict has been possibly due to the fact that most of the domestic conflicts since 1945 have had a strong ethnic/religious component. Indeed, Horowitz (1985) already noted that the Marxian prophecy of an inevitable class struggle has ended up having an ethnic fulfillment. 2 Following the contribution by Easterly and Levine (1997) the attention of empirical research has shifted towards the ethnic and religious social divides as a cause of conflict and low collective action. The index of fractionalization,F, has been the most widely used measure of the ethnic/religious composition of a country. But other indicators such as the Gini-Greenberg 3 index G or the polarization indices by Esteban and Ray (1994) Pand by Reynal-Querol (2002) RQ 4 have been used as well. The empirical works by Alesina et al. (1999), Alesina et al. (2003), Alesina and La Ferrara (2005), Collier (2001), Collier and Hoffler (2004), Desmet et al (2009,2010), Fearon (2003), Fearon and Laitin (2003), Miguel et al. (2004), Montalvo and Reynal-Querol (2005) 1 See the works of Brockett (1992), Midlarski (1988), Muller and Seligson (1987), Muller et al. (1989), and Nagel (1974), among others. 2 See Esteban and Ray (2008) on the salience of ethnicity over class in social conflict. 3 See Desmet et al. (2009,2010) on Greenberg’s index. 4 See Esteban and Ray (1994) –also Wolfson (1994)– for the earliest polarization measure, ER, and Reynal-Querol (2002) for a special case of the Esteban-Ray measure, RQ. Duclos et al. (2004) extend Esteban and Ray’s measure for continuous distributions. See also the special issue of the Journal of Peace Research edited by Esteban and Schneider (2008) devoted to the links between polarization and conflict. 1 and Reynal-Querol (2002) are representative of this literature. Most of the empirical work has been based on the indices of fractionalization Fand polarization RQ. Both share the feature that are based on group sizes only and do not make use of variations in inter-group distances. Fearon (2003) already made the point that ethno-linguistic distances do play a key role in explaining ethnic conflict and computed a measure based on dissimilarity between pairs of languages. Montalvo and Reynal-Querol (2005) (MRQ hereinafter) dismissed the use of such distances arguing that, using Fearon’s data, the correlation between G(that uses distances) and F(that doesn’t) is 0.82. However, this claim has been recently challenged by Desmet et al (2009,2010). They re-examine the point made in various papers by Alesina concerning the lower level of social transfers in ethnically heterogeneous societies and find that the measures that include distances outperform the ones that don’t. Specifically, they obtain that Gis highly significant while Fisn’t and that the same is true for P relative to RQ. 5 Our paper contributes to the literature on two counts. First, our paper studies the link between conflict and ethnic and religious polarization using interpersonal distances driven by the intensity of the ethnic and religious attitudes obtained from the World Values Survey. We derive the intensity of feelings by aggregating the answers to a set of questions related to religious or ethnic attitudes. This permits us to compute a polarization measure that depend on inter-personal distances, such as P, and test whether it performs better than the ones that do not use this information, such as Fand RQ. Second, together with the standard practice of dichotomizing the occurrence of con5 A second relevant feature of this literature is that it has been geared towards finding empirical regularities rather than testing the implications of a specific model of conflict. In Esteban, Mayoral and Ray (2010) we test the empirical implications of the conflict model set up in Esteban and Ray (2010), as a first step towards an explicit link between theory and facts. 2 flict –war/peace– depending on the number of deaths exceeding a given threshold, we also use a second, continuous indicator of social unrest based on political assassinations, demonstrations, strikes, political prisoners, etc. This permits us to overcome two major empirical problems in the literature that we will discuss in Section 4: (i) the results may depend on the choice of the threshold level; and (ii) the dilemma between onset versus incidence as definitions of conflict. Our empirical exercise directly compares the performance of ethnic and religious polarization as measured by RQ and P, using the same controls as in MRQ. In all our estimations we use the continuous index of intensity of conflict as well as the classic binary measure based on a threshold level of casualties. We also check for potential endogeneity in case the intensity of attitudes is the consequence rather than the cause of conflict. We finally perform a series of robustness tests. Our results strongly support the hypothesis that intensity of feelings is highly significant, most especially for ethnic attitudes. More importantly, the distance sensitive indices of religious and ethnic polarization, P, are both independently significant. When we simultaneously consider the two of them, both are significant when we use the discrete measure of conflict, but religious polarization ceases to be significant with the continuous indicator of conflict. In all cases, the use of intensity of attitudes significantly increases the explanatory power of the model, as we obtain levels of R 2 that are much higher than usual in this literature. The paper is organized as follows. In the next section we summarize the main features of the various distributional indices that have been used in the literature. Section discusses our approach to the measurement of inter-personal distances, key to our exercise. Section describes in detail the data used in both the main exercise and the different robustness tests performed. Section presents our empirical results and Section concludes. 3 2. DISTRIBUTIONAL INDICES There are a variety of indicators capturing different features of a distribution. We have already mentioned four different indices that have been used in empirical work: F,G,RQ and P. 6 What are the appropriate indicators to be used if we want to predict conflict? This apparent heterogeneity of measures can be presented as different ways of specifying the measurement of interpersonal distances and the weight of group size. On the first dimension, the specification of distances defines two classes of measures. One class retains the measured inter-personal distances while a second class considers all other groups equally distant –this common distance is normalized to unity. G and Pbelong to the former class and Fand RQ to the latter. The second dimension refers to the treatment of group sizes. We have here too two classes of measures. One class does not take into account the effect of group size on the sense of identity. Fand Gbelong to this class. The second type assigns “returns” to group size and transforms the own group size to a power. This is the feature that makes polarization measures Pand RQ distinctly different from inequality measures. The identity/alienation approach to social antagonisms introduced by Esteban and Ray (1994) may help to establish a taxonomy over the variety of distributional indices that have been used so far. Accordingly with this approach, interpersonal antagonism is the conjoint result of the sense of identification with one’s own group and the alienation felt towards members of other groups. The sense of identity depends on the size of one’s group, n i , and the feeling of alienation on the perceived distance between groups iand j,d ij . More formally, Esteban and Ray (1994) and Duclos, Esteban and Ray (2004) start 6 Collier and Hoeffler (2004) have also used the ratio of the largest over the second group and Desmet et al (2009,2010) introduce the index of peripheral heterogeneity. 4 by defining the general class of measures of societal antagonism, A. The antagonism felt by a member of group ivis-a-vis a member of group j,a(i, j)can be expressed as a(i, j) = φ(n i , d ij ),(1) with φ(n i , d ii ) = 0. Total societal antagonism is defined as the sum of all interpersonal antagonisms: A= i  j n i n j φ(n i , d ij ).(2) Esteban and Ray (1994) embody the concept of polarization of a distribution in a set of five axioms. From these axioms, they uniquely derive the measure of polarization P= i  j n 1+α i n j d ij ,(3) with 1≤α1.6. 7 Note that Pis a specific measure of societal antagonism. The aforementioned axioms imply that interpersonal antagonism is of the form: φ(n i , d ij ) = n α i d ij . Hence, polarization captures both components of interpersonal antagonism: identity and alienation. Suppose now that we simply posit that interpersonal antagonism does not depend on the sense of identity. Then φ(n i , d ij ) = d ij . We obtain the Gini-Greenberg index G= i  j n i n j d ij .(4) 7 Esteban and Ray (2010) introduce an axiom that combined with the axioms for continuous distributions in Duclos et al (2004) pins down α= 1. 5 If in addition we posit that a shift in alienation, ˜ d ij =d ij +δfor all i=j, does not modify interpersonal antagonism, then φ(n i , d ij ) = kfor all i=j. Normalizing k= 1 we obtain the Hirschman-Herfindahl fractionalization index G= i  j=i n i n j = i n i (1 −n i ).(5) Finally, if we make the previous assumption concerning alienation, but continue to retain the role of identification, we obtain that φ(n i , d ij ) = n α i . This gives us the measure proposed by Reynal-Querol (2002): RQ = i  j=i n 1+α i n j = i n 1+α i (1 −n i ). 8 (6) Therefore, the result by MRQ that RQ is significant in explaining conflict while Fis not, can be interpreted as indicating that group concern –hence the group-size effect– is important. Our paper can be seen as a test that not only group size but also alienation both matter for social conflict. To this effect, we shall show that Phas a significantly higher explanatory power for conflict than any of the other distributional measures. Because interpersonal distances d ij do play a significant complementary role we shall show that Poutperforms RQ as an independent explanatory variable for conflict. 8 Note that this measure is conceptually closely related to F. Indeed, while Ftells us the probability that two persons drawn randomly belong to different group, RQ with α= 1 –the empirically relevant variant– tells us the probability that out of three people two belong to the same group and the third to any other group. 6 3. IDENTITY AND ALIENATION IN ETHNIC AND RELIGIOUS POLARIZATION One of the distinct features of our exercise is the use of polarization indices that depend on both, group size and inter-personal distance. As mentioned before, the previous work by MRQ, while emphasizing the role of group sizes, disregarded interpersonal distances. Indeed, the measure RQ only depends on the population size of the different ethnic/religious groups. The classification of the population into ethnic and/or religious groups is not as straightforward as it first appears. It poses the problem of the definition of the groups. Even when the ascription of individuals to groups is unequivocal, there remains the issue of what are the relevant groups. To illustrate the point, let us take the line identifying ethnic group with language. Ethnologue records 6,912 different languages worldwide. This gives an average of thirty five ethnolinguistic groups per country [for the 195 countries existing today]. In India Ethnologue identifies 415 languages. 9 It is plain that one needs to “aggregate" over these micro-groups and focus on broader definitions, merging various “similar" ethnic groups. The same argument can be made of religions. 10 And this takes us to a main problem we wish to underscore. This is the issue of the inter-personal or inter-group distances. The way the grouping is performed in the literature implicitly assumes that inter-group distances are either zero –when the groups are merged into one– or unity –when they are considered alien to each other. This problem can be bypassed by using inter-group distances. The distances among ethnolinguistic groups have been computed by Fearon (2003) and by Desmet et al. (2009,2010) on the basis of different measures of linguistic similarity. The use of these inter-group measures permits to go beyond the dychotomic [0,1] distance and 9 Although the 1991 census recognizes 1,576 “mother tongues". 10 For instance,MRQ treat “christians" as a single group. 7 4.3. Additional Independent Variables While the dependent variable is the incidence of conflict in each period, the control variables refer to the first year of each period or are by their nature invariant in time. The number of countries varies slightly with the exercise and is indicated together with each empirical result. The control variables are as follows. Sociopolitical variables: size of the population, level of democracy. Economic variables: real GDP per capita, LGDP C, share of primary exports on GDP, PRIMEXP , dependence on oil exports, OIL. Geographic variables: percentage of mountainous terrain, MOUNT , noncontinguency of country territory, NONCONT and regional dummies for Latin America, D−AMER, Asia, D−ASIA, and Sub-Saharan Africa, D−AFRICA. While the justification for each control variable can be found elsewhere (Fearon and Laitin, 2003, Collier and Hoeffler, 2004, 2008, Miguel et al., 2004, MRQ) the details are provided in Table B1 included in Appendix B. The data on religious and ethnic composition of countries used for calculation of Fand RQ measures in the main text have been taken from MRQ. 4.4. Sample Size The availability of data from different sources severely conditions the size of our sample. Let us start by the WVS. It has been conducted only for limited number of countries. The sample of countries has expanded from 20 countries participating in the first wave in 1981 to a total of 97 countries being surveyed in the latest fifth wave. Moreover, our analysis is restricted to those countries where a suitable set of questions related to religious or ethnic tolerance were asked and where religious or ethnic identification of each respondent was available . A main point of our paper is that the distance-based Ppolarization index outperforms the RQ index with no variation in inter-personal distances. The WVS sample 14 does not coincide with the sample used by MRQ. The direct comparability of our results with those obtained by MRQ limits even further the number of countries in our sample. While their sample includes 117 countries for the period 1960-1999, the overlap with the WVS yields 51 and 61 countries for ethnic and religious data, respectively. When combined with the number of five-year periods the sample contains 385 and 464 observations. respectively. This subset of countries presents features of conflict similar to the larger set. Using the PRIOCW variable, Montalvo and ReynalQuerol’s (2005) dataset features 159 periods with ongoing civil war, while we have 83 periods for the religious subset of countries and 75 for the ethnic one. If we compute the mean incidence of war per period (the number of war periods divided by the total sample) we obtain very similar results. For the ethnic sample we have a mean incidence of 0.145 and 0.157 for the religious sample, while Montalvo and Reynal-Querol record a mean incidence of 0.168. When we use ISC as the dependent variable we restrict to the same set of countries. Unfortunately, the information on the ISC index is missing for some periods. This reduces the total number of observations to 344 and to 406 for the ethnic and religious data. When performing the robustness checks we shall be using different and larger samples, including more countries and more periods. The specifics are described in detail when needed. 5.EMPIRICAL RESULTS: ETHNIC AND RELIGIOUS POLARIZATION AND CONFLICT 5.1. Polarization and Conflict We use the obtained individual intensity of feelings to compute the Pindex of polarization. We wish to test whether Poutperforms RQ in explaining conflict. We 15 consider the same set of independent variables as in MRQ and two dependent ones: the binary variable of incidence of civil war, PRIOCW, and the continuous measure that captures the intensity of social conflict, ISC. Results are reported in Tables 1 and 2, respectively. Columns 1-3 refer to ethnic polarization, columns 4-6 refer to religious polarization. For each case, the first column reproduces the results when only the RQ index is included in the reduced sample (columns (1) and (4)), the second column uses Pinstead and the third column uses both, RQ and P, permitting a direct contrast between the measures without and with intensity of attitudes. Columns (7) and (8) introduce both religious and ethnic P and the four indices considered in these exercise, respectively. 16 T 1. P  C: P RQ Variable (1) (2) (3) (4) (5) (6) (7) (8) LGDPC −0.457 (0.084) −0.564 (0.050) −0.563 (0.048) −0.290 (0.374) −0.162 (0.523) −0.143 (0.616) −0.450 (0.162) −0.576 (0.083) LPOP 0.124 (0.576) 0.095 (0.650) 0.095 (0.658) 0.164 (0.448) 0.160 (0.457) 0.153 (0.506) 0.045 (0.846) 0.073 (0.738) PRIMEXP −0.632 (0.776) −1.129 (0.595) −1.116 (0.601) −0.600 (0.796) −1.705 (0.564) −1.766 (0.569) −2.848 (0.301) −2.443 (0.373) MOUNT 0.006 (0.700) 0.015 (0.290) 0.015 (0.363) 0.013 (0.225) 0.020 (0.110) 0.020 (0.099) 0.032 (0.021) 0.033 (0.061) NONCONT 0.282 (0.696) 0.613 (0.434) 0.609 (0.452) 0.176 (0.784) 0.449 (0.489) 0.457 (0.475) 0.159 (0.154) 1.178 (0.134) DEM −0.148 (0.768) −0.157 (0.723) −0.157 (0.723) 0.027 (0.954) 0.160 (0.724) 0.157 (0.725) −0.319 (0.463) −0.283 (0.548) RQ Eth 0.737 (0.499) -−0.040 (0.975) - - - - −0.028 (0.983) RQ Rel - - - 1.071 (0.246) -0.164 (0.860) -−1.108 (0.311) P Eth -2.849 (0.013) 2.853 (0.014) - - - 3.920 (0.002) 4.292 (0.002) P Rel - - - - 2.343 (0.009) 2.2858 (0.015) 2.059 (0.038) 2.550 (0.023) Pseudo-R 2 0.066 0.177 0.177 0.084 0.139 0.139 0.253 0.257 Countries 51 51 51 61 61 61 50 50 Observations 385 385 385 464 464 464 377 377 Notes : The dependent variable is PRIOCW and the estimation method is logit. P-values are reported in brackets. Robust standard errors adjusted for clustering have been employed to compute z-statistics. Table 1 shows that while RQ is not significant in our reduced sample of countries (columns (1) and (4)), both ethnic and religious Pare so (columns (2) and (5)). When both P and RQ are included in the regression (columns (3) and (6)), the significance of Premains and the estimated coefficients are very similar as when only Pis considered. These conclusions do not change if both ethnic and religious P are included in the regression (column 7) or when the four indices considered in this 17 exercise are simultaneously introduced (column 8). Furthermore, the introduction of Pincreases the pseudo-R 2 substantially. When both religious and ethnic P are considered, the pseudo-R 2 reaches 0.25. We now replicate the same exercise reported in Table 1 on the performance of P relative to RQ using as dependent variable the continuous indicator of conflict, ISC. The results are displayed in Table 2. This table is organized exactly as Table 1. T 2. P  C: P  RQ  ISC Variable (1) (2) (3) (4) (5) (6) (7) (8) LGDPC −0.483 (0.000) −0.502 (0.000) −0.498 (0.000) −0.501 (0.000) −0.431 (0.000) −0.446 (0.000) −0.481 (0.000) −0.567 (0.000) LPOP 0.139 (0.078) 0.118 (0.120) 0.129 (0.092) 0.191 (0.016) 0.174 (0.019) 0.178 (0.016) 0.109 (0.154) 0.122 (0.104) PRIMEXP −1.144 (0.112) −1.154 (0.097) −1.169 (0.096) −0.986 (0.202) −1.286 (0.091) −1.257 (0.104) −1.341 (0.053) −1.274 (0.062) MOUNT 0.001 (0.874) 0.006 (0.285) 0.004 (0.553) 0.003 (0.521) 0.005 (0.305) 0.005 (0.310) 0.008 (0.182) 0.006 (0.342) NONCONT 0.679 (0.034) 0.698 (0.023) 0.751 (0.021) 0.553 (0.028) 0.637 (0.011) 0.634 (0.011) 0.747 (0.015) 0.829 (0.010) DEM 0.079 (0.694) 0.118 (0.524) 0.106 (0.558) 0.054 (0.793) 0.086 (0.646) 0.092 (0.623) 0.150 (0.441) 0.179 (0.372) RQ Eth 0.680 (0.079) -0.475 (0.198) - - - - 0.534 (0.153) RQ Rel - - - 0.186 (0.631) -−0.121 (0.735) -−0.559 (0.118) P Eth -0.921 (0.000) 0.867 (0.000) - - - 0.878 (0.002) 0.903 (0.002) P Rel - - - - 0.793 (0.026) 0.843 (0.032) 0.337 (0.189) 0.512 (0.075) Pseudo-R 2 0.237 0.286 0.291 0.268 0.281 0.282 0.289 0.304 Countries 51 51 51 61 61 61 50 50 Observations 344 344 344 406 406 406 336 336 Notes : The dependent variable is ISC and the estimation method is pooled OLS. P-values are reported in brackets. Robust standard errors adjusted for clustering have been employed to compute z-statistics. The qualitative conclusions of this table are in line with the ones obtained with 18 the discrete variable for conflict. Results in columns (1) and (4) are very similar as those in the original paper by MRQ: the ethnic dimension RQ is significant (at the 10 percent level) while the religious one is not. Ethnic P turns out to be significant in all the specifications considered and, when both ethnic P and RQ are introduced, ethnic RQ ceases to be so (column (3)). On the religious dimension we obtain that while RQ is never significant, religious P is so when introduced alone or together with religious RQ (columns (5) and (6), respectively). However, considering both the ethnic and the religious P indices in the regression limits the significance of the latter (columns (7) and (8)). Finally, the explanatory power of the model increases substantially when the P indices are considered, as shown by the pseudo R 2 that reaches 0.29. 5.2. Endogeneity of Attitudes: IV regressions We have already mentioned that our previous exercise could be objected on the basis of endogeneity. Diversity measures that do not incorporate interpersonal distances, such as RQ and F, have usually been considered as exogeneous in cross-country regressions, the reason being that group shares are thought to be very stable over time and small changes only have a minor impact on these measures. Ethnic RQ or F could to some extent be endogeneous (conflict can alter group shares or the definitions of ethnic groups can change through time as a function of economic-political variables) but, as pointed out by Alesina et al. (2003), ethnic compositions display tremendous time persistence and thus, the exogeneity assumption could be reasonable at the 2030 year horizon that characterizes cross country regressions in this area. Religious indices, however, may be more problematic. In some repressive regimes, non-official religions might be prosecuted, making it difficult for members of these religions to be counted as such. This may create a spurious correlation between lack of political freedom and religious diversity that could bias the estimates. Nevertheless, there is little doubt that introducing people’s attitudes in our po19 larization measures makes the endogeneity problem more acute. Civil conflict will probably have an immediate impact on people’s attitudes towards the rival groups, making them more intolerant and polarized and, thus, reverse causality cannot be discarded. In the following, we try to overcome this problem by instrumenting the potentially endogenous regressors. Since, for the reasons provided above, religious F or RQ indices could also be endogenous, we focus on the ethnic dimension and instrument ethnic Pconsidering ethnic RQ measures as exogeneous. We instrument ethnic P using averages of the distance between the languages spoken in a country. Language distances are a proxy of the cultural differences among the groups living in a territory. The identification assumption that we adopt is that language distances do not affect conflict directly but only through its correlation with people’s tolerance or intolerance of other groups. We believe that what matters for conflict is not the “objective" cultural differences but the way these differences are perceived and liked or disliked by the different groups, dimension that we are able to capture through our polarization indices. If this assumption holds, variations in P induced by averages of language distances can be considered as exogenous and employed to evaluate the effect of an exogenous change in Pon the level of conflict. In addition, language-based indices are very stable over time and do not present the reverse causality problem that potentially affects ethnic P. There are different ways of measuring distances between languages. Fearon and Laitin (1999, 2000) and Laitin (2000) proposed to use the information provided by language trees. Language trees are genealogical diagrams of languages related by descent of a common ancestor. The distance between two languages iand jis computed as a function of the number of common classifications in the language tree. For instance, Spanish and Basque diverge at the first branch, since they come from structurally unrelated language families. By contrast, Spanish and Catalan share their first 20 7 classifications as Indo-European, Italic, Romance, Italo-Western, Western, GalloIberian and Ibero-Romance languages. We follow Fearon (2003) and Desmet et al. (2009) and define the distance between languages iand jto be d ′ ij = 1 −l m δ , where lis the number of common branches between iand j,mis the maximum number of shared branches between any two languages, and δis a parameter that determines how fast the distance declines as the number of shared languages increases. The weighted average of the distances between any two pair of languages spoken in a country is given by WAD = K  i=1 K  j=1 s i s j d ′ ij , where s h , h={i,j} denotes the share of people that speaks language has a first language and Kis the total number of languages spoken in a particular territory. We will use WAD as instrument for ethnic P. 16 Data on WAD has been taken from Desmet et al. (2009), who have elaborated this variable using the information on language trees provided by the Ethnologue project. 17 The distinctive characteristic of Ethnologue versus other sources is its detail. It provides very disaggregate information of all the languages and dialects spoken in a territory. For instance, as noted by Desmet et al. (2009), the Britannica Book of the year 1990 edition reports 21 living languages for Mexico while Ethnologue lists 291. This means that the shares used to compute WAD are in general very different from the shares employed to compute P, since the WVS classification only considers a small number of categories. Still, the correlation between ethnic P and 16 We have also considered other indices of linguistic distance, including the index of polarization for discrete distributions P, as in (3). We ended up using the Gini-Greenberg index type WAD because it gave the highest correlation with the potentially endogenous variable. 17 Desmet et al. (2009) choose a value of δ= 0.05 to compute d ij. 21 WAD is 0.39. Table 3 presents the results of instrumenting ethnic P by WAD. As pointed out by Angrist and Krueger (2001), estimating a linear probability model by two stage least squares (2SLS) is a robust estimation approach even if the underlying second-stage relationship is nonlinear, as in our case. The linear projection of ethnic P on the rest of the controls and WAD shows that the latter has a positive and highly significant effect on ethnic P, with a p-value of 0.019 and of 0.016, corresponding to the cases where ethnic RQ was included or not in the first-stage regression. Columns (1) to (4) present the results of instrumenting ethnic P in the discrete dependent variable regression while columns (5) and (6) focus on the continuous indicator of conflict. Figures in columns (1), (2), (5) and (6) have been obtained using 2SLS while those in columns (3) and (4) with MLE in a probit specification. The qualitative results do not change with respect to those reported in Tables 1 and 2: P Eth is highly significant while RQ Eth is not. Including or not RQ Eth as an additional regressor does not have any impact on either the estimated coefficient of P Eth or on its significance. Using the rule of thumb suggested by Wooldrigde (2002), it is possible to compare the magnitudes of the linear probability and the probit specifications. To do that, we should divide the probit estimates by 2.5, which turn out to be 1.46 and 1.44 for columns (3) and (4), respectively, and thus, the partial effects implied by the probit specification are similar to those of the linear probability model reported in columns (1) and (2). Using the latter figures to evaluate the partial effect of ethnic polarization on the probability of incidence of civil conflict, we obtain that an increase in 1 standard deviation of P Eth raises the probability of civil war by 0.34. 18 18 Of course, this implication cannot literally be true because continually increasing Pwould eventually drive the probability of conflict to be greater than one. However, these figures can be 22 T 3. IV E! Variable (1) (2) (3) (4) (5) (6) LGDPC −0.098 (0.029) −0.095 (0.035) −0.266 (0.107) −0.261 (0.119) −0.514 (0.000) −0.516 (0.000) LPOP 0.001 (0.982) 0.007 (0.814) −0.003 (0.976) 0.011 (0.910) 0.119 (0.126) 0.113 (0.148) PRIMEXP −0.291 (0.204) −0.299 (0.223) −1.342 (0.175) −1.363 (0.173) −1.195 (0.103) −1.187 (0.105) MOUNT 0.005 (0.119) 0.004 (0.144) 0.014 (0.088) 0.0121 (0.102) 0.006 (0.373) 0.008 (0.214) NONCONT 0.120 (0.363) 0.136 (0.294) 0.522 (0.177) 0.565 (0.130) 0.823 (0.009) 0.797 (0.009) DEM 0.013 (0.864) 0.010 (0.891) −0.067 (0.784) −0.075 (0.757) 0.134 (0.485) 0.141 (0.468) RQ Eth −0.217 (0.326) -−0.431 (0.512) -0.267 (0.541) - P Eth 1.241 (0.008) 1.190 (0.007) 3.660 (0.000) 3.614 (0.000) 1.749 (0.025) 1.808 (0.015) Pseudo-R 2 - - - - 0.254 0.247 Countries 51 51 51 51 51 51 Observations 385 385 385 385 344 344 Notes : The dependent variable PRIOCW (1)-(4) or ISC (5),(6). The estimation method is 2SLS (1),(2),(5),(6) or IV Probit (3),(4). IV is WAD. 5.3. Robustness of Results The previous section has shown that including interpersonal distances in polarization indices is relevant in explaining conflict and that while RQ is not significant, P is so in both the standard and the IV regressions. In this section we explore the robustness of the previous results by considering alternative: a) sample, b) instruments, c) definitions of civil conflict, and d) inclusion of regional dummies. For the sake of briefness we shall concentrate on the estimates for ethnic polarization. good estimates of the partial effects of Pnear the center of the distribution of the covariates. 23 D, J-Y., J. E,  D. R' (2004), “Polarization: Concepts, Measurement, Estimation” Econometrica 72, 1737—1772. E' W.  R. L (1997), ÒAfricaÕs Growth Tragedy: Policies and Ethnic DivisionsÓ, Quarterly Journal of Economics 111, 1203-1250. E, J., L. M'  D. R' (2010), “Ethnicity and Conflict: An Empirical StudyÓ, unpublished manuscript. E, J.  D. R' (1994), “On the Measurement of Polarization” Econometrica 62, 819-852. E, J.  D. R' (2008), “On the Salience of Ethnic Conflict," American Economic Review 98, 2185-2202. E, J.  D. R' (2010), “Linking Conflict to Inequality and Polarization," American Economic Review , forthcoming. E, J.  G. S(2008) ÒPolarization and Conflict: Theoretical and Empirical IssuesÓ (Introduction to the special issue), Journal of Peace Research 45, 131-141. F, J.(2003), “Ethnic and Cultural Diversity by Country", Journal of Economic Growth 8, 195-222. F, J.,  D. L (1999) “Weak States, Rough Terrain, and Large-Scale Ethnic Violence since 1945”, Presented at the Annual Meetings of the American Political Science Association, Atlanta, GA. F, J.,  D. L (2000) “Violence and the Social Construction of Ethnic Identity,” International Organization 54, 845-877. F, J.  D. L (2003), “Ethnicity, Insurgency, and Civil War," American Political Science Review 97, 75—90. 30 H, D.L. (1985), Ethnic Groups in Conflict, Berkeley, CA: University of California Press. L, D. (2000) “What is a Language Community?,” American Journal of Political Science 44,142-155. L, M. I. (1989), “An Evaluation of "Does Economic Inequality Breed Political Conflict?", World Politics 4, 431470. M, A. (2008) “Historical Statistics of the World Economy: 1-2008.AD", http://www.ggdc.net/maddison/. M-, M.I. (1988), “Rulers and the ruled: patterned inequality and the onset of mass political violence,” American Political Science Review 82, 491—509. M*, E., S', S.  E. S* (2004), “Economic Shocks and Civil Conflict: An Instrumental Variables Approach," Journal of Political Economy 112, 725—753. M, J. G.  M. R'-Q (2005), “Ethnic Polarization, Potential Conflict and Civil War," American Economic Review 95, 796—816. M,E.N.  M.A. S* (1987) “Inequality and Insurgency", American Political Science Review 81, 425-451. M, E. N., M.A. S*  H. F (1989), “Land inequality and political violence", American Political Science Review 83, 577—586. N*, J. (1974), “Inequality and discontent: a non-linear hypothesis,” World Politics 26, 453—472. R'-Q, M. (2002), ”Ethnicity, Political Systems, and Civil Wars” Journal of Conflict Resolution 46, 29-54. S-, M. R., F. W. W'!  J. D. S* (2003) “Inter-State, IntraState, and Extra-State Wars: A Comprehensive Look at Their Distribution over Time, 1816-1997," International Studies Quarterly 47, 49-70. 31 S, G. N. W! (2006), “Ethnic Polarization, Potential Conflict, and Civil Wars: Comment", unpublished manuscript, University of Konstanz. W, M. C. (1994), “When Inequalities Diverge", American Economic Review Papers and Proceedings 84 353-358. W*, J. M. (2002), Econometric Analysis of Cross Section and Panel Data. The MIT Press. 32 Appendix A : The Construction of the Aggregate Indices of Attitudes from the World Value Surveys The World Value Surveys (WVS) provide since 1980´s an unique international data on the changing values of the population of different countries, including the opinions on the issues related to religious and ethnic tolerance. We have used the rich WVS data to try to track the inter-group distances that are needed to compute the multi-group DER. The WVS are conduced by means of face-to-face interviews and includes hundrets of questions asked in each participating country. So far, there have been conducted 5 waves of the WVS with participation of 97 countries (at least in one wave). The country sample varies but usually amounts to 1000 respondent and is usually representative with respect to the age and sex structure of each country and often also stratified geographically. We wish mention here the problem of the sampling rules followed by the WVS. Specifically, that neither religion nor ethnicity have been used to obtain a balanced sample. Indeed, the weight of the different religious groups and of ethnicities varies significantly from wave to wave, producing an artificial time variability of all the distributional indices. We have opted for taking the average over all the waves available for each country. The list of questions varies over the waves, and hence we have selected questions that have been asked in all the waves or questions that were very similar over different waves. In case of religious tolerance we use 7 and in case of ethnic 6 questions listed below. The questions have either binary or multiple response with at most 10 possible answers (typically when respondent is asked to mark her opinion between two opposite statements or her degree of aggrement with certain statement). We take the raw survey data and recode all the answers so that 1 is the most tolerant one and 10 stands for the most intolerant one. If less than 10 answers are provided, we use the mean values of the intervals, which number corresponds to the number of possible responses. For instance, for a question "Do you believe in God" we assign value of 3.25 to no and 7.75 to "yes" given that yes can mean somethnig between 33 "totally believe" rather belive than do not" and vise versa for "no". Consequently, we use a Principal Component Analysis (PCA) to collapse the response to numerous questions into a single index number of each respondent (i.e. we retain the first principal component). Comparing density functions of different countries can provide some very basic evidence on the differences between religious and ethnic feelings between different countries. However, the PCA provides by contruction a distributions with the same zero mean that is not bounded. For the purpose of our analysis it is clearly preferable to work with densities that are bounded on the same interval for all the countries while the mean of discribution differ. Therefore, we retain the weights of the first principal component and apply them to the original (not demeaned) values of each variable and bound the index on the same interval of the original responsed (1,10). To see this, note that regular PCA is used to find a vector of weights that maximize the variance of a index zwhere x ij is the answer of person ito question j. z i = j α j (x ij −x j ) If we apply the vector of weights on the original variables x, the distribution of the resulting index y does not have zero mean anymore but its scale varies across countries given the sum of the weighs varies (the sum of the squares of the weights is one): y i = j α j x ij However, we can the minimum and the maximum value of the distribution of each country so as to bound the index distribution on (1,10) interval: y ∗ i =10 (y i −y min ) (y max −y min ) 34 The above distribution bounded on (1, 10) can be simply compared for different countries. However, also for different groups within the same country. In the latter case, we use a variable that idenfies the religious group or ethnicity of respondent. Please note, that identification of the group of the respondent is necesary so as to calculate the DER multigroup polarization index. This identification is available for 81 countries in the case of religion and 67 countries in case of ethnicity. The questions used to obtain the aggregate indices are the following: 1. R* ! •V22 Religious faith: do you consider it to be especially important? •V183. Here are two statements which people sometimes make when discussing good and evil. Which one comes closest to your own point of view? A) There are absolutely clear guidelines about what is good and evil. These always apply to everyone, whatever the circumstances. B) There can never be absolutely clear guidelines about what is good and evil. What is good and evil depends entirely upon the circumstances at the time. •V185. Apart from weddings, funerals and christenings, about how often do you attend religious services these days? •V186. Independently of whether you go to church or not, would you say you are: 1) a religious person; 2) Not a religious person; or 3) a convinced atheist •Which, if any, of the following do you believe in? —V191 Do you believe in God? —V192 Do you believe in life after death? 35 —V194 Do you believe in hell? —V195 Do you believe in heaven? •V196. How important is God in your life? Please use this scale to indicate10 means very important and 1 means not at all important. •V197. Do you find that you get comfort and strength from religion? Given the fact that questions V191, V192, V194, V195 are quite similar, we have computed its average value as a single indicator. 2. E ! •On this list are various groups of people. Could you please sort out any that you would not like to have as neighbors? —V69 People of a different race —V73 Immigrants/foreign workers •V25 Generally speaking, would you say that most people can be trusted or that need to be very careful in dealing with people? •V214 To which of these geographical groups would you say you belong first of all? •V215 And the next? —Locality or town where you live —State of region of country where you live Country as a whole Continent 36 —The world as a whole •V216 How proud are you to be [nationality]? Appendix B: Description of Data C    For both conflict onset and incidence we use the armed conflict dataset from the Upsala Conflict Data Program and the Peace Research Institute Oslo, UCDP/PRIO. This data set covers from 1960 till 2008, subject to data availability. However, to keep consistency with MRQ data tables 1 to 3 restrict to the period 1960-1999 only. The sample is divided into five-year periods. To record whether a country is in conflict shall take the definition of intermediate conflict. Accordingly with PRIO2004 standards the definition for Intermediate conflict is “ PRIOCW : more than 25 battle-related deaths per year and a total conflict history of more than 1000 battle-related deaths, but fewer than 1000 per year." In our robustness checks we shall also work with the notions of “minor conflict" and “war". The corresponding definitions are as follows. “ PRIO25 : between 25 and 999 battle-related deaths in a given year.",“ PRIO1000 : at least 1,000 battle-related deaths in a given year." C I8 This variable is the conflict index computed by The Cross-National Time-Series Data Archive (CNTS). This index of conflict is the weighted average of eight different manifestations of domestic conflict, adopted from Rudolph J. Rummel, "Dimensions of Conflict Behavior Within and Between Nations", General Systems Yearbook, VIII [1963], 1-50). 20 The eight variables included are: 20 The correlation with UCDP/PRIO is 0.45.W 9 PRIO 8? 37 •Assassinations (domestic1): Any politically motivated murder or attempted murder of a high government official or politician. •General Strikes (domestic2): Any strike of 1,000 or more industrial or service workers that involves more than one employer and that is aimed at national government policies or authority. •Guerrilla Warfare (domestic3): Any armed activity, sabotage, or bombings carried on by independent bands of citizens or irregular forces and aimed at the overthrow of the present regime. •Major Government Crises (domestic4): Any rapidly developing situation that threatens to bring the downfall of the present regime - excluding situations of revolt aimed at such overthrow. •Purges (domestic5): Any systematic elimination by jailing or execution of political opposition within the ranks of the regime or the opposition. •Riots (domestic6): Any violent demonstration or clash of more than 100 citizens involving the use of physical force. •Revolutions (domestic7): Any illegal or forced change in the top government elite, any attempt at such a change, or any successful or unsuccessful armed rebellion whose aim is independence from the central government. •Anti-government Demonstrations (domestic8): Any peaceful public gathering of at least 100 people for the primary purpose of displaying or voicing their opposition to government policies or authority, excluding demonstrations of a distinctly anti-foreign nature. The weights used are: Assassinations (25), Strikes (20), Guerrilla Warfare (100), Government Crises (20), Purges (20), Riots (25), Revolutions (150), and Anti-Government 38 Demonstrations (10). The calculation is performed as follows: weighted sum of occurrences of each event divided by 8 (the number of types of events) and multiplied by 100. I9  We summarize all the variables used in our empirical exercises in the following Table. 39