scieee AI-readable full text Open interactive document viewer

Migration Policies and Immigrants’ Language Acquisition in EU‐15: Evidence from Twitter

Gil‐Clavel, Sofia,Grow, André,Bijlsma, Maarten J.

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Gil‐Clavel, Sofia; Grow, André; Bijlsma, Maarten J. Article — Published Version Migration Policies and Immigrants’ Language Acquisition in EU‐15: Evidence from Twitter Population and Development Review Provided in Cooperation with: John Wiley & Sons Suggested Citation: Gil‐Clavel, Sofia; Grow, André; Bijlsma, Maarten J. (2023) : Migration Policies and Immigrants’ Language Acquisition in EU‐15: Evidence from Twitter, Population and Development Review, ISSN 1728-4457, Wiley, Hoboken, NJ, Vol. 49, Iss. 3, pp. 469-497, https://doi.org/10.1111/padr.12574 This Version is available at: https://hdl.handle.net/10419/288067 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/ Migration Policies and Immigrants’ Language Acquisition in EU-15: Evidence from Twitter SOFIA GIL-CLAVEL ,ANDRÉ GROW AND MAARTEN J. BIJLSMA In response to the increasingly complex and heterogeneous immigrant communities settling in Europe, European countries have adopted various civic integration measures. Measures aiming to facilitate language acquisition are considered crucial for integration and cooperation between immigrants and natives. Simultaneously, the rapid expansion of social media usage is believed to change the factors affecting immigrants’ language acquisition. However, only a few previous studies have analyzed whether this is the case. This article uses a novel longitudinal data source derived from Twitter to (1) analyze differences in the pace of immigrants’ language acquisition depending on the migration policies of destination countries and (2) study how the relative sizes of the migrant groups in destination countries, and the linguistic and geographical distances between origin and destination countries, are associated with language acquisition. Results show that immigrants who live in countries with strict language acquisition requirements for immigrants and conservative citizenship policies have the highest median times until language acquisition. Based on Twitter data, we also find that language acquisition is associated with classic explanatory variables, such as the size of the immigrant group in the destination country and the linguistic and geographical distance between origin and destination country similar to the previous studies. Introduction Since the beginning of the 21st century, policymakers across Europe have attempted to enforce the requirement that immigrants learn the Sofia Gil-Clavel, Delft University of Technology, Max Planck Institute for Demographic Research, University of Groningen. E-mail: [email protected]. André Grow, Max Planck Institute for Demographic Research. E-mail: [email protected]. Maarten J. Bijlsma, University of Groningen, Max Planck Institute for Demographic Research. E-mail: [email protected]. POPULATION AND DEVELOPMENT REVIEW 49(3): 469–497 (SEPTEMBER 2023) 469 © 2023 The Authors. Population and Development Review published by Wiley Periodicals LLC on behalf of Population Council. This is an open access article under the terms of the Creative Commons Attribution License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited. 470 MIGRATION POLICIES AND IMMIGRANTS’LANGUAGE ACQUISITION IN EU-15 national language through civic integration policies (Wright and Viggiano 2020). This was a reaction to the settlement of increasingly complex and heterogeneous immigrant communities in Europe, a phenomenon that has been called “superdiversity” (Vertovec 2007). Civic integration policies rest on the assumption that the successful incorporation of immigrants into the host society must go beyond their economic and political incorporation and should rely “also on individual commitments to characteristics typifying national citizenship, specifically country knowledge, language proficiency, and liberal and social values” (Goodman 2010, 754). As language acquisition is often regarded as critical for the integration of immigrants, and for cooperation between immigrants and natives (Eckert 2018; Forrest, Benson, and Siciliano 2018), many integration measures aim to facilitate language acquisition (Duncan 2020). Moreover, it is assumed that migrants who know the country’s language are familiar with its culture and are therefore sufficiently integrated into the country (Goodman 2010). However, little is currently known about how such civic integration measures affect language acquisition. This is in part because of a lack of multinational data that can be used to compare the effects of different civic integration measures on different migrant groups (Frank van Tubergen and Kalmijn 2005). To address this knowledge gap, we use data on language use obtained from Twitter for the period from January 2012 to December 2016. We study the pace of migrants’ acquisition of the destination country language and assess whether and how this pace is associated with different civic integration policies in the EU-15, as categorized by Goodman (2010). Our use of Twitter data enables us to study changes in language use in a longitudinal and nonintrusive way among immigrants to the EU-15 countries with a large number of countries of origin. Unlike traditional data used in migration research, Twitter data can provide researchers with continuous access to transnational and comparable migration data. Because of these properties, Twitter data have been used to study different aspects of migration. For example, Mazzoli et al. (2020) showed that geo-located Twitter data can be used to monitor the migration routes, settlement areas, and mobility of migrants and that the data is correlated with official migration data from international agencies. Similarly, Zagheni et al. (2014) used data from 500,000 geo-located tweets to estimate migration flows from Twitter users in OECD countries, and Hawelka et al. (2014) used geo-located tweets to uncover global patterns of human mobility. While Twitter data have been used less frequently in research on integration, Lamanna et al. (2018) showed that language use patterns on Twitter can be used to study the interplay between migrant integration, social polarization, and spatial segregation in different migrant communities in more than 50 cities. Following Lamanna et al. (2018), we study immigrant integration patterns by analyzing the language they use in their tweets. Studies have shown that there is a positive correlation between language acquisition and language usage (F. van Tubergen and Kalmijn 2008). This is because SOFIA GIL-CLAVEL /ANDRÉ GROW /MAARTEN J. BIJLSMA 471 immigrants who use the host language in day-to-day contexts are better able to learn it and because those who learn the language better use it more in their everyday life. Therefore, our central assumptions are (1) that a switch from tweeting in the language of the country of origin to tweeting in the language of the country of destination is an indicator of language acquisition among migrants; and (2) that the time frame over which this switch happens provides insight into the pace of language acquisition. To develop hypotheses about how different civic integration and citizenship policies affect language acquisition, we draw on the work of Goodman (2010) and Howard (2010). Goodman (2010) and Howard (2010) proposed classifying the EU-15 countries according to their requirements for civic integration and citizenship. Conceptually, we rely on the governmentality framework that theorizes the effects of governmental interventions on individuals (Foucault 1991). In a nutshell, the governmentality framework holds that the government has the power to modify people’s behavior through policy interventions (Foucault 1991). In our analysis, one complicating factor is that the use of social media itself may affect the process of language acquisition. Some scholars have argued that social media makes it easier for migrants to stay in touch with communities in their countries of origin. Therefore, factors that affected language acquisition in the past, such as the geographical distance between the origin and the destination country, may lose their importance (Komito 2011; Wright and Viggiano 2020). To assess this possibility, we also study the effects of factors that have traditionally been considered in studies of language acquisition conditional on civic integration and citizenship policies. Background Language is considered an important factor in the integration process, as acquisition of the host country language facilitates cooperation between immigrants and natives (Eckert 2018; Forrest, Benson, and Siciliano 2018). Indeed, for immigrants, mastering the language of the destination improves their access to education and important institutions and is associated with higher income, more societal recognition, and more social contacts (Duncan 2020). Thus, learning the language of the host country facilitates the acquisition of human capital in the country of destination (Esser 2006). Because language acquisition plays a central role in the integration process, it has always been considered an important variable in the study of immigrant integration (Algan, B et al. 2012; Esser 2006), and it has been the focus of civic and integration policies (De Haas, Castles, and Miller 2020; Wright 2020). The role of civic integration policies Foucault (1991) was among the first to theorize that governmental programs have the capacity to change the behavior of the population. This 472 MIGRATION POLICIES AND IMMIGRANTS’LANGUAGE ACQUISITION IN EU-15 notion is captured in the term governmentality, which refers to the different effects governmental interventions have on individuals depending on their positions in relation to governmental programs (Li 2007). These interventions may be related to poverty, health, and demographic events, such as migration and fertility (Castro-Gómez 2010; Li 2007). Civic integration requirements represent a special category of governmental interventions in which immigrants are the target population. Civic integration requirements usually have a twofold nature. First, they are designed to assist newcomers in acquiring the local language, accessing basic services, and entering the labor market; that is, they promote migrants’ individual autonomy (Duncan 2020; Goodman 2010). Second, civic integration requirements are “intended to protect the host society from the presence of others becoming socially disruptive” (Duncan 2020, 604). Menjívar and Lakhani (2016) found that immigrants gradually adopt new behaviors in response to governmental interventions and that the existence of host country citizenship requirements motivates immigrants to adopt new behaviors and lifestyles in both the short and the long term. According to Menjívar and Lakhani (2016), immigrants may adopt these behaviors in part out of a fear of being deported, and in part because they are seeking to fit into the legal categories through which they can gain admission to the United States. Across countries, many types of civic integration policies have been implemented (Helbling 2013). In the European context, Goodman (2010) systematically examined three relevant policy field outputs (immigration entry, integration, and citizenship) with a special emphasis on language requirements (Helbling 2013). Language requirements are often included in civic integration policies, as it is assumed that knowing the country’s language means that an applicant is familiar with the country’s culture and is therefore sufficiently integrated into the country (Goodman 2010). Goodman’s (2010) classification was done by clustering the EU-15 countries based on their citizenship access and membership content policies. The notion of citizenship access comes from the Citizenship Policy Index (CPI), which evaluates the 2008 citizenship policies of the EU-15 countries (Goodman 2010; Howard 2010). Howard (2010) derived this index by developing theoretical arguments and analyzing cross-national empirical findings. The content of the index has been validated in Helbling (2013). According to Helbling (2013), CPI measures what it intends to measure: namely, the variance of outputs of citizenship policies. Moreover, the CPI is highly correlated with other indexes that cover similar policy components. The notion of membership content is based on the Civic Integration Index (CIVIX), which analyzes the requirements for country knowledge, language, and values (Goodman 2010). Citizenship requirements are the rules that determine the extension of legal status and rights depending on state membership (entrance, settlement, or citizenship), while integration requirements are related to the degree to which newcomers have become SOFIA GIL-CLAVEL /ANDRÉ GROW /MAARTEN J. BIJLSMA 473 integrated into the host society (based on factors such as language acquisition and commitment to values) (Goodman 2010; Howard 2010). Based on the CPI and CIVIX indexes, Goodman (2010) clustered the EU-15 countries into four groups: (1) prohibitive, (2) conditional, (3) enabling, and (4) insular. In the following paragraphs, we use Goodman’s (2010) typology to further elaborate on the civic integration and citizenship measures adopted by the EU-15 countries. The prohibitive group. The prohibitive group is made up of Austria, Denmark, and Germany. The countries in this group have relatively strict citizenship requirements (e.g., no dual nationality and a long period of residence in the country before citizenship can be acquired) (Howard 2010) and integration requirements (e.g., mandatory language requirements and country knowledge) (Goodman 2010). According to Howard (2010), Germany has more liberal citizenship policies than Austria and Denmark. This is because in Austria and Denmark, anti-immigrant attitudes are relatively strong, and there is a lack of economic pressure to liberalize the citizenship requirements (Howard 2010). The language acquisition policies of the countries in the prohibitive group have been characterized by a lack of tolerance of different cultures, which are seen as a threat to the language and the culture of the host society (Beauzamy and Féron 2012; Brochmann and Hagelund 2011; Schierup et al. 2006). Austria did not start offering language training to immigrants until 2002. Prior to that time, immigrants were expected to learn the language on their own (Höhne 2013). In Germany, the government began to finance language courses starting in the mid-1970s, which was before the country had even established integration policies (Höhne 2013). Several studies have highlighted the lack of flexibility of Danish immigration policies. In Denmark, migrants who lack a perfect command of Danish suffer from social and labor market discrimination (Beauzamy and Féron 2012; Lønsmann 2020). Austria, Germany, and Denmark are among the European countries that required a high level of language acquisition (B1 level based on the Common European Framework of Reference for Languages (CEFR)) as a condition for permanent residence and citizenship before 2012 (Höhne 2013). The conditional group. The conditional group consists of France, the United Kingdom, and the Netherlands. The countries in this group combine liberal citizenship criteria with arduous integration requirements. In these countries, citizenship is seen as a reward for integration. Therefore, migrants must acquire the language and country knowledge before obtaining citizenship, or even before moving to the country (Goodman 2010). As these countries are “traditional” immigration countries with a colonial past (Brett 2002), they have—relatively early by European standards—tried to 474 MIGRATION POLICIES AND IMMIGRANTS’LANGUAGE ACQUISITION IN EU-15 incorporate the immigrant population into the host society by promoting an atmosphere of tolerance and cultural diversity (Algan, Landais et al. 2012; Manning and Georgiadis 2012). While France and the United Kingdom are considered historically liberal countries, the Netherlands liberalized its citizenship policies between 1980 and 2008 (Howard 2010). In France, the United Kingdom, and the Netherlands, tolerance of cultural diversity is embedded in law, and citizenship for newcomers is essential to the national identity (Castles, De Haas, and Miller 2013). In the past, the governments of these countries believed that this openness to diversity would lead immigrants to feel that they were part of the wider community. Over time, however, these governments became concerned that they were failing to create common core values; that is, to integrate immigrants into the wider society (Beauzamy and Féron 2012; Manning and Georgiadis 2012). Therefore, before 2012, immigrants to these countries were required to pass a basic language (A1/A2 level based on the CEFR), culture, and history test to become a citizen, or even to gain admission to the country (Höhne 2013; Manning and Georgiadis 2012). The enabling group. The enabling group is made up of Portugal, Finland, Ireland, Belgium, and Sweden. For the countries in this group, citizenship serves as a mechanism for establishing equal status and rights. Hence, it is assumed that citizenship enables integration instead of rewarding it, which is the opposite approach to that of the conditional group (Goodman 2010). While Belgium and Ireland are considered historically liberal countries, in Portugal, Finland, and Sweden, citizenship requirements became more liberal between 1980 and 2008 (Howard 2010). These countries were able to liberalize their citizenship requirements in part because the levels of support for far-right parties in the population were low, and in part because of other factors, including demographic change and the rise of international norms (Howard 2010). Of these countries, only Portugal and Finland required language certification as a condition for citizenship before 2012 (Goodman 2012). Ireland, Belgium, and Sweden required neither national language nor country knowledge as a condition for citizenship or permanent residence (Goodman 2012; Höhne 2013). Sweden is a particular case, as it was among the first countries to implement language courses for immigrants. The Swedish government started financing these courses as early as 1965 (Höhne 2013). However, it was not until 2009 that Swedish became the national language of the country (Bolton and Meierkord 2013). The insular group. The insular group consists of Greece, Spain, Luxembourg, and Italy. In general, these countries have a restrictive approach to granting citizenship to immigrants (Castles, De Haas, and Miller 2013), but often grant citizenship to descendants born abroad (Goodman 2010). This approach may be attributable to the electoral support for far-right parties SOFIA GIL-CLAVEL /ANDRÉ GROW /MAARTEN J. BIJLSMA 475 and the anti-immigrant attitudes held by the populations of these countries (Howard 2010). The countries in this group have complex language landscapes, with linguistically independent languages being spoken in different regions of the countries or being used for different official purposes1(Bruzos, Erdocia, and Khan 2018; Love 2015; Sharma 2018; Skourmalla and Sounoglou 2021). In Italy and Luxembourg, policy mechanisms intended to standardize and regulate official language usage were introduced at the beginning of the 21st century. However, these policies created conflict with the communities that spoke different languages (Love 2015; Sharma 2018) and made linguistic integration more difficult for immigrants (Odero, Karathanasi, and Baumann 2016). In the case of Greece, language policies were introduced in the 1970s to homogenize the linguistic and cultural landscape of the country (Skourmalla and Sounoglou 2021). Before 2012, Greece required migrants to be proficient in Greek to acquire long-term residency (Tsoukalas et al. 2010). In Spain, there were no regulations regarding official language usage or language acquisition by immigrants before 2015 (Bruzos, Erdocia, and Khan 2018). Citizenship policy and civic integration indexes As shown in the previous section, the countries that make up each of the groups are quite heterogeneous. Therefore, we use the raw CPI and CIVIX indexes, as they allow us to account for more variance in the analyses, and to capture differences that are otherwise blurred by a merely categorical variable. We employ both the CPI and CIVIX indexes to characterize the civic integration policies of the countries, as they capture two important macrolevel factors associated with immigrants’ language acquisition (van Tubergen and Kalmijn 2005): the political climate surrounding immigration and language integration policies. The political climate and language acquisition by migrants are related through anti-immigrant attitudes and left-wing majority governments. This is because, in a country where right-wing parties are in the majority, it is more likely that the people in that country hold strong anti-immigrant sentiments. This implies that immigrants to that country will have less exposure to the host language (van Tubergen and Kalmijn 2005). By contrast, in a country where left-wing parties are in the majority, the society and the political climate tend to be more tolerant toward immigrants, and the policies tend to favor linguistic pluralism; that is, there is likely to be more tolerance of other languages (van Tubergen and Kalmijn 2005). In view of both arguments, the election of left-wing parties could (unintentionally) reduce immigrants’ exposure to the second language and the incentives of acquiring that language. Hence, immigrants in countries with a stronger presence 476 MIGRATION POLICIES AND IMMIGRANTS’LANGUAGE ACQUISITION IN EU-15 of left-wing parties in the government might have a lesser command of the destination language (van Tubergen and Kalmijn 2005). Conservative citizenship policies are associated with strong antiimmigrant attitudes because countries whose citizens have relatively low levels of anti-immigrant sentiments are more likely to liberalize their citizenship policies. By contrast, countries with high levels of xenophobia tend to continue their restrictive citizenship policies (Howard 2010). A latent variable underlying these associations may be education (Dennison and Dražanová 2018). This is because the more educated a population is, the more proimmigration and the more democratic it tends to be (Dennison and Dražanová, 2018). The first set of hypotheses concerns the effects that civic integration requirements and citizenship policies have on language acquisition. The first hypothesis consists of two competing alternatives. On the one hand, it is possible that the more conservative a country’s citizenship policies are, the more time immigrants to the country will need to acquire the language (H1.1a). This is because, in countries with conservative citizenship policies, anti-immigrant attitudes tend to be strong, which implies that immigrants will have fewer chances to use the host language. On the other hand, it is also possible that the more liberal a country’s citizenship policies are, the more time immigrants to the country will need to acquire the language (H1.1b). This is because, in countries with liberal policies that favor linguistic pluralism, immigrants may have fewer incentives to learn or to use the host country’s language. The second hypothesis holds that the more integration requirements a country has, the quicker immigrants to the country learn the language of the host country, as there are more incentives for them to learn it (H1.2). Finally, our third hypothesis of this set holds that there is an interaction between civic integration requirements and citizenship policies: that is, the more liberal a country’s citizenship policies are and the more civic integration requirements the country has, the faster immigrants to the country learn the language (H1.3). This is because the imposition of more civic integration requirements provides immigrants with more incentives to learn the language, while more liberalization implies more tolerance of migrants, which should mean that members of the host population are more open to interacting with migrants. Challenges in the study of language acquisition, Twitter as an alternative When analyzing the effects that different civic integration policies across Europe have on language acquisition among immigrants, several difficulties can arise, mainly due to issues of data availability and data quality requirements. First, as highlighted by Beauchemin (2014), comparable databases SOFIA GIL-CLAVEL /ANDRÉ GROW /MAARTEN J. BIJLSMA 483 FIGURE 2 Migration flows from the regions of origin to EU-15 civic integration groups. NOTE: Numbers represent migration flows. The country codes of the receiving countries are in square brackets (AT: Austria, BE: Belgium, DE: Germany, DK: Denmark, ES: Spain, FI: Finland, FR: France, GB: Great Britain, GR: Greece, IE: Ireland, IT: Italy, LU: Luxembourg, NL: Netherlands, PT: Portugal, SE: Sweden). AU and NZ correspond to Australia and New Zealand, respectively. The users’ countries of origin are quite diverse; in total, our sample contains users with 81 countries of origin. Given this considerable diversity, we categorize the users’ origins into five regions in order to visualize them more effectively: Europe; Africa; Asia; Latin America and the Caribbean (LAC); and North America, Australia, and New Zealand (N. America and AU+NZ). The total number of individuals who migrated from these regions is 564, 46, 205, 195, and 200, respectively. Figure 2 shows the migration flows from these regions of origin to the countries in the civic integration clusters described above: prohibitive, enabling, insular, and conditional. While we could not identify any users who could be classified as immigrants to Luxembourg, we kept the code in Figure 2 as part of the insular group. The total number of immigrants each of these clusters received is 276, 152, 484 MIGRATION POLICIES AND IMMIGRANTS’LANGUAGE ACQUISITION IN EU-15 TABLE 1 Percentage of users by gender and current account status Gender (%) Account status (%) Group Total Female Male Unknown Active Deleted Suspended Conditional 539 32.84 58.25 8.90 77.36 18.74 3.89 Insular 243 37.04 53.49 9.46 77.78 20.16 2.06 Enabling 152 29.60 60.52 9.86 75 22.37 2.63 Prohibitive 276 34.42 59.42 6.15 80.43 15.59 3.99 NOTE: Account status corresponds to the information the Twitter API V2.2 returned on August 10, 2021. 243, and 539, respectively. Table T1 of online Appendix B shows the number of immigrants by country of destination and the percentage who started to tweet in an official language of the destination country. We also checked the distribution of the users by gender and account status by civic integration group (Table 1). Users’ gender was inferred from their user names using the databases Social Security Administration (2020) and Demografix ApS (2021). For this purpose, we have built a dictionary with the weighted probability of a name being male or female according to these databases. From this point onward, we expect that the distributions will be similar in each of the groups; if they are not, we can assume that the sample of users that the Twitter API returns is biased to certain regions. Table 1 shows that the percentages of female, male, and unknown users are equally distributed across the groups, as is the current users’ account status. The percentages of female and male users are similar to those reported in previous findings (Zagheni et al. 2014). The percentages of deleted and suspended accounts are also in line with previous estimates (Armstrong et al. 2021). Before performing the analysis, we validated that these users are (were) migrants by performing a qualitative analysis of a 10% sample of users. We analyzed their tweets and their tweets metadata, as suggested by Armstrong et al. (2021). The qualitative analysis shows that some of the users tweeted as a student in a foreign country, while others became a resident in the new country. For the students, their status is deduced from their tweets indicating that they were sharing their experiences as a newcomer in the country. For the residents, their status is deduced from their tweets, and, for some of them, from their current Twitter status profile in which they share that they are from country A and are currently living in country B. For a small proportion of the users, we could not infer their motivations for moving; nonetheless, we kept them for the analysis. No individuals were detected who could with certainty be classified as nonmigrants. Methodology We model the variable T:time until a user mostly tweets in the language of destination for one month using survival models (S(t)). Where mostly means: more than 50% of the tweets the user posted each month were in the host SOFIA GIL-CLAVEL /ANDRÉ GROW /MAARTEN J. BIJLSMA 485 country language. For this analysis, we use the programming language R (R Core Team 2020) and the survival package (Therneau and Grambsch 2000). To choose the best parametric model to fit our data, we test the linearity of the Kaplan–Meier survival values by plotting ln(−ln( ˆ S(t))) vs. ln(t) (Kleinbaum and Klein 2012, 305). This visual test shows that the best model is Weibull, as the values show a linear behavior and the slope of the line is different from one (Figure A2 of online Appendix C). The Weibull parametrization we follow is given by Kleinbaum and Klein (2012) (Equation 1). S(t)=exp (−λtp),(1) To study the factors that enhance language acquisition, we model the accelerated failure time (AFT) ratios of T:time until a user mostly tweets in the language of destination for one month. We decided to use this model because the results are interpreted as the median survival time until the language is acquired, which we consider to be more directly interpretable than proportional hazards. We model the time until a user mostly tweets in the language of destination for one month as a function of the following seven variables. The first variable is the CPI developed by Howard (2010), in which the higher the value is the more liberal the citizenship requirements of the destination country are. Howard (2010) built this index by aggregating the following factors: whether or not a country grants ius soli; the minimum length of residency required for naturalization; and whether or not naturalized immigrants are allowed to hold dual citizenship. In addition, he penalized countries that have added civic integration requirements (such as language and civic tests). The second variable is the CIVIX developed by Goodman (2010), in which the higher the value is the stronger the integration requirements of the country of destination are. Goodman (2010) built this index by giving points to four different categories of requirements: [whether] third-country nationals [are] accountable, specifically family unification; whether civic conditions are required for entry, settlement or citizenship; the number of requirements across the civic targets of country knowledge, language and values, including integration courses, tests, contracts, oath ceremonies and interviews; and, finally, the severity of requirements along the path to citizenship (for example, a "high" level of language proficiency or cost). (Goodman 2010, 759) These two variables are continuous and range from zero to six. The third variable is an interaction term between both the CPI and the CIVIX. This variable is continuous and ranges from zero to 36. We do not transform any of these variables in order to facilitate interpretation, given the interaction term. The fourth variable is the logarithm of the ratio of Twitter users in the country of origin to Twitter users in the country of destination. These 486 MIGRATION POLICIES AND IMMIGRANTS’LANGUAGE ACQUISITION IN EU-15 latter users are all users who tweeted from the same country during their entire Twitter history. Here, a positive value means there are more Twitter users in the origin country than in the destination country, and a negative value means the opposite. The fifth and sixth variables are linguistic distance and geographic distance. These variables come, respectively, from the “Language” and “Gravity” databases from the Centre d’Études Prospectives et d’Informations Internationales.8Linguistic distance is the variable LP2 (Melitz and Toubal 2014). According to Melitz and Toubal (2014), LP2 shows how close two different native languages are based on the similarity of words with identical meanings. It is interpreted as the smaller the value is the closer the languages are in terms of vocabulary and grammar. Geographical distance is the distance in kilometers between the capitals of the origin and the destination countries (Conte, Cotterlaz, and Mayer 2021). The seventh variable is the percentage of the destination population who self-reported knowing at least one foreign language in 2011, except for the United Kingdom, where we interpolated (Eurostat 2021). These variables are continuous and standardized (meaning we subtracted the mean and divided by the standard deviation). In the model, we do not account for the variables gender and age because there is no evidence that leads us to assume that the age and gender distributions of Twitter users differed between the integration regimes. In the case of gender, this is shown in Table 1, which indicates that all of the regimes have around the same percentages of female, male, and unknown gender users. We tested this argument by running the model (Equation 2 beneath) adjusting and without adjusting for gender (Figure A3 of online Appendix D), and found that the results were similar. We do not know of a satisfactory way to impute age.9Furthermore, while migrants who are Twitter users may, on average, be younger and better educated than other migrants, we have no reason to suspect that the overrepresentation of young and highly educated people differs between the integration regimes. We, therefore, assume that the bias caused by this overrepresentation is the same between the integration regimes. If this assumption is correct, the relative comparison (ratio of accelerated failure times) between the regimes will not be affected, and the comparison is valid even when age is not controlled for. The Weibull AFT function is t=[−ln S(t)]1 pλ−1 p,where λ−1/pis parameterized with regression coefficients (Equation 2) (Kleinbaum and Klein 2012, 308). In general, the AFT is a ratio of survival times corresponding to any quantile (q) of survival time (S(t=q)). In this model, an increase in a variable in which the coefficient is positive leads to an increase in the median (or other quantile) survival time until the language is acquired. If the coefficient is negative, then an increase in the variable would lead to a decrease in the median survival time until the language is acquired. SOFIA GIL-CLAVEL /ANDRÉ GROW /MAARTEN J. BIJLSMA 487 FIGURE 3 Exponential of AFT coefficients with their corresponding 95% confidence intervals λ−1 p i=exp (α0+α1CPIi+α2CIVIXi+α3CPIi×CIVIXi+α4log (ratio)i +α5Ling.Dist.i+α6Geo.Dist.i+α7%≥1ForeignLang.i)(2) where CPI is the citizenship Policy Index; CIVIX is the Civic Integration Index; CPI ×CIVIX is the interaction term; log(ratio) is the logarithm of the ratio of the number of Twitter users from the origin country by the number of Twitter users from the destination country; Lin. Dist. is linguistic distance; Geo. Dist. is geographical distance; and %≥1 Foreign Lang. is the percentage of the destination country population who speak more than one foreign language. Results Figure 3 shows the estimated AFT ratios of the Weibull model with 95% confidence intervals. In the case of CIVIX, we find that the stronger the civic integration requirements are the longer the median survival time until the language is acquired; therefore, H1.2 is not supported. In the case of CPI, it does not appear to play a role in the median survival time until the language is acquired conditional on the other variables in the model. However, the interaction variable shows that the greater the CPI and the CIVIX are the lower the median survival time until the language is acquired. This indicates 488 MIGRATION POLICIES AND IMMIGRANTS’LANGUAGE ACQUISITION IN EU-15 FIGURE 4 Contour map of the accelerated failure time (AFT) relative to the civic integration index (CIVIX) and the citizenship policy index (CPI) that CPI does play a role, but only in conjunction with particular CIVIX levels. To clarify the interaction effect, we show the predicted survival time until the language is acquired using a contour map relative to the CIVIX and CPI indexes in Figure 4. These predicted values are obtained by multiplying the model coefficients by the different combinations of CIVIX and CPI values while keeping the rest of the variables constant at their means (which are zero because of the standardization). Figure 4 shows that when the CIVIX index is below 1.75, the CPI index is not associated with the median survival time until immigrants have acquired the language. This can be seen in the median survival values of the insular and enabling groups (excluding Sweden), which range from 85 to 109 months regardless of the CPI values. Once the CIVIX values are higher than one, the CPI index becomes associated with the median survival time until the immigrants’ language ac- SOFIA GIL-CLAVEL /ANDRÉ GROW /MAARTEN J. BIJLSMA 489 quisition, the more liberalized the country is the lower the median length of time it takes for immigrants to acquire the language (conditional group), and the less liberalized the country is the higher the median length of time it takes for immigrants to acquire the language (prohibitive group). In this analysis, Sweden seems to follow a different pattern from the rest of the countries. This could be because, compared to the rest of the countries in the insular and enabling groups, a higher percentage of Sweden’s population speak at least one foreign language. Returning to Figure 3, in line with our secondary hypotheses, the median survival time until the language is adopted increases if the number of Twitter users in the country of origin is larger than in the country of destination (H2.1). This might be because Twitter users may have more incentives to tweet in a language when there is a larger audience with whom they can interact using that language, as previous research has shown (Chiswick and Miller 2001; Esser 2006). Linguistic distance also has a positive association. This means that the larger the distance between the origin and the destination language is, the longer it takes to acquire the destination language (H2.2). A potential explanation for this finding is that when a language is more difficult for immigrants to learn because it differs considerably from their mother tongue, it takes longer for them to use it (Chiswick and Miller 2001; Esser 2006). For geographic distance, we find that the larger the distance between the origin and the destination country is the shorter the median survival time is until the language is acquired (H2.3). This was expected, as geographic distance is associated with greater incentives to invest in language skills (Chiswick and Miller 2001; Esser 2006). Therefore, the classic variables used to explain immigrants’ language acquisition have the same associations with language acquisition as before the advent of social network sites. Finally, for the control variable, an increase of a percentage point in the share of the destination country population who speak a foreign language doubles the mean number of months until the language is acquired. This may be because migrants may lack incentives to learn the destination country language when the destination population can communicate with them in a language migrants might know (such as English) (Bolton and Meierkord 2013; Cromdal 2013). Discussion and conclusions In this work, we studied immigrants’ language acquisition through a longitudinal analysis of the languages they used in their tweets. To do so, we drew on Goodman’s (2010) and Howard’s (2010) work to formulate how citizenship and civic integration policies may have affected immigrants’ language acquisition. Conceptually, we relied on the governmentality framework that theorizes on the effects that governmental interventions have 490 MIGRATION POLICIES AND IMMIGRANTS’LANGUAGE ACQUISITION IN EU-15 on individuals. We used survival models to analyze the pace of immigrants’ language acquisition depending on (1) citizenship and civic integration policies and (2) the relative sizes of migrant groups in the destination country and the linguistic and geographical distances between the countries of origin and destination. Specifically, we analyzed the length of time it took until a user was mostly tweeting in the language of the destination country for one month. We used starting to tweet in the language of the country of destination as a proxy for language acquisition. Our findings point to an interaction effect between civic integration requirements and citizenship policies, whereby immigrants in countries with loose or no civic integration requirements had similar median times until language acquisition regardless of how liberalized the citizenship policies were. This was the case for countries that had heterogeneous citizenship policies but few civic integration requirements. However, among these countries, Sweden appeared to be a particular case. In Sweden, the median time until language acquisition was similar to that of countries with strict civic integration and citizenship requirements. Among the potential explanations for this finding are that a high percentage of the Swedish population speaks at least one foreign language (Bolton and Meierkord 2013) and that Sweden has high levels of multiculturalism, which could discourage immigrants from learning Swedish (van Tubergen and Kalmijn 2005). We found that in the countries with strict civic integration and citizenship requirements (Denmark, Austria, and Germany), the time it took for immigrants to acquire the language was longer than in the other countries. While this may be a consequence of the anti-immigrant attitudes of the majority population, it has also been suggested that strict requirements placed on immigrant groups may be the result of right-wing parties trying to constrain immigrants from having access to rights equal to those of natives (M. B. Jørgensen 2009; Beauzamy and Féron 2012; Bolton and Meierkord 2013; Lønsmann 2020). Research has shown that these types of negative interactions between authorities and migrants can lead to language balkanization and to immigrants rejecting learning the language of the destination country (J. N. Jørgensen 2003). For those countries with onerous civic integration requirements, we found that the more liberalized immigrants’ access to citizenship was the faster they acquired the host country language. This was shown to be the case for France, the Netherlands, and the United Kingdom, which have strict civic integration requirements but are also considered historically liberal countries (Howard 2010). However, this result might also be explained by the early integration requirements that immigrants had to fulfill, such as learning the language before moving to the country (Goodman 2010), which would then be more related to a selection effect than to migration policies. SOFIA GIL-CLAVEL /ANDRÉ GROW /MAARTEN J. BIJLSMA 491 Our results also showed that the evidence of immigrants’ language acquisition on Twitter was associated with the same classic macrolevel explicative variables employed before the advent of social network sites. This result is relevant in two ways. On the one hand, it supports the notion that the Twitter data actually captures migration. On the other hand, it helps to shed light on the question of whether the transnational property of social network sites has affected the association between immigrants’ language acquisition and classic macroexplicative variables. For this sample of Twitter users, the results showed that this has not been the case. However, this may change in the future, as the use of information and communication technologies is becoming more and more pervasive across the globe. Limitations This work has several limitations that we would like to acknowledge. Twitter data are not representative of the general population. Twitter users tend to be young adult men who are highly educated and highly Internet skilled (Hargittai 2020). For this study, this point is especially important, as the results are not representative of all the migrants who moved to the European countries analyzed. As such, the number of possible migrants found was quite small compared to all the terabytes of information analyzed. Furthermore, in the data, certain vulnerable migrant populations, such as illegal migrants, may not be represented, while other populations, such as international students, may be overrepresented. Therefore, caution is advised when interpreting the results. While our sample was a random selection of tweets, the identification of the users’ location depended on longitudinal Voluntarily Geographic Information (Haklay 2016), that is, the analysis was limited to highly active users that contribute to place-based systems and who are considered content producers. Hence, the final sample is subject to selection. However, as far as we know, there is insufficient evidence to ascertain whether content producers, conditional on the other variables adjusted for (the size of the immigrant group in the destination countries and the linguistic and geographical distance between origin and destination countries) differ from other Twitter users in their time to destination language acquisition. If content producers do differ in this way, then this would result in biased estimates of time-to-destination language acquisition. However, it would not necessarily result in biased ratios of time to language acquisition between the country blocs, the variable of interest in this study, unless different “types” of content producers are also more likely to migrate to different country blocs. This could be a subject for future research. Despite these clear data limitations, we showed that Twitter data can be used to study immigrants’ language acquisition and that the data shows 492 MIGRATION POLICIES AND IMMIGRANTS’LANGUAGE ACQUISITION IN EU-15 patterns similar to those found in analyses done with representative samples collected before the advent of social network sites (Chiswick and Miller 2001; Esser 2006). However, it is important to continue studying and developing statistical techniques, training databases, and machine learning algorithms to model data from social network sites to study hard-to-reach populations, such as migrants. Research ethics This work obtained ethical approval from the data protection department of the Max Planck Institute for Demographic Research. For the analyses, we relied on public data from the Internet Archive, and we studied only the language of the users’ tweets. For research purposes only, we also read the public description of the profiles of a small sample of the users. Acknowledgments We would like to thank Lee Fiorio (University of Washington), Clara Mulder (University of Groningen), and Emanuele del Fava (Max Planck Institute for Demographic Research) for their feedback. Conflict of interest The authors declare no conflict of interest. Reproducibility Given Twitter’s terms and conditions, we do not share the final database, as this can lead to the disclosure of user IDs. All of the code needed to replicate this work is available in Gil-Clavel’s GitHub repository: https://github.com/ SofiaG1l Data availability statement The data to reproduce this work are freely available in the Internet Archive: https://archive.org/details/twitterstream Notes 1 This is the case for Luxembourg, where “Luxembourgish is the national language, French the legislative language, and German is the language of instruction in public schools” (Odero, Karathanasi, and Baumann 2016, 4067). 2 Twitter announced that the character limit would be increased to 280 in 2017. https://blog.twitter.com/official/en_ us/topics/product/2017/Giving-you-morecharacters-to-express-yourself.html. Accessed on October 21, 2020.