scieee AI-readable full text Open interactive document viewer

The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators' Data

Bereznyatskiy, Alexander N.; Badina, Svetlana V.; Babkin, Roman A.

Abstract

This study investigates the spatial and temporal characteristics of population distribution in the city of Moscow, with a focus on the constancy of these dynamics over time, using high-frequency data from mobile network operators. The primary unit of analysis is a 500 × 500 meter grid cell covering the city's territory. The analysis spans the period from 2018 to 2020. The stability of population distribution is assessed along two dimensions: the spatial distribution of the permanent population and the interannual variation of daily population gradients across grid cells. The hypothesis of distributional constancy is tested through the construction of regression models, followed by a residual analysis using Geographic Information Systems (GIS). Spatial heterogeneities identified in the residuals are further examined. The findings indicate that, for a number of Moscow districts, the spatial population distribution closely follows a pattern consistent with the Zipf – Mandelbrot law, while other districts deviate significantly from this model. This pattern was further validated using Rosstat data for Moscow districts covering the years 2013–2022, confirming both the functional form and the classification of districts into two distinct groups. The first group primarily includes areas within Moscow's 2012 administrative boundaries and high-density zones in New Moscow, whereas the second group comprises the remaining, less densely populated areas. The proposed methodology demonstrates that the distribution of the permanent population across Moscow remained generally stable in both spatial and temporal terms over the study period. Conversely, the hypothesis of temporal constancy in daily population gradients is rejected, likely due to the influence of complex factors such as the evolving urban transportation infrastructure and cumulative measurement errors.

Full text

RESEARCH ARTICLE Copyright Bereznyatskiy AN, Badina SV, Babkin RA. This isan open access article distributed under the terms ofthe Creative Commons Attribution License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction inany medium, provided the original author and source are credited The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data Alexander N. Bereznyatskiy 1, Svetlana V. Badina 2, Roman A. Babkin 3 1 Central Economics and Mathematics Institute of the Russian Academy of Sciences, Moscow, 117418, Russia 2 Lomonosov Moscow State University, Moscow, 119991, Russia; Bernardo O’Higgins University, Santiago, 8320000, Chile 3 Plekhanov Russian University of Economics, Moscow, 117997, Russia Received 23 March 2024 ♦ Accepted 24 February 2025 ♦ Published 3October 2025 Citation: Bereznyatskiy AN, Badina SV, Babkin RA (2025) The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data. Population and Economics 9(3):80-93. https://doi.org/10.3897/ popecon.9.e123730 Abstract This study investigates the spatial and temporal characteristics ofpopulation distribution inthe city ofMoscow, with afocus onthe constancy ofthese dynamics over time, using high-frequency data from mobile network operators. The primary unit ofanalysis isa 500× 500meter grid cell covering the city’s territory. The analysis spans the period from 2018to2020. The stability ofpopulation distribution isassessed along two dimensions: the spatial distribution ofthe permanent population and the interannual variation ofdaily population gradients across grid cells. The hypothesis ofdistributional constancy istested through the construction ofregression models, followed bya residual analysis using Geographic Information Systems (GIS). Spatial heterogeneities identified inthe residuals are further examined. The findings indicate that, for anumber ofMoscow districts, the spatial population distribution closely follows apattern consistent with the Zipf– Mandelbrot law, while other districts deviate significantly from this model. This pattern was further validated using Rosstat data for Moscow districts covering the years 2013–2022, confirming both the functional form and the classification ofdistricts into two distinct groups. The first group primarily includes areas within Moscow’s 2012administrative boundaries and high-density zones inNew Moscow, whereas the second group comprises the remaining, less densely populated areas. The proposed methodology demonstrates that the distribution ofthe permanent population across Moscow remained generally stable inboth spatial and temporal terms over the study period. Conversely, the hypothesis oftemporal constancy indaily population gradients isrejected, likely due tothe influence ofcomplex factors such asthe evolving urban transportation infrastructure and cumulative measurement errors. Keywords big data, mobile operator data, urban population density, regression analysis, Zipf– Mandelbrot law JEL codes: C01, C21, R14, R23 Population and Economics 9(3): 80–93 DOI 10.3897/popecon.9.e123730 Population and Economics 9(3): 80–93 81 Introduction Understanding the dynamics ofpopulation distribution within urban areas isa critical component ofmodern city modeling. Traditionally, population studies have focused ontotal population changes ordemographic structures, often overlooking the spatial and structural dimensions ofurban settlement patterns (Argunov etal. 2017). However, for avariety ofapplications– including population risk assessment (Tenerelli etal. 2015), the evaluation ofurban policy impacts, and the regulation oftraffic flows (Saghapour etal. 2016)– it isessential toconduct adetailed statistical analysis ofsettlement structures and their temporal dynamics. Such analyses enable the identification ofstable territorial formations and anomalous spatial behaviors, facilitating further modeling ofinfluencing factors (e.g., urban development projects, housing renovation programs; see (Badina etal. 2023)) and forecasting changes insettlement patterns over time. Until recently, the primary limitation inconstructing such detailed models has been the lack ofappropriate data: • with high spatial granularity; • updated athigh temporal frequencies (e.g., minuteorhour-level observations). Recent advances in data availability from cellular network operators and providers ofpublic Wi-Fi infrastructure offer unprecedented opportunities. These data sources enable the tracking ofsubscriber locations atfine spatial and temporal resolutions, effectively merging the spatial precision ofsatellite imagery with the quantitative depth oftraditional demographic statistics. Several studies have already taken advantage ofsuch data inanalyzing settlement dynamics, including notable work focused onthe Moscow region (Makhrova etal. 2016; Makhrova and Babkin 2018; Mayakov 2022). Internationally, comparative analyses ofvarious data sources (e.g. (Lenormand etal. 2014)) have confirmed the high representativeness ofmobile phone data for capturing urban dynamics. Numerous studies have also addressed the broader topic ofurban structure (Babkin etal. 2022b). One ofthe earliest and most prominent initiatives using mobile data was the Austro-American Real-Time Graz project, launched in2005(Ratti 2005). Leveraging data from A1, Austria’s largest mobile network operator, the project successfully mapped labor migration patterns and shifts inpopulation distribution inthe city ofGraz. Similarly, Czech researchers used correlations between human mobility patterns and the spatial distribution ofresidential and commercial real estate toclassify daily rhythm types across Prague (Nemeškal etal. 2020). They demonstrated that mobile phone data– particularly indicators such asduration and timing ofuser presence– can reliably differentiate between residential, work, transportation, and service areas within the urban landscape. Comparable studies have since been conducted ina wide range ofglobal cities, including New York and Los Angeles (Isaacman etal. 2011), Harbin (Yuan etal. 2012), Rome (Calabrese etal. 2013), Singapore (Zhong etal. 2016; Jiang etal. 2017), London (Zhong etal. 2016), Beijing (Zhong etal. 2016), and Jakarta (Ruslani etal. 2019). There are numerous examples ofproject implementations where mobile operator data serve not merely asa supplementary source, but asthe primary and often the only reliable source ofrelevant information. This transition allows researchers and planners tobypass traditional data collection methods and instead leverage real-time information derived from mobile devices and satellite imagery (Giugale 2012). Mobile phone data has proven especially valuable incontexts where official population and mobility statistics are absent, outdated, orunreliable (Lu etal. 2012; Ruslani etal. 2019; Deville etal. 2014). A.N. Bereznyatskiy et al.: The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data 82 Atthe same time, the widespread adoption ofGeographic Information Systems (GIS) inurban studies has significantly enhanced the capacity for spatial analysis ofpopulation distribution (Dolgacheva etal. 2022). When combined with granular big data, GIS technologies offer powerful tools for exploring complex patterns ofurban settlement. Equally important isthe rapid development ofopen-source statistical modeling tools that integrate GIS capabilities. Software environments such asR (notably through packages described in(Bivand etal. 2013)), released under the GNU license, now offer robust functionalities that greatly simplify spatial statistical analyses for researchers. This study presents one such approach for analyzing the “sustainability”– used interchangeably here with the term “constancy”– of the territorial distribution ofthe population inthe city ofMoscow. The stability ofthe underlying processes isinvestigated from two complementary perspectives: • the stability ofthe spatial distribution ofthe permanent population; • the stability ofthe spatial distribution ofpopulation gradients (or pulsations) across the city. Toaddress these objectives, the study employs methods ofapplied statistical analysis and cartographic visualization. Methodology This study isbased ona statistical dataset derived from primary information provided bycellular operators, detailing the positions ofnetwork subscribers within 500× 500meter grid cells across the city ofMoscow, with atemporal resolution of30minutes. The anonymized data were supplied bythe Department ofInformation Technology ofthe City ofMoscow. The full dataset includes approximately 10.000spatial cells, each tagged with unique identifiers that enable geographic referencing tospecific city areas. The data span the period from 2018to2020(for more details, see (Badina etal. 2022)). Prior toanalysis, the dataset underwent preprocessing toeliminate potential double counting– such asinstances where asingle user possesses multiple SIM-cards. Given the focus onterritorial distribution, the accuracy ofsubscriber positioning isa relevant concern. Areview ofthe literature oncellular-based positioning systems (Resch and Romirer-Maierhofer 2005; Babkin etal. 2022a) indicates that the achievable accuracy ranges from afew meters toseveral hundred meters, depending largely onthe density ofbase stations. Maximum precision istypically observed inlarge urban areas. The choice ofa 500× 500meter grid was therefore made toalign with the optimal spatial resolution achievable using the available data. To assess the stability ofthe permanent population distribution, asubsample ofthe dataset was created byextracting records corresponding to2:00a.m.– a time assumed torepresent the typical residential location ofsubscribers. For each month, the median population values were computed for each cell (a similar approach was applied inanalyzing daily population gradients). This assumption rests onthe premise that most individuals are athome during this hour, thus minimizing distortions from intra-urban movement– a finding supported byprevious research (Badina and Babkin 2021; Babkin etal. 2022b). The population distribution and its temporal consistency are illustrated inFig. 1. The analysis ofdaily population fluctuations across the city was based ona derived metric: the daily population gradient, defined asthe ratio between the maximum and minimum population observed ina given cell over a24-hour period. These pulsations are also visualized inFig. 1. Population and Economics 9(3): 80–93 83 Figure 1. The change inthe population ofindividual Moscow cells during the day, 2019, the number at0-00hours istaken as100%. Source: compiled bythe authors based ondata from mobile operators. № 7113 (Akademichesky District) 212 196 180 164 148 132 116 100 84 % 000 100 200 300 400 500 600 700 800 900 1000 1100 1200 1300 1400 1500 1600 1700 1800 1900 2000 2100 2200 2300 Hour № 83998 (Donskoy District) № 23526 (Alexeyevsky District) Moscow, aggregate № 86675 (Akademichesky District) № 34546 (Alexeyevsky District) № 59734 (Bibirevo District) The research methodology follows the sequence outlined below. At each time point t (where tcan represent hours, days, months, years), the population within each spatial cell isrecorded asxi, where i= 1…N, N isthe total number ofcells covering the city ofMoscow. Alternatively, the analysis may beconducted ata more aggregated level, such asadministrative districts. Thus, weobtain apopulation distribution vector (data set): X x x x t t t t N =                   1 2 . . . , which, for each period t, shows the distribution of the population across the territory ofMoscow. Since all cells have astandard size of500 by500meters, the concepts ofpopulation density structure and population distribution across cells can beconsidered mathematically equivalent. Further, ifthe population density structure attimes t1and t2does not undergo major changes during the transition from one period toanother, this means that vector X t 1 and vector X t 2 are consistent with each other, and the stability ofthe settlement structure over A.N. Bereznyatskiy et al.: The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data 84 time isobserved. Otherwise, weare dealing with variability inthe density structure over the period t1, t2. The question ishow toevaluate the consistency ofthe two data sets. In this paper, itis proposed toassess data consistency using aregression model inwhich one data set ismapped onto the other. Subsequently, the hypothesis ofthe invariance ofthe population distribution across the territory ofMoscow istested using this approach, both inthe case ofthe permanent population and inthe analysis ofthe stability ofpulsations over time. Considering both the random component ofthe data (measurement errors, etc.) and the randomness ofpossible observed changes, the hypothesis istested using aregression test model ofthe following type: XX t m t n =+ + 01 αα ε, (1) where α0, α1 – unknown regression parameters (estimated for two test vectors using the least squares method); X tm – the vector ofpopulation distribution across the city ata given time tm; X tn – the vector ofpopulation distribution across the city ata given time tn; ε– arandom residual reflecting the discrepancy between the two data sets. The less consistent the vectors are, the greater the residual component ofthe regression will be when estimating the parameters. For example, inthe case ofself-regression (i.e. tm=tn), the residual component will bezero, which isexpected– asthere are nochanges inthe structure. The hypothesis about the constancy ofthe population distribution over time (for the selected interval) will beconfirmed ifa regression model that isstatistically adequate isobtained. Weare primarily interested inthe coefficient ofdetermination R2, which ranges from 0to1. Avalue of0indicates complete inconsistency between the two data sets, while avalue of1indicates full consistency. Special attention ispaid tothe residual component ε. The values ofthe residuals are recorded inGIS and displayed onmaps for visual analysis, providing more complete information about the phenomenon inquestion. Results and their discussion The methodology described inthe previous section was applied toa sample ofMoscow data for the period 2018–2020. February 2018was taken ast1, and January 2020ast2. The overall appearance ofthe data cloud and the regression line isshown inFig. 2. Figure 2displays apronounced elliptical shape with asmall number ofoutliers. Itis assumed that the constructed regression model for the two vectors will have good statistical properties and, overall, will not reject the hypothesis ofthe constancy ofthe population structure inMoscow for the period from February 2018toJanuary 2020. Table 1presents the results ofthe statistical evaluation ofthe regression model parameters. The test results show that, ingeneral, the hypothesis ofthe stability ofthe distribution ofthe permanent population across Moscow’s cells isnot rejected over atwo-year interval (the coefficient ofdetermination R2) for the regression model was 0.92, indicating good consistency between the two data sets). However, analysis ofthe model residuals and their visualization (Fig. 5) reveals the presence ofoutliers. Although itwould bepossible todisre- Population and Economics 9(3): 80–93 85 gard individual deviations against the background ofthe overall residual pattern, the clear spatial clustering ofthese outliers onthe map suggests that they are not random. Toexplore this result ingreater detail, anadditional analysis was conducted onthe rank patterns ofcells (ranking bythe number ofpermanent residents per cell) and the corresponding population values. This approach– Zipf frequency analysis, modified byMandelbrot– is well known inlinguistics and has been adapted for use inurban demography (PavFigure 2. Visualization ofcomponents and regression lines for vectors ofdistribution ofpermanent population inMoscow cells, people (on the horizontal axis– the number ofcells for February 2018, onthe vertical axis– for January 2020). Source: calculated bythe authors based ondata from mobile operators. 12000 10000 8000 6000 4000 2000 January 2020 0 0 1000 2000 3000 4000 5000 6000 7000 8000 9000 10000 February 2018 Table 1. Regression model testing the stability ofthe permanent population distribution across Moscow’s cells for the period February 2018– January 2020  Estimation ofthecoefficients ofthe model Standard error t-statistics Significance level α094.3 5.3 17.8 *** α11.04 0.003 339 *** R2= 0.92 F (1.10087) = 114679** Number ofobservations = 10089 Source: authors’ calculations based ondata from mobile operators. Note: *10%, **5%, ***1% level ofsignificance A.N. Bereznyatskiy et al.: The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data 86 lov 2020; Jingang etal. 2019; Urzúa 2000). Inthis study, the technique isapplied tospecific territories within the city (Li etal. 2009). The frequency diagram ofthe rank-size distribution ofcells bypopulation ispresented inFig. 3and clearly demonstrates two distinct modes ofpopulation distribution across Moscow’s cells. Experiments with model fitting toapproximate the distribution confirm the results ofthe visual analysis. Toanalyze the result over alonger time horizon, the data were aggregated from individual cells tothe level ofMoscow districts. Population distribution models were constructed for the period from 2013to2022(Fig. 4). Figure 3. The rank-frequency chart for Moscow cells, January 2020. Source: calculated bythe authors based ondata from mobile operators. 13 12.5 12 11 10 9 logarithm of the population of cells 8 01234 56 logarithm of cell rank 11.5 10.5 9.5 8.5 Figure 4. Chart rank-frequency for Moscow districts and model values, left graph– 2022, right graph– models for 2013, 2022incomparison. Source: calculated bythe authors according toRosstat data. 200000 100000 0 observed data 250000 150000 50000 1 12 23 34 45 56 67 78 89 100 111 122 133 144 modelmodel 2013 model 2022 200000 100000 0 250000 150000 50000 1 8 15 22 29 36 43 50 57 64 71 85 92 99 78 106 113 120 127 134 141 Population and Economics 9(3): 80–93 87 In terms ofthe population distribution model, all Moscow districts can begrouped into two classes. InFig. 4, this isshown asthe merging ofa logarithmic and alinear model atan inflection point. Figure 6presents the two identified clusters and the number ofcells onthe map. Ascan beseen, the distribution ofpopulation densities and the two clusters are closely interrelated. Itwas found that asubset ofareas with high population density iswell described bya logarithmic population distribution model, while the remaining areas correspond toa linear model. Importantly, the classification bytype ofpopulation distribution model remains stable across all periods from 2013to2022(calculations based onRosstat data). The parameters ofthe models, calibrated for each period separately, are also stable, asclearly shown bythe right-hand graph inFig. 4. It isworth referring once again tothe results oftesting for the constancy ofthe population distribution across Moscow’s cells and comparing Fig. 5with Fig. 6. Inthis context, the origin ofthe outliers inthe regression model becomes evident. Two different mechanisms within the settlement system give rise topatterns inthe residuals. This mechanism remains stable, atleast throughout the period from 2013to2022. 5.0 Administrative Okrugs: 0.5 0.25 0 -0.25 -0.5 -2.0 Regression model residuals CAO – Central NWAO – North-Western NAO – Northern NEAO – North-Eastern EAO – Eastern SEAO – South-Eastern SAO – Southern SWAO – South-Western WAO – Western NMAO – Novomoskovsky TAO – Troitsky ZelAO – Zelenogradsky ZelAO NEAO NAO NWAO CAO EAO WAO SWAO SEAO SAO NMAO TAO Figure 5. Visualization ofthe residual component ofthe regression model– the stability test ofthe distribution ofthe permanent population ofMoscow for the period February 2018– January 2020. Source: calculated bythe authors based ondata from mobile operators. A.N. Bereznyatskiy et al.: The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data 88 During the 2012 administrative reform, various suburban territories were annexed toMoscow, forming New Moscow. The modeling conducted atthe level ofsmall territorial cells shows that these still constitute two distinct areas, despite their formal unification. This isunsurprising. The territories inquestion have undergone long-term development (in terms ofdemographics, economics, etc.), and itwould benaïve toexpect that their formal merger would immediately alter the demographic and economic systems formed over centuries. This type ofphenomenon isknown asa temporary decline inurban population density over time (Gang etal. 2019), when acity’s territory expands ata pace far exceeding the population growth ofthe newly incorporated areas. Let usnow consider the second part ofthe task– testing the stability ofthe territorial distribution ofpopulation gradients inMoscow. The testing mechanism issimilar tothe one previously discussed, with the only difference being the composition ofthe vector Xt. Inthis case, instead ofthe permanent population, the values ofthe daily gradients calculated for the periods ofFebruary 2018and January 2020are used for each cell. Figure 7shows the data cloud and the regression line for the corresponding values inFebruary 2018and January 2020. In contrast toFig. 2, there isno observable alignment ofpoints along astraight line; asexpected, the hypothesis ofgradient constancy islikely tobe rejected. The results ofthe regression parameter estimation are presented inTable 2. According tothe results ofthe regression assessment, the hypothesis regarding the constancy of the distribution ofdaily pulsations in Moscow’s cells for the period February 2018– January 2020isindeed rejected. The regression model has alow coefficient ofdetermination, R2= 0.3, whereas inthe case ofthe permanent population distribution, itwas 0.92. Apparently, daily population changes inindividual cells are significantly more complex than inthe case ofthe permanent population, which isreflected inthe marked discrepancy inthe distribution ofpulsations byregion during the interval from February 2018toJanuary 2020. During this period, changes may have occurred inthe transport scheme (such asthe development ofmetro lines ormodifications tohighway interchanges), which likely influenced the test results. Figure 6. Number ofpermanent population ofcells (left map) and clustering according tothe law ofdistribution ofcell densities ofthe city ofMoscow. Source: calculated bythe authors based ondata from mobile operators. 10000 + 5000 2500 1000 500 250 0Administrative Okrugs: CAO – Central NWAO – North-Western NAO – Northern NEAO – North-Eastern EAO – Eastern SEAO – South-Eastern SAO – Southern SWAO – South-Western WAO – Western NMAO – Novomoskovsky TAO – Troitsky ZelAO – Zelenogradsky Zipf model Linear model Administrative Okrugs: CAO – Central NWAO – North-Western NAO – Northern NEAO – North-Eastern EAO – Eastern SEAO – South-Eastern SAO – Southern SWAO – South-Western WAO – Western NMAO – Novomoskovsky TAO – Troitsky ZelAO – Zelenogradsky