Full text
Open Access © The Author(s) 2023. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the mate‑ rial. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http:// creat iveco mmons. org/ licen ses/ by/4. 0/. RESEARCH Chenand Dietrich Applied Network Science (2023) 8:60 https://doi.org/10.1007/s41109-023-00585-0 Applied Network Science Normalized closeness centrality ofurban networks: impact ofthelocation ofthecatchment area andevaluation based onanidealized network Hsiao‑Hui Chen1* and Udo Dietrich2 Abstract The decision of where to locate the catchment area of an urban network exerts sig‑ nificant influence on the indicator values and in this research this influence is referred to as the placement effect. Placement effect has significant impact on the stud‑ ies at the neighborhood scale focusing on the structural properties of the network models, the network analysis results and centrality measures, the inferred movement patterns and the accessibility to destination. Placement effect becomes even more sig‑ nificant when multiple catchment areas are sampled to be compared or classified. This research examines placement effect on one of the most affected indicators, closeness centrality, and proposes using an idealized network as a reference to be compared with the real network in order to find a solution to mitigate the placement effects. By comparing the normalized closeness centrality in the real network with that in the ide‑ alized network, we can (1) evaluate the placement effect on the closeness centrality and (2) find the threshold distance in order to mitigate the placement effect. The results show that the closeness centrality of the same node varies remarkably depend‑ ing on its position and how central it is in the chosen catchment area. Specifically, in the selected areas in this research, if the center point of a catchment area is moved by more than 100 m away from the original center point, the closeness central‑ ity of the same node starts to be significantly influenced by the placement effect. The threshold distance of 100 m offers a recommendation that a direct comparison of the closeness centrality between different nodes in the same catchment area should be drawn only if these nodes are less than 100 m away from each other. In other words, when comparing two nodes located further than the threshold distance from each other, it is advisable to create two separate catchment areas, where these nodes serve as the center points. It should be noted that the threshold distance of 100 m derived specifically from the current research should not be generalized to other cases. The threshold distance of different case studies remains open for further investigation in the future as it may vary among cities or areas. Keywords: Street network, Placement effect, Boundary effect, Edge effect, Normalized closeness centrality *Correspondence: hsiao‑hui.chen@tu‑ braunschweig.de 1 Technische Universität Braunschweig, SpACE Lab at ISU – Institute for Sustainable Urbanism, Pockelsstr. 3, 38106 Brunswick, Germany 2 HafenCity Universität Hamburg, Architektur Und REAP, Henning‑Voscherau‑Platz 1, Raum 4.107, 20457 Hamburg, Germany
Page 2 of 14 Chenand Dietrich Applied Network Science (2023) 8:60 Introduction Urban network system is a complex spatial system whose members connect and interact with each other. For any real-world spatial network analysis, it is essential to define an artificial border of the network model (Park 2009). Spatially confined networks or, in other words, local sub-networks of the entire global network system have been termed catchment areas (Chen and Dietrich 2021), contextual areas or bounded systems (Park 2009), subnetworks or regional networks (Rheinwalt etal. 2012). Drawing an arbitrary boundary inevitably cuts the links connecting the catchment area under investigation with the rest of the network outside the boundary (Rheinwalt etal. 2012). However, the events, structures, behavior and dynamics of the entire global network system still affects the local sub-network inside the catchment area (Greenberg etal. 2020). The inevitable arbitrary delineation of the boundary may induce distortion of the results of the measures, which can subsequently induce bias that affects the inferences based on these measures (Paul 2014). Such distortion of the results is found to be more pronounced when the nodes or links are closer to the border of the catchment area (Okabe and Sugihara 2012). These boundary determination problems, which have been termed the edge effect (Crucitti etal. 2006; Gil 2017; Ripley 2004) or the boundary effect (Park 2009; Okabe and Sugihara 2012), have significant impact on the studies focusing on the structural properties of the network models, the network analysis results and centrality measures, the inferred movement patterns and the accessibility to destination. First of all, the choice of the boundaries decides the internal structure of the local network model that is spatially confined within the catchment area. This decision directly influences the members of the sub-structure and topology included in the network model and, therefore, affects our understanding of the spatial structure and functional properties of the network system (Laumann etal. 1989). Secondly, the delineation of the model boundary can cause a certain bias in network analysis results (Ratti 2004; Joutsiniemi 2010) because the analytic algorithms of network analysis are relational (Okabe and Sugihara 2012) and“network data by definition includes dependencies among observations” (Laumann etal. 1983). Similarly, syntactic values are meaningful only with reference to a system boundary that a researcher chooses for his or her analysis (Park 2009). Excluding any elements or members of the entire global network system will affect the characteristics, performance and behavior of the measurement result of the local network models. In particular, path-based measures, such as closeness centrality and betweenness centrality, are very sensitive to the boundary effect. Distortion of the results can be induced when links are cut off by an arbitrary boundary and, therefore, are not included in the calculation. The nodes and links closer to the border are less central and peripheral only because of the presence of the boundary. Usually, the boundary effect on the path-based measures is prevalent in all nodes and links in the entire network model (Rheinwalt etal. 2012) and is particularly pronounced for those at the border of the catchment area. Nodes or links near the center of the catchment area tend to have higher closeness centrality (Gil 2017) and betweenness centrality (Chen and Dietrich 2021) compared with those close to the border. Thirdly, boundary effects can also induce a bias on the inference based on such distorted measure results. For example, Krafta (1994) has carried out tests for the
Page 3 of 14 Chenand Dietrich Applied Network Science (2023) 8:60 correlation between different definitions of boundaries and pedestrian movement. Park (2009) has tested the predictability of human movement patterns under various boundary conditions and found that this predictability reaches its maximum at a certain radius from the boundary, which is also an indication of the presence of the size-independent boundary effects on the internal structure. Finally, boundary definition also affects accessibility analysis. Previous studies (Sharkey and Horel 2008)related to public health issues and to nutrition and food accessibility in rural areas have been criticized for not considering resources outside of the study area, even though resources across the boundary may also affect behavior within the area under investigation (Sadler etal. 2011; van Meter etal. 2010). In response to this methodological deficiency, van Meter etal. (2010) and Sadler etal. (2011) have investigated the boundary effect on reaching the retail shops from the locations within the study area, which is typically within an arbitrary administrative boundary. The results show that the boundary effect has led to considerable bias in mis-identification of food desert communities at the border of the study area, even if there is a source of food right across the border. The actual distance of traveling necessary for buying food has also been over-reported. In order to improve the reliability and the consistency of the network analysis results across locations, a number of procedures and practices have been proposed to mitigate the boundary effect. The first approach, the catchment of catchment (Hillier etal. 1993), adds an additional boundary to create a buffer area outside the actual test area. The size of the boundary of the buffer area is larger than that of the catchment area. Network analysis is then carried out for both the catchment and the buffer area. However, the results of the network measure of the buffer area are not included in the analysis because they are distorted by the boundary effect (Gil 2017; Penn etal. 1998). Secondly, instead of one fixed boundary definition, the other mitigation method, the radius-radius analysis (Hillier 1996) or local radius analysis (Gil 2017), applies various boundary conditions by creating multiple circles around the center of the catchment area under investigation. Although the catchment of catchment method and the radius-radius analysis have proved to be successful in mitigating the boundary effect in many empirical studies, the optimal size of the buffer and the radius remains open to further research. Gil (2017) finds that the results of the network centrality analysis are very unstable in small study areas (e.g. on a neighbourhood scale) and suggests that the study area should be embedded in a larger context. However, there remains the question of how large is a large enough context. In other words, the problem of delineating the boundary becomes the problem of deciding the radius or the size of the study area (Joutsiniemi 2005). In response to this open question, Chen and Dietrich (2021) conducted a series of experiments of the size-related boundary effects, i.e. the size effect, on the indicator values. Based on these experiments, they have suggested that, first of all, the average street length can be one of the indicators for determining the size of the catchment area and, secondly, “the size effect on the indicator is not very significant when the size of the catchment area is larger than 4000 × 4000 m2. Therefore, any size larger than 4000 × 4000 m2 would not be necessary” (Chen and Dietrich 2021). Another mitigating method, namely the moving boundary approach, consists in shifting the center of the circular boundaries with fixed size and shape to calculate network
Page 4 of 14 Chenand Dietrich Applied Network Science (2023) 8:60 measures (Penn etal. 1998; Hillier and Penn 2004; Turner 2007; Gil 2017). However, by keeping the same shape and size of the moving network boundaries, Gil (2017) shows that indicators like closeness centrality still vary in different study areas and are particularly sensitive to the shift of network centers. Since the boundary shape and size are identical, the variation of the indicator values can only be explained by the location of the network. In other words, although the size-related boundary effect can be eliminated by adopting the moving boundary approach, the placement-related boundary problem, which is termed placement effect in this research, still exist. To be more specific, placement effect refers to the phenomenon that the variation of the indicator values depends on where the catchment area is retrieved. From this perspective, a very important question is where the center of the catchment area should be. The current research intends to further investigate this placement effect in order to provide more refined guidelines for deciding the location of the network model. Presenting theplacement effect oncloseness centrality The results of Gil’s experiments (2017) show that one of the indicators that could be most influenced by the placement effect is closeness centrality (Cc), which is the reciprocal of the sum of the shortest distance between the chosen node-v and all other nodes in the catchment area. The Cc of node-v can be formally expressed as the following where Cc(v) refers to closeness centrality of the chosen node-v, S(v, u) refers to the length of the shortest distance between the chosen node-v and other nodes, u, and N refers to the total number of nodes in the chosen catchment area. Closeness centrality measures how fast a node exerts influence on all other nodes. For example, if the target is to spread the information in the network, a node with largecloseness centrality means that it is in a position to spread information quickly. Nodes with higher value of closeness centrality can beimportantinfluencers in the network. Closeness centrality is not an absolute value as it may change depending on the location of the selected catchment area in the entire global network. A node in the center of the catchment area has the advantage of having more influence on other nodes and has higher value of closeness centrality than a node located on the borders of the network (Gil 2017). Therefore, the closeness centrality of a chosen node might not necessarily be small in the entire city street network, but it may be small in the selected catchment area only because it is not close to the center of the catchment area. In order to demonstrate the placement effect, the Plaza Luceros in Alicante has been selected as the center point of the study area. The size of the catchment area is 3000 × 3000 m2 and the unit of closeness centrality is 1/km. We have chosen the area size that is larger than the acceptable walking distance for the pedestrian because this allows more space to move the chosen node further away from the center in order to investigate the placement effect. One of our targets is to foster pedestrians in cities and to help to develop walkability, visibility and accessibility of points of interest. Therefore, we would like to investigate a network for pedestrian and Open Street Network (OSM) (1) C c(v)=1/ N u=1 S(v,u )
Page 5 of 14 Chenand Dietrich Applied Network Science (2023) 8:60 can be a source of data to extract existing pedestrian networks in cities. The following highway tags of the OSM are selected to form the network within the catchment area: primary, secondary, tertiary, residential, pedestrian, steps, path and unclassified. Eight catchment areas with different centers were selected and presented in Table1. The center of each catchment area is indicated by a blue center point. The placement effect on the chosen node, which is indicated by the red node in each catchment area,1 will be investigated by examining the changes in the indicator values of the eight selected study areas in Table1. These eight catchment areas differ in the distance between the blue center point and the red chosen node. In catchment area 1, the blue center point and the red chosen node are overlapping with each other. The centers of the catchment areas offset from 50m in catchment area 2 to 1500m in catchment area 8.That is, starting from catchment area 2 to catchment area 8, the blue center point gradually moves 50m, 100m, 200m, 300m, 500m, 1000m and 1500m to the north from the red chosen node. The results in Table 1 show that the closeness centrality of the red chosen node changes from 0.000678 1/km in catchment area 1 to 0.000438 1/km in catchment area 8. This change shows that the closeness centrality of the red chosen node is affected by its distance to the blue center point in all eight catchment areas. Hence, there is a placement effect on the value of the closeness centrality of the red chosen node. Normalization ofthecloseness centrality inthereal network The processes of normalization has been proposed to connect the number of nodes with the indicator (Masucci and Molinero 2016). The current research also applies the normalization procedure and creates an indicator, normalized closeness centrality, CN, so that the different number of nodes in different catchment areas is balanced through the normalization and the fact that the number of nodes changes with the locations of the catchment area is now taken into the consideration. In this section, the results in Table1 are used to explain the normalization of closeness centrality derived from two different indicators: (1) the closeness centrality and (2) the shortest distance between the chosen node and all other nodes. A common way of determining the normalized closeness centrality is multiplying the closeness centrality of the chosen node with the number of nodes in the catchment area. The normalized closeness centrality can be formally expressed as the following. where CN(v) refers to the normalized closeness centrality of node-v, N refers to total number of nodes in the catchment area, Cc(v) refers to the closeness centrality of the chosen node-v. (N-1) refers the number of connections between a chosen node to all other nodes because the chosen node has no (or zero) connection with itself. In the case of a small urban area, in which N is not very large, (N-1)is used. In the case where N is very large, the ‘1’ can be dropped from (N-1). (2) C N (v) = (N − 1) × C c (v) 1 This red node is also the node that is closest to Plaza Luceros.
Page 6 of 14 Chenand Dietrich Applied Network Science (2023) 8:60 Table 1 Catchment areas with different center (blue) nodes and indicator values of (red) node under investigation in the selected catchment areas Distance between red chosen node and blue center node m Closeness centrality of the red chosen node Average distance from all nodes to chosen node Total number of nodes Normalized closeness centrality C(v) N n=1 S(v,n)/(N−1 ) NCN(v) Unit 1/km m 1/km Catchment area 1 0 0.000678 917 1606 1.089 Catchment area 2 50 0.000674 906 1637 1.103 Catchment area 3 100 0.000668 899 1665 1.112 Catchment area 4 200 0.000649 891 1728 1.12 Catchment area 5 300 0.000635 874 1802 1.145 Catchment area 6 500 0.000611 854 1916 1.170
Page 7 of 14 Chenand Dietrich Applied Network Science (2023) 8:60 To illustrate this, we can use catchment area 1 in Table1 as an example. The total number of nodes, N, is 1606 nodes and the closeness centrality of the red chosen node-v, C(v), is 0.000678 1/km. Based on the formula (2), the normalized closeness centrality of the red chosen node-v, CN(v), accounts to (1606−1) × 0.006784 = 1.089 1/km. Idealized network Knowing the value of the normalized closeness centrality of the chosen node-v, CN(v), in the real network, we need to have a reference for different networks to be compared with in order to evaluate the value of CN(v). In this research we propose to use the idealized network to be the reference. For our purposes, the idealized network is defined to a The length of any link between two nodes can be calculated with a simple Pythagoras’ Theorem Table 1 (continued) Distance between red chosen node and blue center node m Closeness centrality of the red chosen node Average distance from all nodes to chosen node Total number of nodes Normalized closeness centrality C(v) N n=1 S ( v,n ) /(N−1 ) NCN(v) Unit 1/km m 1/km Catchment area 7 1000 0.000528 918 2062 1.088 Catchment area 8 1500 0.000438 1190 1919 0.840 Side length of the catchment area equals to 3000m 1500m Center One of the possible shortest paths to the center Fig. 1 The idealized network
Page 8 of 14 Chenand Dietrich Applied Network Science (2023) 8:60 be mathematically ideal, and its nodes are evenly distributed in a quadratic grid. The idealized network has the star-like pattern, as shown in Fig.1. In this section the center node of the idealized network is the chosen node-v, which is used to explain how the closeness centrality of this chosen node-v changes with (1) different number of nodes in the catchment area, N; and (2) the corresponding sum of the shortest distance between the chosen node-v and all other nodes, N u=1 S(v,u ) . Figure1 presents the idealized network. There is a direct connection between the center node and all other nodes. The nodes and links in such a network form a star-like pattern. The shortest distance from any node to the center node is indicated by the yellow arrow. It should be noted that such a network is only favorable to one node, which is the center node in this case. For all other nodes, such a network is not ideal because the sum of the shortest distances between all nodes and any of the non-center node is larger than the sum of distances between the center node and all other nodes. This type of network is often found in real urban networks, such as Place Charles es de Gaulle in Paris or Connaught Place in New Delhi. It is designed to give the central place a high and exceptional importance. Idealized network asthereference forcomparison This section explains (1) the calculation of the normalized closeness centrality in the idealized network and (2) the comparison of the normalized closeness centrality in the idealized and real networks. Calculation ofnormalized closeness centrality ofidealized network In order to calculate the normalized closeness centrality in the idealized network, the first step is to acquire the sum of the shortest distances from the chosen node to all other nodes. Although the side length of the catchment area (which is indicated by the black rectangular in Fig.2) is 3000m, because of the symmetry it is sufficient to focus on just one quadrant of the catchment area, which has the side length of 1500m and is indicated by the red rectangle in Fig.2. Figure3 presents the relationship between the increasing number of links on each side of the 1500m × 1500m quadrant2 and the average distance between the center node and all other nodes in the idealized network. Table2 presents the average distance between the center node and all other nodes in Fig.2, with an increasing number of nodes on each side of the quadrant. As the number of nodes on each side of the catchment area Side length of the catchment area equals to 3000m 1500m One quadrant of catchment area Boundary of the selected catchment area Fig. 2 Relationship between the catchment area and one quadrant of the catchment area 2 In the idealized network, as the number of links on each side increases, the node density of the network also increases.
Page 9 of 14 Chenand Dietrich Applied Network Science (2023) 8:60 increases, the node density of the network also increases. The results show that, with an increasing number of nodes on the side of the quadrant, the average distance between the center node and all other nodes, i.e. N u=1 S(v,u)/(N−1 ) decreases and reaches saturation at 1148m. Applying formula (1) and (2), one can calculate the normalized closeness centrality of the center node in the idealized network, namely Comparing thenormalized closeness centrality intheidealized andthereal network Our next task is to compare the normalized closeness centrality in the idealized and real network and to examine the relationship between the number of nodes and the closeness centrality. The closeness centrality of the chosen node in the eight catchment areas of the real network in Table1 is plotted in Fig.4a. The normalized closeness centrality of the chosen node in the eight catchment areas of the real network is shown by the red (3) C N(v)=(N−1)/ N u=1 S(v,u)=1/1148m=0.8711/ km Average distance between center point to all nodes (m) Number of links on each side of 1500x1500m 2 quadrant Fig. 3 Relationship between number of links on each side of quadrant and average distance between center and all nodes Table 2 Changes of measurements in quadrant with different number of nodes on each side of the quadrant Number of nodes on each side of the quadrant Average distance between the center node and all other nodes N n=1 S ( v,n )/( N−1 ) Sum of shortest possible distance between each node and the central pointa, N n=1 S(v,n ) Total number of nodes, N Normalized closeness centrality, CN m M 1/km 5 1342.176 33,554 25 0.745 15 1212.577 272,830 225 0.825 25 1186.662 741,664 625 0.843 50 1167.228 2,918,069 2500 0.857 100 1157.511 11,575,105 10,000 0.864 150 1154.272 25,971,110 22,500 0.866 200 1152.652 46,106,082 40,000 0.868 300 1151.033 103,592,930 90,000 0.869 500 1149.737 287,434,240 250,000 0.870 750 1149.089 646,362,654 562,500 0.870 1000 1148.765 1,148,765,266 1,000,000 0.871