scieee AI-readable full text Open interactive document viewer

Identification of exposome clusters based on societal, social, built and natural environment – results of the ABCD cohort study

Telkmann, Klaus; Gudi-Mindermann, Helene; Bogers, Rik; Ahrens, Jenny; Tönnies, Justus; van Kamp, Irene; Vrijkotte, Tanja; Bolte, Gabriele

Abstract

Exposome research has seen a recent increase. The conceptual framework of the Social Exposome extends initial concepts by considering the entirety of societal, social, built and natural environmental exposures which are assumed to holistically impact development and health across the lifecourse. The aim of this study is the identification and characterisation of exposome clusters. Additionally, their relevance for mental health is investigated. To this end 2,850 participants aged 11–12 of the Amsterdam Born Children and their Development (ABCD) Cohort Study were analysed. The exposome was characterized by 60 variables representing the societal, social, built and natural environment. Uniform manifold approximation and projection (UMAP) was applied for dimensionality reduction, and subsequently clustering was performed on the retrieved low-dimensional embedding. Mental health symptoms and behaviour related outcomes were assessed by the Strength and Difficulties Questionnaire (SDQ) as well as the Substance Use Risk Profile Scale (SURPS). The results suggest that exposome clusters are mainly driven by contextual socioeconomic and physical characteristics such as neighborhood income and deprivation rather than social characteristics at the individual level. Moreover, prevalence of children’s mental health problems was more prominent within exposome clusters characterized at the contextual level by more deprived neighborhoods and at the individual level by higher prevalence of maternal mental health problems. This exploratory exposome cluster identification emphasized the relevance of socioeconomic neighborhood characteristics, thus structural inequalities.

Full text

Environment International 197 (2025) 109335 Available online 15 February 2025 0160-4120/© 2025 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/bync/4.0/). Full length article Identification of exposome clusters based on societal, social, built and natural environment – results of the ABCD cohort study Klaus Telkmann a,b,* , Helene Gudi-Mindermann a,b , Rik Bogers c , Jenny Ahrens a,b , Justus T¨ onnies a,b , Irene van Kamp c , Tanja Vrijkotte d , Gabriele Bolte a,b a University of Bremen, Institute of Public Health and Nursing Research, Department of Social Epidemiology, 28359 Bremen, Germany b University of Bremen, Health Sciences Bremen 28359 Bremen, Germany c National Institute for Public Health and the Environment, Centre for Sustainability, Environment and Health, 3720 BA Bilthoven, the Netherlands d University of Amsterdam, Amsterdam UMC, Amsterdam Public Health Research Institute, Department of Public and Occupational Health, 1081 BT Amsterdam, the Netherlands ARTICLE INFO Handling Editor: Dr. Xavier Querol Keywords: Exposome Cluster Mental health Social inequalities Environment ABSTRACT Exposome research has seen a recent increase. The conceptual framework of the Social Exposome extends initial concepts by considering the entirety of societal, social, built and natural environmental exposures which are assumed to holistically impact development and health across the lifecourse. The aim of this study is the identification and characterisation of exposome clusters. Additionally, their relevance for mental health is investigated. To this end 2,850 participants aged 11–12 of the Amsterdam Born Children and their Development (ABCD) Cohort Study were analysed. The exposome was characterized by 60 variables representing the societal, social, built and natural environment. Uniform manifold approximation and projection (UMAP) was applied for dimensionality reduction, and subsequently clustering was performed on the retrieved low-dimensional embedding. Mental health symptoms and behaviour related outcomes were assessed by the Strength and Difficulties Questionnaire (SDQ) as well as the Substance Use Risk Profile Scale (SURPS). The results suggest that exposome clusters are mainly driven by contextual socioeconomic and physical characteristics such as neighborhood income and deprivation rather than social characteristics at the individual level. Moreover, prevalence of children’s mental health problems was more prominent within exposome clusters characterized at the contextual level by more deprived neighborhoods and at the individual level by higher prevalence of maternal mental health problems. This exploratory exposome cluster identification emphasized the relevance of socioeconomic neighborhood characteristics, thus structural inequalities. 1. Introduction There is a steady increase in interest in the exposome perspective on the living environment of individuals and populations, stressing the simultaneous interplay of various exposures, including biological, chemical, physical, and social factors and their holistic effect on health (Wild 2012; Vrijheid 2014; Vermeulen et al. 2020; Vineis et al. 2020; Maitre et al. 2021; Moore et al. 2022). In contrast to epidemiological studies analysing only one or a few exposures at a time, this approach acknowledges the complexity and co-occurrence of a multitude of exposures and their collective relevance for etiology of health and disease. Recent studies again emphasized the importance to assess the complex interplay between a diversity of dynamic environmental exposures (Pearson et al. 2022; Tu et al. 2023; Xu et al. 2023). For different conditions and health problems, specific exposomes have been proposed (Arias et al. 2022; Beulens et al. 2022; Chen et al. 2023; Erzin et al. 2023). These exemplary approaches differ regarding the consideration of condition specific aspects of multiple exposures, such as food environment (diabetes), childhood trauma (schizophrenia) or alcohol/tobacco consumption (colorectal cancer). All these approaches have in common that physical environmental exposures often receive greater prominence compared to social environmental exposures. This focus on physical, biological, and chemical exposures has been criticized with well-founded arguments for a systematic consideration of societal, social and psychosocial exposures in exposome research (Juarez et al. 2014; Senier et al. 2017; Neufcourt et al. 2022; Vineis and Barouki 2022). * Corresponding author at: University of Bremen, Institute of Public Health and Nursing Research, Department of Social Epidemiology, 28359 Bremen, Germany. E-mail address: [email protected] (K. Telkmann). Contents lists available at ScienceDirect Environment International journal homepage: www.elsevier.com/locate/envint https://doi.org/10.1016/j.envint.2025.109335 Received 28 October 2024; Received in revised form 17 January 2025; Accepted 14 February 2025 Environment International 197 (2025) 109335 2 Currently, nine comprehensive collaborative research projects on the human exposome are funded by Horizon 2020 of the European Union (Benjdir et al. 2021; Merino Martinez et al. 2021; Vlaanderen et al. 2021; Vrijheid et al. 2021; Laiho et al. 2022; Pronk et al. 2022; Ronkainen et al. 2022; Ronsmans et al. 2022; van Kamp et al. 2022), building the European Human Exposome Network EHEN (EHEN 2024). One of these research projects, Equal-Life, stands out for addressing the public health priority of tackling health inequities in children (van Kamp et al. 2022; Fayet et al. 2024). Within Equal-Life, we have introduced a novel conceptual framework of the Social Exposome as a paradigm for integrating societal, social and physical exposures at individual and contextual levels and to address the previous underrepresentation of societal and social dimensions in exposome research (Gudi-Mindermann et al. 2023). The key focus of the conceptual framework of the Social Exposome is on understanding the underlying mechanisms that translate societal and social exposures into health outcomes. In particular, insights from research on health equity and environmental justice have been incorporated to uncover how social inequities in health emerge, are maintained, and systematically drive health outcomes. To bring the novel framework of the Social Exposome to practice, in this study, we used data from one of the cohorts comprised in Equal-Life, namely the Amsterdam Children and their Development (ABCD) cohort study (van Eijsden et al. 2011). Our primary aim was to identify exposome clusters with an innovative exploratory statistical approach to be able to comprehensively characterize children’s living environment by integrating societal, social, environmental pollution, built and natural environmental indicators at individual and contextual levels with a theoretical foundation. Importantly, exposome clusters were identified solely based on these factors and independent of health-related outcomes. As an illustration for how the exposome might affect health outcomes, we linked the impact of the holistically assessed living environment and social determinants to mental health. Thus, in a second analysis step, associations between exposome clusters and individual mental health outcomes were investigated. 2. Methods 2.1. Study population The ABCD cohort study is a long-term study of some 8,000 children observed at seven phases from pregnancy to adulthood which started in 2003 (see Fig. S1 in the supplement for a timeline of the different phases). Its aim is the analysis of early life factors potentially impacting later health problems (van Eijsden et al. 2011). The analysis in this paper focused on data on mental health problems collected in phase 4 of the study, i.e. in 2015/2016 when the children were 11–12 years old. During the previous phases around 5,000 children were lost to follow-up, leading to a sample size of 3,253 at phase 4. Data were collected by questionnaires which were filled out by the children themselves as well as both parents, if applicable. The response rates for child, mother and father were 93.1 %, 93.2 % and 69.8 % respectively. Since the paternal response rate was comparably low, we restricted the analysis to the mother and child questionnaires. Consequently, 44 participants for whom only paternal information was available were excluded. Additionally, participants with unknown residential address were removed from the data, reducing the sample size to 2,865. Based on the geocodes of the residential address, the dataset was enriched by factors reflecting the built and natural environment (e. g. landuse, air pollution, noise, greenspaces, availability of supermarkets, hospitals, etc) as well as socioeconomic neighborhood characteristics (e. g. mean income, population composition, age structure, area deprivation) (Lakerveld et al. 2020). For many of the environmental variables of interest, data were only available for phase 3 of data collection (2008/ 09) when the children were 5–6 years old. Therefore, the indicators were linked to residential addresses at phase 3. Some neighborhood characteristics collected at baseline (2004) could also be linked to residential addresses from 2011 assuming that these indicators did not undergo any radical changes over the next following years. These include information on population density, income, percentage of men and women, percentage of households with and without children, unemployment and social benefits. 15 participants with information available for less than 50 % of the items were removed from the data, leading to a final sample size of 2,850 observations. Neighborhood characteristics were available on the spatial resolution of 4-digit postcode areas. The Netherlands can be divided into 4,068 unique 4-digit postcode areas, 85 of which are located in the municipality of Amsterdam. Participants’ residential addresses belonged to 471 distinct neighbourhoods defined by 4-digit postcodes across the Netherlands of which 86 % contained only ten or less observations. 61 of these 471 neighbourhoods were located in the municipality of Amsterdam. Out of the 2,850 participants 1,654 were still residing in Amsterdam whereas the rest was located across all parts of the Netherlands. The maximum number of participants belonging to the same area was 118. 2.2. Variable selection 2.2.1. Exposome variables Clustering variables were selected based on the conceptual framework of the Social Exposome according to Gudi-Mindermann et al. (2023) consisting of several layers. For this analysis, each variable was linked to one of the following exposome dimensions. The inner dimensions “actors”, “relational dynamics” and “places” comprise the individual’s social interactions and network. For instance, the dimension “actors” represents the partners of social interaction, including family members, peers, official figures and digital contacts. The dimension “relational dynamics” captures interactions with other individuals, eg. social climate, social capital and medium of interaction. “Places” contains information on where these interactions occur such as the home environment, places for leisure time and learning as well the digital world. This dimension also entails a direct link to social and physical neighborhood characteristics, thus to contextual social, built and natural exposures. These three major dimensions in turn are embedded in the broader societal context represented by the dimensions “socioeconomic circumstances and sociodemographic characteristics” (SEC/SDC) and “systems, institutions and priorities of the political economy”. However, for the latter no information was present in the data. Finally, “cultural values and social norms” have been conceptualized by Gudi-Mindermann et al. (2023) as part of the broader societal context, which can affect all parts of the concept, therefore influencing individual beliefs, attitudes, and behavior as well as social interactions, and systems, institutions, and priorities of the political economy. Variables were excluded from the analysis if they contained more than 20 % missing values. Additionally, a categorical variable for which more than 97.5 % of the observations fell into a single category, was considered too sparse and ruled out as well. Prior to clustering, correlations and associations between all variables were investigated. Namely, Pearson correlation for continuous, Spearman correlation for ordinal and ordinal with continuous, Cramer’s V for categorical (both ordered and unordered) and intraclass correlations for continuous with unordered categorical features were calculated. If variables showed an absolute correlation greater than 0.8, only one representative was selected. This was done since two highly correlated variables contain similar information and, thus, this information would be assigned a larger weight by the dimension reduction algorithm. High correlation coefficients were mainly observed for neighborhood characteristics included in the dimensions SEC/SDC and Places. Median values of absolute correlations within and between each of the exposome dimensions are depicted in Fig. 1. After checking the aforementioned exclusion criteria, a set of 60 out of 96 initially considered variables remained. For example, a variable containing information of the relation between mother and child (biological mother, adoptive child, etc.) was removed since 99 % were K. Telkmann et al. Environment International 197 (2025) 109335 3 biological mothers. The vast majority of excluded variables were strongly correlated geocoded variables. The largest block of correlated of those variables was information on availability of general practitioners, hospitals, cinemas, swimming pools, schools, etc. within a certain distance. The final selection of variables fit the dimensions of the conceptual framework of the Social Exposome as follows: Relational dynamics: 16 variables containing information on social climate such as parenting styles, psychological distress, safety, harmony / disharmony and media use. Parenting styles of the mother were assessed by the parenting styles and dimensions questionnaire (PSDQ, Robinson et al. 1995). PSDQ sum scores are used to distinguish between authoritative, authoritarian and permissive parenting styles based on use of responsiveness and demandingness. The questionnaire consists of 32 items on a 4-point Likert scale. Higher values on the authoritative scale (score from 15 to 60) are considered a positive outcome whereas higher values on the authoritarian (score from 12 to 48) and permissive scale (score from 5 to 20) are considered negative. As a measure of parental stress experienced by the mother the Nijmeegse Ouderlijke Stress Index (NOSIK) was used, which is the Dutch version of the Parental Stress Index (De Brock et al. 1990). This sum score ranges from 11 to 66 and is based on 11 items on a 6-point Likert scale with higher values representing higher levels of stress. Involvement of the mother was assessed by the child care giving and involvement scale mean score (CCIS). It is based on eight statements rated on a 5-point Likert scale with adjustments to make the statements more applicable to 11–12 year olds. Higher values correspond to more involvement (Larsen et al. 2023). Maternal mental health, assumed to have a strong effect on social climate within the family, has been assessed by the Depression, Anxiety and Stress Scale based on 21 questions (DASS-21, Lovibond and Lovibond 1995). Each item was answered with “(almost) never”, “sometimes”, “often” or “very often” which translates to 0–3 points per statement. The three subscales depression, anxiety, stress consist each of 7 of the 21 items and a score could only be determined if at least 4 of the 7 statements were answered. All subscores were multiplied by 2 to achieve comparability to other literature relying on the extended version of the DASS consisting of 42 items. To generate another variable concerning maternal mental health, mothers were asked to rate their life on a scale from 1 to 10. Safety and social cohesion scores were assessed as neighborhood level indicators by the “Leefbaarometer” provided by the Dutch Ministry of the Interior and Kingdom Relations. These scores indicate the deviation from the average across all neighborhoods in the Netherlands. Moreover, a variable counting the number of stressful events (such as divorce, death of a grandparent, etc.) experienced by the child has been included. Actors: One variable describing with whom the child lives (both parents, only mother, only father, etc). Places: 19 variables reflecting the natural and built environment including social infrastructure. These variables included environmental exposures, such as greenspaces (measured by distance and NDVI values), air pollution (in terms of particulate matter with a diameter of 10 µm or less (PM10) and 2.5 µm or less (PM2.5) and nitrogen oxide (NOx)) or traffic noise. Moreover, the degree of urbanization has been assessed within five categories from very urban to rural. Similarly, the built environment has been evaluated in terms of density of buidings, facilities and infrastructure within 300 m buffers around the residential address. Housing stock, public spaces and resources in the neighborhood were assessed by Leefbaarometer scores. Socioeconomic circumstances and sociodemographic characteristics: 22 variables reflecting these dimensions based on individual socioeconomic position such as self-rated family financial situation, ethnicity of the parents, language primarily spoken at home and education of the mother as well as neighborhood level information such as population density, area deprivation, mean neighborhood income, age composition or general social benefits per 1,000 households. Cultural values / social norms: Two variables assessing whether rules concerning media use and smoking and alcohol consumption were imposed in the families. An overview of all exposome variables with information on the assigned exposome dimension, data source, linkage to residential addresses and whether it was used in dimension reduction can be found in Table S1 in the Supplement. 2.2.2. Mental health-related outcomes Children’s mental health symptoms were assessed with the Strength Fig. 1. Heatmap of median values of absolute correlations within and between the dimensions of the Social Exposome framework. Dimension “actors” not included, since only one variable was available for this analysis. K. Telkmann et al. Environment International 197 (2025) 109335 4 and Difficulty Questionnaire (SDQ, Goodman 1997) as well as the Substance Use Risk Profile Scale (SURPS, Woicik et al. 2009). The SDQ contains five subscales: emotional symptoms, hyperactivity / inattention, conduct problems, peer relationship problems and prosocial behavior. For each subscale a score ranging from 0 to 10 was calculated based on 25 statements which could be answered “not true”, “somewhat true” or “true”. An SDQ sum score between 0–40 was calculated by adding all subscale scores except for the positively connotated prosocial behavior. We investigated prevalence of children’s mental health problems based on two SDQ scores: one based on the mother questionnaire and the other self-assessed by the child. All scores were analysed in its raw version and also in a categorical version consisting of bands representing “normal”, “borderline” and “critical” ranges according to Goodman (1997). The SURPS measures the four personality dimensions hopelessness, anxiety, sensitivity/impulsivity and sensation seeking of the child based on 23 questions on a 4-point Likert scale, self-assessed by the child. Each dimension was assigned the mean value of its subitems. 2.3. Clustering To identify clusters, dimensionality reduction of the covariate space by uniform manifold approximation and projection (UMAP, McInnes et al. 2018) was applied as a preprocessing step. Agglomerative nesting then detected clusters on the two-dimensional embedding of the data (Kaufman and Rousseeuw 2009). The involved steps are described in more detail below. UMAP is a novel machine learning technique which embeds the highdimensional covariate space into low-dimensional Euclidean space. This concept borrows and combines ideas from statistics and algebraic topology such as fuzzy sets. For details we refer to McInnes et al. (2018). The advantages of this algorithm lie in its computational efficiency, underlying mathematical justifications and easy to interpret visual output. Another reason to choose UMAP is that standard techniques such as principal component analysis and latent class analysis are not applicable since data is of mixed type (continuous and categorical). UMAP is often compared to t-distributed stochastic neighbor embedding (t-SNE, Hinton et al., 2003). While UMAP performs similarly well in terms of preserving global structure it outperforms t-SNE in terms of computational efficiency (Becht et al. 2019). Moreover, Allaoui et al. (2020) have demonstrated that applying UMAP prior to clustering can substantially improve accuracy of standard clustering algorithms as k-means, densitybased or agglomerative clustering. UMAP requires as an input the k-nearest neighbors for each observation in the high-dimensional covariate space and rearranges these observations in a two-dimensional space such that near observations still appear close in the new embedding. Two parameters are crucial to this algorithm: k is the number of nearest neighbors used with a default setting of 15. In principle, smaller values tend to set focus on the local structure of the data while larger values emphasize the global structure. The parameter min_dist is a threshold on the minimum distance and determines how close points appear in the final layout. Since the objective is clustering we set this parameter close to zero. After checking several combinations of parameters with k in the range of 10–30 and min_dist from 0.001 to 0.1, we set k =15 and min_dist =0.001. No severe visual differences were detected across all setups (up to rotation and symmetry). Since UMAP operates on nearest neighbors, a metric is needed to define distances between observations. As data is of mixed type (categorical, ordinal and numeric), Euclidean distance is not an appropriate metric. A better suited measure is Gower’s distance (Gower 1971) which takes values in the interval [0,1]. For any two m-dimensional vectors of observations xi= (xi1,⋯,xim)and xj= (xj1,⋯,xjm)this metric is defined by d(xi,xj)=∑m k=1 ω k(xik,xjk)⋅dk(xik,xjk) ∑m k=1 ω k(xik,xjk) where dk individually calculates the distance for each single variable. For instance, the distance between observations i and j on a numerical variable is measured as the absolute difference divided by the range. The distance between categorical features is either 0 if they take the same value and 1 else. Ordinal variables are sorted by rank and then treated as numerical as proposed by Kaufman and Rousseeuw (2009). Moreover, the weights ω k are either 1 if neither xik nor xjk are missing and 0 else. This ensures that for each two observations only the average of the distances between pairwise complete variables is calculated. Another advantage of this method is that it avoids a complete case analysis or needs to rely on imputation methods. Clustering was performed on the two-dimensional representation retrieved from UMAP by agglomerative nesting, a hierarchical clustering algorithm. This is an unsupervised task since no target variable for clustering is defined. The goal was to find a structure in the data by grouping similar observations. 2.4. Analysis of exposome clusters concerning children’s mental health problems In a second analysis step, prevalence of children’s mental health problems was calculated for each identified exposome cluster. Since this is an exploratory approach, we refrain from assessing statistical significance of differences in health-related outcomes between clusters by statistical tests such as ANOVA or post-hoc t-tests. We, thereby, follow the recommendation by Greenland et al. (2016). 2.5. Software All analyses were carried out with the statistical programming software R version 4.3.1 (R Core Team 2023). Cramer’s V was calculated using cramersv in the confintr package (Mayer 2023), intraclass correlations by ICCbare in package ICC (Wolak et al. 2012). Gower’s distance was calculated with gower.dist in the StatMatch package (D’Orazio 2022). UMAP relies on the function umap in the package of the same name (Konopka 2023). Clustering was performed with agnes from the cluster package (Maechler et al. 2022). Maps of cluster distribution across Amsterdam were generated with the sf package (Pebesma 2018). 3. Results 3.1. Clustering The study population comprised 2,850 children aged 11–12 years (50.3 % girls, 49.8 % boys), living in neighborhoods with a mean deprivation score of 3.3 on a scale of 1 (least deprived) to 5 (most deprived). 58 % of the children lived in Amsterdam and around 80 % of the mothers were of Dutch ethnicity. At the neighborhood level, the average number of unemployment assistance recipients per 1,000 inhabitants was 27.67 and the mean annual income per recipient was 20,030 € . An application of UMAP to the 60-dimensional covariate space, representing the exposome information, led to an arrangement of participants in two-dimensional Euclidean space. Clusters retrieved by agglomerative nesting based on this embedding are shown in Fig. 2. Recall that the coordinates of this embedding do not resemble any geographical information. Each point represents one participant. Participants with the same or similar values for the majority of exposures (and therefore having a small Gower’s distance) appear closer in this embedding than those with disparate exposomic information. The representation revealed two separate larger agglomerations of participants at the top and bottom left along with a smaller isolated K. Telkmann et al. Environment International 197 (2025) 109335 5 cluster on the right hand side. Agglomerative nesting with complete linkage was applied to retrieve a clustering tree. A correlation of 0.85 between the cophenetic distances of the clustering tree and the Euclidean distances of observations in the embedding indicated high validity of the clustering tree (Sokal and Rohlf 1962). We then chose to cut the tree into eight clusters based on visual inspection and to ensure that sample sizes within clusters were neither too small nor too big. This led to a finer partitioning of the two large amassments into smaller clusters. The numbering of the clusters was determined by their occurrence in the clustering tree produced by agglomerative nesting. 3.2. Description of clusters 3.2.1. Exposome characterisation of the clusters Closer investigation of the distribution of the Social Exposome variables across the two-dimensional embedding revealed that the most discriminating features were neighborhood characteristics such as address density, degree of urbanization, average income per resident, percentage of immigrants and number of general social assistance benefits per 1,000 households. There was also a clear distinction between children living in the municipality of Amsterdam and those who moved to different municipalities although this information was not used in the UMAP algorithm. Moreover, maternal mental health as measured by the DASS-21 questionnaire showed substantial variation across clusters. The distribution of selected variables across the embedding is shown in Fig. 3. The large agglomeration of observations consisting of clusters 2, 3, 6 and 8 at the bottom of the embedding, mostly consisted of participants living in the municipality of Amsterdam (at least 92–99 % per cluster). However, there were clear differences between the individual clusters, in particular a gradual increase in socioeconomic position from left to right. Starting on the right hand side of the bottom agglomeration, cluster 3 was mostly comprised of children living in densely populated, very urban areas (99 %) with lowest area deprivation (average quintile 1.8 ± 1.0) reflected by the highest income per resident on the neighborhood level (18,060 ±2,750 € ), lowest number of general social assistance benefits per 1,000 households (45.4 ±35.4), highest proportions of western immigrants (24.8 ±2.2 %), good infrastructure and good scores for resources, public spaces, housing and population composition. On the other hand, participants in this cluster were also exposed to the highest levels of air pollution (in terms of PM10, PM2.5 and NOx) as well as noise. On the social individual level, this cluster showed low DASS-21 scores for anxiety (1.1 ±2.0) and stress (5.7 ±5.1). Moreover, the NOSI-K sum score for parental stress (18.2 ±6.4) ranked lowest amongst all clusters and the DASS-21 depression score (2.8 ±4.2) ranked second to lowest. Additionally, the majority of mothers rated their family financial situation as good or very good (78 %) and the highest prevalence of mothers with high education (89 % university) was observed. For children in cluster 3, it was much less common to own a mobile phone (8 % as compared to a population average of 18 %). Cluster 2 was very similar to cluster 3 in terms of degree of urbanization, population density, environmental exposures, infrastructure and scores for public spaces and resources. However, it ranked second highest with respect to area deprivation in quintiles (4.4 ±0.9), which was also reflected by the highest rate of unemployment assistance recipients per 1,000 inhabitants (37.3 ±5.2), high numbers of general social assistance benefits per 1000 households (90.6 ±34.5) and a low housing score. Moreover, DASS-21 levels of maternal anxiety (1.7 ± 2.7), depression (3.5 ±4.6) and stress (7.7 ±6.1) and the NOSI-K sum score (20.6 ±7.5) were substantially higher. Prevalence of self-rated family financial situation (65 % good or very good) and mothers’ high education (85 % university) were slightly lower than in cluster 3. Cluster 6 ranked highest in terms of average area deprivation in quintiles (4.9 ±0.3). This was also substantiated by the highest number of general social assistance benefits per 1000 households (141.6 ±38.4) and the lowest income per resident (10,960 ±1,120 € ). This coincided with the self-rated family financial situation (47 % of the mothers rated their situation as good or very good) and the lowest proportion of highly educated mothers compared to other clusters (45 % with university degree). Moreover, the average percentage of non-western immigrants in the neighborhood ranked by far highest among all clusters (51 ±14 Fig. 2. Clusters of the UMAP embedding determined by agglomerative nesting (the coordinates do not resemble geographical information). K. Telkmann et al. Environment International 197 (2025) 109335 6 %) which was underpinned by the highest proportions of non-Dutch ethnicities amongst the fathers (43 %) and the mothers (48 %). In turn, only 64 % felt that they belong to the Dutch community (compared to 88 % of the study population). Noteworthy were also the lowest scores in safety, housing and population composition. Moreover, the highest proportion of single mothers could be found in this cluster (19 %, which was almost 11 % above the population average). Considering maternal mental health, we observed the highest average DASS-21 scores for anxiety (2.4 ±3.6) and depression (3.7 ±4.9) and the second highest NOSI-K sum score (20.7 ±8.3). Cluster 8 also exhibited a higher area deprivation in quintiles (4.2 ± 1.3) with low scores in resources, housing and population composition. This was also reflected by the self-rated family financial situation (63 % rated as good or very good) and comparably lower education status (63 % of mothers with university degree) in relation to the whole study population. Especially noticeable were the high proportion of residents 65 years and older (22 %) and the highest number of old age assistance recipients per 1,000 inhabitants (129.7 ±41.3). This was also reflected by the highest rate of deaths per 1,000 inhabitants (13.3 ±4.2 as opposed to 7.1 ±4.4 on average). However, this cluster ranked second highest for the social cohesion score. Striking were also the highest average DASS-21 score for stress (7.9 ±6.1), NOSI-K (21.4 ±8.6) and the second highest proportion of children not living with both parents (27 %). It is also worth noting that children in cluster 8 were more than twice as likely to have a TV in their bedroom (33 %) as in the overall population (16 %). The large agglomeration at the top left containing clusters 1, 5 and 7 contained mostly children living in less densely populated areas, exposed to less noise and air pollution and generally greener areas as indicated by higher NDVI values. Moreover, the proportion of married couples was considerably higher in these clusters whereas the number of general social assistance benefits per 1,000 households was much lower. This was also captured by low to medium averages of area deprivation in quintiles. Moreover, the vast majority of children not living in the municipality of Amsterdam anymore was found in these clusters. It is also notable, that the DASS-21 scores for anxiety were quite low compared to Fig. 3. Distribution of selected variables across the embedding. Missing values are shown in grey. K. Telkmann et al. Environment International 197 (2025) 109335 7 other clusters. However, there were still substantial differences between these clusters. Cluster 1 contained participants form mixed degrees of urbanization with 50 % living in urban or very urban areas. Moreover, 60 % did not live in Amsterdam anymore. The neighborhoods represented in this cluster were characterized by low average income per resident (12,270 ±1,100 € ), highest proportion of households with children (44 ±9 %), low scores in rating of resources and a high proportion of nonwestern immigrants (25.9 ±11.9 %). The latter was also reflected by the high proportion of parents of non-Dutch ethnicity (25 % mothers, 25 % fathers). Parental stress as measured by NOSI-K sum score was quite low (19.4 ±7.6) and the DASS-21 score for depression was lowest amongst all clusters (2.8 ±4.2). The largest proportion of children not living in Amsterdam was seen in cluster 5 at 96 %. The degree of urbanization was similarly distributed as in cluster 1 with only 42 % of observations living in urban or very urban areas. The neighborhood characteristics stood out through the highest average proportion of married couples (44.2 ±5 %), highest scores for housing, safety and social cohesion and the second highest for population composition, lowest noise pollution (55.2 ±4.2) and most greenspaces (NDVI mean of 0.5 ±0.1), the lowest number of general social assistance benefits per 1,000 households (22 ±22.9) and high mean income per resident (15,050 ±2,670 € ). The latter was backed by the self-rated family financial situation with 79 % good or very good ratings and the high prevalence of highly educated mothers (86 % of the mothers with university degree). Cluster 7 represented the most urban community within the upper larger agglomeration with 93 % living in either urban or very urban areas. Still, 41 % of participants did not live in Amsterdam. The neighborhoods were on average more deprived as in clusters 1 and 5 (mean quintile of 3.4 ±1.2). This was indicated by a low housing score and a high number of old age assistance recipients per 1,000 inhabitants (90.8 ±23.2). The average income per resident in the neighborhood (13,590 ±1,600 € ) and the self-rated financial situation were close to the study population average. The most notable feature of the isolated cluster 4 was that 96 out of 99 participants lived in the neighborhood “Oostelijk Havengebied” (literal translation: “Eastern Port Area”). As the name suggests, this very urban neighborhood with good infrastructure is close to the water and profits from low levels of air pollution and the highest score for public spaces. Moreover, the area deprivation was second lowest (average quintile 2 ±0.3) amongst all clusters. This area also contained the youngest population (proportion of residents 65 years and older was the lowest at 4.5 ±2.3 %). Although the average income per resident in the neighborhood was the second highest across all clusters (15,600 ± 560 € ), the self-rated family financial situation did not reflect this with only 61 % good or very good ratings. While the NOSIK score of parental stress was quite low (18.9 ±6.8), maternal DASS-21 scores for anxiety (1.7 ±3), depression (3.6 ±4.9) and stress (7.7 ±6.4) were all well above the study population average. For a more detailed exposome characterization of the clusters with emphasis on the dimensions of the conceptual framework of the Social Exposome please see Table S2 in the supplementary material. 3.2.2. Spatial characterization of the clusters Since neighborhood characteristics varied substantially between clusters, we analysed the spatial distribution of clusters. We restricted this analysis to the 1,654 participants living in the municipality of Amsterdam. Fig. 4a) shows the number of study participants per neighborhood. 24 of these 85 neighborhoods did not contain any participant. Furthermore, there was no neighborhood in which every participant was assigned to the same cluster. Thus, only the dominating cluster (highest percentage of cluster membership) per neighborhood is shown in Fig. 4b). The least deprived cluster 3 is dominant in the center of Amsterdam whereas the most deprived clusters 6 and 8 are more prominent in the outskirts of the city. Interesting is also cluster 4, which almost completely contained study participants from the neighborhood “Oostelijk Havengebied” near the city center. With 101 participants, this neighborhood is also the most sampled in the study population. 3.3. Distribution of health-related outcomes across the clusters Health-related outcomes were not included in the dimensionality reduction and clustering step. Thus, these measures were analysed only after exposome clusters have been formed. 3.3.1. Mental health problems assessed with Strengths and Difficulties questionnaire The population mean of SDQ sum scores assessed by the child was at 7.6 (±4.6) and at 6.3 (±4.9) when assessed by the mother. Overall, 94 % fell into the category “normal” as proposed by Goodman (1997) for the former and 92 % for the latter. The highest scores were observed in cluster 6 (child: 8.4 ±5.3, mother 7.7 ±5.7), followed by clusters 7 (child: 8.1 ±4.4, mother: 6.4 ±4.8) and 8 (child: 8 ±4.1, mother: 6.9 ±4.5). Consequently, these clusters also featured the lowest proportion of children in the normal bound. In particular, for cluster 6 we observe 91 % (child) or 86 % (mother) in normal range. However, for cluster 7 (child: 94 %, mother: 92 %) and cluster 8 (child: 96 %, mother: 90 %) these numbers were still close to the population average. Lowest SDQ sum score means were observed in clusters 3 (child: 7 ±4.1 and 97 % in normal range, mother: 5.2 ±4.1 and 96 % in normal range) and 4 (child: 7 ±4.3 and 95 % in normal range, mother: 5.1 ±4.2 and 95 % in normal range). An overview on how SDQ mean scores across clusters compared to the study population mean is given in Fig. 5. For the SDQ subscale conduct problems, population means of 1.3 ± 1.3 and 94 % in normal range or 0.8 ±1.2 and 91 % in normal range were observed for the child and mother, respectively. The highest average was found in cluster 6 (child: 1.6 ±1.4 and 91 % in normal range, mother: 1.1 ±1.4 and 83 % in normal range). As for the sum Fig. 4. A) study participants per neighborhood in amsterdam b) dominating cluster per neighborhood as defined by the highest percentage of cluster membership. the 24 grey areas (labelled na) did not contain any participant. K. Telkmann et al. Environment International 197 (2025) 109335 8 scores, cluster 3 performed very well with below average scores (child: 1.2 ±1.2 and 96 % in normal range, mother: 0.6 ±1 and 95 % in normal range). Emotional symptoms showed a discrepancy between the selfcompleted and mother-completed versions. The scores averaged across the population at 1.9 (±1.8) with 95 % in normal range and 1.7 (±1.9) with 85 % in normal range for the child and mother, respectively. Assessed by the child, well above average scores were seen in clusters 6 (2.1 ±2), 7 (2.1 ±1.8), and 8 (2 ±1.8). The proportion of children in normal range was still close to average within these three clusters with cluster 6 having the smallest share at 92 %. The other clusters scored average or slightly below average. Regarding the mother’s assessment, cluster 6 still scored highest at 2.1 (±2.1) with only 78 % in normal range, whereas clusters 7 and 8 were average or slightly above average. Scores well below average were observed in clusters 3 (1.3 ±1.6, 90 % in normal range) and 4 (1.4 ±1.6, 92 % in normal range). The hyperactivity / inattention subscale scores averaged at 3.4 (± 2.4) and 82 % in normal range (child) and 2.7 (±2.5) and 85 % in normal range (mother). Highest cluster means were observed in clusters 7 (child: 3.7 ±2.4, 79 % normal, mother: 2.9 ±2.5, 85 % normal) and 8 (child: 3.8 ±2.4, 78 % normal, mother: 3.1 ±2.6, 80 % normal). Cluster 6 matched the average when assessed by the child, but was second to worst by the mother’s rating (3 ±2.5, 82 % normal). Best scores were achieved in cluster 3 assessed by the child (3 ±2.2 and 86 % normal) and cluster 4 assessed by the mother (2.1 ±2.1 and 93 % normal). SDQ mean scores for peer relationship problems averaged at 1 ±1.3 and 94 % normal (child) and 1 ±1.46 and 86 % normal (mother). Cluster 6 ranked worst for both assessments (child: 1.4 ±1.4, mother: 1.5 ±1.7) and also showed considerably lower proportions of children in normal range (child: 91 %, mother: 76 %). On the other hand, clusters 3 and 4 showed higher proportions of children in normal range at more than 95 % for the self-completed version and above 91 % for the mothercompleted version. No severe differences could be seen for the SDQ subscale prosocial behavior averaging at 8.5 (±1.4) and 96 % normal as assessed by the child and 8.6 (±1.6) and 95 % normal as assessed by the mother. Neither for child’s assessment nor for mother’s assessment values with strong deviations from the study population mean stood out. Differences between boys and girls were most pronounced in SDQ sum scores assessed by the mother: boys on average score 1.4 points higher than girls. This gender difference was much less severe by the children’s ratings at 0.3. For the subscales conduct problems (child: 0.3, mother: 0.2) and hyperactivity / inattention (child: 0.5, mother: 1.1) boys scored higher on average compared to girls where the opposite is true for emotional symptoms (child: −0.5, mother: −0.1) and prosocial behavior (child: −0.5, mother: −0.5). No notable gender differences were observed for the subscale peer relationship problems (child: 0, mother: 0.1). Cluster 4 showed lower SDQ sum scores for boys compared to girls (gender difference assessed by child: −1.8, mother: −0.6) as opposed to almost all other clusters (exceptions are cluster 5 and 8 in the child questionnaire at 0.4 and 0.1, respectively). On the other hand, sum scores for boys in clusters 1, 3, and 6 were substantially higher than those for girls (child: 0.8 to 0.9, mother: 2 to 2.6). In particular, the percentage of girls falling into normal range in cluster 6 was 11 % higher than for boys. Interestingly, the fact that boys in cluster 4 scored lower than girls applies to all subscales except hyperactivity / inattention assessed by the mother (here, boys score on average 0.3 higher than girls). 3.3.2. Mental health problems assessed with Substance use risk Profile scale SURPS scores were investigated at four subscales. Overall, there were no substantial differences with cluster means only slightly deviating from the population means. The population average for the subscale anxiety sensitivity was 1.9 (±0.6) and the strongest deviation from this value was seen in cluster 6 at 2.1 (±0.7). All other cluster means fluctuated only slightly across the sample mean. In case of hopelessness, cluster means barely deviated from the population mean of 1.3 (±0.4). The highest value was again found in cluster 6 at 1.3 (±0.4). Similarly, for impulsivity cluster 6 deviated at 2 (±0.7) the most from the population mean of 1.9 (±0.6). Noteworthy was also cluster 3 with the lowest average of 1.8 (±0.6). For the dimension sensation seeking, we observed a population mean of 2.7 (±0.7). The largest deviations could be seen in cluster 7 at 2.8 (±0.6) and in the other direction in cluster 4 at 2.7 (±0.7). A heatmap depicting cluster means in comparison to population means is shown in Fig. 6. Fig. 5. Heatmap of standardized mean differences to the study population of SDQ cluster mean scores. Note that for the subscale “prosocial behavior” higher values are regarded as a positive outcome. Consequently, this subscale is not included in the sumscore. K. Telkmann et al. Environment International 197 (2025) 109335 9 For the subscales anxiety sensitivity and hopelessness girls scored on average 0.2 and 0.1 points higher than boys, respectively. Whereas for impulsivity and sensation seeking boys scored 0.1 and 0.3, respectively, higher than girls. No severe deviations from these differences could be observed in the clusters with three exceptions: girls in cluster 4 scored on average 0.5 and 0.2 points higher for anxiety sensitivity and impulsivity than boys, respectively. Additionally, girls in cluster 8 scored 0.1 points higher on average than boys. Descriptive statistics of all health-related outcomes for each cluster can be found in Table S3 in the supplementary material. 4. Discussion Our primary aim was the application of novel statistical methods to identify exposome clusters based on a variety of social and societal characteristics as well as factors of the built and natural environment. In summary, the application of UMAP to dimensionality reduction and clustering resulted in clearly distinct exposome clusters. Cluster formation was mainly driven by contextual socioeconomic and physical neighborhood characteristics, related to the dimensions “socioeconomic circumstances and sociodemographic characteristics” and “places” of the conceptual framework of the Social Exposome (Gudi-Mindermann et al. 2023). For instance, clusters discriminated clearly between areas exhibiting different degrees of urbanization and, in particular, between those living in Amsterdam and those who moved away. A finer distinction could be made in conjunction with socioeconomic neighborhood characteristics. Thus, striking differences in area deprivation were present across clusters expressed through average income per resident, number of general social assistance benefits or ratings of resources, housing, public spaces or social cohesion. Moreover, cluster differences were also present in terms of maternal mental health, which is considered as one crucial aspect of the dimension “relational dynamics” of the domain of social interactions within the Social Exposome concept. A subsequent analysis of the exposome clusters showed considerable differences in children’s mental health. In general, the more deprived the neighborhoods and the higher the prevalence of maternal mental health problems, the higher the prevalence of children with mental health issues, indicated by falling into the borderline and critical range of the SDQ. This is in line with previous studies on risk factors for children’s mental health problems (Ravens-Sieberer et al. 2007; Becker et al. 2015). The SDQ is a widely applied screening instrument in epidemiological studies and has been shown to be valid for detection of emotional and behavioral problems in children/adolescents, especially when taking self-report and parental report into account (Warnick et al. 2008; Theunissen et al. 2019). However, no remarkable prevalence differences in personality traits assessed with the SURPS were observed between the exposome clusters. This might be due to the relatively young age of the study participants. SURPS personality profiles have been shown to be associated with substance use behavior in a study of Dutch adolescents aged 11–15 (Malmberg et al. 2010). 4.1. Relation to other statistical approaches in exposome research With our analysis strategy we identified exposome clusters based on similarity of exposures and then descriptively analysed distributions of mental health-related outcomes across those clusters. In contrast to methods directly linking the exposome to a health-related outcome and then isolating a smaller set of risk factors while adjusting for confounders, this method takes into account the entirety of exposures. To our knowledge, this was the first time in epidemiological research that this innovative approach was used to identify clusters based on exposures of the societal, social, built and natural environment. In particular, the application of UMAP as a dimension reduction technique is a novelty for this kind of research. So far, UMAP has been mostly applied for large scale data such as single-cell data (Becht et al. 2019) or gene and protein data (Dorrity et al. 2020; Lunjani et al. 2024), or to detect urban configuration types (Iungman et al. 2024). Up to now, there is, to our knowledge, only one epidemiological application by Bej et al. (2022) utilizing UMAP to detect clusters based on sociodemographic and socioeconomic characteristics and lifestyle habits to describe heterogeneous Type-2 diabetes subpopulations. However, in our study we applied a unified metric for mixed data, while Bej et al. (2022) employed UMAP on three sets of variables depending on their type (numeric, categorical, ordinal) and only then combined the results. A review of statistical methods by Santos et al. (2020) provided an overview of issues to address and feasible algorithms within the specific context of birth cohort studies. Similarly, Stafoggia et al. (2017) presented a series of applicable methods to analyse multiple environmental exposures. Our approach is in line with the recommended steps of correlation analysis, dimension reduction, and clustering. For instance, Robinson et al. (2015) and Tamayo-Uria et al. (2019) pursued an approach based on investigation of correlations between several families of exposures and then applied principal component analysis. Since our data contained variables of mixed type we opted for a different solution for dimension reduction. Since its introduction, exposome research has seen a variety of different approaches to uncover associations between exposures and health outcomes. Many authors consider methods linking a large set of exposures directly to a specific pre-defined health outcome. In analogy to genome-wide association studies, exposome-wide association studies (ExWAS) have been applied to link the most relevant parts of the exposome to several health outcomes such as BMI in adults (Haddad et al. 2022), birth weight (Nieuwenhuijsen et al. 2019), hypertensive disorders (Hu et al. 2020) and COVID-19 mortality (Hu et al. 2021), among others. This method has also been incorporated into cluster approaches. Guillien et al. (2021, 2022) applied ExWAS to reduce the set of exposures and subsequently define exposome profiles and investigate Fig. 6. Heatmap of standardized mean differences to the study population of SURPS cluster mean scores. K. Telkmann et al.