scieee AI-readable full text Open interactive document viewer

Integrating Artificial Intelligence into Life Cycle Assessment: A Framework for Balancing Automation and Human Expertise

Nwagwu, Chibuikem; Ogorodnyk, Olga; Sølvsberg, Endre; Eleftheriadis, Ragnhild Johnsen; Meskers, Christina

Abstract

Quantifying products’ environmental impacts is essential as policymakers and customers push for transparency and accountability in product sustainability. The applications of machine learning (ML) have extended to many fields, including environmental science. Several studies demonstrate the application of ML in life cycle assessment (LCA), but a gap remains in synthesizing the insights into a unique, deployable framework for rapid and (semi-) automated LCA of products and industrial processes. This is due to outdated, unlabeled, sparse, and irrelevant data and the industry’s poor maturity for AI and LCA. Based on a rapid literature review of the current applications of AI in LCA studies, we find that supervised learning algorithms are most preferred, primarily for data collection and inventory analysis. We evaluated a suite of AI tools for future application in LCA and concluded that the most benefit will be from using large language models (LLMs) and generative algorithms to improve the speed and accuracy of environmental impact assessments. Our developed framework—an AI integration architecture for LCA studies—ensures that human insight and control are retained. The work provides valuable guidance to industry-based sustainability practitioners on the importance of data quality, AI tool selection, cross-domain expertise, and collaboration.

Full text

Vol.:(0123456789) Journal of Sustainable Metallurgy https://doi.org/10.1007/s40831-025-01305-x THEMATIC SECTION: HONORING DIRAN APELIAN'S LEADERSHIP INSUSTAINABLE METALLURGY Integrating Artificial Intelligence intoLife Cycle Assessment: AFramework forBalancing Automation andHuman Expertise ChibuikemC.Nwagwu1 · OlgaOgorodnyk1 · EndreSølvsberg1 · RagnhildJ.Eleftheriadis1 · ChristinaMeskers1 Received: 25 June 2025 / Accepted: 30 September 2025 © The Author(s) 2025 Abstract Quantifying products’ environmental impacts is essential as policymakers and customers push for transparency and accountability in product sustainability. The applications of machine learning (ML) have extended to many fields, including environmental science. Several studies demonstrate the application of ML in life cycle assessment (LCA), but a gap remains in synthesizing the insights into a unique, deployable framework for rapid and (semi-) automated LCA of products and industrial processes. This is due to outdated, unlabeled, sparse, and irrelevant data and the industry’s poor maturity for AI and LCA. Based on a rapid literature review of the current applications of AI in LCA studies, we find that supervised learning algorithms are most preferred, primarily for data collection and inventory analysis. We evaluated a suite of AI tools for future application in LCA and concluded that the most benefit will be from using large language models (LLMs) and generative algorithms to improve the speed and accuracy of environmental impact assessments. Our developed framework—an AI integration architecture for LCA studies—ensures that human insight and control are retained. The work provides valuable guidance to industry-based sustainability practitioners on the importance of data quality, AI tool selection, cross-domain expertise, and collaboration. The contributing editor for this article was U. Pal. * Chibuikem C. Nwagwu chibuikem.nw[email protected] 1 SINTEF Industry, Department ofManufacturing, Trondheim, Norway Journal of Sustainable Metallurgy Graphical Abstract Keywords Data· Algorithms· Environmental impact assessment· Metal industry· Machine learning Introduction Since the Industrial Revolution, the metal-producing and manufacturing industry has taken a prime position in the European economy. It is a major employer and driver of innovation, global trade networks, and economic prosperity. This resource-intensive industry’s development has also led to significant environmental impacts, including about 25% of direct energy-related carbon emissions in 2022 [1] and water, land and resource use and biodiversity impacts. There are also indirect impacts from using manufactured products through design choices and innovation. The European metal-producing and manufacturing industry, therefore, faces a triple challenge of (1) reducing environmental footprint via a circular economy and use of renewable energy, (2) increasing competitiveness and resilience, e.g., via productivity, and (3) collaborating within and across value chains. In response to these challenges, recent European Union (EU) policies like the Eco-design for Sustainable Products Regulation [2], Circular Economy Action Plan [3], Clean Industrial Deal [4], and public procurement legislations have been developed to put pressure on the industry to transition and be fast, agile, and clean. Materials, products, processes, and supply chains are complex and prone to rapid changes, and forecasts have high uncertainties. Therefore, large amounts of good-quality data and dynamic assessment tools are needed that reflect the complexities and remain within technological feasibility and reliability [5]. Thus, industry and society need relevant, complete, accurate, timely, and detailed information on products, technologies, and lifestyles. The industry has therefore adopted data-driven and digital solutions to achieve this need for speed through increased investment in artificial intelligence (AI) to enhance data collection, management, processing, and interpretation possibilities [6]. This has resulted in high interest in the right way of using these tools and culminated in the EU AI Act [7]. Additionally, these digitalization efforts have increased sensor usage and advanced data architecture (e.g., data spaces) for traceability and interoperability. Yet many companies do not fully understand how to leverage the possibilities these technologies offer. As a result, most data collected today is not of good quality as they are not comparable, consistent, Journal of Sustainable Metallurgy accessible, transparent, nonpartisan, or relevant [8]. Additionally, the tools and data themselves, however, good, do not guarantee adequate baseline measurement of environmental sustainability and the impact of changes over time. Life cycle assessment (LCA) done according to ISO 14044 is the preferred tool for environmental impact assessment [5, 9]. To ensure fairness and accountability across players, practitioners use this standard methodology for real problem-solving, comparing alternatives, and decision-making while avoiding burden shifting and greenwashing. The LCA method, databases, and software originated when the world turned slower: before the renewable energy transition, fast-moving consumer goods, and circular economy transition. Today, environmental product declarations (EPDs) are required for all products, and data on energy, material, process, product and systems change rapidly, resulting in (1) a huge competence shortage within the industry, (2) a need for further standardization and automation of LCA methodology, and (3) up-to-date data reflecting industrial reality and complexity. A plethora of software with embedded life cycle impact assessment (LCIA) methods is available [10, 11]. However, (product) LCA requires background inventory data from databases embedded in this software. Most of these databases, however, contain outdated, incomplete, or generic datasets. Thus, they do not necessarily represent technological advancements, the specific sector context, and/ or geographical location. More concerning, there is so much missing data at the foreground or process level [9], especially company-specific data, that replacing aggregates/proxy data is impossible. LCA tools also struggle to incorporate uncertainty and the time dimension in the model design [12]; the complexity of supply chains; and are mostly static in their scope and predictions (Fig.1). This is even more critical for LCAs of a prospective nature that model future technologies and are used in decision-making. To address the data-related challenges, LCA practitioners have over the past 2–3 decades adapted several artificial intelligence (AI) algorithms, developed (Python-based) process simulation models and digital twins which allow seamless integration of process and value chain models with environmental impact assessment, for example, for metal-based products [13, 14]. Some advantages of these approaches include the ability to model complex non-linear relationships, generalize from historical data, retrain models in case of data drift, and address gaps between theory and practical implementation of LCA results. The large industry investment in digitalization should motivate the search for solutions to address the outlined LCA challenges and knowledge gaps. One such solution could be a framework that guides and bridges between AI/ data scientists, industry domain experts and LCA practitioners. To the best of our knowledge, there is no framework for using machine learning across the four LCA stages. Hence, this study develops such a framework. The specific research questions this study addresses, within the context of the European manufacturing industry, are as follows: R1: What is the trend in AI use for LCA today and into the future? R2: What AI algorithms best fit the needs of the different LCA phases? R3: How can companies adopt AI-supported LCA responsibly and ensure human involvement? First, the authors briefly introduce LCA and AI to create a common understanding of these vastly different fields. Thereafter, to address the R1, the authors conduct a rapid literature Fig. 1 Knowledge gaps in the different stages of life cycle assessment Journal of Sustainable Metallurgy review of AI usage in any LCA phase. Based on the literature findings and the authors’ experience/expertise, the authors assess the suitability of current and upcoming AI tools in LCAs. R2 is addressed by developing a conceptual framework for using AI in LCA based on established LCA and AI methodology, and findings from R1. Finally, the authors provide recommendations for ensuring that the choice of AI for LCAs does not neglect the role of human experts. Methodology Life Cycle Assessment (LCA) Life Cycle Assessment is a tool and method for estimating the environmental impacts of products and services across their entire life cycle or specific periods within their life cycle. The ISO 14040 standard defines four key phases of an LCA: goal and scope definition, life cycle inventory, life cycle impact assessment, and interpretation. Data Requirements forLCA Given the data’s strong impact on the LCA results, ISO 14044 sets out strict requirements for data quality, which must be reported along with the study’s results. 1. The data should be complete, including all necessary data to achieve the LCA goals and scope defined. 2. The data must accurately reflect the geographical, temporal, and technological characteristics of the study context. 3. The data sources, methods, and measurement tools must be reliable (i.e., scientifically and technically sound). 4. The data and methodological choices must be consistent throughout the study, including for scenario development. 5. The LCA results must also be clear and comprehensible to the reader, helping them draw meaningful conclusions. 6. The study must be reproducible and achieve the same results. 7. The LCA practitioner must declare the level of detail and exactness in the data, how uncertain and sensitive the data used are themselves and how these affect the LCA results [15]. It is, therefore, essential for the success of any machine learning algorithm applied in LCA to ensure the combination of high accuracyand high precision while maintaining a high level of transparency [5, 16]. Artificial Intelligence AI includes a variety of methods and techniques for creating intelligent systems. These methods can be broadly classified into six main categories, as shown in Fig.2. This classification is neither extensive nor the only suitable one, as some of the methods can be placed in more than one category. Fig. 2 Different categories of AI and some example algorithms within them Journal of Sustainable Metallurgy Machine learning (ML) allows systems to learn patterns from data and make predictions or decisions. It includes supervised learning (trains models on labeled data to predict outcomes), unsupervised learning (finds hidden patterns or structures in unlabeled data), and reinforcement learning (models learn to make decisions through rewards or penalties in dynamic environments). Artificial neural networks (ANNs) are one of the ML methods that can be included in several categories, e.g., supervised and unsupervised, as they can be trained on both labeled (for classification or regression) and unlabeled (for clustering) data. These models are mimicking the human brain by processing data in layers. Please note that even though shallow (1–2 hidden layers) and deep (multiple hidden layers) neural networks are not entirely different methods, we choose to place shallow neural network architectures under the ML category and deep neural network architectures as a separate AI methods category. Rule-based and symbolic AI rely on predefined rules and logical inference to solve problems using knowledge based on facts and rules to emulate human decision-making or model logical reasoning [17]. Deep Learning (DL) is an advanced subfield of AI, which includes deep neural networks such as Recurrent Neural Networks (RNNs), Convolutional Neural Networks (CNNs), and transformers (architectures behind generative AI and Large Language Models (LLMs) like BERT or GPT that excel in natural language processing tasks). Evolutionary and Nature-Inspired Computing like Genetic Algorithms (GAs) and Swarm Intelligence are AI methods inspired by natural processes, emphasizing optimization and adaptability. Probabilistic and Statistical Models like the Bayesian network and Markov models are AI techniques grounded in probability and statistics for uncertainty handling based on probabilistic relationships among variables. Hybrid Systems combine multiple methods for better results in complex tasks. How Does AI Modeling Work? To create an AI model for decision support (Fig. 3), understanding the problem (i.e., the subject area, available data, and data needs) is the first step. This can be done by consulting the process/data experts, as they might provide useful insights both for the data analysis and model development. All unexplained observations and assumptions related to the problem must be clarified as soon as possible to ensure that they are based on facts and do not hinder proper understanding of the concepts and processes being modeled. The problem understanding phase can also reveal if AI is the best method to solve the task or if other tools should be considered. The next step is to obtain the necessary data by installing sensors and establishing proper data acquisition processes. When the data are gathered, it needs to go through preprocessing to be in a format that is easy to handle and analyze. In some cases, simple data exploration can provide a lot of useful information about the problem. Therefore, understanding the data and the related analysis results is crucial. After the data are explored and understood, it is time to choose a suitable AI method based on the amount, structure, and nature of the data. Some methods require feature selection or construction to reduce overfitting and noise in the data, while some have it “built-in” by design. Next, the model is trained, tested and validated, and if successful, it can be used for prediction or forecasting. It is also important to note that the movement between these steps is not linear and might require iteration, especially if the model's performance is unsatisfactory. Data Requirements forAI Modeling AI requires a lot of data to develop meaningful models. Often, available data provide only a limited description of reality, and the model risks drawing incorrect conclusions if the full context is not accounted for. It is essential to recognize that ensuring accurate and relevant data is often the most time-consuming but also the most critical phase in developing useful and meaningful models. These data must be of high quality because a model trained on data with errors will have unsatisfactory performance by default. Moreover, it is necessary to ensure that all the important parameters are included in the dataset when developing a predictive model to avoid bias and inconsistency. These Fig. 3 The steps for AI model development, from understanding the problem to the optimal solution Journal of Sustainable Metallurgy parameters must be checked for autocorrelation and other necessary assumptions for the respective models. Fit forPurpose Beyond the individual data requirements for AI and LCA models detailed above, there are overlapping requirements to fulfill for AI models to be fit for the purpose for which they are created. Therefore, in addition to AI-related data requirements like interpretability, training speed, and robustness to outliers [18], AI models used in LCAs must satisfy the same data requirements as any other LCA and meet ISO 14044. The authors group these overlapping data-related requirements using simplified parameters [19] to address how different AI algorithms fit within the LCA field, as shown in Table2. To ensure that the LCA is complete and fully represents all aspects of the modeled reality, the chosen algorithm must be stable to handle missing information and data without deviating from the LCA’s goal and scope. For sensitivity and uncertainty analysis in the LCA, the algorithm should be able to identify and handle outliers. This affects how the AI model’s predictions change in response to changes in the input data. To ensure the LCA results are reproducible, it is important to choose an algorithm that has the same performance under the same conditions and uses the same input data. In other words, the algorithm should not only be internally consistent but also be reproduced externally without being sensitive to data shuffling and variations in training data. To ensure the LCA results are transparent and verifiable, the algorithm should be interpretable and/or explainable. To increase the trust and reliability in the LCA results, the algorithm should perform defined LCA tasks dependably across different datasets and be robust to noise. To ensure comparability and effective decision-making from LCA results, the algorithm should have good overall predictive power to produce precise results. Finally, the energy needs and environmental impact of using this algorithm should be minimal. Literature Review Methodology This study employed a rapid literature review using abstract proceedings from the ISIE 2023 conference [20] as a starting point, and then snowballing using the Google search engine and cited reference search. Thereafter, a structured literature review search was conducted on Web of Science using the keywords “machine learning,” “artificial intelligence,” and “life cycle.” This returned 134 results, out of which 55 were relevant based on their titles (i.e., after duplicates and studies not having both ML and LCA knitted together were removed). ML is specifically chosen among other AI branches based on its central role in LCA compared to other branches of AI, especially within manufacturing [21, 22]. Results andDiscussion Since no single algorithm exists that handles every type of dataset for all kinds of applications, it is crucial to select algorithms for use in LCAs appropriately and systematically [6]. LCAs are data-reliant, hence, most of today’s applications for ML in LCAs are on data collection and processing using supervised and unsupervised algorithms. There are also a couple of applications of deep learning, reinforcement learning, and generative AI. The list of AI algorithms and their applications provided here is not exhaustive due to the chosen literature search methodology. The algorithms, their applications in LCA and pros and cons are summarized in Table1. Where Is AI Applied Today forLCAs? The probability density function (PDF) is often the go-to algorithm due to its effectiveness in representing uncertainty in data. However, if the product system is very complex and produces little or noisy data, PDFs may struggle or be computationally expensive [23]. Although probabilistic modeling like Bayesian networks may be computationally heavy, they offer an effective way of handling data uncertainty with high interpretability. Since they can incorporate prior knowledge while learning from new evidence, they are useful for modeling complex dependencies and causal relationships [21, 24]. Classification algorithms like Support Vector Machines (SVMs) and decision tree-based ensemble models like Random Forest and Gradient Boosting have been used for inventory collection and data gap filling in the energy sector [25, 26], the chemical sector [27, 28], and the agricultural sector [29–32]. They have also been used for sensitivity and uncertainty analysis [29, 32]. These classification algorithms handle missing numerical and categorical data with high accuracy, especially for structured data with a high number of unique values. Their use of regularization techniques helps them to avoid overfitting the models, although this implies that they can sometimes be computationally heavy and require careful tuning of hyperparameters. Once trained, the prediction is fast. However, their interpretability is low since they do not have an explicit mathematical function [28]. Other classification algorithms, like extreme gradient boosting (XGBoost) have also been applied in similar sectors as above [33, 34] and in construction [35, 36]. They have also been used to predict the characterization Journal of Sustainable Metallurgy Table 1 Overview of AI algorithms, their applications, and limitations in LCA today Algorithm/method Method details Application in LCA Pros Cons Source Generative AI (e.g., LLMs) Large deep learning model Process databases, material safety data sheets, and published LCA results to develop a comprehensive LCI database, extract emissions impact factors for the raw material extraction and manufacturing process phases from any LCI data source, identify environmental hotspots, predicting emission factors Effective at extracting relevant information from unstructured data, can process and understand vast amounts of text data, can generate human-like text as well as realistic synthetic data, useful for augmenting datasets and training models in lowdata scenarios Computationally intensive, may produce bias or incorrect outputs due to training data limitations, quality of generated content may vary widely, low interpretability Tian Gong (n.d.), (Luo etal. (2023), Venugopal and Olivetti, (2023), Cornago etal. (2023) Generative adversarial network Large deep learning model Allocation and estimation of flows and releases by predicting the process yield Zargar etal. (2022) Similarity-based models Machine learning, supervised Using a selection proxy to fill missing LCI dataset based on proximity between nodes in the network Intuitive and easy to understand, very effective for small datasets, requires little computational power The performance depends on the similarity metric chosen, may struggle with high-dimensional data and generalization Meron etal. (2020), (Hou etal. (2018) Decision trees Machine learning, supervised Estimating steam consumption, the highest energy utility consumption in chemical batch plants Simple and easy to interpret, does not require assumptions about data distribution, can handle both numerical and categorical data well Prone to overfitting, especially with deep trees, sensitive to small changes in data Pereira etal. (2018) Decision tree-based ensemble ML (e.g., XGBoost, Random Forest) Machine learning, supervised Estimation of missing/ duplicate ecoinvent data, forecasting future municipal solid waste values, estimating LCI values for energy use intensity (EUI) for commercial office buildings, estimating LCI values for energy use intensity (EUI) for commercial office buildings High accuracy, especially with structured data, can handle missing data and categorical variables well, reduces overfitting with regularization Computationally heavy for large datasets, not easy to interpret, requires careful tuning of hyperparameters, prone to overfitting if not regularized Zhao etal. (2021), (Zhang etal. (2022), Deng etal. (2018), Deng etal. (2018) Journal of Sustainable Metallurgy Table 1 (continued) Algorithm/method Method details Application in LCA Pros Cons Source Artificial neural networks (ANNs) Machine learning, can be both supervised and unsupervised Estimating the pyrolysis yield and gas composition to predict the activation yield, predicting LCI values of energy consumption, solid waste, greenhouse effect, ozone depletion, acidification, eutrophication, and winter and summer smog, grouping existing products according to their environmental characteristics and mapping product attributes into environmental impact driver (EID) index to obtain the LCA results for newly designed products, LCA surrogate model, estimating characterization factors (i.e., effect factor of climatic conditions on respiratory disease incidence and hospital admittance) Can model complex nonlinear relationships, highly flexible and can be used in different applications, can learn directly from raw data Requires large training data, prone to overfitting if not properly regularized, not easy to interpret and understand model behavior Liao etal. (2020), Sousa etal. (2000), Park and Seo (2003), Sousa (2008), Bartie etal. (2021), de Souza Tadano etal. (2016), Araujo etal. (2020) Uncertainty analysis (e.g., Monte Carlo simulation using various PDFs) Can be placed in several AI methods categories: ML (both supervised and unsupervised), Probabilistic and Statistical modeling (Bayesian Networks, Markov models), etc. Estimating steam consumption, the highest energy utility consumption in chemical batch plants Useful for measuring data uncertainty and variability,useful for risk assessment, clear and interpretable Data must fulfill certain assumptions, May be computationally expensive for complex distributions Pereira etal. (2018) Bayesian networks Probabilistic and Statistical modeling Combining publicly available data sources to obtain probability-based information about missing building stock LCI data Provides a probabilistic framework for reasoning under uncertainty, can incorporate prior knowledge and learn from new evidence, useful for modeling complex dependencies and causal relationships Computationally heavy for large and complex networks Dittrich etal. (2024) Support Vector Machine Machine learning, supervised Estimating LCI values for energy use intensity (EUI) for commercial office buildings Effective in high-dimensional spaces with non-linear boundaries, robust to overfitting, works well with small to medium datasets Computationally intensive, especially for large datasets, Requires careful tuning of hyperparameters and kernel selection, Can be difficult to interpret and visualize decision boundaries Deng etal. (2018) Journal of Sustainable Metallurgy factors of chemicals [37]. SHapley Additive exPlanations (SHAP) has also supported the interpretation of the XGBoost model [34, 37]. Compared to other algorithms, classification trees excel due to their ability to capture intricate relationships, their high performance, generalizability, and stability. For example, most tree models surpass neural networks in terms of accuracy [32]. Although deep models outperform them in image and speech recognition, they perform better on tabular datasets [38, 39]. Artificial neural networks (ANNs), on the other hand, excel in their ability to model complex non-linear relationships by learning directly from raw, inaccurate, noisy, or incomplete data (Clarici etal. 1994 in [40]). This is why they have been used to create surrogate LCA models, with the results having similar accuracy as traditional LCA studies [40–44]. Several studies have also applied ANNs to fill missing LCI data and generate data for energy and chemical systems [33, 45–51] and to predict the environmental impacts of different systems [43, 52–54]. ANNs have also been used for characterization factor (CF) generation by predicting the effect of climatic conditions on respiratory disease incidence [55, 56]. However, ANNs often require large training data that are prone to overfitting if not properly regularized, and since they deal with complex relationships, they can be hard to interpret without using algorithms like SHAP [48]. Other neural networks, like CNNs and DNNs have been used to predict building lifespan [57] and to estimate a building’s carbon footprint [58]. RNNs have also been used to perform a multi-objective optimization and scenario analysis of a solar bioenergy system [59]. Graph neural networks (GNNs) can fill in missing data by training them to learn the relationships between the nodes and edges represented in different LCI models [60]. A hybrid algorithm involving combining classification algorithms with statistical algorithms has also been shown to yield the best results compared to other algorithms when predicting biomass feedstock and biomass conversion [45, 61, 62]. Where Can AI Be Applied intheFuture? As described above, AI applications in LCA today have mostly been in LCI and CF generation, LCIA, and a bit for interpretation support, with ANN being the preferred algorithm [12]. Now, the authors assess the trends and expectations for AI in future and how these might fit in the proposed framework. Recent developments of large language models (LLMs) have shown their potential for suggesting chemical processflowsheets, which are useful for LCA flow diagrams and system boundary definition [63].Text-based responses from LLMs are also useful for goal and scope definition and interpretation of results. These LLMs must be trained Table 1 (continued) Algorithm/method Method details Application in LCA Pros Cons Source Reinforcement learning (e.g., Q learning, Deep Q learning) Machine learning, reinforcement Matched LCI data to emissions factors using a semantic-based emission factor matching deep learning model, scenario generation/optimization of energy management systems The model can correct the errors that occurred during the training process, intended to achieve the ideal behavior of a model within a specific context, to maximize its performance, can be useful when the only way to collect information about the environment is to interact with it, improves performance through continuous learning Computationally expensive and time-consuming to train, prone to instability and divergence during training, can be difficult to interpret Luo etal. (2023), Wu etal. (2018) Hierarchical clustering Machine learning, unsupervised LCA surrogate model development Does not require specifying the number of clusters in advance, can handle different types and scales of data Computationally intensive, especially for large datasets, sensitive to noise and outliers, can be difficult to scale for large datasets Sousa and Wallace (2006) Journal of Sustainable Metallurgy 631–632:1279–1294. https:// doi. org/ 10. 1016/j. scito tenv. 2018. 03. 088 44. Sousa I (2025) Part 1: The genesis of Sustainable Minds - The conception of “learning surrogate LCA”, sustainable minds. https:// www. sust a inabl eminds. com/ indus tryblog/ part-1genes issusta inablemindsconce ptionlearn ingsurro gatelca. Accessed 09 Jan 2025 45. Liao M, Kelley S, Yao Y (2020) Generating energy and greenhouse gas inventory data of activated carbon production using machine learning and kinetic based process simulation. ACS Sustain Chem Eng 8(2):1252–1261. https:// doi. org/ 10. 1021/ acssu schem eng. 9b065 22 46. Omidkar A, Alagumalai A, Li Z, Song H (2024) Machine learning assisted techno-economic and life cycle assessment of organic solid waste upgrading under natural gas. Appl Energy 355:122321. https:// doi. org/ 10. 1016/j. apene rgy. 2023. 122321 47. Romeiko XX, Guo Z, Pang Y, Lee EK, Zhang X (2020) Comparing machine learning approaches for predicting spatially explicit life cycle global warming and eutrophication impacts from corn production. Sustainability. https:// doi. org/ 10. 3390/ su120 41481 48. Sun Y, Wang X, Ren N, Liu Y, You S (2023) Improved machine learning models by data processing for predicting life-cycle environmental impacts of chemicals. Environ Sci Technol 57(8):3434–3444. https:// doi. org/ 10. 1021/ acs. est. 2c049 45 49. Venkatraj V, Dixit MK, Yan W, Caffey S, Sideris P, Aryal A (2023) ‘Toward the application of a machine learning framework for building life cycle energy assessment.’ Energy Build 297:113444. https:// doi. org/ 10. 1016/j. enbui ld. 2023. 113444 50. Wan MJ, Phuang ZX, Hoy ZX, Dahlan NY, Azmi AM, Woon KS (2024) Forecasting meteorological impacts on the environmental sustainability of a large-scale solar plant via artificial intelligencebased life cycle assessment. Sci Total Environ 912:168779. https:// doi. org/ 10. 1016/j. scito tenv. 2023. 168779 51. Zhu X, Ho C-H, Wang X (2020) Application of life cycle assessment and machine learning for high-throughput screening of green chemical substitutes. ACS Sustain Chem Eng 8(30):11141– 11151. https:// doi. org/ 10. 1021/ acssu schem eng. 0c022 11 52. Arshad MY etal (2023) Integrating life cycle assessment and machine learning to enhance black soldier fly larvae-based composting of kitchen waste. Sustainability. https:// doi. org/ 10. 3390/ su151 612475 53. Mayol AP etal (2020) Environmental impact prediction of microalgae to biofuels chains using artificial intelligence: A life cycle perspective. IOP Conf Ser Earth Environ Sci 463(1):012011. https:// doi. org/ 10. 1088/ 17551315/ 463/1/ 012011 54. Mohammadi Kashka F, Tahmasebi Sarvestani Z, Pirdashti H, Motevali A, Nadi M, Valipour M (2023) Sustainable systems engineering using life cycle assessment: application of artificial intelligence for predicting agro-environmental footprint. Sustainability. https:// doi. org/ 10. 3390/ su150 76326 55. Araujo LN, Belotti JT, Alves TA, Tadano YdeS, Siqueira H (2020) Ensemble method based on artificial neural networks to estimate air pollution health risks. Environ Model Softw 123:104567. https:// doi. org/ 10. 1016/j. envso ft. 2019. 104567 56. de Souza Tadano Y, Siqueira HV, Alves TA (2016) Unorganized machines to predict hospital admissions for respiratory diseases. In 2016 IEEE Latin American Conference on Computational Intelligence (LA-CCI), pp. 1–6. https:// doi. org/ 10. 1109/ LACCI. 2016. 78856 99 57. Ji S, Lee B, Yi MY (2021) Building life-span prediction for life cycle assessment and life cycle cost using machine learning: a big data approach. Build Environ 205:108267. https:// doi. org/ 10. 1016/j. build env. 2021. 108267 58. Płoszaj-Mazurek M, Ryńska E (2024) Artificial intelligence and digital tools for assisting low-carbon architectural design: merging the use of machine learning, large language models, and building information modeling for life cycle assessment tool development. Energies. https:// doi. org/ 10. 3390/ en171 22997 59. Fang Y, Li X, Wang X, Dai L, Ruan R, You S (2024) Machine learning-based multi-objective optimization of concentrated solar thermal gasification of biomass incorporating life cycle assessment and techno-economic analysis. Energy Convers Manag 302:118137. https:// doi. org/ 10. 1016/j. encon man. 2024. 118137 60. Zargar S, Yao Y, Tu Q (2022) A review of inventory modeling methods for missing data in life cycle assessment. J Ind Ecol. https:// doi. org/ 10. 1111/ jiec. 13305 61. Dinesh A, Rahul Prasad B (2024) Predictive models in machine learning for strength and life cycle assessment of concrete structures. Autom Constr 162:105412. https:// doi. org/ 10. 1016/j. autcon. 2024. 105412 62. Long F, Liu H (2023) An integration of machine learning models and life cycle assessment for lignocellulosic bioethanol platforms. Energy Convers Manag 292:117379. https:// doi. org/ 10. 1016/j. encon man. 2023. 117379 63. Hirtreiter E, Schulze Balhorn L, Schweidtmann AM (2024) Toward automatic generation of control structures for process flow diagrams with large language models. AIChE J 70(1):e18259. https:// doi. org/ 10. 1002/ aic. 18259 64. Cheng A, Calhoun A, Reedy G (2025) Artificial intelligenceassisted academic writing: recommendations for ethical use. Adv Simul 10(1):22. https:// doi. org/ 10. 1186/ s4107702500350-6 65. G. Marcus (2025) Has Grok lost its mind and mind-melded with its owner?, Marcus on AI. https:// garym arcus. subst ack. com/p/ hasgroklostitsmindandmindmelded. Accessed 03 Jun 2025 66. UNEP (2025) AI has an environmental problem. Here’s what the world can do about that. https:// www. unep. org/ newsandstori es/ story/ aihasenvir onmen talprobl emhereswhatworldcandoabout. Accessed 02 Jun 2025 67. E. Mills (2025) How can AI augment rather than dictate human action? This expert explains, World Economic Forum: Emerging Technologies. https:// www. wefor um. org/ stori es/ 2024/ 10/ aiaugme ntratherthandict a tehumanaction/. Accessed 16 Jan 2025 68. The Research Council of Norway (2022) Investment in infrastructures for FAIR research data and public administration data of particular relevance to research. Recommendations from the Data Infrastructure Committee. https:// www. forsk nings radet. no/ conte ntass ets/ 94079 aa751 f94b3 49c15 ffed0 4d5a5 41/ datarireporteng. pdf. Accessed 03 Jun 2025 Publisher's Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.