scieee AI-readable full text Open interactive document viewer

Artificial Intelligence Future Frontiers of Generative AI in Engineering

Şeker, Abdulkadir; YÜKSEK, Ahmet Gürkan

Full text

ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN ENGINEERING Editors Abdulkadir ŞEKER Ahmet Gürkan YÜKSEK Lyon 2025 ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN ENGINEERING Editors Abdulkadir ŞEKER Ahmet Gürkan YÜKSEK Lyon 2025 Artificial Intelligence Future Frontiers of Generative AI in Engineering Editors • Asst. Prof. Abdulkadir ŞEKER • Orcid: 0000-0002-4552-2676 • Assoc. Prof. Ahmet Gürkan YÜKSEK • Orcid: 0000-0001-7709-6360 Cover Design • Motion Graphics Book Layout • Motion Graphics First Published • October 2025, Lyon e-ISBN: 978-2-38236-948-7 DOI: 10.5281/zenodo.17449635 copyright © 2025 by Livre de Lyon All rights reserved. No part of this publication may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, electronic, mechanical, photocopying, recording, or otherwise, without prior written permission from the Publisher. The author or authors of the relevant section are responsible for any copyright infringement that may occur due to the images and graphics used in the book. The editor or publisher does not assume responsibility in this regard. Publisher • Livre de Lyon Address • 37 rue marietton, 69009, Lyon France website • http://www.livredelyon.com e-mail • [email protected] I PREFACE The rapid evolution of Generative Artificial Intelligence (GenAI) has triggered a paradigm shift, profoundly transforming how engineers, scientists, and professionals design, simulate, and innovate. Once largely confined to creative domains, generative models now permeate nearly every discipline— from the intricate virtual replicas of digital twin systems and automated software design to the discovery of novel solar cell materials, advanced cancer research diagnostics, and complex environmental modeling. The convergence of hard engineering principles with generative intelligence has unlocked frontiers where machines not only analyze data but also create, hypothesize, and optimize solutions at a scale and speed previously beyond human imagination. This new era presents both immense opportunities and significant challenges, necessitating a comprehensive guide. This book, Artificial Intelligence: Future Frontiers of Generative AI in Engineering, serves as a critical chronicle of this transformation. It brings together leading-edge, interdisciplinary contributions that explore both the foundational theories and the disruptive real-world applications of GenAI. What sets this volume apart is its ambitious scope. The chapters navigate a spectrum that stretches from the deep theoretical underpinnings of Kolmogorov complexity and conversational AI architectures to urgent, practical implementations in law, digital commerce, and particle physics. Crucially, this book also addresses the technology’s dual-use nature, tackling vital issues of deepfake detection, cyber attacks, and security challenges. By bridging the gap between academia and industry, this work highlights the transformative potential of GenAI for predictive maintenance, intelligent design, and the future of human-AI collaboration. This book has been coordinated under the auspices of the Sivas Cumhuriyet University II   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Artificial Intelligence and Data Science Application and Research Center, reflecting the center’s ongoing commitment to fostering innovative, ethical, and interdisciplinary research in the rapidly advancing fields of artificial intelligence and engineering. Asst. Prof. Abdulkadir ŞEKER Assoc. Prof. Ahmet Gürkan YÜKSEK Editors Asst. Prof. Abdulkadir ŞEKER, Kyrgyz-Turkish Manas University, Assoc. Prof. Ahmet Gürkan YÜKSEK, Sivas Cumhuriyet, University, III CONTENTS PREFACE I CHAPTER I. FOUNDATION AND EVOLUTION OF GENERATIVE ARTIFICIAL INTELLIGENCE 1 Ferhan DEMİRKOPARAN & Metin ZONTUL CHAPTER II. FROM COMPRESSION TO CREATIVITY: KOLMOGOROV COMPLEXITY IN GENERATIVE AI 19 Saliha YEŞİLYURT & Zehra BİLİCİ & Doğa NALCI CHAPTER III. CONVERSATIONAL AI AND QUESTION ANSWERING SYSTEMS: ARCHITECTURES, TRAINING CHALLENGES AND SOLUTIONS 47 Ramazan KATIRCI & Hilal ÇELİK & Taha OĞUZ CHAPTER IV. SOFTWARE DEVELOPMENT SUPPORTED BY GENERATIVE ARTIFICIAL INTELLIGENCE 69 Hakan KEKÜL CHAPTER V. DIGITAL TWIN–DRIVEN PREDICTIVE MAINTENANCE IN INDUSTRY 4.0: EFFECTS OF CONDITIONAL GAN DATA AUGMENTATION ON SVM PERFORMANCE 85 Ahmet Gürkan YÜKSEK CHAPTER VI. LARGE LANGUAGE MODELS IN MATERIALS SCIENCE: APPLICATIONS AND FUTURE DIRECTIONS 107 Nida KATI & Ferhat UÇAR CHAPTER VII. LARGE LANGUAGE MODELS FOR SYSTEMATIC LITERATURE ANALYSIS IN SOLAR CELL MATERIALS RESEARCH: A PRACTICAL FRAMEWORK 121 Nida KATI & Ferhat UÇAR CHAPTER VIII. GENERATIVE AI-ENHANCED APPROACHES FOR PARTICLE PHYSICS APPLICATIONS 149 Mehmet Uğur TÜRKDAMAR , & Celal ÖZTÜRK CHAPTER IX. EMERGING TRENDS IN IOT-DRIVEN GENERATIVE ARTIFICIAL INTELLIGENCE APPLICATIONS FOR AGRICULTURAL ENTERPRISES 177 Şükrü Mustafa KAYA IV   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . CHAPTER X. Q-LEARNING APPLICATIONS IN DIGITAL COMMERCE 191 Havva ABUKAN & Yunus EROĞLU & Selen Yücesoy KAHRAMAN CHAPTER XI. GENERATIVE ARTIFICIAL INTELLIGENCE (GenAI) SUPPORTED DIGITAL MARKETING AND CONTENT STRATEGIES 207 Murat Fatih TUNA & Yunus Emre IŞIK CHAPTER XII. GENERATIVE ARTIFICIAL INTELLIGENCE IN EDUCATIONAL TECHNOLOGIES 233 Mehmet Ali DEVECİ & Yasin GÖRMEZ & Sait BARDAKÇI CHAPTER XIII. A BIBLIOMETRIC ANALYSIS OF CANCER RESEARCH UTILIZING GENERATIVE ARTIFICIAL INTELLIGENCE METHODS BETWEEN 2022 AND 2024 263 Mustafa TEMIZ & Burcu BAKIR-GUNGOR CHAPTER XIV. GENERATIVE ARTIFICIAL INTELLIGENCE APPLICATIONS IN LAW 275 Fatma Betül ŞEKER & Berna DÖNEK CHAPTER XV. GENERATIVE AI IN MEDIA, COMMUNICATION AND VISUAL ARTS 295 Oğuz KAYNAR & Murat Fatih TUNA CHAPTER XVI. DEEPFAKE DETECTION, FAKE CONTENT VERIFICATION, AND CYBER ATTACK SCENARIOS: METHODS, CHALLENGES AND MITIGATION STRATEGIES 327 Saadin OYUCU & Ahmet AKSÖZ CHAPTER XVII. SUSTAINABILITY AND ENVIRONMENTAL MODELING WITH GENERATIVE ARTIFICIAL INTELLIGENCE 341 Emre ÜNSAL & Ahmet Fırat YELKUVAN FOUNDATION AND EVOLUTION OF GENERATIVE ARTIFICIAL INTELLIGENCE   7 Figure 3. The transformer model architecture 3.4. Diffusion Model A diffusion model employs a two-way process: first, it gradually converts the input data into noise, which is called the forward process, and then, a clean output is reconstructed from the noisy data, which is known as the reverse process. A diffusion model is typically designed as a neural network (such as a U-Net) and learns to generate meaningful samples from noise. The noise is iteratively attempted to be approximated towards the target distribution. Diffusion model is successfully used in zero-shot generative tasks and generation of scenes that do not exist in real world environments, but it is computationally expensive (He, et. al., 2025). Besides, it handles the instability and mode collapse problem of GANs and can create high quality, high resolution and diverse samples (Bengesi, et.al., 2024). 8   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . The original mathematical framework that forms the foundation of diffusion models is called Denoising Diffusion Probabilistic Models (DDPM). Since classical DDPM is computationally heavy, some methods have been developed to decrease the cost. Song et. al. decreased the number of steps by using deterministic sampling (Song, et. al., 2020). Rombach et. al. proposed Latent Diffusion Model (LDM) which applies diffusion on a lower dimension latent space (Rombach, et. al., 2022). This forms the scientific foundation of Stable Diffusion. The architectural structure of a text-to-image diffusion model used by Stable Diffusion (Rombach, et. al., 2022) which is an open source LLM can be seen in Figure 4 (Rahmatulloh, 2025). Figure 4. Architectural structure of a text-to-image diffusion model 3.2. Large Language Models (LLMs) Large Language Models (LLMs) are sub-branch of generative AI developed for NLP tasks such as text generation, question answering, natural translation. LLMs require them to be trained with large-scale datasets that consist of a large amount of unstructured, unlabeled data. They can be internet sources such as books, papers, websites or libraries. Even if LLMs were originally designed for text domains, after they become popular a lot of applications in various domains have emerged. They have the ability to create semantically consistent, coherent, realistic text, story, script, advertisement, audio, image, video, speech that cannot be separated from human created content. According to modality LLMs can be classified into two categories: unimodal models which generate same kinds of data as inputs, multimodal models which can operate with data FOUNDATION AND EVOLUTION OF GENERATIVE ARTIFICIAL INTELLIGENCE   9 from various domains. These models are defined as x-to-modality (Bahn & Strobel, 2023). Unimodal models could be image-to-image or text-to-text while unimodal models could be text-to-image, text-to-video etc. General purpose LLMs are initially pretrained with large scale datasets with billions or trillions of numbers of parameters and can be used in multiple domains for many different interdisciplinary tasks. After training process is completed, it can be made domain-specific by fine tuning and additional domain-specific data. This phenomenon is attributable to the principles of transfer learning. It is feasible to develop novel models by leveraging transfer learning in conjunction with opensource LLMs. To specialize in a Large Language Model (LLM) for a particular task, human evaluation can be employed to select sufficient high-quality responses. This process effectively serves to guide the system toward a desired behavioral trajectory in the target domain, which is crucial for fine-tuning or reinforcement learning from human feedback (RLHF) (Christiano, et. al., 2017). On the other hand, the prompts given to LLM system directly affect the quality and direction of generated outputs. So, prompt engineering is crucial for GAI system to adopt certain problems (He, et. al, 2025). The emergence of GPT by OpenAI in 2018 marked the rapid entry of LLMs into mainstream discourse and applications, profoundly impacting various sectors (Radford, et. al., 2018). Generative Pre-trained Transformer (GPT) is a transformer based LLM which employs decoder architecture. Bidirectional Encoder Representation from Transformers (BERT) is released by Google concurrently, which also depends on Transformer. But it adopts a different approach and rather than predicting the next word in the sentence like GPT, it predicts some masked words in the sentence and learns bidirectional semantic context (Devlin, et. al., 2019). Developed as a competitor to ChatGPT in 2023, Gemini (formerly Bard) is based on the Pathways architecture created by Google DeepMind (DeepMind, 2023). It can create not only text, but also audio, video, image and code content. Large Language Model Meta AI (LLaMA) is one of the most important open-source LLMs designed by Meta. The core importance of Llama is rooted in its open-source nature, thereby providing both researchers and developers with direct, unencumbered access to the model’s parameters and architectural blueprints. Open accessibility has led to the development of numerous derivative models such as Alpaca, Vicuna, Mistral, primarily driven by collective community contributions. T5 is Google’s sequence-to-sequence transformer-based model. It approaches all tasks as a text-to-text problem. There are a lot of different LLMs in various domains. Some of them are Meta’s 10   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Make-a-video (Singer, et.al., 2022) for video, Github’s CoPilot (Github, 2023) for code, JukeBox (Dhariwal, et. al., 2020) for music generation. It can be found a summary of major milestones of LLMs in Table 1 and a comprehensive comparison of some LLMs according to accuracy and complexity in Figure 5 (Jovanovic & Campbell, 2022). Figure 5. Accuracy and complexity comparison of LLMs FOUNDATION AND EVOLUTION OF GENERATIVE ARTIFICIAL INTELLIGENCE   11 Table1. A summary of major milestones of LLMs Date Model Key Features Significance 2017 Transformer (Vaswani et al.) Introduced the selfattention mechanism and positional encoding Established the foundation for scalable sequence modeling and parallel training 2018 GPT (Radford et al.) Unsupervised pretraining using transformer decoder blocks Demonstrated transfer learning for natural language understanding 2019 BERT (Devlin et al.) Bidirectional encoding with masked language modeling Enabled deep contextual comprehension in NLP tasks 2019– 2020 GPT-2, RoBERTa, XLNet, T5 Larger datasets, refined training objectives Marked a transition to general-purpose language understanding 2020 GPT-3 (Brown et al.) 175 billion parameters; few-shot learning Enabled coherent opendomain text generation and reasoning 2022 InstructGPT / ChatGPT Reinforcement Learning from Human Feedback (RLHF) Introduced alignment with human intent and safer dialogue generation 2022 Stable Diffusion (Stability AI, LMU Munich, Runway) Latent Diffusion Model (LDM) Enabled high-quality, open-access text-toimage generation using latent space denoising; marked a major step in democratizing generative image synthesis. 2023 Gemini 1 (Google DeepMind) Multimodal (text, vision, audio, code) reasoning Integrated multiple modalities into a unified transformer framework 2023 GitHub Copilot (OpenAI + GitHub) Fine-tuned transformer for software development Popularized AI-assisted programming and code completion 2024 GPT-4 / GPT-4o (OpenAI) Multimodal input (text, image, audio) with improved reasoning Unified perception and reasoning across modalities 12   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . 2024 Claude 3 (Anthropic) Constitutional AI for safety and ethical alignment Advanced transparency and user intent understanding 2024 LLaMA 3 (Meta) Open-weight, highefficiency model family Democratized access to powerful open-source LLMs 2024 Mistral 7B / Mixtral 8x7B Sparse mixture-ofexperts (MoE) design Enhanced efficiency without compromising accuracy 2025 Gemini 1.5 / Gemini 2 (Google DeepMind) Long-context, multimodal reasoning, code understanding Extended context memory and improved symbolic reasoning 2025 Claude 3.5 (Anthropic) Real-time reasoning and factual consistency Enhanced step-by-step reasoning and chain-ofthought reliability 2025 OpenAI o1 (Reasoning Model) Specialized reasoning model with selfverification Represents the frontier in structured logical problemsolving 2025 DeepSeek (Hangzhou DeepSeek AI) Mixture-of-Experts (MoE) with Multihead Latent Attention (MLA) Combines efficient large-scale training and open-weight transparency; competitive with GPT4-level models while maintaining lower computational costs. 4. Ethical Concerns, Limitations, Opportunities and Future Directions Generative AI has quickly become one of the most exciting areas in artificial intelligence, allowing computers to create original and meaningful content. It raised opportunities for human interact with machines and explored many ways of creativity in many different areas such as industry, science, design, architecture, entertainment, art, literature etc. But this comes with ethical concerns, copyright and originality issues. In creativity matters, it is highly vague where generation of users start and where generation of GAI ends. So, this may raise concerns about artists, designers or writers taking credit for their AI assisted works. Despite the remarkable progress in Generative AI, numerous challenges remain regarding ethical concerns, transparency, misuse or misinformation. FOUNDATION AND EVOLUTION OF GENERATIVE ARTIFICIAL INTELLIGENCE   13 Since general purpose, multimodal GAIs require large scale datasets to learn language patterns, they are also highly dependent on the data. Most of data used in GAI are extracted from Internet sources so, the bias, discriminative content and controversial issues in data are directly carried out to model. The bias descended from data is called training bias. Sometimes, even if training data is clean, bias may occur in GAI system which is called algorithmic bias. GAI model being overfitted during training and not be able to learn data distribution appropriately cause algorithmic bias (Koehler, 2024). The deployment of Generative AI (GAI) across the healthcare and education sectors holds the potential to yield substantive and immense benefits. The generation of personalized educational material provides significant advantages for students, teachers, individuals with learning disabilities, and adult learners. Similarly, in the healthcare sector, GAI will become one of the most important tools for improving individuals’ quality of life through its capabilities such as drug design, personalized health monitoring, and customized diagnosis and treatment (Takale, et. al. 2024). Furthermore, given that the dataset holds a critical role in the training of GAI models, and acquiring sufficient, high-quality data is challenging in healthcare and numerous other sectors, the capacity for synthesizing novel data will specifically accelerate academic and scientific studies. On the other hand, GAI applications being unexpensive, accessible and easy to use facilitate malicious use and abuse. Especially with deep fake applications users make highly realistic images, videos or even speeches which can be used for manipulating, deceive or blackmailing people. It can easily ruin an individual’s reputation or impersonate them. This is why the transparency and interpretability of GAI are becoming crucial. Being able to interpret the system’s underlying decision-making process and understand the reasons behind them both encourages responsible use by preventing manipulation and helps users be more aware of inaccurate, incomplete, or fabricated content created by GAI. To use GAI correctly and properly, user awareness must be increased, and training should be offered as needed (Kılınç & Keçecioğlu, 2024). Interest in Generative AI among scientists and academics has been steadily rising, in parallel with its popularization and permanent integration into daily life, particularly over the last decade. The widespread adoption of open-source GAI models is critical for the advancement of this domain. However, due to the substantial requirements for processing power, storage capacity, and specialized hardware, GAI remains significantly more expensive than what an ordinary researcher can afford. For instance, the estimated computational cost for a single training run of GPT-3 (trained with 175 billion parameters on 300 billion tokens) 14   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . is approximately $5 million (Corchado, et. al. 2023). So, it is highly expensive to research as an individual. Conversely, the apprehension that GAI will lead to job displacement by eliminating certain occupations is among the escalating societal concerns. Nevertheless, while GAI’s capability to automate specific tasks may indeed suppress some job roles, it concurrently gives rise to numerous novel job categories, such as prompt engineering. Furthermore, the automated handling of mundane, repetitive, and low-creativity tasks frees up employees in many professions, including software development and accounting, from burdensome chores, allowing them to dedicate more time to higher-level, large-scale, and creativity-intensive assignments. 5. Conclusion The evolution of generative artificial intelligence has redefined the boundaries between human creativity and machine capability. What started as a question about whether machines could think has now become a reality where they can not only process and understand data but also create new, meaningful, and original content. From the foundational architectures such as Autoencoders and Recurrent Neural Networks to the revolutionary Transformer-based models, the progress in generative AI represents one of the most remarkable technological achievements of the modern era. Generative models such as GANs, VAEs, Diffusion Models, and Transformer-based LLMs have each contributed unique mechanisms to learn data distributions and synthesize new content. These models have demonstrated extraordinary success across various modalities, including text, image, audio, and video generation. Their influence extends far beyond the field of computer science, impacting disciplines as diverse as healthcare, education, art, entertainment, and creative design. The multimodal nature of modern architectures such as GPT-4, Gemini, and LLaMA exemplifies the convergence of language, vision, and reasoning within a unified framework of intelligence. Besides numerous benefits of deployment generative AI, ethical, social, and technical challenges remain. Concerns related to bias, misinformation, authenticity, and intellectual property have become increasingly important. The same systems capable of artistic and scientific breakthroughs can also be misused for manipulation or deception if not governed responsibly. Ensuring transparency, fairness, and interpretability within these systems remains one of the central goals of future AI research. Consequently, policymakers, engineers, FOUNDATION AND EVOLUTION OF GENERATIVE ARTIFICIAL INTELLIGENCE   15 ethicists, and philosophers must collaborate to establish the necessary regulatory frameworks that ensure these principles are upheld. In the future, the responsible and transparent use of generative AI will be essential to ensure that it supports human progress, creativity, and ethical innovation. References Feuerriegel, S., Hartmann, J., Janiesch, C., & Zschech, P. (2024). Generative ai. Business & Information Systems Engineering, 66(1), 111-126. Rios-Campos, C., Viteri, J. D. C. L., Batalla, E. A. P., Castro, J. F. C., Núñez, J. B., Calderón, E. V., ... & Tello, M. Y. P. (2023). Generative artificial intelligence. South Florida Journal of Development, 4(6), 2305-2320. Banh, L., & Strobel, G. (2023). Generative artificial intelligence. Electronic Markets, 33(1), 63. Van Engelen, J. E., & Hoos, H. H. (2020). A survey on semi-supervised learning. Machine learning, 109(2), 373-440. Rosenblatt, F. (1957). The perceptron, a perceiving and recognizing automaton project para, report: cornell aeronautical laboratory, cornell aeronautical laboratory. URL: https://books. google. pl/books. Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural computation, 9(8), 1735-1780. Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25. Alec, R., Karthik, N., Tim, S., & Ilya, S. (2018). Improving language understanding with unsupervised learning. Citado, 17, 1-12. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019, June). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) (pp. 4171-4186). Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., ... & Bengio, Y. (2014). Generative adversarial nets. Advances in neural information processing systems, 27. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30. 16   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Kingma, D. P., & Welling, M. (2013). Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114. Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in neural information processing systems, 33, 6840-6851. Bengesi, S., El-Sayed, H., Sarker, M. K., Houkpati, Y., Irungu, J., & Oladunni, T. (2024). Advancements in generative AI: A comprehensive review of GANs, GPT, autoencoders, diffusion model, and transformers. IEEe Access, 12, 69812-69837. Foster, D. (2022). Generative deep learning. “ O’Reilly Media, Inc.”. He, R., Cao, J., & Tan, T. (2025). Generative artificial intelligence: a historical perspective. National Science Review, 12(5), nwaf050. Arjovsky, M., Chintala, S., & Bottou, L. (2017, July). Wasserstein generative adversarial networks. In International conference on machine learning (pp. 214-223). PMLR. Sengar, S. S., Hasan, A. B., Kumar, S., & Carroll, F. (2025). Generative artificial intelligence: a systematic review and applications. Multimedia Tools and Applications, 84(21), 23661-23700. Mao, X., Li, Q., Xie, H., Lau, R. Y., Wang, Z., & Paul Smolley, S. (2017). Least squares generative adversarial networks. In Proceedings of the IEEE international conference on computer vision (pp. 2794-2802). Zhang, H., Goodfellow, I., Metaxas, D., & Odena, A. (2019, May). Selfattention generative adversarial networks. In International conference on machine learning (pp. 7354-7363). PMLR. Odena A (2016) Semi-supervised learning with generative adversarial networks. Odena A, Olah C, Shlens J (2017) Conditional image synthesis with auxiliary classifier gans. Mirza, M., & Osindero, S. (2014). Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784. Radford, A., Metz, L., & Chintala, S. (2015). Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434. Srinivasa Rao, T., Mandava, S. K., Madathala, H., Durgaraju, S., & Dalal, A. (2025). Generative AI. Sarath Krishna and Madathala, Harikrishna and Madathala, Harikrishna and Durgaraju, Sairam and Durgaraju, Sairam and Dalal, Aryendra, Generative AI (January 01, 2025). FROM COMPRESSION TO CREATIVITY: KOLMOGOROV COMPLEXITY IN . . .   23 bits of information. This perspective is purely discrete; information is treated as reduction in uncertainty when one element is selected from a finite set. It applies well to deterministic systems such as codes, games, or logical puzzles, where probabilities are unnecessary. The probabilistic approach, developed most notably by Claude Shannon (1948), introduced entropy as the expected information content of a random variable. Here, information depends on the probability distribution ()px : 2 ( ) ()log () x H X px px =- å This framework enabled engineering breakthroughs such as channel capacity, rate–distortion theory, and source coding theorems. However, it presupposes an ensemble of outcomes—an assumption that becomes artificial when analysing individual objects, such as a particular novel or scientific dataset. These conceptual limitations inspired Kolmogorov to seek a third, more general definition. 2.2. The Algorithmic Turn Kolmogorov’s algorithmic approach addressed the problem of measuring the information content of a single object without assuming any underlying probability distribution. The Kolmogorov complexity ()Kx of a finite string x is defined as the length of the shortest program that produces x and halts on a fixed optimal universal Turing machine (Li & Vitányi, 2008). This definition provides an absolute measure of information that applies to any data type— numeric, textual, or visual—because any computable description can be represented as a program. The algorithmic view unites description and prediction. According to the Algorithmic Coding Theorem, () () (1)K x log m x O=- + where ()mx denotes the universal distribution or algorithmic probability (Solomonoff, 1964). This equation links complexity and probability, objects generated by shorter programs have higher prior probability. The result encapsulates the intuition that simpler explanations are more likely—a mathematical expression of Occam’s razor. Building on this foundation, Levin (1984) introduced resource-bounded variants such as () t Kx , which add runtime penalties to favor both short and 24   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . efficient programs. These forms bridge the theoretical world of unbounded computation and the practical constraints of real machines. They also inspired modern bounded approximations such as the Coding Theorem Method (CTM) and Block Decomposition Method (BDM), which estimate complexity through empirical enumeration or block wise aggregation of local patterns (Zenil et al., 2015). The progression from combinatorial counting to probabilistic entropy and finally to algorithmic complexity is summarized in Figure 3, which visually contrasts these three paradigms and highlights the increasing role of computation in defining information. 2.3. From Incomputable Ideals to Practical Measures Kolmogorov’s framework provides a universal benchmark for all forms of description. Yet ()Kx is incomputable, no algorithm can determine the shortest program that outputs a given string, because that would solve the halting problem. Despite this, the concept remains operational through computable approximations and theoretical bounds (Vitányi, 2006). Researchers distinguish between incomputable measures, which define the ideal, and computable surrogates, which approximate it within limited resources. These relationships are summarized in Table 1, which organizes the main measures in Algorithmic Information Theory (AIT) by their basis, computability, role in prediction, and practical caveats. Figure 3. The combinatorial approach counts possibilities, the probabilistic approach measures average uncertainty, and the algorithmic approach quantifies the complexity of individual objects through computation. Together they chart the evolution from counting to coding to computation. FROM COMPRESSION TO CREATIVITY: KOLMOGOROV COMPLEXITY IN . . .   25 2.4. Inductive Inference and the Birth of Universal Prediction The link between description length and prediction was made explicit by Ray Solomonoff in his theory of algorithmic probability(Solomonoff, 1964). He proposed assigning prior probability to outputs in proportion to the summed weights 2 p-∣∣of all programs that produce them. This formulation yields a universal semimeasure ()mx satisfying the coding theorem identity, ensuring that short programs dominate the distribution. In effect, Solomonoff’s model formalizes inductive reasoning as probabilistic weighting of simple explanations. Later, Hutter extended these principles to define universal sequence prediction, showing that expected loss under this prior converges to the theoretical optimum. Together, these developments established a unified view of intelligence as compression-based inference, where learning corresponds to discovering concise generative mechanisms that balance accuracy and parsimony. In summary, by integrating the combinatorial, probabilistic, and algorithmic approaches, Kolmogorov reframed information theory as a study of computation itself. Description length became not just a measure of uncertainty but a universal language for understanding, predicting, and generating data. The next section builds upon this foundation to examine how compression functions as prediction within modern artificial intelligence systems. 26   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Table 1. Comparative Overview of Algorithmic Information Measures Complexity Variant Basis Computability Predictive / Modelling Role Key Caveats Kolmogorov complexity ()Kx Shortest selfdelimiting program on an optimal UTM Not computable, upper-semi computable Theoretical lower bound on lossless compressibility, via the coding theorem relates to m(x) Depends on choice of optimal UTM up to O(1), not directly estimable Plain complexity ()Cx Shortest program (not self-delimiting) Not computable, upper-semi computable Related to K(x) by () () ( )K x C x O log x£+∣∣ Additive constants matter in practice, lacks prefix-freeness Time-bounded complexity () t Kx Program length + runtime penalty (length–time trade-off) Not computable in general, upper-semi computable, practical upper bounds under resource limits Favors short and fast programs, motivation for efficient prediction Time bound and machine model affect values CTM / BDM approximations Small-machine enumeration, block decomposition Computable heuristics Practical heuristics for short strings & local structure, empirical structure discovery Resolutiondependent, finite-state biases, scaling limits, domain sensitivity Levin’s universal prior ()mx Weighted sum of programs ( ) p p:U p x ( ) (2 )) mx - = = å∣∣ Lower-semi computable, semimeasure Levin’s universal prior m(x) Weighted sum of programs ( The Algorithmic Coding Theorem: K(x) = ( ) (1).log m x O-+ Algorithmic Probability (Solomonoff–Levin): () (2 ) - = å p p:U p x View shift: probabilistic (ensemble) → algorithmic (individual object). 3. Compression as Prediction in AI The link between compression and prediction is not only theoretical—it defines how all modern machine learning systems operate. In essence, a model FROM COMPRESSION TO CREATIVITY: KOLMOGOROV COMPLEXITY IN . . .   27 that predicts well must also compress effectively. This equivalence arises directly from Shannon’s source coding theorem, which proves that the expected length of an optimal code equals the cross-entropy between the true data distribution * ()px and the model’s estimated distribution ()qx : *[ ( ) ] ( * , ) p L E logq x H p q=- = Minimizing cross-entropy, therefore, is identical to minimizing expected code length. Every gradient descent step toward higher likelihood is simultaneously a step toward shorter description. 3.1. The Compression–Prediction Pipeline In practice, arithmetic coding turns this theoretical relationship into a concrete mechanism. Given an autoregressive model () ii qx x < ∣ , arithmetic coding can encode data using approximately 2() ii log q x x< -∣ bits per symbol (Witten et al., 1987). The total message length thus measures predictive accuracy. When this model is trained by maximizing log-likelihood, it implicitly minimizes code length. The same logic extends to latent-variable models, such as Variational Autoencoders (VAEs). When coding both the latent variable z and the data conditional ()px z∣ , the overall cost equals the Evidence Lower Bound (ELBO). Training by maximizing ELBO hence amounts to training for compression (Kingma & Welling, 2013). These relationships unify deep learning objectives with information-theoretic principles, whether through cross-entropy, variational bounds, or mutual information, the goal remains optimal compression. 3.2. Large-Scale Models as Universal Compressors Viewed through this lens, large language models (LLMs) are enormous compression engines. Their training objective—minimizing predictive loss— translates directly into minimizing expected code length. With arithmetic coding, LLMs can achieve lossless compression competitive with or superior to classical algorithms across multiple modalities (Delétang et al., 2023). For instance, text-trained models such as Chinchilla and GPT-4 compress not only text but also images and audio, achieving rates far below standard codecs. This suggests that high-level abstractions learned from textual data generalize across domains, embodying cross-modal compression. Model size, however, complicates this picture. If we treat the parameters themselves as part of the total description, total code length equals the sum of 28   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . data code length and model code length. As model parameters increase, pertoken loss may fall, but total compression efficiency eventually declines—a behaviour visualized earlier in Figure 2. This trade-off forms the foundation of scaling laws, which define optimal model sizes for a given dataset (Hoffmann et al., 2022). 3.3. Diffusion Models and Generative Compression Recent developments extend this principle beyond text to diffusion models, which generate data through iterative denoising. At each timestep, a diffusion model predicts the noise component of a corrupted input; the training loss, usually mean-squared error or KL-divergence, is again equivalent to minimizing expected code length between model predictions and data distributions. In practice, the latent trajectories of diffusion models can be interpreted as progressive compression stages. Each denoising step removes entropy, transforming random noise into structured output. Thus, diffusion models embody compression not only statistically but dynamically compressing disorder into order. 3.4. Tokenization and Representation as Pre-Compression A crucial but often overlooked step in this process is tokenization, which defines the alphabet over which compression occurs. Tokenization determines sequence length, context granularity, and the model’s ability to capture dependencies. Subword methods such as Byte-Pair Encoding (BPE) or SentencePiece serve as pre-compression mechanisms, they balance redundancy removal with expressive capacity (Delétang et al., 2023). Vocabulary size directly affects per-token entropy and therefore achievable bitrates. Large vocabularies increase representational precision but reduce compression efficiency; smaller vocabularies favour compactness at the expense of detail. In this sense, tokenization itself is part of the compression model—it defines what structure is explicit and what must be learned. 3.5. Generalization Through the Minimum Description Length Principle The Minimum Description Length (MDL) principle (Grünwald, 2007; Rissanen, 1978) provides the theoretical bridge between compression and generalization. According to MDL, the best hypothesis is the one that yields the shortest total description of both the model and the data it explains. Learning thus becomes the search for the most concise yet predictive representation. This FROM COMPRESSION TO CREATIVITY: KOLMOGOROV COMPLEXITY IN . . .   29 view aligns naturally with PAC-Bayesian theory, where models that generalize well are precisely those that compress the training data effectively (Dziugaite & Roy, 2017). In empirical studies, neural networks have demonstrated this property even without explicit regularization. When trained on real data, they find internal representations that compress across domains, whereas random labels destroy compressibility and prevent generalization (Zhang et al., 2016). Compression, therefore, is not just a by-product of learning—it is its fundamental precondition. 3.6. The Limits of Statistical Compression Despite their power, current systems achieve only statistical compression, not true algorithmic compression. Statistical compressors approximate probabilities over fixed alphabets; they do not discover the shortest generative program that produces the data. Kolmogorov complexity ()Kx , which defines that ideal, remains incomputable (Vitányi, 2006). Practical methods like CTM or BDM approximate it only for short strings or simple structures (Zenil et al., 2015). Consequently, modern AI operates below the Kolmogorov limit—it captures correlations, not causes. Bridging this gap requires models capable of explicit program induction rather than probability estimation. The next section introduces recent efforts to operationalize this idea, such as the Kolmogorov Test (KT), which evaluates intelligence by asking models to produce minimal programs that recreate observed sequences. These benchmarks move the field from statistical redundancy reduction toward algorithmic reasoning and creative abstraction. 4. Beyond Statistical Compression: Kolmogorov Tests of Intelligence Statistical compression explains how models reduce redundancy by fitting data distributions, yet it does not capture the essence of creative intelligence. True creativity involves discovering the shortest set of generative rules that can reproduce observations, not merely predicting the next symbol. This distinction motivates a growing body of research that evaluates artificial intelligence systems through the lens of algorithmic compression, where success depends on program synthesis rather than likelihood optimization. 4.1. The Kolmogorov Test as a Measure of Intelligence The Kolmogorov Test (KT), introduced by Yoran et al. (2025), evaluates intelligence as compression through code generation. Instead of predicting 30   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . tokens, a model must generate a concise program that reconstructs a target sequence exactly. Each candidate program is executed, and the resulting sequence is compared to the original. The shorter the program that succeeds, the greater the model’s measured intelligence. This framework provides a practical way to approximate the Kolmogorov complexity ()Kx , which defines the length of the shortest program producing an object x . Because ()Kx is incomputable, KT uses program synthesis as a proxy, treating code-generating language models as empirical estimators of algorithmic compression. A model that produces shorter, functionally correct code exhibits deeper abstraction ability. 4.2. Methodology and Design Principles The KT benchmark presents a sequence during inference and asks the model to output a program that, when executed, reproduces that sequence. Each trial measures both correctness and compression ratio, ensuring that mere memorization cannot succeed. The benchmark has several advantages. 1. The compression metric is difficult to exploit because shorter programs are objectively verifiable. 2. It accommodates diverse sequence types such as text, audio, and biological data. 3. Pretraining contamination is unlikely, since training corpora rarely include matching program–sequence pairs. 4. Difficulty can be adjusted by controlling sequence length or the depth of compositional operators. To ensure reproducibility, KT employs a domain-specific language (DSL) with known ground-truth shortest programs. This design allows researchers to quantify how far a model’s generated code diverges from the theoretical minimum, providing an operational measure of algorithmic reasoning. 4.3. Empirical Findings and Interpretations Results reveal that even state-of-the-art large language models perform poorly on KT benchmarks. On natural sequences, LLaMA-3.1-405B fails in about 78% of cases, and GPT-4o fails in about 40%. Smaller code-specialized models, although trained on explicit program–sequence pairs, only partially improve performance, with gains that do not generalize to real data (Yoran et al., 2025). FROM COMPRESSION TO CREATIVITY: KOLMOGOROV COMPLEXITY IN . . .   31 These findings highlight a key limitation. While current models achieve strong Shannon-style compression, they remain far from discovering true Kolmogorov-style programs. They can predict but rarely explain. This gap marks the boundary between statistical redundancy removal and mechanistic understanding, which is central to both creativity and intelligence. 4.4. Creativity as Algorithmic Abstraction In the Kolmogorov framework, creativity can be defined as the ability to find short programs that balance concision with explanatory adequacy. Systems that learn reusable structures or compositional rules are more creative because they achieve deeper compression through generalization. For example, a model that infers the recursive pattern behind a musical sequence demonstrates creativity beyond statistical imitation. Benchmarks such as KT move this notion from theory to measurement. By rewarding models that generate concise, reusable code, KT encourages abductive reasoning, the process of forming hypotheses that explain observations compactly. This approach aligns directly with the algorithmic probability principle, which assigns exponentially higher weight to shorter explanations (Solomonoff, 1964). Creativity, therefore, is not defined by novelty alone but by efficient explanation. A creative model discovers patterns that make the world simpler to describe, bridging compression, understanding, and invention. 4.5. Relationship to Practical Approximations Because exact computation of ()Kx is impossible, researchers use approximations such as the Coding Theorem Method (CTM) and the Block Decomposition Method (BDM)(Zenil et al., 2015). CTM estimates algorithmic probability by enumerating small Turing machines and recording their outputs, while BDM extends this to larger data by decomposing it into smaller blocks and summing their estimated complexities. These methods are computationally expensive but provide valuable upper bounds on ()Kx for short sequences. Within KT and related tests, these approximations serve as calibration tools. They help quantify how far a model’s generated program lies above the theoretical minimum and clarify the extent to which current neural systems achieve algorithmic compression rather than mere probabilistic modelling. 32   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . 4.6. Implications for Diffusion and Transformer Architectures The KT perspective also sheds light on the behaviour of contemporary diffusion and transformer-based architectures. Diffusion models reduce entropy progressively, yet they still rely on learned distributions rather than explicit rules. Their compression is statistical, achieved through noise prediction. Transformers, by contrast, learn flexible internal codes that can in principle represent algorithmic relations, but they seldom generate minimal programs spontaneously. Integrating KT-like objectives into their training may encourage explicit program induction and promote genuine creative reasoning. 4.7. Toward Contamination-Resistant Evaluation One of the strengths of KT is its resistance to benchmark contamination, a persistent issue in modern AI evaluation. Because KT uses procedurally generated data with known minimal programs, models cannot rely on memorization. They must generalize the underlying pattern through abstraction. This makes KT a reliable framework for studying reasoning, creativity, and transfer across modalities. Table 2. Comparative Overview of Algorithmic Information Measures Feature Statistical Compression (Shannon-style) Algorithmic Compression (Kolmogorov-style) LLM Role Likelihood Estimator (predicts next token probability). Program Synthesizer (generates the shortest program). Compression Target Minimizing expected code length (statistical redundancy). Minimizing program length (structural complexity). Key Mechanism Arithmetic Coding based on model probabilities. Code Generation and Execution of the program. Reference Study (Delétang et al., 2023) (Yoran et al., 2025) Flow Data Sequence → Language Model → Arithmetic Coder → Compressed Bit Stream Data Sequence → Code Generating LM → Program Candidate → Execution → Output Sequence As shown in Table 2, the KT distinguishes between statistical compression, typical of large-scale language modelling, and algorithmic compression, which reflects true program discovery. FROM COMPRESSION TO CREATIVITY: KOLMOGOROV COMPLEXITY IN . . .   39 6.3. Algorithmic Probability and Creative Reasoning Algorithmic probability provides a theoretical explanation for creativity in this context. According to Solomonoff , shorter programs have exponentially higher prior probability. This principle unites compression, prediction, and comprehension. A model that finds a compact explanation demonstrates creativity because it discovers rules that generalize beyond the data. The SuperARC framework extends this reasoning. It shows that recursive algorithmic compression can produce accurate predictions and causal explanations at the same time (Hernández-Espinosa et al., 2025). In practice, this means learning reusable subroutines—loops, functions, or grammar rules— that describe tasks compactly. Creative systems build such structures instead of memorizing specific outputs. 6.4. Gradient Descent as a Limited Search Training with gradient descent can be interpreted as a constrained form of compression search. Minimizing log-loss reduces expected code length within a restricted hypothesis class, consistent with the Minimum Description Length (MDL) principle. However, stochastic gradient descent cannot explore the full space of possible programs. It optimizes parameters, not symbolic structures. This limitation explains why models with excellent statistical compression may still fail at true program synthesis. Bridging the gap requires methods that combine continuous optimization with explicit program-space search. Approaches such as guided priors, execution feedback, or hybrid neurosymbolic models (using CTM and BDM) offer promising directions (Hernández-Espinosa et al., 2025). 6.5. The Broader View of Compression and Creativity Generative AI systems generalize well because real-world data are compressible. This property links learning to the physical and cognitive world. Schmidhuber (2008) argued that progress in compression is intrinsically rewarding. It drives curiosity and the pleasure of discovery. When a model or a person finds a shorter explanation for complex information, understanding becomes more elegant and efficient. The same idea connects science and art. Scientists compress the world into concise laws; artists compress experiences into form and rhythm. Both reduce complexity while preserving meaning. In each case, creativity is a search for 40   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . simplicity behind apparent diversity. The more compact the representation, the stronger its explanatory power. In summary, code modelling reveals the intimate relationship between compression and creative reasoning. Statistical compression explains performance gains, but only algorithmic compression explains understanding. Moving toward the latter requires systems that search beyond surface patterns to discover the mechanisms that generate them. The next section outlines the open challenges and future research needed to make that transition. 7. Open Challenges and Future Directions Kolmogorov Complexity (KC) provides a rigorous framework that links compression to prediction, explanation, and creativity. However, the incomputability of K(x) limits its direct use in generative AI, making current approaches only partial and approximate. Several open challenges shape today’s research frontier, as summarized in Table 3. 7.1. Measuring creativity beyond compression. Measuring creativity beyond compression remains difficult. KC formalizes the idea that shorter explanations are better, grounding Occam’s razor in mathematics. This supports the view that creativity can be seen as finding shorter and more explanatory programs. Yet since K(x) is incomputable, no model can confirm that it has found the shortest description. It is still unclear whether description length alone measures novelty or whether other factors, such as mechanistic adequacy and causal sufficiency, are also needed. 7.2. Benchmarking creative reasoning. Benchmarks like the Kolmogorov Test (KT) aim to apply KC by asking models to generate short programs that reproduce sequences across domains such as text, audio, and DNA. In theory, this connects statistical prediction to algorithmic abstraction. In practice, large language models perform poorly, showing that they excel at surface-level compression but not yet at the deeper, program-level creativity implied by KC. 7.3. Recursive compression and contamination-resistant testing. Frameworks like SuperARC extend the KC framework by combining compression, prediction, and comprehension under algorithmic probability FROM COMPRESSION TO CREATIVITY: KOLMOGOROV COMPLEXITY IN . . .   41 and Solomonoff induction. They use evaluation protocols based on program generation and execution, helping to reduce contamination and overfitting. Yet these methods can only approximate KC, never compute it exactly. 7.4. Hybrid neurosymbolic compression frameworks. Approximation tools such as the Coding Theorem Method (CTM) and Block Decomposition Method (BDM) also aim to estimate algorithmic probability. They work well for short strings and local structures but face high computational costs that limit scalability. Future studies will likely merge neural predictors with program-space search to build more efficient KC approximations. 7.5. From compression to simulation and design. KC also highlights the link between compression and simulation. Any good compressor can act as a generator, suggesting that compact programs can serve as simulators for creative synthesis. However, current systems often replicate correlations instead of discovering mechanisms, leaving the path to true abductive compression open. 7.6. Ethical and epistemic limits. Finally, the incomputability of KC sets an epistemic limit. No algorithm can prove that its explanation is minimal, meaning that both human and artificial creativity are “provably approximate.” This raises ethical and philosophical questions about authorship, understanding, and how to evaluate AI-generated works. Table 3 summarizes these open challenges, linking each to its relationship with Kolmogorov Complexity, current limitations, and future research directions. These challenges mark the frontier of research connecting algorithmic information theory and artificial creativity. Addressing them will require both theoretical innovation and computational pragmatism. Progress will depend on how effectively we approximate incomputable ideas within finite systems. 42   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Table 3. Open Challenges in Kolmogorov Complexity and Generative AI Challenge Relation to Kolmogorov Complexity Limitation / Open Question Future Direction Measuring creativity beyond compression KC = shortest description principle KC incomputable; code length ≠ originality Combine compression with mechanistic adequacy Benchmarking creative reasoning (KT) KC as programlength measure LLMs fail on program-level abstraction Develop contaminationfree creative benchmarks SuperARC & contamination resistance Solomonoff induction + algorithmic probability Prevents leakage but still approximate Program-based, execution-verified evaluation Hybrid neurosymbolic frameworks CTM, BDM as KC approximations High cost, not scalable Neural + symbolic hybrid abductive compression From compression to simulation & design Compressor ↔ Generator equivalence Models copy patterns, not mechanisms Use compressed programs as simulators for design Ethical & epistemic limits KC incomputable, only lower bounds No guarantee of “shortest program” Safe approximations, interpretability, credit 8. Conclusion This chapter explored the connection between compression, prediction, and creativity through the lens of Kolmogorov complexity. It showed that all intelligent systems—whether human or artificial—rely on finding short, efficient representations of information. A model that predicts well must also compress effectively, and a model that compresses deeply must, in some sense, understand what it describes. Kolmogorov complexity provides a rigorous way to express this idea. It defines information as the length of the shortest program that can generate a given object. Although this quantity is incomputable, it serves as an ideal benchmark for all learning systems. Every compression algorithm, neural network, or symbolic model can be viewed as an approximation of this principle, striving to capture structure with minimal redundancy. FROM COMPRESSION TO CREATIVITY: KOLMOGOROV COMPLEXITY IN . . .   43 Modern generative models, such as large language models and diffusion systems, have brought this vision closer to practice. Their training objectives, often based on cross-entropy or likelihood, directly minimize expected code length. However, they achieve statistical compression, not algorithmic compression. They predict efficiently but rarely infer the true generative mechanisms behind their data. Bridging this gap requires models that can generate concise, reusable programs rather than long lists of parameters. Recent developments, including the Kolmogorov Test (KT) and the SuperARC framework, move toward this goal. KT measures intelligence as the ability to produce short, functional code that reproduces observed sequences. SuperARC extends this idea through recursive compression, connecting prediction with causal reasoning and explanation. Both frameworks transform Kolmogorov’s abstract theory into operational tests for creativity and understanding. Yet, several open problems remain. Measuring creativity beyond code length, designing scalable recursive compression algorithms, and building hybrid neurosymbolic systems are ongoing challenges. Ethical and epistemic questions also persist. Because no algorithm can confirm that its explanation is truly minimal, every creative act—human or artificial—is necessarily approximate. Even with these limits, the principle remains powerful. Compression links efficiency with meaning, simplicity with insight, and learning with creation. It explains why science and art both seek elegant forms, the shortest explanations are often the most beautiful ones. In this sense, intelligence itself can be seen as the art of compression—the drive to transform complexity into understanding. References Chaitin, G. J. (1966). On the Length of Programs for Computing Finite Binary Sequences. Journal of the ACM (JACM), 13(4), 547–569. https://doi. org/10.1145/321356.321363 Dai, J., Qin, X., Wang, S., Xu, L., Niu, K., & Zhang, P. (2024). Deep Generative Modeling Reshapes Compression and Transmission: From Efficiency to Resiliency. IEEE Wireless Communications, 31(4), 48–56. https:// doi.org/10.1109/MWC.005.2300574 Delétang, G., Ruoss, A., Duquenne, P. A., Catt, E., Genewein, T., Mattern, C., Grau-Moya, J., Wenliang, L. K., Aitchison, M., Orseau, L., Hutter, M., & Veness, J. (2023). Language Modeling Is Compression. 12th International Conference on Learning Representations, ICLR 2024. https://arxiv.org/ pdf/2309.10668 44   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Dziugaite, G. K., & Roy, D. M. (2017). Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data. Uncertainty in Artificial Intelligence - Proceedings of the 33rd Conference, UAI 2017. https://arxiv.org/pdf/1703.11008 Grünwald, P. D. . (2007). The minimum description length principle. 703. Hernández-Espinosa, A., Ozelim, L., Abrahão, F. S., & Zenil, H. (2025). SuperARC: An Agnostic Test for Narrow, General, and Super Intelligence Based On the Principles of Recursive Compression and Algorithmic Probability. https://arxiv.org/pdf/2503.16743v4 Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., … Sifre, L. (2022). Training Compute-Optimal Large Language Models. Advances in Neural Information Processing Systems, 35. https://arxiv.org/pdf/2203.15556 Hutter, M. (2006). On the Foundations of Universal Sequence Prediction. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 3959 LNCS, 408– 420. https://doi.org/10.1007/11750321_39 Kingma, D. P., & Welling, M. (2013). Auto-Encoding Variational Bayes. 2nd International Conference on Learning Representations, ICLR 2014 - Conference Track Proceedings. https://doi.org/10.61603/ceas.v2i1.33 Kolmogorov A. (1965). Three Approaches To The Quantitative Definition of Information. Alexander.Shen.Free.Fr. http://alexander.shen.free.fr/library/ Kolmogorov65_Three-Approaches-to-Information.pdf Levin, L. A. (1984). Randomness conservation inequalities; information and independence in mathematical theories. Information and Control, 61(1), 15–37. https://doi.org/10.1016/S0019-9958(84)80060-1 Li, M., & Vitányi, P. (2008). An Introduction to Kolmogorov Complexity and Its Applications. https://doi.org/10.1007/978-0-387-49820-1 Rissanen, J. (1978). Modeling by shortest data description. Automatica, 14(5), 465–471. https://doi.org/10.1016/0005-1098(78)90005-5 Schmidhuber, J. (2008). Driven by Compression Progress: A Simple Principle Explains Essential Aspects of Subjective Beauty, Novelty, Surprise, Interestingness, Attention, Curiosity, Creativity, Art, Science, Music, Jokes. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 5499 LNAI, 48–76. https://doi.org/10.1007/978-3-642-02565-5_4 FROM COMPRESSION TO CREATIVITY: KOLMOGOROV COMPLEXITY IN . . .   45 Shannon, C. E. (1948). A Mathematical Theory of Communication. Bell System Technical Journal, 27(3), 379–423. https://doi.org/10.1002/J.1538-7305.1948. TB01338.X Solomonoff, R. J. (1964). A formal theory of inductive inference. Part I. Information and Control, 7(1), 1–22. https://doi.org/10.1016/S00199958(64)90223-2 Turing, A. M. (1937). On Computable Numbers, with an Application to the Entscheidungsproblem. Proceedings of the London Mathematical Society, s2-42(1), 230–265. https://doi.org/10.1112/PLMS/S2-42.1.230 Vitányi, P. M. (2006). Meaningful information. IEEE Transactions on Information Theory, 52(10), 4617–4626. https://doi.org/10.1109/ TIT.2006.881729 Witten, I. H., Neal, R. M., & Cleary, J. G. (1987). Arithmetic coding for data compression. Communications of the ACM, 30(6), 520–540. shttps://doi. org/10.1145/214762.214771 Xuyang, S., Luo, X., Cheng, T., Chu, Z., Li, H., wang, ziqi, Huang, S., Zhu, Q., Wang, Q., Zhang, X., Zhou, S., & Che, W. (2025). Is Compression Really Linear with Code Intelligence? https://arxiv.org/pdf/2505.11441 Yoran, O., Zheng, K., Gloeckle, F., Gehring, J., Synnaeve, G., & Cohen, T. (2025). The KoLMogorov Test: Compression by Code Generation. 13th International Conference on Learning Representations, ICLR 2025, 24471– 24501. https://arxiv.org/pdf/2503.13992 Zenil, H., Soler-Toscano, F., Delahaye, J. P., & Gauvrit, N. (2015). Twodimensional Kolmogorov complexity and an empirical validation of the Coding theoremmethod by compressibility. PeerJ Computer Science, 2015(9), e23. https://doi.org/10.7717/PEERJ-CS.23/SUPP-1 Zhang, C., Bengio, S., Hardt, M., Recht, B., & Vinyals, O. (2016). Understanding deep learning requires rethinking generalization. Communications of the ACM, 64(3), 107–115. https://doi.org/10.1145/3446776 47 CHAPTER III CONVERSATIONAL AI AND QUESTION ANSWERING SYSTEMS: ARCHITECTURES, TRAINING CHALLENGES AND SOLUTIONS Ramazan KATIRCI 1 & Hilal ÇELİK 2 & Taha OĞUZ3 1(Prof. Dr.), Sivas University of Science and Technology, Faculty of Engineering and Natural Sciences, Department of Computer Engineering, Sivas/Turkey E-mail: ramazankatir[email protected] ORCID: 0000-0003-2448-011X. 2(Res. Asst.), Sivas University of Science and Technology, Faculty of Engineering and Natural Sciences, Department of Computer Engineering, Sivas/Turkey E-mail: [email protected] ORCID: 0000-0001-5428-3411. 3(Res. Asst.), Sivas University of Science and Technology, Faculty of Engineering and Natural Sciences, Department of Metallurgical and Materials Engineering, Sivas/Turkey E-mail: [email protected] ORCID: 0000-0003-4447-645X. 1. Introduction Artificial intelligence (AI) refers to the development of computational systems capable of performing tasks that typically require human intelligence, such as learning, reasoning, and natural language understanding (Winston, 2017). Today, AI is applied across a wide range of domains, including healthcare (Tuncer et al., 2024), finance (Çınar & Çelik, 2022), computational social science (Çelik & Çınar, 2021), quantum machine learning (Aasim et al., 2024; Katırcı & Oğuz, 2023) and education (Celik & 48   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Cinar, 2021). Within this broad landscape, conversational agents—most notably chatbots—have gained increasing prominence in recent years, evolving from early rule-based systems such as ELIZA (1966) (Adamopoulou & Moussiades, 2020a) to modern machine learning and deep learning models that underpin state-of-the-art dialogue and question-answering systems (Çelik et al., 2024; Dongbo et al., 2023; Nan et al., 2024). Modern chatbot and QA systems aim to understand context, extract relevant information, and generate accurate answers, often framed as a reading comprehension task (Adamopoulou & Moussiades, 2020a) (Banitz, 2020; Choudhary & Chauhan, 2023). However, their development is hindered by learning challenges such as overfitting (Ying, 2019), underfitting (Aliferis & Simon, 2024), the bias–variance (Mavrogiorgos et al., 2024) trade-off, and hallucinations (Bang et al., 2024), all of which compromise reliability and generalization. In response to these challenges, researchers commonly employ strategies such as regularization methods—including early stopping (Ying, 2019), and weight decay (Xie et al., 2023), L1 and L2 regularization (Xu et al., 2010), dropout (Katırcı & Çelik, 2024; Srivastava et al., 2014)—as well as the design of domain-specific architectures that balance efficiency with robustness(R. Zhang et al., 2023). 2. Question Answering (QA) QA is an NLP task that focuses on understanding text and enables machines to provide precise answers to natural language questions (Hao et al., 2022a; İşlek et al., 2024), where QA systems aim to generate responses that resemble natural language for user queries (Çelik et al., 2024). Therefore, QA primarily addresses factoid questions, providing concise and accurate responses to fulfill users’ information needs in real-world applications (Jurafsky & Martin, 2023). These systems range from providing simple yes/no responses to generating complex results synthesized from multiple data sources, through processes such as question analysis, answer retrieval, and ranking (Voorhees & Tice, 2000). Depending on the task definition and the available resources, question answering systems can vary significantly in their capabilities, spanning from basic factbased answering to advanced systems that integrate and synthesize information from diverse sources (Hao et al., 2022a). Many critical downstream tasks, such as QA and paraphrase identification, require an understanding of the relationship between two sentences (Katırcı & Çelik, 2025; Mollá & Vicedo, 2007; Yin & Schütze, 2015), which can be CONVERSATIONAL AI AND QUESTION ANSWERING SYSTEMS: ARCHITECTURES . . .   55 4.1.1.1. Overfitting Overfitting is a fundamental issue in supervised machine learning that prevents a model from generalizing well to unseen data (Liu et al., 2023; Ying, 2019). This phenomenon occurs when model components are evaluated against an inappropriate reference distribution. In such cases, modeling algorithms tend to iteratively select the best among several candidate components and subsequently test whether this component should be incorporated into the model, which may lead to overfitting (Cohen & Jensen, 1997). Overfitting can be understood as the model learning “noise” in the data, meaning it captures idiosyncrasies of the training set that do not exist in the overall population. Machine learning and AI methods, as well as modeling systems or data science protocols, inherently have a propensity to overfit, particularly when model complexity is high relative to the amount of training data. As a result, the model may achieve low training error but fail to generalize to unseen data (Aliferis & Simon, 2024). An overfitted model often achieves near-perfect performance on the training set but fails to adapt to variations in unseen data, leading to poor generalization. This occurs because the model captures dataset-specific noise and peculiarities instead of the true underlying patterns (Ghojogh & Crowley, 2023b). The primary causes of overfitting include excessive model complexity, limited training data, and insufficient use of regularization techniques. In such cases, models with far more parameters than necessary tend to memorize training examples rather than learning generalizable representations. To mitigate this issue, strategies such as early stopping, regularization, cross-validation, and dataset expansion can be applied, all of which encourage the model to focus on patterns that extend beyond the training set (Ying, 2019). 4.1.1.2. Underfitting Underfitting is a fundamental limitation of model training that occurs when a learning algorithm lacks the capacity to capture the essential structure of the data. Consequently, underfitted models struggle to effectively learn data patterns, resulting in higher generalisation error and suboptimal utilisation of available data (Aliferis & Simon, 2024). A related concept is underestimation, whereby predictions consistently underestimate true values, often due to limited model complexity or inadequate training (Cunningham & Delany, 2021). Such underestimation can have practical implications in high-performance computing (HPC), where ignoring the discrepancy between predicted and actual job runtimes may complicate scheduling, prolong prediction or rescheduling processes, and lead to inefficient utilization of computational resources (Yao et al., 2021). 56   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Closely related to underfitting, underestimation occurs when a model systematically predicts values lower than the true outcomes. It often results from limited model capacity and incomplete data coverage, which prevent the model from capturing the full data distribution, and can also be influenced by underfitting as well as underrepresented classes or features (Cunningham & Delany, 2021). 4.1.1.3. The Bias-Variance Tradeoff Bias refers to a systematic tendency in machine learning models to produce errors that favor or disadvantage certain individuals or groups, leading to inequitable outcomes (Cunningham & Delany, 2021; Mavrogiorgos et al., 2024). Such biases can originate from the training data, due to factors like inadequate sampling, labeling errors, or historically embedded discriminatory patterns, or from human influence during algorithm design and data collection. Bias in machine learning is generally categorized into two main types: data bias, arising from imbalances, underrepresentation, or labeling errors in the training set, and algorithmic bias, which stems from the training process, algorithm design, or human influence, increasing the risk of erroneous or unfair decisionmaking in AI systems (Babaeianjelodar et al., 2020; Le et al., 2019). Mitigating bias in machine learning requires interventions at both the data and algorithmic levels (Mikołajczyk-Bareła & Grochowski, 2023). From a data perspective, strategies such as ensuring diverse and representative datasets, resampling underrepresented classes, and careful data collection and labeling can help reduce systematic prejudice (Mavrogiorgos et al., 2024; Y. Zhang et al., 2024). Algorithmic approaches include incorporating fairness-aware learning objectives, applying regularization techniques to prevent the model from favoring specific groups, and using adversarial debiasing methods to minimize unintended correlations between protected attributes and predictions (Dip Bharatbhai Patel, 2023; Yang et al., 2023). Additionally, ongoing evaluation using fairness metrics and explainable AI techniques enables the identification and correction of residual biases, ensuring that model predictions remain equitable and reliable across different populations (Bateni et al., 2022). 4.1.2. Regularization Techniques Regularization methods are techniques designed to prevent overfitting by constraining the capacity of machine learning models, encouraging them to capture underlying patterns rather than memorizing noise or idiosyncrasies in the training data (Kukačka et al., 2017). Common approaches include L1 and L2 weight penalties (Xu et al., 2010), dropout (Katırcı & Çelik, 2024; Srivastava et CONVERSATIONAL AI AND QUESTION ANSWERING SYSTEMS: ARCHITECTURES . . .   57 al., 2014) and weight decay (Ghojogh & Crowley, 2023a), all of which have been shown to enhance generalization across various tasks and model architectures. Deep neural networks, with their multiple non-linear hidden layers, are highly expressive and capable of modeling complex input-output relationships. However, when training data is limited, some learned patterns may reflect sampling noise rather than true relationships, making regularization essential for mitigating overfitting and improving generalization (Srivastava et al., 2014). 4.1.2.1. L1 and L2 Regularization Regularization is a fundamental technique in machine learning used to constrain model parameters, reducing overfitting and improving generalization. L0 regularization is the earliest approach for feature selection, promoting extreme sparsity (Kukačka et al., 2017). L1 regularization (Lasso) offers a computationally feasible alternative by applying a penalty proportional to the absolute values of the parameters, often driving many coefficients to zero and producing sparse models that naturally perform feature selection (Xu et al., 2010). In contrast, L2 regularization (Ridge) penalizes the squared values of the parameters, shrinking them toward smaller magnitudes without forcing exact zeros, which helps prevent overfitting while retaining all features in the model (Ghojogh & Crowley, 2023a; Katırcı & Çelik, 2024; Ng, 2004). 4.1.2.2. Dropout In machine learning, various methods are employed to alleviate the challenges of overfitting and underfitting. Among these, dropout is commonly used as a regularization technique to reduce overfitting. However, when applied during the early stages of training, dropout can also contribute to mitigating underfitting by stabilizing the learning process and improving generalization (Katırcı & Çelik, 2024; Liu et al., 2023). Similarly, underestimation can be reduced by selectively filtering input data based on predefined conditions, ensuring that predictions are made only on sufficiently representative and reliable data. This approach limits the influence of sparse or low-quality data, which might otherwise cause systematic under-prediction, thereby enhancing the overall accuracy of the model’s forecasts (Yao et al., 2021). 4.1.2.3. Weight decay Weight decay is a widely used regularization technique in deep neural networks that penalizes large weights, helping to control model complexity and improve generalization (Xie et al., 2023). In non-linear activation functions 58   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . such as the hyperbolic tangent, very large positive or negative weights can place activations in highly non-linear regions, potentially leading to overfitting. By constraining the weights, weight decay maintains activations in a region that balances linearity and non-linearity, allowing the network to capture complex patterns without excessively fitting the training data. However, care must be taken, as weight decay can interact with large gradient norms in unexpected ways, affecting optimization dynamics. Therefore, monitoring training behavior and adjusting hyperparameters appropriately is essential to avoid unintended consequences while leveraging weight decay for improved generalization (Ghojogh & Crowley, 2023a). 4.1.2.4. Early Stopping Early stopping is a regularization technique used in machine learning to prevent overfitting. It works by monitoring model performance on a validation set and halting training once the performance stops improving or begins to degrade, thereby avoiding excessive noise-learning and unnecessary training iterations (Ying, 2019). By frequently evaluating models during training, it has been empirically observed that worse-performing models can often be distinguished from better ones early in training, which motivates the use of early stopping strategies (Dodge et al., 2020). However, stopping the training too early can lead to underfitting, as the network may fail to fully learn from clean labels while still avoiding overfitting to noisy labels (Bai et al., 2021). 4.2. Hallucination Challenges in LLMs Hallucination in Natural Language Generation (NLG) refers to the generation of content that is not faithful to the source input or contains information that is factually incorrect or fabricated (Bang et al., 2024). In other words, hallucinations occur when a model produces text that cannot be verified against the original data, introducing false or misleading information. This phenomenon undermines factuality and veracity, thereby significantly reducing the reliability of generated text (Lin et al., 2022). 4.2.1. Root Causes of Hallucination Hallucinations in generative AI systems, including LLMs such as ChatGPT, primarily arise when models are trained on large-scale unsupervised corpora. In such settings, the models learn statistical patterns from vast and diverse datasets but may generate content that does not correspond to real-world CONVERSATIONAL AI AND QUESTION ANSWERING SYSTEMS: ARCHITECTURES . . .   59 inputs, producing extrinsic statements that cannot be verified from the source. This phenomenon is influenced by the probabilistic and polysemous nature of language, which allows multiple plausible outputs for a given input, increasing the likelihood of nonfactual or misleading content (Alkaissi & McFarlane, 2023; Ioannidis et al., 2023b; Ji et al., 2023). While commonly viewed as a shortcoming, hallucinations can also be considered an emergent byproduct of language’s inherent characteristics. Mitigating these issues requires rigorous training and evaluation procedures, supported by diverse and representative datasets, to enhance factual grounding and ensure reliable outputs (Widdows, Aboumrad, Kim, Ray, et al., 2024). 4.2.1.1. Data Limitations and Noise Hallucinations can arise even when a model is trained on factually correct data, triggered by data limitations, complex learning processes, and the model’s inference mechanisms. Such errors are more likely when the training data contains gaps, is insufficient, or lacks diversity (Kalai et al., 2025). These observations highlight that uncertainties and irregularities in the learning process—often amplified by noise—are key factors contributing to hallucinations. Techniques such as adaptive noise injection can mitigate these effects, reducing the likelihood of the model generating erroneous or unfounded outputs (Khadangi et al., 2025). 4.2.1.2. Architectural Limitations and Overconfidence Large language models exhibit architectural limitations that restrict their ability to manage uncertainty and perform verification, increasing the likelihood of producing erroneous outputs. Combined with tendencies toward overconfidence, these structural constraints give rise to intrinsic hallucinations, causing the models to generate information with high certainty even when it is not factually grounded (Gumaan, 2025; Kalai et al., 2025). 4.2.1.3. Prompt Ambiguity and Context Limitations Ambiguous or poorly specified prompts can lead LLMs to generate hallucinations, as the model attempts to fill in gaps with plausible but incorrect information. Hallucinations arise when a model’s outputs are not grounded in the prompt or training data; therefore, the quality and specificity of the prompt directly affect the risk of Hallucination (Kalai et al., 2025; Tonmoy et al., 2024). 60   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . 4.2.2. Mitigation Strategies for Reducing Hallucination Mitigating hallucinations in large language models requires a multifaceted approach, combining high-quality training data, factually augmented datasets, architectural enhancements for output verification, and access to external knowledge sources to ensure more accurate and reliable model outputs (Suzuoki & Hatano, 2024). Additionally, employing diverse and representative training datasets, rigorous evaluation procedures, and monitoring techniques such as human review or anomaly detection further reduces the likelihood of generating hallucinated content (Alkaissi & McFarlane, 2023). 4.2.2.1. Reinforcement Learning with Human Feedback (RLHF) Among hallucination mitigation techniques in large language models, a variety of methods are employed, including ensemble approaches, retrievalbased strategies, advanced decoding and reasoning techniques, RLHF, prompt engineering, and post-processing or human-in-the-loop mechanisms. Ensemble methods (Suzuoki & Hatano, 2024) combine the outputs of multiple models to produce a consensus, reducing the likelihood of generating incorrect or fabricated content. 4.2.2.2. Retrieval-Augmented Generation (RAG) Retrieval-Augmented Generation (RAG) (Widdows, Aboumrad, Kim, & Ray, 2024). RAG grounds outputs in external knowledge bases, ensuring factual support beyond the model’s internal parameters. Additional approaches, such as improved decoding strategies like contrastive search, safety RLHF (Suzuoki & Hatano, 2024), parameter editing, and reasoning-based methods like Chain-ofThought, also promote self-verification and alignment with human judgment (Ji et al., 2023). 4.2.2.3. Prompt Engineering and Optimization Prompt engineering and pre-publication summary editing strengthen factual grounding, while post-processing and human-in-the-loop mechanisms provide final validation to ensure output accuracy (Ioannidis et al., 2023a). In addition, prompt tuning, which optimizes learned soft prompts and aligns them with task-specific contexts, helps mitigate hallucination risks by guiding the model to produce more accurate and contextually appropriate outputs (Tonmoy et al., 2024). Collectively, these strategies enhance the reliability and truthfulness of LLM-generated content in practical applications. CONVERSATIONAL AI AND QUESTION ANSWERING SYSTEMS: ARCHITECTURES . . .   61 5. Conclusion This study emphasizes that robust conversational AI requires balancing two critical aspects: optimizing training dynamics to address overfitting, underfitting, and the bias-variance trade-off, as well as directly combating hallucinations with advanced mitigation techniques. Progress in question-answering and chatbot systems depends on integrating reliable generalization strategies with effective hallucination control to ensure accuracy and adaptability. Looking ahead, next-generation systems must enhance their contextual, emotional, and multimodal understanding to enable more natural and inclusive interactions. Supporting these capabilities will necessitate efficient, domainspecific architectures that address generalization challenges while incorporating anti-hallucination strategies tailored to specific domains. These efforts will pave the way for multilingual, culturally adaptive, and ethically aligned AI systems that are scalable, user-centric, and effective across global applications. References Aasim, M., Katırcı, R., Acar, A. Ş., & Ali, S. A. (2024). A comparative and practical approach using quantum machine learning (QML) and support vector classifier (SVC) for Light emitting diodes mediated in vitro micropropagation of black mulberry (Morus nigra L.). Industrial Crops and Products, 213, 118397. https://doi.org/10.1016/j.indcrop.2024.118397 Abu Shawar, B., & Atwell, E. (2007). Chatbots: Are they Really Useful? Journal for Language Technology and Computational Linguistics, 22(1), 29–49. https://doi.org/10.21248/jlcl.22.2007.88 Adamopoulou, E., & Moussiades, L. (2020a). An Overview of Chatbot Technology. In IFIP Advances in Information and Communication Technology: Vol. 584 IFIP (Issue June). Springer International Publishing. https://doi. org/10.1007/978-3-030-49186-4_31 Adamopoulou, E., & Moussiades, L. (2020b). Chatbots: History, technology, and applications. Machine Learning with Applications, 2, 100006. https://doi.org/10.1016/J.MLWA.2020.100006 Al-Amin, M., Ali, M. S., Salam, A., Khan, A., Ali, A., Ullah, A., Alam, M. N., & Chowdhury, S. K. (2024). History of generative Artificial Intelligence (AI) chatbots: past, present, and future development. http://arxiv.org/abs/2402.05122 Aliferis, C., & Simon, G. (2024). Overfitting, Underfitting and General Model Overconfidence and Under-Performance Pitfalls and Best Practices in Machine Learning and AI. https://doi.org/10.1007/978-3-031-39355-6_10 62   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Alkaissi, H., & McFarlane, S. I. (2023). Artificial Hallucinations in ChatGPT: Implications in Scientific Writing. Cureus, 15(2), 2–5. https://doi. org/10.7759/cureus.35179 Babaeianjelodar, M., Lorenz, S., Gordon, J., Matthews, J., & Freitag, E. (2020). Quantifying Gender Bias in Different Corpora. The Web Conference 2020 - Companion of the World Wide Web Conference, WWW 2020, September, 752–759. https://doi.org/10.1145/3366424.3383559 Bai, Y., Yang, E., Han, B., Yang, Y., Li, J., Mao, Y., Niu, G., & Liu, T. (2021). Understanding and Improving Early Stopping for Learning with Noisy Labels. Advances in Neural Information Processing Systems, 29(NeurIPS), 24392–24403. Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu, T., Chung, W., Do, Q. V., Xu, Y., & Fung, P. (2024). A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity. 675–718. https://doi.org/10.18653/v1/2023.ijcnlp-main.45 Banitz, B. (2020). Machine translation: A critical look at the performance of rule-based and statistical machine translation. Cadernos de Traducao, 40(1), 54–71. https://doi.org/10.5007/2175-7968.2020v40n1p54 Bateni, A., Chan, M. C., & Eitel-Porter, R. (2022). AI Fairness: from Principles to Practice. 1–21. http://arxiv.org/abs/2207.09833 Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 2020-Decem. Celik, H., & Cinar, A. (2021). An Application on Ensemble Learning Using KNIME. 2021 International Conference on Data Analytics for Business and Industry, ICDABI 2021, July, 400–403. https://doi.org/10.1109/ ICDABI53623.2021.9655815 Çelik, H., & Çınar, A. (2021). Knime ile CRISP-DM Veri Bilimi Yöntemi Uygulaması. May 2021, 1–6. https://doi.org/10.52460/issc.2021.024 Çelik, H., Katırcı, R., & İşlek, B. (2024). Effect Of Parameters On Performance In Question-Answer Model With Simple Rnn Deep Learning Method. May. https://scholar.google.com/citations?view_ op=view_citation&hl=en&user=7b0QCpsAAAAJ&citation_for_ view=7b0QCpsAAAAJ:WF5omc3nYNoC Chen, D., Fisch, A., Weston, J., & Bordes, A. (2017). Reading Wikipedia to answer open-domain questions. ACL 2017 - 55th Annual Meeting of the CONVERSATIONAL AI AND QUESTION ANSWERING SYSTEMS: ARCHITECTURES . . .   63 Association for Computational Linguistics, Proceedings of the Conference (Long Papers), 1, 1870–1879. https://doi.org/10.18653/v1/P17-1171 Choudhary, P., & Chauhan, S. (2023). An intelligent chatbot design and implementation model using long short-term memory with recurrent neural networks and attention mechanism. Decision Analytics Journal, 9(May), 100359. https://doi.org/10.1016/j.dajour.2023.100359 Çınar, A., & Çelik, H. (2022). Regresyon Analizi Bitcoin Tahmin Uygulamasi Geliştirme. October, 1–6. https://scholar.google.com.tr/scholar?hl=tr&as_sdt=0%2C5&q=REGRESYON+ANALİZİ+BİTCOİN+TAHMİN+UYGULAMASI+GELİŞTİRME&btnG= Cohen, P. R., & Jensen, D. (1997). Overfitting Explained. Preliminary Papers of the Sixth International Workshop on Artificial Intelligence and Statistics, 115–122. Cunningham, P., & Delany, S. J. (2021). Underestimation Bias and Underfitting in Machine Learning. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 12641 LNAI(18), 20–31. https://doi.org/10.1007/978-3-030-73959-1_2 DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., … Zhang, Z. (2025a). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. 500, 1–22. DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., … Zhang, Z. (2025b). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. 500, 1–22. http:// arxiv.org/abs/2501.12948 Dip Bharatbhai Patel. (2023). Ethical AI: Addressing Bias and Fairness in Machine Learning Models for Decision-making. Journal of Computer Science and Technology Studies, 3(1), 13–17. https://doi.org/10.32996/jcsts.2021.3.1.3 Dodge, J., Ilharco, G., Schwartz, R., Farhadi, A., Hajishirzi, H., & Smith, N. (2020). Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping. http://arxiv.org/abs/2002.06305 Dongbo, M., Miniaoui, S., Fen, L., Althubiti, S. A., & Alsenani, T. R. (2023). Intelligent chatbot interaction system capable for sentimental analysis using hybrid machine learning algorithms. Information Processing and Management, 60(5), 103440. https://doi.org/10.1016/j.ipm.2023.103440 Ghojogh, B., & Crowley, M. (2023a). The Theory Behind Overfitting, Cross Validation, Regularization, Bagging, and Boosting: Tutorial. 3. 64   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Ghojogh, B., & Crowley, M. (2023b). The Theory Behind Overfitting, Cross Validation, Regularization, Bagging, and Boosting: Tutorial. 3. http:// arxiv.org/abs/1905.12787 Gumaan, E. (2025). Theoretical Foundations and Mitigation of Hallucination in Large Language Models. http://arxiv.org/abs/2507.22915 Hao, T., Li, X., He, Y., Wang, F. L., & Qu, Y. (2022a). Recent progress in leveraging deep learning methods for question answering. Neural Computing and Applications, 34(4), 2765–2783. https://doi.org/10.1007/s00521-02106748-3 Hao, T., Li, X., He, Y., Wang, F. L., & Qu, Y. (2022b). Recent progress in leveraging deep learning methods for question answering. Neural Computing and Applications, 34(4), 2765–2783. https://doi.org/10.1007/s00521-02106748-3 Higashinaka, R., Imamura, K., Meguro, T., Miyazaki, C., Kobayashi, N., Sugiyama, H., Hirano, T., Makino, T., & Matsuo, Y. (2014). Towards an open-Domain conversational system fully based on natural language processing. COLING 2014 - 25th International Conference on Computational Linguistics, Proceedings of COLING 2014: Technical Papers, 928–939. Hundertmark, D. Z. and S. (2020). Chatbots – An Interactive Technology for Personalized Communication, Transactions and Services. IADIS International Journal on WWW/Internet, 15(February 2018), 96–109. Ioannidis, J., Harper, J., Quah, M. S., & Hunter, D. (2023a). Gracenote. ai: Legal Generative AI for Regulatory Compliance. CEUR Workshop Proceedings, 3423, 20–31. https://doi.org/10.2139/ssrn.4494272 Ioannidis, J., Harper, J., Quah, M. S., & Hunter, D. (2023b). Gracenote . ai : Legal Generative AI for Regulatory Compliance. İşlek, B., Katırcı, R., & Celik, H. (2024). Enhancing Question Answering Systems Through Optimal Hyperparameter Tuning in GRU. 8th International Artificial Intelligence and Data Processing Symposium, IDAP 2024, September. https://doi.org/10.1109/IDAP64064.2024.10710732 Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55(12). https://doi.org/10.1145/3571730 Jurafsky, D., & Martin, J. H. (2023). Speech and Language Processing An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. In The 1870 Ghost Dance. https://doi.org/10.2307/j. ctt1djmg7w.13 SOFTWARE DEVELOPMENT SUPPORTED BY GENERATIVE ARTIFICIAL . . .   71 system operated by a big language model—a descendant of GPT-3—tuned into billions of lines of publicly shared code from GitHub repositories. It was found to be able to coexist with almost any programming language and framework. Similarly, Amazon introduced CodeWhisperer in 2022 that also provided autocode suggestions in virtually every programming language (The Past, Present and Future of AI Coding Tools | TechTarget, n.d.). Thus, code completion has evolved from simple rule-based suggestions into a system capable of generating entirely new lines of code, drawing on the learned knowledge of generative AI models (ML-Enhanced Code Completion Improves Developer Productivity, n.d.). In traditional systems, suggestions were constrained by the compiler or the rules of the programming language; today, however, models trained on massive datasets provide recommendations that are far more flexible and powerful. Today, the most successful approach to code generation and completion relies on large language models (LLMs) built upon the Transformer architecture (ML-Enhanced Code Completion Improves Developer Productivity, n.d.). Transformer models acquire long-range dependencies among tokens within a sequence from the self-attention mechanism. It enables them to learn distant variable references, block structures, or call relationships even in structured and context-sensitive data such as source code. Generative models are first pretrained on huge text corpora and then fine-tuned on code-related data so that they can learn the grammar and patterns of programming languages. For example, the OpenAI Codex model was developed by further training the GPT model initially trained on general text on gigantic amounts of publicly available code aggregated from GitHub repositories. As a result, Codex has learned to code in programming languages such as Python, JavaScript, and Go (Chen et al., 2021). These Transformer-based models process code by segmenting it into smaller units such as subword tokens in the same way natural language is handled, and function as language models that predict the next token. As input, they typically consider the code context surrounding the cursor, often spanning ~1,000 or more tokens, and generate one or several lines of completion suggestions as output. During this process, techniques such as beam search are employed in the background to explore multiple possible continuations, from which the most probable candidates are presented to the developer (ML-Enhanced Code Completion Improves Developer Productivity, n.d.). Generative code-specialized models have a lot to offer over generalpurpose language models. When trained and fine-tuned from a similarly sized model only on programming data, it performs reliably better on code generation 72   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . tasks than its general-purpose version. Recent studies have shown that codespecific models (such as Codex, CodeLlama, and AlphaCode) are much better than the same-sized general models because of their ability to learn programming syntax, structural components, and recurring patterns of library use (Husein et al., 2025). This advantage comes from the point that code-specific models have the ability to learn syntactic rules, programming language structure properties, and typical usage patterns of libraries. For instance, a model dedicated to Python can learn indentation rules or typical usage patterns of particular libraries and therefore produce more accurate completions. Thanks to the Transformer model, this type of model can contextually “remember” how a variable defined early in a function is used several hundred lines later. This is actually the same strategy that human programmers rely on when they consider code. 3. Principles of Generative Code Production Across Different Programming Languages Although generative AI models are language-agnostic in implementation, they operate on the same principles when generating code for any programming language. In essence, across all languages, the model’s objective is to produce the most probable continuation of the code based on context surrounding it—variable names, declared functions, comments, and syntactic styles of the language. Large models are capable of learning numerous programming languages under one model. For example, Google’s in-house research into code completion revealed that training a single Transformer model on eight languages (C++, Java, Python, Go, TypeScript, Kotlin, Dart, and Proto) gave decent suggestions across all of them without requiring separate models for each language (ML-Enhanced Code Completion Improves Developer Productivity, n.d.). This is due to the fact that many programming languages share a set of common features: all of them incorporate similar concepts such as control structures, function calls, and variable assignments. But the performance of these models would vary across programming languages. This is largely due to the quality and availability of training data for each language. For widely used languages such as Python, Java, and JavaScript, the vast volume of open-source code on the internet exposes the models to good examples, thereby improving their performance. On the other hand, for less popular or less readable languages (e.g., Perl), models are more perplexed, indicating more difficulty making accurate predictions. Indeed, it has been found that while Perl perplexity values were high consistently, those for Java, with its SOFTWARE DEVELOPMENT SUPPORTED BY GENERATIVE ARTIFICIAL . . .   73 more controlled syntax, were comparatively lower (Husein et al., 2025). This is in line with the characterization of languages such as Perl as “write-only” (easy to write, difficult to read). On statically type-checked and syntactically strict languages such as C# and Java, models have more accurate predictions than for dynamic languages, since the type information limits the universe of possible completions and hence model uncertainty is lower. Conversely, in dynamic languages such as Python or JavaScript, model proposals must necessarily be more flexible. But modern generative code models are language-agnostic, and with a good quantity of training data, they can generate fair outputs for different programming languages. 4. Training Data, Context Analysis, and Suggestion Strategies in Code Completion Systems Training data plays a critical role in the success of generative code models. These models are typically trained on billions of lines of code collected from large-scale open-source repositories, such as those hosted on GitHub (AI Copilot Tools for Developers - Overview & Comparison of Tools | Zartis, n.d.). This data consists of a huge number of programming languages, the use of various libraries, and an endless variety of coding patterns. Unfiltered code data is not clean or free of bugs; during training, models may also learn from buggy examples in the corpus. Filtering mechanisms such as skipping code that cannot compile or fail tests are therefore employed by some advanced systems to improve the quality of training data. With code completion, context is the most critical input to the system. A highly advanced completion engine takes into account the previous file lines the cursor is in, function and class declarations, and, where possible, relevant declarations in other files in the project. For example, programs like Copilot will typically provide the model with a window of about hundreds to thousands of tokens around the cursor position as input, and then request the model to generate most likely the rest of the code. Here, the model considers what identifiers have been defined, what a method is intended to do (inferred from comments or its name), and what libraries are imported, in an effort to give a pertinent suggestion. With further development of context analysis, the suggestions also become better, since the model gets more cognizant of “where it is” and “what is possible.” Some systems extend the language model by incorporating abstract syntax trees (ASTs) or symbol tables. For instance, Google’s hybrid code completion by Google utilized a traditional compiler-based semantic engine combined with a 74   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Transformer model: while usual one-token completions and ML-based longer completions were calculated in parallel, the compiler ensured correctness of the ML completions and eliminated incorrect completions (e.g., syntax errors or type mismatches) prior to presenting them to the user (ML-Enhanced Code Completion Improves Developer Productivity, n.d.). In terms of suggestion strategies, some tools present multiple options at once (e.g., displaying several completion candidates in the IDE interface). In such cases, the model generates the top-ranked suggestions according to probability. While evaluations often show that the first suggestion is most frequently accepted, the second or third alternatives may sometimes better match the developer’s intent, which is why they are also presented to the user. Another modern strategy involves generating multi-line completions or even entire functions. For example, when a comment line states, “This function sorts the array,” the model can produce not only several lines of code but, in some cases, the complete function body in a single step (From vi to AI: The Incredible Evolution of Coding Tools, n.d.). In such cases, the model emits a long sequence of tokens in a single step and presents it to the developer for review and approval. In essence, code completion systems attempt to comprehend the present context as accurately as they can and replicate code written in similar contexts they have learned earlier. In the process, they can also integrate compilers or syntax checkers to produce results that are more dependable. Lastly, training data diversity and quality, context window width, and ranking and filtering policies for the suggestions are the key factors determining the success of such systems. 5. Advantages and Limitations of Generative AI in Code Generation The integration of generative AI models into software development processes provides significant advantages: Speed and Productivity: Code completion tools greatly accelerate routine coding tasks. Both academic studies and industrial experiments demonstrate that AI-assisted programming enables developers to complete specific tasks significantly faster. For instance, in one experiment, a group of developers using GitHub Copilot completed a designated programming task 55% faster compared to the group working without such assistance (Peng et al., 2023). Similarly, in internal measurements, Google reported that the use of single-line ML-based code completions reduced the time spent between compilation and testing cycles by approximately 6% (ML-Enhanced Code Completion Improves Developer Productivity, n.d.). Overall, AI assistants help developers use their time more efficiently by rapidly suggesting repetitive and tedious code fragments. SOFTWARE DEVELOPMENT SUPPORTED BY GENERATIVE ARTIFICIAL . . .   75 Fewer Errors and Higher Quality: Intelligent suggestion systems help reduce syntax errors and minor programming mistakes by providing structurally correct and compilable code snippets. With the introduction of tools such as IntelliSense, reductions in typographical errors and improvements in code quality have been observed, leading to more consistent and reliable code (From vi to AI: The Incredible Evolution of Coding Tools, n.d.). Generative models go one step further by, for instance, learning appropriate usage patterns of an API and suggesting them in turn. Therefore, even a newbie developer can be made to utilize the library pertaining to it in the most correct manner. This aspect avoids errors even at the coding stage, prior to when they happen in later stages of development. Instant Knowledge Access: Because AI-based tools are founded on vast knowledge bases, they enable developers to learn how to do something without referencing the manual. For example, if a developer wants to sort a list in Python, the assistant will infer the purpose and trigger typical usage patterns. The developer can receive the desired code without switching context by reducing the need for frequent web searching or reference looks-up (ML-Enhanced Code Completion Improves Developer Productivity, n.d.). This phenomenon creates an embedded documentation effect within the code, meaning that generative tools effectively act as a kind of knowledge assistant. Creativity and Focus on Complex Tasks: Generation of dull and repetitive code by AI systems releases the cognitive space of programmers to engage in more imaginative activities. An example is where AI does the mundane job of extracting data from a web form and placing it within a model, the programmer can then focus on coding the business logic or enhancing performance. This division of labor enables the programmers to devote more time to problemsolving on a higher level and system design, thus more development efficiency and innovation (From vi to AI: The Incredible Evolution of Coding Tools, n.d.). In some cases, a generative model may suggest an unconventional solution or an alternative code fragment, which can provide developers with a different perspective on the problem at hand. Alongside these advantages, the limitations and risks associated with code generation through generative AI should not be overlooked: Misleading or Erroneous Clues: The code generated by an AI model can be syntactically correct but semantically wrong. Since the model is probabilistic prediction-based, the code indicated by the model is not always the most accurate solution to a problem. For example, it may give a solution that doesn’t consider an edge case, or it may even invoke functions that don’t exist-a scenario known 76   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . as hallucination. This is why code produced by AI should always be inspected by a developer; in fact, the makers of generative AI systems specifically state that this kind of code should not be executed without human observation and verification (Introducing Codex | OpenAI, n.d.). Overreliance and Loss of Understanding: When developers place excessive trust in AI tools, they may gradually risk losing their ability to understand fine-grained details of the code or to solve problems independently. An academic study notes that developers who work continuously with AI sometimes perceive the tool as a black box, occasionally accepting the suggested code without fully internalizing why it works. Over the long term, this can create a form of cognitive debt, in which the developer becomes increasingly dependent on the system’s generated code and gradually loses mastery over their own work (Husein et al., 2025). Moreover, the code suggested by AI is not always the most minimal or maintainable solution; it may occasionally produce excessive or unnecessarily complex code a phenomenon often referred to as code bloat. Therefore, while AI contributes to faster development, it also imposes on developers the responsibility to refine and fully understand the generated code. Data Privacy and Licensing Concerns: Since generative models are trained using publicly available data, the proposed code by them will sometimes have snippets that raise concerns regarding copyright or licensing. For example, Copilot has found itself surrounded by controversy in some cases, where it was viewed to propose code very similar to parts that were under the GNU GPL license. This issue comes from the nature of training data and is indicative of potential risks in terms of license compliance for ongoing projects. Similarly, confidentiality in sensitive or proprietary code is a matter of concern too. In such cases, it is all the more important to construct models which have never seen such data, or to deploy on-prem solutions (such as the on-prem version of Tabnine) to shield against such risks (AI Copilot Tools for Developers - Overview & Comparison of Tools | Zartis, n.d.) Performance and Resource Intensity: As massive language models possess hundreds of millions or even billions of parameters, offering real-time code completion service could be very intensive in terms of computational resource and memory. Cloud-based products are network-dependent and impose latency, while on-premises products demand top-tier hardware resources. Moreover, very large models can take the entire project context into account, but it is expensive to process such a large input each time. Therefore, a balance between response latency and model capacity is typically unavoidable in practice. As a SOFTWARE DEVELOPMENT SUPPORTED BY GENERATIVE ARTIFICIAL . . .   77 case in point, Google stated that a model of approximately 0.5 billion parameters is an appropriate compromise between low latency and sufficient accuracy (ML-Enhanced Code Completion Improves Developer Productivity, n.d.). Responsibility and Debugging: If an error in coding has been produced by AI, it may be difficult for coders to understand why code has been produced in a specific manner since its logic is built on the internal reasoning of the model. While a coder may usually trace code logic that he himself has created to debug errors, AI-generated code may contain uncertain or unstable traces, and therefore fault diagnosis becomes more difficult. In such circumstances, the developer will have to manually repair the error or even recode completely. So, while production is accelerated through AI software, developers are not exempted but have the responsibility of verifying, authenticating, and correcting the generated code squarely on their shoulders. 6. Academic Studies, Case Analyses, and Evaluation Metrics In recent years, numerous academic studies have been conducted and various industrial case analyses reported to evaluate and compare the success of generative code models. A 2021 study on OpenAI’s Codex model demonstrated remarkable progress in generating correct, executable code from natural language problem statements. In this study, Codex was evaluated on a benchmark called HumanEval, where it successfully solved 28.8% of the given Python programming problems. For comparison, the general-purpose GPT-3 model of similar size achieved 0% accuracy on the same benchmark, while GPT-J, one of Codex’s predecessors, achieved only 11.4%. Furthermore, researchers found that employing a multiplicity strategy, in which the model generated multiple candidate solutions for each problem, significantly increased success rates: when 100 different code samples were produced and tested for each problem, the solution rate rose to approximately 70% (Chen et al., 2021). These findings indicate that while large code models may not always deliver the correct solution on the first attempt, the probability of arriving at the right answer becomes very high when sufficient trials are generated. DeepMind’s AlphaCode program created a lot of buzz for generative model testing in coding competitions. In a 2022 paper, AlphaCode was reported to produce solutions from unseen programming problems that put it at the top 54% of participants in online coding competitions. This milestone indicates that AlphaCode had reached parity with the average human in the activity of competitive programming. To achieve this, AlphaCode employed a fully 78   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Transformer-based model language, which was trained on millions of lines of GitHub code before fine-tuning on the relatively small set of competition problems. When performing tasks, AlphaCode generated tens of thousands of candidate solutions, which were then automatically tested and filtered. Lastly, the system selected the top 10 promising solutions to submit, more or less mimicking the human trial-and-error process in an effort to achieve success (Competitive Programming with AlphaCode - Google DeepMind, n.d.). This approach is significant in that it demonstrates generative models’ potential not only for line-by-line code completion but also for achieving a degree of independent problem-solving capability. On the academic side, a variety of metrics have been developed to evaluate generative code models. In code generation, both text-based similarity measures and functional correctness measures are applied in combination: BLEU (BiLingual Evaluation Understudy): Originally developed to evaluate the similarity of machine translation outputs against reference sentences, the BLEU score was also widely adopted for a long time in the evaluation of code generation (Evtikhiev et al., 2017). Based on n-gram similarity between the generated code and the expected reference code, the BLEU metric yields a numerical score. However, such match-based metrics are known to be limited when applied to code, since multiple syntactically different implementations may accomplish the same functionality, resulting in low surface-level similarity despite semantic equivalence (A Dive into How Pass@k Is Calculated for Evaluation of LLM’s Coding | by Yanan Chen | Medium, n.d.). Therefore, BLEU alone cannot be considered a reliable metric for evaluating the success of code generation. CodeBLEU: This metric was proposed in response to the recognition that BLEU alone may be insufficient for code evaluation. CodeBLEU is calculated as a composite of several sub-criteria: syntactic similarity (e.g., comparison of Abstract Syntax Trees, ASTs), semantic similarity (e.g., analysis of program data-flow graphs), and surface-level textual similarity (Evtikhiev et al., 2017). In this way, the metric provides a richer assessment, capturing not only the literal similarity of the generated code to the reference solution but also its functional closeness. Test-based evaluation (Pass@k): In recent years, the most important evaluation method has been to determine whether the generated code actually executes correctly by subjecting it to testing (Chen et al., 2021; Evtikhiev et al., 2017). The pass@k metric evaluates performance by checking whether, SOFTWARE DEVELOPMENT SUPPORTED BY GENERATIVE ARTIFICIAL . . .   79 after generating up to k candidate solutions for a problem, at least one of them successfully passes the test cases (A Dive into How Pass@k Is Calculated for Evaluation of LLM’s Coding | by Yanan Chen | Medium, n.d.). For example, pass@1 represents the probability that the model produces a correct solution on its first attempt, while pass@10 measures the likelihood that at least one correct solution emerges when the model is allowed to generate ten different attempts. This metric has gained popularity through test-based benchmarks such as OpenAI’s HumanEval, where it has become a standard for evaluating the functional accuracy of code generation models (Evtikhiev et al., 2017). When calculating pass@k, a statistical formula is applied to fairly estimate the likelihood of obtaining at least one correct solution by chance, depending on the trial in which the correct output appears. This ensures that the metric reflects the probabilistic distribution of correct solutions across multiple attempts, rather than relying on raw counts alone (A Dive into How Pass@k Is Calculated for Evaluation of LLM’s Coding | by Yanan Chen | Medium, n.d.). This method is highly valuable because it provides a more accurate reflection of the realworld performance of code generation models specifically, the likelihood that a developer can solve a problem after making several attempts. Other than these, other metrics such as ROUGE-L, METEOR, and Levenshtein distance have been studied for evaluation purposes. The general consensus is that functional correctness is the ultimate measure of performance in code generation. Thus, in recent studies, metrics such as executability and unit test passing rate have appeared as the major comparison metrics for newer models (Evtikhiev et al., 2017). From an industrial perspective, the adoption of AI-assisted coding tools has been increasing rapidly. According to GitHub’s own data, developers using the Copilot extension rely heavily on its suggestions while writing code, and in some companies, it has been reported that 30–40% of newly added code characters are generated through AI recommendations (Dohmke et al., 2023; Octoverse: The State of Open Source and Rise of AI in 2023 - The GitHub Blog, n.d.). Google has reported that, within its internal development environment, the integration of an ML-based code completion system resulted in approximately 3% of newly added code characters being generated through automated suggestions (ML-Enhanced Code Completion Improves Developer Productivity, n.d.). This percentage has been observed to increase over the years, indicating that AI tools are gradually becoming the standard in software development teams. Another case study highlights that AI-powered tools particularly help beginner developers 80   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . lower their learning curve. To illustrate, interns or junior personnel employing Copilot were observed to develop a project faster on a specific task than their peers without the tool, as well as less frequently consulting external sources such as Stack Overflow (Unleash Developer Productivity with Generative AI | McKinsey, n.d.). Nevertheless, some companies restrict the use of AI in sensitive projects or establish internally trained custom models, since the privacy and security risks associated with open models have not yet been fully resolved. 7. Comparison of Human-Driven and AI-Assisted Code Generation There are clear differences between the AI-driven and the conventional human-based code generation method. In the conventional software development process, a programmer analyzes the problem, plans the solution mentally or on paper, and then hand-codes it from scratch. The human maintains full control and creativity; error checking, syntax validity, and determining the optimum solutions are all dependent on the developer’s expertise level. In contrast, in AI-assisted development, the AI and the programmer virtually collaborate. The human developer describes the problem and expresses what is to be accomplished (e.g., by adding a comment line or choosing a descriptive function name), and the AI generates a code snippet in line with this intent. The proposal is then reviewed by the developer, refining or rejecting it as needed. This loop is quite similar to a form of pair programming, with the AI partner taking care of performing the routine tasks and the human programmer retaining the final responsibility for making decisions. This fusion optimizes AI capabilities for handling repetitive and trivial code generation but not holding control in the hands of human developers for processes that require innovative problem-solving and analytical evaluation. For example, in a web application, the AI can create boilerplate functions such as form validation with ease, but it’s up to the human developer to create application-specific business logic and adjust the user interface. The strongest aspect of the AI is recycling patterns it has experienced a thousand times before, while the human being’s strongest point is in perceiving radically new things never seen before. Therefore, in the real world, the best approach is a synergistic mix in which humans and AI complementarily fill out each other’s loopholes. Recent research has demonstrated that coding with AI assistance clearly results in significant time savings, while also emphasizing the risks of relying entirely on AI. For instance, even experienced developers using Copilot have reported that they occasionally struggle to fully understand the generated code, or that during debugging they sometimes feel the need to disable the tool and DIGITAL TWIN–DRIVEN PREDICTIVE MAINTENANCE IN INDUSTRY 4.0: . . .   87 The research entails two primary domains of research work: First, the examination of the role of the generative artificial neural networks in overcoming the challenges from the inadequacy of data and enhancing the prediction capability of digital twins. Second, the exploration of the role the combination plays in enabling self-learning, self-organizations, and adaptability in Industry 4.0 environments. The research objective is the discovery of the extent that the closure of data lacunae using the generative models is impacting the development of smart, independent, and sustainable manufacturing systems. 2. Conceptual Basis and System Architecture In the context of Industry 4.0, digital transformation is reshaping not only production tools but the entire paradigm of production. Digital twin technology is pivotal to this transformation, as it enhances the traceability, predictability, and adaptability of processes through the virtual reflections of physical assets. However, the efficacy of digital twins is contingent upon the continuity of highvolume, diverse, and reliable data sets. In actual industrial practice, acquiring this kind of data is often challenging due to costs, time, and technicalities. During this phase, the use of generative artificial neural networks is vital. The networks can be used in filling missing data through the synthesis of data that cannot be distinguished from actual data, mimicking uncommon events, and improving the forecasting capability of digital twins. Therefore, the combination of digital twins and generative models makes smart production systems more consistent and adaptable through the filling of information gaps. The architecture proposed in this section clarifies how the technologies intersect, explaining how self-learning and self-reconfiguring smart production systems can be established through the preservation of consistent data, correct processes, and correct models. 2.1. The conceptual framework of digital twin technology. A digital twin can be defined as a digital replica of a physical object, process, or system that is regularly updated during its lifecycle. Digital twins differ from historical computer-aided design (CAD) or computer-aided engineering (CAE) models in that they are distinguished by their real-time linkage to a particular physical entity. This linkage allows for the capture of real-time data and ongoing self-updating. Please find below the details of the meeting. (Hartmann, 2021). This feature holds the promise of transfiguring digital twins from mere visualization tools to strategic decision support systems (Walton et al., 2024). The functions of the systems include performance assessment, scenario testing, 88   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . predicting potential failure, and exploring opportunities for improvement (Pires et al., 2023). The digital twin technology has the foundation built on the notion of the digital thread that enables the flow of data between the physical and digital worlds (Mahankali, 2025). Data captured by sensors and Internet of Things (IoT) sensors are transmitted to a centralized bank, processed, and reflected in the twin. The digital thread acts as the key element providing coherence between the physical and digital worlds (Agho et al., 2025). · Initiation of the sensor/IoT data and context information collection. · It integrates, cleanses, and does quality assurance on the stream/data lake. · The twin evaluates various conditions through real-time status forecasting and scenario analysis, thereby pinpointing potential bottlenecks and associated risks. · The measures and alerts are shown on dashboards. · Decisions are applied to devices/workflows via feedback. This enables the digital twin to reflect operating efficiency in various environmental conditions, detect areas of inefficiency, and enhance design or maintenance procedures by integrating artificial intelligence algorithms. The extracted knowledge is then propagated to concerned stakeholders in the form of dashboards, and the resulting decisions are inputs to the physical system such that a continuous feedback mechanism ensues(Central Campus Győr, Széchenyi István University, Győr, Hungary et al., 2025). Figure 1: Components of Digital Twin Systems (Digital Twins: Components, Use Cases, and Implementations Ti) Digital twins can be applied at different scales: Twins can be designed at the component (strength/energy efficiency), asset/product (component interaction and reliability), process/production (time–cost–capacity–automation DIGITAL TWIN–DRIVEN PREDICTIVE MAINTENANCE IN INDUSTRY 4.0: . . .   89 planning) and system/super system (networks, facilities, cities) scales. As the scale increases, modelling dependencies (synchronization, constraints, network effects) and integration costs increase. Twin outputs should be calibrated and validated with experimental/operational data; uncertainty measurement and sensitivity analysis should be performed. Maturity is assessed in stages: descriptive → diagnostic → predictive → prescriptive; performance is monitored using metrics. The production or process twin models are a sequence of steps taken in a digital space. The steps forecast costs, duration, and the degree to which automation can be achieved and display all the intricacies of layered architecture. The collaboration among hardware and software enables the digital twin functions. Components of the Internet of Things (IoT) such as sensors, network appliances, and edge servers collaborate to capture data from the physical world. Middleware controls the subsequent steps of integrating, processing, and validating the quality of such data. From the perspective of the software industry, analytical engines, simulative tools, and visualization panels are significant for converting raw data to meaningful information. The following essay provides a detailed survey of the significant literature on the subject (Hananto et al., 2024). While digital twins promise much in the Industry 4.0 universe, there are many limitations to how they can be utilized and implemented. First, the cost at a high level is a major problem. The necessity for advanced sensors, network hardware, data storage hardware, analysis software, and skilled labor provides a large barrier to investment, especially among small and medium businesses (Onma Enyejo et al., 2024). Second, the digital twin’s effectiveness depends on the quality and consistency of the data. When the data is incomplete, inconsistent, or incorrect, the twin cannot accurately mirror the real world and may lead to incorrect predictions and suboptimal business decisions. Additionally, integrating data across many sources pose significant technological challenges, mainly because there are few standards and many systems do not interoperate well. Third, digital twins implementation and maintenance require expertise spanning many areas. When the data engineers, machine-learning professionals, field engineers, and IT professionals don’t work well together, the viability of the twins may be threatened. This requires new workforce planning and company culture change. Large-scale digital twin deployments have shown to induce additional concerns about cybersecurity and protection of data. It must be realized that systems for collecting and processing data in real time may fall victim to security breaches (Agho et al., 2025). 90   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . 2.2. The conceptual framework of generative artificial neural networks Generative artificial neural networks (GANs) (Sengar et al., 2024) are models that learn the statistical patterns of existing datasets and generate new and consistent examples from the same distribution (Figure 2). These features mean that they meet two critical needs in digital twins: firstly, they can fill data gaps (missing sensor channels, imbalanced classes, rare fault signatures); secondly, they can safely simulate scenarios that are costly or risky to test in the real world. In this section, the contribution of generative models to digital twins will be addressed within a holistic framework through three main families: It has been demonstrated that Generative Adversarial Networks (GANs) (Wang et al., 2024) demonstrate a high level of proficiency in the generation of high realism and the creation of edge cases. By conditioning production to operating conditions and incorporating physical constraints into the model, synthetic data ensures scenario consistency and physical compatibility within the digital twin. The evaluation will not be confined to distribution similarity but will also be based on the task-based performance of the twin. GAN, in its fundamental structure, comprises a generator (G) and a discriminator (D) network that are trained in a concurrent manner with opposing objectives. The generator produces synthetic samples by mapping a lowdimensional latent input (typically Gaussian/Uniform noise, optionally enriched with conditional information) to the data space; the discriminator scores whether an input is real or generated by the generator. This configuration can be regarded as a minimax game: D’s objective is to optimize the real-synthetic distinction, whilst G’s aim is to learn the structure of the distribution to such a degree that it deceives D. Therefore, GAN’s aim is not merely to replicate ‘image similarity’ but to mimic the multidimensional relationships (joint statistics, correlations, structural constraints) within the data distribution. (S & Durgadevi, 2021) (Dan et al., 2020) (Iranmanesh & Nasrabadi, 2021) ( Figure 2). DIGITAL TWIN–DRIVEN PREDICTIVE MAINTENANCE IN INDUSTRY 4.0: . . .   91 Figure 2: Generative Adversarial Networks (GAN) (Dan et al., 2020) From an architectural perspective, the generator is predominantly a bottom-up expanding (upsampling) network. It extracts the latent vector into a high-dimensional tensor using fully connected layers. It then progresses to the target resolution using transpose convolution/upsampling + convolution blocks. Batch normalization and ReLU/LeakyReLU activations are frequently employed; Tanh/Linear is utilized at the output to scale the data appropriately. (Bhagyashree et al., 2020). The discriminator is a discriminative model that operates as follows: convolutional blocks compress the input into the feature space and ultimately produce a reality score. For images, the DCGAN-style 2D convolution is utilized; for time series such as vibration, pressure, and temperature, 1D convolutions, causal convolutions, or sequential layers are preferred. In the context of multichannel sensors, multi-head attention or multi-branch architectures can be employed to preserve cross-channel correlation. (Pu et al., 2022). The training cycle is characterized by the following sequence of steps: • The update of D is executed on a real mini-batch. • The database has been updated once more with the synthetic batch that was produced by G. • It is evident that D is fixed, and consequently, G is compelled to generate outputs that deceive D. The GAN architecture, as outlined in Figure 2, serves as a crucial adjunct to the digital twin framework that is the subject of this study. This is due to the fact 92   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . that uncommon fault signatures, imbalanced classes, and absent sensor channels are prevalent in actual production environments, thereby constraining the twin’s predictive capacity. GANs have been demonstrated to be capable of reliably filling data gaps by generating synthetic samples that preserve conditional (load, speed, temperature, flow rate, etc.) and multi-channel correlations under operating conditions. This is achieved through the competitive training of the generator-discriminator pair. The digital twin is thus enabled to test virtual scenarios that are costly or high-risk to test in the real world (e.g. extreme conditions, failure transitions). The generation of data by architectures adapted for time series (e.g. vibration, pressure, temperature) and multi-sensor streams, as opposed to images, is subject to domain constraints such as bandwidth, energy/ RMS, and spectral shape. These constraints ensure physical consistency and are selectively integrated into the calibration of the twin. It has been demonstrated that GAN-based enrichment enhances the digital twin’s early fault detection, prediction accuracy, and resilience to regime changes. Furthermore, it has been shown to fill in residual behaviors not covered by physics-based models, thereby creating a self-learning and adaptable decision support infrastructure. Consequently, the employment of GANs is regarded as a strategic imperative in all smart manufacturing scenarios characterized by data constraints. 2.3. Data Completion and Integrated Model Interaction: Theoretical Foundations and Application Dynamics of Digital Twin-GAN Integration Digital twin technology is considered to be one of the most critical components of the Industry 4.0 vision, as it facilitates the traceability, predictability, and optimization of processes by creating a digital reflection of physical systems. However, it is important to note that this potential can only be realized with a high-volume, continuous, accurate, and meaningful data flow. The collection of such data in real production environments is challenging for economic and operational reasons. A number of factors must be considered when assessing the capacity of digital twins to operate at full capacity. These include high hardware costs, the inability to observe rare events in statistically significant numbers, and deficiencies in some sensor channels. (Guo et al., 2025). At this juncture, generative artificial neural networks present a novel paradigm for addressing the fundamental data requirements of digital twins and enhancing their modelling accuracy. Generative models, especially Generative Adversarial Networks (GAN) and Variational Autoencoders (VAE), can learn the statistical distributions of data and create synthetic data that is missing, rare or imbalanced. DIGITAL TWIN–DRIVEN PREDICTIVE MAINTENANCE IN INDUSTRY 4.0: . . .   93 These models have a number of advantages over the more traditional data augmentation method. They can be used to conduct digital testing of events that are too costly or dangerous to test in the real world. Consequently, digital twins have the capacity to formulate decisions that are not solely predicated on historical data, but also on likely future scenarios. This transformation of digital twins from reactive systems to proactive decision support systems is a significant development. The methodological process structure is delineated as follows: 1. Data Analysis: The subsequent statistical examination of sensor data, time series and event logs obtained from industrial processes is undertaken for the purpose of identifying data gaps, variance deviations and pattern deficiencies. 2. Identification of Deficiencies: Problems such as underrepresented failure states, missing sensor channels, and imbalanced data classes (e.g., normal–abnormal ratio) are detected using high-dimensional data mining and clustering techniques. 3. The application of the generative model is as follows: GAN and VAEbased models utilize conditional generation logic to generate synthetic data. The generation of this data is informed by considerations of both physical constraints, such as energy consumption and temperature limits, and structural relationships, including multi-channel correlations. 4. Digital Twin Integration: The generated data is integrated into the digital twin’s training and validation sets, enhancing the model’s prediction accuracy, sensitivity, and generalization capacity. 5. Validation: The outputs generated in the simulation environment are then compared with actual production data, and model reliability is assessed using statistical validity tests. This structure is a holistic approach that considers not only data production but also the dynamic interactions between system components: • Digital Twin ↔ Generative Models: Digital twins, which reflect processes by processing real data, compensate for missing or weak data areas by collaborating with generative models. This integration enables digital twins to evolve into predictive systems operating with higher accuracy. • Data Gaps ↔ GAN/VAE: Rare failures, unexpected scenarios, or missing measurements are simulated in terms of data using GAN and VAE models, 94   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . expanding the scope of the system. In particular, the minimax game theory-based learning approach used by GANs recreates complex relationships in the data. • Decision Support ↔ Extended Data Set: Thanks to extended and balanced data sets, decision support systems operate with higher accuracy rates. These systems deliver more stable results in tasks such as regression, classification, and anomaly detection. • Industry 4.0 Ecosystem ↔ Smart Manufacturing: The combination of digital twins and generative artificial intelligence supports the system architecture that will implement the principles of self-learning, self-adaptation, and self-reconfiguration, which form the essence of Industry 4.0. Strategic Advantages and Scientific Value This combined structure enables: • High-accuracy models to be developed with limited data. • Early warning systems to be built for production faults. • Instantaneous prediction even with non-real-time (offline) data. • Model training to be made more balanced and physically consistent. In this context, GAN-supported digital twin structures not only increase efficiency but also redefine the stability, scalability, and sustainability criteria of production systems by developing a deeper epistemological understanding of the system. 3. CASE STUDY: Predictive Maintenance Based on Digital Twin The objective of this section is to establish a decision support structure that accurately represents the operational behavior of equipment on the production line through its ‘digital twin’. This will enable the early and reliable prediction of failure probability, minimize false alarms, and translate maintenance decisions into action. The objective is to develop a predictive maintenance system that is both explainable and calibrated, combining data streams enriched with productive models using signals consistent with physical laws. This will be achieved despite practical constraints such as data scarcity and class imbalance, thereby shifting the digital twin from reactive monitoring to proactive prediction. This synthetic dataset, which was published in the UCI Machine Learning Repository, consists of 10,000 examples and core process variables Table 1. These reflect DIGITAL TWIN–DRIVEN PREDICTIVE MAINTENANCE IN INDUSTRY 4.0: . . .   95 predictive maintenance patterns encountered in industry. The document exhibits a multivariate and ‘time-stamped’ character. The following essay will provide a comprehensive overview of the relevant literature on the subject (S. Matzka, 2020). Table 2 provides descriptive statistics and visualisations of input/output variables in the context of digital twins and predictive maintenance. Units for numerical variables: Air/Process [K], Rotational speed [rpm], Torque [Nm], Tool wear [min]; derivatives: DeltaTemp [K], PowerW [W], Overstrain [min·Nm]. Table 1: Dataset Variables and Characteristics Variable Name Role Type Description Units Missing Values UDI ID Integer Unique identifier for each record – No ProductID ID Categorical Product identifier (L=Low, M=Medium, H=High + serial number) – No Type Feature Categorical Product quality type (L, M, H) – No Airtemperature Feature Continuous Air temperature measured during the process K No Processtemperature Feature Continuous Process temperature measured during the process K No Rotationalspeed Feature Integer Rotational speed of the tool rpm No Torque Feature Continuous Torque applied during the process Nm No Toolwear Feature Integer Tool wear in minutes min No Machinefailure Target Integer Overall machine failure (1 if any failure mode occurs) – No TWF Target Integer Tool wear failure – No HDF Target Integer Heat dissipation failure – No PWF Target Integer Power failure – No OSF Target Integer Overstrain failure – No RNF Target Integer Random failure – No 96   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Table 2: Descriptive Statistical Values of Input Parameters Feature Mean Std Min Q1 Median Q3 Max Air temperature 300.00 2.00 295.30 298.30 300.10 301.50 304.50 Process temperature 310.01 1.48 305.70 308.80 310.10 311.10 313.80 Rotational speed 1538.78 179.28 1168.00 1423.00 1503.00 1612.00 2886.00 Torque 39.99 9.97 3.80 33.20 40.10 46.80 76.60 Tool wear 107.95 63.65 0.00 53.00 108.00 162.00 253.00 The Numerical Variable Correlation Heat Map presented in Figure 3 summarizes the linear relationship between variable pairs and reveals both physical consistency and possible multicollinearity. As expected, Air–Process shows a strong positive relationship (Process ≈ Air+10 K), PowerW shows a high positive relationship with Torque/rpm (Power = Torque·ω), and Overstrain shows a high positive relationship with Tool wear/Torque; DeltaTemp, however, is weak to moderate with other variables. These patterns validate the physical foundations of the HDF/PWF/OSF rules in the digital twin while also guiding feature selection/regularisation decisions. Figure 3: Numerical Variable Correlation Heat Map DIGITAL TWIN–DRIVEN PREDICTIVE MAINTENANCE IN INDUSTRY 4.0: . . .   103 narrowing the positive region, thereby enhancing precision, accuracy, and F1 at a specific threshold. However, this adjustment concurrently weakens PR-AUC and recall, which are metrics independent of the threshold. The operational implications of these two approaches are evident: in scenarios where false alarms incur significant financial consequences, the augmented approach may be preferable; conversely, in situations where the cost of missing failures is high, the base approach is more reliable. The following essay will provide a comprehensive overview of the relevant literature on the subject. Note: The following dataset was utilized to support and validate the theoretical framework of this study. AI4I 2020 Predictive Maintenance Dataset, (2020). UCI Machine Learning Repository. https://doi.org/10.24432/C5HS5C (S. Matzka, 2020) 5. Reference Agho, M. O., Eyo-Udo, N. L., Onukwulu, E. C., Sule, A. K., & Azubuike, C. (2025). Digital Twin Technology for Real-Time Monitoring of Energy Supply Chains. International Journal of Research and Innovation in Applied Science, IX(XII), 564–592. https://doi.org/10.51584/IJRIAS.2024.912049 Bhagyashree, Kushwaha, V., & Nandi, G. C. (2020). Study of Prevention of Mode Collapse in Generative Adversarial Network (GAN). 2020 IEEE 4th Conference on Information & Communication Technology (CICT), 1–6. https:// doi.org/10.1109/CICT51604.2020.9312049 Central Campus Győr, Széchenyi István University, Győr, Hungary, Monek, G. D., Fischer, S., & Central Campus Győr, Széchenyi István University, Győr, Hungary. (2025). Expert Twin: A Digital Twin with an Integrated Fuzzy-Based Decision-Making Module. Decision Making: Applications in Management and Engineering, 8(1), 1–21. https://doi.org/10.31181/dmame8120251181 Chandra, S., Duvinage, M., Prakash, P., Davis, P., & Timmer, S. (2024). Improving Process Yield Through Manufacturing Digital Twin Using Conditional Synthetic Data Engine (COSYNE). In U. Endriss, F. S. Melo, K. Bach, A. Bugarín-Diz, J. M. Alonso-Moral, S. Barro, & F. Heintz (Eds), Frontiers in Artificial Intelligence and Applications. IOS Press. https://doi. org/10.3233/FAIA241063 Dan, Y., Zhao, Y., Li, X., Li, S., Hu, M., & Hu, J. (2020). Generative adversarial networks (GAN) based efficient sampling of chemical composition space for inverse design of inorganic materials. Npj Computational Materials, 6(1), 84. https://doi.org/10.1038/s41524-020-00352-0 104   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Fantozzi, I. C., Santolamazza, A., Loy, G., & Schiraldi, M. M. (2025). Digital Twins: Strategic Guide to Utilize Digital Twins to Improve Operational Efficiency in Industry 4.0. Future Internet, 17(1), 41. https://doi.org/10.3390/ fi17010041 Frontoni, E., Loncarski, J., Pierdicca, R., Bernardini, M., & Sasso, M. (2018). Cyber Physical Systems for Industry 4.0: Towards Real Time Virtual Reality in Smart Manufacturing. In L. T. De Paolis & P. Bourdot (Eds), Augmented Reality, Virtual Reality, and Computer Graphics (Vol. 10851, pp. 422–434). Springer International Publishing. https://doi.org/10.1007/978-3319-95282-6_31 Groumpos, P. P. (2021). A Critical Historical and Scientific Overview of all Industrial Revolutions. IFAC-PapersOnLine, 54(13), 464–471. https://doi. org/10.1016/j.ifacol.2021.10.492 Guo, J., Zhang, Y., & Wang, W. (2025). Predicting hydro turbine failures through digital twin simulations of rare real-world data. Journal of Computational Methods in Sciences and Engineering, 14727978241304260. https://doi.org/10.1177/14727978241304260 Hananto, A. L., Tirta, A., Herawan, S. G., Idris, M., Soudagar, M. E. M., Djamari, D. W., & Veza, I. (2024). Digital Twin and 3D Digital Twin: Concepts, Applications, and Challenges in Industry 4.0 for Digital Twin. Computers, 13(4), 100. https://doi.org/10.3390/computers13040100 Hartmann, D. (2021). Real-Time Digital Twins. Zenodo. https://doi. org/10.5281/ZENODO.5470478 Iranmanesh, S. M., & Nasrabadi, N. M. (2021). HGAN: Hybrid generative adversarial network. Journal of Intelligent & Fuzzy Systems, 40(5), 8927–8938. https://doi.org/10.3233/JIFS-201202 Javaid, M., Haleem, A., & Suman, R. (2023). Digital Twin applications toward Industry 4.0: A Review. Cognitive Robotics, 3, 71–92. https://doi. org/10.1016/j.cogr.2023.04.003 Kampa, A. (2023). Modeling and Simulation of a Digital Twin of a Production System for Industry 4.0 with Work-in-Process Synchronization. Applied Sciences, 13(22), 12261. https://doi.org/10.3390/app132212261 Kusiak, A. (2020). Convolutional and generative adversarial neural networks in manufacturing. International Journal of Production Research, 58(5), 1594–1604. https://doi.org/10.1080/00207543.2019.1662133 Lampropoulos, G., & Siakas, K. (2023). Enhancing and securing cyber‐ physical systems and Industry 4.0 through digital twins: A critical review. Journal DIGITAL TWIN–DRIVEN PREDICTIVE MAINTENANCE IN INDUSTRY 4.0: . . .   105 of Software: Evolution and Process, 35(7), e2494. https://doi.org/10.1002/ smr.2494 Mahankali, R. (2025). DIGITAL TWINS AND ENTERPRISE ARCHITECTURE: A FRAMEWORK FOR REAL-TIME MANUFACTURING DECISION SUPPORT. INTERNATIONAL JOURNAL OF COMPUTER ENGINEERING AND TECHNOLOGY, 16(1), 578–587. https://doi. org/10.34218/IJCET_16_01_049 Mallioris, P., Aivazidou, E., & Bechtsis, D. (2024). Predictive maintenance in Industry 4.0: A systematic multi-sector mapping. CIRP Journal of Manufacturing Science and Technology, 50, 80–103. https://doi.org/10.1016/j. cirpj.2024.02.003 Mannone, G., Hintz, K. D., & Dazer, M. (2025). Generative Twins: A Deep Learning Approach to Next-Gen Digital Twins. 2025 Annual Reliability and Maintainability Symposium (RAMS), 1–6. https://doi.org/10.1109/ RAMS48127.2025.10935021 Mikołajewska, E., Mikołajewski, D., Mikołajczyk, T., & Paczkowski, T. (2025). Generative AI in AI-Based Digital Twins for Fault Diagnosis for Predictive Maintenance in Industry 4.0/5.0. Applied Sciences, 15(6), 3166. https://doi.org/10.3390/app15063166 Onma Enyejo, J., Peter Fajana, O., Sele Jok, I., Judith Ihejirika, C., Olusola Awotiwon, B., & Motilola Olola, T. (2024). Digital Twin Technology, Predictive Analytics, and Sustainable Project Management in Global Supply Chains for Risk Mitigation, Optimization, and Carbon Footprint Reduction through Green Initiatives. International Journal of Innovative Science and Research Technology (IJISRT), 609–630. https://doi.org/10.38124/ijisrt/IJISRT24NOV1344 Parmar Tarun. (2024). Generative Adversarial Networks for Historical Data Generation in Semiconductor Manufacturing: Applications and Challenges. International Journal For Multidisciplinary Research, 6(6), 37230. https://doi. org/10.36948/ijfmr.2024.v06i06.37230 Pires, F., Leitão, P., Moreira, A. P., & Ahmad, B. (2023). Reinforcement learning based trustworthy recommendation model for digital twin-driven decision-support in manufacturing systems. Computers in Industry, 148, 103884. https://doi.org/10.1016/j.compind.2023.103884 Pu, Z., Cabrera, D., Li, C., & De Oliveira, J. V. (2022). VGAN: Generalizing MSE GAN and WGAN-GP for Robot Fault Diagnosis. IEEE Intelligent Systems, 37(3), 65–75. https://doi.org/10.1109/MIS.2022.3168356 106   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Rajesh Lomte. (2025). Industry 4.0 Data Processing Requirements: Use Cases for Big Data Applications. Journal of Information Systems Engineering and Management, 10(16s), 238–244. https://doi.org/10.52783/jisem.v10i16s.2592 Raman, R., Vyas, P., & Vachharajani, H. (2025). Impact of industry 4.0 on supply chain in made to order industries. Annals of Operations Research, 348(3), 1183–1194. https://doi.org/10.1007/s10479-023-05435-x S, Karthika., & Durgadevi, M. (2021). Generative Adversarial Network (GAN): A general review on different variants of GAN and applications. 2021 6th International Conference on Communication and Electronics Systems (ICCES), 1–8. https://doi.org/10.1109/ICCES51350.2021.9489160 S. Matzka, B. S. M. (2020). AI4I 2020 Predictive Maintenance Dataset (Data Set No. The AI4I 2020 Predictive Maintenance Dataset is a synthetic dataset that reflects real predictive maintenance data encountered in industry.; Version 1). UCI Machine Learning Repository. https://doi.org/10.24432/ C5HS5C Selçuk, Ş. Y., Ünal, P., Albayrak, Ö., & Jomâa, M. (2021). A Workflow for Synthetic Data Generation and Predictive Maintenance for Vibration Data. Information, 12(10), 386. https://doi.org/10.3390/info12100386 Sengar, S. S., Hasan, A. B., Kumar, S., & Carroll, F. (2024). Generative artificial intelligence: A systematic review and applications. Multimedia Tools and Applications, 84(21), 23661–23700. https://doi.org/10.1007/s11042-02420016-1 Suthaharan, S. (2016). Support Vector Machine. In S. Suthaharan, Machine Learning Models and Algorithms for Big Data Classification (Vol. 36, pp. 207– 235). Springer US. https://doi.org/10.1007/978-1-4899-7641-3_9 Walton, R. B., Ciarallo, F. W., & Champagne, L. E. (2024). A unified digital twin approach incorporating virtual, physical, and prescriptive analytical components to support adaptive real-time decision-making. Computers & Industrial Engineering, 193, 110241. https://doi.org/10.1016/j.cie.2024.110241 Wang, Y., Zhang, Q., Wang, G.-G., & Cheng, H. (2024). The application of evolutionary computation in generative adversarial networks (GANs): A systematic literature survey. Artificial Intelligence Review, 57(7), 182. https:// doi.org/10.1007/s10462-024-10818-y 107 CHAPTER VI LARGE LANGUAGE MODELS IN MATERIALS SCIENCE: APPLICATIONS AND FUTURE DIRECTIONS Nida KATI1 & Ferhat UÇAR2 1(Assoc. Prof.), Firat University, Faculty of Technology, Department of Metallurgical and Materials Engineering, Elazig, Turkey E-mail: [email protected] ORCID: 0000-0001-7953-1258 2(Assoc. Prof.), Firat University, Faculty of Technology, Department of Software Engineering, Elazig, Turkey E-mail: [email protected] ORCID: 0000-0001-9366-6124 1. Introduction Large Language Models (LLMs), one of the rapidly developing areas of artificial intelligence, have been driving significant changes in materials science in recent years. Models such as GPT, BERT, and Claude have become powerful tools not only for natural language processing tasks but also for scientific research and discovery. Materials science, a field requiring complex structural relationships, multidimensional datasets, and interdisciplinary approaches, significantly benefits from the capabilities offered by LLMs. Traditional materials research requires lengthy processes based on experimental studies and theoretical calculations. The development of a new material, characterization of its properties, and optimization can take years. In this regard, LLMs assist materials scientists in a wide range of tasks, from literature review and material design to property prediction and process optimization. Figure 1 presents a comprehensive visual summary of how large language models are applied across materials science research. The framework illustrates the fundamental shift from traditional research approaches—characterized 108   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . by manual literature searches, limited parameter exploration, and months of iterative experimentation—to AI-driven methodologies that enable systematic multidimensional analysis within significantly compressed timeframes. The diagram organizes LLM applications into three primary domains that reflect the core activities of materials research. Material design and discovery applications leverage inverse design approaches, where desired properties guide the search for optimal material structures rather than following the conventional path from composition to properties. This category encompasses structure-property relationship modeling, composition optimization, and automated hypothesis generation. Property prediction and characterization applications address the computational challenge of estimating material properties—including formation energies, electronic band gaps, mechanical characteristics, and thermal conductivity—without requiring extensive experimental testing or intensive computational simulations. Data analysis and visualization applications tackle the growing challenge of extracting meaningful insights from unstructured scientific literature, performing named entity recognition in complex chemical descriptions, establishing relationships between materials and their properties, and integrating LLM capabilities into laboratory workflows through electronic notebook systems. The framework highlights representative models that demonstrate these capabilities, including AtomGPT for atomistic property prediction and structure generation, MatterChat for multimodal integration of structural and textual information, LLM-Fusion for combining multiple data representations, and domainspecific models like MatBERT trained on materials science literature. The diagram also acknowledges key challenges that researchers must navigate, including ensuring data quality and availability, developing adequate domain-specific understanding in models, managing hallucination risks and verification requirements, addressing interpretability and trust issues, meeting computational resource demands, and establishing standardization protocols for evaluation and comparison. Looking toward future development, the framework emphasizes the need for task-specific model refinement, enhanced multimodal data integration, improved reliability mechanisms, and—most importantly—balanced human-AI collaboration where artificial intelligence augments rather than replaces researcher expertise. LARGE LANGUAGE MODELS IN MATERIALS SCIENCE: APPLICATIONS AND . . .   109 Fig.1. Comprehensive framework of LLM usage in materials science. The integration of artificial intelligence into materials research marks a critical juncture in the field’s evolution, one that demands both enthusiasm for new possibilities and careful consideration of implementation approaches. As materials scientists confront increasingly complex challenges—from designing next-generation energy storage materials to developing sustainable alternatives for critical applications—the volume and complexity of relevant information continue to grow. Large language models emerge in this context not as replacements for experimental validation or theoretical understanding, but as tools that can help researchers navigate vast literature landscapes, identify 110   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . promising research directions, and accelerate the hypothesis-testing cycle. The effectiveness of these tools depends fundamentally on how they are integrated into existing research practices and the extent to which their outputs undergo critical evaluation. This chapter systematically examines the current state of LLM applications in materials science, organizing the discussion around core research activities. In the following section the main branch application areas of the LLM structures in materials science is declarated. Section 2.1 explores material design and discovery, examining how inverse design approaches enable more systematic exploration of compositional spaces and how models like AtomGPT and MatterChat demonstrate practical capabilities for structure generation and property prediction. Section 2.2 addresses property prediction and characterization, reviewing benchmarking efforts such as LLM4MatBench that reveal both the capabilities and limitations of current approaches. Section 2.3 discusses data analysis and visualization, focusing on knowledge extraction from scientific literature and integration into laboratory workflows through systems like electronic notebooks. Section 2.4 examines the challenges that researchers encounter when applying these tools, including data quality constraints, domain-specific understanding requirements, hallucination risks, and the need for standardization protocols. The chapter concludes by synthesizing findings across these domains and identifying priorities for future development, with particular attention to the balance between automation and human expertise that characterizes successful implementation. 2. Application areas of the LLM in materials science Large language models are radically changing traditional research concepts in materials science. This transformation encompasses a broad spectrum, from the way materials researchers address complex problems to the methodologies for processing scientific information and generating new discoveries. The applications of LLMs in materials science can be broadly categorized into different categories: information management and synthesis, design and discovery, characterization and analysis, process optimization, and data analysis. To understand the impact of this technology in materials science, it is useful to compare it to traditional research processes. In the traditional approach, when a researcher wants to develop a new material, they collect existing information through manual literature searches, rely on experience and intuition during hypothesis generation, try a limited number of parameter combinations for experimental design, are limited to statistical tools during data LARGE LANGUAGE MODELS IN MATERIALS SCIENCE: APPLICATIONS AND . . .   111 analysis, and are constrained by their own area of expertise during the results interpretation phase. LLMs, however, revolutionize each of these processes, enabling comprehensive literature analysis in seconds, enabling systematic exploration in a multidimensional parameter space, processing large data sets instantaneously, automatically establishing interdisciplinary connections, and improving performance through continuous learning. This paradigm shift allows materials scientists to think more strategically and focus on creative processes, while automating routine and repetitive tasks. 2.1. Material design and discovery One of the most important and impactful applications of large language models is the comprehensive approaches they offer in new materials design. This technology enables a transition from traditional trial-and-error methodologies to systematic, data-driven design. LLMs demonstrate extraordinary capabilities in modeling structure-property relationships, learning the highly complex and multidimensional relationships between material structure and properties and translating this knowledge into practical applications. This process involves analyzing atomic-level configurations, systematically evaluating crystal structure data, and modeling the effects of these parameters on macroscopic properties. The system is capable of recommending material compositions with desired properties, optimizing element ratios, and improving material performance by developing doping strategies. More importantly, LLMs adopt an inverse design approach; this approach, in contrast to the traditional “material-to-properties” methodology, works from desired properties to material structure. This inverse design strategy can suggest optimal alloy compositions when target mechanical properties are determined; develop optimal crystal structure configurations when desired optical properties are defined; and optimize heat transfer mechanisms through phonon engineering strategies when thermal conductivity targets are set. This comprehensive approach allows materials scientists to reduce research time from months to days, optimize resource utilization, and systematically explore previously unexplored material combinations, leading to significant changes in materials discovery and development. One of the most important studies demonstrating the potential of LLMs in materials design is the AtomGPT model developed by Choudhar. This study, motivated by the recognition that despite the success of large-scale language models like GPT in commercial applications, their applications in materials design have not yet been sufficiently explored, developed a specialized model 112   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . based on transformer architecture. By combining chemical and structural text descriptions, AtomGPT demonstrates both atomistic property prediction and structure generation capabilities. The model can predict critical material properties such as formation energies, electronic band gaps, and superconductor transition temperatures with accuracy comparable to graph neural network models and can generate atomic structures for tasks such as new superconductor design. The obtained predictions are validated with density functional theory calculations, and this study demonstrates the potential of LLMs for forward and inverse material design, offering an efficient alternative approach for material discovery and optimization (Choudhary, 2024). Understanding and predicting the properties of inorganic materials is critical for accelerating advances in materials science and supporting applications in fields such as energy and electronics. MatterChat, developed by Tang et al., is a multimodal large-language model that enhances human-AI interaction by integrating material structural data with language-based information. The key challenge of this study was to integrate atomic structures into LLMs with full resolution, a problem addressed by developing a structure-aware architecture. MatterChat reduces training costs and increases flexibility by using a bridge module that effectively aligns a pre-trained universal machine learning interatomic potential with a pre-trained LLM. The model significantly outperforms generalpurpose LLMs such as GPT-4 in material property prediction and human-AI interaction, and has also proven useful in applications such as advanced scientific reasoning and step-by-step material synthesis (Tang et al., 2025). LLM-Fusion, developed by Boyar et al., is an innovative fusion model that addresses this problem with a multimodal approach. While existing multimodal approaches are known to be promising in integrating different information sources, it has been found that fusion algorithms developed to date are simplistic and lack mechanisms for rich representation of multiple modalities. LLM-Fusion offers a flexible architecture capable of accurate feature prediction by integrating diverse representations such as SMILES, SELFIES, textual descriptions, and molecular fingerprints using large language models. With its LLM-based structure that supports multimodal input processing, the model can perform material property prediction with higher accuracy than traditional methods. Validated on five prediction tasks on two different datasets, the model demonstrates superior performance compared to single-modal and simple fusion methods. This study is significant because it demonstrates how intelligently combining different data sources via LLMs can yield effective LARGE LANGUAGE MODELS IN MATERIALS SCIENCE: APPLICATIONS AND . . .   119 LLM4Mat-Bench have clearly demonstrated the limitations of general-purpose LLMs in materials science. The results demonstrate that fine-tuned task-specific models significantly outperform general-purpose models, highlighting the need to develop domain-specific models for specialized tasks such as material property prediction. Success rates in predicting critical material properties such as formation energies, electronic band gaps, and superconducting transition temperatures demonstrate the potential of this approach. In the field of data analysis and visualization, LLMs have been demonstrated for systematically extracting information from unstructured scientific texts, interpreting experimental data, and automating research workflows. Integration studies into electronic laboratory notebooks have concretely demonstrated how this technology can be incorporated into daily laboratory practices and comprehensively support research processes. However, the challenges encountered in named entity recognition and relationship extraction tasks, particularly when dealing with complex chemical formulas and domain-specific terminology, still demonstrate the existence of significant technical hurdles. One of the most important findings of existing studies is the need for fine-tuning and specialization for LLMs to be effective in materials science. General-purpose models, despite their extensive knowledge bases, fall short of the specialized terminology, complex mathematical relationships, and domainspecific rules of materials science. Therefore, future research should focus on developing task-specific models, improving multimodal data integration, and strengthening reliability mechanisms. The successful application of LLMs in materials science depends on the balanced integration of human expertise and artificial intelligence capabilities. This technology aims not to replace researchers, but to support them and accelerate research processes. Therefore, the validation of LLM outputs, critical evaluation of results, and integration with expert knowledge are fundamental requirements for the responsible and effective use of this technology. Large-language models offer significant opportunities in materials science research. However, fully realizing this potential requires the development of standardized assessment frameworks, the creation of domain-specific models, the strengthening of reliability and validation mechanisms, and the adoption of ethical use principles. In the future, by overcoming these challenges, LLMs are expected to play an even more central role in materials discovery and design, significantly accelerating the development of new materials for critical societal needs. 120   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . References Boyar, O., Priyadarsini, I., Takeda, S., & Hamada, L. (2025). Llm-fusion: A novel multimodal fusion model for accelerated material discovery. arXiv preprint arXiv:2503.01022. Choudhary, K. (2024). Atomgpt: Atomistic generative pretrained transformer for forward and inverse materials design. The Journal of Physical Chemistry Letters, 15(27), 6909-6917. Foppiano, L., Lambard, G., Amagasa, T., & Ishii, M. (2024). Mining experimental data from materials science literature with large language models: an evaluation study. Science and Technology of Advanced Materials: Methods, 4(1), 2356506. Jalali, M., Luo, Y., Caulfield, L., Sauter, E., Nefedov, A., & Wöll, C. (2024). Large language models in electronic laboratory notebooks: Transforming materials science research workflows. Materials Today Communications, 40, 109801. Kumbhar, S., Mishra, V., Coutinho, K., Handa, D., Iquebal, A., & Baral, C. (2025). Hypothesis generation for materials discovery and design using goaldriven and constraint-guided llm agents. arXiv preprint arXiv:2501.13299. Rubungo, A. N., Li, K., Hattrick-Simpers, J., & Dieng, A. B. (2025). LLM4Mat-bench: benchmarking large language models for materials property prediction. Machine Learning: Science and Technology, 6(2), 020501. Tang, Y., Xu, W., Cao, J., Gao, W., Farrell, S., Erichson, B., ... & Yao, Z. (2025). Matterchat: A multi-modal llm for material science. arXiv preprint arXiv:2502.13107. 121 CHAPTER VII LARGE LANGUAGE MODELS FOR SYSTEMATIC LITERATURE ANALYSIS IN SOLAR CELL MATERIALS RESEARCH: A PRACTICAL FRAMEWORK Nida KATI1 & Ferhat UÇAR2 1(Assoc. Prof.), Fırat University, Faculty of Technology, Department of Metallurgy and Materials Engineering, Elazig/Turkiye E-mail: [email protected] ORCID: 0000-0001-7953-1258 2(Assoc. Prof.), Fırat University, Faculty of Technology, Department of Software Engineering, Elazig/Turkiye E-mail: [email protected] ORCID: 0000-0001-9366-6124 1. Introduction Solar cell research has grown substantially over the past decade, driven by global efforts to transition toward renewable energy sources. As research activity intensifies, the volume of scientific literature has increased correspondingly. A search for “solar cell materials” in major academic databases now returns thousands of papers published annually, spanning multiple technology platforms including silicon-based cells, thin-film devices, perovskite solar cells, organic photovoltaics, and emerging tandem configurations. For researchers attempting to stay current with developments in the field, this growth presents a significant challenge. Traditional approaches to literature review rely on manual reading and analysis of individual papers. A comprehensive review of even 100 papers can require several weeks of focused effort, involving careful reading of each paper, extraction of relevant data, categorization by research focus, and synthesis of findings. This process is not only time-consuming but also subject 122   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . to individual interpretation and potential oversight (Schmidt et al., 2019). When research teams need to analyze broader trends across hundreds of papers— such as identifying which materials are gaining attention, tracking performance improvements, or spotting emerging research directions—the manual approach becomes increasingly impractical. Recent developments in artificial intelligence, particularly large language models (LLMs), offer a different approach to handling large volumes of scientific text. These models can process and analyze extensive document collections in a fraction of the time required for manual review. Unlike earlier text analysis tools that relied on simple keyword matching, modern LLMs can understand context, identify relationships between concepts, extract quantitative data, and generate structured summaries (Bai & Zhang, 2025). Several research groups have begun exploring how these capabilities might be applied to materials science, with applications ranging from property prediction to synthesis planning (Fang et al., 2022; Jiang et al., 2025). This chapter focuses on a specific application: using LLMs to conduct systematic literature analysis in solar cell materials research. Rather than attempting to replace human expertise, the approach described here uses LLM capabilities to handle time-intensive tasks such as data extraction, categorization, and initial pattern identification. The goal is to enable researchers to efficiently process large literature datasets while maintaining the critical evaluation and interpretation that require domain knowledge. The framework presented here is built around Claude, an LLM developed by Anthropic, though the general principles apply to other similar models. Claude offers several features that make it suitable for literature analysis: it can process long documents (up to 200,000 tokens in a single interaction), maintains context across multiple uploaded files, and can generate structured outputs including tables and visualizations. Importantly for academic applications, it can be instructed to cite specific sources when making claims, helping to maintain traceability between outputs and original papers. The methodology we describe involves collecting papers in BibTeX format—a standard bibliographic format that includes author names, journal information, publication year, abstracts, and keywords—and using systematic prompts to extract and organize information. This approach does not require programming skills, making it accessible to researchers who want to leverage AI capabilities without learning to code. For those with programming experience, we also describe how the same workflow can be automated using the Claude API. LARGE LANGUAGE MODELS FOR SYSTEMATIC LITERATURE ANALYSIS IN . . .   123 To demonstrate the framework’s practical utility, we present a case study analyzing several hundred recent papers on solar cell materials. The analysis covers publication trends, materials innovation patterns, performance benchmarking, and emerging research directions. This case study illustrates both the capabilities and limitations of LLM-based literature analysis, providing realistic expectations for researchers considering this approach. Why focus specifically on solar cell materials? This field presents characteristics that make it particularly suitable for demonstrating LLM-based analysis. Solar cell research is highly interdisciplinary, involving chemistry, materials science, physics, and engineering. Papers frequently report quantitative performance metrics—power conversion efficiency, open-circuit voltage, short-circuit current density, and fill factor—that can be systematically extracted and compared. The field also exhibits rapid evolution, with new materials and device architectures emerging regularly, creating a genuine need for tools that can quickly identify trends and assess the state of the art. Additionally, the solar cell community has established standardized testing protocols and reporting guidelines, which facilitates consistent data extraction across different studies. The approach described here also addresses a practical limitation faced by many research groups: limited access to specialized AI expertise. While machine learning has become increasingly important in materials science (Maqsood et al., 2024), implementing custom AI solutions often requires programming skills and computational resources that may not be readily available. By demonstrating how a commercially available LLM can be used through a simple web interface or basic API calls, this chapter aims to lower the barrier to entry for researchers interested in AI-assisted literature analysis. It is important to establish realistic expectations about what LLM-based analysis can and cannot accomplish. These tools excel at tasks involving pattern recognition, information extraction, and structured summarization across large document collections. They can quickly identify which materials are being studied most frequently, track how performance metrics have evolved over time, and flag papers that report unusual or noteworthy results. However, they have limitations. LLMs can sometimes generate incorrect information, particularly when asked to extrapolate beyond their training data. They cannot access full-text articles behind paywalls unless explicitly provided with that content. They also lack the deep domain expertise required to evaluate the validity of experimental methods or the significance of subtle technical details. 124   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . For these reasons, the framework we present is designed as a human-AI collaboration rather than full automation. The LLM handles the labor-intensive aspects of literature processing, while human researchers provide the critical evaluation, contextual understanding, and decision-making that require expertise. This division of labor allows researchers to analyze much larger literature collections than would be practical manually, while maintaining the quality standards expected in academic research. Figure 1 provides a visual overview of the framework presented in this chapter. The diagram illustrates the three-phase methodology: data collection from academic databases, systematic analysis using Claude’s natural language processing capabilities, and quality control through human verification. The framework begins with retrieving papers in BibTeX format from databases like ScienceDirect, capturing metadata including titles, abstracts, author information, and keywords. This structured data serves as input for the analysis phase, where systematic prompts guide Claude to extract technology classifications, identify material trends, and compile performance metrics. The process emphasizes human-AI collaboration rather than full automation—the LLM handles laborintensive extraction tasks while researchers provide critical evaluation and domain expertise. Key features highlighted in the diagram include the system’s accessibility (requiring no programming skills for basic use), its large context window enabling analysis of hundreds of papers simultaneously, and its ability to generate structured outputs such as tables and visualizations. The framework produces four categories of insights: bibliometric patterns (journal distribution, temporal trends), technology trends (material classifications, emerging themes), performance benchmarks (efficiency statistics, stability data), and innovation indicators (research gaps, novel approaches). These outputs support various applications including literature reviews for graduate students, competitive intelligence for research groups, and identification of collaboration opportunities. The diagram emphasizes that this approach is designed as a complementary tool—enhancing rather than replacing human expertise—with the LLM managing information processing while researchers maintain responsibility for interpretation and validation. The chapter is organized as follows: Section 2 provides background on solar cell technologies and current applications of LLMs in materials research. Section 3 details the methodology, including data collection procedures, prompt engineering strategies, and quality control measures. Section 4 presents the comprehensive case study with detailed analysis across multiple dimensions. LARGE LANGUAGE MODELS FOR SYSTEMATIC LITERATURE ANALYSIS IN . . .   125 Section 5 offers practical implementation guidance for researchers, including example prompts and code snippets. Section 6 concludes the overall framework, discussing perspectives on future developments in AI-assisted literature analysis. Fig.1 – General Framework of the proposed system 2. Background: LLMs In Materials Science Understanding how large language models can support solar cell materials research requires context on two fronts: the diverse landscape of solar cell technologies and the evolving role of artificial intelligence in materials science. This section provides that foundation, first reviewing the major categories of solar cell technologies and their characteristic materials challenges, then examining current applications of LLMs in materials research more broadly. 126   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . 2.1. Solar Cell Technologies: A Brief Overview Solar cells convert sunlight into electricity through the photovoltaic effect, and different technology platforms have emerged over several decades of development. Understanding the landscape of solar cell research requires familiarity with the major technology categories and their characteristic materials challenges. Silicon-based solar cells, often referred to as first-generation photovoltaics, currently dominate the commercial market. These devices use crystalline or polycrystalline silicon wafers as the light-absorbing material. While silicon solar cells have achieved high efficiencies exceeding 26% in laboratory settings and demonstrate excellent long-term stability, they require high-purity materials and energy-intensive manufacturing processes. Research in this area continues to focus on reducing costs, improving light management, and developing passivation strategies to minimize efficiency losses at material interfaces (Cheng et al.,2025). Thin-film solar cells, representing second-generation technology, use much thinner layers of photoactive materials deposited on glass or flexible substrates. This category includes cadmium telluride (CdTe), copper indium gallium selenide (CIGS), and amorphous silicon devices. Thin-film technologies offer advantages in material usage and manufacturing flexibility, though they generally achieve lower efficiencies than crystalline silicon. Current research emphasizes improving efficiency while addressing concerns about material toxicity and resource availability (Ezihe et al., 2025). Third-generation solar cells encompass several emerging technologies that aim to surpass the theoretical efficiency limits of single-junction devices or offer new functionalities. Perovskite solar cells, which use organic-inorganic hybrid materials with the general formula ABX₃, have attracted intense research interest due to their rapid efficiency improvements and solutionprocessability. Since the first reports of efficient perovskite solar cells in 2012, power conversion efficiencies have increased from around 10% to over 26%, approaching the performance of silicon cells. However, stability concerns related to moisture sensitivity, thermal degradation, and the use of lead remain significant challenges (Moyofola et al., 2025). Organic photovoltaics use carbon-based semiconducting polymers or small molecules as active materials. These devices can be manufactured using low-cost printing techniques and offer advantages such as flexibility and semitransparency. Recent developments in non-fullerene acceptor materials have pushed efficiencies for single-junction devices into competitive ranges. LARGE LANGUAGE MODELS FOR SYSTEMATIC LITERATURE ANALYSIS IN . . .   127 Research continues to address challenges related to efficiency, stability, and large-area manufacturing (Fan et. al, 2025). Quantum dot solar cells employ nanoscale semiconductor particles with size-dependent optical properties. Colloidal quantum dots can be synthesized from earth-abundant materials and processed from solution, offering potential cost advantages. Current research explores optimal quantum dot compositions, surface passivation strategies, and device architectures to improve efficiency and stability (Torres et al., 2025). Tandem solar cells combine multiple absorber materials with different bandgaps to capture a broader range of the solar spectrum. Perovskitesilicon tandems have demonstrated efficiencies exceeding 33%, surpassing the performance of single-junction silicon cells. This approach represents a promising path toward higher efficiencies, though it introduces additional manufacturing complexity (Lin et al., 2026). Across all these technologies, common materials challenges include optimizing charge transport layers (both hole transport materials and electron transport materials), developing effective interfacial passivation strategies, improving stability under operating conditions, and scaling laboratory results to commercially relevant areas (Koten et al., 2025). The diversity of approaches and the rapid pace of development create a research landscape where systematic literature analysis can provide valuable insights into emerging trends and promising directions. 2.2. Generative AI Existence in Materials Research: Current Applications Artificial intelligence has been applied to materials science for several decades, but recent advances in large language models have introduced capabilities that differ fundamentally from earlier approaches. While traditional machine learning in materials science typically focused on predicting specific properties from structural or compositional data, LLMs bring the ability to process and generate natural language, opening possibilities for literature analysis, knowledge synthesis, and experimental planning (Bai & Zhang, 2025). The application of machine learning to materials discovery accelerated significantly in the past decade. Researchers demonstrated that algorithms could learn relationships between composition, structure, and properties from existing datasets, then use those relationships to predict characteristics of hypothetical materials (Schmidt et al., 2019). These approaches proved valuable for screening large numbers of candidates and identifying promising compositions 128   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . for experimental investigation. However, most machine learning applications required carefully curated numerical datasets and domain-specific feature engineering—the process of selecting which material characteristics to include as input variables. Large language models represent a different paradigm. Trained on vast amounts of text data, these models develop statistical representations of how concepts relate to each other in natural language. This capability allows them to work directly with scientific literature in its native format, without requiring manual conversion to structured databases (Fang et al., 2022). An LLM can read a paper abstract describing a new material synthesis, extract relevant information about precursors and processing conditions, and identify connections to similar work reported in other papers—tasks that previously required human expertise. Several applications of LLMs in materials science have emerged in recent years. One prominent area involves literature mining and knowledge extraction. Researchers have used LLMs to automatically extract material properties, synthesis conditions, and performance metrics from large collections of papers. This approach addresses a longstanding challenge in materials informatics: much valuable experimental data exists in scientific publications but remains inaccessible to computational analysis because it is reported in unstructured text rather than standardized databases. Another application involves using LLMs to accelerate the materials design process. By training models on existing relationships between material composition and properties, researchers can query the model to suggest candidate materials with desired characteristics. While this application resembles traditional machine learning approaches, LLMs offer the advantage of being able to explain their suggestions in natural language and consider multiple types of information simultaneously—composition, structure, processing conditions, and even contextual factors like cost or availability (Chen et al., 2024). LLMs have also been applied to experimental planning and protocol optimization. Some systems can suggest synthetic routes for new materials based on their understanding of chemical reactions and processing techniques reported in the literature. Others assist with analyzing experimental data or troubleshooting unexpected results by drawing on knowledge extracted from thousands of papers describing similar systems. In the specific context of solar cell research, AI applications have primarily focused on predicting device performance, optimizing material compositions, and identifying stable material combinations (Jiang et al., 2025). Machine LARGE LANGUAGE MODELS FOR SYSTEMATIC LITERATURE ANALYSIS IN . . .   135 Table.1 – Top 10 publication distribution. Rank Journal Papers % 1 Solar Energy 43 8.8% 2Chemical Engineering Journal 26 5.3% 3 Dyes and Pigments 22 4.5% 4 Journal of Physics and Chemistry of Solids 18 3.7% 5Solar Energy Materials and Solar Cells 11 2.3% 6 Materials Today Energy 9 1.8% 7 Journal of Power Sources 9 1.8% 8 Materials Today Communications 9 1.8% 9 Organic Electronics 9 1.8% 10 Journal of Alloys and Compounds 9 1.8% The dominance of Solar Energy and specialized materials journals reflects the field’s focus on both fundamental materials development and device applications. The wide distribution across 133 journals—with no single venue capturing more than 10% of publications—demonstrates the interdisciplinary nature of the field. 4.2. Research Focus Distribution 4.2.1. Technology Platform Analysis: Based on title and abstract analysis, papers were classified by primary solar cell technology: · Perovskite solar cells: Dominant category, appearing in approximately 35-40% of papers · Organic photovoltaics: Second major category (20-25% of papers) · Silicon-based cells: Established technology, continuing optimization (10-15%) · Dye-sensitized solar cells: Mature field with ongoing research (8-12%) · Quantum dot solar cells: Emerging category (5-8%) · Tandem architectures: Growing interest, particularly perovskite-silicon combinations (5-7%) Note: Categories overlap as some papers address multiple technologies or comparative studies. 136   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . The prominence of perovskite research reflects continued interest in this rapidly advancing field. Organic photovoltaics maintain strong representation, driven by developments in non-fullerene acceptors and ternary blend systems. Silicon research, while proportionally smaller, focuses on advanced passivation techniques and hybrid architectures. 4.2.2. Component Focus: Papers were categorized by the device component they primarily addressed: · Charge transport materials (HTM/ETM): ~35% of papers · Absorber layer optimization: ~30% · Interfacial engineering: ~20% · Stability enhancement: ~15% · Device architecture: ~10% The high proportion of charge transport material research reflects ongoing efforts to replace conventional materials with more stable, cost-effective alternatives. Interface engineering appears frequently, consistent with recognition that defects at layer boundaries significantly impact device performance. 4.3. Materials Innovation Analysis 4.3.1. Perovskite Solar Cells Perovskite research dominates the dataset, with several distinct themes as listed: Lead-Free Alternatives: Approximately 15-20% of perovskite papers investigate lead-free or lead-reduced compositions. Tin-based perovskites (FASnI₃, MASnI₃) appear most frequently, though stability challenges remain. Other approaches include bismuth-based compounds and double perovskites, though these typically achieve lower efficiencies. Hole Transport Materials: Dopant-free HTMs represent a growing trend. Papers report various organic polymers, small molecules, and inorganic materials (NiOₓ, CuSCN) as alternatives to conventional doped spiro-OMeTAD. The motivation centers on improving long-term stability by eliminating hygroscopic dopants like lithium salts. Electron Transport Materials: While TiO₂ and SnO₂ remain standard, papers increasingly explore fullerene derivatives and other organic ETMs for flexible or inverted architectures. Several studies investigate bis(pyrrolidino)fullerenes as shown in one paper reporting 12.3% efficiency with improved stability. LARGE LANGUAGE MODELS FOR SYSTEMATIC LITERATURE ANALYSIS IN . . .   137 4.3.2. Organic Photovoltaics Organic solar cell papers emphasize non-fullerene acceptor (NFA) development. The Y-series acceptors (Y6 and derivatives) appear frequently, paired with various polymer donors (PM6, PTQ10, PBDB-T). Research trends include: · Ternary blend systems to broaden absorption spectra · Molecular engineering to optimize energy levels and morphology · Interface optimization using self-assembled monolayers Several papers report efficiencies exceeding 18% for single-junction devices, with stability improvements through encapsulation and morphology control. 4.3.3. Interfacial Engineering Interface modification emerges as a cross-technology theme. Common approaches include: · Self-assembled monolayers (SAMs) for charge-selective contacts · 2D materials (MoS2, MXenes) for passivation · Aromatic compounds for defect passivation at perovskite interfaces One notable paper systematically reviews aromatic compound applications in perovskite solar cells, analyzing effects on open-circuit voltage, fill factor, and stability. 4.4. Performance and Stability Trends While not all papers report complete device metrics, those that do reveal several patterns can be grouped as follows: Efficiency Progress: Papers reporting device performance typically show efficiencies between 15-25% for perovskite cells, 14-19% for organic cells, and 23-26% for silicon-based systems. Tandem configurations achieve the highest values, with perovskite-silicon tandems reaching 28-33%. Stability Focus: Approximately 40% of papers explicitly address stability testing. Common protocols include: · Thermal stability (85°C aging) · Humidity resistance (controlled RH exposure) · Light soaking (continuous illumination tests) · Operational stability (maximum power point tracking) 138   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Papers increasingly report long-term data (>1000 hours), though standardized testing protocols vary. Perovskite papers often cite ISOS protocols, while organic solar cell papers follow fewer uniform procedures. Scalability Considerations: The dataset includes relatively few papers addressing large-area devices or manufacturing scale-up. Most report lab-scale cells (<0.1 cm²), with only 10-15% discussing devices larger than 1 cm² or rollto-roll processing. 4.5. Emerging Themes Machine Learning Integration: Approximately 8-10% of papers mention machine learning or computational screening. Applications include: · Materials property prediction · Composition optimization using Bayesian methods · Automated literature analysis (including one paper on using LLMs for solar cell database creation) Sustainability Considerations: Papers increasingly address environmental concerns: · Lead-free perovskite research motivated by toxicity reduction · Lifecycle assessment of manufacturing processes · Recyclability of organic materials · Earth-abundant element usage Flexible and Wearable Devices: Approximately 12% of papers investigate flexible substrates or applications in wearable electronics. These papers emphasize mechanical stability, lightweight construction, and compatibility with roll-to-roll processing. 4.6. Geographic Distribution Based on author affiliations, research contributions come from diverse regions: · Asia (China, South Korea, Japan): Dominant contributor (~50% of papers) · Europe (Germany, UK, Italy, Spain): ~25% of papers · North America (USA, Canada): ~15% of papers · Middle East and Other: ~10% of papers LARGE LANGUAGE MODELS FOR SYSTEMATIC LITERATURE ANALYSIS IN . . .   139 Chinese institutions show particularly high publication volumes, consistent with national investment in renewable energy research. 5. Practical Implementation Guide This section provides concrete guidance for researchers who want to apply LLM-based literature analysis to their own work. We present two approaches: a no-code method accessible to all researchers, and a programmatic approach for those comfortable with basic Python scripting. 5.1. No-Code Approach: Using Claude Web Interface 5.1.1. Step 1: Prepare Your Dataset Search your preferred database (ScienceDirect, Web of Science, Scopus) using relevant keywords. Export results in BibTeX format. Most databases allow exporting 100-500 papers at once. If you have more papers, split into multiple files or combine them into a single text file. 5.1.2. Step 2: Upload and Initial Query Visit claude.ai and upload your BibTeX file. Start with a basic overview prompt: “I have uploaded a BibTeX file of solar cell papers. Please tell me: (1) How many papers are included? (2) What years do they cover? (3) Which journals appear most frequently?” This confirms successful upload and provides basic statistics. 5.1.3. Step 3: Technology Classification “Categorize these papers by solar cell technology type: perovskite, organic, silicon, quantum dot, tandem, dye-sensitized, or other. Provide counts and percentages for each category.” 5.1.4. Step 4: Materials Analysis “Identify the most frequently mentioned hole transport materials in perovskite papers. List the top 10 with approximate frequency.” “Which absorber materials are most studied in organic solar cell papers?” 5.1.5. Step 5: Performance Trends “Find papers reporting power conversion efficiency above 20%. Create a table with: Author, Year, Technology, PCE (%), and Key Innovation.” 140   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . 5.1.6. Step 6: Emerging Themes “What are the top 5 research trends or challenges mentioned across these papers? For each trend, list 2-3 representative papers.” 5.2. Programmatic Approach: Python + Claude API For researchers who want to automate repetitive analyses or process very large datasets, the Claude API enables scripting with using Python code design. Also the proposed system can produce the LaTeX code design of the generated report. The basic setup and advanced batch processing multiple queries designs can be defined as follows. 5.2.1. Basic Setup: import anthropic client = anthropic.Anthropic(api_key=”your-api-key-here”) def analyze_literature(bibtex_content, prompt): message = client.messages.create( model=”claude-sonnet-4-20250514”, max_tokens=4096, messages=[{ “role”: “user”, “content”: f”{bibtex_content}\n\n{prompt}” }] ) return message.content[0].text # Load BibTeX file with open(“solar_cells.bib”, “r”) as f: bibtex_data = f.read() # Run analysis result = analyze_literature( bibtex_data, “Categorize by technology type and list top 10 journals.” ) print(result) 5.2.2 Advanced: Batch Processing Multiple Queries pythonqueries = [ LARGE LANGUAGE MODELS FOR SYSTEMATIC LITERATURE ANALYSIS IN . . .   141 “Technology distribution”, “Top 10 journals”, “Most cited papers”, “Emerging material trends” ] results = {} for query in queries: results[query] = analyze_literature(bibtex_data, query) # Save results import json with open(“analysis_results.json”, “w”) as f: json.dump(results, f, indent=2) This approach enables automated processing of multiple datasets during off-hours and automatic generation of time-consuming monthly reports, freeing researchers from repetitive manual tasks. 5.3. Best Practices Verification: Always spot-check extracted data. Randomly select 10-15 papers and verify that reported information matches the original sources. Iterative Refinement: Don’t expect perfect results on the first try. Refine your prompts based on initial outputs. If results are too vague, add more specific criteria. Citation Tracking: When Claude cites specific papers, verify they exist in your dataset. Occasionally, LLMs may conflate information from different sources. Limitations Awareness: Remember that Claude only sees abstracts if working with BibTeX exports. Full-text analysis requires uploading complete PDFs, which may not be practical for large collections. Export and Document: Copy important outputs (tables, insights) into a separate document as you work. Claude conversations have context limits, so documenting key findings prevents information loss. 6. Conclusion This chapter has presented a practical framework for conducting systematic literature analysis in solar cell materials research using large language models. 142   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Through a case study analyzing 488 recent papers, we have demonstrated how LLM-based approaches can efficiently process large document collections while maintaining the scientific rigor expected in academic research. 6.1. Key Findings and Contributions The framework described here addresses a genuine challenge facing materials researchers: the growing volume of scientific literature makes comprehensive manual review increasingly impractical. Our case study revealed several insights that illustrate both the capabilities and appropriate use cases for LLM-based analysis. First, the methodology enables rapid bibliometric analysis across large datasets. Tasks that would require days or weeks of manual effort—categorizing papers by technology type, identifying publication trends, extracting performance metrics—can be completed in hours. The analysis of 488 papers revealed clear patterns: perovskite solar cells dominate current research (35-40% of papers), charge transport material optimization represents a major focus area (35% of papers), and publications span 133 different journals, reflecting the field’s interdisciplinary nature. Second, the approach facilitates identification of emerging research directions. Our analysis detected several trends that merit attention: growing interest in dopant-free hole transport materials, increased emphasis on longterm stability testing, and rising adoption of machine learning for materials screening. These observations emerge naturally from processing large document collections in ways that would be difficult to detect through manual review of individual papers. Third, the framework demonstrates how AI tools can complement rather than replace human expertise. The LLM handles labor-intensive information extraction and pattern recognition, while researchers provide critical evaluation, contextual understanding, and interpretation. This division of labor proves more effective than either fully manual or fully automated approaches. 6.2. Practical Implications For researchers, this framework offers several practical benefits. Graduate students beginning literature reviews can quickly gain overview of their research area, identifying key journals, influential papers, and current research directions. Established researchers can monitor ongoing developments in related fields, detecting potentially relevant work that might otherwise escape notice. Research LARGE LANGUAGE MODELS FOR SYSTEMATIC LITERATURE ANALYSIS IN . . .   143 groups can systematically analyze competitor activities, track technology trends, or identify collaboration opportunities. The accessibility of this approach—requiring no programming skills for basic implementation—lowers barriers to adoption. Researchers can begin using these methods immediately through web interfaces, with the option to automate repetitive tasks later using straightforward API calls. This flexibility accommodates different skill levels and use cases. 6.3. Limitations and Considerations Despite its utility, LLM-based literature analysis has important limitations that users must understand. The most significant constraint involves working with abstracts rather than full texts. BibTeX exports typically contain only paper metadata and abstracts, excluding detailed experimental procedures, comprehensive results, and nuanced discussions found in full articles. This limitation affects the depth of analysis possible, particularly for questions requiring detailed technical information. LLMs can occasionally generate incorrect information, particularly when asked to extrapolate beyond provided data. While this “hallucination” risk can be managed through verification procedures and careful prompt design, it remains a consideration. Users should always verify critical findings by consulting original sources. The framework also cannot fully replace domain expertise. LLMs lack the deep understanding of materials science required to evaluate experimental validity, assess the significance of subtle technical details, or recognize when reported results seem anomalous. These judgments require human expertise informed by years of research experience. The approach depends on the quality of input data. If literature searches miss important papers, or if BibTeX exports contain errors, analysis quality suffers. Careful dataset preparation remains essential. 6.4. Future Directions Several developments promise to enhance LLM-based literature analysis capabilities. Multimodal models that can process both text and images will enable extraction of information from figures, tables, and structural diagrams— addressing a current limitation. Integration with academic databases through APIs would allow real-time literature monitoring, potentially generating automated monthly reports on research developments in specific areas. 144   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Specialized language models trained specifically on scientific literature may offer improved accuracy for domain-specific tasks. Models like MatBERT and ChemBERT demonstrate this potential, though broader adoption requires overcoming data access and computational resource challenges. The development of standardized reporting guidelines for different materials systems would facilitate more systematic data extraction. If papers consistently report performance metrics, synthesis conditions, and stability data using common formats, automated analysis becomes more reliable and comprehensive. Perhaps most significantly, integration of LLM-based literature analysis with other AI tools—property prediction models, synthesis planning systems, experimental automation—could enable more comprehensive research workflows. Imagine a system that monitors recent literature, identifies promising material candidates, suggests experimental protocols, and helps interpret results—all while maintaining human oversight at critical decision points. The framework presented here represents one application of AI in materials science, but the implications extend beyond literature analysis. As AI tools become more sophisticated and accessible, they will increasingly augment human capabilities across research activities. The key to successful integration lies in understanding both capabilities and limitations, using AI tools for tasks where they excel while maintaining human judgment for aspects requiring expertise, creativity, and critical evaluation. Solar cell research provides an ideal testbed for these methods due to its rapid pace of development, quantitative performance metrics, and established reporting conventions. However, the general approach transfers readily to other materials domains: batteries, catalysts, structural materials, and beyond. Researchers in these fields can adapt the framework to their specific needs, adjusting classification schemes and analysis prompts while retaining the core methodology. The integration of AI into scientific research workflows represents neither a threat to human researchers nor a solution to all challenges. Rather, it offers tools that, when used thoughtfully, can help researchers navigate increasingly complex information landscapes more effectively. The framework described in this chapter demonstrates one such application: using LLMs to conduct systematic literature analysis that would be impractical manually while maintaining scientific standards through appropriate verification and human oversight. GENERATIVE ARTIFICIAL INTELLIGENCE IN EDUCATIONAL TECHNOLOGIES   247 ethics-focused educational assistant, is a generative AI assistant designed specifically for education by Khan Academy. It defines its purpose as providing hints and guidance that prompt the student to think, rather than giving direct answers. Announcements and application reports from Khan Academy indicate that Khanmigo focuses on teacher support and student guidance functions and adopts responsible use principles. However, independent, peer-reviewed empirical studies are still increasing (A. Weiss, 2025). Use cases include oneon-one guidance, such as hints encouraging a student to solve a problem stepby-step, and teacher support scenarios like alternative explanations suitable for classroom tasks and monitoring reports. Socratic, provided by Google, is a mobile-supported conceptual guidance application that offers quick explanations and resources by allowing students to upload a photo of a problem or ask a question by voice through their mobile devices. The application is particularly widespread at the high school level and in individual learning, and adaptations for medical and clinical education contexts are also examined in the literature (e.g., adaptation studies for clinical case training) (N. Golchini et al., 2025). Quizizz, an automated quiz tool, through its generative AI plugins, allows teachers to automatically generate question sets from input content (text, video), classify these questions by difficulty level, and provide instant analysis of student performance. While automated question generation and assessment save time, especially for review and testing purposes, the necessity for the pedagogical appropriateness of the created questions to be confirmed by the teacher is emphasized (M. I. Baig et al,. 2024). The visual communication and presentation platform Canva offers a wide range of functions, from AI-supported text summarization and idea generation tools like Magic Write, to automated design suggestions, and various multimedia (video, presentation, infographic) production features. These features have gained great popularity, especially in processes like project-based learning and student presentations, providing students with a powerful digital environment to showcase their creativity and develop their communication skills. However, the adoption of such tools in the classroom requires a rethinking of assessment practices. The ease with which students can create high-quality, professionallooking, and complex multimedia products can render traditional assessment scales inadequate. At this point, it is critical for teachers to shift their assessment focus from the aesthetic quality of the final product to the process and the depth of the content. Elements such as how students manage their creative process, how 248   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . they critically evaluate and structure the design and content suggestions offered by generative AI, research depth, information synthesis skills, and the robustness of the arguments presented should form the basis of authentic assessment. Therefore, it becomes imperative for teachers to adapt their assessment scales and rubrics to consider this new context for the fair and accurate measurement of learning outcomes. Language acquisition is an area that strongly showcases the potential of generative AI to provide personalized and interactive learning environments. Generative AI-supported versions of platforms like Duolingo offer a dynamic language learning experience that goes beyond traditional methods. These systems analyze the user’s performance in real-time to create personalized review schedules; they identify weak grammar rules or vocabulary and aim to increase learning efficiency by focusing on these areas. Furthermore, with exercises for speaking practice and instant error correction (pronunciation, grammar, word choice), they can be highly effective in increasing learning frequency and the amount of fluent speaking practice by offering the learner continuous and immediate feedback. However, a critical perspective on the pedagogical effectiveness and reliability of these tools is necessary. Firstly, the reliability of speech assessment is a significant topic of debate. The extent to which generative AI models can accurately assess factors such as accent diversity, differences in speech rate, and contextual appropriateness remains an area requiring research. While automated systems may be successful in measuring technical accuracy (correct pronunciation of phonemes), they may be limited in assessing higherlevel skills such as communicative fluency and pragmatic competence. 4.3. Sample Lesson Scenarios The pedagogical potential of generative AI becomes more evident when combined with concrete classroom applications and innovative learning models. This section will examine interdisciplinary sample scenarios, as well as emerging integration areas such as micro-learning and virtual/augmented reality. 4.3.1. Sample Lesson Scenarios for Different Disciplines Generative AI provides support to teachers in creating content appropriate to the nature of different disciplines and offers students an active, exploratory learning experience. Looking at sample lesson scenarios, a mathematics teacher could use a large language model with a prompt such as, “Generate 5 different problem scenarios from daily life, such as construction and cartography, using GENERATIVE ARTIFICIAL INTELLIGENCE IN EDUCATIONAL TECHNOLOGIES   249 the Pythagorean theorem for 10th-grade students, and provide a step-bystep solution for each.” While solving these problems, students can request personalized hints from a large language model at steps where they are stuck, such as “Which formula should I use in this step?” or “What should the next operation be?” This process strengthens the internalization of mathematical reasoning and step-by-step problem-solving strategies, rather than focusing solely on the result. A physics teacher can use generative AI with the command “Write a virtual experiment scenario simulating how an astronaut on the International Space Station (a zero-gravity environment) would design a fuel transfer system using fluid pressure principles” to create a comprehensive experiment setup. Then, in a biology class, they can concretize abstract concepts by generating visual material with a tool like Midjourney or DALL-E using the description “a scientifically accurate diagram showing the stages of mitosis—interphase, prophase, metaphase, anaphase, and telophase—with each structure clearly labeled.” In a history class, students can undertake a historical empathy and roleplaying activity by giving a large language model the task: “As an American diplomat in 1945, prepare a draft policy paper containing possible measures to be taken against the Soviet Union’s expansionist policies.” This activity encourages students to examine events from a different perspective, establish cause-effect relationships, and deeply understand complex historical processes. English learners at the A2 level can ask a large language model to “Write a short, clear, everyday dialogue of 4-5 sentences between a tourist and a waiter where the tourist orders coffee and cake in a London café.” Students then have the opportunity to practice this dialogue in pairs, gaining speaking skills, pronunciation practice, and communicative confidence. 4.3.2. Generative AI Integration in Micro-Learning and Gamification Generative AI is in perfect harmony with modern learning trends like microlearning and gamification. Teachers can easily use generative AI to produce targeted micro-learning units of 2-5 minutes, suitable for student attention spans (e.g., a short explanation of a complex formula, a reminder of a grammar rule). Automated daily “on this day in history” notifications, “scientific term of the day,” or “vocabulary” content can be generated; furthermore, quick quizzes and interactive tasks can be designed to reinforce this short content. Gamification elements (points, badges, leveling up, and leaderboards) make these short, 250   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . focused learning sessions generated by generative AI more motivating and sustainable. This combination is ideal for supporting self-directed learning processes. The combination of generative AI with immersive technologies like virtual reality and augmented reality is ushering in an era of integrated learning in education. This integration goes beyond static simulations, holding the potential to create dynamic learning environments that shape themselves according to student responses. For example, when a medical student interacts with a virtual patient in a virtual reality environment, a generative AI model working in the background can generate the patient’s symptoms, responses to questions, and physiological reactions in real-time and dynamically, depending on the student’s interventions. Similarly, in an augmented reality-enhanced history book, when a student scans an ancient Greek statue with their smart device, generative AI can provide not a pre-recorded audio, but an interactive explanation that responds to the student’s specific questions in the voice tone and style of that historical character. These scenarios demonstrate the transformative power of generative AI in turning immersive technologies into personalized, limitless, and deeply interactive learning experiences. 5. Generative Artificial Intelligence in Education: Ethical Boundaries, Privacy, and Risks 5.1. Academic Integrity: Use of Artificial Intelligence in Cheating and Assignments The rapid expansion of generative artificial intelligence (GAI) tools across diverse sectors -including education, academia, defense, and marketinghas become increasingly pronounced. This pervasive adoption has provoked substantial debates, particularly concerning academic integrity. The extensive and direct utilization of Artificial Intelligence (AI) by students in assignments, projects, and theses has the potential to compromise the principles of originality and effort-based learning (Cotton, Cotton, & Shipway, 2023). Advancements in Large Language Models (LLMs) enable students to generate “human-like” text, thereby prompting educators to critically reassess traditional assessment and evaluation practices. While this situation necessitates educators seeking effective solutions, the outright prohibition of Generative AI is neither a practical nor an effective approach. Instead, UNESCO (2023a; 2023b) urges educational institutions to GENERATIVE ARTIFICIAL INTELLIGENCE IN EDUCATIONAL TECHNOLOGIES   251 “regulate AI-assisted tasks in accordance with ethical principles.” Consequently, how students employ AI, their practices in citing sources, and the transparency of their production processes should be systematically incorporated into the teaching and learning framework, fostering responsible and ethical engagement with AI tools. 5.2. Copyright and Content Ownership: Who Owns AI-Generated Content? Creative works are traditionally protected under copyright and intellectual property laws to safeguard the rights of their human creators. However, the emergence of content -such as texts, images, and audiogenerated by contemporary generative artificial intelligence has sparked significant debate at the intersection of law, technology, and education. Unlike human authorship, AI produces a form of “artificial co-creativity” (Samuelson, 2023), raising complex questions about ownership and accountability. Current legal frameworks, however, do not recognize AI as an “author” or rights holder, leaving unresolved challenges regarding the attribution and protection of AI-generated works. In the educational context, uncertainty arises regarding the ownership of content produced by teachers and students. UNESCO (2023a; 2023c) and the OECD (2021) call on educational institutions to develop transparent licensing and sharing policies. Furthermore, it is emphasized that in cases where AI-generated materials undergo human oversight, copyright should remain with the human contributor. 5.3. Data Security and Privacy: Protecting Student Information AI-powered systems process large amounts of personal data, including student profiles, learning habits, and performance levels. This raises critical concerns regarding data privacy, security, and ethical use in education. Floridi and Cowls (2019) provide an ethical framework for AI systems, emphasizing the principles of “transparency, fairness, non-maleficence, accountability, and privacy.” Educational institutions should clearly define data storage procedures when integrating generative AI and ensure compliance with legal standards such as the European Union’s General Data Protection Regulation (EU, 2016). The “Principles for the Use of Artificial Intelligence in Education” guideline published by the Ministry of National Education (MEB, 2024) explicitly prohibits the unauthorized sharing or commercial use of student data. 252   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . 5.4. Artificial Intelligence Literacy: Informed Use by Teachers and Students AI literacy can be defined as the competence of individuals to understand the fundamental functioning, capabilities, limitations, and societal impacts of AI systems, enabling them to interact with these systems in an informed, ethical, and effective manner. Bozkurt (2023) defines AI literacy as a new digital skill area and emphasizes that it should encompass not only technical knowledge but also ethical awareness, verification skills, and elements of critical thinking. Therefore, AI literacy modules should be integrated into teacher education programs. According to researchers such as Ng (2021) and Long & Magerko (2020), AI literacy encompasses the following components: · Conceptual understanding: knowledge of what AI is, how it functions, and the data on which it is trained. · Application skills: the ability to use AI tools consciously and purposefully. · Ethical awareness: sensitivity to issues such as bias, privacy, transparency, and accountability. · Societal awareness: the capacity to assess the impact of AI on individuals, society, and the workforce. In the educational context, AI literacy aims to develop teachers’ and students’ abilities to engage in critical thinking, make ethical decisions, and responsibly manage generative technologies while using AI-based tools for pedagogical purposes. In this regard, it can be considered a contemporary extension of digital literacy (UNESCO, 2023a). 5.5. Responsible Use Guide: Practical Principles and Classroom Examples The responsible use of generative artificial intelligence in education can be defined as the balancing of technological innovation with ethical values. In this regard, OECD (2021) and UNESCO (2023a) propose three fundamental principles: · Transparency: The use of AI should be clearly disclosed. · Human Oversight: Final decisions should be made by human educators. · Equity: The technology should provide fair access to all students. GENERATIVE ARTIFICIAL INTELLIGENCE IN EDUCATIONAL TECHNOLOGIES   253 In the classroom, these principles are implemented when teachers consciously integrate generative AI tools into instructional design. For example, requiring proper citation for texts produced using tools such as ChatGPT or Gemini reinforces students’ ethical awareness. Additionally, analyzing AI-generated erroneous examples during class discussions fosters critical thinking skills. 6. Policy and Future Perspectives 6.1. UNESCO and OECD Frameworks: Ethical and Policy Guidelines in Education UNESCO provides global guidance for the integration of generative artificial intelligence into educational and research processes. In this context, UNESCO’s 2023 document Guidance for Generative AI in Education and Research outlines principles that member countries should consider when implementing both short-term practices and long-term policy planning (UNESCO, 2023a). The guidance advocates for the use of technology for educational purposes within a human-centered vision, adherence to ethical principles, capacity building, and the promotion of technology literacy. The OECD also provides policy principles regarding the use of artificial intelligence in education, emphasizing dimensions such as data privacy, accountability, fairness, and inclusivity, and offering guidance to member countries. These frameworks serve as an “ethical foundation” that enables countries to develop regulations tailored to their own educational systems. An important principle from a policy perspective is continuous monitoring and adaptation; as technology rapidly evolves, policies must remain flexible and be updated to account for emerging risks and opportunities. 6.2. Educational Policies in Türkiye: Projects by the Ministry of National Education (MEB), the Council of Higher Education (YÖK), and The Scientific and Technological Research Council of Türkiye (TÜBİTAK) In recent years, policies regarding artificial intelligence in education have gained momentum in Türkiye. MEB, in its AI in Education Policy Document and Action Plan for the 2025–2029 period, aims to foster an AI culture in education, integrate AI into curricula, implement AI support in management systems, and strengthen infrastructure and data analytics capacity through four strategic objectives, fifteen policy areas, and forty actions. 254   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . The document also prioritizes policies such as the expansion of AI literacy initiatives, the enhancement of teachers’ competencies, and the establishment of international collaborations in the field of educational technologies. TÜBİTAK promotes the integration of educational technologies by supporting infrastructure and R&D projects in the fields of informatics and artificial intelligence. Units such as TÜBİTAK’s Informatics and Information Security Research Center (BİLGEM) conduct research on artificial intelligence, data security, and algorithm development (TÜBİTAK BİLGEM, 2025). In addition, the AI in Education Policy Document prepared by the Artificial Intelligence Policy Association (AIPA, 2023) represents an initiative aimed at incorporating civil society–based approaches and expert perspectives into policy frameworks in Türkiye. In this context, the challenges faced by policymakers in Türkiye include infrastructural disparities, unequal access opportunities, limited teacher competencies, and a lack of awareness. Therefore, during implementation, pilot projects, locally scaled trials, and continuous feedback mechanisms are of critical importance (UNESCO, 2021). 6.3. Generative Artificial Intelligence in Teacher Education: In-Service Trainings and Certificate Programs For policies to be effectively reflected in practice, it is essential to enhance teachers’ competencies. In this context, in-service training, certification programs, and models of continuous professional development support the sustainable use of generative AI. UNESCO’s AI Competency Framework initiative proposes competency domains aimed at developing AI-related skills among both teachers and students. Through this framework, teachers are expected to acquire competencies in technology use, ethical evaluation, guidance, and supervision (UNESCO, 2024a; UNESCO, 2024b). In Türkiye, pilot training programs are being planned through collaborations between the Ministry of National Education (MEB) and universities to enhance teachers’ competencies in AI literacy, data fluency, and the pedagogical use of generative AI tools. Within the policy document’s objective of “Fostering an AI Culture in Education,” particular emphasis is placed on teacher education (MEB, 2025). For teacher training programs to be effective, attention should be given to the following components: GENERATIVE ARTIFICIAL INTELLIGENCE IN EDUCATIONAL TECHNOLOGIES   255 · Modular structure: Modules designed for different skill levels. · Practical examples: Implementation of generative AI tools in real classroom scenarios. · Mentorship and community: Mentoring relationships with experienced teachers and the formation of online professional communities. · Assessment and feedback: Evaluation of participants’ practical outputs and provision of constructive feedback. 6.4. The Future Learning Ecosystem: Human + AI Collaboration It would not be an exaggeration to suggest that future educational environments will evolve into hybrid systems in which humans and artificial intelligence interact and collaborate. In such an ecosystem, AI is expected to assume supportive roles such as instructional assistant, content creator, and feedback tool, while human teachers will continue to carry out critical responsibilities, including pedagogical decision-making, emotional guidance, and fostering learning motivation. UNESCO’s Artificial Intelligence and the Futures of Learning project emphasize the importance of balancing technological and human dimensions, focusing on this collaborative vision. In this approach, AI can be employed not merely as a tool but as a structuring element of the learning environment (UNESCO, 2025). Conceptual models such as HD-AIHED (Human-Driven AI in Higher Education) advocate for the transparent, ethical, and human-supervised use of AI in higher education (Mahajan, 2025). These models aim to establish systems in which AI assumes a supportive role without overshadowing human agency. The requirements for the human + AI collaboration vision in education are as follows (UNESCO, 2023a; OECD, 2021): · Human-centeredness: Decision-making mechanisms should always be aligned with human learning objectives. · Flexibility and adaptability: Systems should respond dynamically to individual student profiles. · Transparency and accountability: The logic behind AI decisions should be clear and understandable. · Continuous improvement: Systems should be capable of self-updating based on performance. 256   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . 6.5. Emerging Trends: AR/VR + Generative AI, Digital Twins, and Metaverse Learning Environments Generative artificial intelligence, increasingly intertwined with emerging technologies, will enable groundbreaking applications in education. · AR/VR + Generative AI: Interactive, real-time content generation in virtual and augmented reality environments; for example, AR-supported lessons, simulation tasks, and virtual laboratories. · Digital Twins: Modeling digital replicas of students’ learning processes, allowing for simulations and personalized guidance based on these replicas. · Metaverse Learning Environments: Lessons, laboratories, and discussion spaces in three-dimensional interactive worlds; dynamic content generation with generative AI and avatar-based instructional scenarios. These trends aim to remove boundaries in education and provide multidimensional learning experiences that encompass both physical and virtual environments. However, the realization of this vision critically depends on infrastructure investments, hardware accessibility, content standards, and ethical frameworks. 7. Conclusion and Recommendations 7.1. Potential: Opportunities Offered by Generative AI Generative artificial intelligence holds significant potential for positive transformation within the context of educational technologies. The benefits that generative artificial intelligence can offer in education are listed below (MEB, 2025; Fullestop, 2025): · Reducing teacher workload through rapid content generation · Providing personalized learning pathways and tailored content for each student · Offering real-time feedback and assessment support · Supporting creativityand project-based learning processes · Gaining insights through data analytics, including student behaviors and areas of difficulty When combined with appropriate policies, infrastructure, and pedagogical approaches, these opportunities can open the door to innovative transformations in education. 263 CHAPTER XIII A BIBLIOMETRIC ANALYSIS OF CANCER RESEARCH UTILIZING GENERATIVE ARTIFICIAL INTELLIGENCE METHODS BETWEEN 2022 AND 2024 Mustafa TEMIZ1 & Burcu BAKIR-GUNGOR2 1(Assist. Prof.) Sivas Cumhuriyet University, Faculty of Economics and Administrative Sciences, Department of Management Information Systems, Sivas, Türkiye E-mail: [email protected] ORCID: 0000-0002-2839-1424 2(Assoc. Prof.) Abdullah Gül University, Faculty of Engineering, Department of Computer Engineering, Kayseri, Türkiye E-mail: bur[email protected] ORCID: 0000-0002-2272-6270 1. Introduction Cancer is one of the most common health problems worldwide, and the burden of disease is increasing every year (Varlamova et al., 2024). According to data from 2022, approximately 20 million new cases and 10 million deaths were recorded, and the number of cases is expected to reach 35 million by 2050 (Bray et al., 2024). Early diagnosis is a key factor in improving patient prognosis, as it is associated with earlier opportunities for intervention and a wider range of treatment options. The prognostic models developed in this context aim to identify individuals at high risk of cancer, diagnose the disease at an early stage or recognise patients at risk of metastasis and recurrence (Moglia et al., 2025). In recent years, particularly over the last five years, methods based on artificial intelligence and machine learning have improved the accuracy rates of histopathological diagnoses. AI-based techniques have increasingly been employed in cancer screening, diagnosis, prognosis prediction, disease 264   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . monitoring, treatment planning, and clinical oncology research, offering powerful tools for uncovering data-driven insights (Varlamova et al., 2024). Generative artificial intelligence (GAI) has initiated a paradigm shift in cancer diagnosis, offering innovative solutions to long-standing challenges. In particular, deep learning-based models, such as generative adversarial networks (GANs), have achieved a high level of accuracy in analyzing and interpreting medical images, in some cases even surpassing the performance of human experts. These technologies offer significant advantages not only in tasks such as image classification or lesion detection, but also in enhancing poor-quality images, completing missing data, and creating synthetic training datasets. Consequently, GAI holds the potential to both improve the reliability of clinical decision support systems and promote more accurate and personalized approaches to early diagnosis and prognosis of cancer (Gautam et al., 2025). By analyzing medical imaging data, biomarkers, and clinical information, artificial intelligence can monitor patients’ responses to treatment, thereby enabling clinicians to update treatment plans in real-time. Machine learning approaches significantly contribute to the identification of critical genetic mutations that drive tumor development, thereby supporting the advancement of targeted therapeutic strategies (Dananjayan & Raj, 2020). Moreover, AI provides valuable insights into tumor genetic heterogeneity, elucidating mechanisms of drug resistance and guiding the design of more effective treatment protocols. Meaningful findings derived from processing clinical records strengthen healthcare professionals’ decision-making processes, while machine learning–based clinical decision-support systems play a guiding role in recommending treatment options and directing patients toward suitable clinical trials (Sakthivel et al., 2024). A growing body of research has investigated the application of generative AI techniques in cancer studies. For example, Sakthivel et al. investigated the implementation of generative models, such as GANs, in medical imaging, focusing on their impact on tasks like anomaly detection and segmentation (Sakthivel et al., 2024). Singh et al. investigated how generative AI is changing oncological imaging, its role in improving diagnostic accuracy, and its integration with multimodal data (Singh et al., 2024). Tan et al. (2025) discussed in detail the applications of generative AI in diagnostic and therapeutic processes in cancer research through a systematic review (Tan et al., 2025). Vadisetty and Polamarasetti (2025) investigated interdisciplinary applications of generative AI in the early detection of oral cancer (Vadisetty & Polamarasetti, 2025).. A BIBLIOMETRIC ANALYSIS OF CANCER RESEARCH UTILIZING GENERATIVE . . .   265 Similarly, Saeed et al. (2023) demonstrated the successful use of augmented datasets created with generative AI in combination with transfer learning models for skin cancer classification (Saeed et al., 2023). In light of these technological developments, the number of studies using generative AI methods for cancer diagnosis, detection, and treatment has increased rapidly. However, no comprehensive study has yet been conducted to investigate the general trends, the structures studied, and their scientific implications. This study aims to analyze research focused on the application of generative AI methods in cancer by examining scientific publications indexed in the Web of Science Core Collection (WoScc) between 2022 and 2024. Specifically, the number of scientific publications, the most prolific authors, pioneering articles, research areas, prominent publishers, and leading universities, as well as the number of citations of these papers, will be analyzed. The study is intended to serve as a guide for researchers interested in the role of generative AI in cancer detection. 2. Materials and Methods This study examines research articles indexed in the Web of Science Core Collection that utilize generative artificial intelligence (GAI) methods and machine learning techniques for cancer detection. Within the scope of this research, the developmental trends of scientific studies are comprehensively analyzed. The dataset was compiled on September 1, 2025, through a search conducted on the WoScc platform using the filters “Cancer” and “Generative Artificial Intelligence.” The analysis primarily focuses on studies published between 2022 and 2024, thereby evaluating research conducted during this three-year period. The included studies employ contemporary generative AI methods in combination with machine learning algorithms for cancer diagnosis. For these studies, the following indicators were identified and assessed: · Number of Scientific Publications · Most prolific authors · Pioneering articles · Research areas · Prominent publishers · Leading universities · Number of Citations 266   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Based on these indicators, the study presents findings related to cancer detection research employing GAI methods. 3. Results a. Number of Scientific Publications Within the scope of this study, a total of 98 scientific publications were retrieved from the Web of Science Core Collection database, based on the predefined search strategy, covering the years 2022–2024. Among these, 58 were research articles, 16 were review papers, 14 were publication materials, 4 were proceedings papers, 3 were letters, and 3 were meeting abstracts. The annual distribution of these publications is presented in Figure 1. Figure 1. Number of publications by year based on the specified search criteria As shown in Figure 1, cancer research employing generative artificial intelligence (GAI) methods has exhibited a steady increase over the years. The highest number of publications was recorded in 2024, with 82 studies. Notably, while only 14 studies were published in 2023, the number increased sharply in 2024, reaching 82 publications —a substantial growth compared to previous years. This remarkable increase can be attributed, in part, to advances in artificial intelligence technologies, which have facilitated the widespread adoption of GAI methods in oncology research. Accordingly, the use of GAI in cancer studies is increasingly recognized as a promising and impactful area of research. A BIBLIOMETRIC ANALYSIS OF CANCER RESEARCH UTILIZING GENERATIVE . . .   267 b. Most Prolific Authors Although the use of generative artificial intelligence (GAI) methods in cancer research represents a relatively new field, it is regarded as a highly reliable area of investigation. Due to its novelty and emerging nature, the number of researchers actively contributing to this domain remains relatively limited. Figure 2 presents the authors with the highest number of publications in this field. Figure 2. Number of publications authored As illustrated in Figure 2, the number of studies conducted in this field remains limited, with the highest contribution being three publications. Among the authors, Lupo Roberto, Ostergaard Soren Dinesen, De Nunzio Giorgio, Conte Luana, Mouayad Masalkhi, Lee Andrew G., Ethan Waisberg, Joshua Ong, and Vitale Elsa stand out with three studies each. The remaining authors have produced two or fewer publications. When the three-year period is considered, it becomes evident that the number of researchers actively engaged in this domain is still relatively low. This suggests that the field has not yet attracted sufficient scholarly attention; however, it also highlights its potential as a promising area for future research. c. Pioneering Articles To identify pioneering contributions, the three most frequently cited studies were considered as seminal works in the field. The first study to lead 268   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . cancer research utilizing GAI methods was conducted by (Lu et al., 2024). Published in the high-impact, Q1-ranked international journal Nature, this article has received 118 citations, thereby establishing itself as a cornerstone in the domain. The second influential contribution was made by (Iannantuono et al., 2023). This study, published in the Q2-ranked international journal Frontiers in Oncology (impact factor: 0.68), has accumulated 56 citations. Finally, the third pioneering article was authored by (Duffourc & Gerke, 2023). Published in the prestigious Q1-ranked journal JAMA – Journal of the American Medical Association, this study has secured 50 citations to date, further underscoring its significance in the emerging literature. d. Research Areas Cancer research, as an inherently interdisciplinary field, holds significant impact in both the domains of information science and healthcare. Relevant studies are published across journals operating within these fields. Figure 3 illustrates the research areas of the journals in which these studies have been published. Figure 3. Publications distributed by research areas As illustrated in Figure 3, the largest proportion of publications in this field falls under Health Care Sciences Services (29.4%). The second most prominent area is Oncology (27.1%), followed by Medical Informatics (25.9%) as another major research domain. Subsequently, Computer Science and A BIBLIOMETRIC ANALYSIS OF CANCER RESEARCH UTILIZING GENERATIVE . . .   269 Public, Environmental, and Occupational Health constitute additional areas of contribution. These findings provide valuable insights for researchers in this field, particularly in guiding the selection of appropriate academic journals and informing decisions regarding the submission of their work. e. Prominent Publishers One of the greatest challenges faced by academic researchers in increasing the visibility of their work is selecting an appropriate journal. When the choice of journal is not aligned with the study’s scope and readership, publications often fail to receive sufficient attention. This outcome is undesirable not only for researchers but also for the scientific community, as it may result in valuable studies remaining underappreciated. In this context, the publishers of journals most frequently selected for cancer research employing GAI methods were examined and presented in Figure 4. As shown in Figure 4, the first two groups of publishers account for the highest number of publications (4 and 2, respectively). In contrast, all remaining journals have published only a single study in this field. Figure 4. Number of publications by publishers As illustrated in Figure 4, the publishing group with the highest number of studies in this field is JMIR Publications, Inc., which accounts for 24 publications. This is followed by Elsevier, MDPI, Nature Portfolio, Springer Nature, Lippincott Williams & Wilkins, Wiley, and IEEE, each contributing 15 publications. An examination of these publishers reveals that they host high- 270   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . impact international journals, underscoring their significance for disseminating research in this area. f. Leading Universities For the topic investigated, universities with the two highest numbers of publications were considered. Including these influential institutions in the analysis provides valuable insights for researchers, enabling them to conduct their searches more effectively and to identify leading centers of expertise in the field. In this context, the universities to which the contributing researchers are affiliated are presented in Figure 5. Figure 5. Universities affiliated with the contributing researchers As illustrated in Figure 5, Harvard University stands out with eleven studies on cancer research utilizing AI and GAI methods. Similarly, Harvard University Medical Affiliates has also contributed to eleven studies, positioning it as another influential institution in this field. Following these, the University of California System ranks next with nine publications, while the University of Texas System has produced eight studies. With seven studies, Harvard Medical School occupies fifth place, followed by Cornell University with six publications. Brigham and Women’s Hospital also emerges as a leading institution with five studies. These universities provide valuable insights for researchers and serve as focal points of interest for scholars aiming to advance cancer research through GAI methods. A BIBLIOMETRIC ANALYSIS OF CANCER RESEARCH UTILIZING GENERATIVE . . .   271 g. Number of Citations The number of citations received by publications serves as an important indicator for researchers in determining areas of focus and research priority. Citation counts can help scholars select topics and shape the direction of their studies. An increase in citation numbers reflects growing interest in the field and indicates a rising number of researchers contributing to the field. In this context, Figure 6 presents the total number of citations over the years. Figure 6. Number of citations received by publications across years As illustrated in Figure 6, a total of 336 citations were recorded across the examined years. Self-citations were included in the analysis. When the results are evaluated on a yearly basis, it is observed that in 2022, during the early stage when artificial intelligence concepts were first being incorporated, 2 studies received a total of 3 citations. In 2023, 14 studies accumulated 30 citations, while in 2024, 82 studies accounted for 303 citations. These findings indicate a growing interest in the research domain, underscoring its increasing importance for scholars in the field. 4. Conclusions Research on cancer using artificial intelligence (AI) methods has become increasingly prevalent in the healthcare domain. Such studies, which are of 272   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . critical importance in diagnostic and therapeutic processes, aim to support physicians in their decision-making. In this context, AI-based approaches have attracted considerable attention. However, research conducted using traditional machine learning methods often fails to achieve the desired level of reliability and remains insufficient. To overcome this limitation, researchers have increasingly turned to innovative AI techniques for disease prediction. This study specifically focuses on cancer-related research employing generative artificial intelligence (GAI) methods. Publications from the period 2022–2024 clearly demonstrate both the timeliness of the topic and the growing academic interest in this field. The growing attention underscores the significance of GAIbased cancer research and suggests that it is likely to become one of the top priority areas for future scientific investigation. Acknowledgement During the preparation of this manuscript, the author used ChatGPT (GPT5, May 2025 version) for the purposes of editing sentence structure and checking for grammatical errors. References Bray, F., Laversanne, M., Sung, H., Ferlay, J., Siegel, R. L., Soerjomataram, I., & Jemal, A. (2024). Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: A Cancer Journal for Clinicians, 74(3), 229–263. https://doi.org/10.3322/caac.21834 Dananjayan, S., & Raj, G. M. (2020). Artificial Intelligence during a pandemic: The COVID‐19 example. The International Journal of Health Planning and Management, 35(5), 1260–1262. https://doi.org/10.1002/ hpm.2987 Duffourc, M., & Gerke, S. (2023). Generative AI in Health Care and Liability Risks for Physicians and Safety Concerns for Patients. JAMA, 330(4), 313–314. https://doi.org/10.1001/jama.2023.9630 Gautam, R., Kaur, P., & Sharma, M. (2025). Predictive Analytics for Early Cancer Detection Using Machine Learning and Generative AI. In Generative Intelligence in Healthcare. CRC Press. Iannantuono, G. M., Bracken-Clarke, D., Floudas, C. S., Roselli, M., Gulley, J. L., & Karzai, F. (2023). Applications of large language models in cancer care: Current evidence and future perspectives. Frontiers in Oncology, 13. https://doi.org/10.3389/fonc.2023.1268915 GENERATIVE ARTIFICIAL INTELLIGENCE APPLICATIONS IN LAW   279 of law related to the decision being sought.10 Similarly, Ross, an AI-powered legal assistant, provides answers by conducting research on court decisions and relevant legislation. Another area where artificial intelligence is utilized in the field of law is the statistical analysis of court decisions. These programs analyze judicial reasoning to determine the factors influencing a judge’s decision to accept or reject a case. Based on these findings, they attempt to predict potential court outcomes. As a result of such analyses, individuals considering filing a lawsuit can learn about their likelihood of winning or losing, and these insights may influence their decision on whether or not to proceed with legal action.11 The use of artificial intelligence in locating and analyzing court decisions also brings with it the risk of AI hallucination. Hallucination, in this context, refers to the phenomenon of artificial intelligence generating incorrect or fabricated information.12 As a result of hallucination, artificial intelligence may fabricate court decisions, which can lead to incorrect judgments, disciplinary liability for the lawyer who presents such information, and delays in judicial proceedings if the erroneous decision is later detected and corrected.13 Indeed, in the United States, there have been instances where lawyers using GenAI were sanctioned for submitting fabricated court decisions produced by AI during legal proceedings.14 In this context, it is primarily the responsibility of lawyers— who rely on such decisions—to verify them, although this duty also extends to judges. According to Article 34 of the Attorneyship Law, lawyers are obligated to perform their duties with due diligence. Within the scope of this duty, a lawyer must also ensure the accuracy and reliability of any auxiliary tools, including artificial intelligence systems, used during a case. Similarly, in an era where AI has become part of legal practice, judges must consider the possibility that a submitted decision may have been generated by artificial intelligence and are therefore required to verify its authenticity. 10. https://www.dejure.ai/nedir (E.T. 14.10.2025) 11. “Yapay Zekanın Karar Verme Süreçlerinde Kullanılması”, s. 325. 12 . Ankara Barosu Avukatlıkta Yapay Zeka Araçlarının Kullanım Rehberi, s. 3 https:// www.ankarabarosu.org.tr/serve/file/bc4de0be-b5fa-11ef-8f94-000c29c9dfce/yapay_zeka_araclarnn_kullanm_rehberi_X1.pdf (E.T. 14.10.2025) 13. Selçuk, Seyhan; “Makûl Sürede Yargılama Yapılmasında Yapay Zekanın Etkisi”, Türkiye Adalet Akademisi Dergisi, Yıl 16, S. 62, Nisan 2025, s. 458. 14.https://www.reuters.com/technology/artificial-intelligence/ai-hallucinations-court-papers-spell-trouble-lawyers-2025-02-18/ (E.T. 14.10.2025) 280   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Legal Research, E-Discovery, and Client Communication In the research process of legal disputes, it is necessary not only to conduct case law research but also to examine legislation and legal doctrine. Reviewing documents, identifying evidence, and determining the applicable laws are among the essential duties of legal professionals. When performed manually, such research and examination can take several days depending on the complexity of the case file. However, with advances in artificial intelligence, systems such as LexisNexis can now carry out these tasks much more efficiently. A key advantage of these AI systems is that they go beyond simple keyword searches and are capable of conducting semantic analysis, enabling a deeper and more accurate understanding of legal texts.15 However, although these systems significantly accelerate the research process, legal interpretation and the unique characteristics of each specific case remain crucial in the practice of law. Furthermore, due to factors such as differences in local legislation and the risk of AI hallucination, human oversight by lawyers is still essential to ensure accuracy and reliability in legal analysis. In large-scale lawsuits, identifying and classifying electronically stored data that may be useful in litigation is a time-consuming process. E-discovery applications facilitate this process by enabling the rapid examination of electronically maintained records of the parties, thereby assisting in the efficient identification and categorization of relevant evidence.16 In the use of e-discovery applications, ensuring data confidentiality is one of the key aspects that lawyers must pay close attention to. As mentioned earlier, the lawyer’s duty of confidentiality and their responsibility as a data controller under the Law on the Protection of Personal Data (LPDP) both come into play in this context. In both pre-trial preparation and trial processes, communication with the client plays a crucial role. Artificial intelligence applications such as Everlaw, AVA, and SANDI, which handle responsibilities like tracking case progress, organizing documents, and communicating with clients—particularly potential clients—enhance both efficiency and speed. However, since these systems lack the human qualities of interaction and empathy, they carry the risk of creating a negative client or prospective client experience compared to direct human communication. 15. “Yapay Zeka Modellerinin Avukatlık Mesleğinin Geleceği Üzerine Olası Etkileri”, s. 7. 16. Efe, Ahmet; “Yargısal ve Hukuki Süreçlerde Yapay Zeka Kullanan Araçlar Üzerine Bir Değerlendirme”, Bilgi Yönetimi Dergisi, C. 5, S. 1, 2022, s. 107 GENERATIVE ARTIFICIAL INTELLIGENCE APPLICATIONS IN LAW   281 Drafting Judicial Decisions and the Concept of the Robot Judge When examining the use of artificial intelligence by judges during legal proceedings, this can be categorized into three main types: AI assisting the judge, AI generating draft judgments, and AI rendering judicial decisions. An AI system that assists judges includes applications such as case law retrieval, identifying relevant legislation, analyzing and summarizing documents, and transcribing statements—as mentioned earlier. One of the countries utilizing such systems is the People’s Republic of China, where Internet Courts employ speech recognition technology to automatically transcribe statements made during proceedings. Additionally, China has implemented a system called “Smart Judge” (Bilge Hakim), designed to promote consistency and uniformity among judicial decisions.17 In India, a system called SUPACE (Supreme Court Portal for Assistance in Court Efficiency) performs functions that assist judges in their work, such as organizing and analyzing case information. Similarly, in Finland, a system known as Anappi supports judges by handling various administrative and analytical tasks to improve judicial efficiency.18 Some of the tasks currently performed by judicial clerks or judge assistants can now be handled by AI systems that assist judges. One of the major advantages of using artificial intelligence in this context is its ability to accelerate the judicial process, thereby contributing to shorter trial durations and overall increased efficiency in case management.19 One of the areas where artificial intelligence is applied is drafting judicial decisions. In this context, AI systems analyze previous court rulings on similar cases and generate a draft judgment based on the identified patterns and reasoning.20 The draft prepared by artificial intelligence in this context serves merely as an assistive tool for the judge and does not carry any binding authority. The judge may review the draft decision and choose to adopt it as is, modify it, or issue a completely different ruling. Among the AI systems that generate such draft judgments are “Predictive Jurisprudence”, used by the Genoa Court in Italy, and the “Robot Judge” application implemented in China.21 This artificial 17. Karadeniz, Salih; “Medeni Yargılamada Yapay Zeka Kullanımı: Hakim Yapay Zeka, Faydaları ve Sakıncaları Üzerine Bir Değerlendirme”, Gelişen Teknolojilerin Medeni Usul Hukukuna Etkileri, Seçkin, 2023, s. 11 18. “Makûl Sürede Yargılama Yapılmasında Yapay Zekanın Etkisi”, s. 440. 19. “Makûl Sürede Yargılama Yapılmasında Yapay Zekanın Etkisi”, s.447. 20. Bilgin, Hikmet; “Yapay Zekânın Mahkeme Kararlarında Kullanımına Uluslararası Bir Bakış ve Robot Hâkimler Hakkında Düşünceler”, İnÜHFD 13(2):405-419 (2022), s.411 21. “Yapay Zekânın Mahkeme Kararlarında Kullanımına Uluslararası Bir Bakış ve Robot Hâkimler Hakkında Düşünceler”, s. 414; “Makûl Sürede Yargılama Yapılmasında Yapay Zekanın Etkisi”, s. 451. 282   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . intelligence application is likely to yield positive outcomes in promoting consistency and uniformity in judicial decisions. One essential component of the right to a fair trial is the requirement that court decisions be reasoned. However, the absence of reasoning in draft judgments generated by artificial intelligence may lead to a violation of this right. Therefore, in order for AI systems to fully assist judges and ensure the reliability and legitimacy of such decisions, it is essential that they possess the ability to provide proper legal reasoning in their outputs.22 Since artificial intelligence generates these drafts based on existing court decisions, one of the main risks of such systems is their inability to produce consistent rulings in cases that are rare, lack sufficient precedents, or involve novel legal issues. Similarly, in dispute types where there are numerous and conflicting precedents, the likelihood of AI producing a coherent and balanced draft judgment decreases. Moreover, in some cases, evolving social and legal needs may lead to shifts in judicial reasoning over time. However, AI systems that rely solely on past decisions are unable to adapt to such changes. In these circumstances, the judge’s human qualities—such as the ability to assess the unique facts of a case, exercise judicial discretion, and render decisions in line with society’s changing values and needs—remain indispensable.23 Using AI systems that generate draft judgments as assistive tools under the supervision and control of judges helps to mitigate potential risks while promoting the development of an application that enhances the consistency and coherence of judicial decisions.24 Another form of artificial intelligence use at the decision-making stage is the robot judge system, in which the decision is rendered directly by the AI itself. In this model, artificial intelligence effectively replaces the human judge and carries out the entire judicial process autonomously. A current example of this can be found in China’s Internet Courts, where such systems have been implemented. 25 There are opinions suggesting that the use of robot judges would be more appropriate in the initial stages of implementation, specifically for simpler, repetitive cases or non-contentious judicial matters. In such proceedings, where 22. “Makûl Sürede Yargılama Yapılmasında Yapay Zekanın Etkisi”, s. 452-453. 23. “Makûl Sürede Yargılama Yapılmasında Yapay Zekanın Etkisi”, s. 452 24. Yılmaz, Oğuz Gökhan, “Yargı Uygulamasında Yapay Zeka Kullanımı-Yapay zeka Hakim Cübbesini Giyebilecek mi?”, Adalet Dergisi, 2021/1, S.66, s. 406. 25. “Medeni Yargılamada Yapay Zeka Kullanımı: Hakim Yapay Zeka, Faydaları ve Sakıncaları Üzerine Bir Değerlendirme”, s. 17. GENERATIVE ARTIFICIAL INTELLIGENCE APPLICATIONS IN LAW   283 the reasoning and discretionary powers of a human judge are less critical, the risk of AI making erroneous decisions is comparatively lower.26 Appeals against decisions rendered by an AI judge would still be possible. However, this raises the question of whether the reviewing authority should be another AI judge or a human judge. In the appellate process, the key issue becomes whether such disputes should be reviewed by humans or machines. At present, the prevailing view is that appeals and judicial reviews should be conducted by human judges, ensuring oversight, accountability, and the preservation of fundamental judicial principles.27 Robot judges are beneficial in terms of delivering fast and consistent decisions during judicial proceedings. However, as mentioned in the previous section, their use raises important concerns regarding how case-related information is processed and the protection of personal data throughout the adjudication process. One of the key risks associated with robot judges is algorithmic bias, which refers to situations where artificial intelligence exhibits unfair or discriminatory behavior. It has been observed that AI systems can sometimes develop biases based on the data sets on which they were trained. Consequently, if a robot judge is trained on biased or unbalanced data, discriminatory tendencies may also manifest in its judicial decisions, raising serious concerns about fairness and equality before the law.28 3. GenAI Applications in Criminal Law Artificial intelligence applications, which can perceive faster, make decisions, and generate new solutions more efficiently than human intelligence, are increasingly used not only in private law but also in criminal law, and their scope of application continues to expand day by day. In the context of criminal law, AI can be utilized both in the criminal procedure process—the stage following the commission of a crime—and in determining the appropriate sentence for the offender after trial. However, the use of AI in criminal law is not limited to these stages; it can also be employed in preventive criminal law practices, that is, in the pre-crime stage, to help prevent offenses before they occur. 26. “Medeni Yargılamada Yapay Zeka Kullanımı: Hakim Yapay Zeka, Faydaları ve Sakıncaları Üzerine Bir Değerlendirme”, s.18 27. Sümer, Seda Yağmur; “Ceza Yargılamasının Geleceği: Robot Hakim” (2021) 23 (2) Dokuz Eylül Üniversitesi Hukuk Fakültesi Dergisi, s. 1583-1584. 28. “Medeni Yargılamada Yapay Zeka Kullanımı: Hakim Yapay Zeka, Faydaları ve Sakıncaları Üzerine Bir Değerlendirme”, s. 23. 284   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . GenAI Systems in Preventive Criminal Law The objective of preventive criminal law systems is based on the belief that the danger and harm caused by a crime cannot be completely eliminated solely through the punishment imposed on the offender29. These systems can be employed not only before a crime is committed but also during the assessment of the likelihood that an offender may reoffend. The ability to predict the probability of recidivism is particularly significant both at the criminal procedure stage and during the execution of a sentence, as it allows for more informed decisions regarding preventive measures, parole, and rehabilitation processes30. In this context, the COMPAS application used in the United States is a risk assessment tool that analyzes an individual’s historical data to provide authorities with a risk score indicating the likelihood of the person reoffending in the future31. However, it is evident that this AI system—which considers factors such as gender, race, and skin color when generating risk scores—poses significant concerns regarding fundamental principles of criminal procedure, particularly the presumption of innocence, which ensures that every individual is deemed innocent until proven guilty by a final court judgment. Similarly, the Static-99 tool, another AI-supported system, has been developed specifically for assessing the risk of sexual offense recidivism. It evaluates factors such as the offender’s relationship with the victim and the timing of the previous offense to estimate the likelihood of the individual committing another sexual crime32. Some preventive predictive systems analyze the locations and other factors of previously committed crimes and provide authorities with data about where 29. İçer, Zafer; “Yapay Zekâ Temelli Önleyici Hukuk Mekanizmaları - Öngörücü Polislik”; Yapay Zeka Temelli Teknolojiler ve Ceza Hukuku, İstanbul Barosu, Yapay Zeka Çalışma Grubu Yıllık Rapor, 2021, s.31, https://www.istanbulbarosu.org.tr/files/komisyonlar/yzcg/2021yzcgyillikrapor.pdf (E.T. 13.10.2025) 30. The belief that the offender may reoffend is a factor considered in the application of many legal institutions. Among the conditions required for the implementation of several institutions— such as the suspension of the execution of a prison sentence regulated under Article 51 of the Turkish Penal Code No. 5237, the deferment of the announcement of the verdict regulated under Article 231 of the Criminal Procedure Code No. 5271, and the conditional release regulated under Article 107 of the Law No. 5275 on the Execution of Sentences and Security Measures—is that the court must reach the conviction that the individual is unlikely to commit another crime 31. Sapan, Oğuzhan; Ceza Muhakemesinde Yapay Zeka Kullanımı, Adalet Yayınevi, İstanbul, 2024, s.70-75. 32. Abanoz Öztürk, Buket; “Suç Davranışını Öngören Üretken Yapay Zekâ Araçlarının Ceza Muhakemesi Hukukunun Temel İlkeleri Bağlamında Değerlendirilmesi”, İstanbul Barosu, Üretken Yapay Zekâ ve Hukuki Meseleler Konferans Bildirisi, https://www.istanbulbarosu.org.tr/ files/komisyonlar/yzcg/yzcg_uretkenyzvehukukimeseleler.pdf s. 18. GENERATIVE ARTIFICIAL INTELLIGENCE APPLICATIONS IN LAW   285 similar offenses are likely to occur in the future. Applications such as PredPol and Precobs make such predictions33. However, it should not be forgotten that using locations of past crimes to predict new crime sites can lead to increased surveillance of people in economically disadvantaged areas and may cause discrimination34. In the context of preventive systems, several AI-supported technologies are utilized within preventive law enforcement. These include MOBESE and similar camera surveillance systems, facial recognition–based security systems developed to ensure secure entry and exit to judicial institutions and public buildings, as well as biometric verification systems. All of these applications employ artificial intelligence to enhance monitoring, identification, and preventive security measures. Using GenAI mechanisms, a general profile can be created from data in previously adjudicated cases, and the pool of suspects in new, similar incidents can be narrowed based on that profile35. In these systems, the physical characteristics of individuals who have previously committed or attempted to commit crimes are stored in memory. Later, when a person enters the coverage area of facial recognition systems, their movements, gestures, and facial expressions are analyzed; if these match the data stored in the system, the goal is to trigger preventive law enforcement actions before a potential crime occurs36. However, alongside certain beneficial outcomes, it should not be forgotten that these systems constitute an intrusion into personal data. Moreover, if people are aware of the locations of facial recognition systems, they may take measures to avoid being recognized or detected, and this should not be overlooked. Naturally, since artificial intelligence is a human creation, it is also susceptible to unauthorized access by malicious third parties, making it possible for the data of individuals captured by the system to be altered or misused. In addition to this risk, the presence of numerous variables—such as a person’s emotional state or environmental conditions—raises another concern: the possibility that the system 33. “Suç Davranışını Öngören Üretken Yapay Zekâ Araçlarının Ceza Muhakemesi Hukukunun Temel İlkeleri Bağlamında Değerlendirilmesi”, s. 17. 34. Ateş, Hüseyin; Ceza Muhakemesinde Kullanılan Yapay Zeka Uygulamalarının Ceza Muhakemesi İlkeleri Açısından Değerlendirilmesi, Yayınlanmamış Doktora Tezi, 2025, s. 90. 35. Dijital Ceza Muhakemesi Hukuku, Ed: Öztürk, Bahri/Tezcan, Durmuş/Erdem, Mustafa Ruhan, Seçkin, Ankara, 2024, s. 229. 36. İçer, Zafer/Dönmez, Elif; “Yüz Tanıma Teknolojilerinin Önleyici Ceza Hukuku ve Ceza Muhakemesi Süreçlerindeki Kullanımı ve Sınırları”, Ceza Hukuku Dergisi, Ağustos 2020, Y.15, S.43, 421-461. 286   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . may record data only for certain types of offenders, thereby operating in a biased and probabilistic manner rather than as an objective and reliable tool. Despite the many positive outcomes of preventive artificial intelligence mechanisms, it is undeniable that they restrict individuals’ freedom of movement. They pose significant risks to the right to privacy in general and to the protection of personal data in particular. Moreover, such systems may also enable malicious uses, such as the increased surveillance of groups that may already be considered socially or economically disadvantaged, thereby exacerbating existing inequalities37. For this reason, AI-supported preventive mechanisms—which involve the monitoring and recording of personal data— must be grounded on solid legal foundations. They should be used solely as supportive tools to assist law enforcement activities, while the final decision must always be rendered by a human authority, not by the AI system itself. In this way, in the event of any legal violation, it will be possible to identify and assign criminal and legal responsibility to the appropriate party. GenAI Applications in Criminal Procedure Law In addition to its use prior to the commission of a crime, artificial intelligence—and particularly GenAI applications—can also be employed during the trial phase, which constitutes the stage following the commission of the offense. At the investigation stage, which marks the initial phase of criminal procedure, the use of AI-supported applications may play a significant role. During an investigation, establishing a complete, accurate, and coherent sequence of events is crucial for uncovering the material truth. At this stage, GenAI applications can assist by compiling the statements taken by law enforcement officers or the public prosecutor together with the interrogation records prepared by the judge, thereby constructing a comprehensive narrative of the case. Of course, in doing so, the system would not rely solely on witness or suspect statements but would also analyze lawfully obtained camera footage and audio recordings to reach conclusions. It could identify inconsistencies between statements or between statements and recorded evidence, and based on these discrepancies, generate alternative narratives for each possible scenario. In this way, during complex and large-scale investigations, the AI system could quickly 37. Özbalçık, Ozan Can; “Ceza Muhakemesinde Yapay Zeka Temelli Risk Değerlendirme Araçları ve Hukuki Etkileri”, Yapay Zeka Temelli Teknolojiler ve Ceza Hukuku, İstanbul Barosu, Yapay Zeka Çalışma Grubu Yıllık Rapor, 2021,s.93, https://www.istanbulbarosu.org.tr/files/komisyonlar/yzcg/2021yzcgyillikrapor.pdf (E.T. 13.10.2025) GENERATIVE ARTIFICIAL INTELLIGENCE APPLICATIONS IN LAW   287 provide criminal procedure practitioners with a comprehensive overview of the case, helping them to visualize the broader context more efficiently. As mentioned above, generative AI-based systems construct the sequence of events by utilizing the available evidence. Artificial intelligence can also be employed both in the generation of digital evidence and in the evaluation of such evidence, along with other types of proof. Indeed, at this stage, criminal procedure seeks to employ the most advanced technologies to uncover the material truth in the fastest and most reliable manner possible38. In criminal procedure, the term digital evidence refers to data that is created, stored, or transmitted in an electronic environment39. In this context, digital evidence includes electronic communications, audio and video materials, computer programs, both hidden and accessible files of any kind, previously visited websites, and even deleted data—all of which constitute data existing in a digital environment40. In terms of the production of digital evidence, it is first possible to restore existing but deleted or corrupted data to transform it into usable and reliable evidence. Additionally, AI systems can generate digital reconstructions or simulations based on the available evidence, thereby creating new digital representations of the event. However, in such cases, it would be more appropriate to treat these AI-generated materials not as conclusive evidence, but rather as interpretations or reconstructions of existing evidence. For digital data generated by artificial intelligence to be considered admissible evidence, several conditions must be met simultaneously. First, the method of generation, including the algorithm used and the data sources, must be transparent and the output must be verifiable through human oversight. Additionally, as with all other types of evidence, digital evidence must be obtained lawfully, and the use of generative AI must not result in violations of personal data, unlawful interception of private communications, or infringement of privacy. Finally, the algorithm employed must be subject to judicial review, and the outcomes must not be random or speculative in nature. Another critical point to consider regarding the production of digital evidence is the creation of fake evidence, commonly known as deepfake. With recent technological advancements, it has become extremely easy to generate manipulated images, fabricated audio recordings, or falsified videos, 38. Ceza Muhakemesinde Yapay Zeka Kullanımı, s.205. 39. Dijital Ceza Muhakemesi, s. 443. 40. Dülger, Murat Volkan; Bilişim Suçları ve İnternet İletişim Hukuku, Seçkin, 2025, s.668 vd. 288   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . highlighting the necessity of adopting a cautious and skeptical approach toward all forms of digital evidence41. In addition to generating digital evidence, AI systems can also be used to evaluate it; in this context, they can sift through hundreds or even thousands of lawfully obtained items—such as camera footage and emails—to determine where individuals were and with whom at the date and time an offense occurred, thereby reducing human error42. GenAI Applications in Sentencing Determination Generative artificial intelligence applications can also be utilized at the sentencing stage, which represents the final outcome of criminal proceedings. In this context, certain AI systems that provide statistical analyses can function as judicial assistants by identifying keywords and factual elements within the case file and presenting the judge with information on types and ranges of penalties imposed in similar past cases43. In another scenario, an AI application could compare the details and judgments of previous case files with the facts of the current case and provide the judge with a recommendation regarding the appropriate type of sentence that could be imposed. Such a recommendation would serve merely as advisory guidance, not as a binding decision. However, regardless of the circumstances, the imposition of a prison sentence, which entails the restriction of an individual’s liberty, should never be left to an AI system. Sentencing decisions must remain the result of the judge’s discretionary authority, grounded not only in legal knowledge but also in judicial experience and human judgment. 4. Results and Conclusion As in many other fields, GenAI applications have made significant positive contributions to the legal domain. Primarily, by saving practitioners time and effort in both branches of law, these systems can enhance human productivity and reduce error rates in situations where factors such as workload or extensive 41. Abanoz, Buket; “Derin Sahte (Deepfake) Teknoloji Karşısında Türk Ceza Hukuku”, Yapay Zeka Temelli Teknolojiler ve Ceza Hukuku, İstanbul Barosu, Yapay Zeka Çalışma Grubu Yıllık Rapor, 2021, s.68, (E.T.13.10.2025.) https://www.istanbulbarosu.org.tr/files/komisyonlar/ yzcg/2021yzcgyillikrapor.pdf 42. Dijital Ceza Muhakemesi Hukuku, s.231. 43. Kabak Yüce, Emine; “Cezanın Belirlenmesinde Yapay Zeka Temelli Sistemlerin Kullanımının Değerlendirilmesi”, Ceza Hukuku Dergisi, 48 (2022), s. 100. 295 CHAPTER XV GENERATIVE AI IN MEDIA, COMMUNICATION AND VISUAL ARTS Oğuz KAYNAR1 & Murat Fatih TUNA2 1(Prof. Dr.) Sivas Cumhuriyet University, Sivas, Türkiye E-mail: [email protected] ORCID: 0000-0003-2387-4053 2(Assoc. Prof. Dr.) Sivas Cumhuriyet University, Sivas, Türkiye E-mail: [email protected] ORCID: 0000-0002-8634-8643 1. Introduction The digital transformation process can be described as a broad restructuring across various fields, from individuals’ daily practices to social structures, industrial processes to economic sustainability, through the integration of innovative technologies such as internet technologies, artificial intelligence, and algorithmic processes (Feher, 2025). Artificial intelligence, in particular, has become a trend of the future, affecting the entire media sector, from the generation of digital content to the distribution of news. Following the changes in digital communication technologies, the sector, now known as the Media and Content Industry (MCI), has become receptive to innovations brought about by related developments (Simon, 2013). Art, highlighted as one of these innovations, has merged with digital media technologies and come to be known as digital media art. The practice of digital media art is a new field of artistic application research and creation that integrates modern science and technology with traditional art fields (Wang, 2022). New media art, an innovative fusion of digital technology and artistic creativity, is restructuring cultural experiences in innovative ways through individual aesthetic perceptions (Y. Zhao, 2024). At this point, artificial intelligence (AI) continues to increase its influence in the field of visual 296   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . communication in new media art, as it does in every area of society (Lyu et al., 2022). Statistics presented in an industry research report published by Grand View Research also point to this effect (Grand View Research, 2025). The report, which highlights the use of AI in the media and entertainment sector between 2025 and 2030, estimates that the relevant market will grow at a CAGR of 24.2% between 2025 and 2030, reaching USD 99.48 billion by the end of the period. Researching the synergistic effect of artificial intelligence technology in the evolution of visual communication in new media art offers broad opportunities for practical application and valuable contributions to theoretical discussions (Hong & Curran, 2019; Y. Zhao, 2024). Artificial intelligence, guided by the science and technology of intelligence, not only enables a renewed understanding of cognitive processes through its capacity to develop and simulate human intelligence, but also offers unlimited possibilities in the field of visual communication in new media art (Lyu et al., 2021). However, AI’s simulation of human intelligence and creativity, while enhancing artistic expression within the text, also allows for the manipulation of reality. This duality carries the risk of misuse of generative models and can pose serious problems regarding the accuracy and reliability of the generated content (Kalantzis & Cope, 2025). Furthermore, as AI-generated content becomes more widespread, it requires a clearer understanding of human-machine interaction, particularly in terms of ensuring efficiency (Lyu et al., 2021; Wu et al., 2020). Creating high-quality images from textual descriptions is a persistent challenge with numerous practical applications. This challenge spans multiple fields, including media, industrial design, e-commerce, education, and many others (Arya et al., 2024). However, its widespread potential for use and the ease of access to digital technologies for individuals have made it difficult to distinguish between authentic images and generated images. While this generated vision has the potential to meet industrial needs, it also introduces several methods that pose privacy and security concerns (Westerlund, 2019). The term Deepfake, derived from the combination of the words “deep learning” and “fake”, is at the forefront of these methods. This technology, which can superimpose and alter images, video clips, and generate convincing fake images, leaves no trace of manipulation (Chawla, 2019). Using Deepfake, users can replace one person’s face with another’s in an image or video, and the original voice and facial expressions in an image or video can also be altered (Chadha et al., 2021). Based on this, the current study aims to investigate the role of GenAI in increasing activity in media, communication, and visual arts. To achieve this goal, the study began with the technical evolution of visual generation models and GENERATIVE AI IN MEDIA, COMMUNICATION AND VISUAL ARTS   297 continued with the role of artificial intelligence in image and video generation. Furthermore, existing technologies have been reviewed, along with supporting literature, to provide insight into possible future research directions. In this regard, opportunities, risks, and threats related to real-time and multi-interactive environment solutions have been addressed within the scope of the study. 2. Technical Evolution of Visual Generation Models People perceive the world and the objects within it through their senses; in a sense, they obtain concrete and cognitive stimuli through their interaction with these objects. The sense of sight, facilitated by images and videos, enables people to get details about the external appearance of objects (such as shape and color) without requiring physical contact (Xi et al., 2024). However, images are also used to describe perceived objects, and the rise of AI technologies in the narratology of these images continues. Takashi & Jumpei (2020) interpret this situation as an increase in the role of artificial intelligence in filling the gap between cognitive science and visual narratology. The use of artificial intelligence is seen as an alternative and low-cost way to enhance the narratology of an object (Gulsoy et al., 2024). Due to their wide range of applications, the generation of visual models using artificial intelligence has become a highly popular topic. Generating high-quality images based on a wide range of styles and levels of realism continuously improves the quality of outputs, thereby convincing the target audience of the image’s realism (Zhang et al., 2019). Despite recent advances in generative image and video modeling (Brock et al., 2019; Wang et al., 2024), achieving AI-based generations from complex datasets remains a challenging goal. Laba (2024) conceptualizes AI-based image generation as a socio-technical practice at the intersection of humans, machines, and culture. Accordingly, visual AI generative structures organize the narratological themes presented to them into a layout based on these three elements. Thus, most image generators working with GenAI technologies enable users to create high-quality images with simple commands, even if they lack advanced visual design skills or detailed artistic expertise (Park et al., 2024). These images are used in various fields, including gaming, media, news, scientific visualization, artistic creation, tourism, advertising, design, engineering, e-commerce, and digital marketing. However, the use of these generative visual models has continued to evolve in line with the needs of the the aforementioned industries. The relevant evolutionary process is shown in Figure 1 (Laba, 2024; Park et al., 2024; Praveen et al., 2025; Si & Bao, 2025; Tatar et al., 2018). 298   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Generative Artificial Intelligence (GenAI) refers to AI structures that utilize deep learning models to generate human-like content based on complex and diverse prompts (Lim et al., 2023). Unlike predictive artificial intelligence, which focuses on analyzing input data to make predictions or decisions, GenAI aims to create new content or data that is similar to, yet distinct from the training data (Heigl, 2025). When examining the development process of Generative Visual Models, which enable GenAI, it becomes clear that autoencoders are at the core of the technical evolution. These structures are one of the widely known deep learning methods that generate new data sampled from a learned latent space (Kingma & Welling, 2022). Variational models that use neural networks for recognition models are at the forefront in this structure (Hinton & Salakhutdinov, 2006). Figure 1. Evolution of Generative Visual Models In the following section, developments in the field of text-to-image and text-to-video generation will be examined within the framework of both proprietary commercial products and fundamental theoretical approaches. Since all of the structures presented are essentially based on an algorithmic foundation, the term “model” will be used as an umbrella term in this study to refer to both commercial products and academic approaches. GENERATIVE AI IN MEDIA, COMMUNICATION AND VISUAL ARTS   299 3. Visual Generation Models According to Generation Paradigms GenAI Visualization Models are divided into three categories based on generation paradigms. These are the VAE (Variational AutoEncoder) probabilistic sampling paradigm from the latent space, the GAN (Generative Adversarial Network) adversarial generation paradigm, and the DM (Diffusion Models) iterative denoising paradigm (Guarnera et al., 2024; Rais et al., 2024; Sordo et al., 2025). These models are compared in Table 1. Table 1. Comparison of VAE, GAN and DM Models MODEL VAE GAN DM Generation Paradigm Probabilistic Sampling Adversarial Generation Iterative Noise Reduction Basic Idea Encodes the data into a compressed latent space and generates images by sampling new points from this space. While a “generator” generates fake images, a “discriminator” attempts to distinguish them from real ones; this competition improves quality. Gradually adds noise to an image and then learns to create a clean image from pure noise by reversing this process. Basic Difference Focuses on encoding and decoding. Prioritizes representation over image quality. Indirectly learns the data distribution through competition based on game theory. Creates the image not in a single step, but through a multi-step purification process. Pros Stable training, meaningful latent space, fast generation Sharp and realistic images, fast generation. Highest image quality, consistent training. Cons Due to reconstruction loss, it generally generates images that are blurrier and less detailed than those generated by others. Unstable training (characterized by problems such as mode collapse) and difficult to tune. Slow generation, high computational cost. Among these models, the VAE only learns the distribution of the data, rather than a compressed image, and can decode and generate new data using this distribution. The two components of the model, the encoder and decoder, are shown in Figure 2a. 300   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . Figure 2a. Encoder-Decoder Structure in the VAE Model (DiShi Zhu, 2020) The running principle of the model is shown in Figure 2b (DiShi Zhu & Vera Tang, 2020). Figure 2b. A Framework for VAE (DiShi Zhu & Vera Tang, 2020) In this model, the encoder attempts to learn parameters φ to compress the input data x into a hidden vector z, and the output encoding z is obtained from a Gaussian density with parameters φ. The input to the decoder is the encoding z, which is the output of the encoder. It parameterizes the reconstructed x via parameters θ, and the output x is obtained from the data distribution. Another model, GAN, is an algorithm that utilizes two neural networks: the generator, denoted as ‘G’ and the discriminator denoted as ‘D’. These two networks compete with each other (hence the term “adversarial”). The framework for the model is shown in Figure 3. GENERATIVE AI IN MEDIA, COMMUNICATION AND VISUAL ARTS   301 Figure 3. A Framework for GAN (Kamat, 2025) According to the framework in the image, the discriminator must master distinguishing whether an image is “real” or “fake”. At the same time, the generator must improve itself to generate images realistic enough to be approved by the discriminator. This continuous competition forces the generator to continually improve itself, ultimately generating extremely realistic and highquality images that are nearly indistinguishable from the original dataset. Diffusion Models perform image generation through a mechanism that gradually converts an image into noise (forward process) and then learns to clean the noise step by step by reversing this process (backward process). This enables the model to synthesize high-quality and diverse new images starting from pure noise (Sordo et al., 2025). The framework for the model is shown in Figure 4. Figure 4. A Framework for Diffusion Models (Sordo et al., 2025) 302   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . The working principle of diffusion models is fundamentally based on two opposing processes: The forward diffusion process and the reverse diffusion process. The forward process involves taking an original, clean image (x₀) and systematically degrading it by adding Gaussian noise (ε) at each step (t) over a time step T. The state of this degradation at any time ‘t’ can be directly calculated using the coefficient ₜ, which depends on the original image and the added noise, and at the end of the process, the image turns into pure noise (xₜ ≈ N(0, I)). The inverse process, where generation occurs, aims to do the exact opposite. Starting from pure noise (xₜ), the model uses a neural network with parameters θ to learn to predict and remove the noise (εθ(xₜ, t)) added to the image at each time step (t). The neural network looks at xₜ and estimates the probability distribution (mean μθ and variance Σθ) of the previous, cleaner state xₜ₋₁. When this noise cleaning process is repeated over T steps, the model successfully generates a new, synthetic image that resembles the original data distribution (pθ(x₀)), starting from the initial meaningless noise. 4. GenAI Visualization According to Functional Differences 4.1. Text-to-Image Models  Variational AutoEncoders (VAEs) and Autoregressive Models (Transformers): Generative models that learn the meaningful and latent representation of data are exemplified by Variational Autoencoders (VAEs), which operate by first compressing high-dimensional inputs into a structured, lower-dimensional space and then sampling from this space to synthesize new data instances. These models, which consist of two components—an encoder and a decoder—convert the high-dimensional data they receive as input into a low-dimensional latent vector, and then attempt to reproduce the original data using this vector sampled from the latent space (Kingma & Welling, 2022). Autoregressive models are AI generation models that create data by generating one element at a time. When generating the next element, they use all previously generated elements as a condition. These models are adept at capturing longrange dependencies in sequences, owing this ability to transformer architectures, which are based entirely on the self-attention mechanism, without requiring recurrent or convolutional layers (Vaswani et al., 2023). Specific models using these model structures can be listed as follows: DALL·E 1: DALL·E is a groundbreaking artificial intelligence (AI) model that can generate images from text descriptions. The model’s name is GENERATIVE AI IN MEDIA, COMMUNICATION AND VISUAL ARTS   303 inspired by Pixar’s animated robot character WALL-E and Spanish surrealist artist Salvador Dalí (GIGAZINE, 2021). Developed by OpenAI, DALL·E utilizes a state-of-the-art deep learning model to generate high-quality, detailed images suitable for various applications, including product design and advertising (Zhou & Nabus, 2023). It can thus generate visual words at low resolution and with high creativity. DALL·E has exciting potential due to the new possibilities it offers for creativity and artistic expression. It is also one of the first models to introduce the ability to generate images from textual descriptions, demonstrating creativity by combining concepts that are difficult to relate to each other in reasonable ways. Technical limitations, such as DALL-E’s 12 billion parameter architecture and intensive training process, necessitated the development of a more efficient approach. Developed in this vein, DALL·E 2 has succeeded in generating higher-resolution and more realistic images through a more efficient training process, thanks to the integration of the CLIP model (Subramanian, 2025), which establishes a stronger semantic bridge between text and visuals. Cogview: It is a powerful generative model that combines computer vision with natural language processing techniques to generate images from text prompts. The model’s core idea is to provide large-scale generative pretraining capabilities for image tokens (such as VQ-VAE) (Hu et al., 2023). It is a model that offers greater fine-tuning capability than DALL·E in the sub-tasks of style learning, resolution, and text-image alignment (Ding et al., 2021). dVAE: Dynamic variational autoencoders (dVAE) combine standard variational encoders with a temporal model, thereby enabling unsupervised representation learning for sequential data (Ramesh et al., 2021). These models can convert continuous data consisting of pixels into discrete text sequences and generate images compatible with transformer model architectures (Oord et al., 2018). VQGAN+CLIP: This model, which combines two separate machine learning algorithms based on text commands, involves VQGAN learning a codebook vector composed of contextually rich visual fragments, and the composition of these codebook vectors is then modeled using an autoregressive transformer. CLIP, on the other hand, is another neural network that can determine how well a caption (or prompt) matches an image (Russell, 2022). Parti (Pathways Autoregressive Text-to-Image): The Pathways Autoregressive Text-to-Image (Parti) model, developed by Yu et al. (2022), generates high-quality photorealistic images and offers the ability to synthesize rich content with complex compositions and digital world knowledge. Parti 304   ARTIFICIAL INTELLIGENCE FUTURE FRONTIERS OF GENERATIVE AI IN . . . treats text-to-image conversion as a type of sequence-to-sequence modeling problem; therefore, image token sequences replace the target output text tokens. VQSEG: VQSEG (Vector Quantized for Segmentation) utilizes a structure called a semantic segmentation map during image generation. This structure enables the image to be synthesized in a semantic infrastructure while also providing guidance on what to generate, thereby making unconditional image generation controllable (Alaniz et al., 2022). Thanks to this control, thematic style and content can be effectively separated to enhance performance while preserving the semantic consistency of the image.  Generative Adversarial Networks (GANs) Based Models GANs (Generative Adversarial Networks) have seen increasing use in image modeling over the past decade. Current text-to-image generation techniques struggle to capture the subtle details and dynamic object components that are discernible in the real world (Arya et al., 2024). GANs, which can be trained dynamically and are sensitive to various factors, including optimization parameters and model architecture, can generate theoretical and empirical insights when they trained in a stable manner (Brock et al., 2019). These models have also transformed image generation by improving the discrimination performance between authentic and generated images (Gulsoy et al., 2024). These models rely on the use of two opposing models (namely, generator and discriminator) (Celard et al., 2024). StackGAN: This model simulates human creativity processes in image generation and occurs in two stages (Arya et al., 2024). In the first stage, GenAI creates a specific shape of the object based on the text. In the second stage, the quality of the existing output is enhanced, and more realistic details are incorporated into the image. The method also demonstrates high performance in identifying faces, facial features, facial identity, facial expressions, and facial emotions in the image generation process (W. Chen et al., 2025). Furthermore, it can prevent face occlusion issues (Jabbar et al., 2022). CycleGAN+BERT: This model combines CycleGAN, which excels at transforming paired generated text sets, with BERT, which deeply understands the semantic structure of text. To ensure model consistency and stabilize data training, it performs transfers from source to target and from target to source (Tsue et al., 2020). Thus, an image-to-text conversion component is added to verify that the generated image is semantically consistent with the input caption text (Zhu et al., 2020). GENERATIVE AI IN MEDIA, COMMUNICATION AND VISUAL ARTS   311 Misleading Journalism and Media: The proliferation of AI-generated images in digital forums has sparked debates in the media regarding legal and ethical issues, such as misinformation risks and copyright protocols for the data underlying generative models (Feher, 2025; Zlateva et al., 2024). Thomson et al. (2024) summarize the potential risks in the media as audience manipulation, objectivity issues, claims of reliable reality, visual mis/disinformation, fake curation, non-policiness, GenAI automation, and ethical violations. Deepfake Technology and Manipulation: Deepfake is fake audio and visual content created with the help of modern AI technologies (Westerlund, 2019). The use of this content causes several common problems and threats, in addition to the specific problems detailed below (Chawla, 2019; Westerlund, 2019; Есимова & Шевякова, 2024): Fake image generation and credibility Misinformation and manipulation Threat to social security Threat to information security Privacy violations Reputation damage Fraud and extortion Unauthorized content Distortion of reality perception The root of the problems is the generation of fake images that cause harm to individuals or institutions, along with the associated legal and social repercussions. 6. What Future Brings? Core Insights from AI GenAI for visualization is expected to gain importance across all sectors besides academia. Table 2 presents projections for the future of GenAI Visual Content across six different dimensions. [Document text truncated for crawler view.]