Available online www.ejaet.com European Journal of Advances in Engineering and Technology, 2025, 12(9):37-52 Research Article ISSN: 2394 - 658X 37 From Attention to Generative AI: A Decade of Architectural Innovations in Large Language Models Tharakesavulu Vangalapat AI Leader — Sr. Principle Data Scientist Broadridge, Austin, Texas, USA Email:
[email protected] _____________________________________________________________________________________________ ABSTRACT Over the past decade, artificial intelligence has undergone a remarkable transformation, particularly in natural language processing (NLP). The field has progressed from recurrent and convolutional models with limited sequence capacity to the attention based Transformer architecture that revolutionized scalability and context modeling. This breakthrough enabled the emergence of foundation models such as BERT, GPT, and T5, which redefined pretraining and transfer learning paradigms. Building on these foundations, the scaling laws formalized in 2020 demonstrated predictable performance gains with larger models, leading to GPT-3 and the era of fewshot and zero shot learning. The mainstream adoption of Generative AI (GenAI) followed in 2022 with ChatGPT and has since expanded to multimodal and instruction-tuned systems such as GPT-4, GPT-4V, Anthropic’s Claude family, and Google DeepMind’s Gemini series. In parallel, the open-source ecosystem has accelerated innovation through Meta’s LLaMA models, Falcon, Mistral, Mixtral, BLOOM, HuggingFace-led initiatives, Stability AI’s StableLM, and recent entrants such as DeepSeek. This survey provides a comprehensive comparative analysis of these developments across model architecture, parameter growth, context length, benchmarks, efficiency methods, and alignment strategies. More than fifty influential works are consolidated, with tables, charts, and metrics illustrating key milestones. I also discuss open challenges including energy efficiency, interpretability, bias, governance, and responsible open-source deployment. By consolidating both proprietary and community-driven contributions, this paper highlights the opportunities and risks that will shape the next generation of AI research and applications. Keywords: Transformers, Attention, Large Language Models, Generative AI, NLP, Scaling Laws, GPT, ChatGPT, GPT-4, Claude, Gemini, LLaMA, Falcon, Mistral, Mixtral, BLOOM, DeepSeek, StableLM, HuggingFace, Multimodality, Instruction Tuning, Reinforcement Learning with Human Feedback (RLHF), Alignment, Efficiency, Interpretability, Responsible AI _____________________________________________________________________________________________ INTRODUCTION Artificial Intelligence (AI) has experienced multiple waves of innovation, but the past decade (2013–2025) has been uniquely transformative in terms of scale, adoption, and architectural evolution. In particular, Natural Language Processing (NLP) shifted from recurrent and convolutional architectures that struggled with long-range dependencies to attention-based Transformers that now underpin modern Generative AI (GenAI). The rise of large language models (LLMs), exemplified by the GPT series [1,2], BERT [3], and LLaMA [4], has reshaped research and industry, catalyzing a new era of intelligent systems. Before 2017, sequence models relied heavily on Recurrent Neural Networks (RNNs) [5], Long Short-Term Memory (LSTM) networks [6], and Gated Recurrent Units (GRUs) [7]. These architectures achieved success in machine translation [8,9] and early dialogue systems but suffered sequential bottlenecks, vanishing gradients, and limited context windows. Even with attention integrated into Seq2Seq [8], scaling to billions of parameters on webscale corpora was infeasible. The Transformer of Vaswani et al. [10] replaced recurrence with multi-head self-attention and positional encoding, enabling parallel training and improved long-range modeling. This set the stage for pretraining and transfer learning: BERT [3], GPT-2 [11], XLNet [12], and T5 [13] showed that scaling data and parameters improves generalization. Scaling laws [14] formalized this trend, relating loss to compute, data, and model size.
Vangalapat T Euro. J. Adv. Engg. Tech., 2025, 12(9):37-52 38 GPT-3 [2] (175B parameters) marked an inflection point with few-shot and zero-shot capabilities. Infrastructure such as Megatron-LM [15] and DeepSpeed [16] enabled distributed training at unprecedented scales. Public adoption accelerated with instruction tuning and RLHF [17], culminating in ChatGPT (late 2022). Multimodal systems (CLIP [18], Flamingo [19], GPT-4V, and Gemini) unify text, vision, and reasoning. Despite success, open problems persist: environmental cost [20], hallucination and bias [21], and the limits of pure scaling. Hybrid approaches such as retrieval-augmented generation (RAG) [22], memory, tool use, and planning are active frontiers. Contributions. (i) A comprehensive survey from RNNs to Transformers and GenAI; (ii) comparative analysis of architectural improvements (context length, efficiency, alignment, benchmarks); (iii) open challenges and future research directions. Update (Aug 2025). This survey incorporates the most recent frontier models: Google’s Gemini 2.5 family (Pro/Flash/Flash-Lite), Anthropic’s Claude Opus 4.1, OpenAI’s GPT-5 and open-weight gpt-oss models, and Meta’s LLaMA-4 family (Scout/Maverick). These additions extend our taxonomy to 2025 and refresh comparisons on long-context, multimodality, agentic behaviors, and open-weight availability, positioning this work as an up-to-date reference for ongoing research and deployments. BACKGROUND: PRE-TRANSFORMER ERA (2010–2016) Representations and Early Neural NLP Prior to deep contextual models, NLP relied on sparse or count-based features (n-grams, TF–IDF) and linear models (SVMs, logistic regression). Dense word embeddings—Word2Vec (CBOW/Skip-gram) and GloVe—provided distributed semantic representations that improved generalization across tasks by capturing distributional similarity in vector spaces. These representations were task agnostic and formed the base for neural architectures. Recurrent Networks: RNNs, LSTMs, and GRUs RNNs [5] process sequences by recurrently updating a hidden state. However, classic RNNs suffer from vanishing/exploding gradients, impeding long-range credit assignment. LSTMs [6] mitigate this via gated memory (input, forget, output gates), enabling retention of longer dependencies. GRUs [7] simplify gating to update and reset gates, reducing parameters while retaining performance. These models achieved breakthroughs in speech recognition, language modeling, and early machine translation. Sequence-to-Sequence Learning The encoder–decoder (Seq2Seq) framework [9] maps input sequences to fixedlength vectors, which decoders transform to outputs (e.g., source-to-target sentences). While effective, compressing an entire source into a single vector degraded with longer inputs. Additive Attention in Neural MT Bahdanau et al. [8] introduced attention over encoder states, letting the decoder learn a soft alignment distribution per target token. Attention alleviated the fixed-bottleneck by conditioning generation on a weighted sum of source states. This improved translation quality and foreshadowed self-attention. Luong attention variants (dot/general) further refined efficiency. CNNs for NLP Convolutional architectures offered parallelism and locality. Kim’s CNN [23] used multiple filter widths and maxpooling for sentence classification. Kalchbrenner et al. proposed dynamic CNNs with k-max pooling for variablelength inputs. CNNs captured local n-gram features with fewer sequential constraints but struggled with long-range dependencies without stacking or dilations. Training Dynamics and Regularization Pre-Transformer models relied on truncated backpropagation through time, gradient clipping, dropout, and schedule sampling to stabilize training. Optimization with SGD/momentum gave way to Adam, which improved convergence for deep RNN/CNN stacks. Despite these, sequential computation limited hardware utilization: tokens had to be processed in order, constraining throughput and batch parallelism. Bottlenecks That Motivated Transformers Three interlocking bottlenecks drove the architectural shift: • Sequential dependency: RNNs could not fully exploit SIMD/GPU parallelism due to time-step recurrence. • Long-context modeling: Even with attention on encoder states, gradient flow over hundreds of tokens remained fragile, and fixed bottlenecks persisted. • Scaling limits: Billion-parameter RNNs on web-scale corpora were impractical; memory and compute costs rose superlinearly with sequence length under recurrence. These limitations motivated architectures where pairwise token interactions are computed in parallel with controllable inductive bias: self-attention with positional encodings. Summary and Transition By 2016, the field had strong ingredients—word embeddings, Seq2Seq, and local/global attention—but lacked a unifying, massively parallelizable sequence learner. The Transformer [10] synthesized these ideas into a fully attentionbased model, removing recurrence, enabling deeper scaling, and setting the stage for pretrained language models and GenAI.
Vangalapat T Euro. J. Adv. Engg. Tech., 2025, 12(9):37-52 39 THE TRANSFORMER BREAKTHROUGH (2017–2019) Attention Is All You Need The introduction of the Transformer architecture by Vaswani et al. in 2017 [10] represented a decisive departure from recurrence and convolution. Instead of processing sequences step by step, Transformers rely entirely on selfattention to compute pairwise interactions between tokens in parallel. This enabled unprecedented scalability and fundamentally altered how context is modeled in sequence data. The Transformer consists of an encoder and decoder, each composed of stacked layers that contain: • Multi-Head Self-Attention: Projects queries, keys, and values into multiple subspaces, allowing the model to jointly attend to information at different representation subspaces and positions. • Position-wise Feedforward Layers: Fully connected layers applied identically to each position, enabling nonlinear transformations. • Residual Connections and Layer Normalization: These stabilize training and enable deeper networks. • Positional Encodings: Since attention lacks recurrence or convolution, sinusoidal position encodings are added to inject order information. The key operation is the scaled dot-product attention: Attention(Q,K,V ) = softmax (1) where Q, K, and V are matrices of queries, keys, and values, and dk is the dimension of the keys. Advantages Over RNNs and CNNs The Transformer solved three major limitations of earlier architectures: 1. Parallelization: Self-attention allows all tokens in a sequence to be processed simultaneously, exploiting GPU/TPU hardware more effectively. 2. Long-Range Dependencies: Every token can directly attend to every other token, avoiding vanishing gradients and bottlenecks. 3. Scalability: Model depth and width can be increased more easily without sequential bottlenecks, enabling scaling to hundreds of layers and billions of parameters. These advantages quickly made the Transformer the default backbone for NLP tasks. Early Impact: Machine Translation and Beyond The original Transformer achieved state-of-the-art BLEU scores in machine translation benchmarks such as WMT’14 English–German and English–French [10], outperforming strong RNN-based baselines while requiring significantly less training time. The architecture was rapidly adopted across tasks: • Language Modeling: GPT-1 (2018) [1] showed that a decoder-only Transformer could be pretrained on large corpora and fine-tuned for downstream NLP tasks. • Contextual Embeddings: BERT (2018) [3] introduced masked language modeling (MLM) and next-sentence prediction, yielding contextualized embeddings that became foundational for NLP pipelines. • Cross-Task Transfer: The GLUE benchmark [24] highlighted Transformers’ ability to generalize across sentiment analysis, entailment, and semantic similarity tasks. Variants and Refinements Between 2018 and 2019, numerous Transformer variants were proposed: • GPT-1 & GPT-2: Demonstrated the power of autoregressive pretraining on large unlabeled corpora. • BERT and RoBERTa: RoBERTa [25] showed that with larger batch sizes, more data, and removal of nextsentence prediction, masked LM pretraining achieves even stronger performance. • XLNet: Generalized autoregressive training to capture bidirectional context while retaining autoregressive advantages [12]. • Transformer-XL: Introduced segment-level recurrence and relative positional encodings to extend context length [26]. These refinements signaled the beginning of the pretraining era, where universal language representations could be reused across tasks. Visualization of the Architecture Figure 1: The Transformer architecture as introduced by Vaswani et al. (2017) [10]. Modern LLMs simplify or adapt components depending on task (e.g., encoder-only for BERT, decoder-only for GPT).
Vangalapat T Euro. J. Adv. Engg. Tech., 2025, 12(9):37-52 40 Figure 1 illustrates the encoder–decoder Transformer architecture, highlighting the self-attention mechanism, positional encoding, and residual connections. In practice, modern large language models (LLMs) often adopt only the encoder (BERT-style) or only the decoder (GPT-style), depending on objectives. Comparison of Early Models Table 1 compares early Transformer-based models, their objectives, parameter scales, and context windows. These paved the way for scaling laws and the era of massive foundation models. Table 1: Comparison of Early Transformer-Based Models (2017–2019) Model Year Params Objective Context Transformer (NMT) 2017 65M Seq2Seq 512 GPT-1 2018 117M LM (AR) 512 BERT (base) 2018 110M MLM+NSP 512 RoBERTa 2019 355M MLM 512 XLNet 2019 340M AR (perm.) 512 Transformer-XL 2019 257M AR 1K+ Summary The Transformer era (2017–2019) redefined NLP. By removing recurrence, introducing parallelizable selfattention, and scaling efficiently, Transformers created a foundation that subsequent innovations—BERT, GPT-2, XLNet, and RoBERTa—would build upon. This period marked the transition from task specific models to universal pretrained architectures, setting the stage for the explosive growth of foundation models and the scaling laws that followed. RISE OF PRETRAINED MODELS AND BERT/GPT (2018–2020) The Shift to Pretraining and Transfer Learning The introduction of the Transformer catalyzed a major paradigm shift: instead of training models from scratch for each task, researchers began to explore pretraining on large unlabeled corpora followed by fine-tuning on taskspecific datasets. This approach, inspired partly by earlier work on word embeddings such as Word2Vec and GloVe, generalized to contextual embeddings and full language models. By decoupling representation learning from task adaptation, pretraining enabled dramatic gains in efficiency and performance across a wide range of NLP tasks. BERT: Bidirectional Representations from Transformers BERT (Bidirectional Encoder Representations from Transformers) introduced by Devlin et al. in 2018 [3] demonstrated the power of large-scale masked language modeling (MLM). Unlike left-to-right autoregressive models, BERT used a bidirectional encoder that simultaneously conditions on both left and right context. Its two main pretraining objectives were: 1. Masked Language Modeling (MLM): Randomly masking tokens in the input and predicting them, forcing the model to capture bidirectional context. 2. Next Sentence Prediction (NSP): Classifying whether two input sentences were consecutive in the original corpus. BERT’s architecture came in two sizes: BERT BASE (110M parameters) and BERT LARGE (340M parameters). The model achieved state-of-the-art results on the GLUE benchmark [24], SQuAD question answering, and other downstream tasks, quickly becoming the standard backbone for NLP research and industry applications. RoBERTa and Refinements Follow-up studies revealed that BERT’s training procedure was under-optimized. RoBERTa (2019) [25] removed NSP, trained on larger corpora (160GB), used longer sequences, and scaled batch sizes. These modifications yielded significant performance improvements without architectural changes. Other refinements included ALBERT (2019) [27], which reduced parameter count through factorized embeddings and cross-layer parameter sharing, enabling more efficient training without sacrificing accuracy. GPT: Autoregressive Pretraining In parallel, OpenAI pursued decoder-only Transformers trained with a simple left-to-right language modeling objective. GPT-1 [1] showed that a model pretrained on BooksCorpus could be fine-tuned for a wide range of NLP tasks. GPT-2 [11], trained on 40GB of internet text, scaled to 1.5B parameters and demonstrated surprisingly coherent multi-paragraph text generation. GPT-2 also revealed the risks of open-ended generation, leading OpenAI to initially restrict full model release. GPT models differed fundamentally from BERT-style encoders: • Objective: GPT used autoregressive prediction, BERT used masked language modeling. • Architecture: GPT was decoder-only, BERT was encoder-only. • Applications: GPT excelled at generation, BERT excelled at classification and understanding tasks.
Vangalapat T Euro. J. Adv. Engg. Tech., 2025, 12(9):37-52 41 XLNet, Transformer-XL, and Unified Approaches To bridge the gap between autoregressive and bidirectional models, XLNet [12] introduced permutation-based autoregressive pretraining, capturing bidirectional context while avoiding the limitations of masking. TransformerXL [26] extended context windows by introducing recurrence across segments with relative positional encodings. These innovations highlighted that both autoregressive and masked objectives could be complementary, and hybrid approaches might be necessary. The GLUE Benchmark and Standardization The release of the General Language Understanding Evaluation (GLUE) benchmark [24] provided a standardized set of tasks including sentiment classification, entailment, paraphrase detection, and semantic similarity. Pretrained Transformers rapidly pushed GLUE scores upward: BERT LARGE surpassed human baselines on several tasks, RoBERTa and XLNet improved further, and ALBERT set efficiency records. GLUE and its successor, SuperGLUE [28], became critical in tracking progress and comparing models. Table 2: Performance of Transformer Models on GLUE Benchmark (dev set). Adapted from [24,28]. Modal Year Avg. GLUE Score (%) BERT BASE 2018 79.6 BERT LARGE 2018 82.1 RoBERTa LARGE 2019 88.5 XLNet LARGE 2019 89.8 ALBERT XXL 2019 89.4 Human Baseline 87.1 T5 and the Text-to-Text Paradigm Google’s T5 (Text-to-Text Transfer Transformer) [13] proposed a unified framework where all NLP tasks are cast as text-to-text transformations. For instance, sentiment classification could be reframed as “input text → positive/negative,” while translation remained “English → French.” This design choice simplified the architecture and training process, enabling multitask learning across diverse tasks. T5 also demonstrated the benefits of largescale pretraining on the Colossal Clean Crawled Corpus (C4). Impact on Industry and Applications The success of pretrained Transformers had an immediate impact: • Search Engines: Google incorporated BERT into search ranking, improving contextual understanding of queries. • Conversational AI: Virtual assistants such as Alexa and Siri integrated BERT-style models for intent classification and slot filling. • Healthcare and Finance: Pretrained LMs improved biomedical text mining (BioBERT [29]) and financial document classification. Summary The period from 2018 to 2020 was defined by the rise of pretrained Transformers. BERT and GPT introduced two complementary paradigms—bidirectional masked language modeling for understanding, and autoregressive pretraining for generation. Refinements such as RoBERTa, XLNet, and T5 pushed performance higher while establishing new benchmarks. This era solidified the notion of foundation models, pretrained once and adapted universally, setting the stage for the scaling laws and massive models that followed. SCALING LAWS AND THE GPT-3 ERA (2020–2022) The Discovery of Scaling Laws Kaplan et al. (2020) [14] empirically demonstrated that language model performance follows predictable power-law relationships with respect to model size (parameters), dataset size, and compute. Specifically, test loss decreases smoothly as these resources increase, until limited by the smallest factor (the so-called choke point). This insight provided a recipe for improving performance: rather than devising new objectives or architectures, one could scale existing Transformer models. Formally, the loss L can be approximated as: L(N,D,C) ≈ AN−α + BD−β + C−γ (2) where N is number of parameters, D is dataset size, C is compute, and α,β,γ are empirical exponents. This formulation implied that predictable gains could be achieved by scaling all three dimensions in balance. GPT-3: 175 Billion Parameters Building on scaling laws, OpenAI released GPT-3 [2], a 175B parameter autoregressive language model trained on 570GB of filtered internet text. GPT-3 marked a watershed moment in NLP: • Few-shot and Zero-shot Learning: GPT-3 demonstrated emergent abilities, solving tasks with minimal or no task-specific supervision. • Broad Capabilities: The model could translate, summarize, answer questions, and even generate code from natural language prompts.
Vangalapat T Euro. J. Adv. Engg. Tech., 2025, 12(9):37-52 42 • Human-like Interaction: Outputs were coherent across multiple paragraphs, making GPT-3 suitable for conversational agents and creative applications. GPT-3’s training required thousands of petaflop/s-days of compute, distributed over GPU/TPU clusters using parallelism strategies such as pipeline parallelism, tensor parallelism, and ZeRO optimizations. Infrastructure Innovations Scaling to 100B+ parameters necessitated advances in distributed systems: • Megatron-LM [15]: Introduced model parallelism via tensor slicing across GPUs, enabling training of 8B+ parameter models. • ZeRO (Zero Redundancy Optimizer) [30]: Reduced memory footprint by partitioning optimizer states, gradients, and parameters across devices. • DeepSpeed [16]: Provided mixed-precision training, optimizer sharding, and communication optimizations to train 100B+ models efficiently. • Pipeline Parallelism: Layers split across GPUs in a pipeline fashion, overlapping computation and communication. These system innovations made trillion-parameter model training feasible, extending scaling laws to new orders of magnitude. Sparse and Efficient Transformers While scaling improved performance, quadratic self-attention (O(n2) with respect to sequence length n) became a bottleneck. Sparse attention architectures addressed this: • Longformer [31] and BigBird [32] used sparse patterns (sliding windows + global tokens) to reduce complexity to O(n) or O(nlogn). • Reformer [33] applied locality-sensitive hashing to approximate attention. • Linformer [34] projected key-value matrices to lower-rank representations, reducing memory. These approaches extended Transformers to longer sequences while maintaining efficiency, crucial for scaling to massive corpora. Emergent Abilities As models scaled, they began to display emergent abilities—qualitative leaps in performance not present in smaller models [35]. Examples include: • Arithmetic reasoning (e.g., multi-digit addition). • Multi-step commonsense reasoning. Table 3: Comparison of Large Transformer Models (2020–2022). *Effective parameters; only a subset is active per forward pass. Model Year Params Training Data (approx.) Objective T5-11B [36] 2020 11B C4 (750GB) Text-to-text (span corruption) GPT-3 [37] 2020 175B 570GB filtered web LM (autoregressive) Switch Transformer [38] 2021 1.6T∗ 750GB Mixture-of-Experts LM GShard [39] 2021 600B∗ Multilingual MoE (translation) Jurassic-1 (AI21) 2021 178B 300B tokens LM (autoregressive) ERNIE 3.0 (Baidu) 2021 260B Multilingual web Hybrid (MLM+tasks) • In-context learning: adapting to new tasks purely from examples in the input prompt. These abilities emerged suddenly at certain parameter scales, reinforcing the validity of scaling laws while suggesting deeper connections to model capacity and representation learning. Ethical and Practical Concerns The GPT-3 era also foregrounded societal concerns: • Bias and Fairness: Large models reflected biases from their training corpora [21]. • Misinformation Risks: Coherent text generation raised fears of automated disinformation campaigns. • Environmental Impact: Training GPT-3 emitted hundreds of tons of CO2 [20]. These issues sparked calls for responsible scaling, transparency, and improved dataset curation. Comparisons of Large Models (2020–2022) Table 3 highlights the growth of LLMs during this period. GPT-3 stood out for scale, while contemporaries like T511B and GShard explored multitask and multilingual training. Visualization of Scaling Laws Figure 2 shows a placeholder for the scaling law plot, illustrating test loss decreasing with parameter size, data, and compute.
Vangalapat T Euro. J. Adv. Engg. Tech., 2025, 12(9):37-52 43 Figure 2: Scaling laws for language models [14]. Test loss follows predictable power-law trends as parameters, data, and compute increase. Summary The period 2020–2022 confirmed that scaling is the dominant driver of LLM performance. GPT-3 exemplified the power of scaling laws, showing emergent fewshot and zero-shot learning abilities. System innovations such as Megatron-LM, ZeRO, and DeepSpeed enabled trillion-parameter models, while sparse attention architectures addressed efficiency. At the same time, ethical concerns and environmental costs underscored the need for responsible development. This era set the stage for the Generative AI revolution, where instruction tuning, multimodality, and alignment transformed LLMs into widely deployed products. THE GENERATIVE AI REVOLUTION (2022–2025) ChatGPT and the Mainstreaming of LLMs The public release of ChatGPT in late 2022, built on OpenAI’s GPT-3.5 and later GPT-4, marked the turning point where large language models became globally recognized tools. ChatGPT reached over 100 million active users within two months, becoming the fastest-growing consumer application in history. This sudden mainstream adoption showcased the utility of conversational interfaces and democratized access to generative AI. The key breakthrough enabling ChatGPT was instruction tuning combined with reinforcement learning from human feedback (RLHF) [17]. Rather than producing raw text, models were fine-tuned to follow user instructions and align outputs with human preferences. This alignment step was essential to transforming foundation models into practical assistants. Alignment and Responsible AI Beyond RLHF, new alignment strategies emerged: • Anthropic’s Constitutional AI [40]: Used a set of guiding principles to reduce reliance on human labelers, improving safety and consistency. • Self-Alignment and Distillation: Techniques like Direct Preference Optimization (DPO) simplified alignment pipelines. These methods improved reliability but raised ongoing debates around transparency, bias, and over-optimization toward human-provided signals. Multimodal Generative AI The post-2022 era also saw the rise of multimodal systems: • CLIP [18]: Bridged vision and language through contrastive training. • Flamingo [19]: Enabled few-shot multimodal learning. • GPT-4V (2023): Added image understanding to GPT-4, supporting captioning, analysis, and multimodal reasoning. • Google DeepMind Gemini (2023): Integrated text, image, and code capabilities, extending context windows up to 128k tokens. • LLaVA (2023): An open-source vision-language assistant trained by fine-tuning LLaMA with CLIP features. These advances signal the convergence of modalities toward unified foundation models. Open-Source Community Models The open-source ecosystem exploded after 2022, spearheaded by Meta’s LLaMA family [4,41]. Although initially released under a research license, leaked checkpoints spurred a flourishing ecosystem of fine-tuned variants: • LLaMA-1 (2023): Compact 7B–65B parameter models trained on curated corpora, competitive with GPT-3 despite smaller scale. • LLaMA-2 (2023): Released with 7B, 13B, and 70B models, optimized for chat, widely adopted in open community.
Vangalapat T Euro. J. Adv. Engg. Tech., 2025, 12(9):37-52 44 • LLaMA-3 (2024): Advanced scaling and alignment, considered competitive with GPT-4-class models. Alongside LLaMA, numerous initiatives proliferated: • Falcon (2023): Developed by TII, trained on Refined Web dataset, topped HuggingFace LLM leaderboards. • Mistral (2023): Small but highly efficient 7B model, followed by Mixtral 8x7B, a Mixture-of-Experts achieving GPT-3.5-level performance. • BLOOM (2022): A multilingual 176B model trained collaboratively by BigScience with HuggingFace support. • StableLM (2023): Stability AI’s open release focused on democratizing LLM research. Competitive Frontier Models By 2024–2025, multiple companies entered the LLM race: • DeepSeek (2024): A Chinese open competitor emphasizing efficiency, multilingual training, and public model access. • xAI Grok (2023): Released by Elon Musk’s xAI, integrated with Twitter/X. • ERNIE 4.0 (2023): Baidu’s foundation model, optimized for Chinese and multilingual tasks. These models highlight that the competitive frontier is no longer limited to a few US-based labs. Ecosystem and Tools HuggingFace played a central role as the de facto community hub for model hosting, evaluation (Open LLM Leaderboard), and deployment pipelines (Transformers library). Frameworks such as LangChain and LlamaIndex enabled LLMs to act as agents, connecting to retrieval systems, tools, and APIs. Retrieval Augmented Generation (RAG) [22] became a standard method for grounding LLM outputs in external knowledge. Comparison of Generative AI Models (2022–2025) Table 4 compares representative models across closed and open communities, highlighting parameters, modality, and availability. Table 4: Comparison of Generative AI Models (2022–2025) Model Year Params Modality Context License Notes ChatGPT (GPT-3.5) 2022 175B Text 4K Proprietary First mainstream LLM GPT-4 2023 N/A Text/Img 32K+ Proprietary Advanced reasoning LLaMA-2 2023 7B–70B Text 4K Open (Meta) Widely adopted OSS Falcon 2023 40B–180B Text 4K Open (TII) HF leaderboard leader Mistral 7B 2023 7B Text 8K Open Efficient dense model Mixtral 8x7B 2023 46.7B Text 32K Open MoE rival to GPT-3.5 BLOOM 2022 176B Text (46L) 2K Open Multilingual model DeepSeek 2024 N/A Text/Multi 32K+ Open/Res. Chinese frontier model Gemini 1.5 2024 N/A Multi 128K Prop. (Google) Multimodal, long-ctx Claude 3 2024 N/A Text 200K+ Prop. (Anthro.) Long-context reasoning Summary The 2022–2025 period transformed LLMs from foundation models into generative assistants and multimodal systems. ChatGPT demonstrated global consumer adoption, while open-source models like LLaMA, Falcon, Mistral, and DeepSeek empowered a community-driven ecosystem. HuggingFace became the hub for open research, while proprietary frontier labs pushed boundaries in multimodality and reasoning. This era firmly established Generative AI as both a commercial product and a global research frontier. COMPARATIVE ARCHITECTURAL IMPROVEMENTS Dimensions of Progress The evolution from RNNs to Transformers and then to Generative AI can be summarized along several key axes: • Parameter Growth: From millions (RNNs, CNNs) to hundreds of billions (GPT-4, LLaMA-3), with some mixture-of-experts models reaching trillion-level effective parameters. • Training Data: From curated corpora such as BooksCorpus and Wikipedia to filtered internet-scale datasets exceeding trillions of tokens. • Context Window Expansion: From under 100 tokens in RNNs to 128k+ tokens in Gemini and Anthropic’s Claude models. • Training Objectives: Shift from supervised sequence labeling and Seq2Seq to masked language modeling (BERT), autoregressive LM (GPT), textto-text (T5), and multimodal objectives. • Efficiency and Hardware Utilization: Sparse attention, quantization, low-rank adaptation, and model-parallel frameworks like Megatron and DeepSpeed. • Alignment: From raw language modeling to instruction tuning, RLHF, Constitutional AI, and preference optimization.
Vangalapat T Euro. J. Adv. Engg. Tech., 2025, 12(9):37-52 45 Table 5: Growth in Parameters and Context Length Across Models Model Year Params Context Seq2Seq + Attention 2015 100M < 100 Transformer (NMT) 2017 65M 512 BERT-LARGE 2018 340M 512 GPT-2 2019 1.5B 1K T5-11B 2020 11B 1K GPT-3 2020 175B 2K Jurassic-1 2021 178B 2K GPT-4 2023 Undisclosed 32K LLaMA-2-70B 2023 70B 4K Mixtral 8x7B 2023 46.7B (MoE) 32K Gemini 1.5 2024 Undisclosed 128K Claude 3 2024 Undisclosed 200K+ Growth in Model Size and Context Table 8 highlights the progression in parameter counts and context lengths across representative models. This trajectory demonstrates exponential scaling both in parameter count and sequence length, directly impacting capabilities such as reasoning, summarization, and long-document analysis. Benchmark Performance Pretrained and generative models have consistently raised the bar on standardized benchmarks such as GLUE, SuperGLUE, MMLU, BIG-bench, and HumanEval. Table 6 shows illustrative performance trends. Table 6: Performance Trends on Key Benchmarks (Illustrative, Adapted from Published Results) Model GLUE (%) S-GLUE (%) MMLU (%) BIG-bench HumanEval BERT-LARGE (2018) 82.1 71.5 – – – RoBERTa-LARGE (2019) 88.5 79.0 – – – GPT-3 (2020) – – 43.9 65.0 21 PaLM-540B (2022) – – 67.5 75.2 36 GPT-4 (2023) – – 86.4 85.5 88 Claude 3 (2024) – – 83.1 84.0 82 LLaMA-2-70B (2023) – – 68.9 72.0 33 Mixtral 8x7B (2023) – – 71.4 73.5 40 DeepSeek (2024) – – 74.2 77.3 46 Benchmarks illustrate three major shifts: 1. Transformer encoders (BERT, RoBERTa) dominated understanding benchmarks (GLUE, SuperGLUE). 2. GPT-style decoders enabled few-shot reasoning and coding tasks (MMLU,HumanEval). 3. Generative AI assistants (GPT-4, Claude, Gemini) surpassed human performance in multiple benchmarks, raising questions about the sufficiency of current evaluations. Efficiency and Adaptation Techniques Scaling to billions of parameters posed efficiency challenges. Several innovations emerged: • Parameter-Efficient Fine-Tuning (PEFT): Techniques like LoRA (LowRank Adaptation) and adapters reduced costs of customizing large models. • Quantization and Pruning: Lower precision (INT8, INT4) enabled deployment on consumer hardware. • Mixture-of-Experts (MoE): Models such as Switch Transformer [38] and Mixtral 8x7B activate only subsets of parameters per inference, balancing scale and efficiency. Alignment Progression Alignment strategies also improved across generations: • GPT-2: No alignment; raw autoregressive LM. • GPT-3: Few-shot prompting, but prone to toxic/unhelpful outputs. • InstructGPT (2022): Instruction tuning + RLHF. • Claude (2023): Constitutional AI for safety principles. • Open-source: Alpaca, Vicuna, and Zephyr applied instruction tuning to LLaMA backbones. Next-Gen Frontier (2024–2025+): Reasoning, Long Context, and Open-Weight Releases Gemini 2.5 (Pro/Flash/Flash-Lite). Google’s Gemini 2.5 introduces native “thinking” capabilities that strengthen code generation and complex multi-step reasoning, with broad multimodal support (text, images, audio/video) and productized availability across Google AI Studio and Vertex AI. The series differentiates for cost/performance (Pro, Flash, Flash-Lite), and introduces enhanced reasoning modes (e.g., “Deep Think”) and computer-use agents.
Vangalapat T Euro. J. Adv. Engg. Tech., 2025, 12(9):37-52 52 [44] Google DeepMind, “Gemini 2.5: Our most intelligent ai model,” Google Blog, Mar 25, 2025, 2025, model availability and reasoning updates. [Online]. Available: https://blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/ [45] Anthropic, “Claude opus 4.1,” Newsroom, Aug 5, 2025, 2025. [Online]. Available: https://www.anthropic.com/news/claude-opus-4-1 [46] OpenAI, “Gpt-5 is here,” Product page, 2025. [Online]. Available: https://openai.com/gpt-5/ [47] ——, “Introducing gpt-oss,” Announcement, Aug 5, 2025, 2025. [Online]. Available: https://openai.com/index/introducing-gpt-oss/ [48] E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots: Can language models be too big?” in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 2021. [49] L. Ouyang, J. Wu, X. Jiang et al., “Training language models to follow instructions with human feedback,” in NeurIPS, 2022. [50] Y. Bai et al., “Constitutional ai: Harmlessness from ai feedback,” arXiv preprint arXiv:2212.08073, 2022. [51] P. Lewis, E. Perez, A. Piktus et al., “Retrieval-augmented generation for knowledge-intensive nlp,” in NeurIPS, 2020. [52] BigScience Workshop, “Bloom: A 176b parameter open-access multilingual language model,” Technical Report, 2022. [53] Anthropic, “Claude opus 4 and 4.1 can now end a rare subset of conversations,” Anthropic research/policy note, 2025. [Online]. Available: https://www.anthropic.com/research/end-subset-conversations