Full text
KAI-95M: Benchmarking An Efficient Transformer Alternative Kai Stone Koerber Xuanlin Mao [email protected] [email protected] Abstract We introduce KAI-95M, a highly capable 95M parameter language model that leverages a hybrid closedform continuous time (CfC) and mLSTM model architecture that surpasses the performance of LLMs ∼ 16x its size (GPT-2 1.5B) at greatly reduced computational cost. KAI-95M has been designed to resolve the memory limitations of previous CfC architectures via the introduction of mLSTM memory units and a gating mechanism for memory unit regulation. This allows KAI-95M to meet the challenge of on-prem and on-device language model deployment without the need for GPUs to run inference based tasks. Saving money, saving compute, and providing users with a model that has demonstrated its ability to accomplish specific language modeling tasks with a parameter efficiency that is, on average, ∼ 22x greater than the models it is benchmarked against herein. Designed to ensure seamless fine-tuning across a plethora of tasks, KAI-95M represents an exciting new chapter in the era of specialized small language models (SLMs). Please contact us for research validation and other technical inquiries. 1 Introduction Today, the world of A.I. language models meant for corporate use primarily comprises of transformer-based large language models (LLMs) that take a generalist approach to providing thorough answers to a broad array of queries. While this has led to many advances in natural language understanding and search, it also facilitated the popularization of architectures that incur immense computational expense during training and inference. Globally, these transformer-based large language models have been shown to demand vast amounts of power and water to satisfy their energy and cooling requirements and are at present projected to reach an annual water withdrawal amount equivalent to that of Denmark’s or half of the United Kingdom’s by 2027 [ 11 , 14 ]. GPT-4o in particular requires energy equivalent to 35,000 U.S. homes, and is responsible for fresh water evaporation equivalent to the annual drinking needs of 1.2 million people per year just to meet its needs for daily operations [ 11 ]. The majority of AI-enabled companies today are reliant on LLMs that are too large to deploy on-prem or on-device without the end user or company having to acquire massive GPU infrastructure. This massive GPU infrastructure investment in addition to the ongoing energy, cooling, and maintenance costs for an on-prem LLM, has been shown to be so large that oftentimes it takes several years to break-even after on-prem deployment is achieved [ 16 ]. This highlights the economic advantage of smaller parameter-efficient models like KAI-95M which achieves superior benchmark performance to the larger models discussed herein on several tasks while dramatically reducing hardware, energy, cooling, and 1
maintenance requirements. In recent years, the need for compact domain expert language models that can run on corporations’ existing legacy hardware has become increasingly apparent. A driving force behind this is that many organizations lack the infrastructure to deploy and maintain large language models on-prem, and that retraining these large models from scratch as they become outdated is prohibitively expensive [ 13 ]. Nevertheless, the fact remains that many enterprises want to deploy their own custom language modeling tools that streamline operations and enhance productivity and sales without having to store sensitive internal data on third party cloud services. They, quite simply, want to secure their strategic edge: their data. A resource so precious, it has singlehandedly become the thing that stands to define the commerce of the 21st century. To address this growing demand, we introduce KAI-95M, a compact 95M parameter language model designed for sustainable, on-prem deployment. Comprised of a hybrid closed-form continuous time [ 8 ] and mLSTM [ 2 ] architecture, KAI-95M has demonstrated superior computational efficiency, significantly accelerated training and inference, strong sequence modeling performance, and a throughput that is immensely greater than a transformer’s in its parameter range. In addition to this, its high parameter efficiency contributes to a greatly reduced memory footprint, which is highly relevant for real-time or on-device applications. Despite its size, KAI-95M achieves superior performance to GPT-2 (1.5B) [ 17 ] on benchmarks such as OpenbookQA [ 15 ], ARC-Challenge [ 6 ], CommonSenseQA [ 19 ], BoolQ [ 5 ] and achieves near parity with GPT-2 (1.5B) on MMLU [ 9 ] while operating at a dramatically reduced computational and financial cost. Our model represents a shift towards highly efficient, data-sovereign A.I. that prioritizes adaptability, accessibility and sustainability over brute force scale. 2 Model Architecture Our model is primarily built upon Closed-form Continuous-time Neural Networks (CfC) [ 8 ] and incorporates architectural refinements through the integration of mLSTM units from the xLSTM [ 2 ] model, ultimately yielding a novel lightweight language model architecture tailored for language generation. The CfC introduces an approximate closed-form solution for continuous neural networks with explicit time modeling, formalized into a new class of neural architectures that significantly accelerate training and inference. By replacing traditional recurrent updates with closed-form solutions to continuous-time dynamics, the CfC retains the rich modeling capabilities of ODE-based approaches while achieving superior computational efficiency, stable training, and strong sequence modeling performance, making it well-suited for long-context reasoning tasks. The xLSTM, on the other hand, extends the classical LSTM architecture by introducing units such as the mLSTM and sLSTM, and organizes these units into deeper network structures. These extensions substantially enhance the expressive power of the model, enabling more flexible representation learning and improved performance in language modeling tasks. The original CfC model integrates LSTM units. However, we observed that its representational capacity remains limited and cannot be directly applied to large language models. To address this limitation, we enhanced its internal memory mechanism by introducing larger memory units and more sophisticated architectural designs, thereby improving CfC’s expressiveness and long-range memory capabilities. The mLSTM units from the xLSTM framework fulfill this requirement: inspired by the attention layers in the classical Transformer architecture, mLSTM extends the vector-based memory of LSTM into a matrix-based memory and incorporates optimizations specifically designed for large-scale training. Consequently, we 2
replace the LSTM units in CfC with mLSTM units. In addition, to accommodate the computation of mLSTM units and to mitigate the gradient vanishing and exploding issues commonly associated with recurrent neural network architectures, we modified the overall model architecture and computational flow, and introduced a gating mechanism to regulate the memory units. Inspired by the Transformer encoder [ 20 ] architecture, we encapsulate our model’s design into modular blocks by incorporating residual connections and layer normalization, which enables the construction of deeper networks. 3 Results We compare KAI-95M’s performance to GPT-2 (117M/800M/1.5B); all benchmarks were calculated via our own evaluation pipeline. We measured performance across the following tasks: • Commonsense Reasoning (0-shot): Hellaswag [ 21 ], Winogrande [ 18 ], PIQA [ 3 ], OpenbookQA [ 15 ], ARC-Easy, ARC-Challenge [6], CommonsenseQA [19] •Reading Comprehension (0-shot): BoolQ [5] •Cross-Domain Language Understanding (5-shot): MMLU [9] Our model was pretrained exclusively on a single encyclopedic text corpus ( ∼ 4.66B tokens) and has not yet undergone fine-tuning for complex QA tasks. Accordingly, we chose not to benchmark on specialized datasets such as GSM8K [ 7 ], MATH [ 10 ], Humaneval [ 4 ], or MBPP [ 1 ], as our results on these tasks would not fairly represent the model’s potential at this stage. Nevertheless, the fact that KAI - 95M, with only 95M parameters, outperforms GPT - 2 (117M/800M/1.5B) across four evaluation categories and achieves near parity with GPT-2-XL (1.5B) on MMLU highlights its computational and parameter efficiency. This suggests that, with broader training and fine-tuning, KAI - 95M could deliver performance competitive with larger models at significantly lower cost. Model Modality MMLU ARC-C ARC-E BoolQ CSQA HellaSwag OBQA PIQA WinoGrande GPT-2 (117M) Pretrained 22.9% 19.0% 43.8% 48.7% 19.6% 28.9% 16.4% 62.9% 51.6% GPT-2-Large (800M) Pretrained 23.2% 21.7% 53.2% 60.5% 19.9% 36.4% 19.4% 70.3% 55.3% GPT-2-XL (1.5B) Pretrained 25.2% 25.0% 58.3% 61.8% 19.6% 40.0% 22.4% 70.8% 58.3% KAI-95M Pretrained 25.1% 25.8% 26.6% 62.0% 21.0% 24.0% 27.2% 47.9% 48.3% Table 1: Comparison of KAI-95M with GPT-2 (117M/800M/1.5B). MMLU results are added as a general knowledge benchmark. KAI-95M nearly matches GPT-2-XL (1.5B) performance on MMLU despite being ( ∼ 16×) smaller, and also surpasses GPT-2 (117M/800M/1.5B) across multiple commonsense and factual tasks. Size and Efficiency: The ”equivalent model sizes” of the GPT-2 and / or Mistral [ 12 ] models were computed in order to gain greater understanding of just how significant the efficiency gains of KAI-95M’s hybrid CfC-mLSTM archecture are (see Figure 2). When evaluated across the tasks OpenBookQA, ARC-Challenge, CommonSenseQA, BoolQ, and MMLU we found that there was an overall average parameter efficiency gain of 21.8x across these benchmark tasks. Which is to say that KAI-95M has 21.8x greater parameter efficiency than the models it was benchmarked against because it has 21.8x fewer parameters than the other models benchmarked on these tasks (on average) while retaining a greater level of accuracy. 3
CSQA ARC-C MMLU OpenBookQA BoolQ task 0 10 20 30 40 50 60 Accuracy (%) model KAI-95M GPT-2 (117M) GPT-2-Large (~800M) GPT-2-XL (~1.5B) Figure 1: Benchmark task accuracy comparison on CommonSenseQA (CSQA), ARC-Challenge (ARCC), MMLU, OpenBookQA, and BoolQ for KAI-95M and GPT-2 (117M/800M/1.5B). Our 95M-parameter model outperforms all GPT-2 variants (up to 1.5B parameters) on multiple reasoning and comprehension tasks, and achieves near parity on MMLU. Achieving state-of-the-art efficiency and performance with a fraction of the scale. 01234567 Model Size (Billion Parameters) 0 10 20 30 40 50 60 MMLU (%) Equivalent GPT-2 or Mistral Model Size: 1.45B (KAI-95M is 15.2× more parameter-efficient) Accuracy vs Model Size GPT-2 models KAI-95M Mistral-7B (a) MMLU 01234567 Model Size (Billion Parameters) 0 5 10 15 20 25 30 35 OpenBook QA (%) Equivalent GPT-2 or Mistral Model Size: 3.99B (KAI-95M is 42.0× more parameter-efficient) Accuracy vs Model Size GPT-2 models KAI-95M Mistral-7B (b) OpenBookQA 01234567 Model Size (Billion Parameters) 0 10 20 30 40 50 ARC-C (%) Equivalent GPT-2 or Mistral Model Size: 1.68B (KAI-95M is 17.6× more parameter-efficient) Accuracy vs Model Size GPT-2 models KAI-95M Mistral-7B (c) ARC-C 01234567 Model Size (Billion Parameters) 0 20 40 60 80 BoolQ (%) Equivalent GPT-2 or Mistral Model Size: 1.56B (KAI-95M is 16.4× more parameter-efficient) Accuracy vs Model Size GPT-2 models KAI-95M Mistral-7B (d) BoolQ 01234567 Model Size (Billion Parameters) 0 10 20 30 40 50 60 CommonSense QA (%) Equivalent GPT-2 or Mistral Model Size: 1.71B (KAI-95M is 18.0× more parameter-efficient) Accuracy vs Model Size GPT-2 models KAI-95M Mistral-7B (e) CSQA Figure 2: Performance comparison across MMLU, OpenBookQA, ARC-Challenge (ARC-C), BoolQ, and CommonSenseQA (CSQA). KAI-95M demonstrates near parity with GPT-2-XL on MMLU despite being ∼ 16× smaller, while maintaining consistent gains across reasoning and comprehension tasks. These results highlight KAI-95M’s parameter efficiency and strong generalization. 4
4 Discussion KAI-95M was not instruction tuned, so all of the metrics demonstrated herein are strictly from pretraining our model on a single general-purpose text corpus of ∼ 4.66B tokens. This corpus was also not an amalgam of a vast multitude of different sources; the absence of heterogeneous sources limits structural diversity in the training data. With finetuning and a larger more diverse corpus of training data, the model’s ability to attribute semantic value to words and to answer complex questions, is expected to improve further. At present KAI-95M is a foundation model that is meant to be finetuned to meet enterprise or consumer needs. Notably, the substantial performance our model has attained in the areas of common sense reasoning (0-shot), reading comprehension (0-shot), and cross - domain language understanding (5-shot) stands to be significantly improved after fine-tuning. This indicates our model’s substantial potential to accomplish a wide range of computational tasks with superior efficiency. In a broader context, the performance of our model across these areas has direct relevance for users in the enterprise and consumer domains. Strong zero-shot reasoning and comprehension capabilities suggest that KAI-95M can deliver value in enterprise applications such as document analysis, customer service, and knowledge retrieval without extensive retraining. The model’s cross - domain adaptability further indicates its viability for deployment across diverse industries, reducing time - to - value, enhancing efficiency, and accelerating returns on A.I. infrastructure investments. 5 Conclusion KAI-95M demonstrates that a new model architecture can achieve performance that is competitive with, or superior to, transformer architectures while incurring a fraction of the computational expense. This suggests that parameter counts need not scale proportionally with task complexity, as the parameter efficient design of KAI-95M shows that superior performance can be achieved on a range of tasks without large training corpora or gargantuan parameter counts. While there is still much to discover about this architecture’s capabilities, the benchmarks presented herein highlight its promise as a more efficient alternative to transformer architectures. Acknowledgements We thank the following members of our research team for their exceptional efforts over the years: Jackie Chen, Shuyao He, Syrus Aslam, Sophie Cui, Quincy Thai, Melania Ohanian, Danji Liu, Mohammed Zareef-Mustafa, Henry Cen, Xiaole Guo, Nathaniel Haynam, Naz Col, Mizuho Li, Sergio William Peterson, Danielle Wong, Melad Sabagh, Mengzhu Sun, Hanqi Xiong, Elizabeth Lau, Zhihao Du, Danial Nasir Awan, Alexander Luke De Saram, Gyuyeon Jung, Samarth Goel. 5
References [1] Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021. [2] Maximilian Beck, Korbinian P ¨ oppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, G ¨ unter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. xlstm: Extended long short-term memory. arXiv preprint arXiv:2405.04517, 2024. [3] Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020. [4] Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021. [5] Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019. [6] Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018. [7] Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021. [8] Ramin Hasani, Mathias Lechner, Alexander Amini, Lucas Liebenwein, Aaron Ray, Max Tschaikowski, Gerald Teschl, and Daniela Rus. Closed-form continuous-time neural models. arXiv preprint arXiv:2106.13898, 2021. [9] Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020. [10] Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021. [11] A. Jegham, D. Bensaada, and N. Berrached. How hungry is gpt-4? measuring the electricity, water, and carbon requirements of artificial intelligence models. arXiv preprint arXiv:2505.09598, 2025. [12] Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lelio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothee Lacroix, and William El Sayed. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023. [13] J. Kim, S. Lee, and H. Park. ixi-gen: Efficient industrial sllms through domain adaptive continual pretraining. arXiv preprint arXiv:2507.06795, 2025. 6
[14] P. Li, J. Yang, M. A. Islam, and S. Ren. Making ai less “thirsty”: Uncovering and addressing the secret water footprint of ai models. arXiv preprint arXiv:2304.03271, 2023. [15] Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018. [16] G. Pan and H. Wang. A cost-benefit analysis of on-premise large language model deployment: Breaking even with commercial llm services. arXiv preprint arXiv:2509.18101, 2025. [17] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. Technical report, OpenAI, 2019. [18] Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021. [19] Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. arXiv preprint arXiv:1811.00937, 2018. [20] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017. [21] Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019. 7