Modeling Energy Consumption in Deep Learning Architectures Using Power Laws Mansour Zoubeirou a Mayakia,*and Victor Charpenaya aMines Saint-Etienne, Université Clermont Auvergne, Clermont Auvergne INP, CNRS, LIMOS Abstract. Modern Deep Learning architectures such as LSTM, GRU, and Transformers achieve remarkable performance in various sequence processing tasks. Yet, their high computational cost and energy consumption have raised concerns about their environmental impact and the sustainability of Deep Learning. In this paper, we present an empirical study assessing the efficiency of training LSTM, GRU, and Transformer models on a GPU. By evaluating these models under various configurations, we characterize the relationship between energy consumption and pre-defined quantities such as hardware efficiency and the number of floating point operations (FLOPs) required for inference. We show that it is possible to derive power laws that make energy consumption predictable, given an architecture and a GPU model. 1 Introduction Recent studies have expressed concerns about the sustainability of Deep Learning (DL) research [26, 13, 27]. The increasing use of larger datasets, more complex models, and prolonged training periods has led to a significant rise in energy usage for training. For instance, training GPT-3 required 3.14 ×1023 floating-point operations (FLOPs) [1], which amounts to at least 20 years of computation on a single H100 GPU platform (one of the fastest Nvidia GPUs). Training DeepSeek-v3, allegedly more cost-efficient than GPT-3 and its successors, was equivalent to 318 years of uninterrupted computation on that same GPU [5]. In reaction, critical questions have been raised about their economic and ecological implications. As DL systems continue to grow in scale, understanding and mitigating their energy demands has become a pressing issue. Yet, comprehensive studies comparing energy consumption across different model architectures, training configurations, and hyperparameter choices remain limited. This paper bridges this gap by conducting a detailed empirical investigation into the energy efficiency of three widely used model architectures: Long Short-Term Memory units (LSTMs), Gated Recurrent Units (GRUs), and Transformers. These architectures are fundamentally different in their computational design. LSTMs and GRUs are recurrent neural networks (RNNs) optimized for processing sequential data through memory mechanisms, while Transformers employ self-attention mechanisms, enabling superior performance on tasks involving long-range dependencies. Despite these differences, all three architectures are widely used in applications ranging from Natural Language Processing (NLP) to time series forecasting, mak- ∗Corresponding Author. Email:
[email protected] ing them ideal candidates for a comparative energy consumption study. Our study focuses on measuring the energy consumption of these architectures during training under varying model configurations, such as the number of layers, hidden dimensions, and attention heads (for Transformers). Beyond raw energy measurements, we also explore the relationship between energy usage, FLOPs count and hardware efficiency, which quantifies actual throughput relative to maximum throughput. Our goal is to show that it is possible to derive a general law that makes energy consumption predictable, given a model architecture. The main contributions of our work can be summarized as follows: •We demonstrate that energy consumption can be estimated using laws that depend on the computational cost of a model, expressed in FLOPs, and a hardware efficiency measure. We empirically found that modeling hardware efficiency was critical to properly model energy consumption of a given model. •We show that hardware efficiency for a given elementary operation, such as matrix multiplication, can be succinctly captured using a power law featuring exponential decay with a saturation offset. •We release a dataset on the real-world energy consumption of Transformer-based models (BERT and GPT-2), trained on two GPU architectures: NVIDIA A100 (80GB) and GeForce RTX 2080 Ti. Despite the scarcity of such data in the field, our dataset includes over 98,000 runs across 24,000 unique configurations, varying in layers , embedding size, attention heads, and sequence lengths. Each run includes energy usage, training time, and other metrics, collected using CodeCarbon. This resource supports benchmarking and research on energy-efficient machine learning. The full data, along with an appendix to the paper, are available online [29]. 2 Related Work This section reviews prior work in three key areas: energy consumption of DL models, computational efficiency, and optimization strategies for energy-efficient training [26, 7, 3, 12, 11, 27, 16]. 2.1 Energy Consumption of Deep Learning Models The growing size and complexity of DL models have led to a dramatic rise in energy consumption, with several studies highlighting the environmental implications of large-scale model training. For instance, Strubell et al. [21] quantified the carbon footprint of designing and training NLP pipelines, showing that such pipelines can emit ECAI 2025 I. Lynce et al. (Eds.) © 2025 The Authors. This article is published online with Open Access by IOS Press and distributed under the terms of the Creative Commons Attribution Non-Commercial License 4.0 (CC BY-NC 4.0). doi:10.3233/FAIA250900 936
as much greenhouse gases as an American household over a year. The LLMCarbon framework by Faiz et al. [6] expand on this by modeling the total carbon footprint of large language models, incorporating both operational and embodied carbon emissions. It provides a predictive tool to estimate emissions across training, inference, and storage, highlighting the importance of architectural optimization, renewable energy integration, and hardware efficiency. Similarly, Patterson et al. [17] highlight the significant energy consumption and carbon footprint of training large NLP models and present strategies for reducing their environmental impact. The study reports that training GPT-3 emitted 552 tCO2e, of the same order of magnitude as a transcontinental flight. It also observes that sparse neural networks, such as mixture-of-experts architectures, may reduce energy consumption by more than 90% compared to dense models while maintaining accuracy. However, the mixture-of-experts approach has been leveraged to increase the size of language models, not to keep it constant [28]. Carbon emissions remain hard to estimate, though. A case study on the Evolved Transformer [20] revealed an 88x overestimation in prior calculations, underscoring the need for accurate reporting and optimized configurations. These results advocate for transparent reporting of energy and carbon metrics in DL research and prioritization of efficiency improvements to address environmental concerns. Tripp et al. [22] contribute to this reporting effort by making the BUTTER-E dataset available, which provides real-world energy consumption measurements of training fully connected neural networks across various architectures, sizes, and hardware configurations. The study reveals complex relationships between dataset size, network structure, and energy use, highlighting the significant impact of cache effects and challenging the assumption that reducing parameters or FLOPs necessarily leads to greater energy efficiency. The authors also propose an energy model that accounts for network size, computing, and memory hierarchy. 2.2 Computational Efficiency of Architectures Comparative analyses of different neural network architectures have provided valuable insights into their computational efficiency. For example, RNNs such as LSTM and GRU have traditionally been favored for sequential data processing due to their ability to capture temporal dependencies. However, the introduction of Transformers [23] has revolutionized the field, offering superior performance on tasks involving long-range dependencies. While Transformers are more computationally intensive, their parallelizable architecture has made them the dominant choice in NLP and other sequential data processing tasks. Recent work [10, 1, 17, 15] has examined the trade-offs between these architectures in terms of computational cost, memory usage, and inference latency, but few studies have focused explicitly on energy consumption during training. 2.3 Optimization Strategies for Energy Efficiency Efforts to reduce energy consumption in DL training have focused on both algorithmic and hardware optimizations. Techniques such as mixed-precision training [14] and gradient checkpointing [2] have been widely adopted to lower computational requirements without sacrificing model performance. On the architectural side, lightweight models such as DistilBERT [19] and MobileNet [8, 18] have been designed to balance performance and efficiency. Some of these techniques, such as distillation, still require to fully train an initial model. 2.4 Positioning of This Study While significant progress has been made in understanding the energy demands of DL models, most studies focus on specific architectures in isolation, lacking comprehensive comparisons under a unified framework. The relationship between energy efficiency and model design choices (number of layers, attention heads, hidden dimensions, etc.) also remains poorly understood. This study addresses these gaps by systematically evaluating the energy consumption of various configurations of LSTM, GRU, and Transformer models during training. We also introduce a methodology to estimate energy consumption prior to training the model, enabling researchers to predict resource requirements and optimize design choices in advance. Contrary to the work by Tripp et al. [22], which relies on detailed measurements from a specific hardware setup, our approach has minimal hardware dependence. Our work also expands its focus by examining RNNs and Transformers, in addition to multi-layer perceptrons, offering a broader perspective on energy consumption across different fundamental neural network architectures. This allows for a more generalized understanding of energy consumption trends and facilitates more informed architectural choices in the context of energy efficiency. 3 Methodology It is well-known that the energy consumption of computation on GPU is not proportional to computational cost, traditionally measured as a FLOPs count. However, energy consumption is highly correlated with execution time which, in turn, depends on the level of parallelization of elementary operations on the GPU. A fully parallelizable operation such as the Hadamard product will take less time than a matrix multiplication, even if they require the same amount of FLOPs. Their exact execution time will depend on the number of cores on the GPU. In addition, GPUs also spend several processing cycles to manage memory, moving tensors across main memory, GPU memory and GPU registers. Memory management operations also depend on hardware characteristics of the GPU, such as memory size and cache size. Parallelization and memory management both tie any energy consumption estimate to a specific hardware platform. Tensor manipulation frameworks such as PyTorch and TensorFlow associate elementary tensor operations (including Hadamard product and matrix multiplication) to kernels, i.e. pieces of code executed on GPU cores. At the hardware level, a training or inference pipeline consists in a sequence of kernel executions, such that the execution time of the entire pipeline can be estimated by providing an estimate for each kernel execution, in a compositional approach1. Our approach consists in providing a hardware efficiency factor (HEF) for each elementary operation. This HEF captures the behavior of a kernel both in terms of parallelization and memory management by comparing the actual compute throughput for that operation, in FLOP/s, with the maximum theoretical throughput on the GPU, achievable for a purely parallel, purely arithmetic operation. The HEF ηΘof operation Θis defined as follows: ηΘ=c/t vmax (1) where cis the computational cost of the operation in FLOPs, tis its execution time in seconds and vmax is the maximum theoretical throughput (or velocity) of the GPU in FLOP/s. 1Compilation techniques for deep learning pipelines have been recently introduced in the Accelerated Linear Algebra (XLA) compiler. For simplicity, we assume that pipelines are not compiled. M. Zoubeirou a Mayaki and V. Charpenay / Modeling Energy Consumption in Deep Learning Architectures Using Power Laws 937
The HEF quantity relates execution time and computational cost in a reliable way. The energy consumption eof training some arbitrary model can then be estimated via a sum over all elementary operations needed to train a single batch, scaled by a hardware-dependent factor. The general expression for eis as follows: e≈h+ n i=1 hi·ti(ci)(2) where ti=ci vmax ·ηΘi (3) and tiis the duration (time) of the i-th operation, which itself is a function of its workload ci. The constant hi(or elasticity) scales energy with respect to that operation’s duration. It measure the increase in total energy (all else equal) due to the increase in operation i. The main challenge to properly estimate eis thus to find an expression for ηΘ1,η Θ2,...,η Θnthat does not depend on the execution duration t. We empirically found that the HEF of most elementary operations follows a power law that only depends on their FLOPs count, with the following parametric expression: ηΘ≈ηmax 1−e−kcα(4) where ηmax is the maximum hardware efficiency factor (saturation limit) for operation Θ,kis a decay rate, determining how quickly the hardware approaches its maximum efficiency as compute increases, and αis an additional law parameter to better fit observations. The exponential decay with a saturation offset effectively models hardware efficiency, capturing the rapid initial improvements in resource utilization followed by diminishing returns as physical or architectural limits are approached (e.g., memory bandwidth or thermal constraints) [25, 24, 9]. Transformer Layer. Figure 1 and Table 1 give examples of power laws for operations inside a Transformer layer, as defined by Vaswani et al. [23]. Attention between matrices Q,Kand Vcan be calculated in these steps: (QKV projections) Q,K,V←QWQ,KWK,VWV(5) (Attention scoring) A←softmax(Q(K)T √d)(6) (Attention output) B←AV (7) (Final projection) O←BWO(8) where Q,K,V,A,B,O are intermediate tensors stored in memory, and WQ,WK,WV,WOare learned projection matrices. Figure 1 shows power laws of the HEF for each of these elementary operations. All operations saturate at a level below 100%, reflecting the fact that memory management imposes a significant overhead compared to purely arithmetic throughput. It is also clear from the figure that different hardware platforms achieve varying efficiency levels. Across all measured Transformer layer components, the highest observed efficiency on the A100 80GB GPU is about 70.4% (Final Projection), while the GeForce RTX 2080 Ti reaches up to 81.9% (Final Projection). Despite the RTX 2080 Ti’s higher peak percentage, the newer A100 remains faster in absolute terms because its maximum theoretical throughput is much greater (156 TFLOP/s vs. 13.45 TFLOP/s). Table 1. Fitted efficiency parameters (ηmax,k,α)for Transformer layer operations across NVIDIA A100 80GB PCIe and GeForce RTX 2080 Ti GPUs. Operation (θ) GPU ηmax kα Attention output A100 80GB PCIe 66.83 8.65 0.80 Attention scoring A100 80GB PCIe 56.47 8.09 0.80 QKV projections A100 80GB PCIe 69.38 10.37 0.78 Final projection A100 80GB PCIe 70.43 6.24 0.77 Attention output RTX 2080 Ti 78.55 21.77 0.56 Attention scores RTX 2080 Ti 71.42 14.25 0.52 QKV projections RTX 2080 Ti 81.45 18.94 0.52 Final projection RTX 2080 Ti 81.93 9.21 0.40 In the above example, QKV projections, attention Output and final Projection can be considered elementary but attention scoring is not. It itself involves a projection, where ηQKV may be reused to estimate its energy consumption. The calculation may be further decomposed into a sequence of four operations: (Transpose) A←KT(9) (Projection) B←QA (10) (Scalar division) C←B/√d(11) (Softmax) D←softmax(C)(12) where dis the number of columns in Q,Kand V, and A, B, C, D are intermediate results stored in memory. To estimate the energy consumption of the full attention mechanism, Equation 2 requires a HEF law for five operations: projection, attention output, transpose, scalar division and softmax. Feed-forward network layers, also present in the Transformer architecture, also consist of matrix multiplication on which the law for projection applies. Strictly speaking, the transpose operation yields no FLOP. We however adopt a generalized definition of FLOPs, which includes pure memory management operations. Other FLOPs counts are given in the online appendix [29, Appendix 7.3]. LSTM and GRU Layer. For LSTM and GRU architectures, the primary elementary operations revolve around their gating mechanisms, which manage information flow within the network. These key elementary operations can be expressed as follows: (Hidden gating)fh,i h,o h,˜ch←hWf hh,... (13) (Input gating)fx,i x,o x,˜cx←xWf ih,... (14) (Activation)f,i,o,˜c←σ(fh+fx),... (15) (Activation)˜c←tanh(˜c)(16) (Cell update)c←fc+i˜c(17) (Hidden update)h←oc(18) where denotes element-wise multiplication of vectors, and his the output hidden state. xis the input vector at the current time step, and his the previous hidden state. cis the cell state, which carries information across time steps, and cis the candidate cell state that will be added to the cell state. The intermediate results (f,i,o,˜c, . . .) are temporarily stored in memory during computation. Given input data x∈RBs×T×d, with Bsrepresenting the batch size, Ttime steps and dthe input dimensionality, and dhrepresenting the hidden dimension of the model, the dimensions of weight matrices are W hh ∈Rdh×dhand W ih ∈Rd×dh(∈{f,i,o,˜c}). All these operations are detailed in the online appendix [29, Appendix 7.4]. Table 2 shows the hardware efficiency trends for various elementary operations across RNN components executed on the NVIDIA A100 M. Zoubeirou a Mayaki and V. Charpenay / Modeling Energy Consumption in Deep Learning Architectures Using Power Laws938
Figure 1. Hardware efficiency factor (HEF) of elementary operations in a Transformer layer. 80GB PCIe GPU. As expected, efficiency saturates at higher compute load, with operations like gating achieving substantially higher peak efficiency (up to 80%) due to their heavier computational demands. In contrast, operations like cell update and activations (σand tanh) exhibit lower peak efficiencies, likely due to their smaller or less parallel workloads. Table 2. Fitted efficiency parameters (ηmax,k,α)for LSTM operations on A100 GPU. Operation (θ)ηmax kα Activations 0.0965 17544.28 0.94 Cell Update 0.0554 12823.88 0.88 Hidden Update 0.0977 10286.70 0.89 Input Gating 41.03 21.57 0.86 Hidden Gating 80.04 10.08 0.79 For a given hardware platform, HEF power law coefficients must be calculated only once per operation. Once a HEF exists for the usual elementary operations, the energy consumption of any deep learning model can be estimated via Equation 2, at design time. In the next section, we evaluate the generality of our approach on Transformers and other well-known model architectures. Note that, for a full Transformer model, one might add the feed-forward network (FFN) component; however, our experiments showed that doing so did not improve predictive accuracy. The reason is twofold: (i) FFN FLOPs are nearly a fixed multiple of the matrix multiplications we already model, making the new term highly collinear and thus redundant for regression; and (ii) in our measurement regime, runtime and energy are primarily driven by memory movement in attention (score/probability tensors and key–value traffic), not by the compute-bound FFN. Consistent with this, augmenting the model with I/O-oriented features (such as approximate bytes moved and precision/FlashAttention switches) yields larger gains than refining FFN arithmetic. We therefore retain the simpler FFN approximation and focus model capacity on I/O-sensitive components. 4 Experiments 4.1 Learning Task and Architectures In our experiments, we used simulated sequential data to systematically evaluate the performance and efficiency of various models under different conditions. For the Transformer-based architectures, we used the pre-trained GPT-2 and BERT models, sourced from the Hugging Face Transformers library. For each architecture, we varied several key hyperparameters and data characteristics, including: the number of training samples (to assess scalability with increasing dataset size); the sequence length, ranging from 128 to 1024 tokens; and the dimensionality of the input data, spanning from 64 to 640 features. Furthermore, we explored the impact of internal model configurations by adjusting the number of layers (2, 4, 6, and 12), the embedding dimensionality (64 to 640), and the number of attention heads (2, 4, and 8). We trained over 2,500 Transformer models with number of parameters varying from 3 millions to 300 millions and 2,000 RNNs (LSTM and GRU) of size varying from 200 thousands to 278 millions parameters. The experimental data and codes are available online [29]. 4.2 Experimental Setup Model Configurations We trained each model for 5 epochs across a diverse range of hyperparameter configurations to thoroughly evaluate their energy consumption and computational efficiency. The hyperparameter configurations included variations in batch sizes, and model-specific parameters such as the number of hidden units, layers and number of heads (see details in the online appendix [29, Table 6]). This approach allowed us to capture how different settings influenced both the training dynamics and resource utilization. Energy Measurement. Energy consumption during training is tracked using a combination of tools that provide detailed metrics for both GPU and CPU usage. These tools allow for precise monitoring and aggregation of energy usage across the entire training process for each model and configuration. The total energy consumption is reported in kilowatt-hours (kWh). The following tools are used: - NVIDIA Nsight Systems. Provides detailed metrics on GPU power consumption, enabling fine-grained monitoring of energy usage at the GPU level. - CodeCarbon [4]. Tracks the energy consumption of the entire training process, including both GPU and CPU usage. Additionally, it estimates the carbon footprint (CO2emissions) associated with the energy consumed. - Used Hardware. The experiments are conducted on an NVIDIA A100 GPU with 80GB memory and GeForce RTX 2080 Ti. The enM. Zoubeirou a Mayaki and V. Charpenay / Modeling Energy Consumption in Deep Learning Architectures Using Power Laws 939
ergy consumption is measured for both the GPU, CPU and RAM to provide a comprehensive analysis. For each configuration, energy consumption was recorded over 5 epochs and averaged. This methodology ensures a consistent and fair comparison across models while providing detailed insights into the factors influencing energy consumption. 4.3 Results For these experiments, we initially estimated the duration tiof each elementary operation within the model using equations (4) and (3). Subsequently, these estimated durations tiserved as input features for our proposed energy consumption model, as described by equation (2). To enhance numerical stability and improve interpretability of regression coefficients, we expressed the duration in microseconds. This was necessary because the original duration values were very small (on the order of microseconds), which can lead to numerical issues in model fitting. Additionally, all energy measurements were converted from watt-hours (Wh) to joules (J) using the standard conversion factor: 1Wh =3,600 J. These transformations do not affect the validity of the regression results but ensure better numerical conditioning and clarity in reporting. Further details of the estimation procedure are described in the online appendix. The coefficients should be interpreted with caution. Each term in the model reflects the marginal effect of its corresponding variable while holding all other variables constant. The negative coefficients from table 3 and table 4 result from the multivariate linear model, where each term reflects the marginal effect of a variable while holding others constant. Due to multicollinearity among features common in neural network operations where components are highly interdependent some coefficients may turn negative to balance overlapping effects and improve overall fit. These values do not necessarily imply that the associated operations are inherently energy-reducing; rather, they indicate how energy consumption changes when that variable increases, assuming all other factors remain fixed. Predictions from our model (2) were compared against the measurements from code carbon using the coefficient of determination (R2). This allowed us to assess both the accuracy and the linear alignment of the model’s predictions with empirical data. The estimated coefficients may also vary with dataset size because the regression is sensitive to the distribution of configurations and runtimes in the sample. Sign changes occur when correlated terms (e.g., projection and attention score FLOPs) redistribute variance between them, a common statistical effect that does not imply a reversal of the underlying physical relationship and does not affect the model’s predictive performance. Energy consumption per elementary operation. Figure 2 illustrates the relationship between operation duration and energy consumption for different attention-related components across two Transformer models: GPT and BERT. The plots reveal a consistent positive correlation between duration and energy consumption, particularly for Attention Scores and Final Projection operations, across both models. Notably, BERT shows a denser clustering at shorter durations, suggesting generally faster execution times compared to GPT. However, both models exhibit similar energy scaling behavior, with QKV Projections consistently consuming less energy relative to the other components. The lines smooth out fluctuations, confirming that longer operation durations lead to higher energy usage, albeit with varying slopes depending on the component and model architecture. Energy estimation model for Transformers. The experimental results for Transformer architectures (Table 3 and Figure 3) confirm that our proposed model accurately predicts energy consumption, with a high coefficient of determination (R2=0.96). Specifically, the estimated constant term (3.63 joules) provides a baseline for energy use independent of computational operations. Among the analyzed components, the duration of the final projection operation (t4) has the strongest positive effect, increasing energy consumption by approximately 0.56 joules per additional microsecond. Attention score computations (t2) and attention output operations (t3) both show consistent positive contributions of about 0.30 joules per microsecond. In contrast, QKV projections (t1) exhibit a small but significant negative coefficient (−0.14 joules per microsecond), indicating that longer runtimes in these operations are associated with reduced incremental energy costs likely reflecting overlap with memory-bound stages or hardware-level optimizations. The final energy consumption model for Transformer architectures is written as follows: Etrans ≈3.6318 −0.1377 ·t1+0.3041 ·t2 +0.3041 ·t3+0.5637 ·t4 (19) Table 3. Estimated coefficients for energy consumption of Transformer models. The last two columns give the lower and upper bounds of the 95% confidence intervals. With a coefficient of determination R2=0.9584. Durations are expressed in microseconds and the energy in joules. Operation Estimate CI0.025 CI0.975 Constant (h) 3.6318 3.4088 3.8858 QKV projections -0.1377 -0.1455 -0.1297 Attention scoring 0.3041 0.3016 0.3065 Final projection 0.5637 0.5396 0.5880 Attention output 0.3041 0.3016 0.3065 Energy estimation model for RNNs. We also tested our method on LSTM and GRU architectures. Initially, we focused on the core gating operations and hidden-state updates. This initial model (Table 4) had a fit (R2) = 0.95), suggesting that these operations explain a significant portion of energy consumption, but accuracy could be improved. We observed that each additional micro-second spent on gating operations reduced energy consumption by approximately 169 Joules, suggesting efficient computational characteristics of gating processes. Each micro second of cell update and hidden update also increases the energy consumption respectively by 4.52×104and 2.95×104. Conversely, activation computations reduces energy usage by about −7.66×104J/μs. Table 4. Estimated coefficients, standard errors, and 95% confidence intervals for RNNs (LSTM and GRU). With a coefficient of determination R2=0.95. Durations are expressed in microseconds and energy in Joules. Operation Estimate Std. Error CI0.025 CI0.975 constant (h) 3367.3304 8.375 3350.882 3383.779 Activations −7.66×1047021.058 −9.04×104−6.28×104 Cell update 4.52×1044221.466 3.70×1045.35×104 Gating 169.0733 9.072 151.255 186.892 Hidden update 2.95×1042813.017 2.39×1043.50×104 Model generalization across GPU architectures. Figure 4 demonstrates the robustness and generalizability of our proposed energy estimation method across different GPU architectures. In this M. Zoubeirou a Mayaki and V. Charpenay / Modeling Energy Consumption in Deep Learning Architectures Using Power Laws940
Figure 2. Plot of energy consumption versus operation duration for key Transformer components, with moving average trend lines. The data is shown separately for GPT and BERT models. The trend lines highlight the average energy scaling behavior as operation duration increases. Figure 3. Quality of energy consumption estimates by architecture. experiment, we evaluated the model using two distinct GPUs: the NVIDIA A100 80GB PCIe and the NVIDIA GeForce RTX 2080 Ti. The plot compares predicted total energy consumption against actual measured values, with each point representing a single model inference. The model achieves an R2score of 0.98 on the test set, indicating that it explains 98% of the variance in energy consumption across both devices. This high level of accuracy across heterogeneous hardware platforms highlights the method’s ability to capture fundamental computational behaviors rather than overfitting to a specific architecture. These results validate the portability and scalability of our approach for estimating energy consumption in real-world, multi-device settings. Figure 4. Quality of energy consumption estimates by GPU Model. 5 Discussion 5.1 Interpretation The results of our experiments demonstrate that the proposed energy estimation model is both accurate and generalizable across diverse neural network architectures and hardware platforms. For Transformer architectures, the model captured key elements of energy use, most notably the attention score computation as reflected by high explanatory power (R2=0.93). Similarly, for RNN architectures, the model effectively identified hidden-state updates as the dominant contributor to energy consumption (R2=0.95). Importantly, our method remained consistent across different GPU architectures, when tested on both the NVIDIA A100 and RTX 2080 Ti platforms. This confirms that the model captures underlying patterns in computational behaviour rather than hardware-specific effects. These findings suggest that the method could be applied reliably in various realM. Zoubeirou a Mayaki and V. Charpenay / Modeling Energy Consumption in Deep Learning Architectures Using Power Laws 941
world deployment settings, offering a lightweight and scalable tool to estimate energy consumption without the need for intrusive hardware measurements. 5.2 Application: Energy-Aware Model Selection This section presents a practical use of our energy estimation method to support model selection based on energy efficiency. We compare two Transformer configurations on a GPU with maximum throughput of 156 TFLOPs/s. Rather than benchmarking, our method uses architectural parameters to estimate energy consumption. Hardware and input settings. All the following steps are encapsulated in modular functions, allowing practitioners to simply define their model configuration and workload parameters ( batch size, sequence length, layers, heads and embedding dimensionality dmodel ). Once these inputs are provided, the system automatically computes the estimated energy consumption for a single training epoch. •Max. Throughput: vmax = 156 ×1012 FLOPs/s •Batch Size= 64, Sequence Length= 320 Model configurations. Let us compare two GPT models. •Model A: 6 layers, dmodel =512, 8 attention heads •Model B: 12 layers, dmodel =768, 12 attention heads Step 1: FLOPs estimation Using the analytical formulas for elementary operations described in Section 3, we compute the number of floating-point operations (FLOPs) required for each component of the Transformer architecture. The third column of Table 5 gives the computational load associated with Transformer operations: projections, attention score computation, and output operations. Step 2: Hardware efficiency and duration Using the fitted parameters from Table 1, the hardware efficiency and the duration for each operation in Table 5 were computed with Equations 3 and 4. Table 5. FLOPs (in teraflops), efficiency values ηΘ(c), and operation durations tΘ(c)(in micro-seconds) for each model and operation. Operation Model FLOPs (TF) ηΘ(c)tΘ(c)(μs) QKV Projections A 0.0322 35.31 35.08 Attention Scores A 0.00674 7.75 33.28 Attention output A 0.00674 9.76 26.43 Final Projection A 0.01074 12.19 33.87 QKV Projections B 0.0724 51.19 108.90 Attention Scores B 0.01 10.43 74.21 Attention output B 0.01 13.11 59.05 Final Projection B 0.024 21.04 88.31 Step 3: Energy Estimation We now apply Equation 19 to estimate the total energy consumption of the two Transformer configurations: EA≈36 J and EB≈79 J. Despite being evaluated under identical input conditions, Model B is estimated to consume more than twice energy per epoch than Model A. While a single-epoch difference of roughly 42 joules may seem modest, the effect becomes pronounced at scale: over one million training epochs, this difference would accumulate to more than 11 kWh. Such a gap is non-negligible in batterypowered, embedded, or high-throughput environments, where energy efficiency directly constrains deployment feasibility. The higher cost of Model B is largely attributable to its increased dimensionality and number of layers, which disproportionately amplify the computational burden of projection and score computations relative to Model A. Beyond highlighting the energy trade-offs between model sizes, this example illustrates how our method enables early-stage architectural comparisons and energy-aware design choices without requiring execution on actual hardware. Further experiments (in the online appendix [29, Tables 10-11]) show that energy consumption is largely insensitive to the number of attention heads, with only negligible variation across configurations. In contrast, the number of layers and the dimensionality have a pronounced impact, with deeper models consuming substantially more energy. This suggests that model depth and the dimensionality of the input and output of all the sub-layers are the dominant factor in Transformer energy efficiency. 5.3 Limitations and Future Directions The main limitation of our study is that all experiments were conducted using single-GPU setups. While this approach provides valuable insights into energy consumption for a controlled environment, it does not fully capture the complexities and efficiencies of training modern state-of-the-art models, which often rely on multiple GPUs or TPUs operating in parallel. These large-scale setups benefit from optimized hardware utilization, distributed computation, and lower energy consumption per unit of computation due to specialized accelerators. Consequently, our results may not fully reflect the energy efficiency or scalability of these systems in real-world, multi-device training scenarios. Future work could extend the experiments to distributed environments with multiple GPUs or TPUs [15]. 6 Conclusion In this work, we proposed and validated a method to estimate the energy consumption of machine learning models using the elementary operations in the model architecture and power laws. The results showed a strong correlation between the estimated and measured energy consumption, with a high R2value, particularly for lower and mid-complexity configurations. Overall, this study provides a practical framework for estimating energy consumption before training, offering valuable insights for optimizing model design and promoting energy-efficient machine learning practices. The proposed method can serve as a useful tool for researchers and practitioners aiming to balance model performance with environmental sustainability. Future work could expand on this study by validating the proposed energy estimation method in distributed training environments, where factors such as inter-node communication, load balancing, and data sharding significantly impact energy usage. Incorporating additional aspects such as parallel efficiency, memory bandwidth limitations, and mixed precision training would further enhance the accuracy and generalizability of the model. Additionally, we plan to apply and adapt our method to emerging architectures that are designed for greater efficiency, such as mixture-of-experts models, which introduce dynamic sparsity and routing mechanisms that pose new challenges for precise energy estimation. Acknowledgements This work was supported by the European Lighthouse to Manifest Trustworthy and Green AI (ENFIELD) project funded by the European Union’s HORIZON Research and Innovation Programme. Grant agreement No: 101120657. The code and the supplementary materials for this paper is available on Zenodo [29] and on GitHub2. 2https://github.com/MansMayaki/MLEnergyConsumption M. Zoubeirou a Mayaki and V. Charpenay / Modeling Energy Consumption in Deep Learning Architectures Using Power Laws942
References [1] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. [2] T. Chen, B. Xu, C. Zhang, and C. Guestrin. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174, 2016. [3] B. Cottier, R. Rahman, L. Fattorini, N. Maslej, and D. Owen. The rising costs of training frontier ai models. arXiv preprint arXiv:2405.21015, 2024. [4] B. Courty, V. Schmidt, S. Luccioni, Goyal-Kamal, MarionCoutarel, B. Feld, J. Lecourt, LiamConnell, A. Saboni, Inimaz, supatomic, M. Léval, L. Blanche, A. Cruveiller, ouminasara, F. Zhao, A. Joshi, A. Bogroff, H. de Lavoreille, N. Laskaris, E. Abati, D. Blank, Z. Wang, A. Catovic, M. Alencon, Michał St˛echły, C. Bauer, L. O. N. de Araújo, JPW, and MinervaBooks. mlco2/codecarbon: v2.4.1, May 2024. URL https://doi.org/10.5281/zenodo.11171501. [5] DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J. Ni, J. Li, J. Wang, J. Chen, J. Chen, J. Yuan, J. Qiu, J. Li, J. Song, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Xu, L. Xia, L. Zhao, L. Wang, L. Zhang, M. Li, M. Wang, M. Zhang, M. Zhang, M. Tang, M. Li, N. Tian, P. Huang, P. Wang, P. Zhang, Q. Wang, Q. Zhu, Q. Chen, Q. Du, R. J. Chen, R. L. Jin, R. Ge, R. Zhang, R. Pan, R. Wang, R. Xu, R. Zhang, R. Chen, S. S. Li, S. Lu, S. Zhou, S. Chen, S. Wu, S. Ye, S. Ye, S. Ma, S. Wang, S. Zhou, S. Yu, S. Zhou, S. Pan, T. Wang, T. Yun, T. Pei, T. Sun, W. L. Xiao, W. Zeng, W. Zhao, W. An, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, X. Q. Li, X. Jin, X. Wang, X. Bi, X. Liu, X. Wang, X. Shen, X. Chen, X. Zhang, X. Chen, X. Nie, X. Sun, X. Wang, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yu, X. Song, X. Shan, X. Zhou, X. Yang, X. Li, X. Su, X. Lin, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. X. Zhu, Y. Zhang, Y. Xu, Y. Xu, Y. Huang, Y. Li, Y. Zhao, Y. Sun, Y. Li, Y. Wang, Y. Yu, Y. Zheng, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Tang, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Wu, Y. Ou, Y. Zhu, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Zha, Y. Xiong, Y. Ma, Y. Yan, Y. Luo, Y. You, Y. Liu, Y. Zhou, Z. F. Wu, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Huang, Z. Zhang, Z. Xie, Z. Zhang, Z. Hao, Z. Gou, Z. Ma, Z. Yan, Z. Shao, Z. Xu, Z. Wu, Z. Zhang, Z. Li, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Gao, and Z. Pan. Deepseek-v3 technical report, 2025. URL https://arxiv.org/abs/2412.19437. [6] A. Faiz, S. Kaneda, R. Wang, R. C. Osi, P. Sharma, F. Chen, and L. Jiang. LLMCarbon: Modeling the end-to-end carbon footprint of large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum? id=aIok3ZD9to. [7] P. Henderson, J. Hu, J. Romoff, E. Brunskill, D. Jurafsky, and J. Pineau. Towards the systematic reporting of the energy and carbon footprints of machine learning. Journal of Machine Learning Research, 21(248): 1–43, 2020. [8] A. G. Howard. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. [9] N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, et al. In-datacenter performance analysis of a tensor processing unit. In Proceedings of the 44th annual international symposium on computer architecture, pages 1–12, 2017. [10] J. Kaplan, S. McCandlish, T. H. OpenAI, T. B. B. OpenAI, B. C. OpenAI, R. C. OpenAI, S. G. OpenAI, A. R. OpenAI, J. W. OpenAI, and D. A. OpenAI. Scaling laws for neural language models, 2020. [11] A. S. Luccioni, S. Viguier, and A.-L. Ligozat. Estimating the carbon footprint of bloom, a 176b parameter language model. Journal of Machine Learning Research, 24(253):1–15, 2023. [12] S. Luccioni, Y. Jernite, and E. Strubell. Power hungry processing: Watts driving the cost of ai deployment? In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 85–99, 2024. [13] V. Mehlin, S. Schacht, and C. Lanquillon. Towards energy-efficient deep learning: An overview of energy-efficient approaches along the deep learning lifecycle. arXiv preprint arXiv:2303.01980, 2023. [14] P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, et al. Mixed precision training. arXiv preprint arXiv:1710.03740, 2017. [15] D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, et al. Efficient large-scale language model training on gpu clusters using megatron-lm. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1– 15, 2021. [16] C. Ogbogu, M. Abernot, C. Delacour, A. Todri-Sanial, S. Pasricha, and P. P. Pande. Energy-efficient machine learning acceleration: From technologies to circuits and systems. In 2023 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED), pages 1–8, 2023. doi: 10.1109/ISLPED58423.2023.10244360. [17] D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021. [18] M. Rhu, N. Gimelshein, J. Clemons, A. Zulfiqar, and S. W. Keckler. vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design. In 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), pages 1–13. IEEE, 2016. [19] V. Sanh. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019. [20] D. So, Q. Le, and C. Liang. The evolved transformer. In International conference on machine learning, pages 5877–5886. PMLR, 2019. [21] E. Strubell, A. Ganesh, and A. McCallum. Energy and policy considerations for modern deep learning research. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 13693–13696, 2020. [22] C. E. Tripp, J. Perr-Sauer, J. Gafur, A. Nag, A. Purkayastha, S. Zisman, and E. A. Bensen. Measuring the energy consumption and efficiency of deep neural networks: An empirical analysis and design recommendations. arXiv preprint arXiv:2403.08151, 2024. [23] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017. URL https://proceedings.neurips.cc/paper/ 2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html. [24] J. Végh. How amdahl’s law limits the performance of large artificial neural networks: why the functionality of full-scale brain simulation on processor-based simulators is limited. Brain informatics, 6(1):4, 2019. [25] S. Williams, A. Waterman, and D. Patterson. Roofline: an insightful visual performance model for multicore architectures. Communications of the ACM, 52(4):65–76, 2009. [26] C.-J. Wu, R. Raghavendra, U. Gupta, B. Acun, N. Ardalani, K. Maeng, G. Chang, F. Aga, J. Huang, C. Bai, et al. Sustainable ai: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems, 4:795–813, 2022. [27] T. Yarally, L. Cruz, D. Feitosa, J. Sallou, and A. Van Deursen. Uncovering energy-efficient practices in deep learning training: Preliminary steps towards green ai. In 2023 IEEE/ACM 2nd International Conference on AI Engineering–Software Engineering for AI (CAIN), pages 25–36. IEEE, 2023. [28] Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. M. Dai, Q. V. Le, J. Laudon, et al. Mixture-of-experts with expert choice routing. Advances in Neural Information Processing Systems, 35:7103–7114, 2022. [29] M. Zoubeirou a Mayaki and V. Charpenay. Dataset, code and online appendix for: Modeling energy consumption in deep learning architectures using power laws (ecai 2025) (v1.0.0), 2025. URL https: //doi.org/10.5281/zenodo.16920231. M. Zoubeirou a Mayaki and V. Charpenay / Modeling Energy Consumption in Deep Learning Architectures Using Power Laws 943