Full text
MARIA GREITZER ENERGY FORECASTING. FOCUS: NATURAL GAS DISSERTATION HAFENCITY UNIVERSITY HAMBURG
A dissertation submitted to HafenCity University Hamburg in fulfilment of the requirements of the “Promotionsordnung der HafenCity Universität Hamburg” for the degree of Doktor-Ingenieur (Dr.-Ing.) Maria Greitzer Supervisor Prof. Dr.-Ing. Ingo Weidlich, HafenCity University Second Supervisor Prof. Michael Bradshaw, The University of Warwick tufte-latex.github.io/tufte-latex/ Licensed under the Apache License, Version 2.0(the “License”); you may not use this file except in compliance with the License. You may obtain a copy of the License at http://www.apache.org/licenses/ LICENSE-2.0. Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "as is" basis, without warranties or conditions of any kind, either express or implied. See the License for the specific language governing permissions and limitations under the License. First printing, April 2021
Contents 1Preface 17 2Introduction 19 2.1Research questions 20 2.2What is forecasting? 21 2.3Scenario modelling 21 2.4Why complexity? 22 2.5Scope 23 3Theoretical Background 25 3.1Energy forecasting vs energy modelling 25 3.1.1Energy forecasting - examples 26 3.1.2Energy modelling - examples 27 3.1.3Analogy with music: An idea experiment 29 3.1.4Interdisciplinary Aspects 31 3.1.5Conclusion 33 3.2Related work 34 3.2.1Forecasting assumptions 34 3.2.2Accuracy criteria 35 3.2.3The essential idea of methods 37 3.2.4Objective methods according to Januschowski et al. (2020)41 3.2.5Forecasting models and applications 44 3.2.6Excluded models 56 3.2.7Econometric models, grey models, genetic algorithms 59 3.2.8Conclusion 59
4 3.3Complexity 62 3.3.1Terminology 62 3.3.2Approximate Entropy and Sample Entropy 64 3.3.3Complexity of forecasting models 65 3.3.4Complexity of a phenomenon in relation to modelling 67 4Study 1: German gas imports 69 4.0.1Data collection 74 4.0.2Material and methods 77 4.0.3Time series models 77 4.0.4Regression analysis 82 4.0.5Artificial Neural Networks 83 4.0.6Modelling conclusions 84 4.0.7Geopolitical implications of gas imports 86 5Study 2: Yearly gas consumption in Germany 91 5.0.1Introduction 91 5.0.2Forecasting literature 92 5.0.3Research setup 93 5.0.4Methodology and Results 97 5.0.5Principal Components Regression 100 5.0.6Support vector machines (SVM) 102 5.0.7Artificial neural networks 104 5.0.8Intersection gas supply security and forecasting 109 5.1Implications for the energy sector 112 5.1.1Hydrogen 112 5.1.2Publicly available estimations of gas consumption in Germany 113 6Conclusion 117 6.1Closing thoughts 119 6.2The value of forecasting 120
5 Appendices 129 7Bibliography 137
List of Figures 3.1Energy forecasting vs energy modelling 28 3.2Parallels between steps for learning jazz piano improvisation and developing new methods for energy forecasting. 31 3.3Illustration of the role of social science in energy forecasting. 31 3.4Visualisation of methods used for energy and water demand. 44 3.5Visualisation of relation of low dimension to high dimension. 53 3.6Basic idea of the backpropagation algorithm 53 3.7SVM - separating data in a two-dimensional space 54 4.1Transmission Capacity Map - Germany, ENTSOG (2019)71 4.2Seasonal plot of gas imports into Germany 77 4.3Decomposition of additive time series 78 4.4Persistence forecast 79 4.5Holt-Winters filtering 80 4.6Forecasts. Holt-Winters Model 81 4.7Autororrelation lag 1-20 (lag = 20)82 4.8Plot of a trained neural network 84 4.9Correlation plot of gas consumption in Germany and its neighbouring countries 87 4.10 Time plot of gas consumption of Germany and neighbouring countries 88 5.1Increasing variability of gas consumption (per cent change from previous year) 94 5.2Variables correlation matrix 97 5.3Prediction. Full linear regression model, 2012-2018 100 5.4Prediction. Reduced linear regression model (2012-2018)100 5.5Cross-validated predictions for gas consumption data 101 5.6One-Sigma strategy for finding optimal model dimension 101 5.7Gas Consumption forecast using PCR with the first 8principal components (2012-2018)102 5.8SVM full model, tuned (2012-2018)102 5.9SVM full model, tuned II (2012-2018)103 5.11 SVM reduced (2012-2018)103
8 5.10 SVM Variable Importance 103 5.12 The influence of weight backtracking on prediction 105 5.13 Prediction (2012-2018), ANN ( backpropagation) 105 5.14 Prediction (2012-2018). "Smaller learning rate" 106 5.16 Gas Consumption forecast using ANN for the reduced model 1106 5.15 Neural network variable importance from method 1106 5.17 Neural network variable importance, method (random forests) 107 5.18 Gas Consumption forecast using ANN for the reduced model 2107 5.19 Gas Consumption forecast (2012 −2018) using ANN for the full model, Approach 1107 5.20 Gas Consumption forecast (2012 −2018) using ANN for the full model, Approach 2108 5.21 Gas consumption forecast 2012-2018 (One year training with a) true values b) forecasted values) 108 5.22 Market prognosis for gas consumption 114 5.23 Paths for German gas imports 115 5.24 Paths for German consumption 115
List of Tables 3.1Energy forecasting vs. energy modelling with a focus on the gas sector 30 3.2Forecast horizons for wind energy, district heat and gas networks 39 3.3Comparison of predictions made by Ma and Li (2010) with the statistical data from the BP, Statistical Review of World Energy 2019, Gas consumption in billion cubic metres. 59 3.4ApEn and Sample Entropy for the data set "Gas imports" 65 4.1LNG Import terminals in Germany. Source: Gas Infrastructure Europe (2021)72 4.2Case studies - short term forecasting 73 4.3Approaches to gas demand forecasting and gas imports’ forecast. Study case: Germany 75 4.4MAPE comparison. 85 5.1List of Abbreviations of all variables in the data set "Gas consumption" 91 5.2Prediction of gas consumption up to 2030 96 5.3Short-term and long-term forecasting of gas imports/consumption 98 5.4Evaluation of main forecasting models 109 1Data key for German gas imports 131 2Data key for German gas consumption 132
energy forecasting.focus:natural gas 16 Abschnitt 3.1steht für den kreativen Teil der Arbeit zu Beginn der Forschung (2018-2019) und erstellt Analogien zwischen Prognosen und Musik. • Die Anwendung von Komplexitätsmaßen wie der ungefähren Entropie (ApEn) und der Probenentropie auf zwei selbst konstruierte Realwelt-Datensätze zu Gasimporten nach Deutschland und Gasverbrauch in Deutschland (Abschnitt 3.3)
1 Preface While writing, statements and expressions with lower information value were left out to keep this thesis compact with respect to the reader’s time. My personal aim was to continue with this research topic only if it still interested me. To my surprise, I enjoyed the process until the last typed sentence. This one.
1In further statements, we refer to gas only, omitting an adjective natural. 2Merkel, G., Povinelli, R., and Brown, R. (2018). Short term load forecasting of natural gas with deep neural network regression. Energies,11(8) 3Tamba, J. G., Essiane, S. N., Sapnken, E. F., Koffi, F. D., Nsouandélé, J. L., Soldo, B., and Njomo Donatien (2018). Forecasting natural gas: A literature survey. International Journal of Energy Economics and Policy, (8(3)):216–249 4Liu, S., editor (2011). Proceedings of 2011 IEEE International Conference on Grey Systems and Intelligent Services (GSIS) with the 15th WOSC International Congress on Cybernetics and Systems, Nanjing, China, 15 -18 September 2011, Piscataway, NJ. IEEE 2 Introduction Gas forecasting1covers the forecasting of gas demand of a city or a country, gas prices, and gas imports. The history of forecasting methods starts with classical mathematical relationship models based on statistics (e.g.linear regression, autoregressive integrated moving average) and ranges to artificial intelligence (e.g. artificial neural network (ANN), deep neural networks). The overview of the latter methods is provided by Merkel et al. (2018),2whereas Tamba et al. (2018)3conducted a literature survey on gas forecasting. Literature on forecasting methods is immense. For grey system theory in China alone, more than 50,000 papers were retrieved from academic periodical databases covering the period from 1982 (the date of the first publication on this topic) to 2010 (Liu (2011)).4Therefore, instead of conducting another literature survey, this work captures the essential idea of the most common gas forecasting methods along with examples from literature. The theory chapter 1) examines energy forecasting and energy modelling, 2) provides theoretical background to understand the methods used in the second section, and 3) discusses various notions of complexity and their eventual use for forecasting. Alternative approaches to the topic would cover: 1) earth science and perspectives on climate change and energy scarcity; 2) a purely economic perspective on imports and exports of energy resources, as well as incentives and sanctions of governments; and 3) energy security perspective from an International Relations (IR) point of view. The experimental chapter applies models to self-constructed data sets, and forecasts gas imports into Germany and gas consumption at the country level. Models are understood as tools used in energy engineering and policy making but they have not become the object of the research per se. If this work intended to introduce new methods for energy forecasting, this work would use standard data sets which
energy forecasting.focus:natural gas 20 5Franco, A. and Fantozzi, F. (2015). Analysis and clustering of natural gas consumption data for thermal energy use forecasting. Journal of Physics: Conference Series,655:012020 6Examples for ML models: support vector machine models, decision trees, artificial neural networks, nonlinear programming. 7Examples for statistical models: autoregressive integrated moving average, linear regression, logistic regression. would enable the comparison of the models’ performance. 2.1Research questions • RQ1. What are current approaches to predict gas imports and gas consumption at the country level? What is the impact of the domain knowledge of the gas sector on forecasting? Gas consumption at country level depends on the amount of gas used to cover residential heat demand as well as on the energy intensiveness of industry. Models may include this type of domain knowledge or work with time series methods. As residential heat demand is correlated with weather, the accuracy of weather forecasts influences the accuracy of forecasting gas load (Franco and Fantozzi (2015)).5 Gas imports, on the other hand, are more related to contracts among exporting and importing countries and, in the long run, on the environmental policies of gas use in future decades. Regarding current approaches, the forecasting community compares machine learning methods6with statistical methods7in terms of accuracy criteria. This work does not contribute to this debate as both terms "statistical" and "machine learning" are ill-defined. Here, structured and unstructured models will be discussed. To understand how forecasting methods work, a few models will be tested on two self-constructed data sets for German gas imports and German gas consumption. The focus lies in constructing data sets and understanding current models instead of developing their more complex versions. • RQ2. Is there any relation between the complexity of the data set, forecasting models and the process modelled (gas flows) to the accuracy of forecasts? As for the complexity of a data set, a few measures such as the Sample Entropy will be tested on the above-mentioned data sets. All data sets represent just a sample providing incomplete information on the process forecasted; thus there is the inherent uncertainty about including the most representative variables that influence the output.
energy forecasting.focus:natural gas 21 8Over-parametrized models will capture not only essential information from the data set but also noise (Pelikán (2014)). 9Makridakis, S. (1995). Forecasting accuracy and system complexity. RAIRO - Operations Research - Recherche Opérationnelle,29(3):259–283.http: //www.numdam.org/item/RO_1995__29_ 3_259_0/ "A forecasting model (predictive model, autoregressive model) is a softwareimplemented model of a system, process, or phenomenon, usable to forecast a value, output, or outcome expected from the system, process, or phenomenon." (Baughman et al. (1988)). 10 Intra-day and day-ahead forecasts assist in scheduling partially flexible conventional generation (i.e. coal power plants, gas power plants). There is no clear distinction between simple and complex models. The organizers of the M3 forecasting competition listed, among others, naïve models and exponential smoothing models as examples of simple models (Green and Armstrong (2015)). Complexity of the models will be expressed in various ways; for example in the computational time required for the model to be run, in the number of parameters8and degrees of freedom. Complexity of the process being modelled increases with time, therefore there are limited possibilities to forecast for a long-term horizon. The essence of the process is also crucial; forecasting the behaviour of economic systems cannot outperform the performance of forecasts for other system with chaotic properties, such as meteorological ones (Makridakis (1995)).9 2.2What is forecasting? Forecasting covers methods for producing forecasts by using historical data to make predictions and to determine trends. This relatively new discipline seeks to gain independence from the fields of statistics and machine learning, although the first generation of forecasters arose from statisticians and econometricians. In renewable energy forecasting, forecasts of solar radiation and wind speeds are used to optimize forecasts for electricity production from renewable energy sources. These natural phenomena affect output to a high degree and are the main source of uncertainty (when it comes to renewable energy production). Predictions,meaning the outcome of the forecasting, support company decisions on daily operations10 (e.g. fuel storage, wood pellet bunkers), strategic decisions on investments and more. Although energy forecasting is regarded as an academic sub-discipline, its potential to impact decisions in the energy sector is expected to grow. 2.3Scenario modelling Scenario modelling is a tool for governments’ decision-makers to introduce policies regarding climate change, for example. Often, a goal
energy forecasting.focus:natural gas 22 11 In German, Deutsches Zentrum für Luft und Raumfahrt, DLR. 12 Friedrichs, J. (2013). The future is not what it used to be: Climate Change and Energy Scarcity. MIT Press, Cambridge, Mass. ISBN 9780262019248 (e.g. the share of renewable energy in a system) is already defined and research questions relate to options for achieving the goal; e.g. how many megawatts of capacity and which type of installed plant is needed to produce almost all electricity from renewable energy sources in 2050? To answer such a question, techno-economic optimization models are used, for example the REMIX-Europe Model of German Aerospace Center.11 This model optimizes capacity, applying a minimum cost function for electricity production under the condition of 100% renewable electricity. Until now, well-known modelling scenarios have underestimated the growth of renewable energy in developed countries and overestimated gas consumption in developing countries. Thus, energy policies have been based on assumptions that have not come true. Installed capacity (and electricity production) for wind and solar had to be adjusted upwards in retrospect for almost every forecasting report. This fact hardly surprises anyone, unless the boom of forecasting and modelling methods is taken into consideration. Friedrichs (2013)12 offers a few reasons for such a development within bodies serving as authorities for estimating energy supply: the International Energy Agency (IEA), the US-Energy Information Administration (EIA), and private entities such as BP and Shell. As for the IEA, the majority of staff have a background in economics and a strong belief in the capability of market forces to reach the optimum balance of supply and demand. Hence, “until 2008 the standard practice of the IEA has been to extrapolate trends in energy demand, and simply to assume that future demand will be met via the market mechanism” (Friedrichs (2013)). Scenario modelling is compared to forecasting in the theoretical section, but not further explored. 2.4Why complexity? Complexity as a term stands for distinct concepts in forecasting, mathematics and information theories; the same statement can also be applied to entropy. Someone considering a career path in science spends years studying one singular concept of complexity or entropy; yet there is no reason that would force them to crosscheck their research object with other disciplines. As a result, the same concept has a different name across various disciplines, or the concepts of the same name mean different things, as will be shown for complexity and entropy. A holistic explanation of these research phenomena is yet to come, as was the case with physics around the end of the 19th century.
energy forecasting.focus:natural gas 23 13 Green, K. C. and Armstrong, J. S. (2015). Simple versus complex forecasting: The evidence. Journal of Business Research,68(8):1678–1685 Forecasting models have become increasingly complex, although accuracy has hardly increased. Still, simplicity repels and complexity lures because "researchers are aware that they can advance their careers by writing in a complex way" (Green and Armstrong (2015)).13 This effort is rewarded by highly-ranked journals that favour complexity over interpretability. However, besides efforts in academia, the primary goal of forecasting is to produce forecasts that support a decision-making process. In the modelling section, models labelled as simple, such as a regression, have been included to check whether the argument by Green and Armstrong (2015) that “most simple methods are more accurate than complex methods” remains valid in this research. 2.5Scope This work explores capabilities of statistical and machine learning forecasting for the energy sector and their relation to complexity. This work also aims to provide methodological and linguistic clarity on forecasting and scenario modelling, especially for the energy sector. The challenge of this thesis was to investigate three unrelated fields of science; understanding the research jargon of institutions such as the School of Management, or the Department of Econometrics and Business Statistics, and combining the acquired knowledge with energy engineering. For this reason, definitions have been included as margin notes.
1Hong, T. (06/30/2014). Global energy forecasting competition. past, present and future. https://forecasters.org/ wp-content/uploads/gravity_forms/ 7-2a51b93047891f1ec3608bdbd77ca58d/ 2014/07/HONG_TAO_ISF2014.pdf Last accessed 2020-12-20 2The first version of this chapter has been published in Grajcar (2019). Grajcar, M. (2019). Energy Forecasting vs Energy Modelling. jazz improvisation vs Symphony. http: //cyseni.com/archives/proceedings/ Proceedings_of_CYSENI_2019.pdf Last accessed 2020-12-20 3Lindley, D. V. (2001). The philosophy of statistics. Journal of the Royal Statistical Society: Series D (The Statistician),49(3):293–337.https: //doi.org/10.1111/1467-9884.00238 4Zhang, W. and Yang, J. (2015). Forecasting natural gas consumption in China by Bayesian Model Averaging. Energy Reports,1:216–220.https://doi. org/10.1016/j.egyr.2015.11.001 3 Theoretical Background 3.1Energy forecasting vs energy modelling Introduction Dr. Tao Hong pointed out four main issues in energy forecasting in his speech on the Global Forecasting Competition in 2014 (Hong (2014))1: impractical research, lack of benchmarking data, a hard-toreproduce process, and limited educational programmes and courses. Here,2this research addresses the first issue - impractical research - since energy forecasters devote their efforts to describing a methodology setup, leaving the interpretation of results aside. Impractical research relates to the broader problem described back in 2000 by David J. Hand in his comment to Lindley (2001)3: "The focus (in statistical journals) seems to be increasingly on narrow technical advance into increasingly specialized areas, with greater merit being awarded to work which is more abstract and more divorced from the realities of data." First, the relationship between energy forecasting and energy modelling is discussed in theory by introducing an analogy with music and by exploring energy models applied to gas forecasting. Most energy models are rooted in econometric theory and statistics, therefore a missing piece - the social sciences - is discussed with regard to the energy model results. Although far from perfect, forecasting models find regular use in policy making and in amending strategies to secure the energy supply of a country or supra-region, as recognised by Zhang and Yang (2015).4Energy forecasting encompasses making predictions for gas and electric load, prices, electricity generation from weather-dependent sources, etc. The term energy determines the use of forecasting
energy forecasting.focus:natural gas 32 19 Sovacool, B. K., Ryan, S. E., Stern, P. C., Janda, K., Rochlin, G., Spreng, D., Pasqualetti, M. J., Wilhite, H., and Lutzenhiser, L. (2015). Integrating social science in energy research. Energy Research & Social Science,6:95–99 20 Jefferson, M. (2016). Energy realities or modelling: Which is more useful in a world of internal contradictions? Energy Research & Social Science,22:1–6.https://doi. org/10.1016/j.erss.2016.08.006 21 Articles of this kind are to be found in Journals such as Energy Research and Social Science, or Energy Policy. 22 Shaikh, F. and Ji, Q. (2016). Forecasting natural gas demand in China: Logistic modelling analysis. International Journal of Electrical Power & Energy Systems,77:25–32 23 DNV-GL (2019). Energy Transition Outlook. Oil and Gas Forecast 2050.https://eto.dnv.com/ 2017/oilgas Last accessed 2021-03-26 24 System dynamics is a branch of systems theory (a way to see the whole as the collection of its interacting parts) "that recognises the role of positive and negative feedback, in which systems can spin out of control, as in virtuous or vicious cycles, and in which systems can be kept within bounds, respectively." (Bale et al. (2015)). Wang et al. (2019) considers DNV GL the only one major energy institute that does not conduct a demand-driven analysis, meaning assuming that gas resources will automatically meet the future demand. 25 Fragkos, P., Kouvaritakis, N., and Capros, P. (2015). Incorporating uncertainty into world energy modelling: the PROMETHEUS model. Environmental Modeling & Assessment,20(5):549–569 Because of the consequences of the status quo shown in figure 3.3,Sovacool et al. (2015)19 suggested a program-centred approach to the energy field instead of the technology-centred approach to encourage interdisciplinary depth. Although this direction is already noticeable in the calls for research projects of national institutions in Germany, energy forecasting and modelling has not been reached by this trend yet. What is needed is shifting the attention from questions such as “how to demonstrate that neural networks are the right tool for prediction of gas demand in Germany in 2030” to “how high will the gas peak demand be in Germany in 2030?” Regarding interdisciplinary aspects, International Energy Relations discuss energy supply security and the reliability of energy models. Jefferson (2016)20 calls the discipline International Political Economy of Energy21 having its roots in the 1970sas a product of the OPEC crisis. Gas crises in Europe in 2009 and 2014 had fewer effects on energy policies in Europe, but initiated interest in energy models’ setup for testing the availability of gas in Europe or the level of dependency on imports to name few. Studies were mostly conducted for vulnerable countries in Central and South-Eastern Europe and for the United Kingdom. Interdisciplinary interdependencies influence the outcome of energy modelling as the following example shows. China’s gas reserves are overestimated due to differences in Chinese and internationally accepted definitions of gas resources (i.e. gas estimates in the ground) and reserves (i.e. gas produced with current prices and technology) (Shaikh and Ji (2016)).22 Changed data of the BP, the US Department of Energy or the IEA serve as inputs in energy models and stimulate discussions of Peak Demand instead of Peak Oil. Some global energy models already included the concept of Peak Demand, e.g. DNV GL already forecasted an oil demand peak for 2023 and gas demand peak for 2035 in their Energy Transition Outlook in 2018 DNV-GL (2019).23 DNV GL characterizes their model as system dynamic modelling of the world energy system.24 Continuation of current technology trends is assumed with one exception being the increased use of hydrogen for energy purposes. Other global energy models follow expectations regarding the use of hydrogen, too. For example Prometheus (in Fragkos et al. (2015))25 identifies 18 hydrogen production technologies in a separate hydrogen module. For the latest list of global gas demand scenarios from IEA, BP, ExxonMobil, Equinor, DNV GL, EIA, and Shell, see Bradshaw and Boersma (2020).
energy forecasting.focus:natural gas 33 26 The base load covers other usages of gas without direct dependence on the temperature, especially in the industry sector. 27 Analysts do recognize shortcomings of their methods when the values violate economic theory (i.e. logic of the models) or common sense. This is usually solved by arbitrary interventions or simply by ignoring these values. 3.1.5Conclusion Most papers follow the same structure: 1. A short reasoning for choosing gas forecasting as the research object, e.g. the rapid growth of gas consumption in a country, or a relevance of the precise forecasting for economic progress; 2. A description of the research methodology used as well as alternative methods; 3. Proof that the chosen methodology outperforms other methods in terms of MAPE and RMSE, or by naming novel qualities of a suggested method (e.g. dynamic, adaptive). There are some peculiarities of forecasting gas demand when compared to oil or electricity demand forecasting. Short-term gas load (consumption) is divided into heat load, dependent mostly on outdoor temperature and base load.26 The relation between heat load and temperature is linear within a certain range of temperatures as shown e.g. in Merkel et al. (2018) for several Midwestern US operating areas or in Franco and Fantozzi (2015) for temperatures below 15◦Cin Italy. Most forecasters treat the whole gas demand as a temperature-dependent output. Research on gas modelling contributes rather to the methodology development27 than to the further knowledge gain for the gas sector. If planners followed forecasts from 2015, China for example would be dealing with a remarkable oversupply of gas right now. However, it is the underestimation of gas consumption that can threaten the economy and the welfare of the population. As for energy security, the impact of research projects is hardly measurable as policy makers form their policies in line with studies assigned by ministries. Research projects become secondary literature for them.
energy forecasting.focus:natural gas 34 28 Training error and its relation to complexity is described in the section on complexity. 29 Lindley, D. V. (2001). The philosophy of statistics. Journal of the Royal Statistical Society: Series D (The Statistician),49(3):293–337.https:// doi.org/10.1111/1467-9884.00238 30 Response variables (statistics) and outputs (machine learning) are considered as synonyms. 31 A model can be understood as "a smooth, low-order polynomial curve fitted to a cloud of points on a plane" as in Zellner et al. (2002). Zellner, A., Keuzenkamp, H. A., and McAleer, M. (2002). Simplicity, inference and modeling: Keeping it sophisticatedly simple. Cambridge University Press, Cambridge and New York. ISBN 0521803616 For parametric models, there is an assumption of uncertainty following a given probability distribution, this is not the case for non-parametric models. 3.2Related work This section reviews literature on gas forecasting with an emphasis on the essential idea of forecasting methods and accuracy criteria. The criteria measure the out-of-sample (or test) error, i.e. the prediction error over an independent test sample. Formally, ydenotes the output (a target variable), Xthe vector of inputs and ˆ f(x)the prediction model from a training data set τ. The loss function is denoted by L(y,ˆ f(X)) (Hastie et al. (2009)). Typical loss functions work with absolute or squared values. For the test error Errτ28 for a fixed training set τwe have Errτ=E[L(y,ˆ f(X))|τ](3.2.1) In simpler terms, a forecast error is defined as Err =y−ˆ f(X)(3.2.2) 3.2.1Forecasting assumptions Lindley (2001)29 stated that: ”a model is merely your reflection of reality and, like probability, it describes neither you nor the world, but only a relationship between you and that world.” This work adds a simplified reflection. All forecasting methods are based on the assumption of underlying relationships among multiple variables, and their discovery assists the estimation of the response variable30 in the future. Moreover, it is assumed that data quality is sufficient and representative of a research object. Additionally, if the model31 predicts accurately in one forecasting period, it will also produce accurate forecasts in further periods. Besides these implicit assumptions, a few more premises set the trend in forecasting and change almost every two decades. In the second half of the 20th century, statisticians believed there could be one superior model that fit almost all forecasting problems. Supposedly, it was just a matter of time before it was found. Regarding complexity, data scientists take clashing positions; either they assume that “more sophisticated models lead to increases in the prediction accuracy” or “simple models predict at least as accurate as complex models and in most cases, they outperform complex models”.The definition of a model sophistication has not been defined in the classical literature on forecasting; non-parametric non-linear models are perceived to be complex. Kaposty et al. (2020) describes more complex methods as models able to take more information into consideration.
energy forecasting.focus:natural gas 35 In the persistence model, forecast equals the latest observation in time series. 32 Kuhn, M. and Johnson, K. (2016). Applied predictive modeling. Springer, New York, Corrected 5th printing edition. ISBN 978-1-4614-6849-3 The forecasting criteria below are mostly used for measuring accuracy of univariate time series forecasts, i.e. forecasts for time series that consists of single scalar observations measured sequentially over equal time steps. Automatic forecasting software may be set to exclude models with results exceeding 50 per cent in MAPE (Katz (2020)). 33 Makridakis, S. (1993). Accuracy measures: theoretical and practical concerns. International Journal of Forecasting, 9(4):527–529.https://doi.org/10. 1016/0169-2070(93)90079-3 3.2.2Accuracy criteria Depending on the task, the accuracy is computed by comparing forecasts to their real values with the aid of commonly used forecast measures as described in the subsection below or - for model selection - to a benchmark model that is usually a persistence model or the mean of previous values. Let rtdenote the relative error, then rt=et/et∗(3.2.3) where etmeans the forecast error of a model compared and et∗the forecast error of the benchmark method (Hyndman and Koehler (2006)). Although overlooked in the literature,Kuhn and Johnson (2016)32 in their textbook Applied Predictive Modelling discuss the upper bound of the accuracy related to the response variable. Systematic error is an inevitable part of any error and includes measurements, summing and reporting final values. As for a self-constructed data set for forecasting gas imports to Germany, imperfections include measurements of amount of gas transported at import nodes, reporting issues and changes in the methodology of calculations. Forecasting criteria a) Mean absolute percentage error (MAPE) is a scale-independent measure used for the comparison of models for forecasting variables with different mean value; with the range of 1-10 (highly accurate), 10-30 (good forecast) and above 50 (inaccurate forecast). MAPE =1 n n ∑ t=1 et yt (3.2.4) where nis the number of data points, ytthe actual value and et the difference between an actual value ytand the predicted value ˆ yt. Since this measure penalises positive errors more than negative ones; Makridakis (1993)33 suggested a symmetric mean absolute percentage error: sMAPE =mean(200|yt−ˆ yt|/(yt+ˆ yt)) (3.2.5) sMAPE criterion can take negative values and it is also not fully symmetric (Hyndman and Koehler (2006)). b) Root mean square error (RMSE) is used in linear models, e.g. for electricity production from wind power plants based on meteorological models, as its minimization (a quadratic problem) can be
energy forecasting.focus:natural gas 36 Gradient descent is a generic approach to minimizing in-sample error R(θ)in neural networks. Neural networks with no hidden layer are actually linear multinomial regression models (Hastie et al. (2009)). 34 Rieck, B. A. (2017). Persistent Homology in Multivariate Data Visualization. PhD thesis, Heidelberg University Library. http://archiv.ub. uni-heidelberg.de/volltextserver/ 22914/1/Dissertation.pdf Last accessed 2021-04-13 35 Browell, J. (2015). Spatio-temporal prediction of wind fields. PhD thesis, University of Strathclyde. http://oleg.lib.strath.ac.uk: 80/R/?func=dbin-jump-full&object_ id=25822 Last accessed 2020-12-20 36 Pelikán, E. (2014). Forecasting of processes in complex systems for realworld problems. Tutorial. Neural Network World,24(6):567–589.http://www. nnw.cz/doi/2014/NNW.2014.24.032.pdf solved by differentiation or iteratively by gradient descent (Browell (2015)). Aggregation and the mean calculation mask the behaviour of various models; scatter plots display the insensitivity of the RMSE (as an aggregate number) as shown in Rieck (2017).34 Besides practice in forecasting, the RMSE is also used as a quality measure for dimensionality reduction methods. RMSE =s1 n n ∑ t=1 e2 t(3.2.6) c) Mean absolute error (MAE) measures the mean absolute difference between the predicted and observed values. The error is representative in situations where “the economic cost of a forecast error is proportional to the magnitude of the error, as opposed to its square” (Browell (2015)).35 However, for this case, no real example from the energy sector was found. The mean absolute error is also used for expressing the forecast error in relation to other errors. MAE =1 n n ∑ t=1|et|(3.2.7) d) Mean square error (MSE), as every error, could be decomposed into a bias term and variance. The bias is the consistent offset of the forecast, the variance represents the variation of the forecast error around its mean. MSE =1 n n ∑ t=1 e2 t(3.2.8) Both, MAE and MSE are generalized by the Minkowski objective function with the exponent R(Pelikán (2014)).36 For R=1 the MAE is computed, for R=2 it is the MSE. All three errors (RMSE, MAE and MSE) are scale-dependent and sensitive to outliers. LM=1 n n ∑ t=1 eR t(3.2.9) e) The coefficient of determination R2measures the proportion of the variance that can be explained by the model. There are many ways to compute R2: one of them, correlation coefficient, measures how well the predicted and real (measured) values are correlated. The most used form of the coefficient is Pearson’s correlation coefficient r.
energy forecasting.focus:natural gas 37 37 Busse, S., Helmholz, P., and Weinmann, M. (2012). Forecasting day ahead spot price movements of natural gas - an analysis of potential influence factors on basis of a NARX neural network. Multikonferenz Wirtschaftsinformatik 2012 - Tagungsband der MKWI 2012.https://publikationsserver. tu-braunschweig.de/servlets/ MCRFileNodeServlet/dbbs_derivate_ 00027726/Beitrag299.pdf Last accessed 2021-04-14 38 Petropoulos, F., Hyndman, R. J., and Bergmeir, C. (2018). Exploring the sources of uncertainty: Why does bagging for time series forecasting work? European Journal of Operational Research,268(2):545–554 39 Data points (cases) with a long distance to all other cases. 40 Ghalehkhondabi, I., Ardjmand, E., Weckman, G. R., and Young, W. A. (2017). An overview of energy demand forecasting methods published in 2005–2015.Energy Systems,8:411–447. 10.1007/s12667-016-0203-y r=n∑XY −(∑X∑Y) p[n∑x2−(∑x2)][n∑y2−(∑y2)] (3.2.10) where ndenotes number of observations in the regression equation, Xthe mean of the independent variable of the regression equation and Ythe mean of the response variable (output). As follows from the formula, the correlation coefficient does not provide any information on systematic overand underpredictions of a model. f) "Hit ratio" reflects the proportion of successful forecasts, defined by the match of their algebraic sign (e.g. a percentage change of gas price compared to its previous value) to the true value. It is rare to find the measure hit ratio in energy forecasting; for example Busse et al. (2012)37 used it for measuring accuracy in gas spot price forecasting. hitratio =f orecasts −errorcount f orecasts (3.2.11) Alternative accuracy measures include the MASE, the scaled variant of the mean absolute error (MAE), proposed by Hyndman and Kochler, where the scaling is equal to the MAE of the seasonal random walk for the in-sample data38 or the overall weighted average (OWA) of chosen accuracy measures. Here, outliers39 are less penalized than in the case of RMSE where the deviation from the actual value is squared. 3.2.3The essential idea of methods Studying forecasting literature of the last 20 years, it is hard to ignore the number of studies devoted to China or Turkey. It is speculated that this is due to their rapidly increasing gas demand due to population increase (Turkey) or high economic expectations in the last decade (China). The paradox lies in the realization that most forecasting methods work best under stable business as usual conditions. Second, the setup of every study with its forecasting models is unique in almost all criteria: a) Training and test data sets in terms of their size and quality. b) Goal of forecasting and various time horizons. Authors of reviews choose arbitrary division lines; Ghalehkhondabi et al. (2017)40 understand short-term forecasting as forecasting up to a month, midterm forecasting up to five years and long-term forecasting from five
energy forecasting.focus:natural gas 38 41 For example, day-ahead electricity price forecasts in Gianfreda et al. (2020). "Reference Scenario" - the continuation of already applied policies, "New Policy Scenario" - reflecting the change in policies. 42 Browell, J. (2015). Spatio-temporal prediction of wind fields. PhD thesis, University of Strathclyde. http://oleg.lib.strath.ac.uk: 80/R/?func=dbin-jump-full&object_ id=25822 Last accessed 2020-12-20 43 Botev, L. and Johnson, P. (2020). Applications of statistical process control in the management of unaccounted for gas. Journal of Natural Gas Science and Engineering,76:103194 to 20 years. Debnath and Mourshed (2018) define a short-term horizon as up to three years, medium-term from three to fifteen years and the long-term as over fifteen years. The short-term forecasts in the energy sector are the matter of interest for energy traders41 and distributors of electricity; therefore, in the electricity load forecasting, the short-term period covers less than one hour and long-term more than one-year. Above 20 years, either a long-term forecast estimate e.g. the return on investment into the technology (new gas turbines, wind power turbines) or the whole energy sector, economy or global economy are modelled by applying the function of the minimal cost of produced and/or distributed energy, published by the IEA, the US EIA and others. Since governments partly control the outcome of demand of energy carriers, the IEA and other organizations publish projections under Scenarios. Geographical coverage affects the choice of a forecasting horizon and required accuracy: the higher the level (district, city, country, global), the higher the tendency to forecast long-term and the lower accuracy is expected given the uncertainty of the input information. In case studies, the length of forecast horizons depends on the characteristics of the object’s field. Table 3.2contains forecast horizons, common temporal resolutions and decisions relevant for wind energy from the dissertation of Browell (2015).42 Two more columns have been added to show if and how forecasts could be applicable for energy networks: district heating (DH) network with the spatial boundary of a city and a gas transmission network with nodes at the countries’ borders. Whereas the wind power plant operators balance their electricity supply in the very-short-term, district heating operators can have up to a few hours to react to changes in demand due to system latency. In large networks at city level, there is a time lag of up to 12 hours for delivering heat from a supplier to the final consumer. Predicting hourly heat demand depends more on behavioral patterns than on buildings’ thermal properties. Similar to district heating networks, gas networks contain linepack, the amount of energy in pipelines, computed with the fixed temperature and the average heating value of gas (Botev and Johnson (2020)).43 c) The choice of variables A variable neglected by one author is of importance in other studies, which is understandable if this work draws a line between in-
Forecast horizon Resolution Decision District heating (DH) networks Prediction: Heat demand for residential sector, based on weather Gas networks Prediction: gas demand Ultra-shortterm <1min, Seconds Wind turbine control Not applicable Not applicable Very-shortterm <1hour 1,5,10,15 min. Balancing, wind farm control, spot markets Not applicable for networks at the city level; a DH network is a shortterm storage. Short-term 1-48 hours 30 min., 1 hour Generation scheduling, day-ahead markets, some spot markets Balancing the supply Balancing the supply Mediumterm 1-10 days 1hour, 3 hours Generation scheduling, maintenance planning Balancing the supply Long-term Monthsyears Daysmonths Maintenance planning, resource assessment/project financing Maintenance planning, fuel purchase Maintenance planning, Strategic planning, Long-term international contracts, LNG contracts. Table 3.2: Forecast horizons, common temporal resolutions, decisions for wind energy (Browell (2015)), district heat and gas networks (M.Greitzer)
energy forecasting.focus:natural gas 40 Models, methods and techniques are treated as synonyms. 44 Šebalj, Dario and Dujak Davor and Mesaric Josip (2017). Predicting natural gas consumption - a literature review. http:// archive.ceciis.foi.hr/app/public/ conferences/2017/08/SPDM-2.pdf 45 Side notes refer to the newest studies of models for gas forecasting. 46 Wang, Z., Li, Y., Feng, Z., and Wen, K. (2019). Natural gas consumption forecasting model based on coal-to-gas project in China. Global Energy Interconnection,2(5):429–435 47 Su, H., Zio, E., Zhang, J., Xu, M., Li, X., and Zhang, Z. (2019). A hybrid hourly natural gas demand forecasting method based on the integration of wavelet transform and enhanced deep-RNN model. Energy,178:585–597 48 Su, H., Zio, E., Zhang, J., Xu, M., Li, X., and Zhang, Z. (2019). A hybrid hourly natural gas demand forecasting method based on the integration of wavelet transform and enhanced deep-RNN model. Energy,178:585–597 49 Szoplik, J. (2015). Forecasting of natural gas consumption with artificial neural networks. Energy,85:208–220 ”The object of statistical methods is the reduction of data. A quantity of data, which usually by its mere bulk is incapable of entering the mind, is to be replaced by relatively few quantities which shall adequately represent the whole, or which, in other words, shall contain as much as possible, ideally the whole, of the relevant information contained in the original data.” in: Fischer, R. A. (1922). On the mathematical foundations of theoretical statistics. Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character,222(594-604):309–368 50 Erdo˘gdu, E. (2010). Natural gas demand in Turkey. Applied Energy, 87(1):211–219.https://doi.org/ 10.1016/j.apenergy.2009.07.006 dustrial and post-industrial countries. The increase of average income/per capita in China could cause increased gas consumption of the country, whereas in Germany, this linear relationship does not exist. In the next section, models are discussed at length. There are no unified criteria for models’ categorization. For comprehensive reviews on models in gas forecasting, see Soldo (2012), Tamba et al. (2018) and Sen et al. (2019). Šebalj, Dario and Dujak Davor and Mesaric Josip (2017)44 reviewed 187 papers on gas consumption forecasting and listed the nine most used methods:45 neural networks, adaptive neuro-fuzzy inference system (ANFIS), grey model, support vector machine (SVM)46, genetic algorithms47, mathematical and statistical models, time series, and hybrid models.48 Szoplik (2015)49 provided an overview of three forecasting methods with a practical meaning for predicting gas consumption: time series methods, regression models, and neural networks. These reviews show that most forecasts in the energy field deal with energy prices (Herrera et al. (2019)), fuel production (Semenychev et al. (2014)) or fuel consumption of natural gas and coal, as both belong to finite energy resources. Tamba et al. (2018) reviewed models up to 2015 and ordered them chronologically from oldest to most recent: • Hubbert model • statistical models (ARIMA, time series models, decomposition models and trend analysis) • regression models • econometric models • AI-expert systems (neural networks) • fuzzy logic • grey prediction models • genetic algorithms • mathematical models • hybrid models • combination or mixed models. Sen et al. (2019) grouped forecasting methodologies by the period of their use into three types: ARIMA modelling, decomposition approaches on a daily basis and heuristic approaches based on economic indicators such as GDP, population, and inflation.
energy forecasting.focus:natural gas 41 51 Makridakis, S., Spiliotis, E., and Assimakopoulos, V. (2018). The M4 competition: Results, findings, conclusion and way forward. International Journal of Forecasting,34(4):802–808 52 Barker, J. (2020). Machine learning in m4: What makes a good unstructured model? International Journal of Forecasting,36(1):150–155 53 Januschowski, T., Gasthaus, J., Wang, Y., Salinas, D., Flunkert, V., BohlkeSchneider, M., and Callot, L. (2020). Criteria for classifying forecasting methods. International Journal of Forecasting, 36(1):167–177 54 Interval and quantile forecasts 55 In German Auftraggeber, usually Ministries and other public institutions. For time series data, Erdo˘gdu (2010)50 listed the following approaches to economic forecasting: a) exponential smoothing methods, b) single-equation regression models, c) simultaneous-equation regression models, d) ARIMA, and e) vector autoregression. Within the forecasting community, various criteria are used to make a distinction between statistical models and machine learning models. In the broad sense, statistical methods are described as “variants of exponential smoothing and ARIMA methods” in Makridakis et al. (2018)51 and machine learning methods are anything else, covering neural networks and random forests. A statistical model can also be understood as a model learning its parameters in one series at a time, whereas the machine learning (ML) model finds the parameters across multiple series. Barker (2020)52 suggests a new division between structured models (such as an autoregressive model) and unstructured models (neural networks) based on the way the data are generated: is the process of generation defined a priori or is it learned from the data? Januschowski et al. (2020)53 formulated objective and subjective dimensions of the classification; the next paragraph shortly summarizes both groups according to their paper, with comments. 3.2.4Objective methods according to Januschowski et al. (2020) Global and local methods. Local methods mean estimating parameters for a model in each time series separately; global methods search across time series. Point vs probabilistic forecasts.54 Point forecasts assign one single value for the amount category of a research object (e.g. wind speed, gas consumption) for a forecasting horizon. Their main advantage is how quickly the information is grasped by the public, though this is without any quantification of forecast uncertainty. Probabilistic forecasts estimate the likelihood of possible outcomes and thus, they are combinations of the two qualities: sharpness and reliability. In gas forecasting, researchers’ results take the form of point forecasts as these are subjects of their contracts with contractors.55 However, interval forecasts would provide higher information value as it is the maximum demand that is relevant for maintaining the security of energy supply. Second, overprediction and underprediction of gas consumption do not have the same negative impact in terms of economic efficiency and security. Therefore, quantile forecasts are more appropriate since they estimate the probability that an observation will exceed the set value.
energy forecasting.focus:natural gas 48 78 Rieck, B. A. (2017). Persistent Homology in Multivariate Data Visualization. PhD thesis, Heidelberg University Library. http://archiv.ub. uni-heidelberg.de/volltextserver/ 22914/1/Dissertation.pdf Last accessed 2021-04-13 Notation on prediction values: f(x) = ˆ y 79 Kuhn, M. and Johnson, K. (2016). Applied predictive modeling. Springer, New York, Corrected 5th printing edition. ISBN 978-1-4614-6849-3 6. Statistical/Linear regression models Linear regression, as all causal models, is based on the assumption that an output responds to the various, up-front chosen variables (predictors). Their number shall be lower than the number of data points and they shall not be correlated. Precisely, given a set of nmeasurements with dattributes, each measurement can be represented as a d-dimensional input vector X. Under the assumption, each vector Xhas an associated scalar property, s∈ R; regression analysis derives a functional relationship between each vector Xand its associated scalar property s(Rieck (2017)).78 f(X) = β0+ p ∑ j=1 Xjˆ βj(3.2.14) The parameter ˆ βis estimated from a training data set (x1,y1)...(xN,yN), by minimizing the sum of squared residuals (a quadratic cost function corresponding to the log-likelihood of the distribution of the error term) (Hastie et al. (2009)): RSS(β) = N ∑ i=1 (yi−f(xi))2(3.2.15) Estimation error refers to ˆ βcoefficients which vary regarding the true regression coefficients. It is assumed the model is built upon the sample of the data and as the sample Nincreases, the estimation error decreases (Tashman (2018)). Thus, it gives more weight to outliers in the data set as these have exponentially large residuals. This can be solved by applying the absolute loss function L(y,f(x)) = |y−f(x)|or Huber loss function (below), which squares ’small’ residuals and takes the absolute value of them when they are larger (Kuhn and Johnson (2016)).79 L(y,ˆ y) = (y−ˆ y)2...|y−ˆ y| ≤ α(3.2.16) |y−ˆ y|...|y−ˆ y|>α(3.2.17) where yis the target variable, ˆ ythe prediction and α∈ R+is a hyperparameter. Multiple linear regression is based on the fact that when all inputs x1,x2,..., xpare orthogonal (uncorrelated), then their multiple
energy forecasting.focus:natural gas 49 80 Beyca, O. F., Ervural, B. C., Tatoglu, E., Ozuyar, P. G., and Zaim, S. (2019). Using machine learning tools for forecasting natural gas consumption in the province of Istanbul. Energy Economics,80:937–949 81 Support vector regression 82 In econometrics, this approach of measuring complexity has been criticized by Keuzenkamp and McAleer (1997). 83 Shaikh, F. and Ji, Q. (2016). Forecasting natural gas demand in China: Logistic modelling analysis. International Journal of Electrical Power & Energy Systems,77:25–32 84 Melikoglu, M. (2013). Vision 2023: Forecasting Turkey’s natural gas demand between 2013 and 2030.Renewable and Sustainable Energy Reviews, 22:393–400 least square estimates ˆ βjare equal to the univariate estimates. That means, the orthogonality of inputs enables parameter estimates to be independent of each other. For inputs x1and x2, the vector x2is regressed on the vector x1, which leaves the residual vector z. Then, yis regressed on z, which gives the coefficient of x2. If some inputs are correlated, the residual vector, representing how much of xpis unexplained by an xk, would be close to zero, which would make the coefficient ˆ βpunstable (Hastie et al. (2009)). Multiple linear regression is used to remove insignificant variables and / or employed to compare and demonstrate performance of other methods, as in Beyca et al. (2019).80 They compared this method with SVR81 and ANN methods for monthly forecasts of gas consumption of up to one year in Istanbul. Similarly to the case study on gas consumption in Germany, Sen et al. (2019) used multiple linear regression to predict Turkey’s yearly gas consumption based on socio-economic variables. The complexity of the linear regression, is measured by degrees of freedom – number of parameters chosen by a model (Kuhn and Johnson (2016)).82 7. Logistic modelling analysis and the logistic-population modelling approach assumes there is a maximum demand related to historic extraction / production of all forms of energy resources, which in turn relates to the availability and depletion of the raw materials (Shaikh and Ji (2016)).83 The following parameters are used: Dmax - the maximum gas demand a country is expected to achieve in the long term, αgrowth parameter, tmax - time in years when half of the Dmax or the Dmax/capita occur. Melikoglu (2013)84 used the logistic equation for forecasting gas demand in Turkey between 2013 and 2030. 8. Random forest is a collection of tree predictors f(x,T,Θk),k= 1,2,...K)where the Θkare independent and identically distributed random vectors. This method has its roots in regression trees and it is constructed by portioning a data set sequentially along values of the explanatory variables (Kaposty et al. (2020)). The stop criterion, i.e. the penalty method, causes the algorithm stop to prevent overfitting. The performance of the model would slightly deteriorate by including extra predictors, as there is a higher chance that a model randomly uses unimportant predictors for splitting (Kuhn and Johnson (2016)).
energy forecasting.focus:natural gas 50 85 Breiman, L. (2001). Random Forests. Machine Learning,45(1):5–32.https: //doi.org/10.1023/A:1010933404324 86 Frequentist methods use the error rate form assessing the quality of the model; models based on Bayesian methods use coherence. 87 Franco, A. and Fantozzi, F. (2015). Analysis and clustering of natural gas consumption data for thermal energy use forecasting. Journal of Physics: Conference Series,655:012020 Šebalj, Dario and Dujak Davor and Mesaric Josip (2017) concluded neural networks are currently the most commonly used models in gas forecasting. Breiman (2001)85 dealt with the problem of overfitting by proposing the mean prediction of many different trees. The method is also used for introducing nonlinearity in the data. The problem of overfitting is typical for frequentist86 methods in contrast to Bayesian approach. In the latter, "the maximization over a subset cannot exceed that over the full set" (Lindley (2001)). Least squares as the method used for fitting "...is equivalent to a Bayesian argument using an improper prio, namely a uniform distribution over the space of the regression parameters" (Lindley (2001)). 9. Temperature correlation models, also called heating degree methods, are used for residential demand forecasts only as the idea lies in the strong negative correlation between the daily outside temperature and gas consumption. Franco and Fantozzi (2015)87 suggested forecasting models for residential gas consumption in the winter as follows: • an additive model for a total consumption Ct C(t) = CN+CW+Cs+Cr(3.2.18) where CNrepresents the standardized load shapes of production, CWthe weather sensitive component, CSa random term. • a multiplicative model C(t) = CN×f(w)×f(s)×f(r)(3.2.19) where CNis the base consumption and f s are correction factors for current weather, special events and fluctuation, respectively. • a model combining an additive and a multiplicative model C(t) = F(d(t)×f(w(t))) + R(t)(3.2.20) where C(t)is the current consumption at time t,d(t)is the day of the week, F(d)is the daily component, w(t)is a function of the weather data (temperature, humidity and wind chill), f(w)is a weather function and R(t)a term of correction. The accuracy of gas forecasting in temperature correlation models depends on the quality of weather forecasts. 10. Artificial Neural network models are non-linear, non-parametric models enabling the forecasting of any subject without knowledge of the specific relationships between variables. Neural networks can search for their parameters locally and globally (see the distinction in Januschowski et al. (2020)) and almost all of them work without any non-convex loss function.
energy forecasting.focus:natural gas 51 88 Werbos, P. J. (1988). Generalization of backpropagation with application to a recurrent gas market model. Neural Networks,1(4):339–356 89 Single-hidden layer neural networks are an example of the learning method with additive expansion, more details on additive models in Hastie et al. (2009), page 341. 90 Livni, R., Shalev-Shwartz, S., and Shamir, O. (2014). On the computational efficiency of training neural networks. https://arxiv.org/pdf/1410.1141 Last accessed 2020-12-20 Polynomial networks are networks using the squared activation function σ2(x) = x2. In unsupervised learning, there is a set of Nobservations (x1,x2, ...xN)of a random p-vector Xhaving joint density Pr(X). Properties of this probability density are deduced without providing any targeted output data for each observation (Hastie et al. (2009). A learning is understood as a density estimation problem if we suppose that (X,Y)are random variables represented by joint probability Pr(X,Y)). The issue of "adapting weights" in a neural network can be seen as a special case for estimating parameters of any other functional model (Werbos (1988)).88 Concretely, a coefficient βjk is the effect of the jth predictor on the kth hidden unit (Hastie et al. (2009)). In the first step, data normalization is required as normalization causes the higher population of data for the same manifold space. Activation function links the data in predictors with the first hidden layer.89 There are several commonly used activation functions, the ReLU (rectified linear unit) σ(a) = max{0, a}is currently the most popular, while the sigmoid function σ(a) = 1 1+eahas been used in the more traditional approach. Both functions are advantageous over the threshold activation (e.g. σ(z) = 1 if z>0 and 0 if otherwise) as they can be trained using gradient based methods (Livni et al. (2014)).90 Stochastic gradient descent involves random shuffling of the training data set before each iteration; this enables different orders of updates to the model parameters. The linear activation function is seen as problematic; the output of the second layer is just a different linear function of the first layer as was shown in the appendix of Barker (2020). In supervised (structured)learning (previous models), input XT= (X1,..., Xp)and output vectors Y= (Y1,...,Ym)are specified and the network tries to minimize an error, characterized by a loss function L(y,ˆ y), for a known set of target values (answer ˆ yi for each xi); in unsupervised learning, targeted output data is not required and this increases degrees of freedom, which can result in overfitting. As Barker (2020) puts it; “the search space of the model increases exponentially with the degrees of freedom of the model, but the number of points to fit only increases linearly with the size of the data set.” The number of parameters pis counted according to the formula: p=H(P+1) + 3+1 (3.2.21) where Hdenotes the number of hidden layers and Pthe number of predictors (inputs). Forecasting is an extrapolation problem, and unstructured models are optimal for interpolation, as the relationship between forecasts and its lags is more flexible (Barker (2020)). In unstructured models, the curse of dimensionality is solved by assuming
energy forecasting.focus:natural gas 52 91 Rieck, B. A. (2017). Persistent Homology in Multivariate Data Visualization. PhD thesis, Heidelberg University Library. http://archiv.ub. uni-heidelberg.de/volltextserver/ 22914/1/Dissertation.pdf Last accessed 2021-04-13 92 Sources: Random Pixel Generator: http://pixelmonkeys. org/, with a pixel size of 4. The left side of the image, "marlene dietrich top hat, for morocco" by carbonated, is licensed with CC BY-NC-SA 2.0. To view a copy of this license, visit https: //creativecommons.org/licenses/ by-nc-sa/2.0/. The online tool used for the pixelated version: https: //onlinepngtools.com/pixelate-png. 93 Günther, F. and Fritsch, S. (2010). Neuralnet: Training of neural networks. R Journal,2(1):30–39.https: //journal.r-project.org/archive/ 2010/RJ-2010-006/RJ-2010-006. pdf Last accessed 2020-12-20 the manifold, hypothesis, i.e. data in a data set are discrete samples of a continuous manifold of some dimension. The manifold hypothesis can also be seen as one of the tools for reducing the complexity. There is no single algorithm for verifying the hypothesis, but manifolds enable a smooth mathematical structure, especially while analyzing natural phenomena (Rieck (2017)).91 In the best possible case, the manifold where the data lie should be as low-dimensional as possible and as densely populated as possible (Barker (2020)). Any low-dimensional entity can be understood as a subset of the higher dimension; this is demonstrated in the two pixeled pictures of identical size. The right side of the merged figure 3.592 depicts the result of using a random pixel generator with a pixel size of 4, the left side shows one possible subset out of many. The left configuration of colours of pixels of the same size is typical for what is identified as “Marlene Dietrich” and any picture of Marlene Dietrich’s portrait shows similarities in the configuration of colours of pixels with the picture on the left. These configurations evoke associations with the Boltzmann’s formula for entropy S S=kBlogW (3.2.22) where kB=1.3807x10−23J/Kdenotes Boltzmann’s constant, the conversion factor between units of temperature and units of energy and Wthe number of real microstates corresponding to the gas’s macrostate. The right part of the figure depicts microstates of presumably equally probable assigning of any colour from the range white-black to a pixel. The left figure does not fulfill this condition; there is a smaller choice of assigning colors to a pixel for producing a picture resembling Marlene Dietrich, and thus it is a subset of the picture on the right. Manifold hypothesis heuristically explains why machine learning techniques work. A model needs to focus on a few key features in a data set to make decisions. The task shall be as specific as possible with lots of data enabling to find these features. In the backpropagation algorithm, the cost function tries to minimize the error, i.e. the sum of the squared residuals by back propagating to the hidden layer, and weights are either increased or decreased until the desired output is achieved. This was the first algorithm that allowed the adaptation of all weights of a neural network. As shown in figure 3.6, a negative partial derivative will increase the weight and a positive partial derivative will decrease it until the local minimum is found (Günther and Fritsch (2010)).93
energy forecasting.focus:natural gas 53 Figure 3.5: Visualisation of relation of low dimension to high dimension. Figure 3.6: Basic idea of the backpropagation algorithm illustrated for a univariate error function E(w) as in Günther and Fritsch (2010). Learning rate ηkshould decrease to zero, as the iteration rapproaches infinity. Weight backtracking adds a smaller value to the weight in the next step. The technique prevents the algorithm from jumping over the minimum by undoing the last iteration (Günther and Fritsch (2010)). 94 Raza, M. Q. and Khosravi, A. (2015). A review on artificial intelligence based load demand forecasting techniques for smart grid and buildings. Renewable and Sustainable Energy Reviews,50:1352–1372 In contrast to classical back-propagation, resilient back propagation enables the change of the learning rate ηkduring the process of searching for the minimum. In shallow areas, for speeding up the convergence, the learning rate ηkwill increase "if the corresponding partial derivative keeps its sign" (Günther and Fritsch (2010)), (see the second formula with the sign below). Changing the sign will cause the learning rate ηkto slow down as it means that the minimum has been missed. For comparison, • the rule for adjusting weights in classical backpropagation: wk(t+1)=wkt−η∂E(t) ∂wk(t)(3.2.23) where tindexes the iteration steps and kthe weights. • the rule for adjusting weights in resilient backpropagation: wk(t+1)=wkt−ηsign ∂E(t) ∂wk(t)!(3.2.24) The neuralnet package, tested in the modelling section, uses classical propagation as described here, resilient backpropagation with or without weight backtracking and the modified globally convergent version. Raza and Khosravi (2015)94 identified four learning issues related to this algorithm: getting trapped in local minima instead of the global minimum i.e. there could be another set of parameters uniformly better, network paralysis, temporal instability and lack of generalization of the network, resulting in overfitting, characterized as unintended memorization of synaptic weight values. To avoid overfitting, several approaches have been developed, such as early
energy forecasting.focus:natural gas 54 Figure 3.7: An example of how to separate training data in a twodimensional space. The support vectors define the margin of largest separation between the two classes. Source: Larhmam (https:// commons.wikimedia.org/wiki/File: SVM_margin.png), „SVM margin“, https://creativecommons.org/ licenses/by-sa/4.0/legalcode 95 Cortes, C. and Vapnik, V. (1995). Support-vector networks. Machine learning,20(3):273–297 stopping (the procedure will stop before approaching the global minimum; when an error estimate starts to increase) or weight decay – adding a penalty to the error function, with λas its tuning parameter. For λestimation, the cross-validation is used (Hastie et al. (2009)). For short-term load forecasting, hybrid models have been proposed such as ANN with fuzzy and genetic algorithm, ANN with wavelet and time series, ANN with genetic algorithms etc. Examples are to be found in Raza and Khosravi (2015). 11. Support vector machines (SVMs) Originally, Cortes and Vapnik (1995)95 developed SVMs for data classification; the main idea is depicted in figure 3.7. To make boundaries more flexible, the feature space is enlarged by using polynomials (expansions). The optimal hyperplane is denoted as: w0·z+b0=0 (3.2.25) The weights w0could be written as the linear combination of support vectors: w0=∑ support vectors αizi(3.2.26) where αiare weights of the output units and ziweights from a hidden unit. As it can be seen from a formula, new samples enter the model as the sum of inner products so the formula can be rewritten with the kernel function K(.). Various kernel functions can be chosen: e.g. linear, polynomial, or the radial basis function. The tuning parameter sigma impacts the smoothness of the decision boundary (Kuhn and Johnson (2016)); its underestimation would cause the boundary to be too sensitive to noise. The cost parameter, (i.e. setting the price for misclassified samples in the training set), is attached to residuals, not to the parameters. It can be manually set or determined by using cross-validation. With large cost parameters the model becomes flexible and likely overfits. Thus, this parameter can be understood as a measure of complexity for SVM. SVMs are one of the methods for performing regression analysis in machine learning; the support vector regression (SVR) is a (linear) regression formulation of the support vector machines. Considering the linear regression model (Hastie et al. (2009)):
energy forecasting.focus:natural gas 55 96 The indication of the accuracy risk is provided by the variance; “how likely a very poor forecast is for a given series” (Lichtendahl and Winkler (2020)). 97 In this section, we do not explore questions 1 and 2 further. Question 3 is discussed in the section on complexity. 98 This method was preferably used in the M4 Competition, organized by the International Institute of Forecasters. 99 Uncertainty types are not of the same magnitude; the effect of parameter uncertainty seems to be of a smaller order of magnitude than that due to other sources (Chatfield (2001)). f(x) = xTβ+β0, (3.2.27) βis estimated by considering minimization of H(β,β0) = N ∑ i=1 V(yi−f(xi)) + λ 2kβk2(3.2.28) where Ve(r)=0 if |r|<e, and Ve(r)=|r|−e, otherwise. From the formula above, it is clear that data points with residuals within threshold eare ignored in the regression fit, analogically to the SVM in the classification problem, where points on the right side of the division and far away form it are also ignored. 12. Hybrid models link causal methods (i.e. input variables are assumed to affect the output) and data-based models. They are composed of structured and unstructured methods, e.g. using machine learning for forecasting the trend component of the decomposition only and are regarded as the next and further developed form than ensemble models. However, the choice of models to be hybridized remains arbitrary. Since they have become a popular method, the forecasting community raised the following questions: • Under which conditions does a model combination perform better? • Would decision makers accept accuracy risk96 (a higher risk of poor forecasts) while combining models? • Do more complex models gain outperform simple models when the criterion of the complexity is measured by computing time?97 de Nadai and van Someren (2015) tested the combination of the ARIMA and ANN model to detect anomalies in gas consumption. 13. Ensemble of models98 One single model with a single benchmark does not address model, data, or parameter uncertainty.99 While combining models, there is a chance a new model will provide information that has not been caught by other models. Thus, a combination of models keeps the bias about the same but decreases the variance (Atiya (2020)). While constructing an ensemble, a practitioner answers two questions:
energy forecasting.focus:natural gas 56 100 Lichtendahl, K. C. and Winkler, R. L. (2020). Why do some combinations perform better than others? International Journal of Forecasting, 36(1):142–149.https://doi.org/ 10.1016/j.ijforecast.2019.03.027 101 Gaba, A., Tsetlin, I., and Winkler, R. L. (2017). Combining interval forecasts. Decision Analysis,14(1):1–20 102 Among forecasting methods, diversity is measured in terms of correlations among their forecasting errors (Lichtendahl and Winkler (2020)). Computable general equilibrium (CGE) models 103 Bale, C. S., Varga, L., and Foxon, T. J. (2015). Energy and Complexity: New ways forward. Applied Energy,138:150–159.https://doi.org/ 10.1016/j.apenergy.2014.10.057 • Which and how many forecasting models are to be combined. A method with poor forecasts contributes to the diversity of the pool of models and therefore it shall not be left out. On the other hand, one of two poor performing methods with highly positively correlated errors can be excluded due to no gain in diversity (Lichtendahl and Winkler (2020)).100 • Which method to use for their combination and how to choose the best weights. Models are aggregated by using mean (the preferred method), median, mode, or trimmed mean. Gaba et al. (2017)101 suggests combining between five and ten forecasts for optimal results. The greater diversity in the pool, i.e. with minor direct relation,102 the better results. In an analogy to complex systems, a practitioner acquires more individual predictions (prediction trajectories) by suitable perturbation of the initial state of the system, to which complex systems are sensitive (Pelikán (2014)). In a similar fashion, various models and their functions fit the data and produce forecasts ("prediction trajectories"). 3.2.6Excluded models Equilibriumand linear optimization models Equilibrium models (e.g. MARKAL, IKARUS, TIMES) assess technology options based on the simple decision rule of cost optimisation (Bale et al. (2015)).103 In these models, agents (e.g. households, energy trading companies) are assumed to act as rational economic actors with the ability of perfect foresight. Scenarios are unique for each case; only their full description captures the essence. The next example illustrates the point. Eser et al. (2019) and Gillessen et al. (2019) share similar objects and time frames in their studies: the impact of the Nord Stream 2pipeline and liquefied natural gas (LNG) on gas trade, the security of supply up to 2030 and infrastructure expansion. However, whereas Eser et al. (2019) assume gas from production countries is annually and hourly limited, Gillessen et al. (2019) examines energy security by modelling the gas flow disruption in one of the transmission pipelines and models the gas flows and velocities of the remaining parts of the transmission system. The study by Eser et al. (2019) simulates a high-pressure gas network with 500 network nodes (i.e. points where the gas flow is measured), and 150 compression stations and gas storage sites within the borders of Germany to identify bottlenecks in the gas system. Input on
energy forecasting.focus:natural gas 57 104 Eser, P., Chokani, N., and Abhari, R. (2019). Impact of Nord Stream 2 and LNG on gas trade and security of supply in the European gas network of 2030.Applied Energy,238:816–830. https://doi.org/10.1016/j.apenergy. 2019.01.068 105 The US Energy Information Administration. 106 Ang, B. W., Choong, W. L., and Ng, T. S. (2015). Energy security: Definitions, dimensions and indexes. Renewable and Sustainable Energy Reviews, 42:1077–1093 107 Bezdek, R. and Wendling, R. (2002). A half century of long-range energy forecasts: Errors made, lessons learned, and implications for forecasting. Journal of Fusion Energy,21.https://doi.org/ 10.1023/A:1026208113925 108 Winebrake, J. J. and Sakva, D. (2006). An evaluation of errors in US energy forecasts: 1982–2003.Energy Policy, 34(18):3475–3483 109 Magdowski and Kaltschmitt (2017) analyzed day-ahead power forecasting from wind turbines in Germany and came to the same conclusion. Moreover, the prediction accuracy depends on the predicted feed-in volume, with small feed-in volumes decreasing the accuracy. In forecasting, statistic properties of data are checked to choose proper models for producing forecasts, and so assumptions on the future development of an energy system are not needed. hourly imports into Germany is produced by using a Monte Carlo approach: “30 optimizations of the gas sourcing, each with stochastically varied gas prices at the boundaries of the network, are solved.”104 Thus, gas imports represent an optimized input, meaning the leastcost gas imports’ mix from probable production countries and world LNG price, and assuming knowledge of events during the simulated year. Authors validate the novel simulation at the annual level although hourly values were produced first. The IEA and the EIA105 work with macroeconomic models with an outcome in the form of scenarios. The IEA applies the World Energy Model (WEM), a partial equilibrium simulation model covering global energy supply, transformed energy and its demand. The energy demand / supply equilibrium is computed for the minimal total cost of providing energy services. World Energy Outlook scenarios are taken as inputs for other models, such as projections on the energy security (Ang et al. (2015)).106 The Annual Energy Outlook of the EIA is produced by an energy-economic model of the US energy system, the ”National Energy Modeling System,” based on the two main drivers of GDP and energy intensity (Bezdek and Wendling (2002)).107 Retrospectively,large deviations from the forecasts and following wrong policy assumptions have been observed. At first sight, this is not obvious as low errors for total energy consumption conceal much larger errors in sectors that offset each other when aggregated (Winebrake and Sakva (2006)).108 They confirmed intuitive assumptions: a) forecasts exhibit increased uncertainty when time horizons are lengthened109; and b) certain sectors (e.g. residential) demonstrate a more accurate level of forecasting than others (e.g. transport). Energy-sector specific errors include • underestimation of oil and gas production, • projecting exhaustion of energy resources. Prediction of oil peak in the time horizon of 10-15 years from the year of estimation, • overestimation of energy consumption, • an assumption that "technically feasible technologies or technologies feasible in an engineering sense will penetrate the market in the future"(Bezdek and Wendling (2002)). The private oil and gas companies Shell, ExxonMobil, Statoil, and the BP (BP (2019)) publish their own projections. Outcomes of all equilibrium models do not equal forecasts as no value is assigned
energy forecasting.focus:natural gas 64 Relative measure for comparing data sets 124 Pincus, S. M. (1991). Approximate entropy as a measure of system complexity. Proceedings of the National Academy of Sciences of the United States of America,88(6):2297–2301 The tolerance is set to the inner product of rDo and standard deviation of the data set to allow comparison on data sets with different amplitudes. Usually, data are normalized to have standard deviation equal to one (Richman and Moorman (2000)). 125 As Pincus (2008) puts it: "“Descriptively, ApEn thus measures the logarithmic likelihood that runs of patterns that are close (within r) for m contiguous observations remain close (within the same tolerance width r) on next incremental comparisons. Opposing extremes are perfectly regular sequences, (e.g. sinusoidal behavior, very low ApEn), and independent sequential processes (very large ApEn)." 126 Richman, J. S. and Moorman, J. R. (2000). Physiological time-series analysis using approximate entropy and sample entropy. American Journal of Physiology. Heart and circulatory physiology,278(6):H2039–49 127 Chen, J. (2002). An entropy theory of value. SSRN Electronic Journal.http: //dx.doi.org/10.2139/ssrn.307442 modified by Grassberger and Procaccia in 1983 to be able to calculate such a rate from time series data. 3.3.2Approximate Entropy and Sample Entropy Approximate Entropy (ApEn) quantifies the regularity in a data set with at least 1,000 data points; the larger value of ApEn denotes greater randomness, and a smaller value corresponds to more cases of patterns in the data (Pincus (1991)).124 Given the data set of Ndata points, "ApEn(m,r,N)is approximately equal to the negative average natural logarithm of the conditional probability that two sequences that are similar for mpoints remain similar (within a tolerance r) at the next point" (Ramesh (2011)) whereas •mis the length of sequences compared and •ris a tolerance window (that is, tolerance for accepting matches). As Pincus (2008) points out, we can imagine “partitioning the state space into uniform width boxes of width r, from which we estimate mth order conditional probabilities. ApEn is unaffected by noise of magnitude below this filter level rand is finite for both stochastic and deterministic processes, in contrast to K-S entropy.125 However, ApEn lacks relative consistency and depends on the length of time series. Sample Entropy (Richman and Moorman (2000))126 is the modification of the approximate entropy as it searches for any repeated patterns of various lengths but does not count self-matches. This was justified by understanding entropy as the rate of information production (Chen (2002)127) and in this context, comparing data with themselves does not make sense (Richman and Moorman (2000)). If Bdenotes the total number of matches of length mand Ais the total number of forward matches of length m+1, then with the notation: A/B= [Am(r)]/[Bm(r)] (3.3.2) Sample Entropy can be expressed as SampEn(m,r,N) = −ln(A/B)(3.3.3) Approximate Entropy and Sample entropy increase as a data set increases. The Sample and Approximate Entropy was computed with the length of reconstructed vector 1and 2(for chaotic processes) and raccording to the recommended formula r=0.2 ·sd(time series) for gas import
energy forecasting.focus:natural gas 65 Ghandar et al. (2016) measured the model complexity in two aspects: a) the number of rules within a rule-base, and 2) the number of inputs utilized within the rule-base. They concluded “the use of overly complex models leads to a greater reduction in performance in a dynamic environment with timevariant relationships than in a static environment where causal relationships do not change substantially.” 128 Hyndman, R. J. and Khandakar, Y. (2008). Automatic time series forecasting: The forecast package for r. Journal of Statistical Software,27(1):1–22 values for Germany, from the data set used for forecasting gas imports, with the results displayed in table 3.4. A low value of entropy indicates a time series is either deterministic and easier to predict; a higher value indicates randomness. Sample Entropy does not depend that much on the length of the time series. Table 3.4: ApEn and Sample Entropy for the data set "Gas imports" Embedding dimension Approximate Entropy Sample Entropy 1 1.635595 - 2 1.037104 1.884978 3.3.3Complexity of forecasting models General associations of complexity in the literature on predictive modelling In statistical learning, complexity of methods is understood as the number of degrees of freedom (also referred to as the number of free parameters in the model), "difficulty in estimation," or as the number of variables. It is specified for every forecasting method. In the section Related work, typical accuracy criteria for measuring out-ofthe sample errors are listed; here we discuss training errors and their complexity. Again, let Ydenote the output (a target variable), Xvector of inputs and ˆ f(X)the prediction model from a training data set τ. The loss function is denoted by L(Y,ˆ f(X)) as in Hastie et al. (2009). The training error is the average loss of the training sample (Hastie et al. (2009)): err =1 N( N ∑ i=1 i=L(yi,ˆ f(xi)) (3.3.4) With higher number of degrees of freedom (for non-parametric methods), the training error decreases as the function of a model describes the data fully, but this overfitting issue results in poor performance of a model for the test-sample. Methods marked as stable are supposed to have the same in-sample and out-of-sample errors. The Akaike Information Criterion is a penalized method based on the in-sample fit (Hyndman and Khandakar (2008)).128 AIC =L∗(ˆ θ,ˆ x0) + 2p(3.3.5)
energy forecasting.focus:natural gas 66 129 Green, K. C. and Armstrong, J. S. (2015). Simple versus complex forecasting: The evidence. Journal of Business Research,68(8):1678–1685 Although it is not the goal of this work to prove that in every case higher complexity will not improve the accuracy of the model, the research in other fields does go in this direction, e.g. AndradeCabrera et al. (2018) found they can reduce the model complexity, defined through computational tractability, by half for Urban Building Energy Modelling (UBEM) and still retain annual energy estimation errors below 10 per cent for a single building. 130 Hyndman, R. J. and Kostenko Andrey V. (2007). Minimum sample size requirements for seasonal forecasting models. Foresight, (6):12–15.https://robjhyndman. com/papers/shortseasonal.pdf 131 The number of layers of the network characterizes the depth of the network. The total number of neurons represents the size of the network (Livni et al. (2014)). where pis the number of parameters in θplus the number of free states in x0, and ˆ θand ˆ x0are the estimates of θand x0. Since the AIC is based on likelihood, the criterion enables the choice of identical point forecasts from two models (Hyndman and Khandakar (2008)). Green and Armstrong (2015)129 defined simplicity as “processes that are understandable to forecast users.” Users, not researchers. In this respect, naïve or no-change models without seasonal adjustment, seasonal adjustment, single-exponential smoothing, Holt’s exponential smoothing, dampened exponential smoothing and simple average of the exponential smoothing forecasts are regarded as simple. They conclude that “complexity beyond the sophisticatedly simple fails to improve accuracy in all but 16 of the 97 comparisons in 32 papers that provide evidence”(Green and Armstrong (2015)). Hyndman and Kostenko Andrey V. (2007) point out the misconception that ARIMA models are more complex than Holt-Winters models and thus need a larger data set. In their article130 it is shown that ARIMA models actually need the minimum number of 16 observations for estimating a seasonal ARIMA model (monthly data), while for the Holt-Winters model it would be 17 observations. With these numbers of observations, prediction intervals would still be finite. For support vector machines, while developing the algorithm, the order of operations has been interchanged. To start, two vectors are compared in the input space and a non-linear transformation takes place while computing the value of the result. Thus, the complexity of the model does not depend on the dimensionality of the feature space, but on the number of support vectors (Cortes and Vapnik (1995)). Changing the parameter Cenables the trade-off between the complexity of decision rule and frequency of error. As for neural networks, let the sample complexity be defined by the number of examples required to learn the class, i.e. the set of all prediction rules obtained by using the same network architecture131 while changing the weights of the network. Until now, theoretical work on neural networks has not been promising, still in practice neural networks yield results thanks to a few tricks used. Livni et al. (2014) listed among them the change of the activation function, overspecified networks and the regularization of weights to speed up the convergence. Changing the activation function to the squared function σ2(x) = x2makes sample complexity grow (the increase of sample complexity is caused by increasing the depth of the network).
energy forecasting.focus:natural gas 67 Besides the natural complexity of processes behind forecasting objects, complexity increases with re-defining of the object; in demand forecasting in general, one may include returns of a product. However, this is not applicable in energy forecasting. Water flow forecasting to optimize electricity generation 132 Tongal, H. and Berndtsson, R. (2017). Impact of complexity on daily and multi-step forecasting of streamflow with chaotic, stochastic, and blackbox models. Stochastic Environmental Research and Risk Assessment,31(3):661– 682 133 Their work defines complexity as a measure of the number of degrees of freedom of a system active locally around a given instantaneous state. 134 Akpinar, M. and Yumusak, N. (2016). Year ahead demand forecast of city natural gas using seasonal time series methods. Energies,9(9):727 135 Potoˇcnik, P., Soldo, B., Šimunovi´c, G., Šari´c, T., Jeromen, A., and Govekar, E. (2014). Comparison of static and adaptive models for short-term residential natural gas forecasting in Croatia. Applied Energy,129:94–103 136 Yu, F. and Xu, X. (2014). A short-term load forecasting model of natural gas based on optimized genetic algorithm and improved BP neural network. Applied Energy,134:102–113.https: //doi.org/10.1016/j.apenergy.2014. 07.104 137 Stock prices are often modelled as a random walk averaged over the number of stocks. By this, fluctuations are reduced. 3.3.4Complexity of a phenomenon in relation to modelling If the process representing the forecast phenomenon is of dynamic (complex) nature, the complexity of this process may influence the choice of a forecasting model. A complex system is a system that consists of many mutually interacting components, and that shows emergent behaviour, i.e. the collective behaviour evidences some traits that cannot be easily derived or explained based on behaviour of individual parts (Pelikán (2014)). Forecasting electricity production from hydro power plants tends to design more complex models to capture complex processes such as precipitation. On the other hand, the lower complexity of river stream flows leads to higher predictability as shown in Tongal and Berndtsson (2017).132 Here, the complexity is measured on the data set as data (i.e. the time series) represent a sample of the phenomenon in question. Tongal and Berndtsson (2017) conclude determining the degree of complexity enables the pre-determination of a suitable model. It is implied artificial neural networks may be a better choice for high complex systems than deterministic models, especially for a one-step forecasting horizon. World weather models are also ranked in terms of their complexity degree, albeit without any clear definition of how to measure the degree of complexity, as in (Scher and Messori (2019)).133 It is unclear whether a model with higher resolutions with fewer components could be assigned a higher complexity degree than a low resolution model with a larger number of processes. Regarding accuracy,combination of models and complexity, the relation between the increasing complexity of forecasting models and accuracy remains ambiguous. Akpinar and Yumusak (2016)134 concluded that with the computation complexity accuracy rates increase. On the other hand, (Potoˇcnik et al. (2014)135 derived that “among the adaptive models, the nonlinear models did not surpass the performance of the adaptive linear models.” Yu and Xu (2014)136 addressed the issue of complexity by observing that “over the years, studies have shown that a combinative model gives better projected results compared to a single model for natural gas prediction” and saw the future of forecasting in hybrid models. Szoplik (2015) shared this vision. However, it is open to discussion whether combinations of models do increase their complexity in each case. Dimitriadou et al. (2018) concluded in their review of oil price forecasting that “machine learning methodologies produce a higher forecasting accuracy in comparison to the typical econometric ones and they typically outperform the random walk (RW) model,137 while econometric approaches often
energy forecasting.focus:natural gas 68 138 Debnath, K. B. and Mourshed, M. (2018). Forecasting methods in energy planning models. Renewable and Sustainable Energy Reviews,88:297–325.https://doi. org/10.1016/j.rser.2018.02.002 fail to do so.” Still, the latest M4competition, organized by the International Institute of Forecasters in 2018 (Makridakis et al. (2020)), did not confirm Dimitriadou’s conclusions. Debnath and Mourshed (2018)138 state “According to reviewed literature, NN (neural network) structure with two hidden layers produced best results for the monthly load forecasting, the peak load forecasting and the daily total load forecasting modules.” However, there is no reasonably explainable connection between the number of hidden layers of the ANN network and the complexity of gas forecasting. Yalcinoz and Eminoglu (2005) state that 1) there is a relation between the correct number of hidden layers, domain (e.g. monthly load forecasting) and the accuracy, or 2) without any speculation of reasons behind it, until now two hidden layers produced the most accurate results for these domains. The short list of self-contradicting statements points out that without holistic consideration of a phenomenon to be forecasted, a data set as a sample of the reality and the complexity of the model, any statement is valid only for the study case used. The current state of the theoretical research is far from being able to generalize.
1Interestingly, LNG gas is not a direct competitor to pipeline gas. Building the entire LNG chain results in further pipeline construction (Bridge and Bradshaw (2017)). 2International Energy Agency (2021). Energy security. https://www.iea.org/ topics/energy-security Last accessed 2021-01-09 3Biresselioglu, M. E., Yelkenci, T., and Oz, I. O. (2015). Investigating the natural gas supply security: A new perspective. Energy,80:168–176 4 Study 1: German gas imports Problem definition The Science Direct search engine found 256 articles on gas demand forecasting (as of March 2020) and three articles on gas imports in relation to either China or India. Gas imports matter to foreign energy policy of any country with a strong industry sector (such as Germany or China), and for the strategy on pipeline/LNG1infrastructure development. This chapter underlines the features of gas imports in comparison to gas demand forecasts, and presents results of models (e.g. Holt-Winters filtering, regression analysis, and ANN) for monthly gas imports based on a self-constructed data set. Moreover, implications for energy engineering and IR are discussed. The International Energy Agency defines energy security as "the uninterrupted availability of energy sources at an affordable price" (International Energy Agency (2021))2; security of energy supply is a more specific term stressing the reliability of gas supply from production countries, routes, transit countries, and the infrastructure of an importing country. When imagining foreign policy as a four-dimensional idea (i.e. a country with time represented by stakeholders thinking and acting in the foreign policy field), the issue of energy security is one element competing with other issues in this space: combating terrorism, climate change, or pandemic diseases, political influence in supranational organizations, humanitarian help, military expenditures, etc. Thus, in countries such as Germany, with this space being overcrowded with a variety of issues, the topic of energy security is less prominent in IR than in countries synonymous with energy supply (e.g. Azerbaijan, Saudi Arabia). In the last decade, various attempts have been made to quantify energy security by creating indices as in Biresselioglu et al. (2015).3In their work, supply security was quantified by the number of supplier
energy forecasting.focus:natural gas 70 4International Energy Agency (2020). Germany 2020. Energy Policy Review. https://www.bmwi. de/Redaktion/DE/Downloads/G/ germany-2020-energy-policy-review. pdf?__blob=publicationFile& v=4 Last accessed 2020-12-27 countries, supplier fragility, and the subject of this research the overall volume of imported gas. To be clear on terminology, the security of gas supply and the reliability of gas supply are not synonymous. The latter is measured in the average time (minutes per year) of the service not being provided for a customer; in Germany’s case: 0.99 minute per year Bundesnetzagentur and Bundeskartellamt (2019). Back-up mechanisms keep the reliability of energy supply high, even if the security of energy supply is lower for a short period of a year. This research shows two novel approaches to this topic: 1) aiming at imports instead of demand/consumption from the data science point of view; and 2) reflecting on the relevance of imports for the discipline of IR. The next subsections compute maximum German gas imports, provide reasons behind a lack of work on import forecasting, and discuss differences in assumptions for import and demand forecasting. After the description of data set construction, the results of selected models are presented. The last subsection gives concluding remarks on forecasting of gas, and on its link to energy security and further to IR. Natural gas imports into Germany. Infrastructure considerations In 2017, Germany was the no.1 ranked country in terms of gas imports worldwide, and was no.13 in the ranking of gas exporting countries (i.e. re-exports). There are three routes for gas imports from the Russian Federation: • The Yamal-Europe system crossing Belarus and Poland, with a capacity of 33 bcm/y, • Nord Stream, connecting Russia and Germany directly via the Baltic Sea, with a total capacity of 55 bcm/y, • The Ukrainian gas transportation system with a total capacity over 100 bcm/y, crossing Slovakia, and the Czech Republic with the border node at Waidhaus, or through Slovakia and Austria (International Energy Agency (2020))4. Figure 4.1shows the excerpt of the Transmission Map 2019 with import, export and virtual nodes for Germany. To answer the question “what maximum amount of gas can be imported into Germany per year within the technical limits of the infrastructure (border nodes)?” this work calculates the upper limit by using maximum technical capacity in GWh/day per transmission pipe at the border provided
energy forecasting.focus:natural gas 71 Figure 4.1: Transmission Capacity Map - Germany, ENTSOG (2019) 5This figure was chosen after recalculating delivered gas to Germany from the Nord Stream pipeline and comparing it to the capacity of the pipeline. The gas market operates with thermal power units (GWh); the capacity of pipelines is defined in mass flows measured in billion cubic metres of transported gas per year. For the conversion, we used the heating value of the gas imported from the Russian Federation. by transmission system operators (TSO). The transmission pipe capacities are defined to satisfy demand on a daily basis. It is assumed that the node can be operated 90% of time.5The largest import node (Greifswald) is an exception; whereas its technical capacity is at 1,570 GWh/d, usually only 618.8 GWh/d (capacity adjusted by the capacities of OPAL pipeline) is considered. The German grid cannot absorb more gas; any excess gas would have to flow to the Czech Republic and from there back to Germany (Federal Ministry for Economic Affairs and Energy (2019)). Still, construction capacities of pipelines are higher than reported technical capacities. Therefore, it would be plausible to use a factor 1 (fully utilized time of a year) while calculating the maximum technical capacity for a year. Another factor could be used to mirror the technical feasibility of the maximum imports
energy forecasting.focus:natural gas 72 6German Federal Office for Economic Affairs and Export Control 7Actual import depends on demand in two market zones for HGas (high-calorific gas) infrastructure imported from Russia and Norway and the L-Gas (low-calorific gas) infrastructure from the Netherlands. 8Real export in 2018 was 1, 563,930 TJ (source: BAFA) and the technically possible calculated export based on the information from the Transmission Capacity map was 3, 758,047 TJ. Abbreviations from the table 4.2: ISSA - Improved singular spectrum analysis, LSTM - long short-term memory, FARX - AutoRegressive model with exogenous variables In theory, the prices of all energy commodities decrease in the long run. 9This is still the case in China. 10 In September 2019, two Saudi Arabian oil facilities were attacked, and their production - accounting for 5% of global production - was interrupted. As a result, Brent crude prices increased by 14.6% to $ 69.02, and US crude oil by 14.7% to $ 62.90 (Wearden (2019)). Ten days after the attack, one of facilities restored its oil production (Astakhova (2019)). per node and a pipeline section by taking into account the capacity of compressor stations. As daily capacities for border nodes are reported at the maximum level, this factor has been disregarded. In the second step, this maximum technical physical capacity for imports to and exports from Germany has been compared with actual imports and exports published by BAFA6for the year 2018.7For 2018, based on the above-mentioned assumptions, Germany used ca. 44% of its technically maximum possible capacity on the import side (10,045,322 TJ); for exports it was 41.6%.8Reported exports from BAFA also include "Ringflüsse" (loops) - the amount of gas leaving Germany at one border crossing point (e.g. Olbernhau) and entering Germany at another point (Waidhaus). Table 4.1shows planned LNG-terminals for Germany. Name of the installation Status StartUp Year Nominal Annual Capacity (in billion cubic metre/year) LNG storage capacity cubic metres LNG Brunsbüttel LNG Terminal planned 2022 8,00 240.000 LNG Stade GmbH planned - 5,00 - Rostock transshipment planned - - - Wilhelmshaven planned 2022 10,00 263.000 Table 4.1: LNG Import terminals in Germany. Source: Gas Infrastructure Europe (2021) Missing forecasting on gas imports Most gas forecasting is conducted on an hourly-basis as in Su et al. (2019); studies on monthly gas consumption forecasting are scarce, and based on the literature review conducted for this research, papers on forecasting gas imports are non-existent. Therefore, the next close field of short-term gas forecasting is investigated with studies summarized in table 4.2. Whereas types of models and criteria remain the same, objects of studies vary in scope and in space, as does the type of the best performing model. Table 4.3compares approaches in forecasting gas demand and gas imports. The variable gas prices deserves attention; as gas prices used to be excluded from data sets as there was no expectation of an increase and some governments strongly regulated them for end customers.9Therefore, gas prices have a low information value for modelling a problem. The occurrence of events such as attacks on
73 Author Forecast subject Models compared Criterion The best outcome Wei et al. (2019) Daily gas consumption London, Melbourne, Karditsa, Hong Kong Back propagation neural network (BPNN), support vector regression (SVR), multiple linear regression (MLR) and other MAE, RMSE, MAPE, mean absolute range normalized error (MARNE) ISSA, LSTM Chen (2018) Day-ahead high-resolution gas demand and supply in Germany Presented model: functional FARX compared to alternative autoregressive models Relative MAPE, relative RMSE, for direction: mean correct prediction (MCP) FARX -for average values of relative MAPE Merkel G. D. et al. (2017) Daily gas load for a utility in the USA 62 operating areas of local distribution companies in the USA 10 years of data Linear regression, artificial neural networks, deep neural network MAPE, RMSE Deep neural networks Akpinar and Yumusak (2016) Gas consumption (commercial and residential consumers, city-level) in Sakarya, Turkey Daily data summarized as monthly January 2011December 2014 Time series decomposition, Holt-Winters exponential smoothing, autoregressive integrated moving average (ARIMA), SARIMA MAPE, R2ARIMA Potoˇcnik et al. (2014) Day-ahead gas consumption, local distribution company in Croatia Data: consumption, weather data for 2heating seasons 5November 2011 –26 April 2012 9November 2012 –31 March 2013 Benchmark models: random-walk, temperature correlation Linear models: regression method, auto-regressive models with exogenous inputs Non-linear models: neural network, support vector regression (SVR) The mean absolute range normalized error (e), the adjusted R2measure Support vector regression Yu and Xu (2014) Day-ahead gas consumption in Shanghai 15 November 2005 –13 October 2008 Three-layer Back Propagation (BP) neural network structure, with a genetic algorithm MAR, MAPE, RMSE CCMGA–Im MBP mode CCMcat chaotic mapping Szoplik (2015) Hourly peak offtake of gas in the following year for the city of Szczecin, Poland 1January 2009 –31 December 2011 ANN network (MultiLayer Perceptrons) MAPE, RMSE nRMSE ANN Table 4.2: Short-term forecasting - case studies
energy forecasting.focus:natural gas 80 21 R Documentation (2020). Holt-Winters function | R documentation. https://www. rdocumentation.org/packages/ stats/versions/3.6.2/topics/ HoltWinters Last accessed 2020-09-06 22 For an additive model: α=0.8293 (the coefficient for the level smoothing), β= 1e−04 (the coefficient for the trend smoothing) and γ= 1e−04 (the coefficient for the seasonal smoothing). at=α(Yt/st−p) + (1−α)(at−1+bt−1) bt=β(at/at−1) + (1−β)bt−1 st=γ(Yt/at) + (1−γ)st−p Special cases are obtained while setting smoothing parameters α, βand γto zero; with α=0 the level is unchanged, with β=0 the slope is constant over time and with γ=0 we fix the seasonal pattern. In the formula at, real values are deseasonalized by dividing real values by the seasonal number (Makridakis et al. (1998)). In the formula stfor seasonality, Ytdenotes real values from the data set, thus containing seasonality and randomness, whereas atis already smoothed (average) value with a seasonality element. The factor γserves to smooth the randomness which is necessary due to the Yt. To start the algorithm, initial components must be set; for the HoltWinters function of the R stats package, start values come from a simple decomposition in trend and seasonal component using moving averages on the start.21 In this study, both trends showed similar results: the values for forecast imports were higher for the multiplicative trend but not more accurate. Parameters are determined by minimizing the squared prediction error and low values for βand γ22 indicate that rather older values of xare weighted more. Values of α(specifying how to smooth the level component) near 1.0 mean that the latest value has more weight. Holt−Winters filtering Time Observed / Fitted 2005 2010 2015 200000 300000 400000 Figure 4.5: Holt-Winters filtering. The black line represents the actual value; the red line represents the filtered time series.
energy forecasting.focus:natural gas 81 Forecasts from Holt-Winters 2005 2010 2015 2020 2e+05 4e+05 . Figure 4.6: Point forecasts and the 80% and 95% prediction intervals obtained using Holt-Winters model 23 Autocorrelation function is one of the diagnostic checks for model uncertainty. For any suitable forecasting model, residuals left over after fitting the model should be white noise (Makridakis et al. (1998)). Figure 4.5depicts forecasted values as well as the actual values. The method performs well in predicting seasonal peaks, i.e. the highest amounts of gas cross-bordered per year. The fan chart (figure 4.6) shows prediction intervals, set to 80% and 95% by default. The autocorrelation function indicates patterns in in-sample forecast errors called residuals (figure 4.7).23 As for figure 4.7:1) the boundaries are set by definition, 2) monthly data are used; the time lag is expressed in years, 3) by default, the ACF at zero time lag is set to 1. The blue lines are set to 95% confidence interval (an estimate of a fixed but unknown parameter value). Correlation values outside of this threshold are likely to mean a correlation. "Likely" is stressed as "a priori approximately 1 out of every 20 correlations will be significant based on chance alone."(E. E. Holmes, M. D. Scheuerell and E. J. Ward (2020)). Second, the Ljung-Box test results have been evaluated with the p-value 0.001293, which indicates the possibility of non-zero autocorrelation within the first 20 lags. To sum up the evaluation, as there is a pattern in the error residuals and due to the result of the hypothesis test, the model would not be the first choice for predictions. Besides the Holt-Winters time series (HWTS) model, we considered the Bayesian Structure Time Series models, however, with less promising results.
energy forecasting.focus:natural gas 82 0.0 0.5 1.0 1.5 −0.2 0.2 0.6 1.0 Lag ACF ACF - patterns in in-sample forecast errors Figure 4.7: Autororrelation lag 1-20 (lag =20) 4.0.4Regression analysis Regression analysis assumes that the output (gas imports) exhibits an explanatory relationship with a few independent variables under the assumption of continuity i.e. the explanatory relationship will not change. First, we included all variables from a data set, however, the variables gas use for electricity production and cooling degree days showed no significance in the model; hence they were excluded. 200000 250000 300000 350000 400000 450000 −50000 0 50000 Fitted values Residuals lm(Import ~ .) Residuals vs Fitted 192 159 179 (a) Residuals vs fitted −3 −2 −1 0 1 2 3 −2 0 2 4 Theoretical Quantiles Standardized residuals lm(Import ~ .) Normal Q−Q 192 159 179 (b) Normal Q-Q Three assumptions for the algorithm have been checked: the linear relationship, normality, and homogeneity. Residuals vs fitted depicted a horizontal line without a distinct pattern (figure 4.0.4, a); Normal QQ confirmed the normality assumption as residual points follow the
energy forecasting.focus:natural gas 83 200000 250000 300000 350000 400000 450000 0.0 0.5 1.0 1.5 2.0 Fitted values Standardized residuals lm(Import ~ .) Scale−Location 192 159 179 (c) Scale Location 0.00 0.05 0.10 0.15 0.20 0.25 −2 0 2 4 Leverage Standardized residuals lm(Import ~ .) Cook's distance 0.5 0.5 1 Residuals vs Leverage 192 20 163 (d) Residuals vs Leverage 24 Heteroscedasticity refers to the situation in which the variability of a variable (residuals) is unequal across the range of values of a second variable (fitted values). 25 Feature selection is an optimization problem of searching for the features’ combination that predict the response optimally (Kuhn and Johnson (2016)). 26 Usually all inputs are standardized to have a mean of zero and standard deviation of one (Hastie et al. (2009)). straight dashed line (figure 4.0.4), b). Scale-location was used to check the homogeneity of variance of the residuals, depicting a heteroscedasticity24 problem (figure 4.0.4, c). As for the feature selection,25 analysis of variance (ANOVA test), has been conducted with two models; model 1 including all of the variables, model 2 with all variables but two above-mentioned inputs. Although feature selection has improved models, the standard residuals overall pattern change is negligible. To solve the problem of non-linearity, we also considered introducing an interaction between some predictors. In the Explanatory Data Analysis (EDA), a relationship between the consumption of gas for electricity production and gas storage balance has been detected. Therefore, an interaction between them has been introduced to determine a possible difference. However, this interaction, as well as the interaction of heating degree days and gas storage, has not contributed to the increased performance of a model. Also, dummy variables have not improved the model and eventually the standard regression models were chosen, excluding price. 4.0.5Artificial Neural Networks An artificial neural network computes its output by multiplying the inputs xby weights (w0)and passing the result through an activation function. The training algorithm selects the weights of the input units by following the goal of minimizing a cost function, e.g. mean squared error (MSE). As the scaling of the inputs26 determines the effective scaling of weights, data has been normalized. Within the R package neuralnet, the network architecture with the logistic activation function has been chosen under the criterion of the least total
energy forecasting.focus:natural gas 84 27 Makridakis, S. G., Wheelwright, S. C., and Hyndman, R. J. (1998). Forecasting: Methods and applications. Wiley, Hoboken, NJ, 3. ed. edition. ISBN 9780471532330 28 We also tested the random forest model. This approach uses the most recent data from the data set as the forecast for the next period but it showed over-fitting and it was left-out of consideration. Hyndman and Koehler (2006) recommends using MAPE if all data is positive and much greater than zero due to its simplicity. squared error (figure 4.8). After the training step with the network learning an approximation of the relationship between inputs and an output (gas imports), predictions have been made and compared with the testing data. 7.90577 −1.98831 −1.64543 Domestic gas productionDE −1.4449 2.06774 2.77064 Cooling Degree Days 0.20998 0.60312 0.90057 Cross-border price −3.16994 −1.07221 5.73664 Gas storage-saldo 0.16016 −3.80498 −5.84787 Gas consumption for electricity production 2.9379 8.00326 0.96169 Heating Degree Days −0.53585 0.71341 −0.7669 Import −1.42099 −0.3592 1.09926 1 0.96007 1 Error: 0.805003 Steps: 1692 Figure 4.8: Plot of a trained neural network including synaptic weights and basic information. The black lines show connections between layers and weights on each connection; the blue lines show the bias term for each step. The training process needed 1692 steps until absolute partial derivatives of the error function were smaller than the default threshold (0.01). The information value of the plot is rather low; one cannot relate inter-steps of the model to making statements on the response variable. Asmall size data set is never ideal for ANN, as neural networks require a much larger number of observations than other models such as linear models.27 Further, it is not clear which parameters (i.e. number of layers and number of nodes) are best for the network; in this case, one layer with ten nodes produced the best result in the mean absolute percentage error (MAPE).28 4.0.6Modelling conclusions The performance of models is checked by the accuracy measure MAPE as it is not scale-dependent, there are no zero values in the data set and the results are not intended to be compared with gas imports of countries with other pattern profiles. Table 4.4provides an overview of the out-of-sample forecasting errors for three chosen models with the
energy forecasting.focus:natural gas 85 29 Szoplik, J. (2015). Forecasting of natural gas consumption with artificial neural networks. Energy,85:208–220 30 Lewis, C. (1982). Industrial and business forecasting methods. A practical guide to exponential smoothing and curve fitting. Butterworth, London. ISBN 0408005599 31 ˇ Ceperi´c, E., Žikovi´c, S., and ˇ Ceperi´c, V. (2017). Short-term forecasting of natural gas prices using machine learning and feature selection algorithms. Energy,140:893–900 32 This is valid especially for neural networks and the field of medicine where the interpretation of models is crucial for the results’ usability. 33 Scarpa, F. and Bianco, V. (2017). Assessing the quality of natural gas consumption forecasting: An application to the Italian residential sector. Energies,10(11):1879 persistence model where the forecast for all months of 2019 equals the value of the last observation from the data set, December 2018. Only results for the ANN with three nodes combat the persistence model and are comparable with results in Szoplik (2015).29 No MAPE outcome is equal to or exceeds 50% of the Lewis Benchmark for inaccuracy (Lewis (1982)).30 Lewis’ benchmark is arbitrary; labelling the model as accurate depends on the object of forecasting, to be more general: on the model domain. Table 4.4: MAPE comparison. Persistence Regression Time Series ANN with 3nodes MAPE 0.0924 0.116 0.278 0.062 Results are reflected upon in the relation to statements from previous studies: • There is still a gap between the rapid pace of algorithms’ development and their real-world application ( ˇ Ceperi´c et al. (2017)).31 Although algorithms work in many cases mentioned in the literature on testing the models (Merkel G. D. et al. (2017)), they are still rarely used as the basis for a decision-making process. As all results are valid for data sets (simplified sample of reality) with approximations describable by mathematical functions at the current stage of data science, it is hardly possible to make one step out of the boundaries of a model and make a statement about the observed reality.32 Authors who dare to make this step, call their statements speculations. Increased complexity of emerging models stresses the relevance of interpretability. • In real-time forecasting, all variables become the subject of uncertainty, and solving the issue by using rough estimates would produce relative confidence bounds so large that they would lose any practical use (Scarpa and Bianco (2017)).33 Therefore, here we restrict a forecasting exercise to the testing data. • Forecasting of gas imports is a more complex task than gas demand forecasting due to factors taking place outside the boundaries of the area (Germany). As an example, rising gas demand of neighbouring countries causes the steady growth of gas exports in the equation: Natural gas imports = natural gas (internal) consumption + exports – natural gas own production +/- storages; going hand in hand with the extension of infrastructure. • An explanatory data analysis shows that the cross-border price does not influence German gas imports. In the short-term, the available
energy forecasting.focus:natural gas 86 34 Most studies on energy security are country-specific (Ang et al. (2015)) although the concept of energy security requires a holistic view due to its interdependent character, especially in Europe. 35 Weather forecast providers use a deterministic model with no randomness, called an operational model and a probabilistic model running at specific hours and providing different weather scenarios, called an ensemble model (Gianfreda et al. (2020)). gas infrastructure (i.e. technical physical capacity of entry points) sets the upper boundary for imported amount of gas. • Yet, the gas market has all the attributes of a slowly-evolving mature sector: a long tradition with detailed legislation, regulation and standardization, a dense gas network, and sector independence. Still, the gas sector experiences pressure from both politics and the public to decrease its carbon dioxide and methane emissions and use its vast infrastructure system for transporting new products (hydrogen, biogas, synthetic methane). This could start a sort of decentralization of gas supply in the network: high pressure pipelines for the transport of the high-calorific gas (H-Gas) will be used for the rising export and distribution lines for local customers will be equipped with new supply entry points for synthetic methane and hydrogen injection. In some regions of Germany, the hydrogen infrastructure is being built as a project-specific island supply solution. 4.0.7Geopolitical implications of gas imports Monthly gas imports matter for the security of energy supply. Thus, the issue is being discussed at state level.34 The BAFA publishes total German gas imports on behalf of the BMWi (German Federal Ministry for Economic Affairs and Energy) and is not involved in forecasting. Several reasons behind not studying gas imports as a subject of short-term forecasting are: • Gas network operators make internal plans for future consumption according to forecasts based on past consumption and weather forecasts.35 It is unknown whether their algorithms are mainly based on ML methods. • Ministries develop long-term strategies, where a one-year ahead forecast, based on monthly data, is too detailed for their goals (i.e. energy security). • Germany’s security of gas supply is high (BMWi (2019)). Due to data protection, the country of gas origin is not disclosed to the public anymore; however, this information is provided in the Bundesnetzagentur and Bundeskartellamt (2019) for 2017 data. Most of the gas has been imported from Russia, Norway, and – with a declining tendency – from the Netherlands. • Missing high-resolution data on imports.
energy forecasting.focus:natural gas 87 In the previous work on gas supplies in the context of the IR discipline back in 2008, two assumptions were stated: gas supplies are concentrated and there will be no physical shortages in gas until 2020. Any shortfall of supply with demand will be due to political or economic obstacles. Geopolitical perspective on supplies distinguishes a) pipeline gas with the geopolitical importance of transit states, b) LNG gas with no direct transit states but “choke points” such as Bosporus, Strait of Hormuz or Strait of Malacca. In the IR discipline (academia), on the other hand, either unique events of the disruption of supply would be set into a story line or theories that enable understanding of main actors’ actions are published. Geopolitical debate on gas supply has undergone a shift in focus in the last twenty years. In the first decade of the 20th century, the discussion in the IR circled around gaining control of new gas fields and transport routes. At Gazprom’s Annual General Shareholder’s Meeting on 30 June 2006, Alexey Miller, Chairman of the Gazprom Management Committee, stated: "Ongoing global competition to gain control over hydrocarbon reserves has shown that state owned and backed companies have considerable advantages in obtaining dominating positions on international markets. Integration of state and commercial approaches enables to ensure a long-term planning based on prospective gas balance on the national and international scale." (Miller (2006)). Figure 4.9: Correlation of gas consumption among Germany and its neighbouring countries (selection). Data source: Eurostat The post-2020 global energy system still counts on rising gas imports but from new exporting regions (the US) and in another form (LNG). The intention for using gas has changed, too. With applying energy-efficiency measures within the residential sector, gas consump-
energy forecasting.focus:natural gas 88 36 Germany is an energy hub for physical gas flows in Europe. Gas consumption in neighbouring countries increases German gas imports while German gas exports also increase. 37 Examples: extreme cold winter in 2010, a warm winter in 2014. tion will rise in the electricity sector to back-up the electricity production from renewable energy sources. To sum up, high fixed costs for pipeline infrastructure and three separate gas consumption regions (Northern-American, Asian-Pacific and European) remained the reality of gas supplies. Changes occur in the gas pricing and the realization of producer countries that gas demand will be curtailed by pursuing environmental policies in Europe. Figure 4.10: Time plot of gas consumption of Germany and neighbouring countries Regarding Germany,36 figure 4.9shows the high correlation of gas consumption among Germany and some of its neighbouring countries. Shared climate and weather conditions are one of the reasons for the correlation. An analysis of of heating degree days revealed a high correlation among all countries in question but Denmark. Gas profiles in figure 4.10 mirror the set-up of the energy sector in a country; e.g. the curve for Germany depicts remarkable differences between winter peaks37 and lows, whereas the Danish profile is flat as heating in Denmark is supplied by district heating based on renewable energy sources rather than by using gas boilers. Furthermore, gas consumption trends of France and Germany resemble each other. In France, industry consumes more gas than the residential sector as of 2016,
energy forecasting.focus:natural gas 89 with the majority of gas consumed in the northern part of France. Germany has been more successful in diversifying gas imports by routes than by product (diminishing L-gas imports from the Netherlands will be replaced by H-gas; low share of other methane-based gases).
energy forecasting.focus:natural gas 96 Production Capacity Year Billion m3Mil.m3/h 2019 6.26 0.80 2020 5.82 0.74 2021 5.72 0.73 2022 5.38 0.68 2023 5.11 0.65 2024 5.76 0.72 2025 5.44 0.68 2026 5.02 0.63 2027 4.61 0.57 2028 4.23 0.52 2029 3.99 0.49 2030 3.73 0.46 Table 5.2: Gas production prediction, source: FNB Gas (2019), original source: BVEG e.V (Bundesverband Erdgas, Erdöl und Geoenergie e. V.) 13 In German: Bundesamt für Wirtschaft and Ausfuhrkontrolle of oil equivalent (TOE) using an average conversion factor, they do not correspond exactly with gas volumes expressed in Germany’s federal statistics. • Data on natural gas imports is taken from the International Energy Agency (IEA) and from the Federal Office for Economic Affairs and Export Control (BAFA).13 • Electricity generation. Scenarios for Germany predict increasing share of gas consumption for electricity production due to the phaseout of nuclear and coal power plants. They balance electricity production from weather-dependent sources; to be precise, the lack of it. There are no data on gas consumption for electricity production back to 1980, therefore the data set includes electricity generation from BP. The value for the last year in the data set (2018) comes from the Federal Environment Agency (Umweltbundesamt). All data are based on the gross electricity production, i.e. power plants’ own consumption is included. • Other data such as oil consumption, hydro electricity, electricity produced in biomass power plants, by solar, wind and nuclear energy and coal consumption are from the BP Statistical Review of World Energy, 2019. There is a lack of high-quality data sets on renewable energy sources, as they only started being collected in the 1990s(this research’s data set starts in 1980). There is no data on LNG gas consumption; at the time of writing, Germany does not operate an LNG
energy forecasting.focus:natural gas 97 14 Busse, S., Helmholz, P., and Weinmann, M. (2012). Forecasting day ahead spot price movements of natural gas - an analysis of potential influence factors on basis of a NARX neural network. Multikonferenz Wirtschaftsinformatik 2012 - Tagungsband der MKWI 2012.https://publikationsserver. tu-braunschweig.de/servlets/ MCRFileNodeServlet/dbbs_derivate_ 00027726/Beitrag299.pdf Last accessed 2021-04-14 15 Other factors influencing gas prices: geopolitical events and political decisions, seasonality/temperature, storage key figures, transport capacity, substitutes (e.g. oil price), spot and front month contracts for baseand peak-load) (Busse et al. (2012)). 16 Previous section on gas imports 17 Outliers could be defined as values differing more than four times the standard deviation from the mean as in Busse et al. (2012). terminal. Still, even without an own terminal, the LNG availability does influence gas prices (and indirectly consumption) according to the survey among several traders and portfolio managers conducted in 2011 by Busse et al. (2012).14,15 After constructing the research data sets for a short-term16 and a long-term period, relevant aspects are summarized in table 5.3. 5.0.4Methodology and Results Exploratory data analysis. Data is highly correlated (the correlation plot in figure 5.2); thus, some of the correlated features have been removed. The data is not normally distributed and contains outliers. Since we keep any input due to the size of the data set, outliers have not been removed.17 For certain models, this research works with log transformations of skewed variables. Some models, such as support vector machines, do not assume the exact shape of distributions. Seasonality in the low frequency (yearly) data is not observed. −1 −0.8 −0.6 −0.4 −0.2 0 0.2 0.4 0.6 0.8 1 GDP Pop NG_Res NG_Prod Brent_Price DE_Oil_Price Elect_Gen NG_Imp Heat_DE Cool_DE NG_Price_In NG_Price_H Oil_Con Hydro_E Biomass Solar Wind Coal_Con Nuclear NG_Con GDP Pop NG_Res NG_Prod Brent_Price DE_Oil_Price Elect_Gen NG_Imp Heat_DE Cool_DE NG_Price_In NG_Price_H Oil_Con Hydro_E Biomass Solar Wind Coal_Con Nuclear NG_Con Figure 5.2: Variables correlation matrix The following forecasts were produced by multiple linear regression (full and reduced form), Principal Components Regression (PCR), Support Vector Machines (SVM), time series analysis, and neural networks. PCR addresses collinearity by reducing the dimensionality of the data set. The time series models applied poorly to the data
energy forecasting.focus:natural gas 98 Criterion Aspect Short-term forecasting horizon: hours, days, months, Frequency: high Long-term with the yearly data Forecasting horizon: years, Frequency: low Purpose of forecasting Use of forecasts Improving operational efficiency of back-up plants for renewable energy sources, further saving energy and reducing emissions (methane and carbon dioxide) Managing supply contracts, indigenous production (measuring gap between domestic production and consumption), infrastructure planning Data set Choice of variables Weather variables (measured values or forecasts) – temperature or heating degree days, wind velocity, solar irradiance, air humidity Periods: weeksdays/public holidays Macro-socio-economic indicators (population, gross national product), gas availability, gas and electricity price, price of competitive fuels/technologies. Decomposition Seasonal pattern Yes, also within a day. No seasonal trend in a classical statistical meaning. Choice of the forecasting methods Statistical methods vs judgemental methods Extrapolation of past trends into the future (times series), Assumption: “patterns/relationships will remain during the forecasting phase. Combination of statistical and judgemental methods; otherwise one conducts judgement only through method selection; extrapolation methods perform suboptimal due to missing seasonal trend. Periodic adaptation of a model by using adaptive models or by retraining the model from time to time. Scenarios Analysis of disruptions of supply Possible if time steps involve hours, days and weeks Insufficient for short-term disruptions analysis Accuracy Criteria Mean absolute error (MAE), the mean absolute percentage error (MAPE), normalized mean square error (NMSE) The same criteria as for the short-term or not defined. Scenarios are rarely evaluated ex post. Table 5.3: Short-term and longterm forecasting of gas imports/consumption
energy forecasting.focus:natural gas 99 18 Hribar, R., Potoˇcnik, P., Šilc, J., and Papa, G. (2019). A comparison of models for forecasting the residential natural gas demand of an urban area. Energy,167:511–522 19 Collinearity implies two variables are near perfect linear combinations of one another, while multicollinearity involves more than two variables. and were abandoned. Neural networks were used in the full model, and two reduced models. The training data set is composed of data from 1980 to 2011, and the validation data for the forecast ranges from 2012 to 2018, unless stated otherwise. The Mean Absolute Percentage Error (MAPE) and Root Mean Square Error (RMSE) were computed and compared for each model. Linear regression Full Model. The linear regression model should be tried first (Hribar et al. (2019)).18 The first trial of a multiple linear regression with all variables shows the correlation coefficient R-squared adjusted at 0.9742 with most variables not significant and most coefficients close to zero. The variable Heating Degree Days has the highest significance level (pvalue for the t-test) of 0.00421, followed by gas produced in Germany and electricity production from wind power plants. The high adjusted R-squared value suggests there might be an issue with collinearity or multicollinearity.19 In the presence of multicollinearity, regression estimates become unstable, with higher standard errors. 60 70 80 90 −4 −2 0 2 4 Fitted values Residuals lm(NG_Con ~ .) Residuals vs Fitted 36 19 25 (a) Residuals vs Fitted plot −2 −1 0 1 2 −3 −1 0 1 2 Theoretical Quantiles Standardized residuals lm(NG_Con ~ .) Normal Q−Q 36 19 25 (b) Normal Q-Q The residuals vs fitted plot appears (figure 5.0.4, a) to have a distinct nonlinear pattern in the scatter, suggesting that the relationship between the explanatory variables and the target is non-linear. The Q-Q (quantile-quantile) plot (5.0.4, right) is non-linear, indicating the assumption of normality is not met either. The full regression model performs poorly on the validation data (figure 5.3).
energy forecasting.focus:natural gas 100 2012 2013 2014 2015 2016 2017 2018 50 60 70 80 90 Time NG Consumption (BCM) True Value Predicted Value Figure 5.3: Prediction. Full linear regression model, 2012-2018 2012 2013 2014 2015 2016 2017 2018 60 65 70 75 80 85 90 Time NG Consumption (BCM) True Value Predicted Value Figure 5.4: Prediction. Reduced linear regression model (2012-2018) PCR is one of the multivariate regresion models. "Multivariate" refers to more than one time-dependent variable. 20 Each vector component is a linear combination of all variables and is orthogonal to other components in the set. Reduced Model. The best subsets select which variables should be kept and which should be dropped from the linear model. A five-variable model with the minimum mean squared error prediction (MSEP) contains [NG_Prod], [Elect_Gen], [Heat_DE], [Wind], and [Coal_Con]. This model has an adjusted R-squared value of 0.9699 and all coefficients are statistically significant. Again, the research team checked diagnostic plots for testing the model’s validity and concluded that model assumptions are not met. The analysis of variance (ANOVA test) of the full and reduced model shows the p-value is 0.249, indicating that the reduced model performs just as well as the full model. Figure 5.4suggests better performance although the p-value is large; for the ANOVA test, there are not large differences between the two models. Removing variables correlated with the target. As the next step, variables [Coal_Con], [Pop], and [NG_Imp] were found to be highly correlated with gas consumption [NG_Con]. These variables were removed and the linear regression model was refit, resulting in an adjusted Rsquared value of 0.9644 with few variables being significant. 5.0.5Principal Components Regression Principal Component Analysis (PCA) converts a set of correlated variables into a set of linearly independent variables (principal components) using orthogonal transformations.20 This technique is widely used when the number of variables exceeds the number of objects (in this research, years) or because of collinearities (Mevik and Cederkvist (2004)). To formulate the problem from the beginning; in multiple linear regression: Y=b0+b1·x1+b2·x2+...bx·xx+e, (5.0.1) where xare predictor variables, ythe prediction and eerror term. The β(beta coefficient, also called regression weight) is given by β= (XTX)−1XTY(5.0.2) Because of the above-mentioned reasons, the XTXcould be singular. Such matrix is caused by linear interdependences among the variables i.e. some variable is a linear combination of other variables causing problems in a data analysis. Therefore we decompose X into orthogonal scores Tand loadings P: X=TP (5.0.3)
energy forecasting.focus:natural gas 101 21 The standard deviation is the square root of the average of the squared deviations from the mean. 40 50 60 70 80 90 100 110 55 65 75 85 NG_Con, 8 comps, validation measured predicted Figure 5.5: Cross-validated predictions for gas consumption data 0 5 10 15 2 4 6 8 10 12 Number of components RMSEP Abs. minimum Selection Figure 5.6: Example of one of strategy for finding optimal model dimension: one-sigma strategy The first principal component accounts for the largest possible variance in data, and each following component has the next highest variance possible provided that it is orthogonal to the preceding components. Whereas in the standard linear regression model a dependent variable is regressed on the explanatory variables directly, the principal components of the explanatory variables function as regressors (independent variables). 0 5 10 15 2 4 6 8 10 12 NG_Con number of components RMSEP (a) Cross-validated RMSEP curves for the gas consumption data 0 5 10 15 0 50 100 150 NG_Con number of components MSEP (b) Cross-validated MSEP curves for gas consumption data Since PCR is sensitive to variable scale, the data were scaled, i.e. standardised by dividing it by its standard deviation,21 before finding the principal components. Using PCR, 91.22% of the variation in the data is accounted for by the first five principal components. The plots show curves of the cross-validated mean squared error (RMSEP) in figure 5.0.5left and the root mean squared error (MSEP) in figure 5.0.5right. If the Partial Least Squares Regression (PLSR), meaning the alternative, was used, few principal components would be sufficient for decreasing the RMSEP error. The PLSR reduces a dimension by not only summarizing the original predictors but also relating them to the outcome. Inspecting a few aspects of the fit, with the first eight principal components chosen, the prediction plot in figure 5.5shows crossvalidated predictions versus measured values. The points follow the target line acceptably, significant anomalies have not beet spotted. Built-in methods in the pls package enable use of a few methods for choosing the suitable number of components. As an example, onesigma heuristic in figure 5.6chose the model with the fewest components; one standard error away from the overall best model.
energy forecasting.focus:natural gas 102 2012 2013 2014 2015 2016 2017 2018 70 75 80 85 90 95 100 Time NG Consumption (BCM) True Value Predicted Value Figure 5.7: Gas Consumption forecast using PCR with the first 8principal components (2012-2018) 22 Examples: back-up power and heat plants, ability of the cold start for combined heat and power (CHP) plants. On electricity production company level, an overprediction can mean a government penalty and an underprediction an increased cost. 2012 2013 2014 2015 2016 2017 2018 60 65 70 75 80 85 90 Time NG Consumption (BCM) True Value Predicted Value Figure 5.8: SVM full model, tuned (2012-2018) The first eight principal components accounted for ca. 98% of the data variation; results are shown in figure 5.7depicting the overprediction of gas consumption. To review, the primary value of forecasts lies in dealing with uncertainty while taking informative decisions. In the gas sector, decisions based on underpredictions may put the security of energy supply at risk, whereas decisions based on forecasts that turn out to be overpredictions decrease economic efficiency of the energy system. However, the energy infrastructure is usually built with a high security and safety margin.22 Thus, while inventing a new accuracy measure for forecasting results for energy infrastructure, it would make sense to penalize the underprediction. By this procedure, forecasters would produce biased forecasts that could be accepted in the forecasting community. Keane and Runkle (1990) wrote on the asymmetric preferences over forecast errors: ”If forecasters have differential costs of overand under-prediction, it could be rational for them to produce biased forecasts. If we were to find that forecasts are biased, it could still be claimed that forecasters were rational if it could be shown that they had such differential costs.” Then, the assumption of symmetric loss function, present in standard regression-based tests would not be appropriate for evaluating forecasts with asymmetric preferences (Keane and Runkle (1990)). 5.0.6Support vector machines (SVM) Classification and regression analysis are two fields for supervised SVM with no requirement for a normality assumption. We chose a non-linear ε-SVM with radial kernel function and tuning with the performance measure mean squared error. Kernel function separates two classes of data by transforming data into higher dimensional feature space which enables a linear separation (the kernel trick). The full model in figure 5.8predicts almost the same consumption for each year. It underfits as the best.tune function chose a very low cost parameter C. This is in line with the literature on SVMs; “when the cost is small, the model will “stiffen” and become less likely to over-fit (but more likely to underfit) because the contribution of the squared parameters is proportionally large in the modified
energy forecasting.focus:natural gas 103 Larger values of C address points near the decision boundary. A decision boundary is the region of a problem space in which the output label of a classifier is ambiguous (Hastie et al. (2009)). 2012 2013 2014 2015 2016 2017 2018 60 65 70 75 80 85 90 Time NG Consumption (BCM) True Value Predicted Value Figure 5.9: SVM full model, tuned II (2012-2018) 2012 2013 2014 2015 2016 2017 2018 60 65 70 75 80 85 90 Time NG Consumption (BCM) True Value Predicted Value Figure 5.11: SVM reduced (2012-2018) error function” (Kuhn and Johnson (2016)). Generally, this is the result of data not being scaled or the parameters not being tuned properly. As we already scaled data, this time we tuned parameters by extending the list of ranges for the cost parameter to 1000. Larger C leads to a smaller margin in the separating hyperplane, training errors will decrease and complexity increase. Obviously, decreasing training errors does not lead to decreasing the testing error (the distinction is explained in the section on complexity). The results of tuning are seen in figure 5.9. The number of variables in the SVM model has been reduced by choosing variable importance greater than 50 ((NG_Imp, GDP, Coal_Con, Wind, Pop, Elect_Gen, Biomass) in figure 5.10. Importance Hydro_E Oil_Con NG_Price_In NG_Prod Nuclear Cool_DE Heat_DE Brent_Price DE_Oil_Price Solar NG_Price_H NG_Res Biomass Elect_Gen Pop Wind Coal_Con GDP NG_Imp 0 20 40 60 80 100 Figure 5.10: SVM Variable Importance Variable importance indicates the most useful predictors for predicting the response variable in this model context. Though, the most important variables overall may not be the ones near the top of the plot. As figure 5.11 displays, overfitting prevents feature selection from producing more accurate predictions. In the terminology used in Fallah et al. (2019), variable selection represents a method whereas the SVM is a model used for the method. Alternative models for variable selection include step-wise refinement, correlation-based methods (searching for variables correlated with the output but not with each other (Fallah et al. (2019)) and others. These methods partially rely on expert judgement.
energy forecasting.focus:natural gas 104 23 Günther, F. and Fritsch, S. (2010). Neuralnet: Training of neural networks. R Journal,2(1):30–39.https: //journal.r-project.org/archive/ 2010/RJ-2010-006/RJ-2010-006. pdf Last accessed 2020-12-20 An activation function is a differentiable function that is used for smoothing the result of the cross product of the covariate or neurons and the weights. 24 SSE =1 2∑L l=1∑H h=1(olh −ylh)2 25 E=−∑L l=1∑H h=1(ylhlog(olh) + (1−ylh)log(1−olh)) "Sag" algorithm failed to predict gas consumption. 5.0.7Artificial neural networks ANN models are unstructured models enabling forecasting of any subject without the knowledge of relationships between variables. However, some working packages have been built in the context of regression analysis and apply a supervised method, i.e. the known output is compared to the predicted input in the iterative cycles unless parameters of the model (called weights in ANN) are optimized, based on the minimization of the error function. Considering the size of the data set, and the fact that the goal is to forecast gas demand in Germany, the neuralnet package was chosen as it is understood as the extension of regression analysis (a supervised method). The first step,data normalization, enables the higher population of data for the same manifold space. Values for weights usually start randomly, drawn from a standard normal distribution (Günther and Fritsch (2010)23) and are adapted according to the chosen algorithm. The activation function transforms the node’s aggregated input into the second layer. This works uses a non-linear logistic function f(x) = 1 1+e−u; the model was also run with the customized softplus function, approximating rectified linear unit (ReLU) function, however with less promising results. Two error functions could be used, either the sum of squared errors SSE24 where l=1,..., Lindexes the observations, i.e. given input-output pairs, and h=1, ..., Hthe output nodes, or, if a classification problem exists, cross-entropy E(Günther and Fritsch (2010)).25 Several algorithms have been tested; the resilient backpropagation with weight backtracking (default) and without backtracking, as well as "sag" and "slr" inducing the usage of the modified globally convergent algorithm, authored by Anastasiadis et al. (2005). The latter further modifies a learning rate, either by associating it with the smallest absolute partial derivative ("sag") or the smallest learning rate ("slr"). Figure 5.12 depicts the forecasting using the full model with and without weight backtracking, figure 5.13 shows prediction with a classical backpropagation algorithm, used as a benchmark. Figure 5.14 shows results for the further modified algorithm by Anastasiadis et al. (2005) with the "smaller learning rate." The default choice (resilient backpropagation with weight backtracking) performs better than back-
energy forecasting.focus:natural gas 105 2012 2013 2014 2015 2016 2017 2018 70 75 80 85 90 95 Time NG Consumption (BCM) True Value Predicted Value (a) Prediction (2012-2018), ANN (the resilient backpropagation with weight backtracking) 2012 2013 2014 2015 2016 2017 2018 70 75 80 85 90 95 Time NG Consumption (BCM) True Value Predicted Value (b) Prediction (2012-2018), ANN (the resilient backpropagation without weight backtracking) Figure 5.12: The influence of weight backtracking on prediction 2012 2013 2014 2015 2016 2017 2018 70 75 80 85 90 95 Time NG Consumption (BCM) True Value Predicted Value Figure 5.13: Prediction (20122018), ANN ( backpropagation) propagation without backtracking; the classical propagation and the method of Anastasiadis et al. (2005) does not follow the line of true consumption in these years. For any identified neural network, its weights follow a multivariate distribution and their confidence interval can be computed if the error function equals the negative log-likelihood (Günther and Fritsch (2010)). An identified neural network does not include irrelevant neurons neither in the input layer nor in the hidden layers; i.e. variables with no effect on the response variable or variables that are a linear combinations of other variables shall be excluded. In the next subsection,
energy forecasting.focus:natural gas 112 Most hydrogen in Germany is produced from natural gas, only 7% (3.85TWh) of demand is met via electrolysis (Ministry for Economic Affairs and Energy (2020). 32 In the simulation of Grüger et al. (2019) hydrogen production costs were reduced by up to 9.2 per cent and wind energy utilization increased by up to 19 per cent. 33 Grüger, F., Hoch, O., Hartmann, J., Robinius, M., and Stolten, D. (2019). Optimized electrolyzer operation: Employing forecasts of wind energy availability, hydrogen demand, and electricity prices. International Journal of Hydrogen Energy,44(9):4387–4397 34 All shares refer to 2017. Source: International Energy Agency (2020). 5.1Implications for the energy sector 5.1.1Hydrogen There are multiple reasons for increasing hydrogen’s share to cover energy demand in Germany and worldwide: • to replace fossil fuels with synthetic fuel produced in power electrolysers that split water into its components, powered by electricity produced by renewable or low carbon dioxide emission sources. This would avoid costs in electricity infrastructure in the long term and stabilize electricity grids in the short term as hydrogen is produced when the spot price of electricity is low,32 which indicates a higher share of electricity produced by renewable energy (Grüger et al. (2019).33 • to increase self-sufficiency of the energy sector. Technologies for utilizing hydrogen have been developed for decades; the issue of right timing or maturing the idea have been assumed as reasons for their late deployment in the energy system. Breaking down segments of gas and hydrogen consumption brings to light issues with their mutual interchangeability. The highest segment in consuming gas in Germany is industry (30.6%), followed by the residential sector (28.8%), heat and power generation (24.4%), commercial (13.8%), other energy (1.8%) and transport (0.6%).34 The total hydrogen consumption amounts to ca. 55 TWh with the highest share of demand in material production processes in industry and production of chemicals (Ministry for Economic Affairs and Energy (2020)). Estimations of hydrogen use in the energy sector Before 2030, none of the German scenarios in Jensterle et al. (2019) assume extensive use of hydrogen. Some scenarios expect hydrogen demand at 133 PJ (36.94 TWh) in transport and industry. As for 2050, German scenarios count on the demand ranging from 300 to 600 PJ (83.3 to 166.7 TWh). By2030, another estimation counts on increasing hydrogen demand by 10 TWh (Ministry for Economic Affairs and Energy (2020)) which does not automatically decrease the amount of natural gas used. Furthermore, planned increase of hydrogen in the transport sector does not impact overall gas consumption.
energy forecasting.focus:natural gas 113 35 Hydrogen will be utilized in all energy sectors: electricity production, transport, heating and industry. The second option is to transport hydrogen through newly-built hydrogen pipelines. 36 Altfeld, K. and Pinchbeck, D. (2013). Admissible hydrogen concentrations in natural gas systems. Gas for Energy, (03). https://www.gerg.eu/wp-content/ uploads/2019/10/HIPS_Final-Report. pdf Last accessed 2020-12-19 37 The Wobbe Index is the measure for the interchangeability of different fuel gasses. 38 The methane number describes the knock behaviour of fuel gas in internal combustion engines. 39 BP (2020). BP Statistical Review of World Energy. https: //www.bp.com/content/dam/bp/ business-sites/en/global/ corporate/pdfs/energy-economics/ statistical-review/ bp-stats-review-2020-natural-gas. pdf Last accessed 2020-12-21 40 Nymoen Strategieberatung (2016). Strategische Marktprognose Erdgas. https://zukunft.erdgas.info/ fileadmin/public/PDF/Politischer_ Rahmen/marktprognose-erdgas-2016. pdf Last accessed 2019-12-27 The IEA reports typical gross calorific (higher) heating values of natural gas in MJ/cubic metres. Russia: 38.23 MJ, Norway 39.24 MJ, UK 39.71.The heating value of hydrogen is 12.7 MJ per cubic metre). Until now, natural gas has been considered to be transported via pipeline infrastructure in Germany and the whole gas consumption has been related to it. To decrease the carbon footprint of gas transported via pipelines and, thus, save the investment in infrastructure, hydrogen is considered as the next energy carrier to fulfill energy demand.35 The first option discussed is injecting hydrogen into the pipeline system, bearing in mind limits in concentration due to technical constraints. Altfeld and Pinchbeck (2013)36 see a mixture up to 10 vol.% as not critical in most cases. In terms of combustion parameters, 10 vol.% of hydrogen decreases the Wobbe Index37 by 3%, as it also decreases the methane number38 and increases turbulent flame speed. Underground storage (bacterial growth), CNG steel tanks (interaction between hydrogen and steel), gas engines (increased combustion and end-gas temperature - higher NOx emissions) and gas turbines (no extra margin for variation in a fuel’s properties) have been identified as sensitive components of the gas system in relation to higher blending of hydrogen with gas (Altfeld and Pinchbeck (2013)). 5.1.2Publicly available estimations of gas consumption in Germany Natural gas consumption at the state level is reported either in billion cubic metres or exo-joules (both in the BP (2020)39) or in the energy unit TWh (Nymoen Strategieberatung (2016)40. For the conversion for past data, one needs to access the share of L-Gas with the heating value at 8.9 kWh/cubic metre and of H-Gas with the heating value of 11.1 kWh/cubic metre. Based on available information on the domestic production in Germany (L-Gas) and imports from the Netherlands (L-Gas), Norway (H-Gas) and other countries (mostly HGas) in 2016, shares have been calculated at 28% for L-Gas and 72% for H-Gas. The difference in 3.7% points rather to different methodologies used in energy statistics. As for 2030, all gas used will be H-Gas. Nymoen Strategieberatung (2016)developed the scenario for final energy consumption until 2035 based on the rise of the GDP, population, forecasts for energy carriers, policy measures as well as on the forecasts of carbon dioxide emission certificates. Figure 5.22 shows scenarios for gas consumption development. Gas consumption (in primary energy) was supposed to decrease by 38 TWh, or 5%, from 2015 to 2020 due to climate protection measures (Nymoen Strategieberatung (2016) which the latest gas statistics do not confirm.
energy forecasting.focus:natural gas 114 Figure 5.22: Market prognosis for gas consumption. Source: Nymoen Strategieberatung (2016) 41 A third option, hydrogen chemically bonded with metal hydrides or liquid organic hydrogen carriers has not entered the market, yet. Federal Ministry for Economic Affairs and Energy (2019) expects a slight to medium-sized decline in gas consumption by 2029, based on the Network Development Plans for Gas Infrastructure. Some German regions will experience an increase in gas consumption; therefore gas infrastructure will continue to expand and became more complex due to injection of synthetic methane, bio-methane and hydrogen. Considering the ambitious goals of the German government in deployment of hydrogen, there are two main options41 for infrastructure: • injection of hydrogen into gas infrastructure • building separate hydrogen pipelines or the conversion of parts of the gas infrastructure into hydrogen infrastructure. Finally, possible paths for German gas imports and gas consumption are presented below. These examples demonstrate one of a few forecasting methods for forecasts into the future, without the possibility of computing accuracy measures. Although paths resemble scenarios in modelling analytical work of EIA or IEA, they were simulated with random value of error terms.
5.1. IMPLICATIONS FOR THE ENERGY SECTOR 115 2e+05 4e+05 6e+05 1990 2000 2010 2020 2030 2040 Time Imports series Series 1 Series 2 Series 3 Series 4 Series 5 Series 6 Series 7 Series 8 . Figure 5.23: Paths for German gas imports for next 25 years. Nnetar function (non-linear autoregressive model) uses lagged values of the time series (Hyndman (2017)). Weights in "repeats" start with random values. 60 70 80 90 1980 1990 2000 2010 2020 2030 2040 Time NG_consumption series Series 1 Series 2 Series 3 Series 4 Series 5 Series 6 Series 7 Series 8 Series 9 Figure 5.24: Nine paths for gas consumption in Germany. Although paths resemble scenarios in modelling analytical work of EIA or IEA, they were simulated with random value of error terms.
1The simple median or the mean of a few experts’ forecasts would be in line with the outcome of the research paper of Mannes et al. (2014) titled “The Wisdom of Select Crowds”. In current trends, forecast combinations make use of their work by considering single models as expert opinions. In the IR field, after decades of efforts to use quantitative methods, expert analysis of events still dominates among methods. 2Commercial forecasting tools have been mentioned but not directly used in these chapters. 3Beyca, O. F., Ervural, B. C., Tatoglu, E., Ozuyar, P. G., and Zaim, S. (2019). Using machine learning tools for forecasting natural gas consumption in the province of Istanbul. Energy Economics,80:937–949 4Akpinar, M. and Yumusak, N. (2016). Year ahead demand forecast of city natural gas using seasonal time series methods. Energies,9(9):727 6 Conclusion The goal of this work was to examine possibilities and trends in gas forecasting while understanding the principles behind commonly used models. First, forecasts based only on expert opinions prevailed.1 Reviewing the research questions • RQ1 What are current approaches to predict gas imports and gas consumption at the country level? What is the impact of the domain knowledge of the gas sector on forecasting? Gas imports have been rarely forecast, with reasons for this listed in chapter 4. Methods for forecasting gas consumption are listed and explained in chapter 3 with modelling exercises in chapters 4 and 5 for gas imports and consumption, respectively. These exercises provide insights into the current possibilities of forecasting packages2and are comparable with studies conducted by Beyca et al. (2019)3and Akpinar and Yumusak (2016).4Previous studies served as an inspiration in terms of searching for suitable input variables and measures of accuracy. The transferability of results of any study is low due to the unique setup of 1) data set (information sample of the forecasted process), 2) choice of models, 3) goal of a study, the length of a forecasting horizon and scope (company, city, state, global level) of forecasting. Domain knowledge applies while simulating gas flow in nodes of a pipeline system and prior to modelling in the process of construction of a data set. Two data sets were necessary as short-term and long-term forecasting require specific input variables, the latter one including GDP, population, and price of alternative fuels, to name a few. For the long-term scenarios, some models in chapter 5 include
energy forecasting.focus:natural gas 118 5This work complies with most of the golden rules in the Forecasting checklist, introduced by Armstrong and Green (2017) but combination of forecasts. Changes in population are small; the expected gas production in Germany up to 2030 is known in FNB Gas (2019), source: BVEG, e.V. 6Standard commercial software does not account for this higher uncertainty and keeps the width of the prediction interval the same over the periods (i.e. months, years). 7"a pattern is simple if it can be generated by a short program or if it can be compressed, which essentially means that the pattern has some "regularity" in it." (Li and Abu-Mostafa (2006)) the simulation of the newly gathered data to the model (ANN, one year data, previous year, true data). The final size of both data sets depended on 1) fulfilling minimum requirements of a data set for a few models, 2) including all relevant variables if data was available from reliable sources, and 3) the selection of a starting point and the last known observations for time series without missing values. No forecasts have been adjusted as this would introduce intentional bias (Armstrong and Green (2017)5. Why were not more ex ante forecasts produced? The forecast error would increase dramatically due to the uncertainty of the future values of inputs such as population and gas production in Germany (lower risk) or weather (higher uncertainty), and this would make prediction intervals too wide to make any statement about the future imports or consumption.6 • RQ2 Is there any relation between the complexity of the data set, forecasting models and the process modelled (gas flows) to the accuracy of forecasts? Section 3.3 discusses various complexity criteria. In accordance with the literature on data complexity7, regularity in data sets has been chosen as a measure for computing complexity of data sets (Sample Entropy, Approximate Entropy). The Approximate Entropy calculates the complexity of the process (model) whereas derived versions of Shannon entropy (information theory) calculate the complexity of a measure. As for the complexity of forecasting models, either the computational complexity (computing time necessary for a model to complete the task of forecasting), or definition based on understanding the model by uses are taken into account. For the latter definition, naïve or no-change model without seasonal adjustment, seasonal adjustment, single-exponential smoothing, Holt’s exponential smoothing,dampened exponential smoothing and simple average of the exponential smoothing forecasts are regarded as simple (Green and Armstrong (2015)). Regarding complex processes to be modelled, in physics, any phenomenon is too complex and therefore physicists work with simplifications, set boundaries, and specified conditions until the
energy forecasting.focus:natural gas 119 8For instance, searching for the inputoutput relationship present in the training data set as noted in Li and Abu-Mostafa (2006)). 9Makridakis, S., Hyndman, R. J., and Petropoulos, F. (2020). Forecasting in social settings: The state of the art. International Journal of Forecasting, 36(1):15–28.https://doi.org/10.1016/ j.ijforecast.2019.05.011 phenomenon is computable. Present trends in forecasting have shown the opposite trend, as more complex models are introduced to probe new ways to forecast. Models of higher complexity go hand in hand with the complexity of modelled processes themselves (e.g. electricity production in hydro power plants). Finally, deep inspection of the relation of complexity towards the data set8, models and processes modelled would require avoiding the connection to energy and forecasting as these relations are supposed to be valid in general terms. 6.1Closing thoughts • More accurate forecasts provide marginal contributions for the gas sector or for energy policy, if they are neglected while making decisions at the gas company or state level. Ministries give assignments to institutes to model situations as if climate goals were achieved by 2050. These aspirational forecasts have to be distinguished from forecasting based on data. Second, policy makers, as most people, prefer a narrower interval, even without the true value included, than the wider interval, which they perceive as too wide (Yaniv and Foster (1995)). In comparison to finely grained judgements, imprecise judgements are less informative and less attention is paid to them from experts outside the forecasting community; however, they are likely to be more accurate. • Contextual information on commodities such as oil, gas or solar technology is strictly bound to the existing available infrastructure and its capacity factors. In turn, this enables the final reality-check of forecasts. • Forecasting, especially with ML methods, is more vital for higher frequencies; higher than hourly values approaching even real-time data forecasting. By this, gas demand forecasting at the state level cannot reap the advantages of new trends as data of higher frequencies is available only to in-house experts in charge of their market area. Data sets used in chapters 4 and 5 contain low resolution data (monthly, yearly) which can decrease accuracy of results (Hong et al. (2014)). The sample (every data set represents just a sample of information on the process) is insufficient for building an appropriate model, therefore features of sophisticated models cannot be utilized to their full extent. Therefore, explanatory variables as well as contextual information was included. In their latest paper, Makridakis et al. (2020)9conclude explanatory variables can be useful if there are accurate forecasts of the explanatory variables
energy forecasting.focus:natural gas 120 10 In classical forecasting problems, researchers compare models as presented in chapters 4 and 5, based on the accuracy measure(s) of their choice, and identify the model that outperforms other models. However, no model can show outstanding performance for all data sets and forecasting problems. available (e.g. in the case of German gas production) and when assumed relationships between an output and explanatory variables probably continue into the future. The latter condition is met for weather forecasts. In electricity load forecasting, the strong correlation between the temperature and load is widely accepted in the field. However, this is valid only for the US or other countries in warm climates with widespread air-conditioning use, meaning an extensive use of electric heating and/or cooling units. Therefore, in energy forecasting, data sets and assumed relationships between a forecast and explanatory variables must be country-specific. The tremendous number of studies on (electricity) load forecasting, hydro power forecasting, and political forecasting shows an imperfect transferability of methods10 and the wish to test current models from new perspectives. Amendments take a form of stacking and shuffling data sets from various countries for forecasting electricity prices, direct forecasting of peak load from data on temperature, etc. Moreover, full-time forecasters prefer producing synthetic data sets to control the quality of probabilistic forecasts (bias, error, variance, covariance) and enhance their understanding of a models’ performances. 6.2The value of forecasting This work has not touched environmental aspects of gas use, especially flaring, methane emissions during the distribution, boil-off during the LNG transport and emissions from gas combustion. Although constraints on emissions are built-in in constructed energy scenarios, they matter in developed countries only. Rather, this research drew attention to assumptions behind forecasting and modelling, which should serve a better understanding of their products (i.e. forecasts, scenarios) and, thus, using them with caution in energy planning. Business forecasting measures value in monetary units and poses the question of "what is the actual economic value of forecasts?" Forecasting in the environmental sciences does not provide straightforward benefits; it does not save the environment nor can it prevent harmful acts due to inefficient energy supply planning. However, an improved ability to accurately anticipate energy consumption, emissions, or inflation rates, would benefit decision-making. For this to happen, ever improving methodologies for the anticipation must be applied. Forecasts assist investment, trading decisions and policies on
energy forecasting.focus:natural gas 121 energy security. Although all forecasts are incorrect, it is better to plan actions now based on incorrect forecasts than to plan without forecasts at all. Post 2030, it will not be macro-economic variables (e.g. GDP or, population) that will determine future gas demand in Germany, but rather policies of phasing-out natural methane-based gases in Germany as well as in neighbouring countries.
energy forecasting.focus:natural gas 128 nnumber of data points. 35,37,48 pthe number of estimated parameters in the model. 47,51,66 rtrelative error. 35 rPearson’s correlation coefficient. 36 τtraining data set. 34 xpredictor variable, an observation in the data set. 45,100 ytactual value. 35 yprediction, outcome, output, target variable. 34,45,48,65,100
Appendices
Input Unit Description Source Heating Degree Days - Germany Eurostat Cooling Degree Days - Germany Eurostat Imported Gas TJ BAFA Website:https://www.bafa.de Cross-border Price EUR/TJ BAFA https://www.bafa.de Gas consumption for Electricity Production GJ Germany Destatis.de, Genesis-Tabelle 43311-0002,https://www-genesis. destatis.de Domestic Natural Gas Production TJ W. E. G. Wirtschaftsverband Erdölund Erdgasgewinnung e. V. Gas - Export TJ Considered but Excluded BAFA Gas Storage Balance TJ BAFA Table 1: Data key for German gas imports
energy forecasting.focus:natural gas 132 Table 2: Data key for German gas consumption Variable Name Meaning Unit Source and Notes Year Year of data collection; denotes a single observation Numerical GDP Gross domestic product (Germany), constant prices Billions of National Currency (Euros) IMF, World Economic Outlook Database (WEO), original source: National Statistics Office. Notes: Data until 1990 refers to German federation only (West Germany). Base year: 2010. The GDP deflator in the base year is not exactly equal to 100 since it is computed based on the quarterly real GDP data, which have been adjusted for seasonal and calendar effects. Type: Expenditure-based GDP. Pop Population of Germany Millions of Persons IMF, WEO, April 2019, original source: National Statistics Office. Latest actual data: 2018 Notes: Data until 1990 refers to German federation only (West Germany). Data from 1991 refer to United Germany. Data last updated: 03/2019 NG_Res Proved Natural gas reserves (Germany) Trillion Cubic Meters BP from the Data product: Energy Production and Consumption. Permalink: https://www. quandl.com/data/BP/GAS_RESERVES_DEU Quandl Code: BP/GAS_RESERVES_DEU Description: Proved reserves of natural gas - Generally taken to be those quantities that geological and engineering information indicates with reasonable certainty can be recovered in the future from known reservoirs under existing economic and operating conditions.For the 2018, data are taken from: LBEG, Landesamt für Bergbau, Energie und Geologie. Oil and gas reserves in Germany on 1st January 2019 (In German: Erdölund Erdgasreserven in der Bundesrepublik Deutschland am 1. Januar 2019, Tab. 5: Rohgasreserven am 1.1.2019 nach Fördergebieten (in Mrd. m³(Vn)), Link: www.niedersachsen.de Last check: 3rd July 2020. Observations from the first source and the second one for 2018 is equal. Continued on next page
energy forecasting.focus:natural gas 133 Table 2– continued from previous page Variable Name Meaning Unit Source and Notes NG_Prod NG production (Germany) Billion Cubic Meters BP from the Data product: Energy Production and Consumption. permalink: https://www. quandl.com/data/BP/GAS_PROD_DEU Quandl Code: BP/GAS_PROD_DEU.As far as possible, the data represents standard cubic metres measured at 15Cand 1013 millibar (mbar); as they are derived directly from tonnes of oil equivalent using an average conversion factor, they do not necessarily equate with gas volumes expressed in specific national terms. Brent_Price Average annual Brent crude oil U.S.dollars/barrel OPEC; IEA. Published by: MWV.de, ID 262860, June 2019. Link: https://www.statista. com/statistics/262860/uk-brent-crude-oil-price-changes-since-1976/ Original source: mwv.de, ID 262860 DE_Oil_Price Price of oil for Germany (Western, change to whole in 1999) Euros per Hectoliter www.destatis.de. Price for light heating oil. In German: Preise. Erzeugerpreise gewerblicher Produkte (Inlandabsatz). Preise für leichtes Heizöl, Motorenbenzin und Dieselkraftstoff. Lange Reihen ab 1976 bis Mai 2019, Artikelnummer: 561240219054 Elect_Gen Electricity generation (Germany) TWh BP from the Data product: Energy Production and Consumption. Permalink: https://www.quandl.com/data/BP/ELEC_GEN_DEU Quandl code: BP/ELEC_GEN_DE Based on the gross output.Value for 2018: https://www.umweltbundesamt. de/sites/_default/files/medien/384/bilder/dateien/2_datentabelle-zur-abb_ entw-bruttostromerzeugung-verbrauch_2019-02-26.pdf NG_Imp Natural gas imports TJ International Energy Agency, from 1999 on: Bundesamt für Wirtschaft und Ausfuhrkontrolle (BAFA) Heat_DE Number of heating degree days in Germany Numerical EUROSTAT. Accessed on 30th June 2019. Until 1990 – BRD. Continued on next page
energy forecasting.focus:natural gas 134 Table 2– continued from previous page Variable Name Meaning Unit Source and Notes Cool_DE Number of cooling degree days in Germany Numerical EUROSTAT. Accessed on 30th June 2019. Until 1990-BRD. Avg_Price_In Average price of NG (industry) (Germany) Cent/kWh www.destatis.de. Excluded: Value-added tax (VAT) Included: usage charges (in German: Nutzungsentgelte), gas taxes Avg_Price_H Average price of NG (household/domestic) (Germany) Cent/kWh www.destatis.de. Excluded: Value-added tax (VAT). Included: usage charges (in German: Nutzungsentgelte), gas taxes Avg_Price Total average price of NG (Germany) Cent/kWh www.destatis.de. Excluded: Value-added tax (VAT). Included: usage charges (in German: Nutzungsentgelte), gas taxes Oil_Con Oil Consumption Consumption in million tonnes oil equivalent* BP, Statistical Review of World Energy, 2019,68th Edition.https://www.bp.com/ content/-dam/bp/business-sites/en/global/corporate/pdfs/energy-economics/ statistical-review/bp-stats-review-2019-full-report.pdf Hydro_E Hydro electricity TWh BP, Statistical Review of World Energy https://www.bp.com/en/global/corporate/ energy-economics/statistical-review-of-world-energy.html Biomass Electricity produced by biomass TWh BP. https://www.bp.com/en/global/corporate/energy-economics/ statistical-review-of-world-energy/downloads.html Continued on next page
energy forecasting.focus:natural gas 135 Table 2– continued from previous page Variable Name Meaning Unit Source and Notes Solar Electricity produced by solar energy TWh BP. https://www.bp.com/en/global/corporate/energy-economics/ statistical-review-of-world-energy/downloads.html Wind Electricity produced from wind energy TWh BP. https://www.bp.com/en/global/corporate/energy-economics/ statistical-review-of-world-energy/downloads.html Coal_Con Coal consumption MTOE BP. https://www.bp.com/en/global/corporate/energy-economics/ statistical-review-of-world-energy/downloads.html Nuclear Electricity from nuclear power plants TWh BP. https://www.bp.com/en/global/corporate/energy-economics/ statistical-review-of-world-energy/downloads.html NG_Con Natural gas consumption (Germany) Billion Cubic Meters BP. https://www.bp.com/en/global/corporate/energy-economics/ statistical-review-of-world-energy/downloads.html
7 Bibliography Akpinar, M. and Yumusak, N. (2016). Year ahead demand forecast of city natural gas using seasonal time series methods. Energies, 9(9):727. Al-Fattah, S. and Startzman, R. (1999). Analysis of world natural gas production. Conference: SPE Eastern Regional Meeting, paper SPE57463. Charleston, West Virginia, USA. https://doi.org/10.2118/ 57463-MS Last accessed 2021-03-26. Altfeld, K. and Pinchbeck, D. (2013). Admissible hydrogen concentrations in natural gas systems. Gas for Energy, (03). https://www.gerg. eu/wp-content/uploads/2019/10/HIPS_Final-Report.pdf Last accessed 2020-12-19. Anagnostis, A., Papageorgiou, E., Dafopoulos, V., and Bochtis, D. D. (2019). Applying long short-term memory networks for natural gas demand prediction. In Bourbakis, N. G., Tsihrintzis, G. A., and Virvou, M., editors, 10th International Conference on Information, Intelligence, Systems and Applications, IISA 2019, Patras, Greece, July 1517,2019, pages 1–7. IEEE. https://doi.org/10.1109/IISA.2019. 8900746 Last accessed 2020-08-31. Anastasiadis, A. D., Magoulas, G. D., and Vrahatis, M. N. (2005). New globally convergent training scheme based on the resilient propagation algorithm. Neurocomputing,64:253–270.http:// www.sciencedirect.com/science/article/pii/S0925231204005168 Last accessed 2020-09-06. Andrade-Cabrera, C., de Rosa, M., Kathirgamanathan, A., Kapetanakis, D.-S., and Finn, D. (2018). A study on the tradeoff between energy forecasting accuracy and computational complexity in lumped parameter building energy models. The 10th Canada Conference of International Building Performance Simulation Association (eSim 2018), Montreal, Canada, 9-10 May 2018. http://hdl.handle.net/10197/9663 Last accessed 2019-03-22.
energy forecasting.focus:natural gas 144 Hong, T. (06/30/2014). Global energy forecasting competition. past, present and future. https://forecasters.org/wp-content/ uploads/gravity_forms/7-2a51b93047891f1ec3608bdbd77ca58d/ 2014/07/HONG_TAO_ISF2014.pdf Last accessed 2020-12-20. Hong, T., Wilson, J., and Xie, J. (2014). Long term probabilistic load forecasting and normalization with hourly information. IEEE Transactions on Smart Grid,5(1):456–462.https://doi.org/10.1109/TSG. 2013.2274373. Hong, W. S. (2013). Intelligent energy demand forecasting, volume 10 of Lecture notes in energy. Springer, London and Heidelberg and New York and Dordrecht. ISBN 9781447149682. Hribar, R., Potoˇcnik, P., Šilc, J., and Papa, G. (2019). A comparison of models for forecasting the residential natural gas demand of an urban area. Energy,167:511–522. Hyndman, R. (2017). nnetar. Neural Network Time Series Forecasts. https://www.rdocumentation.org/packages/\protect\ discretionary{\char\hyphenchar\font}{}{}forecast/versions/ 8.13/topics/nnetar Last accessed 2021-01-10. Hyndman, R. J. and Khandakar, Y. (2008). Automatic time series forecasting: The forecast package for r. Journal of Statistical Software, 27(1):1–22. Hyndman, R. J. and Koehler, A. B. (2006). Another look at measures of forecast accuracy. International Journal of Forecasting,22(4):679–688. Hyndman, R. J. and Kostenko Andrey V. (2007). Minimum sample size requirements for seasonal forecasting models. Foresight, (6):12– 15.https://robjhyndman.com/papers/shortseasonal.pdf. International Energy Agency (2020). Germany 2020. Energy Policy Review. https://www.bmwi.de/Redaktion/DE/ Downloads/G/germany-2020-energy-policy-review.pdf?__blob= publicationFile&v=4 Last accessed 2020-12-27. International Energy Agency (2021). Energy security. https://www. iea.org/topics/energy-security Last accessed 2021-01-09. Januschowski, T., Gasthaus, J., Wang, Y., Salinas, D., Flunkert, V., Bohlke-Schneider, M., and Callot, L. (2020). Criteria for classifying forecasting methods. International Journal of Forecasting,36(1):167– 177. Jefferson, M. (2016). Energy realities or modelling: Which is more useful in a world of internal contradictions? Energy Research & Social Science,22:1–6.https://doi.org/10.1016/j.erss.2016.08.006.
energy forecasting.focus:natural gas 145 Jensterle, M., Narita, J., Piria, R., Samadi, S., Prantner, M., Crone, K., Siegemund, S., Kan, S., Matsumoto, T., Shibata, Y., and Thesen, J. (2019). The role of clean hydrogen in the future energy systems of Japan and Germany. Berlin: adelphi. https://www.adelphi.de Last accessed 2020-12-28. Kaposty, F., Kriebel, J., and Löderbusch, M. (2020). Predicting loss given default in leasing: A closer look at models and variable selection. International Journal of Forecasting,36(2):248–266.https: //doi.org/10.1016/j.ijforecast.2019.05.009. Karabiber, O. A. and Xydis, G. (2020). Forecasting day-ahead natural gas demand in Denmark. Journal of Natural Gas Science and Engineering,76:103193.https://doi.org/10.1016/j.jngse.2020.103193. Katz, J. H. (2020). Monitoring Forecast Models Using Control Charts. Foresight: The International Journal of Applied Forecasting, (56):20–25. Keane, M. and Runkle, D. E. (1990). Testing the rationality of price forecasts: New evidence from panel data. American Economic Review, 80(4):714–35.https://EconPapers.repec.org/RePEc:aea:aecrev: v:80:y:1990:i:4:p:714-35. Keuzenkamp, H. A. and McAleer, M. (1997). The complexity of simplicity. Mathematics and Computers in Simulation,43(3-6):553–561. Kuhn, M. and Johnson, K. (2016). Applied predictive modeling. Springer, New York, Corrected 5th printing edition. ISBN 978-1-4614-6849-3. Lewis, C. (1982). Industrial and business forecasting methods. A practical guide to exponential smoothing and curve fitting. Butterworth, London. ISBN 0408005599. Li, L. and Abu-Mostafa, Y. (2006). Data complexity in machine learning. https://resolver.caltech.edu/CaltechCSTR:2006.004 Computer Science Technical Reports. California Institute Of Technology, Pasedena, USA. Unpublished. Li, Z. and Zhang, Y.-K. (2008). Multi-scale entropy analysis of Mississippi river flow. Stochastic Environmental Research and Risk Assessment, 22(4):507–512.https://doi.org/10.1007/s00477-007-0161-y. Lichtendahl, K. C. and Winkler, R. L. (2020). Why do some combinations perform better than others? International Journal of Forecasting,36(1):142–149.https://doi.org/10.1016/j.ijforecast.2019. 03.027. Lindley, D. V. (2001). The philosophy of statistics. Journal of the Royal Statistical Society: Series D (The Statistician),49(3):293–337.https: //doi.org/10.1111/1467-9884.00238.
energy forecasting.focus:natural gas 146 Liu, S., editor (2011). Proceedings of 2011 IEEE International Conference on Grey Systems and Intelligent Services (GSIS) with the 15th WOSC International Congress on Cybernetics and Systems, Nanjing, China, 15 - 18 September 2011, Piscataway, NJ. IEEE. Livni, R., Shalev-Shwartz, S., and Shamir, O. (2014). On the computational efficiency of training neural networks. https://arxiv.org/ pdf/1410.1141 Last accessed 2020-12-20. Lund, H., Arler, F., Østergaard, P., Hvelplund, F., Connolly, D., Mathiesen, B., and Karnøe, P. (2017). Simulation versus optimisation: Theoretical positions in energy system modelling. Energies,10(7):840. https://doi.org/10.3390/en10070840. Ma, Y. and Li, Y. (2010). Analysis of the supply-demand status of China’s natural gas to 2020.Petroleum Science, 7(1):132–135.https://link.springer.com/content/pdf/10.1007/ s12182-010-0017-9.pdf. Magdowski, A. and Kaltschmitt, M. (2017). Dependency of forecasting accuracy for balancing power supply by weather-dependent renewable energy sources. Jordan Journal of Mechanical and Industrial Engineering,11:209–215. Makridakis, S. (1993). Accuracy measures: theoretical and practical concerns. International Journal of Forecasting,9(4):527–529.https: //doi.org/10.1016/0169-2070(93)90079-3. Makridakis, S. (1995). Forecasting accuracy and system complexity. RAIRO - Operations Research - Recherche Opérationnelle,29(3):259–283. http://www.numdam.org/item/RO_1995__29_3_259_0/. Makridakis, S., Hyndman, R. J., and Petropoulos, F. (2020). Forecasting in social settings: The state of the art. International Journal of Forecasting,36(1):15–28.https://doi.org/10.1016/j.ijforecast. 2019.05.011. Makridakis, S., Spiliotis, E., and Assimakopoulos, V. (2018). The M4 competition: Results, findings, conclusion and way forward. International Journal of Forecasting,34(4):802–808. Makridakis, S. G., Wheelwright, S. C., and Hyndman, R. J. (1998). Forecasting: Methods and applications. Wiley, Hoboken, NJ, 3. ed. edition. ISBN 9780471532330. Mannes, A., Soll, J., and Larrick, R. (2014). The wisdom of select crowds. Journal of personality and social psychology,107:276–299.
energy forecasting.focus:natural gas 147 May, R. M. (1976). Simple mathematical models with very complicated dynamics. Nature,261(5560):459–467. Melikoglu, M. (2013). Vision 2023: Forecasting Turkey’s natural gas demand between 2013 and 2030.Renewable and Sustainable Energy Reviews,22:393–400. Merkel, G., Povinelli, R., and Brown, R. (2018). Short term load forecasting of natural gas with deep neural network regression. Energies, 11(8). Merkel G. D. et al. (2017). Deep neural network regression for shortterm load forecasting of natural gas. Proceedings of the International Symposium on Forecasting. https://epublications.marquette.edu/ electric_fac/287 Electrical and Computer Engineering Faculty Research and Publications. 287. Mevik, B.-H. and Cederkvist, H. R. (2004). Mean squared error of prediction (MSEP) estimates for principal component regression (PCR) and partial least squares regression (PLSR). Journal of Chemometrics, 18(9):422–429. Miller, A. (2006). Gazprom. Strategy for the energy sector leadership. Speech by Alexey Miller at the Annual General Shareholders Meeting . https://www.gazprom.com/press/news/miller-journal/ 2006/99967/ Last accessed 2021-03-25. Ministry for Economic Affairs, F. and Energy (2020). The National Hydrogen Strategy. https://www.bmbf.de/files/bmwi_Nationale$% $20Wasserstoffstrategie_Eng_s01.pdf Last accessed 2020-12-28. Müller-Syring, G., Henel, M., Henel, M., Köppel, W., Mlaker, H., Sterner, M., and Höcher, T. (2013). DVGW. Entwicklung von modularen Konzepten zur Erzeugung, Speicherung und Einspeisung von Wasserstoff und Methan ins Erdgasnetz. https://www.dvgw. de/medien/dvgw/forschung/berichte/g1_07_10.pdf Last accessed 2019-07-25. Nuttall, W. J. and Manz, D. L. (2008). A new energy security paradigm for the twenty-first century. Technological Forecasting and Social Change,75(8):1247–1259. Nymoen Strategieberatung (2016). Strategische Marktprognose Erdgas. https://zukunft.erdgas.info/fileadmin/public/PDF/ Politischer_Rahmen/marktprognose-erdgas-2016.pdf Last accessed 2019-12-27. Overholt, W. H. (2000). Forecasting: An Appraisal for Policymakers and Planners. Policy Sciences,33(1):101–106.
energy forecasting.focus:natural gas 148 Pawlik V. (2019). Gasanschluss im Haushalt in Deutschland 2018 | Statista. https://de.statista.com/statistik/daten/studie/ 181631/umfrage/gasanschluss-im-haushalt-vorhanden/ Last accessed 2019-11-09. Pelikán, E. (2014). Forecasting of processes in complex systems for real-world problems. Tutorial. Neural Network World,24(6):567–589. http://www.nnw.cz/doi/2014/NNW.2014.24.032.pdf. Petropoulos, F., Hyndman, R. J., and Bergmeir, C. (2018). Exploring the sources of uncertainty: Why does bagging for time series forecasting work? European Journal of Operational Research,268(2):545–554. Pincus, S. (2008). Approximate entropy as an irregularity measure for financial data. Econometric Reviews,27(4-6):329–362.https://doi. org/10.1080/07474930801959750. Pincus, S. M. (1991). Approximate entropy as a measure of system complexity. Proceedings of the National Academy of Sciences of the United States of America,88(6):2297–2301. Potoˇcnik, P., Soldo, B., Šimunovi´c, G., Šari´c, T., Jeromen, A., and Govekar, E. (2014). Comparison of static and adaptive models for shortterm residential natural gas forecasting in Croatia. Applied Energy, 129:94–103. Prehn, O. (11/08/2018). Ten steps to improvise jazz. https://www. youtube.com/watch?v=BgvAq1tTnN4&t=1778s Last accessed 2020-0512. Pulhan, A., Yorucu, V., and Sinan Evcan, N. (2020). Global energy market dynamics and natural gas development in the Eastern Mediterranean region. Utilities Policy,64:101040. R Documentation (2020). Holt-Winters function | R documentation. https://www.rdocumentation.org/packages/stats/ versions/3.6.2/topics/HoltWinters Last accessed 2020-09-06. Ramesh, K. S. (2011). The deterministic chaos in heart rate variability signal and analysis techniques. International Journal of Computer Applications,35:39–46. Raza, M. Q. and Khosravi, A. (2015). A review on artificial intelligence based load demand forecasting techniques for smart grid and buildings. Renewable and Sustainable Energy Reviews,50:1352–1372. Richman, J. S. and Moorman, J. R. (2000). Physiological time-series analysis using approximate entropy and sample entropy. American Journal of Physiology. Heart and circulatory physiology,278(6):H2039– 49.
energy forecasting.focus:natural gas 149 Rieck, B. A. (2017). Persistent Homology in Multivariate Data Visualization. PhD thesis, Heidelberg University Library. http://archiv.ub. uni-heidelberg.de/volltextserver/22914/1/Dissertation.pdf Last accessed 2021-04-13. Rácz, L. and Németh, B. (2018). Investigation of dynamic electricity line rating based on neural networks. Energetika,64(2). Scarpa, F. and Bianco, V. (2017). Assessing the quality of natural gas consumption forecasting: An application to the Italian residential sector. Energies,10(11):1879. Scher, S. and Messori, G. (2019). Weather and climate forecasting with neural networks : using general circulation models (GCMs) with different complexity as a study ground. Geoscientific Model Development, 12(7):2797–2809.https://doi.org/10.5194/gmd-12-2797-2019. Semenychev, V. K., Kurkin, E. I., and Semenychev, E. V. (2014). Modelling and forecasting the trends of life cycle curves in the production of non-renewable resources. Energy,75:244–251. Sen, D., Günay, M. E., and Tunç, K. M. (2019). Forecasting annual natural gas consumption using socio-economic indicators for making future policies. Energy,173:1106–1118. Sethna, P. James (2020). Entropy, Order Parameters, and Complexity. Clarendon Press, Oxford. http://pages.physics.cornell.edu/ ~sethna/StatMech/EntropyOrderParametersComplexity20.pdf Last accessed 2021-04-13. Shaikh, F. and Ji, Q. (2016). Forecasting natural gas demand in China: Logistic modelling analysis. International Journal of Electrical Power & Energy Systems,77:25–32. Sharmina, M., Abi Ghanem, D., Browne, A. L., Hall, S. M., Mylan, J., Petrova, S., and Wood, R. (2019). Envisioning surprises: How social sciences could help models represent ‘deep uncertainty’ in future energy and water demand. Energy Research & Social Science,50:18– 28.https://doi.org/10.1016/j.erss.2018.11.008. Siemek, J., Nagy, S., and Rychlicki, S. (2003). Estimation of natural gas consumption in Poland based on the logistic-curve interpretation. Applied Energy,75(1-2):1–7. Smil, V. (2000). Perils of long-range energy forecasting. Technological Forecasting and Social Change,65(3):251–264. Soldo, B. (2012). Forecasting natural gas consumption. Applied Energy, 92:26–37.https://doi.org/10.1016/j.apenergy.2011.11.003.
energy forecasting.focus:natural gas 150 Sovacool, B. K., Ryan, S. E., Stern, P. C., Janda, K., Rochlin, G., Spreng, D., Pasqualetti, M. J., Wilhite, H., and Lutzenhiser, L. (2015). Integrating social science in energy research. Energy Research & Social Science,6:95–99. Su, H., Zio, E., Zhang, J., Xu, M., Li, X., and Zhang, Z. (2019). A hybrid hourly natural gas demand forecasting method based on the integration of wavelet transform and enhanced deep-RNN model. Energy,178:585–597. Szoplik, J. (2015). Forecasting of natural gas consumption with artificial neural networks. Energy,85:208–220. Tamba, J. G., Essiane, S. N., Sapnken, E. F., Koffi, F. D., Nsouandélé, J. L., Soldo, B., and Njomo Donatien (2018). Forecasting natural gas: A literature survey. International Journal of Energy Economics and Policy, (8(3)):216–249. Tashman, L. (Winter 2018). Beware of standard prediction intervals for causal models. Foresight: The International Journal of Applied Forecasting,48:43–48. Theocharis, Z. and Harvey, N. (2019). When does more mean worse? Accuracy of judgmental forecasting is nonlinearly related to length of data series. Omega,87:10–19. Tongal, H. and Berndtsson, R. (2017). Impact of complexity on daily and multi-step forecasting of streamflow with chaotic, stochastic, and black-box models. Stochastic Environmental Research and Risk Assessment,31(3):661–682. Villar, J. A. and Joutz, F. L. (2006). The relationship between crude oil and natural gas prices. http://aceer.uprm.edu/pdfs/CrudeOil_ NaturalGas.pdf Last accessed 2021-04-13. Wang, Z., Li, Y., Feng, Z., and Wen, K. (2019). Natural gas consumption forecasting model based on coal-to-gas project in China. Global Energy Interconnection,2(5):429–435. Wearden, G. (2019). Oil prices spike after Saudi drone attack causes biggest disruption ever – as it happened. https://www.theguardian.com/business/live/2019/sep/16/ oil-price-saudi-arabia-iran-drone-markets-ftse-pound-brexit\ business-live Last accessed 2021-04-13. Wei, N., Li, C., Peng, X., Li, Y., and Zeng, F. (2019). Daily natural gas consumption forecasting via the application of a novel hybrid model. Applied Energy,250:358–368.
energy forecasting.focus:natural gas 151 Werbos, P. J. (1988). Generalization of backpropagation with application to a recurrent gas market model. Neural Networks,1(4):339–356. Wiener, N. (2013). Cybernetics or control and communication in the animal and the machine. MIT Press, Cambridge, Mass., 2. ed. edition. ISBN 9781614275022. Winebrake, J. J. and Sakva, D. (2006). An evaluation of errors in US energy forecasts: 1982–2003.Energy Policy,34(18):3475–3483. Yalcinoz, T. and Eminoglu, U. (2005). Short term and medium term power distribution load forecasting by neural networks. Energy Conversion and Management,46:1393–1405. Yaniv, I. and Foster, D. P. (1995). Graininess of judgment under uncertainty: An accuracy-informativeness trade-off. Journal of Experimental Psychology: General,124(4):424–432. Yu, F. and Xu, X. (2014). A short-term load forecasting model of natural gas based on optimized genetic algorithm and improved BP neural network. Applied Energy,134:102–113.https://doi.org/10.1016/ j.apenergy.2014.07.104. Zambella, D. and Grassberger, P. (1988). Complexity of forecasting in a class of simple models. Complex Syst.,2(3). Zellner, A., Keuzenkamp, H. A., and McAleer, M. (2002). Simplicity, inference and modeling: Keeping it sophisticatedly simple. Cambridge University Press, Cambridge and New York. ISBN 0521803616. Zhang, W. and Yang, J. (2015). Forecasting natural gas consumption in China by Bayesian Model Averaging. Energy Reports,1:216–220. https://doi.org/10.1016/j.egyr.2015.11.001. ˇ Ceperi´c, E., Žikovi´c, S., and ˇ Ceperi´c, V. (2017). Short-term forecasting of natural gas prices using machine learning and feature selection algorithms. Energy,140:893–900. Šebalj, Dario and Dujak Davor and Mesaric Josip (2017). Predicting natural gas consumption - a literature review. http://archive. ceciis.foi.hr/app/public/conferences/2017/08/SPDM-2.pdf. Šumskis, V. and Giedraitis, V. (2015). Economic implications of energy security in the short run. Ekonomika,94(3):119.https://doi.org/ 10.15388/Ekon.2015.3.8791.