Non-extensive entropy econometrics for low frequency series: National accounts-based inverse problems
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Bwanakare, Second Book — Published Version Non-extensive entropy econometrics for low frequency series: National accounts-based inverse problems Provided in Cooperation with: De Gruyter Brill Suggested Citation: Bwanakare, Second (2017) : Non-extensive entropy econometrics for low frequency series: National accounts-based inverse problems, ISBN 978-3-11-055044-3, De Gruyter, Berlin, https://doi.org/10.1515/9783110550443 This Version is available at: https://hdl.handle.net/10419/182286 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/4.0/
Second Bwanakare Non-Extensive Entropy Econometrics for Low Frequency Series National Accounts-Based Inverse Problems
Second Bwanakare Non-Extensive Entropy Econometrics for Low Frequency Series National Accounts-Based Inverse Problems Managing Editor: Maria Laura Parisi Language Editor: Naguib Lallmahomed
Published by De Gruyter Open Ltd, Warsaw/Berlin Part of Walter de Gruyter GmbH, Berlin/Boston The book is published with open access at www.degruyter.com. This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 License. For details go to http://creativecommons.org/licenses/by-nc-nd/4.0/. Library of Congress Cataloging-in-Publication Data A CIP catalog record for this book has been applied for at the Library of Congress. Copyright © 2017 Second Bwanakare ISBN: 978-3-11-055043-6 e-ISBN: 978-3-11-055044-3 Bibliographic information published by the Deutsche Nationalbibliothek. The Deutsche Nationalbibliothek lists this publication in the Deutsche Nationalbibliografie; detailed bibliographic data are available in the Internet at http://dnb.dnb.de. Managing Editor: Maria Laura Parisi Language Editor: Naguib Lallmahomed www.degruyteropen.com Cover illustration: © Thinkstock, Credit: liuzishan
Contents Acknowledgements ix Summary x PART I: Generalities and Scope of the Book 1 Generalities 2 1.1 Information-Theoretic Maximum Entropy Principle and Inverse Problem 2 1.1.1 Information-Theoretic Maximum Entropy Principle 2 1.2 Motivation of the Work 4 1.2.1 Frequent Limitations of Shannon-Gibbs Maximum Entropy Econometrics 4 1.2.2 Rationale of Pl-Related Tsallis Entropy Econometrics and Low Frequency Series 6 1.3 National Accounts-Related Models and the Scope of this Work 10 Bibliography – Part I 11 PART II: Statistical Theory of Information and Generalised Inverse Problem 1 Information and its Main Quantitative Properties 16 1.1 Definition and Generalities 16 1.2 Main Quantitative Properties of Statistical Information 19 2 Ill-posed Inverse Problem Solution and the Maximum Entropy Principle 22 2.1 Introduction 22 2.2 The Inverse Problem and Socio-Economic Phenomena 22 2.2.1 Moore-Penrose Pseudo-Inverse 23 2.2.2 The Gibbs-Shannon Maximum Entropy Principle and the Inverse Problem 25 2.2.3 Kullback-Leibler Cross-Entropy 27 2.3 General Linear Entropy Econometrics 28 2.3.1 Reparametrization of Parameters 28 2.4 Tsallis Entropy and MainProperties 29 2.4.1 Definition and Shannon-Tsallis Entropy Relationships 29 2.4.2 Characterization of Non-Extensive Entropy 31 2.4.2.1 Correlations 31 2.4.2.2 Concavity 32 2.4.3 Tsallis Entropy and Other Forms of Entropy 33
2.4.3.1 Characterization 34 2.4.3.2 Scale of q-Tsallis Index and its Interpretation 35 2.5 Kullback-Leibler-Tsallis Cross-Entropy 35 2.5.1 The q-Generalization of the Kullback-Leibler Relative Entropy 35 2.5.2 Tsallis Versions of the Kullback-Leibler Divergence in Constraining Problems 37 2.6 A Generalized Linear Non-Extensive Entropy Econometric Model 38 2.6.1 A General Model 38 2.6.2 Parameter Confidence Interval Area 40 2.7 An Application Example: a Maximum Tsallis Entropy Econometrics Model for Labour Demand 42 2.7.1 Theoretical Expectation Model 42 2.7.2 A Generalized Non-Extensive Entropy Econometric Model 43 2.7.2.1 General Model 43 2.7.3 Estimated Confidence Area Of Parameters 44 2.7.4 Data and Model Outputs 45 Annex A 48 Annex B: Independence of Events Within q-Generalized Kullback-Leibler Relative Entropy 49 Bibliography – Part II 50 Part III: Updating and Forecasting Input-Output Transaction Matrices 1 Introduction 54 2 The System of National Accounts 55 3 The Input-Output (IO) Table and its Main Application 56 3.1 The I-O Table and Underlying Coefficients 56 3.2 Input-Output Multipliers 58 3.2.1 The principal models 58 3.2.2 A Model of Recovering the Sectorial Greenhouse Gas Emission Structure 61 3.3 Updating and Forecasting I-O Tables 64 3.3.1 Generalities 64 3.3.2 RAS Formalism and its Limits 65 3.3.3 Application: Updating an Aggregated EU I-O Matrix 68 3.3.3.1 The RAS Approach 68 3.3.3.2 The Entropy Approach 72 3.3.3.3 I-O table Forecasting 73 3.3.3.4 The Non-Extensive Cross-Entropy Approach and I-O Table Updating 75 3.3.3.5 Forecasting an I-O Table 77
3.3.4 Application: Forecasting the Aggregated 27 EU IO Coefficients 79 3.3.5 Emission Coefficients Forecasting: A Theoretical Model 81 3.4 Conclusions 83 Bibliography – Part III 83 PART IV: Social Accounting Matrix 1 Position of the Problem 88 2 A SAM as a Walrasian Equilibrium Framework 90 3 The Social Accounting Matrix (SAM) Framework 93 3.1 Generalities 93 3.2 Description of a Standard SAM 94 4 Balancing a SAM 97 4.1 Shannon-Kullback-Leibler Cross-Entropy 98 4.2 Balancing a SAM Through Tsallis-Kullback-Leibler Cross-Entropy 100 4.3 A SAM as a Generalized Input-Output System 101 4.3.1 A Generalized Linear Non-Extensive Entropy Econometric Model 102 4.4 Input-Output Power Law (Pl) Structure 105 4.5 Balancing a SAM of a Developing Country: the Case of the Republic of Gabon 107 4.5.1 Balancing the SAM of GABON by Tsallis Cross-Entropy Formalism 108 4.6 About the Extended SAM 111 5 A SAM and Multiplier Analysis: Economic Linkages and Multiplier Effects 117 5.1 What are the Economic Linkages and Multiplier Effects? 117 5.1.1 A SAM Unconstrained Multiplier 119 5.1.2 Equation System for Constrained SAM Multiplier 121 5.1.3 On Modelling Multiplier Impact for an Ill-Behaved SAM 122 Annex C. Proof of Economy Power Law Properties 122 Bibliography – Part IV 129 PART V: Computable General Equilibrium Models 1 A Historical Perspective 134 2 The CGE Model Among Other Models 137 3 Optimal Behaviour And The General Equilibrium Model 138 3.1 Introduction 138 3.2 Economic Efficiency Prerequisites for a Pareto Optimum 140 4 From a SAM to a CGE Model: a Cobb-Douglas Economy 143 4.1 CGE Building Steps 143 4.2 The Standard Model of Transaction Relationships 146
5 Estimating the CGE Model Through the Maximum Entropy Principle 151 5.1 Introduction 151 5.2 Estimation Approach 153 5.3 Application: Non-Extensive Entropy and Constant Elasticity of Substitution-Based Models 158 5.3.1 Power Law and the Constant Elasticity of Substitution (CES) Function 159 5.3.2 Parameter Outputsof the Tsallis Relative Entropy Model 163 5.4 Conclusions 169 Bibliography – Part V 169 Part VI: From Equilibrium to Real World Disequilibrium: An Environmental Model 1 Introduction 176 2 Extending to an Environmental Model 179 2.1 The Model of Carbon Tax 181 2.2 Carbon Tax Model And Double-Dividend Hypothesis 183 3 Compensatory and Equivalent Variations: Two Types of Welfare Measurement 185 Bibliography – Part VI 189 1 Concluding Remarks 193 Appendix 195 Annex C. Computational Aspects of Using GAMS 195 1.1 Introduction to GAMS 195 1.2 Formulation of a Simple Linear Problem in GAMS 195 1.3 Application to Maximum Entropy Models 196 1.3.1 The Jaynes Unbalanced Dice Problem 196 1.3.2 Nonlinear Non Extensive Entropy Econometric Model 197 Annex D. Recovery of Pollutant Emissions by Industrial Sector and Region: an Instructional Case 201 1.1 Introduction 201 1.2 Recovery of Pollutant Emission by Industrial Sector and Region: the Role of the a Priori 201 Index of Subject 205 Index of Authors 206
Information-Theoretic Maximum Entropy Principle and Inverse Problem 3 (Halmos & Savage, 1979), a generalization of Cramer-Rao inequality (e.g., Kullback, 1959) and the introduction of a general linear model as a consistency restriction (Heckelei et al., 2008) through Bayesian philosophy. Thus, it became possible to unify heterogeneous statistical procedures via the concepts of information theory. Lindley (2008), on the other hand, had provided the interpretation that a statistical sample could be viewed as a noisy channel (Shannon’s terminology) that conveys a message about a parameter (or a set of parameters) with a certain prior distribution. This new interpretation extended application of Shannon’s ideas to statistical theory by referring to the information in a statistical sample rather than in a message. Over the last two decades the literature concerned with applying entropy in social science has grown considerably and disserves closer attention. On one side, Shannon-Jaynes-Kullback-Leibler-based approaches are currently used for modelling economic phenomena competitively with classical econometrics. A new paradigm in econometrical modelling is taking place and finds its roots in the influential work of Golan, Judge, and Miller (1996). The present monograph constitutes an illustration of this. As mentioned above, this approach is particularly useful in the case of solving inverse problems or ill-behaved matrices when we try to estimate parameters of an econometric model on the basis of insufficient information from an observed sample, and this estimation may concern the behaviour of an individual element within the system. Insufficient information implies that we are trying to solve an ill-posed problem, which plausibly can arise in the following cases: — data from sampling design are not sufficient and/or complete due to technical or financial limitations—small area official statistics could illustrate this situation; — non-stationary or non-co-integrating variables are resulting from bad model specification; — data from the statistical sample are linearly dependent or collinear for various reasons; — Gaussian properties of random disturbance are put into question due to, amongst many others things3, systematic errors from the survey process; — the model is not linear and approximate linearization remains the last possibility; — aggregated (in time or space) data observations hide a very complex system represented, for instance, by a PL distribution, and multi-fractal properties of the system may exist. 3 It is not excluded that distribution law may be erroneously applied since, for instance, randomness is dependent on the experimental setup or the sophistication of the apparatus involved in measuring the phenomenon (Smith, 2001).
4 Generalities Using the traditional econometrical approaches in one or more of the above cases— without additional simplifying hypotheses—could lead to various estimation problems owing to the nonexistence of a bounded solution or the instability of estimates. Consequently, outputs from traditional econometrical approaches will display, at best, poor informative parameters. In the literature, there are other well-known techniques to cope with inverse problems or ill-conditioned data. Among them, two popular techniques deserve our attention: the bi-proportional RAS approach (and its variants), particularly used for updating or forecasting input/output matrices (Parikh, 1979) and the Moore-Penrose pseudo-inverse technique, useful for inverting irregular matrices (e.g., Green, 2003, p. 833). In spite of their popularity, both techniques present serious drawbacks in empirical investigations. In fact, the RAS techniques, in spite of their divergence information nature, remain less adapted to solving stochastic problems or to optimizing the information criterion function under a larger number of different prior constraining data. Since Moore-Penrose generalized inverse ensures a minimum distance (Y-BX) only when the matrix B has full rank, it will not reflect an optimal solution in other cases. Golan et al. (1996) have clearly shown higher efficiency of Shannon maximum entropy econometrics over the above cited methods in recovering unknown information when data or model design is poorly conditioned. The suggested superiority stands on the fact that it combines and generalizes maximum entropy philosophy (as in the second law of thermodynamics) and statistical theory of information attributes as a Bayesian information processing rule. As demonstrated convincingly by Golan (1996, 2006), Shannon entropy econometrics formalism may generalize least squares (LS) and the maximum likelihood (ML) approaches and belongs to the class of Bayesian method of moments (BMOM). It is worthwhile to point out that in the coming chapters many cases of cross-entropy (or minimum entropy) formalism will be used in place of maximum entropy. This is because, in this study, many problems to be treated involve information measuring in the context of the Kullback-Leibler framework. This monograph does not intend to treat the case of high frequency series for which a rich literature already exists. We invite readers interested in the case of high frequency series to see, for instance, J.W. Kantelhardt (2008) for testing for the existence of fractal or multi-fractal properties, suggesting the case of a PL distribution. 1.2 Motivation of the Work 1.2.1 Frequent Limitations of Shannon-Gibbs Maximum Entropy Econometrics In spite of a growing interest in the research community, some incisive critics have come forward to address Shannon-based entropy econometrics (e.g., Heckelei et al., 2008). According to some authors, generalized maximum entropy (GME) or cross-
Motivation of the Work 5 entropy (GCE) econometrical techniques face at least three difficulties. The first is related to the specification and interpretation of prior information, imposed via the use of discrete support points, and assigning prior probabilities to them. The authors argue that there are complications that result from the combination of priors and their interaction with the criterion of maximum entropy or minimum cross-entropy in determining the final estimated a posteriori probabilities on the support space. The second group of criticisms questions the sense of the entropy objective function once combined with the prior and data information. The last problem, according to the same authors, refers to computational difficulties owing to the mathematical complexity of the model with an unnecessarily large number of parameters or variables. Concerning the first criticism, the problem—selecting a prior support space and prior probabilities on it—exists since estimation outputs seem to be extremely sensitive to initial conditions. However, when there is a theory or some knowledge about the space on which parameters are supposed to be staying, the problem becomes tractable. In particular, when we have to estimate parameters in the form of ratios, the performance of entropy formalism is high. To this counterargument, it is worthwhile to add that GME or GCE formalism constitutes an approach based on the Bayesian efficient processing rule and, as such, prior values are not fixed constraints of the model; they combine and adapt with respect to other sets of information (e.g., consistency function) added to the model to update a new parameter level in the entropy criterion function. The second problem concerns questioning the sense or interpretability of output probabilities from the maximum entropy criterion function once combined with real world probability-related restrictions. One cannot comment on this problem without making reference to the important contribution of Jaynes (1957, 1957b), who proposed a way to estimate unknown probabilities of a discrete system in the presence of less data point observations than parameters to be estimated through the celebrated example of Jaynes dice. Given a set of all possible ways of distribution resulting from all micro-elements of a system, Jaynes proposed using the one that generates the most “uncertain”4 distribution. To understand this problem, the question becomes a matter of combining philosophical interpretation of the maximum entropy principle with that of Jaynes’ formulation in the context of Shannon entropy. Depending on the type of entropy5 considered, output estimates will have slightly different meaning. However, all interpretations refer to parameter values that assure a long-run, steady4 Here we are in the realm of the second law of thermodynamics, which stipulates, in terms of entropy, that natural equilibrium of any set of events is reached once disorder inside them becomes optimal. This results from their property of having equal (ergodic system) odds to occur. In that state, we reach the maximum uncertainty about which event should occur in the next trial. 5 Later, for comparison, properties of the most well-known types of entropy in the literature will be presented.
6 Generalities state equilibrium of the system (relations defined by the model) with respect to data and other knowledge at hand, usually in the form of moments and/or normalization conditions. Owing to maximum entropy alone, the more consistent moments are or the more other a priori information binds, the more output probabilities will differ from those in a uniform distribution. Considering the above, interpretation of the maximum entropy model is far removed from interpretation of the classical model, especially in the case of the econometric linear model where estimates mean a change in the endogenous variable due to unitary change in an explicative variable, that is, in ceteris paribus conditions. The last criticisms concern the burden arising from the computational and numerical process—a problem common to all complex, nonlinear systems. Thanks to recent developments of computer software, this problem is now less important. In many empirical studies that attempt to solve inverse problems, the Shannon entropy-based approach is relatively efficient in recovering information. However, gaining in parameter precision requires good design of the prior. In particular, the point support space must fit into the space of the true population parameter values. As Golan et al. (1996) have shown, when prior design is weak, outputs of Shannon entropy econometrics will produce approximately the same parameter precision as traditional econometrical methods, such as LS or the ML, which means Shannon entropy could discount information not fitting the maximum entropy principle as expected. The above criticisms of the Shannon entropy econometrics model remain relatively weak as has been shown through the preceding discussion. According to us, the main drawback related to that form of model is due to the analytical function of constraining moments. In fact, as already suggested, longrange correlation and observed time invariant scale structure of high frequency series may still be conserved—in some classes of non-linear models—through a time—or space—aggregation process of statistical data. This raises the question of why this study proposes a new approach of Tsallis non-extensive entropy econometrics. The next section provides a first answer by showing potential theoretical and then empirical drawbacks of the Shannon-Gibbs entropy model and potential advantages from the PL-related Tsallis non-extensive entropy approach. 1.2.2 Rationale of Pl-Related Tsallis Entropy Econometrics and Low Frequency Series This section presents the essence of the scientific contribution of this monograph to econometric modelling. For a few decades, PL has confirmed its central role in describing a large array of systems, natural and manmade. While most scientific fields have integrated this new element into their analytical approaches, econometrics and hence, economics globally, is still dwelling—probably for practical reasons—
Motivation of the Work 7 on the Gaussian fundamentals. This study takes a step forward by introducing Tsallis non-extensive entropy to low frequency series econometric modelling. The potential advantages of this new approach will be presented, in particular, its capacity to analytically solve complex PL-related functions. Since any mathematical function form can be represented by a PL formulation, the importance of the proposed approach becomes clear. To be concrete, one of the complex nonlinear models is the fractionally integrated moving average (ARFIMA) model, which, to our knowledge, has remained non-tractable using traditional statistical instruments. An empirical application to solve such a class of models will be implemented at the end of Part V of this book. According to several studies (Bottazzi & et al, 2007), (Champernowne, 1953), (Gabaix, 2008), a large array of economic laws take the form of a PL, in particular macroeconomic scaling laws, distribution of income, wealth, size of cities and firms6, and distribution of financial variables such as returns and trading volume. Ormerod and Mounfield (2012) underscore a PL distribution of business cycle duration. Stanley et al. (1998) have studied the dynamics of a general system composed of interacting units, each with a complex internal structure comprising many subunits, where the subunits grow in a multiplicative way over a period of twenty years. They found that this system followed a PL distribution. It is worthwhile to note the similarity of such a system with the internal mechanism of national accounts tables, such as a SAM, also composed of interacting economic sectors, each with a complex internal structure defined by firms exercising similar business. Ikeda and Souma (2008) have made an international comparison of labour productivity distribution for manufacturing and non-manufacturing firms. A PL distribution in terms of firms and sector productivity was found in US and Japanese data. Testing the Gibrat's law of proportionate effect, Fujiwara et al. (2004) have found, among others things, that the upper-tail of the distribution of firm size can be fitted with a PL (Pareto-Zipf law). The list of PL evidence here is limited to social science. Since this study focuses on the immense potentiality of PL-related economic models, PL ubiquity in the social sciences will be underscored and a theorem showing the PL character of national accounts in its aggregate form will be presented. In line with the rationale for the proposed methodology detailed below, the following from recent literature is evidence of entropy: – Non-extensive entropy, as such, models the non-ergodic systems which compound Levy7 instable phenomena8 converging in the long range to the Gaussian basin of attraction. In the limiting case, non-extensive entropy converges to Shannon Gibbs entropy. 6 See (Bottazzi & et al, 2007) for different standpoints on the subject. 7 Shlesinger (Shlesinger, Zaslavsky, & Klafter, Strange Kinetics, 1993) 8 (Shlesinger & et al, Lévy Flights and Related Topics in Physics, 1995).
8 Generalities – PL-related Tsallis entropy should remain, even in the case of a low frequency series, a precious device for econometric modelling since the outputs provided by the exponential family law (e.g., the Gibbs-Shannon entropy approach) correspond to the Tsallis entropy limiting case when the Tsallis-q parameter equals unity. – A number of complex phenomena involve long-range correlations which can be seen particularly when data are time scale-aggregated (Drożdż & Kwapień, 2012), (Rak & et al, 2007). This is probably because of the interaction between the functional relationships describing the involved phenomena and the inheritance properties of a PL or because of their nonlinearity. Delimiting the threshold values for a PL transition towards the Gaussian structure (or to the exponential family law) as a function of the data frequency amplitude is difficult since each phenomenon may display its own rate of convergence—if any—towards the central theorem limit attractor. – Systematic errors from statistical data collecting and processing may generate a kind of tail queue distribution. Thus, a systematic application of the ShannonGibbs entropy approach in the above cases—even on the basis of annual data— could be misleading. In the best case, it can lead to unstable solutions. – On the other hand, since non-extensive Tsallis entropy generalizes the exponential family law (Nielsen & Nock, 2012), the Tsallis-q entropy methodology fits well with high or low frequency series. In the class of a few types of entropy displaying higher-order entropy estimators able to generalize the Gaussian law, Tsallis non-extensive entropy has the valuable quality of concavity–and then stability—along the existence interval characterizing most real world phenomena. As far as the q-generalization of the Kullback-Leibler (K-L) relative entropy index is concerned, it conserves the same basic properties as the standard K-L entropy and can be used for the same purpose (Tsallis, 2009). The above-enumerated points imply that in cases where the assumed Levy law complexity is not verified by empirical observation, outputs from the non-extensive entropy model converge with those derived from Shannon entropy. In other words, errors which involve taking a sample as if it were PL-driven has no consequence on outputs if the truth model belongs to the Gaussian basin of attraction. This explains why in most empirical applications—but by no means all—both forms of entropy provide similar results and the entropic Tsallis-q complexity parameter then tends to converge to unity, revealing the case of a normal distribution. Empirical examples will be presented at the end of this document, and the strength of Tsallis maximum entropy econometrics will be demonstrated in different contexts. In summary, the following are entropy function regularities: – The Tsallis entropy model generalizes the Shannon-Gibbs model, which constitutes a converging case of the former for the Tsallis-q parameter equal unity.
Motivation of the Work 9 – The Shannon-Gibbs model fits natural or social phenomena displaying Gaussian properties. – PL high frequency time (space) series scaling—aggregating—does not always lead to Gaussian low frequency time (space) series. Additionally, the rate of convergence from the PL to the Gaussian model, if any, varies according to the form of the function used. Is it judicious to replace Shannon-Gibbs entropy modelling by Tsallis non-extensive entropy for empirical applications? The answer is yes, and this is the motivation for this study. There are at least three expected advantages to introducing Tsallis non-extensive econometric modelling: 1. A data generating system characterized by a low—or no—convergence rate from PL to Gaussian distribution only becomes analytically tractable when using Tsallis entropy formalism. (This will be proven through an econometrical model with constant substitution elasticity and then considered as an inverse problem to be estimated later.) 2. The Tsallis entropy model displays higher stability than the Shannon-Gibbs, particularly when systematic errors affect statistical data. 3. The Tsallis-q parameter presents an expected advantage of monitoring complexity of systems by measuring how far a given random phenomenon is from the Gaussian benchmark. In addition to other advantages, this can help draw attention to the quality of collected data or the distribution involved. The choice of national accounts-related models for testing the new approach of non-extensive entropy econometrics is motivated by the empirical inability of national systems of economic information to provide consistent data according to macroeconomic general equilibrium. As a result, national account tables are generally not balanced unless additional—often contradictory—assumptions are applied to balance them. However, following the principle of not adding (to a hypothetical truth) more than we know, it remains preferable to deal with an unbalanced national accounts table. Trying to balance such a table implies that we are faced with ill-behaved inverse problems. According to the existing literature, and as will be seen through this monograph, entropy formalism remains the best approach to solving such a category of complex problems. The superiority of Tsallis non-extensive entropy econometrics over other known econometrical or statistical procedures results from its capacity to generalize a large category of most known laws, including Gaussian distribution.
10 Generalities 1.3 National Accounts-Related Models and the Scope of this Work Under the high frequency series hypothesis, we postulate that social and economic activities are characterized by complex behavioural interactions between socio-economic agents and/or economic sectors. Recent, Big Data for Official Statistics may illustrate such a complexity. This could mean that the supposed extreme events may appear systematically more (or less) frequently than expected (Gaussian scheme), implying internal and aggregated long-range correlation (over time, space, or both). The maximum entropy principle is best suited to estimating ill-behaved inverse problems and, in particular, models with ratios or elasticity as parameters. In this latter case, as we will see later, the support space area for unknown parameters coincides with the probability area over the space from zero to unity. Fortunately enough, due to its macroeconomic consistency, national account table structure reflects this property. In empirical macroeconomic investigations, the national accounts system of information plays a crucial role for modelling as it guarantees internal coherence of macroeconomic relations. Numerical information is embodied inside comprehensive statistical tables or balance sheets displaying algebraic properties of a matrix. Having in mind an economic or statistical inference investigation, mathematical treatment of information compounded inside these matrices is carried out by economists or statisticians on the basis of a priori information at hand. When such matrices are algebraically regular, traditional inverse methods can be applied to solve the problem of, for instance, estimating parameters that define relationships between the endogenous variable and its covariates. Nevertheless, in the social sciences, causality relationships linking both variables seldom have a one-to-one correspondence. In many cases, two or more different inputs or causes can lead to the same output or effect. Such different causal concomitances for the same output render the social or economic model indeterminate. In such cases, the recovery of a data generating system from the observed finite sample becomes impossible using the traditional statistical or econometric devices, such as the standard maximum likelihood method or the generalized method of moments. On mathematical grounds, this may result from an insufficient number of model data points with respect to the number of parameters to estimate. Such a sample is said to be ill-behaved. This situation leads to the lack of an optimal solution sought. Collinear variables, inadequate size of a small sample, or the poor quality of statistical data may lead to the same difficulties. Finally, taking into account the above deficiencies and anomalies, modellers have to deal with illbehaved inverse problems most of the time. Following what has been said above, this monograph targets developing a robust approach generalizing Kullback-LeiblerShannon entropy for solving inverse problems related to national account models in a way that reflects the complex relationships between economic institutions and/ or agents. Statistical data from such complex interrelations are usually difficult to collect, incomplete, and defective. Additionally—and this may be one of the most important points—modelling national account table-related information involves
Bibliography – Part I 11 some class of nonlinear functions, otherwise only solvable using the PL model; thus, non-ergodic situations are involved. The next area of national accounts modelling to be treated in this monograph is: – Updating an input/output table when the problem is posed as inverse, with the possibility of adding extra sample information to the model in the form of an a priori and without any additional assumption; – Forecasting an input/output table or its extended forms, such as the social accounting matrix (SAM), solely on the basis of yearly published national accounts concerning sectorial elements of final demand and gross domestic product; – Deriving backward or forward multiplier coefficient impact on the basis of insufficient pieces of information; – Demonstrating a method to forecast a sectorial energy final demand and total pollutants emission by producton the basis of an environmentally extended input/ output table when basic information is missing; – Presenting a computable general equilibrium model using the maximum entropy approach instead of calibration techniques to derive the parameters of CES functions, – Estimating other nonlinear economic functions as inverse problems and conducting Monte Carlo experiments to test Tsallis entropy econometrics outputs; – Presenting in detail, across different chapters, national account-related general equilibrium models before coming back to inverse problem solution techniques as suggested above. The reader should be enriched not only by techniques for solving complex inverse problems but also by a thorough examination of different aspects of national account updating and modelling in the Walrasian spirit. To render the models presented here more consistent, emergent elements on an environmentally extended system of accounts will be included along with their impact on the general equilibrium framework and the optimum Pareto or social welfare. Bibliography – Part I Bayes, T. (1763). An essay towards solving a Problem in the Doctrine of Chances, Phil. Trans. 53, 370–418 (1763). http://rstl.royalsocietypublishing.org/content/53/370 Bottazzi, G., & et al. (2007). Invariances and Diversities in the Patterns of Industrial Evolution: Some Evidence from Italian Manufacturing Industries. Small Business Economics 29, pp. 137–159. Bregman, L.M. (1967). The relaxation method of finding the common points of convex sets and its application to the solution of problems in convex programming. Computational Mathematics and Mathematical Physics 7(3), pp. 200–217. Champernowne, D.G. (1953). A Model of Income Distribution. The Economic Journal 63(250), pp.318–351. Cox, R.T. (1946). Probability, Frequency, and Reasonable Expectation. Am. Jour. Phys., 14, pp. 1–13.
12 Generalities Douglas P. et al. (2006). Tunable Tsallis Distributions in Dissipative Optical Lattices. Phys. Rev. Lett. 96, 110601. Drożdż, S., & Kwapień, J. (2012). Physical approach to complex systems. Physica reports, pp.115–226. Fujiwara, & et al. (2004). Gibrat and Pareto–Zipf revisited with European firms. , Physica A 344, 1–2, pp. 112–116. Gabaix, X. (2008, September). Power Laws in Economics and Finance. Retrieved from NBER: http:// www.nber.org/papers/w14299 Gibbs, J.W. (1902). Elementary principles in statistical mechanics. New York, USA: C. Scribner's Sons incl. Golan A. (2006), Information and Entropy Econometrics — A Review and Synthesis, Foundations and Trends in Econometrics, vol 2, no 1–2, pp 1–145. Golan, A., Judge, G., & Miller, D. (1996). Maximum Entropy Econometrics: Robust Estimation with Limited Data. England: Wiley in Chichester. Green, W. (2003). Basics of Econometrics, 5th edition. NY: Prentice Hall. Halmos, P.R., & Savage, L.J. (1949). Application of the Radon-Nikodyma theorem to the theory of sufficient statistics. Annals of Math. Stat. 20, pp. 225–241. Hartley, R.L. (1928, July). Transmission of Information. Bell System Technical Journal. Heckelei, T. et al., (2008). Bayesian alternative to generalized cross entropy solutions for underdetermined econometric models. Bonn: 2008/2, University of Bonn. Ikeda, Y., & Souma, W. (2009). International comparison of Labor Productivity Distribution for Manufacturing and Non-manufacturing Firms. Progress of Theoretical Physics Supplement, 179, pp. 93, Oxford University Press. Jaynes, E.T. (1957). Information Theory and Statistical Mechanics. Physical Review, pp. 620–630. Jaynes, E.T. (1957b). Probability Theory: The Logic Of Science. USA: Washington University. Jeffreys, H. (1946). An Invariant Form for the Prior Probability in Estimation Problems. (R.S. London, Ed.) Proceedings of the Royal Society of London. Series A, Mathematical and Physical Sciences 186 (1007), pp. 453–461. Kantelhardt, J.W. (2008). Fractal and multifractal times series. Wisconsin, USA: Institute of Physics, Martin Luter University. Kullback, S., & Leibler, R.A. (1951). On information and sufficiency. Annals of Mathematical Statistics22, pp. 79–86. Kullback, S. (1959). Information theory and statistics. NY: John Wiley and Sons. Laplace, S.P. (1774). Memoir on the probability of causes of events. (English translation by S.M. Stigler, Ed.) Statist. Sci. 1(19), Mémoires de Mathématique et de Physique, Tome Sixième, 364–378. Lindley, D. (2008). Uncertainty: Einstein, Heisenberg, Bohr, and the Struggle for the Soul of Science. NY: Anchor. Mandelbrot, B. (1967). How Long Is the Coast of Britain? Statistical Self-Similarity and Fractional Dimension. Science, 156(3775), pp. 636–638. Maxwell, J.C. (1871). Theory of Heat. New York: Dover. Nielsen, F., & Nock, R. (2012). A closed-form expression for the Sharma-Mittal entropy of exponential families. Journal of Physics A: Mathematical and Theoretical 45, p. 3. Ormerod, P., & Mounfield, C. (2012). Power law distribution of the duration and magnitude of recessions in capitalist economies: Breakdown of scaling. Physica A293, pp. 573–582. Parikh, A. (1979). Forecasts of Input-Output Matrices Using the R.A.S. Method. The Review of Economics and Statistics 61, 3, pp. 477–481. Rak, R., & et al. (2007). Nonextensive statistical features of the Polish stock market fluctuations. Physica A 374, pp. 315–324.
Main Quantitative Properties of Statistical Information 19 f(x,y) = y yx x yx yxyx 2 2 2 2 2 2 1 2 2 )1(2 1 exp )1(2 1 where hypothesis H2 then represents the product of the normal densities as explained in (2.6), and finally one obtains: )1log( 2 1 ):( 2 21 I , (2.7) which indicates that in the case of bivariate normal distribution, as expected, the mean information is discriminatory in favour of H1 (dependence) against H2 (independence); that is I(μ1 : μ2) is a function of the correlation coefficient ρ alone. Following (Jeffreys, 1946), (Kullback, Information theory and statistics,, 1959), if we define I(2 : 1) as )( )( )( log)():( 1 2 221 xd xf xf xfI (2.8) as the mean information from μ2 for discrimination in favour of H2 against H1, one can define the divergence (noted ∇) by: ∇(H1, H2) = I(μ1 : μ2) + I(μ2 : μ1) = )( )( )( log))()(1( 2 1 2 xd xf xf xfxf = = )( )|( )|( log)( )|( )|( log 2 2 1 1 2 1xd xH xH xd xH xH divergence between hypotheses. (2.9) Thus, ∇(H1, H2) measures the divergence between H1 and H2 or between μ1 andμ2. As such, it constitutes a measure of the difficulty of discriminating between them. 1.2 Main Quantitative Properties of Statistical Information The approach undertaken here is axiomatic (Carter, 2011). It is worthwhile to note that we can apply this axiomatic system in any context where we have an available set of non-negative real numbers. This can be the case, for instance, when we dispose of non-negative coefficients (noted p) of a given set and target the estimation of the related model parameters through their reparametrization (Golan, Judge & Miller, 1996). Naturally, we will come back to such applications, and an estimation approach using probabilities and support space simultaneously will be presented. This underscores an important role to be assigned to the probability form of numbers, which motivated the selection of the axioms below. We will want our information measure I(p) to have several properties:
20 Information and its Main Quantitative Properties 1. Information is a non-negative quantity, i.e., I(p) ≥ 0. Following what has been presented above on information definition (see 2.4), one may generalize this property to convexity in the next theorem: Theorem: I(p1 : p2) is almost positive defined, that is I(p1 : p2) ≥0 with equality if and only if ſ1(x) = ſ2(x) [λ]. We will not demonstrate this theorem (see Kullback, 1959, pp. 14–15); we just provide the reader with the essence channelled through it. The above theorem explains that in the mean, discrimination information from statistical observations is positive. It follows from what has been previously said that no discrimination information will result if the distribution of observations is the same [λ] under hypothesis one and two. A typical example—as we will see later—may constitute maximum entropy and cross-entropy principles. In that case, when noninformative consistency moments from observations are not provided, minimum cross-entropy declines into maximum entropy. 2. If an event has probability 1, certainty follows, and we get no information from the occurrence of the event: I(p = 1) = 0. 3. If two independent events occur (whose joint probability is the product of their individual probabilities), then the information we get from observing the events is the sum of the two pieces of information: I(p1 p2) = I(p1) + I(p2). This property is referred to as additivity. Note that this property presents a valuable feature; it represents the basis of the logarithmic form of information. Intuitively, that means that a sample of n independent observations from the same population provides n times the mean information in a single observation. In the case of non-independent events, the additive property is retained, but in terms of conditional information. 4. Finally, as already stipulated in the preceding section, we will want our information measure to be a continuous (and, in fact, monotonic) function of the probability—slight changes in probability should result in slight changes in information. For consistency with the properties above, it can be useful to show the logarithmic feature of statistical information in the following way: 1. I(p2) = I(pp) = I(p) + I(p) = 2I(p) (2.10) 2. Through inductive reasoning, one can generalize (2.10) and rite, I(pn) = nI(p) 3. I(p) = I((p1/m)m) = m(p1/m) and we have )( 1 )( /1 PI m pI m
Main Quantitative Properties of Statistical Information 21 and, once again, we can generalize in the following way: )()( / pI m n pI mn 4. The property of continuity allows us to write, for 0 < p ≤ 1 and a real numberα: )()( PIpI . From (2.10), one can observe that an operator transforming the probability p at the powern/m, (that is, pn/m) into an information measure I(pn/m) displays a logarithmic property of additivity. This allows us to write a general, useful relation: ) 1 (log)(log)( p ppI bb for base b. (2.11) For other information properties not directly connected with the aim of this work, such as invariance or sufficiency, which will not be presented here, see Jaynes (1994), Kullback (1959). Furthermore, in the coming chapters, additional properties for different forms of entropy will be presented, such as concavity and stability (common for both Shannon-Gibbs and Tsallis entropies) or extensivity (common for both ShannonGibbs and Renyi (1961) entropies). As a final remark of this section, it is important to note that the above logarithmic nature of information as explained in (2.11)—for the case of independent events—is limited to ergodic systems which convey additive-extensive properties of information in the case of independent events.
2 Ill-posed Inverse Problem Solution and the Maximum Entropy Principle 2.1 Introduction As explained in the introduction, many economic relationships are characterized by indeterminacy. This may be because of long-range feedback and complex correlations between source and targets, thus rendering causal relationships more difficult to investigate. In this part of the work, the formal definition of the inverse problem will be discussed. A Moore-Penrose approach will be presented for solving this kind of problem and its limits will be stressed. The next step will be to present the concept of the maximum entropy principle in the context of the Gibbs-Shannon model. Extensions of the model by Jaynes and Kullback-Leibler will be presented and a generalisation of the model will be implemented to take into account random disturbance. The next step will concern the non-ergodic form of entropy known in the literature of thermodynamics as non-extensive entropy or non-additive statistics. There will be a focus on Tsallis entropy, and its main properties will be presented in the context of information theory. To establish a footing in the context of real world problems, non-extensive entropy will be generalized and then random disturbances will be introduced into the model. This part of the work will be concluded with the proposition of a statistical inference in the context of information theory. 2.2 The Inverse Problem and Socio-Economic Phenomena An inverse problem, e.g., Thikonov et al., (1977), Bwanakare (2015), Golan et al., (1996) explains a situation where one tries to capture the causes of phenomena for which experimental observations represent the effect. The essence of the inverse problem is conveyed by the expression: XY (2.12) or its equivalent in continuous form: )(),()()( bdXXBXgY D (2.13) where X represents the state space, Y designates the observation space, D defines the Hilbert support space of the model, B is the transformation kernel linking measures X and Y, b(ζ) displays random error process.
The Inverse Problem and Socio-Economic Phenomena 23 In classical econometrics, when given a state X, an operator B and, as happens most of time, a disturbance term (ζ), what is Y? This is referred to as a forward problem. In social science, one must often cope with the above random (Gaussian or not) disturbance term, and this usually complicates matters in spite of significant, recent developments in econometrics, particularly concerning stochastic time-series analysis (Engle & Granger, 1987). Furthermore, the inverse question is more profound: Given y and a specific B, what is the true state X? If B should also be a functional of X, the problem becomes arbitrarily complex. Correlation between (ζ) and X will be at the base of such additional complexity. Everyday, psychologists cope with such inferential problems. Patients display identical symptoms from different sicknesses. Health practitioners need more historical (a priori) information on patients to try to find the solution. In economics, the same national output growth rate may result from different combinations of factors. One of the main problems encountered by practicing economists is isolating the causes of economic phenomena once they have occurred. In most cases, the economist becomes inventive in finding an appropriate hypothesis before trying to solve the problem. As an example, in the case of a recession or financial turbulence, it is usually difficult to point to principal causes and fix them. Schools of economics suggest different, even contradictory, solutions—the legacy of its inverse problem nature. In empirical research, many techniques exist to try to solve the inverse problem. In the context of the present work, the presentation will be limited to those more applicable to matrix inversion, like the Moore-Penrose pseudo-inverse approach, and, naturally, maximum entropy based approaches. The approach better known in economics for updating national accounts on the basis of bi-proportionalities will then be added to these two techniques. 2.2.1 Moore-Penrose Pseudo-Inverse Let us consider the discrete and determinist case and rewrite (2.12) as follows: Y = XB = Xρ (2.14) In the right equality reflects the case where we have to deal with a ratio or probability parameter, for example, after reparametrizing B. We then have: ρ = YS ⇔ Y = XBS ρ = BY ⇔ Y = XYV Y = Xρ = YXV = XBXρ, (2.15) which means: XBX = X and V, representing the generalized inverse matrix (Golan, 1996), (Kalman, 1960).
24 Ill-posed Inverse Problem Solution and the Maximum Entropy Principle This is a matrix with the symbol B+ that satisfies the following requirements: B B+ B = B, B+B B+ = B, B+B is symmetric, B B+ is symmetric. Following Theil (1967), a unique B+ can be found for any matrix: square, nonsingular or not. When the matrix B+ happens to simultaneously be square and nonsingular, then the generalized inverse will be the ordinary inverse B-. The problem that interests us is the over-determined system of equations Y = XB where B has n rows, K < n columns and column rank equal to R ≤ K. If we retain the particular case when R equals K to ensure the existence of (B’B)-1, then the generalised inverse of B is B+= (B’B)–1B’ as can be easily verified. A solution to the system of equations can be presented as: X= B + Y. Following Green (2003, pp. 833), we note in this case that the length of this vector minimizes the distance between Y and BX, according to the least squares properties method. This distance will naturally remain equal to zero if y lies in the column space of B. If we now retain the more general case where B does not have full rank, the above solution is no longer valid and a spectral decomposition using the reciprocals of the characteristic roots is involved to compute the inverse which becomes: B+ = C1 A1 –1 C1’B’ where C1 are the R characteristic vectors corresponding to the non-zero roots arrayed in the diagonal matrix A1. The next and last case is the one where B is symmetric and singular, that is, with the rank R ≤ K. In such a case, Moore-Penrose inverse is computed as in the preceding case but without pre-multiplying by B’. Thus, for such a symmetric matrix, B+ = C1 A1 –1C1’, (2.16) with A1 –1 being a diagonal matrix of the reciprocals of the non-zero roots of B. It is important to note that only matrix B with full rank ensures a minimum distance between Y and BX. In other cases, there may exist an infinite number of combinations of elements of matrix B or ρ which satisfy (2.14).
The Inverse Problem and Socio-Economic Phenomena 25 To conclude, in spite of strong advantages of the Moore-Penrose generalised inverse, outputs will not always reflect an optimal solution. 2.2.2 The Gibbs-Shannon Maximum Entropy Principle and the Inverse Problem Let us introduce the concept of Shannon entropy by continuing with the case of pure linear inverse problem solution discussed above. The simplest (one dimensional case) example is the Jaynes dice inverse problem. If a dice is fair, and we throw it a large number of times n, with k different output modalities9 (k = 1,..., K), the expected value will be 3.5, as from a uniform distribution with probability fk equal 1/6. How can one infer about pk if we have ‘loaded’ (unfair) dice and the expected value of the trial becomes: 4.5 = K k k kp 1 5.4 (2.17) where frequencies pk is n nk ? In this case, the central question is: Which estimate of the set of frequencies would most likely yield this number? The problem is underdetermined since there are many sets of fk that can be found to fit the single datum of equation (2.17). Here we have to deal with a multinomial distribution where the multinomial coefficient w is given by: !!..! ! !!..! ! 2121 kk kkkppp W Deriving and using the Stirling approximation lnx! ≅ xlnx – x for a large number of N, we get the Shannon entropy formulation: K k kkp pppHMax 1 ln)( (2.18)10 In the case of a die, parameter K equals 6, and W is the multinomial coefficient, i.e., the number yielding a particular set of frequencies among 6N possible outcomes. 9 Generally, if the number of trials is equal to n, we will have nk possible outputs corresponding to each modality k with k k nn . Thus, the frequency pk = n n k is related to each modality k. 10 Note that the generalized form of Shannon entropy in the continuous case has the form: Maxf(y)H(f(y)) = –∫f(y)logf(y)dy.
26 Ill-posed Inverse Problem Solution and the Maximum Entropy Principle We need only find the set of frequencies maximizing W in order to find the set that can be realized in the greatest number of ways. This is the most plausible combination in the case of fair dice. This turns out to convey the same logic as maximizing Shannon Gibbs entropy. Thus, starting from two pieces of information that is, the number k equal to six and N a large number of trialswe are able to derive six probabilities related to a die distribution. Next, Jaynes (1994) maximized the Shannon function through the restriction of consistent information at hand. This opened entropy theory application to many scientific fields, including the social sciences. Thus, if we add to the formulation (2.18) the moment-consistency and the adding up-normalization constraints, we then get: K k kkp pppHMax 1 ln)( (2.19) subject to: K k tktk Ttyxfp 1 1,)( (2.20) K k k p 1 1 (2.21) where {y1, y2,..., yt} denotes a set of observations (e.g., aggregate accounts or their averages) being consistent with a function ft(xk) of explicative variables weighted by a corresponding distribution of probabilities {p1, p2,..., pk}. As usually happens, T is less than K, and the problem is ill-posed (underdetermined). Two main results emerge from the above formulation. First, if all events are independent or quasi-independent (locally dependent) and equally probable, then the above entropy is a linear function of the number of the possible system states and then is extensive11. A second fundamental result is connected with information theory and suggests that a Gaussian variable has the largest entropy among all random variables of equal variance (see Papoulis, 1991 for proof). In the next chapter on non-extensive entropy, a measure to assess the divergence of a given distribution from Gaussian distribution will be presented. 11 For this reason, as earlier alluded to, the Gibbs-Shannon entropy is called extensive. In reverse, as it will be commented on in the coming sections, the hypothesis of long-range correlation between events leads to the concept of non-extensive entropy (e.g., Tsallis entropy) suggesting an entropy no longer being a linear function of data.
The Inverse Problem and Socio-Economic Phenomena 27 Coming back to the dice case, maximization of Shannon entropy in (2.19), that is H(P) = –P΄ ln P under Jaynes consistency, leads to the distribution presented in Table1. To solve this inverse problem of six unknowns, the only two pieces of information available are the expected value—from the experiments in this example—assumed to be equal to 4.5 and the information that the probability of different possibilities adds up to one. However, since we are dealing with unbalanced dice, we have no idea about the distribution. The next chapters extend the Shannon-Gibbs-Jaynes maximum entropy principle with Kullback-Leibler relative entropy. The next to the last targeted presentation will deal with the general linear entropy model, that is, the one with a stochastic component. To conclude, Tsallis power law distribution to generalize Kullback-Leibler crossentropy will be considered. 2.2.3 Kullback-Leibler Cross-Entropy Kullback (1959), Good (1963) extended the Jaynes-Shannon-Gibbs model by formulating the principle of minimum (cross or relative) entropy. Using an a priori piece of information q about unknown parameter p, the resulting formulation is as follows: K 1k qp'lnp')/ln(),(_ pqppqpHMin kkk (2.22) under restrictions: Y = XP (2.23) P'1 = 1 (2.24) where p = (p1,..., pK)ʹ, q = (q1,..., qK). These restrictions are the same as those presented earlier. In the criterion function (2.22), a posteriori and a priori vectors or matrices p and q are confronted with the purpose of measuring entropy reduction resulting from exclusive new content of data information. Table 1: Recovering probability distribution of an unbalanced die through the maximum entropy principle. P1P2P3P4P5P6 0.054 0.079 0.114 0.166 0.240 0.348
28 Ill-posed Inverse Problem Solution and the Maximum Entropy Principle One should note that when q is fully consistent with moments, then p = q and the distribution becomes uniform with qk = 1/K. This leads to the solution of the maximum entropy principle. Thus, the cross-entropy principle stands for a certain form of generalization of maximum entropy. Relation (2.22) above is an illustration of the previous Kullback formulation in (2.8) as a mean information from (2.23) and (2.24) for discrimination in favour of p against q. 2.3 General Linear Entropy Econometrics In social science, it is rare to encounter the situation described by the relation (2.14) where the random term is meaningless as is often encountered in the experimental sciences. Social phenomena are particularly affected by stochastic components. Let us rewrite it below in its generalized form: 𝑦𝑦𝑖𝑖=∑𝐵𝐵𝑗𝑗𝑋𝑋𝑗𝑗+ 𝐾𝐾 𝑗𝑗−1 e i (2.12’) with the random term ζi∈e and i = (1,..., I) (I being the number of observations); K is the number of model parameters to be estimated. 2.3.1 Reparametrization of Parameters Following Golan et al., (1996), we first reparametrize the above generalized entropy model (2.12’). We treat each Bj (j = 1,…, K) as a discrete random variable within a compact support and 2 < M < ∞ possible outcomes. So, we can express Bj as: 1 M k km km m B pv k K (2.25) where pkm is the probability of outcome vkm and the probabilities must be non-negative and sum up to one. Similarly, let us treat each element ζi of e as a finite and discrete random variable with compact support and 2 < M < ∞ possible outcomes centred on zero. We can express ζi as: J j njnji wr 1 . (2.26)
Kullback-Leibler-Tsallis Cross-Entropy 35 NB: R stands for Renyi, and N q LVRA qSS (LVRA and N stand for Landsberg-VedralRajagopal-Abe and normalized, respectively (Gell-Mann & Tsallis, 2004). 2.4.3.2 Scale of q-Tsallis Index and its Interpretation Following the thermodynamic literature built on Lévy-like anomalous diffusion, it has been shown that 2 )( x q exp optimizes 1 )(1 q xpdx S q q under appropriate constraints. If one convolutes n times p(x)(n → ∞), we approach a Gaussian distribution if q .00 qq 5/3, and a Lévy L γ L(x) if 5/3 .00 qq q .00 qq 3. The index γ L of Lévy distribution is related to q as follows: 1 .3 L L q (5/3 .00 qq q .00 qq 3). Thus, in empirical applications, the value of q should vary inside an interval from unity to 5/3, which corresponds to cases of finite variance for phenomena dwelling within the Gaussian basin of attraction. 2.5 Kullback-Leibler-Tsallis Cross-Entropy 2.5.1 The q-Generalization of the Kullback-Leibler Relative Entropy Kullback-Leiber-Tsallis cross-entropy is known in literature as the q-generalization of Kullback-Leibler relative entropy. The Kullback-Leiber-Shannon entropy introduced in Part II can be q-generalized (Tsallis, 2009) in a straightforward manner. The discrete version becomes: 1 1/ ),( 1 q pp pppI q oii ioq (2.40)17 since with any real r .00 qq , one has the following properties: rq r q 1 1 1 1 1 if .00 qq (2.41) 17 In a continuous case, we have: 1 1)(/)( )( )( )( ln)(),( 1 )0()0( )0( q xpxp xdxp xp xp xdxpppI q qq
36 Ill-posed Inverse Problem Solution and the Maximum Entropy Principle = rq r q 1 1 1 1 1 if q = 0 ≤ rq r q 1 1 1 1 1 if q .00 qq 0 Thus, retaining the practical case of , we can write: . )( )( 1 1 1)(/)( 1 xp xp q xpxp o q o Hence18, .11 )( )( 1)( 1 1 )( )( )( 1 xp xp xp q xp xp xp o q o Therefore, coming back again to the generalized K-Ld cross-entropy, we have19: Iq(p,po) ≥0 if q .00 qq 0, =0 if q = 0, ≤0 if q .00 qq 0. (2.42) Thus, as Tsallis (2009) has made us aware, the above q-Kullback-Leibler index has the same basic property as the standard Kullback-Leibler entropy and can be used for the same purpose while having the additional advantage of an adaptive q according to the system with which we are dealing. There exist two different versions of the Kullback-Leibler divergence (K-Ld) in Tsallis statistics, the usual generalized K–Ld shown above and the generalized Bregman K–Ld. According to Venkatesanet et al., (Plastino & Venkatesan, 2011), problems have been encountered in empirical thermodynamics trying to reconcile these two versions. Unfortunately—or fortunately!—the same problems seem to reappear while applying this theory in social science since every version of generalized K-Ld leads to different outputs. Let us try to synthesize what recent literature says about this problem. 18 It is straightforward to derive this property in the case of the continuous case. 19 The same conclusion is obtained by using Jensen's inequality (e.g., Gell-Mann & Tsallis, 2004).
Kullback-Leibler-Tsallis Cross-Entropy 37 2.5.2 Tsallis Versions of the Kullback-Leibler Divergence in Constraining Problems This short section represents the final bridge between theory and the applications in the last parts of this work. In a recent study, Plastino & Venkatesan (2011) lay out interesting aspects of empirical research when q-generalized K-Ld cross-entropy is associated with constraining information. Since, in the social sciences, we particularly need discrete forms of these relative entropies, let us first rewrite these forms before commenting on their conditions of applicability: 1 1/ ),( 1 q pp pppI q oii ioq (2.43) 1 00 1 0 1 0))(()()( 1 1 q i i ii q i q i i iq pppppp q ppI (2.44) The form (2.43) is the one derived directly from Kullback-Leibler formalism and presented in (2.40). The second form is referred to as the generalized Bregman form of K-Ld cross-entropy, and it is more appealing than (2.43) from an information-geometric viewpoint (Plastino & Venkatesan, 2011) even if it does contain certain inherent drawbacks. A study by Abe and Bagci (2005) has demonstrated that the generalized K–Ld defined by (2.44) is jointly convex in terms of both pi and p0i while the form defined by (2.43) is convex only in terms of pi. A further distinction between the two forms of the generalized K–Ld concerns the property of composability. While the form defined by (2.44) is composable, the form defined by (2.43) does not exhibit this property. The second interesting aspect for practitioners concerns the manner in which mean values are computed. Non-extensive statistics has employed a number of forms in which expectations may be defined. The first among these are the linear constraints initially used by Tsallis (2009), also known as normal averages, that is: i i iypy The second is the Curado-Tsallis (C-T) constraints of the form: i i q iq ypy and the normalized Tsallis-Mendes-Plastino (TMP) constraints (also known as q-averages or an escort distribution) of the form: i i i q i q i q y p p y A fourth—less applied by practitioners—constraining procedure is the optimal Lagrange multiplier approach.
38 Ill-posed Inverse Problem Solution and the Maximum Entropy Principle Among these four methods to describe expectations, the most commonly employed by Tsallis practitioners is TMP, referred to as escort distribution. Recent work by Abe (2009) suggest that, in generalized statistics, expectations defined in terms of normal averages, in contrast to those defined by q-averages, seem to display higher consistency in material chaos hypotheses. Recent reformulation of the variational perturbation approximations in non-extensive statistical physics followed from these findings. To my knowledge, application in the social sciences to assess the universality of this finding has not been done yet. Finally, there is the issue of consistency. This stems from the form of the generalized K–Ld defined by (2.43) being consistent with expectations and constraints defined by q-averages (“prominently” the TMP) while, on the other hand, the generalized Bregman K–Ld defined by (2.44) is consistent with expectations defined by normal averages. Thus, through reformulations of an empirical inverse problem, this last point may play a key role since non-appropriated constraints should lead to a non-optimal solution in the best case or to computational problems, as is often the case. 2.6 A Generalized Linear Non-Extensive Entropy Econometric Model 2.6.1 A General Model This section presents a generalized linear non-extensive entropy econometric approach to estimate econometric model. Following Golan et al., (1996), we first reparametrize the generalized linear model of the equation (2.12’) rewritten below: ik K k ki XBy 1 (2.12’) With, once again, the random term ζi∈e and i = (1,...,I) (being the number of observations); K is the number of model parameters to be estimated. Where B values are not necessarily constrained between 0 and 1, and ζ is an unobservable disturbance term with finite variance, owing to the nature of economic data that exhibits error observation from empirical measurement or random shocks. If we treat each Bj (k = 1...K) as a discrete random variable with compact support and 2 < M < ∞ possible outcomes, we can express B as: km M m kmk vpB 1 , Kk ,,1 (2.45) where pkm is the probability of the outcome vkm. The probabilities must be non-negative and add up to one. Similarly, by treating each element ζi of ζ as a finite and dis-
A Generalized Linear Non-Extensive Entropy Econometric Model 39 crete random variable with compact support and 2 < M < ∞ possible outcomes centred around zero, we can express ζi as: J j ijiji wr 1 (2.46) where ri is the probability of outcome wi on the support space j, with j∈{1,...,J} and i∈ {i = 1,...,N}. Note that the term e (an estimator of ζ) can be fixed as a percentage of the explained variable, as an a priori Bayesian hypothesis. Posterior probabilities within the support space may display non-Gaussian distribution. The element vkm constitutes a priori information provided by the researcher while pkm is an unknown probability whose value must be determined by solving a maximum entropy problem. In matrix notation, let us rewrite β = V⋅P with pkm ≥ 0 and K k Mm km p 1 2 1 , where again, K is the number of parameters to be estimated and M the number of data points in the support space. Also, let e = r⋅w, with rij ≥ 0 and N i J Jj ij r 1 2 1 for N the number of observations and J the number of data points on the support space for the error term. Then, the maximum Tsallis Entropy Econometric (MTEE) estimator can be stated as: max 1 1111; qrprpH i j q ij k m q kmq (2.47) subject to J j q j q j J j j M m q m q m M m mik K k ki r r w p p vXeXBy 1 1 1 11 (2.48) K k M mkm p 1 2 1 (2.49) N i J jij r 1 2 1 (2.50) where the real q, as previously stated, stands for the Tsallis parameter. Above, Hq(p,r) weighted by α dual criterion function is nonlinear and measures the entropy in the model. The estimates of the parameters and residual are sensitive
40 Ill-posed Inverse Problem Solution and the Maximum Entropy Principle to the length and position of support intervals of β parameters. When parameters of the proposed mode20 concern elasticity or error correct coefficients, the values of which lie between 0 and 1, then the support space should be defined inside the interval zero and one. In other cases, the support space may be defined between minus and plus infinity, according to the intuitive evaluation of the modeller. Additionally, within the same interval support, the model estimates and their variances should be affected by the number of support values (Golan et al., 1996). Increasing the number of point values inside the support space leads to improving the a priori information about the system. A few years of modelling with the maximum entropy approach seem to show that a well-defined support space is crucial to obtaining better results. The weights α and (1 – α) are introduced into the above dual objective function. The first term “of precision” accounts for deviations of the estimated parameters from the prior (defined under support space). The second, “prediction ex post,” accounts for an empirical error term as a difference between predicted and observed data values of the model. 2.6.2 Parameter Confidence Interval Area In this section, we will propose the normalized Tsallis entropy coefficient S(âk) as an equivalent to a standard error measure in the case of classical econometrics. An equivalent of the determination coefficient R2 will be introduced, also under the entropy symbol S ( Pr ) . The departure point is that the maximum level of entropy-uncertainty is reached when significant information-moment constraints are not enforced. This leads to a uniform distribution of probabilities over the k states of the system. As we add each piece of informative data in the form of a constraint, a departure from the uniform distribution will result, which means a lowering in uncertainty. Thus, the value of the proposed S ( Pr ) below reflects a global departure from the maximum uncertainty for the whole model. Without giving superfluous theoretical details, we follow formulations in, e.g., Bwanakare (2014) and propose a normalized non-extensive entropy measure of S(âk) and S ( Pr ) . From the Tsallis entropy definition, Sq vanishes (for all q) in the case of M = 1; for M > 1, q > 0, whenever one of the pi(i = 1..M) occurrences equals unity, the remaining probabilities, of course, vanish. We get a global, absolute maximum of Sq (for all q) in the case of a uniform distribution, i.e., when all pi = 1/M. This vanishes (for all q) in 20 As already presented, the expression M m q m q m m p p 1 is referred to as escort probabilities, and we have for q=1 (then Pm is normalized to unity), that is, in the case of Gaussian distribution (Gell-Mann i Tsallis, 2004), (Tsallis, 2009).
A Generalized Linear Non-Extensive Entropy Econometric Model 41 the case M = 1; and for M > 1, q >0, whenever one of the pi equals unity, the remaining ones vanish. A global, absolute maximum of Sq (for all q) is obtained in the case of equiprobability, i.e., when all pi = 1/M. Note that we are interested, for our economic analysis, in q values lying inside the interval (1, 5/3). In such an instance, we have for our two systems: Sq(p) = (M1 – q – 1)⋅(1 – q)–1 (2.51) and Sq(r) = (N1–q – 1)⋅(1 – q)–1 (2.52) in the limit when q = 1, relation (2.51) or (2.52) leads to the Boltzmann-Shannon expression (Gell-Mann & Tsallis, 2004). Below, a normalized entropy index is suggested, one in which the numerator stands for the calculated entropy of the system while the denominator displays the highest maximum entropy of the system owing to the equiprobability property: S(âk) = – [1 – ∑ k ∑ m(pkm)q]/[k ⋅(M1 – q – 1)] (2.53) with k varying from 1 to K (number of parameters of the system) and m belonging toM (number of support space points), with M > 2. S(âk) then reporting the accuracy on estimated parameters. Equation (2.48) reflects the non-additivity property of Tsallis entropy for two (probably) independent systems; the first, parameter probability distribution, and the second, error disturbance probability distribution (plausibly with quasi-Gaussian properties): S ( Pr ) = [ S(p + r)] = {[S(p) + S(r)] + (1 – q) ⋅ S(p) ⋅ S(r)} (2.54) where: S(p) = – [1 – ∑ k ∑ m(pkm)q]/[k ⋅(M1 – q – 1)] and S(r)= – [( 1 – ∑ ∑ rq )] /[k ⋅N ⋅(J1 – q – 1)] n j S ( Pr ) is then the sum of normalized entropy related to parameters of the model S(p) and to the disturbance term S(r). Likewise, the latter value S(r) is derived for all observations n, with J the number of data points on the support space of estimated probabilities r related to the error term.
42 Ill-posed Inverse Problem Solution and the Maximum Entropy Principle The values of these normalized entropy indexes S(âk), S ( Pr ) vary between zero and one. Their values, near to one, indicate a poor informative variable while lower values are an indication of better informative parameter estimate âk about the model. The next part of the book will present in detail national accounts tables used for building or forecasting macroeconomic models. The statistical theory will be implemented particularly in the case of the inverse problem, while keeping in line with this work's objective. 2.7 An Application Example: a Maximum Tsallis Entropy Econometrics Model for Labour Demand This example presents, through Monte Carlo simulations, a model for labour demand adjustment for the Polish private sector. It constitutes an extension of an initial model presented by Bwanakare (2010) for the labour demand adjustment by the private sector of Subcarpate province in Poland. The model aims at displaying short-run and long-run relationships between labour demand determinants through an error selfcorrect process. Due to the relatively short period of the sample (fourteen annual data points) and the autoregressive nature of the model, we may have to deal with limited possibilities of statistical inference in the absence of convergence properties or, in the worst case, an inverse ill-behaved problem. Thus, traditional methods of parameter estimation may fail to be effective. We then propose to apply the generalized maximum Tsallis entropy econometric approach—as an extension of Jaynes-Shannon-Gibbs Information theoretic entropy formalism, already applied in econometrics (Golan, Judge & Miller, 1996). Due to an annual data frequency of the sample, the approach proves to be applicable in the case of classical econometrics when a small, lower frequency data sample is available. Such a small data sample should display tail queue Gaussian distribution. Through this application, Monte Carlo experiment outputs seem to confirm the reliability of the Tsallis entropy econometrics approach, which in this particular case performs as well as the generalized least square technique. 2.7.1 Theoretical Expectation Model In the short run, managers decide on the number of employees to be hired (or dismissed) in accordance with the expected long-run optimal level of production. However, because of institutional or economic reasons, that optimal number is not hired (or fired) at once. First, uncertainty remains a predominant characteristic of business. For this reason, employers naturally prefer a moderate and progressive adjustment of recruited workers to the targeted optimal level. Recruitment in some economic sectors could be time-consuming as well, especially when searching for
An Application Example: a Maximum Tsallis Entropy Econometrics Model for Labour Demand 43 good specialists. Second, relatively well organized trade unions could prevent employers from abrupt, large-scale layoffs, or the cost of dismissing a worker may become high, depending on prevailing labour laws at a given period. In both cases, the process of shock correction will be more or less long, depending on its origin and magnitude. Under classical assumptions of constant returns to scale, ex ante and ex post complementarities of factors, and long-run constant rate of labour productivity, the desired level of labour demand Lt* is a function of the output Yt and the technical progress t21: Lt* = α.exp(–β.t).Yt (2.55) Assuming that labour demand adjusts to its targeted level by an error correction model: log(Lt /Lt–1) = λ.log(L*t /L*t–1) +μ. log(L*t–1/Lt–1), (2.56) combining (2.55) and (2.56) leads to: log(Lt /Lt–1) = λ.log(Yt/ Yt–1) +μ. log(Yt –1/Lt–1) + μ. β.t + αo (2.57) The parameter λ is the impact of output on labour demand, and then a short-run elasticity of labour demand with respect to output Yt, μ being the error correction parameter. Since a relation –1 ≤ μ ≤ 0 should prevail, the equilibrium error is only partly adjusted at each period. In other words, this parameter synthesizes employers’ determinants of labour demand adjustment once a shock in sales for the coming period is expected. 2.7.2 A Generalized Non-Extensive Entropy Econometric Model 2.7.2.1 General Model Presently we are interested in the estimation of parameters of a Podkarpacki labour demand model, applying a generalized non-extensive entropy econometric approach. Following Golan, Judge & Miller (1996) and Bwanakare (2014a, 2014b), we reparametrize, in the first step, the generalized linear model before fitting it to equation (2.48). This step allows for including in moment equations-restrictions the same probability variables as those optimized in the criterion function. To reparametrize the model, we follow each equation in (2.45–2.46) where each βk(k = 1,…,K) is treated as a discrete, random variable with compact support and 2<M < ∞ possible outcomes. Next, for the estimation of the model, we maximize the entropy criterion function in (2.47) under moment and normality condition restric21 This is a simplification stipulating that technical progress is a linear function of time.
44 Ill-posed Inverse Problem Solution and the Maximum Entropy Principle tions presented in (2.48–2.50). For confidence area analysis, we need to apply Equations (2.53–2.54). With the purpose of improving estimated parameter quality, one can add additional a priori restrictions to (2.48–2.50) as follows: e = Y – Y = Y – XVp = 0. (2.58) Then we constrain the error term e to sum up to zero22 which provides an additional quality of requiring an unbiased parameter estimator. Efficiency property mainly depends upon informative quality of the prior. When it is poor, the values of the estimated pi from the model tends to be equal for all pi, i.e., the case of a uniform distribution. According to economic theory, we constrain elasticity parameters within a point support space of zero and one. As known (e.g., Golan, Judge & Miller 1996), sharper support area points of a parameter act as increasing quality of the “a priori” information. Furthermore, this allows computations of this nonlinear model to promptly converge to an optimum solution. This is explained as follows: 0 ≤ λ = Vp ≤ 1 (2.59) Likewise, we may add additional economic restrictions to the model (2.57) parameters; this leads to the following formulations: –1 ≤ μ = Vp ≤ 0 (2.60) –∞ ≤ β = Vp ≤ 0 (2.61) 2.7.3 Estimated Confidence Area Of Parameters In classical econometrics, we usually combine the variance of random model error with the co-linearity level of explicative variables to determine the standard error of estimated parameters and to infer their confidence area while assuming a normal distribution law of random errors. This is particularly true in the case of the Least Squares approach for a linear model. In entropy econometrics, the approach is very different. We use the normalized entropy S(âij) (equation 2.53) as an equivalent of the estimate standard error measure in classical linear model econometrics. Likewise, the equivalent of the coefficient of determination R2 is a S ( Pr ) (equation 2.54). Following Golan et al. (1996a, 1996b, 1996c, 2002) and Soofi (1992, 1994), in the case of maximum entropy formulation, the 22 Note that our model has a constant term, suggesting that the economic initial condition may impact the optimal solution.
Bibliography – Part II 51 Bwanakare, S. (2010). Non-Extensive Entropy Econometric Model (NEE): the Case of Labour Demand in the Podkarpackie Province. Acta Physica Polonica A, 117(4):647–651. Bwanakare, S. (2014). Econometric Balancing of a Social Accounting Matrix Under a Power-law Hypothesis. Przegląd Statystyczny 61(3):263–281 Bwanakare, S. (2014). Non-Extensive Entropy Econometrics: New Statistical Features of Constant Elasticity of Substitution-Related Models. Entropy 16(5):2713–2728. Bwanakare, S. (2015). Greenhouse Emission Forecast as an Inverse Stochastic Problem: A CrossEntropy Econometrics Approach. Acta Physica Polonica A 127(3A):13–20. Carter, T. (2011). An introduction to information theory and entropy. Complex Systems Summer School. Retrieved from: http://astarte.csustan.edu/~tom/SFI-CSSS Drożdż, S. & Kwapień, J. (2012). Physical approach to complex systems. Physics Reports 515(3–4):115–226. Engle, R.F., & Granger, C.W.J. (1987). Co-Integration and Error Correction: Representation, Estimation, and Testing. Econometrica, 55(2):251–276. Gell-Mann, M., & Tsallis, C. (Eds.). (2004). Nonextensive Entropy, Interdisciplinary Applications. USA: Oxford University Press. Golan, A., Karp, S.L., & Perloff , J.M. (1996a). Estimating a Mixed Strategy Employing Maximum Entropy. eScholarship, California. Retrieved from http://escholarship.org/uc/item/5rx8z14p Golan, A., Judge, G.G., & Miller, D. (1996b). Maximum Entropy Econometrics: Robust Estimation with Limited Data. Chichester, England: Wiley. Golan, A., Judge, G.G., & Perloff, M.J. (1996c). Estimating the Size Distribution of Firms Using Government Summary Statistics. The Journal of Industrial Economics, 3(1):69–80. Golan, A., & Perloff , J.M. (2002). Comparison of Maximum Entropy and Higher-Order Entropy Estimators. Journal of Econometrics, 107(1):195–211. Good, I.J. (1963). Maximum Entropy for Hypothesis Formulation, Especially for Multidimensional Contingency Tables. The Annals of Mathematical Statistics 34(3):911–934. Grech, D., & Pamula, G. (2013). On the multifractal effects generated by monofractal signals. Physica A, 392(23):5845–5864. Green, W.H. (2003). Econometric Analysis (5th Ed.). NY, USA: Prentice Hall. Halmos, P.R., & Savage, L.J. (1949). Application of the Radon-Nikodym Theorem to the Theory of Sufficient Statistics. The Annals of Mathematical Statistics, 20(2):225–241. Hartley, R.V.L. (1928). Transmission of Information. Bell Labs Technical Journal, 7(3):535–563. Jaynes, E.T. (1994). Probability Theory: The Logic Of Science. Washington DC, USA: Washington University. Jeffreys, H. (1946). An Invariant Form for the Prior Probability in Estimation Problems. Proceedings of the Royal Society A,186(1007):453–461. JSTOR 97883. Kalman, R.E. (1960). A new approach to linear filtering and prediction problems. J. Basic Eng., 82(D):35–45. Khinchin, I. (1957). Mathematical Foundations of Information Theory. NY: Dover. Kullback, S. (1959). Information Theory and Statistics,. NY: John Wiley and Sons. Landsberg, T.P., & Vedral, V. (1998). Distributions and channel capacities in generalized statistical mechanics. Physics Letters A, 247(3):211–217. Loeve, M. (1955). Probability Theory: Foundations, Random Sequences,. NY: Van Nostrand. Papoulis, A. (1991). Probability, Random Variables and Stochastic Processes (3rd Ed.). N.Y., USA: McGraw-Hill. Plastino, A., & Venkatesan, R.C. (2011). Deformed Statistics Kullback-Leibler Divergence Minimization within a Scaled Bregman Framework. Physics Letters A, 375(48):4237–4243. Pukelsheim, F. (1994). The Three Sigma Rule. The American Statistician, 48:88–91. Rajagopal, A.K., & Abe, S. (2000). Justification of Power Law Canonical Distributions Based on the Generalized Central Limit Theorem. Europhysics Letters, 52(610).
52 Ill-posed Inverse Problem Solution and the Maximum Entropy Principle Rényi, A. (1961). On measures of information and entropy. Proc. Fourth Berkeley Symp. on Math. Statist. and Prob., Vol. 1:547–561. (Univ. of Calif. Press, 1961) Savage, L.J. (1954). The Foundations of Statistics. NY, USA: Wiley. Shannon, C.E. (1948). A Mathematical Theory of Communication. Bell System Technical Journal, 27:379–423and623–656. Soofi, E.S. (1992). A Generalizable Formulation of Conditional Logit with Diagnostics. Journal of the American Statistical Association, 87:812–816. Soofi, E.S. (1994). Capturing the Intangible Concept of Information. Journal of the American Statistical Association, 89:1243–1254. Theil, H. (1967). Economics and Information Theory. Chicago, USA: Rand McNally and Company. Tikhonov, A.N., & Arsenin, V. I. (1977). Solutions of Ill-conditioned Problems. New Jersey: John Wiley & Sons. Tsallis, C. (2009). Introduction to Nonextensive Statistical Mechanics: Approaching a Complex World. Berlin: Springer.
Part III: Updating and Forecasting Input-Output Transaction Matrices
1 Introduction The second part of this book was concerned with showing statistical theory of information-based approaches as a basis for solving the ill-behaved inverse problem. This third part deals with the applications of the statistical theory of information in macroeconomics. It introduces ad hoc macroeconomic theory before building and estimating different models in the context of an inverse problem. A system of national accounts with particular emphasis on input-output tables reflecting complex interactions between economic activities, product markets, factors of production, and the behaviour of different economic agents is presented. A section is devoted to inputoutput multipliers. Numerical examples are provided using the RAS approach as well as the non-extensive entropy econometrics procedure. Limited extension of these tables will allow us to treat certain themes of the natural environment through a theoretical model. Then, tables such as input-output tables are treated in detail in the context of modelling or forecasting when prior information is insufficient and the matrix is ill-behaved. Faced with this problem, maximum entropy formalism, in particular, Tsallis non-extensive entropy, will be employed.
2 The System of National Accounts The System of National Accounts (SNA) constitutes the internationally agreed set of standardized conventions on how to compile, in a coherent and consistent way, an integrated set of macroeconomic accounts to measure economic activity (e.g., Eurostat, 2002). Such procedures require a set of internationally agreed-upon concepts, definitions, classifications, and accounting rules. In addition, the SNA provides an overview of economic processes, recording how production is distributed among consumers, businesses, government, and foreign nations. It shows how income originating in production, modified by taxes and transfers, flows to these groups and how they allocate these flows to consumption, saving, and investment. Consequently, the national accounts are building blocks of macroeconomic statistics, forming a basis for economic analysis and policy formulation. The SNA is intended for use by all countries, having been designed to accommodate the needs of countries at different stages of economic development. It also provides an overarching framework for standards in other domains of economic statistics, facilitating the integration of these statistical systems to achieve consistency with the national accounts. However, the complexity of the interrelations that emerge makes the understanding of underlying macroeconomic rules difficult. In 1947, Stone (1955, 1981), then head of the League of Nations Committee of Statistical Experts, submitted for the first time a report from the Subcommittee on National Income Statistics that would constitute the origins of the SNA. During the same year, the United Nations Statistical Commission (UNSC) promoted the evident need for international statistical standards for statistical comparisons in support of a large array of policy implementations. This led in 1953 to the publication of the SNA under the auspices of the UNSC. It consisted of a set of six standard accounts and a set of twelve standard tables presenting details and alternative classifications of flows in the economy. Successive modifications and extensions were implemented in 1960, 1964, 1968, 1993, and 2008. Extensions made in 1968 deserve more attention in the context of the present monograph. Inputoutput accounts and balance sheets were added to the framework of the SNA. Attention was focused on constant price derivation and a comprehensive effort was deployed to bring the SNA and the Material Product System (MPS) closer together. The archetype of the modern SNA was implemented in 1993 (Beutel & De March, 1998). In fact, the 1993 SNA represented a major advance in national accounting and embodied the result of harmonizing the SNA and other international statistical standards more completely than in previous versions. Since the first United Nations Conference on the Human Environment held in Stockholm in mid-1972, increasing needs to incorporate environmental aspects into the SNA started to be fulfilled in 1993 when a system of integrated environmental and economic accounting (SEEA) was introduced by Caticha & Giffin (UN, 1993; USA, 2007). The 2008 SNA update addresses issues brought about by changes in the economic environment, advances in methodological research, and the needs of users.
56 The Input-Output (IO) Table and its Main Application 3 The Input-Output (IO) Table and its Main Application Leontief (1941, 1966), the 1973 Economics Noble laureate and the father of the I-O table approach to economic analysis, began his first book about I-O analysis with the following words: This modest volume describes an attempt to apply the economic theory of general equilibrium— or better, general interdependence—to an empirical study of interrelations among the different parts of a national economy as revealed through covariations of prices, outputs, investments, and incomes. Leontief tried to apply neo-classical (Walras) general equilibrium to practical economic life. This suggests that subsequent analyses based on I-O tables or their extensions could have economic interpretations within the Walrasian framework apart from a few particular cases—e.g., those that consider the environment—violating Pareto optimum conditions. The objective of this chapter is to present a consistent methodology of updating, forecasting, and economic modelling on the basis of I-O tables—for which underlying matrices are ill-behaved or data are not reliable. The proposed maximum entropy methodology can dynamically assess I-O multipliers and update and forecast I-O table information by combining the generalized maximum entropy principle and macroeconomic theory. The procedure remains in line with multiplier-accelerator analysis, assuming that induced investment is a function of expected growth. The only required condition to apply the proposed techniques is the availability of statistical information on final demand or value-added accounts which allow for updating under some constraining information (macroeconomic or not) obtained earlier, according to the traditional approach I-O table. In the following sections, classical structure of an I-O table will be reviewed. The next step will describe I-O multipliers and their usage before trying to solve the more complex aspects of their estimation. Next, the proposed methodology for updating and forecasting an ill-behaved IO table will be described and the model presented. 3.1 The I-O Table and Underlying Coefficients The I-O method (Tomaszewicz 1992, 2005) represents a quantitative research approach that helps to understand how economic global product is created and shared with particular reference to connections within different sectors of production at this intermediate stage of product processing (Almon, & Clopper 2000; Avonds & Luc 2007; Robinson, Cattaneo & El-Said, 2001).
The I-O Table and Underlying Coefficients 57 Use of this method of analysis is built upon I-O tables, the construction of which assumes the existence of constant returns to scale in the production process and the existence of general Walrasian equilibrium in the overall economy. In fact, as suggested earlier, the intention of Wassily Leontief was to apply general equilibrium theory previously proposed by Warlas (Leontief, 1970, 1986). In the context of the methodology of this work, let us first present the familiar I-O table structure, known for many decades, with only a slight modification that allows it to retain its square form. The fundamental concept of the I-O model is the concept of a direct technical financial coefficient: j ij ij X x a for i,j = 1,...,n. (3.1) Matrix A = (aij) is called a matrix of coefficients of inputs. It expresses the proportion of the value of product from sector i to be involved (sold) in the sector j for production of one unit value (e.g., 1 euro). The elements of this table are also called technical coefficients when expressed in quantitative flows between industries. Indications: Xj is the value of the global product of jbranch, j = 1, ..., n. xij is the flow from the branches i to j, i.e., the value of the product manufactured in branch i-th and consumed by a branch of the j-th, i, j = 1, ..., n. Yif is the value of the final demand, i = 1, ..., n and f =1, ...F. Index f represents final demand composition, such as households, investment, exports, etc. The number F depends on the degree of table aggregation. Dj is the value added from branches of the j-th, j = 1, ..., n. Note that the above table can be split into four main parts: The first part (upper-left part of the table) is composed of sub-matrix A showing the structure of intermediary demand of products “i” by sector “j”. Table 4: Simplified inputs-outputs table structure. Sector Flows The demand of the final Total 1x11 x12 x13 …xlnY1fX1 2x21 x22 x23 …x2nY2fX2 … … … … … … … …. Nxn1xn2xn2…xnn Ynf Xn D1D2D3…DnD Total X1X2X3…XnYf Source: own elaboration.
58 The Input-Output (IO) Table and its Main Application In the second part (top right), we have the final demand (Yif). Its traditional elements are: household consumption, government consumption, investment and stocks and the export sector. The third part (lower left) shows the revenue created in the modes of production, i.e., the value added (Dj), i.e., remuneration of labour and gross profit of capital in production sectors. If in the second part the “export” sector is explicitly shown, then in the third part one must, consequently, show the import sector. Part four of the table refers to the secondary division of generated revenues. In the case of the I-O matrix it remains empty. Only construction of the social accounts matrix allows for completing this information. Here, one should recall macroeconomic balance between final demand and value added, i.e.: j i ij YD . A further important equation is the definition of the Leontief model: (I – A)X = Y (3.2a) or (I – A)–1 Y = X (3.2b) where: X = [Xj] n-dimensional relationships to the column vector of global product Y = [Yi] n-dimensional relationships to the row vector of the final product and I is the identity matrix. This relationship between the final product Y, the global product X, and cost structure matrix A is useful for the calculation of one of these elements when information about the other items is available. In the literature, this is known as a forecast of the first type (3.2a) or the second type (3.2b). 3.2 Input-Output Multipliers 3.2.1 The principal models Multiplier coefficients play a central role in economic analysis (Leontieff, 1970). They make it possible to measure the impact of the exogenous variable or a shock on the whole system in which elements are interconnected. On these grounds, systemic models like these based on I-O tables—or their extension, or on computable general equilibrium models—have proved decisively superior to the classical ceteris paribus approach, using the classical econometric models. The above defined relations (3.2a) and (3.2b) are very useful in empirical research. (3.2b) explains, in (input) multiplier terms, responses of producing sectors to a one
Input-Output Multipliers 59 unit value increase in sectorial demand. For example, if the government spends one additional euro for buying a product from a given sector, what will the (backward) production impact be on the whole economy? The response is given by the level of multiplier coefficients along the sector column under consideration and total impact is given by its multiplier sum. The term (I – A)–1 in relation (3.2b) represents the multiplier. This formula can be decomposed as follows: (I – A)–1 = i i i A 0 with A0 = I and Ai = Ai–1 A (3.2c) Such a multiplier displays three impacts: – initial impact equal to one, – direct impact equal to A, and – indirect impact summing up to A2 + A3 +...+ Ak +...+ An. Note that this geometric progression is quickly convergent; after a few steps, the last term vanishes to zero. If we try to make a thorough analysis regarding the probabilistic nature of the multiplier, then we can rewrite (3.2c) as follows: 0 ! i iA Aie = (I – A)–1, (3.2d) after having used the Taylor development. Thus, due to (3.2d), economic multiplier structure displays the exponential family of laws, which may constitute—as indicated in the first part of this book—a transitional law between power law and Gaussian law when power law-related higher frequency data are progressively aggregated towards the low frequency series. The purpose of this short discussion is to draw attention to the use of formula (3.2c). Not only should the frequency level of data have an impact on results, but also the functional relation of matrix A could modify the convergence transition path between the three probabilistic laws above. Equation (3.2b) is generalized by the following, one of the most important relations of I-O modelling theory, referred to as the central model of input-output: Z = B (I – A)–1 Y (3.3) B = matrix of input coefficients for specific variables in economic analysis (intermediates, labour, capital, energy, emissions, etc.) I = Identity matrix A = matrix of technical financial coefficients [aij] Y = diagonal matrix of final demand Z = matrix with results for direct and indirect requirements (intermediates, labour, capital, energy, emissions, etc.)
60 The Input-Output (IO) Table and its Main Application Three of the most frequently used types of multipliers in I-O analysis are those that estimate the effects of the exogenous changes of final demand (consumption, investment, exports) on: a) outputs of the sectors in the economy, b) value added and income earned by the households, and c) employment that is expected to be generated by the new activity levels. However, due to interesting potential applications of this theory, it is worthwhile to be more complete about I-O multipliers. In empirical research, the I-O models used are based on input coefficients. Nevertheless, there is also a family of I-O models that are based on output coefficients. These models are sometimes called Gosh-models (Gosh, 1964). These models can be used to study price and cost effects or forward linkages of industries. Input coefficients reflect production functions or cost structures of activities. In contrast, output coefficients are distribution parameters for products reflecting market shares. Presented below are only the four basic I-O models with input and output multipliers. The four I-O models have a dual character with an underlying symmetry. Each I-O model with input coefficients has a complement with output coefficients. j j j X D d , i.e., input coefficient for value added. (3.4) Input coefficients for intermediates (aij) (3.1) reflect the requirements for the use of product i in industry j for one unit of output of industry j. The capital and labour requirements are defined in the same way. i ij ij X X b , output coefficients for products (3.5) i i i X Y y , output coefficients for final demand (3.6) Output coefficients for intermediates (bij) identify the share of deliveries of sector i for sector j, (xij) in the total output of sector i. Model 1: quantity model with input coefficients This model has been already defined in (3.2a) and (3.2b). Model 2: price model with input coefficients A'p + d = p (3.7) (I – A')p = d (3.8) p = (I – A')–1d (3.9) A' = transposed matrix of input coefficients for intermediates with A =[aij] for i, j=1,...,m. I = identity matrix
Updating and Forecasting I-O Tables 67 For instance, if the new known row and column totals, X and Y, are known with uncertainty—a realistic hypothesis—one will need to add a random component to the model. The RAS procedure is not suited for tackling such problems. The risk of starting with a basic I-O matrix A0 characterized by “spurious consistency” This is the case when the matrix has been updated or balanced on the basis of a theoretical hypothesis, e.g., macroeconomic closure rules. In such a case, the matrix appears well balanced despite possibly containing systematic errors. Because of the above empirical problems, many researchers have tried the extension of the RAS procedure in the hope of rendering it more flexible. Lemelin et al. (2005) show that the RAS procedure can be apprehended as a Bayesian information processing rule with the new known row and column totals X and Y taken for new data in the Bayesian sense. In the process, after having compared the Lagrange multipliers of both procedures, the authors show conditions of equivalence between the RAS procedure and the Kullback-Leibler cross-entropy approach. Similar work has been presented by McDougall (1999). He concluded that the RAS approach corresponds to a generalized Shannon cross-entropy technique, suggesting that the latter cannot supplant the former. Nevertheless, according to McDougal, the cross-entropy approach can extend and adapt the RAS technique to problems that do not fit well with the traditional matrix balancing framework. Interestingly enough, Robinson et al. (2001), have conducted a comprehensive experiment on a 1994 SAM of Mozambique. Starting from the balanced SAM and randomly imposing new row and column totals, the authors have operated a Monte Carlo experiment in which they have simultaneously updated the matrix using RAS and Shannon cross-entropy procedures. They found RAS and CE to be equivalent measures—meaning that RAS is an entropy theoretic—if the CE method uses a single cross-entropy measure as an objective instead of attempting to use the sum of column cross-entropies. They concluded by confirming the findings of many previous researchers according to whom the RAS procedure remains less flexible in the case of new information in comparison with the cross-entropy technique, which is better at processing new information for optimal consistency of the updated SAM. However, due to its popularity, researchers have proposed new extensions of the RAS approach [e.g., http://ec.europa.eu/eurostat/ramon/statmanuals/files/KS-RA07–013-EN.pdf]. One of the interesting extensions has been the Model of Double Proportional Patterns (MODOP) developed by Stäglin (1972:69–81). This model consists of estimating all the existing cells of transaction. The resulting inconsistent matrix is then estimated by the RAS approach. The basic idea is to calculate the geometric mean of row and column multipliers and then to apply this factor to each element of the matrix. Following the above author, outputs from the MODOP are often similar to those from
68 The Input-Output (IO) Table and its Main Application the RAS approach. Some trials that allow the RAS approach to constrain targeted cells inside a matrix or to render the cells stochastic have been undertaken in recent years. In a recent, thorough study on the comparative performance of the cross-entropy and RAS techniques, Chisari et al. (2012) concluded that cross-entropy had a more general character for the following reasons: a) It does not need all the new totals of rows or columns (although prediction will be less accurate). b) It does not need a balanced initial matrix (the sum of rows could be more/less than the sum of columns). c) New rims could contain an error term. d) New rims can be non-fixed parameters. e) Many values on the final matrix could be fixed (not necessarily a parameter). f) It allows non-linear constraints. Referring to their simulation outputs, the authors propose a rule of thumb consisting in preferring the RAS method if and only if any constraint or one constraint is enforced. This seems to explain why the RAS approach continues to be successfully applied in different prediction studies. Furthermore, comparing the starting point to the RAS method, the above authors observe that the purchasing method is preferred to the supplying matrix because the aggregate bias is in this lower case. Furthermore, note that this suggestion does not seem to be consistent with the above investigations done by Robinson et al. (2001) on the Mozambique economy according to which the RAS and Shannon entropy approaches produce the same performance when no additional restriction is imposed. We will come back to this point when we present an example of I-O updating at the end of this section. Trying to assess the cross-entropy approach in a dynamic, stochastic multi-objective optimisation problem, Bekker & Aldrich (2011) have concluded that acceptable results can be obtained while doing relatively few evaluations. Such an empirical fact tends to confirm a large area of cross-entropy application. Finally, the central point to focus on through this section has been the limit—at least according to existing literature—of the RAS method compared to Shannon-Kullback-Leibler cross-entropy. Thus, since the latter is itself a converging case of Tsallis non-extensive entropy, the RAS procedure can be seen as relatively less attractive with respect to both entropy approaches. 3.3.3 Application: Updating an Aggregated EU I-O Matrix 3.3.3.1 The RAS Approach Tables 8 and 9 represent 27 aggregated EU symmetric I-O tables for domestic output at basic prices for years 2006 and 2007. In this simplified example, let us suppose that we have no information—though we do—on the elements of the intermediate
Updating and Forecasting I-O Tables 69 consumption sub-matrices for the period 2007 in Table2. Using both the RAS and cross-entropy approaches, we are asked to predict those elements on the basis of the 2006 I-O matrix and sectorial accounts totals of the period 2007, which are supposed to be known. In the example below, we directly use the transaction matrix in current price value. Thus, instead of coefficient matrix A presented in the above RAS formula, we use the matrix X of transactions. Using the formalism explained in the above section for the RAS approach, we present below the algorithm to solve the problem: Iteration 1 Calculation of row multipliers 1.098105 1.068032 1.066862 1.044765 1.052929 1.037053 Actual Row multipliers 50495.53 184777.3 3012.483 24612.5 3822.238 6453.238 273173.2 280691 1.027522 84702.14 2585662 382824.5 542838.5 230943.7 266151.9 4093122 4128898 1.008741 2941.775 45081.56 361192.4 47169.66 121551.4 47356.15 625293 598576 0.957274 38756.27 770976.1 124401.9 865048.6 264081.2 185633.2 2248897 2228709 0.991023 27106.73 694435.6 18950.7 753990.3 1314546 370177.4 3349764 3236274 0.96612 5269.119 83635.91 11198.4 72001.38 108426.3 267284.2 547815.4 538038 0.982152 where, e.g., starting multiplier 1.098105 to be later multiplied by the first column elements of the initial transaction matrix of 2006 is obtained from the quotient 419340/381876, i.e., the first column output of 2007 divided by the first column output of 2006, both respectively from Tables 9 and 8. With elements of row multipliers on the diagonal matrix premultiplied by X0ij obtained in the above transformation, we get: 0 ~ ij XR equal to: 1.027522 0 0 0 0 0 0 1.008741 0 0 0 0 0 0 0.957274 0 0 0 0 0 0 0.991022853 0 0 0 0 0 0 0.96612 0 0 0 0 0 0 0.982152
70 The Input-Output (IO) Table and its Main Application X 50495.53 184777.3 3012.483 24612.5 3822.238 6453.238 84702.14 2585662 382824.5 542838.5 230943.7 266151.9 2941.775 45081.56 361192.4 47169.66 121551.4 47356.15 38756.27 770976.1 124401.9 865048.6 264081.2 185633.2 27106.73 694435.6 189507.7 753990.3 1314546 370177.4 5269.119 83635.91 11198.4 72001.38 108426.3 267284.2 equal to: 51885.25 189862.6 3095.392 25289.88 3927.432 6630.843 85442.48 2608262 386170.6 547583.2 232962.2 268478.2 2816.084 43155.38 345759.9 45154.27 116358 45332.79 38408.35 764055 123285.2 857282.9 261710.5 183966.7 26188.35 670908 183087.1 728445 1270009 357635.7 5175.075 82143.16 10998.53 70716.28 106491.1 262513.7 209915.6 4358386 1052397 2274472 1991459 1124558 Actual 210592 4373650 1056144 2280581 1974301 1115919 Column multiplier 1.00322 1.003502 1.003561 1.002686 0.991384 0.992318 SXRX ij ~ ~0 1 equal to: 51885.25 189862.6 3095.392 25289.88 3927.432 6630.843 85442.48 2608262 386170.6 547583.2 232962.2 268478.2 2816.084 43155.38 345759.9 45154.27 116358 45332.79 38408.35 764055 123285.2 857282.9 261710.5 183966.7 26188.35 670908 183087.1 728445 1270009 357635.7 5175.075 82143.16 10998.53 70716.28 106491.1 262513.7 X 1.00322 0 0 0 0 0 0 1.003502 0 0 0 0 0 0 1.003561 0 0 0 0 0 0 1.002685984 0 0 0 0 0 0 0.991384 0 0 0 0 0 0 0.992318
Updating and Forecasting I-O Tables 71 equal to: 52052.34 190527.6 3106.414 25357.81 3893.595 6579.903 85717.63 2617396 387545.8 549054 230955.1 266415.7 2825.153 43306.52 346991.2 45275.55 115355.5 44984.53 38532.04 766730.8 123724.2 859585.6 259455.7 182553.5 26272.69 673257.6 183739.1 730401.6 1259068 354888.3 5191.74 82430.84 11037.69 70906.23 105573.6 260497 Iteration 2. …j….: Repeat Iteration 1 algorithm In this example, we reach the optimal value at the seventh iteration and then get the following values of the new supposed unknown transaction matrix of the year 2007: 0.999998 0 0 0 0 0 0 0.999999 0 0 0 0 0 0 1.000001 0 0 0 0 0 0 0.999999944 0 0 0 0 0 0 1.000001 0 0 0 0 0 0 1.000006 X 51929.5 190003.4 3095.379 25249.69 3870.157 6543.312 85635.77 2613871 386712.9 547482.9 229888.1 265307.2 2829.519 43356.44 347112.1 45258.99 115109.9 44909.49 38562.37 767033.6 123673.6 858620.7 258707.4 182111 26405.53 676396.6 184447.4 732693.9 1260792 355538.3 5228.733 82985.65 11103.05 71275.19 105935.8 261511.4 X 1.000001 0 0 0 0 0 0 1.000001 0 0 0 0 001000 0 0 0 0.999999724 0 0 0 0 0 0 0.999999 0 0 0 0 0 0 0.999998
72 The Input-Output (IO) Table and its Main Application equal to: 51929.54 190003.5 3095.379 25249.68 3870.153 6543.302 85635.84 2613873 386712.9 547482.7 229887.9 265306.8 2829.521 43356.47 347112 45258.98 115109.8 44909.42 38562.4 767034 123673.6 858620.4 258707.2 182110.7 26405.55 676397 184447.4 732693.7 1260790 355537.7 5228.738 82985.69 11103.05 71275.17 105935.7 261511 Table 6: Error(%) prediction from the RAS procedure 3 0 1 -4 -7 4 1 0 0 -2 0 0 -7 5 -1 6 -4 3 -6 -1 -2 1 0 3 10200-1 -4 1 4 2 2 -2 The above errors are calculated as the error discrepancy percentage between the true matrix of transaction representing the period of year 2007 and the matrix updated by the RAS procedure on the basis of the 2006 I-O matrix. 3.3.3.2 The Entropy Approach For comparative purposes, let us apply entropy formalism for updating the same IO transactions table as in the above example related to the RAS approach. Thus, using the Tsallis entropy formalism presented in (3.17) and in (2.48–2.50) under the hypothesis that transaction totals of the targeted period are known (and without any additional restriction), we get the outputs presented in Table 7. Comparison of Tables 6 and 7 show slightly higher performance of the RAS approach. This seems to confirm the rule of thumb proposed by Chisari et al. (2012): if and only if any constraint or one constraint is enforced as in the present case. However, such a conclusion, as already mentioned above, is not in line with the one proposed by investigations conducted by Robinson et al. (2001) which lead to equivalent performance in the same conditions between the cross-entropy and RAS approaches. More investigations are needed to contradict or confirm the study results of the above authors. Of course, following the results of several investigations presented above, cross-entropy naturally presents higher performance than the RAS approach, particularly when statistical data are known with uncertainty. In the next section, we are going to propose the forecasting of a higher dimension I-O table through crossentropy formalism. Later, when treating the case of a SAM with higher dimension, we
Updating and Forecasting I-O Tables 73 will present the methodology of updating it in the presence of uncertainty and with a free number of restrictions. 3.3.3.3 I-O table Forecasting By “updating,” we compute operations on table rows and columns with the purpose of balancing its row and column totals, but a forecasting operation implies the use of a certain theoretical model—whether deterministic or not. To forecast an I-O table, the statistical data concerning final demand Yi and the value added Dj are generally available or could be obtained on the basis of existing information. For example, in the case of Eurostat, this information exists for the time horizon year 2013 while Table 7: Error (%) prediction of the I-O table by the Tsallis (or Shannon) cross-entropy procedure Products (CPA) P1 P2 P3 P4 P5 P6 P1 7.985 3.141 1.934 -0.049 -7.602 4.579 P2 8.954 2.214 -0.722 -0.099 -2.772 -0.158 P3 -8.035 2.042 -5.767 3.010 -11.670 -2.940 P4 -4.044 0.691 -4.533 7.906 -4.444 0.523 P5 0.370 -2.219 -2.198 -1.980 -6.118 -0.147 P6 -2.688 0.587 1.827 1.848 -2.704 -4.190 Table 8: Symmetric I-O Tables for domestic output at basic prices (year 2006; EU27, Mio. EUR current prices) Inputs of products Products (CPA) P1 P2 P3 P4 P5 P6 Others Total P1 45984 173007 2824 23558 3630 6223 126650 381876 P2 77135 2420959 358832 519579 219335 256642 2999462 6851944 P3 2679 42210 338556 45149 115441 45664 1130576 1720275 P4 35294 721866 116605 827984 250806 179001 2465270 4596826 P5 24685 650201 177631 721684 1248467 356951 1932724 5112343 P6 4798 78308 10497 68916 102976 257734 3054393 3577624 Others 191301 2765392 715330 2389956 3171688 2475408 1499777 13208852 Total 381876 6851944 1720275 4596826 5112343 3577624 13208852 Source: own calculations. Source: based on http://ec.europa.eu/eurostat/web/esa-supply-use-input-tables/overview
74 The Input-Output (IO) Table and its Main Application neither the value of global product Xj nor the matrix of cost structure [aij] are known for that forecasting period. Here, the last updating of IO tables concerns the period 2007. Building an I-O table is time consuming and a lot of information, particularly concerning the transaction matrix (including import elements), is not easy to assess. Economic information from different enterprises or industries must be updated on the basis of new flows of additional information. As a result, getting a final version of that table comes long after with many lag years. Under such conditions, finding a workable methodology for setting up a robust prediction technique of an I-O table should bear precious advantages. Therefore, the key question is: Based on the I-O table of the previous period, on vectors Yp i and Dp j (bothfrom the forecasting period), is it possible, using the connection (3.2), to estimate the unknown IOtablepof the forecasting period? Of course, the Table 9: Symmetric I-O Tables for domestic output at basic prices(year 2007) Inputs of products Products (CPA) P1 P2 P3 P4 P5 P6 Others Total P1 53529 189315 3126 24311 3626 6784 138648 419340 P2 86528 2624065 386717 535924 229362 266302 3189197 7318095 P3 2656 45671 344941 48062 111100 46147 1236719 1835296 P4 36335 756877 121086 869135 258074 187202 2573895 4802604 P5 26539 673893 188669 730656 1264384 352132 2146659 5382932 P6 5005 83827 11606 72494 107755 257351 3172149 3710187 Others 208748 2944446 779152 2522023 3408631 2594266 1599374,605 14056642 Total 419340 7318096 1835296 4802604 5382933 3710185 14056642 Source: based on http://ec.europa.eu/eurostat/web/esa-supply-use-input-tables/overview where: P1 Products of agriculture, hunting and fishing P2 Industrial products (incl Energy) P3 Construction work P4 Trade, transport and communication services P5 Financial services and business services P6 Other services Others Use of imported products, cif Taxes less subsidies on products Value added at basic prices Total Row or column totals
Updating and Forecasting I-O Tables 75 answer is not affirmative if we use the classical statistical-mathematical approach to solve that inverse problem. On the one hand, we do not have information about matrix Apof the forecasting period to derive the value Xp using relation (3.2) and, in this way, to determine IOtablep. On the other hand, as an effect of the possibility of estimating Xp, for example using a dynamic investment model, the new question that arises is: Disposing of Xp and of Yp of the forecasting period, could one possibly determine the matrix coefficients Apon the basis of the same relation (3.2)? Since the IO table has size NxN, and we have by assumption only information about final demand Yp and global product Xp of each sector, this means that we have (N – 2)x(N – 2) degrees of freedom, where N is, once again, the number of branches in the IO table. Such a problematic belongs then to the category of inverse problems, which suggests that there may be an infinity of matrices Ap that could reproduce the identical values of final demand Yp and the global product Xp. Among them, we will choose the one that best maximizes consistency information between the prior, data, and the posterior. We can retrieve a solution proposal to that problem in the second part of this monograph about the maximum entropy principle or relative (cross) entropy. Let us again present it below in the context of updating and forecasting an I-O table on any other form of its extension. 3.3.3.4 The Non-Extensive Cross-Entropy Approach and I-O Table Updating As suggested above, in the recent literature there are several methods for updating and balancing elements of national accounts balance sheets, for instance, an I-O table when equality of corresponding sums of columns and rows is required. Some of their limits have been emphasised here. Preference is then given to methods based on statistical theory of information for their capacity to adjust information when structural changes affect the economy or when additional consistent information has to be added to what already exists. The most frequently used theoretic-information methods are the Maximum Likelihood, Bayesian method of moments, and methods based on the maximum entropy principle. Through their article, Giffin & Caticha (2007) have proven that the principle of maximum entropy represents a generalization of the Bayesian approach as a method of inference on the basis of an a priori information. Probably for these reasons, application of the cross-entropy approach to balance the social accounting matrix has been widely adopted in empirical application during recent years (e.g., Robinson et al., 2001). As we know from Part II of the monograph, on the basis of the results of Shannon (1948) and Jaynes (1957), Kulback-Leibler (KL) (1951, 1957) and Good (1963) have proposed the principle of minimum (relative) entropy. This principle aims at assessing a posteriori parameters (probabilities P) of the most plausible, shortest divergence in relation to a priori parameters (probabilities Q), under restrictions related to data moments, normalization condition, or any other a priori information presenting con-
76 The Input-Output (IO) Table and its Main Application sistency with probabilities in the criterion function. The formulation in the case of discrete events is like in (3.16) and, thus, we have: Min(rj,r0) → oq rrI 1 00 1 0 1 ))(()()( 1 1 q j j jj q j q j j j rrrrrr q (3.16’) or, as traditionally done in the case of Shannon-Gibbs entropy, QPPP q p pQPMin n ii i iKL lnlnln )( 1 , for i = 1,...,n under restrictions: nj j nj q j q j j X r r X ..1 ..1 . . )( )( (3.18) nj q d j q d j nj j r r D ..1 ' ' ..1 )( )( = ni ni q y i q y iY r r ..1 ..1 ' ' (3.19) Ω(h) = F (3.20) P'I = 1 and P ≥ 0. (3.21) We adopt Shannon-Gibbs symbols in the criterion function above. Correspondence with the Tsallis criterion function in (3.16) is easy to do. The symbol P corresponds to r and Q to ro. Matrix P stands for posterior probabilities guaranteeing the balance of a previously unbalanced table, the elements of which sum up to unity by column (sector). Q is a matrix of coefficients from a known table. In the case of the I-O table (see Table 2) elements of Q are derived by dividing each column element by its total. They then represent input coefficients except the case of the final demand column, where these coefficients explain the structure of product consumption for a given final demand institution. Thus, its elements must satisfy the additivity condition. Equation (3.18) demonstrates that the column total must match with corresponding row elements multiplied by corresponding probabilities (coefficients) Pj. Equation (3.19) states equality between value added and final demand totals, with d j p' and ' y i P being transposed respective vector of sectorial value added components and vector of structural coefficients of final demand. Probabilities are presented in escort distribution formulation presented in footnote 17. Functional h in (3.20) gives a piece of information which has a significant relationship (consistency) with probabilities in the criterion function. This may be, for example, a macroeconomic balance equation or any distribution of a treatment of errors. Equation (3.21) is one of the additivity conditions of probabilities and reminds us that any probability can take a negative value.
Bibliography – Part III 83 1 1 h m fm e with h=f + 2, i.e., additional columns due to the demand of emissions by institutions (one of them having zero value since demand of emissions must have a separate column). All symbols are as before. Note that in this case the environmentally extended I-O matrix has the form presented in Table 10 below. Input coefficients are derived from global column totals, including, thus, quantities (instead of values) representing emissions. 3.4 Conclusions As suggested above, generally, new information should render more homogeneous divergence between the true coefficients and those forecasted. According to the above results, we observe that new restrictions added to the model lead to a significant reduction of errors. This is in accordance with the rule of thumb proposed by Chisari et al. (2012) in section III.3.3.1 about particular conditions explaining the superiority of entropy approaches over the RAS technique. The obtained AEVC coefficients naturally present a random character and different experiments would produce different values. However, as expected through Bayesian formalism, new data evidence will always tend to reduce the level of uncertainty or entropy. The true difficulty in assessing a new methodology to assess a complex information system—like the one represented by an I-O table—is that, due to instruments of measure or adopted economic hypotheses, a part of the observed data may not be accurate. This can be even worse in the case of general equilibrium systems in which the balance of the whole system—or accounts—may be more or less forced. This observation is particularly true in developing countries where statistical information management can be more challenging. What we intend to explain here is that, faced with such circumstances, the output performance of the non-extensive entropy approach should, consequently, be taken with a certain margin of error. The quality of priors and model data will always play a central role. Bibliography – Part III Almon, & Clopper. (2000). Product-to-product tables via product-technology with no negative flows. Economic Systems Research, Vol. 12, No.1,, 12(1), pp. 27–43. Avonds, & Luc. (2007). The input-output framework and modelling assumptions: considered from the point of view of the economic circuit. The 16th International Input-Output Conference. Istanbul: International Input-Output Association.
84 The Input-Output (IO) Table and its Main Application Bacharach, M. (1971). Biproportional matrices and input-output change. Cambridge: Cambridge University Press. Bekker, J., & Aldrich, C. (2011). The cross-entropy method in multi-objective optimisation: An assessment. European Journal of Operational Research, 211(1):112–121. Beutel, J., & De March, M. (1998). Input-output framework of the European System of Accounts (ESA 1995),. Paper presented at the Twelfth International Conference on Input-Output Techniques. New York: International Input-Output Association. Caticha, A., & Giffin, A. (2007). Updating probabilities with data and moments. Workshop on Bayesian Inference and Maximum Entropy Methods in Science and Engineering. NY, USA: Department of Physics, University at Albany–SUNY. Chisari , O.O., Mastronardi , L.J., & Romero, C.A. (2012). Building an input-output Model for Buenos Aires City. MPRA (40028). Deming, W.E., & Stephan, F.F. (1940). On a Least Squares Adjustment of a Sampled Frequency Table When the Expected Marginal Totals are Known. The Annals of Mathematical Statistics, 11(4):427–444. Eurostat. (2002). The ESA 1995 Input-Output Manual - compilation and analysis (draft). Luxembourg. Adom Giffin, A. & Caticha, A. (2007). Updating Probabilities with Data and Moments. Presented at the 27th International Workshop on Bayesian Inference and Maximum Entropy Methods in Science and Engineering, Saratoga Springs, NY, July 8–13, 2007. Gosh, A. (1964). Experiments with Input-Output Models, An Application to the Economy of the United Kingdom, 1948–1955. Cambridge: Cambridge University Press. Greenberg, E., Denzau , A.T., & Gibbons , P.C. (1989). Bayesian Estimation of Proportions with a Cross-Entropy Prior. Communications in Statistics - Theory and Methods, 18(5):1843–1861. Lemelin, A., Fofana, I., & Cockburn, J. (2005). Balancing a social accounting matrix: Theory and application.Centre Interuniversitaire sur le Risque des Politiques Economiques et l‟Emploi (CIPREE), Université de Laval. Leontief, W. (1941). The Structure of American Economy, 1919–1929. New York: Cambridge, (mors): Harvard University Press, (Second Ed. 1951, , Oxford University Press). Leontief, W. (1966). Input-Output Economics. New York: Oxford University Press. Leontief, W. (1970). The Dynamic Inverse. Contributions to Input-Output Analysis, pp. 17–46. Leontief, W. (1986). Input-Output Economics, Second Edition. New York: Oxford University Press. McDougall, R.A. (1999). Entropy Theory and RAS are Friends. Working Paper No. 6, GTAP. Miller, R., & Blair, P. (1985). Input-output analysis: Foundations and extensions. Englewood Cliffs, N.J.: Prentice-Hall. Robinson, S., Cattaneo, A., & El-Said, M. (2001). Updating and Estimating a Social Accounting Matrix Using Cross Entropy Methods,. Economic Systems Research, 13(1):47–64. Snower, D.J. (1990). New methods of updating input-output matrices. Economic System Research, 2:27–38. Stäglin, Reiner (1972), MODOP - Ein Verfahren zur Erstellung empirischer Transaktionsmatrizen, in: Münzner, H. and Wetzel, W.: Anwendung statistischer und mathematischer Methoden auf sozialwissenschaftliche Probleme, Würzburg, pp. 69–81. Stein, C., & James, W. (1961). Estimation with quadratic loss. Proc. Fourth Berkeley Symp.:361–379). Stone, R. (1955). Input-Output and the Social Accounts. Proceedings of an International Conference on Input-Outputs Analysis. New York: University of Pisa. Stone, R. (1981). Input-Output-Analysis and Economic Planning: A Survey. Mathematical Programming and its Applications. Milan: Angeli. Stone, R. (1984). Balancing the National Accounts: The Adjustment of Initial Estimates. Demand, Equilibrium and Trade. Toh, M.H. (1998). The RAS approach in updating input-output matrices: An instrumental variable interpretation and analysis of structural change. Economic Systems Research, 10(1):63–78.
Bibliography – Part III 85 Tomaszewicz , Ł. (1992). Przepływy miedzygałeziowe. Elementy teorii, Wydawnictwo UŁ (rozdz. 1–3, 7,). Łodż: University of Łodż. Tomaszewicz , Ł. (2005). Metody analizy input-output. Przegląd Statystyczny, 52(4), strony 15–22. United Nations. (1993). Integrated Environmental and Economic Accounting, Handbook of National Accounting. New York, USA: United Nations. Eurostat: http://ec.europa.eu/eurostat/web/esa-supply-use-input-tables/overview.
PART IV: Social Accounting Matrix
1 Position of the Problem After presentation, in the previous part of an input-output table in the context of illbehaved inverse problem solution, let us talk about the next national accounts table, known as a Social Accounting Matrix (SAM) (e.g., Stone, 1970; Graham, 1985; Scandizzoa & Ferrareseb, 2015). From a macroeconomic point of view, this table is a generalization of an input-output table. While input-output shows primary distribution of income without saying anything about its secondary distribution, the SAM table fills in this missing information and, at the same time, displays a complete flow of products and values in a general equilibrium framework. In the coming section, the ill-behaved inverse problem will be particularly addressed since SAM table building, more than for input-output tables, requires more information to be gathered from different sources but involving, consequently, higher risk of statistical inconsistencies when this information is aggregated. In this chapter, we present the general aspects of a SAM by introducing the parameter or multiplier estimation approach, and we then discuss the limits of classical econometric methods. For the last two decades, Kullback-Leibler minimum entropy (KLME) formalism has encountered relative success in the social sciences, particularly when a solution was required for inverse problems (e.g., Robinson et al., 2001; Bwanakare, 2013). The objective of the present document is to extend the KLME approach to the non-ergodic Tsallis entropy system, represented in the present study by an initially non-balanced quadratic social accounting matrix (SAM), known to display Walrasian general equilibrium features. The described economic system is then defined by different interactive subsystems, each represented by respective actors and characterized by optimizing behaviour. Households, which tend to maximize a certain utility function, remain the owner of factors of production and are the final consumer of produced commodities while firms maximize profits by optimal renting of these factors from households for the production of goods and services. In this model, government has the passive role of collecting taxes and disbursing tax revenue. Furthermore, the economy analysed is small and open and a ‘price taker’ from the rest of the world. The above optimal behaviour inside subsystems leads to general market equilibrium in all respective sectorial markets. A SAM table is a statistical tool to summarize all the above economic transactions by registering, in respective their rows and columns, income and expenses in accordance with the double-entry book-keeping principle. However, due to different and sometimes contradictory sources of collected statistical information, the SAM is not balanced, i.e., respective column or row totals are not matching. Such statistical data may display, as partially coming from statistical surveys, systematic errors, most of the time evidenced through a tail queue Gaussian distribution.
89 Position of the Problem 89 Since a SAM-based model contains more unknown parameters to estimate than the number of determined equations, updating and balancing such stochastically unbalanced matrices belongs to the category of the generalized inverse problem.
2 A SAM as a Walrasian Equilibrium Framework In this part of the present work, we are going to present different macroeconomic aspects of a SAM in the context of further macroeconomic models to be presented and estimated. It is useful to remind readers about Walrasian general equilibrium features of a SAM. This will increase understanding of its construction and improve interpretation of its post-estimation outputs. This point of interpreting outputs will appear with more acuity in the coming sections when, at the end of the estimation process by the maximum entropy principle, we have to interpret parameters of the estimated model. Though in the next chapter on the computable general equilibrium model we will examine the philosophical underpinnings of the SAM in terms of economic theory, we must immediately be aware that both the input-output table and the SAM were conceived as practical applications of the general equilibrium theory earlier introduced by Walras, one of initiators of the Australian school of thought. Brown & Stone (1962) has provided a definition of Walrasian hypotheses as follows: H1. Observed market demand is the sum of consumers’ demands derived from utility maximization subject to budget constraints at observed market prices. H2. There exists an observable (locally) unique equilibrium price system such that the observable market demand is equal to the observable market supply in every market. H3. The observed equilibrium price system is a (locally) stable equilibrium of trial-and-error price adjustment. The first hypothesis fixes the prerequisites under which Walrasian equilibrium is feasible. The second and third hypotheses specify quantitative relations which lead to equilibrium. Equilibrium in the economic flows results in the conservation of both product and value (Liossatos, 2004). Additionally, the three conditions of market clearance, zero profit, and income balance are employed by CGE modellers to solve simultaneously for the set of prices and the allocation of goods and factors that support general equilibrium. In terms of circular flows in economy, each row total of each economic sector is equal to a corresponding column total, and in that way a general equilibrium of macroeconomic aggregates is ensured. Research contributions in this area are, to our knowledge, very limited. In their pioneering work, Duncan & Smith (Duncan, 1999; Liossatos, 2004) found that Walrasian equilibrium is not guaranteed by a free market existence. Without implausibly strong restrictions on the production sets and preferences (for example, that the production sets do not exhibit increasing returns to scale, a pervasive feature of real tech-
Bibliography – Part III 91 nologies), the demand and supply correspondences may be empty for some prices and may be discontinuous, so that no equilibrium price system can be found. The question of the existence and stability of a Walrasian equilibrium encounters mathematical difficulties and paradoxes. In particular, the question of finding a robust stability in equilibrium prices has remained elusive, and the issue of the existence of Walrasian equilibrium has been settled only by introducing into the argument powerful abstract mathematical principles, which have no real economic foundation (see, e.g., Duncan & Smith, 2009). Taking the above into account, one may think that Walrasian equilibrium, at least on empirical grounds, is a kind of approximation of the Pareto efficiency benchmark. Using the maximum entropy principle, the above authors (Duncan, 1999; Liossatos, 2004) have recently tried to prove the existence, the uniqueness, and the stability of the Pareto optimum. The starting hypothesis is to consider "a set of feasible market transactions as typically large, that is, once the number of types, the number of traders of each type, and the number of points in the offer sets become moderately large. Furthermore, there are many different ways of assigning traders to transactions in their offer sets that clear (or approximately clear) the market. The principle of voluntary market exchange in and of itself is not sufficient to determine the market transaction. Thus, entropy equilibrium is a short-run, temporary equilibrium model of market exchange which replaces the Walrasian picture of the market in equilibrium as a budget hyper plane defined by equilibrium relative prices with a scalar field of transaction probabilities." (Brown & Shannon, 1997). Under these conditions and following the same authors, entropy prices clear the market by distributing agents over their offer sets, rather than moving agents to optimal commodity bundles in their consumption sets, and thus effectively “convexify” the economy. Furthermore, the fact that different traders experience different transaction prices implies that random statistical equilibrium does not exhaust all the potential Pareto-improving transactions in the economy. Thus, such a statistical equilibrium approximates, but does not achieve, Pareto-efficiency. The statistical equilibrium in this market fails to achieve Pareto-efficiency because some potentially mutually advantageous transactions fail to be executed, and there is dispersion in actual transactions prices. To achieve the Pareto optimum, the pioneering work of Foley (1994) has attempted to endogenize the offer sets of economic actors in a rational expectation framework: “If we imagine a given agent repeatedly entering a market in statistical equilibrium, it is tempting to suppose that she will alter her offer set in order to optimize her market outcome given the probabilities that govern transactions in the market equilibrium. This idea gives rise to the concept of endogenous offer sets.”
92 A SAM as a Walrasian Equilibrium Framework According to the same author, the present state of research should be that we rigorously establish that Walrasian equilibrium is the asymptotic outcome of a process in which endogenous offer sets adapt to statistical equilibrium entropy prices, and where the chances to transact in any period become numerous. Walrasian equilibrium is not unique, whereas statistical equilibrium for given offer sets of traders should be unique. The adaptation process sketched out here allows offer sets to change over time, giving rise to a dynamic process which may have multiple equilibria, each corresponding to a Walrasian equilibrium in markets with multiple Walrasian equilibria. We argue that the problem should be placed in the context of non-additive statistics, suggesting that agent behaviours are time or space dependent. In fact, the above Gibbs-Shannon entropy related price suggests an ergodic system in which agent actions are disconnected from long-run memory of past market events and/or confluent—space related—information from surrounding markets. Then, system complexity describing market transactions should lead to non-extensive entropy related prices. Further research is needed to understand to what extent macroeconomic equilibrium and thermodynamic equilibrium are comparable in the context of non-ergodic real world systems. To conclude, there is no assurance that the balanced post-entropy social accounting matrix is achieving (or approximating) the Pareto Optimum. As stated above, some probabilistic distributions of price entropy may be meaningless or not optimal in the convex space of all possible transactions—not even the voluntarily contracted ones. More investigation is required to better appreciate, among other possibilities, the new approach of endogenizing dynamic offer sets, as suggested above. For the time being, entropy econometrics as presented in the coming sections seems to be the best approach to resolve such ill-behaved inverse problems.
Shannon-Kullback-Leibler Cross-Entropy 99 It is assumed that the entropy problem starts with a prior A which plausibly is a SAM from a previous period or, as in this case, a raw and unbalanced SAM. A represents the starting point from which the cross-entropy balancing procedure departs in deriving the new matrix of coefficients A. The entropy problem is to find a new set of A coefficients which minimize the so-called Kullback-Leibler (1951) divergence measure of the ‘cross-entropy’ (CE) between the prior A* and the posteriori coefficients matrixA. ij i j i j ijijij i j ij ij ij AAAA A A AAAI * * * lnlnln)||(min (4.3) subject to ii j ij tottotA (4.4) ii j ij tottotA = 1 and 0 ≤ Aij ≤ 1 (4.5) Note that, according to Walras’s law in general equilibrium theory, one equation can be dropped in the second set of constraints: If all but one column and row sums are equal, the last one must also be equal. The solution of the above problem is obtained by setting up the Lagrangian. The k macro-aggregates can be added to the set of constraints on the problem above as follows: )( )( k i j ij k ij T (4.6) where H is an nxn aggregator matrix with ones for cells that represent the macroconstraints and zeros otherwise, and γ is the value of the aggregate constraint. As mentioned above, in the real world one faces economic data measured with error. The cross-entropy problem can also be formulated as an ‘error-in-variables’ system where the independent variables are measured with noise e. If, for example, we assume the known column sums are measured with error, the row/column consistency constraint can be written as: totj = xi + ei (4.7) where totj is the vector of row sums and xi, the known vector of column sums, is measured with error ei. The prior estimate of the column sums could be, for instance, the initial column sums, the average of the initial column and row sums, or the row sums. Following Golan et al. (1996), the errors are written as weighted averages of known constants v defined over a finite discrete support space m>>1,...,M with points: im Mm imi vfe ,,1 (4.8)
100 Balancing a SAM where fim is a set of weights that fulfil the following constraints: Mm im f ,,1 1 = 1 and 0 ≤ fim ≤ 1 (4.9) In the estimation problem, the weights are thus treated as probabilities to be estimated, and the prior for the error distribution in this case is chosen to be a symmetric distribution around zero with predefined lower and upper bounds, and using either three or five weights. Naturally, not only the column and row sums can be measured with error, the macro-aggregates by which we constrain our estimation problem may also be measured with error, and so we can operate with two sets of errors with separate weights f1’s on the column sum errors, and weights f2’s on the macro-aggregate errors. The optimization problem in the ‘errors-in-variables’ formulation is now the problem of finding A’s, f1’s, and f2’s that minimize the cross-entropy measure, including terms for the error weights: ij i j i j ijijij AAAAffAAI * 21 * lnln)||||(min + i Mm iJim i Mm imim ffff ,, 11 ,, 11 lnln i Mm iJim i Mm imim ffff ,, 22 ,, 22 lnln (4.10) Referring once again to the definition of information provided by Kullback and presented in the second part of this book, cross-entropy measurement reflect how much the information we have introduced has moved our solution estimates away from the inconsistent prior, while also accounting for the imprecision of the moments assumed to be measured with error. Hence, if the information constraints are binding, the distance from the prior will increase. If they are not binding, the cross-entropy distance will be zero. It becomes now clearer why we have proposed, while assessing forecasting performance of the entropy technique, the difference between the average error variance coefficients (AEVC) of the periods 2006 and 2007 as a benchmark, maximum divergence precision measurement. 4.2 Balancing a SAM Through Tsallis-Kullback-Leibler Cross-Entropy In the following, we are going to generalize the Jaynes-Kullbac-Leibler model (4.10) and thus reconsider all the implications of the above theorem on the power law property of economy.
A SAM as a Generalized Input-Output System 101 Let us formally explain Tsallis relative entropy model (2.47-2.50) to be minimized together with the above suggested constraints. In this presentation, the Bregman form of relative entropy (2.47) will be used: 111 111 ))(()()( 1 1 )1( ))(()()( 1 1 ||||min q ih h ihih q ih q ih h ih q iJ i iJ iJ q iJ q iJ i iJq wowowwoww q AAAAAA q wowAAI (4.11) subject to: ij j ij tottotA (4.12) 1 2 N Mi ij A (4.13) 1 ,,1 Ih ih w (4.14) Symbols are as in equ. (4.9), except wih, which takes the place of f2iJ, and both represent disturbance errors on parameters but, this time, of different distribution laws. Empirical, long practice with this class of economy-wide models provides some prior information on relevant ranges for parameter values and likely parameter estimates. Furthermore, while the support of any imposed prior distribution for a parameter is a maintained hypothesis (the estimate must fall within the support), the shape of the prior distribution over that support (e.g., the weights on each support point) is not. Unless the prior is perfect, the data will push the estimated posterior distribution away from the prior. The direction and magnitude of these shifts are, in themselves, informative. Also, note from Equations (4.11– 4.14) that, with increases in the number of data points, the second term of prediction in the objective function increasingly dominates the first term precision. In the limit, the first term in the objective becomes irrelevant. The prior distributions on parameters are only relevant when information is scarce. 4.3 A SAM as a Generalized Input-Output System In the present paragraph, Kullback-Leibler (K-L) information divergence is extended to Tsallis non-ergodic systems and a q-Generalization of the K-L relative entropy criterion function (c.f.), with a priori consistency constraints, is derived for balancing a SAM as a generalized input-output transaction matrix. On the basis of an unbalanced, Gabonese social accounting matrix (SAM) representing a generalized inverse problem input-output system, we propose to update and balance it following the procedure explained through the above section.
102 Balancing a SAM 4.3.1 A Generalized Linear Non-Extensive Entropy Econometric Model This section applies the results of, e.g., Jaynes (1957) and Golan et al. (1996) to present the model to be later implemented for updating and balancing input-output systems. While the argument in the criterion function is already known (see Equation 4.18), we need to reparametrize35 the generalized linear model, to be introduced later into the model as restrictions in the spirit of Bayesian method of moments (e.g., Zellner, 1991). Note that such a linear restriction will be affected by a stochastic term expected to belong to the larger family of power law distribution. Let us succinctly present the general procedure for parameter reparametrization as it follows: Y = X ⋅β + ε (4.15) Parameter β in general bears values not constrained between 0 and 1. When this is the case, reparametrization will no longer be necessary since parameter variation area fits well to probability definition area. The variable ε is an unobservable disturbance term with finite variance, owing to the economic data nature of exhibiting observation errors from empirical measurement or from random shocks. These stochastic errors are assumed to be driven by a large class of PL. As in classical econometrics, variable Y represents the system, the image of which must be recovered, and X accounts for covariates generating the system with unobservable disturbance ε to be estimated through observable error components e. Unlike classical econometric models, no constraining hypothesis is needed. In particular, the number of parameters to be estimated may be higher than the observed data points and the quality of collected information data low. Additionally, as already explained, to increase the accuracy of such estimated parameters from the poor quality of data points, the entropy objective function allows for incorporation of all constraining functions which act as Bayesian a priori information in the model. Let us treat each βk(k = 1,...,K) as a discrete random variable with compact support and 2 < M < ∞ possible outcomes. Thus, we can express βk as: M m kmkmk vpB 1 Kk (4.16) where pkm is the probability of outcome vkm and the probabilities must be non-negative and sum up to one. Similarly, by treating each element ei of e as a finite and discrete 35 Reparametrization aims at treating parameters of the model as outputs of probability distribution to be estimated following the procedure presented by Golan et al. (1996) and later exploited for modelling many entropy econometric models (see, e.g., Bwanakare, et al. (2014, 2015, 2016). Since the same probabilities are related to entropy variable defining the criterion function, optimizing the whole model then leads to outputs taking into account stochastic a priori information owing to model restrictions.
A SAM as a Generalized Input-Output System 103 random variable with compact support and 2 < M < ∞ possible outcomes centred around zero, we can express ei as: Jj njnji zre ,,1 (4.17) where rn is the probability of outcome zn on the support space j. Following Bwanakare (2014), we will use the commonly adopted index n, here and in the remaining mathematical formulations, to set the number of statistical observations. Note that the term e can be initially fixed as a percentage of the explained or endogenous variable, as an a priori Bayesian hypothesis. Posterior probabilities within the support space may display non-Gaussian distribution. The element vkm constitutes an a priori information provided by the researcher while pkm is an unknown probability whose value must be determined by solving a maximum entropy problem. In matrix notation, let us rewrite β = V⋅P with pkm ≥ 0 and K k Mm km p 1 ,,1 1 pkm = 1, where again,is the number of parameters to be estimated and the number of data points over the support space. Also, let e =r ⋅w with rnj ≥ 0 N n J Jj nj r 1 ,,1 1 and rnj = 1 for N the number of observations and the K number of data points over the support space for the error term. Then, the Tsallis cross-entropy econometric estimator can be stated as: 1 1/ 1 1/ 1 1/ )||,||,||( 1 1 1 000 q wow w q ror r q pop pwwrrppMinH q tsts ts q njnj nj q kmkm kmq 1 1/ 1 1/ 1 1/ )||,||,||( 1 1 1 000 q wow w q ror r q pop pwwrrppMinH q tsts ts q njnj nj q kmkm kmq (4.18) S. to J j q nj q nj J j j M m q km q km M m m r r z p p vXeXY 1 1 1 1 (4.19) K k Mm km p 1 ,,2 1 N n J Jj nj r 1 ,,2 1 T t S Ss ts w 1 ,,2 1 K k Mm km p 1 ,,2 1 N n J Jj nj r 1 ,,2 1 T t S Ss ts w 1 ,,2 1
104 Balancing a SAM K k Mm km p 1 ,,2 1 N n J Jj nj r 1 ,,2 1 T t S Ss ts w 1 ,,2 1 (4.20) Additionally, k macro-aggregates can be added to the set of constraints as follows: T t q ts q ts S s s d i j ij ij d w w gT 1 1 )()( , (4.21) where H is an dxd aggregator matrix with ones for cells that represent the macro-constraints and zeros otherwise, and γ is the expected value of the aggregate constraint. Once again, gs stands for a discrete point support space from s = 2,...,s. Probabilities wts stand for point weights over gs. The real q, as previously stated, stands for the Tsallis parameter. Above, Hq(p||p0, r||r0, w||w0) is nonlinear and measures the entropy in the model. Relative entropies of three independent systems (three posteriors p, r, and w and corresponding priors p0, r0, and w0) are then summed up using weights αβδ. These are positive reals summing up to unity under the given restrictions. We need to find the minimum divergence between the priors and the posteriors while the imposed stochastic restrictions and normalization conditions must be fulfilled. As will be the case in the application below, the first component of the criterion function may concern the parameter structure of the table; the second component errors on column (or row) totals and the last component may concern errors around any additional consistency variable, such as the GDP in the case below. As it has been shown by Tsallis (2009), this form of entropy displays the same basic properties as K-L information divergence index or relative entropy. The estimates of the parameters and residual are sensitive to the length and position of support intervals of β parameters (Equations 4.16 and4.17) in the context of the Bayesian prior. When parameters of the proposed model are expressed under the form of elasticity or ratio—as will be the case in the example below—then the support space should be defined inside the interval between zero and one and will fit that of the usual probability variation interval. In such a case, no reparametrization of the model is needed. In general, support space will be defined between minus and plus infinity, according to the prior belief about the parameter area variation by the modeller. Additionally, within the same support space, the model estimates and their variances should be affected by the support space scaling effect, i.e., the number of affected point values (Foley, 1994). The higher the number of these points, the better the prior information about the system. The weights αβδ are introduced into the above dual objective function. The first term of “precision” accounts for deviations of the estimated parameters from the prior (generally defined under a support space). The second and the third terms of “prediction
Input-Output Power Law (Pl) Structure 105 ex-post” account for the empirical error term as a difference between predicted and observed data values of the model. As expected, the presented entropy model is an efficient information processing rule that transforms, according to Bayes’s rule, prior and sample information into posterior information (Ashok, 1979). 4.4 Input-Output Power Law (Pl) Structure It is time now to come back to the fundamental problem concerning the true statistical nature of input-output data used in the above studies or those below. In recent years, as already explained in Part I, many studies (Champernowne, 1953; Gabaix, 2008) have shown that a large array of economic laws take the form of a PL, in particular macroeconomic scaling laws, distribution of income and wealth, size of cities, firms36, and the distribution of financial variables, such as returns and trading volume. Stanley and Mantegna (2007) have studied the dynamics of a general system composed of interacting units each with a complex internal structure comprising many subunits where the latter grow in a multiplicative way over a period of twenty years. They found the system follows a PL distribution. Such outputs should present similarities with the internal mechanism of national accounts tables, such as an input output table or a SAM. A PL displays, besides its well-known scaling law, a set of interesting characterizations related to aggregative properties of a PL according to which a power law is conserved under addition, multiplication, polynomial transformation, and minimum and maximum. As far as the PL hypothesis for a SAM is concerned, taking into consideration the above literature and using PL properties, it should not be difficult to prove the PL character of a SAM, including the Gaussian trivial case. About SAM construction and components, see for example, Pyatt (Pyatt, 1985). General equilibrium (Wing Ian Sue , Sept 2004) implies that respective row and column totals are expected to balance. Conceptually, this model is based on the laws of product and value conservation (Serban and Blake (Serban Scrieciu & Blake., 2005.)) which guarantee conditions of zero profits, market clearance, and income balance. However, different stages of statistical data processing remain concomitant with human errors and the SAM will not balance. It is generally assumed that the main sources of these imbalances remain different sources of documentation and different time of data collecting. This means that an unknown number of economic transaction values within the matrix are inconsistent with the data generating macroeconomic system. For clarity, let us use Table 13 to explain these imbalances, noting, for instance, a difference between the institution row and column totals as follows: (iT + e4) – (iT + ε4) = (e4 – ε4) (4.22) 36 See Bottazzi et al. (2007) for a different standpoint on the subject.
106 Balancing a SAM The term on the left hand side of the above expression represents the difference between two erroneous and unequal totals of institution account. The origin of that difference results from difference between plausibly different stochastic errors e4 and ε4, respectively, on column and row totals. In Table 11, the first alphabetical letter of symbols inside each cell represents the first letter of the row (supply) account, and the second letter represents the first letter of the corresponding (demand) column. In the SAM prototype below, e.g., the symbol “Ca”, explains purchases by the activity sector of goods and services from the commodity sector. The targeted purpose is to find, out of all probability distributions, a set of a posteriori probabilities closest to a priori initial probabilities and insure the balance of the SAM table while satisfying other imposed consistency moments and normalization conditions. Following Shannon terminology, one may consider post-entropy structural coefficients and disturbance errors, respectively, as signal and noise. The first step consists in computing a priori coefficients by column from real data from Table 11 by dividing each cell account by the respective column total. Next, we treat these column coefficients as analogous to probabilities and column totals as expected column sums, weighted by these probabilities (see Equation 4.19). These coefficient values will serve as the starting, best prior estimates of the model. The other two types of priors to initialize the solution concern errors on column totals (Equation4.17) and on gross domestic product (GDP) at factor and market prices (Equation4.21). GDP variables are added to the model with the purpose of binding the latter to meet consistency macroeconomic relationships for different accounts inside the SAM. Other macroeconomic relations like those affecting interior or global consumptions could be added. The proposed approach combines non-ergodic Tsallis entropy with Bayes’ rule to solve a generalized random inverse problem. We may optionally consider only Table 11: General structure of a stochastic non-balanced SAM Activities Commodities Factors Institutions Capital World Total Activities 0 Ac 0 Ai 0 aw aT+ ε1 Commodities Ca 0 0 Ci cc 0 cT+ ε2 Factors Fa 0 0 0 0 0 fT+ ε3 Institutions Ia Ic If ii 0 iw iT+ ε4 Capital 0 0 0 ci 0 cw cT+ ε5 World 0Wc 0wi 0 0 wT+ ε6 Total aT+ e1cT+ e2fT+ e3iT+ e4cT+ e5wT+ e6 Source: own elaboration.
Balancing a SAM of a Developing Country: the Case of the Republic of Gabon 107 some cell values as certain37 while the rest of the accounts are unknown. This is one of the strongest points of the entropy approach over other mechanical techniques of balancing the national accounts table through a stochastic framework. All row and column totals are known with uncertainty. It is straightforward to notice that the potential freedom degree number of parameters to estimate (n – 1) (n – 1) remains significantly higher than n observed data points (column totals). In a particular case of a SAM, and due to empty cells, that number of unknown parameters may be much lower. In any event, that will not generally prevent us from dealing with an ill-behaved inverse stochastic problem. The next important step is initializing the above defined errors through a reparametrizing process. A five point support space symmetric around zero is defined. To scale the error support space to real data, we apply Chebychev’s inequality and three sigma rules (Serban Scrieciu & Blake, 2005). Corresponding optimal probability weights are then computed so as to define the prior noise component (Robinson et al., 2001). 4.5 Balancing a SAM of a Developing Country: the Case of the Republic of Gabon In our analysis of the last cases, we have rather underscored technical aspects of entropy for balancing input-output tables. However, when statistical data from different sources are available and sufficiently consistent, applying a complex procedure as the one relying on entropy formalism can be more time consuming than relatively easier techniques like the RAS (e.g., Pukelsheim, 1994, Bacharach (1970). This is the case for many developed countries where statistical data gathering is generally efficient38. In reverse, as we are going to see in the coming pages, this is not the case for the majority of developing countries in which statistical data are not only scarce but also of bad quality. Thus, to complete an array of empirical advantages of the proposed entropy approach, we are going to analyse the case of developing countries where complete 37 Only transaction accounts with the rest of the world (import, export, external current balance), plus government commodity consumption accounts are concerned. 38 The statistical data gathering system in Poland can been seen as relatively efficient in comparison with those of most of developing countries. Availability of data on a large scale and their quasiconsistency though from various unrelated sources is the criterion retained here for giving such an appraisal. As a result, it should be relatively easier to balance national account tables without using complicated procedures such the entropy-related one. Thus, inconsistencies displayed in Table 1 may not reflect outputs from other publications on the same subject. The purpose of the present example is just to show the performance of the cross-entropy procedure in balancing a system under constraining, a priori information, like different macroeconomic identities characterizing national account tables.
108 Balancing a SAM statistical information is generally unavailable. Not only does such information not fully exist, what does should be approached with a high level of uncertainty. Applying traditional balancing techniques, like the RAS approach, becomes in practice difficult. Based on the Shannon entropy approach, a large number of studies—particularly from developing countries—designed to balance SAM tables have been prepared in the last two decades. The already cited paper of Robinson et al. (2001), consecutive to the publications of Golan et al. (1996), has become a reference work for having shown an algorithm—in GAMS code (General Algebraic Modelling System)—for balancing a SAM in the case of uncertainty. One can list other studies with identical purpose, such as those of of Salem (2004) for Tunisia and Kerwat et al. (2009) for Libya. Murat (2005), using Shannon cross-entropy formalism, has balanced a Turkish Financial Social Accounting Matrix and, more recently, Miller et al. (2011) has built and balanced a disaggregated SAM for Ireland. Note that these last two countries belong, respectively, to the category of intermediary developed and developed countries. Many other entropy-based studies have been presented for various countries like Malawi, South Africa, Zimbabwe, Ghana, Gabon, and Vietnam. The results shown below generalizes, once again, Shannon formalism by applying a non-extensive entropy divergence formalism. 4.5.1 Balancing the SAM of GABON by Tsallis Cross-Entropy Formalism A complete description of data sources or others details concerning the methodology of building the aggregated and disaggregated SAM of Gabon can be found in Bwanakare (2013).39 That methodology has been proposed by Robinson et al. (2001) for balancing the SAM of Mozambique. Briefly, it consists of two steps in building the final SAM. In the first step, an aggregate and unbalanced SAM is built on the basis of official macroeconomic data. The later will serve as a control in building a much more disaggregated SAM in which accounts will be obtained by splitting out aggregated accounts of the balanced40 SAM of the first step. Table 12 below represents the initial aggregated and unbalanced SAM of Gabon. Statistical data come from three sources: the Ministry of Planning, the Ministry of Economy and Finance, and the Bank of Central Africa States. 39 This document was prepared with the help of the Directorate of National Accounting at the Ministry of Planning and Development of the Republic of Gabon. A copy of the outputs of the balancing of this SAM has been transmitted to the Ministry. See the document at http://www.numilog. com/236150/Methodologie-pour-la-balance-d-une-matrice-de-comptabilite-sociale-par-l-approcheeconometrique-de-l-entropie---le-cas-du-Gabon (ebook: Paris: Editions JePublie) 40 We have used the cross-entropy technique for balancing such an aggregated SAM.
About the Extended SAM 115 Table 15: Percentage of Information divergence between tables 13 and 14. aACT aPOLLABAT pCOM pPOLLABAT LABOR CAPITAL POLLFEES HOU ENT GRE CAPAC CAPACENVROW Total aACT 0 0,00 0,09 0,00 0,00 0,00 0,00 3,01 0,00 0,00 0,00 0,00 0,00 0,10 aPOLLABAT 0 0,00 0,00 0,76 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,76 COM -0,23 0,00 0,00 0,00 0,00 0,00 0,00 2,92 0,00 -4,30 -3,35 -0,11 -0,21 0,32 pPOLLABAT -9,27 1,63 0,00 0,00 0,00 0,00 0,00 -1,42 0,00 0,00 0,00 0,00 0,00 -3,61 LABOR -0,65 1,77 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 -0,65 CAPITAL 1,40 1,79 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 1,41 POLLFEES 0,00 0,73 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,73 HOU 0,00 0,00 0,00 0,00 -0,65 -2,49 0,00 0,00 -13,02 -4,27 0,00 0,00 -0,19 -2,74 ENT 0,00 0,00 0,00 0,00 0,00 6,57 0,00 0,00 0,00 3,45 0,00 0,00 7,40 6,54 GRE 1,36 0,00 1,82 0,00 0,00 -1,33 0,73 5,14 0,00 -3,29 0,00 0,00 0,77 2,80 CAPAC 0,00 0,00 0,00 0,00 0,00 0,00 0,00 15,23 -9,83 1,46 0,00 0,00 5,45 -0,33 CAPACENV0,00 0,00 0,00 0,00 0,00 0,00 0,00 0,00 -6,87 6,64 0,00 0,00 0,00 -0,11 ROW 0,00 0,00 0,88 2,82 -0,53 0,00 0,00 3,33 -12,94 -4,12 0,00 0,00 0,00 0,06 Total 0,10 0,76 0,27 0,84 -0,65 1,41 0,73 4,12 -11,02 -3,99 -3,35 -0,11 0,06 Source: own calculations.
116 Balancing a SAM Glossary of table abbreviations: aAct: activity sector aPOLLABAT: abatement ecological activity sector pCom: commodity sector pPOLLABAT: abatement ecological commodity sector Labor: labor sector (factor of production) Capital: capital sector (factor of production) Pollfees: pollution fees sector (factor of production) Hou: households institution Ent: enterprise institution GRE: government institution CapAc: capital accumulation sector (private investment). CAPACENV: abatement -oriented capital accumulation sector (private investment) RoW: rest of the world institution.
5 A SAM and Multiplier Analysis: Economic Linkages and Multiplier Effects 5.1 What are the Economic Linkages and Multiplier Effects? The strongest argument in favour of the Walras equilibrium—as opposed to the Marshall ceteris per ibis approach—will find its momentum once industry linkages and multiplier effects are envisaged. This is so because in these circumstances thinking about partial equilibrium becomes less sustainable. In fact, the effect of a shock from one industry may have direct and indirect impact on the whole system defined by different industries. Let us analyse below a shock generated by the demand side. When we talk of “exogenous demand-side shocks” to an economy, we refer to changes to final control demand aggregates, i.e., export demand, government spending, or net investment demand of stocks. The effects of these shocks are both direct and indirect. The direct effects are to those sectors that affront the shock. For example, an exogenous increase in demand for Polish manufactured exports has a direct impact on the manufacturing industry, which results in increased inputs, production, sales, and value-added. However, the positive consequences of such a shock go beyond the manufacturing industry. It may also have indirect effects stemming from manufactures’ linkages to other industries inside the economy. These indirect linkages can be classified into supply-side and demand-side. When we add up all direct and indirect linkages, we get a measure of the shock’s multiplier effect, or how much an initial effect is amplified or multiplied by indirect linkage effects. Supply-side linkages are determined by industry production technologies, which can be depicted from an input-output table. Next, they are differentiated into backward and forward linkages. Backward production linkages are the demand for additional inputs used by producers to supply additional goods or services. For instance, when production (of manufacturers) expands, it requires additional intermediate goods or services like raw material, machinery, and transport services. This demand then stimulates production of other industries that supply these intermediate goods. Technical coefficients supply information on the input intensity of the production technology used. The more an industry’s production technology is input intensive, the stronger its backward linkages. Forward linkages allude to supply inputs to upstream industries. For instance, increased manufacturer production should lead to increased supply of goods to the construction industry, which, in turn, stimulates, among others, service industries. As in the case of backward linkages, the more important an industry is regarding upstream industries, the stronger its forward linkages will be and multipliers will definitely become larger.
118 A SAM and Multiplier Analysis: Economic Linkages and Multiplier Effects The conceptual structure of the input-output matrix only allows for deriving multipliers that measure the effects of supply linkages. Since the input-output table does not show secondary income distribution, it is not possible to consider consumption linkages, which arise when an expansion of production generates additional incomes for factors and households, which are then used to purchase goods and services. Continuing the same example as above, when manufacturing production expands, it raises households’ incomes, which are used to buy consumer goods. Depending on the share of domestically produced, tradable, and imported goods in households’ consumption baskets, domestic producers benefit from greater demand for their products. The size of consumption linkages depends on various factors, including the share of net factor income distributed to households; for an open economy, the level of gross domestic product per inhabitant, which exercises an influence on the composition of the consumption basket; and the relative price between locally produced and imported goods which determines in Armington fashion the share of domestically supplied goods in consumer demand. Consequently, SAM multipliers tend to be larger than input-output multipliers because they capture both production and consumption/income linkages. Following Breisinger et al. (2009), “while economic linkages are determined by the structural characteristics of an economy (evidenced through technical coefficients and/ or the composition of households’ consumption baskets) and remain thus static, multiplier effects capture the combined dynamic effects of economic linkages over a period of time through different auto-generated rounds”. Three types of multipliers are generally reported in empirical research. First, an output multiplier combines all direct and indirect (consumption and production) effects across multiple rounds and reports the final increase in gross output of all production activities. Second, a GDP multiplier measures the total change value-added or factor incomes caused by direct and indirect effects. Finally, the income multiplier measures the total change in household incomes. The dampening path of multipliers is consecutive to the level of leakage inside economic circular flows. Ultimately, higher leakages stemming from income allocated to imported goods or from government taxes make the round-by-round effects slow down more quickly and reduce the total multiplier effect. In empirical research, one must often deal with two kinds of economic hypotheses. First, still in the context of the above example, one can suppose that demand shock will encounter no constrained response from the supply side. The second case is the one where demand shock is constrained. This can happen when supply is not able to completely satisfy increased demand. In this hypothesis, multipliers will follow a modified dynamic path towards slowing down. Let us still follow Breisinger et al. (2009) and then succinctly analyse both cases.
What are the Economic Linkages and Multiplier Effects? 119 5.1.1 A SAM Unconstrained Multiplier Let us present below a simplified SAM where presented accounts are just those required to derive a multiplier matrix. Table 16: A simplified SAM for multiplier analysis Activities Commodities Factors Households Exogenous demand Total A1 A2 C1 C2 F H E A1 X1 X1 A2 X2 X2 C1 Z11 Z12 C1E1Z1 C2 Z22 Z22 C2E2Z2 F V1V2 V H V1 + V2 Y E L1L2 S E Total X1X2Z1Z2V Y E Source: own elaboration, based on Breisinger, Thomas, and Thurlow (2009). We divide columns by their total to derive the coefficient matrix (M-matrix) excluding the exogenous components of demand. Table 17: Transformed Table 16 Activity Commodities Factors Households Exogenous demand A1 A2 C1 C2 F H E A1 b1=X1/Z1 A2 b2=X2/Z2 C1 a11=Z11/X1a12=Z12/X2c1=C1/Y E1 C2 a21=Z21/X1a22=Z22/X2c2=C2/Y E2 F v1=V1/X1v2=V2/X2 H 1=(V1 + V2)/V E l1=L1/Z1l2=L2/Z2s=S/Y Total 1 1 1 1 1 1 E Source: own elaboration, based on Breisinger, Thomas, and Thurlow (2009).
120 A SAM and Multiplier Analysis: Economic Linkages and Multiplier Effects Symbols: a) Values: X: Gross output of each activity (i.e., X1 and X2) Z: Total demand for each commodity (i.e., Z1 and Z2) V: Total factor income Y: Total household income E: Exogenous components of demand shares b) Share: a: Technical coefficients b: Share of domestic output in total demand v: Share of value-added or factor income in gross output l: Share of the value of total demand from imports or commodity taxes c: Household consumption expenditure shares s: Household savings rate To derive equations representing the relationships in the above SAM, we start by setting up simple demand equations: Z1 = a11X1 + a12X2 + c1Y + E1 Z2 = a21X1 + a22X2 + c2Y + E2 (4.23) Total demand = intermediate demand + household demand + exogenous demand. The next relationships tell us that domestic production X is only part of total demand Z. X1 = b1Z1 X2 = b2Z2 Since household income Y depends on the share each factor earns in each sector, then: Y = v1X1 + v2X2 or, Y = v1b1Z1 + v2b2Z2 Now replacing all X and Y in Equation (4.23), moving everything except for E onto the left-hand side, and grouping Z together, we finally obtain: (I – M)Z = E, (4.24)
What are the Economic Linkages and Multiplier Effects? 121 where 222222 221212 112121 111111 1, ,1 bvcba bvcba bvcba bvcba MI . We note that M is a square matrix, the elements (share values) of which are not negative. Each column sum (see Table 16) is less or equal to unity. Thus, an inverse matrix of (I – M) exists and should display non-negative values, suggesting the non-negativity of the multiplier matrix. Formally, from (4.24) we directly get the final multiplier equation of the form: Z = (I – M)–1E (4.25) Total demand = multiplier matrix × exogenous demand The above formulation tells us that when exogenous demand E increases, one will end up with a final increase in total demand equal to Z, owing to all the direct and indirect multiplier effects (I – M)–1. 5.1.2 Equation System for Constrained SAM Multiplier Often when factor allocation is not optimal, exogenous demand shocks may encounter limited response from producing sectors. Let us analyse below how much a multiplier will change if some producing sectors are unable to correctly respond. The expected issue is that if we fix one of two sectors Z, for instance Z2. In that case, imports should substitute for domestic supply, thus eliminating any growth linkages from this sector. The next equation is related to the non-constrained case and expresses total demand as the sum of its parts. (1 – a11b1 – c1v1b1)Z1 + (– a12b2 – c1v2b2)Z2 = E1 (– a21b1 – c2v1b1)Z1 + (1 – a22b2 – c2v2b2)Z2 = E2 Grouping exogenous terms on the right-hand side (i.e., E1 and Z2) and rearranging42, we finally obtain: 2 1 1 2 1 Z E BMI E Z (4.26) where 1,( 0,1( 112121 111111 bvcba bvcba MI 42 For derivation details, see Breisinger, Thomas, and Thurlow (2009).
122 A SAM and Multiplier Analysis: Economic Linkages and Multiplier Effects and 222222 221212 1,0 ,1 bvcba bvcba B Interpretation of the above equation is the following: an exogenous increase in demand for the unconstrained sectors [E1] leads to final increase in total demand for these sectors [Z1], including all of the forward and backward linkages (I – M*)–1. For the sectors with constrained supply (in our case sector Z2), it is net exports that decline. This means that the current trade balance must worsen if we have to amortize demand shock in the case of constrained supply. If exports remained unchanged, then the alternative of reducing exports would be increasing imports so as to meet additional exogenous demand in the context of this constrained supply. 5.1.3 On Modelling Multiplier Impact for an Ill-Behaved SAM Let us now return back to the central problem of this presentation and suppose that the matrix is unbalanced, which implies that multiplier values are not reliable. The way to avoid this should consist of only estimating parameters of the model without taking into account the obligation that the whole SAM be internally consistent. Thus, we should maximize (or minimize) entropy for probabilities related to the multiplier matrix under traditional restrictions, plus an additional constraint declaring values of an already balanced SAM to be taken as a prior. Remembering about the interpretation of estimated parameters through the maximum entropy principle, it would be easy to make a link between the multiplier effect and maximum entropy modelling. In fact, in a linear model, parameters estimated by entropy formalism are interpreted as the long-run (equilibrium) impact of one unit change of regressor x on regresand y. Thus, long-run impact means that direct and indirect effects of the multiplier are accounted for with respect to the shock. Annex C. Proof of Economy Power Law Properties 1. Definition of Power Law Distribution Since we already know existing relationships between power law function and nonextensive entropy from Part II of this work, let us now present the main properties of the former in the context of a SAM. Using a simplified formulation, a power law is the relation of the form f(x) = Kxα where x .00 qq 0 and K and α are constants. While power laws can appear in many dif-
Annex C. Proof of Economy Power Law Properties 123 ferent contexts, the most common are those where f(x) describes a distribution of random variables or the autocorrelation function of a random process. The formulation above has the advantage of being intuitive. However, it does not show the real attributes of that distribution, which displays asymptotical characteristics. Thus, the notion of a power law as it is used in extreme value theory is an asymptotic scaling relation. Let us first explain what we understand by equivalent scaling. Two functions f and g have equivalent scaling, f(x) ~ g(x) in the limit x → ∞43 if: x xg xfxL 1 )( )()( lim (4.27) x → ∞ with L(x) is a slowly varying function, thus satisfying the relation: x xL txL1 )( )( lim , x → ∞ for any finite constant t > 0. Slowly varying functions are, for example, L(x) = C and L(x) = ln(x), that is, a constant and a logarithmic function, respectively. A power law is defined as any function satisfying f(x) ~ xα. This definition then implies that a power law is not a single function but an asymptotical composite function. The slowly varying function L(x) can be thought of as the deviation from a pure power law for finite x. For f(x) = L(x)xα, taking logarithms of both sides and dividing by log(x) gives log f(x)/log (x) = –α + log L(x)/log (x) (4.28) Remembering that L(x) is a slowly varying function, in the limit, the second term on the right vanishes to zero as x → ∞, and thus we have: log f(x)/log (x) = –α, or equivalently, f(x) = x–α, for x → ∞. This means that the empirical form of the function becomes: f(x) ~ x–α (4.29) and in terms of probabilities, a the cumulative function P(S > x) = kx–α corresponds to a probability density function: f(x) = kαx–(α+1). 43 Note that this limit is not the one possible, but remains a realistic device, e.g., in finance.
124 A SAM and Multiplier Analysis: Economic Linkages and Multiplier Effects 2. Main Properties We list below only properties directly related to two theorems proposed in this annex. a) The property that most interests us and that generally makes power laws special is that they describe scale free phenomena. A variable undergoes a scale transformation of the form x → Cx. If x is transformed, we then obtain: f(x) = kCαxα = Cαf(x) (4.30) provided that given initial power law function is f(x) = kxα. Changing the scale of the independent variable thus preserves the functional form of the solution but with a change in its scale. This is an important property in our case. Scale-free behaviour strongly suggests that the same mechanism is at work across different sectors of the economy, the industrial structure of which remains constant over a relatively long period of time, measured with any time measurement (i.e., seconds, minutes, hours, days, years). A useful example that should be appealing for economists is price. We say that price is a homogenous function of degree zero with respect to income. b) A power law is just a linear relationship between logarithms (Breisinger et al., 2009) of the form: log f(x) = –α log (x) + log k. (4.31) c) Power law also has excellent aggregation properties44. The property of being distributed according to a power law is conserved under addition, multiplication, polynomial transformation, min, and max. The general rule is that when combining two power law variables, the fattest power law (i.e., the one with the smallest exponent) dominates. This property could be helpful for empiricist researchers using this form of function. Let X1,...,Xn be independent random variables, and k, a positive constant. Let αx be also the power law exponent of variable X. Following Gabaix (2008), Jessen and Mikosch (2006), we have the so-called inheritance mechanism for power law: ),....,min( 2121 ..... nn XXXXXX (4.32) ),....,min( 2121 *.....* nn XXXXXX (4.33) ),....,min() 2121 .....,max( nn XXXXXX (4.34) 44 The interested reader is recommended to see the works of Jessen & Mikosch (Jessen & Mikosch, 2006) or Gabaix (Gabaix , 2008). As an example of relative facilities of proofs, if x kxxXP )( kk k x x kxxXPkxxXP )()( 1 k X X k , then x kxxXP )( kk k x x kxxXPkxxXP )()( 1 k X X k , so x kxxXP )( kk k x x kxxXPkxxXP )()( 1 k X X k .
Bibliography – Part IV 131 Wing Ian Sue, (2004). Computable General Equilibrium Models and Their Use in Economy-Wide Policy Analysis, Massachusetts Institute of Technology: Technical Note No. 6, MIT Joint Program on the Science and Policy of Global Change. Zellner, A. (1991). Bayesian Methods and Entropy in Economics and Econometrics. In The Netherlands. (Eds. W.T. Grandy and L.H. Schick), pp. 17–31. Kluwer Academic Publishers, The Netherlands.
PART V: Computable General Equilibrium Models
1 A Historical Perspective Since the 19th century, there has been a methodological discussion in economics about how one should analyse national economies. Walras (1834) introduced a framework he called General Equilibrium Analysis. According to Walras, in a national economy “everything affects everything.” The only correct way to analyse it should be to treat the national economy as an unbroken entity and to use the tools of general equilibrium analysis. Later, many other economists have strengthened the Walras orientation. As an example, one can cite the Edgeworth diagram, which enabled us to explain the Pareto optimum. However, the work of Arrow and Debreu (1954) has constituted the cornerstone of Walrasian equilibrium theory. In fact, by proving the existence of the uniqueness of the optimum of Walrasian equilibrium—plus the two theorems of welfare—Arrow and Debreu combined abstract general equilibrium structure with realistic economic data to solve numerically for the levels of supply, demand, and price that support equilibrium across a specified set of markets. This allowed Walrasian equilibrium to become an applicable theory. It is worthwhile to recall here the contribution of Nash (1950), who introduced anticipation aspects into multi-game equilibrium, thereby achieving something like a quasi-Pareto optimum. This school of thought is the forefather of the Leontief input-output model of production, the social accounting matrix (SAM), and microeconomic-based computable general equilibrium models (CGE). On the other hand, Marshall Marshall (1890) criticized Walras and postulated that general equilibrium analysis is impossible in practice, because it demands too much information. Marshall's claim was that it is enough to separate from the rest of national economy the part under investigation and to analyse it within the framework he called partial equilibrium analysis. As a motivation, Marshall developed the so called ceteris paribus condition, which means “all other things remaining unchanged.” Most post-Keynesian macroeconomic models belong to this school of thought. Criticism against large-scale macro econometric models built in the tradition of the Cowles Commission approach began in the late 1960s. These misgivings were subsequently reflected in the Lucas critique (parameters of models may take into account the reaction effect of agents with respect to expectation—rational or not), Sims’s critique (time series models), and disenchantment with the model’s Keynesian foundations (IS-LM models and the Philips curve) criticised by the Chicago school. In response, classical macro econometric modelling progressed in two parallel ways: one, the improvement of the structure of traditional models, particularly in terms of specifying the supply-side and forward-looking expectations; and the other, strengthening techniques or developing alternative techniques (the so called no economics theory-oriented models), e.g., the LSE approach aided by the advent of co-integration analysis, vector autoregressive (VAR) systems, and dynamic stochastic general equilibrium (DSGE) models.
135 A Historical Perspective 135 Walrasian general equilibrium theory made its resurgence while the Keynesian model started declining. A major stimulus to early CGE modelling was Stone and Brown (1962). As a continuation of Leontiew’s work (1941) — and to a certain degree, F. Quesnay’s tableau économique (18th century) — Stone pioneered the development of the SAM framework with his 1955 article Input-Output and Social Accounts (1962). The general shape of a SAM framework was next described by Pyattand Thorbecke (1976). Then, Pyatt and Roe (1977) published a book giving a detailed description of the example of Sri Lanka. Since then, SAMs have been applied in a wide variety of (developed and developing) countries and regions, and with a wide variety of goals, in particular, as we will see latter, for impact analysis and simulations. While in the early 1960s CGE models were perceived as precious devises for modelling poorer economies (e.g., Adelman et al. (1978), Arrow et al. (1971), de Melo (1988)), CGE modelling of developed economies stems from Leif Johansen's 1960 sectorial growth model (MSG) of Norway as an extension of the Leontiefmodel. The model was later extended by Harberger (1959, 1962). Showen, Scarf, and Walley (1984, 1972, 1992) with the presentation by Scarf (1969) of an algorithm helping to solve the model. Similarly, as far as CGE models for developed countries are concerned, since the early of 1960's, a model was developed by the Cambridge Growth Project under the initiative of Richard Stone in the UK. The Australian MONASH model is the next new generation representative of this class. Both models were dynamic (traced variables through time). Other more recent contributions may draw attention, in particular those of Jorgenson, using an econometric approach, Mc Kenzie (1959, 1981, 1987), Ginsburgh and Waelbroeck (1981, 1976), Ginsburgh and Keyzer (1997), Harrisand Cox (1983), Bourguignon (1983), Decaluwe and Martens (1987, 1988). Today there are many other CGE models from different countries. One of the most well-known CGE models is the GTAP (Global Trade Analysis Project) model of world trade, which involves many researchers around the world. Depending on, among other things, targeted time-scope analysis, the macroeconomics school of thought involved, or the approach to model estimation, nowadays there are large classes of CGE models. Readers interested in the epistemological aspects of CGE models can see—e.g., Xian (1984) Jorgenson (1984, 1998a) , Ginsburgh and Keyzer (1997), McKenzie (1954) or Mansur and Whalley (1984). However, for the clarity of the document, let us now concentrate on two classes of CGE models. The first class models the reactions of the economy over a given perspective of time thus suggesting comparative static and dynamic CGE models. The second focuses on the theoretical aspects of equilibrium, seeing that economic conditions of general equilibrium are not always fulfilled. As far as the first class of models is concerned, many CGE models around the world are static; that is, they model the reactions of the economy at only one point in time. For policy analysis, a simulation analysis is carried out and outputs are often interpreted as showing the reaction of the economy, in some future period, to one or more external shocks or policy changes. From the analytical point of view, the results
136 A Historical Perspective show the difference (usually reported in percent of change) between two conditional alternative future states, that is, “what would happen if the policy shock were implemented.” As opposed to dynamic models, the process of adjustment to the new equilibrium is not explicitly represented in such a model. However, details of the closure rule lead modellers to distinguish between short-run and long-run equilibriums. For example, this will be the case if the hypothesis on whether capital stocks are allowed to adjust or not. Dynamic CGE models, by contrast, explicitly trace each variable through time at regular time steps, generally at annual intervals. While this class of model removes one of the main criticisms of CGE models, that of being unrealistic, as their analysis is based on one-year observations, at the same time, they become more challenging to construct and solve—they require, for instance, that future changes are predicted for all exogenous variables, not just those affected by a possible policy change. Furthermore, dynamic elements may arise from partial adjustment processes or from stock/ flow accumulation relations—between capital stocks and investment and between foreign debt and trade deficits. Recursive-dynamic CGE models are those that can be solved sequentially, over time. They assume that behaviour depends only on current and past states of the economy. The construction of this class of models is less complex and such models are easier to implement in empirical research than dynamic models. Alternatively, if agents' expectations depend on the future state of the economy, it becomes necessary to solve for all periods simultaneously, leading to full multiperiod dynamic CGE models. Recent publications cover this group models known as dynamic stochastic general equilibrium (DSGE) as they explicitly incorporate uncertainty about the future. It is worthwhile to add that the earliest DSGE models were formulated in an attempt to provide an internally consistent framework to investigate real business cycle (RBC) theory48. If we consider the second class of models focusing upon the general equilibrium aspects, one may consider that most CGE models rarely conform to the theoretical general equilibrium model. For instance, the presence of imperfect competition, nonclearing markets, or externalities (e.g., pollution) will lead the economy to disequilibrium conditions. 48 See DSGE: Modern Macroeconomics and Regional Economic Modeling by Dan S. Rickman, Oklahoma State University, prepared for presentation in the JRS 50th Anniversary Symposium at the Federal Reserve Bank of New York.
2 The CGE Model Among Other Models We start this section with the next analytical question: How will a tax increase on petrol/gasoline impact the Polish economy? A tax of this kind would probably affect petrol/gasoline prices and might affect transport costs, car cost, the CPI, and hence, wages and employment. Traditional econometric models would have difficulty answering this question seeing the complexity of response channels tax shock would imply on the whole economic system. CGE models are useful whenever we wish to estimate the effect of changes in one part of the economy upon the rest of the economy. They are being used widely to analyse trade policy. More recently, the CGE has been a popular devise to estimate the economic effects of measures to reduce greenhouse gas emissions. Many research centres and central banks (including the Polish Central Bank) build stochastic dynamic CGE models which encompass the financial sphere of the economy. A CGE model consists of behavioural equations describing model variables consistent with a relatively detailed database of economic information. The standard CGE tends to be neoclassical in spirit by assuming cost-minimizing behaviour by producers, average-cost pricing, and household demands based on optimizing behaviour. Thus, CGE models are built based on the theoretical model of competitive general equilibrium. Its original structure was developed during the second half of the 19th century by neoclassical economists. Among them, four should be mentioned here, the German Gossen (1983), the British Stanley Jevons (1879), the Austrian Menger (1871), and the French Walras. Due to the dominance of his contributions to its conceptualization, the model bears the name of the last author, the Walrasian general system.
3 Optimal Behaviour And The General Equilibrium Model 3.1 Introduction The fact that the model is “computable” means that a numerical solution exists (e.g., Arrow-Debreu, 1954; Mc Kenzie, 1959; Ginsburgh and Keyzer, 1997), and “general equilibrium” refers to simultaneously matching demand and supply on all markets. In the example below, note the difference between a partial and a general equilibrium in the traditional way of analysing a market handed down by the Marshall and Walras schools. Let us suppose a Cobb Douglas two-sector economy with two commodities Xi two sector inputs L1,K1 (labour and capital sectors) and two sector income Yi, with i = 1,2. Then, the partial equilibrium model is defined by the next optimal program: Objective: 2 1 1 21 XXUMax xx Market clearance: Yi = Xi i = 1,2 Production: 1 ii i ii i Y AL K Resource constraints: 12 12 LL L KK K In the case of a general equilibrium, we need to add an income balance restriction to ensure that all inflows and outflows are balanced. Income balance: KrLwXPXP 2211 with pi (i = 1,2), w, r representing the prices of the two commodities, the sectors labour and capital respectively.
Introduction 139 Let us now generalize the above formulation and consider a simple economy with m finite number of producers, n finite number of consumers, r commodities, and let us suppose that the Walras hypotheses are fulfilled. Thus, under these conditions, let us present, below, the behavioural functions of economic representative agents and conditions of market equilibrium. Producer behaviour. Each producer, j (j = 1..m), is confronted with a set of possibilities of production vj, the general element vj of which is a program of production with dimension r, where outputs have a positive sign and inputs a negative sign. The objective of each producer is to select, for a given price p (p = 1..r), an optimal program of profits pvj. Consumer behaviour. Each consumer i (i = 1..n) is supposed to have an initial endowment of goods wi (outputs or inputs) that the consumer is ready to exchange against remuneration by the producer. Thus, the consumer is confronted with Xi possibilities of consumption of which the general element is xi with dimension r. The consumer is never saturated in consuming Xi and his endowment wi allows him to survive. For a given price p of dimension r, consumer i has the objective of maximizing total utility Ui(xi) under his given budgetary constraints: j ijiji pxpvpw with ii Xx with j ijiji pxpvpw with ii Xx where j ijiji pxpvpw with ii Xx (i = 1,2 …, n; j = 1,2, ..m) is a fraction of profits realized by the producer j and transferred to consumer i. Producer function. m producers maximize individual total profits: Max jj vppv ~~ (5.1) subject to: vj∈Vj Consumer function. n consumers maximize individual total utility: Max ) ~ ()( iiii xUxU (5.2) subject to: jijii vpwpxp ~~~~ (5.3) xi∈Xi This is definitely a general equilibrium solution ( ij xvp ~ , ~ , ~ ) from a decentralized system. Market clearance. Excess demand for r goods is not positive: 0 ~~ i i j j i iwvx (5.4)
140 Optimal Behaviour And The General Equilibrium Model Commodities with supply excess, i.e., free commodities, have price zero while other commodities have a positive price: i i j j i iwvxp 0) ~~ ( ~ (5.5) where jj vppv ~~ A general equilibrium solution. This is definitely a general equilibrium solution ( ij xvp ~ , ~ , ~ ) guarantying that each of the markets will have realizable equilibrium. This, too, is an equilibrium for a decentralized economy since it guarantees compatibility of consumer and producer behaviours (Equations 5.1 and 5.2). This is a competitive equilibrium. The price from Equation (5.3) is imposed on all actors of the market. 3.2 Economic Efficiency Prerequisites for a Pareto Optimum The purpose of this section is to clarify the connection between the general equilibrium model and the optimum Pareto state. This will allow us in the next chapter to go beyond such an equilibrium and to analyse impact on social welfare. We must then check whether or not the three conditions below are fulfilled. a) Equality of marginal rates of technical substitution for different producers. Let us limit our generalization to an economy with two goods q1 and q2, and two limited inputs x1 and x1, for two respective producers. q1 = f1(x11, x12), for producer 1, q2 = f2(x21, x22), for producer 2, This means that x1 = x11 + x21 and x1 = x12 + x22 Let us maximize the quantity produced of q1 under restriction of known quantityq2. Using the Lagrange multiplier, we have: L = f1(x11, x12) + λ[f2(x1 – x11, x2 – x12) – q2] Finally, we get: 21 22 2 21 2 12 1 11 1 TmSTTmST x f x f x f x f