scieee AI-readable full text Open interactive document viewer

Micro-Modelling Approaches for Credit Rating and Corporate Survival

Novotná, Martina

Abstract

The essential issue of this book is the term credit, either in the context of credit markets, credit risk or credit rating. Credit markets’ existence is associated with credit risk, which refers to the risk of an economic loss from the failure of a counterparty to meet its contractual obligations. Due to credit risk, suppliers of credit need to assess the creditworthiness of prospective borrowers. Although modern approaches to credit risk analysis have been developed in recent decades, examining borrowers’ ability to repay their funds is one of the oldest lending activities. The main goal of this monograph is to apply and verify certain methods of credit risk modelling to real data from selected CEE countries. For the main purpose of this work, a micro approach is used to measure credit risk based on monitoring basic indicators and allowing creditors to take the necessary actions in time. The book aims at two partial financial and methodological objectives related to credit risk modelling in this context. Both of them are interconnected, and they complement each other throughout the book. In terms of financial application, this book’s principal objective is to analyse credit risk based on real data, assess its main factors, explore mutual relations, and draw conclusions related to risk assessment and market behaviour. The application is focused on two approaches of individual credit risk assessment within the areas of credit rating and corporate survival. The methodological purpose is the application and verification of rating and bankruptcy models. Rating models estimated using conventional approaches such as discriminant analysis or logistic regression are supplemented by an alternative survival analysis approach to determine the probability of rating downgrade over time. Survival analysis is subsequently used in the following empirical studies on corporate bankruptcy. We will investigate the relationship between the rating and corporate bankruptcy rates and estimate the rating assessment depending on the used model, input variables and the company's age. This book is intended for everyone interested in credit risk, particularly rating and corporate survival modelling, mainly for academia and students at all levels of study. This monograph aims to provide complex information on credit risk fundamentals, current trends and rating systems’ principles. However, the primary purpose is the practical application and estimating models using real corporate data. Thus, we can determine the main factors of rating assessment and corporate survival and demonstrate how these models can be developed through different statistical methods. The text is structured into three central parts: The theoretical background on credit risk and the credit rating industry, a description of econometric approaches used in the applications, and empirical studies on credit rating and corporate bankruptcy modelling. If the reader is particularly interested in estimating models and their comparison and interpretation, then it is suggested that they go directly to the practical application. However, reading the book step by step is recommended to understand the essence and main principles and use them in the application.

Full text

Series on Advanced Economic Issues Faculty of Economics, VŠB-TUO www.ekf.vsb.cz/saei [email protected] ZDE ZAČÍNÁME POČÍTAT ČÍSLA STRÁNEK, ŘÍMSKÝM AŽ PO KONEC OBSAHU PRÁZDNÁ STRÁNKA Series on Advanced Economic Issues Faculty of Economics, VŠB-TUO Martina Novotná MICRO-MODELLING APPROACHES FOR CREDIT RATING AND CORPORATE SURVIVAL Ostrava, 2024 Martina Novotná Department of Finance Faculty of Economics VŠB-Technical University Ostrava 17. Listopadu2172/15 708 00 Ostrava-Poruba, CZ [email protected] Reviews Jiří Witzany, Prague University of Economics and Business Lumír Kulhánek, VSB – Technical University of Ostrava INFORMA This publication is the output of research activity by the research team of the project No. SP 2019/132. The text should be cited as follows: Novotná, M. (2024). Micro-Modelling Approaches for Credit Rating and Corporate Survival, SAEI, vol. 69. Ostrava: VSB-TUO. © VŠB-TUO 2024 Printed in VSB-TUO Cover design by EkF VSB-TUO This work is licensed under Creative Commons Uveďte původ 4.0 Mezinárodní. ISBN 978-80-248-4728-3 (print) ISBN 978-80-248-4729-0 (on-line) DOI 10.31490/9788024847290 Preface The essential issue of this book is the term credit, either in the context of credit markets, credit risk or credit rating. Credit markets’ existence is associated with credit risk, which refers to the risk of an economic loss from the failure of a counterparty to meet its contractual obligations. Due to credit risk, suppliers of credit need to assess the creditworthiness of prospective borrowers. Although modern approaches to credit risk analysis have been developed in recent decades, examining borrowers’ ability to repay their funds is one of the oldest lending activities. The main goal of this monograph is to apply and verify certain methods of credit risk modelling to real data from selected CEE countries. For the main purpose of this work, a micro approach is used to measure credit risk based on monitoring basic indicators and allowing creditors to take the necessary actions in time. The book aims at two partial financial and methodological objectives related to credit risk modelling in this context. Both of them are interconnected, and they complement each other throughout the book. In terms of financial application, this book’s principal objective is to analyse credit risk based on real data, assess its main factors, explore mutual relations, and draw conclusions related to risk assessment and market behaviour. The application is focused on two approaches of individual credit risk assessment within the areas of credit rating and corporate survival. The methodological purpose is the application and verification of rating and bankruptcy models. Rating models estimated using conventional approaches such as discriminant analysis or logistic regression are supplemented by an alternative survival analysis approach to determine the probability of rating downgrade over time. Survival analysis is subsequently used in the following empirical studies on corporate bankruptcy. We will investigate the relationship between the rating and corporate bankruptcy rates and estimate the rating assessment depending on the used model, input variables and the company's age. This book is intended for everyone interested in credit risk, particularly rating and corporate survival modelling, mainly for academia and students at all levels of study. This monograph aims to provide complex information on credit risk fundamentals, current trends and rating systems’ principles. However, the primary purpose is the practical application and estimating models using real corporate data. Thus, we can determine the main factors of rating assessment and corporate VI Preface survival and demonstrate how these models can be developed through different statistical methods. The text is structured into three central parts: The theoretical background on credit risk and the credit rating industry, a description of econometric approaches used in the applications, and empirical studies on credit rating and corporate bankruptcy modelling. If the reader is particularly interested in estimating models and their comparison and interpretation, then it is suggested that they go directly to the practical application. However, reading the book step by step is recommended to understand the essence and main principles and use them in the application. Martina Novotná, Ostrava, April 2024 Brief Contents Preface ..................................................................................................... V Brief Contents ...................................................................................... VII Contents .................................................................................................. IX List of Abbreviations ............................................................................. XI Chapter 1 Introduction ........................................................................... 1 Chapter 2 The Essentials of Credit Rating Assessment ....................... 5 2.1 Role of Credit Markets ................................................................. 6 2.2 Classification of Credit Risk ......................................................... 7 2.3 Factors of Credit Risk ................................................................. 11 2.4 Credit Risk Analysis.................................................................... 14 2.5 Obligor-Level Credit Risk .......................................................... 18 2.6 Issue-Specific Credit Risk ........................................................... 24 2.7 Fundamentals of Rating Assessment ......................................... 32 2.8 Description of Credit Rating Process ........................................ 37 2.9 Current Issues in Rating Industry ............................................. 41 2.10 Chapter Summary ....................................................................... 46 Chapter 3 Approaches for Credit Rating and Corporate Bankruptcy Modelling ..................................................................................... 47 3.1 Introduction and Research Background ................................... 47 3.2 Discriminant Analysis ................................................................. 54 3.3 Logistic Regression Analysis ...................................................... 62 3.4 Survival Analysis ......................................................................... 74 3.5 Chapter Summary ....................................................................... 90 VIII Brief Contents Chapter 4 The Effect of Selected Factors on Rating and its Dynamics ....................................................................................................... 93 4.1 Corporate Credit Rating Assessment Models ........................... 94 4.2 Modelling of Rating Downgrades Based on Multiple FailureTime Data ................................................................................... 109 4.3 Chapter Summary ..................................................................... 118 Chapter 5 Relationship Between Rating and Corporate Bankruptcy Rates ........................................................................................... 121 5.1 Association Between Rating and Corporate Defaults ............ 122 5.2 Modelling of Corporate Survival Based on Kaplan-Meier Estimates .................................................................................... 126 5.3 The Relationship Between Bankruptcy Rates and Rating Assessment ................................................................................. 133 5.4 Chapter Summary ..................................................................... 139 Chapter 6 Survival Models with Categorical Variables .................. 141 6.1 Application of the Cox Proportional Hazards Model ............ 141 6.2 Parametric Models .................................................................... 147 6.3 Chapter Summary ..................................................................... 152 Chapter 7 The Use of Financial Performance Indicators in Survival Analysis ...................................................................................... 155 7.1 Estimation of Survival Models ................................................. 155 7.2 Cumulative Bankruptcy Rates and Ratings for Specific Parameters ................................................................................. 167 7.3 Chapter Summary ..................................................................... 173 Chapter 8 Conclusion .......................................................................... 177 Appendix .............................................................................................. 181 List of Figures ...................................................................................... 205 List of Tables ........................................................................................ 207 References ............................................................................................ 211 Index ..................................................................................................... 223 Introduction 3 Micro-Modelling Approaches for Credit Rating and Corporate Survival is essential to understand how rating agencies provide rating assessments. Thus, the main principles of rating systems and the credit rating process are described in the second chapter. Various regulations stimulate the formal quantification of credit risk and the use of credit portfolio models in financial institutions. Thus, many issues must be considered when selecting the appropriate approach for credit risk modelling. For example, suppose the credit risk analyst evaluates credit risk as a discrete event and concentrates merely on a potential default event. In that case, the fundamentalbased models provide a suitable way of assessing credit risk. On the other hand, structural and different quantitative approaches should be applied if the modeller analyses the dynamics of the debt value and the associated credit spread over the whole time interval to maturity. Throughout this book, and especially the application part, credit risk is considered a discrete event, such as a potential default or bankruptcy event represented by rating grade or the probability of survival. In such cases, the main task is to assess the credit risk of a particular issue or issuer, typically through credit risk models developed to discriminate between lower and higher credit risk. Such models are usually based on the statistical analysis of past characteristics of debtors or issuers, mainly quantitative variables such as corporate financial ratios. Since these models are focused on evaluating individual subjects, mostly based on data from financial statements, they are also referred to as micro models or fundamental-based models. In the application part, attention is paid to the procedure and development of such models and their use and interpretation. The econometric approaches selected based on the recent studies and used in the application part are described in Chapter 3. Firstly, we provide some research review, a summary of approaches used in the chronological context, and the main findings of selected studies in this section. Following, we will find the possible extension of the current research and emphasize the contribution of this work in the context of the application part, which is focused on CEE countries and the use of survival analysis. Next, the selected methods used in the application part are described. First, discriminant and logistic regression analyses are described and used in Chapter 4 to estimate rating models. Then, the principles of survival analysis are explained, and selected survival models are described, focusing on the Kaplan-Meier estimates, the Cox proportional hazards and the Weibull model. These approaches are then used in Chapters 5, 6 and 7. The first application study in Chapter 4 is focused on credit rating modelling. Firstly, selected methods are used to estimate rating models based on data from CEE countries. Then, based on the main findings, we determine the main accounting-based variables of the credit rating of non-financial companies from selected CEE countries representing economies with a shorter history and tradition of capital markets. The analysis is based on a corporate rating evaluation known as MORE Rating. This study's methodological objective is to compare models developed by different approaches, such as discriminant and logistic regression analysis and suggest a more suitable method for rating modelling. Finally, this part 4 Chapter 1 2024 Martina Novotná is followed by the modelling of rating downgrade employing survival analysis methods. Therefore, we can compare the main findings and identify the key influential variables on rating assessment and the hazard of rating deterioration based on two different approaches. The aim of Chapter 5 is to assess the relationship between the rating and corporate bankruptcy rates. This study is based exclusively on data from Czech companies and thus complements the main findings from CEE countries. In this section, we compare published default rates with estimated bankruptcy rates. Based on the comparison, we propose the procedure for rating estimation using bankruptcy rates and average spreads. Next, the Cox proportional hazard and the Weibull model are used in Chapter 6 to assess the impact of industry, legal form and company size on the survival probability, followed by evaluating the influence of financial variables on the survival probability in Chapter 7. Both studies identify the effect of selected variables on corporate survival, whether categorical or quantitative. In addition, cumulative bankruptcy rates are estimated using the models and converted to a rating assessment. A significant advantage of this approach is using the time variable in survival models, which allows us to determine the rating not only depending on financial performance or other characteristics but also on the company's age. Thus, this procedure represents a dynamic approach to modelling and predicting individual ratings. Finally, the main findings and overall suggestions are summarized in the conclusion in Chapter 8. This book used two statistical software packages for data analysis: IBM SPSS Software and Stata statistics. All models are estimated based on unique corporate data from the MORE Rating and corporate data on Czech companies from the Magnusweb database. Chapter 2 The Essentials of Credit Rating Assessment The crucial issue in this monograph is the term credit, which refers to credit markets, credit risk, or credit rating. This chapter provides an introduction to credit and credit markets critical role; however, the primary attention will be paid to the explanation of credit risk, its measurement and analysis. This monograph is focused on two areas of application: credit rating and corporate survival modelling. Both topics are associated with credit risk; however, the assessment uses different methods and data and is conducted from different perspectives. Nevertheless, both approaches are used to obtain more information about corporate credit risk and are not mutually exclusive. On the contrary, it is appropriate to use both ways to analyse the credit risk and its dynamic in certain cases. This section aims to provide an introduction and a literature review of credit rating and corporate survival modelling. First, the emphasis will be on the purpose and goals of modelling, some recent research, and a summary of the approaches used. Then, we gradually focus on the theoretical background and the current state of credit rating modelling. Finally, attention will be paid to the overview of the corporate survival problem. The structure of the monograph’s remaining part corresponds with the particular objectives of this work, as they are mentioned in the introduction. First, a literature review will provide the principles and purpose of rating and survival models (Chapter 2). Then, the statistical methods used in the application will be described (Chapter 3). The main part is the application, in which the estimated models, procedures and main results will be presented (Chapters 4–7). Finally, the main findings of the rating and bankruptcy analysis will be compared and used to draw this monograph's main conclusions and recommendations. 6 Chapter 2 2024 Martina Novotná 2.1 Role of Credit Markets Credit markets are markets for credit that can be described as transactions between the creditor (the lender) and the debtor (the borrower). The creditor supplies money or non-monetary assets such as goods, services or securities to the debtor in return for a promise of future payment, which typically includes the amount of interest (Joseph, 2013). The creditors generally have no right of ownership, and the interest represents compensation for undertaken risk. The economic role of credit lies in the fact that borrowers with insufficient resources can get funds from lenders, usually through financial intermediaries. Thus, when used effectively, credit enables the economic growth of borrowers, increasing household consumption and business investment. Credit is used by businesses, individuals, and governments, and typically, it leads to economic growth (Joseph, 2013). In financial markets, credit is supplied by financial institutions such as commercial banks in loans. It can be provided by investors who purchase bonds issued by deficit units. Debt holders are known as creditors or lenders, and they typically grant loans or hold bonds. Conversely, the use of credit by deficit units or borrowers can be considered a primary source of debt financing. While bonds are typically traded, loans are not assumed to be tradable in debt markets (De Servigny and Renault, 2004). The proportion of loans and bonds on the total amount of debt financing can differ in various countries, depending on tradition, legal environment (e.g., property rights system), macroeconomic conditions or the development of capital markets. The proportion of these two ways of financing in the Czech Republic can be seen in Figure 2-1. In this graph, the total loans include short-term (less than one year), medium-term (1 – 5 years) and long-term loans (more than five years) to clients provided in CZK, and total bonds consist of all bonds issued in CZK, including government and corporate bonds. As can be seen, the amount of loans exceeds the number of bonds during the whole period. From 2012 to 2019, we can see relatively stable development, with the share of loans being around 56% on total debt financing and the proportion of bonds moving about 44%. However, since 2019, we have seen that the percentage share of bonds has risen, mainly due to increased government bond issuance during the COVID-19 pandemic. While the ratio of both types of financing remains relatively stable, except in the post-pandemic period, annual changes are somewhat volatile. Bond issuance rose by 21.2% from 2011 to 2012, mainly due to an increase in corporate bonds by 40.4%. The decrease in the bonds issued in 2017 is primarily due to the decline in government bond issuance, with the opposite trend since 2019 (see Figure 2-2). The Essentials of Credit Rating Assessment 7 Micro-Modelling Approaches for Credit Rating and Corporate Survival Figure 2–1 Proportion of loans and bonds in the Czech Republic (end of the year) Source: Czech National Bank (ARAD, 28. 9. 2022), author In most advanced economies, banking intermediation has reduced over the past years due to the increasing breadth of credit markets, and the trend toward market-based finance seems to be very strong (De Servigny and Renault, 2004). However, banks as intermediaries still fulfil a crucial role in three primary functions: liquidity, risk and information intermediation. Figure 2–2 Annual change of loans and bond issues in the Czech Republic (in %) Source: Czech National Bank (28. 9. 2022), Czech Statistical Office (28. 9. 2022), author 2.2 Classification of Credit Risk Risk can generally be defined as the volatility of returns leading to unexpected losses, as defined by Crouhy et al. (2014). Several risk factors can influence this volatility of returns, including: • Market risk that changes in market prices and rates will negatively affect a security or portfolio value. • Credit risk of an economic loss from a counterparty's failure to fulfil their contractual obligations. 0% 20% 40% 60% 80% 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 Bonds Loans -5% 0% 5% 10% 15% 20% 25% 30% 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 Bond Issues Loans to clients GDP 8 Chapter 2 2024 Martina Novotná • Liquidity risk includes the risk that a firm cannot raise the necessary cash (funding risk) or a transaction will not be executed (trading risk). • Operational risk refers to potential losses from operational failures (management, controls, fraud, human factors). It is closely related to legal and regulatory risk or reputation risk. • Business risk refers to uncertainty about the demand for products, prices, and production costs. • Strategic risk is the risk of significant investments with high uncertainty about success and profitability. The existence of credit and credit markets is associated with credit risk. In the text, credit risk refers to the risk of economic loss from a counterparty's failure to meet its contractual obligations, such as interest payments or principal repayment. However, credit risk involves the possibility of non-payment on a future commitment and during a transaction. This type of risk is called settlement risk. It arises from exchanging principals in different currencies or payments in different time zones during a short window, typically a day. Traditionally, credit risk is considered a pre-settlement risk, which arises during the obligation’s life (Jorion, 2011). Overall, credit risk can be decomposed into the following four categories: • Default risk refers to the debtor’s capacity or refusal to meet debt obligations such as interest or principal payments by more than a reasonable relief period from the due date (usually 60 days in the banking industry). • Bankruptcy risk can be considered the risk of taking over a defaulting borrower’s assets or counterparty. In this case, debt holders are taking over the control of the company from the shareholders. • Downgrade risk is the risk that the creditworthiness of the borrower or counterparty might deteriorate in the future when a significant deterioration can be seen as the default. • Settlement risk refers to the risk due to the exchange of cash flows when a transaction is settled. It can be caused by counterparty default, liquidity constraints, or operational issues (Crouhy et al., 2014). Due to all types of credit risk, suppliers must assess their creditworthiness before granting credit to prospective borrowers (Joseph, 2013). In addition, since traditional banks typically hold the loan until maturity, they analyze the riskiness of the borrowers’ activities both before and after the loan is made because they face the risk that the borrower’s credit quality could deteriorate during the life of the loan. Some lenders, such as banks, use financial innovations in various strategies to reduce credit risk and increase returns. These innovations primarily involve securitization, syndication of loans, proprietary trading and investment in nontraditional assets, or increased use of financial derivatives (Saunders and Allen, 2010; Stowell, 2010): The Essentials of Credit Rating Assessment 9 Micro-Modelling Approaches for Credit Rating and Corporate Survival • Securitization represents an innovative way for lenders to raise funds in the capital market by selling their assets’ future receivable cash flows, such as mortgage, student or credit card loans. The loans and other assets are packaged and sold as asset-backed securities. This process transfers credit risk to investors, while banks can free up capital for other lending and investment activities. • Loan syndication is another way how banks can reduce risk exposures. Firstly, a bank originates a loan and then sells parts of the loan to outside investors. The outside investors include other banks, hedge funds, mutual funds, insurance companies and other investors. Banks’ proprietary investment activities involve non-client-related investments in securities or other assets for their accounts; for example, banks establish hedge funds, private equity, or venture capital funds. These subsidiaries are then involved in investment activities that are considered too risky for banks. • The use of financial derivatives covers the use of credit default swaps designed to transfer the credit risk on a portfolio of banks to nonbanks, typically insurance and reinsurance companies. As we can see in the previous text, lenders such as banks can reduce credit risk in different ways. These innovative activities also slightly change the traditional view as an institution that issues short-term deposits and offers long-term loans. Even though these innovative strategies have been increasing in recent years, banks still face a substantial credit risk resulting from their traditional activities. For this reason, they pay considerable attention to credit risk measurement and management. Not only do banks face credit risk from their operations, but also persons placing deposits with banks or investors purchasing corporate bonds. We already defined credit risk as the probability of loss due to the failure or counterparty's unwillingness to meet contractual obligations. According to Joseph (2013), credit risk generally exists whenever a product or service is obtained without paying for it. A single borrower (obligor) exposure is known as firm-credit risk, while credit exposure to a group of borrowers is called a portfolio-credit risk. Credit risk is the product of various events and factors, such as domestic, international or company-specific issues. As we can see from the scheme in Figure 2-3, some of the causes are more controllable than others. 10 Chapter 2 2024 Martina Novotná Figure 2–3 Major sources of credit risk Source: Joseph (2013), p. 16 Uncontrollable risks are called systematic risks, and they are associated with external forces that affect all businesses and households in the country. For instance, the frequency of defaults or bankruptcies typically increases during the economic recession, causing credit losses for the lenders (Joseph, 2013). Figure 24 shows the number of corporate defaults of companies rated by Standard & Poor’s from 1981 to 2015. The bankruptcy number increased during each of three periods of economic downturn: The recession of the early 1990s that came after the Black Monday of October 1987, the first 2000s recession, and the great recession of 2008. Figure 2–4 Total number of corporate defaults (1981–2015) Source: S&P Global Ratings (2015), author Unsystematic risk can be considered controllable because these risks do not affect the entire economy or all businesses or households. On the other hand, these risks are mainly industry or company-specific. Lenders might reduce unsystematic risk through diversification or extending credit to various customers. Credit risk Systematic risk Socio-Political Risks Economic Risks Other Exogenous Risks Unsystematic risk Business Risks Financial Risks 0 50 100 150 200 250 300 Total defaults Year The Essentials of Credit Rating Assessment 11 Micro-Modelling Approaches for Credit Rating and Corporate Survival 2.3 Factors of Credit Risk The principal problem in credit risk measurement is to quantify the risk of losses due to counterparty default. As Jorion (2011) suggests, the distribution of credit risk can be considered a compound process driven by the following three variables: • Default, • loss given default, • credit exposure. Default is the principal issue in credit risk measurement, so it is essential to pay some attention to its explanation. The definition of default of an obligor typically includes the following characteristics: • Days past due criterion for default identification, • indications of unlikeness to pay, • conditions for a return to non-defaulted status. Due to the absence of specific rules and other aspects of the application, various approaches have been adopted across institutions and jurisdictions. Based on the European Banking Authority (EBA, 2016), institutions use differing practices regarding default. As stated in the report 1 , specific rules adopted in most jurisdictions usually focus on counting days past due and applying the material threshold. On the other hand, particular rules on different aspects of the definition of default are much less common. To harmonize a consistent use of default meaning, the EBA suggests guidelines to increase comparability of risk estimates and own funds requirements, especially when using internal rating-based or IRB models. For example, in the Czech Republic, the default subject is regulated by the Act on Bankruptcy and Settlement, known as the Insolvency Act 2 . This Act aims to control the resolution of the debtor’s insolvency and imminent bankruptcy and the debtor’s discharge of debts. According to this Act, a debtor is insolvent if they have several creditors, outstanding financial liabilities overdue for more than 30 days, and cannot fulfil such liabilities. While insolvency is a specific legal term meaning that a debtor cannot pay their debts, default generally means that a debtor has not yet paid a debt as required. We can consider default as a discrete state for the counterparty with some probability of default (PD). The determination of the likelihood of default is the crucial issue in the credit risk management approach and can be achieved through various methods (De Laurentis, 2010): • The observation of historical default frequencies and allocation to different credit classes (ex-post), • the use of mathematical and statistical tools to expect the probability (expost), 1 The guidelines will apply from 1 January 2021 2 Act No. 182/2006 on Bankruptcy and Settlement 12 Chapter 2 2024 Martina Novotná • the combination of judgmental and mechanical approaches or the approach based on market prices. The probability of default can be considered the default risk measure within a specified time horizon, usually one year. Alternatively, when exposures are more than one year, the assessment is typically based on cumulative probabilities (De Laurentis et al., 2010). The typical technique to assess the creditworthiness of retail and commercial loans’ counterparty is scoring models (De Servigny and Renault, 2004). Although the credit scoring method was explored and introduced by Altman (1968) several decades ago, it is still a topical theme for researchers and practitioners. Today, different and more sophisticated methods, such as nonparametric techniques or machine learning methods, can be applied in credit risk management. Another approach to assessing default risk is based on firm-value-based or structural models that describe the default process as the explicit outcome of the firm value’s deterioration. Based on this approach, corporate securities are considered contingent claims or options on the issuing firm’s value. This method was introduced by Merton (1974) as the first example of an application of option pricing methodology to price corporate securities. Credit scoring models can be applied to any borrower, whereas structural models can be primarily used for the largest companies listed on stock exchanges (De Servigny and Renault, 2004). Loss given default (LGD) is the second key variable in a credit risk analysis, and it can be considered the fractional loss due to default, provided that default is given in this case. The complement to one is called recovery rate; for example, if a fractional recovery rate is 30%, 70% of the exposure is LGD. The recovery rate is expressed as a percentage of the par amount recovered on defaulted debts and refers to the amount of money recovered. LGD can be defined as 1i LGD f=− (2.1) where fi is the recovery rate (Jorion, 2011). The main difference between the probability of default (PD) and loss given default (LGD) is that a distribution better represents LGD than a single figure. As De Servigny and Renault (2004) suggest, uncertainty about recovery depends on quantifiable factors and more fuzzy factors such as debtors or creditors’ bargaining power. For example, there is a clear link between seniority and the recovery level, as shown in Table 2-1. The table shows the debt recoveries of companies rated by Moody’s during 1985 – 2016. We can see that recoveries correlate with their priority of claim in the capital structure in most cases, where claims with higher priority have higher average recovery rates. There are small differences in recovery rates between Europe and the rest of the world; however, it must be noted that these results, particularly European recoveries, are based on a relatively small sample of loans and bonds. The Essentials of Credit Rating Assessment 19 Micro-Modelling Approaches for Credit Rating and Corporate Survival determine the probability of default, usually based on their experience. Generally, the main factors affecting an individual credit risk are related to the borrower's personal and economic position, which banks have used for many years. For example, Chapman (1940) specifies two types of aspects related to credit risk in personal lending: • Personal characteristics such as age, sex, family status, and • occupational and economic position, for example, income and borrower's net worth. The ability to pay is primarily determined by the applicant’s employment and the industry in which they are engaged. These factors are related to borrowers' income, assets such as real estate, automobiles, securities, and debts, such as mortgages, credit cards, and other personal loans that can be used to identify the financial capacity. In many countries, the credit history of a borrower’s responsible repayment of debts is recorded and used by lenders as an essential aspect to determine individual creditworthiness or an individual’s ability to repay a debt. The importance of each of the former factors can differ for different banks, and it is usually subject to their assessment. In assessing the credit quality of a loan applicant, lenders look at various measures. The starting point is the applicant’s credit score, a numerical grade of the borrower's credit history (Fabozzi, 2013). The well-known and widely used credit score system by lenders in the United States is the FICO Score, developed by Fair Isaac Corporation and first introduced in 1989. This system is used to assess the credit risk of individual borrowers and determine whether to extend credit. This assessment is based on account payment history, the current level of indebtedness, types of credit used, lengths of credit history or new credit accounts (Fair Isaac Corporation, 2017). FICO scores range from 300-850, with industry-specific scores from 250-900, where the higher the score, the lower the credit risk. The basic scheme of FICO scores and their definitions are in the table below (Table 2-2). According to recent data, American consumers' average FICO score reached 699 in late 2016 (Karimzad, 2015). Table 2–2 FICO Credit Scores FICO Scores Definition 800 + Excellent, an exceptional borrower 749-799 Good, a very dependable borrower 670-739 Average, a good score borrower 580-669 Fair, below the average borrower 579 and lower Poor, a very risky borrower Source: Karimzad (2015), author The system of FICO Scores is based on the following five categories: • Payment history refers to a borrower's historical ability to pay their payment on time; this category represents the most critical factor in the 20 Chapter 2 2024 Martina Novotná credit assessment. Credit history usually includes credit cards, retail accounts, instalment loans, or mortgage loans. • Amounts owed show the amounts owed on specific accounts, including credit card balances, instalment loans, and other revolving credit accounts. • Length of credit history positively affects the credit score; the longer the record, the higher the score. • Credit mix is another crucial determinant of the score, especially the total number and types of borrowers' accounts. • The new credit category suggests that opening several credit accounts in a short period represents a greater risk, especially for people with a brief credit history. The contribution of each category to the total FICO score is shown in Figure 2-6. As we can see, the significant factors in credit assessment are payment history and amounts owed. Figure 2–6 FICO Score categories Source: Fair Isaac Corporation (2017), author The process by which the lender decides whether an applicant is creditworthy and should receive a loan is called underwriting. The requirements specified by the lender to grant the loan are called underwriting standards. The approval process can be judgmental, fully automated, or a combination of the abovementioned types; however, it should consider all necessary information to support loan granting decisions. In the case of secured loans, collateral identification should also be considered (FDIC, 2017). For example, the two primary quantitative underwriting standards for granting residential mortgage loans are: • Payment-to-income ratio (PTI) that refers to the rate of monthly payments to monthly income. PTI is used to measure an applicant's ability to make monthly payments. The higher the ratio, the lower the risk. • Loan-to-value ratio (LTV) is the ratio of the loan amount to the market or appraised property's value. The lower the rate, the lower the risk for a lender (Fabozzi, 2013). Payment history 35% Amounts owed 30% Credit history 15% Account diversity 10% New credit 10% The Essentials of Credit Rating Assessment 21 Micro-Modelling Approaches for Credit Rating and Corporate Survival 2.5.2 Corporate Credit Risk Corporate credit risk assessment examines firm-level credit risk that can be affected by various factors. Firm credit risk analysis typically involves two parts of the evaluation: business and financial risks. Firstly, business or operating risks are associated with risks that originate from other than the company's financial aspects. These risks include outside and inside events with a potential impact on the business credit risk, for example, changes in economic, regulatory, climatic, industry, demographic, geo-political, product innovations, quality of management, or other factors. For example, Joseph (2013) suggests the following three categories of risks from the operating environment: • External, • industry, • internal. External risks can be seen as systematic risks that involve the impact of the business cycle, economic conditions (private consumption, government spending, investment, imports and exports), inflation, the balance of payments, exchange rates, political factors, fiscal policy, monetary policy, demographic factors, regulatory framework, technology, environmental issues, international developments and other types of systematic risks. It should also be considered that these external variables are usually interrelated. Industry analysis focused on industry life stage, composition, nature, or structure is another crucial part of credit risk analysis. In this part of the study, the stage of the industry life cycle, government support, factors of production, the sensitivity of industry to the business cycle and industry profitability should be examined, followed by competitor group analysis. Industry profitability assessment is usually based on the analysis of forces that determine the potential of an industry, known as Porter's model, which provides a basis for analysing the level of competition Figure 2-7. Figure 2–7 Porter’s model Source: Joseph (2013, p. 67), author Bargaining power of suppliers Threat of substitutes Threat of new entrants Bargaining power of buyers Industry rivalry 22 Chapter 2 2024 Martina Novotná Finally, internal or company credit risk analysis is focused on the capabilities, resources strategies, competencies, strengths and weaknesses of the borrowers (Joseph, 2013). In addition, attention is paid to business activities and identifying internal risks, including peer comparison and SWOT analysis. Other internal risks include, for example, production, human resource, product, customer/supplier concentration, legal, reputation or financial risks. The second type of firm credit risk is originated solely from the financial aspects of a business. Because even a successful business may go bankrupt due to inappropriate financial decisions, substantial attention is paid to analysing financial risks. Financial risks are linked to a company's financing policies, strategies, and decision-making that can substantially affect the credit risk level. Financial risk analysis is primarily based on the analysis of financial statements, the balance sheet, income statement, and cash flow statement; however, other financial statement information, such as a statement of stockholders’ equity, can also be useful. For credit risk assessment, a business's comprehensive economic analysis is conducted to identify the financial strengths and weaknesses and warning signals of financial risks. The study involves common size analysis, indexed trend analysis and financial ratio analysis. Typically, the focus is paid to all categories of financial ratios (liquidity, solvency, activity and profitability ratios), including studying their relationships and eventually predicting financial default. The procedure of the analysis of financial statements and their interpretation have already been discussed in many publications, see for example, Fridson and Alvarez (2011), Berk and DeMarzo (2017), Brealey et al. (2014), Megginson et al. (2008), Joseph (2013) or Dluhošová et al. (2014). The analysis of relations among financial ratios is usually examined through the DuPont Model; however, scoring models are typically applied for default prediction. To summarise, business and financial risks should be studied together, and the final credit risk assessment should be based on the company's overall situation. 2.5.3 Corporate Credit Scoring Models We can understand credit scores as statistically derived indicators of risk that indicate the relative risk that a borrower will experience an adverse credit event, for example, delinquency or default. When the credit scoring model is built, the statistical model's output is usually transferred to generate a set number of score points or the probability of a credit event occurrence (Mays and Lynas, 2011). Lenders develop models to assess borrowers' credit risk, both at an individual and corporate level. While an example of a well-known individual credit scoring model used by financial institutions is the FICO model, specific models are developed to assess corporate credit risk. Corporate credit scoring models include both the models developed by banks based on their borrowers’ behaviour and publically available scoring models. Different entities may use the later models to get an overall picture of a counterparty's creditworthiness, for example, business partners. In contrast to individual borrower credit scoring models, corporate credit scoring models use different input variables, typically financial ratios and other corporate financial performance indicators. The Essentials of Credit Rating Assessment 23 Micro-Modelling Approaches for Credit Rating and Corporate Survival Several methods can be used to derive scoring models, as explained further in the text. The proposed models are then used to calculate the score values of entities and can be used to classify them into pre-defined categories. One of the bestknown models in this area was developed by E. I. Altman (1968), whose default model is often known as the Altman’s model or ZScore model as a tool in the financial analysis of a company. This model can identify companies with possible financial problems, namely default risk, and it can be proposed based on multivariate discriminant analysis, whose product is a so-called Z-score, classifying companies. The original version of the Z-Score model can be used to predict the likelihood of a firm going bankrupt, and the score can be calculated using the following formula (Joseph, 2013), 1 2 3 4 5 1.2 1.4 3.3 0.6 0.999Z X X X X X= + + + + , (2.3) where the variables in the formula refer to the following ratios: 1 X working capital/total assets; 2 X retained earnings since inception/total assets; 3 X profit before interest and tax/total assets; 4 X market value of equity/book value of total debt; 5 X sales/total assets. The first Altman’s prediction model is based on a weighting system of five financial ratios. It was developed based on statistical data from sizeable public manufacturing companies with more than $1 million in assets. Its primary purpose is to measure a company’s financial health and predict the probability of bankruptcy within two years. Although some empirical studies show that the model has a 72% – 80% reliability of predicting bankruptcy, we should realise that it can only be used to forecast if a company being analysed can be compared to the database. The resulting scores of the original Z-Score model for public manufacturing companies and their implications can be seen in Table 2-3. Table 2–3 Original Z-Score model Z-Score Forecast Above 3.0 Bankruptcy is not likely 1.8 to 3.0 Bankruptcy cannot be predicted – GREY AREA Below 1.8 Bankruptcy is likely Source: Wilkinson (2013) Although the model is relatively simple, it is still used and is mainly relevant for manufacturing companies. According to Cao (2016), Altman decided on two potentially very powerful variables among all possible financial variables that had not been used yet. One of the variables is the retained earnings because, as Altman explains, “a firm that has grown its assets mainly by reinvesting earnings is healthier than a firm that has grown the assets by using other people’s money. Retained earnings is also a measure of the company's age and leverage” The other 24 Chapter 2 2024 Martina Novotná variable is the market value of the equity relative to the book value of the debt, an indication of the company's ability to raise money from capital markets. Today, equity's market value is a fundamental part of structural models provided, for example, by Merton (1974) or the KMV model by Moody’s Analytics (2017). Since the first version of Altman’s model, several modifications have been suggested, or some comments have been published by Altman, Altman et al. (i.e. 1970, 1977, 2005, 2007, 2010) to other authors. It is necessary to realise that such models' development is highly demanding due to data intensity and modelling specifics. These techniques are difficult to employ without information technologies and specific mathematical-statistical applications. 2.6 Issue-Specific Credit Risk Issue-specific credit risk typically refers to a bond issuer's credit risk, specific bond issues, or other issues of debt securities, which represent a contractual agreement between a lender (investor or bondholder) and a borrower (issuer). However, credit risk can also be associated with innovative contracts such as asset-backed securities or credit derivatives. In this chapter, these three categories of securities will be discussed in more detail, particularly in the context of credit risk. 2.6.1 Credit Analysis of Bonds A bond can be defined as a debt instrument requiring the issuer to repay the investor the amount borrowed plus interest over a specified period (Fabozzi, 2013). Mostly, the principal must be repaid on the maturity date. Thus, we can see an analogy between financial institutions or other entities lending money and the issuance of securities from a credit risk perspective. Similarly, the credit risk assessment will be conducted based on the borrower’s ability to repay all contractual payments when the amounts are due. Therefore, the analysis is focused on the study of issuer business and financial risks, including the analysis of financial statements. On the other hand, bond credit risk analysis should consider some features specific to bonds, such as the study of indenture and covenants. As Fabozzi (2013) suggests, the issuer's nature is a vital feature of a bond. There are three types of issuers of bonds: governments, municipalities, and corporations. Some bonds are issued with an amortisation feature, meaning that the principal repayment can be repaid over the bond's life; these securities are called amortising securities. In addition to simple or ‘plain vanilla’ bonds, there are also bonds with embedded options, for example: • Bonds with a call provision: The issuer has the right to retire the debt before the scheduled maturity date. • Bonds with a put provision: The bondholder has the right to sell the issue back to the issuer at par value on pre-specified dates. • Convertible bonds: The bondholder has the right to exchange the bond for a specified number of shares of common stock. The Essentials of Credit Rating Assessment 25 Micro-Modelling Approaches for Credit Rating and Corporate Survival • Exchangeable bonds: The bondholder can exchange the issue for a specified number of common stock shares of a corporation different from the bond issuer. Investing in bonds is associated with some risks, such as interest rate, reinvestment, call, credit, inflation, exchange, liquidity, or volatility risks. While attention in this chapter will be paid to bond credit risk, the description of other types of risks can be found in a vast literature on this subject, for example, Fabozzi (2013), Bodie et al. (2011), Reilly and Brown (2015), Petitt et al. (2015). Corporate bond credit analysis consists of three areas (Fabozzi, 2013): • Analysis of covenants, • analysis of collateral, and • assessing an issuer’s ability to pay. Analysis of covenants is linked to the study of the indenture provisions that form rules for essential areas of operation for corporate management. These provisions, including bond covenants, can be found in a company’s prospectus for its bond offering. There are generally two covenants: affirmative (promises by the corporation) and negative or restricted (limitations on the borrower). Restrictive covenants may limit the absolute amount of outstanding debt or a fixed charge coverage ratio test. For example, the maintenance test requires the borrower’s earnings ratio to be available for interest or fixed charges at a minimum for a certain period. On the other hand, the debt incurrence test is used to adjust interest or fixed charge coverage when the company takes on additional debt. In some indentures, we can also find limitations on subsidiaries’ borrowing from all other companies except the parent. Analysis of collateral refers to the careful understanding of a corporate debt obligation security when the debt can be secured or unsecured. Generally, if the company is liquidated, proceeds from bankruptcy are preferably distributed to creditors. Secured bonds are collateralized by an asset (i.e. property, equipment). In the event of default, investors claim the issuer’s assets to recover their loss to some extent. However, most corporate bonds are unsecured, and investors have no claim on specific collateral. We can see this fact in Table 2-4, which shows the proportion of secured and unsecured bonds issued by European industrial companies as of the end of 2010 4 . While the balance of unsecured bonds is more than 90% of total bonds, secured bonds represent a minority in both groups of countries. 4 There are 23 countries included and divided into two groups EU-15 (Austria, Belgium, Denmark, Finland, France, Germany, Greece, Ireland, Italy, Luxembourg, Netherlands, Portugal, Spain, Sweden, United Kingdom) and EU-8 (Czech Republic, Estonia, Hungary, Latvia, Lithuania, Poland, Slovakia, Slovenia). 26 Chapter 2 2024 Martina Novotná Table 2–4 Proportion of secured and unsecured bonds EU-15 EU-8 Total proportion (%) Secured/Senior Secured 80 2 7.9 Unsecured/Senior/Subordinated Unsecured 875 85 92.1 Total 955 87 100 Source: Reuters database (accessed 1st December 2010), author’s calculations The third area of the bond credit analysis is focused on assessing an issuer’s ability to make timely payments of interest and principal. Although a substantial part is based on the analysis of financial statements, we should also analyse other factors that may impact the ability to generate cash flow, thus service the debt. This part of the credit analysis is analogical to the research described in Chapter 2.1.2. It involves studying business risk and financial risk, including assessing corporate governance risk with an emphasis on the ownership structure of the corporation, the practices followed by management and policies for financial disclosure. Government bond credit risk is associated with the country's overall situation and institutional strength, especially the banking system's stability and policy credibility. Traditionally, government bonds have been considered relatively riskfree securities; however, we can find significant differences among different countries' risks. Government bond credit risk, so-called sovereign credit risk, is usually assessed by rating agencies in terms of rating. However, financial institutions, other entities or investors may also use their credit scoring models, similar to corporate loans. According to Moody’s rating agency 5 , there are four factors of government credit risk analysis: • Economic strength: GDP (per-capita), economy size and degree of diversification, medium-term trends (productivity, infrastructure). • Institutional strength: Policy predictability (continuity), institutional quality, regulatory framework. • Government financial strength: Fiscal balance, debt indicators (ratios to GDP and revenue), debt affordability (interest/revenue), debt structure, and market access. • Susceptibility to event risk: The impact of economic, financial and political events. The two critical factors in the creditworthiness analysis are government financial strength, monetary policy, and economic power, indicating the economic trend. Susceptibility to event risk is the factor that shows the shock resistance of a country, for example, the impact of the financial crisis, Brexit or a US presidential election on the economy and fiscal outlook. 5 Moody’s Investors Service, Moody’s 7th Annual CEE Credit Risk Conference, Czech National Prague, 16 April 2013. The Essentials of Credit Rating Assessment 27 Micro-Modelling Approaches for Credit Rating and Corporate Survival In addition to government bonds, there are debt securities issued by local governments, districts, or cities called municipal bonds. The credit risk of municipal bonds is also assessed and published by rating agencies to help investors make investment decisions. The main factors of credit risk analysis cover (SEC, 2017; Peterson, 1998): • Sources of funds to pay principal and interest, • purpose of the financing, and • the financial condition of the issuer. In this case, financial condition analysis is focused on the magnitude and structure of local debt, including a proposed borrowing. Economic analysis can be used as an adequate indicator of municipal debt burden and municipal borrower’s ability to service the debt. For example, the most important financial ratios include debt service related to recurring revenues, operating surplus or total income and total debt to the tax base. Finally, credit risk is associated with short-term securities, such as commercial papers and other short-term debts with maturities of up to one year. The credit analysis is slightly different from long-term bonds as it usually does not consider the likely recovery of the debt instruments. The principal factors of the credit analysis include assessing the fundamental long-term credit quality. However, the short-term credit risk is driven predominantly by the issuer’s liquidity position, which indicates the ability to repay the debt from internal or external sources. 2.6.2 Credit Risk of Asset-Backed Securities Asset-backed securities are considered an innovative and alternative way corporations or lenders can raise funds. Through securitization, a corporation pools loans or receivables and uses the pool of assets as collateral for security issuance (Fabozzi, 2013). Since the cash flows are sold in the form of securities backed by the cash flows of the very assets sold, the securities are called assetbacked securities. Compared to traditional ways of debt financing, such as borrowing in the form of a loan or issuing bonds, securitization is associated with specific features and represents a different way of financing. Issuers of assetbacked securities typically raise funds to finance the origination of loans, and from an accounting point of view, this issuance is considered an asset sale. There are various backing assets, such as residential or commercial mortgages, consumer loans, commercial leases, or any financial instruments with predictable and stable receivable cash flows (i.e. credit card receivables, auto loans, student loans). As lenders issue asset-backed securities by structuring future receivable cash flows of underlying assets, they are also called structured finance securities. There are the following specific features of the securitization process (Hu, 2011): • The asset-backed securities are issued through a special purpose entity, • to accounting aspects, the issuing of asset-backed securities is an asset sale (not a debt financing), 28 Chapter 2 2024 Martina Novotná • servicing of the underlying assets for the investor is required, • the credit of the asset-backed security depends on the credit of the underlying asset, • credit enhancement is usually needed. Asset-backed securities are issued in the form of certificates entitling the investors to receive a pre-determined share in a specific pool of assets' cash flows. From this perspective, they are similar to bonds because investors accept regular payments, usually based on a coupon rate. However, in asset-backed securities, the payments depend on the cash flows generated by underlying assets. Thus, principally, we can distinguish between two types of asset-backed securities: • Pass-through securities that are issued as single-class mortgage-backed securities (i.e. agency MBS) and • securities structured in several bond classes called tranches (i.e. nonagency MBS, ABS). A mortgage pass-through security is issued as one bond class, which means that investors are entitled to receive a pro-rata share of the cash flows of the specific mortgage loan pool. When pass-through security is first issued, the principal is known; however, over time, due to regularly scheduled principal payments and prepayments, the amount of the pool’s outstanding loan balance declines. Payments of pass-through security are made each month, and the monthly cash flow is less than the monthly cash flow of the loan pool by an amount equal to servicing and other fees. Agency mortgage pass-through securities, socalled mortgage-backed securities (agency MBS), are issued by US government agencies known as Freddie Mac or Fannie Mae, and Ginnie Mae, which is not the issuer; however, it provides guarantees. Since these pass-through securities carry their warranty and fulfil underwriting standards, they can be considered assetbacked securities with the lowest level of credit risk. Securities structured in bond classes are created to redistribute credit risk using a senior-subordinate structure. While bond classes with the lowest credit risk and the highest rating are referred to as the senior bond classes, the subordinated classes have a lower rating or are not rated. Losses are distributed based on the bond class's position in the structure when losses start from the bottom and move to the senior level. The rules for the cash flow distribution (interest and principal) and losses are explained in the prospectus. They are usually referred to as cash flow waterfall (Fabozzi, 2013; Hu, 2011; Choudry, 2010). The Essentials of Credit Rating Assessment 35 Micro-Modelling Approaches for Credit Rating and Corporate Survival • countries, • credit default swaps, • insurance companies, • municipalities, • structured finance counterparty instruments, • structured finance counterparties, • structured finance interest-only securities. Both agencies also use national scale ratings for the opinions of issuers' relative creditworthiness and financial obligations within a particular country. They are not designed to be compared among states; conversely, they address relative credit risk within a given country. Thus, they are opinions of an obligor’s creditworthiness or overall capacity to meet specific financial obligations relative to other issuers and issues in a given country or region. The notation of national scale rating is based on the characters that indicate the state in which the issuer is located, for example, ‘Aaa.br’ (Moody’s), or ‘brAAA’ (S&P) ratings demonstrate the strongest creditworthiness relative to other domestic issuers in Brazil. National scale systems are maintained only for some countries. We can use the case of the Czech Republic to demonstrate national scale ratings and their use (Table 2-6). There are municipal ratings in the table, both long-term and national scale ratings. Although most long-term ratings are A2, there is a relative difference among issuers according to national scale ratings within the country. Table 2–6 Long-term and national scale municipal ratings in the Czech Republic Issuer Long-term rating National scale rating Ceska Lipa, City of A2 Aa2.cz Klatovy, City of A2 Aa2.cz Liberec, City of Baa1 A3.cz Liberec, Region of A2 Aa3.cz Prostejov, City of A1 Aa1.cz South-Moravian Region A2 Aa3.cz Trebic, City of A2 Aa2.cz Uherske Hradiste, City of A2 Aa3.cz Usti, Region of A2 Aa3.cz Zdar nad Sazavou, City of A2 Aa2.cz Source: Moodys’ Corporation (2017), author Without considering some current issues in the credit rating industry and the problem of misleading some ratings, they generally provide a handy tool for all participants in the financial market. There are many rating users in the financial market; however, the most important rating beneficiaries are investors and issuers. For investors, ratings provide an easily understandable and reliable guide about the likelihood of issuer default on a particular fixed-income instrument. Thus, rating offers the basis for investment decisions. In addition, since rating increases knowledge and transparency, it can reduce uncertainty and provide a benchmark for comparisons. 36 Chapter 2 2024 Martina Novotná 2.7.2 Benefits and Costs of Rating The main roles of rating for investors are as follows (Nye, 2014): • Ratings can help long-term investors (e.g. pension funds, insurance companies) to evaluate various long-term investment options. • Ratings provide the basis for investors to make a more informed investment decision and to match the relative credit risk or debt with their risk tolerance. • The rating system allows the issuer’s credit fundamentals to be compared against the industry peer group. • Ratings reduce investors’ costs of gathering, analysing and monitoring borrowers' financial positions. Ratings also have benefits for other market participants, such as debt issuers. Issuers such as corporations, financial institutions, governments, cities and municipalities use credit ratings to provide independent views of their creditworthiness and the credit quality of their debt issues. Thus, ratings play a useful role in enabling issuers to raise money in the capital markets and facilitate the process of issuing and purchasing bonds by providing a measure of relative credit risk. In addition, debt issuers that acquire ratings may benefit from: • A broader base of investors, lenders, customers or business partners, • alternative financing options, • a standard measure of creditworthiness relative to other issuers in an industry or a country, • lower financing costs and more dependable access to liquidity, • a more accurate estimate of borrowing costs, • a useful discipline on senior management (Nye, 2014). There are additional benefits of ratings for an issuer that is a corporate enterprise. These include lower financing costs, better negotiation power with banks, the basis for offering debt issues in the market, higher visibility and credibility or a peer benchmark. For commercial banks, rating gives confidence to depositors, regulators and may help attract equity investors. If the issuer is a sovereign government, a rating helps attract international capital, support the development of local capital markets, meet international standards of transparency and cooperation or achieve more excellent funding stability (Nye, 2014). Even though ratings play a crucial role in the financial market, we should also consider some rating scales' weaknesses. The disadvantages of rating are related to the issues discussed in Chapter 2.9, primarily the negative influence of the high concentration of the CRA market on financial stability, conflicts of interest, and transparency. The lack of accountability in the rating industry may lead to the limited reliability of rating scales and consequently to increased inefficiency in the financial markets. As mentioned in the previous chapter, rating scales are easy to read and use; however, there might also be a potential problem when unqualified users use and interpret these scales. It is essential to realize that ratings are opinions about credit risk and do not provide an absolute measure of default probability. The Essentials of Credit Rating Assessment 37 Micro-Modelling Approaches for Credit Rating and Corporate Survival The overall credit rating industry is based on the scheme that issuers pay the credit rating agency. From the point of view of new issuers, agency fees should be considered, and the issuer should conduct a cost-benefit analysis of acquiring a rating. Typically, agency fees are negotiable, and issuers can discuss fee options with the agencies. According to Nye (2014), competition among CRAs keeps prices relatively modest, depending on the rated entity. For example, according to the S&P 7 , the minimum fee for industrial corporations and financial service companies is $150,000. 2.8 Description of Credit Rating Process The credit rating process typically involves several steps, requiring close cooperation between the rating agency and the issuer. First, new issuers must choose and contact a rating agency that fits their needs. Second, the rating process begins with an initial meeting to introduce the issuer's approach, methodology, and products. Finally, the steps of the rating process (Figure 2-9) include the following activities: • The issuer signs an application, • rating agency analysts are assigned to the customer to review the relevant information, • analysts meet with the management team to review and discuss information, • analysts evaluate information and propose the rating to a rating committee, • the committee reviews the lead analyst’s rating recommendation and then votes on the credit rating, • the issuer is generally provided with a pre-publication rationale for its credit rating for fact-checking and accuracy purposes, • a press release announcing the public rating is typically published, • the rating is kept current by identifying issues that may result in either an upgrade or a downgrade (S&P Global Ratings, 2017). 7 Global Ratings U.S. Ratings Fee Disclosure [online]. Available at: http://www.standardandpoors.com/, [Accessed 19 September, 2017] 38 Chapter 2 2024 Martina Novotná Figure 2–9 Steps of rating process Source: S&P Global Ratings (2017), author During the rating process, the issuer, an industrial company, provides relevant financial and non-financial information. Then, the discussion at the management meeting generally focuses on the following subjects 8 : • Background and history of the entity, • industry and sector trends, • the national political and regulatory environment, • management policies, experience, track record, attitude toward risk-taking, • management structure, • primary operating and competitive position, corporate strategy and competitive philosophy, • debt structure, including structural subordination and priority of claim, • financial situation and liquidity sources (cash flow, operating margin, a balance sheet analysis of debt profile and maturity). As we can see, assessing companies' creditworthiness is a complex task based on both quantitative and qualitative analysis. CRAs emphasize the qualitative side, arguing that no ratios can capture the full complexity of a company’s financial position, cash available to meet its future obligations, and management's willingness to pay principal to bondholders. Above all, CRAs evaluate long-term fundamentals related to the company’s credit strengths and weaknesses related to its ability to generate cash, plausible stress scenarios, and elements of future cash flow (Nye, 2014). Rating agencies use different assignment methodologies according to the counterparty’s nature (i.e. corporations, countries, public entities) and the nature of products (i.e. bonds, structured finance). Since the empirical analysis in the application part focuses on corporate borrowers, the primary attention is concentrated on risk factors for corporate ratings or individual debt issues specifically. 8 Moody’s Investors Service (2017) The Essentials of Credit Rating Assessment 39 Micro-Modelling Approaches for Credit Rating and Corporate Survival Analysts of CRAs typically start with evaluating the issuer's creditworthiness before assessing the credit quality of a specific debt issue. The main factors of rating assessment described in the previous text consider financial and nonfinancial factors, including key performance indicators, economic, regulatory and geopolitical influences, management and corporate governance attributes, and competitive position. In a rating of an individual debt issue, analysts primarily focus on (S&P Financial Services, 2014): • The terms and conditions of the debt security, including its legal structure, • the relative seniority of the issue concerning the issuer’s other debt issues and priority of repayment in the event of default, • the existence of external support or credit enhancements (i.e. letters of credit, guarantees, insurance or collateral). The critical factors of assessing the corporate rating can be summarized using the rating analytical pyramid (Figure 2-10). Figure 2–10 Rating analysis pyramid Source: Moodys’ Corporation (2017), author Rating agencies can use both top-down and bottom-up approaches to assess the rating, with the main parts of the assessment focusing on the country, industry, company, and indenture analysis. Credit ratings can be upgraded, downgraded, or unchanged, and the percentage of unchanged ratings can be used to measure rating stability. According to the rating changes, we can calculate the transition frequency from one rating class to another. Thus, we can assess the migration risk. The entire possible states a rating can take over a given time horizon is usually called rating transition or migration matrix. Transition matrices are based on time series of rating changes, and they can be estimated using various approaches. For example, Gunnvald (2014) use the Markov chain method and the duration method, Hu et al. (2002) use the combination of a Bayesian approach and ordered probit estimate of sovereign ratings, and Koopman et al. (2008) use an intensity-based duration model. INDENTURE ANALYSIS Issue structure COMPANY SPECIFIC ANALYSIS Financial risk Business risk and management INDUSTRY ANALYSIS Market position Industry trends COUNTRY ANALYSIS Legal and regulatory framework Sovereign macroeconomic analysis 40 Chapter 2 2024 Martina Novotná Frydman and Schuermann (2008) applied Markov mixture models or a mixture of two Markov chains. Their results show that a firm’s rating depends on its current credit ratings and its past rating history. Transition matrices are published by CRAs and can be used to analyse transition rates and rating behaviour. For example, according to S&P (2015), investment-grade rated issuers exhibit greater rating stability (measured by the frequency of rating transition) than speculative-grade issuers. It is confirmed by the one-year transition rates (Table 2-7): The lower the rating, the lower the rating stability. For instance, while 93.26% of AA issuers are still rated as AA one year later, 79.97% of BB issuers stay within the BB rating category in one year. The rating category AAA is relatively stable; however, it should be considered that because the number of issuers with AAA ratings is typically deficient, the downgrade of even one issuer may have a large effect on this category’s rating stability. Table 2–7 One-year corporate global transition rates (2015, in %) From/to AAA AA A BBB BB B CCC/C D NR AAA 100.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 AA 0.29 93.26 4.40 0.00 0.00 0.00 0.00 0.00 2.05 A 0.00 1.43 89.87 5.48 0.00 0.00 0.00 0.00 3.23 BBB 0.00 0.06 3.12 85.52 4.90 0.00 0.00 0.00 6.40 BB 0.00 0.00 0.00 3.63 79.97 6.87 0.24 0.16 9.13 B 0.00 0.00 0.00 0.15 3.58 76.04 4.57 2.39 13.27 C/CCC 0.00 0.00 0.00 0.00 0.00 5.85 49.71 25.73 18.71 Source: S&P (2015), author The level of rating stability depends not only on the rating category but also on the time horizon. Over the long term, rating stability is also consistent with higher ratings; however, the average long-term transition rates are usually higher than one year. For example, from 1981 to 2015, AAA-rated issuers were still rated AAA one year later 87.1% of the time, while CCC/C ratings remained CCC/C 44.2% of the time (Table 2-8). The category denoted as D means payment default on one or more of an obligor’s financial obligations (according to rating agency’ definition of default), or issuer’s filing for bankruptcy, and NR stands for not rated. The analysis of credit migration is the basis from the CreditMetrics approach developed by JP Morgan, the U.S. bank in 1997, subsequently revised as RiskMetrics, Inc. and acquired by MSCI in 2010. In practice, banks widely use this approach to estimate the full one-year forward distribution of any bond or loan portfolio values, where the changes in values are only related to credit migration. The critical assumption is that rated bonds' past migration history accurately describes the probability of migration in the next period (Crouhy et al., 2014). The Essentials of Credit Rating Assessment 41 Micro-Modelling Approaches for Credit Rating and Corporate Survival Table 2–8 Long-term corporate global transition rates (1981 – 2015, in %) From/to AAA AA A BBB BB B CCC/C D NR AAA 87.08 9.00 0.53 0.05 0.08 0.03 0.05 0.00 3.18 AA 0.53 86.69 8.06 0.53 0.06 0.02 0.02 0.02 4.02 A 0.03 1.81 87.65 5.39 0.33 0.02 0.02 0.06 4.58 BBB 0.01 0.11 3.55 85.43 3.82 0.12 0.12 0.19 6.24 BB 0.00 0.03 0.13 5.08 76.78 0.64 0.64 0.73 9.63 B 0.00 0.03 0.09 0.21 5.25 4.39 4.39 3.77 11.99 C/CCC 0.00 0.00 0.14 0.20 0.61 44.19 44.19 26.36 15.66 Source: Source: S&P (2015), author 2.9 Current Issues in Rating Industry CRAs produce ratings in both local and international markets. The credit rating industry is regulated by Regulation (EC) No 462/2013 9 and Directive 2013/14/EU 10 in the European Union. This regulation (hereafter referred to as EC Regulation) aims to regulate the activity of credit rating agencies to protect investors and European financial markets against the risk of malpractice. The first incentives to strengthen the regulatory and supervisory framework of CRAs in the EU and the first set of rules came into effect at the end of 2009, connected with the global financial crisis 2008 – 2009. To be registered in the European Union (EU), credit rating agencies must (EC, 2017): • Avoid conflicts of interest: For example, credit rating analysts must not rate an entity in which they have a holding; • ensure the quality of their ratings and rating methods; • provide a high degree of transparency, for example, by publishing an annual transparency report. In addition to the regulatory framework, the European Securities and Markets Authority (ESMA) was created in 2011 to supervise CRAs registered in the EU. Since July 2011, ESMA has been responsible for registering credit rating agencies and has exclusive supervisory powers concerning such agencies. Under the CRA regulation, it is possible for a rating agency established outside the EU to have its rating recognised and used for regulatory purposes in the EU. It can happen in two ways: certification through the equivalence regime and endorsement. The equivalence certification applies to the CRAs established and supervised outside the EU that have no affiliation in the EU. These CRAs that meet the requirements of the regulation may apply to the ESMA for certification. The endorsement regimes apply to CRAs affiliated with or working closely with EU-registered agencies, and they must also comply with specific legal requirements. The list of registered or certified CRAs in the EU is available in Appendix 1. For example, we can see that as of 24 March 2022, there are 33 CRAs regulated in the EU. 9 Regulation (EC) No 462/2013 of the European Parliament and the Council of 21 May 2013 on credit rating agencies 10 Directive 2013/14/EU of the European Parliament and the Council of 21 May 2013 42 Chapter 2 2024 Martina Novotná Regarding the critical provisions of the last EC Regulation reforms, a recent study on the state of the credit rating market was published in 2016 (EC, 2016). According to this document, hereafter referred to as Report, the current credit rating market has an oligopolistic structure, and it is dominated by the three global CRAs: Moody’s, S&P and Fitch. The EC Regulation on credit rating agencies is aware that CRAs play a crucial role in the global securities and banking markets. A significant number of financial institutions use their credit ratings to estimate their capital requirements for solvency purposes or evaluate risks. It is suggested that these entities make their credit risk assessment and thus less reliance on credit ratings. Concerning credit rating agencies, the main issues that concern the European Commission include (EC, 2016): • Conflicts of interest due to the issuer-pays model, • conflicts of interest due to the remuneration model of credit rating agencies, • disclosure for structured finance instruments, • transparency, • procedural requirements and the timing of publication specifically for a specific period. Thus, specific measures have been adopted in the European Union to improve the situation. For example, the Credit Rating Agency (CRA3) Regulation's last reforms focused on enhancing competition in the credit rating market, further addressing conflicts of interest, enhancing disclosure on structured finance instruments, and the rotation mechanism for CRAs (EC, 2017). 2.9.1 Market Concentration As said in the Report, market shares based on total revenues and the HerfindahlHerfindahl Index (HHI) 11 suggest that the market of CRAs is highly concentrated, both overall and at the individual product category level, with a small increase of concentration between 2012 and 2014. The summary of market shares for credit rating activity and ancillary services provided in the European Union can be seen in Table 2-9. The three largest CRAs cover most of the market, reaching around 96% during the reference period, while the remaining share, about 4%, is filled by the other, much less critical CRAs. 11 According to the report, HHI provides an indication of concentration within markets, with an HHI over 1,000 generally considered to be concentrated. The Essentials of Credit Rating Assessment 43 Micro-Modelling Approaches for Credit Rating and Corporate Survival Table 2–9 Market shares for credit rating activity and ancillary services CRA 2012 2013 2014 Moody’s 36.69% 36.38% 36.99% S&P 32.88% 36.00% 38.43% Fitch 17.67% 16.47% 18.40% Euler Hermes Rating GmbH 0.24% 0.26% 0.25% Feri EuroRating Services AG 0.84% 0.78% 0.76% BCRA-Credit Rating Agency AD 0.02% 0.03% 0.00% Creditreform Rating AG 0.51% 0.54% 0.51% Scope Ratings AG 0.10% 0.20% 0.17% GBB-Rating Gesellschaft für Bonitätsbeurteilung GmbH 0.34% 0.34% 0.33% ASSEKURATA (Assekuranz Rating-Agentur GmbH) 0.30% 0.31% 0.27% ARC Ratings, S.A. (previously Companhia Portuguesa de Rating, S.A) 0.04% 0.03% 0.02% AM Best Europe-Rating Services Ltd. (AMBERS) 0.75% 0.73% 0.47% DBRS Ratings Limited 0.82% 1.23% 1.39% CRIF S.p.A. 0.35% 0.77% 0.07% Capital Intelligence (Cyprus) Ltd 0.00% 0.00% 0.03% European Rating Agency, a.s. 0.00% 0.00% 0.00% Axesor SA 1.85% 1.41% 0.73% The Economist Intelligence Unit Ltd 6.48% 4.41% 1.05% Dagong Europe Credit Rating Srl (Dagong Europe) 0.01% 0.01% 0.01% Spread Research 0.09% 0.09% 0.13% EuroRating Sp. z o.o. 0.01% 0.01% 0.00% HHI-index 2,787 2,916 3,189 Source: EC (2016) 2.9.2 Conflicts of Interest CRAs, similarly to other financial intermediaries, play a crucial role in the financial system because their expertise in interpreting signals and collecting information from their customers gives them a cost advantage in producing information. Thus, CRAs positively affect the problem of asymmetric information in the financial market because borrowers have some information they do not disclose to lenders. Hence, one party in the contract often does not have enough information to make accurate decisions. However, there is a potential problem of conflicts of interest: One party in a financial agreement may have incentives to act in their interest rather than in the interest of the other party. Since CRAs provide very specialized, usually multiple financial services, conflicts of interest may arise in the rating industry; see, for example, Mishkin and Eakins (2009) or Cecchetti and Schoenholtz (2015). Conflicts of interest can substantially reduce the quality of information and even increase asymmetric information problems. For this reason, the next issue of the Report is the problem of conflicts of interest arising from the fact that while 44 Chapter 2 2024 Martina Novotná investors and regulators demand a well-researched credit quality assessment, issuers need a favourable rating. Since the issuers of securities pay a rating firm to have their securities rated, investors and regulators may question their rating quality (Mishkin and Eakins, 2009). Other conflicts of interest might be between CRAs and shareholders and CRAs on a firm level and its employees (e.g. rating analysts). They are linked to ancillary consulting services provided by CRAs in terms of auditing and consultancy services. CRAs may deliver favourable ratings to attract more clients and thus increase asymmetric information in financial markets (EC, 2016). Another problem of conflicts of interest might be connected to structured debt instruments (Efing and Hau, 2015). The Report (EC, 2016) highlights that conflicts of interest are related to competition in the industry. There are two phenomena linked to this problem and discussed in the literature (EC, 2016, p. 72): • Rating shopping: In this situation, issuers solicit ratings from multiple agencies and then choose the most favourable one. This phenomenon can lead to rating inflation. • Rating catering: A situation strictly related to rating shopping. CRAs may be incentivised to loosen their standards to compete with more favourable ratings from other CRAs, usually in booming markets, when CRAs have fewer concerns about their reputation and market shares (Sangiorgi and Spatt, 2013). 2.9.3 Transparency If the operating activities and actions of CRAs are easy for other participants to see and understand, then the CRAs are considered transparent. As Nye (2014) suggests, the rating agencies have become a potent force in demanding transparency in the accounts of issuers they rate since the Asian financial crisis of 1997 – 1998. Thus, CRAs emphasised more precise data, more open accounting, and conforming to international standards to increase their reputation. However, during the financial crisis of 2008 – 2009, rating agencies failed to publish accurate and adequate ratings in many cases, especially in assessing structured products. Due to the Basle rules, there was a strong demand for high ratings by financial institutions that are not permitted to hold assets with lower rating grades. For example, as Crotty (2009) says, in 2005, more than 40% of Moody’s revenue came from a rating of securitized debt. Due to financial boom, inexperience or ignorance, CRAs issued absurdly high ratings to illiquid, non-transparent structured financial products. While the explosion of these innovative securities created large profits at financial institutions, it also destroyed the transparency necessary for any aspect of market efficiency. Although CRAs aim to improve their credibility and transparency, there is still a lack of accountability in the credit rating industry. Some arguments provided (e.g. Pagano and Volpin, 2010) that issuers should disclose all the information relevant to assessing the products' risk instead of requiring CRAs to disclose the information they used. On the other hand, there are also arguments regarding Approaches for Credit Rating and Corporate Bankruptcy Modelling 51 Micro-Modelling Approaches for Credit Rating and Corporate Survival Fisher (1959). Regression analysis became one of the most used methods to estimate ratings in this period. An alternative approach to predicting bond ratings is the multiple discriminant analysis introduced, for example, by Pinches and Mingo (1973), Ang and Patel (1975), Altman and Katz (1976) and Belkaoui (1980). Other research compared particular statistical methods; e.g. Kaplan and Urwitz (1979) compare ordered probit analysis with ordinary least square regression, and Wingler and Watts (1980) compare ordered probit analysis with multiple discriminant analysis. Other studies replicate the process of bond rating model estimation and modify the approaches by considering new variables, such as the study of Chan and Jegadeesh (2004), or examine the impact of financial variables on credit rating for a given country or region, e.g. Gray et al. (2006). Recent studies come from the theoretical framework mentioned above and extend statistical methods for new non-conservative approaches such as neural networks (Dutta and Shekhar, 1988; Surkan and Singleton, 1990). For example, Waagepetersen (2010) assesses the relationship between quantitative models and expert rating evaluation. More recently, Altman et al. (2010) focused on the importance of non-financial information within risk management. Research in rating prediction has shifted mainly to applying machine learning methods in recent years. Hsu et al. (2018) propose a model based on the artificial bee colony approach and support vector machine technique. The authors find that their bio-inspired computing mechanism improves prediction accuracy compared to other statistical methods. Golbayani et al. (2020) compare neural networks, support vector machines and decision trees. The authors find that the decision treebased model achieves the best performance. In their research, they apply conventional accuracy measures and introduce the notch distance approach, which is suitable for comparing the performance of various machine learning methods. As it turns out, all the above techniques are appropriate for rating prediction. The individual models differ mainly in the methodology, used variables, and ability to predict the rating. Therefore, research in this area is focused primarily on improving the predictive accuracy of models. For example, Wang and Ku (2021) developed the parallel artificial neural networks model that creates several independent artificial neural networks. As the authors suggest, their approach achieved competitive results compared to conventional artificial intelligence techniques. Current research shows that conventional approaches and newer methods based on artificial intelligence are widely used to model credit ratings. The huge advantage of these models is their practical applicability and the possibility of using them for potential rating revisions. In addition, research studies show that models' predictive power is sufficient and comparable to other commonly used methods in valuing historical data based on averages or growth rates. For example, Jones et al. (2015) examine the predictive performance of binary classifiers using a large sample of international credit ratings. They apply conventional techniques (logit and probit regression, linear discriminant analysis) and fully nonlinear classifiers (neural networks, support vector machines, general boosting, AdaBoost and random forests). The authors conclude that although the newer classifiers 52 Chapter 3 2024 Martina Novotná outperform others, simpler ones can be viable alternatives to more sophisticated approaches, particularly if interpretability is an important objective of predictive models. The literature review shows that various studies are focused on modelling ratings using fundamental-based approaches. However, from the application point of view, these studies do not pay sufficient attention to Czech entities or companies from CEE countries. At the same time, knowledge of the rating and its influencing factors has utilization in many areas. It is primarily an assessment of the investment quality of unrated bonds. Another application is, for example, corporate finance and the issue of determining the cost of capital. In any case, the absence and unavailability of rating models are the main motivation for this study. The main benefit is applying CEE countries' data, including the application of models corresponding to Czech conditions and needs. The partial goal is to present the models’ estimation process, interpretation, and mutual comparison. An alternative way provided in this work to model ratings is based on time-toevent or survival analysis. For example, Glennon and Nigro (2005) use a discretetime hazard framework to measure the default risk of small business loans. Roa et al. (2009) propose a survival analysis methodology to analyze falling rating duration. They test macroeconomic variables to predict this event in selected countries and find differences between developed and emerging economies. In further research, Zhang and Thomas (2012) compare linear regression and survival analysis for modelling recovery rates. The authors find that linear regression is better for recovery rate modelling; however, they suggest some adjustments and additional validation. Overall, survival analysis methods are typically used for rating or credit transitions and time series rating patterns (e.g. Parnes, 2007; Figlewski et al., 2012; Louis et al., 2013; Leow and Crook, 2014). Thus, we can model the rating behaviour over time and measure, for example, the probability of a certain change in the rating depending on time and other relevant variables. Therefore, survival analysis allows us better to understand the data and its dynamics over time. In addition, it is a method used by rating agencies to estimate default rates, which are regularly published and used by analysts and researchers in the financial market. As part of credit risk measurement, attention is also heavily paid to bankruptcy prediction. The bankruptcy of companies is typically analysed based on credit score models, which are statistically derived models for predicting credit risk. Among all the studies on scoring models, the study by Altman (1968) and the model known as Altman´s or Z-score model are worth mentioning. Since the first publication of this model, extensive research has been conducted in bankruptcy prediction and the application of discriminant analysis, logistic regression, classification trees, and neural networks. In addition, the survival analysis approach can be seen as an alternative way to examine corporate bankruptcy. However, this area has not yet attracted adequate attention compared to the traditional methods mentioned above. Approaches for Credit Rating and Corporate Bankruptcy Modelling 53 Micro-Modelling Approaches for Credit Rating and Corporate Survival Nevertheless, some studies apply survival analysis to predict corporate failure in different countries. For example, the earlier research includes Lane et al. (1986), who employed the Cox model to predict bank failure using a sample of 130 banks. As the authors suggest, the overall accuracy of their model is similar to the discriminant analysis results. Among other studies, Laitinen and Kankaanpää (1999) discuss the six most popular alternative methods to financial failure prediction, including survival analysis. However, their research proposes no only one way, even though the accuracy of failure prediction varies depending on the technique applied. Other empirical results include the study by Agarwal and Audretsch (2001), who focus on the effect of companies’ size on their survival. Their research finds that smaller companies are less likely to survive than larger companies. However, they suggest that general pronouncements are hazardous because the size changes over the industry cycle and with the technological demands of that industry. Similarly, Glennon and Nigro (2005) examined the effect of time on the probability of default on medium–maturity loans under a loan guarantee program for small firms. The authors find that the default behaviour of these loans is timesensitive. As loan seasons, the probability of default initially increases, and it declines after the second year. They also suggest that the likelihood of default is conditional on the borrower, lender, loan characteristics and changes in economic conditions. Finally, De Leonardis and Rocci (2008) used a discrete-time survival analysis approach to assess the default risk of small and medium-sized Italian companies from 1995-1998. The authors suggest that the prediction accuracy of the duration model is better than that provided by a single-period logistic model. In addition to examining the effect of financial and economic factors on corporate failure, some studies assess the impact of other factors. For example, Mokarami and Motefares (2013) examine the effect of the internal mechanisms of corporate governance on the bankruptcy of firms enlisted in the Tehran Stock Exchange. Using the Cox model of survival analysis, the authors claim a significant relationship between CEO replacement and bankruptcy. Other research includes, for example, Pereira (2014), who applied the Cox proportion hazard model in predicting business failure of companies in the textile industry, and Kelly et al. (2015), who focused on corporate liquidations in Ireland. Louzada et al. (2014) modelled the time to default on a personal loan portfolio. They state that survival models are being proposed in financial risk management as alternative tools due to the continuous monitoring of risk over time. Their empirical study is illustrated by credit data from a Brazilian commercial bank. Their results show that attention should be paid to continuously checking the validity of requirements for using the available models. Besides the problems of loan default and bankruptcy, Kristanti and Isynuwardhana (2018) examined the effect of certain predictors on the probability of financial distress of companies enlisted on the Indonesia Stock Exchange. Applying the Cox hazard model of survival analysis, they found evidence that there is an inverse relationship between the control of corruption and the probability of financial distress, except for the predictors such as leverage, operational risk, and size. 54 Chapter 3 2024 Martina Novotná Overall, there is a vast literature on predicting corporate bankruptcy using various techniques. However, there is still little attention to modelling corporate bankruptcy using time-to-event methods. This fact is one of the motivations of our research, which is to compare commonly used approaches with less frequently applied survival analysis methods. The main contribution of this monograph is the expansion of existing research in this area and the application of selected models to specific data from CEE countries, respectively, from the Czech Republic. The main goal is to find a link between rating and survival models, identify the main predictive variables and propose a procedure to convert bankruptcy rates into rating evaluations. As a result, the association between the probability of survival and the rating, or the predicted development of the rating over time, can be better understood. Furthermore, as the rating is widespread and used globally, we believe the interpretation of credit risk using the rating is more suitable for users, especially for individual investors. All participants in credit contracts can use the main findings of this work. In addition, it is useful for analytical departments that create mathematical-statistical models for monitoring and measuring credit risk in banks and other financial institutions. Nevertheless, we see the main use on the part of retail investors, whether individuals or companies, who can use the partial results of the work in several directions, mainly for a better understanding of the factors that significantly influence the survival probability and, thus, the overall rating evaluation. Furthermore, the results of this work can be further used to apply selected models to their own data and their subsequent use to measure credit risk. Finally, this work can also be used in academic research as an example of connecting two different approaches to creating micro credit risk models and possibly expanding further. 3.2 Discriminant Analysis Discriminant analysis is a standard statistical method used to separate groups and, thus, a suitable method for credit scoring or bond rating modelling. The analysis can be used for two primary objectives: first, the description of group separation; second, predicting or allocating observations to groups. Huberty and Olejnik (2006) distinguish between descriptive discriminant analysis (DDA) and predictive discriminant analysis (PDA). The purpose of DDA is usually the study of comparison among a certain number of groups, for each of which we have several outcome variable scores. However, suppose a single set of response variables is used as predictors, and there is a single grouping variable. In that case, the primary purpose is to analyse how well group membership of analysis units may be predicted using PDA. Correspondingly, Rencher (2002) differentiates between discriminant and classification functions. Discriminant functions separate groups, while classification functions assign individual units to one or more groups. In group separation, linear functions of variables describe the differences between two or more groups. The main objective is to identify the relative contribution of p variables to split. The latter problem is focused on the prediction or allocation of observations to groups, which is a common goal of discriminant Approaches for Credit Rating and Corporate Bankruptcy Modelling 55 Micro-Modelling Approaches for Credit Rating and Corporate Survival analysis. A prediction rule then consists of a set of linear combinations of predictors, where the number of combinations reflects the number of groups. Discriminant functions are linear combinations of variables that best separate groups, for example, the k groups of multivariate observations. The description of discriminant analysis and methods can be found, for instance, in Rencher (2002), Manly (2005), Huberty and Olejnik (2006), Tabachnik and Fidell (2007), Harrell (2010) or Hair et al. (2014). As Rencher (2002) suggests, linear functions of variables (discriminant functions) describe the difference between two or more groups for group separation. The goal is to identify the relative contribution of the p variables to separation and derive discriminant functions as linear combinations of variables that best separate groups. Firstly, we assume the discriminant function for two groups (the following definitions and equations are taken from Rencher, 2002): • The two populations to be compared have the same covariance matrix but distinct mean vectors 1  and 2  , • we assume samples 1 11 12 1 , , , n y y y and 2 21 22 2 , , , n y y y from the two populations, • each vector ij y consists of measurement on p variables. The discriminant function is the linear combination of these p variables that maximizes the distance between the two (transformed) group vectors. Thus, a linear combination 'z=ay transforms each observation vector to a scalar: 1 1 1 1 1 2 1 2 1 1, ' ... , 1,2,..., i i i i p ip z a y a y a y i n= = + + + =ay , 1 2 1 2 1 2 2 2 2 2, ' ... , 1,2,..., i i i i p ip z a y a y a y i n= = + + + =ay . Then the 12 nn+ observations in two samples, 12 11 21 12 11 12 , nn yy yy yy are transformed into scalars, 12 11 21 12 11 12 . nn zz zz zz 56 Chapter 3 2024 Martina Novotná We find the means 1 1 1 1 1 1 n i i z z n = == ay and 22 z =ay , where 1 1 1 1 1 n i in = = yy and 2 2 2 2 1 n i in = = yy . Then, we find the vector a that maximises the standard difference 12 ( )/ z z z s− . Finally, we use the squared distance 22 12 ( ) / z z z s− so that the result is positive: ( ) 2 212 12 2pl () z zz s −  − = a y y a S a , (3.1) where Spl is the pooled covariance matrix and n1 + n2 – 2 > p. The maximum of (3.1) occurs when -1 pl 1 1 ( ),=−a S y y (3.2) or when a is any multiple of -1 pl 1 1 ()=−a S y y . The maximizing vector is not unique; however, its direction, or the relative values or ratios of 12 , , , p a a a are unique, and z =ay projects points y onto the line on which 22 12 ( ) / z z z s− is maximized. The linear discriminant analysis (LDA) can be used for more than two groups. Thus, we can extend the previous case for the study of several groups. The objective is to find linear combinations of variables that best separate the k groups of multivariate observations. We assume that for k groups with ni observations in the ith group, we transform each observation vector ij y to obtain ij ij z =ay , where i = 1, 2,..., k and j = 1, 2,..., ni. Then, we find the means ij i z =ay , where 1 i n ij i jn = = i yy . Similarly, to the two-group analysis, we seek the vector a that maximally separates 12 , , , k z z z . In this case, the formula (3.2) will be extended for k-groups. Assuming that ( ) ( ) 1 2 1 2  − = −a y y y y a , then ( ) ( ) ( )( ) 2 2 12 1 2 1 2 1 2 2pl pl z zz s  −   − − −  ==  a y y a y y y y a a S a a S a (3.3) In the case of k-groups, the separation criterion among 12 , , , k z z z can be expressed in terms of matrices, where the H matrix denotes ( )( ) 1 2 1 2  −−y y y y and the matrix E replaces Spl,   = a Ha a Ea , (3.4) Approaches for Credit Rating and Corporate Bankruptcy Modelling 57 Micro-Modelling Approaches for Credit Rating and Corporate Survival Matrix H has a between sum of squares on the diagonal for each of the p variables, and matrix E has a within sum of squares for each variable on the diagonal. Alternatively, the separation criterion can be expressed as SSZ( ) SSE( ) z z  = , (3.5) where SSH (z) and SSE (z) are the between and within sums of squares for z. The formula (3.4) can be rewritten as: ( ) , 0.    = −= a Ha a Ea a Ha Ea (3.6) Next, we examine values of  and a that are solutions of (3.6): 1 0, ( ) 0,   − −= −= Ha Ea E H I a (3.7) where I refers to the inversion matrix. The solutions are the eigenvalues 12 , , , s    and corresponding eigenvectors 12 , , , s a a a of 1− EH . From the s eigenvectors, we obtain s discriminant functions 1 1 2 2 , , , ss z z z    = = =a y a y a y which show the dimensions or directions of differences among 12 , , , . k y y y The ratio of its eigenvalue can calculate the relative importance of each discriminant function as a proportion to the total: 1 i s j j   =  (3.8) The coefficients in discriminant functions can be used to assess the contribution of the y’s to the separation of groups. As Rencher (2002, p. 283) suggests, this comparison is informative if the y’s are measured on the same scale and with comparable differences. For this reason, we use standardized discriminant functions (see, for example, Rencher (2000) for a more detailed description). 3.2.1 Tests of Significance To test hypotheses, we assume multivariate normality. The discriminant criterion (3.8) is maximized by 1  , the largest eigenvalue. The remaining eigenvalues 2,, s  are associated with other discriminant dimensions. The significance test is usually based on the Wilks’ lambda (Wilks’  ), and the eigenvalues are used in the test for significant differences among mean vectors. If the hypothesis H0 is rejected, we conclude that at least one 's  is significantly 58 Chapter 3 2024 Martina Novotná different from zero. Therefore, there is at least one dimension of separation of mean vectors. Each i  is gradually tested until a test fails to reject H0. The test statistic at the mth, step (m = 2, 3, …, s) is: 1 1 s mim i  = = +  , (3.9) which is distributed as 1, , 1p m k m N k m− + − − − +  . The statistic ( ) ( ) ( ) 11 1 ln 1 ln 1 22 s m m i im V N p k N p k  =     = − − − +  = − − + +         (3.10) has an approximate 2  - distribution with (p-m+1)(k-m) degrees of freedom. If more 's  are statistically significant, we may not consider the associated discriminant function if / ij j   is small, even if it is significant (Rencher, 2002). 3.2.2 Interpretation of Discriminant Functions The main purpose of interpretation is to assess the contribution of each variable. However, the signs of the coefficients are considered. According to Rencher (2002), there are three general approaches to assessing the contribution of each variable to separating the groups: • Standardized discriminant function coefficients, • partial F-test for each variable, • correlation between each variable and the discriminant function. Standardized coefficients are useful when the variables are measured on differing scales. In that case, coefficients are adjusted so that they apply to standardized variables. For example, for the observations in the first of two groups, 11 1 1 11 1 2 12 1 1 2 12 , ip p ii ip p yy y y y y z a a a s s s    − −− = + + + (3.11) where i =1, 2,…, n1. The standardized variables 11 / ir r r y y s− are scale-free, and the standardized coefficients r r r a s a = , where r = 1, 2,…, p. This standardization is applied to each of the s discriminant functions. The contribution of the variables to separating the groups is based on the absolute values of the coefficients. The stability of coefficients may vary from sample to sample; for example, if N/p is too small, one sample's important variables may emerge as less important in another sample. The second approach to assessing the contribution of each variable is based on a partial F-value. We can calculate a partial F-test for any variable yr and rank the variables. For example, in the case of two groups, the partial F-value is: Approaches for Credit Rating and Corporate Bankruptcy Modelling 59 Micro-Modelling Approaches for Credit Rating and Corporate Survival ( ) 22 1 21 1, pp p TT Fp T  − − − = − + + (3.12) where 2 p T is the two-sample Hotelling 2 T with all p variables, 21p T− is the 2 T - statistic with all variables except yr, and 12 2nn  = + − . The F-statistic is distributed as 1, 1p F  −+ . Conversely to standardized coefficients, the partial F-values are not associated with a single dimension of group separation. For example, y2 will have a different contribution in each of the s discriminant functions; however, the partial F for y2 creates an overall index of the contribution of y2 to group separation, considering all dimensions. Finally, the correlation between variables and discriminant functions can be used to assess each variable's contribution. These correlations are usually referred to as loadings or structure coefficients. Rencher (2002) points out that these correlations show each variable's contribution in a univariate context rather than in a multivariate one. In addition to LDA, which is one of the well-known methods, we use quadratic and logistic methods of discriminant analysis in the application part. • Quadratic discriminant analysis (QDA) is a variant of LDA that allows for the non-linear separation of data. • Logistic discriminant analysis (LogDA) is when the posterior probabilities are estimated by multi-nominal logistic regression (MLR). 3.2.3 Selection of Variables There are usually a large number of dependent variables available in discriminant analysis applications. For this reason, it is useful to select only some of the variables that will be finally considered for separating groups. For the selection of variables, we can use the following methods of discriminant analysis (Rencher, 2002): • Forward selection is when we begin with a single variable (the one that maximally separates groups). Then, the variable entered at each step is the one that maximises the partial F-statistic based on Wilks’s lambda. • Backward selection, when we begin with all the variables, and then at each step, the variable that contributes least is deleted according to the partial Fstatistic. • Stepwise selection is when we combine the forward and backward approaches. Variables are added one at a time, and at each step, they are reexamined. The procedure ends when the largest partial F among variables available for entry fails to exceed a pre-set threshold value. 60 Chapter 3 2024 Martina Novotná 3.2.4 Classification Analysis The attention in the previous text was paid primarily to a descriptive aspect of discriminant analysis. On the other hand, the discriminant analysis can also solve allocation and group membership prediction. In classification, a sampling unit with unknown group membership is assigned based on the vector of p measured values, y. As Rencher (2002) suggests, the one approach is to compare y with the mean vectors 12 , , , . k y y y of the k samples. Then, the unit is assigned to the group whose i y is closest to y. We can use a classification procedure suggested by Fisher (1936, cited in Rencher, 2002) in two populations. Using Fisher’s approach, we assume that the two populations have the same covariance matrix, 12  =  , normality is not required. Supposing that we get two samples from two populations, we can compute 12 ,yy and Spl. The classification is based on the discriminant function, ( ) 1 1 1 pl z−   = = −a y y y S y , (3.13) where y is the vector of measurements on a new sampling unit that we classify into one of the two groups. To determine the group membership, we compare z with the transformed mean 1 z or 2 z . For each observation 1i y from the first sample, we evaluate (3.13), obtain 1 11 12 1 , , , n z z z and get ( ) 11 1 1 1 1 1 2 pl 1 1/ n i i z z n − =  = = = − a y y y S y , similarly 22 z =ay . Assuming two groups are referred to as G1 and G2, y is assigned to G1 if z =ay is closer to 1 z than to 2 z according to the Fisher’s linear classification procedure, or to G2 if z is closer to 2 z . In the two-group case, the linear classification function is expressed as the discriminant function for group separating. However, the classification functions are different in the several-group case. Supposing classification of several groups, k, we find the sample mean vectors 12 , , , k y y y . We can use a distance function to assign a group membership for a vector y to find the mean vector that y is closest to and set y to the corresponding group. If we assume equal population covariance matrices, 12 k  = = = , then we obtain linear classification functions. A linear function ( ) i Ly can be expressed as: ( ) 0 1 1 2 2 0, i i i i i ip p i L c c y c y c y c  = + = + + + +y c y (3.14) Approaches for Credit Rating and Corporate Bankruptcy Modelling 67 Micro-Modelling Approaches for Credit Rating and Corporate Survival ( ) ( ) ( ) ( ) 2 0 Pr , j k g jg k e Yj e  = = = =  x x xx (3.34) where the vector 00=β and ( ) 00.g=x The variables are coded as follows: • If Y = 0, then Y0 = 1, Y1 = 0, Y2 = 0, • if Y = 1, then Y0 = 0, Y1 = 1, Y2 = 0, • if Y = 2, then Y0 = 0, Y1 = 0, Y2 = 1, and the sum of these variables is 2 01. j jY ==  Then the conditional likelihood function for a sample of n independent observations is: ( ) ( ) ( ) ( ) 0 1 2 0 1 2 1 , i i i ny y y i i i i l    =  =  β x x x (3.35) and the log-likelihood function can be expressed as: ( ) ( ) ( ) ( ) ( ) ( ) 12 1 1 2 2 1 ln 1 . ii ngg i i i i i L y g y g e e = = + − + + xx β x x (3.36) The likelihood equations can be found by taking the first partial derivatives of ( ) Lβ with respect to each of the 2(p + 1) unknown parameters (see for example Hosmer et al. (2013) for more details). Analogically to the binary model, the multivariable model is interpreted by odds ratios. For example, if we assume that the outcome variable 0Y= is the reference outcome, then the odds ratio of the outcome Yj= versus outcome 0Y= for covariate values of x = a versus x = b is: ( ) ( ) ( ) ( ) ( ) Pr Pr 0 ,. Pr Pr 0 j Y j x a Y x a OR a b Y j x b Y x b = = = = == = = = (3.37) The importance of the variable in the model is based on the likelihood ratio test. To test the significance of coefficients, we compare the log-likelihood from the fitted model containing the coefficients ( ) 1 L to the log-likelihood for the model containing only constant terms ( ) 0 L , one for each logit function. The test statistic can be expressed as:   01 2.G L L= −  − (3.38) 68 Chapter 3 2024 Martina Novotná The significance of the coefficients for a variable has degrees of freedom equal to the number of outcome categories minus one times the degrees of freedom for the variable in each logit. Concerning applying multinomial logistic regression to the bond rating prediction, our credit rating analysis to five categories will require four logit functions and determination of the baseline rating category, which is then compared with other logits. A general expression for the conditional probability in the five-category model is: ( ) ( ) ( ) ( ) 5 1 Pr , j k g jg k e Yj e  = = = =  x x xx (3.39) and we form four logits comparing Y = 1, Y = 2, Y = 3 and Y = 4 to it. The four logit functions are then denoted as: ( ) ( ) ( ) 1 10 11 1 12 2 1 1 Pr 1 ln , Pr 5 pp Y g x x x x Y      = = = + + + + =  =   xxβ x (3.40) ( ) ( ) ( ) 2 20 21 1 22 2 2 2 Pr 2 ln , Pr 5 pp Y g x x x x Y      = = = + + + + =  =   xxβ x (3.41) ( ) ( ) ( ) 3 30 31 1 32 2 3 3 Pr 3 ln , Pr 5 pp Y g x x x x Y      = = = + + + + =  =   xxβ x (3.42) ( ) ( ) ( ) 4 40 41 1 42 2 4 4 Pr 4 ln . Pr 5 pp Y g x x x x Y      = = = + + + + =  =   xxβ x (3.43) In addition to multinomial logistic regression analysis, we can also use an ordinal logistic regression approach. 3.3.3 Ordinal Logistic Regression Generally, multinomial logistic regression can be used if the outcome variables are not nominal but ordinal scale. However, it is suggested to use the ordinal logistic regression, which respects the categorical outcome's natural ranking in some cases. Menard (2010) describes ordinal variables as either crude measurement of a variable that could be measured on an interval or ratio scale or measurement of an abstract characteristic for which there is no natural metric or unit of measurement. Hosmer et al. (2013) argue that multinomial logistic regression could be used even in these cases. Still, it must be realised that not Approaches for Credit Rating and Corporate Bankruptcy Modelling 69 Micro-Modelling Approaches for Credit Rating and Corporate Survival considering the natural ordering of outcome variables may lead to the problem that estimated models might not address the analysis's questions. Thus, the ordinal logistic regression seems to be an appropriate method for assessing bond rating that considers the rank ordering of rating categories. As Menard (2010) says, different models make different assumptions about whether the dependent variable is intrinsically ordinal or reflects an underlying continuous interval or ratio variable. Various models can be proposed to analyse ordinal dependent variables, such as the cumulative logit model, the continuation ratio logit model, the adjacent categories logit model or the stereotype model. The cumulative logit model is the most widely used logistic regression model, and it is described in more detail in this chapter. According to Hosmer et al. (2013), the ordinal logistic model can be described using the following definitions. We assume that the ordinal outcome variable, Y, can take on K + 1 values coded 0,1,…, K. The probability that the outcome is equal to k conditional on a vector x of p covariates is denoted ( ) ( ) Pr k Yk  ==xx . If the model is assumed to be multinomial, then ( ) ( ) kk  =xx , where the model is given in equations (3.31 – 3.33) for K = 2. The multinomial model is called the baseline logit model in the case of ordinal logistic regression. When we assume an ordinal model, we must decide what outcomes to compare and the most reasonable model for the logit. There are three following main alternatives: • Compare each response to the next larger response (adjacent-category logistic model), • compare each response to all lower responses (continuation-ratio logistic model), • compare the probability of an equal or smaller response to the likelihood of a larger response (proportional odds model). As the proportional odds model will be used to model bond rating in the application part, we describe this approach's principles in the following text. Using the proportional odds model, we compare the probability of an equal or smaller response, Yk , to the likelihood of a larger response, Yk , ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) 01 12 Pr ln ln , Pr k kk k k K Yk cYk        ++   + + +  = = = −   + + +     xx x x xxβ x x x x (3.44) for k = 0, 1, …, K – 1. Supposing 1K= , the model is simplified to the usual logistic regression model in that it yields odds ratios of 0Y= versus 1Y= . The method used to fit the ordinal model is based on adapting the multinomial likelihood and log (3.36) for K = 2. The basic procedure involves the following steps (Hosmer et al., 2013, p. 292): 70 Chapter 3 2024 Martina Novotná a) The expressions defining the model-specific logits are used to create an equation defining ( ) k  x as a function of the unknown parameters. b) The values of a K + 1-dimensional multinomial outcome, ( ) 01 , , , k z z z =z are created from the ordinal outcome as 1 k z= if yk= and 0 k z= otherwise, where only one value of z equals 1. The general form of the likelihood for a sample of n independent observations ( ) , , 1,2, , , ii y i n=x is: ( ) ( ) ( ) ( ) 01 01 1 , i i Ki nz z z i i K i i l    =  =     β x x x (3.45) where β denotes both the p slope coefficients and the K model-specific intercept coefficients. Then, the log-likelihood function can be expressed as: ( ) ( ) ( ) ( ) 0 1 1 1 ln ln ln . n i o i i i Ki K i i L z z z    = = + + +              β x x x (3.46) When applying the method, it should be checked whether the data support the assumption of proportional odds. Tests for assessing the goodness of fit are based on the comparison of the model to an augmented model in which the coefficients for the model covariates are allowed to be different: ( ) ( ) ( ) Pr ln , Pr Yk k ckk Yk    = = −     β x xx x (3.47) where 1kk  +  for 1, , .kK= Norušis (2012, p. 70) explains a minus sign before the coefficients for the predictor variables, which is done so that larger coefficients indicate an association with larger scores. For a continuous variable, positive coefficients suggest that the likelihood of a larger score increases as the variable's values increase. Each logit has its k  term but the same coefficient, which means that the independent variable's effect is the same for different logit functions. The k  terms are called threshold values, and they are used in the calculations of predicted values. In the application section, the assigned rating category is considered as the outcome variable in the model. For instance, in our analysis (Chapter 4), the ratings are classified into five ordinal categories, ranging from the lowest (BB) to the highest (A) in terms of bond investment quality. Therefore, we account for five distinct rating grades, with the event of interest being the observation of a specific rating grade or lower. In this context, the ordinal model will evaluate a series of dichotomies between successive rating categories: • grade (1) versus grades (2, 3, 4 or 5), Approaches for Credit Rating and Corporate Bankruptcy Modelling 71 Micro-Modelling Approaches for Credit Rating and Corporate Survival • grades (1 or 2) versus grades (3, 4 or 5), • grades (1 or 2 or 3) versus grades (4 or 5), • grades (1 or 2 or 3 or 4) versus grade (5). Thus, in terms of bond rating groups, we will estimate the following odds based on the proportional odds model: ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) 1 2 3 4 Pr rating 1 , Pr rating greater than 1 Pr rating 1 or 2 , Pr rating greater than 2 Pr rating 1 or 2 or 3 , Pr rating greater than 3 Pr rating 1 or 2 or 3 or 4 . Pr rating greater than 4     = = = = (3.48) The highest category, 5, does not have associated odds because the probability of being in that category or any lower category is 1. Unlike the multinomial approach, ordinal logistic regression requires estimating only one set of regression coefficients. 3.3.4 ROC Analysis The estimated models represent a classification rule used to assign objects into classes. Thus, we need to know how effectively this classification rule works, preferably using the validation sample. It means that the available data are split into two datasets: an experimental sample (training) used for constructing the rule and the validation sample used for assessing the performance. There are other ways to split the data, for example, the leave-one-out method, when only one data point is put in the validation sample and others in the experimental sample, or the bootstrap methods. To understand how the performance is measured, we need to specify some terms required for further analysis. According to Krzanowski and Hand (2009), a classification rule yields a score s(X) for each object. It will result in distribution ( P)ps for objects in the positive group, P, and distribution ( N)ps for objects in the negative group, N. Then, the classifications are given by comparing the scores with a threshold, T. The authors claim that if we can find a threshold Tt= such that all members of class P have scores that are all greater than t , and all members of class N have scores all less or equal to t , we attain the perfect classification. However, the two sets of scores typically overlap to some extent, and perfect classification is impossible. In this case, performance is measured by the extent to which scores for objects in class P tend to take large values, and scores for objects in class N tend to take small values. The methods used for these measurements are 72 Chapter 3 2024 Martina Novotná based on the two-by-two classification table resulting from cross-classifying the true class of each object by its predicted class. Krzanowski and Hand (2009) state that the proportions of the validation set that fall in this table's cells are empirical realisations of joint probabilities ( , ), ( , ), ( , ), ( , )p s t P p s t N p s t P p s t N    . Then, different ways of summarising these four joint probabilities yield various measures of classification performance. Generally, we can use the misclassification or error rate measure, which is the probability of a class N object having a score greater than t or a class P object having a score less than t . The misclassification rate is a widely used criterion; however, it weights two kinds of classification (class N misclassified as P, and vice versa) equally important. We use the following two conditional probabilities and one marginal probability in the evaluation of classification ability (Krzanowski and Hand, 2009): • The false positive rate ()fp – the probability that an object from class N yields a score greater than : ( N)t p s t , • the true positive rate ()tp – the probability that an object from class P yields a score greater than : ( P)t p s t , • the marginal probability that an object belongs to class P: (P)p . Next, we use two complementary conditional rates and one complementary marginal probability (Krzanowski and Hand, 2009): • The true negative rate ( tn ), ( N)p s t - the proportion of class N objects which are correctly classified as class N, equal to 1fp− , • the false negative rate ( fn ), ( P)p s t - the proportion of class N objects which are correctly classified as class N, equal to 1fp− , • the marginal probability that an object belongs to class : (N) 1 (P).N p p=− The true positive rate is typically called the Sensitivity ( Se ), and the true negative rate is the Specificity ( Sp ). The rates described above are all conditional probabilities of having a particular predicted class given the true class. As Krzanowski and Hand (2009) point out, there are obvious relationships between the various conditional, marginal, and joint probabilities. For example, the misclassification rate e of a classification rule can be expressed as a weighted sum of the true positive and false positive rate: (1 ) (P) (N).e tp p fp p= −  +  (3.49) Since a classification rule's true positive and negative rates are complementary, they are typically used together as joint performance measures. Generally, the true positive rate increases as t decreases, while the true negative Approaches for Credit Rating and Corporate Bankruptcy Modelling 73 Micro-Modelling Approaches for Credit Rating and Corporate Survival rate decreases with a lower t . Thus, we can find the misclassification rate as the value of t which leads to the overall minimum of the weighted sum e in the formula (3.49). Krzanowski and Hand (2009) suggest another way to determine the threshold by choosing the maximum tp fp− , or 1tp tn+− (Sensitivity + Specifity – 1). The maximum value is called the Youden index (YI). Generally, the performance measures are based on comparing the distributions of the scores for the positive and negative populations. A good classification rule tends to produce high scores for the positive population and low scores for the negative population. The larger the extent to which these distributions differ, the better the classifier. The graphical depiction of both two distributions is presented by the ROC (Receiver Operating Characteristic) curve. The method based on the ROC curve is a commonly used way for assessing the performance of classification rules. First, the graph shows the true positive rate ( tp ) on the vertical axis and the false positive rate ( fp ) on the horizontal axis, as the classification threshold t varies. Then, the misclassification rate is the minimum distance between the curve and the upper left corner of the square containing the ROC plot. If we develop a classification rule for more than two classes, we face a more complex problem. For example, the rating models assign objects to several rating categories. In this case, we combine multiple ROC curves and use different approaches to assess performance. Krzanowski and Hand (2009) suggest treating the situation using two-class analyses. There are two main approaches to how the ROC analyses can be achieved: • Assuming k classes, we produce k different ROC curves by considering each class in turn as population P and the union of all other classes as population N, • we have all ( 1)kk− distinct pairwise-class ROC curves. Both approaches are suitable for summary statistics, such as the AUC (Area Under the Curve). In the case of perfect separation of P and N, AUC is the area under the upper borders of the ROC (the area of a square of side one, so the upper bound is 1). In random allocation, AUC is the area under the chance diagonal (the area of a triangle whose base and height are equal to 1, so the lower band is 0.5). Based on Krzanowski and Hand (2009), the AUC can be generally expressed as 1 0( ) .AUC y x dx= (3.50) The AUC can be defined as the average positive rate, taken uniformly over all possible false-positive rates in the range (0,1). A frequently used interpretation of AUC is that it is a probability that the classifier will allocate a higher score to a randomly chosen individual from population P than it will to a randomly and independently chosen individual from population N. 74 Chapter 3 2024 Martina Novotná 3.4 Survival Analysis The latter application study focuses on applying the most popular survival analysis techniques. It is suggested to use the regression models appropriate for survivor data to analyse time to event. As Hosmer et al. (2008, p. 3) claim, the most important differences between the outcome variables modelled via linear and logistic regression analyses and the time variable are that we may only partially observe the survival time. If the event's occurrence is unimportant, the event can be analysed as a binary outcome using the logistic regression model. As Harrell (2010) points out, survival analysis is used to analyse the data in which the time until the event is of interest. The input variable is the time until the event or duration time. The survival analysis allows the response to be incompletely determined for some subjects; perhaps we cannot follow all observations in the dataset. For example, some companies are still alive after the observation time or lost to follow-up. As we face incomplete information, we need to analyse the data using specialised survival techniques. The analysis involves a censoring mechanism when we define the censored and uncensored observations. For example, Hosmer et al. (2008, p. 18) explain a censored observation as one whose value is incomplete due to random factors for each subject. If we analyse data using the survival procedure, the dates of start and finish are not dealt with because they are different. Rather, we consider the length of time before the initial event, t = 0, and the terminal event or date of the last information about the object, t = 1. Usually, when an observation begins at the defined time t = 0 and terminates before the outcome of interest, it is assumed to be a censored observation. If no responses are censored, standard regression models for continuous responses could analyse the failure times (Harrell, 2010). Based on the distribution of failure times, we use parametric, semiparametric, and nonparametric methods. Survival analysis is the approach that allows working with incomplete data and modelling the time to an event, such as a corporate failure or default. Two time points must be clearly defined for time to event modelling: the beginning point and an endpoint when the event of interest occurs. In this context, survival time refers to the distance on the time scale between these two points (Hosmer et al., 2008). 3.4.1 Censoring When applying the survival analysis, we deal with censoring the data that comes from the fact that we can face the problem of incomplete observation of time. Two mechanisms can lead to incomplete observation in time: censoring and truncation. These two terms can be defined as follows: • A censored observation: the value of an observation is incomplete due to random factors for each subject. • A truncated observation: the value of an observation is incomplete due to a selection process inherent in the study design. Approaches for Credit Rating and Corporate Bankruptcy Modelling 75 Micro-Modelling Approaches for Credit Rating and Corporate Survival According to Hosmer et al. (2008), there are several types of incomplete observations: • Right censoring, • left censoring, • interval censoring, • left truncation, • right truncation. For survival analysis, we have to specify a point when observation ends on all subjects. Thus, subjects may enter the study at different times; however, they will have variable lengths of maximum follow-up time. For example, in Figure 3-1, we can see a hypothetical study of four subjects, the end of the study is September 2016. The bond issuer 1 entered the study in January 2015 and defaulted in February 2016. The bond issuer 2 joined the study in February 2015 and was lost to follow up in December 2015. The bond issuer 3 entered the study in May 2015, and there was no default until September 2016, the end of the study. Finally, bond issuer 4 joined the study in August 2015 and defaulted in April 2016. Figure 3–1 Line plot in calendar time (follow-up study) Source: Hosmer et al. (2008, p. 7), author For the practical reasons of survival analysis, we can assume that all subjects entered the study at the same calendar time and were followed until their respective endpoint. Thus, we must convert the collecting data from calendar time to analysis time (Figure 3-2). Calendar time Subject 01/15 02/15 05/15 12/15 02/16 09/16 1 2 3 4 08/15 04/16 76 Chapter 3 2024 Martina Novotná Figure 3–2 Line plot in the time scale (follow-up study) Source: Hosmer et al. (2008, p. 7), author The previous example is the case of the most common type of censoring, right censoring. The incomplete observations occur in the right tail of the time axis, usually when the observation begins at the defined time and terminates before the outcome of interest is observed. In some cases, left censoring can be used if the event of interest has already occurred when observation begins. If the time is not observable continuously, we can use interval censoring. It is a special type of failure data that include the right-censored failure time data but have a much more complex structure; for more details, you can see, for example, Sun and Li (2014). Since the left and right truncation are less common forms of incomplete data, they will not be considered in this text, and for practical reasons, the focus will be especially on right censoring. 3.4.2 Survival and Hazard Functions In the case of right censoring, two random variables need to be defined (Houwelingen and Stijnen, 2014): • The survival time (Tsurv), • the censoring time (Tcens), where the termination of the study usually determines the censoring time. The necessary condition for statistical analysis is that survival time and censoring time are independent. This condition can be weakened to the independence of survival time and censoring time conditional on the explanatory variables in the presence of explanatory variables. Then, we can define cumulative distribution functions for both random variables: ( ) Pr( ), surv surv F t T t= (3.51) Time in months Subject 1 2 3 4 810 13 16 Approaches for Credit Rating and Corporate Bankruptcy Modelling 83 Micro-Modelling Approaches for Credit Rating and Corporate Survival The regression coefficients can be estimated by the maximum likelihood method (see Gourieroux and Jasiak, 2007, p. 99; Hosmer et al., 2008). After fitting the model, the significance of the model and the formation of a confidence interval for key estimated parameters should follow. The latter mentioned authors suggest three related tests to assess the significance of the coefficient: • The partial likelihood ratio test, • the Wald test, • the score test. The partial likelihood test is based on the following statistic:   ˆ 2 ( ) (0) , pp G L L  =− (3.74) where p L refers to the log partial likelihood of the model containing the covariate and 0 L is the log partial likelihood for the model not containing the covariate. The log partial likelihood evaluated at 0  = is: 1 (0) ln( ), m pi i Ln = =−  (3.75) where i n denotes the number of subjects in the risk set at observed survival time i t . Under the null hypothesis that the coefficient is equal to zero, this statistic will follow a 2  -distribution with 1 degree of freedom and thus can be used to obtain p-values to test the significance of a coefficient (Hosmer et al., 2008). The Wald statistic test is based on the ratio of the estimated coefficient to its estimated standard error, assuming that the statistic follows a standard normal distribution. Unlike the linear regression, the Wald and log partial likelihood ratio test are not numerically related. The Wald statistic is given by ˆ. ˆ () zSE   = (3.76) The third approach, the score test, is based on the ratio of the derivative of the log partial likelihood to the square root of the observed information all evaluated at 0  = (see Hosmer, et al., 2008). Cleves et al. (2010) use the term relative hazard for x e  , and the log relative hazard, or risk score, for x  . To verify the model's specification x  and adequate parameterisation, we can use tests called tests of the proportional-hazard assumptions (P-H assumptions). In the application, the tests will be based on the analysis of residuals. As to the fact that the proportional hazards model to censored survival data is fit using the partial likelihood, the calculation of residuals differs from the usual regression models. For this reason, various approaches have been 84 Chapter 3 2024 Martina Novotná developed for Cox proportional model. The residuals used in the empirical study will be based on Schoenfeld residuals. For more details, see, for example, Hosmer et al. (2008), Cleves et al. (2010), Harrel (2010). Since the survival models estimate the time to event, the explained variation should also be assessed after fitting the model. The measures of explained variation for use with censored survival data differ from the traditional concept of variation using the index of determination. In our case, we apply the measure proposed by Royston (2006) with the character of explained variation in proportional hazards models, which can be used as an adjusted index of determination in PH models. 3.4.5 Parametric Models While nonparametric analysis is a useful tool for describing our survivor data, semiparametric models are used especially for the estimation of hazard ratios and their further explanations. In many cases, semiparametric analysis based on the Cox model can sufficiently analyse our data. However, if we aim to predict the time to failure, some parametric assumption is necessary. Parametric models generally provide smooth estimates of the hazard and survival functions and enable us to model also a nonproportional effect on the hazard scale (Royston and Lambert, 2011). Compared to semiparametric models, parametric models are used when the distribution of survival time has a known parametric form. In this case, the fully parametric model enables us to better analyze survival data. Cleves et al. (2010) describe six standard parametric survival models: exponential, Weibull, Gompertz, lognormal, and generalized gamma. According to Hosmer et al. (2008), using these models may have the following advantages: • Full maximum likelihood may be used to estimate the parameters, • the estimated coefficients or their transformations can provide clinically meaningful estimates of effect, • fitted values from the model can provide estimates of survival time, • residuals can be computed as differences between observed and predicted values of the time. In parametric models, we assume that the distribution of time to event (T) can be described as a function of a single covariate: 01 . x Te   + = (3.77) Since the time must always be positive, the equation (3.77) can be expressed as the product of a positive systematic component, 01 exp( )x  + , and an error component,  , that also takes only positive values. a) Exponential Regression Model The exponential survival model is the simplest parametric model with a constant hazard function. We denote the exponential distribution with survival function Approaches for Credit Rating and Corporate Bankruptcy Modelling 85 Micro-Modelling Approaches for Credit Rating and Corporate Survival ( ) exp( )S t t=− as (1).E The model can be expressed by taking the natural log of each side of the equation (3.77): * 01 ln( ) ,Tx    = + + (3.78) where *ln( ).  = If the error component  follows the exponential distribution, then the error component *  follows the extreme minimum value distribution denoted as (0,1)G . The model in (3.76) is called the exponential regression model. If this model is generalized by allowing the shape parameter to be different from 1 by using (0, )G  distribution: * 01 ln( ) ,Tx     = + +  (3.79) we get the Weibull regression model (3.79). Survival time models that are linearized by taking logs are called accelerated failure time models. The covariate effect in these models is multiplicative on the time scale, as we can see in (3.77). In other words, the impact of the covariate is said to accelerate survival time. To describe the method of the exponential regression model, firstly, we assume the single covariate model (3.77), where the error distribution is log-exponential. Then, the survival function for the model can be expressed as: 01 ( , , ) exp( 1 ). x S t x e  + =−β (3.80) If we set the right-hand side of this equation equal to 0.5 and solve the resulting equation, we get an equation for the covariate specific median survival time of 01 50 ( , ) ln(0.5). x t x e  + = − β (3.81) Assuming the dichotomous covariate in (3.79) coded 0 or 1, then the ratio of the median survival time for the group with 1x= to the group with 0x= is: 01 1 0 50 50 ( 1, ) ln(0.5) TR( 1, 0) , ( 0, ) ln(0.5) tx e x x e tx e    + =− = = = = = =− β β (3.82) where TR denotes time ratio. The relationship between the two median times can be written as: 1 50 50 ( 1, ) ( 0, ).t x e t x  = = =ββ (3.83) For example, if 1 exp( ) 2  = , then the median survival time in the group with 1x= is twice the median survival time in the group with 0x= . The quantity 1 exp( )  is usually called the acceleration factor, although it can accelerate or decelerate survival time. 86 Chapter 3 2024 Martina Novotná The multiplicative covariate effect can be presented using the following form of survival function, 1 ( , 1, ) ( , 0, ).S t x S te x  − = = =ββ (3.84) The equation shows that the value of the survival function at time t for the group with 1x= can be obtained by evaluating the survival function for the group with 1x= at time 1 exp( ).t  − In addition to the survival function, the model in (3.80) can be expressed in terms of the hazard function as follows, 01 ( , , ) , x h t x e  + =β (3.85) that is constant over time because it depends only on model coefficients and covariate values. Thus, on the one hand, the hazard function is relatively simple. However, it may be too simple to provide a realistic description of the survival data. The hazard ratio for a dichotomous covariate is 1 HR( 1, 0) .x x e  − = = = (3.86) The model can be estimated using the maximum likelihood method. The fitted values are predictions of values from a censored exponential distribution. The estimator of variances and covariances of the estimator of the coefficients are obtained using the second partial derivative of the log-likelihood function. The influence of individual subjects on the values of the estimated parameters is based on the score residuals (see Hosmer et al., 2008). In parametric models, the assumption of proportional hazards is replaced by the procedure that determines whether the data support the particular parametric form of the hazard function. Hosmer et al. (2008) suggest using the model-based estimate of the cumulative hazard function to form the Cox-Snell residuals. The estimated cumulative hazard function values can be considered observations from a censored sample from an exponential distribution with a parameter equal to one. Then, the model diagnosis is based on the plot that compares the model-based cumulative hazard to the hazard obtained from a nonparametric estimator (KaplanMeier, Nelson-Aalen). As the authors say, the nonparametric estimator uses the model-based estimates of the cumulative hazard at each observed time as the time variable and the censoring indicator from the survival time variable as the censoring variable. Thus, this plot should follow a line through the origin with a slope equal to one if the parametric model is correct. The estimator of Cox-Snell residuals can be obtained by exponentiating the additive residuals on the log time scale (see Hosmer et al., 2008, p. 257 for more details). Similar to Cox proportional hazards model, the significance of variables can be assessed using the score test, likelihood ratio or Wald test. Approaches for Credit Rating and Corporate Bankruptcy Modelling 87 Micro-Modelling Approaches for Credit Rating and Corporate Survival b) Weibull Regression Model We consider the Weibull distribution a natural generalization of the exponential (Roysten and Lambert, 2011). The Weibull regression model can be expressed using the natural log as in equation (3.80). Compared to the exponential model, the Weibull model allows the shape parameter  to be different from 1 by using (0, )G  distribution. The hazard function for the single covariate model is 01 1 () ( , , , ) , x t h t x e      − + =β (3.87) and we assume that 1.  = The proportional hazards form of the function can be expressed as 0 1 0 1 () 11 ( , , , ) , or xx h t x t e t e e          − + − − −− ==β (3.88) 011 10 ( , , , ) ( ) , xx h t x t e e h t e       −− − ==β (3.89) where 0 0 1 1 exp( ) exp( ),        = − = = − and the baseline hazard function is 1 0( ) ,h t t   − = (3.90) and 1  = is usually called the shape parameter. Although the parameter  is a variance-like parameter on the log-scale, we can refer to  as the shape parameter in this text, as suggested by Hosmer et al. (2008, p. 261). The parameter  is called the scale parameter. For example, the expression in (3.80) leads to a hazard ratio interpretation of the parameter 1.  The Weibull distribution can provide variety of shapes of the hazard function determined by the estimated parameter  . When 1  = , the hazard is constant, and the Weibull model reduces to the exponential model. When 1   , the hazard in monotone decreasing, and when 1   , it is monotone increasing. Thus, the Weibull model is suitable for modelling data that exhibit monotone hazard rates (Cleves et al., 2010). The accelerated failure-time form of the hazard function can be expressed as 01 11 () 11 ( , , , ) ( ) . xxx h t x t e te e         −+ −− −− ==β (3.91) The relationship between the two sets of estimated coefficients using the proportional hazards form and the accelerated failure-time form of the hazard function is .θ = -β σ The survival function that corresponds to the accelerated failure-time form of the hazard function in (3.92) is 88 Chapter 3 2024 Martina Novotná     01 ( , , , ) exp exp ( 1 )( ) . x S t x t      = − − +β (3.92) Then, the median survival time can be obtained by setting the survival function equal to 0.5 and solving for time,   01 50 ( , , ) ln(0.5) . x S x e   + =−β (3.93) If the covariate is dichotomous and coded 0/1, then the time ratio at the median survival time is (Hosmer et al., 2008, p. 262):     01 1 0 50 50 ln(0.5) ( 1, , ) TR( 1, 0) ( 0, , ) ln(0.5) e tx x x e tx e      + − = = = = = = =− β β . (3.94) The interpretation of the β form of the coefficients is the same as in the exponential regression model. The assessment of model fit is based on the score residuals, similarly to the exponential regression model. See, for example, Hosmer et al. (2008) for a detailed description. c) Flexible Parametric Models Royston and Lambert (2011) suggest that simpler parametric models may not be flexible enough to represent the hazard function adequately and thus fit our data well. For example, the hazard function of a Weibull model always goes in the same direction with time. Therefore, the authors propose new parametric models that include flexible PH models, flexible proportional odds (PO), and probit-scale models. Thus, we get alternative models which extend the range of survival distributions that can be estimated. Furthermore, these models allow nonlinearity functions and thus increase their practical use and applications. Royston and Lambert (2011) propose Royston-Parmar (RP) models, which have considerably greater flexibility concerning the shapes of the survival distributions they can model. In this case, the baseline distribution function is a restricted cubic spline function of log time instead of simply as a linear function of log time. The complexity of models with spline functions is determined by the number and the positions of the connection points in log time (knots) of the spline’s cubic polynomial segments. The parameters of models are estimated based on maximum likelihood; for more details and description of models, see, for example, Royston and Lambert (2011). The authors consider RP models as an extension of the Weibull, loglogistic, and lognormal models. While we assume that the effect of covariates is proportional on the appropriate scale (hazard, odds of failure, or probit of failure probability) in these models, the assumption of linearity is relaxed in RP models. The generalisation of the Weibull model using spline functions is described by the authors as follows. Firstly, we express the Weibull cumulative hazard function in logarithmic form: 1 0 1 ln ( ) ln ln lnH t t t     = + = + , (3.95) Approaches for Credit Rating and Corporate Bankruptcy Modelling 89 Micro-Modelling Approaches for Credit Rating and Corporate Survival where ln ( )Ht is a sum of two components: a constant ( 0  ) and a linear function of log time ( 1lnt  ). In the case when the latter component does not correctly capture the shape of the (log) cumulative hazard function, we might need a more flexible model: ln ( ) ( ; )H t f t  = , (3.96) where ( ; )ft  represents some general family of a nonlinear function of time t , having some parameter vector  . Royston and Lambert (2011) suggest fractional polynomials and splines as suitable functions. For example, they describe a restricted cubic spline function as (ln ; )st  with s standing for splines and lnt to show that we use the scale of log time: 0 1 2 1 3 2 ln ( ) (ln ; ) ln (ln ) (ln ) ...,H t s t t z t z t      = = + + + + (3.97) where 12 ln , (ln ), (ln )t z t z t and so on are the basis functions of the restricted cubic spline. Thus, if there are no knots, then 0 1 0 1 (ln ; , ) lns t t     =+ , which is the Weibull model. The parameters  are estimated by maximum likelihood, as proposed by Lambert and Royston (2009). In practical application and estimation of models, the chosen number of interior knots specifies the degrees of freedom (one plus the number of knots). So then, the PH(d) model is a PH model whose spline function has d degrees of freedom (d – 1 interior knots and when d > 1, two boundary knots). The fit of estimated parametric models can be compared based on the Akaike information criterion (AIC), defined as the deviance plus 2k , where k is the dimension of the model (the number of fitted parameters). Alternatively, we can use the Bayes information criterion (BIC), which is the deviance penalized by adding logkn , where n is the sample size. Because parametric models are estimated by the maximum likelihood method, both criteria, AIC and BIC, can be used to compare fitted models. However, since the Cox model is estimated by the maximum partial likelihood method, this model cannot be compared with parametric models based on AIC and BIC criteria (Royston and Lambert, 2011). 3.4.6 Multiple Failure-Time Data Cleves (2000) describes multiple failure-time data as data when two or more events (failures) occur for the same subject or from identical events occurring to related subjects. The typical feature is that failure times are correlated within a cluster (subject or group), violating the independence of failure times assumption required in traditional survival analysis. As the author points out, failure events should be classified according to whether they have a natural order and recurrences of the same type of events. The events are supposed to be ordered when the second 90 Chapter 3 2024 Martina Novotná event cannot occur before the first event. On the contrary, unordered events can happen in any sequence. There are more approaches to examining multiple failure-time data. Firstly, we can examine the time to the first event, ignoring additional failures. However, it means we do not use all available data. The second method is based on the available data analysis while accounting for the lack of independence of the failure times. Cleves (2000) suggests corresponding procedures for estimating these models using the Cox proportional hazard model. Under the proportional hazard assumption, the hazard function (3.72) of the ith cluster for the kth failure type is as follows: , 0 ( , ) ( ) , i Z k ki h t Z h t e  = (3.98) where ki Z is a p-vector of possibly time-dependent covariates for ith cluster to the kth failure type. While we presume in equation (3.98) that the baseline hazard function is equal for every failure type, the baseline hazard function is allowed to differ by failure type in the following formula: , 0 ( , ) ( ) . i Z k ki k h t Z h t e  = (3.99) As Cleves (2000) suggests, the maximum likelihood estimates for the models (3.98) and (3.99) are obtained from Cox's partial likelihood function ()L  , assuming independence of failure times. Concerning the analysis of multiple failure-time data, Cleves (2000) emphasizes the need to determine whether it is ordered or unordered data and select a suitable method for estimating models accordingly. In the case of unordered times, which is the case for rating analysis, it is first necessary to determine whether the events are of the same or different types. Similarly, deciding whether the baseline hazard is the same or different for all event failures is necessary. In any case, it is essential to implement the methods for correctly structuring the data, including identifying individual failure events. 3.5 Chapter Summary The purpose of this chapter was to clarify the main motives for conducting an application study in this work. Therefore, the basic characteristics and meaning of micro approaches for measuring credit risk were provided at the beginning of this section. Subsequently, attention was paid to the research review, based on which it was determined which procedures are suitable for modelling individual credit risk and which results the selected authors reached. Hence, the motivation and the main goal were specified, and the main contribution to the current research was outlined. The next part of the chapter was devoted to describing the econometric models used in the application part of this work, namely discriminant, regression and survival analysis. Since plenty of professional publications deal with these Approaches for Credit Rating and Corporate Bankruptcy Modelling 91 Micro-Modelling Approaches for Credit Rating and Corporate Survival approaches in great detail and professionally, the selected methods were only briefly described in this section. The main attention was paid to understanding the main principles and the possible use of models. Finally, the methods described in this chapter are applied in the following four sections. The first study aims at modelling the influence of selected factors on ratings and their downgrades. Next, we find out whether there is a relationship between rating and corporate bankruptcy rates. Then, we analyse the effect of selected corporate characteristics on the survival probability. Finally, we formulate parametric survival models based on the previous results and estimate bankruptcy rates and ratings. The Effect of Selected Factors on Rating and its Dynamics 99 Micro-Modelling Approaches for Credit Rating and Corporate Survival Table 4–6 Standardized canonical discriminant function coefficients (model 1) Variable Function 1 2 3 4 roa 0.8919 1.2621 0.8018 0.2660 roe -0.3216 -0.4274 -0.0586 0.0642 eqta 0.7286 -0.5229 0.3506 0.2142 lnta -0.3680 1.6715 1.7264 0.7928 lnintcov 0.4202 -0.0856 -0.6401 -0.2755 lncf 0.3627 -1.6822 -2.1066 -0.8115 lnliqr 0.2756 -0.0771 0.5208 -0.7117 lncurr 0.0955 -0.0001 -0.7138 1.0340 lnltdta -0.0644 -0.0112 0.0435 0.1917 ebitdar 0.0371 0.1531 -0.5950 0.3354 b) Classification ability The coefficients of classification functions (Fisher’s linear discriminant functions) of the estimated models are shown in Table 4-7. The classification functions are used to classify individual cases: First, the values of five functions are computed and compared. Then, the group corresponding to the function with the highest value is selected as the target rating category. Table 4–7 Classification function coefficients (model 1) Variable 1 2 3 4 5 roa 1.2600 1.4491 1.6511 2.0131 2.7745 roe 0.1019 0.0861 0.0694 0.0346 -0.0262 eqta 9.2790 22.2835 37.4865 51.4774 62.9647 lnta 34.2175 33.1949 31.7998 30.5193 32.5883 lnintcov 0.3915 1.0540 1.9416 3.3279 4.1108 lncf -26.6527 -25.8531 -24.6153 -23.3636 -25.4225 lnliqr -4.5765 -2.4185 -1.2134 0.0995 1.8487 lncurr -0.3275 -1.6194 -0.9198 0.0149 0.0925 lnltdta 0.4275 0.3187 0.2539 0.1220 0.0150 ebitdar -0.3804 -0.4616 -0.4653 -0.4115 -0.3934 constant -77.2415 -73.9966 -77.9575 -93.3122 -126.6945 For example, using the mean, minimum and maximum values of input variables representing three hypothetical companies, the classification functions assign them to different rating groups (Table 4-8). The average company is given to the middle group 3 – BBB. This result is not surprising because, as mentioned above, most companies have the middle rating assessment. Since the groups are not equally sized, the classification functions are weighted more heavily to classify group three in our case. The hypothetical company with the minimum (maximum) values is classified as group 1 – B (5 – AA). This result is also not surprising, as 100 Chapter 4 2024 Martina Novotná the higher value of most variables is generally associated with a better company's financial situation. Table 4–8 Example of classification 1 2 3 4 5 Mean 65.48 74.35 78.48 74.84 60.93 Minimum 33.05 28.52 13.21 -10.44 -31.02 Maximum 121.33 122.26 143.74 175.39 191.59 The main characteristics of the classification ability of models are summarised in Table 4-9 . The criteria for comparison and ranking the models are classification tables (confusion matrices) and error rates: • The resubstitution classification table, obtained by classifying the observations used to build the discriminant model, is Class. (ES). • The classification table, based on the hold-out sample's estimation ability, is called Class. (hold). • The error rate, which represents the overall error rate and the error rate for each group and corresponds to the classification table, is based on the count-based estimate. Table 4–9 Percentage correctly classified (PCC) – 5cat models Method Model Experimental sample (ES) Number of obs. (ES) Class. (ES) Class. (hold) LDA 1 non-random 3518 0.8511 0.8775 LDA 2 random 3274 0.8525 0.8600 Although the classification ability of models is similar (Table 4-9), it differs across the rating groups (Table 4-10). For example, while only 8.35% of subjects from rating 3 are misclassified, it is 47.01% from rating 1. The least accurate is, therefore, the classification into rating 1, i.e. category B. From the lender's point of view, the model tends to reduce the classification ability of the companies with the worst ratings. The results of the second model are proportionally similar. Table 4–10 Error rate (5 cat) Model Misclas. Rating Total 1 2 3 4 5 1 Error rate 0.4701 0.1978 0.0835 0.1816 0.2123 0.1489 Priors 0.3326 0.2041 0.4801 0.2317 0.0509 2 Error rate 0.4205 0.2750 0.0677 0.1745 0.2473 0.1476 Priors 0.0269 0.1677 0.5100 0.2398 0.0559 According to the overall error rate criterion, both five-category discriminant models achieve a similar classification ability. For example, the total error rate of model 1 shows that the proportion of misclassified observations is 14.89%. The Effect of Selected Factors on Rating and its Dynamics 101 Micro-Modelling Approaches for Credit Rating and Corporate Survival Overall, the non-random LDA model has a better classification ability on the hold-out sample. However, the general results of both models do not allow a clear choice of a more suitable model. Furthermore, the classification accuracy of these models is strongly influenced by the fact that the boundary categories are not evenly represented in the sample. Therefore, we eliminate this shortcoming by deriving models for only three internal rating categories: 2, 3, and 4 (Table 4-11). Table 4–11 Percentage correctly classified (PCC) – 3cat models Method Model Experimental sample (ES) Number of obs. (ES) Class. (ES) Class. (hold) LDA 3 non-random 3222 0.8858 0.8986 LDA 4 random 3274 0.8878 0.8882 Unsurprisingly, both the classification ability and the error rate provide better results. The detailed error rates are summarized in Table 4-12. Table 4–12 Error rate (3 cat) Model Misclas. Rating Total 2 3 4 3 Error rate 0.1671 0.0787 0.1411 0.1142 Priors 0.2228 0.5242 0.2529 4 Error rate 0.2332 0.0641 0.1300 0.1122 Priors 0.1828 0.5560 0.2613 In this case, the classification ability of the 3-category models is similar. However, model 3 (non-random) achieves a lower misclassification at the lowest rating and higher classification accuracy on the hold-out sample. For this reason, this model can be considered more suitable for rating prediction. 4.1.4 Multinomial Logistic Models Since we use two logistic regression analysis methods, we estimate eight models (Appendix 3). The models are summarized in Table 4-13. Table 4–13 Overview of logistic models Model Method No. of predictors No. of categories Sample Model 5 MLR 10 5 Non-random Model 6 MLR 10 5 Random Model 7 OLR 10 5 Non-random Model 8 OLR 10 5 Random Model 9 MLR 10 3 Non-random Model 10 MLR 10 3 Random Model 11 OLR 10 3 Non-random Model 12 OLR 10 3 Random 102 Chapter 4 2024 Martina Novotná a) Main Results and Interpretation The overall results based on the logistic regression suggest that the estimated models do not differ significantly. Therefore, similarly to discriminant analysis, only one model will be described in more detail in the following text. The interpretation for other models is analogous. The following text focuses on model 5 (MLR, non-random, five categories). However, we will pay attention to both approaches because we use two methods. In addition, some tables from Appendix 3 are rewritten in the next sections for better clarification. First, four logit functions are estimated because it is a fivecategory model. We arbitrarily use the middle rating category 3 (BBB) as a reference value. Thus, we form four logits, 1 2 4 5 ( ), ( ), ( ), ( ),g x g x g x g x comparing each group to the reference category. Equations (3.40) – (3.43) are used to estimate the unknown parameters based on the maximum likelihood method to fit the model. The parameter estimates are summarized in Table 4-14, and the details are provided in Appendix 3. The assessment of parameter estimates and their significance is based on the Wald test used to test the null hypothesis that each of the individual coefficients is zero for each logit. Table 4–14 MLR parameter estimates (model 5) Variable Rating 1 2 4 5 roa -0.8719* -0.3489* 0.4810* 0.8939* roe 0.0288* 0.0253* -0.1126* -0.2060* eqta -50.6921* -25.1934* 7.2455* 1.7041 lnta 2.3379* 1.9669* -1.1749* -1.0128 lnintcov -1.8130* -1.3457* 2.2807* 4.3186* lncf -.19509* -1.8375* 1.1257* 1.3755 lnliqr -4.5343* -2.4685* 1.4070* 2.7621* lncurr -3.6769* -2.6221* 0.9934* 0.4867 lnltdta 0.3388* 0.1937* -0.1992* -0.3323* ebitdar -0.0629* -0.0588* -0.9185* -3.4539* constant 12.2038* 9.1524* -12.0212* -26.3962* *significant at .05 level The parameter estimates compare pairs of outcome variables (note that Rating 3 is a reference category). Thus, for example, the first part of the table labelled Rating 1 compares this category against Rating 3, Rating 2 compares this category against Rating 3 and so on. Therefore, the interpretation is similar to the binary logistic regression. The coefficients in Table 4-14 are expressed in terms of the log odds. For example, the coefficient -0.8719 (Rating 1, roa) implies that a oneunit change in roa results in a -0.8719 unit change in the log of odds. However, we typically prefer using the odds ratios for interpretation, computed by exponentiating the coefficients (StataCorp, 2021; UCLA, 2021a; UCLA, 2021b). The odds ratios of model 5 are summarized in Table 4-15. The Effect of Selected Factors on Rating and its Dynamics 103 Micro-Modelling Approaches for Credit Rating and Corporate Survival Table 4–15 MLR odds ratios (model 5) Variable Rating 1 2 4 5 roa 0.4182 0.7055 1.6177 2.4446 roe 1.0292 1.0256 0.8935 0.8138 eqta 0.0000 0.0000 1401.7826 5.4964 lnta 10.3595 7.1485 0.3088 0.3632 lnintcov 0.1632 0.2604 9.7835 75.0834 lncf 0.8228 0.1592 3.0824 3.9571 lnliqr 0.0107 0.0847 4.0837 15.8331 lncurr 0.0253 0.0727 2.7004 1.6269 lnltdta 1.4033 1.2137 0.8194 0.7173 ebitdar 0.9390 0.9429 0.3991 0.0316 The multinomial logistic model is a simple extension of the binary model. However, the interpretation is more difficult because of relevant binary comparisons. For example, with five outcomes (rating groups), we have ten binary comparisons for one model: R1 versus R2, R1 versus R3, R1 versus R4, R1 versus R5, R2 versus R3, R2 versus R4, R2 versus R5, R3 versus R4, R3 versus R5, and R4 versusR5. To demonstrate the interpretation of odds ratios, we compare Rating 1 (B) against the reference category, Rating 3 (BBB): • The odds ratio of roa is 0.4182, meaning that as roa increases, the company is likely to get a rating 3, assuming that all other variables in the model are held constant. Similarly, for eqta, lnintcov, lncf, lnliqr, lncurr and ebitdar (their odds ratios are less than one). • On the other hand, the odds ratio of roe is 1.0292. Thus, as roe increases, the company is likely to get rating 1 relatively to rating 3, similarly, for lnta and lnltdta (their odds ratios are more than one). Based on the estimated coefficients, the four logit functions can be written as: 1 2 0.87 0.03 50.69 2.34 1.81 0.2 4.53 3.68 0.34 0.06 12.2, 0.35 0.03 25.19 1.97 1.35 1.84 2.47 2.62 0.19 g roa roe eqta lnta lnintcov lncf lnliqr lncurr lnltdta ebitdar g roa roe eqta lnta lnintcov lncf lnliqr lncurr = − + − + − − − − − + − + = − + − + − − − − − + 0.06 9.15,lnltdta ebitdar−+ 4 5 0.48 0.11 7.25 1.17 2.28 1.26 1.41 0.99 0.2 0.92 12.02, 0.89 0.21 1.7 1.01 4.32 1.38 2.76 0.49 0.33 g roa roe eqta lnta lnintcov lncf lnliqr lncurr lnltdta ebitdar g roa roe eqta lnta lnintcov lncf lnliqr lncurr lnlt = − + − + + + + + − − − = − + − + + + + + − 3.45 26.4.dta ebitdar−− 104 Chapter 4 2024 Martina Novotná The conditional probability for each rating category can be expressed as: ( ) ( ) ( ) ( ) ( ) 1 5 1 2 4 2 5 1 2 4 5 1 2 4 4 5 1 2 4 5 1 2 4 () 1() ( ) ( ) ( ) () 2() ( ) ( ) ( ) 3() ( ) ( ) ( ) () 4() ( ) ( ) ( ) () 5( ) ( ) ( 1, 1 2, 11 3, 1 4, 1 51 g g g g g g g g g g g g g g g g g g g g g g g e Ye e e e e Ye e e e Ye e e e e Ye e e e e Ye e e      == + + + + == + + + + == + + + + == + + + + == + + + x x xxx x x xxx x xxx x x xxx x xx x x x x x5() ). g e+x x b) Model fitting The overview and main characteristics of estimated models are summarized in Table 4-16, including: • Deviance (D=-2LLM): The value of a likelihood-ratio chi-squared for the test of the null hypothesis that all the coefficients associated with independent variables are simultaneously equal to zero (including degrees of freedom as the number of constrained parameters), • pseudo R2 measured as McFadden’s, • AIC and BIC. Likelihood ratio tests for the overall models test the null hypothesis that all coefficients in the model are zero. Firstly, we calculate the value -2 log-likelihood for models with only an intercept term and all variables. Then, we get the difference between these values, Chi-square. If the observed significance is small, we can reject the null hypothesis that all coefficients are zero and conclude that the final model is significantly better than the intercept-only. According to the values in Table 4-16, we conclude that all models are better than the intercept-only models. Table 4–16 Model-fitting Model Deviance LR Chi2(df) Pseudo R2 AIC BIC Model 5 2176.797 6830.33 (40) 0.7583 2264.794 2536.082 Model 6 2004.445 6135.13 (40) 0.7537 2092.445 2360.571 Model 7 2683.156 6323.97 (10) 0.7021 2711.156 2797.475 Model 8 2450.653 5688.92 (10) 0.6989 2478.653 2563.965 Model 9 1509.218 5068.91 (20) 0.7706 1553.218 1686.928 Model 10 1347.236 4586.88 (20) 0.7730 1391.236 1523.406 Model 11 1812.961 4765.17 (10) 0.7244 1836.961 1909.894 Model 12 1602.200 4331.91 (10) 0.7300 1626.200 1698.292 The Effect of Selected Factors on Rating and its Dynamics 105 Micro-Modelling Approaches for Credit Rating and Corporate Survival Next, we consider the information criteria AIC and BIC (for the description of the measure, see, for example, Long and Freese, 2014). Based on the criterion BIC, we prefer model 6 among 5-cat models and model 10 among 3-cat models (they have the smallest value of BIC). These results are also suggested by the criterion AIC. c) Classification ability Similar to the discriminant analysis, we use the model to determine the rating of a hypothetical firm with average, minimum and maximum values of input variables. However, unlike the discriminant analysis, we calculate the probability of belonging to a group and classify it into the group with the highest probability (Table 4-17). Table 4–17 Example of classification 1 2 3 4 5 Mean 4.89E-05 9.98E04 9.94E-01 4.89E03 7.37E-10 Min 1.00E+00 1.81E10 2.02E-24 2.63E24 1.98E-23 Max 5.56E-45 2.42E26 1.00E+00 4.94E114 0.00E+00 The average company is assigned to the middle rating 3 – BBB. The hypothetical company with the minimum (maximum) values is rated 1 – B (3 – BBB). Compared with the discriminant model (Table 4-8), the logistic model predicts a different rating group using the minimum values of predictors. Next, we examine the classification accuracy of estimation and hold-out samples (Table 4-18). All models achieve a relatively high overall classification accuracy, and their differences are small. For example, the highest classification accuracy on a hold-out sample is achieved by Model 9. Overall, it is clear that MLR models have a higher classification accuracy than OLR models. Also, 3-cat models perform better than 5-cat models. Table 4–18 Percentage correctly classified (PCC) Model Class. (ES) Class. (hold) Model 5 0.8801 0.8850 Model 6 0.8861 0.8851 Model 7 0.8649 0.8468 Model 8 0.8494 0.8600 Model 9 0.9096 0.9060 Model 10 0.9105 0.8977 Model 11 0.8818 0.8882 Model 12 0.8986 0.8811 4.1.5 Comparison of Estimated Rating Models Our application developed classification rules for five and three rating classes, which is more complex than the binary task. Therefore, we will combine multiple ROC curves to assess the performance of estimated models. As mentioned in 106 Chapter 4 2024 Martina Novotná Chapter 3.2.4, there are two main approaches to multiple ROC analysis: Each class versus the union of other classes, or distinct pairwise-class ROC curves. Both methods are suitable for summary statistics, such as the AUC (Area Under the Curve). We use the first approach in our application to compare the ability to predict a category versus a union of other categories. This way is sufficient for our purposes and effective for overall comparison. Thus, for each 5-cat model, we produce ROC curves and calculate the AUC as follows: • Cat1 versus the union of other categories (cat2 + cat3 + cat4 + cat5), • cat2 versus the union of other categories (cat1 + cat3 + cat4 + cat5), • cat3 versus the union of other categories (cat1 + cat2 + cat4 + cat5), • cat4 versus the union of other categories (cat1 + cat2 + cat3 + cat5), • cat5 versus the union of other categories (cat1 + cat2 + cat3 + cat4). The procedure is analogical for 3-cat models when considering only three rating categories. The ROC analysis is based on a parametric model, using the maximum likelihood estimation. We analyse the whole experimental and hold-out sample to determine whether the classification ability varies with the used sample selection. Preferably, we focus on the classification ability of the hold-out sample that is not used to estimate models. Table 4–19 AUC (5-cat models) Model/AUC Cat1 Cat2 Cat3 Cat4 Cat5 LDA (Model 1) Est. sample Hold-out 0.9658 0.9647 ++ 0.9876 0.9887 ++ 0.9689 0.9793 0.9199 0.9803 0.9840 0.9696 0.9836 0.9797 0.9931 LDA (Model 2) Est. sample Hold-out 0.9444 0.9556 0.9200 0.9815 0.9814 0.9816 0.9689 0.9676 0.9730 0.9804 0.9813 0.9779 0.9836 0.9721 0.9998 MLR (Model 5) Est. sample Hold-out 0.9950 0.9947 ++ 0.9942 0.9956 ++ 0.9807 0.9892 0.9253 0.9902 0.9940 0.9741 0.9955 0.9957 0.9967 MLR (Model 6) Est. sample Hold-out 0.9946 0.9926 0.9987 0.9927 0.9917 0.9953 0.9807 0.9794 0.9849 0.9902 0.9912 0.9866 0.9955 0.9924 0.9998 OLR (Model 7) Est. sample Hold-out 0.9890 0.9884 ++ 0.9867 0.9884 ++ 0.9654 0.9749 0.9130 0.9798 0.9823 0.9701 0.9800 0.9746 0.9829 OLR (Model 8) Est. sample Hold-out 0.9879 0.9854 0.9944 0.9832 0.9817 0.9878 0.9654 0.9635 0.9713 0.9798 0.9808 0.9772 0.9800 0.9712 0.9975 ++ not sufficient data to perform ROC analysis; random models (white), nonrandom models (grey) Firstly, we compare the 5-cat models. Based on the AUC values of 5-cat models (Table 4-19), the classification ability is sufficient, and there are only minor differences among the models. Nevertheless, we conclude that the best classification ability is performed by model 6 (MLR, random). We can also see The Effect of Selected Factors on Rating and its Dynamics 107 Micro-Modelling Approaches for Credit Rating and Corporate Survival that the best model is estimated by multivariate regression analysis no matter what sample we use (random, nonrandom). Table 4–20 AUC (3-cat models) Model/AUC Cat2 Cat3 Cat4 LDA (Model 3) Est. sample Hold-out 0.9913 0.9934 ++ 0.9604 0.9732 0.8992 0.9927 0.9951 0.9803 LDA (Model 4) Est. sample Hold-out 0.9870 0.9858 0.9911 0.9600 0.9589 0.9633 0.9923 0.9935 0.9877 MLR (Model 9) Est. sample Hold-out 0.9900 0.9918 ++ 0.9711 0.9822 0.9121 0.9960 0.9979 0.9837 MLR (Model 10) Est. sample Hold-out 0.9936 0.9929 0.9957 0.9753 0.9738 0.9798 0.9966 0.9974 0.9937 OLR (Model 11) Est. sample Hold-out 0.9912 0.9931 ++ 0.9584 0.9718 0.8945 0.9930 0.9954 0.9803 OLR (Model 12) Est. sample Hold-out 0.9875 0.9862 0.9911 0.9586 0.9571 0.9635 0.9923 0.9942 0.9894 ++ not sufficient data to perform ROC analysis; random models (white), nonrandom models (grey) Next, we compare 3-cat models (Table 4-20 ). Overall, the classification ability is slightly higher compared to the 5-cat models. However, all models perform sufficient classification ability, and there are minor differences among the AUC. The results support the main findings from the 5-cat models because the best classification ability is performed by model 10 (MLR, random). The ROC curves are presented only for 5-cat models (Appendix 4). They reflect the data on AUC shown in Table 4-19. The classification of models is similar and relatively high because we test the ability to predict one category against the union of other categories. Overall, we prefer model 6, which provides the best results. 4.1.6 Summary of Results We estimated twelve rating models in this study. We assumed ten financial variables and five or three output categories. In addition, we used two samples to determine whether sample selection affects the models and their prediction ability. All estimated models suggest that all used financial variables are good predictors of rating. The main findings of this study provide evidence that accounting-based variables significantly impact corporate rating in our sample. However, the direct effect on particular ratings is not easily interpretable and must be explained in the context of estimated models. 108 Chapter 4 2024 Martina Novotná The impact of selected variables on rating in discriminant models is measured by estimating the coefficients' contributions. However, since the contribution of the variables for the other variables in the model differs in each discriminant function (standardised coefficients), it is not conceivable to draw clear conclusions. Thus, ignoring other variables in the models, the association of individual variables with each discriminant function is compared through the correlation coefficients (structure matrix). According to the structure matrix, the highest average association with discriminant functions apply for roa, eqta, lnintcov, lnliqr and lncurr. The impact of the variables on rating in the logistic regression models is determined through their statistical significance, particularly in logit functions. For example, variables eqta, lnta, lncurr and ebitdar are not statistically significant in some logistic functions; thus, they are not considered key rating factors. Furthermore, based on the logistic models, the higher the value of roa, lnintcov, lnliqr and lncf, the greater the probability of a better rating assessment. Contrary, greater values of roe and lnltdta increase the likelihood of a lower rating. To summarise all partial results of the role of financial variables, we conclude that roa, roe, lnintcov, and lnliqr achieve the highest correlations with discriminant functions and are statistically significant in all logit functions. Thus, they are considered as the primary factors of rating prediction, followed by lncf and lnltdta. We provide evidence that the following financial variables are the main factors of rating assessment: • Return on total assets, • return on equity, • interest cover, • liquidity ratio, • cash flow, • long-term debt to total assets. The classification accuracy of all estimated models was determined based on the overall percentage correctly classified (PCC). The PCC, or hit ratio, ranges from 88% to 90.6% for a hold-out sample for our models. To assess the overall classification ability, we must set an acceptable level and compare the hit ratio to the standard. Firstly, we determine the percentage that could be determined correctly by chance. Since we compare the hit ratio for unequal group sizes, we consider just the largest group. For example, the largest random sample group in our study is represented by rating BBB (2353 observations in the experimental sample and 719 observations in the hold-out sample). Thus, we can arbitrarily assign all the subjects to the largest group. Therefore, if we classify each observation into this largest group in the case of random models, we would achieve a classification accuracy of 46.8% (experimental sample) and 44.4% (hold-out sample). Assuming nonrandom models, we get 43.4% (experimental sample) and 61.63% (hold-out sample). This approach is referred to as the maximum chance criterion, and unless a model achieves accuracy more than the computed values, The Effect of Selected Factors on Rating and its Dynamics 115 Micro-Modelling Approaches for Credit Rating and Corporate Survival constructed for models with null values of variables (h0) and for a hypothetical company with i) mean and ii) medium values (h1). Comparing Figures 4-3, 4-4 and 4-5, or specifically the cumulative hazard with smoothed hazard curves, we can see that plotting ranges are narrower. As Cleves et al. (2010) explain, this is because kernel smoothing requires averaging values over a moving window of data. Figure 4–4 Smoothed hazard functions 4.2.4 Model Verification Verifying whether the hazard functions are multiplicatively related is advisable when using the Cox model. Therefore, we assess the proportional hazards assumption by plotting the estimated hazards on a log scale. The lines in all graphs (Figure 4-5) seem parallel. Thus, we conclude that the proportionality assumption in both models is not violated. Figure 4–5 Smoothed hazard functions (log scale) 116 Chapter 4 2024 Martina Novotná Table 4–24 Test of PH assumptions Indep. variable Single rho (Chi2) Multiple rho (Chi2) tag 0.0217 (0.17) -0.0060 (0.02) roag 0.0387 (0.59) 0.0270 (0.39) roeg 0.0048 (0.01) 0.0260 (0.40) cfg 0.0060 (0.03) 0.0131 (0.20) intcovg -0.0716* (6.55) -0.0734* (10.27) liqrg 0.0598 (2.05) 0.0403 (1.24) Global 8.86 12.42 *significant at 0.05, standard error adjusted for 737 clusters The test of the proportional-hazards specification is based on the Schoenfeld residuals after fitting the model. It is used to test the independence between residuals and time. The test results in Table 4-24 suggest that the hazard assumption is not proportional to the variable intcovg. Furthermore, we find no evidence that our specification violates the proportional-hazard assumption regarding other variables. Therefore, our specification does not violate the proportional-hazards assumption in both models based on the global test. The models are evaluated by the overall model fit using Cox-Snell residuals. Figure 4-6 shows the Nelson-Aalen cumulative hazard estimator plots for CoxSnell residuals for both models. We can see some variability around the 45°line, particularly in the right-hand tail. Cleves et al. (2010) argue that some variability is expected due to the reduced effective sample caused by prior failures and censoring. However, we can see that both graphs fit the data adequately based on the charts. (a) Single (b) Multiple Figure 4–6 Cumulative hazard of Cox-Snell residuals The Effect of Selected Factors on Rating and its Dynamics 117 Micro-Modelling Approaches for Credit Rating and Corporate Survival However, we cannot choose a better model based on the graphical illustration. For this reason, we evaluate the predictive power by computing the Harrell’s C concordance statistics, which measures the agreement of predictions with observed failure order. The statistics can be defined as the proportion of all usable subject pairs in which the predictions and outcomes are concordant (Cleves et al., 2010). The values of C range between 0 and 1. Additionally, we can use Somers’D, which reports the rank correlation value, ranging from -1 to 1. Both measures are related as 2( 0.5).DC=− As Cleves et al. (2010) state, a value of 0.5 Harrell’s C and 0 of Somer’s D indicate no predictive ability of the model. The values of Harrell’s C and Somer’s D and their calculation procedure are shown in Table 4-25. Harrell’s C values are 0.8586 (single) and 0.8705 (multiple). Thus, we can correctly identify the order of the survival times for pairs of subjects 85.86%, or 87.05% of the time. Although the results are similar, the values are slightly higher for the multiple model. These results are supported by the value of Somer’s D, which is 0.7172 (single) and 0.7411 (multiple). The results provide evidence that both models have sufficient predictive accuracy. We prefer the multiple model over the single one based on the values. Table 4–25 Harrell’s C and Somer’s D Single model Multiple model Number of subjects (N) 3554 3554 Number of comparison pairs (P) 924134 984815 Number of orderings as expected (E) 793479 857329 Number of tied predictions (T) 0 0 Harrell’s C = (E + T/2) / P 0.8586 0.8705 Somer’s D 0.7172 0.7411 To conclude, the multiple failure-time data analysis leads to a more suitable model based on the statistical significance of the estimated coefficients and goodness of fit. On the other hand, it should be noted that both survival models are very similar based on estimated coefficients and used criteria. 4.2.5 Summary of Results This study aimed to develop rating models using survival analysis methods. Specifically, we applied the Cox proportional hazards model to analyze the survival time until the event. In our case, we focused on using survival analysis to model the time to a rating downgrade. As a part of the analysis, we examined the effect of financial variables on the probability of negative annual rating change. We provide evidence that annual changes in some variables are related to the rating downgrade, specifically using covariates tag, roag, roeg, cfg, intcovg and liqrg. On the contrary, variables eqtag, ebitdarg, ltdtag and currg are not statistically significant. Thus, we can conclude that annual changes in the following financial variables are good predictors of potential rating deterioration measured as rating downgrade in our sample: 118 Chapter 4 2024 Martina Novotná • Total assets, • return on total assets, • return on equity, • cash flow, • interest cover, • liquidity ratio. Two different approaches were used to estimate the models, depending on whether we considered only one or more events (rating downgrades) for one company. First, the single model was derived, assuming that the event can occur only once for each subject. On the other hand, the multiple models accept that the event can occur repeatedly. Due to these different assumptions, the input data and structure also had to be adjusted. The resulting models are presented based on the estimated coefficients for the variables used in the analysis. Both models are statistically significant, as are the estimated coefficients of the individual variables in the multiple model. We used baseline hazard and the hazard of the so-called average (medium) company to interpret the models based on mean (medium) values of variables. The fit of both models was assessed using Cox residuals. Based on the main findings of this study, we conclude that the multiple model approach is more suitable for events that might occur repeatedly. The simple model should be preferably used when the survival time until the first event is a matter of interest. In other cases, we should use multiple failure-time data analyses, making data use better. Overall, the findings of this study show that survival analysis is, in addition to typical financial problems, suitable for other types of tasks, such as the analysis of survival to bankruptcy or default. However, it is necessary to consider the specific data structure when applying it, especially whether the event can repeatedly occur for one subject or whether more events can occur for a given subject. In these cases, it is appropriate to use multiple failure-time analysis, which better corresponds to the problem. 4.3 Chapter Summary The fourth chapter's main goal was to analyse corporate data of selected CEE countries and understand the influence of selected financial variables on the MORE rating. Therefore, we first focused on estimating rating models using discriminant analysis and two logistic regression methods. In total, we obtained twelve rating models, which were compared with each other. Next, we identified financial variables that can be considered predictive factors of the rating evaluation. Even though the individual models differed slightly in terms of their formulation and classification ability, the overall conclusions confirm the main role of used financial variables in rating assessment. In the next section, we used the same data set and examined the relationship between financial variables and rating downgrades. In this case, we applied the The Effect of Selected Factors on Rating and its Dynamics 119 Micro-Modelling Approaches for Credit Rating and Corporate Survival method of survival analysis using the Cox model. Survival analysis allows us to estimate the survival probability of subjects to a predefined event. Since rating deterioration is quite a fundamental problem, especially for lenders, it is certainly important to recognize an impending rating change early. For this reason, the rating downgrade was chosen as the event in the survival analysis. Using the Cox model, we subsequently found six financial variables that have a fundamental connection with the annual deterioration of the rating assessment. The main influential financial variables based on both studies are summarized in Table 4-26. The results confirm that common financial indicators influence the rating assessment and its annual deterioration, regardless of the method used or the output variable. These are five commonly used financial indicators in financial analysis and evaluation of the financial performance of companies. Even if these indicators seem basic and simple, they still play a major role in assessing the borrower's credit quality and, thus, the rating. Of course, it is necessary to take them in the context of their change and its possible impact on the rating, or rather on its deterioration. Table 4–26 Main factors of rating and its downgrade Rating assessment Rating downgrade Total assets x ✓ Return on total assets ✓ ✓ Return on equity ✓ ✓ Liquidity ratio ✓ ✓ Cash flow ✓ ✓ Interest cover ✓ ✓ Long-term debt to total assets ✓ x In general, it has been shown in this chapter that the first approach, based on discriminant analysis and logistic regression, and the second method, using survival analysis, lead to similar results. In addition, with the help of survival analysis, we could estimate the hazard and survival functions, which can be used to predict the probability of survival or the time until the rating deteriorates. Furthermore, it means we can look deeper into the development of the monitored variable over time and its dynamics. For this reason, survival analysis will be used in the following section, where the survival time of firms until bankruptcy will be modelled. Moreover, as this is a specific event related to the borrower's credit quality, a sub-goal of the following chapter will be to understand and link the relationship between corporate survival probability and credit rating. Chapter 5 Relationship Between Rating and Corporate Bankruptcy Rates This chapter provides an alternative view on measuring and predicting firm-based credit risk. While we estimated rating models in the previous section, the focus is on bankruptcy models in the following application. Bankruptcy, as a terminal state of the company, and rating are related because the worst rating assessment is typically issued to insolvent companies in financial trouble, often close to bankruptcy. Thus, we also examine and model corporate bankruptcy in this research to extend the previous findings of the main predictors of rating assessment. The knowledge and understanding of rating and bankruptcy factors are essential for credit risk management. The evidence of corporate survival and nondefault rates helps investors and lenders assess the credit quality of borrowers and predict potential problems of default, insolvency or even corporate bankruptcy. For example, the analysis of time to default, which CRAs conduct, is based on the cumulative distribution of defaulters by the time of default and survival rates. According to the historical default rates published by CRAs, there is a clear association between rating and default rates. Thus, assuming this relationship, we can estimate the survival rates of a certain sample of companies and compare them with historical non-default rates published by CRAs. The main goal of this chapter is to determine whether there is a measurable relationship between estimated bankruptcy rates and default rates published by rating agencies using empirical data from Czech companies. In a positive case, this relationship can be described in a certain way. Then, a procedure can be suggested in which the detected bankruptcy rates could be used for a rating assessment corresponding to the rating agencies' assessment. This procedure has meaning and main application, especially in cases where we have data on corporate bankruptcies. We can process them statistically, and our goal is to translate them in a certain way into the “language” of rating agencies. 122 Chapter 5 2024 Martina Novotná As mentioned in Chapter 3.1, there is a vast literature on predicting corporate bankruptcy using various techniques. However, little attention is paid to estimating corporate bankruptcy using survival analysis methods compared to discriminant, logistic, neural network methods or classification trees. Therefore, the partial aim of this study is to analyse survivor data of Czech companies and assess the impact of various factors on corporate survival. During the observed time interval, our analysis considers bankruptcy the failure event. Thus, the time between the start of the business and bankruptcy is used to estimate survival and hazard functions. Eventually, the estimated survival rates can be compared with CRAs’ historical default rates and corresponding rating assessments. Hence, we consider three credit risk measures in the following study: bankruptcy rates, default rates, and rating. The structure is as follows. Firstly, the association between rating and corporate defaults observed and published by rating agencies is studied. Next, the Kaplan-Meier method is used to assess the effect of selected characteristics of Czech companies on the survival probability. Then, the relationship between estimated cumulative bankruptcy rates and published default rates is explored. Finally, based on the main findings, a procedure is proposed for converting the bankruptcy rates into rating assessments. 5.1 Association Between Rating and Corporate Defaults This chapter provides an overview of the relationship between rating and default rates based on historical data from rating agencies. Next, we describe the approach rating agencies use to calculate cumulative default rates. 5.1.1 CRA Annual Default Rates The purpose of this section is to explore the relationship between credit rating grades and corporate default rates. According to the historical occurrence of defaults within rating grades observed and published by CRAs, there is an evident correlation between the initial rating of a firm and its time to default. Typically, the historical number of defaults within an investment grade is substantially lower when compared to a speculative grade (Table 5-1). The highest number of defaults, or the highest default rates, occurred during the financial crisis of 2008 – 2009, even within the investment grade. During the last 30 years, from 1990 to 2020, the investment grade achieved the highest default rate, 0.42%, during the two financial crises in 2002 and 2008 (S&P, 2021). Relationship Between Rating and Corporate Bankruptcy Rates 123 Micro-Modelling Approaches for Credit Rating and Corporate Survival Table 5–1 Corporate default summary, 2005 – 2020 Year Total defaults Inv.- grade defaults Spec.- grade defaults Default rate (%) Inv.-grade default rate (%) Spec.-grade default rate (%) 2005 40 1 31 0.60 0.03 1.51 2006 30 0 26 0.48 0.00 1.19 2007 24 0 21 0.37 0.00 0.91 2008 127 14 89 1.80 0.42 3.71 2009 268 11 224 4.18 0.33 9.95 2010 83 0 64 1.21 0.00 3.02 2011 53 1 44 0.80 0.03 1.85 2012 83 0 66 1.14 0.00 2.59 2013 81 0 64 1.06 0.00 2.31 2014 60 0 45 0.69 0.00 1.44 2015 113 0 94 1.36 0.00 2.78 2016 163 1 143 2.09 0.03 4.24 2017 95 0 83 1.21 0.00 2.47 2018 82 0 72 1.03 0.00 2.10 2019 118 2 92 1.30 0.06 2.54 2020 226 0 198 2.74 0.00 5.50 Source: S&P (2021) CRAs also calculate and publish annual default rates of rating grades, showing the historical trend of corporate defaults within rating categories (Table 5-2). Most rated defaulters come from the lowest categories, B and CCC/C, while the default rates of the AAA category are zero. Table 5–2 Global annual default rates by rating category (%), 2005 – 2020 Year AAA AA A BBB BB B CCC/C 2005 0.00 0.00 0.00 0.07 0.31 1.74 9.09 2006 0.00 0.00 0.00 0.00 0.30 0.82 13.33 2007 0.00 0.00 0.00 0.00 0.20 0.25 15.24 2008 0.00 0.38 0.39 0.49 0.81 4.08 27.27 2009 0.00 0.00 0.22 0.55 0.75 10.91 49.46 2010 0.00 0.00 0.00 0.00 0.58 0.85 22.73 2011 0.00 0.00 0.00 0.07 0.00 1.66 16.42 2012 0.00 0.00 0.00 0.00 0.30 1.56 27.33 2013 0.00 0.00 0.00 0.00 0.10 1.63 24.34 2014 0.00 0.00 0.00 0.00 0.00 0.77 17.03 2015 0.00 0.00 0.00 0.00 0.16 2.39 25.73 2016 0.00 0.00 0.00 0.06 0.47 3.76 33.17 2017 0.00 0.00 0.00 0.00 0.08 1.00 26.56 2018 0.00 0.00 0.00 0.00 0.00 0.99 27.18 2019 0.00 0.00 0.00 0.11 0.00 1.49 29.76 2020 0.00 0.00 0.00 0.00 0.93 3.52 47.48 Source: S&P (2021) 124 Chapter 5 2024 Martina Novotná The historical trend of the default rates is relatively stable and shows a clear association between rating grade and the general default rate or credit risk. There is a negative correlation between the initial rating of a firm and its time to default. For example, the average time to a default of entities initially rated A grade (the time between first rating and date of default) is 14.1 years. In comparison, the average time to default among entities originally B was 5.1 years based on 1981 – 2020 in the study by S&P Global Ratings (S&P, 2021). Next, the average time for AAA rating is 18 years, BBB is 9.2 years, and CCC/C is only 2.2 years. In addition to calculating time to default, the cumulative distribution of defaulters by the time of default and survival, or non-default rates, can also be used to get more information on the rating dynamics and defaults. For example, the survival rate or the percentage of B corporate issuers still alive for one, three or five years was 97.6%, 93.4% and 89.6% for 2011 – 2015 (Table 5-3). Table 5–3 Corporate defaults and survival rates (2011 – 2015) Rating One-year pool (2015) Three-year pool (2013 – 2015) Five-year pool (2011 – 2015) Number of defaults Nondefault rate Number of defaults Nondefault rate Number of defaults Nondefault rate AAA 0 100.0% 0 100.0% 0 100.0% AA 0 100.0% 0 100.0% 0 100.0% A 0 100.0% 0 100.0% 1 99.8% BBB 0 100.0% 0 100.0% 0 100.0% BB 2 99.8% 7 99.1% 22 97.1% B 42 97.6% 93 93.4% 122 89.6% CCC/C 38 73.8% 56 57.9% 47 59.8% Source: Source: S&P (2015) For comparison, Table 5-4 summarises default rates for 2016 – 2020, partly already reflecting the consequences of the COVID-19 pandemic. Table 5–4 Corporate defaults and survival rates (2016 – 2020) Rating One-year pool (2020) Three-year pool (2018 – 2020) Five-year pool (2016 – 2020) Number of defaults Nondefault rate Number of defaults Nondefault rate Number of defaults Nondefault rate AAA 0 100.0% 0 100.0% 0 100.0% AA 0 100.0% 0 100.0% 0 100.0% A 0 100.0% 2 99.9% 0 100.0% BBB 0 100.0% 2 99.9% 13 99.3% BB 12 99.1% 17 98.7% 28 97.8% B 73 96.5% 175 90.9% 283 85.0% CCC/C 113 52.5% 103 47.2% 103 48.2% Source: S&P (2021) Relationship Between Rating and Corporate Bankruptcy Rates 131 Micro-Modelling Approaches for Credit Rating and Corporate Survival industrial companies is the lowest among all groups, followed by utility, services and agriculture (see Figure 5-2). The tests of equality of overall survival functions across groups based on the log-rank test (Chi2(3)=421.24), Wilcoxon (Chi2(3)=384.24) and Peto-Peto test (Chi2(3)=416.54) reject the hypothesis that the survival functions are the same. Figure 5–2 Survival and cumulative functions by sectors b) The Effect of Legal Form The mean estimated times are summarised in Table 5-12. Joint-stock and limitedliability companies have the lowest estimated survival time, followed by cooperatives and other legal forms. Table 5–12 Estimated mean survival time by legal status Group Legal status Mean Standard error Confidence interval (95%) 1 Joint-stock company 8517.66 40.5545 8438.17 8597.15 2 Cooperatives 8726.66 86.3793 8557.36 8895.96 3 Limited-liab. 8552.97 18.1951 8517.29 8588.62 4 Other 8996.49 65.7681 5567.59 9125.39 Total 8568.40 16.2197 8536.61 8600.19 The survival and cumulative functions of each legal form are depicted in Figure 5-3. The tests of equality of overall survival functions across groups based on the log-rank test (Chi2(3)=14.31), Wilcoxon (Chi2(3)=14.54) and Peto-Peto test (Chi2(3)=15.96) reject the hypothesis that the survival functions are the same. 132 Chapter 5 2024 Martina Novotná Figure 5–3 Survival and cumulative hazard functions by legal form c) The Effect of Business Size While the lowest estimated mean survival time is associated with small companies, the mean survival times of other categories are similar (Table 5-13). Micro companies have the highest estimated mean survival time, followed by medium and large companies. The visual results suggest that small companies have the lowest probability of survival (Figure 5-4). Table 5–13 Estimated mean survival time by size Group Business category Mean Standard error Confidence interval (95%) 1 Micro 8595.66 21.6411 8553.25 8638.08 2 Small 8488.22 29.7732 8429.86 8546.57 3 Medium 8568.99 44.0947 8482.56 8655.41 4 Large 8556.97 109.99 8341.39 8772.54 Total 8644.55 16.4998 8612.21 8676.89 Figure 5–4 Survival and cumulative hazard functions by size Relationship Between Rating and Corporate Bankruptcy Rates 133 Micro-Modelling Approaches for Credit Rating and Corporate Survival The tests of equality of overall survival functions across groups based on the log-rank test (Chi2(3)=16.58), Wilcoxon (Chi2(3)=12.54) and Peto-Peto test (Chi2(3)=15.86) reject the hypothesis that the survival functions are the same. 5.3 The Relationship Between Bankruptcy Rates and Rating Assessment Based on the results of our prior analysis, the estimated bankruptcy rates are compared with CRAs’ historical default rates to assess the average credit quality of the observed firms. Since bankruptcy can be considered a legal procedure for liquidating a business that cannot fully pay its debts, a strong correlation between bankruptcy and default rates is assumed. Thus, we can compare default rates with bankruptcy rates without much impact on the overall findings and their interpretation. The aim of this section is to examine whether there is an association between the estimated corporate bankruptcy rates and rating cumulative default rates. First, the estimated survival functions are used to determine cumulative bankruptcy rates at the end of a particular year. Next, they are compared with CRAs’ historical cumulative default rates and, eventually, corresponding rating assessments. Finally, through peer comparison, we propose a way that bankruptcy rates can be translated into the rating. Hence, the main purpose of this study is to link the bankruptcy rates and rating assessment. 5.3.1 Association Between Bankruptcy Rates and Rating To assess the average rating quality of companies in our data sample, we examine the association between corporate bankruptcy rates and S&P rating grades. Firstly, cumulative bankruptcy rates (CBR) of the sample are calculated using the full data set according to the time horizon. Then, the bankruptcy rates are compared with long-term average global cumulative default rates of S&P rating (CDR, 19812020); see Table 5-14. The association between corporate bankruptcy rates and S&P cumulative default rates is graphically presented in Figure 5-5. It can be seen from this figure that there are substantial differences between default rates of rating groups. For example, cumulative default rate CCC/C, B, and BB curves lie above all other curves, reflecting the greatest credit risk. It is also evident that the bankruptcy rates lie between BBB and BB's rating categories, approaching BB with a longer time horizon. Hence, according to the rating definition, our companies' investment quality changes with time. 134 Chapter 5 2024 Martina Novotná Table 5–14 Cumulative bankruptcy rates and average cumulative default rates Time horiz. CBR CDR AAA CDR AA CDR A CDR BBB CDR BB CDR B CDR CCC/C 1 0.0072 0.0000 0.0002 0.0005 0.0016 0.0063 0.0334 0.2830 2 0.0129 0.0003 0.0006 0.0013 0.0043 0.0193 0.0780 0.3833 3 0.0184 0.0013 0.0011 0.0022 0.0075 0.0346 0.1175 0.4342 4 0.0234 0.0024 0.0021 0.0033 0.0114 0.0499 0.1489 0.4636 5 0.0281 0.0034 0.0030 0.0046 0.0154 0.0643 0.1735 0.4858 6 0.0350 0.0045 0.0041 0.0060 0.0194 0.0775 0.1936 0.4961 7 0.0428 0.0051 0.0049 0.0076 0.0227 0.0889 0.2099 0.5075 8 0.0494 0.0059 0.0056 0.0900 0.0261 0.0990 0.2231 0.5149 9 0.0556 0.0064 0.0063 0.0105 0.0293 0.1082 0.2350 0.5216 10 0.0629 0.0070 0.0070 0.0120 0.0324 0.1164 0.2462 0.5276 11 0.0685 0.0072 0.0076 0.0134 0.0355 0.1233 0.2558 0.5321 12 0.0753 0.0075 0.0082 0.0146 0.0380 0.1299 0.2631 0.5368 13 0.0856 0.0078 0.0088 0.0159 0.0403 0.1359 0.2699 0.5423 14 0.0985 0.0084 0.0093 0.0171 0.0428 0.1409 0.2763 0.5469 15 0.1095 0.0090 0.0099 0.0184 0.0454 0.1465 0.2824 0.5476 Source: S&P (2021), author Figure 5–5 Cumulative bankruptcy and default rates However, it is important to compare cumulative bankruptcy rates with default rates. Unlike a default event that refers to the debtor’s incapacity or refusal to meet their debt obligations when due, bankruptcy is the legal status of an entity that cannot repay debts to creditors. Based on their definitions, we assume default rates to be generally higher than bankruptcy rates. Nevertheless, this comparison can give us an interesting look at the association between bankruptcy and rating default rates and the overall credit risk of corporates in the data sample. In addition to the average values, we summarize the bankruptcy rates by sector, legal form and size (Appendix 5). The rates correspond to the main findings in Chapter 5.2.2, specifically Figure 5-2, Figure 5-3 and Figure 5-4, showing survival and hazard functions. Relationship Between Rating and Corporate Bankruptcy Rates 135 Micro-Modelling Approaches for Credit Rating and Corporate Survival Next, we examine credit default rates (CDRs) and their relationship with estimated corporate bankruptcy rates. Table 5-15 summarizes the average credit default rates for the time horizon of ten years based on the statistics by S&P (2021). We can see the global CDRs and the rates from Europe and emerging countries. In addition to the average rates by rating groups, note the overall rates for investment and speculative grades. Table 5–15 Average CDRs (t = 10 years) Global CDR Europe CDR Emerging CDR AAA 0.0036 0.0000 0.0000 AA 0.0035 0.0017 0.0000 A 0.0057 0.0025 0.0003 BBB 0.0170 0.0070 0.0146 BB 0.0664 0.0363 0.0452 B 0.1659 0.1281 0.1191 CCC/C 0.4618 0.4599 0.2800 Investment 0.0097 0.0036 0.0109 Speculative 0.1418 0.1065 0.0888 Source: S&P (2021) Based on Table 5-14, we calculate the average cumulative bankruptcy rate for a time horizon of 1–10 years ( 0.0336CBR = ). Then, we find the differences between the average CDRs (Table 5-15 ) and the calculated average CBR. Finally, we can see the differences between CDRs and CBR, referred to as average spreads (AS), in Table 5-16. Table 5–16 Average spreads CDR-CBR (t = 10 years) Global AS Europe AS Emerging AS AAA -0.0299 -0.0336 -0.0336 AA -0.0301 -0.0319 -0.0336 A -0.0279 -0.0310 -0.0333 BBB -0.0166 -0.0265 -0.0189 BB 0.0329 0.0027 0.0117 B 0.1323 0.0945 0.0855 CCC/C 0.4282 0.4263 0.2464 Investment -0.0239 -0.0300 -0.0227 Speculative 0.1083 0.0729 0.0552 The relationship between the calculated average spreads and ratings is graphically presented in Figure 5-6. The higher the spread, the lower the average rating assessment, coded from 1 (AAA) to 7 (CCC/C). Hence, based on the results, bankruptcy rates seem to be good indicators of rating quality. 136 Chapter 5 2024 Martina Novotná Figure 5–6 Rating and average spreads In the next section, a method is proposed for how the calculated average spreads can be used to assign the probable corporate rating. As shown above, the CBRs are based on the N-A estimates of cumulative hazard rates in this study. Therefore, for rating estimation, we need average spreads by rating grades. Furthermore, since we model the data of Czech companies, the spreads are based on European CDRs (Table 5-17). Table 5–17 Spreads by rating and time horizon Time horiz. AAA AA A BBB BB B CCC/ C Inv. Spec. 1 -0.0072 -0.0072 -0.0069 -0.0066 -0.0036 0.0150 0.2792 -0.0069 0.0216 2 -0.0129 -0.0127 -0.0122 -0.0112 -0.0016 0.0442 0.3743 -0.0120 0.0430 3 -0.0184 -0.0179 -0.0174 -0.0154 0.0005 0.0705 0.4091 -0.0169 0.0594 4 -0.0234 -0.0224 -0.0219 -0.0191 0.0034 0.0910 0.4429 -0.0212 0.0730 5 -0.0281 -0.0265 -0.0258 -0.0227 0.0078 0.1075 0.4614 -0.0251 0.0842 6 -0.0350 -0.0329 -0.0321 -0.0272 0.0080 0.1170 0.4635 -0.0309 0.0889 7 -0.0428 -0.0404 -0.0390 -0.0332 0.0070 0.1222 0.4611 -0.0378 0.0906 8 -0.0494 -0.0467 -0.0453 -0.0383 0.0048 0.1242 0.4612 -0.0438 0.0904 9 -0.0556 -0.0526 -0.0513 -0.0429 0.0022 0.1265 0.4550 -0.0494 0.0897 10 -0.0629 -0.0599 -0.0585 -0.0487 -0.0011 0.1270 0.4557 -0.0562 0.0881 Source: S&P (2021); author 5.3.2 Transmission of Bankruptcy Rates to Rating Assessment In the previous section, we observed the relationship between the estimated bankruptcy rates and the rating using the spreads between the average bankruptcy rates and default rates for different time horizons. In this part, we propose a procedure for converting the bankruptcy rates to ratings using the calculated average spreads. The procedure steps are as follows: Relationship Between Rating and Corporate Bankruptcy Rates 137 Micro-Modelling Approaches for Credit Rating and Corporate Survival 1. Estimate a cumulative bankruptcy rate (ECBR), for example, for the particular combination of variables (based on the table in Appendix 5). 2. Calculate the estimated spread between CDR and ECBR for each rating category (based on Table 5-14) using the equation (5.1), ES CDR ECBR=− . (5.1) 3. Compare ES with the average spread AS summarized in Table 5-16 . Then, choose the rating category with the lowest absolute value of deviance based on equation (5.2). D ES AS=− . (5.2) Although this procedure is based on the European credit default rates published by the S&P agency, by analogy, it can be used for different geographical locations, rating agencies and time horizons. An application example of this procedure is shown below in the text. Assuming that bankruptcy rates depend only on the sector, legal form and corporate size, we can estimate a rating based on the following procedure. For example, the above technique will be used to determine the probable rating for the following combinations of corporate characteristics: a) Industrials, limited-liability, micro (i = 1), b) services, joint-stock, large (i = 2), c) agriculture, cooperative, medium (i = 3). We assume time horizons 15t= and 210t= years to assess the effect of time on the development of rating quality. Firstly, we find the estimated cumulative bankruptcy rates (ECBRs) as simple averages of 5-year and 10-year CBRs. Since our cumulative hazard rate estimates are based on a nonparametric analysis, we use the same weight for each categorical variable. Then, for ith combination of variables (i = 1, 2, 3) and time horizon t (t = 5, 10), the estimated cumulative bankruptcy rate is found as: , 1( _ _ _ ). 3 it ECBR CBR sector CBR legal CBR size= + + (5.3) For example, using the equation (5.3), we get the following ECBRs: ( ) ( ) ( ) 1, 5 2, 5 3, 5 10.045+0.0304 0.0223 0.0326, 3 10.0216+0.0161 0.0097 0.0158, 3 10.0145+0.0306 0.159 0.0203. 3 t t t ECBR ECBR ECBR = = = = + = = + = = + = 138 Chapter 5 2024 Martina Novotná ( ) ( ) ( ) 1, 10 2, 10 3, 10 10.105 0.0647 0.0503 0.0733, 3 10.0448 0.0618 0.05 0.0522, 3 10.0275 0.0402 0.0526 0.0401. 3 t t t ECBR ECBR ECBR = = = = + + = = + + = = + + = The results suggest differences between the estimated cumulative bankruptcy rates of two time periods. For example, assuming the time horizon of 5 years, the greatest cumulative hazard rate is associated with a combination of i = 1 (industrials, limited-liability, micro), followed by i = 3 (agriculture, cooperative, medium) and i = 2 (services, joint-stock, large). In the 10-year time horizon, the greatest cumulative hazard rate is again associated with the scenario i = 1. It is followed by i = 2, suggesting that the cumulative hazard rates vary with time for our combinations of variables. Next, we calculate ES and D using the equations (5.1) and (5.2) and find the corresponding rating categories with the minimum absolute values of D. The results are summarized in Table 5-18. Thus, assuming the time horizon of 5 years (since the company's founding), we assign the middle rating BBB to the companies with the combination of variables i = 2 and i = 3. In the longer time horizon, the assigned rating is the same for i = 3; however, it is estimated to be BB for i = 2. The first corporate characteristics (i = 1) indicate a BB rating category regardless of the time horizon. Nevertheless, the credit rating seems to have worsened for a long time based on the suggested speculative rating grade. Table 5–18 Rating estimation (K-M model) AAA AA A BBB BB B CCC/C AS -0.0336 -0.0319 -0.0310 -0.0265 0.0027 0.0945 0.4263 ES1,t=5 -0.0326 -0.0310 -0.0303 -0.0272 0.0033 0.1030 0.4569 ES1,t=10 -0.0733 -0.0703 -0.0689 -0.0591 -0.0115 0.1166 0.4453 ES2,t=5 -0.0158 -0.0142 -0.0135 -0.0104 0.0201 0.1198 0.4737 ES2,t=10 -0.0522 -0.0492 -0.0478 -0.0380 0.0096 0.1377 0.4664 ES3,t=5 -0.0203 -0.0187 -0.0180 -0.0149 0.0156 0.1153 0.4692 ES3,t=10 -0.0401 -0.0371 -0.0357 -0.0259 0.0217 0.1498 0.4785 D1,t=5 0.0010 0.0010 0.0008 0.00064 0.00059 0.0085 0.0306 D1,t=10 0.0398 0.0384 0.0379 0.0326 0.0143 0.0221 0.0189 D2,t=5 0.0178 0.0177 0.0175 0.0161 0.0174 0.0253 0.0474 D2,t=10 0.0186 0.0173 0.0168 0.0115 0.0069 0.0432 0.0401 D3,t=5 0.0132 0.0132 0.0130 0.0116 0.0128 0.0208 0.0428 D3,t=10 0.0065 0.0052 0.0047 0.0006 0.0190 0.0553 0.0522 The overall results suggest that the first combination of variables (industrials, limited-liability, micro) is the riskiest. The second case (services, joint-stock, large) is associated with the average risk, potentially worsening with a longer time Relationship Between Rating and Corporate Bankruptcy Rates 139 Micro-Modelling Approaches for Credit Rating and Corporate Survival horizon. Finally, based on the estimated rating, we found the combination with the lowest and most stable level of credit risk (agriculture, cooperative, medium). The main findings show that based on estimated bankruptcy rates, we can estimate the future rating development according to the time horizon. 5.4 Chapter Summary This chapter focused on understanding the relationship between bankruptcy rates, default rates published by CRAs, and credit ratings. As was shown in the introductory part of this chapter, methods based on survivorship analysis are used by rating agencies to determine the default rates of rated subjects, both according to rating categories and to the considered time horizon. This section used the survival method to assess the influence of selected corporate characteristics on the survival probability of Czech companies. For this purpose, the basic procedure used was the Kaplan-Meier method. It is a nonparametric analysis with which we can discover the basic relationships and understand the data survivorship. We focused on assessing the influence of industry, legal form and size on survival time. The results indicate that all used factors are related to the probability of corporate survival. Two hypotheses were confirmed, suggesting that small, industrial firms are riskier compared to the other considered sectors and sizes. On the other hand, our findings suggest that jointstock companies are riskier compared to other legal forms, unlike the assumption. The overall results are summarized in Table 5-19. For example, a small, industrial joint-stock company can be considered the riskiest compared to other cases. Table 5–19 Summary of results Corporate characteristics Category Mean survival time Industry Services Industrials Agriculture Utility ••• • •••• •• Legal form Joint-stock Cooperatives Limited-liability Other • ••• •• •••• Size Micro Small Medium Large •••• • ••• •• • (lowest) •••• (greatest) Using the procedures described above in Chapter 5.3, we derived survival and cumulative hazard functions, which were subsequently used to estimate the rating. They were determined using the proposed approach based on the average spread between cumulative bankruptcy and default rates. Based on the average spread, it 140 Chapter 5 2024 Martina Novotná was confirmed that the higher the spread, the lower the rating. Thus, the proposed procedure was built on this finding. The relevant rating category was determined using the absolute deviation between the estimated and average spread. This procedure was subsequently used to determine the rating of three hypothetical companies with different combinations of the considered characteristics. Even though this is a greatly simplified approach to determining the rating based on only three categorical variables, this method clearly shows that the risk associated with the probability of survival can be better understood with the help of the variables used. The calculation above considered only three categories of corporate characteristics, ignoring the potential effect of other factors. Although we cannot accurately determine the rating based on the sector, legal form and corporate size, converting the cumulative bankruptcy rate into a rating will be further examined in the following chapters. The procedure will be additionally used in the next section based on survival and cumulative hazard functions estimated by other survival analysis methods. The Cox model will be used first, and then the Weibull model. The individual steps are analogous to the last part. First, the survival and cumulative hazard functions will be estimated. Then, the cumulative bankruptcy rates for the selected time horizons will be determined, and the proposed procedure will be used to determine the rating.