scieee AI-readable full text Open interactive document viewer

Combining data and text mining techniques for automatic analysis of financial reports

Pinto, Marcelo Queirós

Abstract

The application of Data Science techniques, specifically Natural Language Process ing (NLP) and Machine Learning, in financial markets is of immense interest to in vestors, as these techniques can have a potential economic impact. In particular, stock markets represent an opportunity that has been exploited in several ways, such as us ing market opinions (e.g., news, blogs) to predict the direction of price movement or even volatility. This study analyses the 10-K documents of the S&P 100 index for 10 years (2008-2017), which contains the 102 largest companies in the United States of America. The 10-K is an annual financial report required by the United States Securities and Ex change Commission (SEC), which describes the financial performance of a company. Recent research suggests that the readability of a company’s 10-K text document may influence its future financial performance, since the way the market perceives textual information also depends on the readability of that text. In this sense, this work aims to understand the relationship between 48 readability metrics applied to these reports and the corresponding future financial performance of these companies. A clustering approach was applied over these readability metrics, aiming to identify distinct and valuable readability clusters. As an external evaluation, we assessed the information value of the clusters by analyzing 3 future crash risk metrics, that are often used to assess the companies’ financial performance.

Full text

University of Minho School of Engineering Department of Informatics Marcelo Queir´ os Pinto Combining Data and Text Mining techniques for automatic analysis of Financial Reports October 2019 University of Minho School of Engineering Department of Informatics Marcelo Queir´ os Pinto Combining Data and Text Mining techniques for automatic analysis of Financial Reports Master dissertation Master Degree in Computer Science and Engineering Dissertation supervised by Professor Doctor Paulo Cortez Professor Doctor Nelson Areal October 2019 Despacho RT - 31 /2019 - Anexo 3 Declaração a incluir na Tese de Doutoramento (ou equivalente) ou no trabalho de Mestrado DIREITOS DE AUTOR E CONDIÇÕES DE UTILIZAÇÃO DO TRABALHO POR TERCEIROS Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Licença concedida aos utilizadores deste trabalho Atribuição CC BY https://creativecommons.org/licenses/by/4.0/ ACKNOWLEDGEMENTS I consider that the personal stability is fundamental to professional success, and in this sense, I give my deep gratitude to the people I can always count on in my personal life, and who were undoubtedly fundamental throughout this process. Thank you for listening and giving feedback, for enduring me, for motivating me even on the toughest days and essentially for all your concern. I would like to thank Professor Paulo Cortez and Professor Nelson Areal for the fantastic orientation of the project, the meetings, the suggestions and especially for all the learning I have gathered. Both were two of the best professors I have ever dealt with. Since during this thesis I worked in 2companies, I want to thank them for all the flexibility and support they provided whenever I needed it. They have given me tremendous technical knowledge important for this thesis as well as important personal skills for my future. The colleagues and friends I made in these companies were always available to give me any opinion on this thesis and I really appreciate it. To my master and bachelor colleagues at the University of Minho and University of Tr´ as-os-Montes and Alto-Douro, who at the same time many of them became personal friends, I greatly appreciate all your willingness to help me and give good opinions. Marcelo Queir´ os i Despacho RT - 31 /2019 - Anexo 4 Declaração a incluir na Tese de Doutoramento (ou equivalente) ou no trabalho de Mestrado STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho. ABSTRACT The application of Data Science techniques, specifically Natural Language Processing (NLP) and Machine Learning, in financial markets is of immense interest to investors, as these techniques can have a potential economic impact. In particular, stock markets represent an opportunity that has been exploited in several ways, such as using market opinions (e.g., news, blogs) to predict the direction of price movement or even volatility. This study analyses the 10-K documents of the S&P 100 index for 10 years (20082017), which contains the 102 largest companies in the United States of America. The 10-K is an annual financial report required by the United States Securities and Exchange Commission (SEC), which describes the financial performance of a company. Recent research suggests that the readability of a company’s 10-K text document may influence its future financial performance, since the way the market perceives textual information also depends on the readability of that text. In this sense, this work aims to understand the relationship between 48 readability metrics applied to these reports and the corresponding future financial performance of these companies. A clustering approach was applied over these readability metrics, aiming to identify distinct and valuable readability clusters. As an external evaluation, we assessed the information value of the clusters by analyzing 3future crash risk metrics, that are often used to assess the companies’ financial performance. Keywords: Data Science; Financial performance prediction; Natural Language Processing; Readability evaluation; Stock markets. ii RESUMO A aplicac¸˜ ao das t´ ecnicas de Ciˆ encia de Dados, especificamente Processamento de Linguagem Natural e Machine Learning, nos mercados financeiros ´ e de imenso interesse para os investidores, uma vez que podem ter um potencial impacto econ´ omico. Em particular, os mercados de ac¸ ˜ oes representam uma oportunidade que tem sido explorada de v´ arias formas, como no uso de informac¸ ˜ oes de mercado (por exemplo not´ ıcias, blogs) para prever a direc¸˜ ao do movimento dos prec¸os ou mesmo o movimento da volatilidade. Este estudo analisa os documentos 10-K do ´ ındice S&P 100 durante 10 anos (20082017), que cont´ em as 102 maiores empresas dos Estados Unidos da Am´ erica. O 10-K ´ e um relat´ orio financeiro anual exigido pela Comiss˜ ao de Valores Mobili´ arios dos Estados Unidos (SEC), que descreve o desempenho financeiro de uma empresa. Pesquisas recentes sugerem que a legibilidade do documento de texto 10-K de uma empresa pode influenciar o seu desempenho financeiro futuro, uma vez que a forma como o mercado perceciona as informac¸ ˜ oes textuais tamb´ em depende da legibilidade desse texto. Neste sentido, este trabalho visa compreender a relac¸˜ ao entre 48 m´ etricas de legibilidade aplicadas a esses relat´ orios e o desempenho financeiro futuro correspondente dessas empresas. Uma abordagem de agrupamento de dados foi aplicada nestas m´ etricas de legibilidade, com o objetivo de identificar grupos de legibilidade distintos e relevantes. Com uma avaliac¸ ˜ ao externa, avaliamos o valor das informac¸ ˜ oes desses grupos analisando trˆ es m´ etricas de crash risk futuro, que s˜ ao frequentemente usadas para avaliar o desempenho financeiro das empresas. Palavras-Chave: Avaliac¸˜ ao da legibilidade; Ciˆ encia de Dados; Mercado de ac¸ ˜ oes; Previs˜ ao do desempenho financeiro; Processamento de Linguagem Natural. iii CONTENTS Acronyms 1 1 introduction 2 1.1Context and motivation 2 1.2Objectives 3 1.3Document organization 4 2 state of the art 6 2.1Defining readability 6 2.2Readability measures 6 2.2.1Word count and file size 7 2.2.2Fog Index 7 2.2.3Bog Index 8 2.2.4All 48 readability measures used in this study 9 2.3Evidence from financial reports readability research in finance indicators 18 2.4Text Mining and NLP 19 2.4.1Related data analysis areas 19 2.4.2Topic modelling 20 2.4.3Classification 21 2.5Future crash risk 22 2.5.1What is a crash? 22 2.5.2Reasons for a higher future crash risk 22 2.5.3Future crash risk measures 23 3 development 26 3.1Methodology 26 3.1.1The problem and its challenges 26 3.1.2Proposed Approach 27 3.2Business understanding 29 3.3Data understanding 29 3.3.1Quality assurance 29 3.3.2Importance of ensuring accurate results 30 3.3.3Data collection 30 iv Contents v 3.3.4Describing the data 32 3.4Data preparation 38 3.4.1Select the data 38 3.4.2Financial reports cleaning 43 3.4.3Get readability measures 44 3.4.4Get future crash risk measures 44 3.5External analysis with clustering 45 4 experiments 47 4.1Case study 1: evaluate 6clusters of 48 readability measures 48 4.1.1Optimal cluster number evaluation - case study 149 4.1.2Clustering results - case study 152 4.1.3External analysis results - case study 155 4.2Case study 2: evaluate 5clusters of 43 readability measures 63 4.2.1Optimal cluster number evaluation - case study 263 4.2.2Clustering results - case study 267 4.2.3External analysis results - case study 270 4.3Discussion 78 5 conclusion 82 5.1Conclusions 82 5.2Prospect for future work 83 1.2. Objectives 3 companies’ future crash risk. Future crash risk metrics are often used to assess the companies’ financial performance. Realizing the relationship between present business components such as financial reporting readability, it is possible to build a crash risk forecasting model using readability indicators. 1.2 objectives As Li [2008], Lawrence [2013] and Yu and Miller [2010] research indicates, it is possible that readability measures can reveal insights about the current and future financial state of companies. These studies indicate, for example, that financial report texts can reveal components as the information asymmetry, the phenomenon that occurs when two or more economic agents hold different qualitative or quantitative information, defined in microeconomics as market failure, as well as it can describe the uncertainty of the information environment and the financial stress of companies. In this particular study, annual financial reports (10-K) will be analysed, concluding their readability from 48 different readability metrics (within our knowledge, there is no other study that analyses such large amount of readability measures). Afterwards, this work aims to evaluate the relationship of the readability of these reports and the future crash risk, obtained from the three measures. Readability metrics measure how readable the financial reports are and the future crash risk allows to evaluate the future company’s performance. In this project, future crash risk is obtained from the day following the publication of the financial report until the publication of the next report (next year). This study is intended to contribute to the increase of knowledge about the mentioned relationship. The following are the particular set of objectives that should be achieved. Table 1details the designed plan. 1. Business understanding: study the different readability and future crash risk metrics by reviewing their literature. 2. Data preparation: Financial reports will be obtained from US companies using the SEC EDGAR platform and financial information (used to calculate future crash risk metrics) will be obtained from the Center for Research in Security Prices (CRSP) databases and Professor K. French website [Kenneth,2019]. 3. Analyse all S&P 100 index companies, for 10 years. 4. Obtain 48 readability measures from this financial reports, analysing different components of their text. 1.3. Document organization 4 5. Obtain future crash risk measures from the day after the financial report is published until the day before the publication of the next financial report (in the next year). 6. Analyse the relationship between readability measures and future crash risk measures of the same company for the same year (relate the readability of a report to next year’s financial performance, starting the day after the financial report is published). The following is the plan to achieve the objectives of this work: 2019 12345678910 Business Understanding Data Understanding Study and obtain readability measures Study and obtain future crash risk measures Evaluate the relationship/methodology Writing of the dissertation project Table 1: Objectives in a Gantt chart. 1.3 document organization This document is organized in terms of the chapters: 1. Chapter 2: State of the art, which organizes and exposes the research found as well as it is intended to explain all 48 readability measures and their components, also showing their calculation formula, as seen in Section 2.2.4. It also refers and explains future crash risk measures and how financial information is used in their formulas. This section analyses recent research that demonstrates evidences of how readability can be related to financial components. 1.3. Document organization 5 2. Chapter 3: Development. This phase goes thorough the development process, explaining the necessary steps to analyses the relationship between the readability and future crash risk. It includes business understanding, data understanding, data preparation as well as the finally analysis, to evaluate the relationship, in terms of technical issues. 3. Chapter 4: Experiments, which is intended to evaluate the mentioned relationship, making it possible to draw conclusions about the contribution of this project in the financial and Data Science fields. 4. Chapter 5: Conclusions, which presents the main conclusions of the study and the various directions for future work. 2 S TAT E O F T H E A RT This chapter presents the main concepts approached in this work, also disclosing the relevant works in this domain. 2.1 defining readability There is no universal definition of readability [Loughran and McDonald,2014]. Some definitions focus solely on the style of writing, content, coherence and organization. For example: Klare et al. [1963] mentions it is ”the ease of understanding or comprehension due to the style of writing”. This definition focuses the majority on the clarity, transparency and accessibility of the text as well as the ease with which the readers can understand the text. On the other hand, Mc Laughlin [1969] and DuBay [2004] define readability as ”the degree to which a given class of people finds a reading attractive and understandable” focusing on the public-target, insofar as readability takes into account the characteristics of the readers and their limitations. Davison and Kantor [1982] emphasize that ”background knowledge assumed in the reader” is more important than ”trying to make a text fit a readability level defined by a formula”. 2.2 readability measures A readability measure aims to probabilistically measure how readable a text is. It allows to know the readability of a text and thus be aware of its complexity. This section will cover dozens of redability measures, describing them and discussing how they are calculated. In this project, a total of 48 readability measures were used to increase the efficiency of the approach. 6 2.2. Readability measures 7 2.2.1Word count and file size Some measures are based on the notion of overwriting: when documents are written with too much detail, using dispensable and superfluous words that makes the document too long and consequently difficult to process and understand, influencing readability [Bonsall IV et al.,2017]. Therefore, the use of attributes such as the number of words in the document [You and Zhang,2009] and the file size, (e.g., number of megabytes) are examples of overwriting measures [Loughran and McDonald,2014]. These measures may be inefficient because they capture other factors that do not affect readability: using the file size may consider parts of the document that are independent of financial information and that vary in all financial reports (e.g., pictures, bond indentures, compensation, supplier/customer agreements) [Loughran and McDonald, 2014]. Furthermore, these measures will be associated with a lack of interpretation of semantics, which impairs readability. 2.2.2Fog Index Gunning [1952] developed the Fog Index, considered by many researchers as a primary measure of readability, since it has a good measure of difficult complexity text and is easy to calculate and adaptable to the computational level through the Fog Index formula. The Fog Index has 2elements: the average sentence size and the percentage of complex words (i.e., words with three or more syllables). The final measure will be a sum of 2elements multiplied by a scalar (0.4). 0.4 hwords sentences +100 complexWords words i (1) Fog Index is a good sign of hard-to-read text but it has limitations: There are words that have three or more syllables (definition for complex words), which are not really difficult. For example, ”interesting” is not generally thought to be a difficult word, although it has four syllables (then classified as a complex word). Also, a short word can be difficult even if it is not used very often by most people, like the word “bilk”. The frequency with which words are in normal use affects the readability of text [Seely, 2013]. 2.2. Readability measures 8 2.2.3Bog Index Bonsall IV et al. [2017], recognizing the limitations of the Fog Index, developed a methodology capable of capturing the plain English attributes of disclosure, recommended by linguistics experts and highlighted in the SEC’s Plain English Handbook [SEC,1998]. An example of attributes to capture would be a sentence that contains ”I made an application”, which has a readability mistake termed by linguistics as a hidden verb, which should have been replaced by ”I applied”. Bog Index is computed as the sum of three multifaceted components: hBogIndex =SentenceBog +WordBog −Pepi(2) 1. Sentence Bog identifies subjects related to sentence size, such as the average sentence size, which is squared and scaled by a standard long sentence limit of 35 words per sentence. 2. Word Bog has 2main subcomponents: (a) plain English style problems and (b) word difficulty. The calculation of Word Bog is the sum of (a) plain English style problems and (b) word difficulty multiplied by 250 and divided by the number of words. The (a) plain English style problems is a combination of issues highlighted in the SEC’s Plain English Handbook [SEC,1998]: passive verbs, hidden verbs, overwriting, legal terms, cliches, abstract words and wordy phrases. The (b) word difficulty is calculated based on the difficulty for general vocabulary (i.e., heavy words), abbreviations and specialist terms [Bonsall IV et al.,2017]. 3. Pep identifies writing attributes that facilitate understanding of texts by readers, as interesting and easy words. Pep is calculated as the sum of this components multiplied by 25 (Word Bog are multiplied by 250) and scaled by the number of words in the document plus sentence variety (i.e., the standard deviation of sentence length multiplied by ten and scaled by the average sentence length) [Bonsall IV et al.,2017]. In contrast to the Fog Index, which considers a word with three or more syllables, the Bog Index measures word difficulty using a proprietary list of over 200,000 words based on familiarity and assesses penalties between zero and four points based on 2.2. Readability measures 9 a combination of the word’s familiarity and precision (abstract words receive higher scores). 2.2.4All 48 readability measures used in this study For this study, 48 readability measures were used, which analyze several components of textual complexity. From Table 2to Table 9, it is presented the 48 readability measures: a description about them, their calculation formula and the main bibliographic reference for the measure. These measures are not uniform, i.e., some measures suggest a higher readability when their values are higher, while others produce numbers in the opposite direction, with higher values indicating more difficult texts. Thus, we also present this information in the tables. Nevertheless, to facilitate the analysis of the results, all measures were rescaled in an uniform fashion, assuring that larger values are related with more readable documents. The rescaling was based on the simpler symmetrical transformation, where only the measures that are lower for more readable texts invert the signal. We adopted the following notation to define the readability measures: 1.nw= number of words; 2.nc= number of characters; 3.nst = number of sentences; 4.nsy = number of syllables; 5.nw f = number of words matching the Dale-Chall List of 3000 ”familiar words”; 6. ASL = Average Sentence Length: number of words / number of sentences; 7. AWL = Average Word Length: number of characters / number of words; 8. AFW = Average Familiar Words: count of words matching the Dale-Chall list of 3000 ”familiar words” / number of all words; and 9.nwd = number of ”difficult” words not matching the Dale-Chall list of ”familiar” words. 2.2. Readability measures 10 Measure name Description More readable when: ARI Automated Readability Index [Smith and Senter, 1967]. 0.5ASL +4.71AWL −21.34 Lower ARI.Simple A simplified version of Smith and Senter [1967] Automated Readability Index. ASL +9AWL Lower ARI.NRI Automated Readability Index [Smith and Senter,1967] with revised parameters from the Navy Readability Indexes. 0.4ASL +6AWL −27.4 Lower Bormuth.MC Bormuth [1969] Mean Cloze Formula. 0.886593 −0.03640 ∗AWL +0.161911 ∗ AFW −0.21401 ∗ASL −0.000577 ∗ ASL2−0.000005 ∗ASL3 Higher Bormuth.GP Bormuth [1969] Grade Placement score. 4.275 +12.881M−34.934M2+20.388M3+ 26.194CCS −2.046CCS2−11.767CCS3− 42.285(M∗CCS) + 97.620(M∗CCS)2− 59.538(M∗CCS)2 where M is the Bormuth Mean Cloze Formula as in ”Bormuth” above and CCS is the Cloze Criterion Score [Bormuth,1968]. Lower Coleman Coleman [1971] Readability Formula 1. 1.29 ∗(100 ∗nwsy=1/nw)−38.45 where nwsy=1=the number of words with 1syllable Higher Coleman.C2Coleman [1971] Readability Formula 2. 1.16 ∗(100 ∗nwsy=1/nw) + 1.48 ∗(100 ∗ nst/nw)−37.95 Higher Table 2: All 48 readability measures used in this study - Part 1of 8. 2.2. Readability measures 11 Measure name Description More readable when: Coleman.Liau .ECP Coleman-Liau Estimated Cloze Percent (ECP) [Coleman and Liau,1975]. 141.8401 −(0.214590 ∗100 ∗AWL) + (1.079812 ∗nst ∗100/nw) Higher Coleman.Liau .grade Coleman-Liau Grade Level [Coleman and Liau, 1975]. −27.4004 ∗Coleman.Liau.ECP/100 + 23.06395 Lower Coleman.Liau .short Coleman-Liau Index [Coleman and Liau,1975]. 5.88 ∗AWL + (0.296 ∗nst/nw)−15.8 Lower Dale.Chall The New Dale-Chall Readability formula [Chall and Dale,1995]. 64 −(0.95 ∗100 ∗nwd/nw)−(0.69 ∗ASL) Higher Dale.Chall.old The original Dale-Chall Readability formula [Dale and Chall,1948]. 0.1579 ∗100 ∗nwd/nw+0.0496 ∗ ASL[+3.6365] The additional constant 3.6365 is only added if (nwd/nw)>0.05. Lower Dale.Chall.PSK The Powers-Sumner-Kearl Variation of the Dale and Chall Readability formula [Powers et al., 1958]. (0.1155 ∗100 ∗nwd/nw) + (0.0596 ∗ASL) + 3.2672 Lower Danielson .Bryan Danielson and Bryan [1963] Readability Measure 1. (1.0364 ∗nc/nblank) + (0.0194 ∗nc/nst)− 0.6059 where nblank =the number of blanks. Lower Table 3: All 48 readability measures used in this study - Part 2of 8. 2.2. Readability measures 12 Measure name Description More readable when: Danielson .Bryan.2 Danielson and Bryan [1963] Readability Measure 2. 131.059 −(10.364 ∗nc/nblank)+(0.0194 ∗ nc/nst) where nblank =the number of blanks. Higher Dickes.Steiwer Dickes-Steiwer Index [Dickes and Steiwer,1977]. 235.95993 −(73.021 ∗AWL)−(12.56438 ∗ ASL)−(50.03293 ∗TTR) where TTR is the Type-Token Ratio. Higher DRP Degrees of Reading Power Bormuth [1969]. (1−Bormuth.MC)∗100 where Bormuth.MC refers to Bormuth [1969] Mean Cloze Formula. Lower ELF Easy Listening Formula [Fang,1966]. nwsy>=2/nst . where nwsy>=2the number of words with 2syllables or more. Lower Farr.Jenkins .Paterson Farr-Jenkins-Paterson’s Simplification of Flesch’s Reading Ease Score [Farr et al.,1951]. −31.517 −(1.015 ∗ASL) + (1.599 ∗ nwsy=1/nw) where nwsy=1=the number of one-syllable words Higher. Flesch Flesch’s Reading Ease Score [Flesch,1948]. 206.835 −(1.015 ∗ASL)−(84.6 ∗(nsy/nw)) Higher Flesch.PSK The Powers-Sumner-Kearl’s Variation of Flesch Reading Ease Score [Powers et al.,1958]. (0.0078 ∗ASL) + (4.55 ∗nsy/nw)−2.2029 Lower Table 4: All 48 readability measures used in this study - Part 3of 8. 2.4. Text Mining and NLP 19 2.4 text mining and nlp The growing volume of data generated and stored today is largely unstructured or semi-structured, with no clear organization and therefore cannot be used properly to impact a business. In this sense lies the importance of Text Mining, which aims to exploit a great slice of unstructured or semi-structured data: the textual data. As a result of the need to extract standards and knowledge of textual data, the concept of Text Mining has emerged. The nature of human communication is not trivial for algorithmic development [Gupta et al.,2009]: a word may have different meaning when applied in different contexts or even semantically similar phrases may have words with opposite meanings are some examples. Thus, when the goal is to apply Text Mining to understand the meaning of textual data, the challenge increases substantially in the case of NLP.NLP is the area responsible for understanding the natural language, extracting its computational meaning and generating natural language. However, [Gupta et al.,2009] points out that, despite these difficulties, computers have the advantage of being able to process large volumes of text, unlike humans, which gives many opportunities to explore and retrieve relevant information. 2.4.1Related data analysis areas This work is set within the context of several data analysis areas, namely: 1. Data Science is an interdisciplinary field that involves areas such as Data Mining, Data Visualization and Data Analytics. The term covers any set of techniques that aims at extracting data insights, with different purposes: descriptive, predictive, or prescriptive analysis [O’Neil and Schutt,2013] and understands the whole process from the acquisition of unstructured data to the formation of data products or models [Larson and Chang,2016]. 2. Data Mining is a subprocess in the Data Science pipeline and requires the understanding of the business and data, its preparation, as well as the development and evaluation of the model [Witten et al.,2016]. The purpose of the model is the discovery of new information or relationship between existing information [Sharma et al.,2018] and oftentimes uses Statistics and Machine Learning approaches [Larson and Chang,2016]. 2.4. Text Mining and NLP 20 Figure 1: Related areas of Text Mining and NLP. 3. Machine Learning involves the study of algorithms that can extract information automatically (i.e., without direct human guidance). A Machine Learning program learns some task from experience: its performance improves with new data according to some performance measure. Some of these procedures include ideas derived from statistics. Fig. 1attempts to explain intuitively the relationship between these areas, showing their interceptions. Check that dashed lines are intended to mean an unclear boundary. 2.4.2Topic modelling Topic modeling is a type of statistical modeling for discovering similarities in unlabelled texts. Unsupervised Latent Dirichlet Allocation [Blei et al.,2003] is a hierarchical Bayesian model, where topic proportions for a document are drawn from a Dirichlet distribution and words in the document are repeatedly sampled from a topic which itself is drawn from those topic proportions [Zhu et al.,2009]. Usually, studies about topic modelling follow unsupervised approaches, but not all. For example, Ramage et al. [2010] improved the model and made it an supervised approach, learning from training data. Conditional Random Fields [Nikfarjam et al.,2015] is a method that can take context into account, whereas a discrete classifier predicts a label for a single sample without considering neighboring samples. 2.4. Text Mining and NLP 21 2.4.3Classification Text classification is the Text Mining step that addresses Machine Learning algorithms that aims to discriminate or characterize a piece of text in a particular format value. This value can vary from a number (sentiment analysis), labels (multi-labeling tasks), classes (binary or multi-class tasks). Some examples are: Sentiment analysis, whose goal is to identify the polarity of text content: the type of opinion it expresses. This may take the form of a binary like/dislike rating, or a more granular set of options, such as a star rating from 1to 5. Some commonly adopted Machine Learning algorithms for text classification are: 1. A Naive Bayes Classifier makes assumptions about how the data (in this case words in documents) is generated and proposes a probabilistic model based on these assumptions. It will then use a set of training examples to estimate the parameters of the model. Bayes rule is used to classify new examples and select the class that most likely has generated the example [Chakrabarti et al.,1997, Allahyari et al.,2017]. 2. Nearest Neighbor Classifier is a proximity classifier which uses distance measures to perform the classification. The idea is that documents which belong to the same class are more close to each other based on the similarity measures. The classification of the test document is inferred from the class labels of the similar documents in the training set. If is considered the k-nearest neighbor in the training data set, the approach is called k-nearest neighbor classification and the most common class from these k neighbors is reported as the class label [Han et al.,2001,Allahyari et al.,2017]. 3. Decision tree is essentially a hierarchical tree of the training instances, in which a condition on the attribute value is used to divide the data hierarchically. In other words, the decision tree recursively partitions the training data set into smaller subdivisions based on a set of tests defined at each node or branch. Each node of the tree is a test of some attribute, and each branch descending from the node corresponds to one the value of this attribute. An instance is classified by beginning at the root node, testing the attribute by this node and moving down the tree branch corresponding to the value of the attribute in the given instance. This process is then recursively repeated [Allahyari et al.,2017]. 2.5. Future crash risk 22 4. Support Vector Machines (SVM) are a supervised learning classification algorithms where have been extensively used in text classification problems. SVM are a form of linear classifiers. Linear classifiers in the context of text documents are models that making a classification decision is based on the value of the linear combinations of the documents features [Allahyari et al.,2017]. 2.5 future crash risk Future crash risk aims to probabilistically measure the risk of a crash. It is essential for assessing the financial health of companies and therefore helps investors make investment decisions. 2.5.1What is a crash? A crash is a sudden and significant decline in the price of an asset. It is both an economic and a psychological phenomenon, i.e., economically there may be a fall in value perceived by some investors that causes a sudden fall in value that psychologically affects other investors, leading them to follow the trend. Most investors, even if they have no economic reason to sell their shares, are driven to sell for fear that they will lose even more value and are clearly influenced psychologically by the sudden decline. As investors who prefer to sell their shares rather than hold them, investors who could buy also decide not to buy, for the same reasons, thus leading to a growing volume of available stocks with downward demand [Johansen et al.,1999]. For the psychological reasons already mentioned, lack of demand generates more lack of demand: investors will sell their shares and precipitate other investors to sell their shares as well. These investors, who sell their shares after noticing their rapid devaluation, act out of concern that their prices will fall further. Thus, this phenomenon leads to the possibility of a vicious cycle marked by negative crowd behavior. 2.5.2Reasons for a higher future crash risk The causes of a crash are not deterministic, i.e., there is no set of patterns that clearly identify a crash. However, there are a number of factors where future crash risk is greatest, including: 2.5. Future crash risk 23 1. Hide negative information that when released to the market causes extremely negative reactions [Kim et al.,2011]. 2. The stock price is inflated [Kim and Zhang,2016], leading to the creation of a bubble, i.e., the stock price is rising and when the market realizes that it is inflated, the price will go down dramatically, possibly causing a crash. 3. Emerging economies, as emerging equity markets are characterized by excessive volatility and are more likely to have weak corporate governance and may be more susceptible to crashes [Vo,2019]. 4. Companies where institutional investors are most distracted and ignore the attention-grabbing exogenous events [Xiang et al.,2019]. There is an aggravation when companies are state-owned, when CEOs control directors and when there is less coverage by analysts. Some authors will also identify features in stock returns that can lead to a crash, among them the variance of conditional volatility and the fat tailed distribution of the return series [Bates,2012]. There is also evidence that some factors can considerably mitigate the risk of a crash, such as factors linked to the reputations of senior management. 2.5.3Future crash risk measures Since future crash risk is a probabilistic value that measures the likelihood of a crash happening, there are ways to measure it. There are measures created by various authors that allow to estimate future crash risk, for example [Callen and Fang,2015]: 1.NCSKEW: The Negative Coefficient of Skewness of firm-specific daily returns 2.DUVOL: The Down-to-Up Volatility of firm-specific daily returns. 3. Crash Count: The difference between the number of days with negative extreme firm-specific daily returns and positive extreme firm-specific daily returns. To calculate these measures, it is necessary to calculate the firm-specific residual daily returns using the expanded market and industry index model for each firm and year [Hutton et al.,2009]: rj,t=αj+β1,jrm,t−1+β2,jri,t−1+β3,jrm,t+β4,jri,t+εj,t(3) 2.5. Future crash risk 24 where ri,tis the return on the value-weighted industry index based on 2-digit Standard Industrial Classification SIC codes on day t, rm,tis the return on the CRSP value-weighted market index on day t and rj,tis the return on stock j on day t. The firm-specific daily return, Rj,t, is the natural log of (1plus the residual return from equation 3). Log transforming raw residual returns is used to reduce the positive skew in the return distribution and to help ensure symmetry [Chen et al.,2001]. Thus, the mathematical approach of the measures emerges as: 1.NCSKEW is calculated as the negative of the third moment of each stock’s firm-specific daily returns, divided by the cubed standard deviation [Callen and Fang,2015]. So, for any stock j over the fiscal year T, NCSKEWj,T=−(n(n−1)3 2∑R3 j,t) ((n−1)(n−2)(∑R2 j,t)3 2)(4) where n is the number of observations of firm-specific daily returns during the fiscal year T. An increase in the value of NCSKEW corresponds to more likely crashes. 2.DUVOL, is calculated as [Callen and Fang,2015]: DUVOLj,T=log ((nu−1)∑DOWN R2 j,t (nd−1)∑UP R2 j,t)(5) where nuand ndare the number of up and down days over the fiscal year T, respectively. For this measure, it is necessary to separate the days with firm-specific daily returns above and below the period average (e.g., 1year), shown in the formula as ”up” and ”down”. It is also necessary to calculate the standard deviation for the up and down samples separately and then calculate the logarithmic ratio of the standard deviation of the ”down” sample to the standard deviation of the ”up” sample. A higher value of DUVOL corresponds to more crash-prone stock. 3. Crash Count, the downside frequencies minus the upside frequencies, uses the number of firm-specific daily returns exceeding standard deviations of 3.09 above and below the firm-specific average daily return over the fiscal year. Stan- 2.5. Future crash risk 25 dard deviation 3,09 was used by Hutton et al. [2009] to generate frequencies of 0.1% in the normal distribution. A higher value of Crash Count corresponds to a higher frequency of crashes. 3 DEVELOPMENT This chapter details all the methodology and development steps to achieve the objectives presented in Section 1.2. 3.1 methodology 3.1.1The problem and its challenges In this project, it is intended to focus on the relationship between the readability measures of financial reports and the correspondent companies’ crash risk for each year, which brings several challenges, which are: 1. Understand the financial reports structure and financial information needed to achieve future crash risk. 2. All the data preprocessing, from its unstructured state to its complete structuring. Because the source data is textual, contains formations and a set of anomalies, this step is very laborious and therefore appropriate methodologies should be chosen. 3. Get readability and future crash risk measures. 4. Understanding the relationship between the readability measures of financial reports and the correspondent companies’ future crash risk. 26 3.1. Methodology 27 3.1.2Proposed Approach The proposed solution is based on a Cross-Industry Standard Process for Data Mining (CRISP-DM) methodology adapted to this project. CRISP-DM originally has 6phases [Chapman et al.,2000]: 1. Business Understanding; 2. Data Understanding; 3. Data Preparation; 4. Modeling; 5. Evaluation; and 6. Deployment. In this project, it is not intended to reach the deployment phase but essentially to evaluate the relationship between the readability measures of financial reports and the correspondent companies’ future crash risk. The adapted CRISP-DM methodology is depicted in Fig. 2: 3.1. Methodology 28 Figure 2: Proposed approach. 3.3. Data understanding 35 4. Industry SIC codes In order to identify the industry of a company and for each day, it is necessary to convert the SIC code to the industry, which is made from this dataset. Thus, the variable ”min” and ”max” identify the minimum and maximum SIC to be able to belong to one industry and the variable ”number” joins all sub-industries of a larger industry, for example, the first 5industries belong to a bigger industry. termed ”Agric”. This dataset was collected from the CRSP database. A sample of this dataset is shown in Table 13. Table 13: Sample of the ”industry SIC codes” data. 3.3. Data understanding 36 5. Industry values The SIC for a company and day, obtained in with the ”SIC codes per period” dataset, is used to identify the industry (the translation is made using the previous dataset) and thus access the return on the value-weighted industry index, obtained from this dataset. The matrix is accessed by checking the day and the name of the industry. This dataset was collected from the CRSP database. A sample of this dataset is presented in Table 14. Table 14: Sample of the ”industry values” data. Showing 8of 49 industries. 3.3. Data understanding 37 6. Fama and French factors The Fama and French factors dataset has market indicators, which will be used to calculate the value-weighted market index, which is essential for calculating future crash risk metrics, as we will see later. This dataset is taken from Professor K. French website [Kenneth,2019]. A sample of this dataset is in Table 15. Table 15: Sample of the ”Fama and French factors” data. 3.4. Data preparation 38 3.4 data preparation This step is of utmost importance as data preparation defines much of the success of the next steps and is one of the most laborious steps. The datasets collected in the data collection phase are: 1. Financial reports; 2. Stocks daily; 3. Stocks sample; 4.SIC codes per period; 5. Industry SIC codes; 6. Industry values; and 7. Fama and French factors. 3.4.1Select the data At the beginning of this phase, all the financial reports of all US listed companies from 2006 to 2018 were available. Since a methodology is being tested, a portion of this data was selected in order to obtain an interesting coherent experimental setup. The annual financial reports (10-K) of the S&P 100 companies were selected during a ten year period, related to the years of 2008 to 2017. The S&P 100 index is a subset of the S&P 500 index and measures the performance of the largest capitalized companies in the United States. This index includes about 100 of the largest US companies from various industries [Ishares,2019]. Several investors decide to invest in specific indexes, thus investing in all the companies contained in it, being the S&P 100 one of the most popular US indices. The indexes are quite useful for some investors, depending on their investment strategy, because it allow to mitigate the risk (but also decrease the earnings) since if a company crashes or falls in the value of its shares, remaining index companies can mitigate this loss. However, if one company has a high appreciation the others will mitigate these gains too. In Fig. 3 and Fig. 4, we can see the value of the shares of the S&P 100 index in the last 10 and 30 years, respectively, with monthly plots. 3.4. Data preparation 39 Figure 3:10 years S&P 100 Stocks Chart (monthly plots), adapted from BarChart [2019]. Figure 4:30 years S&P 100 Stocks Chart (quarter plots), adapted from BarChart [2019]. 3.4. Data preparation 40 There are several types of financial reports, varying in terms of time frame and purpose, including: 1.10-K reports: These are annual financial information and are mandatory [Investor.gov,2019]. 2. Report 10-Q: Quarter financial information that is not audited [sec,2019a]. 3.10-KSB Report: Similar to type 10-K, it differs only by targeting small businesses (under $ 75M in stock and $ 50 million in revenue). This type of report has been terminated by the SEC, forcing entities to report in 10-K [sec,2019b]. Initially, the financial reporting dataset consisted of folders sorted in (1) years, (2) companies and (4) quarters where each quarter’s reports were. This setting is shown in Fig. 5. Since the sample of this project uses, as explained, the annual financial reports (10-K) of the S&P 100 companies for 10 years, from 2008 to 2017 being chosen, Python scripts were used to implement the following tasks: 1. The directory structure has been changed, as seen in Fig. 6.10-K reports could be in any quarter folder. 2. Select only annual reports (10-K), which is the focus of the investigation. This filter was obtained by comparing a section of the name of each report where it indicates the type of report. 3. Select only financial reports from S&P 100 constituent companies. This filter was obtained by obtaining the CIK in the name of each financial report document, subsequently identifying whether the CIK belonged to the S&P 100 index company and thus moving to the new folder. Thus, these scripts facilitated the automation of the data extraction procedure. For future work, it is possible to look at other years or other companies using the same scripts. Thus, the reproduction of research is greatly facilitated. 3.4. Data preparation 41 Figure 5: Initial directory of the financial reporting dataset. Figure 6: Final directory of the financial reporting dataset. 3.4. Data preparation 42 Data selection by selecting 10 years (2008-2017) of the S&P 100 (total of 102 companies) should yield 1020 (102 *10) financial reports. Each company should have one annual report per year, totaling 10 financial reports per company. However, there are a total of only 921 financial reports, as seen in Table 16 and Table 17, where it can be seen the number of financial reports by company and by year, respectively. The difference is due to the recent inclusion of some companies in the index or to the stock exchange. This fact was expected and was taken into account. Table 16: Number of financial reports per company. Table 17: Number of financial reports per year. 3.4. Data preparation 43 3.4.2Financial reports cleaning Since the purpose of this project is to analyze the readability of the financial reports text, it is relevant to clean these reports, eliminating all material not necessary for the readability measurement. Much of the content of financial reporting files consists of HTML code, embedded PDF and other artifacts that are of no interest to our purposes [McDonald,2019]. This content is not only unnecessary and may even prejudice the results. Moreover it also requires a high memory. For example, some records ed initially more than 400MB, but after cleaning the size of the data was reduced to just 5KB. The steps to clean the financial reports are listed below [McDonald,2019]: 1. Elimination of document segment <TYPE>tags of ZIP, EXCEL GRAPHIC, JSON and PDF. 2. Elimination of <FONT>,<DIV>,<TR>and <TD>tags. 3. Elimination of all XML code. 4. Elimination of all XBRL segments. 5. Elimination of SEC Header and Footer. 6.\&NBSP and \ replacement with a blank space. 7.\& and \& replacement with ”&”. 8. Elimination of all remaining extended character references. 9. Elimination of all tables. Sometimes some paragraphs are defined with table tags, so each table segment is first stripped of all HTML and then the number of numeric characters against alphabetic characters is compared. Numeric characters / (alphabetic + numeric characters) are calculated and table segments are removed where the result is greater than 10. 10. Translation of exhibits with original tags of ”<TYPE>EX-##” to: <EX-##> . . . original text </Ex-##> 11. Elimination of all remaining markup tags. 3.4. Data preparation 44 12. Elimination of the excess linefeeds. 13. The HEADER component is eliminated from our analysis. Clean financial reports were downloaded from the McDonald [2019] source. 3.4.3Get readability measures At this stage, the aim is to obtain reliable readability measures for financial reports. The cleanup of the financial reports has been completed in Section 3.4.2and is essential for this phase since no factors outside the meaning of the text are intended to interfere with this assessment. The following are the steps that were required to obtain readability measures: 1. Read company financial reports for variables, one company at a time, not overloading RAM, by doing the following items and then loading new reports, repeating the process, greatly increasing the speed and decreasing the waste of RAM. 2. Obtain the 48 readability measures for each financial report from the Quanteda library for the R language. These measures are explained in Section 2.2. The name and description of the measures can be seen, for example, in Section 2.2.4. 3. For each row containing the readability measures of a report, also assign the CIK, company name, year and full date, obtaining this information directly from the report name. 4. Save file to rds and csv. 3.4.4Get future crash risk measures To examine the effect of readability on corporate financial health, the future crash risk indicator discussed in Section 2.5was used. Thus, several firm-specific future crash risk measures for each year were constructed which are technically explained in Section 2.5.3. This section aims to explain the steps for obtaining these metrics. To this end, it was essential to collect and use several datasets, explained in Section 3.3.3. The algorithmic steps to obtain future crash risk measures include: 4.1. Case study 1: evaluate 6clusters of 48 readability measures 51 Figure 7: Selection of the number of clusters using the Gap Statistic - case study 1. 4.1. Case study 1: evaluate 6clusters of 48 readability measures 52 4.1.2Clustering results - case study 1 This section presents the results for the 6readability clusters. For better visualization and discussion, two approaches were made, one focusing on all readability measures, using radar charts (Fig. 9) and another focused on decreasing the complexity of the 48 readability measures, reducing them to only 2measures, using PCA method (Fig. 8). While radar charts show the results of all 48 readability measures, allowing to discuss and draw accurate conclusions, PCA allows to have a more brief and simpler visualization of the readability clusters. PCA is a statistical procedure that uses an orthogonal transformation to convert a set of observations from possibly correlated variables (several of the 48 measures are correlated) into a set of linearly uncorrelated variable values called principal components, decreasing the initial number of features [Wold et al.,1987]. Analyzing Fig. 8, it can be seen the result of the PCA method, visualizing the result of the clustering method for readability measures in a simpler way. Each of the two variables (Dim1and Dim2) represents a certain amount of information contained in the 43 readability measures. The proportion of which is represented by the percentage in the axis labels. Dim1has 68.6% of the information contained in the initial 48 readability measures and Dim2 contains only 22.1%, noting that Dim1is responsible for more information. It is also noted that 68.6% + 22.1% = 90.7%, meaning that 9.3% (100% - 90.7%) of the information in this graph has been lost, it is natural since the reduction in the size of the data was large, from 43 features to just 2features. However, visually it is very worthwhile, since it is clear that at least for the 90.7% of information, the clusters are distinctly different. Fig. 9shows the 6clusters on 4radar charts (12 readability measures per chart, 48 in total), allowing to analyze the 4graphs at the same time. Ideally, different clusters should be found, each one corresponding to a particular set of readability measures that, as will be assessed in the next section (where external analysis results are analyzed), would match to a specific future crash risk level (e.g., it would be interesting to find a readability cluster that is often associated with a high degree of future crash risk). Thus, in this section we first want to show and analyze if the clustering method was effective. Taking this into consideration, it appears that the 6clusters have obvious differences and there is no cluster that has a similar pattern to others for the 43 readability measures analysed in Fig. 9. The clusters analyzed in Fig. 9show a differentiation of them, considering the 43 readability measures visualized in the 4radar charts of this figure. However, due to 4.1. Case study 1: evaluate 6clusters of 48 readability measures 53 Figure 8:PCA results. The points inside the clusters represent the financial reports - case study 1(with 6clusters). 4.1. Case study 1: evaluate 6clusters of 48 readability measures 54 Figure 9: Radar charts - case study 1(with 6clusters). The closer to the outer edge of the chart a measure is, the more readable the cluster is. 4.1. Case study 1: evaluate 6clusters of 48 readability measures 55 the number of features, the view in Fig. 8becomes more favorable, clearly seeing totally distinct clusters. In this sense, the PCA was of great importance, reducing the initial spectrum from 48 characteristics to only 2, Dim1, containing 68.6% of the initial information of the 43 readability measures and Dim2, containing 22.1% of this information, meaning that Dim1represents most of the information obtained in this graph. This view proves the success of the k-means clustering method by splitting the clusters properly. 4.1.3External analysis results - case study 1 After observing and analyzing the clusters for the 48 readability measures in Section 4.1.2, it was concluded that it created suitable groups. Taking this into consideration, it is now possible to analyze the clusters in terms of future crash risk, evaluating future crash risk measures in each cluster, externally. Readability metrics measure how readable the financial reports are and the future crash risk allows to evaluate the future company’s performance. In this project, future crash risk is obtained from the day following the publication of the financial report until the publication of the next report (next year). Thus, this phase goes through characterizing the readability clusters with future crash risk and realizing what distinguishes them. As explained, the financial characteristic analyzed in this project is the future crash risk. The financial graphs in Fig. 10, Fig. 11 and Fig. 12 show the future crash risk measures of the 6clusters analyzing 2future crash risk measures per graph. The size of the circles is proportional to the size of the clusters. Therefore, it is possible to have the notion of the size of each cluster from these views, also using this feature for their comparison. It is also possible to visualize the correlation between the different future crash risk measures in Table 25 and by cluster in Table 20. The following points list commonalities between readability and future crash risk measures, taking into account the PCA method. 1. Focusing on the readability clusters depicted in Fig. 8, representing the clusters obtained from readability measures, it can be seen that Cluster 2is isolated from others, in Dim1. As already explained (Section 4.1.2), the method PCA, represented in this figure (Fig. 8), contains 2features (resulting from the initial 48 readability measures), Dim1and Dim2, as depicted in the image. Dim1contains 68.6% of the initial information of the 48 readability measures and Dim2contains 22.1% of the information, meaning that Dim1represents most of the information 4.1. Case study 1: evaluate 6clusters of 48 readability measures 56 Figure 10:NCSKEW and DUVOL per cluster - case study 1(with 6clusters). Circles are proportional to the size of the cluster. Figure 11:NCSKEW and Crash Count per cluster - case study 1(with 6clusters). Circles are proportional to the size of the cluster. 4.1. Case study 1: evaluate 6clusters of 48 readability measures 57 Figure 12:DUVOL and Crash Count per cluster - case study 1(with 6clusters). Circles are proportional to the size of the cluster. obtained in this graph. Therefore, this cluster is isolated in the most of the information, Dim1. Analysing the 3future crash risk metrics (NCSKEW,DUVOL and Crash Count), in Fig. 10, Fig. 11 and Fig. 12, Cluster 2always appears further away from all other clusters. Cluster 2has the lowest future crash risk mean (-0.54, as seen in Table 19) and the lowest NCSKEW,DUVOL and Crash Count. This is the cluster that identifies companies with the lower future crash risk. As mentioned, this cluster also has a differentiation in most readability information (68.6%), as seen in Fig.8, therefore this may indicate a relationship between some readability measures (68.6%) and future crash risk measures, but it does not show what type of relationship is (how do the measures affect each other). 2. It can also be seen from Fig. 8that clusters 1and 3, represented by red and yellow, respectively, are more or less aligned in the Dim1. As explained, Dim1 contains more information than Dim2and therefore it is more representative: Dim1contains 68.6% of the information and Dim2contains 22.1% of the information. These clusters (1and 3) are similar in Dim1(most of the information). 4.1. Case study 1: evaluate 6clusters of 48 readability measures 58 Table 19: Cluster size and external features mean - case study 1(with 6clusters). Table 20: Correlation of future crash risk measures per cluster - case study 1(with 6clusters). Analysing future crash risk, in Table 19 and specifically the Crash Count metric, these clusters have a very close value too (both have -0.2and in this metric all clusters values vary between -0.1and -1.1). Other relationships can be seen by analyzing the cluster rankings for readability measures (these can be grouped into similar groups) and then analyzing the cluster rankings for future crash risk measures. Comparing both rankings is important, since a similar clustering order between several readability measures and some future crash risk measures can be a great indicator of the relationship between these measures. Thus, since 48 readability measures are analysed, it is intended to group the measures in similar groups and considering their cluster order. After this, this groups will be compared with the results of future crash risk rankings by measure. These groups can be compared in terms of readability and future crash risk at the same time for some clusters, as it will be seen. Then, it is presented the readability groups and analyzed the points that these groups have in common considering their calculation formula, using the information from the Table 2to the Table 9: 1. Group 1(measures sorted in descending order in Table 21): 33 metrics including ARI family (3metrics), Bormuth family (2metrics), Dale.Chall, Dale.Chall.old, 4.1. Case study 1: evaluate 6clusters of 48 readability measures 59 Danielson.Bryan, Dickes.Steiwer, DRP, ELF, Flesch.PSK, Flesch.Kincaid, FOG family (3metrics), Farr.Jenkins.Paterson, Fucks, Linsear.Write, LIW, nWS.3, nWS.4, RIX, SMOG family (4metrics), Spache family (2metrics), Strain, Traenkle.Bailer, Wheeler.Smith and meanSentenceLength. This group contains major emphasis on Average Sentence Length (ASL) and Average Word Length (AWL), also considering the size of syllables of a sentence, nwsy, (e.g., ELF, Farr.Jenkins.Paterson, Flesch, Flesch.PSK, FOG), with different variations (e.g., word count with 1,2or 3syllables per word, depending on the metric). Less frequently, another components are also used as: number of ”difficult” words not matching the Dale-Chall list of ”familiar” words, nwd (e.g., Dale.Chall.PSK), number of words, nw(e.g., Dale.Chall. PSK, Flesch, Flesch.PSK), number of characters, nc(e.g., Danielson.Bryan), the number of blanks, nblank (e.g., Danielson.Bryan) and number of sentences, nst (e.g., Danielson.Bryan, ELF, FOG). See notation in Section 2.2.4. 2. Group 2(measures sorted in descending order in Table 22): 4metrics including Dale.Chall.PSK, Flesch, nWS and nWS.2. The components of this group are similar to that of Group 1, but also using the number of characters per word, nwchar, (e.g., nWS and nWS.2using the number of words with 6characters or more, nwchar>=6) and without using Average Word Length (AWL), in contrast to Group 1which has a great use of this component. 3. Group 3(measures sorted in descending order in Table 23): 11 metrics including Coleman family (5metrics), Danielson.Bryan.2, FORCAST family (2metrics), Scrabble, Traenkle.Bailer.2and meanWordSyllables. This group does not use Average Sentence Length (ASL), unlike previous groups. This group is based on nwsy (e.g., Coleman, Coleman.C2, FORCAST), nw(e.g., Traenkle.Bailer2, Coleman.Liau.ECP, FORCAST), nst (e.g., Danielson.Bryan2). Other less commonly used components are the number of prepositions, nconj (e.g., Traenkle.Bailer2) and the number of conjunctions, nnconj (e.g., Traenkle.Bailer2). See notation in Section 2.2.4. Looking at Table 19, it is now possible to ranking the clusters regarding the future crash risk metrics. The three future crash risk metrics have some consistency in the order they define clusters, since they are partially correlated, as shown in Fig. 20. However, they all have a slightly different order. The cluster rankings of the future crash risk measures and their mean are shown: 4.1. Case study 1: evaluate 6clusters of 48 readability measures 60 Table 21: Readability measures Group 1for case study 1(with 6clusters) Table 22: Readability measures Group 2for case study 1(with 6clusters) Table 23: Readability measures Group 3for case study 1(with 6clusters) 4.2. Case study 2: evaluate 5clusters of 43 readability measures 67 4.2.2Clustering results - case study 2 This section presents the results for the 5readability clusters. For better visualization and discussion, two approaches were made, one focusing on the 43 readability measures, using radar charts (Fig. 15) and another focused on decreasing the complexity of the 43 readability measures, reducing them to only 2measures, using PCA method (Fig. 14). While radar charts show the results of the 43 readability measures, allowing to discuss and draw accurate conclusions, PCA allows to have a more brief and simpler visualization of the readability clusters. Analyzing Fig. 14, it can be seen the result of the PCA method, visualizing the result of the clustering method for readability measures in a simpler way. Each of the two variables (Dim1and Dim2) represents a certain amount of information contained in the 43 readability measures. The proportion of which is represented by the percentage in the axis labels. Dim1has 68.6% of the information contained in the initial 48 readability measures and Dim2contains only 22.5%, noting that Dim1is responsible for more information. It is also noted that 68.6% + 22.5% = 91.1%, meaning that 8.9% (100% - 91.1%) of the information in this graph has been lost, it is natural since the reduction in the size of the data was large, from 48 features to just 2. However, visually it is very worthwhile, since it is clear that at least for the 91.1% of information, the clusters are distinctly different. Fig. 15 shows the 5clusters on 4radar charts (11 readability measures in 3charts and one chart with 10 measures, 43 in total), allowing to analyze the 4graphs at the same time. The closer to the outer edge of the chart a measure is, the more readable the cluster is. Ideally, different clusters should be found, each one corresponding to a particular set of readability measures that, as will be assessed in the next section (where external analysis results are analyzed), would match to a specific future crash risk level (e.g., it would be interesting to find a readability cluster that is often associated with a high degree of future crash risk). Thus, in this section, it is first shown and analyzed if the clustering method was effective. Taking this into consideration, it appears that the 5clusters have obvious differences and there is no cluster that has a similar pattern to others for the 48 readability measures analysed in Fig. 15. 4.2. Case study 2: evaluate 5clusters of 43 readability measures 68 Figure 14:PCA results. The points inside the clusters represent the financial reports - case study 2(with 5clusters). 4.2. Case study 2: evaluate 5clusters of 43 readability measures 69 Figure 15: Radar charts with 43 readability measures - case study 2(with 5clusters). No FOG, sentence length, or word syllables measures. The closer to the outer edge of the chart a measure is, the more readable the cluster is. 4.2. Case study 2: evaluate 5clusters of 43 readability measures 70 The clusters analyzed in Fig. 15 show a differentiation of them, considering the 43 readability measures visualized in the 4radar charts of this figure. However, due to the number of features, the view in Fig. 14 becomes more favorable, clearly seeing totally distinct clusters. In this sense, the PCA was of great importance, reducing the initial spectrum from 48 characteristics to only 2, Dim1, containing 68.6% of the initial information of the 48 readability measures and Dim2, containing 22.5% of this information, meaning that Dim1represents most of the information obtained in this graph. This view proves the success of the k-means clustering method by splitting the clusters properly. 4.2.3External analysis results - case study 2 After observing and analyzing the clusters for the 43 readability measures in Section 4.2.2, it was concluded that it created suitable groups, it is intended to compare these clusters also in terms of future crash risk, externally. Taking this into consideration, it is now possible to analyze the clusters in terms of future crash risk, evaluating future crash risk measures in each cluster (external analysis). Readability metrics measure how readable the financial reports are and the future crash risk allows to evaluate the future company’s performance. In this project, future crash risk is obtained from the day following the publication of the financial report until the publication of the next report (next year). Thus, this phase goes through characterizing the readability clusters with future crash risk and realizing what distinguishes them. As explained, the financial characteristic analyzed in this project is the future crash risk. The financial graphs in Fig. 16, Fig. 17 and Fig. 18 show the future crash risk measures of the 5clusters analyzing 2future crash risk measures per graph. The size of the circles is proportional to the size of the clusters. Therefore, it is possible to have the notion of the size of each cluster from these views, also using this feature for their comparison. It is also possible to visualize the correlation between the different future crash risk measures in Table 25 and by cluster in Table 20. The following points list commonalities between readability and future crash risk measures, taking into account the PCA method. 1. Focusing on the readability clusters depicted in Fig. 14, representing the clusters obtained from readability measures, it can be seen that Cluster 3is isolated from others, in Dim1. As already explained, (Section 4.2.2and previous experiment, Section 4.1), the method PCA, represented in this figure (Fig. 14), contains two 4.2. Case study 2: evaluate 5clusters of 43 readability measures 71 Table 25: Correlation between future crash risk measures. Figure 16:NCSKEW and DUVOL per cluster - case study 2(with 5clusters). Circles are proportional to the size of the cluster. 4.2. Case study 2: evaluate 5clusters of 43 readability measures 72 Figure 17:NCSKEW and Crash Count per cluster - case study 2(with 5clusters). Circles are proportional to the size of the cluster. Figure 18:DUVOL and Crash Count per cluster - case study 2(with 5clusters). Circles are proportional to the size of the cluster. 4.2. Case study 2: evaluate 5clusters of 43 readability measures 73 features (resulting from the initial 43 readability measures analysed in this experiment), Dim1and Dim2, as depicted in the image. Dim1contains 68.6% of the initial information of the 43 readability measures and Dim2contains 22.5% of the information, meaning that Dim1represents most of the information obtained in this graph. Therefore, this cluster is isolated in the most of the information, Dim1. Analysing the 3future crash risk metrics (NCSKEW,DUVOL and Crash Count), in Fig. 16, Fig. 17 and Fig. 18, Cluster 2always appears further away from all other clusters. Cluster 2has the lowest future crash risk mean (-0.45, as seen in Table 26) and the lowest NCSKEW,DUVOL and Crash Count. As mentioned, this cluster also has a differentiation in most readability information (68.6%), as seen in Fig.14, therefore this may indicate a relationship between some readability measures (68.6% of the information) and future crash risk measures, but it does not show what type of relationship is (how do the measures affect each other). 2. Accordingly, and for the same reasons as the previous paragraph, cluster 4and cluster 5are also about the same value as Dim1in Fig. 14, and these clusters are therefore quite similar (similar in 68.6% of readability information, as explained in the previous point). In Fig. 16, Fig. 17 and Fig. 18, it can be noticed that these clusters are also very close in future crash risk measures, so these clusters are very similar in readability and future crash risk. Other relationships can be seen by analyzing the cluster rankings for readability measures (these can be grouped into similar groups) and then analyzing the cluster rankings for future crash risk measures. Comparing both rankings is important, since a similar clustering order between several readability measures and some future crash risk measures can be a great indicator of the relationship between these measures. Thus, since 43 readability measures are analysed, it is intended to group the measures in similar groups and considering their cluster order. After this, this groups will be compared with the results of future crash risk rankings by measure. These groups can be compared in terms of readability and future crash risk at the same time for some clusters, as it will be seen. Then, it is presented the readability groups and analyzed the points that these groups have in common considering their calculation formula, using the information from the Table 2to the Table 9: 1. Group 1(measures sorted in descending order in Table 28): 37 measures including ARI family (3metrics), Bormuth family (2metrics), 4.2. Case study 2: evaluate 5clusters of 43 readability measures 74 Table 26: Cluster size and external features mean - case study 2(with 5clusters) Table 27: Correlation of future crash risk measures per cluster - case study 2(with 5clusters) Coleman.C2, Coleman.Liau.ECP, Coleman.Liau.grade, Coleman.Liau.short, Dale.Chall, Dale.Chall.old, Dale.Chall.PSK, Danielson.Bryan, Dickes.Steiwer, DRP, ELF, Farr.Jenkins.Paterson, Flesch, Flesch.PSK, Flesch.Kincaid, Fucks, Linsear.Write, LIW, nWS, nWS.2, nWS.3, nWS.4, Scrabble, SMOG family (4metrics), Spache family (2metrics), Strain, Traenkle.Bailer and Wheeler.Smith. This group contains major emphasis on Average Sentence Length (ASL) and Average Word Length (AWL), also considering the size of syllables of a sentence, nwsy, (e.g., ELF, Farr.Jenkins.Paterson, Flesch, Flesch.PSK), with different variations (e.g., word count with one, two or three syllables per word, depending on the metric). Less frequently, another components are also used as: number of ”difficult” words not matching the Dale-Chall list of ”familiar” words, nwd (e.g., Dale.Chall.PSK), number of words, nw(e.g., Dale.Chall. PSK, Flesch, Flesch.PSK), number of characters, nc(e.g., Danielson.Bryan), the number of blanks, nblank (e.g., Danielson.Bryan) and number of sentences, nst (e.g., Danielson.Bryan, ELF). See notation in Section 2.2.4. 2. Group 2(measures sorted in descending order in Table 29): 10 measures including Coleman family (5measures), Danielson.Bryan.2, FORCAST family (2 metrics), Scrabble and Traenkle.Bailer.2. 4.2. Case study 2: evaluate 5clusters of 43 readability measures 75 This group does not use Average Sentence Length (ASL), unlike previous groups. This group is based on nwsy (e.g., Coleman, Coleman.C2, FORCAST), nw(e.g., Traenkle.Bailer2, Coleman.Liau.ECP, FORCAST), nst (e.g., Danielson.Bryan2). Other less commonly used components are the number of prepositions, nconj 3(e.g., Traenkle.Bailer2) and the number of conjunctions, nnconj (e.g., Traenkle.Bailer2). See notation in Section 2.2.4. 3. Group 3(measures sorted in descending order in Table 30): 5measures including Bormuth family (2measures), DRP, Dickes.Steiwer and Farr.Jenkins.Paterson. These measures have the same main characteristics as Group 1measures. However, since the components are used differently, these metrics have different results in this case study, in one of the cases that will be analyzed. Looking at Table 26, it is now possible to ranking the clusters regarding the future crash risk metrics: The three future crash risk metrics have some consistency in the order they define clusters, since they are partially correlated, as shown in Fig. 27. However, Crash Count has a slightly different order (in cluster 1and 2). The cluster rankings of the future crash risk measures and their mean are shown: 1. Sorting clusters by NCSKEW mean, ascending form: [3,1,2,4,5]. Color order: [yellow, red, orange, green, blue]. 2. Sorting clusters by DUVOL mean, ascending form: [3,1,2,4,5]. Color order: [yellow, red, orange, green, blue]. 3. Sorting clusters by Crash Count mean, ascending form: [3,2,1,4,5]. Color order: [yellow, orange, red, green, blue]. To show the evidence of the relationship of these readability measures groups with future crash risk measures, it will be analyzed by the order of clusters for the future crash risk measures mentioned above and their relationship with the proposed groups. As already listed, the ascending order of clusters with respect to future crash risk measures is again mentioned: (NCSKEW: [3,1,2,4,5]), (DUVOL: [3,1,2,4,5]) and (Crash Count: [3,2,1,4,5]). These sortings will be used as follows, analysing the relationship of readability measures groups and future crash risk measures: 1. Part of Group 2(Coleman.C2, Coleman.Liau.ECP, Coleman.Liau.grade, Coleman.Liau.short and Scrabble) with Crash Count: Considering the clusters order 4.2. Case study 2: evaluate 5clusters of 43 readability measures 76 Table 28: Readability measures Group 1for case study 2(with 5clusters) Table 29: Readability measures Group 2for case study 2(with 5clusters) Table 30: Readability measures Group 3for case study 2(with 5clusters) 5.2. Prospect for future work 83 highlight Coleman family (5metrics), Danielson.Bryan.2, FORCAST family (2 metrics), Scrabble and Traenkle.Bailer.2metrics (Group 3of case study 1and Group 2of case study 2), which in addition to having several relationships with future crash risk metrics, they include all the future crash risk metrics (NCSKEW, DUVOL and Crash Count). These metrics point to a more readable financial reporting ratio being related to less crash-prone companies. 4. The previous point suggests that Coleman family (5metrics), Danielson.Bryan.2, FORCAST family (2metrics), Scrabble and Traenkle.Bailer.2metrics are related to future crash risk metrics. In this sense, and since these metrics share similar components, these components may be related to future crash risk. These metrics have in common not using Average Sentence Length (ASL) and they are based on nwsy (e.g., Coleman, Coleman.C2, FORCAST), nw(e.g., Traenkle.Bailer2, Coleman.Liau.ECP, FORCAST) and nst (e.g., Danielson.Bryan2). Other less commonly used components are the number of prepositions, nconj (e.g., Traenkle.Bailer2) and the number of conjunctions, nnconj (e.g., Traenkle.Bailer2). See notation in Section 2.2.4. This study, using S&P 100 from 2008 to 2017, suggests that these components may be more related to crash risk metrics NCSKEW,DUVOL and Crash Count. 5.2 prospect for future work This work allowed us to understand some relationships between readability and crash risk metrics, highlighting some metrics and components that may be more related to crash risk. With this in mind, the future work has several directions that can be followed: 1. Retry clustering with the most related metrics: Coleman family (5metrics), Danielson.Bryan.2, FORCAST family (2metrics), Scrabble and Traenkle.Bailer.2 metrics. Understand more details of these metrics and why they might be linked to future crash risk metrics. Test if they are related to other financial indicators. 2. Expand data size. For example, using the S&P 500 index or analysing more than 10 years. 3. Test other financial indicators, in addition to future crash risk. 4. Finding other readability metrics or components that are related to future crash risk. 5.2. Prospect for future work 84 5. With the knowledge gathered, develop new readability measures, that may be more related to companies’ financial performance.. 6. Develop prediction algorithms, building various Machine Learning models using different readability measures to predict different indicators of financial performance (one in each model). In this process, several models are created, where each model predicts a different financial indicator. Analyze the results, seeking conclusions on a set of metrics that can predict a financial indicator. BIBLIOGRAPHY Mehdi Allahyari, Seyed Amin Pouriyeh, Mehdi Assefi, Saied Safaei, Elizabeth D. Trippe, Juan B. Gutierrez, and Krys Kochut. A brief survey of text mining: Classification, clustering and extraction techniques. CoRR, abs/1707.02919,2017. URL http://arxiv.org/abs/1707.02919. Jonathan Anderson. Lix and rix: Variations on a little-known readability index. Journal of Reading,26(6):490–496,1983. Adam Atkins, Mahesan Niranjan, and Enrico Gerding. Financial news predicts stock market volatility better than close price. The Journal of Finance and Data Science,4(2): 120–137,2018. Richard Bamberger and Erich Vanecek. Lesen, verstehen, lernen, schreiben: die Schwierigkeitsstufen von Texten in deutscher Sprache. Jugend und Volk, 1984. BarChart. Provider of real-time or delayed intraday stock and commodities charts and quotes, 2019. URL https://www.barchart.com. David S Bates. Us stock market crash risk, 1926–2010.Journal of Financial Economics, 105(2):229–259,2012. CH Bj¨ ornsson. L¨ asbarhet, liber. Stockholm, Sweden,1968. David M Blei, Andrew Y Ng, and Michael I Jordan. Latent dirichlet allocation. Journal of machine Learning research,3(Jan):993–1022,2003. Samuel B Bonsall IV, Andrew J Leone, Brian P Miller, and Kristina Rennekamp. A plain english measure of financial reporting readability. Journal of Accounting and Economics,63(2-3):329–357,2017. John R Bormuth. Cloze test readability: Criterion reference scores. Journal of educational measurement,5(3):189–196,1968. John R Bormuth. Development of readability analysis. 1969. 85 Bibliography 86 Jeffrey L Callen and Xiaohua Fang. Religion and stock price crash risk. Journal of Financial and Quantitative Analysis,50(1-2):169–195,2015. John S Caylor and Thomas G Sticht. Development of a simple readability index for job reading material. 1973. Soumen Chakrabarti, Byron Dom, Rakesh Agrawal, and Prabhakar Raghavan. Using taxonomy, discriminants, and signatures for navigating in text databases. In VLDB, volume 97, pages 446–455,1997. Jeanne Sternlicht Chall and Edgar Dale. Readability revisited: The new Dale-Chall readability formula. Brookline Books, 1995. Pete Chapman, Julian Clinton, Randy Kerber, Thomas Khabaza, Thomas Reinartz, Colin Shearer, Rudiger Wirth, et al. Crisp-dm 1.0: Step-by-step data mining guide. SPSS inc,16,2000. Joseph Chen, Harrison Hong, and Jeremy C Stein. Forecasting crashes: Trading volume, past returns, and conditional skewness in stock prices. Journal of financial Economics,61(3):345–381,2001. Edmund B Coleman. Developing a technology of written instruction: Some determiners of the complexity of prose. Verbal learning research and the technology of written instruction, pages 155–204,1971. Meri Coleman and Ta Lin Liau. A computer readability formula designed for machine scoring. Journal of Applied Psychology,60(2):283,1975. CRSP. Center for research in security prices, 2019. URL http://www.crsp.com/. Edgar Dale and Jeanne S Chall. A formula for predicting readability: Instructions. Educational research bulletin, pages 37–54,1948. Wayne A Danielson and Sam Dunn Bryan. Computer automation of two readability formulas. Journalism Quarterly,40(2):201–206,1963. Alice Davison and Robert N Kantor. On the failure of readability formulas to define readable texts: A case study from adaptations. Reading research quarterly, pages 187– 209,1982. Gus De Franco, Ole-Kristian Hope, Dushyantkumar Vyas, and Yibin Zhou. Analyst report readability. Contemporary Accounting Research,32(1):76–104,2015. Bibliography 87 Paul Dickes and Laure Steiwer. Ausarbeitung von lesbarkeitsformeln f¨ ur die deutsche sprache. Zeitschrift f¨ ur Entwicklungspsychologie und P¨ adagogische Psychologie,9(1):20– 28,1977. William H DuBay. The principles of readability. Online Submission,2004. Irving E Fang. The “easy listening formula”. Journal of Broadcasting & Electronic Media, 11(1):63–68,1966. James N Farr, James J Jenkins, and Donald G Paterson. Simplification of flesch reading ease formula. Journal of applied psychology,35(5):333,1951. Rudolph Flesch. A new readability yardstick. Journal of applied psychology,32(3):221, 1948. Wilhelm Fucks. Unterschied des Prosastils von Dichtern und anderen Schriftstellern: ein Beispiel mathematischer Stilanalyse. Bouvier, 1955. John Gantz and David Reinsel. The digital universe in 2020: Big data, bigger digital shadows, and biggest growth in the far east. IDC iView: IDC Analyze the future,2007 (2012):1–16,2012. Robert Gunning. The technique of clear writing. 1952.New York, NYMcGraw-Hill,1952. Vishal Gupta, Gurpreet S Lehal, et al. A survey of text mining techniques and applications. Journal of emerging technologies in web intelligence,1(1):60–76,2009. Eui-Hong Sam Han, George Karypis, and Vipin Kumar. Text categorization using weight adjusted k-nearest neighbor classification. In Pacific-asia conference on knowledge discovery and data mining, pages 53–65. Springer, 2001. Amy P Hutton, Alan J Marcus, and Hassan Tehranian. Opaque financial reports, r2, and crash risk. Journal of financial Economics,94(1):67–86,2009. Investor.gov. Form 10-k, 2019. URL https://www.investor.gov/ additional-resources/general-resources/glossary/form-10-k. Ishares. Sp 100 etf, 2019. URL https://www.ishares.com/us/products/239723/ ishares-sp-100-etf. Anders Johansen, Didier Sornette, and Olivier Ledoit. Predicting financial crashes using discrete scale invariance. arXiv preprint cond-mat/9903321,1999. Bibliography 88 Kenneth. Kenneth r. french site, 2019. URL http://mba.tuck.dartmouth.edu/pages/ faculty/ken.french/data_library.html. Jeong-Bon Kim and Liandong Zhang. Accounting conservatism and stock price crash risk: Firm-level evidence. Contemporary Accounting Research,33(1):412–441,2016. Jeong-Bon Kim, Yinghua Li, and Liandong Zhang. Corporate tax avoidance and stock price crash risk: Firm-level analysis. Journal of Financial Economics,100(3):639–662, 2011. J Peter Kincaid, Robert P Fishburne Jr, Richard L Rogers, and Brad S Chissom. Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel. 1975. George R Klare. Assessing readability. Reading research quarterly, pages 62–102,1974. George Roger Klare et al. Measurement of readability. 1963. Genell L Knatterud, Frank W Rockhold, Stephen L George, Franca B Barton, CE Davis, William R Fairweather, Tom Honohan, Richard Mowery, and Robert O’Neill. Guidelines for quality assurance in multicenter trials: a position paper. Controlled clinical trials,19(5):477–493,1998. Deanne Larson and Victor Chang. A review and future direction of agile, business intelligence, analytics and data science. International Journal of Information Management, 36(5):700–710,2016. Alastair Lawrence. Individual investors and financial disclosure. Journal of Accounting and Economics,56(1):130–147,2013. Reuven Lehavy, Feng Li, and Kenneth Merkley. The effect of annual report readability on analyst following and the properties of their earnings forecasts. The Accounting Review,86(3):1087–1115,2011. Feng Li. Annual report readability, current earnings, and earnings persistence. Journal of Accounting and economics,45(2-3):221–247,2008. Tim Loughran and Bill McDonald. Measuring readability in financial disclosures. The Journal of Finance,69(4):1643–1671,2014. G Harry Mc Laughlin. Smog grading-a new readability formula. Journal of reading,12 (8):639–646,1969. Bibliography 89 Bill McDonald. Bill mcdonald, 2019. URL https://sraf.nd.edu/data/ stage-one-10-x-parse-data/. Marlene M Most, Shirley Craddick, Staci Crawford, Susan Redican, Donna Rhodes, Fran Rukenbrod, Reesa Laws, Dash-Sodium Collaborative Research Group, et al. Dietary quality assurance processes of the dash-sodium controlled diet study. Journal of the American Dietetic Association,103(10):1339–1346,2003. Azadeh Nikfarjam, Abeed Sarker, Karen O’connor, Rachel Ginn, and Graciela Gonzalez. Pharmacovigilance from social media: mining adverse drug reaction mentions using sequence labeling with word embedding cluster features. Journal of the American Medical Informatics Association,22(3):671–681,2015. Cathy O’Neil and Rachel Schutt. Doing data science: Straight talk from the frontline. ” O’Reilly Media, Inc.”, 2013. Richard D Powers, William A Sumner, and Bryant E Kearl. A recalculation of four adult readability formulas. Journal of Educational Psychology,49(2):99,1958. Daniel Ramage, Susan T. Dumais, and Daniel J. Liebling. Characterizing microblogs with topic models. In Proceedings of the Fourth International Conference on Weblogs and Social Media, ICWSM 2010, Washington, DC, USA, May 23-26,2010,2010. SEC. Plain English Handbook. Office of Investor Education and Assistance U.S. Securities and Exchange Commission, 1998. sec. Form 10-q, 2019a. URL https://www.sec.gov/fast-answers/ answersform10qhtm.html. sec. Changeover to the sec’s new smaller reporting company system by small business issuers and non-accelerated filer companies, 2019b. URL https://www.sec.gov/ info/smallbus/secg/smrepcosysguid.pdf. SEC.gov. Sec.gov, 2019. URL https://www.sec.gov/edgar/searchedgar/ companysearch.html. John Seely. Oxford AZ of grammar and punctuation. Oxford University Press, 2013. Jyoti Sharma, Mrs Shashi Sharma, and Ruchi Pandey. A complete review of concept of data mining. International Journal For Technological Research In Engineering,5(6): 3143–3146,2018. Bibliography 90 Edgar A Smith and RJ Senter. Automated readability index. AMRL-TR. Aerospace Medical Research Laboratories (US), pages 1–14,1967. George Spache. A new readability formula for primary-grade reading materials. The Elementary School Journal,53(7):410–413,1953. Ulrich Tr¨ ankle and Harald Bailer. Kreuzvalidierung und neuberechnung von lesbarkeitsformeln f¨ ur die deutsche sprache. Zeitschrift f¨ ur Entwicklungspsychologie und P¨ adagogische Psychologie,1984. Xuan Vinh Vo. Foreign investors and stock price crash risk: Evidence from vietnam. International Review of Finance,2019. Lester R Wheeler and Edwin H Smith. A practical readability formula for the classroom teacher in the primary grades. Elementary English,31(7):397–399,1954. Ian H Witten, Eibe Frank, Mark A Hall, and Christopher J Pal. Data Mining: Practical machine learning tools and techniques. Morgan Kaufmann, 2016. Svante Wold, Kim Esbensen, and Paul Geladi. Principal component analysis. Chemometrics and intelligent laboratory systems,2(1-3):37–52,1987. Cheng Xiang, Fengwen Chen, and Qian Wang. Institutional investor inattention and stock price crash risk. Finance Research Letters,2019. Haifeng You and Xiao-jun Zhang. Financial reporting complexity and investor underreaction to 10-k information. Review of Accounting studies,14(4):559–586,2009. Chen-Hsiang Yu and Robert C Miller. Enhancing web page readability for non-native readers. In Proceedings of the sIGCHI conference on human factors in computing systems, pages 2523–2532. ACM, 2010. Jun Zhu, Amr Ahmed, and Eric P Xing. Medlda: maximum margin supervised topic models for regression and classification. In Proceedings of the 26th annual international conference on machine learning, pages 1257–1264. ACM, 2009.