scieee AI-readable full text Open interactive document viewer

FinTextSim: Enhancing Financial Text Analysis with BERTopic

Jehnen, Simon; Ordieres-Meré, Joaquín; Villalba-Diez, Javier

Full text

FinTextSim: Enhancing Financial Text Analysis with BERTopic Simon Jehnena,d, Javier Villalba-D´ıezb,c, Joaqu´ın Ordieres-Mer´ea aUniversidad Polit´ecnica de Madrid, DEGIN doctoral program, Department of Industrial Management. ETSII, c. Jos´e Guti´errez Abascal 2, Madrid, 28006, Madrid, Spain bFakult¨at f¨ur Wirtschaft, Hochschule Heilbronn, Bildungscampus, Max-Planck-Straße 39, Heilbronn, 74081, State Two, Germany cEscuela T´ecnica Superior de Ingenier´ıa Industrial, Universidad de la Rioja, C. Luis de Ulloa, 4, Logro˜no, 26004, La Rioja, Spain dBeta Klinik GmbH, Joseph-Schumpeter-Allee 15, Bonn, 53227, Germany Abstract Recent advancements in information availability and computational capabilities have transformed the analysis of annual reports, integrating traditional financial metrics with insights from textual data. To extract valuable insights from this wealth of textual data, automated review processes, such as topic modeling, are crucial. This study examines the effectiveness of BERTopic, a state-of-the-art topic model relying on contextual embeddings, for analyzing Item 7 and Item 7A of 10-K filings from S&P 500 companies (2016–2022). Moreover, we introduce FinTextSim, a finetuned sentencetransformer model optimized for clustering and semantic search in financial contexts. Compared to all-MiniLM-L6-v2, the most widely used sentencetransformer, FinTextSim increases intratopic similarity by 81% and reduces intertopic similarity by 100%, significantly enhancing organizational clarity. We assess BERTopic’s performance using embeddings from both FinTextSim and all-MiniLM-L6-v2. Our findings reveal that BERTopic only forms clear and distinct economic topic clusters when paired with FinTextSim’s embeddings. Without FinTextSim, BERTopic struggles with misclassification and overlapping topics. Thus, FinTextSim is pivotal for advancing financial text analysis. FinTextSim’s enhanced contextual embeddings, tailored for the financial domain, elevate the quality of future research and financial information. This improved quality of financial information will enable stakeholders to gain a competitive advantage, streamlining resource allocation and decision-making processes. Moreover, the improved insights have Preprint submitted to International Review of Economics and Finance. April 23, 2025 arXiv:2504.15683v1 [cs.CL] 22 Apr 2025 the potential to leverage business valuation and stock price prediction models. Keywords: Topic Modeling, 10-K, Artificial Intelligence, BERTopic, MD&A, FinTextSim, Sentence Transformers Acknowledgement The second and third authors want to acknowledge the partial support by the Spanish “Agencia Estatal de Investigaci´on” through the grant PID2022137748OB-C31 funded by MCIN/AEI/10.13039/501100011033 and ”ERDF A way of making Europe”. 1. Introduction In recent years, the increasing availability of information (Gupta et al., 2020; Abukari et al., 2024) and advances in computational capabilities have transformed the way we analyze annual reports, including 10-K filings. 10-K filings are among the most critical annual reports (Griffin, 2003), providing a standardized snapshot of a company’s financial situation through both numerical and textual data (Masson and Paroubek, 2020). The information contained carries predictive power for future profitability (You and Zhang, 2009). Hence, various stakeholders, such as investors and financial analysts, rely on 10-K reports to make informed decisions (Liu, 2022). Traditionally, evaluation focused on retrospective quantitative financial metrics. However, there is a growing recognition of the value embedded in qualitative textual data. For example, Cohen et al. (2020); Li (2010a) demonstrated that language and tone in financial reports correlate with future company returns. Therefore, integrating retrospective financial metrics with textual analysis offers a fuller picture of a company, improving various decision-making processes (Hsieh and Hristova, 2022). Item 7 and Item 7A of 10-K filings contain valuable information regarding companies listed in the Standard and Poor’s 500 (S&P 500). Among the 15 items included in 10-K reports, they stand out as particularly crucial. Item 7 is the Management Discussion & Analysis (MD&A) section. In this section, the management presents the company’s perspective on various aspects, including operations, performance, risks, opportunities, and strategies to address future challenges (Cohen et al., 2020). Item 7A contains qualitative and quantitative disclosures about market risk. As the submission of a 2 10-K filing is mandatory for publicly traded companies, there is a wealth of information. We need clearly defined review processes to evaluate and use this information (Dutta et al., 2019). Automated review processes offer various advantages over manual approaches. First, manual review of textual data is time-consuming and prone to subjectivity bias (Li, 2010a,b). Moreover, the vast and growing amount of data (Gupta et al., 2020) can lead to information overload, particularly for stakeholders with limited attention (Lu, 2022). Thus, attention must be efficiently allocated (Liu, 2022). Topic modeling, a technique from Natural Language Processing (NLP), addresses these challenges by uncovering latent topics within textual datasets. Hence, they help to organize and summarize large text corpora (Blei et al., 2003). Recently developed neural topic models address the limitations of classical topic modeling approaches, which still dominate applied topic modeling. Classical topic modeling approaches have been discussed in the literature since the introduction of Latent Semantic Indexing in 1990 (Deerwester et al., 1990). Until 2015, Bayesian probabilistic models, most notably Latent Dirichlet Allocation (LDA), were considered state-of-the-art. Relying on the bag-of-words (BoW) assumption, each document is treated as a collection of words, disregarding their sequential order. However, this approach limits the model’s ability to capture the semantic meaning of text. Neural topic modeling approaches address this issue by employing contextual embeddings (Blair et al., 2020), allowing them to capture richer semantic and contextual relationships within the data (Booker et al., 2024; Bhattacharya and Mickovic, 2024). Recent advances in contextual embeddings have transformed NLP, driven by key innovations such as the transformer architecture, encoder-only models, and sentence-transformers. The transformer architecture, introduced by Vaswani et al. (2017), revolutionized NLP by relying entirely on attention mechanisms, allowing models to capture long-range dependencies and rich contextual information. This made transformers the state-of-the-art approach for Natural Language Understanding tasks (Caron and M¨uller, 2020; Han et al., 2024). Building on this foundation, Bidirectional Encoder Representations from Transformers (BERT) established a new standard for deep contextualized language modeling (Devlin et al., 2019). BERT relies on the encoder from the transformers architecture, allowing it to processes text bidirectionally and capture nuanced semantic relationships. More recently, Warner et al. (2024) introduced improvements to this architecture, increasing 3 efficiency and surpassing BERT in classification and retrieval tasks. Despite their strengths, encoder-only models are less effective for large-scale semantic similarity comparison and clustering (Reimers and Gurevych, 2019). To address these limitations, sentence-transformers refine encoder-only models using siamese or triplet network architectures, enabling efficient and precise similarity assessments (Reimers and Gurevych, 2019). Sentence-transformers encode text into dense vector representations—contextual embeddings—that quantify semantic similarity by mapping similar texts closer in a shared vector space. Neural topic models, such as BERTopic, leverage these contextual embeddings to enhance topic modeling by clustering semantically related documents, enabling more accurate topic discovery (Grootendorst, 2022). While these advancements have demonstrated significant improvements in general text processing, it remains unclear how these sophisticated methods perform when applied to tasks specifically relevant to the finance and accounting domains (Bhattacharya and Mickovic, 2024). Additionally, there is little evidence that the customization of financial text leads to benefits in performance (Huang et al., 2023). Traditional algorithms continue to dominate applied topic modeling, hindering the generation of new knowledge (Egger and Yu, 2021; Blair et al., 2020; Huang et al., 2023). Although extensive studies on various topic modeling approaches (Albalawi et al., 2020; Fu et al., 2021; Egger and Yu, 2022; Farzadnia et al., 2024) have been conducted, the domain of Management Accounting and Finance, particularly Item 7 and Item 7A of 10-K reports from S&P 500 companies, remains significantly under-researched. This presents a critical opportunity to integrate Machine Learning (ML)-based methods to fully exploit the value hidden in financial textual data (Ranta et al., 2022). Furthermore, it opens avenues for refining domain-specific contextual embeddings by incorporating domain expertise (Murphy et al., 2024). Such an approach aims to enhance the quality of contextual embeddings, allowing them to capture terminology and concepts unique to finance with greater precision (Dong et al., 2024). To address this gap, we introduce FinTextSim, a finetuned sentence transformer leveraging the capabilities of contextual embeddings for the financial domain. We benchmark FinTextSim against all-miniLM-L6-v2 (AM), the most widely used general-purpose sentence transformer. To isolate the effect of the selected sentence-transformer, we generate contextual embeddings for our dataset using both FinTextSim and AM. We apply these embeddings to BERTopic, a state-of-the-art neural topic modeling approach, while keeping all other model parameters identical. This comparison allows us to determine 4 whether domain-specific fine-tuning enhances topic modeling and financial text interpretation. We will explore the following research questions based on Item 7 and Item 7A from S&P500 companies between 2016 and 2022: RQ1 How can we leverage the capabilities of contextual embeddings for the financial domain? RQ2 Which embedding model — FinTextSim or AM — produces more qualitative and coherent topics when used as input for BERTopic? RQ3 Which embedding model — FinTextSim or AM — better supports BERTopic in organizing and summarizing large-scale financial text corpora? By addressing the research questions, our work makes the following significant contributions: 1. We identify the most effective approach for extracting meaningful topics from Item 7 and Item 7A of S&P500 companies. This comparison provides valuable insights for researchers and practitioners selecting embedding models for NLP tasks. 2. We introduce FinTextSim, a finetuned sentence-transformer, improving the analysis of financial text for various downstream tasks. Hence, FinTextSim will boost future research quality. 3. By enhancing BERTopic for financial text with FinTextSim, we generate higher quality financial information regarding companies, sectors as well as whole markets and economies. Thus, we enable managers, financial analysts, investors, regulators and other stakeholders to gain a competitive advantage which aids in allocating resources more efficiently and making rational operational and strategic decisions. Moreover, the improved insights have the potential to leverage business valuation and stock price prediction models. The rest of the paper hereinafter is organized as follows. Section 2 reviews the state-of-the-art literature and methodologies. Section 3 describes our study’s materials and methods, including the training procedure of FinTextSim. Section 4 presents and discusses the main findings. Finally, Section 5 provides the conclusion. This structure ensures a clear and logical progression, enabling a thorough understanding of our study’s contributions. 5 2. State of the Art The following subsections provide an overview of the evolution of contemporary topic modeling techniques and a detailed examination of BERTopic, a state-of-the-art topic modeling approach. 2.1. Evolution of Contemporary Topic Modeling Approaches Recent advancements in topic modeling have seen the integration of contextual embeddings, offering both significant benefits and challenges. Modern methodologies address the limitations of classical models by utilizing advanced text embedding techniques, moving beyond simple BoW representations. This enables them to better capture semantic relationships within text (Blair et al., 2020). While classical models require extensive customization and become increasingly complex with larger datasets, modern approaches offer enhanced flexibility and scalability (Zhao et al., 2021; Wu et al., 2024). Moreover, neural topic models simplify the inference problem, enabling parallelization (Abdelrazek et al., 2023; Wu et al., 2024). Among other innovative methods (Wu et al., 2024), contextual vector representations are combined with centroid-based clustering techniques (Sia et al., 2020; Angelov, 2020). They assume that the centroid of a cluster represents the topic. Words closest to the centroid are considered as the most representative ones for the topic. However, this assumption is fragile as clusters may not always conform to a spherical distribution around the centroid. As a result, misrepresentation of topics may occur (Grootendorst, 2022). A promising approach for topic modeling based on contextual embeddings, addressing centroid-based issues, is BERTopic. 2.2. BERTopic BERTopic structures topic modeling into five sequential steps. First, document embeddings are generated using a pre-trained sentence transformer. A method that offers enduring benefits by leveraging advancements in language models (Grootendorst, 2022; Gu et al., 2024a). Second, the dimensionality of these embeddings is reduced. Subsequently, the reduced embeddings are clustered into semantically similar groups, i.e., topics. The reduction of dimensionality is deliberately placed before clustering to increase computational efficiency as well as clustering accuracy (Allaoui et al., 2020). In the fourth step, topics are tokenized. Finally, tokens are weighted. To enhance the quality of extracted topic representations, Grootendorst (2022) 6 introduces class-based tfidf (c-tfidf), which weighs the importance of tokens within topics, enabling a more efficient extraction of topic representations. Despite its significant advantages, BERTopic also faces several drawbacks. It tends to produce a manifold of closely interconnected topics which may vary upon repeated modeling attempts (Egger and Yu, 2022). This variability contributes to inconsistency in producing meaningful results, further complicated by the complexity of interpreting hyperparameters, hindering troubleshooting and diminishing the reliability of results (Abdelrazek et al., 2023). Moreover, BERTopic assumes that each document relates to a single topic, potentially oversimplifying real-world document complexity (Grootendorst, 2022). Additionally, sentence-transformer models used for document embedding perform optimally with sentences or paragraphs (Reimers and Gurevych, 2019). Furthermore, high computation times can result from processing large amounts of data (Grootendorst, 2022). Due to its novelty, applications and enhancements of BERTopic are still in their infancy. In a financial context, Kim et al. (2022) utilized BERTopic on Item 1A from 10-K filings. They assessed whether identified topics can enhance the accuracy of ESG rating predictions and quantify each topic’s relative contribution to the final rating prediction. In other contexts, BERTopic has been applied in various studies: S´anchez-Franco and Rey-Moreno (2022) analyzed customer reviews, Abuzayed and Al-Khalifa (2021) explored its application with pre-trained Arabic language models, Egger and Yu (2022) evaluated its performance on Twitter data, and Grigore and Pintilie (2023) extended BERTopic to predict individual’s responses to a questionnaire based on their social media activity. 2.3. Topic Modeling of Item 7 and Item 7A Our research is driven by several motivations regarding the choice of documents and analysis techniques. Item 7 and Item 7A stand out as particularly crucial sections in 10-K reports (Bhattacharya and Mickovic, 2024). The MD&A section (Item 7) provides a narrative that contextualizes the presented numbers, covering topics such as performance, liquidity, risks, and operations. In this section, management offers its individual perspective, which is essential for understanding the company’s strategic direction and potential challenges. Additionally, the MD&A section offers the most leeway and flexibility, making it rich with insights and indicative of future performance (Cohen et al., 2020). Item 7A focuses on market risks, containing valuable 7 information regarding the company’s prospective performance. Hence, analyzing Item 7 and Item 7A allows us to uncover hidden textual information that has the potential to support the prediction of a company’s future performance. A technique that shows promise in addressing this challenge is topic modeling (Ranta et al., 2022). Despite advancements in computational capabilities and the emergence of new topic modeling techniques, there remains a gap in applying topic modeling methods to financial texts, particularly Item 7 and Item 7A. Furthermore, LDA continues to dominate applied topic modeling, although newer approaches like BERTopic offer potential improvements (Egger and Yu, 2021; Blair et al., 2020). In this paper, we aim to demonstrate how FinTextSim, a finetuned sentence transformer, outperforms the most widely used sentence transformer AM on text from Item 7 and Item 7A from 10-K reports of S&P 500 companies. Additionally, we hypothesize that combining BERTopic with FinTextSim will significantly enhance the quality of financial information, providing better insights for stakeholders. We foresee that FinTextSim will also facilitate the application of aspect-based sentiment analysis, eventually improving business valuation and stock price prediction models. 3. Materials and Methods In the following subsections, we outline the materials and methods of our study. This section is divided into several parts: sourcing the dataset, creating an enhanced financial keyword list, training FinTextSim, creating the topic models, and presenting the metrics used to evaluate the performance of the topic models. 3.1. Dataset Our study focuses exclusively on Item 7 and Item 7A of 10-K reports while avoiding survivorship bias and ensuring the highest possible document comparability. Given their greater significance, we deliberately choose 10-K over 10-Q reports (Griffin, 2003). We source our data from the Notre Dame Software Repository for Accounting and Finance in text-file format, which underwent a ’Stage One Parse’ to remove all HTML tags.1 To prevent survivorship bias, we filter 10-K filings of all companies that have been listed in the S&P 500 index between 2015 and 2022. Using a 1The data can be found at: https://sraf.nd.edu/data/stage-one-10-x-parse-data/. 8 regular expression-based extractor, we isolate the text from the start of Item 7 to the start of Item 8. Through this endeavor, we obtain the raw text of Item 7 and Item 7A. We refer to this combination of Item 7 and Item 7A as ’documents’. In order to maintain comparability, documents containing fewer than 250 words are discarded.2 Subsequently, we remove further outlier documents, identified by z-score. The z-score is a statistical measure that quantifies the distance of data points to the mean of a dataset, taking the standard deviation into account. We define data points as outliers if they deviate more than two standard deviations from the mean. Subsequently, we focus on documents from the period between 2016 and 2022. Finally, we apply classical text processing techniques to identify relevant documents, ensuring that only meaningful texts are included in the analysis. As part of this process, we normalize documents by removing stopwords, applying lemmatization, and performing tokenization. Subsequently, documents with low cosine similarity to others are filtered out. This procedure ensures that both classical (evaluated in Appendix D) and contemporary models are applied to the same set of documents for a direct comparison. The text fed into the sentence transformers remains unprocessed by classical techniques. The number of documents retained at each preprocessing step is shown in Table 1. Table 1: Dataset. Preprocessing-Step # documents Documents extracted related to S&P 500 companies 4,600 Documents with less than 250 words 1,019 Documents with z-score greater than 2 165 Documents outside the timeframe of 2016-2022 373 Documents with low cosine similarity 1,604 Remaining documents in database 1,439 2Paragraphs typically consist of 100–200 words. Moreover, sentence-transformers, such as AM and FinTextSim are designed to capture the semantic information of sentences and short paragraphs. Input texts longer than 256-word pieces (approximately 170-210 words) are truncated by default. The 250-word threshold ensures that each document includes at least two paragraphs, enhancing relevance, as shorter texts often lack substantive or complete ideas. 9 •Topic Embedding Calculation: Compute the mean of all sentence embeddings for each topic to obtain topic embeddings. •Cosine Similarity Matrix Computation: Calculate the cosine similarity between each topic embedding and all other topic embeddings, forming a pairwise similarity matrix. •Extraction of Upper Triangle Values: Extract the upper triangle of the similarity matrix, excluding the self-similarity values in the diagonal. •Mean Calculation for all Topics: The mean of these cosine similarities results in the model’s intertopic similarity. For the evaluation of FinTextSim, we use the topic labels from the test dataset. As we do not have any ground-truth labels for BERTopic, we employ its topic assignments. 3.5.3. Applicability for the Financial Domain To evaluate the topic models specifically for the financial domain, we incorporate a topic-precision-based weighting into the metrics for topic quality and organizing power, using the previously described keyword list (see Section 3.2). Specifically, we weigh NPMI Coherence and intratopic similarity by multiplying them with topic-precision, while intertopic similarity is adjusted by dividing it by topic-precision. We calculate topic-precision as follows: •Determine the dominant topic: A topic is classified as dominant if it contains at least two keywords from a single topic domain and no more than one keyword from another domain. •Count True Positives (TP): The number of keywords belonging to the dominant topic’s original domain. •Count False Positives (FP): The number of keywords from other topic domains. Keywords not present in the keyword list are ignored. •Compute topic-precision: Topic −Precision =TP (TP +FP)(1) •Penalty for missing topics: If a topic from the keyword list is not captured, its topic-precision is set to 0. 16 •Assess overall performance: The model’s ability to capture financial topics is evaluated by averaging topic-precision across all topics. This approach ensures that the detection of financial topics is the most crucial aspect of the topic models. If the identified topics lack financial relevance, their structural organization becomes obsolete. Moreover, if no financial topics are detected, the scores are penalized, reflecting the model’s unsuitability for financial text analysis. 4. Results and Discussion We structure the results and discussion section according to our research questions: RQ1 FinTextSim: Leveraging the quality of contextual embeddings for the financial domain RQ2 Topic Quality: Creating qualitative, coherent topic representations RQ3 Organizing Power: Organizing large textual datasets For FinTextSim, we highlight its result on the test dataset as well as its value in combination with BERTopic for topic modeling tasks. The results are presented and contextualized in the following subsections. 4.1. FinTextSim - Leveraging Contextual Embeddings for the Financial Domain FinTextSim generates improved clusters and notably reduces the number of outliers in comparison to AM. As illustrated in Figure 1 and Table 2, FinTextSim leads to a significant increase of intratopic similarity, while simultaneously reducing intertopic similarity compared to AM on the test dataset. Specifically, FinTextSim increases intratopic similarity by 81%, achieving a score of 0.9972, compared to 0.5498 for AM. Moreover, FinTextSim reduces intertopic similarity by 100%, resulting in a score of 0.002, whereas AM has a score of 0.4647. When combining BERTopic with AM, we identified 226,605 outliers. In contrast, FinTextSim generated 184,470 outliers, reducing the number of outliers by 19%. The results indicate that FinTextSim creates significantly better clusters of semantically similar concepts compared to AM. FinTextSim’s clusters are characterized by high intratopic similarity and low intertopic similarity. In 17 Table 2: FinTextSim vs. AM: Intraand Intertopic Similarity on test dataset. Model Intratopic Similarity Intertopic Similarity Outliers within BERTopic FinTextSim 0.9972 0.0002 184,470 AM 0.5498 0.4647 226,605 (a) UMAP reduced sentence embeddings - FinTextSim. (b) UMAP reduced sentence embeddings - AM. Figure 1: FinTextSim vs. AM on the test dataset. The colors of the datapoints represent a topic from the keyword list. contrast, AM fails to accurately detect topic specifics, resulting in an indistinguishable mash of data points (see Figure 1). This demonstrates that OTS sentence transformers are unsuitable for financial text. Figure 2: Topic representations - FinTextSim vs. AM - HR. Original cleaned sentence: ’a majority of employees belong to labor unions’. Turning to a practical example, BERTopic with FinTextSim is able to accurately capture the underlying topic of a sentence, while the results with AM are misleading. Figure 2 displays word clouds for topics assigned to 18 the same sentence using BERTopic with FinTextSim and AM. FinTextSim correctly identifies that the sentence focuses on ’HR and Employment’. In contrast, AM fails to detect the topic, leading the analyst to associate it incorrectly with revenue and products. 4.2. Topic Quality As displayed in Section 3.5, we use NPMI coherence weighted by topicprecision to objectively measure the topic quality of each topic model. Table 3 displays the topic-precision scores for each topic model. Table 3: Topic-Precision Scores. Type Model Sentences Refined Sentences AM 0.311 0.684 FinTextSim 1.0 1.0 Using BERTopic with FinTextSim significantly enhances the identification of economic topics. For sentence-level input, AM achieves a topicprecision of 0.311 while failing to detect ten out of 14 economic topics. Refining sentence input improves AM’s topic-precision to 0.684, reducing the number of missed topics to six. However, even with refinement, AM struggles to identify key financial themes such as Operations, Liquidity and Solvency, and Investment. In contrast, FinTextSim correctly identifies all 14 economic topics, achieving a perfect topic-precision of 1.0.5 Although refining sentence input improves AM’s topic precision, generalpurpose embedding models are insufficient for comprehensive financial text analysis. AM models fail to capture the full scope of financial topics, limiting their practical application. In contrast, FinTextSim not only identifies all relevant topics but also minimizes conceptual overlap, ensuring clearer topic distinctions and more effective document organization. These findings underscore the necessity of domain-specific embeddings for financial text processing, as general-purpose models fail to capture the nuances of economic language. 5The wordclouds for each model are displayed in Figures A.6 to A.9. 19 Table 4: Coherence Scores. Type Model Sentences Refined Sentences AM 0.106 (0.341) 0.257 (0.376) FinTextSim 0.279 (0.279) 0.301 (0.301) Table 4 presents the coherence scores, with non-topic-precision weighted coherence shown in parentheses. When incorporating topic-precision weighting, we find that BERTopic generates higher-quality financial topics using FinTextSim compared to AM. For sentence-level input, FinTextSim achieves a topic-precision weighted coherence score of 0.279, surpassing AM’s 0.106 by 163%. With refined sentence input, coherence scores improve across both models. For AM, this increase is primarily attributed to its increased topicprecision. However, FinTextSim still outperforms AM by 17%. Contrary to our expectations, BERTopic achieves higher coherence with AM than with FinTextSim when topic-precision-based weighting is not applied. Although FinTextSim better captures the distinct structure of financial text, as reflected by its high topic-precision, it yields lower coherence scores. However, this result is misleading in the context of financial text analysis. AM generates significantly more outliers and fails to capture key economic topics, leading to a loss of valuable information, potentially compromising studies. We suspect that the higher coherence of BERTopic with AM results from the increased number of outliers, which simplifies the compression and generation of topics. Another factor is the vocabulary of the financial domain. Financial terms are often standalone words that do not necessarily co-occur within a sliding window. Therefore, coherence scores do not always align with human judgment of topic quality.6Relying solely on standard coherence metrics without domain-specific weighting can be problematic. In financial text analysis, ensuring high topic-precision is crucial as the meaningful organization of domain-specific knowledge must take precedence over raw coherence scores. Since topic evaluation requires domain expertise and subjective interpretation (Grootendorst, 2022), standard coherence metrics alone are insufficient. Without domain-specific weighting, 6This trend is also observed for classical models (see Appendix D). 20 coherence fails to reflect practical applicability, making topic-precision essential for capturing meaningful financial insights. Figure 3: Topic representations - FinTextSim vs. AM - Cost. Original cleaned sentence: ’business is vulnerable to fluctuations in fuel costs and disruptions in fuel supplies’. To illustrate this in a practical example, we refer to Figure 3. Once again, FinTextSim correctly identifies the underlying ’Cost’ topic, while AM misclassifies it, associating it with currency and exchange rates. Despite this misclassification, AM receives a higher coherence score of 0.448 compared to FinTextSim’s 0.324. Figure 4: Topic representations - FinTextSim vs. AM - Operations. Original cleaned sentence: ’the company manufactures markets and distributes spices seasoning mixes condiments and other flavorful products to the entire food industry retailers food manufacturers and foodservice businesses’. For further clarity, Figure 4 presents another practical example. FinTextSim correctly identifies the ’Operations’ topic, whereas AM misclassifies it as a non-economic concept. Yet, FinTextSim receives a lower coherence score of 0.205 compared to AM’s 0.262. These discrepancies underscore the limitations of coherence scores in evaluating financial text, as they fail to account for domain relevance and topic-precision. 21 4.3. Organizing Power To efficiently organize and structure large collections of documents, maximizing intratopic similarity while simultaneously minimizing intertopic similarity is desirable. The results for intratopic similarity of our models are displayed in Table 5. Non-topic-precision weighted scores are illustrated in parentheses. Table 5: Intratopic Similarity. Type Model Sentences Refined Sentences AM 0.206 (0.661) 0.441 (0.644) FinTextSim 0.925 (0.925) 0.929 (0.929) FinTextSim consistently achieves higher intratopic similarity than AM, regardless of input type or the application of topic-precision weighting. When topic-precision weighting is considered, FinTextSim outperforms AM by a wide margin, achieving an intratopic similarity of 0.925 compared to 0.206 for sentence input, representing a 350% improvement. For refined sentence input, FinTextSim reaches 0.929, outperforming AM’s 0.441 by 111%. Even without topic-precision weighting, FinTextSim maintains a strong advantage, with a 40% increase for sentence input, scoring 0.925 compared to AM’s 0.661. The trend remains consistent for refined sentence input, where FinTextSim achieves 0.929, exceeding AM’s 0.644 by 44%. These results further highlight FinTextSim’s ability to generate more cohesive topic clusters, reinforcing its suitability for financial text analysis. The intertopic similarities of the models are displayed in Table 6. Nontopic-precision weighted scores are illustrated in parentheses. A good separation of the generated topics is represented by low intertopic similarities. FinTextSim significantly reduces intertopic similarity compared to AM, indicating enhanced topic separation. When topic-precision weighting is applied, FinTextSim achieves a 95% reduction in intertopic similarity, scoring 0.064 compared to AM’s 1.315. With refined sentence input, intertopic similarity remains 90% lower, with FinTextSim at 0.066 compared to AM’s 0.631. Neglecting topic-precision weighting, FinTextSim continues to outperform AM by a wide margin, reducing intertopic similarity by 84% for 22 Table 6: Intertopic Similarity. Type Model Sentences Refined Sentences AM 1.315 (0.409) 0.631 (0.432) FinTextSim 0.064 (0.064) 0.066 (0.066) sentence input, with scores of 0.064 versus 0.409. For refined sentence input, FinTextSim outperforms AM by 85%, where it achieves 0.066 compared to AM’s 0.432. These results underscore FinTextSim’s effectiveness in minimizing topic overlap, leading to clearer and more distinct financial topic clusters. Figure 5: Topic representations - FinTextSim vs. AM - Accounting. Original cleaned sentence: ’critical accounting policies estimates and judgments our consolidated financial statements are based on gaap which requires us to make estimates and assumptions about future events that affect the amounts reported in our consolidated financial statements’. Figure 5 illustrates this situation. In this practical case, BERTopic with FinTextSim can accurately capture the underlying topic of the sentence, while the results with AM are misleading. FinTextSim correctly recognizes that the sentence pertains to ’Accounting,’ ensuring a precise and domain-relevant topic assignment. In contrast, AM fails to detect the specific topic, causing the analyst to misinterpret the sentence as relating to multiple themes, such as cost, tax, and HR. This inability to distinguish between economic topics results in high intertopic similarity and low intratopic similarity, highlighting the limitations of AM for financial text analysis. 23 4.4. Wrapup of Results and Discussion We find that BERTopic is highly on financial text when combined with FinTextSim. AM, on the other hand, generates more general topics and fails to capture economic topics, leading to significant gaps in coverage. The quality of topics improves significantly when FinTextSim is used. Only in combination with FinTextSim, BERTopic produce clear, distinct clusters of economic topics. In contrast, AM leads to frequent misclassifications. Our findings support the hypothesis of Dong et al. (2024), demonstrating that a fine-tuned model grounded in a specialized dataset significantly improves both performance and domain-specific understanding. Furthermore, as Gu et al. (2024a) indicates, finetuning a foundational base model enhances performance on complex tasks. Relying solely on OTS models may compromise reliability and introduce systematic errors, highlighting the importance of integrating fine-tuned models like FinTextSim for extracting meaningful and reliable insights. However, the extent to which FinTextSim generalizes beyond 10-K reports remains an open question. A small-scale experiment on Item 1 yielded similar results as described in Section 4.1, suggesting that FinTextSim’s effectiveness extends to other sections of 10-K filings. Expanding its training data to include diverse financial sources, such as news articles, conference call transcripts, and analyst reports, could further enhance its generalization capabilities. Additionally, incorporating researcher-labeled data may provide further improvements in FinTextSim’s adaptability and robustness across financial contexts. Regarding our results, it is important to note that the displayed metrics should never be considered in isolation, especially those regarding organizing power. For instance, even if AM would show high intratopic and low intertopic similarity, it does not necessarily produce ’good’ clusters, as the quality of the generated topics may remain low. In such cases, its ability to enhance organizational clarity would still be limited. Therefore, evaluating topics requires looking beyond the raw metrics to consider their true quality. Hence, evaluating topic models remains challenging (Zhao et al., 2021). Our analysis reveals the limitations of coherence as a measure. For instance, BERTopic with AM achieves higher coherence scores than FinTextSim. Yet, we identified low topic-precision scores, indicating numerous missing economic topics and/or overlapping concepts. This suggests that higher coherence does not necessarily correlate with higher topic quality. Our findings underscore the necessity for new coherence or topic quality measures, particularly for domain-specific texts like finance. In such texts, topic words often 24 stand alone and may not co-occur within a sliding window. Hence, traditional coherence metrics cannot capture the ’true’ quality of the generated topics. While BERTopic enhances topic modeling compared to the classical approaches, there is still significant room for improvement. The transformer architecture, which BERTopic heavily relies on, may not be fully optimized yet. Thus, more sophisticated and computationally efficient alternatives should be explored (Karami and Ghodsi, 2024). Further advancements in encoder-only models could enhance sentence transformers by improving their contextual understanding of language (Warner et al., 2024). Moreover, applying domainspecific pre-training methods to optimized BERT variants may deepen the model’s understanding of financial language, leading to more effective downstream task performance (Huang et al., 2023). Additionally, BERTopic’s inconsistency in producing meaningful results, compounded by the complex hyperparameters of the underlying models, compromises reliability (Abdelrazek et al., 2023). Hence, future research should focus on developing an objective standard for selecting models and tuning hyperparameters. Specifically, we plan to investigate the impact of hyperparameter tuning for both dimensionality reduction and clustering techniques on contextual embeddings. This approach aims to streamline the process of topic modeling, objectively determining hyperparameters. Eventually, this will improve topic clustering and extraction, thus enhancing the analysis of textual data. 5. Conclusion Increased availability of information and enhanced computational capabilities have transformed the analysis of annual reports, recognizing the value embedded within qualitative textual data. Automated review processes, such as topic modeling, are crucial for analyzing this data. However, in our domain, the use of those ML-based methods, including contextual embeddings, remains underexplored (Ranta et al., 2022). We address these issues by introducing FinTextSim, a finetuned sentence transformer enhancing analysis of financial text with BERTopic. Our study reveals the significant advantages of FinTextSim over OTS sentence-transformer models. FinTextSim excels in generating distinct clusters of topics, substantially outperforming OTS sentence-transformer models on financial text. This highlights the need for domain-specifically finetuned sentence-transformer models. Additionally, FinTextSim allows BERTopic to 25 pii/S1059056024003484, doi:https://doi.org/10.1016/j.iref.2024. 05.050. Gupta, A., Dengre, V., Kheruwala, H.A., Shah, M., 2020. Comprehensive review of text-mining applications in finance. Financial Innovation 6, 1–25. Han, D., Guo, W., Chen, H., Wang, B., Guo, Z., 2024. Lest: Large language models and spatio-temporal data analysis for enhanced sino-us exchange rate forecasting. International Review of Economics & Finance 96, 103508. URL: https://www.sciencedirect.com/science/article/ pii/S1059056024005008, doi:https://doi.org/10.1016/j.iref.2024. 103508. Hong, W., Zheng, X., Qi, J., Wang, W., Zheng, N., Weng, Y., 2018. Financialflow: Visual analytics of financial news based on hierarchical dirichlet process, in: 2018 33rd Youth Academic Annual Conference of Chinese Association of Automation (YAC), IEEE. pp. 375–380. Hsieh, H.T., Hristova, D., 2022. Transformer-based summarization and sentiment analysis of sec 10-k annual reports for company performance prediction, in: Proceedings of the 55th Hawaii International Conference on System Sciences, Hawaii International Conference on System Sciences. pp. 1759–1768. URL: https://hdl.handle.net/10125/79550, doi:10.24251/hicss.2022.218. Huang, A.H., Wang, H., Yang, Y., 2023. Finbert: A large language model for extracting information from financial text. Contemporary Accounting Research 40, 806–841. Jegadeesh, N., Wu, D.A., 2017. Deciphering Fedspeak: The Information Content of FOMC Meetings. SSRN Electronic Journal doi:10.2139/ssrn. 2939937. Karami, M., Ghodsi, A., 2024. Orchid: Flexible and data-dependent convolution for sequence modeling. arXiv preprint arXiv:2402.18508 . Kim, M.G., Kim, K.S., Lee, K.C., 2022. Analyzing the effects of topics underlying companies’ financial disclosures about risk factors on prediction of esg risk ratings: Emphasis on bertopic, in: 2022 IEEE International Conference on Big Data (Big Data), IEEE. pp. 4520–4527. 32 Lee, D.D., Seung, H.S., 1999. Learning the parts of objects by non-negative matrix factorization. Nature 401, 788–791. Li, F., 2010a. The Information Content of Forward-Looking Statements in Corporate Filings—A Na¨ıve Bayesian Machine Learning Approach. Journal of Accounting Research 48, 1049–1102. doi:10.1111/j.1475-679X. 2010.00382.x. Li, F., 2010b. Textual analysis of corporate disclosures: A survey of the literature. Journal of Accounting Literature 29, 143–165. Li, T., Chen, H., Liu, W., Yu, G., Yu, Y., 2023. Understanding the role of social media sentiment in identifying irrational herding behavior in the stock market. International Review of Economics & Finance 87, 163–179. URL: https://www.sciencedirect.com/science/article/ pii/S1059056023001326, doi:https://doi.org/10.1016/j.iref.2023. 04.016. Lin, H., Hwang, Y., 2021. The effects of personal information management capabilities and social-psychological factors on accounting professionals’ knowledge-sharing intentions: Pre and post covid-19. International Journal of Accounting Information Systems 42, 100522. Liu, M., 2022. Assessing human information processing in lending decisions: A machine learning approach. Journal of Accounting Research 60, 607– 651. Loughran, T., McDonald, B., 2011. When is a liability not a liability? textual analysis, dictionaries, and 10-ks. The Journal of finance 66, 35–65. Lowry, M., Michaely, R., Volkova, E., 2020. Information revealed through the regulatory process: Interactions between the sec and companies ahead of their ipo. The Review of Financial Studies 33, 5510–5554. Lu, J., 2022. Limited attention: Implications for financial reporting. Journal of Accounting Research 60, 1991–2027. Maier, D., Waldherr, A., Miltner, P., Wiedemann, G., Niekler, A., Keinert, A., Pfetsch, B., Heyer, G., Reber, U., H¨aussler, T., Schmid-Petri, H., 33 Adam, S., 2018. Applying LDA Topic Modeling in Communication Research: Toward a Valid and Reliable Methodology. Communication Methods and Measures 12, 93–118. doi:10.1080/19312458.2018.1430754. Masson, C., Paroubek, P., 2020. Nlp analytics in finance with dore: a french 257m tokens corpus of corporate annual reports, in: Language Resources and Evaluation Conference (LREC 2020), ELRA. pp. 2261–2267. Mayasari, R.W., Fithriasari, K., Prastyo, D.D., 2021. Text mining to analyse publication topics of covid-19 using hdp and lda methods, in: AECon 2020: Proceedings of The 6th Asia-Pacific Education And Science Conference, AECon 2020, 19-20 December 2020, Purwokerto, Indonesia, European Alliance for Innovation. p. 374. McInnes, L., Healy, J., 2017. Accelerated Hierarchical Density Clustering, in: 2017 IEEE International Conference on Data Mining Workshops (ICDMW), pp. 33–42. doi:10.1109/ICDMW.2017.12,arXiv:1705.07321. Murphy, B., Feeney, O., Rosati, P., Lynn, T., 2024. Exploring accounting and ai using topic modelling. International Journal of Accounting Information Systems 55, 100709. O’Callaghan, D., Greene, D., Carthy, J., Cunningham, P., 2015. An analysis of the coherence of descriptors in topic modeling. Expert Systems with Applications 42, 5645–5657. doi:10.1016/j.eswa.2015.02.055. Ranta, M., Ylinen, M., J¨arvenp¨a¨a, M., 2022. Machine Learning in Management Accounting Research: Literature Review and Pathways for the Future. European Accounting Review , 1–30doi:10.1080/09638180.2022. 2137221. Reimers, N., Gurevych, I., 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084 . R¨oder, M., Both, A., Hinneburg, A., 2015. Exploring the space of topic coherence measures, in: Proceedings of the eighth ACM international conference on Web search and data mining, pp. 399–408. S´anchez-Franco, M.J., Rey-Moreno, M., 2022. Do travelers’ reviews depend on the destination? an analysis in coastal and urban peer-to-peer lodgings. Psychology & marketing 39, 441–459. 34 Sia, S., Dalmia, A., Mielke, S.J., 2020. Tired of Topic Models? Clusters of Pretrained Word Embeddings Make for Fast and Good Topics too!, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics, Online. pp. 1728–1736. doi:10.18653/v1/2020.emnlp-main.135. Sun, Y., Cheng, C., Zhang, Y., Zhang, C., Zheng, L., Wang, Z., Wei, Y., 2020. Circle loss: A unified perspective of pair similarity optimization, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 6398–6407. Teh, Y.W., Jordan, M.I., Beal, M.J., work(s):, D.M.B.R., 2006. Hierarchical Dirichlet Processes. Journal of the American Statistical Association 101, 1566–1581. arXiv:27639773. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I., 2017. Attention Is All You Need, in: 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA. pp. 1–15. arXiv:1706.03762. Wang, J., Zhang, X.L., 2023. Deep nmf topic modeling. Neurocomputing 515, 157–173. Warner, B., Chaffin, A., Clavi´e, B., Weller, O., Hallstr¨om, O., Taghadouini, S., Gallagher, A., Biswas, R., Ladhak, F., Aarsen, T., et al., 2024. Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference. arXiv preprint arXiv:2412.13663 . Wu, X., Nguyen, T., Luu, A.T., 2024. A survey on neural topic models: methods, applications, and challenges. Artificial Intelligence Review 57, 18. Yau, C.K., Porter, A., Newman, N., Suominen, A., 2014. Clustering scientific documents with topic modeling. Scientometrics 100, 767–786. You, H., Zhang, X.j., 2009. Financial reporting complexity and investor underreaction to 10-k information. Review of Accounting studies 14, 559– 586. 35 Zhao, H., Phung, D., Huynh, V., Jin, Y., Du, L., Buntine, W., 2021. Topic modelling meets deep neural networks: A survey. arXiv preprint arXiv:2103.00498 . 36 Appendix A. Wordclouds BERTopic models Figure A.6: Wordcloud - FinTextSim - Sentences. The color of each word represents its associated unique topic from the keyword list. Words colored in black are not present in the keyword list. Words colored in darkred are bigrams containing words from multiple keyword domains. 37 Figure A.7: Wordcloud - FinTextSim - Refined Sentences. The color of each word represents its associated unique topic from the keyword list. Words colored in black are not present in the keyword list. Words colored in darkred are bigrams containing words from multiple keyword domains. 38 Figure A.8: Wordcloud - AM - Sentences. The color of each word represents its associated unique topic from the keyword list. Words colored in black are not present in the keyword list. Words colored in darkred are bigrams containing words from multiple keyword domains. 39 Figure A.9: Wordcloud - FinTextSim - Refined Sentences. The color of each word represents its associated unique topic from the keyword list. Words colored in black are not present in the keyword list. Words colored in darkred are bigrams containing words from multiple keyword domains. Appendix B. Algorithms - Classical Topic Modeling Approaches For a deeper understanding of classical topic modeling approaches, we highlight two Bayesian models: LDA and Hierarchical Dirichlet Process (HDP), and one algebraic model: Non-Negative Matrix Factorization (NMF). The Bayesian models define a hypothetical generative process for documents, then work backwards to infer the topics that could have generated the documents (Abdelrazek et al., 2023). In contrast, NMF factorizes a documentterm matrix into term-topic and topic-document matrices (Lee and Seung, 1999). All three models operate under the BoW assumption, viewing each document as a mixture of underlying topics and each topic as a mixture of words. Hence, the algorithms assign prevalence of terms to topics (β) and topics to documents (γ) (Blei et al., 2003; Teh et al., 2006). To ensure and enhance the performance of the presented classical models, several preprocessing steps are essential. These include tokenizing documents, removing stopwords, and lemmatization or stemming of words (Bellstam et al., 2021; Fu et al., 2021; Albalawi et al., 2020). The following subsections provide an 40 overview of classical topic modeling techniques. Appendix B.1. Latent Dirichlet Allocation The most prevalent topic modeling approach in literature is LDA. LDA is a three-level parametric hierarchical Bayesian model, relying mainly on three hyperparameters (Blei et al., 2003): the number of topics (k), the concentration parameter of the Dirichlet prior of the document-topic distribution (α), and the parameter controlling the distribution of words across topics (η) (Fernandes et al., 2020). Each hyperparameter significantly influences the quality and stability of the generated topics. Yet, their selection remains challenging due to the inherent complexity of textual data (Maier et al., 2018; Agrawal et al., 2018). Despite its widespread use, LDA has notable limitations. Agrawal et al. (2018) discovered that LDA is sensitive to the order of training data, i.e., the topics vary when the training data is shuffled. Hence, systematic errors are incorporated into studies. By tuning the model’s hyperparameters, these effects can be mitigated (Agrawal et al., 2018). Moreover, since LDA extracts topics from word distributions independently, overlapping topics can occur (Campbell et al., 2015). LDA has been used in various fields. Bao and Datta (2014) pioneered the integration of unsupervised learning methods into Management Accounting and Finance using LDA to analyze risk disclosures from 10-K reports. Dyer et al. (2017) examined topics contributing to the lengthening of 10-K reports over time, while Brown et al. (2020) identified topics predicting financial misreporting. In additional financial research studies, LDA has been used to quantify the economic content in communications, identify central subjects, or estimate innovation capabilities, among other applications (Jegadeesh and Wu, 2017; Lowry et al., 2020; Bellstam et al., 2021; Garc´ıa-M´endez et al., 2023; Gao et al., 2025). Appendix B.2. Hierarchical Dirichlet Process HDP offers the flexibility of inferring the number of topics from the data itself, accommodating potentially infinite topics. Hence, this non-parametric hierarchical Bayesian model is particularly advantageous when the number of topics is uncertain (Teh et al., 2006). However, the non-parametric nature of HDP increases the computational complexity (Fu et al., 2021; Fernandes et al., 2020). Additionally, hyperparameter tuning is essential to enhance the quality and stability of topics. Each hyperparameter significantly influences 41 Combining guidance and tuning yields the best topic-precision when using tf-idf weighting, while having no effect with tf weighting. Guided and tuned LDA with tf-idf weighting achieves the highest topic-precision among classical approaches with a score of 0.786. Yet, the model misses three financial topics, namely Liquidity, Financing and Tax and Regulation. Hence, we conclude that none of the classical models reach the performance of contemporary topic modeling techniques, particularly when paired with FinTextSim (see Table 3). Table D.9: Coherence Scores - Classical Approaches. Type Model Whole document Sentences Refined Sentences OTS LDA tf 0 (0.008) 0.009 (0.055) 0.002 (0.011) LDA tfidf -0.172 0.004 (0.021) 0 (0.002) HDP tf (-0.330) (-0.327) (-0.330) HDP tfidf (-0.330) (-0.327) (-0.330) NMF tf 0 (0.031) 0.025 (0.178) 0.056 (0.168) NMF tfidf (-0.151) 0.041 (0.185) 0.062 (0.199) Guided LDA tf 0 (0.008) 0.002 (0.029) 0.006 (0.026) LDA tfidf (-0.045) 0.006 (0.024) 0.043 (0.017) Tuned LDA tf 0 (0.023) - - LDA tfidf 0.006 (0.083) - - HDP tf (-0.328) - - HDP tfidf (-0.328) - - NMF tf 0 (0.031) 0.089 (0.201) 0.069 (0.203) NMF tfidf (-0.046) 0.093 (0.247) 0.073 (0.251) Guided - Tuned LDA tf 0 (0.008) - - LDA tfidf 0.049 (0.063) - - Table D.9 presents the coherence scores for classical topic modeling techniques, with non-topic-precision-weighted NPMI coherence shown in paren48 theses. Incorporating topic-precision weighting reveals a critical weakness across all classical models: they fail to generate precise and meaningful topics within our financial dataset. Even the best-performing models exhibit severely diminished coherence when accounting for topic quality, demonstrating that none are truly effective for our domain. While tuning and guiding approaches improve raw coherence scores, they fail to address the core issue of generating topics with economic relevance and precision. Even refined-sentence based tfidf-weighted NMF, which achieves the highest raw coherence, remains unreliable when topic-precision is considered. This highlights a fundamental limitation of classical techniques in handling financial disclosures. Neglecting topic-precision weighting, NMF with sentence-based input achieves the highest coherence. NMF’s superior performance indicates its suitability for non-mainstream text, producing more coherent topics than LDA and HDP for our domain-specific dataset (O’Callaghan et al., 2015). LDA produces topics with lower coherence than NMF, consistent with Egger and Yu (2022); O’Callaghan et al. (2015); Chen et al. (2019), but contrary to observations of Albalawi et al. (2020) and Farzadnia et al. (2024). HDP ranks last in terms of topic coherence, coinciding with Mayasari et al. (2021) and Farzadnia et al. (2024). Contrary to Altaweel et al. (2019), our results do not support the assertion that both HDP and LDA are capable of finding good topics, particularly for HDP. For whole-document OTS models, tf-weighting consistently yields higher coherence than tfidf-weighting. LDA achieves the highest coherence, followed by NMF, while HDP remains ineffective. Using sentence-based input for OTS models, we find varying impacts: NMF and LDA show a notable enhancement, whereas HDP’s coherence grows only slightly. The effect on NMF and LDA is even more significant for tfidfweighting. For both weighting methods, NMF continues to outperform the other classical OTS models. The significant leap for NMF shows its superior performance for short-text modeling, aligning with Chen et al. (2019). Further refining input sentences of OTS models leads to growing coherence only for tfidf-weighted NMF. All other models face a reduction. Based on the observed reduction of coherence scores between regular sentence-based and refined sentence-based models, we find that the models not only handle noise but leverage it to increase topic coherence. We assume that the increased sparsity of the input, induced by tfidf-weighting, helps the topic allocation, eventually improving coherence (Lee and Seung, 1999). Tuning the models leads to increased coherence scores with more significant improvement ob49 served for tfidf-weighting. We attribute this higher leverage to the nature of the financial vocabulary. In Finance, topic words tend to occur frequently within documents. Hence, tfidf-weighting downweighs their importance, resulting in increased sparseness and improved allocation of topics (Lee and Seung, 1999). Guiding the models has only marginal effects on coherence, with the strongest impact observed for tfidf-weighted LDA. Hence, we only partially concur with Chen et al. (2019) that incorporating domain knowledge enhances coherence. Incorporating domain knowledge alongside hyperparameter tuning results in reduced coherence compared to purely tuned LDA models. This is contrary to our initial expectations; Guided-Tuned LDA has lower coherence than tuned LDA. We assume that guiding the model restricts the range of possible hyperparameters, limiting optimization capabilities. Overall, we conclude that BERTopic models significantly outperform classical models (see Tables D.9 and 4), aligning with the findings of Abuzayed and Al-Khalifa (2021) and Egger and Yu (2022). While non-weighted coherence scores suggest that tuned tfidf-weighted NMF performs best among classical models, incorporating topic-precision reveals their fundamental limitations. No classical approach is capable of generating both precise and coherent topics for financial disclosures.13 The reliance on modern techniques such as BERTopic is, therefore, essential for producing reliable financial topic models. Given that classical models fail to produce meaningful economic topics consistently and reliably, we shift our focus toward enhancing contemporary topic modeling techniques. To this end, we introduce FinTextSim, designed to capture the unique characteristics of financial language and improve topic modeling for the financial domain. Furthermore, as the topics generated by classical models are inherently weak, their separation becomes irrelevant. Consequently, we deliberately exclude the comparisons of topic similarity for classical models. Appendix E. Excerpt of Keyword List •Sales: sale, revenue, market, consumer, demand, competition, pricing, contract, price 13Wordclouds for a selection of models are displayed in Appendix F. 50 •Cost: cost, expense, liability, goodwill, impairment, depreciate, depreciation •Profit/Loss: profit, performance, result, margin, income, earnings, loss, management, ebitda, ebda, ebit •Operations: operations, production, business, produce, supply, operational, producer, process, processing, manufacturing, manufacture, supplychain, logistics, transport, marketing, advertising, advertisement, advertise •Liquidity: liquidity, interest, coverage, cash, capital, balance, cashflow, excess, working capital, work capital, excess cash •Investment: investment, expenditure, m&a, divestiture, invest, asset, disposal, divestment, •Financing: financing, finance, debt, equity, dividend, repurchase, share, funding, security, indebtness, indebtedness, borrowing, credit •Litigation: litigation, lawsuit, legal, matter, dispute, complaint, arbitration, patent •HR: employee, retention, hiring, hire, union, consultant, staff, recruiting, recruit, recruitment, labor, incentive, insurance, team, training, salary, wage, job, work •Regulation: regulations, tax, government, tax expense, government affair, legislation, federal, regulator, regulate •Accounting: accounting, account, audit, auditing, control, adjustment, filing, auditor, report •Energy: energy, coal, solar, fuel, wind, water, electric, oil, megawatts, mwh, megawatthours, kilowatts, kwh, kilowatthours, gigawatts, gwh, gigawatthours •ESG: plastic, recycle, waste, carbon, emission, renewable, environment, sustainable, sustain, ecologic, ecological •Covid-19: covid, cov-19, pandemic, disease, corona, covid-19, sars-cov 51 Appendix F. Topic Representations - Classical Approaches Figure F.10: Wordcloud - LDA - Whole Document - tf-weighting. The color of each word represents its associated unique topic from the keyword list. Words colored in black are not present in the keyword list. 52 Figure F.11: Wordcloud - LDA - Whole Document - Guided - tfidf-weighting. The color of each word represents its associated unique topic from the keyword list. Words colored in black are not present in the keyword list. 53 Figure F.12: Wordcloud - LDA - Whole Document - Guided-Tuned - tfidf-weighting. The color of each word represents its associated unique topic from the keyword list. Words colored in black are not present in the keyword list. 54 Figure F.13: Wordcloud - HDP - Whole Document - Tuned - tf-weighting. The color of each word represents its associated unique topic from the keyword list. Words colored in black are not present in the keyword list. 55 Figure F.14: Wordcloud - NMF - Sentences - tf-weighting. The color of each word represents its associated unique topic from the keyword list. Words colored in black are not present in the keyword list. 56 Figure F.15: Wordcloud - NMF - Sentences - Tuned - tfidf-weighting. The color of each word represents its associated unique topic from the keyword list. Words colored in black are not present in the keyword list. 57