Full text
Universidade do Minho Escola de Engenharia António Manuel Almeida Pinto Bandeira Santos Natural Language Processing Applied to Human Resources October, 2023
Universidade do Minho Escola de Engenharia António Manuel Almeida Pinto Bandeira Santos Natural Language Processing Applied to Human Resources Master Thesis Master in Informatics Engineering Work developed under the supervision of: José João Antunes Guimarães Dias Almeida Luís Filipe Costa Cunha October, 2023
COPYRIGHT AND TERMS OF USE OF THIS WORK BY A THIRD PARTY This is academic work that can be used by third parties as long as internationally accepted rules and good practices regarding copyright and related rights are respected. Accordingly, this work may be used under the license provided below. If the user needs permission to make use of the work under conditions not provided for in the indicated licensing, they should contact the author through the RepositoriUM of Universidade do Minho. License granted to the users of this work Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International CC BY-NC-SA 4.0 https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en iii
Acknowledgements The conclusion of this dissertation marks the end of a stage in my life that spanned the last 18 years. This stage, and everyone involved in it, has profoundly shaped my life, and I wouldn’t have had it any other way. Firstly, I would like to express my gratitude to Daniel Oliveira and everyone at Konkconsulting for providing the means and support for the completion of this document. I would also like to thank Prof. José João Almeida and Filipe Cunha for supervising and assisting in making this thesis the best it could be. Special thanks go to my godfather, Jorge Simões, for generously offering his time, insight, and expertise to ensure the proper refinement of my thesis. Additionally, I’m grateful to ChatGPT for rectifying the broken English present in this document. I want to extend my appreciation to my entire family for supporting me throughout this journey, even when I diverged from the path and became the first person in two generations of my family not to pursue medicine. I’d like to offer a special thanks to my parents, grandparents, and my sister for playing a pivotal role in shaping the person I am today and ensuring my success in life. I would also like to acknowledge all my friends and close people in my life, ’À equipa,’ for providing me with great times and also keeping me on track during moments of heavy procrastination. To Né and Sara, who stood by my side throughout the last academic years, making this journey much brighter. Additionally, I would like to express my gratitude to Sara for being the person she is and being my mental safe haven this last year, always being there for me whenever I needed support, truly a beacon of happiness in my life. Last but not least, I want to thank my furry little friend, Jagu, for simply existing and being the best cat any owner could dream of. Every single one of you who had an impact in my life indirectly had a part in the inception of this dissertation and for that i would like to thank you all. Thank you! v
STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the Universidade do Minho. Universidade do Minho, Braga, October 2023 António Manuel Almeida Pinto Bandeira Santos vii
3.3 Preprocessing .................................. 27 3.3.1 Tokenization............................... 27 3.3.2 Stopwords................................ 28 3.3.3 Lemmatization.............................. 28 3.3.4 Embedding ............................... 29 3.4 Workingagainstbias ............................... 30 3.4.1 Undersampling.............................. 30 3.4.2 Oversampling .............................. 31 3.4.3 Weightdistribution............................ 32 3.5 Models...................................... 32 3.5.1 Classification or Regression - What to use? . . . . . . . . . . . . . . . . 32 3.5.2 LogisticRegression............................ 33 3.5.3 Support Vector Machine (SVM) . . . . . . . . . . . . . . . . . . . . . . 33 3.5.4 Long Short-Term Memory (LSTM) . . . . . . . . . . . . . . . . . . . . . 33 3.5.5 Transformers .............................. 34 3.5.6 Hyperparameters............................. 34 4 Solution Implementation 37 4.1 Usedtechnologies ................................ 37 4.2 UsedHardware.................................. 38 4.3 DatasetAnalysis ................................. 38 4.3.1 Dataset Implementation . . . . . . . . . . . . . . . . . . . . . . . . . 38 4.3.2 DatasetComparison ........................... 39 4.4 Pre-processing .................................. 43 4.5 ResultAnalysis.................................. 44 4.5.1 ModelComparison............................ 44 4.6 FinalArchitecture................................. 49 4.6.1 Hyperparameter tuning . . . . . . . . . . . . . . . . . . . . . . . . . . 50 4.6.2 FinalResults............................... 51 4.7 API........................................ 51 5 Conclusion and Future Work 55 5.1 Conclusion.................................... 55 5.2 Limitations and Future Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56 5.3 FinalConsiderations ............................... 56 Bibliography 59 xiv
List of Figures 1 Workplan ...................................... 4 2 EmotionCircumplexModel .............................. 9 3 Plutchik’sModel ................................... 9 4 KeywordSpotting................................... 12 5 RuleBasedApproach................................. 12 6 A Classical Learning Approach (SVM) . . . . . . . . . . . . . . . . . . . . . . . . . 12 7 LongShort-TermMemory............................... 13 8 Solution simple architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 9 DailyDialogbarplot.................................. 24 10 Meldemotioncount.................................. 25 11 Meldpolaritycount.................................. 26 12 Goemotionsemotioncount .............................. 27 13 VectorizationGraph.................................. 30 14 Filtered Daily Dialog bar plot . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 15 FinalEmotionDistribution............................... 42 16 Daily Dialog Disgust WordCloud . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 17 Daily Dialog Happiness WordCloud . . . . . . . . . . . . . . . . . . . . . . . . . . 43 18 Daily Dialog Neutral WordCloud . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43 19 Preprocessingpipeline ................................ 44 20 Testedmodels .................................... 45 21 F1-ScoreComparison................................. 47 22 Trainingtimecomparison............................... 48 23 ModelArchitecture .................................. 49 24 APIWebpage..................................... 52 25 API Web page multiple emotions example . . . . . . . . . . . . . . . . . . . . . . . 53 xv
List of Tables 1 ResearchSources................................... 6 2 SearchTerms..................................... 6 3 HardwareUsed.................................... 38 4 EmotionDistribution ................................. 39 5 Baseline Algorithms Comparison . . . . . . . . . . . . . . . . . . . . . . . . . . . 45 6 RNNComparison................................... 46 7 TestedTransformers ................................. 46 8 AccuracyperLabeltable ............................... 47 9 TransformerComparison ............................... 48 10 OptimalHyperparameters............................... 51 11 Performance of final solution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51 12 Emotion-EmojiMapping................................ 52 xvii
List of Listings 1 Datasetclassstructure .............................. 39 2 APIresponse................................... 53 xix
Acronyms AI Artificial Intelligence ix, 2 API Application Programming Interface 3 BERT Bidirecional Encoder Representations from Transformers ix, 18 BiGRU Bidirectional Gated Recurrent Unit 17 BiLSTM Bidirectional Long-Short Term Memory 16 CBET Cleaned Balanced Emotional Tweets 13 CNN Convolutional Neural Network ix, 16 DCNN Deep Convolutional Neural Network 13 distilBERT Distilled BERT 18, 19 EARL Emotion Annotation and Representation Language 8 EWE Emotion Word Embeddings 16 GloVe Global Vectors for Word Representation 14, 16 GPT Generative Pre-trained Transformer 2 GRU Gated Recurrent Unit 16 HR Human Resources ix, 1, 2, 10, 17 IEMOCAP Interactive Emotional Dyadic Motion Capture 10 ISEAR International Survey on Emotion Antecedents and Reactions 11, 18 LSTM Long Short-Term Memory 13 MELD Multimodal EmotionLines Dataset 10 xxi
NLP Natural Language Processing ix, 2 RNN Recurrent Neural Network ix, 13 RoBERTa Robustly Optimized BERT Pretraining Approach 18, 19 SEER Semantic Emoticon Emotion Recognition 17 SENN Semantic-Emotion Neural Network 16 SVM Support Vector Machine xv, 12, 16, 17 xxii
Chapter 1 Introduction 1.1 Context Currently, having employees experienced in a certain domain is essential for any company to thrive in the competitive business world. As a result, companies have to be at the top of their game to keep their workforce happy in order to assure the retention and evolution of the workers, which in turn can provide better work efficiency and growth of the organization. These concepts, known as talent management and employee retention, are key concepts for the human resources department, and it is its role to handle those correctly. According to the U.S. Bureau of Labor Statistics [1] an employee stays on average 4.1 years at a given company. This value reflects the difficulty companies currently have in retaining its workforce, which in turn drastically impacts the company’s efficiency and its ability to grow, mainly due to dissatisfaction of one of the parties. In addition, employee retention is crucial for a company’s economic growth. Research indicates that the cost of replacing an employee is roughly one third of their annual salary [2]. This includes the decreased productivity of the workforce, as well as the expenses of recruiting and training a replacement. As time passes, the cost of employee turnover is steadily increasing. In 2021, this cost had reached over 700 billion dollars in the United States alone [3]. One of the main reasons for employee departure is dissatisfaction with their job. In 2021, 8% of all employment contract terminations were related to dissatisfaction with management [3]. This can be caused by poor communication between management and the workforce. A satisfied employee is more productive and 87% less likely to leave their job compared to an unsatisfied one, according to studies [4]. The domain of Human Resources (HR) is an area that relies heavily on communication between a company and its employees to ensure that both parties are satisfied. This means that it is essential 1
CHAPTER 2. STATE OF THE ART Emotions can be classified discretely or dimensionally. As a rule of thumb discrete models tend to be simpler but can only classify basic emotions while dimensional models are able to define emotions that are a result of a mix of more primitive emotions. 2.3.1.1 Discrete classification models According to William James [11] emotion can be categorized into four basic sections, fear, grief, love, and rage. His concept relies on the study of organic reverberation or in simpler terms body language. Ekman [12] expands this model by defining an emotion as 6 possible concrete types which are anger, disgust, fear, happiness, sadness, and surprise. This model is based on the facial expressions made when reacting to an event. Richard and Bernice Lazarus [13] classify emotions with a deck of 15 possible types of aesthetic experience, anger, anxiety, compassion, depression, envy, fright, gratitude, guilt, happiness, hope, jealousy, love, pride, relief, sadness, and shame. There are many more discrete classification models with a varying number of possible emotion types, some, like the Emotion Annotation and Representation Language (EARL) [14] contain over 45 different types of emotion. However, depending on the use case, some emotion types may be discarded. In this study only a few emotions might be useful for an efficient solution, so simpler models should be used 2.3.1.2 Dimensional classification models A general consensus is that any emotion can be classified by assigning a value to each of the three following dimensions: • Valence: Represents how positive or negative the emotion is. • Arousal: Represents how much energy is in the emotion • Power: Represents the degree of power the emotion has. Posner, Russel, and Peterson’s circumplex model [15] classifies emotions bi-dimensionally by regarding only the valence and arousal dimensions. 8
2.3. RESULTS DISCUSSION Figure 2: Emotion Circumplex Model Plutchik [16] handles the concept of emotions multidimensionally by creating different models defined by 8 main emotion types, each of which blend and create more precise subtypes. Figure 3: Plutchik’s Model Depending on how precise emotions need to be defined, different models can be used. While better emotion precision may seem desirable it makes the task harder for the computer to process, so it might be wise to prioritize machine learning model efficiency over emotion precision. 9
CHAPTER 2. STATE OF THE ART For this study’s goal, a polarity concept might also be viable since it is enough to know if an emotion is positive or negative without knowing which type it is. This simplifies the problem and makes it trivial for a computer to recognize the polarity. 2.3.2 How can a computer recognize emotion behind a message accurately? In Poria et al. (2019) [17] a study was made about the environment of emotion recognition in conversations at the time. It highlights some challenges one might face when tackling a problem in this domain. One of those challenges is the definition and categorization of an emotion, which is already addressed in the previous section. Another challenge is the importance of context in a conversation. Emotions can be hidden under the context of a conversation which makes it impossible to extract them if one only has the grasp of each message individually. This affects the viability of models trained with context-independent data such as tweets since this kind of messages are pretty self-contained and are not influenced by past messages. It also highlights some nuances that are usually present in conversations, such as emotional shifts, the presence of sarcasm among others. Since this study focuses on conversations between an HR representative and an employee, we can expect a more formal way of conducting conversations. However there can always be edge cases where those nuances appear that need to be addressed. The study also enumerates and compares some datasets with relevant data that can be used to train machine learning models, capable of extracting emotions from conversations. Some datasets like the Interactive Emotional Dyadic Motion Capture (IEMOCAP) [18], SEMAINE [19] and Multimodal EmotionLines Dataset (MELD) [20] contain data in multiple formats (audio, video, and text) while DailyDialog [21], EmotionLines [22] and EmoContext [23] only contain data in text form. For this study’s scope, only data in text form will be relevant. Out of the datasets mentioned, only SEMAINE defines emotions dimensionally by labeling the data with four values: arousal, valence, expectancy, and power, the other datasets only contain categorical emotion labels. MELD is a dataset dedicated to multiparty conversations, something that will not be considered in the current project’s scope. The study concludes by comparing different state-of-the-art models by using the mentioned datasets in the training phase and testing their performance against each other. In Alswaidan et al. (2020) [24] a survey regarding state-of-the-art approaches for emotion recognition in text was made. Here, a lot of useful information for someone looking to start a project in the area is condensed. Like the previous study, it starts by enumerating the different ways that an emotion can be defined. While this has already been discussed in this study, the survey expands by adding a third emotion modeling approach, the Appraisal modeling approach. This approach can be seen as an extension of the dimensional approaches. Its model bases itself on the appraisal theory, which states that an emotion can be extracted by evaluating the events a person has been subjected to and basing the emotion by studying 10
2.3. RESULTS DISCUSSION the person’s experience, goals, and opportunities for actions. In this scenario, emotions are extracted by studying changes in many components of the human mind including cognition, physiology, feelings, expressions, etc. Most of these components cannot be evaluated in this study’s scope, so this approach is not viable to use in the current context. The survey then starts listing useful resources to train models to extract emotion. It lists and compares six datasets. The first one is Alm which consists of 185 children’s stories, where each one is labeled as neutral, anger-disgust, sadness, fear, happiness, positive surprise, and negative surprise. Aman consists of blog posts labelled with one of the six emotions from Ekman’s model [12]. International Survey on Emotion Antecedents and Reactions (ISEAR) [25], a survey conducted, where people labeled some of their experiences as joy, fear, anger, sadness, disgust, shame, and guilt. SemEval-2007 [26] compiles headlines from different news sources and labels them using Ekman’s model [12] as well. SemEval-2018 [27] consists of tweets labeled as anger, disgust, fear, joy, love, optimism, pessimism, sadness, surprise, trust, or neutral. SemEval-2019 is another name given to the EmoContext dataset mentioned in [17]. It consists of three turns of exchanges between two individuals where each exchange is either labeled as joy, anger, sadness, or others. For the current context this is the datasets that seem the most viable to use among the mentioned ones. The survey then lists the amount of data available in each dataset and inserts it into a table for comparison. The survey also enumerates some lexicons related to the problem in question. Lexicons are collections of words and/or phrases along with their associated information, such as their part of speech, meaning, and context. The lexicons mentioned add additional emotional information regarding certain words, simplifying the learning process of emotion recognition models and increasing their overall performance. It lists and briefly explains sixteen lexicons. Some approaches used to recognize emotion are then listed along with some studies that use them. The first approach, the keyword-based approach, relies on finding occurrences of certain keywords related to a specific label. The most common technique is the keyword-spotting technique. Here each emotion is defined by a list of keywords using lexicons and then after some text pre-processing each remaining word from a sentence is iterated and checked against the keyword lists. Depending on how many words from the sentence are inside each emotion’s keyword list, the emotion label for the sentence is determined. Figure 4 illustrates an example of a keyword-based approach. 11
CHAPTER 2. STATE OF THE ART Figure 4: Keyword Spotting The second approach, the rule-based approach, consists of manipulating the information present in order to be able to extract a conclusion. The information is pre-processed, then, the emotion rules are extracted using linguistic, statistics, and computational concepts and the best ones are selected. Finally, those rules are applied to the dataset in order to extract the emotion labels. Figure 5: Rule Based Approach The third approach, the classical learning based approach, provides the system with the ability to automatically learn and improve from experience. These use the more basic forms of machine learning algorithms, the most used one is Support Vector Machine (SVM),shown in Figure 6, which is a supervised algorithm. In this approach, a dataset’s content is processed, and its most useful features are extracted, after that it is fed into the algorithm which will then output a model capable of predicting the labels of unseen data. Figure 6: A Classical Learning Approach (SVM) 12
2.3. RESULTS DISCUSSION The fourth approach, the Deep-learning approach, is a branch of machine learning where programs learn from experience and understand what’s around them in terms of a hierarchy of concepts. With this method, each concept is created in terms of the relation of simpler concepts. This allows programs to understand complicated concepts by relating them with simpler ones. The most common deep learning model for emotion recognition is the Long Short-Term Memory (LSTM), found in Figure 7. This model expands on Recurrent Neural Network (RNN) by giving it the capability of handling long term dependencies. This means that this model is able to extract emotions by using the whole conversation as a medium, instead of just being able to relate the emotion to the message in which it is embedded. To use these models first the dataset is processed and an embedding layer is created and fed into one or more LSTM layers. Then its output is fed into a dense neural network properly configured to perform the emotion classification. Figure 7: Long Short-Term Memory In [28], a multi-labeled Deep Convolutional Neural Network (DCNN) was introduced to detect at least two emotions in text. The research concluded that most emotion datasets contain multiple extractable emotions. Since many existing solutions only detect the most prominent emotion in their input, this study suggests potential methods for enhancing the overall ecosystem. This approach consists of an architecture composed of multiple components, each one with a distinct task. The first component is the dataset. In order for the model to learn, it needs a large amount of correctly labeled data. This enables the model to extract the information present in the data and extract its own conclusions about what influences the data in order to give its respective label. The method used in this solution was to extract and combine all data with more than one label from two already available datasets (Cleaned Balanced Emotional Tweets (CBET) [29] and semEval). After composing the dataset, it is necessary to process the data for the purpose of optimizing the learning phase of the solution. This approach manipulated the dataset in the following ways: •Data cleanup - All irrelevant data such as usernames, links, etc., only add noise to the dataset which makes it harder and more computationally intensive to analyze it and as such are removed. 13
CHAPTER 2. STATE OF THE ART •Deleting Stop-Words - These are some of the most common words used in standard conversations and, like the data referred above, add little to no value for the learning algorithms. Some examples of stop words are conjunctions, suffixes, and pronouns which are also removed from the whole dataset. •Tokenizing - The process of tokenizing consists of converting all text into tokens which is a more refined way of representing words. These tokens can then be transformed into a vector and fed into the algorithm, improving its learning efficiency. Not all words have the same impact when defining an underlying emotion in a piece of text, so even with all relevant words stored in a way that is easily readable by the machine might not be enough to produce an efficient model. The study’s approach to this problem is adding an extra layer, which they called ”Attention Layer”, to the architecture capable of calculating the weight each word has by creating ” a distributed representation of the entire document with respect to the highlighted words ”. The final result will represent the importance the word has in regard to the definition of the text’s label. After allocating an importance value to each word data is fed to the ”Word Embedding Layer”.To tackle this layer the study used two different approaches, one that works in tune with all other layers in order to learn the semantic vector of the words itself by training and giving weights to the words . The other uses two pre-trained models, fastText [30] and Global Vectors for Word Representation (GloVe) [31], in order to compare their performance with the first approach. The last layer of the whole solution’s architecture is the deep learning model itself. This layer is composed of many components whose task is to perform mathematic operations on the vectors given by the previous layers and, finally, to be able to produce an accurate prediction of the emotion present in the given input. The whole architecture is too complex to provide a brief summary of its inner workings. In order to evaluate the performance of the developed method, the study used the following metrics: •Jaccard index: This metric calculates the accuracy of the multi-labeling process by dividing the number of correctly predicted labels by the total number of predicted labels. 𝐽𝑎𝑐𝑐𝑎𝑟𝑑 𝑖𝑛𝑑𝑒𝑥 = 1 𝑁 𝑁 Õ 𝑖=1 𝑌𝑖∩ˆ 𝑌𝑖 𝑌𝑖∪ˆ 𝑌𝑖 (2.1) Where: 𝑌Correct Label ˆ 𝑌Predicted Label 𝑁Number of samples •Hamming Loss: The Hamming Loss metric computes the fraction of the mislabeled samples by the total number of labels. 14
2.3. RESULTS DISCUSSION 𝐻𝑎𝑚𝑚𝑖𝑛𝑔 𝐿𝑜𝑠𝑠 = 1 𝑑𝑙 𝑑 Õ 𝑖=1 𝑙 Õ 𝑗=1ℎ𝑖 𝑗 Δ𝑦𝑖 𝑗 (2.2) Where: 𝑦Correct Label set ℎPredicted Label set 𝑑Sample 𝑙Label For the rest of the metrics, we can assume that: 𝑇𝑃 True Positive 𝑇 𝑁 True Negative 𝐹𝑃 False Positive 𝐹𝑁 False Negative •Micro Average Precision: The micro average formulas denote the behavior of individual classes. This formula calculates the precision of each class by dividing the sum of all true positive predictions by the sum of all positive predictions, regardless of being correct or not. 𝑀𝑖𝑐𝑟𝑜 𝑎𝑣𝑒𝑟𝑎𝑔𝑒 𝑝𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 =Í𝑖 𝑗=1𝑇𝑃𝑗 Í𝑖 𝑗=1𝑇𝑃𝑗+Í𝑖 𝑗=1𝐹𝑃𝑗 (2.3) •Micro Average Recall: Like the previous formula, this one also denotes the behavior of each individual class. It calculates the recall of each class by dividing the sum of all true positive predictions by the sum of all positive results. 𝑀𝑖𝑐𝑟𝑜 𝑎𝑣𝑒𝑟𝑎𝑔𝑒 𝑟𝑒𝑐𝑎𝑙𝑙 =Í𝑖 𝑗=1𝑇𝑃𝑗 Í𝑖 𝑗=1𝑇𝑃𝑗+Í𝑖 𝑗=1𝐹𝑁𝑗 (2.4) •Macro Average Precision: The macro average precision is the arithmetic mean of all the precision values for the different classes. 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 = 𝑇𝑃 𝑇𝑃 +𝐹𝑃 (2.5) 𝑀𝑎𝑐𝑟𝑜 𝐴𝑣𝑒𝑟𝑎𝑔𝑒 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 =Í𝑘 𝑚=1𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛𝑘 𝑘(2.6) •Macro Average Recall: The macro average recall is the arithmetic mean of all the recall values for the different classes. 𝑅𝑒𝑐𝑎𝑙𝑙 = 𝑇𝑃 𝑇𝑃 +𝐹𝑁 (2.7) 15
CHAPTER 2. STATE OF THE ART 𝑀𝑎𝑐𝑟𝑜 𝐴𝑣𝑒𝑟𝑎𝑔𝑒 𝑅𝑒𝑐𝑎𝑙𝑙 =Í𝑘 𝑚=1𝑅𝑒𝑐𝑎𝑙𝑙𝑘 𝑘(2.8) •F1 Score: The F1 score is the harmonic mean of all the precision and recall values. 𝐹1𝑆𝑐𝑜𝑟𝑒 =2∗𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 ∗𝑅𝑒𝑐𝑎𝑙𝑙 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 +𝑅𝑒𝑐𝑎𝑙𝑙 (2.9) The evaluation was made using 5-fold cross-validation, using the model by itself and in conjunction with both previously mentioned word embedding models separately. By the time this study came out, the values provided by its metrics showed an all around better performance in comparison to other studies that were available at the time. In Batbaatar et al (2019) [32], a new Neural Network architecture was proposed, the SemanticEmotion Neural Network (SENN), a model capable of obtaining and utilizing both semantic and emotional information by adopting already trained word representations. This leads to better quality results in comparison to other state-of-the-art approaches, and can be improved even further with other emotional word embeddings. The architecture is mainly composed of two subnetworks. The first subnetwork utilizes Bidirectional Long-Short Term Memory (BiLSTM) in order to uncover the context of each word by analyzing the information from both the forward and backward directions. The second subnetwork uses a Convolutional Neural Network (CNN) to extract emotional features and concentrates on the emotional connections between words and the text as a whole. To train the model, various publicly available datasets were used separately, allowing for a comparison of the model performance between them. All datasets were filtered in order to only include sentences that contained just one label and where the emotion associated with it was included in Ekman’s [12] model of emotions (joy, fear, sadness, surprise, anger, and disgust). In those sentences a cleanup was executed where all numbers, special characters and usernames were removed and all uppercase characters were changed to their lower case variant. The optimal parameters of the models were also obtained using the grid search algorithm in order to achieve the best performance out of each model tested. The performance of each model was evaluated using the Recall (2.7), Precision (2.5) and F1-Score (2.9) metrics. The models were then compared against some baseline models, like Naive-Bayes classifier, Random Forest Classifier, Support Vector Machine, among others. These used basic methods like Bagof-Words or Term Frequency-Inverse Document Frequency as the class of each text. They were also compared to more complex deep learning models like Convolutional Neural Network, Gated Recurrent Unit (GRU), and others, using frameworks like Word2Vec [33], GloVe [31] or FastText [30] for semantic word embeddings and Emotion Word Embeddings (EWE)[34] for emotion word embedding to convert the data into a more refined state. 16
2.3. RESULTS DISCUSSION The study concluded that its proposed approach performed better than all the other tested models while having a comparable if not lower execution time, which made their solution the best one to use in many cases. In Liu et al (2021)[35] a model architecture was presented. This new model named Semantic Emoticon Emotion Recognition (SEER) expands on the Bidirectional Gated Recurrent Unit (BiGRU) network by adding an attention mechanism to give weight to more impactful words. SEER also utilizes the presence of emojis to give more accurate predictions instead of discarding them completely like most developed models do. Since emojis can appear in text conversations between HR representatives and employees, this functionality can be beneficial to have in order to obtain the best predictions possible. The architecture starts with a sentence classification layer. This layer is tasked to parse and divide the sentences given into four groups: sentences that contain explicit emotional words, sentences that contain explicit emotional words as well as emoticons, sentences that have words where the emotion is implicit, and sentences that have implicit emotional words and emoticons. All the organized data is then fed to the Word Embedding layer. This layer refines the words given into a word vector representation using the framework Word2Vec which is sent to the next layer of the model. With the words processed into a word vector, the BiGRU layer is capable of extracting semantic features from the word aspect. In order to be able to understand and extract meaning from emojis, this solution utilizes an emoji embedding layer. This layer contains an emoji distribution where each emoji is labeled with six values which represent its weight for each of the 6 emotions present in Ekman’s model [12]. Each weight is computed by gathering sentences that contain emojis from existing labeled datasets and dividing the number of sentences of each emotion where the emoji is present by the total number of sentences where the emoji has appeared. Like the previous models, this architecture also calculates the impact each word has in defining an emotion by using an attention layer. This layer utilizes a self-attention mechanism to compute the weight of the information present in the BiGRU’s output without the need for external data. The output of the emoji embedding and attention layer is then merged in the connection layer and a final vector containing all the relevant data from the previous steps is sent to the final layer, the emotion recognition layer that, with the input given, is able to predict the emotion behind the initial message. In the training phase, the study used a dataset that contained posts from the Chinese platform ”Weibo”. As a result, the model was trained using Chinese characters and all English characters as well as numbers, special characters, Weibo usernames had to be removed in order to eliminate any possible noise in the data. The model’s performance was studied by calculating its macro precision (2.6), macro recall (2.8), and macro F1-Score (2.10) and tested against baseline methods like Word Lexicons, SVMs, BiGRU without emoji support, among others. In the end SEER, outperformed every other method, even the BiGRU implementation without emoji support. This makes it safe to conclude that BiGRU models are a good 17
CHAPTER 3. SOLUTION PLANNING Figure 9: Daily Dialog bar plot 3.2.2 MELD MELD is a dataset composed of 2 160 dialogues from the television series ”Friends”. Each line of each dialogue is composed by: • Utterance • Speaker • Emotion –Neutral –Surprise –Fear –Sadness –Joy –Disgust –Anger • Polarity 24
3.2. DATASET –Negative –Neutral –Positive • Dialogue_ID The dialogue identifier • Utterance_ID The utterance identifier inside each dialogue The remaining columns are not relevant for the goal of this study. Similarly to other datasets dialogues will be flattened and only the utterances will be used. This results in 12,840 unique utterances with their respective emotion and polarity. The column that represented the speaker of each utterance was also removed to generalize the model. Figure 10: Meld emotion count 25
CHAPTER 3. SOLUTION PLANNING Figure 11: Meld polarity count 3.2.3 ScenarioSA ScenarioSA is a dyadic database used for sentiment analysis. It is composed of 2 214 dialogues. Each line of each dialogue is composed by: • Utterance • Speaker (A or B) • Polarity –Negative (-1) –Neutral (0) –Positive (1) It also contains each speaker’s polarity at the end of the dialogue and the topic however these are not relevant for the study. Since this dataset only contains information regarding polarity and not individual emotions, it can be concluded that its usage is not optimal for this study. 26
3.3. PRE PROCESSING 3.2.4 goEmotions GoEmotions comprises 76 536 comments manually labeled by humans from the platform Reddit, with each comment assigned one or more of 27 unique emotions. This approach allows the development of a model capable of predicting finer-grained emotions, however, to maintain consistency with other tested datasets, the labels were streamlined to encompass only the six emotions outlined in Ekman’s model. Figure 12: Goemotions emotion count 3.3 Pre processing Machine learning models work better when handling pure numbers. This results in the necessity to process the data so as to have it in a more appropriate format to feed the models. 3.3.1 Tokenization Tokenization is the process of splitting data into smaller pieces (called tokens). Some models that handle text use words as tokens while others use individual letters. In the context of this study, it involves separating each word from their respective sentence. > I’m sorry that you are mad ! -> [”I’m”,”sorry”,”that”,”you”,”are”,”mad”,”!”] 27
CHAPTER 3. SOLUTION PLANNING The number of tokens present in the training phase should always be kept at a minimum to ensure an optimized process. So, to accomplish that without losing valuable data, some considerations can be made when analysing the dataset. • *You’re* and *You are* will result in different tokens. • *You* and *you* will also result in different tokens due to the different casing. • *You!* and *You* result on different tokens because of *!*. • There are some words that add little relevance to the meaning of a sentence and its removal **may** be beneficial. (Stopwords) • There are some words that can be simplified in order to reduce the number of different tokens in the training data. (Lemmatization) 3.3.2 Stopwords Stopwords are words that are commonly used in any given language. These words usually have no relevance in the development of machine learning models since they are present in almost every sentence and their presence adds little insight for the prediction of the desired label. > removeStopwords(”I am sorry that you are mad !”) -> ”sorry mad !” By removing stopwords we are decreasing the amount of data used to train the model, which in return reduces the time needed for the solution to be fully trained. However, the removal of stopwords can lead to the sentence completely changing its meaning. Stopwords such as ” not ”are crucial for the context of a sentence and thus are essential for the underlying emotion in the sentence. For example: > removeStopwords(”I am not having a great day!”) -> ”great day!” In this example by removing the stopwords the sentence changes from being a Neutral/Sadness sentence to being related to the Happiness emotion, something that would severely impact negatively the model. So, for this study’s domain, this technique is usually not recommended. 3.3.3 Lemmatization The lemmatization of a sentence is the process of converting each word into its *lemma*. Lemmas are the most primitive form of a word, by using lemmas we can ensure that the words derived from the same source will result in the same token. > lemmatize(”Golden foot is the king of dancing! He danced his feet off.”) -> ”Golden foot be the king of dance! He dance his foot off.”) 28
3.3. PRE PROCESSING Lemmatization may not be an easy task to execute, there are some words where their lemma depends on the context of the sentence. Let’s say for example: > lemmatize(”He always leaves the tree leaves scattered on the balcony ”) -> ”he always leave the tree leave scatter on the balcony” This lemmatization does not represent the original sentence correctly. It can be seen that both occurrences of the word ”leaves”should result in different lemmas, in this case in leave and leaf. 3.3.4 Embedding In machine learning, embedding refers to the process of representing data - such as words, sentences, or images, as continuous vectors or numerical representations in a lower-dimensional space. The most basic way of embedding a sentence is by assigning a number to each distinct word and representing the sentence as an array with said numbers. Frequently, each word is represented by its frequency (how common it is) in the dataset, mainly by being represented by its occurrence ranking (most common word is 1, 2nd is 2 and so on...). Since machines don’t know any linguistic properties by default, this method is the most basic and simplest way of converting data to a more readable state, being the most commonly used method when there is no additional information about words available. However, models trained with data in this state will only be able to reach conclusions about the properties and relationships between words by analyzing their position in relation to each other. 3.3.4.1 Vectorization By vectorizing, data is refined in a way that makes additional information more accessible for the model. For example, with vectors, the relationships between words can be easily represented. Considering the following 5 words as 2-dimensional vectors: • man = (-1,0) • woman = (1,0) • royalty = (0,1) • king = (-1,1) • queen = (1,1) It can be inferred that: > man + royalty = king > woman + royalty = queen > The word queen is more closely related to the word woman than to the word man. 29
CHAPTER 3. SOLUTION PLANNING Figure 13: Vectorization Graph With 2-dimensional vectors, the model has access to additional information. However, due to machines not having any problem handling multiple dimensions in contrast to humans, most models represent data as vectors with dimensions ranging between 100 and 300, which can give magnitudes more information than its raw counterpart. As a result, the training process gets more efficient, since the model has more useful information to extract and to reach correct conclusions. 3.4 Working against bias When training a model with an imbalanced dataset, the most probable outcome is that the model will lean heavily towards the most common labels. In order to counteract this event, several solutions may be applied. These solutions mostly fit into one of these categories: 3.4.1 Undersampling Undersampling is a balancing technique that involves the removal of data related to the most common labels. This results in a more balanced dataset. However, one must be cautious, since the removal of useful data may be harmful to the model’s development. For this study, since the dataset that is being used is so unbalanced an undersampling technique is used where it limits the quantity of entries for each class to a certain point. If that limit is, for example, 30
3.4. WORKING AGAINST BIAS 1000 then every class will have 1000 entries at most. This makes the solution better at generalizing since it doesn’t rely as much on the majored classes to predict the output. 3.4.2 Oversampling In contrast to the previous technique, oversampling battles imbalance by adding data to the less common labels. This can be done by cloning already existing data or by producing new synthetic data that is automatically generated by ”mutating”the rows that belong to the less common classes. 3.4.2.1 Cloning data This method extracts information related to the less frequent classes and duplicates it in order to balance the data. This augmentation technique can be applied to any type of data. However, it is 100% artificial, since no new useful information is being added, which may help the model deal with imbalance but will make the model less capable of generalizing results. By cloning data the model had a more balanced dataset, however since the minority classes are so outnumbered in relation to the majored class the dataset became composed of mostly duplicates which increased the time needed for the training phase without a big improvement for the model itself. 3.4.2.2 Synthetic data A dataset can be balanced by creating small changes to existing information in order to generate ”new”data to aid the model’s training process. This method relies on how the input is represented and may need several tweaks in order to be viable in broader use cases. Synthetic data can be exemplified as follows: • When the input is a sentence, new data can be generated by replacing words with one of its own synonyms. • When the input is formed by numerical values, new information may be generated by randomly tweaking its numbers. This technique reduces the presence of possible duplicates that were present in the previous technique, however, its implementation is more complex and depends on the type of information present in the dataset. 3.4.2.3 Data Collection While it may seem obvious, one of the most effective ways to augment a dataset is by manually adding new annotated data. This approach is considered the best solution because it ensures that all the added data is both organic and correctly annotated. However, it is also the method that demands the most effort and investigation, making it potentially infeasible in some cases. 31
CHAPTER 3. SOLUTION PLANNING 3.4.3 Weight distribution If modifying the dataset’s content is not desired then it is possible to still achieve a more balanced and efficient training by assigning different weights to the existing data. This means that some information may have more influence on how the model behaves than others. So, by assigning a larger weight to data related to the less common classes then a more balanced training phase can be achieved. Weights can be assigned in one of two ways: 3.4.3.1 Class weights In this method, the same weight is applied to data regarding the same class. For example, for a dataset with two classes, and one class is 10 times more common than the other, a weight of 10 could be assigned to the minority class and a weight of 1 to the majority class. This would tell the model that the minority class is more important, and it would help the model to learn to predict the minority class more accurately. 3.4.3.2 Sample weights This method goes one step above and assigns a different weight to each sample of the dataset. This can be accomplished if a more in-depth fine-tuning is necessary, as it allows for the adjustment of the importance of each individual data point within the cluster of information. 3.5 Models 3.5.1 Classification or Regression - What to use? Before choosing the proper models to implement in the study one must research the domain of the problem in order to use the right tools for the job. The first step is to determine if it is a regression or classification problem by identifying the format of the expected output. Classification models are machine learning algorithms used to categorize or assign discrete labels or classes to input data based on patterns and features. They are employed to make predictions about which category or class new data points belong to. Some examples of Classification problems are: • Identifying if a certain message is spam; • Predicting the sentiment behind a piece of information; • Classifying images in order to identify objects or detect anomalies. Regression models are statistical techniques used in machine learning to analyze the relationship between one or more independent variables and a dependent variable, aiming to predict or estimate numerical outcomes. Some examples of Regression problems are: 32
3.5. MODELS • Predicting stock prices of certain companies; • Predicting prices based on supply/demand; • Predicting future climate patterns based on previous behaviour. By comparing the main uses for each category, it is easy to conclude that Classification models are the more appropriate solution for this study. 3.5.2 Logistic Regression Despite having regression in its name, Logistic Regression is actually mainly used for classification problems. This technique models the relationship between the input features and the probability of belonging to the positive class (class 1) using a logistic function. This function also known as the sigmoid function is an S-Shaped curve that maps any value to a value between 0-1. This means that Logistic Regression really shines in binary classification problems, however, it can also be used for multiclass predictions. 3.5.3 Support Vector Machine (SVM) Support Vector Machine is a technique used for classification and regression problems. It maps all available data into an n-dimensional space and finds an hyperplane capable of isolating all data belonging to each class. In a binary classification, it finds a function capable of separating data belonging to the two classes. If there is no proper function capable of isolating the information then all information can be mapped into a higher dimensional to allow for such function to exist. Natively, this technique only supports binary classification, however, it is possible to break down multi-class problems into several binary classifications. Therefore, by merging several SVMs it is possible to use it in a multiclass problem. 3.5.4 Long Short-Term Memory (LSTM) Long Short-Term Memory [42] is a type of Recurrent Neural Network (RNN) architecture that is designed to address limitations in other RNN architecture such has the incapacity to handle long term dependencies between sequential data. This makes LSTM ideal for datasets where data can be influenced by previous entries such as conversations where context matters. In addition to the standard architecture, a variant exists, the Bidirectional Long Short-Term Memory, or BiLSTM for short, is able to handle dependencies in both forward and backward directions. This means that it is also capable of handling datasets where an entry can be influenced by the previous entries and the following ones. 33
CHAPTER 4. SOLUTION IMPLEMENTATION 4.3.2.1 MELD The MELD dataset consists of dialogues originally scripted for a comedy television series. Consequently, the data does not accurately represent real-life conversations, especially in the context of HR-employee communication. However, it’s worth noting that the text is well-composed, as the dataset comprises transcriptions from the TV show, devoid of uncommon abbreviations and internet slang. It’s important to recognize that the utterances in the dataset are attributed to one of the show’s six main characters. This suggests that the sentences are mostly tailored to reflect their personalities, making it challenging to generalize the dataset for training models in the HR context. Despite its imbalance, with the most prevalent emotion being neutral, the dataset provides labels for individual emotions and the polarity of those emotions within each utterance. This labeling facilitates the identification of ambiguous emotions; for instance, the ”surprise”emotion that can be either positive or negative. 4.3.2.2 goEmotions The goEmotions dataset stands out from the other datasets due to its composition of informal language and sentences. Comprising Reddit comments, this dataset has no constraints on the use of slang and uncommon abbreviations. While this informality may not be suitable for accurately representing HRemployee communication, it can offer certain valuable insights. Notably, some utterances within the dataset incorporate emojis, a common feature in chats between HR departments and employees. Emojis can provide a direct means of conveying emotions through text, making them highly relevant for emotion extraction purposes. Additionally, it’s worth mentioning that the dataset is somewhat unbalanced. However, it has the most evenly distributed range of emotions among all the collected datasets. 4.3.2.3 DailyDialog The DailyDialog dataset contains several favorable characteristics. Upon close examination of its content, both advantages and disadvantages in the context of this topic can be identified. Firstly, it’s worth noting that the text within this dataset is well-written, devoid of uncommon abbreviations or slang words. The dialogues are also diverse, encompassing a wide range of topics and conversational styles. However, as mentioned earlier, a notable drawback is the dataset’s considerable imbalance concerning various emotions. The majority of sentences are labeled with the ”neutral”emotion, constituting approximately 83% of the entire dataset. Additionally, the conversations in this dataset are not directly related to HR-employee communication. 40
4.3. DATASET ANALYSIS 4.3.2.4 Chosen dataset After a comprehensive analysis and comparison of various datasets, it becomes evident that the DailyDialog dataset stands out as the most suitable choice in terms of aligning with the problem’s scope and context. Although it is not specifically tailored to workplace conversations, its content represents more common dialogs and so is the one that closely mirrors the kinds of conversations that employees might have the most. This inherent similarity makes it the prime candidate for delivering the best results when integrated into the workplace environment. One of the biggest problems regarding the chosen dataset is how imbalanced it is. Over 83% of the data is related to the neutral emotion. By analyzing the dataset it can be concluded that there are some dialogs that only contain the neutral emotion. Those dialogs can be removed in order to have a more balanced dataset without removing much relevant information. Figure 14: Filtered Daily Dialog bar plot While still heavily biased toward the neutral emotion, this filtration reduced the model’s reliance on the neutral emotion which returned better results. However the model wasn’t still able to generalize well enough to be useful with sentences outside the dataset’s genre, so additional manipulation was required. Further oversampling, undersampling and weight distribution techniques were tested and mixed in order to counteract dataset imbalance. The technique that yielded the best results was a truncation undersampling technique where each emotion was limited to having 2000 sentences at most for each. This resulted in the following distribution: 41
CHAPTER 4. SOLUTION IMPLEMENTATION Figure 15: Final Emotion Distribution It is possible to gain insight into the influence of some words in the emotion labeling process by generating word clouds for each label of the dataset. Words with a high amount of emotional charge are represented in a large part of sentences belonging to a certain emotion. Words such as hate, sick, awful and horrible are prominent in sentences labeled with disgust while words that have a more positive connotation such as great, nice, love, good and beautiful are common in sentences labeled with joy. Figure 16: Daily Dialog Disgust WordCloud 42
4.4. PRE-PROCESSING Figure 17: Daily Dialog Happiness WordCloud Neutral labelled sentences do not contain emotionally charged words, these sentences are the only ones where the word ”Oh”is not prominently present, this adds relevance to this word since it can be a good way to exclude the neutral emotion when labelling a sentence. Figure 18: Daily Dialog Neutral WordCloud 4.4 Pre-processing In order to sanitize all data and optimize the training process for efficiency, several transformations were applied and a pipeline was formed. First, all special characters were isolated from words to allow the tokenizer to recognize the word and special characters as separate tokens. This results in a reduced number of tokens and enhances the importance of special characters, such as punctuation, in the training process. 43
CHAPTER 4. SOLUTION IMPLEMENTATION Next, each word is individually tagged with its grammatical form (adjective, noun, verb, etc.). This tag assists the lemmatization process by providing additional information, which is particularly helpful for lemmatizing ambiguous words, such as ’leaves’. ’Leaves’ can be interpreted as the plural of ’leaf’ or as the conjugated form of the verb ’leave’. Each word is then lemmatized, reducing the number of distinct tokens and highlighting the significance of words with multiple derivations. Subsequently, the results are converted to lowercase, ensuring that identical words are represented by the same token, regardless of their capitalization. The sentence is then tokenized at the word level, resulting in an array where each element represents a word from the sentence. Finally, the results are vectorized. The format of each token depends on the model used. For transformer models like BERT, each token is represented by a numerical value, while for others, each token is vectorized into a 200-dimensional vector using GLoVe embedding. The size of the vectorized result varies based on the dataset used. Since models require consistent input dimensions, the ideal approach is to pad all lists of vectors to match the size of the largest sentence in terms of the number of tokens - this ensures that all data from the dataset is used. However, due to hardware limitations, training models with large arrays of vectors was not possible. As a result, the maximum size of the vectorized results is determined by finding a size that is able to handle whole sentences for 75% of the dataset, with the remainder being truncated. Figure 19: Preprocessing pipeline In this pipeline Stopwords are not removed due to their important influence in the underlying emotion of some sentences. 4.5 Result Analysis 4.5.1 Model Comparison In order to find the best solution several models were tested. These models range from basic algorithms such as Logistic Regression and Support Vector Machines which will act as a baseline to State of the Art 44
4.5. RESULT ANALYSIS transformers such as BERT and XLNet, there will be some Recurrent Neural Networks tested as well. The diagram in figure 20 shows the compared models: Figure 20: Tested models 4.5.1.1 Baseline Models In this section, an analysis, and comparison of a Logistic Regression Model and a Support Vector Machine is done. These two supervised learning algorithms will be included as baselines for comparing all others. Due to their simplicity, their implementation is straightforward. Although their performance may not match that of the rest, they can serve as a means to validate the effectiveness of the other solutions. In order to also test the effectiveness of different embeddings, both algorithms will be trained with tokens represented as single values using a Count Vectorizer and represented as 200-Dimensional vectors using GLoVe embeddings, table 5 illustrates the efficiency of the different setups. Table 5: Baseline Algorithms Comparison Setup Time to train F1_Score Accuracy Loss CountVectorizer + Logistic Regression 1.70s 0.5619 0.6519 1.0537 GLoVe + Logistic Regression 25.02s 0.3537 0.4884 1.3589 CountVectorizer + SVM 19.30s 0.5536 0.6441 1.0745 GLoVe+ SVM 34.56s 0.3400 0.4845 1.3600 From this table, it is evident that the Logistic Regression algorithm outperforms the Support Vector Machine and is faster as well. Additionally, the use of more complex token representations has a negative impact on the model’s performance. Therefore, in this case, the simpler solution is recommended. 45
CHAPTER 4. SOLUTION IMPLEMENTATION 4.5.1.2 Recurrent Neural Networks Recurrent Neural Network architectures have been a primary choice in various machine learning applications, including classification problems. In this section, two specific architectures: Long Short-Term Memory and Gated Recurrent Unit (GRU) are tested. Like the models described in the previous section, these architectures do not incorporate a specialized embedding process. Therefore, their performance using embeddings from GLoVe and simple tokenization is evaluated. Table 6: RNN Comparison F1_Score Accuracy Loss LSTM 0.4508 0.6231 1.2397 LSTM + GLoVe 0.4140 0.5792 1.3085 GRU 0.4563 0.6325 1.1760 GRU + GLoVe 0.3762 0.5139 1.0041 Much like the previous section, the utilization of GLoVe embeddings yields subpar results. While GLoVe embeddings typically have a positive impact on RNN architectures, as observed in the literature review, it’s noteworthy that these models unexpectedly exhibit poorer performance compared to the baseline. One possible explanation may be an improper implementation of these architectures. 4.5.1.3 Transformers Finally, a comparison of several transformers was conducted. These models represent state-of-the-art solutions with great potential in various areas of natural language processing, including text classification. Table 7 lists each tested transformer, along with its corresponding number of parameters, reflecting the complexity of each model. Table 7: Tested Transformers Transformer Parameter Count BERT 109 482 240 DeBERTa 183 831 552 DistilBERT 66 362 880 RoBERTa 124 645 632 XLNet 116 718 336 By analyzing the table it is evident that DistilBERT has by far the fewest parameters, while DeBERTa has the most. The remaining models exhibit relatively similar complexities. All five transformers were trained and tested under identical conditions, using the same parameters and data. This approach ensures that any variations in performance metrics are solely attributable to the transformer itself, without any external influences. 46
4.5. RESULT ANALYSIS The testing environment comprises the processed DailyDialog dataset along with the finalized model architecture featuring 128 units per dense layer, a batch size of 16 samples, and a dropout rate of 50%. The number of epochs and the learning rate are dynamically adjusted, as detailed in the following sections. The plot in Figure 21 illustrates the evolution of each model, displaying the mean of all validation F1-Scores at the end of each epoch. Figure 21: F1-Score Comparison It is evident that performance remains consistently similar across all transformers, indicating that, in this case, the dataset is the bottleneck. Consequently, employing more complex models does not yield superior results. For this reason, opting for more lightweight and faster models proves to be the better choice. Table 8: Accuracy per Label table Disgust Neutral Joy Anger Surprise Sadness Fear BERT 0.4642 0.7400 0.8519 0.6407 0.8354 0.7612 0.4634 DeBERTa 0.4537 0.7467 0.8666 0.6310 0.8385 0.7652 0.2816 DistilBERT 0.5510 0.7313 0.8189 0.6393 0.8265 0.7466 0.4383 RoBERTa 0.4947 0.7301 0.8337 0.6289 0.8271 0.7694 0.4938 XLNet 0.5205 0.7240 0.8519 0.6262 0.8254 0.7380 0.3958 Upon examining the emotion distribution, as shown in Figure 15, and referencing table 8, it becomes evident that the model’s performance is more significantly influenced by the available data rather than the choice of transformer. Disgust and fear emotions notably represent the minority classes within the dataset, and these are the ones that show the lowest accuracy. This highlights the necessity for a more balanced and improved dataset, however, given the rarity of these types of messages and the detrimental 47
CHAPTER 4. SOLUTION IMPLEMENTATION impact of oversampling on the prediction of more common labels, it is not advisable to upsample these classes. The plot in Figure 22 helps illustrate which model should be chosen based on their respective speeds. Figure 22: Training time comparison Upon examining the plot in Figure 22, it becomes apparent that BERT requires the fewest epochs to reach the optimal solution. However, DistilBERT significantly outpaces the others in terms of time per epoch, making it the fastest to converge to the optimal solution. A small overview of the training process for each model is presented in the table 9, which consists of six rows. The first row indicates the total number of epochs performed until the training met the stopping condition. The second row highlights the epoch when the best solution was achieved, while the third row displays the time in seconds taken to reach that solution. The subsequent rows depict the mean of all F1_Scores, accuracy, and loss values for the best solution of each model. Table 9: Transformer Comparison Total Epochs Optimal Epoch Time to reach solution F1_Score Accuracy Loss BERT 87 1336 0.6796 0.7622 1.0533 DeBERTa 12 12 3553 0.6548 0.7665 1.1003 DistilBERT 9 6 607 0.6732 0.7509 1.1257 RoBERTa 12 9 1701 0.6826 0.7560 1.1070 XLNet 14 14 2925 0.6689 0.7528 1.5074 Based on this information, it is apparent that all models possess similar capabilities and can effectively handle the task with acceptable performance. Each model, except for XLNet, outperformed others in at least one metric. Given the overall similarity in performance and the desire for efficiency, DistilBERT was 48
4.6. FINAL ARCHITECTURE selected as the final transformer solution due to its significantly lighter weight and faster processing speed compared to the competition. 4.6 Final Architecture The finalized model is illustrated in Figure 23 comprises the selected transformer model, accompanied by additional layers responsible for processing the transformer’s outputs and converting them into classification predictions. These layers are designed to be compatible with every tested transformer, simplifying the testing and comparison process by allowing for easy swapping of transformers. Figure 23: Model Architecture The initial layer within this architectural design is the transformer model. This fundamental component serves as the backbone of the neural network. It takes two inputs: ’input_ids’ and ’attention_mask.’ The ’input_ids’ input holds the tokens represented as a vector. Depending on the input’s size, it may either be truncated or padded to fit within the vector, whose size is defined during model construction. The ’attention_mask’ input consists of a vector of the same size, with each element being either 0 or 1, indicating whether the model should ignore the token or incorporate it during the training phase. This mechanism prevents padding from affecting the training process. Following the transformer layer, the output is directed to the GlobalAveragePooling1D layer. This layer transforms the data into a 1-dimensional array by computing the average of each pooled value. This step is 49
CHAPTER 5. CONCLUSION AND FUTURE WORK To achieve this, an extensive review of the current state-of-the-art literature was conducted, encompassing the proposal of various potential machine learning architectures. This stage spanned five months. Subsequently, over the following eight months, a thorough analysis of available data, tools, and techniques was carried out to ensure the creation of an optimal solution whose performance is illustrated in Table 11. Finally, a proof-of-concept tool was designed and implemented, resulting in a functional prototype with satisfactory results. This study and prototype successfully met the company’s requirements and are now ready to be included in KonkConsulting’s portfolio of solutions. 5.2 Limitations and Future Work While the finalized solution accomplishes the goal given and can be considered a success, there are several points that can be improved. Firstly, no dataset was provided, as the solution is intended to operate within a specialized work environment, and conversation dynamics can significantly impact its predictions the lack of data relating to HR-employee interactions imposes a considerable limitation on its effectiveness, considering the specific problem scope. To mitigate this challenge, the solution was designed with a strong emphasis on generalization, prioritizing adaptability over specialization to the provided data. This approach aims to create a model capable of achieving acceptable accuracy across diverse conversational contexts and environments. The hardware employed for this study’s development fell significantly short of optimal performance, being noticeably underpowered. This deficiency led to prolonged training durations and a substantial shortage of GPU memory, which had a severe effect on the study’s developmental phase. As a result of these hardware constraints, several limitations had to be imposed, including a preference for faster and more lightweight models over their more performant counterparts. Additionally, a multitude of hyperparameter configurations and architectural choices had to be discarded due to the hardware’s incapacity to efficiently handle training or complete it within a reasonable timeframe. In terms of future work, addressing the most significant bottleneck in the solution involves performing extraction and labeling of a specialized dataset. As a contingency plan, further study of dataset manipulation can be conducted to ensure proper training without discarding so much valuable information. Additionally, the API developed is currently rudimentary and not fully implemented, as it was created solely as a proof of concept. Should this solution be considered for actual use, it will require substantial API development. 5.3 Final Considerations The development of this dissertation has proven to be my most significant project to date, and its conclusion marks a crucial milestone in my life. The fact that this study was a long-term venture, entirely dependent on my own efforts, demanded that I acquire autonomy and delve into an area where I had little 56
5.3. FINAL CONSIDERATIONS prior expertise, particularly in the field of machine learning. This experience also enabled me to construct a complete solution from the ground up, something I had never undertaken on such a large scale before. I take great pride in the work accomplished over the past year, which has resulted in the creation of a functional tool with potential real-world applicability. 57
Bibliography [1] U. B. of Labor Statistics. EMPLOYEE TENURE IN 2022 . 2022. url: https://www.bls.gov/ news.release/pdf/tenure.pdf (visited on 12/05/2022). [2] L. Andre. 112 Employee Turnover Statistics: 2023 Causes, Cost Prevention Data . en. 2021. url: https://financesonline.com/employee-turnover-statistics/ (visited on 02/03/2023). [3] W. Institute. 2022 Retention Report: How Employers Caused the Great Resignation . en. url: https: //info.workinstitute.com/hubfs/2022RetentionReport/2022RetentionReportWorkInstitute.pdf (visited on 02/03/2023). [4] S. Ariella. en. url: https://www.zippia.com/advice/employee-turnover-statistics/ (visited on 02/03/2023). [5] url: https://www.marketsandmarkets.com/Market-Reports/natural-languageprocessing-nlp-825.html (visited on 02/03/2023). [6] K. Linly. An introduction of NLP and how it’s changing the future of HR . 2021. url: https:// www.plugandplaytechcenter. com/resources/introductionnlp-and-howits-changing-future-hr/ (visited on 12/05/2022). [7] I. Future Market Insights. Natural Language Processing (NLP) Market Size, Share Forecast | US$ 45 billion by 2032 . 2022. url: https://finance.yahoo.com/news/natural-languageprocessing-nlp-market-070000969.html (visited on 12/05/2022). [8] J. Franz. How GPT-3 Unlocks Deeper Listening at inVibe . 2021. url: https://www.invibe.co/ blog/how-gpt-3-unlocks-deeper-listening-at-invibe (visited on 12/02/2022). [9] T. Brown et al. “Language Models are Few-Shot Learners”. In: Advances in Neural Information Processing Systems . Vol. 33. Curran Associates, Inc., 2020, pp. 1877–1901. url: https:// papers . nips . cc / paper / 2020 / hash / 1457c0d6bfcb4967418bfb8ac142f64a - Abstract.html. [10] M.-W. Dictionary. Emotion Definition & Meaning . url: https://www.merriam-webster.com/ dictionary/emotion (visited on 12/18/2022). [11] W. James. The Principles of Psychology . Henry Holt and Company, 1890. 59
BIBLIOGRAPHY [12] P. Ekman. “Facial expression and emotion”. eng. In: The American Psychologist 48.4 (1993), pp. 384–392. issn: 0003-066X. doi: 10.1037//0003-066x.48.4.384. [13] R. Lazarus and B. Lazarus. Passion and Reason: Making Sense of Our Emotions . Oxford University Press, 1994. [14] M. Schröder, H. Pirker, and M. Lamolle. “First Suggestions for an Emotion Annotation and Representation Language”. en. In: (). [15] J. Posner, J. A. Russell, and B. S. Peterson. “The circumplex model of affect: an integrative approach to affective neuroscience, cognitive development, and psychopathology”. eng. In: Development and Psychopathology 17.3 (2005), pp. 715–734. issn: 0954-5794. doi: 10.1017/S09545794050 50340. [16] R. Plutchik. “A psychoevolutionary theory of emotions”. en. In: Social Science Information 21.4–5 (1982), pp. 529–553. issn: 0539-0184, 1461-7412. doi: 10.1177/053901882021004003. [17] S. Poria et al. “Emotion Recognition in Conversation: Research Challenges, Datasets, and Recent Advances”. en. In: IEEE Access 7 (2019), pp. 100943–100953. issn: 2169-3536. doi: 10.1109 /ACCESS.2019.2929050. [18] C. Busso et al. “IEMOCAP: interactive emotional dyadic motion capture database”. en. In: Language Resources and Evaluation 42.4 (Dec. 2008), pp. 335–359. issn: 1574-0218. doi: 10.1007/s1 0579-008-9076-6. [19] G. McKeown et al. “The SEMAINE Database: Annotated Multimodal Records of Emotionally Colored Conversations between a Person and a Limited Agent”. In: IEEE Transactions on Affective Computing 3.1 (Jan. 2012), pp. 5–17. issn: 1949-3045. doi: 10.1109/T-AFFC.2011.20. [20] S. Poria et al. “MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations”. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Florence, Italy: Association for Computational Linguistics, July 2019, pp. 527–536. doi: 10.18653 /v1/P19-1050. url: https://aclanthology.org/P19-1050. [21] Y. Li et al. “DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset”. In: Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Taipei, Taiwan: Asian Federation of Natural Language Processing, Nov. 2017, pp. 986–995. url: https://aclanthology.org/I17-1099. [22] S.-Y. Chen et al. “EmotionLines: An Emotion Corpus of Multi-Party Conversations”. In: arXiv:1802.08379 (May 2018). arXiv:1802.08379 [cs]. doi: 10.48550/arXiv.1802.08379. url: http:// arxiv.org/abs/1802.08379. 60
BIBLIOGRAPHY [23] A. Chatterjee et al. “SemEval-2019 Task 3: EmoContext Contextual Emotion Detection in Text”. In: Proceedings of the 13th International Workshop on Semantic Evaluation . Minneapolis, Minnesota, USA: Association for Computational Linguistics, June 2019, pp. 39–48. doi: 10.18653/v1/S1 9-2005. url: https://aclanthology.org/S19-2005. [24] N. Alswaidan and M. E. B. Menai. “A survey of state-of-the-art approaches for emotion recognition in text.” In: Knowledge Information Systems 62.8 (2020), pp. 2937–2937–2987. issn: 0219-1377. doi: 10.1007/s10115-020-01449-0. [25] K. R. Scherer and H. G. Wallbott. ““Evidence for universality and cultural variation of differential emotion response patterning”: Correction”. In: Journal of Personality and Social Psychology 67.1 (1994), pp. 55–55. issn: 1939-1315. doi: 10.1037/0022-3514.67.1.55. [26] C. Strapparava and R. Mihalcea. “SemEval-2007 Task 14: Affective Text”. In: Proceedings of the Fourth International Workshop on Semantic Evaluations (SemEval-2007) . Prague, Czech Republic: Association for Computational Linguistics, June 2007, pp. 70–74. url: https://aclanthology. org/S07-1013. [27] S. Mohammad et al. “SemEval-2018 Task 1: Affect in Tweets”. In: Proceedings of the 12th International Workshop on Semantic Evaluation . New Orleans, Louisiana: Association for Computational Linguistics, June 2018, pp. 1–17. doi: 10.18653/v1/S18-1001. url: https:// aclanthology.org/S18-1001. [28] H. Izadkhah. “Detection of multiple emotions in texts using a new deep convolutional neural network”. en. In: 2022 9th Iranian Joint Congress on Fuzzy and Intelligent Systems (CFIS) . Bam, Iran, Islamic Republic of: IEEE, 2022, pp. 1–6. isbn: 978-1-66547-872-4. doi: 10.1109/CFIS54774 .2022.9756494. url: https://ieeexplore.ieee.org/document/9756494/. [29] A. G. Shahraki and O. R. Zaïane. “Lexical and learning-based emotion mining from text”. In: Proceedings of the International Conference on Computational Linguistics and Intelligent Text Processing . 2017. [30] P. Bojanowski et al. “Enriching Word Vectors with Subword Information”. In: Transactions of the Association for Computational Linguistics 5 (2017), pp. 135–146. doi: 10.1162/tacl_a_000 51. [31] J. Pennington, R. Socher, and C. Manning. “GloVe: Global Vectors for Word Representation”. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Doha, Qatar: Association for Computational Linguistics, Oct. 2014, pp. 1532–1543. doi: 10.311 5/v1/D14-1162. url: https://aclanthology.org/D14-1162. [32] E. Batbaatar, M. Li, and K. H. Ryu. “Semantic-Emotion Neural Network for Emotion Recognition From Text”. In: IEEE Access 7 (2019), pp. 111866–111878. issn: 2169-3536. doi: 10.1109 /ACCESS.2019.2934529. 61
BIBLIOGRAPHY [33] T. Mikolov et al. “Efficient Estimation of Word Representations in Vector Space”. In: 1st International Conference on Learning Representations, ICLR 2013, Scottsdale, Arizona, USA, May 2-4, 2013, Workshop Track Proceedings . Ed. by Y. Bengio and Y. LeCun. 2013. url: http://arxiv.org/ abs/1301.3781. [34] A. Agrawal, A. An, and M. Papagelis. “Learning Emotion-enriched Word Representations”. In: Proceedings of the 27th International Conference on Computational Linguistics . Santa Fe, New Mexico, USA: Association for Computational Linguistics, Aug. 2018, pp. 950–961. url: https:// aclanthology.org/C18-1081. [35] C. Liu et al. “Individual Emotion Recognition Approach Combined Gated Recurrent Unit With Emoticon Distribution Model”. In: IEEE Access 9 (2021), pp. 163542–163553. issn: 2169-3536. doi: 10.1109/ACCESS.2021.3124585. [36] A. F. Adoma, N.-M. Henry, and W. Chen. “Comparative Analyses of Bert, Roberta, Distilbert, and Xlnet for Text-Based Emotion Recognition”. en. In: 2020 17th International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP) . Chengdu, China: IEEE, Dec. 2020, pp. 117–121. isbn: 978-1-66540-503-4. doi: 10.1109/ICCWAMTIP51612.2 020.9317379. url: https://ieeexplore.ieee.org/document/9317379/. [37] J. Devlin et al. “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”. In: arXiv:1810.04805 (May 2019). arXiv:1810.04805 [cs]. doi: 10.48550/arXiv.1810 .04805. url: http://arxiv.org/abs/1810.04805. [38] Y. Liu et al. “RoBERTa: A Robustly Optimized BERT Pretraining Approach”. In: arXiv:1907.11692 (July 2019). arXiv:1907.11692 [cs]. doi: 10.48550/arXiv .1907.11692. url: http:// arxiv.org/abs/1907.11692. [39] V. Sanh et al. “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter”. In: arXiv:1910.01108 (Feb. 2020). arXiv:1910.01108 [cs]. doi: 10.48550/arXiv.1910.01108. url: http://arxiv.org/abs/1910.01108. [40] Z. Yang et al. “XLNet: Generalized Autoregressive Pretraining for Language Understanding”. In: arXiv:1906.08237 (Jan. 2020). arXiv:1906.08237 [cs]. doi: 10.48550/arXiv.1906.08237. url: http://arxiv.org/abs/1906.08237. [41] A. Vaswani et al. “Attention is All you Need”. In: Advances in Neural Information Processing Systems . Vol. 30. Curran Associates, Inc., 2017. url: https : / /papers .nips .cc / paper_ files/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html. [42] S. Hochreiter and J. Schmidhuber. “Long Short-term Memory”. In: Neural computation 9 (Dec. 1997), pp. 1735–80. doi: 10.1162/neco.1997.9.8.1735. 62
BIBLIOGRAPHY [43] X. Ying. “An Overview of Overfitting and its Solutions”. en. In: Journal of Physics: Conference Series 1168 (Feb. 2019), p. 022022. issn: 1742-6588, 1742-6596. doi: 10.1088/1742-6596/116 8/2/022022. [44] L. Li et al. “Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization”. In: arXiv:1603.06560 (June 2018). arXiv:1603.06560 [cs, stat]. doi: 10.48550/arXiv.1603.06 560. url: http://arxiv.org/abs/1603.06560. Thisdocumentwascreatedwiththe(pdf/Xe/Lua)L A T EXprocessorandtheNOVAthesistemplate(v6.7.0)[0]. [0] J.M.Lourenço.TheNOVAthesisL A T EXTemplateUser’sManual.NOVAUniversityLisbon.2021.URL:https://github.com/joaomlourenco/novathesis/raw/master/template.pdf. 63