Journal of Theoretical and Applied Information Technology 31st May 2025. Vol.103. No.10 © Little Lion Scientific ISSN: 1992-8645 www.jatit.org E-ISSN: 1817-3195 4356 ENHANCING ROBUSTNESS IN MEDICAL QUESTION ANSWERING SYSTEMS WITH NOVEL DEFENSE MODELS AGAINST ADVERSARIAL ATTACKS ATRAB A. ABD EL-AZIZ1, REDA A EL-KHORIBI2, AND NOUR ELDEEN KHALIFA2* 1Department Of Information Technology, Faculty of Computers and Information, Kafrelsheikh University, Egypt 2Department Of Information Technology, Faculty of Computers and Artificial Intelligence, Cairo University, Egypt E-mail:
[email protected], 2
[email protected] ABSTRACT Medical Question Answering (MQA) systems play a critical role in supporting accurate medical diagnoses and healthcare decision-making. However, they are increasingly vulnerable to adversarial text attacks. These attacks subtly alter input questions and lead to incorrect outputs. While prior research has extensively explored adversarial defenses for medical images, there remains a significant gap in protection strategies for text-based MQA systems. To the best of our knowledge, this paper is the first to propose and evaluate defense mechanisms specifically designed to secure MQA systems against these attacks. We introduce three novel defense models that address both word-level (synonym substitution, word deletion) and character-level (random character insertion) attacks targeting the BERT model. The Synonym Substitution Embedding (SSE) Defense Framework combines TF-IDF ranking with transformer-based synonym embeddings to resist synonym substitution attacks. CosineDefender leverages cosine similarity to detect and neutralize perturbed inputs, while JaccardDefender applies Jaccard similarity to provide robust protection across multiple attack vectors. To validate our approach, we conduct experiments on two medical datasets (Symptom2Disease and Medical Symptoms Text and Audio Classification) and a natural language dataset (AG’s News) for comparative analysis. Our results show that the SSE model reduces the attack success rate on AG’s News from 8.7% to just 0.4%. On medical datasets, CosineDefender significantly lowers attack success rates to 3.4%, 4.3%, and 12.8%, while JaccardDefender consistently achieves the best performance, reducing all attack success rates to around 3.4% and maintaining high classification accuracy. This work introduces a new line of defense for MQA systems. It establishes a baseline for adversarial robustness in the medical NLP domain. It also contributes the first comprehensive evaluation of targeted defense models in this critical area. Keywords: Adversarial Attacks, BERT, Medical Question Answer (MQA), Term Frequency-Inverse Document Frequency (TFIDF). 1. INTRODUCTION The integration of natural language processing (NLP) systems in healthcare has revolutionized MQA and diagnostic support. Nevertheless, the increasing reliance on these systems has made them susceptible to adversarial attacks. These attacks previously explored in computer vision and now applicable to NLP. Adversarial attacks manipulate input data subtly to deceive deep learning models. Within the context of MQA, such attacks can introduce incorrect diagnoses and raise ethical concerns [1]. Unlike general NLP applications, robust defenses for intelligent MQA systems are crucial due to the potentially fatal impact of medical errors. Yet, defending these systems remains underexplored compared to general NLP and medical imaging. This research addresses this gap by developing novel defense techniques to protect medical decision-making. Textual data must adhere to multiple properties including grammatical, lexical, and semantic constraints. As a result, numerous efficient adversarial image attack techniques such as gradient-based methods can't be easily transferred to text data. This is due to the risk of generating incorrect characters and non-existent terms. Recent investigations have illuminated the susceptibility of medical NLP models to adversarial attacks. Tactics like synonym substitution, character insertions, and deletions can significantly alter model outputs while maintaining linguistic coherence. Consequently, a demand for robust
Journal of Theoretical and Applied Information Technology 31st May 2025. Vol.103. No.10 © Little Lion Scientific ISSN: 1992-8645 www.jatit.org E-ISSN: 1817-3195 4357 defense mechanisms has arisen to ensure the precision and security of the MQA systems [2]. While adversarial attacks are well-studied in general NLP, the unique challenges of MQA such as medical terminology, patient safety, and coherence remain underexplored. This paper highlights the adversarial attacks targeting DL-based MQA systems and examines their potential ramifications on patient care and clinical decision-making. Drawing inspiration from the broader realm of adversarial DL and the distinctive challenges presented by healthcare applications. This paper introduces three new defense frameworks tailored to mitigate the impact of adversarial attacks. The SSE defense model against word substitution attacks is designed with a focus on applications such as news and MQA Bertbased text classification. This model dynamically enhances the robustness of deep learning-powered text classifiers. Our method's effectiveness is demonstrated through experiments utilizing a BERT pre-trained classification model and the widely recognized AG's news dataset. AG's news is a common benchmark for text classification. The results indicate that our approach surpasses existing word substitution adversarial defense methods in terms of both attack success rate and model accuracy. Furthermore, we assess the SSE model's performance using two different MQA datasets. The outcomes of the experiments reveal not only its efficacy on newsrelated data but also its transferability to medical datasets. The other two proposed defense models protect against three types of character and word-level attacks. These models are based on the Cosine and Jaccard similarity techniques to identify the most similar attributes between original and adversarial instances. These models are tested on the same three datasets with three different attack types. The study also explores the limitations of the proposed models and their impact on transformer models. Novel Contributions: 1. Development of Defense Models for MQA Systems. This paper introduces three innovative defense models specifically designed to counter adversarial attacks in MQA systems. These models address both word-level and character-level perturbations. This helps fill a significant gap in existing research on MQA system defenses. 2. Synonym Substitution Embedding (SSE) Defense Framework: The SSE model employs Term Frequency-Inverse Document Frequency (TFIDF) and pre-trained transformers to refine synonym embeddings by effectively defending against word synonym substitution attacks. 3. CosineDefender and JaccardDefender Models: These models utilize cosine and Jaccard similarities, respectively, to enhance robustness against the three mentioned adversarial attacks. While JaccardDefender demonstrates superior performance with the lowest attack success rates across datasets. 4. Comprehensive Evaluation on Diverse Datasets: The proposed models are rigorously evaluated on three diverse datasets including medical and natural language datasets and provide a comparative analysis of their effectiveness. 5. Highlighting a Research Gap: This paper addresses a critical gap in the current literature by focusing on MQA systems that have been overlooked in previous adversarial defense research. To our knowledge, this is the first work to tackle adversarial attacks specifically in the context of MQA systems. This paper is organized as follows: Section 2 presents a comprehensive literature review and summarizes relevant research supporting the proposed approach. Section 3 introduces the proposed models and outlines the methodology in detail. Section 4 describes the experimental setup and evaluates the results, followed by an in-depth discussion of the findings in Section 5. Finally, Section 6 concludes the paper with key insights and potential future work. 2. LITERATURE REVIEWS Adversarial NLP text-based attacks and defenses have emerged as dynamic areas of research in recent times. Within the domain of medical text, numerous tasks are encountering the risk of adversarial attacks. For instance, tasks like machine translation, text classification, medical question and answer (MQA), and others. These models are particularly susceptible to malicious adversaries. The initial focus of this section is on addressing the issue of adversarial attacks and their corresponding defense mechanisms in the context of text classification tasks. Subsequently, an initial overview of attack models is explained, specifically delving into the realm of several commonplace word-level synonym adversarial attack strategies. Figure 1 shows a medical scenario for a text adversarial attack.
Journal of Theoretical and Applied Information Technology 31st May 2025. Vol.103. No.10 © Little Lion Scientific ISSN: 1992-8645 www.jatit.org E-ISSN: 1817-3195 4358 Definition of the Problem: Considering a text classifier denoted as C: X → Y, where X represents the input space and Y signifies the output space, let's assume there exists an input text denoted as x ∈ X. With this setup, the classifier is capable of generating a predicted true label 𝑦through a posterior probability denoted as P [3] as illustrated in Equation (1). 𝑎𝑟𝑔𝑚𝑎𝑥∈𝑃((𝑦|𝑥)= 𝑦 (1) Adversarial attack Definition: an adversary's capability to generate an adversarial example 𝑥 by incorporating a perturbation that is imperceptible to human observation on the classifier C as illustrated in Equation (2). 𝑎𝑟𝑔𝑚𝑎𝑥∈𝑃((𝑦|𝑥)≠ 𝑦 (2) 𝑥= 𝑥 + 𝑥, ‖∆𝑥‖𝑝 <∈ Here, ∈ is a parameter that regulates the magnitude of the small perturbation in a way that ensures the crafted example remains imperceptible to human senses. The notation ‖∆𝑥‖𝑝 represents the pnorm. 2.1 Classification of Text Adversarial Attacks Given that textual data varies from data in image or audio domains, attack types also differ. Depending on the components altered within the text, adversarial attack techniques can be categorized into four distinct types: character-level, word-level, sentence-level, and multi-level attacks. In these attack categories, manipulations typically involve the insertion, removal, swapping/replacement, or flipping of text data. However, it's important to note that not all of these options are necessarily employed across different levels of attacks [4]. Current advanced adversarial attacks on text classification can be classified into the following categories: Character-level Attack: In this attack, modifications are made to individual characters within the text. This can involve altering characters by replacing them with new characters, special symbols, or numbers. The attack techniques include adding new characters to the text, swapping characters with neighboring ones, removing characters from words, or flipping them. These manipulations aim to create subtle changes in the text while attempting to maintain its overall structure and coherence. Gao et al. [5] also introduced DeepWordBug, a technique that introduces minor character perturbations to create adversarial examples by creating typos and grammatical inconsistencies in the sentence against DNN classifiers. Ebrahimi et al. [6] proposed an efficient method named Hotflip, which generates white-box adversarial texts to deceive character-level neural networks. Additionally, text adversarial samples were generated in both white-box and black-box scenarios. However, these approaches are susceptible to defense by incorporating a word recognition model before inputting data into the neural network. Random Character Insertion (RCI) Attack: this technique involves introducing random characters into an input text. The objective here is to disturb the model's comprehension of the text's meaning while keeping the text's overall structure and coherence intact. The attacker initiates by selecting a clean input text for modification. The attacker identifies random positions within the input text and inserts arbitrary characters. These characters can include letters, digits, symbols, or a combination of these elements. The inserted characters are intended to interfere with the semantic coherence and contextual flow of the text. Nonetheless, the attacker takes care to ensure that the inserted characters do not render the text conspicuously altered or nonsensical. the altered text is fed into the NLP model. The anticipation is that the model's interpretation of the text is altered sufficiently to lead to misclassification or inaccurate outcomes. Defenses against this type of attack have included methods such as spell checkers [7,8]. However, these same defenses prove to be particularly susceptible to word-level attacks that maintain language coherence. Against syntactically accurate attacks, Dirichlet Neighborhood Ensemble (DNE) [9], effective strategies encompass adversarial training (AT) [10], Adversarial Sparse Convex Combination (ASCC) [11], and Synonym Encoding Method Figure 1:Example Of General Adversarial Attack
Journal of Theoretical and Applied Information Technology 31st May 2025. Vol.103. No.10 © Little Lion Scientific ISSN: 1992-8645 www.jatit.org E-ISSN: 1817-3195 4359 (SEM) [12]. The first three methods utilize some form of data augmentation by training the model on perturbed samples. Conversely, the last approach introduces an encoding step before the input layer of the target model and trains it to eliminate potential perturbations. Moreover, there are methods for detecting adversarial inputs. In contrast to other defense strategies, these methods possess the capability to explicitly identify manipulated inputs and generate alert signals. there are two available methods: learning to discriminate perturbation (DISP) [13] and Frequency-Guided Word Substitution (FGWS) [14]. The former approach leverages the frequency characteristics of adversarial words and represents the latest and most accurate technique in this regard. Beyond the security concerns highlighted earlier and the methods for text adversarial attacks, it's important to acknowledge that the field of healthcare security encompasses a broader range of issues. Numerous challenges to healthcare security, adversarial attacks, and defense strategies have been explored and put forth as well. The integration of NLP systems in healthcare has brought about transformative advancements, yet the regulatory landscape and industry standards governing their security and robustness remain largely undefined. This lack of specificity poses challenges in ensuring the consistent and reliable performance of these systems, particularly in the face of adversarial attacks. Establishing clear guidelines and benchmarks for security measures, testing protocols, and model validation would be instrumental in fostering trust and mitigating potential risks associated with the deployment of NLP technologies in critical healthcare applications. Traditional static defense mechanisms often rely on fixed rules or predetermined patterns to identify and mitigate attacks, making them vulnerable to adaptive adversaries who can easily circumvent these defenses. In contrast, the SSE defense model employs a dynamic approach by leveraging contextual information and semantic relationships between words to detect and neutralize adversarial perturbations. This adaptability allows the SSE model to effectively counter a wider range of attacks, including those that may not conform to predefined patterns. Word-Level Attack: involve the substitution of words within original texts using synonyms, antonyms, or by simulating typing errors. Alternatively, words might be entirely removed from the text to create variations. This is a strategy employed to maintain semantic coherence while altering the content. Liang et al. [15] present a technique that involves identifying suitable terms for insertion, replacement, and deletion based on the calculation of the most substantial gradient magnitude of the cost function and the word frequency. However, their approach necessitates a notable degree of human involvement in the creation of adversarial instances. In order to sustain semantic consistency and minimize the likelihood of human detection, their method demands manual efforts. Word-level attacks pose a greater challenge to detect due to their ability to preserve the semantic meaning and grammatical correctness of the original text. Synonym substitutions, in particular, can seamlessly replace words while maintaining contextual coherence, making the attack less conspicuous. The vast number of potential synonyms further expands the attack space, making it difficult to anticipate and defend against all possible variations. Additionally, these attacks can be effectively executed in blackbox scenarios, where the attacker lacks knowledge of the target model's internal workings, enhancing their versatility and applicability. For examples, Word Synonym Substitution (WSS) Attack: this involves the replacement of words in a sentence with their synonyms, to retain the overall meaning and context of the text. Initially, a system must identify suitable synonyms for the words present in the input text. This task can be accomplished using resources such as WordNet or pre-trained word embeddings. The selected synonyms should ensure that the intended meaning and syntactic structure of the sentence are preserved. Furthermore, the substituted word should seamlessly integrate into the sentence, maintaining coherence and readability. Adversarial examples generated through word synonym substitution are designed to deceive machine learning models. To assess the efficacy of this technique, one can measure its impact on model performance and evaluate how well the substituted text maintains the original meaning and context. Random Word Deletion (RWD) Attack: is a method that involves the removal of words from an input text in a randomized manner, aiming to disrupt the model's comprehension of the text's meaning while striving to uphold its grammatical structure. The primary aim of this attack is to introduce alterations to the input while maintaining an appearance of innocence. The process initiates with a clean input text that they intend to manipulate. This input could encompass a sentence, paragraph, or an extended piece of text. The attacker randomly eliminates words from the input. This random selection contributes to unpredictability and minimizes the chances of detection. The challenge
Journal of Theoretical and Applied Information Technology 31st May 2025. Vol.103. No.10 © Little Lion Scientific ISSN: 1992-8645 www.jatit.org E-ISSN: 1817-3195 4360 for the attacker lies in perturbing the semantic flow of the text while adhering to grammatical correctness, thereby reducing suspicion. Following the deletion of words, the modified text is inputted into the NLP model. The expectation is that the model's interpretation of the text is distorted sufficiently to lead to misclassification or inaccurate results. The success of the attack is determined by whether the model produces an output that deviates from the desired outcome. As a result, other papers primarily focus on the approach of word substitution to achieve automated generation. The pivotal distinction among these subsequent methodologies lies in their methods of generating alternative words. Samanta et al. [16] proposed the construction of a candidate pool containing synonyms employing the Fast Gradient Sign Method (FGSM), genre-specific keywords, and typographical errors. In contrast, Papernot et al. [17] perturb a word vector by computing its forward derivative and subsequently mapping this perturbed word vector to the nearest word within the word embedding space. The adversarial attack technique "DeepWordBug" introduced in [5] also tested word-level attacks for crafting adversarial instances. It relies on a scoring strategy to identify and modify the most significant words, leading to substantial changes in the classification outcome. This method successfully manipulated classification results to a considerable degree. Building upon the synonym substitution technique, Ren, et al. [18] introduced a greedy algorithm called PWWS, designed specifically for text adversarial attacks. The method is devised to perturb an initial text example into an adversarial one. They introduced a novel word replacement sequence that takes into account both the saliency of words and the classification probability. By ensuring that the altered example retains a semantic similarity to the original, it becomes challenging for humans to detect any anomalies in the modified text. This algorithm focuses on word-level adversarial examples, which tend to be less noticeable to humans and present greater challenges for deep neural networks (DNNs) to counteract or defend against. Yang et al. [19] introduce two distinct techniques: The Greedy Attack, which relies on perturbation, and the Gumbel Attack, which is built upon scalable learning. To reinstate the interpretability of adversarial attacks utilizing the word substitution approach, Sato et al. [20] confine the direction of perturbations to existing words within the input embedding space. In [21], the susceptibility of Deep Learning-based Text Understanding techniques to adversarial text attacks is thoroughly examined. The authors devised a comprehensive attack framework, TextBugger, to create adversarial texts. Within this framework, they adopted distinct approaches for character-level and word-level perturbations. For character-level perturbation, their method involves introducing misspelled versions of significant words, leading to efficient misclassification of models; however, these misspellings are easily detectable. Conversely, for word-level perturbation, they opted to select substitute words from the word embedding space, utilizing the GloVe model [22] as the basis for their choice. Defenses against word-level text adversarial attacks have seen limited research activity. As far as our knowledge extends, the work of [23] stands out as the sole attempt at countering attacks based on synonym substitution. They introduced the Synonym Encoding Method (SEM), which involves encoding synonyms using identical word embeddings to counteract adversarial perturbations. Nevertheless, SEM requires an additional encoding step before regular training and is constrained by fixed synonym substitution. Our framework, in contrast, employs a unified training approach and offers a versatile synonym substitution encoding strategy. The distinction between the random word substitution attack and other adversarial attacks lies in its specific method of manipulating text input. While other attacks might involve character-level changes, sentence insertions, or combinations of different techniques, random word substitution focuses solely on replacing words within the original text with their synonyms. This targeted approach aims to maintain the overall meaning and grammatical structure of the sentence while subtly altering the content to mislead machine learning models. The random nature of the substitution adds an element of unpredictability, making it harder for defense mechanisms to anticipate and counteract the attack. Sentence-Level Attack: adversarial examples are created by inserting entirely new sentences. While other approaches have been less explored, this method focuses on introducing new sentence structures to manipulate the model's predictions. Cao et al. [24] introduced the black-box adversarial attack named Twin Answer Sentences Attack (TASA). This attack designed to alter the context of a question without compromising its fluency or the accuracy of the correct answer. TASA identifies the relevant answer sentence in the context and generates two modified sentences by replacing
Journal of Theoretical and Applied Information Technology 31st May 2025. Vol.103. No.10 © Little Lion Scientific ISSN: 1992-8645 www.jatit.org E-ISSN: 1817-3195 4361 the key terms with synonyms to exploit biases in question-answering models. This lead to produce a misleading answer that directs the model toward an incorrect response by introducing irrelevant entities. In [25], two techniques are proposed for generating adversarial examples targeting Math Word Problem (MWP) solvers. The "Question Reordering" is the first approach involves rearranging the question portion to appear at the beginning of the problem text. The "Sentence Paraphrasing" is the second technique focuses on rephrasing each sentence in the problem while preserving both its semantic meaning and numerical information. Multi-Level Attack: encompasses a combination of character, word, and sentence-level techniques. These attacks leverage various levels of manipulation to create more complex and impactful adversarial examples. Authors in [26] introduced the Visuo-Adaptive DualStrike (VADS) attack which is a novel method that combines transfer-based and query-based strategies to exploit vulnerabilities in Visual Question Answering (VQA) systems. VADS employs a momentum-like ensemble method to identify potential attack targets and compress perturbations. It then uses a query-based strategy to dynamically adjust the perturbation weights for each surrogate model. Evaluation across two datasets demonstrated that VADS surpasses existing adversarial techniques in both efficiency and success rate. In another study, authors in [27] are the first to investigate and successfully execute attacks on a multilingual Question Answering (MLQA) system pre-trained with multilingual BERT using various adversarial strategies. Their approach demonstrates that these attacks can degrade system performance by up to 85%. They reveal that the model exhibits a preference for English and the language of the question, often neglecting other languages present in the QA pair. Additionally, the authors show that incorporating these attack strategies during the training phase can help mitigate the impact of such adversarial attacks. 2.2 Defense Techniques against Adversarial Text Research on text adversarial defense techniques has predominantly explored three primary approaches [28]: Adversarial Training: In these strategies, the model's training process is altered to acquire robust features and heightened resilience against adversarial attacks. During testing, inputs are also adjusted to prevent the introduction of adversarial perturbations. A significant challenge in adversarial training arises from the requirement to be aware of various attack strategies during the training process. The limitation stems from the fact that adversaries typically do not disclose their attack techniques, rendering adversarial training constrained by the user's awareness. If a user attempts to incorporate adversarial training to counteract all known attacks within their knowledge, the resulting model's capability to carry out accurate classification could be severely compromised. This is due to the model acquiring minimal information about the genuine data, which ultimately hampers its classification performance. Adversarial training, while enhancing robustness against attacks, can sometimes negatively impact the model's ability to classify clean, authentic data. This is because the model learns to focus on features that distinguish adversarial examples from clean ones, potentially overlooking subtle nuances important for accurate classification of real-world data. Modifying Networks: This approach involves enhancing the model's architecture by incorporating additional layers, and sub-networks, or modifying loss and activation functions to bolster its defensive capabilities. Network Add-on: External networks are integrated into the system as supplementary components for classifying previously unseen data, thus augmenting the defense mechanisms against adversarial inputs. A defense model is considered successful when the generated adversarial example 𝑥fails to deceive the classifier C∗ or given an input text example x, the attacker is unable to create an adversarial example 𝑥. In this paper, we introduce many defense models against three types of natural and medical text attacks depending on the modifying network's direction which involves enhancing the model's architecture by adding some preprocessing steps in the test phase to reduce the attack impact on the clean model. Problem Statement: As noted from the literature review, there is a critical gap in defending MQA systems used for clinical decision-making. Existing defenses focus mainly on general NLP models or medical image processing. They overlook the unique challenges of MQA such as handling sensitive patient data, ensuring interpretability, and maintaining accuracy in medical contexts. This paper addresses this gap by developing and evaluating three novel defense mechanisms to protect MQA systems from wordand character-level adversarial perturbations. These
Journal of Theoretical and Applied Information Technology 31st May 2025. Vol.103. No.10 © Little Lion Scientific ISSN: 1992-8645 www.jatit.org E-ISSN: 1817-3195 4362 models are the first to enhance MQA's robustness in clinical settings. 3. PROPOSED FRAMEWORKS The proposed models defend against text adversarial attacks by examining changes in the classification outcome once the text modification module is applied. This method operates under the assumption that the intended target of potential attacks is a BERT model which is the cutting-edge model for text recognition. 3.1 First: Synonym Substitution Embedding (SSE) Defense Framework Synonym Replacement is a defense mechanism that entails the identification of words within the input text that can be substituted with synonyms, maintaining the text's overall meaning. Synonyms are words that possess similar meanings but may exhibit distinct linguistic forms. For instance, in the sentence "The weather is nice," the word "nice" could be substituted with "pleasant," all the while preserving the sentence's intended meaning. Figure 2 shows the overall processes of the SSE defense model. This defense mechanism is implemented through a structured seven-steps process, as outlined below: 1. Preprocessing: The three datasets undergo a sequence of preprocessing steps designed to eliminate inconsistent, missing, and redundant values. The text in all datasets is converted to lowercase. Next, tokenization is carried out, involving the segmentation of the provided text into distinct linguistic units known as tokens. Following this, padding is applied by inserting specific tokens, often represented by a PAD token, at the end of sequences to ensure consistent lengths. Lastly, text truncation is applied, which involves removing tokens from sequences that exceed a predetermined maximum length. To ensure uniformity, the maximum sequence length is determined based on the longest text in the dataset. Questions with shorter lengths are padded with zeros to align them appropriately. In the realm of deep neural models, we investigate BERT as a state-of-theart model which is an attention mechanism that captures contextual relationships between words in a sentence. It is designed for text classification, accommodating both word-level and character-level data processing. 2. TF-IDF is an acronym for Term FrequencyInverse Document Frequency, encompassing two interconnected metrics for gauging the relevance of a word within a document. Each word is assigned distinct Term Frequency and Inverse Document Frequency values. The TF*IDF score results from the multiplication of these individual weights. A higher TF*IDF score signifies infrequent occurrence. Specifically, TF corresponds to Term Frequency, reflecting how often a term appears in a document, while IDF stands for Inverse Document Frequency, indicating the significance of the term across the entire collection of documents [29]. 𝑇𝐹 = 𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑡𝑖𝑚𝑒𝑠 𝑡ℎ𝑖𝑠 𝑤𝑜𝑟𝑑 𝑜𝑐𝑢𝑟𝑒𝑠 𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑤𝑜𝑟𝑑𝑠 𝑖𝑛 𝑠𝑒𝑛𝑡𝑒𝑛𝑐𝑒 𝐼𝐷𝐹 = log 𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑆𝑒𝑛𝑡𝑒𝑛𝑐𝑒𝑠 𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑠𝑒𝑛𝑡𝑒𝑛𝑐𝑒𝑠 𝑤ℎ𝑒𝑟𝑒 𝑡ℎ𝑖𝑠 𝑤𝑜𝑟𝑑 𝑜𝑐𝑐𝑢𝑟𝑒𝑑 Figure 2:SSE Proposed Model Structure Against Word Level Attack. 3. Sentence Transformer: the pre-trained MiniLML6-v2 [30] sentence transformer model is employed to extract N-gram key-phrases from real-time datasets. It is designed to encode sentences or text passages into fixeddimensional vectors that capture semantic information. It has a size of 80MB and consists of 384 hidden layers. This model is known for its efficient encoding speed, capable of processing 14,200 sentences per second on a V100 Graphics Processing Unit (GPU).
Journal of Theoretical and Applied Information Technology 31st May 2025. Vol.103. No.10 © Little Lion Scientific ISSN: 1992-8645 www.jatit.org E-ISSN: 1817-3195 4363 4. Fusion of Sentence Transformers and TF-IDF: The combined use of sentence transformers and TF-IDF enhanced the embeddings’ ability to capture semantic meaning and word importance. This integrated approach results in more accurate semantic similarity assessments and improved defense against adversarial attacks. 5. Deep Neural BERT Model: The BERT-BaseUncased model is fine-tuned using the Adam optimizer, and different learning rates are tested over 3 epochs. Categorical cross-entropy is employed as the loss function to minimize during the training process. The parameters of our BERT-Base-Uncased model are evaluated throughout the training phase. The optimal model is identified when the validation loss is minimized, achieved through adjusting hyperparameters. Notably, altering hyperparameters significantly influences the model’s performance. A learning rate of 2e-5 yields superior results in comparison to other learning rates [31]. 6. Adversarial Text Generation: This step involves the creation of adversarial text examples using the word synonym substitution attack. This process modifies the original text to introduce perturbations aimed at challenging the model’s robustness. The objective is to generate adversarial examples that test the resilience of the model against various forms of textual manipulation. Once the adversarial text is generated, it undergoes transformation into numerical vectors using the Term FrequencyInverse Document Frequency (TF-IDF) method. TF-IDF serves to identify the most salient words within the text by quantifying their importance based on their frequency of occurrence. 7. Similarity Calculation Using Sentence Transformer: In the final step, a sentence transformer model is employed to compute semantic similarity between words using Euclidean distance. The pre-trained MiniLML6-v2 model is used to encode text into fixeddimensional vectors that capture the semantic meaning of sentences. The model’s efficiency in processing, with a capacity to handle 14,200 sentences per second on a V100 GPU, allows for rapid encoding and comparison. The MiniLML6-v2 model with 384 hidden layers and a size of 80MB generates dense vector embeddings that encapsulate the semantic content of the sentences. This approach enhances the model’s ability to detect and counteract word substitution attacks by evaluating the contextual similarity of words. The fusion of Sentence Transformers and TF-IDF results in enhanced embeddings and improved performance when assessing semantic similarity. Transformers are employed to create dense vector representations, or embeddings, for sentences. These embeddings effectively encapsulate the semantic meaning of sentences within a continuous vector space. This integrated approach yields enhanced embeddings that not only capture the semantic intricacies of sentences but also the importance of individual words within those sentences. 3.2 Second: CosineDefender and JaccardDefinder Defense Frameworks The objectives of both CosineDefender and JaccardDefinder defense techniques are to introduce controlled alterations to input text, creating difficulties for adversarial attacks to influence the model's predictions. Nevertheless, maintaining a careful equilibrium in the amount of noise added is crucial to prevent adverse effects on the model's accuracy with clean data. Furthermore, the efficiency of these defense models can differ depending on the particular NLP architecture, the characteristics of the attacks, and the extent of robustness testing they undergo. These two models protect against three distinct types of attacks. Among these, two are at the word level which involving attacks such as word substitution and random word deletion. The third attack is conducted at the character level and is known as Noise Injection. As shown in Figure 3, This defense model shares similarities with the first model in terms of its clean model training structure. However, there are two notable distinctions: firstly, this framework is designed to defend against three distinct types of attacks. Secondly, the defense techniques employed have been altered, as two additional methods cosine similarity and Jaccard similarity are tested as defense mechanisms against these three attack types. The CosineDefender and JaccardDefender frameworks share the same initial preprocessing and BERT model steps as the SSE Defense model. However, instead of using TF-IDF and Sentence Transformers, these frameworks utilize Cosine Similarity and Jaccard Similarity to detect adversarial perturbations and defend against adversarial attacks. Below is a detailed description of each step involved: As with the SSE Defense model, the preprocessing and BERT stages remain consistent: 1. Preprocessing includes converting text to lowercase, tokenizing it, applying padding, and
Journal of Theoretical and Applied Information Technology 31st May 2025. Vol.103. No.10 © Little Lion Scientific ISSN: 1992-8645 www.jatit.org E-ISSN: 1817-3195 4364 truncating sequences. These steps standardize the input for model processing. 2. BERT Model: The tokenized text is passed through the BERT model which generates contextual embeddings for each token in the input sequence. BERT’s bidirectional attention mechanism helps capture semantic relationships between words and outputs dense vector representations of the text. 3. Cosine similarity: In this framework, we replace TF-IDF and Sentence Transformers with Cosine Similarity to measure the semantic similarity between the original input and its adversarially perturbed version. Cosine Similarity is a prevalent measurement used in the field of NLP to gauge the similarity between two vectors. Its applicability extends to the realm of adversarial defense, particularly in assessing semantic similarity. This metric operates by computing the cosine of the angle formed between vectors, which can represent embeddings of words, phrases, or even entire documents within NLP. When the cosine similarity between two vectors is high, it signifies that these vectors align closely in direction, implying shared semantic meanings. The importance of cosine similarity stems from its capability to reveal the semantic connections present in textual elements. Particularly in the context of adversarial attacks, where alterations are frequently inconspicuous yet uphold semantic consistency, cosine similarity assumes a pivotal function. Through the computation of cosine similarity between the initial input and its perturbed version, it becomes viable to measure the extent of semantic diversion between these instances. A significant reduction in cosine similarity might signal the potential occurrence of an adversarial attack [29]. 𝑐𝑜𝑠𝑖𝑛𝑒_𝑠𝑖𝑚𝑖𝑙𝑎𝑟𝑖𝑡𝑦(𝐴, 𝐵)=. ‖‖∗‖‖ (3) Where A ⋅ B represents the dot product of vectors A and B. ||A|| and ||B|| represent the magnitudes (Euclidean norms) of vectors A and B. The values of cosine similarity fall within the range of -1 to 1. A score of 1 indicates perfect similarity, 0 indicates no similarity, and -1 signifies perfect dissimilarity. In summary, cosine similarity proves to be a versatile tool, proficient in quantifying semantic similarity and adept at detecting semantic alterations induced by adversarial perturbations. 4. Jaccard similarity: The JaccardDefender defense model replaces the TF-IDF and Sentence Transformers with Jaccard Similarity. Jaccard Similarity is another widely used metric in NLP, for quantifying the similarity between sets. Jaccard similarity can play a role in defending against adversarial attacks. Like cosine similarity, Jaccard similarity can be harnessed to detect changes in the semantic distance between the original input and its perturbed counterpart caused by adversarial attacks. A significant drop in Jaccard similarity between the perturbed gradient and the original gradient could indicate the presence of adversarial modifications. By setting a Jaccard similarity threshold, it becomes feasible to identify highly similar text between adversarial input and the testing data. Text inputs exceeding the threshold might be potentially recoverable and subjected to additional analysis. The Jaccard similarity coefficient is defined as the size of the intersection of the sets divided by the size of the union of the sets [33]. Mathematically, it can be expressed as: Figure 3:The Proposed Jaccarddefender And Cosinedefender Models Against Word Deletion, Substituation And Character Insertion Attacks
Journal of Theoretical and Applied Information Technology 31st May 2025. Vol.103. No.10 © Little Lion Scientific ISSN: 1992-8645 www.jatit.org E-ISSN: 1817-3195 4371 [4] Han, X., Zhang, Y., Wang, W. & Wang, B, "Text adversarial attacks and defenses: Issues, taxonomy, and perspectives " . Secur. Commun. Networks 2022, vol.6458488 ,2022. [5] Gao, J., Lanchantin, J., Soffa, M. L. & Qi, Y," Black-box generation of adversarial text sequences to evade deep learning classifiers." In 2018 IEEE Security and Privacy Workshops (SPW),IEEE, 2018 ,pp.50–56 . [6] Ebrahimi, J., Rao, A., Lowd, D. & Dou, D. Hotflip, "White-box adversarial examples for text classification." arXiv preprint arXiv:1712.06751,2017, pp.16-18 [7] Pruthi, D., Dhingra, B. & Lipton, Z. C. "Combating adversarial misspellings with robust word recognition." arXiv preprint arXiv:1905.11268 ,2019. [8] Huang, P.-S. et al. "Achieving verified robustness to symbol substitutions via interval bound propagation." arXiv preprint arXiv:1909.014,2019. [9] Zhou, Y., Zheng, X., Hsieh, C.-J., Chang, K.-w. & Huang, X. "Defense against adversarial attacks in nlp via Dirichlet neighborhood ensemble." arXiv preprint arXiv:2006.11627,2020. [10] Goodfellow, I. J., Shlens, J. & Szegedy, C. "Explaining and harnessing adversarial examples." arXiv preprint arXiv:1412.6572,2014. [11] Dong, X., Luu, A. T., Ji, R. & Liu, H. "Towards robustness against natural language word substitutions." arXiv preprint arXiv:2107.13541 ),2021. [12] Chen, Y. et al. "Why should adversarial perturbations be imperceptible? rethink the research paradigm in adversarial nlp." arXiv preprint arXiv:2210.10683 ) ,2022. [13] Zhou, Y., Jiang, J.-Y., Chang, K.-W. & Wang, W. "Learning to discriminate perturbations for blocking adversarial attacks in text classification." arXiv preprint arXiv:1909.03084,2019. [14] Mozes, M., Stenetorp, P., Kleinberg, B. & Griffin, L. D. "Frequency-guided word substitutions for detecting textual adversarial examples." arXiv preprint arXiv:2004.05887 ),2020. [15] Liang, B. et al. "Deep text classification can be fooled." arXiv preprint arXiv:1704.08006 ),2017. [16] Samanta, S. & Mehta, S. "Towards crafting text adversarial samples." arXiv preprint arXiv:1707.02812,2017. [17] Papernot, N., McDaniel, P., Swami, A. & Harang, R. "Crafting adversarial input sequences for recurrent neural networks." In MILCOM 2016-2016 IEEE Military Communications Conference, IEEE,2016, pp.49–54 . [18] Ren, S., Deng, Y., He, K. & Che, W. "Generating natural language adversarial examples through probability weighted word saliency." In Proceedings of the 57th annual meeting of the association for computational linguistics,2019, pp.1085–1097 ,2019. [19] Yang, P., Chen, J., Hsieh, C.-J., Wang, J.-L. & Jordan, M. I. "Greedy attack and gumbel attack: Generating adversarial examples for discrete data." J. Mach. Learn. Res, vol.21,2020, pp.1-36. [20] Sato, M., Suzuki, J., Shindo, H. & Matsumoto, Y. "Interpretable adversarial perturbation in input embedding space for text.arXiv preprint arXiv:1805.02917 ,2018. [21] Li, J., Ji, S., Du, T., Li, B. & Wang, T. "Textbugger: Generating adversarial text against real-world applications." arXiv preprint arXiv:1812.05271 ,2018. [22] Pennington, J., Socher, R. & Manning, C. D. "Glove: Global vectors for word representation." In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 2014, pp.1532–1543 . [23] Wang, X., Hao, J., Yang, Y. & He, K. "Natural language adversarial defense through synonym encoding." In Uncertainty in Artificial Intelligence, PMLR ,2021, pp.823–833 . [24] Cao, Y. et al. Tasa, "Deceiving question answering models by twin answer sentences attack." arXiv preprint arXiv:2210.15221, 2022. [25] Kumar, V., Maheshwary, R. & Pudi, V. "Adversarial examples for evaluating math word problem solvers." arXiv preprint arXiv:2109.05925 ,2021. [26] Zhang, B., Li, J., Shi, Y., Han, Y. & Hu, Q. Vads,"Visuo-adaptive dualstrike attack on visual question answer." Comput. Vis.Image Underst. Vol.104137 ,2024. [27] Karra, R. & Lasfar, A. "Analysis of qa system behavior against context and question changes." Int. Arab. J. Inf. Technol. Vol.21,2024, pp.191– 200. [28] Wang, W., Wang, R., Wang, L., Wang, Z. & Ye, A. "Towards a robust deep neural network against adversarial texts: A survey." ieee
Journal of Theoretical and Applied Information Technology 31st May 2025. Vol.103. No.10 © Little Lion Scientific ISSN: 1992-8645 www.jatit.org E-ISSN: 1817-3195 4372 transactions on knowledge data engineering, no.35, 2021, pp.3159–3179 ) . [29] Kumar, V. & Subba, B. "A tfidfvectorizer and svm based sentiment analysis framework for text data corpus." In 2020 national conference on communications (NCC), IEEE, 2020 pp.1–6 . [30] Wilianto, D. & Girsang, A. S. "Automatic short answer grading on high school’s e-learning using semantic similarity methods." TEM J. 12 )2023 .( [31] Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. Bert. "Pre-training of deep bidirectional transformers for language understanding." arXiv preprint arXiv:1810.04805 )2018 .( [32] Lahitani, A. R., Permanasari, A. E. & Setiawan, N. A. "Cosine similarity to determine similarity measure: Study case in online essay assessment". In 2016 4th International conference on cyber and IT service management,IEE,2016, pp.1–6 . [33] Soto, J. E., Hernández, C. & Figueroa, M. Jaccfpga, "A hardware accelerator for jaccard similarity estimation using fpgas in the cloud." Futur. Gener. Comput. Syst. 138, 2023,pp.26– 42 . [34] AGâ˘Zsnews,https://www.kaggle.com/datasets/ amananandrai/ag-news-classification-dataset, available online(27Aug2023). [Accessed 22-072024] . [35] Symptom2Disease,https://www.kaggle.com/dat asets/niyarrbarman/symptom2disease,availableo nline(27Aug2023). [Accessed 22-07-2024 . [36] MedicalSymptomsTextandAudioClassification, https://www.kaggle.com/code/paultimothymoon ey/ medical-symptoms-text-and-audioclassification,availableonline(27Aug2023). [Accessed 22-07-2024] . [37] Cresswell, W. & Quinn, J. L, "Attack frequency, attack success and choice of prey group size for two predators with contrasting hunting strategies." Animal Behav. Vol.80, 2010, pp.643–648.