Are LLMs Enough for Hyperpartisan, Fake, Polarized and Harmful Content Detection? Evaluating In-Context Learning vs. Fine-Tuning
Full text
Are LLMs Enough for Hyperpartisan, Fake, Polarized and Harmful Content Detection? Evaluating In-Context Learning vs. Fine-Tuning Michele Joshua Maggini1, Dhia Merzougui2, Rabiraj Bandyopadhyay3, Ga¨ el Dias2, Fabrice Maurel2, Pablo Gamallo1 1Centro Singular de Investigaci´ on en Tecnolox´ ıas Intelixentes da USC, Universidade de Santiago de Compostela 2UNICAEN, ENSICAEN, CNRS, GREYC, Normandie Univ 3GESIS Leibniz Institute for the Social Sciences [email protected] Abstract The spread of fake news, polarizing, politically biased, and harmful content on online platforms has been a serious concern. With large language models becoming a promising approach, however, no study has properly benchmarked their performance across different models, usage methods, and languages. This study presents a comprehensive overview of different Large Language Models adaptation paradigms for the detection of hyperpartisan and fake news, harmful tweets, and political bias. Our experiments spanned 10 datasets and 5 different languages (English, Spanish, Portuguese, Arabic and Bulgarian), covering both binary and multiclass classification scenarios. We tested different strategies ranging from parameter efficient Fine-Tuning of language models to a variety of different In-Context Learning strategies and prompts. These included zero-shot prompts, codebooks, few-shot (with both randomly-selected and diversely-selected examples using Determinantal Point Processes), and Chain-of-Thought. We discovered that In-Context Learning often underperforms when compared to Fine-Tuning a model. This main finding highlights the importance of Fine-Tuning even smaller models on task-specific settings even when compared to the largest models evaluated in an In-Context Learning setup - in our case LlaMA3.1-8b-Instruct, Mistral-Nemo-Instruct-2407 and Qwen2.5-7B-Instruct. Code and Dataset — https://github.com/HikariLight/ hyperpartisanship classification/tree/main Introduction Politically biased (PB), hyperpartisan (HP) and fake news (FN) as well as harmful (HF) social media content when covering divisive topics (e.g. politics, COVID-19) present a significant challenge to public discourse and democratic integrity, and most of those phenomena of our interest can fall under the misinformation category (Wardle and Derakhshan 2017). FN refers to fabricated stories that mimic legitimate news formats (Lazer et al. 2018). HP, on the other hand, involves misleading coverage of real events presented with a strong partisan bias (Potthast et al. 2018; Maggini et al. 2025). Both often contain politically charged messages that distort facts and polarize audiences. PB reporting further complicates the media landscape by subtly influencing and Copyright © 2026, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved. polarizing public opinion and eroding the impartiality expected of journalism (Zhou and Zafarani 2020). While the concept of harmful content (HF) is broad and its content can vary based on political priorities (Eyuboglu et al. 2023), its diffusion on social media can rely on the spread of hate speech (Yang et al. 2023a) or misinformation (Eyuboglu et al. 2023). Indeed, harmful forms like polarizing content are frequently fueled by false or misleading narratives to incite frustration (Cinelli et al. 2021; Osmundsen et al. 2021). Consequently, the detection of such content necessitates robust fact-checking and verification (Nakov et al. 2022). This conceptual overlap is reflected in recent research challenges. For example, Azizov and Nakov (2023) considered harmful tweet detection a subtask of identifying relevant claims in tweets containing COVID-19 information. By defining a tweet as harmful when it potentially contained misinformation on COVID-19, they effectively demonstrated how, in a specific and critical context, the tasks of detecting harmful content and misinformation are intrinsically linked. Accurate detection of these diverse forms of problematic content is crucial. Large Language Models (LLMs) are valuable tools for this task. The dominant approaches have been fine-tuning (FT) both encoder-only models (Howard and Ruder 2018) and decoder-only LLMs (Aman 2024). However, a systematic comparison of these models’ performance on the specific tasks of fake and hyperpartisan news, politically biased news, and harmful content detection, especially across multiple languages, is still missing from the literature. Our work fills this gap with a comprehensive comparison of encoder-only and decoder-only LLMs using both finetuning (FT) and various in-context learning (ICL) settings. Our ICL methods include: zero-shot prompts with different degrees of task specifications and with rule-based approaches (e.g., codebooks); Few-shot (FS) using both randomly selected examples, and diversity-optimized examples selected through Determinantal Point Process (DPP); Chain-of-Thought (CoT). We conducted experiments on 10 datasets in five different languages, moving beyond the common limitation of using only English or U.S.-centric data (Maggini et al. 2025). This study addresses three key research questions: •RQ1: How do different LLM adaptation paradigms (FT vs. ICL) compare for these tasks, considering model architecture, size, and pre-training data? arXiv:2509.07768v1 [cs.CL] 9 Sep 2025
•RQ2: What is the impact of various ICL strategies—including few-shot examples and rule-based methods—on performance and stability? •RQ3: How do performance and optimal strategies for these tasks vary across different languages, especially for midand low-resource languages? Our extensive experiments reveal that FT remains a highly effective technique, often outperforming ICL strategies. Specifically, fine-tuned decoders performed better for PB and FN detection, while encoders were more effective for HP and HF tweets. Within ICL, the codebook approach was generally the effective, outperforming CoT for classification. The use of DPP for few-shot example selection sometimes reduced classification variance, though it did not consistently boost performance. Related Work Fine-tuning and Political Text Classification Text classification is an NLP task that assigns a label to a given text. FT adapts a pre-trained model to a specific task by further training the model together with a newly added classification head using task-specific labeled data. (Howard and Ruder 2018). While effective, this process can be sensitive to the classifier head’s initialization (Yang et al. 2022). In political text classification, FT has been successfully applied to various tasks. For instance, Liu et al. (2022) finetuned RoBERTa to create POLITICS, achieving state-ofthe-art performance on SemEval 2019 for hyperpartisan news detection. Other works have explored stylistic features for hyperpartisan content discrimination (Potthast et al. 2018), created new datasets for multiclass hyperpartisan detection using fine-tuned BERT models (Lyu et al. 2023), or combined BERT with ELMo to enhance FT (Naredla and Adedoyin 2022). More recently, LLMs like Llama 2 have been fine-tuned for tasks such as FN detection, leveraging their understanding and analytical capabilities (Aman 2024; Pavlyshenko 2023). In-Context Learning With the advent of recent decoderonly LLMs, ICL has emerged as a valuable technique in NLP. Users interact with models directly through prompts, specific textual templates containing instructions and optionally examples (Brown et al. 2020). This approach allows models to perform tasks without prior task-specific fine-tuning (Efrat and Levy 2020), leveraging a single pretrained model for various downstream tasks and enabling desired behavior specification via natural language. ICL has shown remarkable performance on challenging reasoning tasks (Brown et al. 2020; Wei et al. 2022). However, ICL is highly sensitive to input format and order (Lu et al. 2022; Min et al. 2022), and can lead to irreproducible outcomes as slight prompt changes significantly impact performance (Lu, Schuff, and Gurevych 2024; Sun, Shaib, and Wallace 2023). To overcome these limitations, our work comprehensively tested various prompt strategies and demonstration selection methods, including DPP for more stable performance, and introduced codebook prompting. Research on prompt design aims to elicit better performance and reasoning. Notable approaches include CoT prompting (Kojima et al. 2022), which encourages step-by-step reasoning, and its zero-shot variant (Wei et al. 2022). Lu et al. (2022) also highlighted the importance of careful prompt format and example selection in few-shot learning. Regarding the comparison of ICL and fine-tuning, Labrak, Rouvier, and Dufour (2024) and Edwards and Camacho-Collados (2024) demonstrated that fine-tuning smaller models often outperforms ICL in larger language models across various NLP and text classification tasks. Aligned with these findings, our study expands this comparison by evaluating a broader range of prompt strategies, models (including ModernBERT (Warner et al. 2024) beyond Llama models), and by focusing on specialized domains to rigorously compare the efficacy of these tools. Codebook A codebook provides definitions of categories, including examples and classifying instructions. For instance, Vincent and Mestre (2018) developed a codebook to classify hyperpartisan news on a 5-point scale, and Hughes et al. (2021) crafted one for content analysis of COVID-19 articles, classifying tropes and rhetorical strategies to detect misinformation. Codebooks offer a method to explicitly prompt LLMs with structured contexts, eliciting their rule-based reasoning capabilities, which goes beyond standard ICL that often relies primarily on examples. Hu et al. (2024) explored codebook application in zero-shot settings with closed models for political phenomena classification, using the codebook as a structured framework for interpretation. Similarly, Halterman and Keith (2025) utilized codebooks to evaluate open models, demonstrating how these guidelines facilitate assessing an LLM’s adherence to predefined classification logic. While these specific codebooks were tailored to different political phenomena, NLP tasks, or modeled tasks differently from our dataset intentions, thus not directly applicable to our study, they still served as valuable inspiration for our experimental setup and underscore the potential of integrating structured rule sets into LLM prompts. This enhances adherence to specific classification schemes, aligning with broader prompt engineering efforts to elicit precise and controlled LLM outputs, especially in domains requiring nuanced categorization criteria. Misinformation and Bias Detection LLMs demonstrate reasoning capabilities across various applications, including misinformation detection (Li et al. 2023; Leite et al. 2025). However, LLMs pose a dual challenge: they can be misused to spread misinformation, and their detection capabilities may diminish with implicit or newly crafted content (Chen and Shu 2024). Consequently, the efficacy of detection methods relative to the rapid updating rate of misinformation is a significant concern (Jiang et al. 2024). In misinformation studies, several works have explored LLM capabilities. Jose and Greenstadt (2024) compared proprietary models (GPT, Claude) for zero-shot propaganda detection, finding their performance inferior to RoBERTa-CRF. For hyperpartisan detection, Maggini and Gamallo Otero (2024) showed that increased prompt complexity and external knowledge usually improved Llama-3.1-8b-Instruct’s performance. Conversely, Omidi Shayegan et al. (2024) found encoder models like RoBERTa generally outperformed generative LLMs (GPT-3.5) for Persian hyperpartisan content. In Fake News
Detection, Anirudh, Srikanth, and Shahina (2023) observed gpt-3.5-turbo’s superiority over a bi-directional transformer for Tamil classification. Notably, while these works explore various aspects, none of them have focused on a comprehensive benchmarking across different ICL strategies, models, and multilingual contexts, which is a key contribution of our study. Experimental Setting Task Formulation The core objective of this study is to evaluate the effectiveness of various NLP models in HP, FN, PB and HF detection, by comparing two widely used approaches: FT and ICL, where models are tested as off-the-shelf tools without additional training. Specifically, we focus on tasks such as identifying hyperpartisan and fake news, harmful tweets and a news’ political leaning, recognizing the distinct linguistic and contextual challenges each task presents. Our approach encompasses both binary and multi-class classification scenarios, leveraging a variety of datasets and multilingual contexts. While the ultimate goal is not to develop productionready models, we prioritize thorough experimentation with various transformer-based architectures, prompt strategies, and learning techniques. This exploration serves to highlight the strengths and limitations of the tested architectures, contributing to the broader effort of refining misinformation detection methodologies within the NLP community. We acknowledge the fact that hyperpartisan news shows peculiar stylistic traits, rather than fake news (Potthast et al. 2018). Datasets For our experiments, we selected datasets for both binary and multiclass classification tasks. The 10 datasets focus on articles, headlines, tweets on COVID-19 and political news and different topics including TV, politics, sports, and health. They cover two types of domains: news and Twitter, and include four specific-oriented classification tasks: hyperpartisan and fake news, harmful tweet and political bias detection. For hyperpartisan detection, we selected the SemEval-2019 Task 4 by-article dataset (Kiesel et al. 2019), which contains articles from hyperpartisan and mainstream websites annotated by three annotators. They mostly cover the first term of Trump, Gun Control and other U.S.-centric related topics. The dataset’s strength lies in its article-level annotations that allow for analysis of extended argumentative structures and narrative techniques typical of hyperpartisan content, rather than just isolated claims. The VIStA-H dataset (Lyu et al. 2023) includes hyperpartisan and neutral news headlines from right, center and left U.S. newspapers. Its focus on headlines rather than full articles complements the SemEval dataset. The dataset’s temporal coverage, from 2014 to 2023, is particularly valuable for tracking how hyperpartisan news evolved through multiple election cycles. The Fake News Net dataset (Shu et al. 2017) contains news articles shared on Twitter, allowing for analysis of how hyperpartisan content circulates in social media environments. The Spanish Fake News Corpus (G´ omez-Adorno et al. 2021) gathers news from Spanish newspapers and media company websites, along with factchecking websites. This corpus was selected to broaden the linguistic and cultural scope of the research beyond Englishlanguage and U.S.-centric media. The incorporation of factchecking websites provides an additional layer of verification that strengthens the dataset’s reliability. The Fake.br Corpus (Monteiro et al. 2018) focuses on Brazilian Portuguese manually collected and checked news. The inclusion of this Brazilian Portuguese corpus further expands the cross-cultural and multilingual dimensions of the research. Brazil’s distinct political landscape and media environment provide an important comparative case for testing the generalizability of the detection approaches. The manual verification process strengthens the dataset’s reliability as a benchmark for testing detection methods in a language where NLP resources might be less abundant than for English or Spanish. CLEF 2022 CheckThat! Lab Subtask 1C (Nakov et al. 2022) was selected for its focus on COVID-19 misinformation tweets across multiple languages (Arabic, Bulgarian, and English), providing an opportunity to study fake news content around a global crisis that transcended national boundaries. Regarding political bias detection, we considered the Qbias dataset (Haak and Schaer 2023), which contains articles from AllSides, a news aggregator with an established methodology for evaluating political leaning. This provides a more nuanced and professionally curated ground truth for political bias than many other available resources. This nuanced approach helps move beyond binary classifications of political content and supports more sophisticated analysis of bias indicators. Lastly, CLEF 2023 CheckThat! Lab Task 3A dataset (Azizov and Nakov 2023) provides contemporary examples that reflect the current state of political communication and media bias. This diverse collection of datasets provides a comprehensive foundation for developing and evaluating models across multiple languages, cultural contexts, media formats and misinformation tasks. The URLs to retrieve the datasets can be found in the Appendix. To produce input for the classifier in the Spanish Fake News Corpus and Qbias datasets, we concatenated the headline and the body of the article. The datasets are summarized in Table 1. Models For our experiment, we compared two types of model architectures: encoder-only and decoder-only models. Encoder-Only Masked Language Models: Following Edwards and Camacho-Collados (2024), we selected BERTderived models such as RoBERTa-base (125 million parameters) and RoBERTa-large (354 million parameters) (Liu 2019), XLM-RoBERTa (Conneau et al. 2019), POLITICS (Liu et al. 2022), which has been adapted for the political domain using continuous pre-training, and mDeBERTaV3 (He, Gao, and Chen 2021). RoBERTa is pre-trained on English data, while XLM-RoBERTa is trained on 100 different languages, making it suitable for evaluating the impact of multilingual training on performance. Decoder-Only Large Language Models: We experiment using different small-size open-weight LLMs: LlaMA3.18B and LlaMA3.1-8B-Instruct, Mistral-Nemo-Instruct-2407
Dataset Abbr. Lang. Timeframe Train Size Test Size Avg. Tkn Train Len. Avg. Tkn Test Len. Domain Type Task Label Ratios (Train / Test) VIStA-H (Lyu et al. 2023) HV en 2014–2023 1999 201 13 13 News S HP HP: 0.39/0.63; N: 0.50/0.50 SemEval-2019 by-article (Kiesel et al. 2019) SH en 2007–N/A 645 628 735 757 News D HP HP: 0.50/0.50; N: 0.50/0.50 Spanish Fake News Corpus (G´ omezAdorno et al. 2021) SFN es 2020–2021 676 572 607 843 News D FN T: 0.50/0.50; F: 0.50/0.50 Fake News Net (Shu et al. 2017) FNN en N/A 18556 4640 17 16 News D FN T: 0.74/0.74; F: 0.26/0.26 Fake.br Corpus (Monteiro et al. 2018) FBC pt 2016–2018 5760 1440 688 698 News D FN T: 0.50/0.50; F: 0.50/0.50 CLEF 2022 1C (Nakov et al. 2022) C1A ar N/A 3624 1201 73 68 Twitter S HT NH: 0.81/0.84; H: 0.19/0.16 C1B bu N/A 708 325 62 68 Twitter S HT NH: 0.87/0.97; H: 0.13/0.03 C1E en N/A 3323 251 60 51 Twitter S HT NH: 0.91/0.84; H: 0.09/0.16 CLEF 2023 3A (Azizov and Nakov 2023) C3A en N/A 45066 5198 90 110 News D PB R: 0.39/0.13; C: 0.34/0.38; L: 0.27/0.50 Qbias (Haak and Schaer 2023) QB en 2012–2022 17403 4351 97 96 News D PB R: 0.39/0.13; C: 0.34/0.38; L: 0.27/0.50 Table 1: Description of datasets used in our experiments. Average token length (Train/Test) is computed with the Llama3.1-8b tokenizer. Dataset types: S = Sentence, D = Document. Tasks: HP = Hyperpartisan News Detection; FN = Fake News Detection; HT = Harmful Tweet Detection; PB = Political Bias Detection. Labels: HP = Hyperpartisan, N = Neutral; T/F = True/Fake; NH/H = Non-harmful/Harmful; R/C/L = Right/Center/Left. Abbr. contains the abbreviation we will use in this paper to refer to the datasets. (Mistral 2024), Qwen2.5-7B-Instruct (Qwen et al. 2025). These models are decoder-only and testing them allows us for generalizable effects across model families and tasks. The temperature was set to 0 for all the experiments. Generally, for non-English datasets, we evaluated models exclusively trained on multilingual datasets to ensure appropriate language coverage and performance. Further details are given in the Appendix . Prompt design Earlier studies like (Wei et al. 2022), (Jung et al. 2022) and (Mishra et al. 2022) have demonstrated the effectiveness of using task-specific prompts. Therefore, following (Edwards and Camacho-Collados 2024) and (Labrak, Rouvier, and Dufour 2024), we constructed the prompts concatenating the following elements: 1) an instruction detailing the task, domain, and describing the meaning of the label; 2) the input argument, supplying essential information for the task; 3) the constraints on the output space, guiding the model during output generation. To improve the coherence, the specificity of the prompts, the instructions to follow in the codebook, and the fine-grained reasoning in CoT for the political domain, we collaborated with an expert in Political Science. In particular, to structure and develop our codebooks, we were inspired by Vincent and Mestre (2018) and (Halterman and Keith 2025), which introduced clear task definitions, explicit and exhaustive rules to determine the label of a data point, as well as provided examples covering both correct and borderline cases. We tested different prompting and ICL strategies such as zero-shot, Few-Shot, codebook and a variant of guided CoT, intending the reasoning as a multi-task evaluation (Lee et al. 2024; Duan et al. 2024), to provide explainable results like in (Yang et al. 2023a). We compare the random selection of Few-Shot exemplars with a more diversified selection using DPP, for which we provide a brief introduction in the following section. We also compare the results given by prompting the models with instructions containing different levels of complexity: general instructions, specific definitions of political phenomena or specialized instructions with more context provided. During the prompt optimization phase, we placed particular emphasis on ensuring that the model adhered to a consistent label format. This was crucial to ensure the outputs were reliably parseable. For instance, we discovered that the models when asked to provide string labels generated few unparseable outputs, considered wrong at the time of inference. This behavior happened across all the models and configurations. Moreover, at the begininnig we tested Mistral-7B and most of the time it did not follow the instructions regarding the template. This is why we did not introduce it in the experimental setting. Lastly, in CoT, to ensure its right functioning, we made sure the models generated the thoughts before the final output. Please, see the Appendix for further details. Determinantal Point Process Determinantal Point Process is a probability distribution over cloud of points that are used as computational tools across the fields of physics, statistics and machine learning (Gautier et al. 2019). DPP has been used to select diverse and representative set of datapoints for in-context learning (Yang et al. 2023b), data annotation (Wang et al. 2024b), instruction tuning (Wang et al. 2024a) and pre-training (Yang et al. 2024). DPP has been preferred for these tasks because it helps in promoting efficiency while maintaining a diversity of the selected subset from a large set. For this let us take 2 sets called index set A={1,2,...N}and its corresponding item set IA={x1, x2, . . . , xN}. Then the problem of subset selection becomes evaluating 2Msubsets, which is computationally intractable and combinatorially explosive as the size of the super set grows. In order to approximately solve this problem DPP first uses the representation of the data xi. We use Sentence-BERT (Reimers and Gurevych 2019) to calculate the representations or embeddings. After the representations are computed, DPP algorithm works by calculating a Kernel Kij =k(xi,xj)where kernel function k can be any similarity or distance metric between 2 points. Based on this we select a subset Y∈A, the probability selection of Y is given by. P(Y) = det(KY) det(K+I)(1) Here KYis the subset of the matrix Kand consists of Kij for i, j ∈Y.Iis the identity matrix and det(·)represents the determinant of a matrix. Under these conditions the selection of the best subset can be formulated as the optimization problem as follows: Ybest =argmaxY⊂A,|Y|=kdet(LY)(2)
Algorithms exist to select such subsets of size k from the superset by sampling from the posterior distribution. For the task of selection we use the exact sampler (Gautier et al. 2019; Mazoyer, Coeurjolly, and Amblard 2020) which is the part of DPPy package by (Gautier et al. 2019) and is faster than Monte-Carlo sampler (Bardenet and Hardy 2019). Since the algorithm works on sampling at the kernel level it helps in selecting datapoints which are diverse in the representation space. ICL Setting Our aim is to compare the capabilities of different learning techniques, namely FT and ICL, and model architectures for hyperpartisan, fake news, harmful tweet and political bias classification. To investigate the ability of LLMs on those tasks in ICL, we used LlaMA3.1-8b-Instruct, MistralNemo-Instruct-2407 and Qwen2.5-7B-Instruct by prompting them with different setups: 0-shot with General Prompt, 0-shot with Specific Prompt, CoT and Few-Shot with kshot where k¿0. In the k-shot configuration, we adopted the General Prompt along with random examples and the respective labels of the dataset. To further test the stability of the prompting, we used Determinantal Point Process (Gautier et al. 2019) to select a diverse set of datapoints for each of the k-shot settings. The general prompt template is: <Role><Task description><Definition or Instructions><Text to classify or examples followed by text to classify><Response format>. We provide examples of different our prompts in the Appendix (see Tables 4, 5, 6 and 7). In order to maintain a balanced pool of examples, for multi-class datasets, we sampled from 1 to 3 examples per label; otherwise, we extracted the same k-shot per class. Finally, for non-English data, the corresponding roles and instructions were provided in the respective language to ensure accuracy and contextual relevance. The type of prompts we used are the following: General Prompt By providing the model with task-specific context (e.g., a headline, article, or tweet), we prompted it to classify the input text with the appropriate task label. With this configuration, we leverage the internal knowledge of the model to predict the answer, while being aware that it can suffer from political bias (Bang et al. 2024). We used it in 0-shot and Few-Shot. Specific Prompt We slightly changed the previous template, introducing in the instruction the political definition of the phenomenon analyzed and some knowledge regarding the biases in partisan texts and asked the model to classify the text with the correct label. These political definitions were provided by a domain expert. Thus, we insert external knowledge and introduce a political definition to maximize model’s understanding capability and improve its outputs’ quality. We tested its efficacy only in zero-shot. Codebook Based on previous works, with the help of a Political Scientist we crafted a specific codebook for each task. These codebooks contain a definition of the phenomenon, a description of the task’s characteristics considering several aspects (e.g., style, narrative) and particular linguistic features (e.g., use of hashtags, tone, source credibility). Furthermore, the detection criteria contain examples to help the LLM in understanding the specific rules for each task. Crucially, this detailed codebook information was directly embedded within the ¡Definition or Instructions¿ component of our prompt template, allowing the LLM to perform rule-based reasoning for classification. This approach enables the models to leverage explicit, structured knowledge during inference, directly addressing the complexities of the tasks. Guided CoT Prompt We guided the model to break down its reasoning step by step before making a final classification on the specified context, ensuring it produced explanations for all the steps before the final prediction. Specifically, we divided the hyperpartisan classification task into different sub-tasks: sentiment analysis, rhetorical bias, framing bias, ideology detection (Maggini and Gamallo Otero 2024). Moreover, by asking the model to identify itself with a specific political leaning, we are introducing recursive thinking (Duan et al. 2024). This method encourages the model to consider multiple factors and clearly articulate its reasoning, potentially leading to more robust and explainable classifications. Lastly, our approach allows us to consider the multidimensionality of each misinformation phenomenon analyzed. By guiding the model through this structured reasoning process, we aimed to reduce misclassification and promote a more nuanced analysis. This approach also enabled us to observe how the model weighs different textual elements in its decision-making process, which helps in identifying any inherent biases or limitations within the model’s reasoning. We conducted preliminary tests with various prompts and configurations to refine the ones used in this experiment, ultimately selecting the configurations that yielded the best results on the training set. The optimization of these prompts was done manually rather than through automated methods. As a result, our prompts vary in terms of length, complexity, task specificity, and domain relevance, providing a comprehensive range of settings for evaluation. This structured and manually-optimized approach not only enhances the model’s classification performance but also provides deeper insights into the model’s interpretability and decision-making process across different political contexts. Main Results and Discussion Fine-Tuning Table 2 presents the results for fine-tuning. On average across datasets, decoder-based models tend to outperform encoder-based models on tasks that require factual world knowledge, such as fake news detection and political bias identification. For instance, in fake news detection (Macro Avg. FN F1 score), the LlaMA3.1-8b (decoder) achieves .907, while the best performing encoder with a directly comparable macro average, ModernBERT-base, scores .854. Similarly, for political leaning detection (Macro Avg. PL F1 score), the Mistral-Nemo-Instruct (decoder) reaches .849, significantly surpassing the top encoder, POLITICS, which scores .675. Conversely, encoders achieve better results on linguistically oriented tasks, specifically harmful tweet detection and hyperpartisan language identification. RoBERTa-large (encoder) records an F1 score of .850 on
Model HV SH Macro Avg. HP FNN SFN FB Macro Avg. FN C1A C1B C1E Macro Avg. HF QB C3A Macro Avg. PL RoBERTa-base Acc .822 .865 .843 .879 – – – – – .919 – .622 .659 .640 F1 .818 .865 .841 .880 – – – – – .915 – .604 .660 .632 RoBERTa-large Acc .852 .865 .858 .893 – – – – – .925 – .683 .663 .673 F1 .850 .865 .857 .893 – – – – – .923 – .674 .660 .667 XLM-RoBERTa Acc .827 .801 .814 .842 .623 .957 .807 .917 .917 .917 .917 .582 .632 .607 F1 .825 .798 .811 .844 .577 .957 .793 .883 .884 .885 .884 .553 .629 .591 POLITICS Acc .831 .854 .842 .867 – – – – – .916 – .682 .679 .680 F1 .826 .854 .840 .868 – – – – – .911 – .673 .678 .675 mDeBERTaV3 Acc .819 .776 .797 .893 – – – – – .915 – .539 .590 .564 F1 .816 .772 .794 .841 – – – – – .874 – .534 .570 .552 ModernBERT-large Acc .829 .854 .839 .858 .863 .941 .883 .815 .965 .835 .872 .658 .654 .653 F1 .824 .853 .839 .846 .863 .941 .815 .815 .949 .803 – .649 .657 .667 ModernBERT-base Acc .764 .780 .767 .852 .782 .942 .854 .830 .966 .830 .875 .584 .617 .592 F1 .755 .779 .767 .840 .781 .942 .854 .792 .966 .774 .844 .571 .612 .592 LlaMA3.1-8b Acc .830 .810 .820 .945 .812 .975 .911 .864 .869 .921 .884 .788 .762 .775 F1 .830 .801 .815 .945 .801 .975 .907 .858 .832 .920 .870 .786 .763 .774 LlaMA3.1-8b-Instruct Acc .784 .820 .801 .875 .823 .976 .890 .878 .915 .862 .885 .786 .796 .791 F1 .782 .820 .801 .869 .823 .976 .889 .867 .928 .829 .875 .781 .802 .792 Mistral-Nemo-Instruct-2407 Acc .834 .736 .783 .867 .725 .976 .851 .850 .943 .838 .877 .790 .846 .819 F1 .833 .733 .783 .855 .722 .976 .851 .859 .946 .825 .877 .787 .789 .849 Qwen2.5-7B-Instruct Acc .819 .686 .745 .864 .685 .974 .835 .856 .913 .824 .864 .735 .698 .711 F1 .812 .677 .745 .855 .675 .974 .835 .863 .928 .806 .866 .728 .693 .711 Table 2: Performance of models in the FT setting. The reported weighted Accuracy and weighted F1 scores are the averages obtained by running each model five times on the same dataset, reporting standard deviation. HV, compared to the best decoder, Mistral-Nemo-Instruct, at .833. For SH, RoBERTa-large again leads with an F1 score of .865, while the best performing decoder on this task, LlaMA3.1-8b-Instruct, achieves .820. We hypothesize that this difference arises because the bidirectional attention mechanism of encoders may be better at capturing nuanced linguistic features, whereas decoders might excel at tasks more reliant on content or semantic understanding. Surprisingly, continuous pretraining of RoBERTa-base aimed at adapting it to either the political domain (POLITICS) or multilingual contexts (XLM-RoBERTa) has, in several instances, not led to improved performance over the original RoBERTa-base model and sometimes resulted in a reduction. For example, on the Macro Avg. HP (Hyperpartisan) task, RoBERTa-base achieves an F1 score of .841, whereas POLITICS scores .840 and XLM-RoBERTa scores .811. Regarding decoder models, we observe that a larger parameter count or expanded training corpus does not necessarily equate to superior results. This is demonstrated by LlaMA 3.1-8B outperforming Mistral Nemo-Instruct2407 in hyperpartisan detection (SH F1 score of .801 for LlaMA 3.1-8b versus .733 for Mistral Nemo). Meanwhile, the decoder models exhibit relatively comparable high performance in harmful text detection (Macro Avg. HF F1 scores): LlaMA3.1-8b (.870), LlaMA3.1-8b-Instruct (.875), and Mistral-Nemo-Instruct-2407 (.877). During the FT experiment, we reached the SOTA in HV, FNN, SFN and FB. We provide the comparison with the previous research in Table 9 in the Appendix. In-Context Learning For the results discussed in this section, please refer to Figures 1 and 2 for the zero-shot configurations and CoT; and 3 for FS. Detailed results are reported in Table 8 in the Appendix HP zero-shot-general vs zero-shot-specific In the hyperpartisan (HP) task, moving from generic to specific zeroshot prompting results in moderate improvements across all three models, especially for LLaMA 3.1-8B Instruct and Mistral-Nemo-Instruct, whose F1 scores increase from 0.678 and 0.686 to 0.738 and 0.740, respectively. This performance gap indicates that these models have learned robust representations of hyperpartisan content during pretraining, and that providing more detailed task descriptions helps further refine their predictions. Lastly, SH predictions are generally stronger, largely because this dataset includes full articles rather than just headlines, as in HV. The richer contextual information in SH provides models with more linguistic and semantic cues, enabling a deeper understanding of the content and improving their ability to detect hyperpartisan narratives. In contrast, the limited context in headlines offers fewer signals for accurate classification. Codebook The codebook approach yields improvements for SH, where Llama and Qwen demonstrate the most significant gains, particularly on the SH dataset, where their performance approaches F1 .810 and .748 respectively. This suggests that providing explicit criteria for identifying partisan language can help address edge cases where the models depend solely on their internal task representations. The performance gap between models narrows with codebook prompting, indicating that structured guidance can help equalize performance differences stemming from model’s architecture, that may rely on different definitions of the phenomenon investigated. FN zero-shot-general vs zero-shot-specific Fake news detection tasks show variable performance under zero-shot conditions. The FNN task proves more challenging are the multilingual datasets: SFN and FBC. The transition from generic to specific prompting yields minimal gains, with Mistral-Nemo-Instruct-2407 maintaining the highest performance on 2 out of 3 fake news (FN) datasets. We observed a 0.275-point drop in F1 score for Qwen on the Spanish dataset, which may be attributed to a misalignment between the fake news definitions—specifically, the model’s internally assumed definition in the zero-shot generic prompt versus the expert-crafted definition used in the zero-shot spe-
Figure 1: Results for zero-shot and CoT grouped by models. Figure 2: Results zero-shot and CoT grouped by configuration. cific prompt. This limited improvement suggests that, beyond clear definitions, models need additional contextual or world knowledge to effectively differentiate factual from fabricated content. Codebook The codebook approach yields minimal improvements for FNN and SFN. indicates that providing explicit criteria for evaluating factual claims, source credibility markers, and stylistic indicators of fabricated content may help models overcome the inherent complexity of fact verification. The codebook’s little effectiveness in this domain suggests that FN does not completely benefit from structured evaluation frameworks. Regarding FN, with the different prompt strategies tested in ICL, we found out that the small LLMs are not effectively capable of detecting this kind of disinformation because they can not rely properly on the ontological structures encoded in their world-knowledge. Thus, employing these models as out-of-the-box tools with an ICL setup proves to be inefficient for this task. PL zero-shot-general vs zero-shot-specific Political leaning (PL) classification tasks exhibit low performance across all models under zero-shot conditions. It is important to note that this is a multiclass classification task, which adds complexity. Generally, specific prompting produces a slight decrease for this task. Llama maintains consistently higher performance (on average F1 .416) than the other models. The limited effectiveness with a specific definition of the political wings suggests that PL requires more than definitional refinement to overcome the inherent subjectivity involved. Codebook Providing more detailed and specific knowledge through the codebook generally resulted in decreased performance across all models. This highlights the complexity of the task and suggests that, despite offering explicit rules to interpret the U.S. political context, agendas, the cultural nuances and linguistic factors involved are insufficient for effectively addressing the task. This result reveals the complexity underlying political bias detection. HF zero-shot-general vs zero-shot-specific HF covered three languages and Qwen reached the best results in C1A with zero-shot-generic prompts (F1 .851). However, when prompted with zero-shot-specific prompts, its performance decreased. On the other hand, Llama particularly benefitted from the introduction of the specific knowledge for C1B (from F1 .507 to .676) and C1E (from F1 .730 to .764). Codebook With the introduction of specific dimensions to frame the task, all the models across the HF datasets (except for Qwen in C1A) improved their performances. Specifically, Mistral reached F1 .864 in C1B. This marked improvement highlights the value of explicit harm taxonomies and classification criteria for this sensitive domain. The codebook’s effectiveness for harmful content detection
Figure 3: FS DPP vs Random results. suggests that these tasks require clearly articulated boundaries and examples to overcome potential ambiguities in what constitutes harmful material. Summary for Zero-shot configurations Across all classification domains, the transition from generic to specific zero-shot prompting lead only to marginal improvements. This pattern suggests that merely elaborating on task definitions provides insufficient guidance for these models to significantly alter their classification behavior. The codebook approach demonstrates modest effectiveness across all task categories, with particularly improvements for HF and FN classification. This pattern indicates that providing structured classification criteria helps models overcome the inherent complexities of these judgment tasks. The codebook’s effectiveness stems from its ability to bridge the gap between abstract classification concepts and concrete textual indicators, providing models with clearer decision boundaries for ambiguous cases. Lastly, we noticed that more subjective and nuanced tasks like HF and PL, that require extensive knowledge of the facts and their truthfulness, show improvement with advanced prompting strategies than more stylistic-based tasks (e.g. HP detection), suggesting that prompt optimization benefits may correlate with task complexity. Indeed, the rule based approach is the best ICL configuration for 3 out of 10 datasets: SH, C1B and FBC. FS: DPP-selected vs Random examples To test the Few-Shot capacity of the model, we decided to compare the performances using random datapoints against a representative set of examples, maintaining the dataset diversity using DPP. Table 3 reports the results for this comparison. Random selection risks subsets that lack diversity or fail to represent edge cases, while DPP ensures prompt stability by challenging the model with dissimilar patterns. In Few-Shot Learning, where models rely on limited examples to generalize, diverse subsets prevent overfitting to specific features and improve generalization. DPP-selected examples enhance prompt informativeness by showcasing varied cases, enabling the model to better understand nuanced relationships. This diversity reduces classification errors and improves accuracy by covering a broader range of inputs. A key observation across both FS Random and FS DPP is that increasing the number of shots does not consistently or monotonically improve performance for all models and datasets. Generally, performance often peaks at an intermediate number of shots (e.g., 1-shot, 5-shot, 6-shot) and can then plateau, fluctuate, or even decline as more shots are added (e.g., in the range 7-10 shots). This implies that simply providing more examples is not always the best strategy, regardless the datapoints representativeness of the selected examples. There is not any ideal threshold for the n-shots, since the performances vary across models and datasets. For instance, while using FS Random examples in Hyperpartisan Detection on the SH dataset, Llama-3.1-8b-Instruct’s F1 score varies from .751 (1-shot) down to .623 (4-shot) and then to .587 (10-shot). However, the results do not indicate a clear best method across all scenarios. Peak performance for a given model and dataset can be achieved by either method, often at different n-shot values. For example, with Llama-3.1-8B-Instruct on the HV dataset, Random 9shot yields an F1 of .811, while DPP 10-shot gives .801. Conversely, Mistral-Nemo-Instruct-2407 on C1B achieves its highest few-shot F1 of .851 with DPP 10-shot. FS Random can achieve high scores, possibly when the random selection happens to include particularly effective examples. However, its performance can be inherently more variable. FS DPP, by design, selects for diversity, which might be expected to lead to more robust or consistent improvements. While it achieves strong results in some cases, it also exhibits fluctuations and doesn’t always outperform Random FS. Chain of Thought This setting provided reasoning steps with increasing levels of abstraction modeled as successively sub-task steps. Llama-3.1-8B-Instruct shows good performance with CoT on HP tasks, achieving the best F1/Accuracy in its section for both SH (F1 .792, Acc .795) and HV (F1 .757, Acc .764) datasets. It also performs well on PL Detection for C3A (F1 .465, Acc .459) and QB (F1 .416, Acc .405), again leading in its section. Mistral with CoT achieves the overall best scores for FN Detection across all models and configurations on the FNN dataset (F1 .623, Acc .680) and across sections in FBC dataset (F1 .315, Acc .397), and SFN (F1 .167, Acc .168). This suggests CoT might be particularly beneficial for this
model on these specific complex tasks. Regarding, Qwen, CoT helps it in achieving strong results on C1A (F1 .804, Acc .783), but only F1 .646 in C1B and F1 .575 in C1E, leading this section for these datasets. Nevertheless, in most of the cases, it revealed to be suboptimal. The CoT prompting proved largely suboptimal across our experiments, showing significant improvement only for the FNN task. Our analysis suggests this underperformance stems primarily from language representation issues in the model’s training data. When prompted in underrepresented languages with insufficient training tokens, the models struggled to process the unfamiliar linguistic patterns. Rather than aiding reasoning, these novel tokens appeared to confuse the models, disrupting their inference capabilities. Additionally, we observed that even in zero-shot-specific, codebook, and few-shot configurations, the models sometimes generated explanations unprompted, suggesting they were trained to occasionally provide reasoning alongside their answers. This built-in explanatory behavior likely accounts for why explicit CoT prompting offered minimal additional benefits despite its theoretical advantages and added complexity. Insights from Fine-Tuning vs. In-Context Learning in LLMs Across all the configurations tested in our experiments, FT emerged as the most effective method to apply models to the political domain, and, in particular, the disinformation subdomain. Specifically, for LLMs, the FT configuration demonstrated its efficacy in 28 out of 33 cases. Notably, Qwen particularly benefited from ICL achieving its best performance with the following configurations: few-shot random for C1E (F1: 0.833) and SFN (F1: 0.678), and zero-shot codebook for SH (F1: 0.810). These results suggest that updating model parameters through FT is generally the most reliable way to optimize performance for downstream tasks in this domain. However, ICL remains a valid and convenient strategy for probing a model’s task-specific knowledge without parameter updates. Despite our efforts to optimize prompts—by incorporating external domain-specific knowledge, employing rule-based approaches, and eliciting reasoning capabilities—ICL configurations still showed more limited effectiveness compared to FT. Lastly, both model architectures benefited from fine-tuning, with encoder-based models achieving superior performance on 6 out of 10 datasets, and smaller LLMs performing better on the remaining 4—particularly in tasks such as fake news and political leaning detection, which require deeper world knowledge. It is important to note, however, that fine-tuning —especially when applied to LLMs—demands significant computational resources, making it a considerably resource-intensive approach. Conclusion and Future Work This paper provides a comprehensive benchmark for FT and ICL methods across classification tasks in several misinformation domains. Indeed, we largely compared different model architectures, learning techniques and sets of prompts in several classification tasks. We evaluated performance on 10 diverse datasets, spanning binary and multiclass contexts in English, Spanish, Brazilian Portuguese, Arabic, and Bulgarian. Particularly, we compared nine models covering the following transformer family’s architecture: (1) encoders: RoBERTa-base, RoBERTalarge, XLM-RoBERTa, POLITICS, Modern-BERT-base and -large, mDeBERTaV3; (2) decoders: LlaMA3.1-8b-Instruct, Mistral-Nemo-Instruct-2407 and Qwen2.5-7B-Instruct. ICL consistently proved to be less effective than fine-tuning across most settings. Indeed, results showed that in FT, decoders were better at PB and FN detection, whereas encoders were better at HP and HF tweet detection. Regarding ICL, we applied different levels of prompt optimization, testing all the main ICL techniques. In Few-Shot, we found that DPP sometimes reduces variations in classification rather than randomly sampling the shots, though it does not systematically increase the performance. Furthermore, the adaptation methods in ICL did not behave the same depending on the LLM used. Lastly, except for fake news detection tasks, eliciting the model with a codebook was generally the best approach in ICL, making the CoT unreliable for classification task. Future work could investigate the application of a RAG system to incorporate up-to-date information and improve the factual verification of news. Limitations LLMs In CoT, to overcome the different templates generated by the model in the initial phase of our experiment, we crafted a template for the different tasks (see the prompt Tables 4,5,6,7 in the Appendix). Moreover, in few cases the model produced irregular outputs. When this occurred, we considered those cases to be incorrectly labeled. Practical applications Our work is a comprehensive overview of architectures and methods. Nevertheless, some datasets do not contain up-to-date information that can be used effectively to tackle fake news propagation, since this kind of misinformation does not rely only on linguistic clues but also on pre-existence knowledge of political facts. Indeed, the temporal limitation can affect how the perception of a fact - a general one - can be perceived and or discussed, because the language is subjected to changes over time. However, those dataset could be used for continual pre-training or as part of a RAG system that also incorporates up-to-date information. Domain We acknowledge that our experiments focus on a subset of political NLP tasks, specifically misinformationrelated tasks across different languages. As such, our findings should not be generalized to the full range of NLP tasks. Open LLMs and size We limited our model selection to open models, while discarding closed one (e.g. Claude, OpenAI) for the sake of reproducibility and budget limitations. Furthermore, our GPUs could not host larger models. This fact limits our findings. Ethics Statement One of the primary objectives of this work is to address the challenge of misinformation spreading — a critical issue in today’s society. Tackling misinformation is both ben-
Table 5: Table showing different examples of instruction for HF Detection. Prompt Text zero shot generic ”You are a fake news expert. You will be provided with a text, and your task is to determine, to the best of your judgment, whether the text is harmful to society or not. If you are unsure, make the most informed decision you can based on the content. If the news is harmful, the label is ’harmful’, if the text is neutral, the label is ’neutral’. Make sure to structure your answer in the following manner:\n\n Final Answer ==>harmful\n\nor \n\nFinal Answer ==>neutral\n\n Here’s the text: ”, zero shot generic ”You are a fake news expert. Harmful tweet is a post on the social media platform Twitter that can cause distress, harm, or damage to individuals, groups, or society. This type of tweet may include content that is abusive, offensive, defamatory, threatening, or incites violence against the policies to contain the spread of COVID-19. Additionally, it may spread false information, contribute to cyberbullying, or perpetuate hate speech and discrimination. Harmful tweets can negatively impact mental health, reputations, and community safety. You will be provided with labeled examples and a text, and your task is to determine, to the best of your judgment, whether the text is harmful to society or not. If you are unsure, make the most informed decision you can based on the content. If the news is harmful, the label is ’harmful’, if the text is a neutral, the label is ’neutral’. Make sure to structure your answer in the following manner:\n\n Final Answer ==>harmful\n\nor\n\nFinal Answer ==>neutral\n\n Here’s the text: ”, few shot ”You are a fake news expert. You will be provided with labeled examples and a text, and your task is to determine, to the best of your judgment, whether the text is harmful to society or not. If you are unsure, make the most informed decision you can based on the content. If the news is harmful, the label is ’harmful’, if the text is neutral, the label is ’neutral’. Make sure to structure your answer in the following manner:\n\n Final Answer ==>harmful\n\nor\n\nFinal Answer ==>neutral\n\nHere’s the text: ”, codebook ”You are a fake news expert. **Definition: Harmful tweet is a post on the social media platform Twitter that can cause distress, harm, or damage to individuals, groups, or society. This type of tweet may include content that is abusive, offensive, defamatory, threatening, or incites violence against the policies to contain the spread of COVID-19. Additionally, it may spread false information, contribute to cyberbullying, or perpetuate hate speech and discrimination. Harmful tweets can negatively impact mental health, reputations, and community safety. **Task: You will be provided with a codebook and a text, and your task is to determine, to the best of your judgment, whether the text is harmful to society or not. If you are unsure, make the most informed decision you can based on the content. If the news is harmful, the label is ’harmful’, if the text is a neutral, the label is ’neutral’. Make sure to structure your answer in the following manner:\n\n Final Answer ==>harmful\n\nor \n\nFinal Answer ==>neutral\n\n Here’s the text: \n\nThis is the codebook: — Label — Class — Definition — Typical Content Patterns — Neutral — Neutral— The tweet provides factual, supportive, or neutral information about COVID-19. — - Promotes vaccinations, masking, or safety —\n—- Shares news from reliable sources (WHO, CDC) —\n—- Describes personal experiences without disinformation —\n— Harmful — Harmful — The tweet contains disinformation, conspiracy theories, or medically incorrect claims. — - Claims vaccines are deadly or part of a conspiracy —\n—- Encourages rejection of public health guidance —\n—- Spreads pseudoscience or political hoaxes —\n\n— Aspect — Description of Harmful Characteristics Conspiracy Theory — References to population control, “pandemic”, 5G causes COVID, microchips in vaccines —\n— Distrust in Science — Discredits CDC/WHO, claims that mainstream medicine is false or corrupt —\n— Alternative Cures — Promotes ivermectin, bleach, natural remedies as vaccine alternatives without evidence —\n— Anti-Vaccine — Claims vaccines are unsafe, rushed, or part of a control agenda —\n— Hoax Narrative — Declares COVID-19 is fake, exaggerated, or invented by governments —\n— Call to Action — Urges followers to ignore mandates, avoid vaccination, or resist masking —\n— Language Style — Emotionally charged, sarcastic, or mocking tone; often uses ALL CAPS, emojis, hashtags like **NoVax **Plandemic —\n\n— Linguistic Features — Harmful — Not Harmful — Use of hashtags — **plandemic, **NoVax, **scamdemic — **GetVaccinated, **StaySafe ——\n—— Source credibility — Anonymous sources, influencer claims — WHO, CDC, news outlets ——\n— Personal stories — Claims of side effects or hoaxes — Stories of recovery, gratitude for vaccines ——\n— Tone — Alarmist, rebellious, sarcastic — Informative, cautious, or appreciative —. \n\nMake sure to structure your answer in the following manner:\n\n Final Answer ==>harmful\n\nor\n\nFinal Answer ==>neutral\n\n Here’s the text: ”, cot ”You are a fake news expert. You will be provided with labeled examples and a text, and your task is to determine, to the best of your judgment, whether the text is harmful to society or not. If you are unsure, make the most informed decision you can based on the content. Analyze the text for fake news using step-by-step reasoning. If the news is neutral, the label is ’neutral’, if the text is harmful, the label is ’harmful’. You have to choose one label and the output is only one label. Make sure to structure your answer in the following manner: Explanation: [Eplanation for each reasoning step]. \n\n Final Answer ==>harmful —\n— or —\n— Final Answer ==>eutral.—\n—****** Reason step by step:—\n—1. *Lexical Analysis**\n\nCheck for vague sourcing, absolutist language, logical fallacies—\n—. **Sentiment Analysis**\n\nIdentify polarizing/emotional language—\n—3. **Content Authenticity**\n\nDetect bot-like patterns, verify data—\n—4. **Framing Bias**\n\nAssess selective presentation—\n— 5. Audience Analysis**\n\nDetermine target audience and intent—\n—6. **Information Adequacy**\n\nEvaluate comprehensiveness and reliability. \n\n Make sure to structure your answer in the following manner:Explanation: [Explanation for each reasoning step]. Final Answer ==>harmful\n\nor\n\nFinal Answer ==>neutral. \n\n Here’s the text: ”
Table 5 – continued from previous page Prompt Text Table 5: Table showing different examples of instruction for HF Detection. For the multilingual dataset, we translated this prompt to Arabic and Bulgarian.
Table 6: Table showing different examples of instruction for HP Detection. Prompt Text zero shot generic ”You are a hyperpartisan news expert. You will be provided with a text, and your task is to determine, to the best of your judgment, whether the text is hyperpartisan or not. If you are unsure, make the most informed decision you can based on the content. The possible choices are: ’neutral’ if the article is neutral, ’hyperpartisan’ if the article is hyperpartisan. You have to choose one label and the output is only one label. \n\n Make sure to structure your answer in the following manner:\n\n Final Answer ==>hyperpartisan\n\nor \n\nFinal Answer ==>neutral \n\n This is the text: ” ”zero shot generic” ”You are a hyperpartisan news expert. You will be provided with a text, and your task is to determine, to the best of your judgment, whether the text is hyperpartisan or not. If you are unsure, make the most informed decision you can based on the content. Hyperpartisan articles contain biases, particularly ad hominem attack, loaded language, and evidence of political ideology. Sometimes they rely on cherry-picking strategy. The possible choices are: ’neutral’ if the article is neutral, ’hyperpartisan’ if the article is hyperpartisan. You have to choose one label and the output is only one label. Make sure to structure your answer in the following manner:\n\n Final Answer ==>hyperpartisan\n\nor \n\nFinal Answer ==>neutral \n\n This is the text: ” few shot ”You are a hyperpartisan news expert. You will be provided with a text, and your task is to determine, to the best of your judgment, whether the text is hyperpartisan or not. If you are unsure, make the most informed decision you can based on the content. You will be provided with some examples of labeled text. The possible choices are: ’neutral’ if the article is neutral, ’hyperpartisan’ if the article is hyperpartisan. You have to choose one label and the output is only one label. Make sure to structure your answer in the following manner:\n\n Final Answer ==>hyperpartisan\n\nor \n\nFinal Answer ==>neutral \n\n******Examples of labelled articles: \n\n This is the text: ” codebook ”You are a hyperpartisan news expert. You will be provided with a text, and your task is to determine, to the best of your judgment, whether the text is hyperpartisan or not. If you are unsure, make the most informed decision you can based on the content. You will be provided with some examples of labeled text. \n\n Definition: Hyperpartisan news detection is the process of identifying news articles that exhibit extreme one-sidedness, characterized by a pronounced use of bias. The prefix hyperhighlights the exaggerated application of at least one specific type of bias—such as spin, ad hominem attacks, opinionated statements, ideological slants, framing, selective coverage, political leaning, or slant bias—to promote a particular ideological perspective. This strong ideological alignment is conveyed through amplified linguistic elements that reinforce one of these bias types within the text.\n\n****** Task: Read carefully the codebook provided and assign a label to the text. You can choose only one label. If the text is neutral, you will write ’neutral’, if it is hyperpartisan ’hyperpartisan’.\n\nFollow the output template given as an example. Under no circumstances we are asking to provide or generate harmful content. Please, provide only the label.\n\n****** Make sure to structure your answer in the following manner:\n\n Final Answer ==>hyperpartisan\n\nor \n\nFinal Answer ==>neutral \n\n **This is the codebook: Linguistic Features:\n\n******** 1. Lexical Features —\n Feature — Hyperpartisan — Neutral — Lexical Polarity — Frequent use of emotionally charged words (e.g., disaster, outrageous) — Neutral and precise language ——\n—— Modality and Certainty — Strong modal verbs (e.g., will destroy) — Hedging markers (e.g., may, might) ——\n—— Repetition — Repeating claims or slogans — Minimal repetition ——\n—— Pronouns — Frequent us vs them language — Focus on third-person objectivity —\n\n******** 2. Rhetorical Devices—\n—— Feature — Hyperpartisan — Neutral ——\n—— Appeal to Emotion — Frequent appeals to fear/anger — Logical/factual appeal ——\n—— Hyperbole — Common exaggeration — Proportional statements ——\n—— Metaphors — Politically loaded metaphors — Literal language preferred —\n\n******** 3. Discourse Structure—\n—— Feature — Hyperpartisan — Neutral ——\n—— Framing — Blame/conflict framing — Balanced framing —\n— Source Attribution — Partisan sources only — Multiple reputable sources ——\n—— Balance of Views — One-sided presentation — Multiple perspectives —\n\n******** 4. Ideological Markers —\n—— Feature — Hyperpartisan — Neutral —\n— Us vs Them — Strong binary division — Avoids binaries —\n— Ideological Alignment — Clear left/right alignment — Issue-focused —\n\n******** 5. Pragmatic Features—\n—— Feature — Hyperpartisan — Neutral ——\n—— Intent — Persuade/convert — Inform/explain ——\n—— Tone — Confrontational/accusatory — Formal/detached —\n\n****** \n\n Final Answer ==>hyperpartisan\n\nor \n\nFinal Answer ==>neutral\n\n This is the text: ”
Table 6 – continued from previous page Prompt Text cot ”You are a hyperpartisan news expert. You will be provided with a text, and your task is to determine, to the best of your judgment, whether the text is hyperpartisan or not. If you are unsure, make the most informed decision you can based on the content. You have to choose one label and the output will be the explanation and the determined label for that article. Make sure to structure your answer in the following manner:\n\n Explanation: [Eplanation for each reasoning step]. \n\n Final Answer ==>hyperpartisan\n\nor \n\nFinal Answer ==>neutral. \n\n Reason step by step:\n\n1. **Lexical and Sentiment analysis**: Analyze the tone and language. Does the article use polarizing, emotionally charged, or exaggerated language? Are there strong positive/negative sentiments toward a group, ideology, or issue?\n\n2. **Rhetorical bias**: Does the text use loaded language, name-calling, or manipulative rhetoric? Are there oversimplifications, strawman arguments, or exaggerated claims?\n\n3. **Framing bias**: Is information presented selectively to favor one perspective? Does it emphasize certain aspects while downplaying others to shape perception?\n\n4. **Ideological bias**: Does it emphasize certain aspects while downplaying others to shape perception? Does the article vilify opposing views rather than engaging with them fairly?\n\n5. **Unilateral coverage**: Does the article present multiple viewpoints (neutral) or only one side (hyperpartisan)? Are opposing arguments ignored, misrepresented, or dismissed?\n\n6. **Intent and Purpose**: Is the primary goal to inform objectively (neutral) or to persuade/mislead (hyperpartisan)? Does it present facts fairly, or does it push a clear agenda?\n\n7. Final prediction: Based on your previous considerations, classify the input as:\n\n- ’neutral’ (Neutral): Balanced, factual, and objective.\n\n- ’hyperpartisan’: Biased, one-sided, or manipulative.\n\n While generating the explanation for each reasoning step be coincise.. \n\n Remember to follow the output template for the label. \n\n Final Answer ==>hyperpartisan\n\nor \n\nFinal Answer ==>neutral\n\n This is the text: ” Table 6: Table showing different examples of instruction for HP Detection.
Table 7: Table showing different examples of instruction for FN Detection. Prompt Text zero shot generic ”You are a fake news expert. You will be provided with a text, and your task is to determine, to the best of your judgment, whether the text is fake news or not. If you are unsure, make the most informed decision you can based on the content. If the news is true, the label is ’true’, if the text is a fake news, the label is ’fake’. Make sure to structure your answer in the following manner:\n\n Final Answer ==>fake\n\nor \n\nFinal Answer ==>true\n\n Here’s the text: ” zero shot generic ”You are a fake news expert. We define fake news as: ¨ purposefully crafted, sensational, emotionally charged, misleading or totally fabricated information that mimics mainstream news¨ . You will be provided with a text and your task is to determine, to the best of your judgment, whether the text is fake news or not. If you are unsure, make the most informed decision you can based on the content. If the news is true, the label is ’true’, if the text is a fake news, the label is ’fake’. Make sure to structure your answer in the following manner:\n\n Final Answer ==>fake\n\nor \n\nFinal Answer ==>true\n\n Here’s the text: ”, few shot ”You are a fake news expert. You will be provided with a text and labeled examples, and your task is to determine, to the best of your judgment, whether the text is fake news or not. If you are unsure, make the most informed decision you can based on the content. If the news is true, the label is ’true’, if the text is a fake news, the label is ’fake’. Make sure to structure your answer in the following manner:\n\n Final Answer ==>fake\n\nor \n\nFinal Answer ==>true\n\n Here’s the text: ”, codebook ”You are a fake news expert. You will be provided with a text, and your task is to determine, to the best of your judgment, whether the text is fake news or not. If you are unsure, make the most informed decision you can based on the content. **Definition: Fake news detection identifies intentionally misleading content characterized by sensationalism, lack of credible sources, or manipulative language.\n\nIf the news is true, the label is ’true’, if the text is a fake news, the label is ’fake’. \n\n****** Task: To structure the output, follow the template in the example. \n\n Make sure to structure your answer in the following manner:\n\n Final Answer ==>fake\n\nor \n\nFinal Answer ==>true. This is the codebook you have to use to perform the classification: \n\n****** Detection Criteria:\n1. **Source Origin**\n\nReal: Established media, government sources\n\nFake: Anonymous/unverifiable sources\n2. **Event Reporting**\n\nReal: Specific references, quantitative data\n\nFake: Broad generalizations, unverifiable claims\n3. **Language Style**\n\nReal: Neutral, professional\n\nFake: Sensationalized, lickbait\n4. **Entity Authenticity**\n\nReal: Named real-world entities\n\nFake: Fictional/misspelled names\n5. **Claim Reliability**\n\nReal: Evidencebased\n\nFake: Absurd/absolute claims\n6. **Emotional Tone**\n\nReal: Objective\n\nFake: Emotional intensifiers\n7. **Source Credibility**\n\nReal: Verifiable\n\nFake: Unknown/misleading domains \n8. **Political Balance**\n\nReal: Multi-perspective\n\nFake: One-sided\n9. **Satire Markers**\n\nReal: No satire \nFake: Absurd content\n10. *Conspiracy Indicators* \nReal: Supported theories\n\nFake: Fringe conspiracy phrases\n\n Make sure to structure your answer in the following manner:\n\n Final Answer ==>fake\n\nor \n\nFinal Answer ==>true \n\n Here’s the text: ”, cot ”You are a fake news expert. You will be provided with a text, and your task is to determine, to the best of your judgment, whether the text is fake news or not. If you are unsure, make the most informed decision you can based on the content. Analyze the text for fake news using step-by-step reasoning. If the news is true, the label is ’true’, if the text is a fake news, the label is ’fake’. You have to choose one label and the output is only one label. To structure the output, follow the template in the example. \n\nOutput example: Explanation: [Eplanation for each reasoning step]. Final Prediction: Final Answer ==>[true or fake]. \n\n****** Reasoning step by step:\n1. **Lexical Analysis**\n\nCheck for vague sourcing, absolutist language, logical fallacies\n2. **Sentiment Analysis**\n\nIdentify polarizing/emotional language\n3. **Content Authenticity**\n\nDetect bot-like patterns, verify data\n4. **Framing Bias**\n\nAssess selective presentation\n5. Audience Analysis**\n\nDetermine target audience and intent\n6. **Information Adequacy**\n\nEvaluate comprehensiveness and reliability. \n\n Make sure to structure your answer in the following manner:\n\n Final Answer ==>fake\n\nor \n\nFinal Answer ==>true \n\n Here’s the text: ” Table 7: Table showing different examples of instruction for FN Detection. For the multilingual dataset, we translated this prompt to Spanish and Portugues. Results and SOTA
Hyperpartisan Det. Political Leaning Det. Harmful Text Det. Fake News Det. SH HV C3A QB C1E C1A C1B FNN FBC SFN F1 Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 Acc Zero-Shot Generic Llama-3.1-8B-Instruct 0.678 0.701 0.742 0.739 0.437 0.425 0.426 0.425 0.730 0.689 0.722 0.679 0.507 0.372 0.364 0.329 0.240 0.256 0.208 0.210 Mistral-Nemo-Instruct-2407 0.686 0.707 0.607 0.623 0.422 0.421 0.360 0.360 0.720 0.677 0.565 0.516 0.793 0.695 0.551 0.542 0.304 0.327 0.144 0.145 Qwen2.5-7B-Instruct 0.780 0.782 0.621 0.678 0.404 0.399 0.336 0.335 0.783 0.761 0.851 0.857 0.755 0.643 0.328 0.321 0.263 0.265 0.469 0.335 Specific Llama-3.1-8B-Instruct 0.738 0.748 0.764 0.764 0.431 0.419 0.419 0.417 0.764 0.737 0.721 0.678 0.676 0.548 0.412 0.371 0.269 0.326 0.182 0.182 Mistral-Nemo-Instruct-2407 0.740 0.750 0.574 0.593 0.409 0.408 0.352 0.356 0.719 0.677 0.613 0.562 0.818 0.732 0.549 0.540 0.315 0.354 0.150 0.150 Qwen2.5-7B-Instruct 0.789 0.790 0.585 0.668 0.395 0.391 0.340 0.335 0.721 0.685 0.827 0.818 0.738 0.622 0.257 0.278 0.275 0.276 0.194 0.205 Codebook Llama-3.1-8B-Instruct 0.748 0.753 0.715 0.722 0.382 0.407 0.324 0.344 0.762 0.729 0.728 0.686 0.670 0.538 0.388 0.349 0.304 0.405 0.165 0.166 Mistral-Nemo-Instruct-2407 0.694 0.713 0.706 0.704 0.372 0.385 0.318 0.340 0.753 0.717 0.636 0.585 0.864 0.806 0.580 0.595 0.335 0.440 0.144 0.145 Qwen2.5-7B-Instruct 0.810 0.811 0.636 0.688 0.377 0.395 0.311 0.323 0.785 0.769 0.832 0.838 0.820 0.735 0.341 0.320 0.256 0.256 0.422 0.295 Few-Shot Random 1-shot Llama-3.1-8B-Instruct 0.751 ±0.037 0.755 ±0.034 0.758 ±0.015 0.756 ±0.015 0.463 ±0.022 0.457 ±0.024 0.411 ±0.003 0.411 ±0.006 0.725 ±0.026 0.684 ±0.032 0.745 ±0.036 0.708 ±0.044 0.580 ±0.047 0.446 ±0.047 0.288 ±0.033 0.270 ±0.023 0.257 ±0.028 0.305 ±0.036 0.184 ±0.006 0.185 ±0.006 Mistral-Nemo-Instruct-2407 0.716 ±0.017 0.728 ±0.014 0.776 ±0.021 0.774 ±0.021 0.406 ±0.011 0.390 ±0.013 0.393 ±0.014 0.395 ±0.009 0.679 ±0.024 0.629 ±0.028 0.511 ±0.110 0.473 ±0.091 0.827 ±0.067 0.752 ±0.096 0.570 ±0.028 0.580 ±0.049 0.298 ±0.011 0.334 ±0.022 0.155 ±0.005 0.156 ±0.006 Qwen2.5-7B-Instruct 0.718 ±0.041 0.730 ±0.035 0.615 ±0.025 0.683 ±0.013 0.394 ±0.020 0.402 ±0.014 0.356 ±0.007 0.355 ±0.006 0.834 ±0.010 0.838 ±0.015 0.820 ±0.009 0.854 ±0.005 0.810 ±0.035 0.723 ±0.051 0.248 ±0.031 0.273 ±0.011 0.306 ±0.011 0.332 ±0.021 0.629 ±0.016 0.640 ±0.012 Random 2-shot Llama-3.1-8B-Instruct 0.664 ±0.041 0.685 ±0.031 0.767 ±0.020 0.765 ±0.019 0.457 ±0.021 0.451 ±0.018 0.412 ±0.007 0.416 ±0.008 0.746 ±0.012 0.708 ±0.015 0.712 ±0.034 0.668 ±0.039 0.422 ±0.066 0.302 ±0.052 0.300 ±0.027 0.278 ±0.020 0.241 ±0.018 0.269 ±0.019 0.179 ±0.011 0.179 ±0.011 Mistral-Nemo-Instruct-2407 0.725 ±0.013 0.735 ±0.009 0.641 ±0.069 0.653 ±0.057 0.421 ±0.011 0.407 ±0.011 0.404 ±0.018 0.400 ±0.017 0.626 ±0.030 0.574 ±0.032 0.619 ±0.066 0.571 ±0.066 0.797 ±0.032 0.703 ±0.045 0.523 ±0.028 0.505 ±0.041 0.276 ±0.020 0.296 ±0.018 0.153 ±0.007 0.155 ±0.007 Qwen2.5-7B-Instruct 0.788 ±0.018 0.789 ±0.017 0.661 ±0.026 0.711 ±0.014 0.408 ±0.012 0.411 ±0.010 0.368 ±0.009 0.361 ±0.009 0.814 ±0.011 0.804 ±0.013 0.838 ±0.007 0.855 ±0.003 0.758 ±0.012 0.648 ±0.016 0.252 ±0.030 0.277 ±0.013 0.282 ±0.004 0.299 ±0.013 0.619 ±0.033 0.640 ±0.022 Random 3-shot Llama-3.1-8B-Instruct 0.654 ±0.020 0.678 ±0.017 0.780 ±0.018 0.778 ±0.019 0.466 ±0.016 0.461 ±0.019 0.434 ±0.009 0.428 ±0.011 0.732 ±0.024 0.692 ±0.029 0.701 ±0.034 0.655 ±0.039 0.544 ±0.086 0.412 ±0.078 0.285 ±0.030 0.268 ±0.021 0.270 ±0.015 0.331 ±0.028 0.179 ±0.005 0.180 ±0.005 Mistral-Nemo-Instruct-2407 0.733 ±0.019 0.740 ±0.016 0.678 ±0.062 0.684 ±0.056 0.415 ±0.008 0.404 ±0.009 0.375 ±0.019 0.374 ±0.015 0.633 ±0.018 0.581 ±0.019 0.613 ±0.040 0.563 ±0.039 0.789 ±0.035 0.693 ±0.050 0.523 ±0.031 0.507 ±0.043 0.303 ±0.014 0.309 ±0.017 0.156 ±0.005 0.157 ±0.005 Qwen2.5-7B-Instruct 0.777 ±0.014 0.780 ±0.013 0.679 ±0.010 0.720 ±0.007 0.417 ±0.012 0.421 ±0.005 0.370 ±0.006 0.359 ±0.005 0.809 ±0.013 0.799 ±0.018 0.846 ±0.007 0.852 ±0.004 0.748 ±0.025 0.634 ±0.033 0.245 ±0.022 0.271 ±0.010 0.308 ±0.009 0.323 ±0.013 0.652 ±0.040 0.663 ±0.032 Random 4-shot Llama-3.1-8B-Instruct 0.623 ±0.037 0.656 ±0.028 0.762 ±0.035 0.760 ±0.035 0.463 ±0.020 0.456 ±0.022 0.425 ±0.010 0.421 ±0.007 0.739 ±0.020 0.700 ±0.024 0.688 ±0.024 0.640 ±0.027 0.509 ±0.090 0.380 ±0.081 0.275 ±0.035 0.264 ±0.023 0.234 ±0.018 0.267 ±0.047 0.180 ±0.012 0.183 ±0.013 Mistral-Nemo-Instruct-2407 0.741 ±0.025 0.747 ±0.020 0.614 ±0.064 0.631 ±0.052 0.417 ±0.005 0.404 ±0.007 0.396 ±0.017 0.392 ±0.013 0.610 ±0.017 0.558 ±0.017 0.642 ±0.043 0.593 ±0.044 0.790 ±0.033 0.694 ±0.049 0.500 ±0.043 0.475 ±0.059 0.290 ±0.014 0.291 ±0.014 0.168 ±0.003 0.172 ±0.004 Qwen2.5-7B-Instruct 0.779 ±0.009 0.780 ±0.009 0.681 ±0.006 0.721 ±0.002 0.418 ±0.007 0.421 ±0.003 0.384 ±0.010 0.372 ±0.010 0.784 ±0.028 0.763 ±0.039 0.846 ±0.009 0.851 ±0.005 0.749 ±0.023 0.636 ±0.030 0.257 ±0.017 0.276 ±0.008 0.322 ±0.013 0.335 ±0.013 0.656 ±0.026 0.667 ±0.020 Random 5-shot Llama-3.1-8B-Instruct 0.650 ±0.011 0.676 ±0.008 0.769 ±0.027 0.767 ±0.027 0.467 ±0.013 0.460 ±0.016 0.435 ±0.011 0.432 ±0.011 0.727 ±0.016 0.686 ±0.019 0.672 ±0.022 0.623 ±0.024 0.587 ±0.071 0.455 ±0.076 0.270 ±0.024 0.261 ±0.016 0.260 ±0.025 0.327 ±0.051 0.187 ±0.007 0.192 ±0.008 Mistral-Nemo-Instruct-2407 0.761 ±0.009 0.762 ±0.009 0.676 ±0.060 0.682 ±0.053 0.425 ±0.010 0.411 ±0.011 0.397 ±0.014 0.392 ±0.013 0.611 ±0.022 0.559 ±0.021 0.653 ±0.030 0.604 ±0.032 0.783 ±0.037 0.684 ±0.052 0.512 ±0.030 0.490 ±0.042 0.309 ±0.027 0.314 ±0.031 0.166 ±0.005 0.169 ±0.006 Qwen2.5-7B-Instruct 0.786 ±0.014 0.788 ±0.013 0.686 ±0.017 0.723 ±0.011 0.413 ±0.010 0.417 ±0.006 0.379 ±0.014 0.369 ±0.016 0.802 ±0.023 0.786 ±0.032 0.845 ±0.004 0.842 ±0.008 0.747 ±0.029 0.634 ±0.038 0.242 ±0.015 0.270 ±0.006 0.332 ±0.018 0.339 ±0.019 0.679 ±0.029 0.684 ±0.023 Random 6-shot Llama-3.1-8B-Instruct 0.619 ±0.071 0.657 ±0.047 0.785 ±0.025 0.783 ±0.026 0.476 ±0.014 0.471 ±0.016 0.441 ±0.010 0.438 ±0.010 0.730 ±0.014 0.689 ±0.017 0.627 ±0.039 0.576 ±0.040 0.567 ±0.074 0.434 ±0.076 0.266 ±0.028 0.259 ±0.017 0.265 ±0.027 0.327 ±0.059 0.198 ±0.010 0.202 ±0.011 Mistral-Nemo-Instruct-2407 0.754 ±0.022 0.757 ±0.020 0.645 ±0.049 0.655 ±0.043 0.419 ±0.009 0.408 ±0.012 0.384 ±0.022 0.382 ±0.017 0.600 ±0.029 0.548 ±0.029 0.667 ±0.033 0.618 ±0.036 0.780 ±0.035 0.679 ±0.050 0.521 ±0.028 0.500 ±0.038 0.302 ±0.030 0.305 ±0.032 0.167 ±0.006 0.170 ±0.007 Qwen2.5-7B-Instruct 0.786 ±0.010 0.786 ±0.009 0.692 ±0.013 0.727 ±0.007 0.416 ±0.010 0.421 ±0.004 0.379 ±0.009 0.369 ±0.009 0.794 ±0.023 0.775 ±0.034 0.846 ±0.007 0.843 ±0.009 0.743 ±0.037 0.630 ±0.048 0.247 ±0.026 0.273 ±0.011 0.341 ±0.014 0.348 ±0.011 0.658 ±0.023 0.667 ±0.017 Random 7-shot Llama-3.1-8B-Instruct 0.622 ±0.097 0.661 ±0.066 0.794 ±0.020 0.792 ±0.020 0.464 ±0.015 0.457 ±0.018 0.434 ±0.018 0.432 ±0.017 0.716 ±0.013 0.672 ±0.015 0.643 ±0.023 0.592 ±0.024 0.610 ±0.078 0.479 ±0.082 0.261 ±0.024 0.257 ±0.013 0.268 ±0.017 0.346 ±0.048 0.194 ±0.012 0.199 ±0.013 Mistral-Nemo-Instruct-2407 0.784 ±0.014 0.785 ±0.014 0.670 ±0.051 0.676 ±0.047 0.419 ±0.010 0.407 ±0.011 0.384 ±0.023 0.383 ±0.017 0.601 ±0.029 0.549 ±0.028 0.664 ±0.045 0.616 ±0.048 0.788 ±0.035 0.691 ±0.051 0.532 ±0.016 0.516 ±0.022 0.336 ±0.017 0.345 ±0.023 0.164 ±0.004 0.168 ±0.004 Qwen2.5-7B-Instruct 0.789 ±0.009 0.790 ±0.009 0.683 ±0.033 0.724 ±0.018 0.411 ±0.016 0.417 ±0.008 0.383 ±0.009 0.373 ±0.010 0.803 ±0.015 0.788 ±0.023 0.844 ±0.006 0.838 ±0.010 0.742 ±0.030 0.627 ±0.039 0.214 ±0.019 0.259 ±0.008 0.342 ±0.013 0.344 ±0.013 0.659 ±0.017 0.667 ±0.013 Random 8-shot Llama-3.1-8B-Instruct 0.590 ±0.089 0.639 ±0.053 0.793 ±0.026 0.791 ±0.026 0.466 ±0.011 0.459 ±0.014 0.435 ±0.015 0.433 ±0.013 0.703 ±0.016 0.657 ±0.018 0.616 ±0.043 0.565 ±0.044 0.531 ±0.101 0.402 ±0.098 0.246 ±0.025 0.248 ±0.014 0.258 ±0.022 0.321 ±0.050 0.225 ±0.022 0.230 ±0.026 Mistral-Nemo-Instruct-2407 0.770 ±0.014 0.771 ±0.013 0.653 ±0.049 0.661 ±0.043 0.414 ±0.011 0.400 ±0.012 0.392 ±0.017 0.389 ±0.014 0.608 ±0.030 0.556 ±0.030 0.661 ±0.028 0.612 ±0.030 0.787 ±0.041 0.690 ±0.059 0.522 ±0.014 0.500 ±0.019 0.342 ±0.022 0.355 ±0.025 0.177 ±0.006 0.182 ±0.006 Qwen2.5-7B-Instruct 0.780 ±0.009 0.780 ±0.009 0.685 ±0.034 0.724 ±0.022 0.412 ±0.013 0.416 ±0.007 0.375 ±0.004 0.364 ±0.004 0.795 ±0.014 0.775 ±0.018 0.844 ±0.005 0.837 ±0.009 0.748 ±0.030 0.635 ±0.040 0.208 ±0.031 0.259 ±0.013 0.364 ±0.032 0.366 ±0.032 0.646 ±0.029 0.656 ±0.024 Random 9-shot Llama-3.1-8B-Instruct 0.631 ±0.060 0.665 ±0.041 0.811 ±0.013 0.809 ±0.013 0.468 ±0.017 0.462 ±0.022 0.440 ±0.021 0.439 ±0.021 0.711 ±0.022 0.667 ±0.025 0.645 ±0.034 0.595 ±0.036 0.578 ±0.092 0.447 ±0.097 0.235 ±0.027 0.242 ±0.014 0.290 ±0.014 0.384 ±0.019 0.219 ±0.016 0.229 ±0.024 Mistral-Nemo-Instruct-2407 0.780 ±0.030 0.781 ±0.028 0.690 ±0.052 0.694 ±0.046 0.413 ±0.010 0.400 ±0.011 0.388 ±0.017 0.385 ±0.014 0.607 ±0.025 0.555 ±0.024 0.671 ±0.037 0.623 ±0.040 0.796 ±0.047 0.704 ±0.067 0.523 ±0.009 0.501 ±0.013 0.326 ±0.029 0.341 ±0.026 0.170 ±0.008 0.174 ±0.008 Qwen2.5-7B-Instruct 0.780 ±0.014 0.781 ±0.013 0.696 ±0.021 0.732 ±0.014 0.417 ±0.009 0.416 ±0.009 0.384 ±0.008 0.373 ±0.008 0.794 ±0.011 0.773 ±0.016 0.846 ±0.008 0.840 ±0.014 0.743 ±0.040 0.630 ±0.051 0.208 ±0.016 0.256 ±0.008 0.359 ±0.026 0.362 ±0.027 0.659 ±0.031 0.668 ±0.025 Random 10-shot Llama-3.1-8B-Instruct 0.587 ±0.045 0.635 ±0.029 0.798 ±0.019 0.796 ±0.019 0.462 ±0.013 0.457 ±0.017 0.432 ±0.019 0.431 ±0.019 0.694 ±0.019 0.647 ±0.021 0.612 ±0.036 0.561 ±0.037 0.582 ±0.100 0.453 ±0.103 0.243 ±0.026 0.245 ±0.012 0.285 ±0.026 0.367 ±0.048 0.254 ±0.012 0.273 ±0.019 Mistral-Nemo-Instruct-2407 0.771 ±0.016 0.772 ±0.015 0.691 ±0.053 0.695 ±0.049 0.414 ±0.012 0.401 ±0.012 0.391 ±0.019 0.388 ±0.016 0.615 ±0.028 0.563 ±0.028 0.679 ±0.027 0.631 ±0.030 0.790 ±0.060 0.697 ±0.084 0.500 ±0.007 0.470 ±0.009 0.352 ±0.045 0.364 ±0.050 0.183 ±0.005 0.190 ±0.008 Qwen2.5-7B-Instruct 0.783 ±0.010 0.783 ±0.010 0.704 ±0.024 0.736 ±0.016 0.425 ±0.008 0.421 ±0.007 0.384 ±0.014 0.373 ±0.014 0.786 ±0.011 0.760 ±0.016 0.847 ±0.006 0.841 ±0.010 0.750 ±0.048 0.640 ±0.063 0.208 ±0.018 0.256 ±0.008 0.361 ±0.014 0.362 ±0.013 0.236 ±0.020 0.237 ±0.021 DPP 2-shot Llama-3.1-8B-Instruct 0.685 ±0.063 0.704 ±0.048 0.787 ±0.010 0.784 ±0.010 - - - - 0.749 ±0.010 0.712 ±0.012 0.695 ±0.053 0.650 ±0.057 0.409 ±0.054 0.292 ±0.042 0.274 ±0.025 0.261 ±0.017 0.276 ±0.032 0.309 ±0.035 0.209 ±0.024 0.226 ±0.041 Mistral-Nemo-Instruct-2407 0.625 ±0.039 0.660 ±0.027 0.626 ±0.046 0.640 ±0.038 - - - - 0.664 ±0.030 0.614 ±0.032 0.608 ±0.093 0.562 ±0.090 0.785 ±0.022 0.686 ±0.031 0.524 ±0.056 0.511 ±0.074 0.264 ±0.006 0.276 ±0.017 0.155 ±0.009 0.158 ±0.011 Qwen2.5-7B-Instruct 0.782 ±0.010 0.784 ±0.009 0.633 ±0.019 0.694 ±0.010 - - - - 0.818 ±0.010 0.814 ±0.017 0.836 ±0.007 0.848 ±0.010 0.746 ±0.019 0.633 ±0.025 0.205 ±0.039 0.258 ±0.012 0.278 ±0.017 0.298 ±0.011 0.205 ±0.020 0.213 ±0.025 DPP 3-shot Llama-3.1-8B-Instruct - - - - 0.460 ±0.011 0.453 ±0.013 0.422 ±0.010 0.416 ±0.011 - - - - - - - - - - - - Mistral-Nemo-Instruct-2407 - - - - 0.422 ±0.008 0.417 ±0.010 0.394 ±0.008 0.386 ±0.006 - - - - - - - - - - - - Qwen2.5-7B-Instruct - - - - 0.401 ±0.003 0.407 ±0.012 0.359 ±0.012 0.350 ±0.010 - - - - - - - - - - - - DPP 4-shot Llama-3.1-8B-Instruct 0.708 ±0.056 0.721 ±0.043 0.771 ±0.021 0.768 ±0.021 - - - - 0.730 ±0.018 0.690 ±0.022 0.669 ±0.022 0.619 ±0.024 0.415 ±0.079 0.298 ±0.061 0.262 ±0.032 0.253 ±0.021 0.246 ±0.019 0.264 ±0.025 0.219 ±0.044 0.238 ±0.075 Mistral-Nemo-Instruct-2407 0.713 ±0.021 0.724 ±0.015 0.563 ±0.031 0.590 ±0.022 - - - - 0.649 ±0.013 0.598 ±0.014 0.687 ±0.012 0.639 ±0.013 0.781 ±0.041 0.681 ±0.057 0.516 ±0.018 0.496 ±0.026 0.274 ±0.019 0.279 ±0.019 0.160 ±0.008 0.164 ±0.008 Qwen2.5-7B-Instruct 0.781 ±0.005 0.782 ±0.006 0.640 ±0.013 0.695 ±0.008 - - - - 0.806 ±0.013 0.793 ±0.019 0.840 ±0.014 0.841 ±0.023 0.749 ±0.025 0.636 ±0.033 0.220 ±0.029 0.259 ±0.010 0.330 ±0.024 0.343 ±0.017 0.224 ±0.014 0.231 ±0.017 DPP 6-shot Llama-3.1-8B-Instruct 0.665 ±0.047 0.686 ±0.032 0.793 ±0.023 0.791 ±0.023 0.461 ±0.022 0.455 ±0.021 0.433 ±0.013 0.429 ±0.017 0.728 ±0.029 0.688 ±0.034 0.658 ±0.035 0.608 ±0.037 0.466 ±0.041 0.337 ±0.035 0.230 ±0.022 0.235 ±0.013 0.277 ±0.031 0.291 ±0.024 0.219 ±0.017 0.226 ±0.025 Mistral-Nemo-Instruct-2407 0.682 ±0.026 0.698 ±0.021 0.530 ±0.041 0.566 ±0.028 0.410 ±0.010 0.410 ±0.015 0.390 ±0.012 0.381 ±0.011 0.630 ±0.013 0.578 ±0.015 0.689 ±0.030 0.643 ±0.032 0.747 ±0.052 0.636 ±0.070 0.510 ±0.039 0.490 ±0.056 0.302 ±0.029 0.304 ±0.028 0.162 ±0.007 0.170 ±0.007 Qwen2.5-7B-Instruct 0.779 ±0.010 0.780 ±0.010 0.632 ±0.013 0.691 ±0.008 0.406 ±0.006 0.413 ±0.009 0.354 ±0.009 0.345 ±0.008 0.795 ±0.022 0.777 ±0.033 0.843 ±0.011 0.840 ±0.017 0.724 ±0.024 0.604 ±0.030 0.190 ±0.028 0.249 ±0.009 0.338 ±0.028 0.352 ±0.024 0.235 ±0.003 0.242 ±0.006 DPP 8-shot Llama-3.1-8B-Instruct 0.692 ±0.013 0.703 ±0.011 0.787 ±0.011 0.786 ±0.012 - - - - 0.707 ±0.033 0.663 ±0.036 0.619 ±0.034 0.568 ±0.035 0.498 ±0.057 0.367 ±0.052 0.216 ±0.027 0.227 ±0.012 0.292 ±0.016 0.303 ±0.020 0.257 ±0.043 0.287 ±0.079 Mistral-Nemo-Instruct-2407 0.702 ±0.029 0.710 ±0.025 0.536 ±0.047 0.571 ±0.033 - - - - 0.626 ±0.012 0.575 ±0.013 0.721 ±0.023 0.678 ±0.026 0.823 ±0.033 0.742 ±0.050 0.531 ±0.041 0.521 ±0.063 0.345 ±0.039 0.352 ±0.041 0.177 ±0.007 0.187 ±0.006 Qwen2.5-7B-Instruct 0.767 ±0.012 0.767 ±0.012 0.627 ±0.012 0.688 ±0.007 - - - - 0.793 ±0.014 0.772 ±0.021 0.847 ±0.014 0.841 ±0.020 0.715 ±0.035 0.593 ±0.045 0.170 ±0.025 0.243 ±0.008 0.349 ±0.019 0.362 ±0.028 0.234 ±0.015 0.236 ±0.016 DPP 9-shot Llama-3.1-8B-Instruct - - - - 0.459 ±0.017 0.454 ±0.019 0.433 ±0.014 0.428 ±0.014 - - - - - - - - - - - - Mistral-Nemo-Instruct-2407 - - - - 0.405 ±0.006 0.407 ±0.009 0.383 ±0.011 0.376 ±0.011 - - - - - - - - - - - - Qwen2.5-7B-Instruct - - - - 0.412 ±0.007 0.416 ±0.009 0.362 ±0.013 0.353 ±0.011 - - - - - - - - - - - - DPP 10-shot Llama-3.1-8B-Instruct 0.696 ±0.017 0.708 ±0.013 0.801 ±0.020 0.800 ±0.021 - - - - 0.681 ±0.042 0.634 ±0.044 0.653 ±0.036 0.603 ±0.037 0.523 ±0.060 0.390 ±0.059 0.202 ±0.016 0.220 ±0.005 0.269 ±0.029 0.280 ±0.034 0.247 ±0.028 0.267 ±0.042 Mistral-Nemo-Instruct-2407 0.687 ±0.028 0.694 ±0.020 0.542 ±0.064 0.576 ±0.046 - - - - 0.643 ±0.017 0.592 ±0.018 0.753 ±0.022 0.715 ±0.026 0.851 ±0.041 0.788 ±0.066 0.519 ±0.036 0.502 ±0.052 0.332 ±0.046 0.346 ±0.058 0.178 ±0.008 0.186 ±0.008 Qwen2.5-7B-Instruct 0.759 ±0.011 0.760 ±0.011 0.627 ±0.021 0.688 ±0.012 - - - - 0.794 ±0.017 0.774 ±0.024 0.846 ±0.015 0.841 ±0.022 0.724 ±0.036 0.604 ±0.044 0.159 ±0.023 0.239 ±0.005 0.356 ±0.041 0.368 ±0.050 0.229 ±0.018 0.230 ±0.019 CoT CoT Llama-3.1-8B-Instruct 0.792 0.795 0.757 0.764 0.465 0.459 0.416 0.405 0.557 0.502 0.503 0.454 0.226 0.160 0.377 0.340 0.280 0.358 0.161 0.161 Mistral-Nemo-Instruct-2407 0.713 0.728 0.732 0.729 0.464 0.465 0.401 0.386 0.252 0.275 0.584 0.530 0.528 0.394 0.623 0.680 0.315 0.397 0.167 0.168 Qwen2.5-7B-Instruct 0.711 0.725 0.714 0.729 0.422 0.428 0.355 0.349 0.575 0.522 0.804 0.783 0.646 0.511 0.491 0.456 0.277 0.340 0.147 0.149 Table 8: Comparison of F1 and accuracy across all models, configurations, and datasets. Each cell shows the mean F1 and accuracy (±one standard deviation for Few-Shot only, since Zero-Shot and CoT use greedy decoding and have no standard deviation). Boldface = best overall; underlined = best within each section.
Dataset Reference Model Performance Our Best Model Configuration Performance Acc F1 Acc F1 VIStA-H (Lyu et al. 2023) BERT-base 0.84 0.78 RoBERTa-large FT .852 .850 Fake News Net (Jin et al. 2022) Graph-based Reasoning 0.870 0.892 Llama3.1-8b FT .945 .945 Spanish Fake News Corpus (G´ omez-Adorno et al. 2021) BERT 0.766 N/A Modern-BERT-large FT .863 .863 Fake.br (Monteiro et al. 2018) SVM 0.89 0.89 Llama3-8b-Instruct FT .979 .979 Table 9: SOTA results.