scieee AI-readable full text Open interactive document viewer

(Poster) The Trilemma of Truth in Large Language Models @ Mech Interp Workshop (NeurIPS 2025)

Savcisens, Germans; Eliassi-Rad, Tina

Abstract

This poster is presented at the Mechanistic Interpretability Workshop as part of the NeurIPS 2025.

Full text

Mechanistic Interpretability Workshop, NeurIPS 2025 Paper Data Code DOI: 10.5281/zenodo.17725915 Table 1: Data sets used in this work. The first column specifies the domain. The last column displays example statements with ground truth labels. Data Set Examples City Locations (True) The city of Mˆacon is located in France. (False) The city of Dhar¯an is located in Ecuador. (Neither) The city of Staakess is located in Marbate. Word Definitions (True) Corsage is a synonym of a nosegay. (False) Towner is not a type of a resident. (Neither) Kharter is not a synonym of a greging Medical Indications (True) PR-104 is indicated for the treatment of tumors. (False) Zolpidem is indicated for the treatment of angina. (Neither) Alostat is indicated for the treatment of candigemia. We introduce three novel benchmarks covering City Locations, Medical Indications, and Word Definitions. Each dataset comprises statements labeled as true, false, or neither. Neither-valued statements contain synthetic entities, intentionally constructed so that the LLM has no prior knowledge about them. To create these synthetic entities, we use Markov-chain generation, which replicates linguistic structures. DATA SETS MOTIVATION How can we reliably probe what an LLM “knows” about a statement and quantify its certainty? LLMs generate remarkably coherent text, yet sometimes produce factually incorrect content. For example, when a model says “Agadir is located in Morocco,” we need to know: 1. Does the model actually encode that fact? 2. With what level of confidence does it make that claim? Zero-shot Prompt Question: Is the following statement correct? The city of Agadir is located in Morocco. Select one of the following options: 1. The statement is correct. 2. The statement is incorrect. 3. I do not have sufficient knowledge. 4. The statement is too ambiguous. 5. All of the above. 6. None of the above. Please respond with the corresponding number. The final answer is LARGE LANGUAGE MODEL Input to the model Output of the model (next token prediction) LARGE LANGUAGE MODEL Representation-based Probe Input to the model DECODER I DECODER II The city of Agadir is located in Morocco. DECODER N 1 2 3 4 5 6 Probability of a token P(True) P(False) P(Neither) Representation of the last token P(True) P(False) P(Neither) Truthfulness Probe Falsehood Probe Neither Probe Representations of all tokens (bag) A B B1 B2 Veracity Probe P(True) P(False) Mean-difference probe Multiclass sAwMIL probe (ours) Used to sanity check the model [PERIOD] The city of Agadir is located in [PERIOD] Morocco aan zit probabilities of all other tokens Figure 1: Overview of methods for probing veracity in LLMs. (A) Zero-shot prompting. (B) Probes that use internal activations. (B1) Linear classifier that accounts only for True and False cases (B2) Our (multiclass) multipleinstance classifier considers all the tokens in the statement. Germans Savcisens and Tina Eliassi-Rad The Trilemma of Truth in Large Language Models LLMs encode truth, falsehood, and neither as separate signals, rather than a single true–false axis. ` TL;DR Common probing methods fail to provide a reliable and transferable veracity direction and, in some settings, perform worse than zero-shot prompting • Truth and falsehood are not encoded symmetrically in LLMs; they appear as related but distinct representations. • LLMs encode a third signal that is different from both true and false, corresponding to neither-valued statements • We introduce novel representation-based probe (sAwMIL) that reliably detects these three signals and generalizes across models and domains. MD with CP Zero-shot with CP TTPD with CP sPCA with CP SVM sAwMIL Multiclass Evaluation Setting Last Token Bag/Sentence Mean MCC (with Standard Error) Mean Performance (Correlation Criterion) A B Mean Performance (Generalization Criterion) Mean MCC (with Standard Error) Multiclass MD with CP TTPD with CP sPCA with CP SVM sAwMIL Evaluation Setting Last Token Bag/Sentence MD+CP (Word Definitions) Qwen-2.5-14b-Instruct (Decoder 32) Multiclass sAwMIL (Word Definitions) Qwen-2.5-14b-Instruct (Decoder 24) True False Neither Predicted Labels TrueFalseNeither True Labels 0.88 0.09 0.03 0.06 0.92 0.02 0.38 0.59 0.03 CDE True False Neither Abstained Predicted Labels TrueFalseNeither True Labels 0.55 0.32 0.13 0.00 0.06 0.89 0.06 0.00 0.01 0.37 0.62 0.00 Zero-shot Prompt (Word Definitions) Qwen-2.5-14b-Instruct True False Neither Abstained Predicted Labels TrueFalseNeither True Labels 0.86 0.07 0.01 0.05 0.06 0.86 0.00 0.08 0.00 0.01 0.98 0.01 Fraction of samples 0.0 0.2 0.4 0.6 0.8 1.0 METHOD AND EVALUTION We use 16 open-source LLMs (3–14B parameters), including default and chat-tuned variants. Our method builds on multiple-instance learning and SVMs, analyzing all token embeddings to pinpoint veracity signals, while conformal prediction provides calibrated uncertainty. Figure 2: Mean performance across 16 LLMs and 3 datasets, evaluated under two settings: using only the last token vs. the full sentence (all token embeddings). Results are shown for both (A) correlation (in-domain) and (B) generalization (cross-domain). Binary probes (trained on last token representation) perform poorly and fail to generalize, while sAwMIL remains robust in both settings. EXAMPLES Score Intepretation: 0: Token contains falsehood signal 1: Token contains truthful signal Score Intepretation: 0: Token does not contain truthful signal 1: Token contains truthful signal