LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
Fuente:
arXiv
Saved in:
| Main Authors: | Haller, Patrick, Ibrahim, Mark, Kirichenko, Polina, Sagun, Levent, Bell, Samuel J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Effective Theory of Bias Amplification
by: Subramonian, Arjun, et al.
Published: (2024)
by: Subramonian, Arjun, et al.
Published: (2024)
Reassessing the Validity of Spurious Correlations Benchmarks
by: Bell, Samuel J., et al.
Published: (2024)
by: Bell, Samuel J., et al.
Published: (2024)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
by: Lavoie, Samuel, et al.
Published: (2024)
by: Lavoie, Samuel, et al.
Published: (2024)
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
by: Christoph, Daniel, et al.
Published: (2025)
by: Christoph, Daniel, et al.
Published: (2025)
TruthFlow: Truthful LLM Generation via Representation Flow Correction
by: Wang, Hanyu, et al.
Published: (2025)
by: Wang, Hanyu, et al.
Published: (2025)
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
by: Godbole, Ameya, et al.
Published: (2025)
by: Godbole, Ameya, et al.
Published: (2025)
Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers
by: Xu, Ruichen, et al.
Published: (2026)
by: Xu, Ruichen, et al.
Published: (2026)
What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
by: Ross, Candace, et al.
Published: (2025)
by: Ross, Candace, et al.
Published: (2025)
On the Role of Speech Data in Reducing Toxicity Detection Bias
by: Bell, Samuel J., et al.
Published: (2024)
by: Bell, Samuel J., et al.
Published: (2024)
Non-Linear Inference Time Intervention: Improving LLM Truthfulness
by: Hoscilowicz, Jakub, et al.
Published: (2024)
by: Hoscilowicz, Jakub, et al.
Published: (2024)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
by: Pangakis, Nicholas, et al.
Published: (2024)
by: Pangakis, Nicholas, et al.
Published: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
by: Fu, Yao, et al.
Published: (2025)
by: Fu, Yao, et al.
Published: (2025)
Chained Tuning Leads to Biased Forgetting
by: Ung, Megan, et al.
Published: (2024)
by: Ung, Megan, et al.
Published: (2024)
Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link Prediction
by: Subramonian, Arjun, et al.
Published: (2023)
by: Subramonian, Arjun, et al.
Published: (2023)
Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning
by: Damirchi, Hamed, et al.
Published: (2026)
by: Damirchi, Hamed, et al.
Published: (2026)
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
by: Jang, Yeonwoo, et al.
Published: (2025)
by: Jang, Yeonwoo, et al.
Published: (2025)
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Context Management
by: Cui, Guanyu, et al.
Published: (2026)
by: Cui, Guanyu, et al.
Published: (2026)
Revisiting the Superficial Alignment Hypothesis
by: Raghavendra, Mohit, et al.
Published: (2024)
by: Raghavendra, Mohit, et al.
Published: (2024)
Beyond Simple Averaging: Improving NLP Ensemble Performance with Topological-Data-Analysis-Based Weighting
by: Proskura, Polina, et al.
Published: (2024)
by: Proskura, Polina, et al.
Published: (2024)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
The Truth Lies Somewhere in the Middle (of the Generated Tokens)
by: Wang, Sophie L., et al.
Published: (2026)
by: Wang, Sophie L., et al.
Published: (2026)
Advancing Academic Knowledge Retrieval via LLM-enhanced Representation Similarity Fusion
by: Dai, Wei, et al.
Published: (2024)
by: Dai, Wei, et al.
Published: (2024)
Unintended Impacts of LLM Alignment on Global Representation
by: Ryan, Michael J., et al.
Published: (2024)
by: Ryan, Michael J., et al.
Published: (2024)
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
by: Wei, Zhepei, et al.
Published: (2025)
by: Wei, Zhepei, et al.
Published: (2025)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
by: Jaiswal, Ajay, et al.
Published: (2023)
by: Jaiswal, Ajay, et al.
Published: (2023)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
by: Du, Hongzhe, et al.
Published: (2025)
by: Du, Hongzhe, et al.
Published: (2025)
LLM Flow Processes for Text-Conditioned Regression
by: Biggs, Felix, et al.
Published: (2026)
by: Biggs, Felix, et al.
Published: (2026)
Bridge: A Unified Framework to Knowledge Graph Completion via Language Models and Knowledge Representation
by: Qiao, Qiao, et al.
Published: (2024)
by: Qiao, Qiao, et al.
Published: (2024)
Enhancing LLM Knowledge Learning through Generalization
by: Zhu, Mingkang, et al.
Published: (2025)
by: Zhu, Mingkang, et al.
Published: (2025)
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
by: Zhang, Shaolei, et al.
Published: (2024)
by: Zhang, Shaolei, et al.
Published: (2024)
Multimodal Contrastive Representation Learning in Augmented Biomedical Knowledge Graphs
by: Dang, Tien, et al.
Published: (2025)
by: Dang, Tien, et al.
Published: (2025)
Representation Consistency for Accurate and Coherent LLM Answer Aggregation
by: Jiang, Junqi, et al.
Published: (2025)
by: Jiang, Junqi, et al.
Published: (2025)
Advancing LLM Safe Alignment with Safety Representation Ranking
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
Improving LLM Final Representations with Inter-Layer Geometry
by: Ulanovski, Tom, et al.
Published: (2026)
by: Ulanovski, Tom, et al.
Published: (2026)
Understanding the Detrimental Class-level Effects of Data Augmentation
by: Kirichenko, Polina, et al.
Published: (2023)
by: Kirichenko, Polina, et al.
Published: (2023)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
Mitigating Heterogeneous Token Overfitting in LLM Knowledge Editing
by: Liu, Tianci, et al.
Published: (2025)
by: Liu, Tianci, et al.
Published: (2025)
Controlled LLM Decoding via Discrete Auto-regressive Biasing
by: Pynadath, Patrick, et al.
Published: (2025)
by: Pynadath, Patrick, et al.
Published: (2025)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
by: Lei, Ge, et al.
Published: (2025)
by: Lei, Ge, et al.
Published: (2025)
Similar Items
-
An Effective Theory of Bias Amplification
by: Subramonian, Arjun, et al.
Published: (2024) -
Reassessing the Validity of Spurious Correlations Benchmarks
by: Bell, Samuel J., et al.
Published: (2024) -
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
by: Lavoie, Samuel, et al.
Published: (2024) -
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
by: Christoph, Daniel, et al.
Published: (2025) -
TruthFlow: Truthful LLM Generation via Representation Flow Correction
by: Wang, Hanyu, et al.
Published: (2025)