Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hanna, Michael, Pezzelle, Sandro, Belinkov, Yonatan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Are formal and functional linguistic mechanisms dissociated in language models?
par: Hanna, Michael, et autres
Publié: (2025)
par: Hanna, Michael, et autres
Publié: (2025)
Position-aware Automatic Circuit Discovery
par: Haklay, Tal, et autres
Publié: (2025)
par: Haklay, Tal, et autres
Publié: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
par: Ashuach, Tomer, et autres
Publié: (2025)
par: Ashuach, Tomer, et autres
Publié: (2025)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
par: Nikankin, Yaniv, et autres
Publié: (2025)
par: Nikankin, Yaniv, et autres
Publié: (2025)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
par: Ashuach, Tomer, et autres
Publié: (2024)
par: Ashuach, Tomer, et autres
Publié: (2024)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
par: Orgad, Hadas, et autres
Publié: (2026)
par: Orgad, Hadas, et autres
Publié: (2026)
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
par: Höth, Max Henning, et autres
Publié: (2026)
par: Höth, Max Henning, et autres
Publié: (2026)
Do Models Know Why They Changed Their Mind? Interpretability and Faithfulness of Chain-of-Thought Under Knowledge Conflict
par: Venkata, Pruthvinath Jeripity
Publié: (2026)
par: Venkata, Pruthvinath Jeripity
Publié: (2026)
Distinguishing Ignorance from Error in LLM Hallucinations
par: Simhi, Adi, et autres
Publié: (2024)
par: Simhi, Adi, et autres
Publié: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
par: Simhi, Adi, et autres
Publié: (2024)
par: Simhi, Adi, et autres
Publié: (2024)
Aligning Large Language Models for Faithful Integrity Against Opposing Argument
par: Zhao, Yong, et autres
Publié: (2025)
par: Zhao, Yong, et autres
Publié: (2025)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
par: Toker, Michael, et autres
Publié: (2024)
par: Toker, Michael, et autres
Publié: (2024)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
par: Parcalabescu, Letitia, et autres
Publié: (2023)
par: Parcalabescu, Letitia, et autres
Publié: (2023)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
par: Simhi, Adi, et autres
Publié: (2025)
par: Simhi, Adi, et autres
Publié: (2025)
Where Should LoRA Go? Component-Type Placement in Hybrid Language Models
par: Borobia, Hector, et autres
Publié: (2026)
par: Borobia, Hector, et autres
Publié: (2026)
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
par: Ashuach, Tomer, et autres
Publié: (2026)
par: Ashuach, Tomer, et autres
Publié: (2026)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
par: Nikankin, Yaniv, et autres
Publié: (2024)
par: Nikankin, Yaniv, et autres
Publié: (2024)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
par: Simhi, Adi, et autres
Publié: (2025)
par: Simhi, Adi, et autres
Publié: (2025)
Scaling Laws for Forgetting When Fine-Tuning Large Language Models
par: Kalajdzievski, Damjan
Publié: (2024)
par: Kalajdzievski, Damjan
Publié: (2024)
Reasoning Models Know What's Important, and Encode It in Their Activations
par: Nikankin, Yaniv, et autres
Publié: (2026)
par: Nikankin, Yaniv, et autres
Publié: (2026)
When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
par: Young, Richard J., et autres
Publié: (2025)
par: Young, Richard J., et autres
Publié: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
par: Oketunji, Abiodun Finbarrs
Publié: (2023)
par: Oketunji, Abiodun Finbarrs
Publié: (2023)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
par: He, Yanjin, et autres
Publié: (2025)
par: He, Yanjin, et autres
Publié: (2025)
Correctness is not Faithfulness in RAG Attributions
par: Wallat, Jonas, et autres
Publié: (2024)
par: Wallat, Jonas, et autres
Publié: (2024)
Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
par: Gu, Hao, et autres
Publié: (2025)
par: Gu, Hao, et autres
Publié: (2025)
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
par: Siegel, Noah Y., et autres
Publié: (2025)
par: Siegel, Noah Y., et autres
Publié: (2025)
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
par: Szilvasy, Gergely, et autres
Publié: (2026)
par: Szilvasy, Gergely, et autres
Publié: (2026)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
par: Arad, Dana, et autres
Publié: (2023)
par: Arad, Dana, et autres
Publié: (2023)
Engineering A Large Language Model From Scratch
par: Oketunji, Abiodun Finbarrs
Publié: (2024)
par: Oketunji, Abiodun Finbarrs
Publié: (2024)
Large Language Model (LLM) Bias Index -- LLMBI
par: Oketunji, Abiodun Finbarrs, et autres
Publié: (2023)
par: Oketunji, Abiodun Finbarrs, et autres
Publié: (2023)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
par: Balasubramanian, Sriram, et autres
Publié: (2025)
par: Balasubramanian, Sriram, et autres
Publié: (2025)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
par: Simhi, Adi, et autres
Publié: (2026)
par: Simhi, Adi, et autres
Publié: (2026)
HACK: Hallucinations Along Certainty and Knowledge Axes
par: Simhi, Adi, et autres
Publié: (2025)
par: Simhi, Adi, et autres
Publié: (2025)
Language Models Are Implicitly Continuous
par: Marro, Samuele, et autres
Publié: (2025)
par: Marro, Samuele, et autres
Publié: (2025)
Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
par: Bianchessi, Arthur S., et autres
Publié: (2025)
par: Bianchessi, Arthur S., et autres
Publié: (2025)
Ensemble Language Models for Multilingual Sentiment Analysis
par: Hasan, Md Arid
Publié: (2024)
par: Hasan, Md Arid
Publié: (2024)
On the Effect of (Near) Duplicate Subwords in Language Modelling
par: Schäfer, Anton, et autres
Publié: (2024)
par: Schäfer, Anton, et autres
Publié: (2024)
RBCorr: Response Bias Correction in Language Models
par: Bhatt, Om, et autres
Publié: (2026)
par: Bhatt, Om, et autres
Publié: (2026)
Compact Example-Based Explanations for Language Models
par: Schoenegger, Loris, et autres
Publié: (2026)
par: Schoenegger, Loris, et autres
Publié: (2026)
The Open Source Advantage in Large Language Models (LLMs)
par: Manchanda, Jiya, et autres
Publié: (2024)
par: Manchanda, Jiya, et autres
Publié: (2024)
Documents similaires
-
Are formal and functional linguistic mechanisms dissociated in language models?
par: Hanna, Michael, et autres
Publié: (2025) -
Position-aware Automatic Circuit Discovery
par: Haklay, Tal, et autres
Publié: (2025) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
par: Ashuach, Tomer, et autres
Publié: (2025) -
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
par: Nikankin, Yaniv, et autres
Publié: (2025) -
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
par: Ashuach, Tomer, et autres
Publié: (2024)