Interpretability Illusions in the Generalization of Simplified Models
Fuente:
arXiv
Saved in:
| Main Authors: | Friedman, Dan, Lampinen, Andrew, Dixon, Lucas, Chen, Danqi, Ghandeharioun, Asma |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models
by: Bhaskar, Adithya, et al.
Published: (2024)
by: Bhaskar, Adithya, et al.
Published: (2024)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
by: Ghandeharioun, Asma, et al.
Published: (2024)
by: Ghandeharioun, Asma, et al.
Published: (2024)
Representing Rule-based Chatbots with Transformers
by: Friedman, Dan, et al.
Published: (2024)
by: Friedman, Dan, et al.
Published: (2024)
Extracting Rule-based Descriptions of Attention Features in Transformers
by: Friedman, Dan, et al.
Published: (2025)
by: Friedman, Dan, et al.
Published: (2025)
Think Before You Lie: How Reasoning Leads to Honesty
by: Yuan, Ann, et al.
Published: (2026)
by: Yuan, Ann, et al.
Published: (2026)
When Can Transformers Count to n?
by: Yehudai, Gilad, et al.
Published: (2024)
by: Yehudai, Gilad, et al.
Published: (2024)
How to Train Long-Context Language Models (Effectively)
by: Gao, Tianyu, et al.
Published: (2024)
by: Gao, Tianyu, et al.
Published: (2024)
QuRating: Selecting High-Quality Data for Training Language Models
by: Wettig, Alexander, et al.
Published: (2024)
by: Wettig, Alexander, et al.
Published: (2024)
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
by: Zhong, Zexuan, et al.
Published: (2024)
by: Zhong, Zexuan, et al.
Published: (2024)
Towards Unifying Interpretability and Control: Evaluation via Intervention
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
SimPO: Simple Preference Optimization with a Reference-Free Reward
by: Meng, Yu, et al.
Published: (2024)
by: Meng, Yu, et al.
Published: (2024)
Linear representations in language models can change dramatically over a conversation
by: Lampinen, Andrew Kyle, et al.
Published: (2026)
by: Lampinen, Andrew Kyle, et al.
Published: (2026)
The broader spectrum of in-context learning
by: Lampinen, Andrew Kyle, et al.
Published: (2024)
by: Lampinen, Andrew Kyle, et al.
Published: (2024)
How do language models learn facts? Dynamics, curricula and hallucinations
by: Zucchet, Nicolas, et al.
Published: (2025)
by: Zucchet, Nicolas, et al.
Published: (2025)
The Illusion of Stochasticity in LLMs
by: Gu, Xiangming, et al.
Published: (2026)
by: Gu, Xiangming, et al.
Published: (2026)
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
by: Chen, Howard, et al.
Published: (2025)
by: Chen, Howard, et al.
Published: (2025)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
Latent learning: episodic memory complements parametric learning by enabling flexible reuse of experiences
by: Lampinen, Andrew Kyle, et al.
Published: (2025)
by: Lampinen, Andrew Kyle, et al.
Published: (2025)
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
The Illusion of State in State-Space Models
by: Merrill, William, et al.
Published: (2024)
by: Merrill, William, et al.
Published: (2024)
Who's asking? User personas and the mechanics of latent misalignment
by: Ghandeharioun, Asma, et al.
Published: (2024)
by: Ghandeharioun, Asma, et al.
Published: (2024)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
by: Lee, Jin Hwa, et al.
Published: (2025)
by: Lee, Jin Hwa, et al.
Published: (2025)
Evaluating Large Language Models at Evaluating Instruction Following
by: Zeng, Zhiyuan, et al.
Published: (2023)
by: Zeng, Zhiyuan, et al.
Published: (2023)
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
by: Xia, Mengzhou, et al.
Published: (2023)
by: Xia, Mengzhou, et al.
Published: (2023)
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
by: Rizvi-Martel, Michael, et al.
Published: (2026)
by: Rizvi-Martel, Michael, et al.
Published: (2026)
Racing Thoughts: Explaining Contextualization Errors in Large Language Models
by: Lepori, Michael A., et al.
Published: (2024)
by: Lepori, Michael A., et al.
Published: (2024)
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)
by: Malladi, Sadhika, et al.
Published: (2023)
The Leaderboard Illusion
by: Singh, Shivalika, et al.
Published: (2025)
by: Singh, Shivalika, et al.
Published: (2025)
The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
by: Tomov, Tim, et al.
Published: (2025)
by: Tomov, Tim, et al.
Published: (2025)
LoRA vs Full Fine-tuning: An Illusion of Equivalence
by: Shuttleworth, Reece, et al.
Published: (2024)
by: Shuttleworth, Reece, et al.
Published: (2024)
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
by: Zhu, Xinyu, et al.
Published: (2025)
by: Zhu, Xinyu, et al.
Published: (2025)
The Illusion of Readiness in Health AI
by: Gu, Yu, et al.
Published: (2025)
by: Gu, Yu, et al.
Published: (2025)
Continual Memorization of Factoids in Language Models
by: Chen, Howard, et al.
Published: (2024)
by: Chen, Howard, et al.
Published: (2024)
Are UFOs Driving Innovation? The Illusion of Causality in Large Language Models
by: Carro, María Victoria, et al.
Published: (2024)
by: Carro, María Victoria, et al.
Published: (2024)
Putting It All into Context: Simplifying Agents with LCLMs
by: Jiang, Mingjian, et al.
Published: (2025)
by: Jiang, Mingjian, et al.
Published: (2025)
Finding Transformer Circuits with Edge Pruning
by: Bhaskar, Adithya, et al.
Published: (2024)
by: Bhaskar, Adithya, et al.
Published: (2024)
Improving Language Understanding from Screenshots
by: Gao, Tianyu, et al.
Published: (2024)
by: Gao, Tianyu, et al.
Published: (2024)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
Simplifying Outcomes of Language Model Component Analyses with ELIA
by: Eidt, Aaron Louis, et al.
Published: (2026)
by: Eidt, Aaron Louis, et al.
Published: (2026)
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Similar Items
-
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models
by: Bhaskar, Adithya, et al.
Published: (2024) -
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
by: Ghandeharioun, Asma, et al.
Published: (2024) -
Representing Rule-based Chatbots with Transformers
by: Friedman, Dan, et al.
Published: (2024) -
Extracting Rule-based Descriptions of Attention Features in Transformers
by: Friedman, Dan, et al.
Published: (2025) -
Think Before You Lie: How Reasoning Leads to Honesty
by: Yuan, Ann, et al.
Published: (2026)