Can Interpretation Predict Behavior on Unseen Data?
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Victoria R., Kaufmann, Jenny, Wattenberg, Martin, Alvarez-Melis, David, Saphra, Naomi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization
by: Qin, Tian, et al.
Published: (2024)
by: Qin, Tian, et al.
Published: (2024)
Mechanistic?
by: Saphra, Naomi, et al.
Published: (2024)
by: Saphra, Naomi, et al.
Published: (2024)
When Bad Data Leads to Good Models
by: Li, Kenneth, et al.
Published: (2025)
by: Li, Kenneth, et al.
Published: (2025)
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
by: Li, Victoria R., et al.
Published: (2024)
by: Li, Victoria R., et al.
Published: (2024)
Dialogue Action Tokens: Steering Language Models in Goal-Directed Dialogue with a Multi-Turn Planner
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
by: Li, Kenneth, et al.
Published: (2023)
by: Li, Kenneth, et al.
Published: (2023)
Language Models Can Predict Their Own Behavior
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
by: Li, Kenneth, et al.
Published: (2022)
by: Li, Kenneth, et al.
Published: (2022)
Using Shapley interactions to understand how models use structure
by: Singhvi, Divyansh, et al.
Published: (2024)
by: Singhvi, Divyansh, et al.
Published: (2024)
Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
by: Pham, Bao, et al.
Published: (2026)
by: Pham, Bao, et al.
Published: (2026)
Does visualization help AI understand data?
by: Li, Victoria R., et al.
Published: (2025)
by: Li, Victoria R., et al.
Published: (2025)
Adapting Language Models via Token Translation
by: Feng, Zhili, et al.
Published: (2024)
by: Feng, Zhili, et al.
Published: (2024)
Tag-LLM: Repurposing General-Purpose LLMs for Specialized Domains
by: Shen, Junhong, et al.
Published: (2024)
by: Shen, Junhong, et al.
Published: (2024)
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
by: Goffinet, Etienne, et al.
Published: (2025)
by: Goffinet, Etienne, et al.
Published: (2025)
Instruction Diversity Drives Generalization To Unseen Tasks
by: Zhang, Dylan, et al.
Published: (2024)
by: Zhang, Dylan, et al.
Published: (2024)
MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses
by: Yang, Zonglin, et al.
Published: (2024)
by: Yang, Zonglin, et al.
Published: (2024)
Learning to Generalize Unseen Domains via Multi-Source Meta Learning for Text Classification
by: Hu, Yuxuan, et al.
Published: (2024)
by: Hu, Yuxuan, et al.
Published: (2024)
ICLR: In-Context Learning of Representations
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
Think Before You Lie: How Reasoning Leads to Honesty
by: Yuan, Ann, et al.
Published: (2026)
by: Yuan, Ann, et al.
Published: (2026)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
by: Gu, Xinran, et al.
Published: (2025)
by: Gu, Xinran, et al.
Published: (2025)
Towards Reducing Diagnostic Errors with Interpretable Risk Prediction
by: McInerney, Denis Jered, et al.
Published: (2024)
by: McInerney, Denis Jered, et al.
Published: (2024)
Large Language Models for Travel Behavior Prediction
by: Mo, Baichuan, et al.
Published: (2023)
by: Mo, Baichuan, et al.
Published: (2023)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Does Pre-trained Language Model Actually Infer Unseen Links in Knowledge Graph Completion?
by: Sakai, Yusuke, et al.
Published: (2023)
by: Sakai, Yusuke, et al.
Published: (2023)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
by: Huang, Vincent, et al.
Published: (2025)
by: Huang, Vincent, et al.
Published: (2025)
Interpretable Predictability-Based AI Text Detection: A Replication Study
by: Skurla, Adam, et al.
Published: (2026)
by: Skurla, Adam, et al.
Published: (2026)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
by: Nainani, Jatin, et al.
Published: (2024)
by: Nainani, Jatin, et al.
Published: (2024)
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022)
by: Lopardo, Gianluigi, et al.
Published: (2022)
From Data to Behavior: Predicting Unintended Model Behaviors Before Training
by: Wang, Mengru, et al.
Published: (2026)
by: Wang, Mengru, et al.
Published: (2026)
How Much Can We Forget about Data Contamination?
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
by: Chen, Jianhui, et al.
Published: (2026)
by: Chen, Jianhui, et al.
Published: (2026)
Can LLMs Follow Simple Rules?
by: Mu, Norman, et al.
Published: (2023)
by: Mu, Norman, et al.
Published: (2023)
Inner Speech as Behavior Guides: Steerable Imitation of Diverse Behaviors for Human-AI coordination
by: Trivedi, Rakshit, et al.
Published: (2026)
by: Trivedi, Rakshit, et al.
Published: (2026)
Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs
by: Pepper, Keenan, et al.
Published: (2026)
by: Pepper, Keenan, et al.
Published: (2026)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
Data Scaling as Progressive Coverage of a Predictive Contribution Spectrum
by: Song, Zihui, et al.
Published: (2026)
by: Song, Zihui, et al.
Published: (2026)
You Can Generate It Again: Data-to-Text Generation with Verification and Correction Prompting
by: Ren, Xuan, et al.
Published: (2023)
by: Ren, Xuan, et al.
Published: (2023)
Similar Items
-
Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization
by: Qin, Tian, et al.
Published: (2024) -
Mechanistic?
by: Saphra, Naomi, et al.
Published: (2024) -
When Bad Data Leads to Good Models
by: Li, Kenneth, et al.
Published: (2025) -
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
by: Li, Kenneth, et al.
Published: (2024) -
ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
by: Li, Victoria R., et al.
Published: (2024)