Enhancing Automated Interpretability with Output-Centric Feature Descriptions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gur-Arieh, Yoav, Mayan, Roy, Agassy, Chen, Geiger, Atticus, Geva, Mor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
Constructing Interpretable Features from Compositional Neuron Groups
von: Shafran, Or, et al.
Veröffentlicht: (2025)
von: Shafran, Or, et al.
Veröffentlicht: (2025)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2026)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2026)
Disentangling MLP Neuron Weights in Vocabulary Space
von: Avrahamy, Asaf, et al.
Veröffentlicht: (2026)
von: Avrahamy, Asaf, et al.
Veröffentlicht: (2026)
Precise In-Parameter Concept Erasure in Large Language Models
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
von: Huang, Jing, et al.
Veröffentlicht: (2024)
von: Huang, Jing, et al.
Veröffentlicht: (2024)
LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
von: Gottesman, Daniela, et al.
Veröffentlicht: (2025)
von: Gottesman, Daniela, et al.
Veröffentlicht: (2025)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
von: Shafran, Or, et al.
Veröffentlicht: (2026)
von: Shafran, Or, et al.
Veröffentlicht: (2026)
Eliciting Textual Descriptions from Representations of Continuous Prompts
von: Ramati, Dana, et al.
Veröffentlicht: (2024)
von: Ramati, Dana, et al.
Veröffentlicht: (2024)
Estimating Knowledge in Large Language Models Without Generating a Single Token
von: Gottesman, Daniela, et al.
Veröffentlicht: (2024)
von: Gottesman, Daniela, et al.
Veröffentlicht: (2024)
Inferring Functionality of Attention Heads from their Parameters
von: Elhelo, Amit, et al.
Veröffentlicht: (2024)
von: Elhelo, Amit, et al.
Veröffentlicht: (2024)
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2023)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP
von: Mosbach, Marius, et al.
Veröffentlicht: (2024)
von: Mosbach, Marius, et al.
Veröffentlicht: (2024)
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)
von: Zur, Amir, et al.
Veröffentlicht: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
von: Yona, Gal, et al.
Veröffentlicht: (2024)
von: Yona, Gal, et al.
Veröffentlicht: (2024)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
von: Yona, Gal, et al.
Veröffentlicht: (2024)
von: Yona, Gal, et al.
Veröffentlicht: (2024)
Hallucinations Undermine Trust; Metacognition is a Way Forward
von: Yona, Gal, et al.
Veröffentlicht: (2026)
von: Yona, Gal, et al.
Veröffentlicht: (2026)
Preventing Rogue Agents Improves Multi-Agent Collaboration
von: Barbi, Ohav, et al.
Veröffentlicht: (2025)
von: Barbi, Ohav, et al.
Veröffentlicht: (2025)
Detecting (Un)answerability in Large Language Models with Linear Directions
von: Lavi, Maor Juliet, et al.
Veröffentlicht: (2025)
von: Lavi, Maor Juliet, et al.
Veröffentlicht: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
von: Ahrac, Sagi, et al.
Veröffentlicht: (2026)
von: Ahrac, Sagi, et al.
Veröffentlicht: (2026)
How Causal Abstraction Underpins Computational Explanation
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
How Do Transformers Learn Variable Binding in Symbolic Programs?
von: Wu, Yiwei, et al.
Veröffentlicht: (2025)
von: Wu, Yiwei, et al.
Veröffentlicht: (2025)
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
von: Cohen, Ido, et al.
Veröffentlicht: (2024)
von: Cohen, Ido, et al.
Veröffentlicht: (2024)
Rethinking Selective Knowledge Distillation
von: Tavor, Almog, et al.
Veröffentlicht: (2026)
von: Tavor, Almog, et al.
Veröffentlicht: (2026)
Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex
von: Grosbard, Idan Daniel, et al.
Veröffentlicht: (2026)
von: Grosbard, Idan Daniel, et al.
Veröffentlicht: (2026)
Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
von: Yalon, Noam Steinmetz, et al.
Veröffentlicht: (2026)
von: Yalon, Noam Steinmetz, et al.
Veröffentlicht: (2026)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
von: Din, Alexander Yom, et al.
Veröffentlicht: (2023)
von: Din, Alexander Yom, et al.
Veröffentlicht: (2023)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
von: Yona, Itay, et al.
Veröffentlicht: (2026)
von: Yona, Itay, et al.
Veröffentlicht: (2026)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
von: Ivgi, Maor, et al.
Veröffentlicht: (2024)
von: Ivgi, Maor, et al.
Veröffentlicht: (2024)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
von: Zur, Amir, et al.
Veröffentlicht: (2025)
von: Zur, Amir, et al.
Veröffentlicht: (2025)
MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
von: Wolfson, Tomer, et al.
Veröffentlicht: (2025)
von: Wolfson, Tomer, et al.
Veröffentlicht: (2025)
Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?
von: Yang, Sohee, et al.
Veröffentlicht: (2024)
von: Yang, Sohee, et al.
Veröffentlicht: (2024)
Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries
von: Biran, Eden, et al.
Veröffentlicht: (2024)
von: Biran, Eden, et al.
Veröffentlicht: (2024)
Do Large Language Models Latently Perform Multi-Hop Reasoning?
von: Yang, Sohee, et al.
Veröffentlicht: (2024)
von: Yang, Sohee, et al.
Veröffentlicht: (2024)
Activation Steering via Generative Causal Mediation
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
von: Puyin, Li, et al.
Veröffentlicht: (2026)
von: Puyin, Li, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025) -
Constructing Interpretable Features from Compositional Neuron Groups
von: Shafran, Or, et al.
Veröffentlicht: (2025) -
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2026) -
Disentangling MLP Neuron Weights in Vocabulary Space
von: Avrahamy, Asaf, et al.
Veröffentlicht: (2026) -
Precise In-Parameter Concept Erasure in Large Language Models
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)