Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gur-Arieh, Yoav, Geva, Mor, Geiger, Atticus |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
par: Gur-Arieh, Yoav, et autres
Publié: (2026)
par: Gur-Arieh, Yoav, et autres
Publié: (2026)
Disentangling MLP Neuron Weights in Vocabulary Space
par: Avrahamy, Asaf, et autres
Publié: (2026)
par: Avrahamy, Asaf, et autres
Publié: (2026)
Precise In-Parameter Concept Erasure in Large Language Models
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
par: Gur-Arieh, Yoav, et autres
Publié: (2025)
Constructing Interpretable Features from Compositional Neuron Groups
par: Shafran, Or, et autres
Publié: (2025)
par: Shafran, Or, et autres
Publié: (2025)
LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
par: Gottesman, Daniela, et autres
Publié: (2025)
par: Gottesman, Daniela, et autres
Publié: (2025)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
par: Huang, Jing, et autres
Publié: (2024)
par: Huang, Jing, et autres
Publié: (2024)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
par: Shafran, Or, et autres
Publié: (2026)
par: Shafran, Or, et autres
Publié: (2026)
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
par: Cohen, Ido, et autres
Publié: (2024)
par: Cohen, Ido, et autres
Publié: (2024)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
par: Yona, Itay, et autres
Publié: (2026)
par: Yona, Itay, et autres
Publié: (2026)
Estimating Knowledge in Large Language Models Without Generating a Single Token
par: Gottesman, Daniela, et autres
Publié: (2024)
par: Gottesman, Daniela, et autres
Publié: (2024)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
par: Yona, Gal, et autres
Publié: (2024)
par: Yona, Gal, et autres
Publié: (2024)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
par: She, Jingyuan Selena, et autres
Publié: (2023)
par: She, Jingyuan Selena, et autres
Publié: (2023)
Detecting (Un)answerability in Large Language Models with Linear Directions
par: Lavi, Maor Juliet, et autres
Publié: (2025)
par: Lavi, Maor Juliet, et autres
Publié: (2025)
Inferring Functionality of Attention Heads from their Parameters
par: Elhelo, Amit, et autres
Publié: (2024)
par: Elhelo, Amit, et autres
Publié: (2024)
How Causal Abstraction Underpins Computational Explanation
par: Geiger, Atticus, et autres
Publié: (2025)
par: Geiger, Atticus, et autres
Publié: (2025)
How Do Transformers Learn Variable Binding in Symbolic Programs?
par: Wu, Yiwei, et autres
Publié: (2025)
par: Wu, Yiwei, et autres
Publié: (2025)
Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
par: Yalon, Noam Steinmetz, et autres
Publié: (2026)
par: Yalon, Noam Steinmetz, et autres
Publié: (2026)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
par: Ivgi, Maor, et autres
Publié: (2024)
par: Ivgi, Maor, et autres
Publié: (2024)
Do Large Language Models Latently Perform Multi-Hop Reasoning?
par: Yang, Sohee, et autres
Publié: (2024)
par: Yang, Sohee, et autres
Publié: (2024)
Eliciting Textual Descriptions from Representations of Continuous Prompts
par: Ramati, Dana, et autres
Publié: (2024)
par: Ramati, Dana, et autres
Publié: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
par: Yona, Gal, et autres
Publié: (2024)
par: Yona, Gal, et autres
Publié: (2024)
Hallucinations Undermine Trust; Metacognition is a Way Forward
par: Yona, Gal, et autres
Publié: (2026)
par: Yona, Gal, et autres
Publié: (2026)
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
par: Wu, Zhengxuan, et autres
Publié: (2023)
par: Wu, Zhengxuan, et autres
Publié: (2023)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
par: Katz, Shahar, et autres
Publié: (2024)
par: Katz, Shahar, et autres
Publié: (2024)
Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries
par: Biran, Eden, et autres
Publié: (2024)
par: Biran, Eden, et autres
Publié: (2024)
Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?
par: Yang, Sohee, et autres
Publié: (2024)
par: Yang, Sohee, et autres
Publié: (2024)
How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?
par: Yang, Sohee, et autres
Publié: (2025)
par: Yang, Sohee, et autres
Publié: (2025)
Preventing Rogue Agents Improves Multi-Agent Collaboration
par: Barbi, Ohav, et autres
Publié: (2025)
par: Barbi, Ohav, et autres
Publié: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
par: Ahrac, Sagi, et autres
Publié: (2026)
par: Ahrac, Sagi, et autres
Publié: (2026)
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
par: Gekhman, Zorik, et autres
Publié: (2026)
par: Gekhman, Zorik, et autres
Publié: (2026)
The Hidden Space of Transformer Language Adapters
par: Alabi, Jesujoba O., et autres
Publié: (2024)
par: Alabi, Jesujoba O., et autres
Publié: (2024)
NER Retriever: Zero-Shot Named Entity Retrieval with Type-Aware Embeddings
par: Shachar, Or, et autres
Publié: (2025)
par: Shachar, Or, et autres
Publié: (2025)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
par: Ghandeharioun, Asma, et autres
Publié: (2024)
par: Ghandeharioun, Asma, et autres
Publié: (2024)
How do Language Models Bind Entities in Context?
par: Feng, Jiahai, et autres
Publié: (2023)
par: Feng, Jiahai, et autres
Publié: (2023)
Rethinking Selective Knowledge Distillation
par: Tavor, Almog, et autres
Publié: (2026)
par: Tavor, Almog, et autres
Publié: (2026)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
par: Din, Alexander Yom, et autres
Publié: (2023)
par: Din, Alexander Yom, et autres
Publié: (2023)
Language Models use Lookbacks to Track Beliefs
par: Prakash, Nikhil, et autres
Publié: (2025)
par: Prakash, Nikhil, et autres
Publié: (2025)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
par: Zur, Amir, et autres
Publié: (2025)
par: Zur, Amir, et autres
Publié: (2025)
MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
par: Wolfson, Tomer, et autres
Publié: (2025)
par: Wolfson, Tomer, et autres
Publié: (2025)
Documents similaires
-
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
par: Gur-Arieh, Yoav, et autres
Publié: (2025) -
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
par: Gur-Arieh, Yoav, et autres
Publié: (2026) -
Disentangling MLP Neuron Weights in Vocabulary Space
par: Avrahamy, Asaf, et autres
Publié: (2026) -
Precise In-Parameter Concept Erasure in Large Language Models
par: Gur-Arieh, Yoav, et autres
Publié: (2025) -
Constructing Interpretable Features from Compositional Neuron Groups
par: Shafran, Or, et autres
Publié: (2025)