Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Venhoff, Constantin, Khakzar, Ashkan, Joseph, Sonia, Torr, Philip, Nanda, Neel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Visual Representations Map to Language Feature Space in Multimodal LLMs
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Understanding Reasoning in Thinking Language Models via Steering Vectors
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Base Models Know How to Reason, Thinking Models Learn When
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Towards Understanding Multimodal Fine-Tuning: Spatial Features
by: Naghashyar, Lachin, et al.
Published: (2026)
by: Naghashyar, Lachin, et al.
Published: (2026)
Reasoning-Finetuning Repurposes Latent Representations in Base Models
by: Ward, Jake, et al.
Published: (2025)
by: Ward, Jake, et al.
Published: (2025)
Mixture of Experts Made Intrinsically Interpretable
by: Yang, Xingyi, et al.
Published: (2025)
by: Yang, Xingyi, et al.
Published: (2025)
RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching
by: Jafari, Farnoush Rezaei, et al.
Published: (2025)
by: Jafari, Farnoush Rezaei, et al.
Published: (2025)
SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders
by: Venhoff, Constantin, et al.
Published: (2024)
by: Venhoff, Constantin, et al.
Published: (2024)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
by: Chughtai, Bilal, et al.
Published: (2024)
by: Chughtai, Bilal, et al.
Published: (2024)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
by: Lan, Michael, et al.
Published: (2024)
by: Lan, Michael, et al.
Published: (2024)
Latent Guard: a Safety Framework for Text-to-image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
AttributionLab: Faithfulness of Feature Attribution Under Controllable Environments
by: Zhang, Yang, et al.
Published: (2023)
by: Zhang, Yang, et al.
Published: (2023)
Minimalist Concept Erasure in Generative Models
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
How to use and interpret activation patching
by: Heimersheim, Stefan, et al.
Published: (2024)
by: Heimersheim, Stefan, et al.
Published: (2024)
The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms
by: Davies, Adam, et al.
Published: (2024)
by: Davies, Adam, et al.
Published: (2024)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
On Discprecncies between Perturbation Evaluations of Graph Neural Network Attributions
by: Rezaei, Razieh, et al.
Published: (2024)
by: Rezaei, Razieh, et al.
Published: (2024)
Explorations of Self-Repair in Language Models
by: Rushing, Cody, et al.
Published: (2024)
by: Rushing, Cody, et al.
Published: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
by: Zhang, Fred, et al.
Published: (2023)
by: Zhang, Fred, et al.
Published: (2023)
Difficulties with Evaluating a Deception Detector for AIs
by: Smith, Lewis, et al.
Published: (2025)
by: Smith, Lewis, et al.
Published: (2025)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
by: Makelov, Aleksandar, et al.
Published: (2024)
by: Makelov, Aleksandar, et al.
Published: (2024)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
by: Ferrando, Javier, et al.
Published: (2024)
by: Ferrando, Javier, et al.
Published: (2024)
A Dual-Perspective Approach to Evaluating Feature Attribution Methods
by: Li, Yawei, et al.
Published: (2023)
by: Li, Yawei, et al.
Published: (2023)
Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models
by: Leask, Patrick, et al.
Published: (2025)
by: Leask, Patrick, et al.
Published: (2025)
BatchTopK Sparse Autoencoders
by: Bussmann, Bart, et al.
Published: (2024)
by: Bussmann, Bart, et al.
Published: (2024)
Transcoders Find Interpretable LLM Feature Circuits
by: Dunefsky, Jacob, et al.
Published: (2024)
by: Dunefsky, Jacob, et al.
Published: (2024)
Interpreting Attention Layer Outputs with Sparse Autoencoders
by: Kissane, Connor, et al.
Published: (2024)
by: Kissane, Connor, et al.
Published: (2024)
Eliciting Secret Knowledge from Language Models
by: Cywiński, Bartosz, et al.
Published: (2025)
by: Cywiński, Bartosz, et al.
Published: (2025)
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
by: Cywiński, Bartosz, et al.
Published: (2025)
by: Cywiński, Bartosz, et al.
Published: (2025)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
by: Casademunt, Helena, et al.
Published: (2026)
by: Casademunt, Helena, et al.
Published: (2026)
Empowering Tuberculosis Screening with Explainable Self-Supervised Deep Neural Networks
by: Patel, Neel, et al.
Published: (2024)
by: Patel, Neel, et al.
Published: (2024)
FullRecall: A Semantic Search-Based Ranking Approach for Maximizing Recall in Patent Retrieval
by: Ali, Amna, et al.
Published: (2025)
by: Ali, Amna, et al.
Published: (2025)
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
by: Deb, Oishi, et al.
Published: (2025)
by: Deb, Oishi, et al.
Published: (2025)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
by: Bussmann, Bart, et al.
Published: (2025)
by: Bussmann, Bart, et al.
Published: (2025)
Convergent Linear Representations of Emergent Misalignment
by: Soligo, Anna, et al.
Published: (2025)
by: Soligo, Anna, et al.
Published: (2025)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
by: Kramár, János, et al.
Published: (2024)
by: Kramár, János, et al.
Published: (2024)
RADAR: Recall Augmentation through Deferred Asynchronous Retrieval
by: Jaspal, Amit, et al.
Published: (2025)
by: Jaspal, Amit, et al.
Published: (2025)
Retrieval-enhanced Knowledge Editing in Language Models for Multi-Hop Question Answering
by: Shi, Yucheng, et al.
Published: (2024)
by: Shi, Yucheng, et al.
Published: (2024)
A Two-Phase Recall-and-Select Framework for Fast Model Selection
by: Cui, Jianwei, et al.
Published: (2024)
by: Cui, Jianwei, et al.
Published: (2024)
Similar Items
-
How Visual Representations Map to Language Feature Space in Multimodal LLMs
by: Venhoff, Constantin, et al.
Published: (2025) -
Understanding Reasoning in Thinking Language Models via Steering Vectors
by: Venhoff, Constantin, et al.
Published: (2025) -
Base Models Know How to Reason, Thinking Models Learn When
by: Venhoff, Constantin, et al.
Published: (2025) -
Towards Understanding Multimodal Fine-Tuning: Spatial Features
by: Naghashyar, Lachin, et al.
Published: (2026) -
Reasoning-Finetuning Repurposes Latent Representations in Base Models
by: Ward, Jake, et al.
Published: (2025)