Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Yibo, Rajendran, Goutham, Ravikumar, Pradeep, Aragam, Bryon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Origins of Linear Representations in Large Language Models
by: Jiang, Yibo, et al.
Published: (2024)
by: Jiang, Yibo, et al.
Published: (2024)
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
by: Rajendran, Goutham, et al.
Published: (2024)
by: Rajendran, Goutham, et al.
Published: (2024)
Markov Equivalence and Consistency in Differentiable Structure Learning
by: Deng, Chang, et al.
Published: (2024)
by: Deng, Chang, et al.
Published: (2024)
iSCAN: Identifying Causal Mechanism Shifts among Nonlinear Additive Noise Models
by: Chen, Tianyu, et al.
Published: (2023)
by: Chen, Tianyu, et al.
Published: (2023)
An Interventional Perspective on Identifiability in Gaussian LTI Systems with Independent Component Analysis
by: Rajendran, Goutham, et al.
Published: (2023)
by: Rajendran, Goutham, et al.
Published: (2023)
Identifying General Mechanism Shifts in Linear Causal Representations
by: Chen, Tianyu, et al.
Published: (2024)
by: Chen, Tianyu, et al.
Published: (2024)
Greedy equivalence search for nonparametric graphical models
by: Aragam, Bryon
Published: (2024)
by: Aragam, Bryon
Published: (2024)
Model-free Estimation of Latent Structure via Multiscale Nonparametric Maximum Likelihood
by: Aragam, Bryon, et al.
Published: (2024)
by: Aragam, Bryon, et al.
Published: (2024)
Differentiable Structure Learning and Causal Discovery for General Binary Data
by: Deng, Chang, et al.
Published: (2025)
by: Deng, Chang, et al.
Published: (2025)
LLM-Select: Feature Selection with Large Language Models
by: Jeong, Daniel P., et al.
Published: (2024)
by: Jeong, Daniel P., et al.
Published: (2024)
Towards Interpretable Deep Generative Models via Causal Representation Learning
by: Moran, Gemma E., et al.
Published: (2025)
by: Moran, Gemma E., et al.
Published: (2025)
Optimal structure learning and conditional independence testing
by: Gao, Ming, et al.
Published: (2025)
by: Gao, Ming, et al.
Published: (2025)
Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
by: Tsaknakis, Ioannis, et al.
Published: (2025)
by: Tsaknakis, Ioannis, et al.
Published: (2025)
Beyond identifiability: Learning causal representations with few environments and finite samples
by: Lee, Inbeom, et al.
Published: (2026)
by: Lee, Inbeom, et al.
Published: (2026)
Latent learning: episodic memory complements parametric learning by enabling flexible reuse of experiences
by: Lampinen, Andrew Kyle, et al.
Published: (2025)
by: Lampinen, Andrew Kyle, et al.
Published: (2025)
Learning general conditional independence structures via the neighbourhood lattice
by: Amini, Arash A., et al.
Published: (2022)
by: Amini, Arash A., et al.
Published: (2022)
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
by: Chang, Shuochen, et al.
Published: (2026)
by: Chang, Shuochen, et al.
Published: (2026)
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
by: Zhao, Siyan, et al.
Published: (2025)
by: Zhao, Siyan, et al.
Published: (2025)
Latent Logic Tree Extraction for Event Sequence Explanation from LLMs
by: Song, Zitao, et al.
Published: (2024)
by: Song, Zitao, et al.
Published: (2024)
Ambiguity in LLMs is a concept missing problem
by: Hu, Zhibo, et al.
Published: (2025)
by: Hu, Zhibo, et al.
Published: (2025)
Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying Probes
by: Maiya, Sharan, et al.
Published: (2025)
by: Maiya, Sharan, et al.
Published: (2025)
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
by: Wang, Yinjie, et al.
Published: (2025)
by: Wang, Yinjie, et al.
Published: (2025)
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
by: Pan, Alexander, et al.
Published: (2024)
by: Pan, Alexander, et al.
Published: (2024)
ATLAS: Adaptive Test-Time Latent Steering with External Verifiers for Enhancing LLMs Reasoning
by: Nguyen, Tuc, et al.
Published: (2026)
by: Nguyen, Tuc, et al.
Published: (2026)
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
by: Wei, Rongzhe, et al.
Published: (2025)
by: Wei, Rongzhe, et al.
Published: (2025)
TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
by: Nagaraj, Manish, et al.
Published: (2025)
by: Nagaraj, Manish, et al.
Published: (2025)
Do Multilingual LLMs Think In English?
by: Schut, Lisa, et al.
Published: (2025)
by: Schut, Lisa, et al.
Published: (2025)
GeoReasoner: Reasoning On Geospatially Grounded Context For Natural Language Understanding
by: Yan, Yibo, et al.
Published: (2024)
by: Yan, Yibo, et al.
Published: (2024)
SOMP: Scalable Gradient Inversion for Large Language Models via Subspace-Guided Orthogonal Matching Pursuit
by: Li, Yibo, et al.
Published: (2026)
by: Li, Yibo, et al.
Published: (2026)
Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Debiasing Multilingual LLMs in Cross-lingual Latent Space
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?
by: Zverev, Egor, et al.
Published: (2024)
by: Zverev, Egor, et al.
Published: (2024)
Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
by: Yuan, Yu, et al.
Published: (2024)
by: Yuan, Yu, et al.
Published: (2024)
Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification
by: AlMarri, Saeed, et al.
Published: (2025)
by: AlMarri, Saeed, et al.
Published: (2025)
Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
by: Fernandes, Patrick, et al.
Published: (2025)
by: Fernandes, Patrick, et al.
Published: (2025)
Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction
by: Wu, Qinyuan, et al.
Published: (2024)
by: Wu, Qinyuan, et al.
Published: (2024)
Thinking in Latents: Adaptive Anchor Refinement for Implicit Reasoning in LLMs
by: Sheshanarayana, Disha, et al.
Published: (2026)
by: Sheshanarayana, Disha, et al.
Published: (2026)
Enhancing Latent Computation in Transformers with Latent Tokens
by: Sun, Yuchang, et al.
Published: (2025)
by: Sun, Yuchang, et al.
Published: (2025)
Reinforcement Learning for Latent-Space Thinking in LLMs
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
On-Device Fine-Tuning via Backprop-Free Zeroth-Order Optimization
by: Katti, Prabodh, et al.
Published: (2025)
by: Katti, Prabodh, et al.
Published: (2025)
Similar Items
-
On the Origins of Linear Representations in Large Language Models
by: Jiang, Yibo, et al.
Published: (2024) -
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
by: Rajendran, Goutham, et al.
Published: (2024) -
Markov Equivalence and Consistency in Differentiable Structure Learning
by: Deng, Chang, et al.
Published: (2024) -
iSCAN: Identifying Causal Mechanism Shifts among Nonlinear Additive Noise Models
by: Chen, Tianyu, et al.
Published: (2023) -
An Interventional Perspective on Identifiability in Gaussian LTI Systems with Independent Component Analysis
by: Rajendran, Goutham, et al.
Published: (2023)