How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Im, Shawn, Oh, Changdae, Fang, Zhen, Li, Sharon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach
by: Oh, Changdae, et al.
Published: (2025)
by: Oh, Changdae, et al.
Published: (2025)
General Exploratory Bonus for Optimistic Exploration in RLHF
by: Li, Wendi, et al.
Published: (2025)
by: Li, Wendi, et al.
Published: (2025)
A Unified Understanding and Evaluation of Steering Methods
by: Im, Shawn, et al.
Published: (2025)
by: Im, Shawn, et al.
Published: (2025)
Mechanistic Interpretability of Binary and Ternary Transformers
by: Li, Jason
Published: (2024)
by: Li, Jason
Published: (2024)
How Well Can Preference Optimization Generalize Under Noisy Feedback?
by: Im, Shawn, et al.
Published: (2025)
by: Im, Shawn, et al.
Published: (2025)
Can DPO Learn Diverse Human Values? A Theoretical Scaling Law
by: Im, Shawn, et al.
Published: (2024)
by: Im, Shawn, et al.
Published: (2024)
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023)
by: Shen, Lingfeng, et al.
Published: (2023)
Triangulation as an Acceptance Rule for Multilingual Mechanistic Interpretability
by: Long, Yanan
Published: (2025)
by: Long, Yanan
Published: (2025)
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
by: Sun, Alan, et al.
Published: (2026)
by: Sun, Alan, et al.
Published: (2026)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
by: Sung, Yi-Lin, et al.
Published: (2025)
by: Sung, Yi-Lin, et al.
Published: (2025)
How Do Transformers Learn Variable Binding in Symbolic Programs?
by: Wu, Yiwei, et al.
Published: (2025)
by: Wu, Yiwei, et al.
Published: (2025)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
by: Shin, Seungjun, et al.
Published: (2025)
by: Shin, Seungjun, et al.
Published: (2025)
I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
Thinking Makes LLM Agents Introverted: How Mandatory Thinking Can Backfire in User-Engaged Agents
by: Li, Jiatong, et al.
Published: (2026)
by: Li, Jiatong, et al.
Published: (2026)
MIB: A Mechanistic Interpretability Benchmark
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning
by: Li, Wendi, et al.
Published: (2026)
by: Li, Wendi, et al.
Published: (2026)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
Enhancing Latent Computation in Transformers with Latent Tokens
by: Sun, Yuchang, et al.
Published: (2025)
by: Sun, Yuchang, et al.
Published: (2025)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
by: Yan, Lecheng, et al.
Published: (2026)
by: Yan, Lecheng, et al.
Published: (2026)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
by: Mihaila, George
Published: (2026)
by: Mihaila, George
Published: (2026)
Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning
by: Hellwig, Philipp, et al.
Published: (2026)
by: Hellwig, Philipp, et al.
Published: (2026)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
by: Mishra, Anurag
Published: (2025)
by: Mishra, Anurag
Published: (2025)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
by: Yang, Yushi, et al.
Published: (2024)
by: Yang, Yushi, et al.
Published: (2024)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
by: García-Carrasco, Jorge, et al.
Published: (2024)
by: García-Carrasco, Jorge, et al.
Published: (2024)
Beyond Transcription: Mechanistic Interpretability in ASR
by: Glazer, Neta, et al.
Published: (2025)
by: Glazer, Neta, et al.
Published: (2025)
On the Duality between Gradient Transformations and Adapters
by: Torroba-Hennigen, Lucas, et al.
Published: (2025)
by: Torroba-Hennigen, Lucas, et al.
Published: (2025)
Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
by: Long, Lin, et al.
Published: (2025)
by: Long, Lin, et al.
Published: (2025)
Visual Instruction Bottleneck Tuning
by: Oh, Changdae, et al.
Published: (2025)
by: Oh, Changdae, et al.
Published: (2025)
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling
by: Huang, Hongzhi, et al.
Published: (2025)
by: Huang, Hongzhi, et al.
Published: (2025)
BRIDO: Bringing Democratic Order to Abstractive Summarization
by: Lee, Junhyun, et al.
Published: (2025)
by: Lee, Junhyun, et al.
Published: (2025)
Do Sentence Transformers Learn Quasi-Geospatial Concepts from General Text?
by: Ilyankou, Ilya, et al.
Published: (2024)
by: Ilyankou, Ilya, et al.
Published: (2024)
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
by: Chen, Jianhui, et al.
Published: (2026)
by: Chen, Jianhui, et al.
Published: (2026)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024)
by: Kim, Geonhee, et al.
Published: (2024)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
by: Bai, Xiaoyan, et al.
Published: (2026)
by: Bai, Xiaoyan, et al.
Published: (2026)
Circuit Fingerprints: How Answer Tokens Encode Their Geometrical Path
by: Saurez, Andres, et al.
Published: (2026)
by: Saurez, Andres, et al.
Published: (2026)
Similar Items
-
Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach
by: Oh, Changdae, et al.
Published: (2025) -
General Exploratory Bonus for Optimistic Exploration in RLHF
by: Li, Wendi, et al.
Published: (2025) -
A Unified Understanding and Evaluation of Steering Methods
by: Im, Shawn, et al.
Published: (2025) -
Mechanistic Interpretability of Binary and Ternary Transformers
by: Li, Jason
Published: (2024) -
How Well Can Preference Optimization Generalize Under Noisy Feedback?
by: Im, Shawn, et al.
Published: (2025)