Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Jianhui, Luo, Yuzhang, Pan, Liangming |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Beyond Behavioural Trade-Offs: Mechanistic Tracing of Pain-Pleasure Decisions in an LLM
par: Bianco, Francesca, et autres
Publié: (2026)
par: Bianco, Francesca, et autres
Publié: (2026)
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
par: Yang, Minglai, et autres
Publié: (2025)
par: Yang, Minglai, et autres
Publié: (2025)
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
par: Gao, Yutong, et autres
Publié: (2026)
par: Gao, Yutong, et autres
Publié: (2026)
Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation
par: Li, Ruizhe, et autres
Publié: (2025)
par: Li, Ruizhe, et autres
Publié: (2025)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
par: Chen, Jianhui, et autres
Publié: (2024)
par: Chen, Jianhui, et autres
Publié: (2024)
MIB: A Mechanistic Interpretability Benchmark
par: Mueller, Aaron, et autres
Publié: (2025)
par: Mueller, Aaron, et autres
Publié: (2025)
Investigating the Transferability of Code Repair for Low-Resource Programming Languages
par: Wong, Kyle, et autres
Publié: (2024)
par: Wong, Kyle, et autres
Publié: (2024)
Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
par: Li, Zhe, et autres
Publié: (2025)
par: Li, Zhe, et autres
Publié: (2025)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
par: Xu, Bingxin, et autres
Publié: (2025)
par: Xu, Bingxin, et autres
Publié: (2025)
MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
par: Deng, Xinle, et autres
Publié: (2026)
par: Deng, Xinle, et autres
Publié: (2026)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
par: Sun, Jiuding, et autres
Publié: (2025)
par: Sun, Jiuding, et autres
Publié: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
par: Mishra, Anurag
Publié: (2025)
par: Mishra, Anurag
Publié: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
par: Cho, Hakaze, et autres
Publié: (2025)
par: Cho, Hakaze, et autres
Publié: (2025)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
par: Méloux, Maxime, et autres
Publié: (2025)
par: Méloux, Maxime, et autres
Publié: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
par: Méloux, Maxime, et autres
Publié: (2025)
par: Méloux, Maxime, et autres
Publié: (2025)
LLM4DistReconfig: A Fine-tuned Large Language Model for Power Distribution Network Reconfiguration
par: Christou, Panayiotis, et autres
Publié: (2025)
par: Christou, Panayiotis, et autres
Publié: (2025)
Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation
par: Wang, Xinyi, et autres
Publié: (2024)
par: Wang, Xinyi, et autres
Publié: (2024)
Towards a Mechanistic Understanding of Propositional Logical Reasoning in Large Language Models
par: Chen, Danchun, et autres
Publié: (2026)
par: Chen, Danchun, et autres
Publié: (2026)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
par: Song, Xiangchen, et autres
Publié: (2025)
par: Song, Xiangchen, et autres
Publié: (2025)
Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
par: Liang, Jia, et autres
Publié: (2026)
par: Liang, Jia, et autres
Publié: (2026)
A Training-free Method for LLM Text Attribution
par: Radvand, Tara, et autres
Publié: (2025)
par: Radvand, Tara, et autres
Publié: (2025)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
par: Kim, Geonhee, et autres
Publié: (2024)
par: Kim, Geonhee, et autres
Publié: (2024)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
par: Zhou, Sifan, et autres
Publié: (2025)
par: Zhou, Sifan, et autres
Publié: (2025)
ControlMath: Controllable Data Generation Promotes Math Generalist Models
par: Chen, Nuo, et autres
Publié: (2024)
par: Chen, Nuo, et autres
Publié: (2024)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
par: Wang, Xu, et autres
Publié: (2026)
par: Wang, Xu, et autres
Publié: (2026)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
par: Lee, Yu-Ting, et autres
Publié: (2025)
par: Lee, Yu-Ting, et autres
Publié: (2025)
DETAIL: Task DEmonsTration Attribution for Interpretable In-context Learning
par: Zhou, Zijian, et autres
Publié: (2024)
par: Zhou, Zijian, et autres
Publié: (2024)
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
par: Ayonrinde, Kola, et autres
Publié: (2025)
par: Ayonrinde, Kola, et autres
Publié: (2025)
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
par: Tsujimura, Hikaru, et autres
Publié: (2025)
par: Tsujimura, Hikaru, et autres
Publié: (2025)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
par: Bai, Xiaoyan, et autres
Publié: (2026)
par: Bai, Xiaoyan, et autres
Publié: (2026)
Mechanistic Fine-tuning for In-context Learning
par: Cho, Hakaze, et autres
Publié: (2025)
par: Cho, Hakaze, et autres
Publié: (2025)
LLM-Powered Ensemble Learning for Paper Source Tracing: A GPU-Free Approach
par: Chen, Kunlong, et autres
Publié: (2024)
par: Chen, Kunlong, et autres
Publié: (2024)
Fast Training Dataset Attribution via In-Context Learning
par: Fotouhi, Milad, et autres
Publié: (2024)
par: Fotouhi, Milad, et autres
Publié: (2024)
LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines
par: Gao, Jiechao, et autres
Publié: (2026)
par: Gao, Jiechao, et autres
Publié: (2026)
Beyond Accuracy: Introducing a Symbolic-Mechanistic Approach to Interpretable Evaluation
par: Habibi, Reza, et autres
Publié: (2026)
par: Habibi, Reza, et autres
Publié: (2026)
Mechanistic?
par: Saphra, Naomi, et autres
Publié: (2024)
par: Saphra, Naomi, et autres
Publié: (2024)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
par: Du, Hongzhe, et autres
Publié: (2025)
par: Du, Hongzhe, et autres
Publié: (2025)
Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information
par: Huang, Xin, et autres
Publié: (2026)
par: Huang, Xin, et autres
Publié: (2026)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
par: Zhang, Kangning, et autres
Publié: (2025)
par: Zhang, Kangning, et autres
Publié: (2025)
Muon is Scalable for LLM Training
par: Liu, Jingyuan, et autres
Publié: (2025)
par: Liu, Jingyuan, et autres
Publié: (2025)
Documents similaires
-
Beyond Behavioural Trade-Offs: Mechanistic Tracing of Pain-Pleasure Decisions in an LLM
par: Bianco, Francesca, et autres
Publié: (2026) -
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
par: Yang, Minglai, et autres
Publié: (2025) -
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
par: Gao, Yutong, et autres
Publié: (2026) -
Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation
par: Li, Ruizhe, et autres
Publié: (2025) -
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
par: Chen, Jianhui, et autres
Publié: (2024)