Towards Long-Horizon Interpretability: Efficient and Faithful Multi-Token Attribution for Reasoning LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Pan, Wenbo, Liu, Zhichao, Wang, Xianlong, Yu, Haining, Jia, Xiaohua |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
por: Pan, Wenbo, et al.
Publicado: (2025)
por: Pan, Wenbo, et al.
Publicado: (2025)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
por: Yang, Zhicheng, et al.
Publicado: (2026)
por: Yang, Zhicheng, et al.
Publicado: (2026)
Faithful or Just Plausible? Evaluating the Faithfulness of Closed-Source LLMs in Medical Reasoning
por: Afolabi, Halimat, et al.
Publicado: (2026)
por: Afolabi, Halimat, et al.
Publicado: (2026)
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
por: Li, Yingru, et al.
Publicado: (2026)
por: Li, Yingru, et al.
Publicado: (2026)
TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs
por: Zhang, Yuxiang, et al.
Publicado: (2025)
por: Zhang, Yuxiang, et al.
Publicado: (2025)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
por: Jia, Jinghan, et al.
Publicado: (2026)
por: Jia, Jinghan, et al.
Publicado: (2026)
AttributionLab: Faithfulness of Feature Attribution Under Controllable Environments
por: Zhang, Yang, et al.
Publicado: (2023)
por: Zhang, Yang, et al.
Publicado: (2023)
Faithful Interpretation for Graph Neural Networks
por: Hu, Lijie, et al.
Publicado: (2024)
por: Hu, Lijie, et al.
Publicado: (2024)
MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
por: Zhao, Guojiang, et al.
Publicado: (2025)
por: Zhao, Guojiang, et al.
Publicado: (2025)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
por: Wu, Wei, et al.
Publicado: (2024)
por: Wu, Wei, et al.
Publicado: (2024)
Faithful and Efficient Explanations for Neural Networks via Neural Tangent Kernel Surrogate Models
por: Engel, Andrew, et al.
Publicado: (2023)
por: Engel, Andrew, et al.
Publicado: (2023)
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
por: Hong, Ilgee, et al.
Publicado: (2025)
por: Hong, Ilgee, et al.
Publicado: (2025)
TracLLM: A Generic Framework for Attributing Long Context LLMs
por: Wang, Yanting, et al.
Publicado: (2025)
por: Wang, Yanting, et al.
Publicado: (2025)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
por: Liu, Hanbing, et al.
Publicado: (2025)
por: Liu, Hanbing, et al.
Publicado: (2025)
Recursive Models for Long-Horizon Reasoning
por: Yang, Chenxiao, et al.
Publicado: (2026)
por: Yang, Chenxiao, et al.
Publicado: (2026)
LEGO: A Lightweight and Efficient Multiple-Attribute Unlearning Framework for Recommender Systems
por: Yu, Fengyuan, et al.
Publicado: (2025)
por: Yu, Fengyuan, et al.
Publicado: (2025)
Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL
por: Wu, Ian, et al.
Publicado: (2026)
por: Wu, Ian, et al.
Publicado: (2026)
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
por: Liu, Gabrielle Kaili-May, et al.
Publicado: (2025)
por: Liu, Gabrielle Kaili-May, et al.
Publicado: (2025)
WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation
por: Liu, Zhichao, et al.
Publicado: (2026)
por: Liu, Zhichao, et al.
Publicado: (2026)
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
por: Motwani, Sumeet Ramesh, et al.
Publicado: (2026)
por: Motwani, Sumeet Ramesh, et al.
Publicado: (2026)
Improving Interpretation Faithfulness for Vision Transformers
por: Hu, Lijie, et al.
Publicado: (2023)
por: Hu, Lijie, et al.
Publicado: (2023)
Ideal Attribution and Faithful Watermarks for Language Models
por: Song, Min Jae, et al.
Publicado: (2025)
por: Song, Min Jae, et al.
Publicado: (2025)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
por: Wu, Wenbo, et al.
Publicado: (2025)
por: Wu, Wenbo, et al.
Publicado: (2025)
Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
por: Oldfield, James, et al.
Publicado: (2025)
por: Oldfield, James, et al.
Publicado: (2025)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
por: Zarch, Hossein Entezari, et al.
Publicado: (2025)
por: Zarch, Hossein Entezari, et al.
Publicado: (2025)
RSAT: Structured Attribution Makes Small Language Models Faithful Table Reasoners
por: Gajjar, Jugal, et al.
Publicado: (2026)
por: Gajjar, Jugal, et al.
Publicado: (2026)
Toward a Theory of Tokenization in LLMs
por: Rajaraman, Nived, et al.
Publicado: (2024)
por: Rajaraman, Nived, et al.
Publicado: (2024)
Towards Interpretable and Trustworthy Time Series Reasoning: A BlueSky Vision
por: Ning, Kanghui, et al.
Publicado: (2025)
por: Ning, Kanghui, et al.
Publicado: (2025)
FaithLM: Towards Faithful Explanations for Large Language Models
por: Chuang, Yu-Neng, et al.
Publicado: (2024)
por: Chuang, Yu-Neng, et al.
Publicado: (2024)
How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines
por: Deng, Junwei, et al.
Publicado: (2026)
por: Deng, Junwei, et al.
Publicado: (2026)
Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences
por: Bekman, Stas, et al.
Publicado: (2025)
por: Bekman, Stas, et al.
Publicado: (2025)
Toward Efficient Membership Inference Attacks against Federated Large Language Models: A Projection Residual Approach
por: Deng, Guilin, et al.
Publicado: (2026)
por: Deng, Guilin, et al.
Publicado: (2026)
Long Input Sequence Network for Long Time Series Forecasting
por: Ma, Chao, et al.
Publicado: (2024)
por: Ma, Chao, et al.
Publicado: (2024)
Incorporating Attribution Importance for Improving Faithfulness Metrics
por: Zhao, Zhixue, et al.
Publicado: (2023)
por: Zhao, Zhixue, et al.
Publicado: (2023)
CRAFT: Calibrated Reasoning with Answer-Faithful Traces via Reinforcement Learning for Multi-Hop Question Answering
por: Liu, Yu, et al.
Publicado: (2026)
por: Liu, Yu, et al.
Publicado: (2026)
Planning Transformer: Long-Horizon Offline Reinforcement Learning with Planning Tokens
por: Clinton, Joseph, et al.
Publicado: (2024)
por: Clinton, Joseph, et al.
Publicado: (2024)
Adaptive Discovery of Interpretable Audio Attributes with Multimodal LLMs for Low-Resource Classification
por: Yoshimura, Kosuke, et al.
Publicado: (2026)
por: Yoshimura, Kosuke, et al.
Publicado: (2026)
Do Contemporary Causal Inference Models Capture Real-World Heterogeneity? Findings from a Large-Scale Benchmark
por: Yu, Haining, et al.
Publicado: (2024)
por: Yu, Haining, et al.
Publicado: (2024)
TokenShapley: Token Level Context Attribution with Shapley Value
por: Xiao, Yingtai, et al.
Publicado: (2025)
por: Xiao, Yingtai, et al.
Publicado: (2025)
Beyond Output Faithfulness: Learning Attributions that Preserve Computational Pathways
por: Zhang, Siyu, et al.
Publicado: (2025)
por: Zhang, Siyu, et al.
Publicado: (2025)
Ejemplares similares
-
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
por: Pan, Wenbo, et al.
Publicado: (2025) -
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
por: Yang, Zhicheng, et al.
Publicado: (2026) -
Faithful or Just Plausible? Evaluating the Faithfulness of Closed-Source LLMs in Medical Reasoning
por: Afolabi, Halimat, et al.
Publicado: (2026) -
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
por: Li, Yingru, et al.
Publicado: (2026) -
TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs
por: Zhang, Yuxiang, et al.
Publicado: (2025)