Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Yiju, Yang, Wenkai, Sun, Zexu, Ding, Ning, Liu, Zhiyuan, Lin, Yankai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
von: Guo, Yiju, et al.
Veröffentlicht: (2026)
von: Guo, Yiju, et al.
Veröffentlicht: (2026)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
von: Guo, Yiju, et al.
Veröffentlicht: (2024)
von: Guo, Yiju, et al.
Veröffentlicht: (2024)
Distilling Rule-based Knowledge into Large Language Models
von: Yang, Wenkai, et al.
Veröffentlicht: (2023)
von: Yang, Wenkai, et al.
Veröffentlicht: (2023)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
Exploring Backdoor Vulnerabilities of Chat Models
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2024)
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2024)
EGAD: Entropy-Guided Adaptive Distillation for Token-Level Knowledge Transfer
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
von: Liu, Weihao, et al.
Veröffentlicht: (2025)
von: Liu, Weihao, et al.
Veröffentlicht: (2025)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2026)
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2026)
SmartTrim: Adaptive Tokens and Attention Pruning for Efficient Vision-Language Models
von: Wang, Zekun, et al.
Veröffentlicht: (2023)
von: Wang, Zekun, et al.
Veröffentlicht: (2023)
Representation Learning for Natural Language Processing
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2020)
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2020)
AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation
von: Fang, Yixiong, et al.
Veröffentlicht: (2025)
von: Fang, Yixiong, et al.
Veröffentlicht: (2025)
The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE Training
von: Chen, Weize, et al.
Veröffentlicht: (2025)
von: Chen, Weize, et al.
Veröffentlicht: (2025)
LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models
von: Sun, Yizheng, et al.
Veröffentlicht: (2025)
von: Sun, Yizheng, et al.
Veröffentlicht: (2025)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)
MBTSAD: Mitigating Backdoors in Language Models Based on Token Splitting and Attention Distillation
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
DeepCritic: Deliberate Critique with Large Language Models
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
Token Distillation: Attention-aware Input Embeddings For New Tokens
von: Dobler, Konstantin, et al.
Veröffentlicht: (2025)
von: Dobler, Konstantin, et al.
Veröffentlicht: (2025)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
CATP: Cross-Attention Token Pruning for Accuracy Preserved Multimodal Model Inference
von: Liao, Ruqi, et al.
Veröffentlicht: (2024)
von: Liao, Ruqi, et al.
Veröffentlicht: (2024)
Causal-Guided Active Learning for Debiasing Large Language Models
von: Du, Li, et al.
Veröffentlicht: (2024)
von: Du, Li, et al.
Veröffentlicht: (2024)
The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
von: Jiang, Titong, et al.
Veröffentlicht: (2025)
von: Jiang, Titong, et al.
Veröffentlicht: (2025)
Distilling to Hybrid Attention Models via KL-Guided Layer Selection
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
Towards Codable Watermarking for Injecting Multi-bits Information to LLMs
von: Wang, Lean, et al.
Veröffentlicht: (2023)
von: Wang, Lean, et al.
Veröffentlicht: (2023)
Optimizing Korean-Centric LLMs via Token Pruning
von: Kim, Hoyeol, et al.
Veröffentlicht: (2026)
von: Kim, Hoyeol, et al.
Veröffentlicht: (2026)
SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning
von: Long, Lingkun, et al.
Veröffentlicht: (2025)
von: Long, Lingkun, et al.
Veröffentlicht: (2025)
SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
WuNeng: Hybrid State with Attention
von: Xiao, Liu, et al.
Veröffentlicht: (2025)
von: Xiao, Liu, et al.
Veröffentlicht: (2025)
ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
von: Hou, Bairu, et al.
Veröffentlicht: (2025)
von: Hou, Bairu, et al.
Veröffentlicht: (2025)
HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models
von: Guo, Yansong, et al.
Veröffentlicht: (2026)
von: Guo, Yansong, et al.
Veröffentlicht: (2026)
Predicting Emergent Abilities with Infinite Resolution Evaluation
von: Hu, Shengding, et al.
Veröffentlicht: (2023)
von: Hu, Shengding, et al.
Veröffentlicht: (2023)
IG-Pruning: Input-Guided Block Pruning for Large Language Models
von: Qiao, Kangyu, et al.
Veröffentlicht: (2025)
von: Qiao, Kangyu, et al.
Veröffentlicht: (2025)
Sentence-Level or Token-Level? A Comprehensive Study on Knowledge Distillation
von: Wei, Jingxuan, et al.
Veröffentlicht: (2024)
von: Wei, Jingxuan, et al.
Veröffentlicht: (2024)
GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
von: Tang, Yixuan, et al.
Veröffentlicht: (2025)
von: Tang, Yixuan, et al.
Veröffentlicht: (2025)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
Winning the Pruning Gamble: A Unified Approach to Joint Sample and Token Pruning for Efficient Supervised Fine-Tuning
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
LongAttn: Selecting Long-context Training Data via Token-level Attention
von: Wu, Longyun, et al.
Veröffentlicht: (2025)
von: Wu, Longyun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
von: Guo, Yiju, et al.
Veröffentlicht: (2026) -
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
von: Yang, Wenkai, et al.
Veröffentlicht: (2025) -
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
von: Guo, Yiju, et al.
Veröffentlicht: (2024) -
Distilling Rule-based Knowledge into Large Language Models
von: Yang, Wenkai, et al.
Veröffentlicht: (2023) -
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)