Distilling to Hybrid Attention Models via KL-Guided Layer Selection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yanhong, Yang, Songlin, Tan, Shawn, Mishra, Mayank, Panda, Rameswar, Zhou, Jiawei, Kim, Yoon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
Chunk-Distilled Language Modeling
von: Li, Yanhong, et al.
Veröffentlicht: (2024)
von: Li, Yanhong, et al.
Veröffentlicht: (2024)
Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
von: Shen, Yikang, et al.
Veröffentlicht: (2024)
von: Shen, Yikang, et al.
Veröffentlicht: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
Data Engineering for Scaling Language Models to 128K Context
von: Fu, Yao, et al.
Veröffentlicht: (2024)
von: Fu, Yao, et al.
Veröffentlicht: (2024)
Gated Linear Attention Transformers with Hardware-Efficient Training
von: Yang, Songlin, et al.
Veröffentlicht: (2023)
von: Yang, Songlin, et al.
Veröffentlicht: (2023)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
von: Brandon, William, et al.
Veröffentlicht: (2024)
von: Brandon, William, et al.
Veröffentlicht: (2024)
TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2024)
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2024)
On the Predictive Power of Representation Dispersion in Language Models
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs Distillation
von: Dai, Chengwei, et al.
Veröffentlicht: (2024)
von: Dai, Chengwei, et al.
Veröffentlicht: (2024)
Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation
von: Dai, Chengwei, et al.
Veröffentlicht: (2024)
von: Dai, Chengwei, et al.
Veröffentlicht: (2024)
LangNav: Language as a Perceptual Representation for Navigation
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
KL for a KL: On-Policy Distillation with Control Variate Baseline
von: Oh, Minjae, et al.
Veröffentlicht: (2026)
von: Oh, Minjae, et al.
Veröffentlicht: (2026)
WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models
von: Lee, Hanna, et al.
Veröffentlicht: (2026)
von: Lee, Hanna, et al.
Veröffentlicht: (2026)
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
von: Guo, Zhen, et al.
Veröffentlicht: (2024)
von: Guo, Zhen, et al.
Veröffentlicht: (2024)
Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
LLM-MRD: LLM-Guided Multi-View Reasoning Distillation for Fake News Detection
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2025)
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2025)
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
von: Li, Shuaiyi, et al.
Veröffentlicht: (2026)
von: Li, Shuaiyi, et al.
Veröffentlicht: (2026)
ReAD: Reinforcement-Guided Capability Distillation for Large Language Models
von: Cheng, Xueqi, et al.
Veröffentlicht: (2026)
von: Cheng, Xueqi, et al.
Veröffentlicht: (2026)
What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Sensory-Aware Sequential Recommendation via Review-Distilled Representations
von: Yoon, Yeo Chan, et al.
Veröffentlicht: (2026)
von: Yoon, Yeo Chan, et al.
Veröffentlicht: (2026)
Selective Self-Rehearsal: A Fine-Tuning Approach to Improve Generalization in Large Language Models
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
von: Mishra, Mayank, et al.
Veröffentlicht: (2026)
von: Mishra, Mayank, et al.
Veröffentlicht: (2026)
Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
von: Krishnakumar, Arjun, et al.
Veröffentlicht: (2025)
von: Krishnakumar, Arjun, et al.
Veröffentlicht: (2025)
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
RoSA: Enhancing Parameter-Efficient Fine-Tuning via RoPE-aware Selective Adaptation in Large Language Models
von: Pan, Dayan, et al.
Veröffentlicht: (2025)
von: Pan, Dayan, et al.
Veröffentlicht: (2025)
$Δ$-AttnMask: Attention-Guided Masked Hidden States for Efficient Data Selection and Augmentation
von: Hu, Jucheng, et al.
Veröffentlicht: (2025)
von: Hu, Jucheng, et al.
Veröffentlicht: (2025)
Hybrid Policy Distillation for LLMs
von: Zhu, Wenhong, et al.
Veröffentlicht: (2026)
von: Zhu, Wenhong, et al.
Veröffentlicht: (2026)
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
ELAD: Explanation-Guided Large Language Models Active Distillation
von: Zhang, Yifei, et al.
Veröffentlicht: (2024)
von: Zhang, Yifei, et al.
Veröffentlicht: (2024)
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
von: Gao, Yizhao, et al.
Veröffentlicht: (2026)
von: Gao, Yizhao, et al.
Veröffentlicht: (2026)
Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models
von: Zhou, Yue, et al.
Veröffentlicht: (2024)
von: Zhou, Yue, et al.
Veröffentlicht: (2024)
Scaling Bidirectional Spans and Span Violations in Attention Mechanism
von: Kim, Jongwook, et al.
Veröffentlicht: (2025)
von: Kim, Jongwook, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
von: Tan, Shawn, et al.
Veröffentlicht: (2024) -
PaTH Attention: Position Encoding via Accumulating Householder Transformations
von: Yang, Songlin, et al.
Veröffentlicht: (2025) -
Chunk-Distilled Language Modeling
von: Li, Yanhong, et al.
Veröffentlicht: (2024) -
Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
von: Wang, Xinyi, et al.
Veröffentlicht: (2025) -
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
von: Shen, Yikang, et al.
Veröffentlicht: (2024)