On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Ro, Yeonju, Zhang, Zhenyu, Kundu, Souvik, Wang, Zhangyang, Akella, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving the Throughput of Diffusion-based Large Language Models via a Training-Free Confidence-Aware Calibration
by: Shen, Jucheng, et al.
Published: (2025)
by: Shen, Jucheng, et al.
Published: (2025)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
by: Cai, Ruisi, et al.
Published: (2024)
by: Cai, Ruisi, et al.
Published: (2024)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
by: Jaiswal, Ajay, et al.
Published: (2024)
by: Jaiswal, Ajay, et al.
Published: (2024)
Beyond Static Cutoffs: One-Shot Dynamic Thresholding for Diffusion Language Models
by: Shen, Jucheng, et al.
Published: (2025)
by: Shen, Jucheng, et al.
Published: (2025)
Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
by: Wang, Penghao, et al.
Published: (2025)
by: Wang, Penghao, et al.
Published: (2025)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
by: Yin, Lu, et al.
Published: (2023)
by: Yin, Lu, et al.
Published: (2023)
When Do Graph Foundation Models Transfer? A Data-Centric Theory
by: Zhu, Jiajun, et al.
Published: (2026)
by: Zhu, Jiajun, et al.
Published: (2026)
A Teacher-Free Graph Knowledge Distillation Framework with Dual Self-Distillation
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
Linearizing Models for Efficient yet Robust Private Inference
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model Training
by: Kim, Geon-Woo, et al.
Published: (2025)
by: Kim, Geon-Woo, et al.
Published: (2025)
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
by: Zhu, Jianing, et al.
Published: (2026)
by: Zhu, Jianing, et al.
Published: (2026)
FlyKD: Graph Knowledge Distillation on the Fly with Curriculum Learning
by: Ku, Eugene
Published: (2024)
by: Ku, Eugene
Published: (2024)
On the Fundamental Limitations of Decentralized Learnable Reward Shaping in Cooperative Multi-Agent Reinforcement Learning
by: Akella, Aditya
Published: (2025)
by: Akella, Aditya
Published: (2025)
E2Former-V2: On-the-Fly Equivariant Attention with Linear Activation Memory
by: Huang, Lin, et al.
Published: (2026)
by: Huang, Lin, et al.
Published: (2026)
COBRA: Catastrophic Bit-flip Reliability Analysis of State-Space Models
by: Das, Sanjay, et al.
Published: (2025)
by: Das, Sanjay, et al.
Published: (2025)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
by: Zhang, Zhenyu, et al.
Published: (2024)
by: Zhang, Zhenyu, et al.
Published: (2024)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
by: Nishikawa, Naoki, et al.
Published: (2025)
by: Nishikawa, Naoki, et al.
Published: (2025)
Projection-Free Transformers via Gaussian Kernel Attention
by: Kundu, Debarshi, et al.
Published: (2026)
by: Kundu, Debarshi, et al.
Published: (2026)
Block Selective Reprogramming for On-device Training of Vision Transformers
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
Cottention: Linear Transformers With Cosine Attention
by: Mongaras, Gabriel, et al.
Published: (2024)
by: Mongaras, Gabriel, et al.
Published: (2024)
Adaptive Memory Decay for Log-Linear Attention
by: Amin, Yaxita, et al.
Published: (2026)
by: Amin, Yaxita, et al.
Published: (2026)
Sherlock: Reliable and Efficient Agentic Workflow Execution
by: Ro, Yeonju, et al.
Published: (2025)
by: Ro, Yeonju, et al.
Published: (2025)
InAttention: Linear Context Scaling for Transformers
by: Eisner, Joseph
Published: (2024)
by: Eisner, Joseph
Published: (2024)
Scaling Linear Attention with Sparse State Expansion
by: Pan, Yuqi, et al.
Published: (2025)
by: Pan, Yuqi, et al.
Published: (2025)
Comparative Analysis of Time Series Foundation Models for Demographic Forecasting: Enhancing Predictive Accuracy in US Population Dynamics
by: Akella, Aditya, et al.
Published: (2025)
by: Akella, Aditya, et al.
Published: (2025)
Hybrid Dual-Path Linear Transformations for Efficient Transformer Architectures
by: Khasia, Vladimer
Published: (2026)
by: Khasia, Vladimer
Published: (2026)
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
TwinFormer: A Dual-Level Transformer for Long-Sequence Time-Series Forecasting
by: Kumavat, Mahima, et al.
Published: (2025)
by: Kumavat, Mahima, et al.
Published: (2025)
PRISM: Enhancing Protein Inverse Folding through Fine-Grained Retrieval on Structure-Sequence Multimodal Representations
by: Mahbub, Sazan, et al.
Published: (2025)
by: Mahbub, Sazan, et al.
Published: (2025)
LaMDA: Large Model Fine-Tuning via Spectrally Decomposed Low-Dimensional Adaptation
by: Azizi, Seyedarmin, et al.
Published: (2024)
by: Azizi, Seyedarmin, et al.
Published: (2024)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
by: Wu, Ziyang, et al.
Published: (2024)
by: Wu, Ziyang, et al.
Published: (2024)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
by: Hu, Wenjie, et al.
Published: (2025)
by: Hu, Wenjie, et al.
Published: (2025)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning
by: Wang, Peihao, et al.
Published: (2025)
by: Wang, Peihao, et al.
Published: (2025)
On the Interpolation Error of Nonlinear Attention versus Linear Regression
by: Liao, Zhenyu, et al.
Published: (2025)
by: Liao, Zhenyu, et al.
Published: (2025)
State Rank Dynamics in Linear Attention LLMs
by: Sun, Ao, et al.
Published: (2026)
by: Sun, Ao, et al.
Published: (2026)
Similar Items
-
Improving the Throughput of Diffusion-based Large Language Models via a Training-Free Confidence-Aware Calibration
by: Shen, Jucheng, et al.
Published: (2025) -
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
by: Cai, Ruisi, et al.
Published: (2024) -
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
by: Jaiswal, Ajay, et al.
Published: (2024) -
Beyond Static Cutoffs: One-Shot Dynamic Thresholding for Diffusion Language Models
by: Shen, Jucheng, et al.
Published: (2025) -
Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
by: Wang, Penghao, et al.
Published: (2025)