Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yingfa, Thai, Zhen Leng, Zhou, Zihan, Zhang, Zhu, Shen, Xingyu, Wang, Shuo, Xiao, Chaojun, Han, Xu, Liu, Zhiyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
StateX: Enhancing RNN Recall via Post-training State Expansion
von: Shen, Xingyu, et al.
Veröffentlicht: (2025)
von: Shen, Xingyu, et al.
Veröffentlicht: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
$\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens
von: Zhang, Xinrong, et al.
Veröffentlicht: (2024)
von: Zhang, Xinrong, et al.
Veröffentlicht: (2024)
Student-in-the-Loop Chain-of-Thought Distillation via Generation-Time Selection
von: He, Chaoqun, et al.
Veröffentlicht: (2026)
von: He, Chaoqun, et al.
Veröffentlicht: (2026)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
Linear Algebra Done Right
von: Axler, Sheldon
Veröffentlicht: (2023)
von: Axler, Sheldon
Veröffentlicht: (2023)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
von: Ling Team, et al.
Veröffentlicht: (2025)
von: Ling Team, et al.
Veröffentlicht: (2025)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
von: Sun, Ao, et al.
Veröffentlicht: (2024)
von: Sun, Ao, et al.
Veröffentlicht: (2024)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
von: Song, Chenyang, et al.
Veröffentlicht: (2025)
von: Song, Chenyang, et al.
Veröffentlicht: (2025)
MoRight: Motion Control Done Right
von: Liu, Shaowei, et al.
Veröffentlicht: (2026)
von: Liu, Shaowei, et al.
Veröffentlicht: (2026)
Monocle: Hybrid Local-Global In-Context Evaluation for Long-Text Generation with Uncertainty-Based Active Learning
von: Wang, Xiaorong, et al.
Veröffentlicht: (2025)
von: Wang, Xiaorong, et al.
Veröffentlicht: (2025)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
Drupal Done Right
von: Coombs, Karen
Veröffentlicht: (2009)
von: Coombs, Karen
Veröffentlicht: (2009)
LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
Robust and Scalable Model Editing for Large Language Models
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
von: Fu, Zizhuo, et al.
Veröffentlicht: (2025)
von: Fu, Zizhuo, et al.
Veröffentlicht: (2025)
Fine-tuning Done Right in Model Editing
von: Yang, Wanli, et al.
Veröffentlicht: (2025)
von: Yang, Wanli, et al.
Veröffentlicht: (2025)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
Fovea Transformer: Efficient Long-Context Modeling with Structured Fine-to-Coarse Attention
von: He, Ziwei, et al.
Veröffentlicht: (2023)
von: He, Ziwei, et al.
Veröffentlicht: (2023)
Effective Distillation to Hybrid xLSTM Architectures
von: Hauzenberger, Lukas, et al.
Veröffentlicht: (2026)
von: Hauzenberger, Lukas, et al.
Veröffentlicht: (2026)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
Squid: Long Context as a New Modality for Energy-Efficient On-Device Language Models
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers
von: Horton, Mark, et al.
Veröffentlicht: (2025)
von: Horton, Mark, et al.
Veröffentlicht: (2025)
Batch Speculative Decoding Done Right
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
Test-Time Training Done Right
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2025)
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
von: Huang, Yuxiang, et al.
Veröffentlicht: (2026)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2026)
Panorama Generation From NFoV Image Done Right
von: Zheng, Dian, et al.
Veröffentlicht: (2025)
von: Zheng, Dian, et al.
Veröffentlicht: (2025)
CFDBench: A Large-Scale Benchmark for Machine Learning Methods in Fluid Dynamics
von: Luo, Yining, et al.
Veröffentlicht: (2023)
von: Luo, Yining, et al.
Veröffentlicht: (2023)
Kimi Linear: An Expressive, Efficient Attention Architecture
von: Kimi Team, et al.
Veröffentlicht: (2025)
von: Kimi Team, et al.
Veröffentlicht: (2025)
Algorithm Support for Graph Databases, Done Right
von: de Graaf, Daan, et al.
Veröffentlicht: (2026)
von: de Graaf, Daan, et al.
Veröffentlicht: (2026)
Consensus Under Adversary Majority Done Right
von: Sridhar, Srivatsan, et al.
Veröffentlicht: (2024)
von: Sridhar, Srivatsan, et al.
Veröffentlicht: (2024)
LeaseGuard: Raft Leases Done Right
von: Davis, A. Jesse Jiryu, et al.
Veröffentlicht: (2025)
von: Davis, A. Jesse Jiryu, et al.
Veröffentlicht: (2025)
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
von: Deng, Weishu, et al.
Veröffentlicht: (2025)
von: Deng, Weishu, et al.
Veröffentlicht: (2025)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
von: Ai, Xuan, et al.
Veröffentlicht: (2026)
von: Ai, Xuan, et al.
Veröffentlicht: (2026)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
An Efficient Long-Context Ranking Architecture With Calibrated LLM Distillation: Application to Person-Job Fit
von: Jouanneau, Warren, et al.
Veröffentlicht: (2026)
von: Jouanneau, Warren, et al.
Veröffentlicht: (2026)
ProEdit: Inversion-based Editing From Prompts Done Right
von: Ouyang, Zhi, et al.
Veröffentlicht: (2025)
von: Ouyang, Zhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
von: Chen, Yingfa, et al.
Veröffentlicht: (2025) -
StateX: Enhancing RNN Recall via Post-training State Expansion
von: Shen, Xingyu, et al.
Veröffentlicht: (2025) -
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
von: MiniCPM Team, et al.
Veröffentlicht: (2026) -
$\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens
von: Zhang, Xinrong, et al.
Veröffentlicht: (2024) -
Student-in-the-Loop Chain-of-Thought Distillation via Generation-Time Selection
von: He, Chaoqun, et al.
Veröffentlicht: (2026)