InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Weilin, Zhou, Zihan, Su, Zhou, Xiao, Chaojun, Li, Yuxuan, Li, Yanghao, Zhang, Yudi, Zhao, Weilun, Li, Zhen, Huang, Yuxiang, Sun, Ao, Han, Xu, Liu, Zhiyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
NOSA: Native and Offloadable Sparse Attention
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
by: Huang, Yuxiang, et al.
Published: (2026)
by: Huang, Yuxiang, et al.
Published: (2026)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
by: Song, Chenyang, et al.
Published: (2026)
by: Song, Chenyang, et al.
Published: (2026)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
by: Zhao, Weilin, et al.
Published: (2025)
by: Zhao, Weilin, et al.
Published: (2025)
Densing Law of LLMs
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
by: Sun, Ao, et al.
Published: (2024)
by: Sun, Ao, et al.
Published: (2024)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
by: Chen, Yingfa, et al.
Published: (2026)
by: Chen, Yingfa, et al.
Published: (2026)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
by: Song, Chenyang, et al.
Published: (2025)
by: Song, Chenyang, et al.
Published: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)
by: MiniCPM Team, et al.
Published: (2026)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
by: Huang, Yuxiang, et al.
Published: (2024)
by: Huang, Yuxiang, et al.
Published: (2024)
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
by: Gao, Yizhao, et al.
Published: (2025)
by: Gao, Yizhao, et al.
Published: (2025)
C-TLSAN: Content-Enhanced Time-Aware Long- and Short-Term Attention Network for Personalized Recommendation
by: Liang, Siqi, et al.
Published: (2025)
by: Liang, Siqi, et al.
Published: (2025)
InfLoRA: Interference-Free Low-Rank Adaptation for Continual Learning
by: Liang, Yan-Shuo, et al.
Published: (2024)
by: Liang, Yan-Shuo, et al.
Published: (2024)
MultiMax: Sparse and Multi-Modal Attention Learning
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
ST-LoRA: Low-rank Adaptation for Spatio-Temporal Forecasting
by: Ruan, Weilin, et al.
Published: (2024)
by: Ruan, Weilin, et al.
Published: (2024)
BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
by: Sun, Ao, et al.
Published: (2025)
by: Sun, Ao, et al.
Published: (2025)
A Dense Fluorinated Composite Polymer Electrolyte Featuring Fluoride‐Rich Interface for Long‐Life Lithium Batteries
by: Sisi Liu, et al.
Published: (2026)
by: Sisi Liu, et al.
Published: (2026)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
by: Zhao, Weilin, et al.
Published: (2024)
by: Zhao, Weilin, et al.
Published: (2024)
Hydrodynamic Evolution and Detectability of Nova Remnants in the Galactic Center
by: Su, Zhao, et al.
Published: (2025)
by: Su, Zhao, et al.
Published: (2025)
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
Efficient Low Rank Attention for Long-Context Inference in Large Language Models
by: Li, Tenghui, et al.
Published: (2025)
by: Li, Tenghui, et al.
Published: (2025)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
InfRS: Incremental Few-Shot Object Detection in Remote Sensing Images
by: Li, Wuzhou, et al.
Published: (2024)
by: Li, Wuzhou, et al.
Published: (2024)
Sparse Meets Dense: Unified Generative Recommendations with Cascaded Sparse-Dense Representations
by: Yang, Yuhao, et al.
Published: (2025)
by: Yang, Yuhao, et al.
Published: (2025)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
by: Xu, Hongtao, et al.
Published: (2026)
by: Xu, Hongtao, et al.
Published: (2026)
The Binary Quantized Neural Network for Dense Prediction via Specially Designed Upsampling and Attention
by: Ding, Xingyu, et al.
Published: (2024)
by: Ding, Xingyu, et al.
Published: (2024)
InfMem: Learning System-2 Memory Control for Long-Context Agent
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO
by: Fang, Xueji, et al.
Published: (2025)
by: Fang, Xueji, et al.
Published: (2025)
PENCIL: Long Thoughts with Short Memory
by: Yang, Chenxiao, et al.
Published: (2025)
by: Yang, Chenxiao, et al.
Published: (2025)
Dense Point Clouds Matter: Dust-GS for Scene Reconstruction from Sparse Viewpoints
by: Chen, Shan, et al.
Published: (2024)
by: Chen, Shan, et al.
Published: (2024)
Lag-Relative Sparse Attention In Long Context Training
by: Liang, Manlai, et al.
Published: (2025)
by: Liang, Manlai, et al.
Published: (2025)
Wind-fed Supermassive Black Hole Accretion by the Nuclear Star Cluster: the Case of M31*
by: Su, Zhao, et al.
Published: (2025)
by: Su, Zhao, et al.
Published: (2025)
AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
Short-Long Convolutions Help Hardware-Efficient Linear Attention to Focus on Long Sequences
by: Liu, Zicheng, et al.
Published: (2024)
by: Liu, Zicheng, et al.
Published: (2024)
DBA-Fusion: Tightly Integrating Deep Dense Visual Bundle Adjustment with Multiple Sensors for Large-Scale Localization and Mapping
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
Flexible Seamless 2-in-1 Design with Sample Size Adaptation
by: Li, Runjia, et al.
Published: (2022)
by: Li, Runjia, et al.
Published: (2022)
Enhancing One-shot Pruned Pre-trained Language Models through Sparse-Dense-Sparse Mechanism
by: Li, Guanchen, et al.
Published: (2024)
by: Li, Guanchen, et al.
Published: (2024)
Similar Items
-
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024) -
NOSA: Native and Offloadable Sparse Attention
by: Huang, Yuxiang, et al.
Published: (2025) -
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
by: Huang, Yuxiang, et al.
Published: (2026) -
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
by: Song, Chenyang, et al.
Published: (2026) -
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025)