Native Hybrid Attention for Efficient Sequence Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Jusen, Hu, Jiaxi, Zhang, Tao, Sun, Weigao, Cheng, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoM: Linear Sequence Modeling with Mixture-of-Memories
von: Du, Jusen, et al.
Veröffentlicht: (2025)
von: Du, Jusen, et al.
Veröffentlicht: (2025)
Liger: Linearizing Large Language Models to Gated Recurrent Structures
von: Lan, Disen, et al.
Veröffentlicht: (2025)
von: Lan, Disen, et al.
Veröffentlicht: (2025)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Comba: Improving Bilinear RNNs with Closed-loop Control
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
NOSA: Native and Offloadable Sparse Attention
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
von: Ling Team, et al.
Veröffentlicht: (2025)
von: Ling Team, et al.
Veröffentlicht: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
Speed Always Wins: A Survey on Efficient Architectures for Large Language Models
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Star Attention: Efficient LLM Inference over Long Sequences
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
Linear Attention Sequence Parallelism
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
Demystifying the Slash Pattern in Attention: The Role of RoPE
von: Cheng, Yuan, et al.
Veröffentlicht: (2026)
von: Cheng, Yuan, et al.
Veröffentlicht: (2026)
Synergy-of-Thoughts: Eliciting Efficient Reasoning in Hybrid Language Models
von: Shang, Yu, et al.
Veröffentlicht: (2024)
von: Shang, Yu, et al.
Veröffentlicht: (2024)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
When Attention Sink Emerges in Language Models: An Empirical View
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
Characterizing Model-Native Skills
von: Kang, Feiyang, et al.
Veröffentlicht: (2026)
von: Kang, Feiyang, et al.
Veröffentlicht: (2026)
Sliding Window Attention Training for Efficient Large Language Models
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
von: Dentamaro, Vincenzo
Veröffentlicht: (2025)
von: Dentamaro, Vincenzo
Veröffentlicht: (2025)
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
von: Lv, Xingtai, et al.
Veröffentlicht: (2024)
von: Lv, Xingtai, et al.
Veröffentlicht: (2024)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2024)
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2024)
Block-Attention for Efficient Prefilling
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
von: Zhang, Tianyi, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2024)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
von: Sharma, Agniv, et al.
Veröffentlicht: (2024)
von: Sharma, Agniv, et al.
Veröffentlicht: (2024)
Sequence-to-Sequence Spanish Pre-trained Language Models
von: Araujo, Vladimir, et al.
Veröffentlicht: (2023)
von: Araujo, Vladimir, et al.
Veröffentlicht: (2023)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
von: Hu, Xing, et al.
Veröffentlicht: (2024)
von: Hu, Xing, et al.
Veröffentlicht: (2024)
HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization
von: Zhuo, Zhijian, et al.
Veröffentlicht: (2025)
von: Zhuo, Zhijian, et al.
Veröffentlicht: (2025)
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU
von: Chen, Weizhe, et al.
Veröffentlicht: (2026)
von: Chen, Weizhe, et al.
Veröffentlicht: (2026)
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2025)
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2025)
Parameter-Efficient Fine-Tuning for Foundation Models
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
Parallax: Parameterized Local Linear Attention for Language Modeling
von: Zuo, Yifei, et al.
Veröffentlicht: (2026)
von: Zuo, Yifei, et al.
Veröffentlicht: (2026)
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
Interpretable by AI Mother Tongue: Native Symbolic Reasoning in Neural Models
von: Liu, Hung Ming
Veröffentlicht: (2025)
von: Liu, Hung Ming
Veröffentlicht: (2025)
ArabianGPT: Native Arabic GPT-based Large Language Model
von: Koubaa, Anis, et al.
Veröffentlicht: (2024)
von: Koubaa, Anis, et al.
Veröffentlicht: (2024)
FusionBench: A Unified Library and Comprehensive Benchmark for Deep Model Fusion
von: Tang, Anke, et al.
Veröffentlicht: (2024)
von: Tang, Anke, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MoM: Linear Sequence Modeling with Mixture-of-Memories
von: Du, Jusen, et al.
Veröffentlicht: (2025) -
Liger: Linearizing Large Language Models to Gated Recurrent Structures
von: Lan, Disen, et al.
Veröffentlicht: (2025) -
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
von: Sun, Weigao, et al.
Veröffentlicht: (2025) -
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025) -
Comba: Improving Bilinear RNNs with Closed-loop Control
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)