Enhanced Structured State Space Models via Grouped FIR Filtering and Attention Sink Mechanisms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Meng, Tian, Tao, Yang, Yin, Wuliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
von: Yu, Zhongzhi, et al.
Veröffentlicht: (2024)
von: Yu, Zhongzhi, et al.
Veröffentlicht: (2024)
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
von: Zhang, Stephen, et al.
Veröffentlicht: (2025)
von: Zhang, Stephen, et al.
Veröffentlicht: (2025)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
von: Binkowski, Jakub, et al.
Veröffentlicht: (2026)
von: Binkowski, Jakub, et al.
Veröffentlicht: (2026)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
On the Existence and Behavior of Secondary Attention Sinks
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
When Attention Sink Emerges in Language Models: An Empirical View
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
von: Shang, Bingqi, et al.
Veröffentlicht: (2025)
von: Shang, Bingqi, et al.
Veröffentlicht: (2025)
Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation
von: Mutisya, Hillary, et al.
Veröffentlicht: (2026)
von: Mutisya, Hillary, et al.
Veröffentlicht: (2026)
Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective
von: Bai, Xueying, et al.
Veröffentlicht: (2024)
von: Bai, Xueying, et al.
Veröffentlicht: (2024)
Sessa: Selective State Space Attention
von: Horbatko, Liubomyr
Veröffentlicht: (2026)
von: Horbatko, Liubomyr
Veröffentlicht: (2026)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2024)
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2024)
Sink-Aware Pruning for Diffusion Language Models
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026)
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026)
Rethinking Token Reduction for State Space Models
von: Zhan, Zheng, et al.
Veröffentlicht: (2024)
von: Zhan, Zheng, et al.
Veröffentlicht: (2024)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
Private Language Models via Truncated Laplacian Mechanism
von: Huang, Tianhao, et al.
Veröffentlicht: (2024)
von: Huang, Tianhao, et al.
Veröffentlicht: (2024)
ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
von: Ye, Lu, et al.
Veröffentlicht: (2024)
von: Ye, Lu, et al.
Veröffentlicht: (2024)
Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
von: Ren, Liliang, et al.
Veröffentlicht: (2024)
von: Ren, Liliang, et al.
Veröffentlicht: (2024)
On Structured State-Space Duality
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
von: Chen, Feiyang, et al.
Veröffentlicht: (2025)
von: Chen, Feiyang, et al.
Veröffentlicht: (2025)
Retrieval Backward Attention without Additional Training: Enhance Embeddings of Large Language Models via Repetition
von: Duan, Yifei, et al.
Veröffentlicht: (2025)
von: Duan, Yifei, et al.
Veröffentlicht: (2025)
Generalized Probabilistic Attention Mechanism in Transformers
von: Heo, DongNyeong, et al.
Veröffentlicht: (2024)
von: Heo, DongNyeong, et al.
Veröffentlicht: (2024)
Detection vs. Execution: Single-Bucket Probes Miss Half the Mamba-2 State Sink
von: Jiang, Yuhang
Veröffentlicht: (2026)
von: Jiang, Yuhang
Veröffentlicht: (2026)
The Illusion of State in State-Space Models
von: Merrill, William, et al.
Veröffentlicht: (2024)
von: Merrill, William, et al.
Veröffentlicht: (2024)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
DenseMamba: State Space Models with Dense Hidden Connection for Efficient Large Language Models
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning of State Space Models
von: Galim, Kevin, et al.
Veröffentlicht: (2024)
von: Galim, Kevin, et al.
Veröffentlicht: (2024)
PICASO: Permutation-Invariant Context Composition with State Space Models
von: Liu, Tian Yu, et al.
Veröffentlicht: (2025)
von: Liu, Tian Yu, et al.
Veröffentlicht: (2025)
Fast Multipole Attention: A Scalable Multilevel Attention Mechanism for Text and Images
von: Kang, Yanming, et al.
Veröffentlicht: (2023)
von: Kang, Yanming, et al.
Veröffentlicht: (2023)
One Pass Streaming Algorithm for Super Long Token Attention Approximation in Sublinear Space
von: Addanki, Raghav, et al.
Veröffentlicht: (2023)
von: Addanki, Raghav, et al.
Veröffentlicht: (2023)
Semantic Structure of Feature Space in Large Language Models
von: Kozlowski, Austin C., et al.
Veröffentlicht: (2026)
von: Kozlowski, Austin C., et al.
Veröffentlicht: (2026)
LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering
von: Wong, Sing Hieng, et al.
Veröffentlicht: (2026)
von: Wong, Sing Hieng, et al.
Veröffentlicht: (2026)
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
von: Shin, Seungjun, et al.
Veröffentlicht: (2025)
von: Shin, Seungjun, et al.
Veröffentlicht: (2025)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
von: Liu, Di, et al.
Veröffentlicht: (2024)
von: Liu, Di, et al.
Veröffentlicht: (2024)
Enhancing One-shot Pruned Pre-trained Language Models through Sparse-Dense-Sparse Mechanism
von: Li, Guanchen, et al.
Veröffentlicht: (2024)
von: Li, Guanchen, et al.
Veröffentlicht: (2024)
Node Classification via Semantic-Structural Attention-Enhanced Graph Convolutional Networks
von: Zhu, Hongyin
Veröffentlicht: (2024)
von: Zhu, Hongyin
Veröffentlicht: (2024)
Distil-xLSTM: Learning Attention Mechanisms through Recurrent Structures
von: Thiombiano, Abdoul Majid O., et al.
Veröffentlicht: (2025)
von: Thiombiano, Abdoul Majid O., et al.
Veröffentlicht: (2025)
Structural Rationale Distillation via Reasoning Space Compression
von: Yang, Jialin, et al.
Veröffentlicht: (2026)
von: Yang, Jialin, et al.
Veröffentlicht: (2026)
MambaByte: Token-free Selective State Space Model
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026) -
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
von: Yu, Zhongzhi, et al.
Veröffentlicht: (2024) -
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
von: Zhang, Stephen, et al.
Veröffentlicht: (2025) -
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
von: Binkowski, Jakub, et al.
Veröffentlicht: (2026) -
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
von: Peng, Runyu, et al.
Veröffentlicht: (2026)