Stuffed Mamba: Oversized States Lead to the Inability to Forget
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yingfa, Zhang, Xinrong, Hu, Shengding, Han, Xu, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StateX: Enhancing RNN Recall via Post-training State Expansion
by: Shen, Xingyu, et al.
Published: (2025)
by: Shen, Xingyu, et al.
Published: (2025)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
by: Chen, Yingfa, et al.
Published: (2025)
by: Chen, Yingfa, et al.
Published: (2025)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
by: Song, Chenyang, et al.
Published: (2024)
by: Song, Chenyang, et al.
Published: (2024)
LEGENT: Open Platform for Embodied Agents
by: Cheng, Zhili, et al.
Published: (2024)
by: Cheng, Zhili, et al.
Published: (2024)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
by: Chen, Yingfa, et al.
Published: (2026)
by: Chen, Yingfa, et al.
Published: (2026)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
by: Zhang, Xinrong, et al.
Published: (2024)
by: Zhang, Xinrong, et al.
Published: (2024)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
by: Xu, Zukang, et al.
Published: (2025)
by: Xu, Zukang, et al.
Published: (2025)
MemMamba: Rethinking Memory Patterns in State Space Model
by: Wang, Youjin, et al.
Published: (2025)
by: Wang, Youjin, et al.
Published: (2025)
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
by: Zhao, Weilin, et al.
Published: (2025)
by: Zhao, Weilin, et al.
Published: (2025)
Configurable Foundation Models: Building LLMs from a Modular Perspective
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
by: Song, Chenyang, et al.
Published: (2025)
by: Song, Chenyang, et al.
Published: (2025)
Robust and Scalable Model Editing for Large Language Models
by: Chen, Yingfa, et al.
Published: (2024)
by: Chen, Yingfa, et al.
Published: (2024)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
Wings: Learning Multimodal LLMs without Text-only Forgetting
by: Zhang, Yi-Kai, et al.
Published: (2024)
by: Zhang, Yi-Kai, et al.
Published: (2024)
AgentRM: Enhancing Agent Generalization with Reward Modeling
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
by: Diao, Muxi, et al.
Published: (2026)
by: Diao, Muxi, et al.
Published: (2026)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
by: Liu, Chi, et al.
Published: (2026)
by: Liu, Chi, et al.
Published: (2026)
Adaptive Computation Pruning for the Forgetting Transformer
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
Unveiling and Addressing Pseudo Forgetting in Large Language Models
by: Sun, Huashan, et al.
Published: (2024)
by: Sun, Huashan, et al.
Published: (2024)
Graceful Forgetting in Generative Language Models
by: Jiang, Chunyang, et al.
Published: (2025)
by: Jiang, Chunyang, et al.
Published: (2025)
Hidden State Poisoning Attacks against Mamba-based Language Models
by: Mercier, Alexandre Le, et al.
Published: (2026)
by: Mercier, Alexandre Le, et al.
Published: (2026)
LLM Unlearning via Loss Adjustment with Only Forget Data
by: Wang, Yaxuan, et al.
Published: (2024)
by: Wang, Yaxuan, et al.
Published: (2024)
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
by: Lai, Song, et al.
Published: (2025)
by: Lai, Song, et al.
Published: (2025)
When Does Multimodality Lead to Better Time Series Forecasting?
by: Zhang, Xiyuan, et al.
Published: (2025)
by: Zhang, Xiyuan, et al.
Published: (2025)
Offline Learning and Forgetting for Reasoning with Large Language Models
by: Ni, Tianwei, et al.
Published: (2025)
by: Ni, Tianwei, et al.
Published: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)
by: MiniCPM Team, et al.
Published: (2026)
FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning
by: Feng, Yujie, et al.
Published: (2026)
by: Feng, Yujie, et al.
Published: (2026)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
Differential Mamba
by: Schneider, Nadav, et al.
Published: (2025)
by: Schneider, Nadav, et al.
Published: (2025)
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
by: Li, Junzhuo, et al.
Published: (2025)
by: Li, Junzhuo, et al.
Published: (2025)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
by: Chen, Yifang, et al.
Published: (2024)
by: Chen, Yifang, et al.
Published: (2024)
$V_0$: A Generalist Value Model for Any Policy at State Zero
by: Zhang, Yi-Kai, et al.
Published: (2026)
by: Zhang, Yi-Kai, et al.
Published: (2026)
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
by: Liu, Yibai, et al.
Published: (2025)
by: Liu, Yibai, et al.
Published: (2025)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
by: Ren, Ruifeng, et al.
Published: (2024)
by: Ren, Ruifeng, et al.
Published: (2024)
UltraFeedback: Boosting Language Models with Scaled AI Feedback
by: Cui, Ganqu, et al.
Published: (2023)
by: Cui, Ganqu, et al.
Published: (2023)
The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning
by: He, Bingxiang, et al.
Published: (2024)
by: He, Bingxiang, et al.
Published: (2024)
Tool Learning with Foundation Models
by: Qin, Yujia, et al.
Published: (2023)
by: Qin, Yujia, et al.
Published: (2023)
Similar Items
-
StateX: Enhancing RNN Recall via Post-training State Expansion
by: Shen, Xingyu, et al.
Published: (2025) -
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
by: Chen, Yingfa, et al.
Published: (2025) -
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
by: Luo, Kairong, et al.
Published: (2025) -
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
by: Song, Chenyang, et al.
Published: (2024) -
LEGENT: Open Platform for Embodied Agents
by: Cheng, Zhili, et al.
Published: (2024)