ReMamba: Equip Mamba with Effective Long-Sequence Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Danlong, Liu, Jiahao, Li, Bei, Zhang, Huishuai, Wang, Jingang, Cai, Xunliang, Zhao, Dongyan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Parallel Decoding via Hidden Transfer for Lossless Large Language Model Acceleration
by: Wu, Pengfei, et al.
Published: (2024)
by: Wu, Pengfei, et al.
Published: (2024)
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
by: Liu, Jiahao, et al.
Published: (2024)
by: Liu, Jiahao, et al.
Published: (2024)
FIRP: Faster LLM inference via future intermediate representation prediction
by: Wu, Pengfei, et al.
Published: (2024)
by: Wu, Pengfei, et al.
Published: (2024)
Shorten After You're Right: Lazy Length Penalties for Reasoning RL
by: Yuan, Danlong, et al.
Published: (2025)
by: Yuan, Danlong, et al.
Published: (2025)
Graph-Structured Speculative Decoding
by: Gong, Zhuocheng, et al.
Published: (2024)
by: Gong, Zhuocheng, et al.
Published: (2024)
Libra: Assessing and Improving Reward Model by Learning to Think
by: Zhou, Meng, et al.
Published: (2025)
by: Zhou, Meng, et al.
Published: (2025)
Evidence-Enhanced Triplet Generation Framework for Hallucination Alleviation in Generative Question Answering
by: Du, Haowei, et al.
Published: (2024)
by: Du, Haowei, et al.
Published: (2024)
Dynamic Fisher-weighted Model Merging via Bayesian Optimization
by: Lee, Sanwoo, et al.
Published: (2025)
by: Lee, Sanwoo, et al.
Published: (2025)
Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes
by: Gong, Zhuocheng, et al.
Published: (2025)
by: Gong, Zhuocheng, et al.
Published: (2025)
FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training
by: Xu, Liangyu, et al.
Published: (2025)
by: Xu, Liangyu, et al.
Published: (2025)
NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables
by: Wang, Lanrui, et al.
Published: (2025)
by: Wang, Lanrui, et al.
Published: (2025)
Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay Perspective
by: Shao, Ruichen, et al.
Published: (2025)
by: Shao, Ruichen, et al.
Published: (2025)
ML-Mamba: Efficient Multi-Modal Large Language Model Utilizing Mamba-2
by: Huang, Wenjun, et al.
Published: (2024)
by: Huang, Wenjun, et al.
Published: (2024)
Efficient Continual Pre-training by Mitigating the Stability Gap
by: Guo, Yiduo, et al.
Published: (2024)
by: Guo, Yiduo, et al.
Published: (2024)
LightMamba: Efficient Mamba Acceleration on FPGA with Quantization and Hardware Co-design
by: Wei, Renjie, et al.
Published: (2025)
by: Wei, Renjie, et al.
Published: (2025)
Bio-Inspired Mamba: Temporal Locality and Bioplausible Learning in Selective State Space Models
by: Qin, Jiahao
Published: (2024)
by: Qin, Jiahao
Published: (2024)
LongMamba: Enhancing Mamba's Long Context Capabilities via Training-Free Receptive Field Enlargement
by: Ye, Zhifan, et al.
Published: (2025)
by: Ye, Zhifan, et al.
Published: (2025)
Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient Learning
by: Li, Bei, et al.
Published: (2024)
by: Li, Bei, et al.
Published: (2024)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
by: Tang, Hongyin, et al.
Published: (2024)
by: Tang, Hongyin, et al.
Published: (2024)
Causal Autoregressive Diffusion Language Model
by: Ruan, Junhao, et al.
Published: (2026)
by: Ruan, Junhao, et al.
Published: (2026)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
MemMamba: Rethinking Memory Patterns in State Space Model
by: Wang, Youjin, et al.
Published: (2025)
by: Wang, Youjin, et al.
Published: (2025)
Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-context Models
by: Liu, Xinyu, et al.
Published: (2024)
by: Liu, Xinyu, et al.
Published: (2024)
SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models
by: Diao, Muxi, et al.
Published: (2024)
by: Diao, Muxi, et al.
Published: (2024)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text
by: Xu, Zhihao, et al.
Published: (2026)
by: Xu, Zhihao, et al.
Published: (2026)
What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation
by: Gong, Zhuocheng, et al.
Published: (2024)
by: Gong, Zhuocheng, et al.
Published: (2024)
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
by: Yuan, Danlong, et al.
Published: (2026)
by: Yuan, Danlong, et al.
Published: (2026)
Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
by: Ma, Junyu, et al.
Published: (2025)
by: Ma, Junyu, et al.
Published: (2025)
Venturing into Uncharted Waters: The Navigation Compass from Transformer to Mamba
by: Zou, Yuchen, et al.
Published: (2024)
by: Zou, Yuchen, et al.
Published: (2024)
GeneMamba: An Efficient and Effective Foundation Model on Single Cell Data
by: Qi, Cong, et al.
Published: (2025)
by: Qi, Cong, et al.
Published: (2025)
LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Large-Scale Diverse Synthesis for Mid-Training
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
by: Xu, Zukang, et al.
Published: (2025)
by: Xu, Zukang, et al.
Published: (2025)
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
by: Ruan, Junhao, et al.
Published: (2026)
by: Ruan, Junhao, et al.
Published: (2026)
BioMamba: Domain-Adaptive Biomedical Language Models
by: Yue, Ling, et al.
Published: (2024)
by: Yue, Ling, et al.
Published: (2024)
Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
RankMamba: Benchmarking Mamba's Document Ranking Performance in the Era of Transformers
by: Xu, Zhichao
Published: (2024)
by: Xu, Zhichao
Published: (2024)
Similar Items
-
Parallel Decoding via Hidden Transfer for Lossless Large Language Model Acceleration
by: Wu, Pengfei, et al.
Published: (2024) -
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
by: Liu, Jiahao, et al.
Published: (2024) -
FIRP: Faster LLM inference via future intermediate representation prediction
by: Wu, Pengfei, et al.
Published: (2024) -
Shorten After You're Right: Lazy Length Penalties for Reasoning RL
by: Yuan, Danlong, et al.
Published: (2025) -
Graph-Structured Speculative Decoding
by: Gong, Zhuocheng, et al.
Published: (2024)