Bi-Mamba: Towards Accurate 1-Bit State Space Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Shengkun, Ma, Liqun, Li, Haonan, Sun, Mingjie, Shen, Zhiqiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation
by: Ma, Liqun, et al.
Published: (2024)
by: Ma, Liqun, et al.
Published: (2024)
Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models
by: Das, Rocktim Jyoti, et al.
Published: (2023)
by: Das, Rocktim Jyoti, et al.
Published: (2023)
Sink-Aware Pruning for Diffusion Language Models
by: Myrzakhan, Aidar, et al.
Published: (2026)
by: Myrzakhan, Aidar, et al.
Published: (2026)
DocMamba: Efficient Document Pre-training with State Space Model
by: Hu, Pengfei, et al.
Published: (2024)
by: Hu, Pengfei, et al.
Published: (2024)
MemMamba: Rethinking Memory Patterns in State Space Model
by: Wang, Youjin, et al.
Published: (2025)
by: Wang, Youjin, et al.
Published: (2025)
Making Pre-trained Language Models Better Continual Few-Shot Relation Extractors
by: Ma, Shengkun, et al.
Published: (2024)
by: Ma, Shengkun, et al.
Published: (2024)
MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark
by: Ma, Shengkun, et al.
Published: (2025)
by: Ma, Shengkun, et al.
Published: (2025)
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
by: Cheng, Wenhua, et al.
Published: (2025)
by: Cheng, Wenhua, et al.
Published: (2025)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
by: Chen, Yifang, et al.
Published: (2024)
by: Chen, Yifang, et al.
Published: (2024)
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training
by: Tang, Shengkun, et al.
Published: (2026)
by: Tang, Shengkun, et al.
Published: (2026)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
BlackMamba: Mixture of Experts for State-Space Models
by: Anthony, Quentin, et al.
Published: (2024)
by: Anthony, Quentin, et al.
Published: (2024)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
by: Liang, Weixin, et al.
Published: (2025)
by: Liang, Weixin, et al.
Published: (2025)
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
by: Su, Zunhai, et al.
Published: (2025)
by: Su, Zunhai, et al.
Published: (2025)
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Revealing and Mitigating the Local Pattern Shortcuts of Mamba
by: You, Wangjie, et al.
Published: (2024)
by: You, Wangjie, et al.
Published: (2024)
R2Gen-Mamba: A Selective State Space Model for Radiology Report Generation
by: Sun, Yongheng, et al.
Published: (2024)
by: Sun, Yongheng, et al.
Published: (2024)
EMBRE: Entity-aware Masking for Biomedical Relation Extraction
by: Li, Mingjie, et al.
Published: (2024)
by: Li, Mingjie, et al.
Published: (2024)
R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
by: Chen, Yingfa, et al.
Published: (2024)
by: Chen, Yingfa, et al.
Published: (2024)
SlimPajama-DC: Understanding Data Combinations for LLM Training
by: Shen, Zhiqiang, et al.
Published: (2023)
by: Shen, Zhiqiang, et al.
Published: (2023)
ECMNet:Lightweight Semantic Segmentation with Efficient CNN-Mamba Network
by: Du, Feixiang, et al.
Published: (2025)
by: Du, Feixiang, et al.
Published: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
by: Wang, Dongwei, et al.
Published: (2026)
by: Wang, Dongwei, et al.
Published: (2026)
TransXSSM: A Hybrid Transformer State Space Model with Unified Rotary Position Embedding
by: Wu, Bingheng, et al.
Published: (2025)
by: Wu, Bingheng, et al.
Published: (2025)
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4
by: Bsharat, Sondos Mahmoud, et al.
Published: (2023)
by: Bsharat, Sondos Mahmoud, et al.
Published: (2023)
Automated Classification of Tutors' Dialogue Acts Using Generative AI: A Case Study Using the CIMA Corpus
by: He, Liqun, et al.
Published: (2025)
by: He, Liqun, et al.
Published: (2025)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts
by: Li, Dengchun, et al.
Published: (2024)
by: Li, Dengchun, et al.
Published: (2024)
BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation
by: Li, Minchong, et al.
Published: (2024)
by: Li, Minchong, et al.
Published: (2024)
An Exploration of Mamba for Speech Self-Supervised Models
by: Lin, Tzu-Quan, et al.
Published: (2025)
by: Lin, Tzu-Quan, et al.
Published: (2025)
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
by: Muñoz, J. Pablo, et al.
Published: (2025)
by: Muñoz, J. Pablo, et al.
Published: (2025)
ML-Mamba: Efficient Multi-Modal Large Language Model Utilizing Mamba-2
by: Huang, Wenjun, et al.
Published: (2024)
by: Huang, Wenjun, et al.
Published: (2024)
Hidden State Poisoning Attacks against Mamba-based Language Models
by: Mercier, Alexandre Le, et al.
Published: (2026)
by: Mercier, Alexandre Le, et al.
Published: (2026)
Affective-NLI: Towards Accurate and Interpretable Personality Recognition in Conversation
by: Wen, Zhiyuan, et al.
Published: (2024)
by: Wen, Zhiyuan, et al.
Published: (2024)
A Survey on Diffusion Language Models
by: Li, Tianyi, et al.
Published: (2025)
by: Li, Tianyi, et al.
Published: (2025)
Detoxification of Large Language Models through Output-layer Fusion with a Calibration Model
by: Tian, Yuanhe, et al.
Published: (2025)
by: Tian, Yuanhe, et al.
Published: (2025)
Discovering Agentic Safety Specifications from 1-Bit Danger Signals
by: Gallego, Víctor
Published: (2026)
by: Gallego, Víctor
Published: (2026)
Towards Goal-oriented Prompt Engineering for Large Language Models: A Survey
by: Li, Haochen, et al.
Published: (2024)
by: Li, Haochen, et al.
Published: (2024)
Building Accurate Translation-Tailored LLMs with Language Aware Instruction Tuning
by: Zan, Changtong, et al.
Published: (2024)
by: Zan, Changtong, et al.
Published: (2024)
Similar Items
-
FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation
by: Ma, Liqun, et al.
Published: (2024) -
Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models
by: Das, Rocktim Jyoti, et al.
Published: (2023) -
Sink-Aware Pruning for Diffusion Language Models
by: Myrzakhan, Aidar, et al.
Published: (2026) -
DocMamba: Efficient Document Pre-training with State Space Model
by: Hu, Pengfei, et al.
Published: (2024) -
MemMamba: Rethinking Memory Patterns in State Space Model
by: Wang, Youjin, et al.
Published: (2025)