MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Zukang, Yue, Yuxuan, Hu, Xing, Yuan, Zhihang, Jiang, Zixu, Chen, Zhixuan, Yu, Jiangyong, Xu, Chen, Zhou, Sifan, Yang, Dawei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
di: Xu, Chen, et al.
Pubblicazione: (2025)
di: Xu, Chen, et al.
Pubblicazione: (2025)
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
di: Hu, Xing, et al.
Pubblicazione: (2025)
di: Hu, Xing, et al.
Pubblicazione: (2025)
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
di: Hu, Xing, et al.
Pubblicazione: (2025)
di: Hu, Xing, et al.
Pubblicazione: (2025)
TORQ: Two-Level Orthogonal Rotation for MXFP4 Quantization
di: Xu, Zukang, et al.
Pubblicazione: (2026)
di: Xu, Zukang, et al.
Pubblicazione: (2026)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
di: Hu, Xing, et al.
Pubblicazione: (2024)
di: Hu, Xing, et al.
Pubblicazione: (2024)
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
di: Xu, Zukang, et al.
Pubblicazione: (2026)
di: Xu, Zukang, et al.
Pubblicazione: (2026)
PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling
di: Yue, Yuxuan, et al.
Pubblicazione: (2025)
di: Yue, Yuxuan, et al.
Pubblicazione: (2025)
MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization
di: Yu, JiangYong, et al.
Pubblicazione: (2025)
di: Yu, JiangYong, et al.
Pubblicazione: (2025)
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
di: Zhao, Zhixiong, et al.
Pubblicazione: (2026)
di: Zhao, Zhixiong, et al.
Pubblicazione: (2026)
SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression
di: Hu, Xing, et al.
Pubblicazione: (2026)
di: Hu, Xing, et al.
Pubblicazione: (2026)
RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
di: Xu, Zukang, et al.
Pubblicazione: (2025)
di: Xu, Zukang, et al.
Pubblicazione: (2025)
MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization
di: Wang, Zhong, et al.
Pubblicazione: (2026)
di: Wang, Zhong, et al.
Pubblicazione: (2026)
FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection
di: Yu, Jiangyong, et al.
Pubblicazione: (2025)
di: Yu, Jiangyong, et al.
Pubblicazione: (2025)
FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection
di: Yu, Jiangyong, et al.
Pubblicazione: (2025)
di: Yu, Jiangyong, et al.
Pubblicazione: (2025)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
di: Zhao, Zhixiong, et al.
Pubblicazione: (2026)
di: Zhao, Zhixiong, et al.
Pubblicazione: (2026)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
di: Li, Yan, et al.
Pubblicazione: (2024)
di: Li, Yan, et al.
Pubblicazione: (2024)
AlignMamba-2: Enhancing Multimodal Fusion and Sentiment Analysis with Modality-Aware Mamba
di: Li, Yan, et al.
Pubblicazione: (2026)
di: Li, Yan, et al.
Pubblicazione: (2026)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
di: Yue, Yuxuan, et al.
Pubblicazione: (2024)
di: Yue, Yuxuan, et al.
Pubblicazione: (2024)
Information Entropy Guided Height-aware Histogram for Quantization-friendly Pillar Feature Encoder
di: Zhou, Sifan, et al.
Pubblicazione: (2024)
di: Zhou, Sifan, et al.
Pubblicazione: (2024)
LightMamba: Efficient Mamba Acceleration on FPGA with Quantization and Hardware Co-design
di: Wei, Renjie, et al.
Pubblicazione: (2025)
di: Wei, Renjie, et al.
Pubblicazione: (2025)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
di: Zhou, Sifan, et al.
Pubblicazione: (2025)
di: Zhou, Sifan, et al.
Pubblicazione: (2025)
OuroMamba: A Data-Free Quantization Framework for Vision Mamba
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration
di: Chen, Yujie, et al.
Pubblicazione: (2025)
di: Chen, Yujie, et al.
Pubblicazione: (2025)
Rotation Equivariant Mamba for Vision Tasks
di: Zhao, Zhongchen, et al.
Pubblicazione: (2026)
di: Zhao, Zhongchen, et al.
Pubblicazione: (2026)
OTARo: Once Tuning for All Precisions toward Robust On-Device LLMs
di: Chen, Shaoyuan, et al.
Pubblicazione: (2025)
di: Chen, Shaoyuan, et al.
Pubblicazione: (2025)
NLI:Non-uniform Linear Interpolation Approximation of Nonlinear Operations for Efficient LLMs Inference
di: Yu, Jiangyong, et al.
Pubblicazione: (2026)
di: Yu, Jiangyong, et al.
Pubblicazione: (2026)
miMamba: EEG-based Emotion Recognition with Multi-scale Inverted Mamba Models
di: Zhou, Xin, et al.
Pubblicazione: (2024)
di: Zhou, Xin, et al.
Pubblicazione: (2024)
Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution
di: Shi, Shijun, et al.
Pubblicazione: (2025)
di: Shi, Shijun, et al.
Pubblicazione: (2025)
MobileMamba: Lightweight Multi-Receptive Visual Mamba Network
di: He, Haoyang, et al.
Pubblicazione: (2024)
di: He, Haoyang, et al.
Pubblicazione: (2024)
SMILES-Mamba: Chemical Mamba Foundation Models for Drug ADMET Prediction
di: Xu, Bohao, et al.
Pubblicazione: (2024)
di: Xu, Bohao, et al.
Pubblicazione: (2024)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
FlatQuant: Flatness Matters for LLM Quantization
di: Sun, Yuxuan, et al.
Pubblicazione: (2024)
di: Sun, Yuxuan, et al.
Pubblicazione: (2024)
TransMamba: Fast Universal Architecture Adaption from Transformers to Mamba
di: Chen, Xiuwei, et al.
Pubblicazione: (2025)
di: Chen, Xiuwei, et al.
Pubblicazione: (2025)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
di: Liang, Yesheng, et al.
Pubblicazione: (2025)
di: Liang, Yesheng, et al.
Pubblicazione: (2025)
MambaVesselNet++: A Hybrid CNN-Mamba Architecture for Medical Image Segmentation
di: Xu, Qing, et al.
Pubblicazione: (2025)
di: Xu, Qing, et al.
Pubblicazione: (2025)
ST-Mamba: Spatial-Temporal Mamba for Traffic Flow Estimation Recovery using Limited Data
di: Yuan, Doncheng, et al.
Pubblicazione: (2024)
di: Yuan, Doncheng, et al.
Pubblicazione: (2024)
RI-Mamba: Rotation-Invariant Mamba for Robust Text-to-Shape Retrieval
di: Nguyen, Khanh, et al.
Pubblicazione: (2026)
di: Nguyen, Khanh, et al.
Pubblicazione: (2026)
Protein-Mamba: Biological Mamba Models for Protein Function Prediction
di: Xu, Bohao, et al.
Pubblicazione: (2024)
di: Xu, Bohao, et al.
Pubblicazione: (2024)
MMR-Mamba: Multi-Modal MRI Reconstruction with Mamba and Spatial-Frequency Information Fusion
di: Zou, Jing, et al.
Pubblicazione: (2024)
di: Zou, Jing, et al.
Pubblicazione: (2024)
RankMamba: Benchmarking Mamba's Document Ranking Performance in the Era of Transformers
di: Xu, Zhichao
Pubblicazione: (2024)
di: Xu, Zhichao
Pubblicazione: (2024)
Documenti analoghi
-
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
di: Xu, Chen, et al.
Pubblicazione: (2025) -
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
di: Hu, Xing, et al.
Pubblicazione: (2025) -
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
di: Hu, Xing, et al.
Pubblicazione: (2025) -
TORQ: Two-Level Orthogonal Rotation for MXFP4 Quantization
di: Xu, Zukang, et al.
Pubblicazione: (2026) -
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
di: Hu, Xing, et al.
Pubblicazione: (2024)