SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Jaewoo, Lin, Fangzhen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
von: Li, Ke, et al.
Veröffentlicht: (2026)
von: Li, Ke, et al.
Veröffentlicht: (2026)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
Verification of Bit-Flip Attacks against Quantized Neural Networks
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
Hetero-SplitEE: Split Learning of Neural Networks with Early Exits for Heterogeneous IoT Devices
von: Oda, Yuki, et al.
Veröffentlicht: (2025)
von: Oda, Yuki, et al.
Veröffentlicht: (2025)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
von: Chen, Han, et al.
Veröffentlicht: (2025)
von: Chen, Han, et al.
Veröffentlicht: (2025)
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
von: Zhang, Xi, et al.
Veröffentlicht: (2025)
von: Zhang, Xi, et al.
Veröffentlicht: (2025)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory
von: Chee, Jerry, et al.
Veröffentlicht: (2025)
von: Chee, Jerry, et al.
Veröffentlicht: (2025)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
von: Lee, Deokjae, et al.
Veröffentlicht: (2025)
von: Lee, Deokjae, et al.
Veröffentlicht: (2025)
PolarQuant: Quantizing KV Caches with Polar Transformation
von: Han, Insu, et al.
Veröffentlicht: (2025)
von: Han, Insu, et al.
Veröffentlicht: (2025)
Adversarial Attacks to Latent Representations of Distributed Neural Networks in Split Computing
von: Zhang, Milin, et al.
Veröffentlicht: (2023)
von: Zhang, Milin, et al.
Veröffentlicht: (2023)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
von: Tao, Qian, et al.
Veröffentlicht: (2024)
von: Tao, Qian, et al.
Veröffentlicht: (2024)
Why Do Some Inputs Break Low-Bit LLM Quantization?
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2025)
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2025)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
Anomaly Detection Based on Critical Paths for Deep Neural Networks
von: Zhao, Fangzhen, et al.
Veröffentlicht: (2025)
von: Zhao, Fangzhen, et al.
Veröffentlicht: (2025)
BaB-prob: Branch and Bound with Preactivation Splitting for Probabilistic Verification of Neural Networks
von: Wang, Fangji, et al.
Veröffentlicht: (2025)
von: Wang, Fangji, et al.
Veröffentlicht: (2025)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
von: Becking, Daniel, et al.
Veröffentlicht: (2021)
von: Becking, Daniel, et al.
Veröffentlicht: (2021)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization
von: Zhao, Jiayu, et al.
Veröffentlicht: (2026)
von: Zhao, Jiayu, et al.
Veröffentlicht: (2026)
Quantization Meets Reasoning: Exploring and Mitigating Degradation of Low-Bit LLMs in Mathematical Reasoning
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026)
QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals
von: Zhang, Nan, et al.
Veröffentlicht: (2026)
von: Zhang, Nan, et al.
Veröffentlicht: (2026)
Calibration and Transformation-Free Weight-Only LLMs Quantization via Dynamic Grouping
von: Zheng, Xinzhe, et al.
Veröffentlicht: (2025)
von: Zheng, Xinzhe, et al.
Veröffentlicht: (2025)
SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling
von: Guo, Yi, et al.
Veröffentlicht: (2025)
von: Guo, Yi, et al.
Veröffentlicht: (2025)
ASFL: An Adaptive Model Splitting and Resource Allocation Framework for Split Federated Learning
von: Meng, Chuiyang, et al.
Veröffentlicht: (2026)
von: Meng, Chuiyang, et al.
Veröffentlicht: (2026)
Revisiting RaBitQ and TurboQuant: A Symmetric Comparison of Methods, Theory, and Experiments
von: Gao, Jianyang, et al.
Veröffentlicht: (2026)
von: Gao, Jianyang, et al.
Veröffentlicht: (2026)
FP=xINT:Representing Neural Networks via Low-Bit Series Basis Functions
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
von: Liu, Wenyuan, et al.
Veröffentlicht: (2024)
von: Liu, Wenyuan, et al.
Veröffentlicht: (2024)
Tiled Bit Networks: Sub-Bit Neural Network Compression Through Reuse of Learnable Binary Vectors
von: Gorbett, Matt, et al.
Veröffentlicht: (2024)
von: Gorbett, Matt, et al.
Veröffentlicht: (2024)
HealSplit: Towards Self-Healing through Adversarial Distillation in Split Federated Learning
von: Xie, Yuhan, et al.
Veröffentlicht: (2025)
von: Xie, Yuhan, et al.
Veröffentlicht: (2025)
Low-bit Model Quantization for Deep Neural Networks: A Survey
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
von: Savkin, Semyon, et al.
Veröffentlicht: (2025)
von: Savkin, Semyon, et al.
Veröffentlicht: (2025)
Split Federated Learning Architectures for High-Accuracy and Low-Delay Model Training
von: Papageorgiou, Yiannis, et al.
Veröffentlicht: (2026)
von: Papageorgiou, Yiannis, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
von: Song, Jaewoo, et al.
Veröffentlicht: (2025) -
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
von: Li, Ke, et al.
Veröffentlicht: (2026) -
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025) -
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
von: Zhang, Tianao, et al.
Veröffentlicht: (2025) -
Verification of Bit-Flip Attacks against Quantized Neural Networks
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)