BitNet a4.8: 4-bit Activations for 1-bit LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Hongyu, Ma, Shuming, Wei, Furu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
BitNet b1.58 2B4T Technical Report
von: Ma, Shuming, et al.
Veröffentlicht: (2025)
von: Ma, Shuming, et al.
Veröffentlicht: (2025)
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
von: Ma, Shuming, et al.
Veröffentlicht: (2024)
von: Ma, Shuming, et al.
Veröffentlicht: (2024)
BitNet Distillation
von: Wu, Xun, et al.
Veröffentlicht: (2025)
von: Wu, Xun, et al.
Veröffentlicht: (2025)
1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs
von: Wang, Jinheng, et al.
Veröffentlicht: (2024)
von: Wang, Jinheng, et al.
Veröffentlicht: (2024)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity
von: Zhang, Di, et al.
Veröffentlicht: (2026)
von: Zhang, Di, et al.
Veröffentlicht: (2026)
Q-Sparse: All Large Language Models can be Fully Sparsely-Activated
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
AAAC: Activation-Aware Adaptive Codebooks for 4-bit LLM Weight Quantization
von: IslamBouli, Beshr, et al.
Veröffentlicht: (2026)
von: IslamBouli, Beshr, et al.
Veröffentlicht: (2026)
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
Auto-ICL: In-Context Learning without Human Supervision
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
A Word is Worth 4-bit: Efficient Log Parsing with Binary Coded Decimal Recognition
von: Srivastava, Prerak, et al.
Veröffentlicht: (2025)
von: Srivastava, Prerak, et al.
Veröffentlicht: (2025)
Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
von: Blumenberg, Patrick, et al.
Veröffentlicht: (2025)
von: Blumenberg, Patrick, et al.
Veröffentlicht: (2025)
Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
von: Wang, Jinheng, et al.
Veröffentlicht: (2025)
von: Wang, Jinheng, et al.
Veröffentlicht: (2025)
AMAQ: Adaptive Mixed-bit Activation Quantization for Collaborative Parameter Efficient Fine-tuning
von: Song, Yurun, et al.
Veröffentlicht: (2025)
von: Song, Yurun, et al.
Veröffentlicht: (2025)
iFairy: the First 2-bit Complex LLM with All Parameters in $\{\pm1, \pm i\}$
von: Wang, Feiyu, et al.
Veröffentlicht: (2025)
von: Wang, Feiyu, et al.
Veröffentlicht: (2025)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
von: Jia, Jinda, et al.
Veröffentlicht: (2024)
von: Jia, Jinda, et al.
Veröffentlicht: (2024)
1-Bit FQT: Pushing the Limit of Fully Quantized Training to 1-bit
von: Gao, Chang, et al.
Veröffentlicht: (2024)
von: Gao, Chang, et al.
Veröffentlicht: (2024)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
von: Dong, Peijie, et al.
Veröffentlicht: (2024)
von: Dong, Peijie, et al.
Veröffentlicht: (2024)
A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms
von: Gong, Ruihao, et al.
Veröffentlicht: (2024)
von: Gong, Ruihao, et al.
Veröffentlicht: (2024)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
von: Yankun, Hong, et al.
Veröffentlicht: (2025)
von: Yankun, Hong, et al.
Veröffentlicht: (2025)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
any4: Learned 4-bit Numeric Representation for LLMs
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2025)
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2025)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
von: Li, Jinhao, et al.
Veröffentlicht: (2023)
von: Li, Jinhao, et al.
Veröffentlicht: (2023)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
von: Liu, Zechun, et al.
Veröffentlicht: (2025)
von: Liu, Zechun, et al.
Veröffentlicht: (2025)
OneBit: Towards Extremely Low-bit Large Language Models
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2024)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2024)
PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
von: Zhu, Dawei, et al.
Veröffentlicht: (2023)
von: Zhu, Dawei, et al.
Veröffentlicht: (2023)
MAGNET: Autonomous Expert Model Generation via Decentralized Autoresearch and BitNet Training
von: Kim, Yongwan, et al.
Veröffentlicht: (2026)
von: Kim, Yongwan, et al.
Veröffentlicht: (2026)
Memory-Efficient 4-bit Preconditioned Stochastic Optimization
von: Li, Jingyang, et al.
Veröffentlicht: (2024)
von: Li, Jingyang, et al.
Veröffentlicht: (2024)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
1.58-bit FLUX
von: Yang, Chenglin, et al.
Veröffentlicht: (2024)
von: Yang, Chenglin, et al.
Veröffentlicht: (2024)
Matmul or No Matmul in the Era of 1-bit LLMs
von: Malekar, Jinendra, et al.
Veröffentlicht: (2024)
von: Malekar, Jinendra, et al.
Veröffentlicht: (2024)
Mixture of LoRA Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs
von: Zhou, Zhaojing, et al.
Veröffentlicht: (2025)
von: Zhou, Zhaojing, et al.
Veröffentlicht: (2025)
bit2bit: 1-bit quanta video reconstruction via self-supervised photon prediction
von: Liu, Yehe, et al.
Veröffentlicht: (2024)
von: Liu, Yehe, et al.
Veröffentlicht: (2024)
Thinking Augmented Pre-training
von: Wang, Liang, et al.
Veröffentlicht: (2025)
von: Wang, Liang, et al.
Veröffentlicht: (2025)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
von: Wang, YiFeng, et al.
Veröffentlicht: (2026)
von: Wang, YiFeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
von: Wang, Hongyu, et al.
Veröffentlicht: (2025) -
BitNet b1.58 2B4T Technical Report
von: Ma, Shuming, et al.
Veröffentlicht: (2025) -
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
von: Ma, Shuming, et al.
Veröffentlicht: (2024) -
BitNet Distillation
von: Wu, Xun, et al.
Veröffentlicht: (2025) -
1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs
von: Wang, Jinheng, et al.
Veröffentlicht: (2024)