1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jinheng, Zhou, Hansong, Song, Ting, Mao, Shaoguang, Ma, Shuming, Wang, Hongyu, Xia, Yan, Wei, Furu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BitNet b1.58 2B4T Technical Report
by: Ma, Shuming, et al.
Published: (2025)
by: Ma, Shuming, et al.
Published: (2025)
BitNet a4.8: 4-bit Activations for 1-bit LLMs
by: Wang, Hongyu, et al.
Published: (2024)
by: Wang, Hongyu, et al.
Published: (2024)
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity
by: Zhang, Di, et al.
Published: (2026)
by: Zhang, Di, et al.
Published: (2026)
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
by: Ma, Shuming, et al.
Published: (2024)
by: Ma, Shuming, et al.
Published: (2024)
BitNet Distillation
by: Wu, Xun, et al.
Published: (2025)
by: Wu, Xun, et al.
Published: (2025)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
by: Nielsen, Jacob, et al.
Published: (2024)
by: Nielsen, Jacob, et al.
Published: (2024)
BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
by: Nielsen, Jacob, et al.
Published: (2024)
by: Nielsen, Jacob, et al.
Published: (2024)
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
by: Nielsen, Jacob, et al.
Published: (2025)
by: Nielsen, Jacob, et al.
Published: (2025)
Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
by: Wang, Jinheng, et al.
Published: (2025)
by: Wang, Jinheng, et al.
Published: (2025)
ViT-1.58b: Mobile Vision Transformers in the 1-bit Era
by: Yuan, Zhengqing, et al.
Published: (2024)
by: Yuan, Zhengqing, et al.
Published: (2024)
1.58-bit FLUX
by: Yang, Chenglin, et al.
Published: (2024)
by: Yang, Chenglin, et al.
Published: (2024)
BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
by: Zhang, Wenlun, et al.
Published: (2025)
by: Zhang, Wenlun, et al.
Published: (2025)
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
by: Kawamura, Masaya, et al.
Published: (2025)
by: Kawamura, Masaya, et al.
Published: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
by: Geens, Robin, et al.
Published: (2026)
by: Geens, Robin, et al.
Published: (2026)
Ultra-Quantisation: Efficient Embedding Search via 1.58-bit Encodings
by: Connor, Richard, et al.
Published: (2025)
by: Connor, Richard, et al.
Published: (2025)
MAGNET: Autonomous Expert Model Generation via Decentralized Autoresearch and BitNet Training
by: Kim, Yongwan, et al.
Published: (2026)
by: Kim, Yongwan, et al.
Published: (2026)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
by: Steinmetz, Cody, et al.
Published: (2025)
by: Steinmetz, Cody, et al.
Published: (2025)
Q-Sparse: All Large Language Models can be Fully Sparsely-Activated
by: Wang, Hongyu, et al.
Published: (2024)
by: Wang, Hongyu, et al.
Published: (2024)
1-Bit FQT: Pushing the Limit of Fully Quantized Training to 1-bit
by: Gao, Chang, et al.
Published: (2024)
by: Gao, Chang, et al.
Published: (2024)
Nearly Lossless Adaptive Bit Switching
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Hybrid Gated Flow (HGF): Stabilizing 1.58-bit LLMs via Selective Low-Rank Correction
by: Pizzo, David Alejandro Trejo
Published: (2026)
by: Pizzo, David Alejandro Trejo
Published: (2026)
BitsFusion: 1.99 bits Weight Quantization of Diffusion Model
by: Sui, Yang, et al.
Published: (2024)
by: Sui, Yang, et al.
Published: (2024)
IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs
by: Mao, Yuzhen, et al.
Published: (2024)
by: Mao, Yuzhen, et al.
Published: (2024)
Massive MIMO-ISAC System With 1-Bit ADCs/DACs
by: Wang, Bowen, et al.
Published: (2024)
by: Wang, Bowen, et al.
Published: (2024)
Latent-Space Mean-Field Theory for Deep BitNet-like Training: Constrained Gradient Flows with Smooth Quantization and STE Limits
by: Kim, Dongwon, et al.
Published: (2025)
by: Kim, Dongwon, et al.
Published: (2025)
K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic Reasoning
by: Zhang, Yadong, et al.
Published: (2024)
by: Zhang, Yadong, et al.
Published: (2024)
Documentation of sediment core M39/1_58-1 (M39058-1)
by: Hsieh, Chia-Yun, et al.
Published: (2016)
by: Hsieh, Chia-Yun, et al.
Published: (2016)
Distributed Inference Performance Optimization for LLMs on CPUs
by: He, Pujiang, et al.
Published: (2024)
by: He, Pujiang, et al.
Published: (2024)
Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration
by: Wang, Zhenhailong, et al.
Published: (2023)
by: Wang, Zhenhailong, et al.
Published: (2023)
ALYMPICS: LLM Agents Meet Game Theory -- Exploring Strategic Decision-Making with AI Agents
by: Mao, Shaoguang, et al.
Published: (2023)
by: Mao, Shaoguang, et al.
Published: (2023)
Fast and Compact Tsetlin Machine Inference on CPUs Using Instruction-Level Optimization
by: Zeng, Yefan, et al.
Published: (2025)
by: Zeng, Yefan, et al.
Published: (2025)
Empirical Study of Large Language Models as Automated Essay Scoring Tools in English Composition__Taking TOEFL Independent Writing Task for Example
by: Xia, Wei, et al.
Published: (2024)
by: Xia, Wei, et al.
Published: (2024)
Enhancing Language Model Rationality with Bi-Directional Deliberation Reasoning
by: Zhang, Yadong, et al.
Published: (2024)
by: Zhang, Yadong, et al.
Published: (2024)
Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models
by: Wu, Wenshan, et al.
Published: (2024)
by: Wu, Wenshan, et al.
Published: (2024)
Meta Reasoning for Large Language Models
by: Gao, Peizhong, et al.
Published: (2024)
by: Gao, Peizhong, et al.
Published: (2024)
RAWIC: Bit-Depth Adaptive Lossless Raw Image Compression
by: Zheng, Chunhang, et al.
Published: (2026)
by: Zheng, Chunhang, et al.
Published: (2026)
Inference Performance Optimization for Large Language Models on CPUs
by: He, Pujiang, et al.
Published: (2024)
by: He, Pujiang, et al.
Published: (2024)
The State, Market Economy, and Transition
by: Wang Shaoguang
by: Wang Shaoguang
Similar Items
-
BitNet b1.58 2B4T Technical Report
by: Ma, Shuming, et al.
Published: (2025) -
BitNet a4.8: 4-bit Activations for 1-bit LLMs
by: Wang, Hongyu, et al.
Published: (2024) -
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
by: Wang, Hongyu, et al.
Published: (2025) -
Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity
by: Zhang, Di, et al.
Published: (2026) -
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
by: Ma, Shuming, et al.
Published: (2024)