BitNet b1.58 2B4T Technical Report
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Shuming, Wang, Hongyu, Huang, Shaohan, Zhang, Xingxing, Hu, Ying, Song, Ting, Xia, Yan, Wei, Furu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs
von: Wang, Jinheng, et al.
Veröffentlicht: (2024)
von: Wang, Jinheng, et al.
Veröffentlicht: (2024)
BitNet Distillation
von: Wu, Xun, et al.
Veröffentlicht: (2025)
von: Wu, Xun, et al.
Veröffentlicht: (2025)
BitNet a4.8: 4-bit Activations for 1-bit LLMs
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity
von: Zhang, Di, et al.
Veröffentlicht: (2026)
von: Zhang, Di, et al.
Veröffentlicht: (2026)
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
von: Ma, Shuming, et al.
Veröffentlicht: (2024)
von: Ma, Shuming, et al.
Veröffentlicht: (2024)
BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
MH-MoE: Multi-Head Mixture-of-Experts
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
Q-Sparse: All Large Language Models can be Fully Sparsely-Activated
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
Adapting Large Language Models to Domains via Reading Comprehension
von: Cheng, Daixuan, et al.
Veröffentlicht: (2023)
von: Cheng, Daixuan, et al.
Veröffentlicht: (2023)
Mixture of LoRA Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
VibeVoice Technical Report
von: Peng, Zhiliang, et al.
Veröffentlicht: (2025)
von: Peng, Zhiliang, et al.
Veröffentlicht: (2025)
Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
von: Wang, Jinheng, et al.
Veröffentlicht: (2025)
von: Wang, Jinheng, et al.
Veröffentlicht: (2025)
Auto-ICL: In-Context Learning without Human Supervision
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
You Only Cache Once: Decoder-Decoder Architectures for Language Models
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
Textual Aesthetics in Large Language Models
von: Jiang, Lingjie, et al.
Veröffentlicht: (2024)
von: Jiang, Lingjie, et al.
Veröffentlicht: (2024)
Multi-Head Mixture-of-Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
On-Policy Context Distillation for Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
RedStone: Curating General, Code, Math, and QA Data for Large Language Models
von: Chang, Yaoyao, et al.
Veröffentlicht: (2024)
von: Chang, Yaoyao, et al.
Veröffentlicht: (2024)
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
Thinking Augmented Pre-training
von: Wang, Liang, et al.
Veröffentlicht: (2025)
von: Wang, Liang, et al.
Veröffentlicht: (2025)
Instruction Pre-Training: Language Models are Supervised Multitask Learners
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
Multilingual E5 Text Embeddings: A Technical Report
von: Wang, Liang, et al.
Veröffentlicht: (2024)
von: Wang, Liang, et al.
Veröffentlicht: (2024)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
von: Steinmetz, Cody, et al.
Veröffentlicht: (2025)
von: Steinmetz, Cody, et al.
Veröffentlicht: (2025)
LongReasonArena: A Long Reasoning Benchmark for Large Language Models
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
Online Experiential Learning for Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
Universal YOCO for Efficient Depth Scaling
von: Sun, Yutao, et al.
Veröffentlicht: (2026)
von: Sun, Yutao, et al.
Veröffentlicht: (2026)
QueST: Incentivizing LLMs to Generate Difficult Problems
von: Hu, Hanxu, et al.
Veröffentlicht: (2025)
von: Hu, Hanxu, et al.
Veröffentlicht: (2025)
Think Only When You Need with Large Hybrid-Reasoning Models
von: Jiang, Lingjie, et al.
Veröffentlicht: (2025)
von: Jiang, Lingjie, et al.
Veröffentlicht: (2025)
Black-Box On-Policy Distillation of Large Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2025)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2025)
On-Policy RL with Optimal Reward Baseline
von: Hao, Yaru, et al.
Veröffentlicht: (2025)
von: Hao, Yaru, et al.
Veröffentlicht: (2025)
Bootstrap Your Own Context Length
von: Wang, Liang, et al.
Veröffentlicht: (2024)
von: Wang, Liang, et al.
Veröffentlicht: (2024)
Kosmos-G: Generating Images in Context with Multimodal Large Language Models
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
Reward Reasoning Model
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
von: Jiang, Lingjie, et al.
Veröffentlicht: (2025)
von: Jiang, Lingjie, et al.
Veröffentlicht: (2025)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
MAGNET: Autonomous Expert Model Generation via Decentralized Autoresearch and BitNet Training
von: Kim, Yongwan, et al.
Veröffentlicht: (2026)
von: Kim, Yongwan, et al.
Veröffentlicht: (2026)
Code Aesthetics with Agentic Reward Feedback
von: Xiao, Bang, et al.
Veröffentlicht: (2025)
von: Xiao, Bang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs
von: Wang, Jinheng, et al.
Veröffentlicht: (2024) -
BitNet Distillation
von: Wu, Xun, et al.
Veröffentlicht: (2025) -
BitNet a4.8: 4-bit Activations for 1-bit LLMs
von: Wang, Hongyu, et al.
Veröffentlicht: (2024) -
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
von: Wang, Hongyu, et al.
Veröffentlicht: (2025) -
Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity
von: Zhang, Di, et al.
Veröffentlicht: (2026)