An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
Fuente:
arXiv
Saved in:
| Main Authors: | Steinmetz, Cody, Childress, Gavin, Herbst, Aaron, Jones, Gavin, Singh, Jasdeep, Vang, Eli, Weinstock, Keagan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synergistic Simulations: Multi-Agent Problem Solving with Large Language Models
by: Sprigler, Asher, et al.
Published: (2024)
by: Sprigler, Asher, et al.
Published: (2024)
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
by: Ma, Shuming, et al.
Published: (2024)
by: Ma, Shuming, et al.
Published: (2024)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
Density Adaptive Attention is All You Need: Robust Parameter-Efficient Fine-Tuning Across Multiple Modalities
by: Ioannides, Georgios, et al.
Published: (2024)
by: Ioannides, Georgios, et al.
Published: (2024)
BitNet b1.58 2B4T Technical Report
by: Ma, Shuming, et al.
Published: (2025)
by: Ma, Shuming, et al.
Published: (2025)
Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity
by: Zhang, Di, et al.
Published: (2026)
by: Zhang, Di, et al.
Published: (2026)
The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
by: Ji, Ke, et al.
Published: (2025)
by: Ji, Ke, et al.
Published: (2025)
Not All Documents Are What You Need for Extracting Instruction Tuning Data
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
by: Nielsen, Jacob, et al.
Published: (2024)
by: Nielsen, Jacob, et al.
Published: (2024)
1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs
by: Wang, Jinheng, et al.
Published: (2024)
by: Wang, Jinheng, et al.
Published: (2024)
Contrast Is All You Need
by: Kilic, Burak, et al.
Published: (2023)
by: Kilic, Burak, et al.
Published: (2023)
COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning
by: Bai, Yuelin, et al.
Published: (2024)
by: Bai, Yuelin, et al.
Published: (2024)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
BitDelta: Your Fine-Tune May Only Be Worth One Bit
by: Liu, James, et al.
Published: (2024)
by: Liu, James, et al.
Published: (2024)
All You Need is One: Capsule Prompt Tuning with a Single Vector
by: Liu, Yiyang, et al.
Published: (2025)
by: Liu, Yiyang, et al.
Published: (2025)
Rho-1: Not All Tokens Are What You Need
by: Lin, Zhenghao, et al.
Published: (2024)
by: Lin, Zhenghao, et al.
Published: (2024)
Training on the Benchmark Is Not All You Need
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
by: Nielsen, Jacob, et al.
Published: (2024)
by: Nielsen, Jacob, et al.
Published: (2024)
Not All Tokens Are What You Need In Thinking
by: Yuan, Hang, et al.
Published: (2025)
by: Yuan, Hang, et al.
Published: (2025)
Increasing the Thinking Budget is Not All You Need
by: Iacobacci, Ignacio, et al.
Published: (2025)
by: Iacobacci, Ignacio, et al.
Published: (2025)
EMA Is Not All You Need: Mapping the Boundary Between Structure and Content in Recurrent Context
by: Singh, Arth
Published: (2026)
by: Singh, Arth
Published: (2026)
More Agents Is All You Need
by: Li, Junyou, et al.
Published: (2024)
by: Li, Junyou, et al.
Published: (2024)
Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations
by: Chandra, Mohit, et al.
Published: (2025)
by: Chandra, Mohit, et al.
Published: (2025)
Agents Are All You Need for LLM Unlearning
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Setting Targets is All You Need:Improved Order Competitive Ratio for Online Selection
by: Chen, Liyan, et al.
Published: (2024)
by: Chen, Liyan, et al.
Published: (2024)
Tensor Product Attention Is All You Need
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026)
by: Zade, Saleh Zare, et al.
Published: (2026)
Grimoire is All You Need for Enhancing Large Language Models
by: Chen, Ding, et al.
Published: (2024)
by: Chen, Ding, et al.
Published: (2024)
Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)
by: Lall, Supriya, et al.
Published: (2025)
by: Lall, Supriya, et al.
Published: (2025)
Addition is All You Need for Energy-efficient Language Models
by: Luo, Hongyin, et al.
Published: (2024)
by: Luo, Hongyin, et al.
Published: (2024)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
by: Uzan, Omri, et al.
Published: (2024)
by: Uzan, Omri, et al.
Published: (2024)
Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need
by: Zhang, Bo-Wen, et al.
Published: (2024)
by: Zhang, Bo-Wen, et al.
Published: (2024)
Beyond Fixed Length: Bucket Pre-training is All You Need
by: Yang, Qing, et al.
Published: (2024)
by: Yang, Qing, et al.
Published: (2024)
Selected Languages are All You Need for Cross-lingual Truthfulness Transfer
by: Liu, Weihao, et al.
Published: (2024)
by: Liu, Weihao, et al.
Published: (2024)
From Performance to Purpose: A Sociotechnical Taxonomy for Evaluating Large Language Model Utility
by: Levinson, Gavin, et al.
Published: (2026)
by: Levinson, Gavin, et al.
Published: (2026)
Block Rotation is All You Need for MXFP4 Quantization
by: Shao, Yuantian, et al.
Published: (2025)
by: Shao, Yuantian, et al.
Published: (2025)
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
by: Wu, Zongqian, et al.
Published: (2025)
by: Wu, Zongqian, et al.
Published: (2025)
Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models
by: Cantini, Riccardo, et al.
Published: (2025)
by: Cantini, Riccardo, et al.
Published: (2025)
Is Grep All You Need? How Agent Harnesses Reshape Agentic Search
by: Sen, Sahil, et al.
Published: (2026)
by: Sen, Sahil, et al.
Published: (2026)
Similar Items
-
Synergistic Simulations: Multi-Agent Problem Solving with Large Language Models
by: Sprigler, Asher, et al.
Published: (2024) -
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
by: Ma, Shuming, et al.
Published: (2024) -
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
by: Li, Pengyi, et al.
Published: (2025) -
Density Adaptive Attention is All You Need: Robust Parameter-Efficient Fine-Tuning Across Multiple Modalities
by: Ioannides, Georgios, et al.
Published: (2024) -
BitNet b1.58 2B4T Technical Report
by: Ma, Shuming, et al.
Published: (2025)