BitDelta: Your Fine-Tune May Only Be Worth One Bit
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, James, Xiao, Guangxuan, Li, Kai, Lee, Jason D., Han, Song, Dao, Tri, Cai, Tianle |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
por: Cai, Tianle, et al.
Publicado: (2024)
por: Cai, Tianle, et al.
Publicado: (2024)
OneBit: Towards Extremely Low-bit Large Language Models
por: Xu, Yuzhuang, et al.
Publicado: (2024)
por: Xu, Yuzhuang, et al.
Publicado: (2024)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
por: Steinmetz, Cody, et al.
Publicado: (2025)
por: Steinmetz, Cody, et al.
Publicado: (2025)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
por: Lee, Banseok, et al.
Publicado: (2025)
por: Lee, Banseok, et al.
Publicado: (2025)
BitNet Distillation
por: Wu, Xun, et al.
Publicado: (2025)
por: Wu, Xun, et al.
Publicado: (2025)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
por: Zhou, Zikai, et al.
Publicado: (2025)
por: Zhou, Zikai, et al.
Publicado: (2025)
BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
por: Du, Dayou, et al.
Publicado: (2024)
por: Du, Dayou, et al.
Publicado: (2024)
Optimizing Mixture of Block Attention
por: Xiao, Guangxuan, et al.
Publicado: (2025)
por: Xiao, Guangxuan, et al.
Publicado: (2025)
Bit Blasting Probabilistic Programs
por: Garg, Poorva, et al.
Publicado: (2023)
por: Garg, Poorva, et al.
Publicado: (2023)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
por: Hu, Xing, et al.
Publicado: (2024)
por: Hu, Xing, et al.
Publicado: (2024)
Multi-Bit Distortion-Free Watermarking for Large Language Models
por: Boroujeny, Massieh Kordi, et al.
Publicado: (2024)
por: Boroujeny, Massieh Kordi, et al.
Publicado: (2024)
Mixed-Precision Graph Neural Quantization for Low Bit Large Language Models
por: Liu, Wanlong, et al.
Publicado: (2025)
por: Liu, Wanlong, et al.
Publicado: (2025)
Bit-Vector CHC Solving for Binary Analysis and Binary Analysis for Bit-Vector CHC Solving
por: Bembenek, Aaron, et al.
Publicado: (2026)
por: Bembenek, Aaron, et al.
Publicado: (2026)
Equational Bit-Vector Solving via Strong Gröbner Bases
por: Song, Jiaxin, et al.
Publicado: (2024)
por: Song, Jiaxin, et al.
Publicado: (2024)
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
por: Chee, Jerry, et al.
Publicado: (2023)
por: Chee, Jerry, et al.
Publicado: (2023)
Bit-level BPE: Below the byte boundary
por: Moon, Sangwhan, et al.
Publicado: (2025)
por: Moon, Sangwhan, et al.
Publicado: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
por: Du, Dayou, et al.
Publicado: (2025)
por: Du, Dayou, et al.
Publicado: (2025)
Marking: Visual Grading with Highlighting Errors and Annotating Missing Bits
por: Sonkar, Shashank, et al.
Publicado: (2024)
por: Sonkar, Shashank, et al.
Publicado: (2024)
Reward Collapse in Aligning Large Language Models
por: Song, Ziang, et al.
Publicado: (2023)
por: Song, Ziang, et al.
Publicado: (2023)
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
por: Ben-Zaken, Elad, et al.
Publicado: (2021)
por: Ben-Zaken, Elad, et al.
Publicado: (2021)
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
por: Daliri, Majid, et al.
Publicado: (2024)
por: Daliri, Majid, et al.
Publicado: (2024)
BitMar: Low-Bit Multimodal Fusion with Episodic Memory for Edge Devices
por: Aman, Euhid, et al.
Publicado: (2025)
por: Aman, Euhid, et al.
Publicado: (2025)
To be Continuous, or to be Discrete, Those are Bits of Questions
por: Wang, Yiran, et al.
Publicado: (2024)
por: Wang, Yiran, et al.
Publicado: (2024)
Efficient Streaming Language Models with Attention Sinks
por: Xiao, Guangxuan, et al.
Publicado: (2023)
por: Xiao, Guangxuan, et al.
Publicado: (2023)
Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning
por: Fang, Yangui, et al.
Publicado: (2025)
por: Fang, Yangui, et al.
Publicado: (2025)
Dr. SoW: Density Ratio of Strong-over-weak LLMs for Reducing the Cost of Human Annotation in Preference Tuning
por: Xu, Guangxuan, et al.
Publicado: (2024)
por: Xu, Guangxuan, et al.
Publicado: (2024)
BitLM: Unlocking Multi-Token Language Generation with Bitwise Continuous Diffusion
por: Zhuang, Shaobin, et al.
Publicado: (2026)
por: Zhuang, Shaobin, et al.
Publicado: (2026)
XAttention: Block Sparse Attention with Antidiagonal Scoring
por: Xu, Ruyi, et al.
Publicado: (2025)
por: Xu, Ruyi, et al.
Publicado: (2025)
Majority Bit-Aware Watermarking For Large Language Models
por: Xu, Jiahao, et al.
Publicado: (2025)
por: Xu, Jiahao, et al.
Publicado: (2025)
FrameQuant: Flexible Low-Bit Quantization for Transformers
por: Adepu, Harshavardhan, et al.
Publicado: (2024)
por: Adepu, Harshavardhan, et al.
Publicado: (2024)
Learning to Prioritize IT Tickets: A Comparative Evaluation of Embedding-based Approaches and Fine-Tuned Transformer Models
por: LÊ, Minh Tri, et al.
Publicado: (2025)
por: LÊ, Minh Tri, et al.
Publicado: (2025)
BitNet b1.58 2B4T Technical Report
por: Ma, Shuming, et al.
Publicado: (2025)
por: Ma, Shuming, et al.
Publicado: (2025)
H1B-KV: Hybrid One-Bit Caches for Memory-Efficient Large Language Model Inference
por: Vejendla, Harshil
Publicado: (2025)
por: Vejendla, Harshil
Publicado: (2025)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
por: Yang, Haoqi, et al.
Publicado: (2025)
por: Yang, Haoqi, et al.
Publicado: (2025)
A New Pipeline For Generating Instruction Dataset via RAG and Self Fine-Tuning
por: Song, Chih-Wei, et al.
Publicado: (2024)
por: Song, Chih-Wei, et al.
Publicado: (2024)
AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models
por: Li, Zeyu, et al.
Publicado: (2025)
por: Li, Zeyu, et al.
Publicado: (2025)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
por: Zandieh, Amir, et al.
Publicado: (2024)
por: Zandieh, Amir, et al.
Publicado: (2024)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
por: Zhang, Chao, et al.
Publicado: (2025)
por: Zhang, Chao, et al.
Publicado: (2025)
STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
por: Dong, Peijie, et al.
Publicado: (2024)
por: Dong, Peijie, et al.
Publicado: (2024)
Hardware-Efficient Attention for Fast Decoding
por: Zadouri, Ted, et al.
Publicado: (2025)
por: Zadouri, Ted, et al.
Publicado: (2025)
Ejemplares similares
-
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
por: Cai, Tianle, et al.
Publicado: (2024) -
OneBit: Towards Extremely Low-bit Large Language Models
por: Xu, Yuzhuang, et al.
Publicado: (2024) -
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
por: Steinmetz, Cody, et al.
Publicado: (2025) -
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
por: Lee, Banseok, et al.
Publicado: (2025) -
BitNet Distillation
por: Wu, Xun, et al.
Publicado: (2025)