DB-LLM: Accurate Dual-Binarization for Efficient LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Hong, Lv, Chengtao, Ding, Liang, Qin, Haotong, Zhou, Xiabin, Ding, Yifu, Liu, Xuebo, Zhang, Min, Guo, Jinyang, Liu, Xianglong, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BWTA: Accurate and Efficient Binarized Transformer by Algorithm-Hardware Co-design
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
PTQ4SAM: Post-Training Quantization for Segment Anything
von: Lv, Chengtao, et al.
Veröffentlicht: (2024)
von: Lv, Chengtao, et al.
Veröffentlicht: (2024)
BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models
von: Zheng, Xingyu, et al.
Veröffentlicht: (2024)
von: Zheng, Xingyu, et al.
Veröffentlicht: (2024)
A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms
von: Gong, Ruihao, et al.
Veröffentlicht: (2024)
von: Gong, Ruihao, et al.
Veröffentlicht: (2024)
BiVM: Accurate Binarized Neural Network for Efficient Video Matting
von: Qin, Haotong, et al.
Veröffentlicht: (2025)
von: Qin, Haotong, et al.
Veröffentlicht: (2025)
Progressive Binarization with Semi-Structured Pruning for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment
von: Yang, Ge, et al.
Veröffentlicht: (2024)
von: Yang, Ge, et al.
Veröffentlicht: (2024)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
QVGen: Pushing the Limit of Quantized Video Generative Models
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
Dynamic Parallel Tree Search for Efficient LLM Reasoning
von: Ding, Yifu, et al.
Veröffentlicht: (2025)
von: Ding, Yifu, et al.
Veröffentlicht: (2025)
Self-Powered LLM Modality Expansion for Large Speech-Text Models
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
First-Order Error Matters: Accurate Compensation for Quantized Large Language Models
von: Zheng, Xingyu, et al.
Veröffentlicht: (2025)
von: Zheng, Xingyu, et al.
Veröffentlicht: (2025)
An empirical study of LLaMA3 quantization: from LLMs to MLLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
ARB-LLM: Alternating Refined Binarizations for Large Language Models
von: Li, Zhiteng, et al.
Veröffentlicht: (2024)
von: Li, Zhiteng, et al.
Veröffentlicht: (2024)
Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
QVD: Post-training Quantization for Video Diffusion Models
von: Tian, Shilong, et al.
Veröffentlicht: (2024)
von: Tian, Shilong, et al.
Veröffentlicht: (2024)
Building Accurate Translation-Tailored LLMs with Language Aware Instruction Tuning
von: Zan, Changtong, et al.
Veröffentlicht: (2024)
von: Zan, Changtong, et al.
Veröffentlicht: (2024)
Exploring and Enhancing the Transfer of Distribution in Knowledge Distillation for Autoregressive Language Models
von: Rao, Jun, et al.
Veröffentlicht: (2024)
von: Rao, Jun, et al.
Veröffentlicht: (2024)
Revisiting Demonstration Selection Strategies in In-Context Learning
von: Peng, Keqin, et al.
Veröffentlicht: (2024)
von: Peng, Keqin, et al.
Veröffentlicht: (2024)
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
von: Qin, Haotong, et al.
Veröffentlicht: (2024)
von: Qin, Haotong, et al.
Veröffentlicht: (2024)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
von: Gong, Ruihao, et al.
Veröffentlicht: (2024)
von: Gong, Ruihao, et al.
Veröffentlicht: (2024)
QuantSR+: Pushing the Limit of Quantized Image Super-Resolution Networks
von: Qin, Haotong, et al.
Veröffentlicht: (2026)
von: Qin, Haotong, et al.
Veröffentlicht: (2026)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Entropy-Guided Watermarking for LLMs: A Test-Time Framework for Robust and Traceable Text Generation
von: Cai, Shizhan, et al.
Veröffentlicht: (2025)
von: Cai, Shizhan, et al.
Veröffentlicht: (2025)
AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration
von: Wang, Zhexuan, et al.
Veröffentlicht: (2025)
von: Wang, Zhexuan, et al.
Veröffentlicht: (2025)
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
Binarized Diffusion Model for Image Super-Resolution
von: Chen, Zheng, et al.
Veröffentlicht: (2024)
von: Chen, Zheng, et al.
Veröffentlicht: (2024)
RaBiT: Residual-Aware Binarization Training for Accurate and Efficient LLMs
von: You, Youngcheon, et al.
Veröffentlicht: (2026)
von: You, Youngcheon, et al.
Veröffentlicht: (2026)
BiDense: Binarization for Dense Prediction
von: Yin, Rui, et al.
Veröffentlicht: (2024)
von: Yin, Rui, et al.
Veröffentlicht: (2024)
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
von: Lu, Qingyu, et al.
Veröffentlicht: (2025)
von: Lu, Qingyu, et al.
Veröffentlicht: (2025)
AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaboration
von: Tian, Chunhao, et al.
Veröffentlicht: (2025)
von: Tian, Chunhao, et al.
Veröffentlicht: (2025)
BiDM: Pushing the Limit of Quantization for Diffusion Models
von: Zheng, Xingyu, et al.
Veröffentlicht: (2024)
von: Zheng, Xingyu, et al.
Veröffentlicht: (2024)
Intention Analysis Makes LLMs A Good Jailbreak Defender
von: Zhang, Yuqi, et al.
Veröffentlicht: (2024)
von: Zhang, Yuqi, et al.
Veröffentlicht: (2024)
VORTA: Efficient Video Diffusion via Routing Sparse Attention
von: Sun, Wenhao, et al.
Veröffentlicht: (2025)
von: Sun, Wenhao, et al.
Veröffentlicht: (2025)
Edit Once, Update Everywhere: A Simple Framework for Cross-Lingual Knowledge Synchronization in LLMs
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
PANDA: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation
von: Zhong, Qihuang, et al.
Veröffentlicht: (2022)
von: Zhong, Qihuang, et al.
Veröffentlicht: (2022)
Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language Models
von: Xiao, Yisong, et al.
Veröffentlicht: (2025)
von: Xiao, Yisong, et al.
Veröffentlicht: (2025)
Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BWTA: Accurate and Efficient Binarized Transformer by Algorithm-Hardware Co-design
von: Ding, Yifu, et al.
Veröffentlicht: (2026) -
PTQ4SAM: Post-Training Quantization for Segment Anything
von: Lv, Chengtao, et al.
Veröffentlicht: (2024) -
BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models
von: Zheng, Xingyu, et al.
Veröffentlicht: (2024) -
A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms
von: Gong, Ruihao, et al.
Veröffentlicht: (2024) -
BiVM: Accurate Binarized Neural Network for Efficient Video Matting
von: Qin, Haotong, et al.
Veröffentlicht: (2025)