Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Hong, Wu, Decheng, Hu, Qiangqiang, Yu, Guanghua, Yang, Jinhai, Zhu, Jianchen, Liu, Xue, Wu, Dapeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tequila: Trapping-free Ternary Quantization for Large Language Models
von: Huang, Hong, et al.
Veröffentlicht: (2025)
von: Huang, Hong, et al.
Veröffentlicht: (2025)
Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis
von: Huang, Hong, et al.
Veröffentlicht: (2025)
von: Huang, Hong, et al.
Veröffentlicht: (2025)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
MSQ: Memory-Efficient Bit Sparsification Quantization
von: Han, Seokho, et al.
Veröffentlicht: (2025)
von: Han, Seokho, et al.
Veröffentlicht: (2025)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
von: Peng, Yanxin, et al.
Veröffentlicht: (2025)
von: Peng, Yanxin, et al.
Veröffentlicht: (2025)
EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diffusion Models
von: He, Yefei, et al.
Veröffentlicht: (2023)
von: He, Yefei, et al.
Veröffentlicht: (2023)
Time-Series Analysis on Edge-AI Hardware for Healthcare Monitoring
von: Hu, Jinhai
Veröffentlicht: (2025)
von: Hu, Jinhai
Veröffentlicht: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
One-Bit Quantization and Sparsification for Multiclass Linear Classification with Strong Regularization
von: Ghane, Reza, et al.
Veröffentlicht: (2024)
von: Ghane, Reza, et al.
Veröffentlicht: (2024)
AngelSlim: A more accessible, comprehensive, and efficient toolkit for large model compression
von: Cen, Rui, et al.
Veröffentlicht: (2026)
von: Cen, Rui, et al.
Veröffentlicht: (2026)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
von: Xie, Xilong, et al.
Veröffentlicht: (2025)
von: Xie, Xilong, et al.
Veröffentlicht: (2025)
Efficient Asynchronous Federated Learning with Sparsification and Quantization
von: Jia, Juncheng, et al.
Veröffentlicht: (2023)
von: Jia, Juncheng, et al.
Veröffentlicht: (2023)
JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration
von: Wang, Mingzi, et al.
Veröffentlicht: (2025)
von: Wang, Mingzi, et al.
Veröffentlicht: (2025)
IDPruner: Harmonizing Importance and Diversity in Visual Token Pruning for MLLMs
von: Tan, Yifan, et al.
Veröffentlicht: (2026)
von: Tan, Yifan, et al.
Veröffentlicht: (2026)
Rethinking Cross-Domain Evaluation for Face Forgery Detection with Semantic Fine-grained Alignment and Mixture-of-Experts
von: Luo, Yuhan, et al.
Veröffentlicht: (2026)
von: Luo, Yuhan, et al.
Veröffentlicht: (2026)
TernaryLM: Memory-Efficient Language Modeling via Native 1.5-Bit Quantization with Adaptive Layer-wise Scaling
von: Nargund, Nisharg, et al.
Veröffentlicht: (2026)
von: Nargund, Nisharg, et al.
Veröffentlicht: (2026)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
Sherry Ortner
von: Sergio Daniel López
Veröffentlicht: (2006)
von: Sergio Daniel López
Veröffentlicht: (2006)
DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network Inference
von: Zhou, Jiajun, et al.
Veröffentlicht: (2023)
von: Zhou, Jiajun, et al.
Veröffentlicht: (2023)
GLT-PEFT: Gated Lie-Tucker Parameter-Efficient Fine-Tuning for Alzheimer's Disease Diagnosis with Hippocampal Segmentation Pretraining
von: He, Guanghua, et al.
Veröffentlicht: (2026)
von: He, Guanghua, et al.
Veröffentlicht: (2026)
Training Language Models to Generate Text with Citations via Fine-grained Rewards
von: Huang, Chengyu, et al.
Veröffentlicht: (2024)
von: Huang, Chengyu, et al.
Veröffentlicht: (2024)
LoTA-QAF: Lossless Ternary Adaptation for Quantization-Aware Fine-Tuning
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
Efficient Distributed Training through Gradient Compression with Sparsification and Quantization Techniques
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
von: Lu, Yu-Chen, et al.
Veröffentlicht: (2025)
von: Lu, Yu-Chen, et al.
Veröffentlicht: (2025)
From Bits to Chips: An LLM-based Hardware-Aware Quantization Agent for Streamlined Deployment of LLMs
von: Deng, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Deng, Kaiyuan, et al.
Veröffentlicht: (2026)
Privacy-preserving formal concept analysis: A homomorphic encryption-based concept construction
von: Chen, Qiangqiang, et al.
Veröffentlicht: (2025)
von: Chen, Qiangqiang, et al.
Veröffentlicht: (2025)
SBVR: Summation of BitVector Representation for Efficient LLM Quantization
von: Bang, Wonjun, et al.
Veröffentlicht: (2025)
von: Bang, Wonjun, et al.
Veröffentlicht: (2025)
Invited Paper: BitMedViT: Ternary-Quantized Vision Transformer for Medical AI Assistants on the Edge
von: Walczak, Mikolaj, et al.
Veröffentlicht: (2025)
von: Walczak, Mikolaj, et al.
Veröffentlicht: (2025)
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
Scalable Graph Sparsification Techniques for Efficient Fraud Detection using GNNs
von: Wu, Ko-Wen
Veröffentlicht: (2026)
von: Wu, Ko-Wen
Veröffentlicht: (2026)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
von: Zhao, Maosen, et al.
Veröffentlicht: (2025)
von: Zhao, Maosen, et al.
Veröffentlicht: (2025)
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
von: Feng, Weilun, et al.
Veröffentlicht: (2025)
von: Feng, Weilun, et al.
Veröffentlicht: (2025)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
von: He, Junhui, et al.
Veröffentlicht: (2024)
von: He, Junhui, et al.
Veröffentlicht: (2024)
Distribution-Aware Hadamard Quantization for Hardware-Efficient Implicit Neural Representations
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
Stem: Rethinking Causal Information Flow in Sparse Attention
von: Niu, Lin, et al.
Veröffentlicht: (2026)
von: Niu, Lin, et al.
Veröffentlicht: (2026)
Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
von: Guo, Hang, et al.
Veröffentlicht: (2025)
von: Guo, Hang, et al.
Veröffentlicht: (2025)
Efficient Unbiased Sparsification
von: Barnes, Leighton, et al.
Veröffentlicht: (2024)
von: Barnes, Leighton, et al.
Veröffentlicht: (2024)
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
von: Lin, Haokun, et al.
Veröffentlicht: (2026)
von: Lin, Haokun, et al.
Veröffentlicht: (2026)
Gaussian Entropy Fields: Driving Adaptive Sparsity in 3D Gaussian Optimization
von: Kuang, Hong, et al.
Veröffentlicht: (2025)
von: Kuang, Hong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Tequila: Trapping-free Ternary Quantization for Large Language Models
von: Huang, Hong, et al.
Veröffentlicht: (2025) -
Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis
von: Huang, Hong, et al.
Veröffentlicht: (2025) -
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
von: Tang, Hanlin, et al.
Veröffentlicht: (2024) -
MSQ: Memory-Efficient Bit Sparsification Quantization
von: Han, Seokho, et al.
Veröffentlicht: (2025) -
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
von: Peng, Yanxin, et al.
Veröffentlicht: (2025)