DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Jiajun, Wu, Jiajun, Gao, Yizhao, Ding, Yuhao, Tao, Chaofan, Li, Boyu, Tu, Fengbin, Cheng, Kwang-Ting, So, Hayden Kwok-Hay, Wong, Ngai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
von: Gao, Yizhao, et al.
Veröffentlicht: (2024)
von: Gao, Yizhao, et al.
Veröffentlicht: (2024)
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
von: Wu, Jiajun, et al.
Veröffentlicht: (2024)
von: Wu, Jiajun, et al.
Veröffentlicht: (2024)
Co-designing a Sub-millisecond Latency Event-based Eye Tracking System with Submanifold Sparse CNN
von: Zhang, Baoheng, et al.
Veröffentlicht: (2024)
von: Zhang, Baoheng, et al.
Veröffentlicht: (2024)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
Beyond Outliers: A Data-Free Layer-wise Mixed-Precision Quantization Approach Driven by Numerical and Structural Dual-Sensitivity
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2026)
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models
von: Feng, Weilun, et al.
Veröffentlicht: (2024)
von: Feng, Weilun, et al.
Veröffentlicht: (2024)
BPDQ: Bit-Plane Decomposition Quantization on a Variable Grid for Large Language Models
von: Chen, Junyu, et al.
Veröffentlicht: (2026)
von: Chen, Junyu, et al.
Veröffentlicht: (2026)
Outlier-Aware Training for Low-Bit Quantization of Structural Re-Parameterized Networks
von: Niu, Muqun, et al.
Veröffentlicht: (2024)
von: Niu, Muqun, et al.
Veröffentlicht: (2024)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023)
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023)
Mixed-Precision Graph Neural Quantization for Low Bit Large Language Models
von: Liu, Wanlong, et al.
Veröffentlicht: (2025)
von: Liu, Wanlong, et al.
Veröffentlicht: (2025)
6Bit-Diffusion: Inference-Time Mixed-Precision Quantization for Video Diffusion Models
von: Su, Rundong, et al.
Veröffentlicht: (2026)
von: Su, Rundong, et al.
Veröffentlicht: (2026)
AIS: Adaptive Importance Sampling for Quantized RL
von: Zhou, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhou, Jiajun, et al.
Veröffentlicht: (2026)
Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models
von: Ding, Xin, et al.
Veröffentlicht: (2025)
von: Ding, Xin, et al.
Veröffentlicht: (2025)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
Getting Free Bits Back from Rotational Symmetries in LLMs
von: He, Jiajun, et al.
Veröffentlicht: (2024)
von: He, Jiajun, et al.
Veröffentlicht: (2024)
GuiLoMo: Allocating Expert Number and Rank for LoRA-MoE via Bilevel Optimization with GuidedSelection Vectors
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2025)
Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2025)
Quantization Meets Reasoning: Exploring and Mitigating Degradation of Low-Bit LLMs in Mathematical Reasoning
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation
von: Feng, Weilun, et al.
Veröffentlicht: (2025)
von: Feng, Weilun, et al.
Veröffentlicht: (2025)
True 4-Bit Quantized Convolutional Neural Network Training on CPU: Achieving Full-Precision Parity
von: Tathe, Shivnath
Veröffentlicht: (2026)
von: Tathe, Shivnath
Veröffentlicht: (2026)
Bits for Privacy: Evaluating Post-Training Quantization via Membership Inference
von: Zhang, Chenxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Chenxiang, et al.
Veröffentlicht: (2025)
Verification of Bit-Flip Attacks against Quantized Neural Networks
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
von: Kang, Haidong, et al.
Veröffentlicht: (2025)
von: Kang, Haidong, et al.
Veröffentlicht: (2025)
Distributed Optimization with Finite Bit Adaptive Quantization for Efficient Communication and Precision Enhancement
von: Rikos, Apostolos I., et al.
Veröffentlicht: (2024)
von: Rikos, Apostolos I., et al.
Veröffentlicht: (2024)
LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
von: Mirzaei, Amir Reza, et al.
Veröffentlicht: (2025)
von: Mirzaei, Amir Reza, et al.
Veröffentlicht: (2025)
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
von: Kim, Jinhee, et al.
Veröffentlicht: (2025)
von: Kim, Jinhee, et al.
Veröffentlicht: (2025)
Mixed-Precision Quantization: Make the Best Use of Bits Where They Matter Most
von: Fang, Yiming, et al.
Veröffentlicht: (2024)
von: Fang, Yiming, et al.
Veröffentlicht: (2024)
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
von: Patarlapalli, Sai Babu, et al.
Veröffentlicht: (2026)
von: Patarlapalli, Sai Babu, et al.
Veröffentlicht: (2026)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2024)
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2024)
TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review
von: Chang, Yuan, et al.
Veröffentlicht: (2025)
von: Chang, Yuan, et al.
Veröffentlicht: (2025)
Why Do Some Inputs Break Low-Bit LLM Quantization?
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2025)
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2025)
ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models
von: Zeng, Chao, et al.
Veröffentlicht: (2024)
von: Zeng, Chao, et al.
Veröffentlicht: (2024)
CondiQuant: Condition Number Based Low-Bit Quantization for Image Super-Resolution
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
Every Bit Counts: A Theoretical Study of Precision-Expressivity Tradeoffs in Quantized Transformers
von: Chakrabarti, Sayak, et al.
Veröffentlicht: (2026)
von: Chakrabarti, Sayak, et al.
Veröffentlicht: (2026)
eXmY: A Data Type and Technique for Arbitrary Bit Precision Quantization
von: Agrawal, Aditya, et al.
Veröffentlicht: (2024)
von: Agrawal, Aditya, et al.
Veröffentlicht: (2024)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
von: Gao, Yizhao, et al.
Veröffentlicht: (2024) -
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
von: Wu, Jiajun, et al.
Veröffentlicht: (2024) -
Co-designing a Sub-millisecond Latency Event-based Eye Tracking System with Submanifold Sparse CNN
von: Zhang, Baoheng, et al.
Veröffentlicht: (2024) -
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
von: Huang, Xijie, et al.
Veröffentlicht: (2023) -
Beyond Outliers: A Data-Free Layer-wise Mixed-Precision Quantization Approach Driven by Numerical and Structural Dual-Sensitivity
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2026)