TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jinhee, Yoon, Seoyeon, Lee, Taeho, Lee, Joo Chan, Jeon, Kang Eun, Ko, Jong Hwan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MSQ: Memory-Efficient Bit Sparsification Quantization
by: Han, Seokho, et al.
Published: (2025)
by: Han, Seokho, et al.
Published: (2025)
Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
by: Kim, Jinhee, et al.
Published: (2025)
by: Kim, Jinhee, et al.
Published: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
by: Wang, Dongwei, et al.
Published: (2026)
by: Wang, Dongwei, et al.
Published: (2026)
Row-Column Hybrid Grouping for Fault-Resilient Multi-Bit Weight Representation on IMC Arrays
by: Jeon, Kang Eun, et al.
Published: (2025)
by: Jeon, Kang Eun, et al.
Published: (2025)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing Architectures
by: Kang, Do Yeong, et al.
Published: (2025)
by: Kang, Do Yeong, et al.
Published: (2025)
FrameQuant: Flexible Low-Bit Quantization for Transformers
by: Adepu, Harshavardhan, et al.
Published: (2024)
by: Adepu, Harshavardhan, et al.
Published: (2024)
RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNs
by: Ko, Hanum, et al.
Published: (2026)
by: Ko, Hanum, et al.
Published: (2026)
Optimized Minimal 3D Gaussian Splatting
by: Lee, Joo Chan, et al.
Published: (2025)
by: Lee, Joo Chan, et al.
Published: (2025)
Single-step Diffusion for Image Compression at Ultra-Low Bitrates
by: Park, Chanung, et al.
Published: (2025)
by: Park, Chanung, et al.
Published: (2025)
Low-Rank Compression for IMC Arrays
by: Jeon, Kang Eun, et al.
Published: (2025)
by: Jeon, Kang Eun, et al.
Published: (2025)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
by: Zhao, Zhixiong, et al.
Published: (2025)
by: Zhao, Zhixiong, et al.
Published: (2025)
TruncFormer: Private LLM Inference Using Only Truncations
by: Yubeaton, Patrick, et al.
Published: (2024)
by: Yubeaton, Patrick, et al.
Published: (2024)
Continuous Memory Representation for Anomaly Detection
by: Lee, Joo Chan, et al.
Published: (2024)
by: Lee, Joo Chan, et al.
Published: (2024)
FAMES: Fast Approximate Multiplier Substitution for Mixed-Precision Quantized DNNs--Down to 2 Bits!
by: Ren, Yi, et al.
Published: (2024)
by: Ren, Yi, et al.
Published: (2024)
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
by: Chong, Hyochan, et al.
Published: (2026)
by: Chong, Hyochan, et al.
Published: (2026)
Compact 3D Gaussian Splatting for Static and Dynamic Radiance Fields
by: Lee, Joo Chan, et al.
Published: (2024)
by: Lee, Joo Chan, et al.
Published: (2024)
Compact 3D Gaussian Representation for Radiance Field
by: Lee, Joo Chan, et al.
Published: (2023)
by: Lee, Joo Chan, et al.
Published: (2023)
ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models
by: Yoon, Junho, et al.
Published: (2025)
by: Yoon, Junho, et al.
Published: (2025)
Scheduling Weight Transitions for Quantization-Aware Training
by: Lee, Junghyup, et al.
Published: (2024)
by: Lee, Junghyup, et al.
Published: (2024)
AccuQuant: Simulating Multiple Denoising Steps for Quantizing Diffusion Models
by: Lee, Seunghoon, et al.
Published: (2025)
by: Lee, Seunghoon, et al.
Published: (2025)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
by: Lee, Namyoon, et al.
Published: (2026)
by: Lee, Namyoon, et al.
Published: (2026)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
The Effects of Environmental Policy Stringency on Firms' Performances: Focusing on Korean Manufacturing Industry
by: Jinhee Lee, et al.
Published: (2024)
by: Jinhee Lee, et al.
Published: (2024)
Synthetic Data Generation for Phrase Break Prediction with Large Language Model
by: Lee, Hoyeon, et al.
Published: (2025)
by: Lee, Hoyeon, et al.
Published: (2025)
DuAL-Net: A Hybrid Framework for Alzheimer's Disease Prediction from Whole-Genome Sequencing via Local SNP Windows and Global Annotations
by: Lee, Eun Hye, et al.
Published: (2025)
by: Lee, Eun Hye, et al.
Published: (2025)
A novel deep learning model with transformer architectures to enable multi‐scale whole genome sequence analysis for Alzheimer's disease dementia prediction
by: Eun Hye Lee, et al.
Published: (2025)
by: Eun Hye Lee, et al.
Published: (2025)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
by: Lee, Banseok, et al.
Published: (2025)
by: Lee, Banseok, et al.
Published: (2025)
Efficient Reprogramming of Memristive Crossbars for DNNs: Weight Sorting and Bit Stucking
by: Farias, Matheus, et al.
Published: (2024)
by: Farias, Matheus, et al.
Published: (2024)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
Enhancing Empathy in Virtual Reality: An Embodied Approach to Mindset Modulation
by: Bae, Seoyeon, et al.
Published: (2024)
by: Bae, Seoyeon, et al.
Published: (2024)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
by: Becking, Daniel, et al.
Published: (2021)
by: Becking, Daniel, et al.
Published: (2021)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
by: Choi, Kanghyun, et al.
Published: (2024)
by: Choi, Kanghyun, et al.
Published: (2024)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
by: Park, Jungwoo, et al.
Published: (2025)
by: Park, Jungwoo, et al.
Published: (2025)
MASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision Quantization
by: Kim, Seeyeon, et al.
Published: (2026)
by: Kim, Seeyeon, et al.
Published: (2026)
F-3DGS: Factorized Coordinates and Representations for 3D Gaussian Splatting
by: Sun, Xiangyu, et al.
Published: (2024)
by: Sun, Xiangyu, et al.
Published: (2024)
Flexible Bit-Truncation Memory for Approximate Applications on the Edge
by: Oswald, William, et al.
Published: (2025)
by: Oswald, William, et al.
Published: (2025)
Attention-Propagation Network for Egocentric Heatmap to 3D Pose Lifting
by: Kang, Taeho, et al.
Published: (2024)
by: Kang, Taeho, et al.
Published: (2024)
Similar Items
-
MSQ: Memory-Efficient Bit Sparsification Quantization
by: Han, Seokho, et al.
Published: (2025) -
Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
by: Kim, Jinhee, et al.
Published: (2025) -
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
by: Wang, Dongwei, et al.
Published: (2026) -
Row-Column Hybrid Grouping for Fault-Resilient Multi-Bit Weight Representation on IMC Arrays
by: Jeon, Kang Eun, et al.
Published: (2025) -
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)