SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Muyang, Lin, Yujun, Zhang, Zhekai, Cai, Tianle, Li, Xiuyu, Guo, Junxian, Xie, Enze, Meng, Chenlin, Zhu, Jun-Yan, Han, Song |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
by: Xie, Enze, et al.
Published: (2024)
by: Xie, Enze, et al.
Published: (2024)
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
by: He, Wenkun, et al.
Published: (2025)
by: He, Wenkun, et al.
Published: (2025)
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
by: Xie, Enze, et al.
Published: (2025)
by: Xie, Enze, et al.
Published: (2025)
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
by: Xu, Binxing, et al.
Published: (2026)
by: Xu, Binxing, et al.
Published: (2026)
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
by: Xi, Haocheng, et al.
Published: (2026)
by: Xi, Haocheng, et al.
Published: (2026)
LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
by: Lin, Yujun, et al.
Published: (2025)
by: Lin, Yujun, et al.
Published: (2025)
Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
by: Chen, Junyu, et al.
Published: (2024)
by: Chen, Junyu, et al.
Published: (2024)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
by: Xi, Haocheng, et al.
Published: (2025)
by: Xi, Haocheng, et al.
Published: (2025)
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
BitDelta: Your Fine-Tune May Only Be Worth One Bit
by: Liu, James, et al.
Published: (2024)
by: Liu, James, et al.
Published: (2024)
DistriFusion: Distributed Parallel Inference for High-Resolution Diffusion Models
by: Li, Muyang, et al.
Published: (2024)
by: Li, Muyang, et al.
Published: (2024)
Outlier-Aware Training for Low-Bit Quantization of Structural Re-Parameterized Networks
by: Niu, Muqun, et al.
Published: (2024)
by: Niu, Muqun, et al.
Published: (2024)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
OFA-Diffusion Compression: Compressing Diffusion Model in One-Shot Manner
by: Jiang, Haoyang, et al.
Published: (2026)
by: Jiang, Haoyang, et al.
Published: (2026)
DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Simple KNN-Based Outlier Detection Achieves Robust Clustering
by: Jiang, Tianle, et al.
Published: (2026)
by: Jiang, Tianle, et al.
Published: (2026)
Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
by: Lou, Aaron, et al.
Published: (2023)
by: Lou, Aaron, et al.
Published: (2023)
NIRECP: An ECC Processor Over 256‐Bit NIST Prime Field Using New Iterative Reduction
by: Yujun Xie, et al.
Published: (2025)
by: Yujun Xie, et al.
Published: (2025)
LRDUN: A Low-Rank Deep Unfolding Network for Efficient Spectral Compressive Imaging
by: Huang, He, et al.
Published: (2025)
by: Huang, He, et al.
Published: (2025)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
by: Zhang, Tianao, et al.
Published: (2025)
by: Zhang, Tianao, et al.
Published: (2025)
CRAFT-LoRA: Content-Style Personalization via Rank-Constrained Adaptation and Training-Free Fusion
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
DQ-LoRe: Dual Queries with Low Rank Approximation Re-ranking for In-Context Learning
by: Xiong, Jing, et al.
Published: (2023)
by: Xiong, Jing, et al.
Published: (2023)
6D-Diff: A Keypoint Diffusion Framework for 6D Object Pose Estimation
by: Xu, Li, et al.
Published: (2023)
by: Xu, Li, et al.
Published: (2023)
From Cloud-Native to Trust-Native: A Protocol for Verifiable Multi-Agent Systems
by: Li, Muyang
Published: (2025)
by: Li, Muyang
Published: (2025)
Exploring the Material Basis of Taxillus chinensis in the Treatment of Hyperuricemia Nephropathy Through Absorbed Into Blood Component Analysis and Network Pharmacology
by: Jiemei Liang, et al.
Published: (2025)
by: Jiemei Liang, et al.
Published: (2025)
StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
by: Feng, Tianrui, et al.
Published: (2025)
by: Feng, Tianrui, et al.
Published: (2025)
One-Bit Sensing of Low-Rank and Bisparse Matrices
by: Foucart, Simon, et al.
Published: (2019)
by: Foucart, Simon, et al.
Published: (2019)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
by: Li, Xingyang, et al.
Published: (2025)
by: Li, Xingyang, et al.
Published: (2025)
Analysis of Chemical Components and Blood‐Absorbed Components in Youjing Granules by UHPLC‐Q‐Orbitrap‐MS
by: Mingxin Guo, et al.
Published: (2025)
by: Mingxin Guo, et al.
Published: (2025)
HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models
by: Mao, Shizhuo, et al.
Published: (2025)
by: Mao, Shizhuo, et al.
Published: (2025)
State-Action Inpainting Diffuser for Continuous Control with Delay
by: Han, Dongqi, et al.
Published: (2026)
by: Han, Dongqi, et al.
Published: (2026)
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
by: Zhang, Stephen, et al.
Published: (2024)
by: Zhang, Stephen, et al.
Published: (2024)
Robust Randomized Low-Rank Approximation with Row-Wise Outlier Detection
by: Tiruvan, Aidan
Published: (2025)
by: Tiruvan, Aidan
Published: (2025)
LCQ: Low-Rank Codebook based Quantization for Large Language Models
by: Cai, Wen-Pu, et al.
Published: (2024)
by: Cai, Wen-Pu, et al.
Published: (2024)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
by: Lin, Yujun, et al.
Published: (2024)
by: Lin, Yujun, et al.
Published: (2024)
SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation
by: Chen, Junsong, et al.
Published: (2025)
by: Chen, Junsong, et al.
Published: (2025)
Low-Rank Plus Sparse Matrix Transfer Learning under Growing Representations and Ambient Dimensions
by: Chai, Jinhang, et al.
Published: (2026)
by: Chai, Jinhang, et al.
Published: (2026)
Similar Items
-
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
by: Xie, Enze, et al.
Published: (2024) -
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
by: He, Wenkun, et al.
Published: (2025) -
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
by: Xie, Enze, et al.
Published: (2025) -
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
by: Xu, Binxing, et al.
Published: (2026) -
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
by: Xi, Haocheng, et al.
Published: (2026)