SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Ziyi, Jiang, Nan, Lin, Guang, Song, Qifan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Generalized Discrete Diffusion with Self-Correction
por: Wang, Linxuan, et al.
Publicado: (2026)
por: Wang, Linxuan, et al.
Publicado: (2026)
LLM Safety Alignment is Divergence Estimation in Disguise
por: Haldar, Rajdeep, et al.
Publicado: (2025)
por: Haldar, Rajdeep, et al.
Publicado: (2025)
On the Compressibility of Quantized Large Language Models
por: Mao, Yu, et al.
Publicado: (2024)
por: Mao, Yu, et al.
Publicado: (2024)
Hardware-Aware DNN Compression for Homogeneous Edge Devices
por: Zhang, Kunlong, et al.
Publicado: (2025)
por: Zhang, Kunlong, et al.
Publicado: (2025)
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
por: Zhang, Xi, et al.
Publicado: (2025)
por: Zhang, Xi, et al.
Publicado: (2025)
OCGEC: One-class Graph Embedding Classification for DNN Backdoor Detection
por: Jiang, Haoyu, et al.
Publicado: (2023)
por: Jiang, Haoyu, et al.
Publicado: (2023)
RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
por: Zhong, Zhengjia, et al.
Publicado: (2026)
por: Zhong, Zhengjia, et al.
Publicado: (2026)
DNN Memory Footprint Reduction via Post-Training Intra-Layer Multi-Precision Quantization
por: Ghavami, Behnam, et al.
Publicado: (2024)
por: Ghavami, Behnam, et al.
Publicado: (2024)
EGG-SR: Embedding Symbolic Equivalence into Symbolic Regression via Equality Graph
por: Jiang, Nan, et al.
Publicado: (2025)
por: Jiang, Nan, et al.
Publicado: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
por: Yu, Xiaoming, et al.
Publicado: (2026)
por: Yu, Xiaoming, et al.
Publicado: (2026)
SDQ: Sparse Decomposed Quantization for LLM Inference
por: Jeong, Geonhwa, et al.
Publicado: (2024)
por: Jeong, Geonhwa, et al.
Publicado: (2024)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
por: Liu, Wenyuan, et al.
Publicado: (2024)
por: Liu, Wenyuan, et al.
Publicado: (2024)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
por: Song, Jaewoo, et al.
Publicado: (2025)
por: Song, Jaewoo, et al.
Publicado: (2025)
UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
por: Chiang, Hung-Yueh, et al.
Publicado: (2025)
por: Chiang, Hung-Yueh, et al.
Publicado: (2025)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
por: Swain, Kabir, et al.
Publicado: (2026)
por: Swain, Kabir, et al.
Publicado: (2026)
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
por: Song, Jaewoo, et al.
Publicado: (2025)
por: Song, Jaewoo, et al.
Publicado: (2025)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
por: Xiao, Hanqi, et al.
Publicado: (2025)
por: Xiao, Hanqi, et al.
Publicado: (2025)
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
por: Xie, Jianhang, et al.
Publicado: (2025)
por: Xie, Jianhang, et al.
Publicado: (2025)
Hierarchical Sparse Plus Low Rank Compression of LLM
por: Kumar, Pawan, et al.
Publicado: (2025)
por: Kumar, Pawan, et al.
Publicado: (2025)
DeltaDQ: Ultra-High Delta Compression for Fine-Tuned LLMs via Group-wise Dropout and Separate Quantization
por: Jiang, Yanfeng, et al.
Publicado: (2024)
por: Jiang, Yanfeng, et al.
Publicado: (2024)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
por: Becking, Daniel, et al.
Publicado: (2021)
por: Becking, Daniel, et al.
Publicado: (2021)
E4GEN: Event-level Explainable Extreme-Enhanced Time-series Generation
por: Jiang, Lin, et al.
Publicado: (2026)
por: Jiang, Lin, et al.
Publicado: (2026)
HealthMamba: An Uncertainty-aware Spatiotemporal Graph State Space Model for Effective and Reliable Healthcare Facility Visit Prediction
por: Yu, Dahai, et al.
Publicado: (2026)
por: Yu, Dahai, et al.
Publicado: (2026)
Prototype-Guided Classification Sub-Task Decoupling Framework: Enhancing Generalization and Interpretability for Multivariate Time Series
por: Song, Xianhao, et al.
Publicado: (2026)
por: Song, Xianhao, et al.
Publicado: (2026)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
por: Peng, Hanyang, et al.
Publicado: (2025)
por: Peng, Hanyang, et al.
Publicado: (2025)
DNN-GDITD: Out-of-distribution detection via Deep Neural Network based Gaussian Descriptor for Imbalanced Tabular Data
por: Chudasama, Priyanka, et al.
Publicado: (2024)
por: Chudasama, Priyanka, et al.
Publicado: (2024)
Harnessing Neuron Stability to Improve DNN Verification
por: Duong, Hai, et al.
Publicado: (2024)
por: Duong, Hai, et al.
Publicado: (2024)
SCOUT: Toward Sub-Quadratic Attention via Segment Compression for Optimized Utility in Transformers
por: Jafari, Aref, et al.
Publicado: (2025)
por: Jafari, Aref, et al.
Publicado: (2025)
D-QRELO: Training- and Data-Free Delta Compression for Large Language Models via Quantization and Residual Low-Rank Approximation
por: Li, Junlin, et al.
Publicado: (2026)
por: Li, Junlin, et al.
Publicado: (2026)
Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
por: Bereska, Leonard, et al.
Publicado: (2025)
por: Bereska, Leonard, et al.
Publicado: (2025)
Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression
por: Cao, Tue M., et al.
Publicado: (2026)
por: Cao, Tue M., et al.
Publicado: (2026)
Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
por: Yang, Xu, et al.
Publicado: (2026)
por: Yang, Xu, et al.
Publicado: (2026)
Quantization through Piecewise-Affine Regularization: Optimization and Statistical Guarantees
por: Ma, Jianhao, et al.
Publicado: (2025)
por: Ma, Jianhao, et al.
Publicado: (2025)
TurboAngle: Near-Lossless KV Cache Compression via Uniform Angle Quantization
por: Patel, Dipkumar
Publicado: (2026)
por: Patel, Dipkumar
Publicado: (2026)
Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models
por: Chen, Kejia, et al.
Publicado: (2025)
por: Chen, Kejia, et al.
Publicado: (2025)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
por: Qiao, Ye, et al.
Publicado: (2026)
por: Qiao, Ye, et al.
Publicado: (2026)
Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series Forecasting
por: Zhao, Yanjun, et al.
Publicado: (2024)
por: Zhao, Yanjun, et al.
Publicado: (2024)
Quantization Dominates Rank Reduction for KV-Cache Compression
por: Salfati, Samuel
Publicado: (2026)
por: Salfati, Samuel
Publicado: (2026)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
por: Zhang, Peiyuan, et al.
Publicado: (2026)
por: Zhang, Peiyuan, et al.
Publicado: (2026)
Selecting Belief-State Approximations in Simulators with Latent States
por: Jiang, Nan
Publicado: (2025)
por: Jiang, Nan
Publicado: (2025)
Ejemplares similares
-
Generalized Discrete Diffusion with Self-Correction
por: Wang, Linxuan, et al.
Publicado: (2026) -
LLM Safety Alignment is Divergence Estimation in Disguise
por: Haldar, Rajdeep, et al.
Publicado: (2025) -
On the Compressibility of Quantized Large Language Models
por: Mao, Yu, et al.
Publicado: (2024) -
Hardware-Aware DNN Compression for Homogeneous Edge Devices
por: Zhang, Kunlong, et al.
Publicado: (2025) -
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
por: Zhang, Xi, et al.
Publicado: (2025)