Guardado en:
| Autores principales: | Lee, Dongyoung, Choi, Seungkyu, Chang, Ik Joon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2501.13331 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
por: Park, Yeonsik, et al.
Publicado: (2026)
por: Park, Yeonsik, et al.
Publicado: (2026)
Post-Training Quantization via Residual Truncation and Zero Suppression for Diffusion Models
por: Kim, Donghoon, et al.
Publicado: (2025)
por: Kim, Donghoon, et al.
Publicado: (2025)
Beta-Sigma VAE: Separating beta and decoder variance in Gaussian variational autoencoder
por: Kim, Seunghwan, et al.
Publicado: (2024)
por: Kim, Seunghwan, et al.
Publicado: (2024)
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
por: Zhao, Yilong, et al.
Publicado: (2023)
por: Zhao, Yilong, et al.
Publicado: (2023)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
por: Jia, Jinda, et al.
Publicado: (2024)
por: Jia, Jinda, et al.
Publicado: (2024)
LO-BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference
por: Elangovan, Reena, et al.
Publicado: (2025)
por: Elangovan, Reena, et al.
Publicado: (2025)
AAAC: Activation-Aware Adaptive Codebooks for 4-bit LLM Weight Quantization
por: IslamBouli, Beshr, et al.
Publicado: (2026)
por: IslamBouli, Beshr, et al.
Publicado: (2026)
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
por: Lee, Geonho, et al.
Publicado: (2024)
por: Lee, Geonho, et al.
Publicado: (2024)
4bit-Quantization in Vector-Embedding for RAG
por: Jeong, Taehee
Publicado: (2025)
por: Jeong, Taehee
Publicado: (2025)
OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization
por: Li, Zhikai, et al.
Publicado: (2026)
por: Li, Zhikai, et al.
Publicado: (2026)
ICQuant: Index Coding enables Low-bit LLM Quantization
por: Li, Xinlin, et al.
Publicado: (2025)
por: Li, Xinlin, et al.
Publicado: (2025)
Training-free LLM Verification via Recycling Few-shot Examples
por: Lee, Dongseok, et al.
Publicado: (2025)
por: Lee, Dongseok, et al.
Publicado: (2025)
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs
por: Zhou, Zhaojing, et al.
Publicado: (2025)
por: Zhou, Zhaojing, et al.
Publicado: (2025)
Occam's Razor is Only as Sharp as Your ELBO
por: Harvey, Ethan, et al.
Publicado: (2026)
por: Harvey, Ethan, et al.
Publicado: (2026)
SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization
por: Bai, Runsheng, et al.
Publicado: (2024)
por: Bai, Runsheng, et al.
Publicado: (2024)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
por: Kim, Dongyoung, et al.
Publicado: (2024)
por: Kim, Dongyoung, et al.
Publicado: (2024)
Toward Architecture-Agnostic Local Control of Posterior Collapse in VAEs
por: Song, Hyunsoo, et al.
Publicado: (2025)
por: Song, Hyunsoo, et al.
Publicado: (2025)
EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
por: Zhang, Shu-Hao, et al.
Publicado: (2026)
por: Zhang, Shu-Hao, et al.
Publicado: (2026)
A Geometric Modeling of Occam's Razor in Deep Learning
por: Sun, Ke, et al.
Publicado: (2019)
por: Sun, Ke, et al.
Publicado: (2019)
Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
por: Blumenberg, Patrick, et al.
Publicado: (2025)
por: Blumenberg, Patrick, et al.
Publicado: (2025)
LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization
por: Bouquet, Yann, et al.
Publicado: (2026)
por: Bouquet, Yann, et al.
Publicado: (2026)
Effortless, Simulation-Efficient Bayesian Inference using Tabular Foundation Models
por: Vetter, Julius, et al.
Publicado: (2025)
por: Vetter, Julius, et al.
Publicado: (2025)
yProv4ML: Effortless Provenance Tracking for Machine Learning Systems
por: Padovani, Gabriele, et al.
Publicado: (2025)
por: Padovani, Gabriele, et al.
Publicado: (2025)
OTTER: Effortless Label Distribution Adaptation of Zero-shot Models
por: Shin, Changho, et al.
Publicado: (2024)
por: Shin, Changho, et al.
Publicado: (2024)
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
por: Zhang, Xuan, et al.
Publicado: (2024)
por: Zhang, Xuan, et al.
Publicado: (2024)
RL's Razor: Why Online Reinforcement Learning Forgets Less
por: Shenfeld, Idan, et al.
Publicado: (2025)
por: Shenfeld, Idan, et al.
Publicado: (2025)
1-Bit FQT: Pushing the Limit of Fully Quantized Training to 1-bit
por: Gao, Chang, et al.
Publicado: (2024)
por: Gao, Chang, et al.
Publicado: (2024)
FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
por: Baek, Daehyeon, et al.
Publicado: (2025)
por: Baek, Daehyeon, et al.
Publicado: (2025)
On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
por: Fan, Chenghao, et al.
Publicado: (2024)
por: Fan, Chenghao, et al.
Publicado: (2024)
SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
por: Xia, Junhao, et al.
Publicado: (2025)
por: Xia, Junhao, et al.
Publicado: (2025)
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
por: Heo, Jung Hwan, et al.
Publicado: (2023)
por: Heo, Jung Hwan, et al.
Publicado: (2023)
DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers
por: Sharify, Sayeh, et al.
Publicado: (2026)
por: Sharify, Sayeh, et al.
Publicado: (2026)
MergeQuant: Accurate 4-bit Static Quantization of Large Language Models by Channel-wise Calibration
por: Wang, Jinguang, et al.
Publicado: (2025)
por: Wang, Jinguang, et al.
Publicado: (2025)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
por: Xu, Bingxin, et al.
Publicado: (2025)
por: Xu, Bingxin, et al.
Publicado: (2025)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
por: Liu, Zechun, et al.
Publicado: (2025)
por: Liu, Zechun, et al.
Publicado: (2025)
Effortless Active Labeling for Long-Term Test-Time Adaptation
por: Wang, Guowei, et al.
Publicado: (2025)
por: Wang, Guowei, et al.
Publicado: (2025)
FLARE: FP-Less PTQ and Low-ENOB ADC Based AMS-PiM for Error-Resilient, Fast, and Efficient Transformer Acceleration
por: Yi, Donghyeon, et al.
Publicado: (2024)
por: Yi, Donghyeon, et al.
Publicado: (2024)
Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural Networks
por: Shahverdi, Vahid, et al.
Publicado: (2025)
por: Shahverdi, Vahid, et al.
Publicado: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
por: Tang, Hanlin, et al.
Publicado: (2024)
por: Tang, Hanlin, et al.
Publicado: (2024)
Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks Using the Marginal Likelihood
por: Dhahri, Rayen, et al.
Publicado: (2024)
por: Dhahri, Rayen, et al.
Publicado: (2024)
Ejemplares similares
-
SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
por: Park, Yeonsik, et al.
Publicado: (2026) -
Post-Training Quantization via Residual Truncation and Zero Suppression for Diffusion Models
por: Kim, Donghoon, et al.
Publicado: (2025) -
Beta-Sigma VAE: Separating beta and decoder variance in Gaussian variational autoencoder
por: Kim, Seunghwan, et al.
Publicado: (2024) -
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
por: Zhao, Yilong, et al.
Publicado: (2023) -
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
por: Jia, Jinda, et al.
Publicado: (2024)