Gespeichert in:
| Hauptverfasser: | Zhou, Yuli, Chen, Qingxuan, Benini, Luca, Sun, Guolei, Li, Yawei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.02151 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
von: Guo, Hang, et al.
Veröffentlicht: (2025)
von: Guo, Hang, et al.
Veröffentlicht: (2025)
CamSAM2: Segment Anything Accurately in Camouflaged Videos
von: Zhou, Yuli, et al.
Veröffentlicht: (2025)
von: Zhou, Yuli, et al.
Veröffentlicht: (2025)
Direct Quantized Training of Language Models with Stochastic Rounding
von: Zhao, Kaiyan, et al.
Veröffentlicht: (2024)
von: Zhao, Kaiyan, et al.
Veröffentlicht: (2024)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
When SAM2 Meets Video Camouflaged Object Segmentation: A Comprehensive Evaluation and Adaptation
von: Zhou, Yuli, et al.
Veröffentlicht: (2024)
von: Zhou, Yuli, et al.
Veröffentlicht: (2024)
Robust Training of Vector Quantized Bottleneck Models
von: Łańcucki, Adrian, et al.
Veröffentlicht: (2020)
von: Łańcucki, Adrian, et al.
Veröffentlicht: (2020)
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
von: You, Haoran, et al.
Veröffentlicht: (2024)
von: You, Haoran, et al.
Veröffentlicht: (2024)
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
von: Zhang, Junkai, et al.
Veröffentlicht: (2026)
von: Zhang, Junkai, et al.
Veröffentlicht: (2026)
A Reparameterized Discrete Diffusion Model for Text Generation
von: Zheng, Lin, et al.
Veröffentlicht: (2023)
von: Zheng, Lin, et al.
Veröffentlicht: (2023)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
von: Cheng, Wenhua, et al.
Veröffentlicht: (2023)
von: Cheng, Wenhua, et al.
Veröffentlicht: (2023)
End-to-End Training for Back-Translation with Categorical Reparameterization Trick
von: Heo, DongNyeong, et al.
Veröffentlicht: (2022)
von: Heo, DongNyeong, et al.
Veröffentlicht: (2022)
AAAC: Activation-Aware Adaptive Codebooks for 4-bit LLM Weight Quantization
von: IslamBouli, Beshr, et al.
Veröffentlicht: (2026)
von: IslamBouli, Beshr, et al.
Veröffentlicht: (2026)
FlatQuant: Flatness Matters for LLM Quantization
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2024)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2024)
SqueezeLLM: Dense-and-Sparse Quantization
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
Robust and Efficient Fine-tuning of LLMs with Bayesian Reparameterization of Low-Rank Adaptation
von: Sengupta, Ayan, et al.
Veröffentlicht: (2024)
von: Sengupta, Ayan, et al.
Veröffentlicht: (2024)
Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
von: Zhang, Xiaoyun, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyun, et al.
Veröffentlicht: (2025)
RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
HiM2SAM: Enhancing SAM2 with Hierarchical Motion Estimation and Memory Optimization towards Long-term Tracking
von: Chen, Ruixiang, et al.
Veröffentlicht: (2025)
von: Chen, Ruixiang, et al.
Veröffentlicht: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
von: Wang, YiFeng, et al.
Veröffentlicht: (2026)
von: Wang, YiFeng, et al.
Veröffentlicht: (2026)
GPTVQ: The Blessing of Dimensionality for LLM Quantization
von: van Baalen, Mart, et al.
Veröffentlicht: (2024)
von: van Baalen, Mart, et al.
Veröffentlicht: (2024)
Round and Round We Go! What makes Rotary Positional Encodings useful?
von: Barbero, Federico, et al.
Veröffentlicht: (2024)
von: Barbero, Federico, et al.
Veröffentlicht: (2024)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
Vector Quantized Latent Concepts: A Scalable Alternative to Clustering-Based Concept Discovery
von: Yu, Xuemin, et al.
Veröffentlicht: (2026)
von: Yu, Xuemin, et al.
Veröffentlicht: (2026)
Adaptive Task Vectors for Large Language Models
von: Kang, Joonseong, et al.
Veröffentlicht: (2025)
von: Kang, Joonseong, et al.
Veröffentlicht: (2025)
Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A Benchmark
von: Zhang, Yihua, et al.
Veröffentlicht: (2024)
von: Zhang, Yihua, et al.
Veröffentlicht: (2024)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
von: Liu, Di, et al.
Veröffentlicht: (2024)
von: Liu, Di, et al.
Veröffentlicht: (2024)
QUAD: Quantization and Parameter-Efficient Tuning of LLM with Activation Decomposition
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
Transformer-VQ: Linear-Time Transformers via Vector Quantization
von: Lingle, Lucas D.
Veröffentlicht: (2023)
von: Lingle, Lucas D.
Veröffentlicht: (2023)
Evaluation Hallucination in Multi-Round Incomplete Information Lateral-Driven Reasoning Tasks
von: Dong, Wenhan, et al.
Veröffentlicht: (2025)
von: Dong, Wenhan, et al.
Veröffentlicht: (2025)
Scaling Law for Quantization-Aware Training
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
von: Wang, Shida, et al.
Veröffentlicht: (2023)
von: Wang, Shida, et al.
Veröffentlicht: (2023)
Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
von: Gao, Bin, et al.
Veröffentlicht: (2024)
von: Gao, Bin, et al.
Veröffentlicht: (2024)
A Framework for Cost-Effective and Self-Adaptive LLM Shaking and Recovery Mechanism
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
von: Li, Pengyi, et al.
Veröffentlicht: (2026)
von: Li, Pengyi, et al.
Veröffentlicht: (2026)
BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search
von: Sun, Linzhuang, et al.
Veröffentlicht: (2024)
von: Sun, Linzhuang, et al.
Veröffentlicht: (2024)
From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization
von: Zhou, Chenxi, et al.
Veröffentlicht: (2026)
von: Zhou, Chenxi, et al.
Veröffentlicht: (2026)
Trust in One Round: Confidence Estimation for Large Language Models via Structural Signals
von: Yang, Pengyue, et al.
Veröffentlicht: (2026)
von: Yang, Pengyue, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
von: Huang, Wei, et al.
Veröffentlicht: (2024) -
Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
von: Guo, Hang, et al.
Veröffentlicht: (2025) -
CamSAM2: Segment Anything Accurately in Camouflaged Videos
von: Zhou, Yuli, et al.
Veröffentlicht: (2025) -
Direct Quantized Training of Language Models with Stochastic Rounding
von: Zhao, Kaiyan, et al.
Veröffentlicht: (2024) -
Reparameterized LLM Training via Orthogonal Equivalence Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)