Calibration and Transformation-Free Weight-Only LLMs Quantization via Dynamic Grouping
Fuente:
arXiv
Guardado en:
| Autores principales: | Zheng, Xinzhe, Yang, Zhen-Qun, Liu, Zishan, Xie, Haoran, Qin, S. Joe, Chen, Arlene, Lin, Fangzhen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
por: Song, Jaewoo, et al.
Publicado: (2025)
por: Song, Jaewoo, et al.
Publicado: (2025)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
por: Luo, Yingsong, et al.
Publicado: (2024)
por: Luo, Yingsong, et al.
Publicado: (2024)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
por: Song, Jaewoo, et al.
Publicado: (2025)
por: Song, Jaewoo, et al.
Publicado: (2025)
Emotion-Enhanced Multi-Task Learning with LLMs for Aspect Category Sentiment Analysis
por: Chai, Yaping, et al.
Publicado: (2025)
por: Chai, Yaping, et al.
Publicado: (2025)
Predicting Multi-Type Talented Students in Secondary School Using Semi-Supervised Machine Learning
por: Zheng, Xinzhe, et al.
Publicado: (2025)
por: Zheng, Xinzhe, et al.
Publicado: (2025)
Unsupervised Accelerated MRI Reconstruction via Ground-Truth-Free Flow Matching
por: Luo, Xinzhe, et al.
Publicado: (2025)
por: Luo, Xinzhe, et al.
Publicado: (2025)
Predicting Student Dropout Risk With A Dual-Modal Abrupt Behavioral Changes Approach
por: Cheng, Jiabei, et al.
Publicado: (2025)
por: Cheng, Jiabei, et al.
Publicado: (2025)
SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
por: Müller, Lorenz K., et al.
Publicado: (2025)
por: Müller, Lorenz K., et al.
Publicado: (2025)
Imagine a City: CityGenAgent for Procedural 3D City Generation
por: Liu, Zishan, et al.
Publicado: (2026)
por: Liu, Zishan, et al.
Publicado: (2026)
When LLMs Team Up: The Emergence of Collaborative Affective Computing
por: Lai, Wenna, et al.
Publicado: (2025)
por: Lai, Wenna, et al.
Publicado: (2025)
DAST: Context-Aware Compression in LLMs via Dynamic Allocation of Soft Tokens
por: Chen, Shaoshen, et al.
Publicado: (2025)
por: Chen, Shaoshen, et al.
Publicado: (2025)
Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
por: Xie, Zhuofan, et al.
Publicado: (2026)
por: Xie, Zhuofan, et al.
Publicado: (2026)
Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective
por: Ma, Xiaorui, et al.
Publicado: (2025)
por: Ma, Xiaorui, et al.
Publicado: (2025)
Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities
por: Chai, Yaping, et al.
Publicado: (2025)
por: Chai, Yaping, et al.
Publicado: (2025)
Semantic-preserved Augmentation with Confidence-weighted Fine-tuning for Aspect Category Sentiment Analysis
por: Chai, Yaping, et al.
Publicado: (2025)
por: Chai, Yaping, et al.
Publicado: (2025)
Task-Routed Mixture-of-Experts with Cognitive Appraisal for Implicit Sentiment Analysis
por: Chai, Yaping, et al.
Publicado: (2026)
por: Chai, Yaping, et al.
Publicado: (2026)
Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
por: Wang, Haozhe, et al.
Publicado: (2025)
por: Wang, Haozhe, et al.
Publicado: (2025)
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
por: Qin, Haotong, et al.
Publicado: (2024)
por: Qin, Haotong, et al.
Publicado: (2024)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
por: Cheng, Wenhua, et al.
Publicado: (2023)
por: Cheng, Wenhua, et al.
Publicado: (2023)
AdpQ: A Zero-shot Calibration Free Adaptive Post Training Quantization Method for LLMs
por: Ghaffari, Alireza, et al.
Publicado: (2024)
por: Ghaffari, Alireza, et al.
Publicado: (2024)
Constant-Memory Strategies in Stochastic Games: Best Responses and Equilibria
por: Zhu, Fengming, et al.
Publicado: (2025)
por: Zhu, Fengming, et al.
Publicado: (2025)
Computing Universal Plans for Partially Observable Multi-Agent Routing Using Answer Set Programming
por: Zhu, Fengming, et al.
Publicado: (2023)
por: Zhu, Fengming, et al.
Publicado: (2023)
Single-Agent Planning in a Multi-Agent System: A Unified Framework for Type-Based Planners
por: Zhu, Fengming, et al.
Publicado: (2025)
por: Zhu, Fengming, et al.
Publicado: (2025)
Quantized Delta Weight Is Safety Keeper
por: Liu, Yule, et al.
Publicado: (2024)
por: Liu, Yule, et al.
Publicado: (2024)
CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering
por: Li, Zongxi, et al.
Publicado: (2025)
por: Li, Zongxi, et al.
Publicado: (2025)
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
por: Choi, Euntae, et al.
Publicado: (2025)
por: Choi, Euntae, et al.
Publicado: (2025)
UPMRI: Unsupervised Parallel MRI Reconstruction via Projected Conditional Flow Matching
por: Luo, Xinzhe, et al.
Publicado: (2025)
por: Luo, Xinzhe, et al.
Publicado: (2025)
HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?
por: Ji, Sijie, et al.
Publicado: (2024)
por: Ji, Sijie, et al.
Publicado: (2024)
TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling
por: Gu, Hongyaoxing, et al.
Publicado: (2026)
por: Gu, Hongyaoxing, et al.
Publicado: (2026)
Rethinking ChatGPT's Success: Usability and Cognitive Behaviors Enabled by Auto-regressive LLMs' Prompting
por: Li, Xinzhe, et al.
Publicado: (2024)
por: Li, Xinzhe, et al.
Publicado: (2024)
Weight Group-wise Post-Training Quantization for Medical Foundation Model
por: Chen, Yineng, et al.
Publicado: (2026)
por: Chen, Yineng, et al.
Publicado: (2026)
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization
por: Deng, Yanxia, et al.
Publicado: (2025)
por: Deng, Yanxia, et al.
Publicado: (2025)
Training-Free Vector Quantization via Gaussian VAEs
por: Xu, Tongda, et al.
Publicado: (2025)
por: Xu, Tongda, et al.
Publicado: (2025)
Integer Scale: A Free Lunch for Faster Fine-grained Quantization of LLMs
por: Li, Qingyuan, et al.
Publicado: (2024)
por: Li, Qingyuan, et al.
Publicado: (2024)
Fine-grained Verification via Diagnostic Reasoning Supervision for Aspect Sentiment Triplet Extraction
por: Lai, Wenna, et al.
Publicado: (2026)
por: Lai, Wenna, et al.
Publicado: (2026)
Scaling Image Tokenizers with Grouped Spherical Quantization
por: Wang, Jiangtao, et al.
Publicado: (2024)
por: Wang, Jiangtao, et al.
Publicado: (2024)
GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration
por: Li, Yuhang, et al.
Publicado: (2025)
por: Li, Yuhang, et al.
Publicado: (2025)
Drive-Only Interaction Engineering via Dynamical Freezing
por: Xie, Songbo, et al.
Publicado: (2026)
por: Xie, Songbo, et al.
Publicado: (2026)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
por: Qiao, Ye, et al.
Publicado: (2026)
por: Qiao, Ye, et al.
Publicado: (2026)
CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
por: Sun, Ziteng, et al.
Publicado: (2025)
por: Sun, Ziteng, et al.
Publicado: (2025)
Ejemplares similares
-
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
por: Song, Jaewoo, et al.
Publicado: (2025) -
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
por: Luo, Yingsong, et al.
Publicado: (2024) -
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
por: Song, Jaewoo, et al.
Publicado: (2025) -
Emotion-Enhanced Multi-Task Learning with LLMs for Aspect Category Sentiment Analysis
por: Chai, Yaping, et al.
Publicado: (2025) -
Predicting Multi-Type Talented Students in Secondary School Using Semi-Supervised Machine Learning
por: Zheng, Xinzhe, et al.
Publicado: (2025)