PoTPTQ: A Two-step Power-of-Two Post-training for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xinyu, Nia, Vahid Partovi, Lu, Peng, Huang, Jerry, Chang, Xiao-Wen, Chen, Boxing, Cui, Yufei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdpQ: A Zero-shot Calibration Free Adaptive Post Training Quantization Method for LLMs
by: Ghaffari, Alireza, et al.
Published: (2024)
by: Ghaffari, Alireza, et al.
Published: (2024)
OAC: Output-adaptive Calibration for Accurate Post-training Quantization
by: Edalati, Ali, et al.
Published: (2024)
by: Edalati, Ali, et al.
Published: (2024)
Resona: Improving Context Copying in Linear Recurrence Models with Retrieval
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models
by: Ghaffari, Alireza, et al.
Published: (2023)
by: Ghaffari, Alireza, et al.
Published: (2023)
Rethinking Post-Training Quantization: Introducing a Statistical Pre-Calibration Approach
by: Ghaffari, Alireza, et al.
Published: (2025)
by: Ghaffari, Alireza, et al.
Published: (2025)
Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning
by: Gu, Yu, et al.
Published: (2026)
by: Gu, Yu, et al.
Published: (2026)
BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMs
by: Ghaddar, Abbas, et al.
Published: (2026)
by: Ghaddar, Abbas, et al.
Published: (2026)
EvoEdit: Evolving Null-space Alignment for Robust and Efficient Knowledge Editing
by: Lyu, Sicheng, et al.
Published: (2025)
by: Lyu, Sicheng, et al.
Published: (2025)
Power-of-Two Quantization-Aware-Training (PoT-QAT) in Large Language Models (LLMs)
by: Elgenedy, Mahmoud
Published: (2026)
by: Elgenedy, Mahmoud
Published: (2026)
Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models
by: Metel, Michael R., et al.
Published: (2026)
by: Metel, Michael R., et al.
Published: (2026)
Mamba Modulation: On the Length Generalization of Mamba
by: Lu, Peng, et al.
Published: (2025)
by: Lu, Peng, et al.
Published: (2025)
Understanding Neural Network Binarization with Forward and Backward Proximal Quantizers
by: Lu, Yiwei, et al.
Published: (2024)
by: Lu, Yiwei, et al.
Published: (2024)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
by: Heuillet, Maxime, et al.
Published: (2025)
by: Heuillet, Maxime, et al.
Published: (2025)
Enterprise Resource Planning Using Multi-type Transformers in Ferro-Titanium Industry
by: Yazdanpourmoghadam, Samira, et al.
Published: (2026)
by: Yazdanpourmoghadam, Samira, et al.
Published: (2026)
InfMem: Learning System-2 Memory Control for Long-Context Agent
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
MoKA: Mixture of Kronecker Adapters
by: Sadeghi, Mohammadreza, et al.
Published: (2025)
by: Sadeghi, Mohammadreza, et al.
Published: (2025)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Transferable Post-training via Inverse Value Learning
by: Lu, Xinyu, et al.
Published: (2024)
by: Lu, Xinyu, et al.
Published: (2024)
On Predicting the Post-training Potential of Pre-trained LLMs
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
Two Heads are Better than One: Nested PoE for Robust Defense Against Multi-Backdoors
by: Graf, Victoria, et al.
Published: (2024)
by: Graf, Victoria, et al.
Published: (2024)
Self-controller: Controlling LLMs with Multi-round Step-by-step Self-awareness
by: Peng, Xiao, et al.
Published: (2024)
by: Peng, Xiao, et al.
Published: (2024)
Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization
by: Xing, Tiancheng, et al.
Published: (2025)
by: Xing, Tiancheng, et al.
Published: (2025)
ReGLA: Refining Gated Linear Attention
by: Lu, Peng, et al.
Published: (2025)
by: Lu, Peng, et al.
Published: (2025)
Power-of-Two (PoT) Weights in Large Language Models (LLMs)
by: Elgenedy, Mahmoud
Published: (2025)
by: Elgenedy, Mahmoud
Published: (2025)
A Unified Understanding of Offline Data Selection and Online Self-refining Generation for Post-training LLMs
by: Xiao, Quan, et al.
Published: (2025)
by: Xiao, Quan, et al.
Published: (2025)
Beyond Hard Writes and Rigid Preservation: Soft Recursive Least-Squares for Lifelong LLM Editing
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
by: Metel, Michael R., et al.
Published: (2024)
by: Metel, Michael R., et al.
Published: (2024)
Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization
by: Tseng, Yu-Min, et al.
Published: (2024)
by: Tseng, Yu-Min, et al.
Published: (2024)
Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
by: Lin, Haokun, et al.
Published: (2025)
by: Lin, Haokun, et al.
Published: (2025)
Beyond Mode-Seeking RL: Trajectory-Balance Post-Training for Diffusion Language Models
by: Ahmadi, Saba, et al.
Published: (2026)
by: Ahmadi, Saba, et al.
Published: (2026)
Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Asymmetric Conflict and Synergy in Post-training for LLM-based Multilingual Machine Translation
by: Zheng, Tong, et al.
Published: (2025)
by: Zheng, Tong, et al.
Published: (2025)
Two Failures of Self-Consistency in the Multi-Step Reasoning of LLMs
by: Chen, Angelica, et al.
Published: (2023)
by: Chen, Angelica, et al.
Published: (2023)
Do Localization Methods Actually Localize Memorized Data in LLMs? A Tale of Two Benchmarks
by: Chang, Ting-Yun, et al.
Published: (2023)
by: Chang, Ting-Yun, et al.
Published: (2023)
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
by: Metel, Michael R., et al.
Published: (2024)
by: Metel, Michael R., et al.
Published: (2024)
Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning
by: Saparkhan, Raman, et al.
Published: (2026)
by: Saparkhan, Raman, et al.
Published: (2026)
Two-Stage Regularization-Based Structured Pruning for LLMs
by: Feng, Mingkuan, et al.
Published: (2025)
by: Feng, Mingkuan, et al.
Published: (2025)
SkillVerse : Assessing and Enhancing LLMs with Tree Evaluation
by: Tian, Yufei, et al.
Published: (2025)
by: Tian, Yufei, et al.
Published: (2025)
Can Post-Training Transform LLMs into Causal Reasoners?
by: Chen, Junqi, et al.
Published: (2026)
by: Chen, Junqi, et al.
Published: (2026)
Similar Items
-
AdpQ: A Zero-shot Calibration Free Adaptive Post Training Quantization Method for LLMs
by: Ghaffari, Alireza, et al.
Published: (2024) -
OAC: Output-adaptive Calibration for Accurate Post-training Quantization
by: Edalati, Ali, et al.
Published: (2024) -
Resona: Improving Context Copying in Linear Recurrence Models with Retrieval
by: Wang, Xinyu, et al.
Published: (2025) -
Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models
by: Ghaffari, Alireza, et al.
Published: (2023) -
Rethinking Post-Training Quantization: Introducing a Statistical Pre-Calibration Approach
by: Ghaffari, Alireza, et al.
Published: (2025)