AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Jing, Koike-Akino, Toshiaki, Wang, Ye, Mansour, Hassan, Brand, Matthew |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2026)
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2026)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
LatentLLM: Attention-Aware Joint Tensor Compression
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
TuneComp: Joint Fine-tuning and Compression for Large Foundation Models
von: Chen, Xiangyu, et al.
Veröffentlicht: (2025)
von: Chen, Xiangyu, et al.
Veröffentlicht: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
Directional Embedding Smoothing for Robust Vision Language Models
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
Quantum Diffusion Models for Few-Shot Learning
von: Wang, Ruhan, et al.
Veröffentlicht: (2024)
von: Wang, Ruhan, et al.
Veröffentlicht: (2024)
Embedding Morphology into Transformers for Cross-Robot Policy Learning
von: Suzuki, Kei, et al.
Veröffentlicht: (2026)
von: Suzuki, Kei, et al.
Veröffentlicht: (2026)
Random Channel Ablation for Robust Hand Gesture Classification with Multimodal Biosignals
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024)
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024)
SuperLoRA: Parameter-Efficient Unified Adaptation of Multi-Layer Attention Modules
von: Chen, Xiangyu, et al.
Veröffentlicht: (2024)
von: Chen, Xiangyu, et al.
Veröffentlicht: (2024)
GPT Sonograpy: Hand Gesture Decoding from Forearm Ultrasound Images via VLM
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024)
von: Bimbraw, Keshav, et al.
Veröffentlicht: (2024)
Efficient Differentially Private Fine-Tuning of Diffusion Models
von: Liu, Jing, et al.
Veröffentlicht: (2024)
von: Liu, Jing, et al.
Veröffentlicht: (2024)
Analyzing Inference Privacy Risks Through Gradients in Machine Learning
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
Variational Randomized Smoothing for Sample-Wise Adversarial Robustness
von: Hase, Ryo, et al.
Veröffentlicht: (2024)
von: Hase, Ryo, et al.
Veröffentlicht: (2024)
Quantum Implicit Neural Compression
von: Fujihashi, Takuya, et al.
Veröffentlicht: (2024)
von: Fujihashi, Takuya, et al.
Veröffentlicht: (2024)
Exploring User-level Gradient Inversion with a Diffusion Prior
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
AutoHLS: Learning to Accelerate Design Space Exploration for HLS Designs
von: Ahmed, Md Rubel, et al.
Veröffentlicht: (2024)
von: Ahmed, Md Rubel, et al.
Veröffentlicht: (2024)
Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents
von: Lewis, Ashley, et al.
Veröffentlicht: (2025)
von: Lewis, Ashley, et al.
Veröffentlicht: (2025)
Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2024)
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2024)
Why Does Differential Privacy with Large Epsilon Defend Against Practical Membership Inference Attacks?
von: Lowy, Andrew, et al.
Veröffentlicht: (2024)
von: Lowy, Andrew, et al.
Veröffentlicht: (2024)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
von: Cheng, Wenhua, et al.
Veröffentlicht: (2023)
von: Cheng, Wenhua, et al.
Veröffentlicht: (2023)
Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2026)
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2026)
Probabilistic Forecasting for Building Energy Systems using Time-Series Foundation Models
von: Park, Young Jin, et al.
Veröffentlicht: (2025)
von: Park, Young Jin, et al.
Veröffentlicht: (2025)
Quantum-PEFT: Ultra parameter-efficient fine-tuning
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
A Granger-Causal Perspective on Gradient Descent with Application to Pruning
von: Shah, Aditya, et al.
Veröffentlicht: (2024)
von: Shah, Aditya, et al.
Veröffentlicht: (2024)
Smoothed Embeddings for Robust Language Models
von: Hase, Ryo, et al.
Veröffentlicht: (2025)
von: Hase, Ryo, et al.
Veröffentlicht: (2025)
GWQ: Gradient-Aware Weight Quantization for Large Language Models
von: Shao, Yihua, et al.
Veröffentlicht: (2024)
von: Shao, Yihua, et al.
Veröffentlicht: (2024)
Revisiting Gradient Descent: A Dual-Weight Method for Improved Learning
von: Wang, Xi
Veröffentlicht: (2025)
von: Wang, Xi
Veröffentlicht: (2025)
Geo-ADAPT-VQE: Quantum Information Metric-Aware Circuit Optimization for Quantum Chemistry
von: Sohail, Mohammad Aamir, et al.
Veröffentlicht: (2026)
von: Sohail, Mohammad Aamir, et al.
Veröffentlicht: (2026)
Adjacent Leader Decentralized Stochastic Gradient Descent
von: He, Haoze, et al.
Veröffentlicht: (2024)
von: He, Haoze, et al.
Veröffentlicht: (2024)
Attacking Large Language Models with Projected Gradient Descent
von: Geisler, Simon, et al.
Veröffentlicht: (2024)
von: Geisler, Simon, et al.
Veröffentlicht: (2024)
Weighted Averaged Stochastic Gradient Descent: Asymptotic Normality and Optimality
von: Wei, Ziyang, et al.
Veröffentlicht: (2023)
von: Wei, Ziyang, et al.
Veröffentlicht: (2023)
AAAC: Activation-Aware Adaptive Codebooks for 4-bit LLM Weight Quantization
von: IslamBouli, Beshr, et al.
Veröffentlicht: (2026)
von: IslamBouli, Beshr, et al.
Veröffentlicht: (2026)
Efficient Search for Customized Activation Functions with Gradient Descent
von: Strack, Lukas, et al.
Veröffentlicht: (2024)
von: Strack, Lukas, et al.
Veröffentlicht: (2024)
Preconditioning for Accelerated Gradient Descent Optimization and Regularization
von: Ye, Qiang
Veröffentlicht: (2024)
von: Ye, Qiang
Veröffentlicht: (2024)
On the Inherent Privacy of Zeroth Order Projected Gradient Descent
von: Gupta, Devansh, et al.
Veröffentlicht: (2025)
von: Gupta, Devansh, et al.
Veröffentlicht: (2025)
Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent
von: Wei, Yongxian, et al.
Veröffentlicht: (2025)
von: Wei, Yongxian, et al.
Veröffentlicht: (2025)
Stochastic Gradient Descent with Adaptive Data
von: Che, Ethan, et al.
Veröffentlicht: (2024)
von: Che, Ethan, et al.
Veröffentlicht: (2024)
Fisher-Orthogonal Projected Natural Gradient Descent for Continual Learning
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
Generalized Gradient Descent is a Hypergraph Functor
von: Hanks, Tyler, et al.
Veröffentlicht: (2024)
von: Hanks, Tyler, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2026) -
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025) -
LatentLLM: Attention-Aware Joint Tensor Compression
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025) -
TuneComp: Joint Fine-tuning and Compression for Large Foundation Models
von: Chen, Xiangyu, et al.
Veröffentlicht: (2025) -
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
von: Wang, Ye, et al.
Veröffentlicht: (2026)