Automatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and Compression
Fuente:
arXiv
Salvato in:
| Autori principali: | Qu, Xiaoyi, Aponte, David, Banbury, Colby, Robinson, Daniel P., Ding, Tianyu, Koishida, Kazuhito, Zharkov, Ilya, Chen, Tianyi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HESSO: Towards Automatic Efficient and User Friendly Any Neural Network Training and Pruning
di: Chen, Tianyi, et al.
Pubblicazione: (2024)
di: Chen, Tianyi, et al.
Pubblicazione: (2024)
WinClick: GUI Grounding with Multimodal Large Language Models
di: Hui, Zheng, et al.
Pubblicazione: (2025)
di: Hui, Zheng, et al.
Pubblicazione: (2025)
Zero-Shot Text-to-Speech from Continuous Text Streams
di: Dang, Trung, et al.
Pubblicazione: (2024)
di: Dang, Trung, et al.
Pubblicazione: (2024)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
di: Dang, Trung, et al.
Pubblicazione: (2024)
di: Dang, Trung, et al.
Pubblicazione: (2024)
Learned Image Compression with Text Quality Enhancement
di: Lai, Chih-Yu, et al.
Pubblicazione: (2024)
di: Lai, Chih-Yu, et al.
Pubblicazione: (2024)
CaesarNeRF: Calibrated Semantic Representation for Few-shot Generalizable Neural Rendering
di: Zhu, Haidong, et al.
Pubblicazione: (2023)
di: Zhu, Haidong, et al.
Pubblicazione: (2023)
FORA: Fast-Forward Caching in Diffusion Transformer Acceleration
di: Selvaraju, Pratheba, et al.
Pubblicazione: (2024)
di: Selvaraju, Pratheba, et al.
Pubblicazione: (2024)
Fast Data Aware Neural Architecture Search via Supernet Accelerated Evaluation
di: Njor, Emil, et al.
Pubblicazione: (2025)
di: Njor, Emil, et al.
Pubblicazione: (2025)
Integrating Pruning with Quantization for Efficient Deep Neural Networks Compression
di: Makenali, Sara, et al.
Pubblicazione: (2025)
di: Makenali, Sara, et al.
Pubblicazione: (2025)
AdaContour: Adaptive Contour Descriptor with Hierarchical Representation
di: Ding, Tianyu, et al.
Pubblicazione: (2024)
di: Ding, Tianyu, et al.
Pubblicazione: (2024)
Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model Compression
di: Kim, Minjun, et al.
Pubblicazione: (2026)
di: Kim, Minjun, et al.
Pubblicazione: (2026)
Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
di: Amizadeh, Saeed, et al.
Pubblicazione: (2025)
di: Amizadeh, Saeed, et al.
Pubblicazione: (2025)
Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression
di: Zhou, Longsheng, et al.
Pubblicazione: (2026)
di: Zhou, Longsheng, et al.
Pubblicazione: (2026)
AppSelectBench: Application-Level Tool Selection Benchmark
di: Chen, Tianyi, et al.
Pubblicazione: (2025)
di: Chen, Tianyi, et al.
Pubblicazione: (2025)
Single-channel speech enhancement using learnable loss mixup
di: Chang, Oscar, et al.
Pubblicazione: (2023)
di: Chang, Oscar, et al.
Pubblicazione: (2023)
DREAM: Diffusion Rectification and Estimation-Adaptive Models
di: Zhou, Jinxin, et al.
Pubblicazione: (2023)
di: Zhou, Jinxin, et al.
Pubblicazione: (2023)
Physics Inspired Criterion for Pruning-Quantization Joint Learning
di: Xie, Weiying, et al.
Pubblicazione: (2023)
di: Xie, Weiying, et al.
Pubblicazione: (2023)
Joint Pruning and Channel-wise Mixed-Precision Quantization for Efficient Deep Neural Networks
di: Motetti, Beatrice Alessandra, et al.
Pubblicazione: (2024)
di: Motetti, Beatrice Alessandra, et al.
Pubblicazione: (2024)
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
di: Tabassum, Afrina, et al.
Pubblicazione: (2024)
di: Tabassum, Afrina, et al.
Pubblicazione: (2024)
S3Editor: A Sparse Semantic-Disentangled Self-Training Framework for Face Video Editing
di: Wang, Guangzhi, et al.
Pubblicazione: (2024)
di: Wang, Guangzhi, et al.
Pubblicazione: (2024)
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
di: Ko, Jongwoo, et al.
Pubblicazione: (2025)
di: Ko, Jongwoo, et al.
Pubblicazione: (2025)
ProCrop: Learning Aesthetic Image Cropping from Professional Compositions
di: Zhang, Ke, et al.
Pubblicazione: (2025)
di: Zhang, Ke, et al.
Pubblicazione: (2025)
Instruction Agent: Enhancing Agent with Expert Demonstration
di: Li, Yinheng, et al.
Pubblicazione: (2025)
di: Li, Yinheng, et al.
Pubblicazione: (2025)
Weakly-supervised Audio Separation via Bi-modal Semantic Similarity
di: Mahmud, Tanvir, et al.
Pubblicazione: (2024)
di: Mahmud, Tanvir, et al.
Pubblicazione: (2024)
Structured Pruning and Quantization for Learned Image Compression
di: Hossain, Md Adnan Faisal, et al.
Pubblicazione: (2025)
di: Hossain, Md Adnan Faisal, et al.
Pubblicazione: (2025)
GETA-3DGS: Automatic Joint Structured Pruning and Quantization for 3D Gaussian Splatting
di: Zhang, Baobing, et al.
Pubblicazione: (2026)
di: Zhang, Baobing, et al.
Pubblicazione: (2026)
Cat-AIR: Content and Task-Aware All-in-One Image Restoration
di: Jiang, Jiachen, et al.
Pubblicazione: (2025)
di: Jiang, Jiachen, et al.
Pubblicazione: (2025)
C-SWAP: Explainability-Aware Structured Pruning for Efficient Neural Networks Compression
di: Bauvin, Baptiste, et al.
Pubblicazione: (2025)
di: Bauvin, Baptiste, et al.
Pubblicazione: (2025)
Joint Quantization and Pruning Neural Networks Approach: A Case Study on FSO Receivers
di: Obeed, Mohanad, et al.
Pubblicazione: (2025)
di: Obeed, Mohanad, et al.
Pubblicazione: (2025)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
di: Bai, Yatong, et al.
Pubblicazione: (2023)
di: Bai, Yatong, et al.
Pubblicazione: (2023)
Data Generation Using Large Language Models for Text Classification: An Empirical Case Study
di: Li, Yinheng, et al.
Pubblicazione: (2024)
di: Li, Yinheng, et al.
Pubblicazione: (2024)
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
di: Chen, Tianyi, et al.
Pubblicazione: (2026)
di: Chen, Tianyi, et al.
Pubblicazione: (2026)
Pruning and Quantization Impact on Graph Neural Networks
di: Khedri, Khatoon, et al.
Pubblicazione: (2025)
di: Khedri, Khatoon, et al.
Pubblicazione: (2025)
Shapley Pruning for Neural Network Compression
di: Adamczewski, Kamil, et al.
Pubblicazione: (2024)
di: Adamczewski, Kamil, et al.
Pubblicazione: (2024)
Efficient Training with Denoised Neural Weights
di: Gong, Yifan, et al.
Pubblicazione: (2024)
di: Gong, Yifan, et al.
Pubblicazione: (2024)
WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
di: Chen, Sihan, et al.
Pubblicazione: (2025)
di: Chen, Sihan, et al.
Pubblicazione: (2025)
The Efficiency Spectrum of Large Language Models: An Algorithmic Survey
di: Ding, Tianyu, et al.
Pubblicazione: (2023)
di: Ding, Tianyu, et al.
Pubblicazione: (2023)
OFER: Occluded Face Expression Reconstruction
di: Selvaraju, Pratheba, et al.
Pubblicazione: (2024)
di: Selvaraju, Pratheba, et al.
Pubblicazione: (2024)
QP-SNN: Quantized and Pruned Spiking Neural Networks
di: Wei, Wenjie, et al.
Pubblicazione: (2025)
di: Wei, Wenjie, et al.
Pubblicazione: (2025)
Unified Stochastic Framework for Neural Network Quantization and Pruning
di: Zhang, Haoyu, et al.
Pubblicazione: (2024)
di: Zhang, Haoyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
HESSO: Towards Automatic Efficient and User Friendly Any Neural Network Training and Pruning
di: Chen, Tianyi, et al.
Pubblicazione: (2024) -
WinClick: GUI Grounding with Multimodal Large Language Models
di: Hui, Zheng, et al.
Pubblicazione: (2025) -
Zero-Shot Text-to-Speech from Continuous Text Streams
di: Dang, Trung, et al.
Pubblicazione: (2024) -
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
di: Dang, Trung, et al.
Pubblicazione: (2024) -
Learned Image Compression with Text Quality Enhancement
di: Lai, Chih-Yu, et al.
Pubblicazione: (2024)