ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
Fuente:
arXiv
Saved in:
| Main Authors: | You, Haoran, Guo, Yipin, Fu, Yichao, Zhou, Wei, Shi, Huihong, Zhang, Xiaofan, Kundu, Souvik, Yazdanbakhsh, Amir, Lin, Yingyan Celine |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision Transformer
by: You, Haoran, et al.
Published: (2023)
by: You, Haoran, et al.
Published: (2023)
ShiftAddNAS: Hardware-Inspired Search for More Accurate and Efficient Neural Networks
by: You, Haoran, et al.
Published: (2022)
by: You, Haoran, et al.
Published: (2022)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
by: You, Haoran, et al.
Published: (2024)
by: You, Haoran, et al.
Published: (2024)
ShiftAddNet: A Hardware-Inspired Deep Network
by: You, Haoran, et al.
Published: (2020)
by: You, Haoran, et al.
Published: (2020)
ShiftAddAug: Augment Multiplication-Free Tiny Neural Network with Hybrid Computation
by: Guo, Yipin, et al.
Published: (2024)
by: Guo, Yipin, et al.
Published: (2024)
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design
by: You, Haoran, et al.
Published: (2022)
by: You, Haoran, et al.
Published: (2022)
GCoD: Graph Convolutional Network Acceleration via Dedicated Algorithm and Accelerator Co-Design
by: You, Haoran, et al.
Published: (2021)
by: You, Haoran, et al.
Published: (2021)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
by: Yu, Zhongzhi, et al.
Published: (2024)
by: Yu, Zhongzhi, et al.
Published: (2024)
Early-Bird GCNs: Graph-Network Co-Optimization Towards More Efficient GCN Training and Inference via Drawing Early-Bird Lottery Tickets
by: You, Haoran, et al.
Published: (2021)
by: You, Haoran, et al.
Published: (2021)
Instant-3D: Instant Neural Radiance Field Training Towards On-Device AR/VR 3D Reconstruction
by: Li, Sixu, et al.
Published: (2023)
by: Li, Sixu, et al.
Published: (2023)
NeRFool: Uncovering the Vulnerability of Generalizable Neural Radiance Fields against Adversarial Perturbations
by: Fu, Yonggan, et al.
Published: (2023)
by: Fu, Yonggan, et al.
Published: (2023)
SLoPe: Double-Pruned Sparse Plus Lazy Low-Rank Adapter Pretraining of LLMs
by: Mozaffari, Mohammad, et al.
Published: (2024)
by: Mozaffari, Mohammad, et al.
Published: (2024)
DNA: Differentiable Network-Accelerator Co-Search
by: Zhang, Yongan, et al.
Published: (2020)
by: Zhang, Yongan, et al.
Published: (2020)
Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
by: Yazdanbakhsh, Amir
Published: (2025)
by: Yazdanbakhsh, Amir
Published: (2025)
FracTrain: Fractionally Squeezing Bit Savings Both Temporally and Spatially for Efficient DNN Training
by: Fu, Yonggan, et al.
Published: (2020)
by: Fu, Yonggan, et al.
Published: (2020)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
Max-Affine Spline Insights Into Deep Network Pruning
by: You, Haoran, et al.
Published: (2021)
by: You, Haoran, et al.
Published: (2021)
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
P$^2$-ViT: Power-of-Two Post-Training Quantization and Acceleration for Fully Quantized Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
NASH: Neural Architecture and Accelerator Search for Multiplication-Reduced Hybrid Models
by: Xu, Yang, et al.
Published: (2024)
by: Xu, Yang, et al.
Published: (2024)
SuperTickets: Drawing Task-Agnostic Lottery Tickets from Supernets via Jointly Architecture Searching and Parameter Pruning
by: You, Haoran, et al.
Published: (2022)
by: You, Haoran, et al.
Published: (2022)
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
MixGCN: Scalable GCN Training by Mixture of Parallelism and Mixture of Accelerators
by: Wan, Cheng, et al.
Published: (2025)
by: Wan, Cheng, et al.
Published: (2025)
Agentic AI Workload Characteristics
by: Yuan, Yichao, et al.
Published: (2026)
by: Yuan, Yichao, et al.
Published: (2026)
Drawing Early-Bird Tickets: Towards More Efficient Training of Deep Networks
by: You, Haoran, et al.
Published: (2019)
by: You, Haoran, et al.
Published: (2019)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
by: Qiu, Zeju, et al.
Published: (2025)
by: Qiu, Zeju, et al.
Published: (2025)
Auto-NBA: Efficient and Effective Search Over the Joint Space of Networks, Bitwidths, and Accelerators
by: Fu, Yonggan, et al.
Published: (2021)
by: Fu, Yonggan, et al.
Published: (2021)
Gen-NeRF: Efficient and Generalizable Neural Radiance Fields via Algorithm-Hardware Co-Design
by: Fu, Yonggan, et al.
Published: (2023)
by: Fu, Yonggan, et al.
Published: (2023)
2-in-1 Accelerator: Enabling Random Precision Switch for Winning Both Adversarial Robustness and Efficiency
by: Fu, Yonggan, et al.
Published: (2021)
by: Fu, Yonggan, et al.
Published: (2021)
A3C-S: Automated Agent Accelerator Co-Search towards Efficient Deep Reinforcement Learning
by: Fu, Yonggan, et al.
Published: (2021)
by: Fu, Yonggan, et al.
Published: (2021)
Post-Training Quantization for Vision Mamba with k-Scaled Quantization and Reparameterization
by: Shi, Bo-Yun, et al.
Published: (2025)
by: Shi, Bo-Yun, et al.
Published: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
by: Mozaffari, Mohammad, et al.
Published: (2024)
by: Mozaffari, Mohammad, et al.
Published: (2024)
EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training
by: Dorkin, Aleksei, et al.
Published: (2026)
by: Dorkin, Aleksei, et al.
Published: (2026)
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
by: Liang, Yanbiao, et al.
Published: (2025)
by: Liang, Yanbiao, et al.
Published: (2025)
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
MG-Verilog: Multi-grained Dataset Towards Enhanced LLM-assisted Verilog Generation
by: Zhang, Yongan, et al.
Published: (2024)
by: Zhang, Yongan, et al.
Published: (2024)
GauRast: Enhancing GPU Triangle Rasterizers to Accelerate 3D Gaussian Splatting
by: Li, Sixu, et al.
Published: (2025)
by: Li, Sixu, et al.
Published: (2025)
Similar Items
-
ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision Transformer
by: You, Haoran, et al.
Published: (2023) -
ShiftAddNAS: Hardware-Inspired Search for More Accurate and Efficient Neural Networks
by: You, Haoran, et al.
Published: (2022) -
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
by: You, Haoran, et al.
Published: (2024) -
ShiftAddNet: A Hardware-Inspired Deep Network
by: You, Haoran, et al.
Published: (2020) -
ShiftAddAug: Augment Multiplication-Free Tiny Neural Network with Hybrid Computation
by: Guo, Yipin, et al.
Published: (2024)