Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916963354148864 |
|---|---|
| author | Ho, Siu Hang Ganesan, Prasad Duong, Nguyen Schlabig, Daniel |
| author_facet | Ho, Siu Hang Ganesan, Prasad Duong, Nguyen Schlabig, Daniel |
| contents | Efficient inference is a critical challenge in deep generative modeling, particularly as diffusion models grow in capacity and complexity. While increased complexity often improves accuracy, it raises compute costs, latency, and memory requirements. This work investigates techniques such as pruning, quantization, knowledge distillation, and simplified attention to reduce computational overhead without impacting performance. The study also explores the Mixture of Experts (MoE) approach to further enhance efficiency. These experiments provide insights into optimizing inference for the state-of-the-art Fast Diffusion Transformer (fast-DiT) model. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_17894 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark Ho, Siu Hang Ganesan, Prasad Duong, Nguyen Schlabig, Daniel Machine Learning 68T07 I.2.6; I.5.1 Efficient inference is a critical challenge in deep generative modeling, particularly as diffusion models grow in capacity and complexity. While increased complexity often improves accuracy, it raises compute costs, latency, and memory requirements. This work investigates techniques such as pruning, quantization, knowledge distillation, and simplified attention to reduce computational overhead without impacting performance. The study also explores the Mixture of Experts (MoE) approach to further enhance efficiency. These experiments provide insights into optimizing inference for the state-of-the-art Fast Diffusion Transformer (fast-DiT) model. |
| title | Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark |
| topic | Machine Learning 68T07 I.2.6; I.5.1 |
| url | https://arxiv.org/abs/2509.17894 |