Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ho, Siu Hang, Ganesan, Prasad, Duong, Nguyen, Schlabig, Daniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916963354148864
author Ho, Siu Hang
Ganesan, Prasad
Duong, Nguyen
Schlabig, Daniel
author_facet Ho, Siu Hang
Ganesan, Prasad
Duong, Nguyen
Schlabig, Daniel
contents Efficient inference is a critical challenge in deep generative modeling, particularly as diffusion models grow in capacity and complexity. While increased complexity often improves accuracy, it raises compute costs, latency, and memory requirements. This work investigates techniques such as pruning, quantization, knowledge distillation, and simplified attention to reduce computational overhead without impacting performance. The study also explores the Mixture of Experts (MoE) approach to further enhance efficiency. These experiments provide insights into optimizing inference for the state-of-the-art Fast Diffusion Transformer (fast-DiT) model.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17894
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
Ho, Siu Hang
Ganesan, Prasad
Duong, Nguyen
Schlabig, Daniel
Machine Learning
68T07
I.2.6; I.5.1
Efficient inference is a critical challenge in deep generative modeling, particularly as diffusion models grow in capacity and complexity. While increased complexity often improves accuracy, it raises compute costs, latency, and memory requirements. This work investigates techniques such as pruning, quantization, knowledge distillation, and simplified attention to reduce computational overhead without impacting performance. The study also explores the Mixture of Experts (MoE) approach to further enhance efficiency. These experiments provide insights into optimizing inference for the state-of-the-art Fast Diffusion Transformer (fast-DiT) model.
title Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
topic Machine Learning
68T07
I.2.6; I.5.1
url https://arxiv.org/abs/2509.17894