TorchAO: PyTorch-Native Training-to-Serving Model Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Or, Andrew, Jain, Apurva, Vega-Myhre, Daniel, Cai, Jesse, Hernandez, Charles David, Zheng, Zhenrui, Guessous, Driss, Kuznetsov, Vasiliy, Puhrsch, Christian, Saroufim, Mark, Rao, Supriya, Tran, Thien, Samardžić, Aleksandar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908460479676416
author Or, Andrew
Jain, Apurva
Vega-Myhre, Daniel
Cai, Jesse
Hernandez, Charles David
Zheng, Zhenrui
Guessous, Driss
Kuznetsov, Vasiliy
Puhrsch, Christian
Saroufim, Mark
Rao, Supriya
Tran, Thien
Samardžić, Aleksandar
author_facet Or, Andrew
Jain, Apurva
Vega-Myhre, Daniel
Cai, Jesse
Hernandez, Charles David
Zheng, Zhenrui
Guessous, Driss
Kuznetsov, Vasiliy
Puhrsch, Christian
Saroufim, Mark
Rao, Supriya
Tran, Thien
Samardžić, Aleksandar
contents We present TorchAO, a PyTorch-native model optimization framework leveraging quantization and sparsity to provide an end-to-end, training-to-serving workflow for AI models. TorchAO supports a variety of popular model optimization techniques, including FP8 quantized training, quantization-aware training (QAT), post-training quantization (PTQ), and 2:4 sparsity, and leverages a novel tensor subclass abstraction to represent a variety of widely-used, backend agnostic low precision data types, including INT4, INT8, FP8, MXFP4, MXFP6, and MXFP8. TorchAO integrates closely with the broader ecosystem at each step of the model optimization pipeline, from pre-training (TorchTitan) to fine-tuning (TorchTune, Axolotl) to serving (HuggingFace, vLLM, SGLang, ExecuTorch), connecting an otherwise fragmented space in a single, unified workflow. TorchAO has enabled recent launches of the quantized Llama 3.2 1B/3B and LlamaGuard3-8B models and is open-source at https://github.com/pytorch/ao/.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16099
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TorchAO: PyTorch-Native Training-to-Serving Model Optimization
Or, Andrew
Jain, Apurva
Vega-Myhre, Daniel
Cai, Jesse
Hernandez, Charles David
Zheng, Zhenrui
Guessous, Driss
Kuznetsov, Vasiliy
Puhrsch, Christian
Saroufim, Mark
Rao, Supriya
Tran, Thien
Samardžić, Aleksandar
Machine Learning
We present TorchAO, a PyTorch-native model optimization framework leveraging quantization and sparsity to provide an end-to-end, training-to-serving workflow for AI models. TorchAO supports a variety of popular model optimization techniques, including FP8 quantized training, quantization-aware training (QAT), post-training quantization (PTQ), and 2:4 sparsity, and leverages a novel tensor subclass abstraction to represent a variety of widely-used, backend agnostic low precision data types, including INT4, INT8, FP8, MXFP4, MXFP6, and MXFP8. TorchAO integrates closely with the broader ecosystem at each step of the model optimization pipeline, from pre-training (TorchTitan) to fine-tuning (TorchTune, Axolotl) to serving (HuggingFace, vLLM, SGLang, ExecuTorch), connecting an otherwise fragmented space in a single, unified workflow. TorchAO has enabled recent launches of the quantized Llama 3.2 1B/3B and LlamaGuard3-8B models and is open-source at https://github.com/pytorch/ao/.
title TorchAO: PyTorch-Native Training-to-Serving Model Optimization
topic Machine Learning
url https://arxiv.org/abs/2507.16099