Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Santavas, Nicholas, Eissa, Kareem, Cieplicka, Patrycja, Florek, Piotr, Nulli, Matteo, Vasilev, Stefan, Hashemi, Seyyed Hadi, Gasteratos, Antonios, Khadivi, Shahram
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911404307513344
author Santavas, Nicholas
Eissa, Kareem
Cieplicka, Patrycja
Florek, Piotr
Nulli, Matteo
Vasilev, Stefan
Hashemi, Seyyed Hadi
Gasteratos, Antonios
Khadivi, Shahram
author_facet Santavas, Nicholas
Eissa, Kareem
Cieplicka, Patrycja
Florek, Piotr
Nulli, Matteo
Vasilev, Stefan
Hashemi, Seyyed Hadi
Gasteratos, Antonios
Khadivi, Shahram
contents Enterprise LLM deployment faces a critical scalability challenge: organizations must optimize models systematically to scale AI initiatives within constrained compute budgets, yet the specialized expertise required for manual optimization remains a niche and scarce skillset. This challenge is particularly evident in managing GPU utilization across heterogeneous infrastructure while enabling teams with diverse workloads and limited LLM optimization experience to deploy models efficiently. We present OptiKIT, a distributed LLM optimization framework that democratizes model compression and tuning by automating complex optimization workflows for non-expert teams. OptiKIT provides dynamic resource allocation, staged pipeline execution with automatic cleanup, and seamless enterprise integration. In production, it delivers more than 2x GPU throughput improvement while empowering application teams to achieve consistent performance improvements without deep LLM optimization expertise. We share both the platform design and key engineering insights into resource allocation algorithms, pipeline orchestration, and integration patterns that enable large-scale, production-grade democratization of model optimization. Finally, we open-source the system to enable external contributions and broader reproducibility.
format Preprint
id arxiv_https___arxiv_org_abs_2601_20408
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT
Santavas, Nicholas
Eissa, Kareem
Cieplicka, Patrycja
Florek, Piotr
Nulli, Matteo
Vasilev, Stefan
Hashemi, Seyyed Hadi
Gasteratos, Antonios
Khadivi, Shahram
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Enterprise LLM deployment faces a critical scalability challenge: organizations must optimize models systematically to scale AI initiatives within constrained compute budgets, yet the specialized expertise required for manual optimization remains a niche and scarce skillset. This challenge is particularly evident in managing GPU utilization across heterogeneous infrastructure while enabling teams with diverse workloads and limited LLM optimization experience to deploy models efficiently. We present OptiKIT, a distributed LLM optimization framework that democratizes model compression and tuning by automating complex optimization workflows for non-expert teams. OptiKIT provides dynamic resource allocation, staged pipeline execution with automatic cleanup, and seamless enterprise integration. In production, it delivers more than 2x GPU throughput improvement while empowering application teams to achieve consistent performance improvements without deep LLM optimization expertise. We share both the platform design and key engineering insights into resource allocation algorithms, pipeline orchestration, and integration patterns that enable large-scale, production-grade democratization of model optimization. Finally, we open-source the system to enable external contributions and broader reproducibility.
title Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2601.20408