Saved in:
Bibliographic Details
Main Authors: Rajbhandari, Samyam, Hidayetoglu, Mert, Qiao, Aurick, Wang, Ye, Yang, Juncheng, Rasley, Jeff, Wyatt, Michael, He, Yuxiong
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.11830
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Inference is now the dominant AI workload, yet existing systems force trade-offs between latency, throughput, and cost. Arctic Inference, an open-source vLLM plugin from Snowflake AI Research, introduces Shift Parallelism, a dynamic parallelism strategy that adapts to real-world traffic while integrating speculative decoding, SwiftKV compute reduction, and optimized embedding inference. It achieves up to 3.4 times faster request completion, 1.75 times faster generation, and 1.6M tokens/sec per GPU for embeddings, outperforming both latency- and throughput-optimized deployments. Already powering Snowflake Cortex AI, Arctic Inference delivers state-of-the-art, cost-effective inference for enterprise AI and is now available to the community.