Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914229945106432 |
|---|---|
| author | Agrawal, Amey Yadav, Mayank Kumar, Sukrit Agrawal, Anirudha Ghai, Garv Bera, Souradeep Pinto, Elton Gambhira, Sirish Adain, Mohammad Sohrab, Kasra Antonanzas, Chus Tumanov, Alexey |
| author_facet | Agrawal, Amey Yadav, Mayank Kumar, Sukrit Agrawal, Anirudha Ghai, Garv Bera, Souradeep Pinto, Elton Gambhira, Sirish Adain, Mohammad Sohrab, Kasra Antonanzas, Chus Tumanov, Alexey |
| contents | Deploying LLMs efficiently requires testing hundreds of serving configurations, but evaluating each one on a GPU cluster takes hours and costs thousands of dollars. Discrete-event simulators are faster and cheaper, but they require re-implementing the serving system's control logic -- a burden that compounds as frameworks evolve.
We present Revati, a time-warp emulator that enables performance modeling by directly executing real serving system code at simulation-like speed. The system intercepts CUDA API calls to virtualize device management, allowing serving frameworks to run without physical GPUs. Instead of executing GPU kernels, it performs time jumps -- fast-forwarding virtual time by predicted kernel durations. We propose a coordination protocol that synchronizes these jumps across distributed processes while preserving causality. On vLLM and SGLang, Revati achieves less than 5% prediction error across multiple models and parallelism configurations, while running 5-17x faster than real GPU execution. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_00397 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving Agrawal, Amey Yadav, Mayank Kumar, Sukrit Agrawal, Anirudha Ghai, Garv Bera, Souradeep Pinto, Elton Gambhira, Sirish Adain, Mohammad Sohrab, Kasra Antonanzas, Chus Tumanov, Alexey Distributed, Parallel, and Cluster Computing Machine Learning Deploying LLMs efficiently requires testing hundreds of serving configurations, but evaluating each one on a GPU cluster takes hours and costs thousands of dollars. Discrete-event simulators are faster and cheaper, but they require re-implementing the serving system's control logic -- a burden that compounds as frameworks evolve. We present Revati, a time-warp emulator that enables performance modeling by directly executing real serving system code at simulation-like speed. The system intercepts CUDA API calls to virtualize device management, allowing serving frameworks to run without physical GPUs. Instead of executing GPU kernels, it performs time jumps -- fast-forwarding virtual time by predicted kernel durations. We propose a coordination protocol that synchronizes these jumps across distributed processes while preserving causality. On vLLM and SGLang, Revati achieves less than 5% prediction error across multiple models and parallelism configurations, while running 5-17x faster than real GPU execution. |
| title | Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving |
| topic | Distributed, Parallel, and Cluster Computing Machine Learning |
| url | https://arxiv.org/abs/2601.00397 |