SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhifei, Xia, Tian, Mao, Ziming, Zhou, Zihan, Jackson, Ethan J., Kerney, Jamison, Wu, Zhanghao, Mishra, Pratik, Xu, Yi, Qiao, Yifan, Shenker, Scott, Stoica, Ion |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
by: Xia, Tian, et al.
Published: (2025)
by: Xia, Tian, et al.
Published: (2025)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
by: Mao, Ziming, et al.
Published: (2024)
by: Mao, Ziming, et al.
Published: (2024)
Revisiting Cache Freshness for Emerging Real-Time Applications
by: Mao, Ziming, et al.
Published: (2024)
by: Mao, Ziming, et al.
Published: (2024)
Extracting Database Access-control Policies From Web Applications
by: Zhang, Wen, et al.
Published: (2024)
by: Zhang, Wen, et al.
Published: (2024)
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
by: Cao, Shiyi, et al.
Published: (2026)
by: Cao, Shiyi, et al.
Published: (2026)
Pie: Pooling CPU Memory for LLM Inference
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
Some Present-Day Problems of Romanian Library Science
by: Stoica, Ion
Published: (1973)
by: Stoica, Ion
Published: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
by: Stoica, Ion
Published: (1972)
by: Stoica, Ion
Published: (1972)
Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI
by: Xia, Tian, et al.
Published: (2026)
by: Xia, Tian, et al.
Published: (2026)
Managing Bandwidth: The Key to Cloud-Assisted Autonomous Driving
by: Krentsel, Alexander, et al.
Published: (2024)
by: Krentsel, Alexander, et al.
Published: (2024)
UCCL-EP: Portable Expert-Parallel Communication
by: Mao, Ziming, et al.
Published: (2025)
by: Mao, Ziming, et al.
Published: (2025)
TURBO: Utility-Aware Bandwidth Allocation for Cloud-Augmented Autonomous Control
by: Schafhalter, Peter, et al.
Published: (2025)
by: Schafhalter, Peter, et al.
Published: (2025)
MPC-Minimized Secure LLM Inference
by: Rathee, Deevashwer, et al.
Published: (2024)
by: Rathee, Deevashwer, et al.
Published: (2024)
Nomad: Autonomous Exploration and Discovery
by: Jia, Bokang, et al.
Published: (2026)
by: Jia, Bokang, et al.
Published: (2026)
The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution
by: Luan, Frank Sifei, et al.
Published: (2025)
by: Luan, Frank Sifei, et al.
Published: (2025)
SpotVista: Availability-Aware Recommendation System for Reliable and Cost-Efficient Multi-Node Spot Instances
by: Kim, Taeyoon, et al.
Published: (2026)
by: Kim, Taeyoon, et al.
Published: (2026)
SkyStore: Cost-Optimized Object Storage Across Regions and Clouds
by: Liu, Shu, et al.
Published: (2025)
by: Liu, Shu, et al.
Published: (2025)
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
by: Mishra, Mayank, et al.
Published: (2026)
by: Mishra, Mayank, et al.
Published: (2026)
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
by: Guo, Wentao, et al.
Published: (2025)
by: Guo, Wentao, et al.
Published: (2025)
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
by: Wu, Yongji, et al.
Published: (2025)
by: Wu, Yongji, et al.
Published: (2025)
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching
by: Zhao, Yilong, et al.
Published: (2024)
by: Zhao, Yilong, et al.
Published: (2024)
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
by: Chen, Lingjiao, et al.
Published: (2026)
by: Chen, Lingjiao, et al.
Published: (2026)
Nomadic Peoples
Published: (2023)
Published: (2023)
Nomad Properties
Published: (2026)
Published: (2026)
Nomadic Camera
Published: (2026)
Published: (2026)
Nomadic Connectivity
by: Butter, Inge
Published: (2025)
by: Butter, Inge
Published: (2025)
Dance of the Nomad
by: McCulloch, Ann
Published: (2013)
by: McCulloch, Ann
Published: (2013)
A field guide to the land snails of Britain and north-west Europe / M.P. Kerney and R.A.D. Cameron ; illustrated by Gordon Riley
by: Kerney, M. P. (Michael P.)
Published: (1916)
by: Kerney, M. P. (Michael P.)
Published: (1916)
Beyond Possessions: Nomadic Living Sparks Minimalism Tendency and Preference for Experiential Purchases
by: Hanyu (Yuki) Chen, et al.
Published: (2025)
by: Hanyu (Yuki) Chen, et al.
Published: (2025)
RedunCut: Measurement-Driven Sampling and Accuracy Performance Modeling for Low-Cost Live Video Analytics
by: Sela, Gur-Eyal, et al.
Published: (2025)
by: Sela, Gur-Eyal, et al.
Published: (2025)
LEANN: A Low-Storage Vector Index
by: Wang, Yichuan, et al.
Published: (2025)
by: Wang, Yichuan, et al.
Published: (2025)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
by: Yu, Shan, et al.
Published: (2025)
by: Yu, Shan, et al.
Published: (2025)
KubePACS: Kubernetes Cluster Using Performant, Highly Available, and Cost Efficient Spot Instances
by: Kim, Taeyoon, et al.
Published: (2026)
by: Kim, Taeyoon, et al.
Published: (2026)
Nomads, Empires, States
by: van der Pijl, Kees
Published: (2018)
by: van der Pijl, Kees
Published: (2018)
Problems of Nomadism in the Sahara
by: Robert Capot-Rey
Published: (1964)
by: Robert Capot-Rey
Published: (1964)
Formulation Correction for Spot Color Inks Under Batch‐to‐Batch Variation of Substrates and Inks
by: Junfeng Li, et al.
Published: (2025)
by: Junfeng Li, et al.
Published: (2025)
Efficient LLM Scheduling by Learning to Rank
by: Fu, Yichao, et al.
Published: (2024)
by: Fu, Yichao, et al.
Published: (2024)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025)
by: Xing, Jiarong, et al.
Published: (2025)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
by: Griggs, Tyler, et al.
Published: (2024)
by: Griggs, Tyler, et al.
Published: (2024)
Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration
by: Xiang, Lingfeng, et al.
Published: (2024)
by: Xiang, Lingfeng, et al.
Published: (2024)
Similar Items
-
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
by: Xia, Tian, et al.
Published: (2025) -
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
by: Mao, Ziming, et al.
Published: (2024) -
Revisiting Cache Freshness for Emerging Real-Time Applications
by: Mao, Ziming, et al.
Published: (2024) -
Extracting Database Access-control Policies From Web Applications
by: Zhang, Wen, et al.
Published: (2024) -
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
by: Cao, Shiyi, et al.
Published: (2026)