SpotVista: Availability-Aware Recommendation System for Reliable and Cost-Efficient Multi-Node Spot Instances
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Taeyoon, Kim, Kyumin, Kim, Kyunghwan, Kim, Hayoung, Jeong, Seungwoo, Song, Moohyun, Lee, Kyungyong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Ding-Dong Ditch: Peeking Into Spot Instance Availability
por: Kim, Kyumin, et al.
Publicado: (2026)
por: Kim, Kyumin, et al.
Publicado: (2026)
KubePACS: Kubernetes Cluster Using Performant, Highly Available, and Cost Efficient Spot Instances
por: Kim, Taeyoon, et al.
Publicado: (2026)
por: Kim, Taeyoon, et al.
Publicado: (2026)
DeepVM: Integrating Spot and On-Demand VMs for Cost-Efficient Deep Learning Clusters in the Cloud
por: Kim, Yoochan, et al.
Publicado: (2024)
por: Kim, Yoochan, et al.
Publicado: (2024)
Why Do AI Agents Systematically Fail at Cloud Root Cause Analysis?
por: Kim, Taeyoon, et al.
Publicado: (2026)
por: Kim, Taeyoon, et al.
Publicado: (2026)
SpotKube: Cost-Optimal Microservices Deployment with Cluster Autoscaling and Spot Pricing
por: Edirisinghe, Dasith, et al.
Publicado: (2024)
por: Edirisinghe, Dasith, et al.
Publicado: (2024)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
por: Duan, Jiaang, et al.
Publicado: (2025)
por: Duan, Jiaang, et al.
Publicado: (2025)
FedCostAware: Enabling Cost-Aware Federated Learning on the Cloud
por: Sinha, Aditya, et al.
Publicado: (2025)
por: Sinha, Aditya, et al.
Publicado: (2025)
AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
por: Park, Seongyeon, et al.
Publicado: (2024)
por: Park, Seongyeon, et al.
Publicado: (2024)
SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
por: Li, Zhifei, et al.
Publicado: (2026)
por: Li, Zhifei, et al.
Publicado: (2026)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
por: Kim, Kihyun, et al.
Publicado: (2025)
por: Kim, Kihyun, et al.
Publicado: (2025)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
por: Mao, Ziming, et al.
Publicado: (2024)
por: Mao, Ziming, et al.
Publicado: (2024)
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
por: Lee, Sunjung, et al.
Publicado: (2026)
por: Lee, Sunjung, et al.
Publicado: (2026)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
por: Wang, Shaoyu, et al.
Publicado: (2025)
por: Wang, Shaoyu, et al.
Publicado: (2025)
DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization
por: An, Hyeonjun, et al.
Publicado: (2026)
por: An, Hyeonjun, et al.
Publicado: (2026)
CusADi: A GPU Parallelization Framework for Symbolic Expressions and Optimal Control
por: Jeon, Se Hwan, et al.
Publicado: (2024)
por: Jeon, Se Hwan, et al.
Publicado: (2024)
Optimizing Spot Instance Reliability and Security Using Cloud-Native Data and Tools
por: Saqib, Muhammad, et al.
Publicado: (2025)
por: Saqib, Muhammad, et al.
Publicado: (2025)
Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
por: Wang, Yuxiao, et al.
Publicado: (2025)
por: Wang, Yuxiao, et al.
Publicado: (2025)
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
por: Park, Seongyeon, et al.
Publicado: (2025)
por: Park, Seongyeon, et al.
Publicado: (2025)
AI-Driven Multi-Region Provisioning for Cloud Services Using Spot Fleets
por: Fabra, Javier, et al.
Publicado: (2026)
por: Fabra, Javier, et al.
Publicado: (2026)
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
por: Oh, Hyungjun, et al.
Publicado: (2024)
por: Oh, Hyungjun, et al.
Publicado: (2024)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
por: Kim, Joon Ha, et al.
Publicado: (2026)
por: Kim, Joon Ha, et al.
Publicado: (2026)
Deadline-Aware Online Scheduling for LLM Fine-Tuning with Spot Market Predictions
por: Kong, Linggao, et al.
Publicado: (2025)
por: Kong, Linggao, et al.
Publicado: (2025)
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
por: Wang, Yidi, et al.
Publicado: (2024)
por: Wang, Yidi, et al.
Publicado: (2024)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
por: Lee, Sanghyeon, et al.
Publicado: (2025)
por: Lee, Sanghyeon, et al.
Publicado: (2025)
A Case Study of API Design for Interoperability and Security of the Internet of Things
por: Kim, Dongha, et al.
Publicado: (2024)
por: Kim, Dongha, et al.
Publicado: (2024)
Traversal Learning: A Lossless And Efficient Distributed Learning Framework
por: Batbaatar, Erdenebileg, et al.
Publicado: (2025)
por: Batbaatar, Erdenebileg, et al.
Publicado: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
por: Park, Gunho, et al.
Publicado: (2022)
por: Park, Gunho, et al.
Publicado: (2022)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
por: Yoon, Dongha, et al.
Publicado: (2025)
por: Yoon, Dongha, et al.
Publicado: (2025)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
por: Chang, Dali, et al.
Publicado: (2026)
por: Chang, Dali, et al.
Publicado: (2026)
GraNNDis: Efficient Unified Distributed Training Framework for Deep GNNs on Large Clusters
por: Song, Jaeyong, et al.
Publicado: (2023)
por: Song, Jaeyong, et al.
Publicado: (2023)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
por: Li, Shiju, et al.
Publicado: (2025)
por: Li, Shiju, et al.
Publicado: (2025)
Pairbot: A Novel Model for Autonomous Mobile Robot Systems Consisting of Paired Robots
por: Kim, Yonghwan, et al.
Publicado: (2020)
por: Kim, Yonghwan, et al.
Publicado: (2020)
Are Your Epochs Too Epic? Batch Free Can Be Harmful
por: Kim, Daewoo, et al.
Publicado: (2024)
por: Kim, Daewoo, et al.
Publicado: (2024)
Contextual Chain: Single-State Ledger Design for Mobile/IoT Networks with Frequent Partitions
por: Kim, Song-Ju
Publicado: (2026)
por: Kim, Song-Ju
Publicado: (2026)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
por: Zhang, Han, et al.
Publicado: (2026)
por: Zhang, Han, et al.
Publicado: (2026)
Efficient Coordination for Distributed Discrete-Event Systems
por: Jun, Byeonggil, et al.
Publicado: (2024)
por: Jun, Byeonggil, et al.
Publicado: (2024)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
por: Du, Jiangsu, et al.
Publicado: (2025)
por: Du, Jiangsu, et al.
Publicado: (2025)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
por: Griggs, Tyler, et al.
Publicado: (2024)
por: Griggs, Tyler, et al.
Publicado: (2024)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
por: Kim, Heehoon, et al.
Publicado: (2026)
por: Kim, Heehoon, et al.
Publicado: (2026)
TAPAAL SMC: Statistical Model Checking of Stochastic Timed-Arc Petri Nets
por: Dubois, Tanguy, et al.
Publicado: (2026)
por: Dubois, Tanguy, et al.
Publicado: (2026)
Ejemplares similares
-
Ding-Dong Ditch: Peeking Into Spot Instance Availability
por: Kim, Kyumin, et al.
Publicado: (2026) -
KubePACS: Kubernetes Cluster Using Performant, Highly Available, and Cost Efficient Spot Instances
por: Kim, Taeyoon, et al.
Publicado: (2026) -
DeepVM: Integrating Spot and On-Demand VMs for Cost-Efficient Deep Learning Clusters in the Cloud
por: Kim, Yoochan, et al.
Publicado: (2024) -
Why Do AI Agents Systematically Fail at Cloud Root Cause Analysis?
por: Kim, Taeyoon, et al.
Publicado: (2026) -
SpotKube: Cost-Optimal Microservices Deployment with Cluster Autoscaling and Spot Pricing
por: Edirisinghe, Dasith, et al.
Publicado: (2024)