Minions: Cost-efficient Collaboration Between On-device and Cloud Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Narayan, Avanika, Biderman, Dan, Eyuboglu, Sabri, May, Avner, Linderman, Scott, Zou, James, Re, Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
by: Wang, Siqi, et al.
Published: (2024)
by: Wang, Siqi, et al.
Published: (2024)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
ECC-SNN: Cost-Effective Edge-Cloud Collaboration for Spiking Neural Networks
by: Yu, Di, et al.
Published: (2025)
by: Yu, Di, et al.
Published: (2025)
FedCostAware: Enabling Cost-Aware Federated Learning on the Cloud
by: Sinha, Aditya, et al.
Published: (2025)
by: Sinha, Aditya, et al.
Published: (2025)
SimDC: A High-Fidelity Device Simulation Platform for Device-Cloud Collaborative Computing
by: Pei, Ruiguang, et al.
Published: (2025)
by: Pei, Ruiguang, et al.
Published: (2025)
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
by: Chang, Tzu-Tao, et al.
Published: (2025)
by: Chang, Tzu-Tao, et al.
Published: (2025)
SkyStore: Cost-Optimized Object Storage Across Regions and Clouds
by: Liu, Shu, et al.
Published: (2025)
by: Liu, Shu, et al.
Published: (2025)
Adaptive K-PackCache: Cost-Centric Data Caching in Cloud
by: Sarkar, Suvarthi, et al.
Published: (2025)
by: Sarkar, Suvarthi, et al.
Published: (2025)
Propius: A Platform for Collaborative Machine Learning across the Edge and the Cloud
by: Ding, Eric
Published: (2025)
by: Ding, Eric
Published: (2025)
Kubernetes in the Cloud vs. Bare Metal: A Comparative Study of Network Costs
by: Redoli, Rodrigo Mompo, et al.
Published: (2025)
by: Redoli, Rodrigo Mompo, et al.
Published: (2025)
Possible Futures for Cloud Cost Models
by: Sochat, Vanessa, et al.
Published: (2025)
by: Sochat, Vanessa, et al.
Published: (2025)
Collaborative Multi-Agent Reinforcement Learning Approach for Elastic Cloud Resource Scaling
by: Fang, Bruce, et al.
Published: (2025)
by: Fang, Bruce, et al.
Published: (2025)
ML-ECS: A Collaborative Multimodal Learning Framework for Edge-Cloud Synergies
by: Liu, Yuze, et al.
Published: (2026)
by: Liu, Yuze, et al.
Published: (2026)
KCES: A Workflow Containerization Scheduling Scheme Under Cloud-Edge Collaboration Framework
by: Shan, Chenggang, et al.
Published: (2024)
by: Shan, Chenggang, et al.
Published: (2024)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
by: Zhang, Yida, et al.
Published: (2026)
by: Zhang, Yida, et al.
Published: (2026)
Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Optimization
by: Gao, Luyao, et al.
Published: (2024)
by: Gao, Luyao, et al.
Published: (2024)
Collaborative State Machines: A Better Programming Model for the Cloud-Edge-IoT Continuum
by: Etheredge, Marlon, et al.
Published: (2025)
by: Etheredge, Marlon, et al.
Published: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
by: Han, Yunhe, et al.
Published: (2026)
by: Han, Yunhe, et al.
Published: (2026)
A Deep Reinforcement Learning Approach for Cost Optimized Workflow Scheduling in Cloud Computing Environments
by: Jayanetti, Amanda, et al.
Published: (2024)
by: Jayanetti, Amanda, et al.
Published: (2024)
DeepVM: Integrating Spot and On-Demand VMs for Cost-Efficient Deep Learning Clusters in the Cloud
by: Kim, Yoochan, et al.
Published: (2024)
by: Kim, Yoochan, et al.
Published: (2024)
Agora: Bridging the GPU Cloud Resource-Price Disconnect
by: McDougall, Ian, et al.
Published: (2025)
by: McDougall, Ian, et al.
Published: (2025)
Edge-Cloud Collaborative Pothole Detection via Onboard Event Screening and Federated Temporal Segmentation
by: Wu, Yingjie, et al.
Published: (2026)
by: Wu, Yingjie, et al.
Published: (2026)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
by: Tang, Lujie, et al.
Published: (2024)
by: Tang, Lujie, et al.
Published: (2024)
Revisiting the Time Cost Model of AllReduce
by: Xiong, Dian, et al.
Published: (2024)
by: Xiong, Dian, et al.
Published: (2024)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2026)
by: Yang, Zheming, et al.
Published: (2026)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
by: Li, Yuchen, et al.
Published: (2026)
by: Li, Yuchen, et al.
Published: (2026)
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
by: Yang, Zheming, et al.
Published: (2025)
by: Yang, Zheming, et al.
Published: (2025)
Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
by: Li, Senyao, et al.
Published: (2025)
by: Li, Senyao, et al.
Published: (2025)
EES-CND: Collaborative Neural Decision-Making for Drift-Aware Fault-Tolerant Edge-Cloud Service Placement
by: Herabad, Mohammadsadeq Garshasbi, et al.
Published: (2026)
by: Herabad, Mohammadsadeq Garshasbi, et al.
Published: (2026)
FASTEN: Towards a FAult-tolerant and STorage EfficieNt Cloud: Balancing Between Replication and Deduplication
by: Ahmed, Sabbir, et al.
Published: (2023)
by: Ahmed, Sabbir, et al.
Published: (2023)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2025)
by: Yang, Zheming, et al.
Published: (2025)
Efficient Local-to-Global Collaborative Perception via Joint Communication and Computation Optimization
by: Zhang, Hui, et al.
Published: (2026)
by: Zhang, Hui, et al.
Published: (2026)
Performance Cost Tradeoffs in Intelligent Load Balancing for Multi Data Center Cloud Systems: From Static Policies to Adaptive Resource Distribution
by: Najafabadi, Saeid Aghasoleymani, et al.
Published: (2025)
by: Najafabadi, Saeid Aghasoleymani, et al.
Published: (2025)
Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
by: Yao, Chenxuan, et al.
Published: (2025)
by: Yao, Chenxuan, et al.
Published: (2025)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
by: Tharwani, Jay, et al.
Published: (2025)
by: Tharwani, Jay, et al.
Published: (2025)
Cloud Revolution: Tracing the Origins and Rise of Cloud Computing
by: Gurung, Deepa, et al.
Published: (2025)
by: Gurung, Deepa, et al.
Published: (2025)
LaissezCloud: Continuous Resource Renegotiation for the Public Cloud
by: Harith, Tejas, et al.
Published: (2026)
by: Harith, Tejas, et al.
Published: (2026)
Huawei Cloud Model-as-a-Service on the CloudMatrix384 SuperPod
by: Xiao, Ao, et al.
Published: (2025)
by: Xiao, Ao, et al.
Published: (2025)
FourierCompress: Layer-Aware Spectral Activation Compression for Efficient and Accurate Collaborative LLM Inference
by: Ma, Jian, et al.
Published: (2025)
by: Ma, Jian, et al.
Published: (2025)
XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing
by: Hoefler, Torsten, et al.
Published: (2024)
by: Hoefler, Torsten, et al.
Published: (2024)
Similar Items
-
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
by: Wang, Siqi, et al.
Published: (2024) -
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
by: Jiang, Youhe, et al.
Published: (2025) -
ECC-SNN: Cost-Effective Edge-Cloud Collaboration for Spiking Neural Networks
by: Yu, Di, et al.
Published: (2025) -
FedCostAware: Enabling Cost-Aware Federated Learning on the Cloud
by: Sinha, Aditya, et al.
Published: (2025) -
SimDC: A High-Fidelity Device Simulation Platform for Device-Cloud Collaborative Computing
by: Pei, Ruiguang, et al.
Published: (2025)