RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Zheng, Zihao, Tian, Sicheng, Cao, Hangyu, Li, Chenyue, Chen, Jiayu, Li, Maoliang, Sun, Xinhao, Zou, Hailong, Luo, Guojie, Chen, Xiang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models
por: Zheng, Zihao, et al.
Publicado: (2026)
por: Zheng, Zihao, et al.
Publicado: (2026)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
por: Li, Maoliang, et al.
Publicado: (2026)
por: Li, Maoliang, et al.
Publicado: (2026)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
por: Dong, Jiangwen, et al.
Publicado: (2025)
por: Dong, Jiangwen, et al.
Publicado: (2025)
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
por: Wei, Xinming, et al.
Publicado: (2025)
por: Wei, Xinming, et al.
Publicado: (2025)
Resilience Evaluation of Kubernetes in Cloud-Edge Environments via Failure Injection
por: Chen, Zihao, et al.
Publicado: (2025)
por: Chen, Zihao, et al.
Publicado: (2025)
Learning the Optimal Path and DNN Partition for Collaborative Edge Inference
por: Huang, Yin, et al.
Publicado: (2024)
por: Huang, Yin, et al.
Publicado: (2024)
Where to Split? A Pareto-Front Analysis of DNN Partitioning for Edge Inference
por: Masud, Adiba, et al.
Publicado: (2026)
por: Masud, Adiba, et al.
Publicado: (2026)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
por: Karfakis, George, et al.
Publicado: (2025)
por: Karfakis, George, et al.
Publicado: (2025)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
por: Zhang, Yida, et al.
Publicado: (2026)
por: Zhang, Yida, et al.
Publicado: (2026)
Administrative Decentralization in Edge-Cloud Multi-Agent for Mobile Automation
por: Li, Senyao, et al.
Publicado: (2026)
por: Li, Senyao, et al.
Publicado: (2026)
DSPE: Profit Maximization in Edge-Cloud Storage System using Dynamic Space Partitioning with Erasure Code
por: Roy, Shubhradeep, et al.
Publicado: (2025)
por: Roy, Shubhradeep, et al.
Publicado: (2025)
Cooperative Inference with Interleaved Operator Partitioning for CNNs
por: Liu, Zhibang, et al.
Publicado: (2024)
por: Liu, Zhibang, et al.
Publicado: (2024)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
por: Zhan, Huiyou, et al.
Publicado: (2025)
por: Zhan, Huiyou, et al.
Publicado: (2025)
Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
por: Li, Senyao, et al.
Publicado: (2025)
por: Li, Senyao, et al.
Publicado: (2025)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2026)
por: Yang, Zheming, et al.
Publicado: (2026)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
por: Raj, Suman, et al.
Publicado: (2024)
por: Raj, Suman, et al.
Publicado: (2024)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
por: Yu, Shibo, et al.
Publicado: (2025)
por: Yu, Shibo, et al.
Publicado: (2025)
H-EYE: Holistic Resource Modeling and Management for Diversely Scaled Edge-Cloud Systems
por: Dagli, Ismet, et al.
Publicado: (2024)
por: Dagli, Ismet, et al.
Publicado: (2024)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
por: Li, Rui, et al.
Publicado: (2024)
por: Li, Rui, et al.
Publicado: (2024)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
por: Han, Yunhe, et al.
Publicado: (2026)
por: Han, Yunhe, et al.
Publicado: (2026)
SynergAI: Edge-to-Cloud Synergy for Architecture-Driven High-Performance Orchestration for AI Inference
por: Stathopoulou, Foteini, et al.
Publicado: (2025)
por: Stathopoulou, Foteini, et al.
Publicado: (2025)
Large Language Model Partitioning for Low-Latency Inference at the Edge
por: Kafetzis, Dimitrios, et al.
Publicado: (2025)
por: Kafetzis, Dimitrios, et al.
Publicado: (2025)
AgentFlow: Resilient Adaptive Cloud-Edge Framework for Multi-Agent Coordination
por: Chen, Ching Han, et al.
Publicado: (2025)
por: Chen, Ching Han, et al.
Publicado: (2025)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
por: Li, Yuchen, et al.
Publicado: (2026)
por: Li, Yuchen, et al.
Publicado: (2026)
ECCENTRIC: Edge-Cloud Collaboration Framework for Distributed Inference Using Knowledge Adaptation
por: Kamani, Mohammad Mahdi, et al.
Publicado: (2025)
por: Kamani, Mohammad Mahdi, et al.
Publicado: (2025)
Matrix representation and GPU-optimized parallel B-spline computing
por: Wu, Jiayu, et al.
Publicado: (2025)
por: Wu, Jiayu, et al.
Publicado: (2025)
Distributed Edge Analytics in Edge-Fog-Cloud Continuum
por: Srirama, Satish Narayana
Publicado: (2024)
por: Srirama, Satish Narayana
Publicado: (2024)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
por: Liu, Zhibang, et al.
Publicado: (2025)
por: Liu, Zhibang, et al.
Publicado: (2025)
Edge-Cloud Collaborative Pothole Detection via Onboard Event Screening and Federated Temporal Segmentation
por: Wu, Yingjie, et al.
Publicado: (2026)
por: Wu, Yingjie, et al.
Publicado: (2026)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
por: Bournias, Ilias, et al.
Publicado: (2024)
por: Bournias, Ilias, et al.
Publicado: (2024)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
por: Yang, Zheming, et al.
Publicado: (2024)
por: Yang, Zheming, et al.
Publicado: (2024)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2025)
por: Yang, Zheming, et al.
Publicado: (2025)
Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities
por: Chen, Zhixiong, et al.
Publicado: (2026)
por: Chen, Zhixiong, et al.
Publicado: (2026)
MIRAGE: Runtime Scheduling for Multi-Vector Image Retrieval with Hierarchical Decomposition
por: Li, Maoliang, et al.
Publicado: (2025)
por: Li, Maoliang, et al.
Publicado: (2025)
GeoFaaS: An Edge-to-Cloud FaaS Platform
por: Malekabbasi, Mohammadreza, et al.
Publicado: (2024)
por: Malekabbasi, Mohammadreza, et al.
Publicado: (2024)
DTVM: Revolutionizing Smart Contract Execution with Determinism and Compatibility
por: Zhou, Wei, et al.
Publicado: (2025)
por: Zhou, Wei, et al.
Publicado: (2025)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
por: Zhang, Mingjin, et al.
Publicado: (2024)
por: Zhang, Mingjin, et al.
Publicado: (2024)
Bridge the Present and Future: A Cross-Layer Matching Game in Dynamic Cloud-Aided Mobile Edge Networks
por: Qi, Houyi, et al.
Publicado: (2023)
por: Qi, Houyi, et al.
Publicado: (2023)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
por: Sun, Mingyu, et al.
Publicado: (2025)
por: Sun, Mingyu, et al.
Publicado: (2025)
GenAI at the Edge: Comprehensive Survey on Empowering Edge Devices
por: Navardi, Mozhgan, et al.
Publicado: (2025)
por: Navardi, Mozhgan, et al.
Publicado: (2025)
Ejemplares similares
-
RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models
por: Zheng, Zihao, et al.
Publicado: (2026) -
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
por: Li, Maoliang, et al.
Publicado: (2026) -
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
por: Dong, Jiangwen, et al.
Publicado: (2025) -
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
por: Wei, Xinming, et al.
Publicado: (2025) -
Resilience Evaluation of Kubernetes in Cloud-Edge Environments via Failure Injection
por: Chen, Zihao, et al.
Publicado: (2025)