HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Dong, Jiangwen, Li, Jiayu, Zheng, Tianhang, Lin, Wanyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
por: Yu, Shibo, et al.
Publicado: (2025)
por: Yu, Shibo, et al.
Publicado: (2025)
HybridFlow: A Flexible and Efficient RLHF Framework
por: Sheng, Guangming, et al.
Publicado: (2024)
por: Sheng, Guangming, et al.
Publicado: (2024)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2026)
por: Yang, Zheming, et al.
Publicado: (2026)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
por: Zhang, Yida, et al.
Publicado: (2026)
por: Zhang, Yida, et al.
Publicado: (2026)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2025)
por: Yang, Zheming, et al.
Publicado: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
por: Lin, Mao, et al.
Publicado: (2026)
por: Lin, Mao, et al.
Publicado: (2026)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
por: Lin, Haoran, et al.
Publicado: (2025)
por: Lin, Haoran, et al.
Publicado: (2025)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
por: Zhang, Mingjin, et al.
Publicado: (2024)
por: Zhang, Mingjin, et al.
Publicado: (2024)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
por: Raj, Suman, et al.
Publicado: (2024)
por: Raj, Suman, et al.
Publicado: (2024)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
por: Xu, Chuhao, et al.
Publicado: (2025)
por: Xu, Chuhao, et al.
Publicado: (2025)
AgentFlow: Resilient Adaptive Cloud-Edge Framework for Multi-Agent Coordination
por: Chen, Ching Han, et al.
Publicado: (2025)
por: Chen, Ching Han, et al.
Publicado: (2025)
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
por: Zheng, Zihao, et al.
Publicado: (2026)
por: Zheng, Zihao, et al.
Publicado: (2026)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
por: Zhan, Huiyou, et al.
Publicado: (2025)
por: Zhan, Huiyou, et al.
Publicado: (2025)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
por: Yang, Yi, et al.
Publicado: (2025)
por: Yang, Yi, et al.
Publicado: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
por: Han, Yunhe, et al.
Publicado: (2026)
por: Han, Yunhe, et al.
Publicado: (2026)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
por: Lee, Sanghyeon, et al.
Publicado: (2025)
por: Lee, Sanghyeon, et al.
Publicado: (2025)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
por: Xie, Jincheng, et al.
Publicado: (2026)
por: Xie, Jincheng, et al.
Publicado: (2026)
A Cloud-Based Spatio-Temporal GNN-Transformer Hybrid Model for Traffic Flow Forecasting with External Feature Integration
por: Zheng, Zhuo, et al.
Publicado: (2025)
por: Zheng, Zhuo, et al.
Publicado: (2025)
Token Level Routing Inference System for Edge Devices
por: She, Jianshu, et al.
Publicado: (2025)
por: She, Jianshu, et al.
Publicado: (2025)
Distributed Resource Selection for Self-Organising Cloud-Edge Systems
por: Renau, Quentin, et al.
Publicado: (2025)
por: Renau, Quentin, et al.
Publicado: (2025)
Adaptive AI-based Decentralized Resource Management in the Cloud-Edge Continuum
por: Li, Lanpei, et al.
Publicado: (2025)
por: Li, Lanpei, et al.
Publicado: (2025)
Cloud Native System for LLM Inference Serving
por: Xu, Minxian, et al.
Publicado: (2025)
por: Xu, Minxian, et al.
Publicado: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
por: Wu, Yu, et al.
Publicado: (2025)
por: Wu, Yu, et al.
Publicado: (2025)
Covariance-Guided Resource Adaptive Learning for Efficient Edge Inference
por: Nabhaan, Ahmad N. L., et al.
Publicado: (2026)
por: Nabhaan, Ahmad N. L., et al.
Publicado: (2026)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
por: Arya, Mayank, et al.
Publicado: (2025)
por: Arya, Mayank, et al.
Publicado: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
por: Rajashekar, Kolichala, et al.
Publicado: (2025)
por: Rajashekar, Kolichala, et al.
Publicado: (2025)
REACH: Reinforcement Learning for Adaptive Microservice Rescheduling in the Cloud-Edge Continuum
por: Bai, Xu, et al.
Publicado: (2025)
por: Bai, Xu, et al.
Publicado: (2025)
CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration
por: Jin, Hongpeng, et al.
Publicado: (2024)
por: Jin, Hongpeng, et al.
Publicado: (2024)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
por: Khoshsirat, Aria, et al.
Publicado: (2024)
por: Khoshsirat, Aria, et al.
Publicado: (2024)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
por: Yu, Minchen, et al.
Publicado: (2023)
por: Yu, Minchen, et al.
Publicado: (2023)
H-EYE: Holistic Resource Modeling and Management for Diversely Scaled Edge-Cloud Systems
por: Dagli, Ismet, et al.
Publicado: (2024)
por: Dagli, Ismet, et al.
Publicado: (2024)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
por: Sun, Mingyu, et al.
Publicado: (2025)
por: Sun, Mingyu, et al.
Publicado: (2025)
Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
por: Li, Senyao, et al.
Publicado: (2025)
por: Li, Senyao, et al.
Publicado: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
por: He, Yiyuan, et al.
Publicado: (2024)
por: He, Yiyuan, et al.
Publicado: (2024)
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
por: Behera, Adarsh Prasad, et al.
Publicado: (2024)
por: Behera, Adarsh Prasad, et al.
Publicado: (2024)
Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing
por: Sheng, Jinhao, et al.
Publicado: (2025)
por: Sheng, Jinhao, et al.
Publicado: (2025)
Adaptive, Efficient and Fair Resource Allocation in Cloud Datacenters leveraging Weighted A3C Deep Reinforcement Learning
por: Kumari, Suchi, et al.
Publicado: (2025)
por: Kumari, Suchi, et al.
Publicado: (2025)
SynergAI: Edge-to-Cloud Synergy for Architecture-Driven High-Performance Orchestration for AI Inference
por: Stathopoulou, Foteini, et al.
Publicado: (2025)
por: Stathopoulou, Foteini, et al.
Publicado: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
por: Chow, Will
Publicado: (2025)
por: Chow, Will
Publicado: (2025)
ARRC: Explainable, Workflow-Integrated Recommender for Sustainable Resource Optimization Across the Edge-Cloud Continuum
por: Jahnke, Brian-Frederik, et al.
Publicado: (2025)
por: Jahnke, Brian-Frederik, et al.
Publicado: (2025)
Ejemplares similares
-
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
por: Yu, Shibo, et al.
Publicado: (2025) -
HybridFlow: A Flexible and Efficient RLHF Framework
por: Sheng, Guangming, et al.
Publicado: (2024) -
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2026) -
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
por: Zhang, Yida, et al.
Publicado: (2026) -
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2025)