ECCENTRIC: Edge-Cloud Collaboration Framework for Distributed Inference Using Knowledge Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kamani, Mohammad Mahdi, Cheng, Zhongwei, Chen, Lin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Distributed Inference on Mobile Edge and Cloud: A Data-Cartography based Clustering Approach
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
DGRAG: Distributed Graph-based Retrieval-Augmented Generation in Edge-Cloud Systems
von: Zhou, Wenqing, et al.
Veröffentlicht: (2025)
von: Zhou, Wenqing, et al.
Veröffentlicht: (2025)
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
Edge-Cloud Collaborative Computing on Distributed Intelligence and Model Optimization: A Survey
von: Liu, Jing, et al.
Veröffentlicht: (2025)
von: Liu, Jing, et al.
Veröffentlicht: (2025)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
von: Deng, Xiumei, et al.
Veröffentlicht: (2025)
von: Deng, Xiumei, et al.
Veröffentlicht: (2025)
Distributed Inference on Mobile Edge and Cloud: An Early Exit based Clustering Approach
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
von: Wang, Li, et al.
Veröffentlicht: (2024)
von: Wang, Li, et al.
Veröffentlicht: (2024)
Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI
von: Jin, Jianli, et al.
Veröffentlicht: (2025)
von: Jin, Jianli, et al.
Veröffentlicht: (2025)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
Adaptive AI-based Decentralized Resource Management in the Cloud-Edge Continuum
von: Li, Lanpei, et al.
Veröffentlicht: (2025)
von: Li, Lanpei, et al.
Veröffentlicht: (2025)
AI Inference as Relocatable Electricity Demand: A Latency-Constrained Energy-Geography Framework
von: Luo, Xubin, et al.
Veröffentlicht: (2026)
von: Luo, Xubin, et al.
Veröffentlicht: (2026)
Large Language Model Partitioning for Low-Latency Inference at the Edge
von: Kafetzis, Dimitrios, et al.
Veröffentlicht: (2025)
von: Kafetzis, Dimitrios, et al.
Veröffentlicht: (2025)
Reinforcement Learning-driven Data-intensive Workflow Scheduling for Volunteer Edge-Cloud
von: Mounesan, Motahare, et al.
Veröffentlicht: (2024)
von: Mounesan, Motahare, et al.
Veröffentlicht: (2024)
An AI-Driven Framework for Energy-Efficient Environmental Monitoring in Smart Cities Using Edge Intelligence
von: Liu, Yichen, et al.
Veröffentlicht: (2026)
von: Liu, Yichen, et al.
Veröffentlicht: (2026)
Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
von: Wang, Zhanghan, et al.
Veröffentlicht: (2025)
von: Wang, Zhanghan, et al.
Veröffentlicht: (2025)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
Quantifying Energy and Cost Benefits of Hybrid Edge Cloud: Analysis of Traditional and Agentic Workloads
von: Alamouti, Siavash
Veröffentlicht: (2025)
von: Alamouti, Siavash
Veröffentlicht: (2025)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
Hardware Utilization and Inference Performance of Edge Object Detection Under Fault Injection
von: Pasandideh, Faezeh, et al.
Veröffentlicht: (2026)
von: Pasandideh, Faezeh, et al.
Veröffentlicht: (2026)
Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
von: Zhang, Zongshun, et al.
Veröffentlicht: (2025)
von: Zhang, Zongshun, et al.
Veröffentlicht: (2025)
EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model
von: Pinnock, Alyssa, et al.
Veröffentlicht: (2025)
von: Pinnock, Alyssa, et al.
Veröffentlicht: (2025)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
CloudEval-YAML: A Practical Benchmark for Cloud Configuration Generation
von: Xu, Yifei, et al.
Veröffentlicht: (2023)
von: Xu, Yifei, et al.
Veröffentlicht: (2023)
Backpropagation-Free Multi-modal On-Device Model Adaptation via Cloud-Device Collaboration
von: Ji, Wei, et al.
Veröffentlicht: (2024)
von: Ji, Wei, et al.
Veröffentlicht: (2024)
KAITIAN: A Unified Communication Framework for Enabling Efficient Collaboration Across Heterogeneous Accelerators in Embodied AI Systems
von: Lin, Jieke, et al.
Veröffentlicht: (2025)
von: Lin, Jieke, et al.
Veröffentlicht: (2025)
Research on Edge Computing and Cloud Collaborative Resource Scheduling Optimization Based on Deep Reinforcement Learning
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Intelligent Autonomous Orchestration for Distributed Cloud Resources using Complex-Stability Analysis
von: Shyam, Gopal Krishna, et al.
Veröffentlicht: (2026)
von: Shyam, Gopal Krishna, et al.
Veröffentlicht: (2026)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
von: Li, Wanqian, et al.
Veröffentlicht: (2026)
von: Li, Wanqian, et al.
Veröffentlicht: (2026)
Learning Provably Correct Distributed Protocols Without Human Knowledge
von: Hui, Yujie, et al.
Veröffentlicht: (2026)
von: Hui, Yujie, et al.
Veröffentlicht: (2026)
Leveraging Large Language Model for Intelligent Log Processing and Autonomous Debugging in Cloud AI Platforms
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
von: Taufique, Zain, et al.
Veröffentlicht: (2024)
von: Taufique, Zain, et al.
Veröffentlicht: (2024)
Cloud-Based AI Systems: Leveraging Large Language Models for Intelligent Fault Detection and Autonomous Self-Healing
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
von: Yu, Shibo, et al.
Veröffentlicht: (2025)
von: Yu, Shibo, et al.
Veröffentlicht: (2025)
Optimized Cloud Resource Allocation Using Genetic Algorithms for Energy Efficiency and QoS Assurance
von: Panggabean, Caroline, et al.
Veröffentlicht: (2025)
von: Panggabean, Caroline, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Distributed Inference on Mobile Edge and Cloud: A Data-Cartography based Clustering Approach
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024) -
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026) -
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026) -
DGRAG: Distributed Graph-based Retrieval-Augmented Generation in Edge-Cloud Systems
von: Zhou, Wenqing, et al.
Veröffentlicht: (2025) -
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)