Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Koch, Fernando, Djuhera, Aladin, Binotto, Alecio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Joint Partitioning and Placement of Foundation Models for Real-Time Edge AI
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
Priority-Aware Model-Distributed Inference at Edge Networks
von: Li, Teng, et al.
Veröffentlicht: (2024)
von: Li, Teng, et al.
Veröffentlicht: (2024)
A Survey on Collaborative DNN Inference for Edge Intelligence
von: Ren, Weiqing, et al.
Veröffentlicht: (2022)
von: Ren, Weiqing, et al.
Veröffentlicht: (2022)
Designing Large Foundation Models for Efficient Training and Inference: A Survey
von: Liu, Dong, et al.
Veröffentlicht: (2024)
von: Liu, Dong, et al.
Veröffentlicht: (2024)
Distributed Load Orchestration for Vision Computing in Multi-Access Edge Computing
von: Boing, Ricardo N., et al.
Veröffentlicht: (2022)
von: Boing, Ricardo N., et al.
Veröffentlicht: (2022)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
Towards Integrated Fine-tuning and Inference when Generative AI meets Edge Intelligence
von: Chen, Ning, et al.
Veröffentlicht: (2024)
von: Chen, Ning, et al.
Veröffentlicht: (2024)
On Harnessing Idle Compute at the Edge for Foundation Model Training
von: Xue, Leyang, et al.
Veröffentlicht: (2025)
von: Xue, Leyang, et al.
Veröffentlicht: (2025)
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
Leveraging Foundation Models for Efficient Federated Learning in Resource-restricted Edge Networks
von: Atapour, S. Kawa, et al.
Veröffentlicht: (2024)
von: Atapour, S. Kawa, et al.
Veröffentlicht: (2024)
DisCEdge: Distributed Context Management for Large Language Models at the Edge
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
Decentralized Orchestration Architecture for Fluid Computing: A Secure Distributed AI Use Case
von: Cajaraville-Aboy, Diego, et al.
Veröffentlicht: (2026)
von: Cajaraville-Aboy, Diego, et al.
Veröffentlicht: (2026)
Adaptive Stream Processing on Edge Devices through Active Inference
von: Sedlak, Boris, et al.
Veröffentlicht: (2024)
von: Sedlak, Boris, et al.
Veröffentlicht: (2024)
Learning the Optimal Path and DNN Partition for Collaborative Edge Inference
von: Huang, Yin, et al.
Veröffentlicht: (2024)
von: Huang, Yin, et al.
Veröffentlicht: (2024)
AI-in-the-Loop Sensing and Communication Joint Design for Edge Intelligence
von: Cai, Zhijie, et al.
Veröffentlicht: (2025)
von: Cai, Zhijie, et al.
Veröffentlicht: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
Edge-Cloud Collaborative Computing on Distributed Intelligence and Model Optimization: A Survey
von: Liu, Jing, et al.
Veröffentlicht: (2025)
von: Liu, Jing, et al.
Veröffentlicht: (2025)
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
von: Han, Xueyuan, et al.
Veröffentlicht: (2024)
von: Han, Xueyuan, et al.
Veröffentlicht: (2024)
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
Distributed Convolutional Neural Network Training on Mobile and Edge Clusters
von: Rama, Pranav, et al.
Veröffentlicht: (2024)
von: Rama, Pranav, et al.
Veröffentlicht: (2024)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
von: Fu, Yao, et al.
Veröffentlicht: (2024)
von: Fu, Yao, et al.
Veröffentlicht: (2024)
MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines
von: Gao, Lei, et al.
Veröffentlicht: (2024)
von: Gao, Lei, et al.
Veröffentlicht: (2024)
CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration
von: Jin, Hongpeng, et al.
Veröffentlicht: (2024)
von: Jin, Hongpeng, et al.
Veröffentlicht: (2024)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
von: Ghosh, Himel
Veröffentlicht: (2024)
von: Ghosh, Himel
Veröffentlicht: (2024)
MultiTASC++: A Continuously Adaptive Scheduler for Edge-Based Multi-Device Cascade Inference
von: Nikolaidis, Sokratis, et al.
Veröffentlicht: (2024)
von: Nikolaidis, Sokratis, et al.
Veröffentlicht: (2024)
AIGC-assisted Federated Learning for Edge Intelligence: Architecture Design, Research Challenges and Future Directions
von: Qiang, Xianke, et al.
Veröffentlicht: (2025)
von: Qiang, Xianke, et al.
Veröffentlicht: (2025)
GOGH: Correlation-Guided Orchestration of GPUs in Heterogeneous Clusters
von: Raeisi, Ahmad, et al.
Veröffentlicht: (2025)
von: Raeisi, Ahmad, et al.
Veröffentlicht: (2025)
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
von: Hagemann, Johannes, et al.
Veröffentlicht: (2023)
von: Hagemann, Johannes, et al.
Veröffentlicht: (2023)
Deal: Distributed End-to-End GNN Inference for All Nodes
von: Chen, Shiyang, et al.
Veröffentlicht: (2025)
von: Chen, Shiyang, et al.
Veröffentlicht: (2025)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
von: Tummalapalli, Pranay, et al.
Veröffentlicht: (2026)
von: Tummalapalli, Pranay, et al.
Veröffentlicht: (2026)
iDDS: Intelligent Distributed Dispatch and Scheduling for Workflow Orchestration
von: Guan, Wen, et al.
Veröffentlicht: (2025)
von: Guan, Wen, et al.
Veröffentlicht: (2025)
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
von: Lu, Yishun, et al.
Veröffentlicht: (2026)
von: Lu, Yishun, et al.
Veröffentlicht: (2026)
ADAPT: A Self-Calibrating Proactive Autoscaler for Container Orchestration
von: Baghel, Himanshu Singh
Veröffentlicht: (2026)
von: Baghel, Himanshu Singh
Veröffentlicht: (2026)
Minder: Faulty Machine Detection for Large-scale Distributed Model Training
von: Deng, Yangtao, et al.
Veröffentlicht: (2024)
von: Deng, Yangtao, et al.
Veröffentlicht: (2024)
Accelerating Local LLMs on Resource-Constrained Edge Devices via Distributed Prompt Caching
von: Matsutani, Hiroki, et al.
Veröffentlicht: (2026)
von: Matsutani, Hiroki, et al.
Veröffentlicht: (2026)
BatchWeave: A Consistent Object-Store-Native Data Plane for Large Foundation Model Training
von: Sun, Ting, et al.
Veröffentlicht: (2026)
von: Sun, Ting, et al.
Veröffentlicht: (2026)
InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
von: Lee, Wonbeom, et al.
Veröffentlicht: (2024)
von: Lee, Wonbeom, et al.
Veröffentlicht: (2024)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025)
von: Gond, Raja, et al.
Veröffentlicht: (2025)
EarthSight: A Distributed Framework for Low-Latency Satellite Intelligence
von: Erol, Ansel Kaplan, et al.
Veröffentlicht: (2025)
von: Erol, Ansel Kaplan, et al.
Veröffentlicht: (2025)
Incentivised Orchestrated Training Architecture (IOTA): A Technical Primer for Release
von: Quinque, Felix, et al.
Veröffentlicht: (2025)
von: Quinque, Felix, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Joint Partitioning and Placement of Foundation Models for Real-Time Edge AI
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025) -
Priority-Aware Model-Distributed Inference at Edge Networks
von: Li, Teng, et al.
Veröffentlicht: (2024) -
A Survey on Collaborative DNN Inference for Edge Intelligence
von: Ren, Weiqing, et al.
Veröffentlicht: (2022) -
Designing Large Foundation Models for Efficient Training and Inference: A Survey
von: Liu, Dong, et al.
Veröffentlicht: (2024) -
Distributed Load Orchestration for Vision Computing in Multi-Access Edge Computing
von: Boing, Ricardo N., et al.
Veröffentlicht: (2022)