FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Runhua, Jiang, Hongxu, Geng, Jinkun, Ma, Yuhang, Zhu, Chenhui, Wang, Haojie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
di: Wu, Tian, et al.
Pubblicazione: (2025)
di: Wu, Tian, et al.
Pubblicazione: (2025)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
di: Tayal, Mumuksh, et al.
Pubblicazione: (2025)
di: Tayal, Mumuksh, et al.
Pubblicazione: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
di: Sun, Mingyu, et al.
Pubblicazione: (2025)
di: Sun, Mingyu, et al.
Pubblicazione: (2025)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
Accelerating Distributed MoE Training and Inference with Lina
di: Li, Jiamin, et al.
Pubblicazione: (2022)
di: Li, Jiamin, et al.
Pubblicazione: (2022)
Tiga: Accelerating Geo-Distributed Transactions with Synchronized Clocks [Technical Report]
di: Geng, Jinkun, et al.
Pubblicazione: (2025)
di: Geng, Jinkun, et al.
Pubblicazione: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
di: Arya, Mayank, et al.
Pubblicazione: (2025)
di: Arya, Mayank, et al.
Pubblicazione: (2025)
Optimizing Distributed Protocols with Query Rewrites [Technical Report]
di: Chu, David, et al.
Pubblicazione: (2024)
di: Chu, David, et al.
Pubblicazione: (2024)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
di: Tran, Phuong, et al.
Pubblicazione: (2025)
di: Tran, Phuong, et al.
Pubblicazione: (2025)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
di: K., Prashanthi S., et al.
Pubblicazione: (2023)
di: K., Prashanthi S., et al.
Pubblicazione: (2023)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
di: Chow, Will
Pubblicazione: (2025)
di: Chow, Will
Pubblicazione: (2025)
Pie: Pooling CPU Memory for LLM Inference
di: Xu, Yi, et al.
Pubblicazione: (2024)
di: Xu, Yi, et al.
Pubblicazione: (2024)
FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism
di: Wang, Yujie, et al.
Pubblicazione: (2024)
di: Wang, Yujie, et al.
Pubblicazione: (2024)
Characterizing the Performance of Accelerated Jetson Edge Devices for Training Deep Learning Models
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
Distributed On-Device LLM Inference With Over-the-Air Computation
di: Zhang, Kai, et al.
Pubblicazione: (2025)
di: Zhang, Kai, et al.
Pubblicazione: (2025)
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
di: Behera, Adarsh Prasad, et al.
Pubblicazione: (2024)
di: Behera, Adarsh Prasad, et al.
Pubblicazione: (2024)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
di: Wang, Zhibin, et al.
Pubblicazione: (2025)
di: Wang, Zhibin, et al.
Pubblicazione: (2025)
FlexFL: Heterogeneous Federated Learning via APoZ-Guided Flexible Pruning in Uncertain Scenarios
di: Chen, Zekai, et al.
Pubblicazione: (2024)
di: Chen, Zekai, et al.
Pubblicazione: (2024)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
di: Li, Yuchen, et al.
Pubblicazione: (2026)
di: Li, Yuchen, et al.
Pubblicazione: (2026)
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
di: Wang, Li, et al.
Pubblicazione: (2024)
di: Wang, Li, et al.
Pubblicazione: (2024)
Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Optimization
di: Gao, Luyao, et al.
Pubblicazione: (2024)
di: Gao, Luyao, et al.
Pubblicazione: (2024)
A Preliminary Study on Accelerating Simulation Optimization with GPU Implementation
di: He, Jinghai, et al.
Pubblicazione: (2024)
di: He, Jinghai, et al.
Pubblicazione: (2024)
Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital Twin-Assisted Approach
di: Hu, Shisheng, et al.
Pubblicazione: (2024)
di: Hu, Shisheng, et al.
Pubblicazione: (2024)
Barycentric Coded Distributed Computing with Flexible Recovery Threshold for Collaborative Mobile Edge Computing
di: Qiu, Houming, et al.
Pubblicazione: (2025)
di: Qiu, Houming, et al.
Pubblicazione: (2025)
Accelerating Local LLMs on Resource-Constrained Edge Devices via Distributed Prompt Caching
di: Matsutani, Hiroki, et al.
Pubblicazione: (2026)
di: Matsutani, Hiroki, et al.
Pubblicazione: (2026)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
di: Li, Rui, et al.
Pubblicazione: (2024)
di: Li, Rui, et al.
Pubblicazione: (2024)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
di: Zhu, Juan, et al.
Pubblicazione: (2026)
di: Zhu, Juan, et al.
Pubblicazione: (2026)
Distributed Inference Performance Optimization for LLMs on CPUs
di: He, Pujiang, et al.
Pubblicazione: (2024)
di: He, Pujiang, et al.
Pubblicazione: (2024)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
di: Xia, Mengchun, et al.
Pubblicazione: (2026)
di: Xia, Mengchun, et al.
Pubblicazione: (2026)
GenAI at the Edge: Comprehensive Survey on Empowering Edge Devices
di: Navardi, Mozhgan, et al.
Pubblicazione: (2025)
di: Navardi, Mozhgan, et al.
Pubblicazione: (2025)
Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices
di: Shen, Tao, et al.
Pubblicazione: (2025)
di: Shen, Tao, et al.
Pubblicazione: (2025)
Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative Caching
di: Liang, Wenyi, et al.
Pubblicazione: (2024)
di: Liang, Wenyi, et al.
Pubblicazione: (2024)
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
di: Wang, Tinglue, et al.
Pubblicazione: (2025)
di: Wang, Tinglue, et al.
Pubblicazione: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
di: Wu, Panlong, et al.
Pubblicazione: (2025)
di: Wu, Panlong, et al.
Pubblicazione: (2025)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
di: Wang, Liujianfu, et al.
Pubblicazione: (2025)
di: Wang, Liujianfu, et al.
Pubblicazione: (2025)
Infer-EDGE: Dynamic DNN Inference Optimization in 'Just-in-time' Edge-AI Implementations
di: Mounesan, Motahare, et al.
Pubblicazione: (2025)
di: Mounesan, Motahare, et al.
Pubblicazione: (2025)
SkyMemory: A LEO Edge Cache for Transformer Inference Optimization and Scale Out
di: Sandholm, Thomas, et al.
Pubblicazione: (2025)
di: Sandholm, Thomas, et al.
Pubblicazione: (2025)
Token Level Routing Inference System for Edge Devices
di: She, Jianshu, et al.
Pubblicazione: (2025)
di: She, Jianshu, et al.
Pubblicazione: (2025)
INTELLECT-1 Technical Report
di: Jaghouar, Sami, et al.
Pubblicazione: (2024)
di: Jaghouar, Sami, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
di: Wu, Tian, et al.
Pubblicazione: (2025) -
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
di: Tayal, Mumuksh, et al.
Pubblicazione: (2025) -
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
di: Sun, Mingyu, et al.
Pubblicazione: (2025) -
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
di: K., Prashanthi S., et al.
Pubblicazione: (2025) -
Accelerating Distributed MoE Training and Inference with Lina
di: Li, Jiamin, et al.
Pubblicazione: (2022)