Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Paramanayakam, Varatheepan, Karatzas, Andreas, Stamoulis, Dimitrios, Anagnostopoulos, Iraklis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Less is More: Optimizing Function Calling for LLM Execution on Edge Devices
von: Paramanayakam, Varatheepan, et al.
Veröffentlicht: (2024)
von: Paramanayakam, Varatheepan, et al.
Veröffentlicht: (2024)
RankMap: Priority-Aware Multi-DNN Manager for Heterogeneous Embedded Devices
von: Karatzas, Andreas, et al.
Veröffentlicht: (2024)
von: Karatzas, Andreas, et al.
Veröffentlicht: (2024)
CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
von: Paramanayakam, Varatheepan, et al.
Veröffentlicht: (2025)
von: Paramanayakam, Varatheepan, et al.
Veröffentlicht: (2025)
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
LLM-dCache: Improving Tool-Augmented LLMs with GPT-Driven Localized Data Caching
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
Execution time budget assignment for mixed criticality systems
von: Khelassi, Mohamed Amine, et al.
Veröffentlicht: (2023)
von: Khelassi, Mohamed Amine, et al.
Veröffentlicht: (2023)
Node Compass: Multilevel Tracing and Debugging of Request Executions in JavaScript-Based Web-Servers
von: Kabamba, Herve Mbikayi, et al.
Veröffentlicht: (2023)
von: Kabamba, Herve Mbikayi, et al.
Veröffentlicht: (2023)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
von: Dutta, Akash, et al.
Veröffentlicht: (2024)
von: Dutta, Akash, et al.
Veröffentlicht: (2024)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
When Less is More: Achieving Faster Convergence in Distributed Edge Machine Learning
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
The Energy Cost of Execution-Idle in GPU Clusters
von: Lei, Yiran, et al.
Veröffentlicht: (2026)
von: Lei, Yiran, et al.
Veröffentlicht: (2026)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
cedar: Optimized and Unified Machine Learning Input Data Pipelines
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
von: Chu, Ruifan, et al.
Veröffentlicht: (2025)
von: Chu, Ruifan, et al.
Veröffentlicht: (2025)
Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge Platforms
von: Taufique, Zain, et al.
Veröffentlicht: (2023)
von: Taufique, Zain, et al.
Veröffentlicht: (2023)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
von: Lin, Zhongyi, et al.
Veröffentlicht: (2024)
von: Lin, Zhongyi, et al.
Veröffentlicht: (2024)
Ariel-ML: Computing Parallelization with Embedded Rust for Neural Networks on Heterogeneous Multi-core Microcontrollers
von: Huang, Zhaolan, et al.
Veröffentlicht: (2025)
von: Huang, Zhaolan, et al.
Veröffentlicht: (2025)
Multi-Dimensional Autoscaling of Stream Processing Services on Edge Devices
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
A Delta-Aware Orchestration Framework for Scalable Multi-Agent Edge Computing
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2026)
Adaptive DNN Partitioning and Offloading in Heterogeneous Edge-Cloud Continuum
von: Deng, Akuen Akoi, et al.
Veröffentlicht: (2026)
von: Deng, Akuen Akoi, et al.
Veröffentlicht: (2026)
Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
von: Dagli, Ismet, et al.
Veröffentlicht: (2023)
von: Dagli, Ismet, et al.
Veröffentlicht: (2023)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
von: Zarkadas, Ioannis, et al.
Veröffentlicht: (2025)
von: Zarkadas, Ioannis, et al.
Veröffentlicht: (2025)
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
von: Wu, Dan, et al.
Veröffentlicht: (2023)
von: Wu, Dan, et al.
Veröffentlicht: (2023)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
KVDirect: Distributed Disaggregated LLM Inference
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
GPU Cluster Scheduling for Network-Sensitive Deep Learning
von: Sharma, Aakash, et al.
Veröffentlicht: (2024)
von: Sharma, Aakash, et al.
Veröffentlicht: (2024)
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
von: Ghafouri, Saeid, et al.
Veröffentlicht: (2023)
von: Ghafouri, Saeid, et al.
Veröffentlicht: (2023)
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
von: Anik, Md Saidul Hoque, et al.
Veröffentlicht: (2024)
von: Anik, Md Saidul Hoque, et al.
Veröffentlicht: (2024)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
von: Gupta, Ahan, et al.
Veröffentlicht: (2026)
von: Gupta, Ahan, et al.
Veröffentlicht: (2026)
ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
Is Intelligence the Right Direction in New OS Scheduling for Multiple Resources in Cloud Environments?
von: Dou, Xinglei, et al.
Veröffentlicht: (2025)
von: Dou, Xinglei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Less is More: Optimizing Function Calling for LLM Execution on Edge Devices
von: Paramanayakam, Varatheepan, et al.
Veröffentlicht: (2024) -
RankMap: Priority-Aware Multi-DNN Manager for Heterogeneous Embedded Devices
von: Karatzas, Andreas, et al.
Veröffentlicht: (2024) -
CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
von: Paramanayakam, Varatheepan, et al.
Veröffentlicht: (2025) -
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026) -
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)