Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ng, Nathan, Hanafy, Walid A., Kadambi, Prashanthi, Sunil, Balachandra, Gupta, Ayush, Irwin, David, Simmhan, Yogesh, Shenoy, Prashant |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
von: Wu, Li, et al.
Veröffentlicht: (2025)
von: Wu, Li, et al.
Veröffentlicht: (2025)
CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters
von: Hanafy, Walid A., et al.
Veröffentlicht: (2025)
von: Hanafy, Walid A., et al.
Veröffentlicht: (2025)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
Characterizing the Performance of Accelerated Jetson Edge Devices for Training Deep Learning Models
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
CarbonEdge: Leveraging Mesoscale Spatial Carbon-Intensity Variations for Low Carbon Edge Computing
von: Wu, Li, et al.
Veröffentlicht: (2025)
von: Wu, Li, et al.
Veröffentlicht: (2025)
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
von: Ng, Nathan, et al.
Veröffentlicht: (2025)
von: Ng, Nathan, et al.
Veröffentlicht: (2025)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2025)
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2025)
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
von: Paramanayakam, Varatheepan, et al.
Veröffentlicht: (2025)
von: Paramanayakam, Varatheepan, et al.
Veröffentlicht: (2025)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
von: Raj, Suman, et al.
Veröffentlicht: (2024)
von: Raj, Suman, et al.
Veröffentlicht: (2024)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
Quantifying the Carbon Reduction of DAG Workloads: A Job Shop Scheduling Perspective
von: Bostandoost, Roozbeh, et al.
Veröffentlicht: (2025)
von: Bostandoost, Roozbeh, et al.
Veröffentlicht: (2025)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
Ripple: Scalable Incremental GNN Inferencing on Large Streaming Graphs
von: Naman, Pranjal, et al.
Veröffentlicht: (2025)
von: Naman, Pranjal, et al.
Veröffentlicht: (2025)
ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural Networks
von: Naman, Pranjal, et al.
Veröffentlicht: (2026)
von: Naman, Pranjal, et al.
Veröffentlicht: (2026)
Pagoda: An Energy and Time Roofline Study for DNN Workloads on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
Bridding OT and PaaS in Edge-to-Cloud Continuum
von: Barrios, Carlos J, et al.
Veröffentlicht: (2025)
von: Barrios, Carlos J, et al.
Veröffentlicht: (2025)
Scalable GPU Performance Variability Analysis framework
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
Active Inference-Based Adaptive Routing for Heterogeneous Edge AI Services
von: Wang, Zihang, et al.
Veröffentlicht: (2026)
von: Wang, Zihang, et al.
Veröffentlicht: (2026)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
CASPER: Carbon-Aware Scheduling and Provisioning for Distributed Web Services
von: Souza, Abel, et al.
Veröffentlicht: (2024)
von: Souza, Abel, et al.
Veröffentlicht: (2024)
PowerTrain: Fast, Generalizable Time and Power Prediction Models to Optimize DNN Training on Accelerated Edges
von: K., Prashanthi S., et al.
Veröffentlicht: (2024)
von: K., Prashanthi S., et al.
Veröffentlicht: (2024)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
von: Proaño, Andrès Rubio, et al.
Veröffentlicht: (2024)
von: Proaño, Andrès Rubio, et al.
Veröffentlicht: (2024)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
von: Wu, Li, et al.
Veröffentlicht: (2025) -
CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters
von: Hanafy, Walid A., et al.
Veröffentlicht: (2025) -
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2023) -
Characterizing the Performance of Accelerated Jetson Edge Devices for Training Deep Learning Models
von: K., Prashanthi S., et al.
Veröffentlicht: (2025) -
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
von: Arya, Mayank, et al.
Veröffentlicht: (2025)