Agentic AI Workload Characteristics
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Yichao, Nayak, Ankita, Kundu, Souvik, Talati, Nishil |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion Models
by: Xia, Yuchen, et al.
Published: (2025)
by: Xia, Yuchen, et al.
Published: (2025)
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
by: Yuan, Yichao, et al.
Published: (2026)
by: Yuan, Yichao, et al.
Published: (2026)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
by: Yuan, Yichao, et al.
Published: (2025)
by: Yuan, Yichao, et al.
Published: (2025)
Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data Analytics
by: Yuan, Yichao, et al.
Published: (2025)
by: Yuan, Yichao, et al.
Published: (2025)
BlazingAML: High-Throughput Anti-Money Laundering (AML) via Multi-Stage Graph Mining
by: Ye, Haojie, et al.
Published: (2026)
by: Ye, Haojie, et al.
Published: (2026)
ZKProphet: Understanding Performance of Zero-Knowledge Proofs on GPUs
by: Verma, Tarunesh, et al.
Published: (2025)
by: Verma, Tarunesh, et al.
Published: (2025)
Toward Cross-Layer Energy Optimizations in AI Systems
by: Chung, Jae-Won, et al.
Published: (2024)
by: Chung, Jae-Won, et al.
Published: (2024)
Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining
by: Singapuram, Sanjay Sri Vallabh, et al.
Published: (2025)
by: Singapuram, Sanjay Sri Vallabh, et al.
Published: (2025)
Eventually-Consistent Federated Scheduling for Data Center Workloads
by: Thiyyakat, Meghana, et al.
Published: (2023)
by: Thiyyakat, Meghana, et al.
Published: (2023)
Profiling and Modeling of Power Characteristics of Leadership-Scale HPC System Workloads
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
AI Surrogate Model for Distributed Computing Workloads
by: Park, David K., et al.
Published: (2024)
by: Park, David K., et al.
Published: (2024)
A Deep Dive into the Google Cluster Workload Traces: Analyzing the Application Failure Characteristics and User Behaviors
by: Bappy, Faisal Haque, et al.
Published: (2023)
by: Bappy, Faisal Haque, et al.
Published: (2023)
Accelerating Compound LLM Training Workloads with Maestro
by: Yuan, Xiulong, et al.
Published: (2026)
by: Yuan, Xiulong, et al.
Published: (2026)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026)
by: Jain, Rutwik, et al.
Published: (2026)
NMP-PaK: Near-Memory Processing Acceleration of Scalable De Novo Genome Assembly
by: Kim, Heewoo, et al.
Published: (2025)
by: Kim, Heewoo, et al.
Published: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
by: Zhang, Yuning, et al.
Published: (2026)
by: Zhang, Yuning, et al.
Published: (2026)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
by: Stoyanov, Radostin, et al.
Published: (2025)
by: Stoyanov, Radostin, et al.
Published: (2025)
Optimal Workload Placement on Multi-Instance GPUs
by: Turkkan, Bekir, et al.
Published: (2024)
by: Turkkan, Bekir, et al.
Published: (2024)
Union: An Automatic Workload Manager for Accelerating Network Simulation
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Crossword: Adaptive Consensus for Dynamic Data-Heavy Workloads
by: Hu, Guanzhou, et al.
Published: (2025)
by: Hu, Guanzhou, et al.
Published: (2025)
Distributed Load Balancing with Workload-Dependent Service Rates
by: Zhang, Wenxin, et al.
Published: (2024)
by: Zhang, Wenxin, et al.
Published: (2024)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
Data Management System Analysis for Distributed Computing Workloads
by: Hsu, Kuan-Chieh, et al.
Published: (2025)
by: Hsu, Kuan-Chieh, et al.
Published: (2025)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
by: Huang, Lexiang, et al.
Published: (2024)
by: Huang, Lexiang, et al.
Published: (2024)
Towards Cloud Efficiency with Large-scale Workload Characterization
by: Parayil, Anjaly, et al.
Published: (2024)
by: Parayil, Anjaly, et al.
Published: (2024)
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
by: Wei, Xinming, et al.
Published: (2025)
by: Wei, Xinming, et al.
Published: (2025)
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
Evaluating Container Orchestration for Neuromorphic Workloads in Virtual Edge Environments
by: Pham, Huyen, et al.
Published: (2026)
by: Pham, Huyen, et al.
Published: (2026)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
by: Ye, Fanjiang, et al.
Published: (2026)
by: Ye, Fanjiang, et al.
Published: (2026)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
by: Zhu, Botao, et al.
Published: (2025)
by: Zhu, Botao, et al.
Published: (2025)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
by: Nguyen, Thanh-Tung, et al.
Published: (2025)
by: Nguyen, Thanh-Tung, et al.
Published: (2025)
Orchestrating Mixed-Criticality Cloud Workloads in Reconfigurable Manufacturing Systems
by: Barletta, Marco, et al.
Published: (2024)
by: Barletta, Marco, et al.
Published: (2024)
GreenFaaS: Maximizing Energy Efficiency of HPC Workloads with FaaS
by: Kamatar, Alok, et al.
Published: (2024)
by: Kamatar, Alok, et al.
Published: (2024)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
by: Chazapis, Antony, et al.
Published: (2024)
by: Chazapis, Antony, et al.
Published: (2024)
Knowledge-driven Reasoning for Mobile Agentic AI: Concepts, Approaches, and Directions
by: Liu, Guangyuan, et al.
Published: (2026)
by: Liu, Guangyuan, et al.
Published: (2026)
Quantifying Energy and Cost Benefits of Hybrid Edge Cloud: Analysis of Traditional and Agentic Workloads
by: Alamouti, Siavash
Published: (2025)
by: Alamouti, Siavash
Published: (2025)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
MERBIT: A GPU-Based SpMV Method for Iterative Workloads
by: Zhang, Qi, et al.
Published: (2026)
by: Zhang, Qi, et al.
Published: (2026)
Similar Items
-
MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion Models
by: Xia, Yuchen, et al.
Published: (2025) -
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
by: Yuan, Yichao, et al.
Published: (2026) -
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
by: Yuan, Yichao, et al.
Published: (2025) -
Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data Analytics
by: Yuan, Yichao, et al.
Published: (2025) -
BlazingAML: High-Throughput Anti-Money Laundering (AML) via Multi-Stage Graph Mining
by: Ye, Haojie, et al.
Published: (2026)