TailBench++: Flexible Multi-Client, Multi-Server Benchmarking for Latency-Critical Workloads
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhilin, Pons, Lucia, Petit, Salvador, Sahuquillo, Julio, Pons, Julio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prefetching in Deep Memory Hierarchies with NVRAM as Main Memory
von: Lurbe, Manel, et al.
Veröffentlicht: (2025)
von: Lurbe, Manel, et al.
Veröffentlicht: (2025)
A New Family of Thread to Core Allocation Policies for an SMT ARM Processor
von: Navarro, Marta, et al.
Veröffentlicht: (2025)
von: Navarro, Marta, et al.
Veröffentlicht: (2025)
Optimal Workload Placement on Multi-Instance GPUs
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
von: Hidayetoglu, Mert, et al.
Veröffentlicht: (2025)
von: Hidayetoglu, Mert, et al.
Veröffentlicht: (2025)
Dynamic Client Clustering, Bandwidth Allocation, and Workload Optimization for Semi-synchronous Federated Learning
von: Yu, Liangkun, et al.
Veröffentlicht: (2024)
von: Yu, Liangkun, et al.
Veröffentlicht: (2024)
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
von: Seo, Eunil, et al.
Veröffentlicht: (2025)
von: Seo, Eunil, et al.
Veröffentlicht: (2025)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
Workload Distribution with Rateless Encoding: A Low-Latency Computation Offloading Method within Edge Networks
von: Guo, Zhongfu, et al.
Veröffentlicht: (2023)
von: Guo, Zhongfu, et al.
Veröffentlicht: (2023)
Areon: Latency-Friendly and Resilient Multi-Proposer Consensus
von: Castro-Castilla, Álvaro, et al.
Veröffentlicht: (2025)
von: Castro-Castilla, Álvaro, et al.
Veröffentlicht: (2025)
Orchestrating Mixed-Criticality Cloud Workloads in Reconfigurable Manufacturing Systems
von: Barletta, Marco, et al.
Veröffentlicht: (2024)
von: Barletta, Marco, et al.
Veröffentlicht: (2024)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
von: Ting, Hsu-Tzu, et al.
Veröffentlicht: (2025)
von: Ting, Hsu-Tzu, et al.
Veröffentlicht: (2025)
Asynchronous Multi-Server Federated Learning for Geo-Distributed Clients
von: Zuo, Yuncong, et al.
Veröffentlicht: (2024)
von: Zuo, Yuncong, et al.
Veröffentlicht: (2024)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
Understanding Server-Assisted Federated Learning in the Presence of Incomplete Client Participation
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
SWARM+: Scalable and Resilient Multi-Agent Consensus for Fully-Decentralized Data-Aware Workload Management
von: Thareja, Komal, et al.
Veröffentlicht: (2026)
von: Thareja, Komal, et al.
Veröffentlicht: (2026)
Scene-Aware Latency Estimation for Microservices via Multi-Scale Graph Fusion
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
Design and Implementation of a Java-Based Client-Server Application
von: Patil, Omkar, et al.
Veröffentlicht: (2024)
von: Patil, Omkar, et al.
Veröffentlicht: (2024)
Are Bus-Mounted Edge Servers Feasible?
von: Li, Xuezhi, et al.
Veröffentlicht: (2025)
von: Li, Xuezhi, et al.
Veröffentlicht: (2025)
CycleSL: Server-Client Cyclical Update Driven Scalable Split Learning
von: Wang, Mengdi, et al.
Veröffentlicht: (2025)
von: Wang, Mengdi, et al.
Veröffentlicht: (2025)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
FaaS Is Not Enough: Serverless Handling of Burst-Parallel Jobs
von: Barcelona-Pons, Daniel, et al.
Veröffentlicht: (2024)
von: Barcelona-Pons, Daniel, et al.
Veröffentlicht: (2024)
Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative Caching
von: Liang, Wenyi, et al.
Veröffentlicht: (2024)
von: Liang, Wenyi, et al.
Veröffentlicht: (2024)
Agentic AI Workload Characteristics
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
Practical Federated Learning without a Server
von: Dhasade, Akash, et al.
Veröffentlicht: (2025)
von: Dhasade, Akash, et al.
Veröffentlicht: (2025)
Accelerating Compound LLM Training Workloads with Maestro
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
ElastiBench: Scalable Continuous Benchmarking on Cloud FaaS Platforms
von: Schirmer, Trever, et al.
Veröffentlicht: (2024)
von: Schirmer, Trever, et al.
Veröffentlicht: (2024)
Experimental Analysis of Server-Side Caching for Web Performance
von: Umar, Mohammad, et al.
Veröffentlicht: (2026)
von: Umar, Mohammad, et al.
Veröffentlicht: (2026)
SCARIF: Towards Carbon Modeling of Cloud Servers with Accelerators
von: Ji, Shixin, et al.
Veröffentlicht: (2024)
von: Ji, Shixin, et al.
Veröffentlicht: (2024)
Delay-Aware Multi-Stage Edge Server Upgrade with Budget Constraint
von: Wihidayat, Endar Suprih, et al.
Veröffentlicht: (2025)
von: Wihidayat, Endar Suprih, et al.
Veröffentlicht: (2025)
SECO: Secure Inference With Model Splitting Across Multi-Server Hierarchy
von: Chen, Shuangyi, et al.
Veröffentlicht: (2024)
von: Chen, Shuangyi, et al.
Veröffentlicht: (2024)
Ocularone-Bench: Benchmarking DNN Models on GPUs to Assist the Visually Impaired
von: Raj, Suman, et al.
Veröffentlicht: (2025)
von: Raj, Suman, et al.
Veröffentlicht: (2025)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
von: Stoyanov, Radostin, et al.
Veröffentlicht: (2025)
von: Stoyanov, Radostin, et al.
Veröffentlicht: (2025)
AI Surrogate Model for Distributed Computing Workloads
von: Park, David K., et al.
Veröffentlicht: (2024)
von: Park, David K., et al.
Veröffentlicht: (2024)
Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
von: Wu, Yebo, et al.
Veröffentlicht: (2024)
von: Wu, Yebo, et al.
Veröffentlicht: (2024)
Risk-Aware and Stable Edge Server Selection Under Network Latency SLOs
von: Liyanage, Mohan, et al.
Veröffentlicht: (2026)
von: Liyanage, Mohan, et al.
Veröffentlicht: (2026)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
Analysis of Server Throughput For Managed Big Data Analytics Frameworks
von: Anagnostakis, Emmanouil, et al.
Veröffentlicht: (2025)
von: Anagnostakis, Emmanouil, et al.
Veröffentlicht: (2025)
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
von: Liu, Guangyi, et al.
Veröffentlicht: (2026)
von: Liu, Guangyi, et al.
Veröffentlicht: (2026)
Crossword: Adaptive Consensus for Dynamic Data-Heavy Workloads
von: Hu, Guanzhou, et al.
Veröffentlicht: (2025)
von: Hu, Guanzhou, et al.
Veröffentlicht: (2025)
Data Management System Analysis for Distributed Computing Workloads
von: Hsu, Kuan-Chieh, et al.
Veröffentlicht: (2025)
von: Hsu, Kuan-Chieh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Prefetching in Deep Memory Hierarchies with NVRAM as Main Memory
von: Lurbe, Manel, et al.
Veröffentlicht: (2025) -
A New Family of Thread to Core Allocation Policies for an SMT ARM Processor
von: Navarro, Marta, et al.
Veröffentlicht: (2025) -
Optimal Workload Placement on Multi-Instance GPUs
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024) -
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
von: Hidayetoglu, Mert, et al.
Veröffentlicht: (2025) -
Dynamic Client Clustering, Bandwidth Allocation, and Workload Optimization for Semi-synchronous Federated Learning
von: Yu, Liangkun, et al.
Veröffentlicht: (2024)