Tetris: Efficient Intra-Datacenter Calls Packing for Large Conferencing Services
Fuente:
arXiv
Saved in:
| Main Authors: | Gandhi, Rohan, Mallick, Ankur, Sueda, Ken, Liang, Rui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
by: Gandhi, Rohan, et al.
Published: (2024)
by: Gandhi, Rohan, et al.
Published: (2024)
Capsule: Efficient Player Isolation for Datacenters
by: Du, Zhouheng, et al.
Published: (2025)
by: Du, Zhouheng, et al.
Published: (2025)
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
by: Qin, Ruoyu, et al.
Published: (2026)
by: Qin, Ruoyu, et al.
Published: (2026)
Datacenter Energy Optimized Power Profiles
by: Narayanaswamy, Sreedhar, et al.
Published: (2025)
by: Narayanaswamy, Sreedhar, et al.
Published: (2025)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
by: Lu, Runyu, et al.
Published: (2025)
by: Lu, Runyu, et al.
Published: (2025)
Serving Compound Inference Systems on Datacenter GPUs
by: Devata, Sriram, et al.
Published: (2026)
by: Devata, Sriram, et al.
Published: (2026)
CFP: Efficient Optimization of Intra-Operator Parallelism Plans for Large Model Training
by: Hu, Weifang, et al.
Published: (2025)
by: Hu, Weifang, et al.
Published: (2025)
Characterization of Large Language Model Development in the Datacenter
by: Hu, Qinghao, et al.
Published: (2024)
by: Hu, Qinghao, et al.
Published: (2024)
Adaptive, Efficient and Fair Resource Allocation in Cloud Datacenters leveraging Weighted A3C Deep Reinforcement Learning
by: Kumari, Suchi, et al.
Published: (2025)
by: Kumari, Suchi, et al.
Published: (2025)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
by: Chen, Tiancheng, et al.
Published: (2025)
by: Chen, Tiancheng, et al.
Published: (2025)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
by: Xue, Chunyu, et al.
Published: (2026)
by: Xue, Chunyu, et al.
Published: (2026)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
by: Mehboob, Talha, et al.
Published: (2025)
by: Mehboob, Talha, et al.
Published: (2025)
The Ghost in the Datacenter: Link Flapping, Topology Knowledge Failures, and the FITO Category Mistake
by: Borrill, Paul
Published: (2026)
by: Borrill, Paul
Published: (2026)
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
by: Niewenhuis, Dante, et al.
Published: (2026)
by: Niewenhuis, Dante, et al.
Published: (2026)
DCGen 1.1 Technical Report: Generating Datacenter Configurations (including IT, Power, Cooling)
by: Gnibga, Wedan Emmanuel, et al.
Published: (2026)
by: Gnibga, Wedan Emmanuel, et al.
Published: (2026)
OpenDT: Exploring Datacenter Performance and Sustainability with a Self-Calibrating Digital Twin
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
Uncertainty-Aware Decarbonization for Datacenters
by: Li, Amy, et al.
Published: (2024)
by: Li, Amy, et al.
Published: (2024)
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
Distribution and Management of Datacenter Load Decoupling
by: Lin, Liuzixuan, et al.
Published: (2025)
by: Lin, Liuzixuan, et al.
Published: (2025)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
by: Ning, Rui, et al.
Published: (2026)
by: Ning, Rui, et al.
Published: (2026)
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
Hotspot-Aware Scheduling of Virtual Machines with Overcommitment for Ultimate Utilization in Cloud Datacenters
by: Wu, Jiaxi, et al.
Published: (2026)
by: Wu, Jiaxi, et al.
Published: (2026)
QAOA in Quantum Datacenters: Parallelization, Simulation, and Orchestration
by: Liaqat, Amana, et al.
Published: (2025)
by: Liaqat, Amana, et al.
Published: (2025)
Coordinated Cooling and Compute Management for AI Datacenters
by: Abera, Nardos Belay, et al.
Published: (2026)
by: Abera, Nardos Belay, et al.
Published: (2026)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
Designing Datacenter Power Delivery Hierarchies for the AI Era
by: Wilkins, Grant, et al.
Published: (2026)
by: Wilkins, Grant, et al.
Published: (2026)
Cost-aware Duration Prediction for Software Upgrades in Datacenters
by: Ding, Yi, et al.
Published: (2022)
by: Ding, Yi, et al.
Published: (2022)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
by: Lettich, Francesco, et al.
Published: (2024)
by: Lettich, Francesco, et al.
Published: (2024)
OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
Low-Latency Video Conferencing via Optimized Packet Routing and Reordering
by: Xiao, Yao, et al.
Published: (2023)
by: Xiao, Yao, et al.
Published: (2023)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
by: Papavasileiou, Ioannis, et al.
Published: (2026)
by: Papavasileiou, Ioannis, et al.
Published: (2026)
Designing Dense Satellite Clusters for Distributed Space-based Datacenters
by: Pénot, Jules, et al.
Published: (2026)
by: Pénot, Jules, et al.
Published: (2026)
Telepathic Datacenters: Fast RPCs using Shared CXL Memory
by: Mahar, Suyash, et al.
Published: (2024)
by: Mahar, Suyash, et al.
Published: (2024)
A Performance Analyzer for a Public Cloud's ML-Augmented VM Allocator
by: Bostandoost, Roozbeh, et al.
Published: (2025)
by: Bostandoost, Roozbeh, et al.
Published: (2025)
Adaptive K-PackCache: Cost-Centric Data Caching in Cloud
by: Sarkar, Suvarthi, et al.
Published: (2025)
by: Sarkar, Suvarthi, et al.
Published: (2025)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
Priority Matters: Optimising Kubernetes Clusters Usage with Constraint-Based Pod Packing
by: Christensen, Henrik Daniel, et al.
Published: (2025)
by: Christensen, Henrik Daniel, et al.
Published: (2025)
Declarative Data Pipeline for Large Scale ML Services
by: Yang, Yunzhao, et al.
Published: (2025)
by: Yang, Yunzhao, et al.
Published: (2025)
Parallax: Efficient LLM Inference Service over Decentralized Environment
by: Tong, Chris, et al.
Published: (2025)
by: Tong, Chris, et al.
Published: (2025)
FailSafe: High-performance Resilient Serving
by: Xu, Ziyi, et al.
Published: (2025)
by: Xu, Ziyi, et al.
Published: (2025)
Similar Items
-
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
by: Gandhi, Rohan, et al.
Published: (2024) -
Capsule: Efficient Player Isolation for Datacenters
by: Du, Zhouheng, et al.
Published: (2025) -
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
by: Qin, Ruoyu, et al.
Published: (2026) -
Datacenter Energy Optimized Power Profiles
by: Narayanaswamy, Sreedhar, et al.
Published: (2025) -
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
by: Lu, Runyu, et al.
Published: (2025)