The Merit of Simple Policies: Buying Performance With Parallelism and System Architecture
Fuente:
arXiv
Saved in:
| Main Authors: | Yildiz, Mert, Rolich, Alexey, Baiocchi, Andrea |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dispatching Odyssey: Exploring Performance in Computing Clusters under Real-world Workloads
by: Yildiz, Mert, et al.
Published: (2025)
by: Yildiz, Mert, et al.
Published: (2025)
"Two-Stagification": Job Dispatching in Large-Scale Clusters via a Two-Stage Architecture
by: Yildiz, Mert, et al.
Published: (2025)
by: Yildiz, Mert, et al.
Published: (2025)
Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption
by: Yildiz, Mert, et al.
Published: (2026)
by: Yildiz, Mert, et al.
Published: (2026)
MeritRank: Sybil Tolerant Reputation for Merit-based Tokenomics
by: Nasrulin, Bulat, et al.
Published: (2022)
by: Nasrulin, Bulat, et al.
Published: (2022)
Parallel DNA Sequence Alignment on High-Performance Systems with CUDA and MPI
by: Zwaka, Linus
Published: (2024)
by: Zwaka, Linus
Published: (2024)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
by: Hidayetoglu, Mert, et al.
Published: (2025)
by: Hidayetoglu, Mert, et al.
Published: (2025)
The Entropy of Parallel Systems
by: Adefemi, Temitayo
Published: (2025)
by: Adefemi, Temitayo
Published: (2025)
STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems
by: Egersdoerfer, Chris, et al.
Published: (2026)
by: Egersdoerfer, Chris, et al.
Published: (2026)
Performance-Driven Optimization of Parallel Breadth-First Search
by: Bhaskar, Marati, et al.
Published: (2025)
by: Bhaskar, Marati, et al.
Published: (2025)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
by: Zhang, Hongbin, et al.
Published: (2025)
by: Zhang, Hongbin, et al.
Published: (2025)
Linear Complexity $\mathcal{H}^2$ Direct Solver for Fine-Grained Parallel Architectures
by: Boukaram, Wajih, et al.
Published: (2025)
by: Boukaram, Wajih, et al.
Published: (2025)
A New Execution Model and Executor for Adaptively Optimizing the Performance of Parallel Algorithms Using HPX Runtime System
by: Mohammadiporshokooh, Karame, et al.
Published: (2025)
by: Mohammadiporshokooh, Karame, et al.
Published: (2025)
High-Performance Parallelization of Dijkstra's Algorithm Using MPI and CUDA
by: Song, Boyang
Published: (2025)
by: Song, Boyang
Published: (2025)
Design Principles of Dynamic Resource Management for High-Performance Parallel Programming Models
by: Huber, Dominik, et al.
Published: (2024)
by: Huber, Dominik, et al.
Published: (2024)
Balancing Pipeline Parallelism with Vocabulary Parallelism
by: Yeung, Man Tsung, et al.
Published: (2024)
by: Yeung, Man Tsung, et al.
Published: (2024)
Parallel Reduced Order Modeling for Digital Twins using High-Performance Computing Workflows
by: de Parga, S. Ares, et al.
Published: (2024)
by: de Parga, S. Ares, et al.
Published: (2024)
Efficient Parallel Implementation of the Pilot Assignment Problem in Massive MIMO Systems
by: Alqudah, Eman, et al.
Published: (2025)
by: Alqudah, Eman, et al.
Published: (2025)
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
by: Kang, Xueze, et al.
Published: (2025)
by: Kang, Xueze, et al.
Published: (2025)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
by: Knorr, Fabian, et al.
Published: (2025)
by: Knorr, Fabian, et al.
Published: (2025)
Modular Architecture for High-Performance and Low Overhead Data Transfers
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
CkIO: Parallel File Input for Over-Decomposed Task-Based Systems
by: Jacob, Mathew, et al.
Published: (2024)
by: Jacob, Mathew, et al.
Published: (2024)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
by: Ramesh, Risshab Srinivas
Published: (2024)
by: Ramesh, Risshab Srinivas
Published: (2024)
Automated Programmatic Performance Analysis of Parallel Programs
by: Cankur, Onur, et al.
Published: (2024)
by: Cankur, Onur, et al.
Published: (2024)
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
by: Shan, Baodi, et al.
Published: (2024)
by: Shan, Baodi, et al.
Published: (2024)
New Kids: An Architecture and Performance Investigation of Second-Generation Serverless Platforms
by: Schirmer, Trever, et al.
Published: (2026)
by: Schirmer, Trever, et al.
Published: (2026)
Parallel Data Object Creation: Towards Scalable Metadata Management in High-Performance I/O Library
by: Li, Youjia, et al.
Published: (2025)
by: Li, Youjia, et al.
Published: (2025)
Arctic Inference with Shift Parallelism: Fast and Efficient Open Source Inference System for Enterprise AI
by: Rajbhandari, Samyam, et al.
Published: (2025)
by: Rajbhandari, Samyam, et al.
Published: (2025)
Lectures on Parallel Computing
by: Träff, Jesper Larsson
Published: (2024)
by: Träff, Jesper Larsson
Published: (2024)
Optimizing CDN Architectures: Multi-Metric Algorithmic Breakthroughs for Edge and Distributed Performance
by: Absur, Md Nurul, et al.
Published: (2024)
by: Absur, Md Nurul, et al.
Published: (2024)
A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture
by: Yi, Xinyao
Published: (2024)
by: Yi, Xinyao
Published: (2024)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
Performance Cost Tradeoffs in Intelligent Load Balancing for Multi Data Center Cloud Systems: From Static Policies to Adaptive Resource Distribution
by: Najafabadi, Saeid Aghasoleymani, et al.
Published: (2025)
by: Najafabadi, Saeid Aghasoleymani, et al.
Published: (2025)
Tight Bounds on Channel Reliability via Generalized Quorum Systems (Extended Version)
by: Naser-Pastoriza, Alejandro, et al.
Published: (2025)
by: Naser-Pastoriza, Alejandro, et al.
Published: (2025)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
by: Chen, Jiefei, et al.
Published: (2026)
by: Chen, Jiefei, et al.
Published: (2026)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
by: Tang, Ding, et al.
Published: (2024)
by: Tang, Ding, et al.
Published: (2024)
Assessing Redundancy Strategies to Improve Availability in Virtualized System Architectures
by: Silva, Alison, et al.
Published: (2025)
by: Silva, Alison, et al.
Published: (2025)
EvoSort: A Genetic-Algorithm-Based Adaptive Parallel Sorting Framework for Large-Scale High Performance Computing
by: Raj, Shashank, et al.
Published: (2025)
by: Raj, Shashank, et al.
Published: (2025)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
by: Li, Zixuan, et al.
Published: (2026)
by: Li, Zixuan, et al.
Published: (2026)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
by: Svedas, Jonas, et al.
Published: (2026)
by: Svedas, Jonas, et al.
Published: (2026)
Similar Items
-
Dispatching Odyssey: Exploring Performance in Computing Clusters under Real-world Workloads
by: Yildiz, Mert, et al.
Published: (2025) -
"Two-Stagification": Job Dispatching in Large-Scale Clusters via a Two-Stage Architecture
by: Yildiz, Mert, et al.
Published: (2025) -
Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption
by: Yildiz, Mert, et al.
Published: (2026) -
MeritRank: Sybil Tolerant Reputation for Merit-based Tokenomics
by: Nasrulin, Bulat, et al.
Published: (2022) -
Parallel DNA Sequence Alignment on High-Performance Systems with CUDA and MPI
by: Zwaka, Linus
Published: (2024)