Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Anjiang, Nie, Allen, Teixeira, Thiago S. F. X., Yadav, Rohan, Lee, Wonchan, Wang, Ke, Aiken, Alex |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mapple: A Domain-Specific Language for Mapping Distributed Programs
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Composing Distributed Computations Through Task and Kernel Fusion
by: Yadav, Rohan, et al.
Published: (2024)
by: Yadav, Rohan, et al.
Published: (2024)
Automatic Tracing in Task-Based Runtime Systems
by: Yadav, Rohan, et al.
Published: (2024)
by: Yadav, Rohan, et al.
Published: (2024)
On the Duality of Task and Actor Programming Models
by: Yadav, Rohan, et al.
Published: (2025)
by: Yadav, Rohan, et al.
Published: (2025)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Parallelized Multi-Agent Bayesian Optimization in Lava
by: Snyder, Shay, et al.
Published: (2024)
by: Snyder, Shay, et al.
Published: (2024)
Task-Based Programming for Adaptive Mesh Refinement in Compressible Flow Simulations
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Performance-Driven Optimization of Parallel Breadth-First Search
by: Bhaskar, Marati, et al.
Published: (2025)
by: Bhaskar, Marati, et al.
Published: (2025)
Design Principles of Dynamic Resource Management for High-Performance Parallel Programming Models
by: Huber, Dominik, et al.
Published: (2024)
by: Huber, Dominik, et al.
Published: (2024)
Efficient Fault Tolerance for Pipelined Query Engines via Write-ahead Lineage
by: Wang, Ziheng, et al.
Published: (2024)
by: Wang, Ziheng, et al.
Published: (2024)
Automated Programmatic Performance Analysis of Parallel Programs
by: Cankur, Onur, et al.
Published: (2024)
by: Cankur, Onur, et al.
Published: (2024)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
by: Li, Cong, et al.
Published: (2025)
by: Li, Cong, et al.
Published: (2025)
Dynamic Contract Analysis for Parallel Programming Models
by: Oraji, Yussur Mustafa, et al.
Published: (2026)
by: Oraji, Yussur Mustafa, et al.
Published: (2026)
STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems
by: Egersdoerfer, Chris, et al.
Published: (2026)
by: Egersdoerfer, Chris, et al.
Published: (2026)
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
by: Gandhi, Rohan, et al.
Published: (2024)
by: Gandhi, Rohan, et al.
Published: (2024)
Maxing Out the SVM: Performance Impact of Memory and Program Cache Sizes in the Agave Validator
by: Vural, Turan, et al.
Published: (2025)
by: Vural, Turan, et al.
Published: (2025)
Exploring DAOS Interfaces and Performance
by: Manubens, Nicolau, et al.
Published: (2024)
by: Manubens, Nicolau, et al.
Published: (2024)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
A New Execution Model and Executor for Adaptively Optimizing the Performance of Parallel Algorithms Using HPX Runtime System
by: Mohammadiporshokooh, Karame, et al.
Published: (2025)
by: Mohammadiporshokooh, Karame, et al.
Published: (2025)
Deferred Objects to Enhance Smart Contract Programming with Optimistic Parallel Execution
by: Mitenkov, George, et al.
Published: (2024)
by: Mitenkov, George, et al.
Published: (2024)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
by: Knorr, Fabian, et al.
Published: (2025)
by: Knorr, Fabian, et al.
Published: (2025)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
by: Li, Haoyang, et al.
Published: (2024)
by: Li, Haoyang, et al.
Published: (2024)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
by: Chen, Haoyu, et al.
Published: (2025)
by: Chen, Haoyu, et al.
Published: (2025)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
by: Srivatsa, Vikranth, et al.
Published: (2026)
by: Srivatsa, Vikranth, et al.
Published: (2026)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
The Merit of Simple Policies: Buying Performance With Parallelism and System Architecture
by: Yildiz, Mert, et al.
Published: (2025)
by: Yildiz, Mert, et al.
Published: (2025)
High-Performance Parallelization of Dijkstra's Algorithm Using MPI and CUDA
by: Song, Boyang
Published: (2025)
by: Song, Boyang
Published: (2025)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
by: Ma, Tenghui, et al.
Published: (2026)
by: Ma, Tenghui, et al.
Published: (2026)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
by: Zhao, Alan, et al.
Published: (2026)
by: Zhao, Alan, et al.
Published: (2026)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
by: Ma, Mulei, et al.
Published: (2025)
by: Ma, Mulei, et al.
Published: (2025)
Fusionize++: Improving Serverless Application Performance Using Dynamic Task Inlining and Infrastructure Optimization
by: Schirmer, Trever, et al.
Published: (2023)
by: Schirmer, Trever, et al.
Published: (2023)
Optimizing View Change for Byzantine Fault Tolerance in Parallel Consensus
by: Xie, Yifei, et al.
Published: (2026)
by: Xie, Yifei, et al.
Published: (2026)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
by: Luo, Jiajun, et al.
Published: (2024)
by: Luo, Jiajun, et al.
Published: (2024)
WRATH: Workload Resilience Across Task Hierarchies in Task-based Parallel Programming Frameworks
by: Zhou, Sicheng, et al.
Published: (2025)
by: Zhou, Sicheng, et al.
Published: (2025)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
by: Nie, Chengyi, et al.
Published: (2024)
by: Nie, Chengyi, et al.
Published: (2024)
Parallel DNA Sequence Alignment on High-Performance Systems with CUDA and MPI
by: Zwaka, Linus
Published: (2024)
by: Zwaka, Linus
Published: (2024)
GTaP: A GPU-Resident Fork-Join Task-Parallel Runtime with a Pragma-Based Interface
by: Maeda, Yuki, et al.
Published: (2026)
by: Maeda, Yuki, et al.
Published: (2026)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
by: Wei, Jinhui, et al.
Published: (2025)
by: Wei, Jinhui, et al.
Published: (2025)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Similar Items
-
Mapple: A Domain-Specific Language for Mapping Distributed Programs
by: Wei, Anjiang, et al.
Published: (2025) -
Composing Distributed Computations Through Task and Kernel Fusion
by: Yadav, Rohan, et al.
Published: (2024) -
Automatic Tracing in Task-Based Runtime Systems
by: Yadav, Rohan, et al.
Published: (2024) -
On the Duality of Task and Actor Programming Models
by: Yadav, Rohan, et al.
Published: (2025) -
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
by: Wei, Anjiang, et al.
Published: (2025)