ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mai, Haohui, Guo, Xiaoyan, Ding, Xiangyun, Li, Daifeng, Yu, Qiuchu, Guo, Chenzhun, Wang, Cong, Zhao, Jiacheng, Kozyrakis, Christos, Yuan, Binhang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
FailSafe: High-performance Resilient Serving
von: Xu, Ziyi, et al.
Veröffentlicht: (2025)
von: Xu, Ziyi, et al.
Veröffentlicht: (2025)
cedar: Optimized and Unified Machine Learning Input Data Pipelines
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
LoHan: Low-Cost High-Performance Framework to Fine-Tune 100B Model on a Consumer GPU
von: Liao, Changyue, et al.
Veröffentlicht: (2024)
von: Liao, Changyue, et al.
Veröffentlicht: (2024)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
von: Peng, You, et al.
Veröffentlicht: (2026)
von: Peng, You, et al.
Veröffentlicht: (2026)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2025)
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2025)
ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
von: Liu, Rixin, et al.
Veröffentlicht: (2026)
von: Liu, Rixin, et al.
Veröffentlicht: (2026)
FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
von: Mei, Junyi, et al.
Veröffentlicht: (2024)
von: Mei, Junyi, et al.
Veröffentlicht: (2024)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
Strata: Hierarchical Context Caching for Long Context Language Model Serving
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
VDCores: Resource Decoupled Programming and Execution for Asynchronous GPU
von: He, Zijian, et al.
Veröffentlicht: (2026)
von: He, Zijian, et al.
Veröffentlicht: (2026)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
von: Lin, Shouxu, et al.
Veröffentlicht: (2026)
von: Lin, Shouxu, et al.
Veröffentlicht: (2026)
Cephalo: Harnessing Heterogeneous GPU Clusters for Training Transformer Models
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2024)
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2024)
DeepOps & SLURM: Your GPU Cluster Guide
von: Majee, Arindam
Veröffentlicht: (2024)
von: Majee, Arindam
Veröffentlicht: (2024)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026)
von: Liang, Yan, et al.
Veröffentlicht: (2026)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
Optimizing Bloom Filters for Modern GPU Architectures
von: Jünger, Daniel, et al.
Veröffentlicht: (2025)
von: Jünger, Daniel, et al.
Veröffentlicht: (2025)
BandPilot: Towards Performance- and Contention-Aware GPU Dispatching in AI Clusters
von: Zhang, Kunming, et al.
Veröffentlicht: (2025)
von: Zhang, Kunming, et al.
Veröffentlicht: (2025)
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
Heat: Satellite's meat is GPU's poison
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
A Preliminary Study on Accelerating Simulation Optimization with GPU Implementation
von: He, Jinghai, et al.
Veröffentlicht: (2024)
von: He, Jinghai, et al.
Veröffentlicht: (2024)
EXaCTz: Guaranteed Extremum Graph and Contour Tree Preservation for Distributed- and GPU-Parallel Lossy Compression
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
Towards Fast Setup and High Throughput of GPU Serverless Computing
von: Zhao, Han, et al.
Veröffentlicht: (2024)
von: Zhao, Han, et al.
Veröffentlicht: (2024)
Regulating Branch Parallelism in LLM Serving
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
von: Siavashi, Ahmad, et al.
Veröffentlicht: (2025)
von: Siavashi, Ahmad, et al.
Veröffentlicht: (2025)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
Cascadia: An Efficient Cascade Serving System for Large Language Models
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
ACC Saturator: Automatic Kernel Optimization for Directive-Based GPU Code
von: Matsumura, Kazuaki, et al.
Veröffentlicht: (2023)
von: Matsumura, Kazuaki, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024) -
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
von: Yan, Ran, et al.
Veröffentlicht: (2025) -
FailSafe: High-performance Resilient Serving
von: Xu, Ziyi, et al.
Veröffentlicht: (2025) -
cedar: Optimized and Unified Machine Learning Input Data Pipelines
von: Zhao, Mark, et al.
Veröffentlicht: (2024) -
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)