Characterization-Guided GPU Fault Resilience in NVIDIA MPS
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Rixin, Cui, Xingqi, Wang, Kaijian, Ding, Xinheng, Liu, Zirui, Wang, Yuke, Xing, Jiarong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Efficient and Practical GPU Multitasking in the Era of LLM
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
From Models to Operators: Rethinking Autoscaling Granularity for Large Generative Models
von: Cui, Xingqi, et al.
Veröffentlicht: (2025)
von: Cui, Xingqi, et al.
Veröffentlicht: (2025)
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies
von: Vanecek, Stepan, et al.
Veröffentlicht: (2025)
von: Vanecek, Stepan, et al.
Veröffentlicht: (2025)
Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
TurboFFT: A High-Performance Fast Fourier Transform with Fault Tolerance on GPU
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
To Repair or Not to Repair: Assessing Fault Resilience in MPI Stencil Applications
von: Rocco, Roberto, et al.
Veröffentlicht: (2024)
von: Rocco, Roberto, et al.
Veröffentlicht: (2024)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
A GPU accelerated mixed-precision Smoothed Particle Hydrodynamics framework with cell-based relative coordinates
von: Mao, Zirui, et al.
Veröffentlicht: (2023)
von: Mao, Zirui, et al.
Veröffentlicht: (2023)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
von: Sun, Desen, et al.
Veröffentlicht: (2025)
von: Sun, Desen, et al.
Veröffentlicht: (2025)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
von: Tramm, John, et al.
Veröffentlicht: (2024)
von: Tramm, John, et al.
Veröffentlicht: (2024)
DeepOps & SLURM: Your GPU Cluster Guide
von: Majee, Arindam
Veröffentlicht: (2024)
von: Majee, Arindam
Veröffentlicht: (2024)
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
Characterizing Production GPU Workloads using System-wide Telemetry Data
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
FT K-means: A High-Performance K-means on GPU with Fault Tolerance
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
Combining GPU and CPU for accelerating evolutionary computing workloads
von: Eynaliyev, Rustam, et al.
Veröffentlicht: (2025)
von: Eynaliyev, Rustam, et al.
Veröffentlicht: (2025)
Heat: Satellite's meat is GPU's poison
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
Byzantine Fault-Tolerant Min-Max Optimization
von: Liu, Shuo, et al.
Veröffentlicht: (2022)
von: Liu, Shuo, et al.
Veröffentlicht: (2022)
FastMPS: Revisit Data Parallel in Large-scale Matrix Product State Sampling
von: Chen, Yaojian, et al.
Veröffentlicht: (2025)
von: Chen, Yaojian, et al.
Veröffentlicht: (2025)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
von: Zuo, Jingwei, et al.
Veröffentlicht: (2026)
von: Zuo, Jingwei, et al.
Veröffentlicht: (2026)
Resilient Packet Forwarding: A Reinforcement Learning Approach to Routing in Gaussian Interconnected Networks with Clustered Faults
von: Charrwi, Mohammad Walid, et al.
Veröffentlicht: (2025)
von: Charrwi, Mohammad Walid, et al.
Veröffentlicht: (2025)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
Approximate Byzantine Fault-Tolerance in Distributed Optimization
von: Liu, Shuo, et al.
Veröffentlicht: (2021)
von: Liu, Shuo, et al.
Veröffentlicht: (2021)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
CROWDio: A Practical Mobile Crowd Computing Framework with Developer-Oriented Design, Adaptive Scheduling, and Fault Resilience
von: Manamperi, Lakshani, et al.
Veröffentlicht: (2026)
von: Manamperi, Lakshani, et al.
Veröffentlicht: (2026)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
Optimizing Bloom Filters for Modern GPU Architectures
von: Jünger, Daniel, et al.
Veröffentlicht: (2025)
von: Jünger, Daniel, et al.
Veröffentlicht: (2025)
Communication-Computation Pipeline Parallel Split Learning over Wireless Edge Networks
von: Liu, Chenyu, et al.
Veröffentlicht: (2025)
von: Liu, Chenyu, et al.
Veröffentlicht: (2025)
GAS: Generative Activation-Aided Asynchronous Split Federated Learning
von: Yang, Jiarong, et al.
Veröffentlicht: (2024)
von: Yang, Jiarong, et al.
Veröffentlicht: (2024)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
von: Park, Seongyeon, et al.
Veröffentlicht: (2024)
von: Park, Seongyeon, et al.
Veröffentlicht: (2024)
Six Times to Spare: Characterizing GPU-Accelerated 5G LDPC Decoding for Edge-RSU Communications
von: Barker, Ryan, et al.
Veröffentlicht: (2026)
von: Barker, Ryan, et al.
Veröffentlicht: (2026)
DuaLip-GPU Technical Report
von: Dexter, Gregory, et al.
Veröffentlicht: (2026)
von: Dexter, Gregory, et al.
Veröffentlicht: (2026)
A Preliminary Study on Accelerating Simulation Optimization with GPU Implementation
von: He, Jinghai, et al.
Veröffentlicht: (2024)
von: He, Jinghai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Efficient and Practical GPU Multitasking in the Era of LLM
von: Xing, Jiarong, et al.
Veröffentlicht: (2025) -
From Models to Operators: Rethinking Autoscaling Granularity for Large Generative Models
von: Cui, Xingqi, et al.
Veröffentlicht: (2025) -
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
von: Li, Junjie, et al.
Veröffentlicht: (2024) -
MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies
von: Vanecek, Stepan, et al.
Veröffentlicht: (2025) -
Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)