Microsecond-scale Dynamic Validation of Idempotency for GPU Kernels
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Mingcong, Shen, Weihang, Peng, Guanwen, Chen, Rong, Chen, Haibo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024)
by: Wei, Xingda, et al.
Published: (2024)
Performance Isolation and Semantic Determinism in Efficient GPU Spatial Sharing
by: Yang, Zhenyuan, et al.
Published: (2026)
by: Yang, Zhenyuan, et al.
Published: (2026)
Nanvix: A Multikernel OS Design for High-Density Serverless Deployments
by: Segarra, Carlos, et al.
Published: (2026)
by: Segarra, Carlos, et al.
Published: (2026)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
Flexible Swapping for the Cloud
by: Pandurov, Milan, et al.
Published: (2024)
by: Pandurov, Milan, et al.
Published: (2024)
Shattering the Ephemeral Storage Cost Barrier for Data-Intensive Serverless Workflows
by: Ustiugov, Dmitrii, et al.
Published: (2023)
by: Ustiugov, Dmitrii, et al.
Published: (2023)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
by: Zhang, Dingyan, et al.
Published: (2024)
by: Zhang, Dingyan, et al.
Published: (2024)
VUDA: Breaking CUDA-Vulkan Isolation for Spatial Sharing of Compute and Graphics on the Same GPU
by: Xu, Bin, et al.
Published: (2026)
by: Xu, Bin, et al.
Published: (2026)
"Range as a Key" is the Key! Fast and Compact Cloud Block Store Index with RASK
by: Zhao, Haoru, et al.
Published: (2026)
by: Zhao, Haoru, et al.
Published: (2026)
LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling
by: Zhang, Dingyan, et al.
Published: (2026)
by: Zhang, Dingyan, et al.
Published: (2026)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025)
by: Xing, Jiarong, et al.
Published: (2025)
Peformance Isolation for Inference Processes in Edge GPU Systems
by: Martín, Juan José, et al.
Published: (2026)
by: Martín, Juan José, et al.
Published: (2026)
NCCLbpf: Verified, Composable Policy Execution for GPU Collective Communication
by: Zheng, Yusheng
Published: (2026)
by: Zheng, Yusheng
Published: (2026)
GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion
by: Yang, Yiwei, et al.
Published: (2026)
by: Yang, Yiwei, et al.
Published: (2026)
Serverless Cold Starts and Where to Find Them
by: Joosen, Artjom, et al.
Published: (2024)
by: Joosen, Artjom, et al.
Published: (2024)
RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU
by: Yu, Weiping, et al.
Published: (2025)
by: Yu, Weiping, et al.
Published: (2025)
Optimizing Task Scheduling in Heterogeneous Computing Environments: A Comparative Analysis of CPU, GPU, and ASIC Platforms Using E2C Simulator
by: Mohammadjafari, Ali, et al.
Published: (2024)
by: Mohammadjafari, Ali, et al.
Published: (2024)
GPUVM: GPU-driven Unified Virtual Memory
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
SwitchFS: Asynchronous Metadata Updates for Distributed Filesystems with In-Network Coordination
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
Rhea: Detecting Privilege-Escalated Evasive Ransomware Attacks Using Format-Aware Validation in the Cloud
by: Kim, Beom Heyn, et al.
Published: (2026)
by: Kim, Beom Heyn, et al.
Published: (2026)
SteelDB: Diagnosing Kernel-Space Bottlenecks in Cloud OLTP Databases
by: Kondo, Mitsumasa
Published: (2026)
by: Kondo, Mitsumasa
Published: (2026)
EdgeFlow: Fast Cold Starts for LLMs on Mobile Devices
by: Yan, Yongsheng, et al.
Published: (2026)
by: Yan, Yongsheng, et al.
Published: (2026)
Agent Centric Operating System -- a Comprehensive Review and Outlook for Operating System
by: Jia, Shian, et al.
Published: (2024)
by: Jia, Shian, et al.
Published: (2024)
Nixie: Efficient, Transparent Temporal Multiplexing for Consumer GPUs
by: Xu, Yechen, et al.
Published: (2026)
by: Xu, Yechen, et al.
Published: (2026)
Telepathic Datacenters: Fast RPCs using Shared CXL Memory
by: Mahar, Suyash, et al.
Published: (2024)
by: Mahar, Suyash, et al.
Published: (2024)
LEFT-RS: A Lock-Free Fault-Tolerant Resource Sharing Protocol for Multicore Real-Time Systems
by: Chen, Nan, et al.
Published: (2025)
by: Chen, Nan, et al.
Published: (2025)
Online Job Scheduler for Fault-tolerant Quantum Multiprogramming
by: Nishio, Shin, et al.
Published: (2025)
by: Nishio, Shin, et al.
Published: (2025)
Planetary computing for data-driven environmental policy-making
by: Ferris, Patrick, et al.
Published: (2023)
by: Ferris, Patrick, et al.
Published: (2023)
TrEnv-X: Transparently Share Serverless Execution Environments Across Different Functions and Nodes
by: Huang, Jialiang, et al.
Published: (2025)
by: Huang, Jialiang, et al.
Published: (2025)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)
by: Yuan, Renzhong, et al.
Published: (2026)
Locked In, Leaked Out: Measuring Isolation via Kernel Locks
by: Anjali, et al.
Published: (2025)
by: Anjali, et al.
Published: (2025)
EROICA: Online Performance Troubleshooting for Large-scale Model Training
by: Guan, Yu, et al.
Published: (2025)
by: Guan, Yu, et al.
Published: (2025)
Mitigating context switching in densely packed Linux clusters with Latency-Aware Group Scheduling
by: Isstaif, Al Amjad Tawfiq, et al.
Published: (2025)
by: Isstaif, Al Amjad Tawfiq, et al.
Published: (2025)
DPC: A Distributed Page Cache over CXL
by: Bergman, Shai, et al.
Published: (2026)
by: Bergman, Shai, et al.
Published: (2026)
Unlocking True Elasticity for the Cloud-Native Era with Dandelion
by: Kuchler, Tom, et al.
Published: (2025)
by: Kuchler, Tom, et al.
Published: (2025)
A Periodic Space of Distributed Computing: Vision & Framework
by: Salehi, Mohsen Amini, et al.
Published: (2026)
by: Salehi, Mohsen Amini, et al.
Published: (2026)
THEMIS: Time, Heterogeneity, and Energy Minded Scheduling for Fair Multi-Tenant Use in FPGAs
by: Karabulut, Emre, et al.
Published: (2024)
by: Karabulut, Emre, et al.
Published: (2024)
Mewz: Lightweight Execution Environment for WebAssembly with High Isolation and Portability using Unikernels
by: Ueda, Soichiro, et al.
Published: (2024)
by: Ueda, Soichiro, et al.
Published: (2024)
Similar Items
-
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024) -
Performance Isolation and Semantic Determinism in Efficient GPU Spatial Sharing
by: Yang, Zhenyuan, et al.
Published: (2026) -
Nanvix: A Multikernel OS Design for High-Density Serverless Deployments
by: Segarra, Carlos, et al.
Published: (2026) -
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024) -
Flexible Swapping for the Cloud
by: Pandurov, Milan, et al.
Published: (2024)