Methodology for GPU Frequency Switching Latency Measurement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Velicka, Daniel, Vysocky, Ondrej, Riha, Lubomir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
von: Homola, Jakub, et al.
Veröffentlicht: (2025)
von: Homola, Jakub, et al.
Veröffentlicht: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
von: Chen, Jiesong, et al.
Veröffentlicht: (2026)
von: Chen, Jiesong, et al.
Veröffentlicht: (2026)
Optimizing Bloom Filters for Modern GPU Architectures
von: Jünger, Daniel, et al.
Veröffentlicht: (2025)
von: Jünger, Daniel, et al.
Veröffentlicht: (2025)
Revisiting Speculative Leaderless Protocols for Low-Latency BFT Replication
von: Qian, Daniel, et al.
Veröffentlicht: (2026)
von: Qian, Daniel, et al.
Veröffentlicht: (2026)
GPU Sharing with Triples Mode
von: Byun, Chansup, et al.
Veröffentlicht: (2024)
von: Byun, Chansup, et al.
Veröffentlicht: (2024)
Asynchronous Latency and Fast Atomic Snapshot
von: Bezerra, João Paulo, et al.
Veröffentlicht: (2024)
von: Bezerra, João Paulo, et al.
Veröffentlicht: (2024)
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
von: Seo, Eunil, et al.
Veröffentlicht: (2025)
von: Seo, Eunil, et al.
Veröffentlicht: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
DAG it off: Latency Prefers No Common Coins
von: Amores-Sesar, Ignacio, et al.
Veröffentlicht: (2025)
von: Amores-Sesar, Ignacio, et al.
Veröffentlicht: (2025)
Accelerating Biclique Counting on GPU
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
GPU-Accelerated Syndrome Decoding for Quantum LDPC Codes below the 63 $μ$s Latency Threshold
von: Ferraz, Oscar, et al.
Veröffentlicht: (2025)
von: Ferraz, Oscar, et al.
Veröffentlicht: (2025)
Towards Exascale Computing for Astrophysical Simulation Leveraging the Leonardo EuroHPC System
von: Shukla, Nitin, et al.
Veröffentlicht: (2025)
von: Shukla, Nitin, et al.
Veröffentlicht: (2025)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
Areon: Latency-Friendly and Resilient Multi-Proposer Consensus
von: Castro-Castilla, Álvaro, et al.
Veröffentlicht: (2025)
von: Castro-Castilla, Álvaro, et al.
Veröffentlicht: (2025)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
Predictable LLM Serving on GPU Clusters
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
DuaLip-GPU Technical Report
von: Dexter, Gregory, et al.
Veröffentlicht: (2026)
von: Dexter, Gregory, et al.
Veröffentlicht: (2026)
Incidence Constraints in Hypergraph Partitioning on GPU
von: Ronzani, Marco, et al.
Veröffentlicht: (2026)
von: Ronzani, Marco, et al.
Veröffentlicht: (2026)
GPU Accelerated Sparse Cholesky Factorization
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
Heat: Satellite's meat is GPU's poison
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
von: Yuan, Zhehu, et al.
Veröffentlicht: (2024)
Coordinating GPU Data Centers and Power Grid Regulation Service for Exogenous Carbon Benefits
von: Jahanshahi, Ali, et al.
Veröffentlicht: (2026)
von: Jahanshahi, Ali, et al.
Veröffentlicht: (2026)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
Hiding Latencies in Network-Based Image Loading for Deep Learning
von: Versaci, Francesco, et al.
Veröffentlicht: (2025)
von: Versaci, Francesco, et al.
Veröffentlicht: (2025)
Low Latency, High Bandwidth Streaming of Experimental Data with EJFAT
von: Baldin, Ilya, et al.
Veröffentlicht: (2025)
von: Baldin, Ilya, et al.
Veröffentlicht: (2025)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
von: Stoyanov, Radostin, et al.
Veröffentlicht: (2025)
von: Stoyanov, Radostin, et al.
Veröffentlicht: (2025)
Understanding the Landscape of Ampere GPU Memory Errors
von: Zhu, Zhu, et al.
Veröffentlicht: (2025)
von: Zhu, Zhu, et al.
Veröffentlicht: (2025)
Adaptive Multidimensional Quadrature on Multi-GPU Systems
von: Tonarelli, Melanie, et al.
Veröffentlicht: (2025)
von: Tonarelli, Melanie, et al.
Veröffentlicht: (2025)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
von: Siavashi, Ahmad, et al.
Veröffentlicht: (2025)
von: Siavashi, Ahmad, et al.
Veröffentlicht: (2025)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
Falcon: Advancing Asynchronous BFT Consensus for Lower Latency and Enhanced Throughput
von: Dai, Xiaohai, et al.
Veröffentlicht: (2025)
von: Dai, Xiaohai, et al.
Veröffentlicht: (2025)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
CD-Raft: Reducing the Latency of Distributed Consensus in Cross-Domain Sites
von: Wang, Yangyang, et al.
Veröffentlicht: (2026)
von: Wang, Yangyang, et al.
Veröffentlicht: (2026)
A New Approach for Evaluating the Performance of Distributed Latency-Sensitive Services
von: Theodoropoulos, Theodoros, et al.
Veröffentlicht: (2024)
von: Theodoropoulos, Theodoros, et al.
Veröffentlicht: (2024)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
von: Homola, Jakub, et al.
Veröffentlicht: (2025) -
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2023) -
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
von: Chen, Jiesong, et al.
Veröffentlicht: (2026) -
Optimizing Bloom Filters for Modern GPU Architectures
von: Jünger, Daniel, et al.
Veröffentlicht: (2025) -
Revisiting Speculative Leaderless Protocols for Low-Latency BFT Replication
von: Qian, Daniel, et al.
Veröffentlicht: (2026)