Revealing NVIDIA Closed-Source Driver Command Streams for CPU-GPU Runtime Behavior Insight
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Yuang, Karlin, Ian, Grant, Ryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmark-based Study of CPU/GPU Power-Related Features through JAX and TensorFlow
von: Tchakoute, Roblex Nana, et al.
Veröffentlicht: (2025)
von: Tchakoute, Roblex Nana, et al.
Veröffentlicht: (2025)
Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
von: Prieto, Pablo, et al.
Veröffentlicht: (2025)
von: Prieto, Pablo, et al.
Veröffentlicht: (2025)
MambaCPU: Enhanced Correlation Mining with State Space Models for CPU Performance Prediction
von: Liu, Xiaoman
Veröffentlicht: (2024)
von: Liu, Xiaoman
Veröffentlicht: (2024)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
Comparing CPU and GPU compute of PERMANOVA on MI300A
von: Sfiligoi, Igor
Veröffentlicht: (2025)
von: Sfiligoi, Igor
Veröffentlicht: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
Achieving Consistent and Comparable CPU Evaluation
von: Wang, Chenxi, et al.
Veröffentlicht: (2024)
von: Wang, Chenxi, et al.
Veröffentlicht: (2024)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
von: Chrapek, Marcin, et al.
Veröffentlicht: (2025)
von: Chrapek, Marcin, et al.
Veröffentlicht: (2025)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
CAPSim: A Fast CPU Performance Simulator Using Attention-based Predictor
von: Xu, Buqing, et al.
Veröffentlicht: (2025)
von: Xu, Buqing, et al.
Veröffentlicht: (2025)
Architectural Trade-offs in the Energy-Efficient Era: A Comparative Study of power-capping NVIDIA H100 and H200
von: Ujeniya, Aditya, et al.
Veröffentlicht: (2026)
von: Ujeniya, Aditya, et al.
Veröffentlicht: (2026)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
CPU Optimization of a Monocular 3D Biomechanics Pipeline for Low-Resource Deployment
von: Zhang, Yan, et al.
Veröffentlicht: (2026)
von: Zhang, Yan, et al.
Veröffentlicht: (2026)
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison
von: Li, Ruihao, et al.
Veröffentlicht: (2026)
von: Li, Ruihao, et al.
Veröffentlicht: (2026)
How to Increase Energy Efficiency with a Single Linux Command
von: Jelvani, Alborz, et al.
Veröffentlicht: (2025)
von: Jelvani, Alborz, et al.
Veröffentlicht: (2025)
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2026)
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2026)
Towards CPU Performance Prediction: New Challenge Benchmark Dataset and Novel Approach
von: Liu, Xiaoman
Veröffentlicht: (2024)
von: Liu, Xiaoman
Veröffentlicht: (2024)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
Unikernels vs. Containers: A Runtime-Level Performance Comparison for Resource-Constrained Edge Workloads
von: Dinh-Tuan, Hai
Veröffentlicht: (2025)
von: Dinh-Tuan, Hai
Veröffentlicht: (2025)
AMD MI300X GPU Performance Analysis
von: Ambati, Chandrish, et al.
Veröffentlicht: (2025)
von: Ambati, Chandrish, et al.
Veröffentlicht: (2025)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
von: Luo, Weile, et al.
Veröffentlicht: (2025)
von: Luo, Weile, et al.
Veröffentlicht: (2025)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
von: Wang, Haoxin, et al.
Veröffentlicht: (2025)
von: Wang, Haoxin, et al.
Veröffentlicht: (2025)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
von: Tharwani, Jay, et al.
Veröffentlicht: (2025)
von: Tharwani, Jay, et al.
Veröffentlicht: (2025)
DSO: A GPU Energy Efficiency Optimizer by Fusing Dynamic and Static Information
von: Wang, Qiang, et al.
Veröffentlicht: (2024)
von: Wang, Qiang, et al.
Veröffentlicht: (2024)
Streaming Data in HPC Workflows Using ADIOS
von: Eisenhauer, Greg, et al.
Veröffentlicht: (2024)
von: Eisenhauer, Greg, et al.
Veröffentlicht: (2024)
Elastic Sketch under Random Stationary Streams: Limiting Behavior and Near-Optimal Configuration
von: Mazziane, Younes Ben, et al.
Veröffentlicht: (2026)
von: Mazziane, Younes Ben, et al.
Veröffentlicht: (2026)
AGIPC: Adaptive In-Solve Algebraic Coarsening for GPU IPC
von: Wang, Xuan, et al.
Veröffentlicht: (2026)
von: Wang, Xuan, et al.
Veröffentlicht: (2026)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
von: Nasr-Esfahany, Arash, et al.
Veröffentlicht: (2025)
von: Nasr-Esfahany, Arash, et al.
Veröffentlicht: (2025)
Runtime Verification on Abstract Finite State Models
von: Jevitha, KP, et al.
Veröffentlicht: (2024)
von: Jevitha, KP, et al.
Veröffentlicht: (2024)
Towards Multi-dimensional Elasticity for Pervasive Stream Processing Services
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
cuTeSpMM: Accelerating Sparse-Dense Matrix Multiplication using GPU Tensor Cores
von: Xiang, Lizhi, et al.
Veröffentlicht: (2025)
von: Xiang, Lizhi, et al.
Veröffentlicht: (2025)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
StiffGIPC: Advancing GPU IPC for stiff affine-deformable simulation
von: Huang, Kemeng, et al.
Veröffentlicht: (2024)
von: Huang, Kemeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Benchmark-based Study of CPU/GPU Power-Related Features through JAX and TensorFlow
von: Tchakoute, Roblex Nana, et al.
Veröffentlicht: (2025) -
Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
von: Prieto, Pablo, et al.
Veröffentlicht: (2025) -
MambaCPU: Enhanced Correlation Mining with State Space Models for CPU Performance Prediction
von: Liu, Xiaoman
Veröffentlicht: (2024) -
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026) -
Comparing CPU and GPU compute of PERMANOVA on MI300A
von: Sfiligoi, Igor
Veröffentlicht: (2025)