GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | VG, Jithin, PS, Ditto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Inference Acceleration for Large Language Models on CPUs
von: PS, Ditto, et al.
Veröffentlicht: (2024)
von: PS, Ditto, et al.
Veröffentlicht: (2024)
MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
von: Hwang, Changho, et al.
Veröffentlicht: (2025)
von: Hwang, Changho, et al.
Veröffentlicht: (2025)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems
von: Wawdhane, Sourish, et al.
Veröffentlicht: (2026)
von: Wawdhane, Sourish, et al.
Veröffentlicht: (2026)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
A Parallel CPU-GPU Framework for Batching Heuristic Operations in Depth-First Heuristic Search
von: Futuhi, Ehsan, et al.
Veröffentlicht: (2025)
von: Futuhi, Ehsan, et al.
Veröffentlicht: (2025)
UCCL-Zip: Lossless Compression Supercharged GPU Communication
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
Beyond the GPU: The Strategic Role of FPGAs in the Next Wave of AI
von: Jiménez, Arturo Urías
Veröffentlicht: (2025)
von: Jiménez, Arturo Urías
Veröffentlicht: (2025)
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
Towards Scalable GPU-Accelerated SNN Training via Temporal Fusion
von: Li, Yanchen, et al.
Veröffentlicht: (2024)
von: Li, Yanchen, et al.
Veröffentlicht: (2024)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
von: Xu, Lang, et al.
Veröffentlicht: (2024)
von: Xu, Lang, et al.
Veröffentlicht: (2024)
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
von: Yang, Ruijia, et al.
Veröffentlicht: (2026)
von: Yang, Ruijia, et al.
Veröffentlicht: (2026)
Accelerated Digital Twin Learning for Edge AI: A Comparison of FPGA and Mobile GPU
von: Xu, Bin, et al.
Veröffentlicht: (2025)
von: Xu, Bin, et al.
Veröffentlicht: (2025)
Speeding up Local Optimization in Vehicle Routing with Tensor-based GPU Acceleration
von: Lei, Zhenyu, et al.
Veröffentlicht: (2025)
von: Lei, Zhenyu, et al.
Veröffentlicht: (2025)
Accurate GPU Memory Prediction for Deep Learning Jobs through Dynamic Analysis
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
Reducing Fragmentation and Starvation in GPU Clusters through Dynamic Multi-Objective Scheduling
von: Mamirov, Akhmadillo
Veröffentlicht: (2025)
von: Mamirov, Akhmadillo
Veröffentlicht: (2025)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
von: Spoczynski, Marcin, et al.
Veröffentlicht: (2026)
von: Spoczynski, Marcin, et al.
Veröffentlicht: (2026)
PolyKAN: Efficient Fused GPU Operators for Polynomial Kolmogorov-Arnold Network Variants
von: Yu, Mingkun, et al.
Veröffentlicht: (2025)
von: Yu, Mingkun, et al.
Veröffentlicht: (2025)
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference
von: Zhao, Bingzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Bingzhe, et al.
Veröffentlicht: (2025)
Practical offloading for fine-tuning LLM on commodity GPU via learned sparse projectors
von: Chen, Siyuan, et al.
Veröffentlicht: (2024)
von: Chen, Siyuan, et al.
Veröffentlicht: (2024)
ProbSelect: Stochastic Client Selection for GPU-Accelerated Compute Devices in the 3D Continuum
von: Stanisic, Andrija, et al.
Veröffentlicht: (2025)
von: Stanisic, Andrija, et al.
Veröffentlicht: (2025)
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
von: Li, Yufei, et al.
Veröffentlicht: (2025)
von: Li, Yufei, et al.
Veröffentlicht: (2025)
Debunking the CUDA Myth Towards GPU-based AI Systems
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
Fine-Tuning and Serving Gemma 4 31B on Google Cloud TPU: A Technical Comparison with GPU Baselines
von: Kishnani, Jatin, et al.
Veröffentlicht: (2026)
von: Kishnani, Jatin, et al.
Veröffentlicht: (2026)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
ODIN-Based CPU-GPU Architecture with Replay-Driven Simulation and Emulation
von: Dorairaj, Nij, et al.
Veröffentlicht: (2026)
von: Dorairaj, Nij, et al.
Veröffentlicht: (2026)
GPU Memory Prediction for Multimodal Model Training
von: Jeong, Jinwoo, et al.
Veröffentlicht: (2025)
von: Jeong, Jinwoo, et al.
Veröffentlicht: (2025)
Fine-Tuning GPT-5 for GPU Kernel Generation
von: Tehrani, Ali, et al.
Veröffentlicht: (2026)
von: Tehrani, Ali, et al.
Veröffentlicht: (2026)
Hybrid Learning and Optimization-Based Dynamic Scheduling for DL Workloads on Heterogeneous GPU Clusters
von: Dongare, Shruti, et al.
Veröffentlicht: (2025)
von: Dongare, Shruti, et al.
Veröffentlicht: (2025)
Lobster: A GPU-Accelerated Framework for Neurosymbolic Programming
von: Biberstein, Paul, et al.
Veröffentlicht: (2025)
von: Biberstein, Paul, et al.
Veröffentlicht: (2025)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
von: Zhao, Dan, et al.
Veröffentlicht: (2024)
von: Zhao, Dan, et al.
Veröffentlicht: (2024)
ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants
von: Mai, Haohui, et al.
Veröffentlicht: (2026)
von: Mai, Haohui, et al.
Veröffentlicht: (2026)
Thousand-GPU Large-Scale Training and Optimization Recipe for AI-Native Cloud Embodied Intelligence Infrastructure
von: Guo, Yongjian, et al.
Veröffentlicht: (2026)
von: Guo, Yongjian, et al.
Veröffentlicht: (2026)
AEGIS: Scaling Long-Sequence Homomorphic Encrypted Transformer Inference via Hybrid Parallelism on Multi-GPU Systems
von: Gong, Zhaoting, et al.
Veröffentlicht: (2026)
von: Gong, Zhaoting, et al.
Veröffentlicht: (2026)
GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
von: Chu, Ruifan, et al.
Veröffentlicht: (2025)
von: Chu, Ruifan, et al.
Veröffentlicht: (2025)
Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
von: Zhang, Biyao, et al.
Veröffentlicht: (2025)
von: Zhang, Biyao, et al.
Veröffentlicht: (2025)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
von: Zhu, Wenbin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenbin, et al.
Veröffentlicht: (2025)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
Making Room for AI: Multi-GPU Molecular Dynamics with Deep Potentials in GROMACS
von: Pennati, Luca, et al.
Veröffentlicht: (2026)
von: Pennati, Luca, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Inference Acceleration for Large Language Models on CPUs
von: PS, Ditto, et al.
Veröffentlicht: (2024) -
MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
von: Hwang, Changho, et al.
Veröffentlicht: (2025) -
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026) -
GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems
von: Wawdhane, Sourish, et al.
Veröffentlicht: (2026) -
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)