GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jiaqi, Sun, Jingwei, Luo, Jiyu, Li, Han, Sun, Guangzhong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Graph neural networks with configuration cross-attention for tensor compilers
by: Khizbullin, Dmitrii, et al.
Published: (2024)
by: Khizbullin, Dmitrii, et al.
Published: (2024)
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
by: Zhang, Kaixuan, et al.
Published: (2026)
by: Zhang, Kaixuan, et al.
Published: (2026)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
by: Chrapek, Marcin, et al.
Published: (2025)
by: Chrapek, Marcin, et al.
Published: (2025)
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
by: Palaniappan, Kathiravan
Published: (2026)
by: Palaniappan, Kathiravan
Published: (2026)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
by: Zhou, Cyrus, et al.
Published: (2023)
by: Zhou, Cyrus, et al.
Published: (2023)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
by: Bi, Zhen, et al.
Published: (2026)
by: Bi, Zhen, et al.
Published: (2026)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
by: Ibrahim, Muhammad Sohail, et al.
Published: (2024)
by: Ibrahim, Muhammad Sohail, et al.
Published: (2024)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
by: Saha, Rappy, et al.
Published: (2026)
by: Saha, Rappy, et al.
Published: (2026)
Toward A Formalized Approach for Spike Sorting Algorithms and Hardware Evaluation
by: Zhang, Tim, et al.
Published: (2022)
by: Zhang, Tim, et al.
Published: (2022)
Search Your Block Floating Point Scales!
by: Gupta, Tanmaey, et al.
Published: (2026)
by: Gupta, Tanmaey, et al.
Published: (2026)
A2Q+: Improving Accumulator-Aware Weight Quantization
by: Colbert, Ian, et al.
Published: (2024)
by: Colbert, Ian, et al.
Published: (2024)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024)
by: Karami, Rachid, et al.
Published: (2024)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
by: Atmer, Hannah, et al.
Published: (2025)
by: Atmer, Hannah, et al.
Published: (2025)
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
by: Zhang, Kaixuan, et al.
Published: (2026)
by: Zhang, Kaixuan, et al.
Published: (2026)
LightningSimV2: Faster and Scalable Simulation for High-Level Synthesis via Graph Compilation and Optimization
by: Sarkar, Rishov, et al.
Published: (2024)
by: Sarkar, Rishov, et al.
Published: (2024)
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
by: Liao, Gang, et al.
Published: (2025)
by: Liao, Gang, et al.
Published: (2025)
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
by: Yang, Peiming, et al.
Published: (2025)
by: Yang, Peiming, et al.
Published: (2025)
Characterizing and Understanding HGNN Training on GPUs
by: Han, Dengke, et al.
Published: (2024)
by: Han, Dengke, et al.
Published: (2024)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
by: Choudhary, Mansi, et al.
Published: (2025)
by: Choudhary, Mansi, et al.
Published: (2025)
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression
by: Ltaief, Hatem, et al.
Published: (2024)
by: Ltaief, Hatem, et al.
Published: (2024)
Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures
by: Anik, Shafayat Mowla, et al.
Published: (2026)
by: Anik, Shafayat Mowla, et al.
Published: (2026)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
by: Patwari, Rajeev, et al.
Published: (2025)
by: Patwari, Rajeev, et al.
Published: (2025)
Design Space Exploration of Approximate Computing Techniques with a Reinforcement Learning Approach
by: Saeedi, Sepide, et al.
Published: (2023)
by: Saeedi, Sepide, et al.
Published: (2023)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
by: Zhang, Hang, et al.
Published: (2025)
by: Zhang, Hang, et al.
Published: (2025)
Accelerating Transistor-Level Simulation of Integrated Circuits via Equivalence of RC Long-Chain Structures
by: Tang, Ruibai, et al.
Published: (2025)
by: Tang, Ruibai, et al.
Published: (2025)
Fast NF4 Dequantization Kernels for Large Language Model Inference
by: Qi, Xiangbo, et al.
Published: (2026)
by: Qi, Xiangbo, et al.
Published: (2026)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
by: Yang, Hanchen, et al.
Published: (2025)
by: Yang, Hanchen, et al.
Published: (2025)
Recurrent CircuitSAT Sampling for Sequential Circuits
by: Ardakani, Arash, et al.
Published: (2025)
by: Ardakani, Arash, et al.
Published: (2025)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
by: Lübeck, Konstantin, et al.
Published: (2024)
by: Lübeck, Konstantin, et al.
Published: (2024)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
by: Zhou, Zikai, et al.
Published: (2025)
by: Zhou, Zikai, et al.
Published: (2025)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
by: Hendria, Willy Fitra
Published: (2026)
by: Hendria, Willy Fitra
Published: (2026)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
by: Müller, Mika Markus, et al.
Published: (2025)
by: Müller, Mika Markus, et al.
Published: (2025)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
by: Chhugani, Jatin, et al.
Published: (2026)
by: Chhugani, Jatin, et al.
Published: (2026)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
by: Ali, Wajid, et al.
Published: (2025)
by: Ali, Wajid, et al.
Published: (2025)
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
by: Ham, Hyungkyu, et al.
Published: (2024)
by: Ham, Hyungkyu, et al.
Published: (2024)
Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry
by: Chen, Ziji, et al.
Published: (2025)
by: Chen, Ziji, et al.
Published: (2025)
OmniSim: Simulating Hardware with C Speed and RTL Accuracy for High-Level Synthesis Designs
by: Sarkar, Rishov, et al.
Published: (2025)
by: Sarkar, Rishov, et al.
Published: (2025)
Similar Items
-
Graph neural networks with configuration cross-attention for tensor compilers
by: Khizbullin, Dmitrii, et al.
Published: (2024) -
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
by: Zhang, Kaixuan, et al.
Published: (2026) -
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
by: Chrapek, Marcin, et al.
Published: (2025) -
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
by: Palaniappan, Kathiravan
Published: (2026) -
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
by: Zhou, Cyrus, et al.
Published: (2023)