SIP: Autotuning GPU Native Schedules via Stochastic Instruction Perturbation
Fuente:
arXiv
Saved in:
| Main Authors: | He, Guoliang, Yoneki, Eiko |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning
by: He, Guoliang, et al.
Published: (2025)
by: He, Guoliang, et al.
Published: (2025)
GPU Performance Portability needs Autotuning
by: Ringlein, Burkhard, et al.
Published: (2025)
by: Ringlein, Burkhard, et al.
Published: (2025)
Qute: Towards Quantum-Native Database
by: Chen, Muzhi, et al.
Published: (2026)
by: Chen, Muzhi, et al.
Published: (2026)
Differentiable Initialization-Accelerated CPU-GPU Hybrid Combinatorial Scheduling
by: Liu, Mingju, et al.
Published: (2026)
by: Liu, Mingju, et al.
Published: (2026)
LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
by: Zhang, Tao, et al.
Published: (2026)
by: Zhang, Tao, et al.
Published: (2026)
Efficient LLM inference solution on Intel GPU
by: Wu, Hui, et al.
Published: (2023)
by: Wu, Hui, et al.
Published: (2023)
Analysis of LLM Vulnerability to GPU Soft Errors: An Instruction-Level Fault Injection Study
by: Chai, Duo, et al.
Published: (2025)
by: Chai, Duo, et al.
Published: (2025)
ChipAlign: Instruction Alignment in Large Language Models for Chip Design via Geodesic Interpolation
by: Deng, Chenhui, et al.
Published: (2024)
by: Deng, Chenhui, et al.
Published: (2024)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
by: Ma, Songchen, et al.
Published: (2026)
by: Ma, Songchen, et al.
Published: (2026)
SALSA: Simulated Annealing based Loop-Ordering Scheduler for DNN Accelerators
by: Jung, Victor J. B., et al.
Published: (2023)
by: Jung, Victor J. B., et al.
Published: (2023)
Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning Workloads
by: Pelke, Rebecca, et al.
Published: (2025)
by: Pelke, Rebecca, et al.
Published: (2025)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
by: Kim, Dowon, et al.
Published: (2025)
by: Kim, Dowon, et al.
Published: (2025)
Learning in Log-Domain: Subthreshold Analog AI Accelerator Based on Stochastic Gradient Descent
by: Tageldeen, Momen K, et al.
Published: (2025)
by: Tageldeen, Momen K, et al.
Published: (2025)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
by: Shin, Yongwon, et al.
Published: (2024)
by: Shin, Yongwon, et al.
Published: (2024)
GauRast: Enhancing GPU Triangle Rasterizers to Accelerate 3D Gaussian Splatting
by: Li, Sixu, et al.
Published: (2025)
by: Li, Sixu, et al.
Published: (2025)
Differentiable Combinatorial Scheduling at Scale
by: Liu, Mingju, et al.
Published: (2024)
by: Liu, Mingju, et al.
Published: (2024)
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
by: Ma, Shaobo, et al.
Published: (2024)
by: Ma, Shaobo, et al.
Published: (2024)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
by: He, Zifan, et al.
Published: (2025)
by: He, Zifan, et al.
Published: (2025)
BoolSkeleton: Boolean Network Skeletonization via Homogeneous Pattern Reduction
by: Ni, Liwei, et al.
Published: (2025)
by: Ni, Liwei, et al.
Published: (2025)
Enhancing Computational Efficiency in Intensive Domains via Redundant Residue Number Systems
by: Mousavi, Soudabeh, et al.
Published: (2024)
by: Mousavi, Soudabeh, et al.
Published: (2024)
Accelerating Post-Quantum Cryptography via LLM-Driven Hardware-Software Co-Design
by: Liao, Yuchao, et al.
Published: (2026)
by: Liao, Yuchao, et al.
Published: (2026)
ChatSVA: Bridging SVA Generation for Hardware Verification via Task-Specific LLMs
by: Fu, Lik Tung, et al.
Published: (2026)
by: Fu, Lik Tung, et al.
Published: (2026)
Shavette: Low Power Neural Network Acceleration via Algorithm-level Error Detection and Undervolting
by: Rinkinen, Mikael, et al.
Published: (2024)
by: Rinkinen, Mikael, et al.
Published: (2024)
MapTune: Advancing ASIC Technology Mapping via Reinforcement Learning Guided Library Tuning
by: Liu, Mingju, et al.
Published: (2024)
by: Liu, Mingju, et al.
Published: (2024)
Optimizing Neural Networks with Learnable Non-Linear Activation Functions via Lookup-Based FPGA Acceleration
by: Yin, Mengyuan, et al.
Published: (2025)
by: Yin, Mengyuan, et al.
Published: (2025)
ChatEDA: A Large Language Model Powered Autonomous Agent for EDA
by: He, Zhuolun, et al.
Published: (2023)
by: He, Zhuolun, et al.
Published: (2023)
EvoVerilog: Large Langugage Model Assisted Evolution of Verilog Code
by: Guo, Ping, et al.
Published: (2025)
by: Guo, Ping, et al.
Published: (2025)
Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement
by: Fang, Wenji, et al.
Published: (2026)
by: Fang, Wenji, et al.
Published: (2026)
TL-nvSRAM-CIM: Ultra-High-Density Three-Level ReRAM-Assisted Computing-in-nvSRAM with DC-Power Free Restore and Ternary MAC Operations
by: Wang, Dengfeng, et al.
Published: (2023)
by: Wang, Dengfeng, et al.
Published: (2023)
AMSnet-q: Unsupervised Circuit Identification and Performance Labeling for AMS Circuits
by: Zhang, Ze, et al.
Published: (2026)
by: Zhang, Ze, et al.
Published: (2026)
ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing
by: You, Kang, et al.
Published: (2026)
by: You, Kang, et al.
Published: (2026)
NVR: Vector Runahead on NPUs for Sparse Memory Access
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
Instruction Scheduling in the Saturn Vector Unit
by: Zhao, Jerry, et al.
Published: (2024)
by: Zhao, Jerry, et al.
Published: (2024)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
by: Pan, Yudong, et al.
Published: (2026)
by: Pan, Yudong, et al.
Published: (2026)
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
by: Guan, Yue, et al.
Published: (2026)
by: Guan, Yue, et al.
Published: (2026)
StoX-Net: Stochastic Processing of Partial Sums for Efficient In-Memory Computing DNN Accelerators
by: Rogers, Ethan G, et al.
Published: (2024)
by: Rogers, Ethan G, et al.
Published: (2024)
A Heterogeneous RISC-V based SoC for Secure Nano-UAV Navigation
by: Valente, Luca, et al.
Published: (2024)
by: Valente, Luca, et al.
Published: (2024)
RISC-V R-Extension: Advancing Efficiency with Rented-Pipeline for Edge DNN Processing
by: Kim, Won Hyeok, et al.
Published: (2024)
by: Kim, Won Hyeok, et al.
Published: (2024)
HAPM -- Hardware Aware Pruning Method for CNN hardware accelerators in resource constrained devices
by: Peccia, Federico Nicolas, et al.
Published: (2024)
by: Peccia, Federico Nicolas, et al.
Published: (2024)
Pragmatic Formal Verification Methodology for Clock Domain Crossing (CDC)
by: Kumar, Aman, et al.
Published: (2024)
by: Kumar, Aman, et al.
Published: (2024)
Similar Items
-
CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning
by: He, Guoliang, et al.
Published: (2025) -
GPU Performance Portability needs Autotuning
by: Ringlein, Burkhard, et al.
Published: (2025) -
Qute: Towards Quantum-Native Database
by: Chen, Muzhi, et al.
Published: (2026) -
Differentiable Initialization-Accelerated CPU-GPU Hybrid Combinatorial Scheduling
by: Liu, Mingju, et al.
Published: (2026) -
LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
by: Zhang, Tao, et al.
Published: (2026)