Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Hongzheng, Fan, Bin, Collins, Alexander, Hagedorn, Bastian, Gaburov, Evghenii, Masuda, Masahiro, Brookhart, Matthew, Sullivan, Chris, Knight, Jason, Zhang, Zhiru, Grover, Vinod |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
by: Soi, Rupanshu, et al.
Published: (2025)
by: Soi, Rupanshu, et al.
Published: (2025)
Control Flow Management in Modern GPUs
by: Shoushtary, Mojtaba Abaie, et al.
Published: (2024)
by: Shoushtary, Mojtaba Abaie, et al.
Published: (2024)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
by: Zhang, Niansong, et al.
Published: (2025)
by: Zhang, Niansong, et al.
Published: (2025)
Allo: A Programming Model for Composable Accelerator Design
by: Chen, Hongzheng, et al.
Published: (2024)
by: Chen, Hongzheng, et al.
Published: (2024)
Dato: A Task-Based Programming Model for Dataflow Accelerators
by: Fang, Shihan, et al.
Published: (2025)
by: Fang, Shihan, et al.
Published: (2025)
Stream-HLS: Towards Automatic Dataflow Acceleration
by: Basalama, Suhail, et al.
Published: (2025)
by: Basalama, Suhail, et al.
Published: (2025)
Evaluation of Hardware-based Video Encoders on Modern GPUs for UHD Live-Streaming
by: Arunruangsirilert, Kasidis, et al.
Published: (2025)
by: Arunruangsirilert, Kasidis, et al.
Published: (2025)
Privacy-Preserving Performance Profiling of In-The-Wild GPUs
by: McDougall, Ian, et al.
Published: (2025)
by: McDougall, Ian, et al.
Published: (2025)
A Systematic Characterization of LLM Inference on GPUs
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
by: Hong, Jeongmin, et al.
Published: (2024)
by: Hong, Jeongmin, et al.
Published: (2024)
Hierarchical Resource Partitioning on Modern GPUs: A Reinforcement Learning Approach
by: Saroliya, Urvij, et al.
Published: (2024)
by: Saroliya, Urvij, et al.
Published: (2024)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
by: Chen, Hongzheng, et al.
Published: (2023)
by: Chen, Hongzheng, et al.
Published: (2023)
SAMIPS: A Synthesised Asynchronous Processor
by: Zhang, Qianyi, et al.
Published: (2024)
by: Zhang, Qianyi, et al.
Published: (2024)
Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
by: Chowdhary, Sangeeta, et al.
Published: (2026)
by: Chowdhary, Sangeeta, et al.
Published: (2026)
GauS: Differentiable Scheduling Optimization via Gaussian Reparameterization
by: Cai, Yaohui, et al.
Published: (2026)
by: Cai, Yaohui, et al.
Published: (2026)
Automatic Hardware Pragma Insertion in High-Level Synthesis: A Non-Linear Programming Approach
by: Pouget, Stéphane, et al.
Published: (2024)
by: Pouget, Stéphane, et al.
Published: (2024)
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
by: Pu, Huanzhi, et al.
Published: (2025)
by: Pu, Huanzhi, et al.
Published: (2025)
WaSP: Warp Scheduling to Mimic Prefetching in Graphics Workloads
by: Joseph, Diya, et al.
Published: (2024)
by: Joseph, Diya, et al.
Published: (2024)
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
by: Kim, Hansung, et al.
Published: (2024)
by: Kim, Hansung, et al.
Published: (2024)
CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs
by: Hu, Jiajun, et al.
Published: (2025)
by: Hu, Jiajun, et al.
Published: (2025)
ZipFlow: a Compiler-based Framework to Unleash Compressed Data Movement for Modern GPUs
by: Yeo, Gwangoo, et al.
Published: (2026)
by: Yeo, Gwangoo, et al.
Published: (2026)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
by: Xu, Haocheng, et al.
Published: (2024)
by: Xu, Haocheng, et al.
Published: (2024)
TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
by: Guo, Licheng, et al.
Published: (2022)
by: Guo, Licheng, et al.
Published: (2022)
Study on the Particle Sorting Performance for Reactor Monte Carlo Neutron Transport on Apple Unified Memory GPUs
by: Liu, Changyuan
Published: (2024)
by: Liu, Changyuan
Published: (2024)
GPIR: Enabling Practical Private Information Retrieval with GPUs
by: Ji, Hyesung, et al.
Published: (2026)
by: Ji, Hyesung, et al.
Published: (2026)
Hidden Risks of Unmonitored GPUs in Intelligent Transportation Systems
by: Puspa, Sefatun-Noor, et al.
Published: (2026)
by: Puspa, Sefatun-Noor, et al.
Published: (2026)
Cicero: Addressing Algorithmic and Architectural Bottlenecks in Neural Rendering by Radiance Warping and Memory Optimizations
by: Feng, Yu, et al.
Published: (2024)
by: Feng, Yu, et al.
Published: (2024)
Iceberg: Enhancing HLS Modeling with Synthetic Data
by: Ding, Zijian, et al.
Published: (2025)
by: Ding, Zijian, et al.
Published: (2025)
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory Access
by: Wang, Luming, et al.
Published: (2024)
by: Wang, Luming, et al.
Published: (2024)
How Much Progress Has There Been in NVIDIA Datacenter GPUs?
by: Del Sozzo, Emanuele, et al.
Published: (2026)
by: Del Sozzo, Emanuele, et al.
Published: (2026)
Sustainable Hardware Specialization
by: Dangi, Pranav, et al.
Published: (2024)
by: Dangi, Pranav, et al.
Published: (2024)
Warp-Cortex: An Asynchronous, Memory-Efficient Architecture for Million-Agent Cognitive Scaling on Consumer Hardware
by: Williams, Jorge L. Ruiz
Published: (2026)
by: Williams, Jorge L. Ruiz
Published: (2026)
Five-Minute Rule 40 Years Later: A First-Principles Revisit for Modern Memory Hierarchy
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
ANCoEF: Asynchronous Neuromorphic Algorithm/Hardware Co-Exploration Framework with a Fully Asynchronous Simulator
by: Zhang, Jian, et al.
Published: (2024)
by: Zhang, Jian, et al.
Published: (2024)
A Performance Model for Warp Specialization Kernels
by: Liu, Zhengyang, et al.
Published: (2025)
by: Liu, Zhengyang, et al.
Published: (2025)
e-boost: Boosted E-Graph Extraction with Adaptive Heuristics and Exact Solving
by: Yin, Jiaqi, et al.
Published: (2025)
by: Yin, Jiaqi, et al.
Published: (2025)
Kitsune: Enabling Dataflow Execution on GPUs
by: Davies, Michael, et al.
Published: (2025)
by: Davies, Michael, et al.
Published: (2025)
Analyzing Modern NVIDIA GPU cores
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
Less is More: Hop-Wise Graph Attention for Scalable and Generalizable Learning on Circuits
by: Deng, Chenhui, et al.
Published: (2024)
by: Deng, Chenhui, et al.
Published: (2024)
Similar Items
-
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
by: Soi, Rupanshu, et al.
Published: (2025) -
Control Flow Management in Modern GPUs
by: Shoushtary, Mojtaba Abaie, et al.
Published: (2024) -
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
by: Zhang, Niansong, et al.
Published: (2025) -
Allo: A Programming Model for Composable Accelerator Design
by: Chen, Hongzheng, et al.
Published: (2024) -
Dato: A Task-Based Programming Model for Dataflow Accelerators
by: Fang, Shihan, et al.
Published: (2025)