Control Flow Management in Modern GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Shoushtary, Mojtaba Abaie, Murgadas, Jordi Tubella, Gonzalez, Antonio |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyzing and Improving Hardware Modeling of Accel-Sim
by: Huerta, Rodrigo, et al.
Published: (2024)
by: Huerta, Rodrigo, et al.
Published: (2024)
Analyzing Modern NVIDIA GPU cores
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
Privacy-Preserving Performance Profiling of In-The-Wild GPUs
by: McDougall, Ian, et al.
Published: (2025)
by: McDougall, Ian, et al.
Published: (2025)
A Systematic Characterization of LLM Inference on GPUs
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
Towards Efficient LUT-based PIM: A Scalable and Low-Power Approach for Modern Workloads
by: Khabbazan, Bahareh, et al.
Published: (2025)
by: Khabbazan, Bahareh, et al.
Published: (2025)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
by: Hong, Jeongmin, et al.
Published: (2024)
by: Hong, Jeongmin, et al.
Published: (2024)
Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
by: Chowdhary, Sangeeta, et al.
Published: (2026)
by: Chowdhary, Sangeeta, et al.
Published: (2026)
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
by: Chen, Hongzheng, et al.
Published: (2025)
by: Chen, Hongzheng, et al.
Published: (2025)
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
by: Kim, Hansung, et al.
Published: (2024)
by: Kim, Hansung, et al.
Published: (2024)
CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs
by: Hu, Jiajun, et al.
Published: (2025)
by: Hu, Jiajun, et al.
Published: (2025)
ZipFlow: a Compiler-based Framework to Unleash Compressed Data Movement for Modern GPUs
by: Yeo, Gwangoo, et al.
Published: (2026)
by: Yeo, Gwangoo, et al.
Published: (2026)
Study on the Particle Sorting Performance for Reactor Monte Carlo Neutron Transport on Apple Unified Memory GPUs
by: Liu, Changyuan
Published: (2024)
by: Liu, Changyuan
Published: (2024)
Evaluation of Hardware-based Video Encoders on Modern GPUs for UHD Live-Streaming
by: Arunruangsirilert, Kasidis, et al.
Published: (2025)
by: Arunruangsirilert, Kasidis, et al.
Published: (2025)
CIS: Composable Instruction Set for Data Streaming Applications
by: Yang, Yu, et al.
Published: (2024)
by: Yang, Yu, et al.
Published: (2024)
Decoupled Control Flow and Data Access in RISC-V GPGPUs
by: Sarda, Giuseppe M., et al.
Published: (2025)
by: Sarda, Giuseppe M., et al.
Published: (2025)
Loop Control Management in Tightly Coupled Processor Arrays (TCPAs)
by: Walter, Dominik, et al.
Published: (2026)
by: Walter, Dominik, et al.
Published: (2026)
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
by: Chang, Kaiyan, et al.
Published: (2025)
by: Chang, Kaiyan, et al.
Published: (2025)
GPIR: Enabling Practical Private Information Retrieval with GPUs
by: Ji, Hyesung, et al.
Published: (2026)
by: Ji, Hyesung, et al.
Published: (2026)
Hidden Risks of Unmonitored GPUs in Intelligent Transportation Systems
by: Puspa, Sefatun-Noor, et al.
Published: (2026)
by: Puspa, Sefatun-Noor, et al.
Published: (2026)
Rethinking the Producer-Consumer Relationship in Modern DRAM-Based Systems
by: Patel, Minesh, et al.
Published: (2024)
by: Patel, Minesh, et al.
Published: (2024)
Compromising the Intelligence of Modern DNNs: On the Effectiveness of Targeted RowPress
by: Zhou, Ranyang, et al.
Published: (2024)
by: Zhou, Ranyang, et al.
Published: (2024)
How Much Progress Has There Been in NVIDIA Datacenter GPUs?
by: Del Sozzo, Emanuele, et al.
Published: (2026)
by: Del Sozzo, Emanuele, et al.
Published: (2026)
Hierarchical Resource Partitioning on Modern GPUs: A Reinforcement Learning Approach
by: Saroliya, Urvij, et al.
Published: (2024)
by: Saroliya, Urvij, et al.
Published: (2024)
FireBridge: Cycle-Accurate Hardware + Firmware Co-Verification for Modern Accelerators
by: Abarajithan, G, et al.
Published: (2026)
by: Abarajithan, G, et al.
Published: (2026)
AERO: Adaptive Erase Operation for Improving Lifetime and Performance of Modern NAND Flash-Based SSDs
by: Cho, Sungjun, et al.
Published: (2024)
by: Cho, Sungjun, et al.
Published: (2024)
STAR: Improving Lifetime and Performance of High-Capacity Modern SSDs Using State-Aware Randomizer
by: Kwon, Omin, et al.
Published: (2025)
by: Kwon, Omin, et al.
Published: (2025)
Five-Minute Rule 40 Years Later: A First-Principles Revisit for Modern Memory Hierarchy
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
ARAS: An Adaptive Low-Cost ReRAM-Based Accelerator for DNNs
by: Sabri, Mohammad, et al.
Published: (2024)
by: Sabri, Mohammad, et al.
Published: (2024)
Hamun: An Approximate Computation Method to Prolong the Lifespan of ReRAM-Based Accelerators
by: Sabri, Mohammad, et al.
Published: (2025)
by: Sabri, Mohammad, et al.
Published: (2025)
QUADOL: A Quality-Driven Approximate Logic Synthesis Method Exploiting Dual-Output LUTs for Modern FPGAs
by: Shi, Jian, et al.
Published: (2024)
by: Shi, Jian, et al.
Published: (2024)
SwiftSpatial: Spatial Joins on Modern Hardware
by: Jiang, Wenqi, et al.
Published: (2023)
by: Jiang, Wenqi, et al.
Published: (2023)
Evaluating Computing Platforms for Sustainability: A Comparative Analysis of FPGAs against ASICs, GPUs, and CPUs
by: Sudarshan, Chetan Choppali, et al.
Published: (2026)
by: Sudarshan, Chetan Choppali, et al.
Published: (2026)
FastFlow in FPGA Stacks of Data Centers
by: Paul, Rourab, et al.
Published: (2024)
by: Paul, Rourab, et al.
Published: (2024)
ControlPULP: A RISC-V On-Chip Parallel Power Controller for Many-Core HPC Processors with FPGA-Based Hardware-In-The-Loop Power and Thermal Emulation
by: Ottaviano, Alessandro, et al.
Published: (2023)
by: Ottaviano, Alessandro, et al.
Published: (2023)
Addressing memory bandwidth scalability in vector processors for streaming applications
by: Altayo, Jordi, et al.
Published: (2025)
by: Altayo, Jordi, et al.
Published: (2025)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
by: Mao, Shunyu, et al.
Published: (2024)
by: Mao, Shunyu, et al.
Published: (2024)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
by: Xu, Haocheng, et al.
Published: (2024)
by: Xu, Haocheng, et al.
Published: (2024)
Kitsune: Enabling Dataflow Execution on GPUs
by: Davies, Michael, et al.
Published: (2025)
by: Davies, Michael, et al.
Published: (2025)
DEER: Deep Runahead for Instruction Prefetching on Modern Mobile Workloads
by: Vahdatniya, Parmida, et al.
Published: (2025)
by: Vahdatniya, Parmida, et al.
Published: (2025)
A Scalable Resource Management Layer for FPGA SoCs in 6G Radio Units
by: Bartzoudis, Nikolaos, et al.
Published: (2025)
by: Bartzoudis, Nikolaos, et al.
Published: (2025)
Similar Items
-
Analyzing and Improving Hardware Modeling of Accel-Sim
by: Huerta, Rodrigo, et al.
Published: (2024) -
Analyzing Modern NVIDIA GPU cores
by: Huerta, Rodrigo, et al.
Published: (2025) -
Privacy-Preserving Performance Profiling of In-The-Wild GPUs
by: McDougall, Ian, et al.
Published: (2025) -
A Systematic Characterization of LLM Inference on GPUs
by: Wang, Haonan, et al.
Published: (2025) -
Towards Efficient LUT-based PIM: A Scalable and Low-Power Approach for Modern Workloads
by: Khabbazan, Bahareh, et al.
Published: (2025)