Inside VOLT: Designing an Open-Source GPU Compiler
Fuente:
arXiv
Salvato in:
| Autori principali: | Jeong, Shinnung, Ahn, Chihyo, Pu, Huanzhi, Zhao, Jisheng, Kim, Hyesoon, Tine, Blaise |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CuLifter: Lifting GPU Binaries to Typed IR
di: Zhao, Jisheng, et al.
Pubblicazione: (2026)
di: Zhao, Jisheng, et al.
Pubblicazione: (2026)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
di: Chung, Euijun, et al.
Pubblicazione: (2026)
di: Chung, Euijun, et al.
Pubblicazione: (2026)
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
di: Pu, Huanzhi, et al.
Pubblicazione: (2025)
di: Pu, Huanzhi, et al.
Pubblicazione: (2025)
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
di: Li, Bingyao, et al.
Pubblicazione: (2024)
di: Li, Bingyao, et al.
Pubblicazione: (2024)
Switchboard: An Open-Source Framework for Modular Simulation of Large Hardware Systems
di: Herbst, Steven, et al.
Pubblicazione: (2024)
di: Herbst, Steven, et al.
Pubblicazione: (2024)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2024)
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2024)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
HetGPU: The pursuit of making binary compatibility towards GPUs
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
di: Agrawal, Anirudha, et al.
Pubblicazione: (2024)
di: Agrawal, Anirudha, et al.
Pubblicazione: (2024)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
di: Elwasif, Wael, et al.
Pubblicazione: (2022)
di: Elwasif, Wael, et al.
Pubblicazione: (2022)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods
di: Fatima, Amel, et al.
Pubblicazione: (2026)
di: Fatima, Amel, et al.
Pubblicazione: (2026)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -- A Critical Review
di: Majeed, Ashiyana Abdul, et al.
Pubblicazione: (2025)
di: Majeed, Ashiyana Abdul, et al.
Pubblicazione: (2025)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
di: McDaniel, Adam, et al.
Pubblicazione: (2026)
di: McDaniel, Adam, et al.
Pubblicazione: (2026)
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
di: Eleftherakis, Panagiotis-Eleftherios, et al.
Pubblicazione: (2026)
di: Eleftherakis, Panagiotis-Eleftherios, et al.
Pubblicazione: (2026)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
di: Torres, L. A., et al.
Pubblicazione: (2024)
di: Torres, L. A., et al.
Pubblicazione: (2024)
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
di: Zhang, Yichao, et al.
Pubblicazione: (2024)
di: Zhang, Yichao, et al.
Pubblicazione: (2024)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
di: Zhou, Zhuoshan, et al.
Pubblicazione: (2026)
di: Zhou, Zhuoshan, et al.
Pubblicazione: (2026)
GigaAPI for GPU Parallelization
di: Suvarna, M., et al.
Pubblicazione: (2025)
di: Suvarna, M., et al.
Pubblicazione: (2025)
Debunking the CUDA Myth Towards GPU-based AI Systems
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
Parallelizing a modern GPU simulator
di: Huerta, Rodrigo, et al.
Pubblicazione: (2025)
di: Huerta, Rodrigo, et al.
Pubblicazione: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
di: Szafarczyk, Robert, et al.
Pubblicazione: (2025)
di: Szafarczyk, Robert, et al.
Pubblicazione: (2025)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
di: Pati, Suchita, et al.
Pubblicazione: (2024)
di: Pati, Suchita, et al.
Pubblicazione: (2024)
RapidOMS: FPGA-based Open Modification Spectral Library Searching with HD Computing
di: Pinge, Sumukh, et al.
Pubblicazione: (2024)
di: Pinge, Sumukh, et al.
Pubblicazione: (2024)
ZipFlow: a Compiler-based Framework to Unleash Compressed Data Movement for Modern GPUs
di: Yeo, Gwangoo, et al.
Pubblicazione: (2026)
di: Yeo, Gwangoo, et al.
Pubblicazione: (2026)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
di: Prakriya, Neha, et al.
Pubblicazione: (2023)
di: Prakriya, Neha, et al.
Pubblicazione: (2023)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
di: Qu, Huanyu, et al.
Pubblicazione: (2025)
di: Qu, Huanyu, et al.
Pubblicazione: (2025)
Simopt-Power: Leveraging Simulation Metadata for Low-Power Design Synthesis
di: Wadhwa, Eashan, et al.
Pubblicazione: (2025)
di: Wadhwa, Eashan, et al.
Pubblicazione: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
di: Zhang, Yichao, et al.
Pubblicazione: (2026)
di: Zhang, Yichao, et al.
Pubblicazione: (2026)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
di: Noh, Si Ung, et al.
Pubblicazione: (2024)
di: Noh, Si Ung, et al.
Pubblicazione: (2024)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
di: Shen, Aofeng, et al.
Pubblicazione: (2025)
di: Shen, Aofeng, et al.
Pubblicazione: (2025)
Muchisim: A Simulation Framework for Design Exploration of Multi-Chip Manycore Systems
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
di: Mo, Zhiwen, et al.
Pubblicazione: (2026)
di: Mo, Zhiwen, et al.
Pubblicazione: (2026)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
Atomique: A Quantum Compiler for Reconfigurable Neutral Atom Arrays
di: Wang, Hanrui, et al.
Pubblicazione: (2023)
di: Wang, Hanrui, et al.
Pubblicazione: (2023)
Documenti analoghi
-
CuLifter: Lifting GPU Binaries to Typed IR
di: Zhao, Jisheng, et al.
Pubblicazione: (2026) -
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
di: Chung, Euijun, et al.
Pubblicazione: (2026) -
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
di: Pu, Huanzhi, et al.
Pubblicazione: (2025) -
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
di: Li, Bingyao, et al.
Pubblicazione: (2024) -
Switchboard: An Open-Source Framework for Modular Simulation of Large Hardware Systems
di: Herbst, Steven, et al.
Pubblicazione: (2024)