Towards a high-performance AI compiler with upstream MLIR
Fuente:
arXiv
Guardado en:
| Autores principales: | Golin, Renato, Chelini, Lorenzo, Siemieniuk, Adam, Madhu, Kavitha, Hasabnis, Niranjan, Pabst, Hans, Georganas, Evangelos, Heinecke, Alexander |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
por: Torres, L. A., et al.
Publicado: (2024)
por: Torres, L. A., et al.
Publicado: (2024)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
por: Colagrande, Luca, et al.
Publicado: (2026)
por: Colagrande, Luca, et al.
Publicado: (2026)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
por: Asquini, Lorenzo, et al.
Publicado: (2025)
por: Asquini, Lorenzo, et al.
Publicado: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
por: Zhang, Chen, et al.
Publicado: (2026)
por: Zhang, Chen, et al.
Publicado: (2026)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
por: Kim, Hyeseong, et al.
Publicado: (2026)
por: Kim, Hyeseong, et al.
Publicado: (2026)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
por: McDaniel, Adam, et al.
Publicado: (2026)
por: McDaniel, Adam, et al.
Publicado: (2026)
Forge-UGC: FX optimization and register-graph engine for universal graph compiler
por: Kumar, Satyam, et al.
Publicado: (2026)
por: Kumar, Satyam, et al.
Publicado: (2026)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
por: Kwak, Hyunseok, et al.
Publicado: (2025)
por: Kwak, Hyunseok, et al.
Publicado: (2025)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
por: Kreuzer, Anke, et al.
Publicado: (2019)
por: Kreuzer, Anke, et al.
Publicado: (2019)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
por: Qiu, Tong Dong, et al.
Publicado: (2023)
por: Qiu, Tong Dong, et al.
Publicado: (2023)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
por: Xu, Weihong, et al.
Publicado: (2025)
por: Xu, Weihong, et al.
Publicado: (2025)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
por: Negi, Shubham, et al.
Publicado: (2025)
por: Negi, Shubham, et al.
Publicado: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
por: Kubo, Tatsuya, et al.
Publicado: (2025)
por: Kubo, Tatsuya, et al.
Publicado: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
por: Liu, Lian, et al.
Publicado: (2026)
por: Liu, Lian, et al.
Publicado: (2026)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
por: Chen, Yanru, et al.
Publicado: (2025)
por: Chen, Yanru, et al.
Publicado: (2025)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
por: Kurzynski, Marco, et al.
Publicado: (2025)
por: Kurzynski, Marco, et al.
Publicado: (2025)
Efficient deadlock avoidance for 2D mesh NoCs that use OQ or VOQ routers
por: Papaphilippou, Philippos, et al.
Publicado: (2023)
por: Papaphilippou, Philippos, et al.
Publicado: (2023)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
por: Orenes-Vera, Marcelo, et al.
Publicado: (2023)
por: Orenes-Vera, Marcelo, et al.
Publicado: (2023)
LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity Multicores
por: García-García, Adrián, et al.
Publicado: (2024)
por: García-García, Adrián, et al.
Publicado: (2024)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
por: Li, Bohan, et al.
Publicado: (2026)
por: Li, Bohan, et al.
Publicado: (2026)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
por: Font, Martí Llopart, et al.
Publicado: (2026)
por: Font, Martí Llopart, et al.
Publicado: (2026)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
por: Muntaka, Siddique Abubakr, et al.
Publicado: (2026)
por: Muntaka, Siddique Abubakr, et al.
Publicado: (2026)
NetSmith: An Optimization Framework for Machine-Discovered Network Topologies
por: Green, Conor, et al.
Publicado: (2024)
por: Green, Conor, et al.
Publicado: (2024)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
por: Zhang, Zhekai, et al.
Publicado: (2020)
por: Zhang, Zhekai, et al.
Publicado: (2020)
Navigating the Landscape of Distributed File Systems: Architectures, Implementations, and Considerations
por: Pan, Xueting, et al.
Publicado: (2024)
por: Pan, Xueting, et al.
Publicado: (2024)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
por: Ma, Ke, et al.
Publicado: (2025)
por: Ma, Ke, et al.
Publicado: (2025)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
por: Yu, Yanpeng, et al.
Publicado: (2025)
por: Yu, Yanpeng, et al.
Publicado: (2025)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
por: Shen, Aofeng, et al.
Publicado: (2025)
por: Shen, Aofeng, et al.
Publicado: (2025)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
por: Li, Jiamin, et al.
Publicado: (2025)
por: Li, Jiamin, et al.
Publicado: (2025)
Evaluating Rapid Makespan Predictions for Heterogeneous Systems with Programmable Logic
por: Wilhelm, Martin, et al.
Publicado: (2025)
por: Wilhelm, Martin, et al.
Publicado: (2025)
Context-aware Simopt-Power: Using structural data with simulation metadata to optimise FPGA designs
por: Wadhwa, Eashan, et al.
Publicado: (2026)
por: Wadhwa, Eashan, et al.
Publicado: (2026)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
por: Yuksel, Ismail Emir, et al.
Publicado: (2023)
por: Yuksel, Ismail Emir, et al.
Publicado: (2023)
Tascade: Hardware Support for Atomic-free, Asynchronous and Efficient Reduction Trees
por: Orenes-Vera, Marcelo, et al.
Publicado: (2023)
por: Orenes-Vera, Marcelo, et al.
Publicado: (2023)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
por: Sharma, Harsh, et al.
Publicado: (2023)
por: Sharma, Harsh, et al.
Publicado: (2023)
xPUE: Extending Power Usage Effectiveness Metrics for Cloud Infrastructures
por: Fieni, Guillaume, et al.
Publicado: (2025)
por: Fieni, Guillaume, et al.
Publicado: (2025)
Kitsune: Enabling Dataflow Execution on GPUs
por: Davies, Michael, et al.
Publicado: (2025)
por: Davies, Michael, et al.
Publicado: (2025)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
por: Garg, Raveesh, et al.
Publicado: (2023)
por: Garg, Raveesh, et al.
Publicado: (2023)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
por: Elwasif, Wael, et al.
Publicado: (2022)
por: Elwasif, Wael, et al.
Publicado: (2022)
RapidOMS: FPGA-based Open Modification Spectral Library Searching with HD Computing
por: Pinge, Sumukh, et al.
Publicado: (2024)
por: Pinge, Sumukh, et al.
Publicado: (2024)
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
por: Cao, Yingqi, et al.
Publicado: (2024)
por: Cao, Yingqi, et al.
Publicado: (2024)
Ejemplares similares
-
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
por: Torres, L. A., et al.
Publicado: (2024) -
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
por: Colagrande, Luca, et al.
Publicado: (2026) -
Accelerating Triangle Counting with Real Processing-in-Memory Systems
por: Asquini, Lorenzo, et al.
Publicado: (2025) -
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
por: Zhang, Chen, et al.
Publicado: (2026) -
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
por: Kim, Hyeseong, et al.
Publicado: (2026)