Data Transfer Optimizations for Host-CPU and Accelerators in AXI4MLIR
Fuente:
arXiv
Saved in:
| Main Authors: | Haris, Jude, Agostini, Nicolas Bohm, Tumeo, Antonino, Kaeli, David, Cano, José |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AXI4MLIR: User-Driven Automatic Host Code Generation for Custom AXI-Based Accelerators
by: Agostini, Nicolas Bohm, et al.
Published: (2023)
by: Agostini, Nicolas Bohm, et al.
Published: (2023)
HEC: Equivalence Verification Checking for Code Transformation via Equality Saturation
by: Yin, Jiaqi, et al.
Published: (2025)
by: Yin, Jiaqi, et al.
Published: (2025)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
by: Saha, Rappy, et al.
Published: (2026)
by: Saha, Rappy, et al.
Published: (2026)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
by: Zhang, Weichuang, et al.
Published: (2024)
by: Zhang, Weichuang, et al.
Published: (2024)
Practical Formal Verification for MLIR Programs
by: Tucker, Emily, et al.
Published: (2026)
by: Tucker, Emily, et al.
Published: (2026)
Building Bridges: Julia as an MLIR Frontend
by: Merckx, Jules
Published: (2025)
by: Merckx, Jules
Published: (2025)
MLIR-Forge: A Modular Framework for Language Smiths
by: Ates, Berke, et al.
Published: (2026)
by: Ates, Berke, et al.
Published: (2026)
Nice to Meet You: Synthesizing Practical MLIR Abstract Transformers
by: Peng, Xuanyu, et al.
Published: (2025)
by: Peng, Xuanyu, et al.
Published: (2025)
WAMI: Compilation to WebAssembly through MLIR without Losing Abstraction
by: Kang, Byeongjee, et al.
Published: (2025)
by: Kang, Byeongjee, et al.
Published: (2025)
An MLIR pipeline for offloading Fortran to FPGAs via OpenMP
by: Rodriguez-Canal, Gabriel, et al.
Published: (2025)
by: Rodriguez-Canal, Gabriel, et al.
Published: (2025)
The MLIR Transform Dialect. Your compiler is more powerful than you think
by: Lücke, Martin Paul, et al.
Published: (2024)
by: Lücke, Martin Paul, et al.
Published: (2024)
Accelerating Transposed Convolutions on FPGA-based Edge Devices
by: Haris, Jude, et al.
Published: (2025)
by: Haris, Jude, et al.
Published: (2025)
An MLIR Lowering Pipeline for Stencils at Wafer-Scale
by: Stawinoga, Nicolai, et al.
Published: (2026)
by: Stawinoga, Nicolai, et al.
Published: (2026)
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
by: Haris, Jude, et al.
Published: (2025)
by: Haris, Jude, et al.
Published: (2025)
Demonstrating a Future for MLIR-native DSL Compilers on a NumPy-like Example
by: Friebel, Karl F. A., et al.
Published: (2026)
by: Friebel, Karl F. A., et al.
Published: (2026)
Analyzing Latency Hiding and Parallelism in an MLIR-based AI Kernel Compiler
by: Absar, Javed, et al.
Published: (2026)
by: Absar, Javed, et al.
Published: (2026)
Fully integrating the Flang Fortran compiler with standard MLIR
by: Brown, Nick
Published: (2024)
by: Brown, Nick
Published: (2024)
Hexagon-MLIR: An AI Compilation Stack For Qualcomm's Neural Processing Units (NPUs)
by: Absar, Mohammed Javed, et al.
Published: (2026)
by: Absar, Mohammed Javed, et al.
Published: (2026)
MLIR-Smith: A Novel Random Program Generator for Evaluating Compiler Pipelines
by: Ates, Berke, et al.
Published: (2026)
by: Ates, Berke, et al.
Published: (2026)
Is It a Good Idea to Build an HLS Tool on Top of MLIR? Experience from Building the Dynamatic HLS Compiler
by: Xu, Jiahui, et al.
Published: (2026)
by: Xu, Jiahui, et al.
Published: (2026)
Hardware.jl - An MLIR-based Julia HLS Flow (Work in Progress)
by: Short, Benedict, et al.
Published: (2025)
by: Short, Benedict, et al.
Published: (2025)
Accelerating PoT Quantization on Edge Devices
by: Saha, Rappy, et al.
Published: (2024)
by: Saha, Rappy, et al.
Published: (2024)
Towards a high-performance AI compiler with upstream MLIR
by: Golin, Renato, et al.
Published: (2024)
by: Golin, Renato, et al.
Published: (2024)
DSP-MLIR: A MLIR Dialect for Digital Signal Processing
by: Kumar, Abhinav, et al.
Published: (2024)
by: Kumar, Abhinav, et al.
Published: (2024)
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
by: Shivdikar, Kaustubh, et al.
Published: (2024)
by: Shivdikar, Kaustubh, et al.
Published: (2024)
Designing Efficient LLM Accelerators for Edge Devices
by: Haris, Jude, et al.
Published: (2024)
by: Haris, Jude, et al.
Published: (2024)
Mojo: MLIR-Based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem
by: Godoy, William F., et al.
Published: (2025)
by: Godoy, William F., et al.
Published: (2025)
LLM-Driven Design Space Exploration of FPGA-based Accelerators
by: Sharma, Vinamra, et al.
Published: (2026)
by: Sharma, Vinamra, et al.
Published: (2026)
CPU-less parallel execution of lambda calculus in digital logic
by: Fitchett, Harry, et al.
Published: (2026)
by: Fitchett, Harry, et al.
Published: (2026)
A Data-driven Analysis of Code Optimizations
by: Hakimi, Yacine, et al.
Published: (2025)
by: Hakimi, Yacine, et al.
Published: (2025)
Expression Acceleration: Seamless Parallelization of Typed High-Level Languages
by: Hummelgren, Lars, et al.
Published: (2022)
by: Hummelgren, Lars, et al.
Published: (2022)
Massimult: A Novel Parallel CPU Architecture Based on Combinator Reduction
by: Nicklisch-Franken, Jurgen, et al.
Published: (2024)
by: Nicklisch-Franken, Jurgen, et al.
Published: (2024)
An Optimizing Just-In-Time Compiler for Rotor
by: Trindade, João H., et al.
Published: (2024)
by: Trindade, João H., et al.
Published: (2024)
Cage: Hardware-Accelerated Safe WebAssembly
by: Fink, Martin, et al.
Published: (2024)
by: Fink, Martin, et al.
Published: (2024)
Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving
by: Zhao, Juntao, et al.
Published: (2025)
by: Zhao, Juntao, et al.
Published: (2025)
HPVM-HDC: A Heterogeneous Programming System for Accelerating Hyperdimensional Computing
by: Arbore, Russel, et al.
Published: (2024)
by: Arbore, Russel, et al.
Published: (2024)
Filling the Gaps of Polarity: Implementing Dependent Data and Codata Types with Implicit Arguments
by: Liesnikov, Bohdan, et al.
Published: (2025)
by: Liesnikov, Bohdan, et al.
Published: (2025)
Pushing Tensor Accelerators Beyond MatMul in a User-Schedulable Language
by: Zhang, Yihong, et al.
Published: (2025)
by: Zhang, Yihong, et al.
Published: (2025)
Exploiting Multiple Abstract Call Patterns for Optimizing Run-Time Checks
by: Ferreiro, Daniela, et al.
Published: (2026)
by: Ferreiro, Daniela, et al.
Published: (2026)
Automated Profile-Guided Replacement of Data Structures to Reduce Memory Allocation
by: Makor, Lukas, et al.
Published: (2025)
by: Makor, Lukas, et al.
Published: (2025)
Similar Items
-
AXI4MLIR: User-Driven Automatic Host Code Generation for Custom AXI-Based Accelerators
by: Agostini, Nicolas Bohm, et al.
Published: (2023) -
HEC: Equivalence Verification Checking for Code Transformation via Equality Saturation
by: Yin, Jiaqi, et al.
Published: (2025) -
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
by: Saha, Rappy, et al.
Published: (2026) -
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
by: Zhang, Weichuang, et al.
Published: (2024) -
Practical Formal Verification for MLIR Programs
by: Tucker, Emily, et al.
Published: (2026)