Accelerating Transposed Convolutions on FPGA-based Edge Devices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Haris, Jude, Cano, José |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
von: Haris, Jude, et al.
Veröffentlicht: (2025)
von: Haris, Jude, et al.
Veröffentlicht: (2025)
Exploring Parallelism in FPGA-Based Accelerators for Machine Learning Applications
von: Centeno, Sed, et al.
Veröffentlicht: (2025)
von: Centeno, Sed, et al.
Veröffentlicht: (2025)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
von: Zhu, Yu, et al.
Veröffentlicht: (2025)
von: Zhu, Yu, et al.
Veröffentlicht: (2025)
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
von: Peccia, Federico Nicolas, et al.
Veröffentlicht: (2024)
von: Peccia, Federico Nicolas, et al.
Veröffentlicht: (2024)
The Feasibility of Implementing Large-Scale Transformers on Multi-FPGA Platforms
von: Gao, Yu, et al.
Veröffentlicht: (2024)
von: Gao, Yu, et al.
Veröffentlicht: (2024)
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
von: Russo, Enrico, et al.
Veröffentlicht: (2026)
von: Russo, Enrico, et al.
Veröffentlicht: (2026)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
Integrated Hardware Architecture and Device Placement Search
von: Wang, Irene, et al.
Veröffentlicht: (2024)
von: Wang, Irene, et al.
Veröffentlicht: (2024)
Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator Systems
von: Blanco, Francesco G., et al.
Veröffentlicht: (2024)
von: Blanco, Francesco G., et al.
Veröffentlicht: (2024)
MANOJAVAM: A Scalable, Unified FPGA Accelerator for Matrix Multiplication and Singular Value Decomposition in Principal Component Analysis
von: Ramasubramanian, Srivaths, et al.
Veröffentlicht: (2026)
von: Ramasubramanian, Srivaths, et al.
Veröffentlicht: (2026)
A Survey on Graph Neural Network Acceleration: Algorithms, Systems, and Customized Hardware
von: Zhang, Shichang, et al.
Veröffentlicht: (2023)
von: Zhang, Shichang, et al.
Veröffentlicht: (2023)
Accelerated Execution of Bayesian Neural Networks using a Single Probabilistic Forward Pass and Code Generation
von: Klein, Bernhard, et al.
Veröffentlicht: (2025)
von: Klein, Bernhard, et al.
Veröffentlicht: (2025)
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
von: Zhu, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2024)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
von: Hsia, Samuel, et al.
Veröffentlicht: (2023)
von: Hsia, Samuel, et al.
Veröffentlicht: (2023)
Analyzing a Two-Tier Disaggregated Memory Protection Scheme Based on Memory Replication
von: Volos, Haris, et al.
Veröffentlicht: (2025)
von: Volos, Haris, et al.
Veröffentlicht: (2025)
Towards Fair and Firm Real-Time Scheduling in DNN Multi-Tenant Multi-Accelerator Systems via Reinforcement Learning
von: Russo, Enrico, et al.
Veröffentlicht: (2024)
von: Russo, Enrico, et al.
Veröffentlicht: (2024)
RapidOMS: FPGA-based Open Modification Spectral Library Searching with HD Computing
von: Pinge, Sumukh, et al.
Veröffentlicht: (2024)
von: Pinge, Sumukh, et al.
Veröffentlicht: (2024)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
FPGA Innovation Research in the Netherlands: Present Landscape and Future Outlook
von: Alachiotis, Nikolaos, et al.
Veröffentlicht: (2025)
von: Alachiotis, Nikolaos, et al.
Veröffentlicht: (2025)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
EPOCH: Enabling Preemption Operation for Context Saving in Heterogeneous FPGA Systems
von: Malik, Arsalan Ali, et al.
Veröffentlicht: (2025)
von: Malik, Arsalan Ali, et al.
Veröffentlicht: (2025)
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
von: Pal, Shagnik, et al.
Veröffentlicht: (2025)
von: Pal, Shagnik, et al.
Veröffentlicht: (2025)
Enabling Time-Aware Priority Traffic Management over Distributed FPGA Nodes
von: Scionti, Alberto, et al.
Veröffentlicht: (2025)
von: Scionti, Alberto, et al.
Veröffentlicht: (2025)
Spira: Exploiting Voxel Data Structural Properties for Efficient Sparse Convolution in Point Cloud Networks
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
Context-aware Simopt-Power: Using structural data with simulation metadata to optimise FPGA designs
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2026)
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2026)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
von: Elwasif, Wael, et al.
Veröffentlicht: (2022)
von: Elwasif, Wael, et al.
Veröffentlicht: (2022)
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2024)
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2024)
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
von: Zou, An, et al.
Veröffentlicht: (2025)
von: Zou, An, et al.
Veröffentlicht: (2025)
FlashMoE: Fast Distributed MoE in a Single Kernel
von: Aimuyo, Osayamen Jonathan, et al.
Veröffentlicht: (2025)
von: Aimuyo, Osayamen Jonathan, et al.
Veröffentlicht: (2025)
ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
von: Liu, Yiqi, et al.
Veröffentlicht: (2025)
von: Liu, Yiqi, et al.
Veröffentlicht: (2025)
SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
von: Zhang, Hengrui, et al.
Veröffentlicht: (2025)
von: Zhang, Hengrui, et al.
Veröffentlicht: (2025)
FedUHD: Unsupervised Federated Learning using Hyperdimensional Computing
von: Lee, You Hak, et al.
Veröffentlicht: (2025)
von: Lee, You Hak, et al.
Veröffentlicht: (2025)
An End-to-End DNN Inference Framework for the SpiNNaker2 Neuromorphic MPSoC
von: Jobst, Matthias, et al.
Veröffentlicht: (2025)
von: Jobst, Matthias, et al.
Veröffentlicht: (2025)
JExplore: Design Space Exploration Tool for Nvidia Jetson Boards
von: Kutukcu, Basar, et al.
Veröffentlicht: (2025)
von: Kutukcu, Basar, et al.
Veröffentlicht: (2025)
Efficient, VRAM-Constrained xLM Inference on Clients
von: Ukarande, Aditya, et al.
Veröffentlicht: (2026)
von: Ukarande, Aditya, et al.
Veröffentlicht: (2026)
Hierarchical Resource Partitioning on Modern GPUs: A Reinforcement Learning Approach
von: Saroliya, Urvij, et al.
Veröffentlicht: (2024)
von: Saroliya, Urvij, et al.
Veröffentlicht: (2024)
Llumnix: Dynamic Scheduling for Large Language Model Serving
von: Sun, Biao, et al.
Veröffentlicht: (2024)
von: Sun, Biao, et al.
Veröffentlicht: (2024)
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
von: Ding, Jianru, et al.
Veröffentlicht: (2026)
von: Ding, Jianru, et al.
Veröffentlicht: (2026)
WWW: What, When, Where to Compute-in-Memory
von: Sharma, Tanvi, et al.
Veröffentlicht: (2023)
von: Sharma, Tanvi, et al.
Veröffentlicht: (2023)
MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from Microwatts to Megawatts for Sustainable AI
von: Tschand, Arya, et al.
Veröffentlicht: (2024)
von: Tschand, Arya, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
von: Haris, Jude, et al.
Veröffentlicht: (2025) -
Exploring Parallelism in FPGA-Based Accelerators for Machine Learning Applications
von: Centeno, Sed, et al.
Veröffentlicht: (2025) -
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
von: Zhu, Yu, et al.
Veröffentlicht: (2025) -
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
von: Peccia, Federico Nicolas, et al.
Veröffentlicht: (2024) -
The Feasibility of Implementing Large-Scale Transformers on Multi-FPGA Platforms
von: Gao, Yu, et al.
Veröffentlicht: (2024)