GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
Fuente:
arXiv
Saved in:
| Main Authors: | Ramachandran, Anvitha, Parikh, Dhruv, Prasanna, Viktor |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
by: Ramachandran, Anvitha, et al.
Published: (2025)
by: Ramachandran, Anvitha, et al.
Published: (2025)
Accelerating ViT Inference on FPGA through Static and Dynamic Pruning
by: Parikh, Dhruv, et al.
Published: (2024)
by: Parikh, Dhruv, et al.
Published: (2024)
VTR: An Optimized Vision Transformer for SAR ATR Acceleration on FPGA
by: Wickramasinghe, Sachini, et al.
Published: (2024)
by: Wickramasinghe, Sachini, et al.
Published: (2024)
GCV-Turbo: End-to-end Acceleration of GNN-based Computer Vision Tasks on FPGA
by: Zhang, Bingyi, et al.
Published: (2024)
by: Zhang, Bingyi, et al.
Published: (2024)
ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
by: Zhang, Zuoning, et al.
Published: (2024)
by: Zhang, Zuoning, et al.
Published: (2024)
Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
by: Gupta, Neelesh, et al.
Published: (2025)
by: Gupta, Neelesh, et al.
Published: (2025)
A Unified CPU-GPU Protocol for GNN Training
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
PipeDiT: Accelerating Diffusion Transformers in Video Generation with Task Pipelining and Model Decoupling
by: Wang, Sijie, et al.
Published: (2025)
by: Wang, Sijie, et al.
Published: (2025)
Partially Conditioned Patch Parallelism for Accelerated Diffusion Model Inference
by: Zhang, XiuYu, et al.
Published: (2024)
by: Zhang, XiuYu, et al.
Published: (2024)
DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
by: Tan, Xin, et al.
Published: (2025)
by: Tan, Xin, et al.
Published: (2025)
Personalized Federated Fine-Tuning of Vision Foundation Models for Healthcare
by: Tupper, Adam, et al.
Published: (2025)
by: Tupper, Adam, et al.
Published: (2025)
COALA: A Practical and Vision-Centric Federated Learning Platform
by: Zhuang, Weiming, et al.
Published: (2024)
by: Zhuang, Weiming, et al.
Published: (2024)
Minuet: Accelerating 3D Sparse Convolutions on GPUs
by: Yang, Jiacheng, et al.
Published: (2023)
by: Yang, Jiacheng, et al.
Published: (2023)
Ask the Expert: Collaborative Inference for Vision Transformers with Near-Edge Accelerators
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Can Graphs Help Vision SSMs See Better?
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Incremental GNN Embedding Computation on Streaming Graphs
by: Wang, Qiange, et al.
Published: (2026)
by: Wang, Qiange, et al.
Published: (2026)
Code Generation for a Variety of Accelerators for a Graph DSL
by: Kumar, Ashwina, et al.
Published: (2024)
by: Kumar, Ashwina, et al.
Published: (2024)
SOLANET: Distributed Neighbor Graph Construction on GPU-Accelerated Systems
by: Iwabuchi, Keita, et al.
Published: (2026)
by: Iwabuchi, Keita, et al.
Published: (2026)
CondenseGraph: Communication-Efficient Distributed GNN Training via On-the-Fly Graph Condensation
by: Zhang, Zizhao, et al.
Published: (2026)
by: Zhang, Zizhao, et al.
Published: (2026)
Covariance-Guided Resource Adaptive Learning for Efficient Edge Inference
by: Nabhaan, Ahmad N. L., et al.
Published: (2026)
by: Nabhaan, Ahmad N. L., et al.
Published: (2026)
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
by: Chen, Yukang, et al.
Published: (2026)
by: Chen, Yukang, et al.
Published: (2026)
DSFedMed: Dual-Scale Federated Medical Image Segmentation via Mutual Distillation Between Foundation and Lightweight Models
by: Zhang, Hanwen, et al.
Published: (2026)
by: Zhang, Hanwen, et al.
Published: (2026)
SC-MII: Infrastructure LiDAR-based 3D Object Detection on Edge Devices for Split Computing with Multiple Intermediate Outputs Integration
by: Noguchi, Taisuke, et al.
Published: (2026)
by: Noguchi, Taisuke, et al.
Published: (2026)
SwiftFusion: Scalable Sequence Parallelism for Distributed Inference of Diffusion Transformers on GPUs
by: Yang, Jiacheng, et al.
Published: (2026)
by: Yang, Jiacheng, et al.
Published: (2026)
Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction
by: Zhang, Jinxiao, et al.
Published: (2026)
by: Zhang, Jinxiao, et al.
Published: (2026)
Improving Zero-shot ADL Recognition with Large Language Models through Event-based Context and Confidence
by: Fiori, Michele, et al.
Published: (2026)
by: Fiori, Michele, et al.
Published: (2026)
Federated Unsupervised Visual Representation Learning via Exploiting General Content and Personal Style
by: Yang, Yuewei, et al.
Published: (2022)
by: Yang, Yuewei, et al.
Published: (2022)
A Scalable Distributed Framework for Multimodal GigaVoxel Image Registration
by: Jena, Rohit, et al.
Published: (2025)
by: Jena, Rohit, et al.
Published: (2025)
SlimEdge: Performance and Device Aware Distributed DNN Deployment on Resource-Constrained Edge Hardware
by: Kumar, Mahadev Sunil, et al.
Published: (2025)
by: Kumar, Mahadev Sunil, et al.
Published: (2025)
Benchmarking Federated Learning Frameworks for Medical Imaging Deployment: A Comparative Study of NVIDIA FLARE, Flower, and Owkin Substra
by: Gupta, Riya, et al.
Published: (2025)
by: Gupta, Riya, et al.
Published: (2025)
Investigation of Federated Learning Algorithms for Retinal Optical Coherence Tomography Image Classification with Statistical Heterogeneity
by: Amgain, Sanskar, et al.
Published: (2024)
by: Amgain, Sanskar, et al.
Published: (2024)
Argus: Quality-Aware High-Throughput Text-to-Image Inference Serving System
by: Agarwal, Shubham, et al.
Published: (2025)
by: Agarwal, Shubham, et al.
Published: (2025)
Multi-modal video data-pipelines for machine learning with minimal human supervision
by: Pîrvu, Mihai-Cristian, et al.
Published: (2025)
by: Pîrvu, Mihai-Cristian, et al.
Published: (2025)
Radiant: Large-scale 3D Gaussian Rendering based on Hierarchical Framework
by: Peng, Haosong, et al.
Published: (2024)
by: Peng, Haosong, et al.
Published: (2024)
Cross-architecture universal feature coding via distribution alignment
by: Gao, Changsheng, et al.
Published: (2025)
by: Gao, Changsheng, et al.
Published: (2025)
Similar Items
-
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
by: Ramachandran, Anvitha, et al.
Published: (2025) -
Accelerating ViT Inference on FPGA through Static and Dynamic Pruning
by: Parikh, Dhruv, et al.
Published: (2024) -
VTR: An Optimized Vision Transformer for SAR ATR Acceleration on FPGA
by: Wickramasinghe, Sachini, et al.
Published: (2024) -
GCV-Turbo: End-to-end Acceleration of GNN-based Computer Vision Tasks on FPGA
by: Zhang, Bingyi, et al.
Published: (2024) -
ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning
by: Parikh, Dhruv, et al.
Published: (2025)