Code Generation for a Variety of Accelerators for a Graph DSL
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kumar, Ashwina, Krishna, M. Venkata, Bartakke, Prasanna, Kumar, Rahul, M, Rajesh Pandian, Behera, Nibedita, Nasre, Rupesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generating Dynamic Graph Algorithms for Multiple Backends for a Graph DSL
von: Behera, Nibedita, et al.
Veröffentlicht: (2025)
von: Behera, Nibedita, et al.
Veröffentlicht: (2025)
Scalable Maxflow Processing for Dynamic Graphs
von: Kannappan, Shruthi, et al.
Veröffentlicht: (2025)
von: Kannappan, Shruthi, et al.
Veröffentlicht: (2025)
Efficient Dynamic MaxFlow Computation on GPUs
von: Kannappan, Shruthi, et al.
Veröffentlicht: (2025)
von: Kannappan, Shruthi, et al.
Veröffentlicht: (2025)
StarDist: A Code Generator for Distributed Graph Algorithms
von: Nandy, Barenya Kumar, et al.
Veröffentlicht: (2025)
von: Nandy, Barenya Kumar, et al.
Veröffentlicht: (2025)
Morphling: Fast, Fused, and Flexible GNN Training at Scale
von: Anubhab, et al.
Veröffentlicht: (2025)
von: Anubhab, et al.
Veröffentlicht: (2025)
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
von: Ramachandran, Anvitha, et al.
Veröffentlicht: (2025)
von: Ramachandran, Anvitha, et al.
Veröffentlicht: (2025)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
Efficient Task Graph Scheduling for Parallel QR Factorization in SLSQP
von: Chatterjee, Soumyajit, et al.
Veröffentlicht: (2025)
von: Chatterjee, Soumyajit, et al.
Veröffentlicht: (2025)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
von: Kumar, Abhishek Vijaya, et al.
Veröffentlicht: (2024)
von: Kumar, Abhishek Vijaya, et al.
Veröffentlicht: (2024)
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
von: Ramachandran, Anvitha, et al.
Veröffentlicht: (2026)
von: Ramachandran, Anvitha, et al.
Veröffentlicht: (2026)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
Accelerating Loading WebGraphs in ParaGrapher
von: Esfahani, Mohsen Koohi
Veröffentlicht: (2025)
von: Esfahani, Mohsen Koohi
Veröffentlicht: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
Efficient Accelerated Graph Edit Distance Computation on GPU
von: Dabah, Adel, et al.
Veröffentlicht: (2026)
von: Dabah, Adel, et al.
Veröffentlicht: (2026)
Parallel Online Directed Acyclic Graph Exploration for Atlasing Soft-Matter Assembly Configuration Spaces
von: Prabhu, Rahul, et al.
Veröffentlicht: (2024)
von: Prabhu, Rahul, et al.
Veröffentlicht: (2024)
SOLANET: Distributed Neighbor Graph Construction on GPU-Accelerated Systems
von: Iwabuchi, Keita, et al.
Veröffentlicht: (2026)
von: Iwabuchi, Keita, et al.
Veröffentlicht: (2026)
Pagoda: An Energy and Time Roofline Study for DNN Workloads on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
Graph-based Gossiping for Communication Efficiency in Decentralized Federated Learning
von: Nguyen, Huong, et al.
Veröffentlicht: (2025)
von: Nguyen, Huong, et al.
Veröffentlicht: (2025)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
Transforming Agriculture: Exploring Diverse Practices and Technological Innovations
von: Kumar, Ramakant
Veröffentlicht: (2024)
von: Kumar, Ramakant
Veröffentlicht: (2024)
UniPar: A Unified LLM-Based Framework for Parallel and Accelerated Code Translation in HPC
von: Bitan, Tomer, et al.
Veröffentlicht: (2025)
von: Bitan, Tomer, et al.
Veröffentlicht: (2025)
City-Scale Visibility Graph Analysis via GPU-Accelerated HyperBall
von: Hodge, Alex, et al.
Veröffentlicht: (2026)
von: Hodge, Alex, et al.
Veröffentlicht: (2026)
A Unified CPU-GPU Protocol for GNN Training
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
Zen-Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs
von: Deshmukh, Aadesh, et al.
Veröffentlicht: (2025)
von: Deshmukh, Aadesh, et al.
Veröffentlicht: (2025)
GPU Accelerated Sparse Cholesky Factorization
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
von: Zhang, Zuoning, et al.
Veröffentlicht: (2024)
von: Zhang, Zuoning, et al.
Veröffentlicht: (2024)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
An Efficient Approach for Energy Conservation in Cloud Computing Environment
von: Pande, Sohan Kumar, et al.
Veröffentlicht: (2025)
von: Pande, Sohan Kumar, et al.
Veröffentlicht: (2025)
Combining Performance and Productivity: Accelerating the Network Sensing Graph Challenge with GPUs and Commodity Data Science Software
von: Samsi, Siddharth, et al.
Veröffentlicht: (2025)
von: Samsi, Siddharth, et al.
Veröffentlicht: (2025)
AnTi-MiCS: Analytical Framework for Bounding Time in Embedded Mixed-Criticality Systems
von: Ranjbar, Behnaz, et al.
Veröffentlicht: (2026)
von: Ranjbar, Behnaz, et al.
Veröffentlicht: (2026)
QEIL v2: Heterogeneous Computing for Edge Intelligence via Roofline-Derived Pareto-Optimal Energy Modeling and Multi-Objective Orchestration
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
von: Behera, Adarsh Prasad, et al.
Veröffentlicht: (2024)
von: Behera, Adarsh Prasad, et al.
Veröffentlicht: (2024)
Analytical Performance Estimation during Code Generation on Modern GPUs
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
von: Ernst, Dominik, et al.
Veröffentlicht: (2022)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
von: Zhao, Zhan, et al.
Veröffentlicht: (2026)
von: Zhao, Zhan, et al.
Veröffentlicht: (2026)
Fog Device-as-a-Service (FDaaS): A Framework for Service Deployment in Public Fog Environments
von: Battula, Sudheer Kumar, et al.
Veröffentlicht: (2023)
von: Battula, Sudheer Kumar, et al.
Veröffentlicht: (2023)
Computing Tree Structures in Anonymous Graphs via Mobile Agents
von: Chand, Prabhat Kumar, et al.
Veröffentlicht: (2025)
von: Chand, Prabhat Kumar, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generating Dynamic Graph Algorithms for Multiple Backends for a Graph DSL
von: Behera, Nibedita, et al.
Veröffentlicht: (2025) -
Scalable Maxflow Processing for Dynamic Graphs
von: Kannappan, Shruthi, et al.
Veröffentlicht: (2025) -
Efficient Dynamic MaxFlow Computation on GPUs
von: Kannappan, Shruthi, et al.
Veröffentlicht: (2025) -
StarDist: A Code Generator for Distributed Graph Algorithms
von: Nandy, Barenya Kumar, et al.
Veröffentlicht: (2025) -
Morphling: Fast, Fused, and Flexible GNN Training at Scale
von: Anubhab, et al.
Veröffentlicht: (2025)