DF-GNN: Dynamic Fusion Framework for Attention Graph Neural Networks on GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jiahui, Cai, Zhenkun, Chen, Zhiyong, Wang, Minjie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Graph Knowledge Distillation from GNNs to Kolmogorov--Arnold Networks via Self-Attention Dynamic Sampling
by: Cui, Can, et al.
Published: (2025)
by: Cui, Can, et al.
Published: (2025)
Benchmarking GPUs on SVBRDF Extractor Model
by: Kandel, Narayan, et al.
Published: (2023)
by: Kandel, Narayan, et al.
Published: (2023)
AutoSAGE: Input-Aware CUDA Scheduling for Sparse GNN Aggregation (SpMM/SDDMM) and CSR Attention
by: Stankovic, Aleksandar
Published: (2025)
by: Stankovic, Aleksandar
Published: (2025)
msf-CNN: Patch-based Multi-Stage Fusion with Convolutional Neural Networks for TinyML
by: Huang, Zhaolan, et al.
Published: (2025)
by: Huang, Zhaolan, et al.
Published: (2025)
Can Graph Reordering Speed Up Graph Neural Network Training? An Experimental Study
by: Merkel, Nikolai, et al.
Published: (2024)
by: Merkel, Nikolai, et al.
Published: (2024)
Prefetching Cache Optimization Using Graph Neural Networks: A Modular Framework and Conceptual Analysis
by: Qowy, F. I.
Published: (2025)
by: Qowy, F. I.
Published: (2025)
Exchangeability in Neural Network and its Application to Dynamic Pruning
by: Pu, et al.
Published: (2025)
by: Pu, et al.
Published: (2025)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025)
by: Chen, Feiyang, et al.
Published: (2025)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
by: Choudhary, Mansi, et al.
Published: (2025)
by: Choudhary, Mansi, et al.
Published: (2025)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
by: Georganas, Evangelos, et al.
Published: (2025)
by: Georganas, Evangelos, et al.
Published: (2025)
A Structure-Aware Framework for Learning Device Placements on Computation Graphs
by: Duan, Shukai, et al.
Published: (2024)
by: Duan, Shukai, et al.
Published: (2024)
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
by: Sarkar, Aishwarya, et al.
Published: (2024)
by: Sarkar, Aishwarya, et al.
Published: (2024)
SAfEPaTh: A System-Level Approach for Efficient Power and Thermal Estimation of Convolutional Neural Network Accelerator
by: Chen, Yukai, et al.
Published: (2024)
by: Chen, Yukai, et al.
Published: (2024)
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
by: Wu, Dan, et al.
Published: (2023)
by: Wu, Dan, et al.
Published: (2023)
Characterizing and Understanding HGNN Training on GPUs
by: Han, Dengke, et al.
Published: (2024)
by: Han, Dengke, et al.
Published: (2024)
BitLogic: Training Framework for Gradient-Based FPGA-Native Neural Networks
by: Bührer, Simon, et al.
Published: (2026)
by: Bührer, Simon, et al.
Published: (2026)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
by: You, Bozhi, et al.
Published: (2025)
by: You, Bozhi, et al.
Published: (2025)
CPINN-ABPI: Physics-Informed Neural Networks for Accurate Power Estimation in MPSoCs
by: Elshamy, Mohamed R., et al.
Published: (2025)
by: Elshamy, Mohamed R., et al.
Published: (2025)
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
by: Knoop, Jonathan, et al.
Published: (2026)
by: Knoop, Jonathan, et al.
Published: (2026)
Profiling LoRA/QLoRA Fine-Tuning Efficiency on Consumer GPUs: An RTX 4060 Case Study
by: Avinash, MSR
Published: (2025)
by: Avinash, MSR
Published: (2025)
Accelerating AI Performance using Anderson Extrapolation on GPUs
by: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Published: (2024)
by: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Published: (2024)
Estudio de la eficiencia en la escalabilidad de GPUs para el entrenamiento de Inteligencia Artificial
by: Cortes, David, et al.
Published: (2025)
by: Cortes, David, et al.
Published: (2025)
Distributed Matrix-Based Sampling for Graph Neural Network Training
by: Tripathy, Alok, et al.
Published: (2023)
by: Tripathy, Alok, et al.
Published: (2023)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
by: Zhao, Yushang, et al.
Published: (2025)
by: Zhao, Yushang, et al.
Published: (2025)
DiskGNN: Bridging I/O Efficiency and Model Accuracy for Out-of-Core GNN Training
by: Liu, Renjie, et al.
Published: (2024)
by: Liu, Renjie, et al.
Published: (2024)
MuseGNN: Forming Scalable, Convergent GNN Layers that Minimize a Sampling-Based Energy
by: Jiang, Haitian, et al.
Published: (2023)
by: Jiang, Haitian, et al.
Published: (2023)
The Power of Training: How Different Neural Network Setups Influence the Energy Demand
by: Geißler, Daniel, et al.
Published: (2024)
by: Geißler, Daniel, et al.
Published: (2024)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
by: Li, Jianhui, et al.
Published: (2023)
by: Li, Jianhui, et al.
Published: (2023)
MQ-GNN: A Multi-Queue Pipelined Architecture for Scalable and Efficient GNN Training
by: Ullah, Irfan, et al.
Published: (2026)
by: Ullah, Irfan, et al.
Published: (2026)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
by: Dong, Juechu, et al.
Published: (2024)
by: Dong, Juechu, et al.
Published: (2024)
The Hidden Power of Pure 16-bit Floating-Point Neural Networks
by: Yun, Juyoung, et al.
Published: (2023)
by: Yun, Juyoung, et al.
Published: (2023)
Towards an Integrated Performance Framework for Fire Science and Management Workflows
by: Ahmed, H., et al.
Published: (2024)
by: Ahmed, H., et al.
Published: (2024)
Light Differentiable Logic Gate Networks
by: Rüttgers, Lukas, et al.
Published: (2025)
by: Rüttgers, Lukas, et al.
Published: (2025)
LiveTune: Dynamic Parameter Tuning for Feedback-Driven Optimization
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
Quantitative Analysis of Deeply Quantized Tiny Neural Networks Robust to Adversarial Attacks
by: Zakariyya, Idris, et al.
Published: (2025)
by: Zakariyya, Idris, et al.
Published: (2025)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
Application Research On Real-Time Perception Of Device Performance Status
by: Wang, Zhe, et al.
Published: (2024)
by: Wang, Zhe, et al.
Published: (2024)
SynthEval: A Framework for Detailed Utility and Privacy Evaluation of Tabular Synthetic Data
by: Lautrup, Anton Danholt, et al.
Published: (2024)
by: Lautrup, Anton Danholt, et al.
Published: (2024)
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
by: Ding, Shiwei, et al.
Published: (2025)
by: Ding, Shiwei, et al.
Published: (2025)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
by: Ibrahim, Muhammad Sohail, et al.
Published: (2024)
by: Ibrahim, Muhammad Sohail, et al.
Published: (2024)
Similar Items
-
Efficient Graph Knowledge Distillation from GNNs to Kolmogorov--Arnold Networks via Self-Attention Dynamic Sampling
by: Cui, Can, et al.
Published: (2025) -
Benchmarking GPUs on SVBRDF Extractor Model
by: Kandel, Narayan, et al.
Published: (2023) -
AutoSAGE: Input-Aware CUDA Scheduling for Sparse GNN Aggregation (SpMM/SDDMM) and CSR Attention
by: Stankovic, Aleksandar
Published: (2025) -
msf-CNN: Patch-based Multi-Stage Fusion with Convolutional Neural Networks for TinyML
by: Huang, Zhaolan, et al.
Published: (2025) -
Can Graph Reordering Speed Up Graph Neural Network Training? An Experimental Study
by: Merkel, Nikolai, et al.
Published: (2024)