Accelerating Mini-batch HGNN Training by Reducing CUDA Kernels
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wu, Meng, Qiu, Jingkai, Yan, Mingyu, Li, Wenming, Zhang, Yang, Zhang, Zhimin, Ye, Xiaochun, Fan, Dongrui |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SiHGNN: Leveraging Properties of Semantic Graphs for Efficient HGNN Acceleration
par: Xue, Runzhen, et autres
Publié: (2024)
par: Xue, Runzhen, et autres
Publié: (2024)
ADE-HGNN: Accelerating HGNNs through Attention Disparity Exploitation
par: Han, Dengke, et autres
Publié: (2024)
par: Han, Dengke, et autres
Publié: (2024)
HiHGNN: Accelerating HGNNs through Parallelism and Data Reusability Exploitation
par: Xue, Runzhen, et autres
Publié: (2023)
par: Xue, Runzhen, et autres
Publié: (2023)
TLV-HGNN: Thinking Like a Vertex for Memory-efficient HGNN Inference
par: Han, Dengke, et autres
Publié: (2025)
par: Han, Dengke, et autres
Publié: (2025)
GDR-HGNN: A Heterogeneous Graph Neural Networks Accelerator Frontend with Graph Decoupling and Recoupling
par: Xue, Runzhen, et autres
Publié: (2024)
par: Xue, Runzhen, et autres
Publié: (2024)
Characterizing and Understanding HGNN Training on GPUs
par: Han, Dengke, et autres
Publié: (2024)
par: Han, Dengke, et autres
Publié: (2024)
Survey on Characterizing and Understanding GNNs from a Computer Architecture Perspective
par: Wu, Meng, et autres
Publié: (2024)
par: Wu, Meng, et autres
Publié: (2024)
Accelerating GNN Training through Locality-aware Dropout and Merge
par: Sun, Gongjian, et autres
Publié: (2025)
par: Sun, Gongjian, et autres
Publié: (2025)
StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal Transformer
par: Qin, Shantian, et autres
Publié: (2025)
par: Qin, Shantian, et autres
Publié: (2025)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
par: Wu, Haibin, et autres
Publié: (2024)
par: Wu, Haibin, et autres
Publié: (2024)
MetaDSE: A Few-shot Meta-learning Framework for Cross-workload CPU Design Space Exploration
par: Xue, Runzhen, et autres
Publié: (2025)
par: Xue, Runzhen, et autres
Publié: (2025)
Multi-objective Optimization in CPU Design Space Exploration: Attention is All You Need
par: Xue, Runzhen, et autres
Publié: (2024)
par: Xue, Runzhen, et autres
Publié: (2024)
AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices
par: Zirui, Ma, et autres
Publié: (2026)
par: Zirui, Ma, et autres
Publié: (2026)
A Systematic Characterization of LLM Inference on GPUs
par: Wang, Haonan, et autres
Publié: (2025)
par: Wang, Haonan, et autres
Publié: (2025)
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
par: Lin, Chenqi, et autres
Publié: (2025)
par: Lin, Chenqi, et autres
Publié: (2025)
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
par: Lu, Jinming, et autres
Publié: (2025)
par: Lu, Jinming, et autres
Publié: (2025)
Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators
par: Xie, Tong, et autres
Publié: (2026)
par: Xie, Tong, et autres
Publié: (2026)
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
par: Li, Meng, et autres
Publié: (2026)
par: Li, Meng, et autres
Publié: (2026)
Efficient yet Accurate End-to-End SC Accelerator Design
par: Li, Meng, et autres
Publié: (2024)
par: Li, Meng, et autres
Publié: (2024)
Hyft: A Reconfigurable Softmax Accelerator with Hybrid Numeric Format for both Training and Inference
par: Xia, Tianhua, et autres
Publié: (2023)
par: Xia, Tianhua, et autres
Publié: (2023)
Efficient Kernel Mapping and Comprehensive System Evaluation of LLM Acceleration on a CGLA
par: Ando, Takuto, et autres
Publié: (2025)
par: Ando, Takuto, et autres
Publié: (2025)
SoMa: Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN Accelerators
par: Cai, Jingwei, et autres
Publié: (2025)
par: Cai, Jingwei, et autres
Publié: (2025)
RAS: A Bit-Exact rANS Accelerator For High-Performance Neural Lossless Compression
par: Qin, Yuchao, et autres
Publié: (2025)
par: Qin, Yuchao, et autres
Publié: (2025)
Squire: A General-Purpose Accelerator to Exploit Fine-Grain Parallelism on Dependency-Bound Kernels
par: Langarita, Rubén, et autres
Publié: (2025)
par: Langarita, Rubén, et autres
Publié: (2025)
Composing Mini Oscilloscope on Embedded Systems
par: Romero, Brennan, et autres
Publié: (2025)
par: Romero, Brennan, et autres
Publié: (2025)
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
par: Huang, Sixiao, et autres
Publié: (2025)
par: Huang, Sixiao, et autres
Publié: (2025)
Sparsity-Aware Streaming SNN Accelerator with Output-Channel Dataflow for Automatic Modulation Classification
par: Yang, Kuilian, et autres
Publié: (2026)
par: Yang, Kuilian, et autres
Publié: (2026)
FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
par: Han, Meng, et autres
Publié: (2023)
par: Han, Meng, et autres
Publié: (2023)
AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
par: Cheng, Feng, et autres
Publié: (2025)
par: Cheng, Feng, et autres
Publié: (2025)
Large Processor Chip Model
par: Chang, Kaiyan, et autres
Publié: (2025)
par: Chang, Kaiyan, et autres
Publié: (2025)
GOMA: Geometrically Optimal Mapping via Analytical Modeling for Spatial Accelerators
par: Yang, Wulve, et autres
Publié: (2026)
par: Yang, Wulve, et autres
Publié: (2026)
Trinity: A General Purpose FHE Accelerator
par: Deng, Xianglong, et autres
Publié: (2024)
par: Deng, Xianglong, et autres
Publié: (2024)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
par: Ye, Hanchen, et autres
Publié: (2025)
par: Ye, Hanchen, et autres
Publié: (2025)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
par: Huang, Wei-Hsing, et autres
Publié: (2024)
par: Huang, Wei-Hsing, et autres
Publié: (2024)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale Systems
par: Huang, Wei-Hsing, et autres
Publié: (2025)
par: Huang, Wei-Hsing, et autres
Publié: (2025)
MiniFloat-NN and ExSdotp: An ISA Extension and a Modular Open Hardware Unit for Low-Precision Training on RISC-V cores
par: Bertaccini, Luca, et autres
Publié: (2022)
par: Bertaccini, Luca, et autres
Publié: (2022)
An FPGA-Based Accelerator Enabling Efficient Support for CNNs with Arbitrary Kernel Sizes
par: Wang, Miaoxin, et autres
Publié: (2024)
par: Wang, Miaoxin, et autres
Publié: (2024)
Trimma: Trimming Metadata Storage and Latency for Hybrid Memory Systems
par: Li, Yiwei, et autres
Publié: (2024)
par: Li, Yiwei, et autres
Publié: (2024)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
par: Li, Zhengke, et autres
Publié: (2025)
par: Li, Zhengke, et autres
Publié: (2025)
SigDLA: A Deep Learning Accelerator Extension for Signal Processing
par: Fu, Fangfa, et autres
Publié: (2024)
par: Fu, Fangfa, et autres
Publié: (2024)
Documents similaires
-
SiHGNN: Leveraging Properties of Semantic Graphs for Efficient HGNN Acceleration
par: Xue, Runzhen, et autres
Publié: (2024) -
ADE-HGNN: Accelerating HGNNs through Attention Disparity Exploitation
par: Han, Dengke, et autres
Publié: (2024) -
HiHGNN: Accelerating HGNNs through Parallelism and Data Reusability Exploitation
par: Xue, Runzhen, et autres
Publié: (2023) -
TLV-HGNN: Thinking Like a Vertex for Memory-efficient HGNN Inference
par: Han, Dengke, et autres
Publié: (2025) -
GDR-HGNN: A Heterogeneous Graph Neural Networks Accelerator Frontend with Graph Decoupling and Recoupling
par: Xue, Runzhen, et autres
Publié: (2024)