Gensor: A Graph-based Construction Tensor Compilation Method for Deep Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Hangda, Diao, Boyu, Yang, Yu, Chen, Wenxin, Peng, Xiaohui, Xu, Yongjun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Nonlinear Hash-based Optimization Method for SpMV on GPUs
von: Yan, Chen, et al.
Veröffentlicht: (2025)
von: Yan, Chen, et al.
Veröffentlicht: (2025)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
SW-TNC : Reaching the Most Complex Random Quantum Circuit via Tensor Network Contraction
von: Chen, Yaojian, et al.
Veröffentlicht: (2025)
von: Chen, Yaojian, et al.
Veröffentlicht: (2025)
A Survey on Model-heterogeneous Federated Learning: Problems, Methods, and Prospects
von: Fan, Boyu, et al.
Veröffentlicht: (2023)
von: Fan, Boyu, et al.
Veröffentlicht: (2023)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
von: Ai, Xin, et al.
Veröffentlicht: (2024)
von: Ai, Xin, et al.
Veröffentlicht: (2024)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
von: Liu, Zhibang, et al.
Veröffentlicht: (2025)
von: Liu, Zhibang, et al.
Veröffentlicht: (2025)
Federated Learning Using Coupled Tensor Train Decomposition
von: Zhang, Xiangtao, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangtao, et al.
Veröffentlicht: (2024)
Synergistic Tensor and Pipeline Parallelism
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Deep Reinforcement Learning-based Methods for Resource Scheduling in Cloud Computing: A Review and Future Directions
von: Zhou, Guangyao, et al.
Veröffentlicht: (2021)
von: Zhou, Guangyao, et al.
Veröffentlicht: (2021)
Graph for Science: From API based Programming to Graph Engine based Programming for HPC
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
Towards the Distributed Large-scale k-NN Graph Construction by Graph Merge
von: Zhang, Cheng, et al.
Veröffentlicht: (2025)
von: Zhang, Cheng, et al.
Veröffentlicht: (2025)
Optimizing High-Throughput Distributed Data Pipelines for Reproducible Deep Learning at Scale
von: Mittal, Kashish, et al.
Veröffentlicht: (2026)
von: Mittal, Kashish, et al.
Veröffentlicht: (2026)
Fast Iterative Graph Computing with Updated Neighbor States
von: Zhou, Yijie, et al.
Veröffentlicht: (2024)
von: Zhou, Yijie, et al.
Veröffentlicht: (2024)
Deep Reinforcement Learning (DRL)-based Methods for Serverless Stream Processing Engines: A Vision, Architectural Elements, and Future Directions
von: Read, Maria R., et al.
Veröffentlicht: (2024)
von: Read, Maria R., et al.
Veröffentlicht: (2024)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
von: Zheng, Size, et al.
Veröffentlicht: (2025)
von: Zheng, Size, et al.
Veröffentlicht: (2025)
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
von: Ramachandran, Anvitha, et al.
Veröffentlicht: (2025)
von: Ramachandran, Anvitha, et al.
Veröffentlicht: (2025)
SOLANET: Distributed Neighbor Graph Construction on GPU-Accelerated Systems
von: Iwabuchi, Keita, et al.
Veröffentlicht: (2026)
von: Iwabuchi, Keita, et al.
Veröffentlicht: (2026)
Predictive Performance of Photonic SRAM-based In-Memory Computing for Tensor Decomposition
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
HGraphScale: Hierarchical Graph Learning for Autoscaling Microservice Applications in Container-based Cloud Computing
von: Fang, Zhengxin, et al.
Veröffentlicht: (2025)
von: Fang, Zhengxin, et al.
Veröffentlicht: (2025)
A Survey of Distributed Graph Algorithms on Massive Graphs
von: Meng, Lingkai, et al.
Veröffentlicht: (2024)
von: Meng, Lingkai, et al.
Veröffentlicht: (2024)
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
von: Lee, Sunjung, et al.
Veröffentlicht: (2026)
von: Lee, Sunjung, et al.
Veröffentlicht: (2026)
Scalable Distributed Vector Search via Accuracy Preserving Index Construction
von: Xu, Yuming, et al.
Veröffentlicht: (2025)
von: Xu, Yuming, et al.
Veröffentlicht: (2025)
Graph-based Gossiping for Communication Efficiency in Decentralized Federated Learning
von: Nguyen, Huong, et al.
Veröffentlicht: (2025)
von: Nguyen, Huong, et al.
Veröffentlicht: (2025)
Bingo: Radix-based Bias Factorization for Random Walk on Dynamic Graphs
von: Wang, Pinhuan, et al.
Veröffentlicht: (2025)
von: Wang, Pinhuan, et al.
Veröffentlicht: (2025)
Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling
von: Zhang, Xinyi, et al.
Veröffentlicht: (2024)
von: Zhang, Xinyi, et al.
Veröffentlicht: (2024)
Leveraging Neural Graph Compilers in Machine Learning Research for Edge-Cloud Systems
von: Furutanpey, Alireza, et al.
Veröffentlicht: (2025)
von: Furutanpey, Alireza, et al.
Veröffentlicht: (2025)
Towards Communication-Efficient Decentralized Federated Graph Learning over Non-IID Data
von: Wang, Shilong, et al.
Veröffentlicht: (2025)
von: Wang, Shilong, et al.
Veröffentlicht: (2025)
Efficient Parallel Compilation and Profiling of Quantum Circuits at Large Scales
von: Moore, Jane, et al.
Veröffentlicht: (2026)
von: Moore, Jane, et al.
Veröffentlicht: (2026)
LAPIS: A Performance Portable, High Productivity Compiler Framework
von: Kelley, Brian, et al.
Veröffentlicht: (2025)
von: Kelley, Brian, et al.
Veröffentlicht: (2025)
Selection of Supervised Learning-based Sparse Matrix Reordering Algorithms
von: Tang, Tao, et al.
Veröffentlicht: (2025)
von: Tang, Tao, et al.
Veröffentlicht: (2025)
CIR: Lightweight Container Image for Cross-Platform Deployment
von: Li, Fengzhi, et al.
Veröffentlicht: (2026)
von: Li, Fengzhi, et al.
Veröffentlicht: (2026)
TensorSocket: Shared Data Loading for Deep Learning Training
von: Robroek, Ties, et al.
Veröffentlicht: (2024)
von: Robroek, Ties, et al.
Veröffentlicht: (2024)
HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
von: Mei, Junyi, et al.
Veröffentlicht: (2024)
von: Mei, Junyi, et al.
Veröffentlicht: (2024)
Scale: Deep Reinforcement Learning for Container Scheduling in Serverless Edge Computing
von: Chen, Chen, et al.
Veröffentlicht: (2026)
von: Chen, Chen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Nonlinear Hash-based Optimization Method for SpMV on GPUs
von: Yan, Chen, et al.
Veröffentlicht: (2025) -
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025) -
SW-TNC : Reaching the Most Complex Random Quantum Circuit via Tensor Network Contraction
von: Chen, Yaojian, et al.
Veröffentlicht: (2025) -
A Survey on Model-heterogeneous Federated Learning: Problems, Methods, and Prospects
von: Fan, Boyu, et al.
Veröffentlicht: (2023) -
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
von: Ai, Xin, et al.
Veröffentlicht: (2024)