Vortex: Efficient Sample-Free Dynamic Tensor Program Optimization via Hardware-aware Strategy Space Hierarchization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yangjie, Zhu, Honglin, Qiu, Qian, Cui, Weihao, Liu, Zihan, Guo, Cong, Feng, Siyuan, Meng, Jintao, Lan, Haidong, Leng, Jingwen, Zhu, Wenxi, Deng, Minwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
von: Zhu, Honglin, et al.
Veröffentlicht: (2026)
von: Zhu, Honglin, et al.
Veröffentlicht: (2026)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
Accelerating Sparse DNNs Based on Tiled GEMM
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
von: Huang, Ziyu, et al.
Veröffentlicht: (2025)
von: Huang, Ziyu, et al.
Veröffentlicht: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
Towards Fast Setup and High Throughput of GPU Serverless Computing
von: Zhao, Han, et al.
Veröffentlicht: (2024)
von: Zhao, Han, et al.
Veröffentlicht: (2024)
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
von: Luo, Xinhao, et al.
Veröffentlicht: (2025)
von: Luo, Xinhao, et al.
Veröffentlicht: (2025)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
von: Ma, Cong, et al.
Veröffentlicht: (2025)
von: Ma, Cong, et al.
Veröffentlicht: (2025)
Approximated Coded Computing: Towards Fast, Private and Secure Distributed Machine Learning
von: Qiu, Houming, et al.
Veröffentlicht: (2024)
von: Qiu, Houming, et al.
Veröffentlicht: (2024)
Synergistic Tensor and Pipeline Parallelism
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
KernelFoundry: Hardware-aware evolutionary GPU kernel optimization
von: Wiedemann, Nina, et al.
Veröffentlicht: (2026)
von: Wiedemann, Nina, et al.
Veröffentlicht: (2026)
Hardware-Aware Reformulation of Convolutions for Efficient Execution on Specialized AI Hardware: A Case Study on NVIDIA Tensor Cores
von: Bikshandi, Ganesh
Veröffentlicht: (2026)
von: Bikshandi, Ganesh
Veröffentlicht: (2026)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
Federated Learning Using Coupled Tensor Train Decomposition
von: Zhang, Xiangtao, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangtao, et al.
Veröffentlicht: (2024)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
Hardware-aware Circuit Cutting and Distributed Qubit Mapping for Connected Quantum Systems
von: Du, Zefan, et al.
Veröffentlicht: (2024)
von: Du, Zefan, et al.
Veröffentlicht: (2024)
Barycentric Coded Distributed Computing with Flexible Recovery Threshold for Collaborative Mobile Edge Computing
von: Qiu, Houming, et al.
Veröffentlicht: (2025)
von: Qiu, Houming, et al.
Veröffentlicht: (2025)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
von: Shi, Long, et al.
Veröffentlicht: (2025)
von: Shi, Long, et al.
Veröffentlicht: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
von: Patke, Archit, et al.
Veröffentlicht: (2025)
von: Patke, Archit, et al.
Veröffentlicht: (2025)
veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
von: Li, Youjie, et al.
Veröffentlicht: (2025)
von: Li, Youjie, et al.
Veröffentlicht: (2025)
Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Hierarchical Observe-Orient-Decide-Act Enabled UAV Swarms in Uncertain Environments: Frameworks, Potentials, and Challenges
von: Jia, Ziye, et al.
Veröffentlicht: (2026)
von: Jia, Ziye, et al.
Veröffentlicht: (2026)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
von: Tang, Ding, et al.
Veröffentlicht: (2024)
von: Tang, Ding, et al.
Veröffentlicht: (2024)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
On Principled Local Optimization Methods for Federated Learning
von: Yuan, Honglin
Veröffentlicht: (2024)
von: Yuan, Honglin
Veröffentlicht: (2024)
DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training
von: Qiang, Xinwei, et al.
Veröffentlicht: (2026)
von: Qiang, Xinwei, et al.
Veröffentlicht: (2026)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
Lumos: Heterogeneity-aware Federated Graph Learning over Decentralized Devices
von: Pan, Qiying, et al.
Veröffentlicht: (2023)
von: Pan, Qiying, et al.
Veröffentlicht: (2023)
When MoE Meets Blockchain: A Trustworthy Distributed Framework of Large Models
von: Zhu, Weihao, et al.
Veröffentlicht: (2025)
von: Zhu, Weihao, et al.
Veröffentlicht: (2025)
Efficient Model Compression for Hierarchical Federated Learning
von: Zhu, Xi, et al.
Veröffentlicht: (2024)
von: Zhu, Xi, et al.
Veröffentlicht: (2024)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
Topology-aware Federated Learning in Edge Computing: A Comprehensive Survey
von: Wu, Jiajun, et al.
Veröffentlicht: (2023)
von: Wu, Jiajun, et al.
Veröffentlicht: (2023)
HiRL: Hierarchical Reinforcement Learning for Coordinated Resource Management in Heterogeneous Edge Computing
von: Zhu, Jianyong, et al.
Veröffentlicht: (2026)
von: Zhu, Jianyong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
von: Zhu, Honglin, et al.
Veröffentlicht: (2026) -
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
von: Liu, Zihan, et al.
Veröffentlicht: (2025) -
Accelerating Sparse DNNs Based on Tiled GEMM
von: Guo, Cong, et al.
Veröffentlicht: (2024) -
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2024) -
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
von: Huang, Ziyu, et al.
Veröffentlicht: (2025)