Scaling Deep Learning Computation over the Inter-Core Connected Intelligence Processor with T10
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yiqi, Xue, Yuqi, Cheng, Yu, Ma, Lingxiao, Miao, Ziming, Xue, Jilong, Huang, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
von: Liu, Yiqi, et al.
Veröffentlicht: (2025)
von: Liu, Yiqi, et al.
Veröffentlicht: (2025)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
von: Huang, Ziyu, et al.
Veröffentlicht: (2025)
von: Huang, Ziyu, et al.
Veröffentlicht: (2025)
WaferLLM: Large Language Model Inference at Wafer Scale
von: He, Congjie, et al.
Veröffentlicht: (2025)
von: He, Congjie, et al.
Veröffentlicht: (2025)
Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
von: Feng, Dahu, et al.
Veröffentlicht: (2025)
von: Feng, Dahu, et al.
Veröffentlicht: (2025)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
Deep Reinforcement Learning-based Methods for Resource Scheduling in Cloud Computing: A Review and Future Directions
von: Zhou, Guangyao, et al.
Veröffentlicht: (2021)
von: Zhou, Guangyao, et al.
Veröffentlicht: (2021)
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
von: Duan, Jiaang, et al.
Veröffentlicht: (2024)
von: Duan, Jiaang, et al.
Veröffentlicht: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2024)
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2024)
Scale: Deep Reinforcement Learning for Container Scheduling in Serverless Edge Computing
von: Chen, Chen, et al.
Veröffentlicht: (2026)
von: Chen, Chen, et al.
Veröffentlicht: (2026)
Co-designing a Programmable RISC-V Accelerator for MPC-based Energy and Thermal Management of Many-Core HPC Processors
von: Ottaviano, Alessandro, et al.
Veröffentlicht: (2025)
von: Ottaviano, Alessandro, et al.
Veröffentlicht: (2025)
Neutron particle transport 3D method of characteristic Multi GPU platform Parallel Computing
von: Zhou, Faguo, et al.
Veröffentlicht: (2025)
von: Zhou, Faguo, et al.
Veröffentlicht: (2025)
Edge Intelligence in Satellite-Terrestrial Networks with Hybrid Quantum Computing
von: Huang, Siyue, et al.
Veröffentlicht: (2024)
von: Huang, Siyue, et al.
Veröffentlicht: (2024)
Mitigating Interference of Microservices with a Scoring Mechanism in Large-scale Clusters
von: Yang, Dingyu, et al.
Veröffentlicht: (2024)
von: Yang, Dingyu, et al.
Veröffentlicht: (2024)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
von: Miksits, Samuel, et al.
Veröffentlicht: (2024)
von: Miksits, Samuel, et al.
Veröffentlicht: (2024)
HGraphScale: Hierarchical Graph Learning for Autoscaling Microservice Applications in Container-based Cloud Computing
von: Fang, Zhengxin, et al.
Veröffentlicht: (2025)
von: Fang, Zhengxin, et al.
Veröffentlicht: (2025)
Do We Need Tensor Cores for Stencil Computations?
von: Gu, Qiqi, et al.
Veröffentlicht: (2026)
von: Gu, Qiqi, et al.
Veröffentlicht: (2026)
A New Family of Thread to Core Allocation Policies for an SMT ARM Processor
von: Navarro, Marta, et al.
Veröffentlicht: (2025)
von: Navarro, Marta, et al.
Veröffentlicht: (2025)
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
Energy-Aware Scheduling Strategies for Partially-Replicable Task Chains on Heterogeneous Processors
von: Idouar, Yacine, et al.
Veröffentlicht: (2025)
von: Idouar, Yacine, et al.
Veröffentlicht: (2025)
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
von: Mao, Ying, et al.
Veröffentlicht: (2020)
von: Mao, Ying, et al.
Veröffentlicht: (2020)
Optimizing High-Throughput Distributed Data Pipelines for Reproducible Deep Learning at Scale
von: Mittal, Kashish, et al.
Veröffentlicht: (2026)
von: Mittal, Kashish, et al.
Veröffentlicht: (2026)
Collaborative Evolution of Intelligent Agents in Large-Scale Microservice Systems
von: Li, Yilin, et al.
Veröffentlicht: (2025)
von: Li, Yilin, et al.
Veröffentlicht: (2025)
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
von: Hua, Qin, et al.
Veröffentlicht: (2024)
von: Hua, Qin, et al.
Veröffentlicht: (2024)
Exploring Uncore Frequency Scaling for Heterogeneous Computing
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
Air-FedGA: A Grouping Asynchronous Federated Learning Mechanism Exploiting Over-the-air Computation
von: Ma, Qianpiao, et al.
Veröffentlicht: (2025)
von: Ma, Qianpiao, et al.
Veröffentlicht: (2025)
Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
von: Kinkead, Shannon, et al.
Veröffentlicht: (2026)
von: Kinkead, Shannon, et al.
Veröffentlicht: (2026)
Leveraging Core and Uncore Frequency Scaling for Power-Efficient Serverless Workflows
von: Tzenetopoulos, Achilleas, et al.
Veröffentlicht: (2024)
von: Tzenetopoulos, Achilleas, et al.
Veröffentlicht: (2024)
Ray Tracing Cores for General-Purpose Computing: A Literature Review
von: Meneses, Enzo, et al.
Veröffentlicht: (2026)
von: Meneses, Enzo, et al.
Veröffentlicht: (2026)
High Performance Unstructured SpMM Computation Using Tensor Cores
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural Networks
von: Naman, Pranjal, et al.
Veröffentlicht: (2026)
von: Naman, Pranjal, et al.
Veröffentlicht: (2026)
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
von: GU, Qiqi, et al.
Veröffentlicht: (2025)
von: GU, Qiqi, et al.
Veröffentlicht: (2025)
Task Scheduling in Geo-Distributed Computing: A Survey
von: Wu, Yujian, et al.
Veröffentlicht: (2025)
von: Wu, Yujian, et al.
Veröffentlicht: (2025)
Differentially Private Perturbed Push-Sum Protocol and Its Application in Non-Convex Optimization
von: Zhou, Yiming, et al.
Veröffentlicht: (2026)
von: Zhou, Yiming, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
von: Liu, Yiqi, et al.
Veröffentlicht: (2025) -
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
von: Liu, Yiqi, et al.
Veröffentlicht: (2026) -
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
von: Huang, Ziyu, et al.
Veröffentlicht: (2025) -
WaferLLM: Large Language Model Inference at Wafer Scale
von: He, Congjie, et al.
Veröffentlicht: (2025) -
Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
von: Feng, Dahu, et al.
Veröffentlicht: (2025)