Dual-pronged deep learning preprocessing on heterogeneous platforms with CPU, Accelerator and CSD
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Jia, Zhang, Xingjun, Pedrycz, Witold, Wang, Longxiang, Zhao, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Service Level Agreements and Security SLA: A Comprehensive Survey
von: Nicolazzo, Serena, et al.
Veröffentlicht: (2024)
von: Nicolazzo, Serena, et al.
Veröffentlicht: (2024)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
A Unified Programming Model for Heterogeneous Computing with CPU and Accelerator Technologies
von: Xiong, Yuqing
Veröffentlicht: (2022)
von: Xiong, Yuqing
Veröffentlicht: (2022)
Regent based parallel meshfree LSKUM solver for heterogenous HPC platforms
von: Salil, Sanath, et al.
Veröffentlicht: (2024)
von: Salil, Sanath, et al.
Veröffentlicht: (2024)
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
von: Schieffer, Gabin, et al.
Veröffentlicht: (2026)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2026)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
Towards CXL Resilience to CPU Failures
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
Efficient Column-Wise N:M Pruning on RISC-V CPU
von: Chu, Chi-Wei, et al.
Veröffentlicht: (2025)
von: Chu, Chi-Wei, et al.
Veröffentlicht: (2025)
MMStencil: Optimizing High-order Stencils on Multicore CPU using Matrix Unit
von: Wang, Yinuo, et al.
Veröffentlicht: (2025)
von: Wang, Yinuo, et al.
Veröffentlicht: (2025)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
HeteroSTA: A CPU-GPU Heterogeneous Static Timing Analysis Engine with Holistic Industrial Design Support
von: Guo, Zizheng, et al.
Veröffentlicht: (2025)
von: Guo, Zizheng, et al.
Veröffentlicht: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
von: He, Yongchao, et al.
Veröffentlicht: (2025)
von: He, Yongchao, et al.
Veröffentlicht: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
A Unified CPU-GPU Protocol for GNN Training
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
WindVE: Collaborative CPU-NPU Vector Embedding
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
Combining GPU and CPU for accelerating evolutionary computing workloads
von: Eynaliyev, Rustam, et al.
Veröffentlicht: (2025)
von: Eynaliyev, Rustam, et al.
Veröffentlicht: (2025)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
von: Qiao, Tong, et al.
Veröffentlicht: (2025)
von: Qiao, Tong, et al.
Veröffentlicht: (2025)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
Justin: Hybrid CPU/Memory Elastic Scaling for Distributed Stream Processing
von: Schmitz, Donatien, et al.
Veröffentlicht: (2025)
von: Schmitz, Donatien, et al.
Veröffentlicht: (2025)
KAITIAN: A Unified Communication Framework for Enabling Efficient Collaboration Across Heterogeneous Accelerators in Embodied AI Systems
von: Lin, Jieke, et al.
Veröffentlicht: (2025)
von: Lin, Jieke, et al.
Veröffentlicht: (2025)
Flotilla: A scalable, modular and resilient federated learning framework for heterogeneous resources
von: Banerjee, Roopkatha, et al.
Veröffentlicht: (2025)
von: Banerjee, Roopkatha, et al.
Veröffentlicht: (2025)
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
von: Brown, Nick, et al.
Veröffentlicht: (2024)
von: Brown, Nick, et al.
Veröffentlicht: (2024)
A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture
von: Yi, Xinyao
Veröffentlicht: (2024)
von: Yi, Xinyao
Veröffentlicht: (2024)
Distributed Retrieval-Augmented Generation
von: Xu, Chenhao, et al.
Veröffentlicht: (2025)
von: Xu, Chenhao, et al.
Veröffentlicht: (2025)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
von: Bridges, Patrick G., et al.
Veröffentlicht: (2026)
von: Bridges, Patrick G., et al.
Veröffentlicht: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
Aging-aware CPU Core Management for Embodied Carbon Amortization in Cloud LLM Inference
von: Hewage, Tharindu B., et al.
Veröffentlicht: (2025)
von: Hewage, Tharindu B., et al.
Veröffentlicht: (2025)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
Cost-Performance Analysis: A Comparative Study of CPU-Based Serverless and GPU-Based Training Architectures
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
Orchestrated Co-scheduling, Resource Partitioning, and Power Capping on CPU-GPU Heterogeneous Systems via Machine Learning
von: Saba, Issa, et al.
Veröffentlicht: (2024)
von: Saba, Issa, et al.
Veröffentlicht: (2024)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge
von: Wei, Jianyu, et al.
Veröffentlicht: (2024)
von: Wei, Jianyu, et al.
Veröffentlicht: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Service Level Agreements and Security SLA: A Comprehensive Survey
von: Nicolazzo, Serena, et al.
Veröffentlicht: (2024) -
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
von: Chen, Kefu, et al.
Veröffentlicht: (2026) -
A Unified Programming Model for Heterogeneous Computing with CPU and Accelerator Technologies
von: Xiong, Yuqing
Veröffentlicht: (2022) -
Regent based parallel meshfree LSKUM solver for heterogenous HPC platforms
von: Salil, Sanath, et al.
Veröffentlicht: (2024) -
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
von: Schieffer, Gabin, et al.
Veröffentlicht: (2026)