HeteroSTA: A CPU-GPU Heterogeneous Static Timing Analysis Engine with Holistic Industrial Design Support
Fuente:
arXiv
Guardado en:
| Autores principales: | Guo, Zizheng, Liu, Haichuan, Shi, Xizhe, Hua, Shenglu, Zhang, Zuodong, Zhao, Chunyuan, Wang, Runsheng, Lin, Yibo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
por: Schieffer, Gabin, et al.
Publicado: (2026)
por: Schieffer, Gabin, et al.
Publicado: (2026)
Cephalo: Harnessing Heterogeneous GPU Clusters for Training Transformer Models
por: Guo, Runsheng Benson, et al.
Publicado: (2024)
por: Guo, Runsheng Benson, et al.
Publicado: (2024)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
por: Guo, Runsheng Benson, et al.
Publicado: (2025)
por: Guo, Runsheng Benson, et al.
Publicado: (2025)
Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
por: Qiao, Tong, et al.
Publicado: (2025)
por: Qiao, Tong, et al.
Publicado: (2025)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
por: Bridges, Patrick G., et al.
Publicado: (2026)
por: Bridges, Patrick G., et al.
Publicado: (2026)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
por: Zhong, Shuzhang, et al.
Publicado: (2025)
por: Zhong, Shuzhang, et al.
Publicado: (2025)
A Unified CPU-GPU Protocol for GNN Training
por: Lin, Yi-Chien, et al.
Publicado: (2024)
por: Lin, Yi-Chien, et al.
Publicado: (2024)
Combining GPU and CPU for accelerating evolutionary computing workloads
por: Eynaliyev, Rustam, et al.
Publicado: (2025)
por: Eynaliyev, Rustam, et al.
Publicado: (2025)
Orchestrated Co-scheduling, Resource Partitioning, and Power Capping on CPU-GPU Heterogeneous Systems via Machine Learning
por: Saba, Issa, et al.
Publicado: (2024)
por: Saba, Issa, et al.
Publicado: (2024)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
por: Ding, Zhimin, et al.
Publicado: (2024)
por: Ding, Zhimin, et al.
Publicado: (2024)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
por: Li, Zhonggen, et al.
Publicado: (2025)
por: Li, Zhonggen, et al.
Publicado: (2025)
A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture
por: Yi, Xinyao
Publicado: (2024)
por: Yi, Xinyao
Publicado: (2024)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
por: Fan, Jiakun, et al.
Publicado: (2025)
por: Fan, Jiakun, et al.
Publicado: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
por: Mo, Zizhao, et al.
Publicado: (2026)
por: Mo, Zizhao, et al.
Publicado: (2026)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
por: Maurya, Avinash, et al.
Publicado: (2024)
por: Maurya, Avinash, et al.
Publicado: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
por: Wahlgren, Jacob, et al.
Publicado: (2025)
por: Wahlgren, Jacob, et al.
Publicado: (2025)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
por: Huang, En-Ming, et al.
Publicado: (2025)
por: Huang, En-Ming, et al.
Publicado: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
por: He, Yongchao, et al.
Publicado: (2025)
por: He, Yongchao, et al.
Publicado: (2025)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
por: Schieffer, Gabin, et al.
Publicado: (2024)
por: Schieffer, Gabin, et al.
Publicado: (2024)
A Unified Programming Model for Heterogeneous Computing with CPU and Accelerator Technologies
por: Xiong, Yuqing
Publicado: (2022)
por: Xiong, Yuqing
Publicado: (2022)
Cost-Performance Analysis: A Comparative Study of CPU-Based Serverless and GPU-Based Training Architectures
por: Barrak, Amine, et al.
Publicado: (2025)
por: Barrak, Amine, et al.
Publicado: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
por: Zhang, WenZheng, et al.
Publicado: (2024)
por: Zhang, WenZheng, et al.
Publicado: (2024)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
por: Liang, Antian, et al.
Publicado: (2025)
por: Liang, Antian, et al.
Publicado: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
por: Lin, Mao, et al.
Publicado: (2026)
por: Lin, Mao, et al.
Publicado: (2026)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
por: Chung, Euijun, et al.
Publicado: (2026)
por: Chung, Euijun, et al.
Publicado: (2026)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
por: Chen, Kefu, et al.
Publicado: (2026)
por: Chen, Kefu, et al.
Publicado: (2026)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
por: Zhao, Xinkui, et al.
Publicado: (2025)
por: Zhao, Xinkui, et al.
Publicado: (2025)
Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
por: Mo, Zizhao, et al.
Publicado: (2024)
por: Mo, Zizhao, et al.
Publicado: (2024)
Towards CXL Resilience to CPU Failures
por: Psistakis, Antonis, et al.
Publicado: (2026)
por: Psistakis, Antonis, et al.
Publicado: (2026)
Optimizing Task Scheduling in Heterogeneous Computing Environments: A Comparative Analysis of CPU, GPU, and ASIC Platforms Using E2C Simulator
por: Mohammadjafari, Ali, et al.
Publicado: (2024)
por: Mohammadjafari, Ali, et al.
Publicado: (2024)
Comparing CPU and GPU compute of PERMANOVA on MI300A
por: Sfiligoi, Igor
Publicado: (2025)
por: Sfiligoi, Igor
Publicado: (2025)
Optimizing Allreduce Operations for Modern Heterogeneous Architectures with Multiple Processes per GPU
por: Adams, Michael, et al.
Publicado: (2025)
por: Adams, Michael, et al.
Publicado: (2025)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
por: Mo, Zizhao, et al.
Publicado: (2025)
por: Mo, Zizhao, et al.
Publicado: (2025)
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
por: Zhang, Shiwei, et al.
Publicado: (2024)
por: Zhang, Shiwei, et al.
Publicado: (2024)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
por: Shi, Jiabo, et al.
Publicado: (2025)
por: Shi, Jiabo, et al.
Publicado: (2025)
WindVE: Collaborative CPU-NPU Vector Embedding
por: Huang, Jinqi, et al.
Publicado: (2025)
por: Huang, Jinqi, et al.
Publicado: (2025)
A Parallel CPU-GPU Framework for Batching Heuristic Operations in Depth-First Heuristic Search
por: Futuhi, Ehsan, et al.
Publicado: (2025)
por: Futuhi, Ehsan, et al.
Publicado: (2025)
Parallel CPU- and GPU-based connected component algorithms for event building for hybrid pixel detectors
por: Čelko, Tomáš, et al.
Publicado: (2024)
por: Čelko, Tomáš, et al.
Publicado: (2024)
Warp-STAR: High-performance, Differentiable GPU-Accelerated Static Timing Analysis through Warp-oriented Parallel Orchestration
por: Huang, En-Ming, et al.
Publicado: (2026)
por: Huang, En-Ming, et al.
Publicado: (2026)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
por: Liu, Yunzhao, et al.
Publicado: (2025)
por: Liu, Yunzhao, et al.
Publicado: (2025)
Ejemplares similares
-
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
por: Schieffer, Gabin, et al.
Publicado: (2026) -
Cephalo: Harnessing Heterogeneous GPU Clusters for Training Transformer Models
por: Guo, Runsheng Benson, et al.
Publicado: (2024) -
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
por: Guo, Runsheng Benson, et al.
Publicado: (2025) -
Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
por: Qiao, Tong, et al.
Publicado: (2025) -
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
por: Bridges, Patrick G., et al.
Publicado: (2026)