Warp-STAR: High-performance, Differentiable GPU-Accelerated Static Timing Analysis through Warp-oriented Parallel Orchestration
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, En-Ming, Hung, Shih-Hao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
por: Agrawal, Amey, et al.
Publicado: (2026)
por: Agrawal, Amey, et al.
Publicado: (2026)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
por: Liang, Antian, et al.
Publicado: (2025)
por: Liang, Antian, et al.
Publicado: (2025)
WarpSpeed: A High-Performance Library for Concurrent GPU Hash Tables
por: McCoy, Hunter, et al.
Publicado: (2025)
por: McCoy, Hunter, et al.
Publicado: (2025)
Hive Hash Table: A Warp-Cooperative, Dynamically Resizable Hash Table for GPUs
por: Polak, Md Sabbir Hossain, et al.
Publicado: (2025)
por: Polak, Md Sabbir Hossain, et al.
Publicado: (2025)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
por: Huang, En-Ming, et al.
Publicado: (2025)
por: Huang, En-Ming, et al.
Publicado: (2025)
Multi-GPU Acceleration of PALABOS Fluid Solver using C++ Standard Parallelism
por: Latt, Jonas, et al.
Publicado: (2025)
por: Latt, Jonas, et al.
Publicado: (2025)
A large-scale distributed parallel discrete event simulation engines based on Warped2 for Wargaming simulation
por: Jia, Xiaoning, et al.
Publicado: (2025)
por: Jia, Xiaoning, et al.
Publicado: (2025)
HeteroSTA: A CPU-GPU Heterogeneous Static Timing Analysis Engine with Holistic Industrial Design Support
por: Guo, Zizheng, et al.
Publicado: (2025)
por: Guo, Zizheng, et al.
Publicado: (2025)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
por: Xia, Mengchun, et al.
Publicado: (2026)
por: Xia, Mengchun, et al.
Publicado: (2026)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
por: Wu, Hao, et al.
Publicado: (2024)
por: Wu, Hao, et al.
Publicado: (2024)
Accelerating Microswimmer Simulations via a Heterogeneous Pipelined Parallel-in-Time Framework
por: Huang, Ruixiang, et al.
Publicado: (2026)
por: Huang, Ruixiang, et al.
Publicado: (2026)
Large Scale Multi-GPU Based Parallel Traffic Simulation for Accelerated Traffic Assignment and Propagation
por: Jiang, Xuan, et al.
Publicado: (2024)
por: Jiang, Xuan, et al.
Publicado: (2024)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
por: Knorr, Fabian, et al.
Publicado: (2025)
por: Knorr, Fabian, et al.
Publicado: (2025)
Accelerating Cloud-Based Transcriptomics: Performance Analysis and Optimization of the STAR Aligner Workflow
por: Kica, Piotr, et al.
Publicado: (2025)
por: Kica, Piotr, et al.
Publicado: (2025)
Accelerating Biclique Counting on GPU
por: Qiu, Linshan, et al.
Publicado: (2024)
por: Qiu, Linshan, et al.
Publicado: (2024)
GPZ: GPU-Accelerated Lossy Compressor for Particle Data
por: Li, Ruoyu, et al.
Publicado: (2025)
por: Li, Ruoyu, et al.
Publicado: (2025)
Dataflow-Oriented Classification and Performance Analysis of GPU-Accelerated Homomorphic Encryption
por: Nozaki, Ai, et al.
Publicado: (2026)
por: Nozaki, Ai, et al.
Publicado: (2026)
On Orchestrating Parallel Broadcasts for Distributed Ledgers
por: Sheng, Peiyao, et al.
Publicado: (2024)
por: Sheng, Peiyao, et al.
Publicado: (2024)
GPU Accelerated Sparse Cholesky Factorization
por: Karsavuran, M. Ozan, et al.
Publicado: (2024)
por: Karsavuran, M. Ozan, et al.
Publicado: (2024)
City-Scale Visibility Graph Analysis via GPU-Accelerated HyperBall
por: Hodge, Alex, et al.
Publicado: (2026)
por: Hodge, Alex, et al.
Publicado: (2026)
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
por: Schieffer, Gabin, et al.
Publicado: (2026)
por: Schieffer, Gabin, et al.
Publicado: (2026)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
por: Ting, Hsu-Tzu, et al.
Publicado: (2025)
por: Ting, Hsu-Tzu, et al.
Publicado: (2025)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
por: Huang, Jiajun, et al.
Publicado: (2023)
por: Huang, Jiajun, et al.
Publicado: (2023)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
por: Stoyanov, Radostin, et al.
Publicado: (2025)
por: Stoyanov, Radostin, et al.
Publicado: (2025)
GPU-Accelerated Batch-Dynamic Subgraph Matching
por: Qiu, Linshan, et al.
Publicado: (2024)
por: Qiu, Linshan, et al.
Publicado: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
por: Sojoodi, Amirhossein, et al.
Publicado: (2026)
por: Sojoodi, Amirhossein, et al.
Publicado: (2026)
GPU-Based Parallel Computing Methods for Medical Photoacoustic Image Reconstruction
por: Yi, Xinyao, et al.
Publicado: (2024)
por: Yi, Xinyao, et al.
Publicado: (2024)
Six Times to Spare: Characterizing GPU-Accelerated 5G LDPC Decoding for Edge-RSU Communications
por: Barker, Ryan, et al.
Publicado: (2026)
por: Barker, Ryan, et al.
Publicado: (2026)
Efficient Accelerated Graph Edit Distance Computation on GPU
por: Dabah, Adel, et al.
Publicado: (2026)
por: Dabah, Adel, et al.
Publicado: (2026)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
por: Wijeratne, Sasindu, et al.
Publicado: (2025)
por: Wijeratne, Sasindu, et al.
Publicado: (2025)
PICO: Accelerating All k-Core Paradigms on GPU
por: Zhao, Chen, et al.
Publicado: (2024)
por: Zhao, Chen, et al.
Publicado: (2024)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
por: Jiang, Youhe, et al.
Publicado: (2026)
por: Jiang, Youhe, et al.
Publicado: (2026)
Orchestrated Co-scheduling, Resource Partitioning, and Power Capping on CPU-GPU Heterogeneous Systems via Machine Learning
por: Saba, Issa, et al.
Publicado: (2024)
por: Saba, Issa, et al.
Publicado: (2024)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
por: Mo, Zizhao, et al.
Publicado: (2025)
por: Mo, Zizhao, et al.
Publicado: (2025)
Heimdall++: Optimizing GPU Utilization and Pipeline Parallelism for Efficient Single-Pulse Detection
por: Xia, Bingzheng, et al.
Publicado: (2025)
por: Xia, Bingzheng, et al.
Publicado: (2025)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
por: McFarland, Thomas, et al.
Publicado: (2025)
por: McFarland, Thomas, et al.
Publicado: (2025)
SOLANET: Distributed Neighbor Graph Construction on GPU-Accelerated Systems
por: Iwabuchi, Keita, et al.
Publicado: (2026)
por: Iwabuchi, Keita, et al.
Publicado: (2026)
A Preliminary Study on Accelerating Simulation Optimization with GPU Implementation
por: He, Jinghai, et al.
Publicado: (2024)
por: He, Jinghai, et al.
Publicado: (2024)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
por: Xu, Zhihao, et al.
Publicado: (2025)
por: Xu, Zhihao, et al.
Publicado: (2025)
PilotANN: Memory-Bounded GPU Acceleration for Vector Search
por: Gui, Yuntao, et al.
Publicado: (2025)
por: Gui, Yuntao, et al.
Publicado: (2025)
Ejemplares similares
-
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
por: Agrawal, Amey, et al.
Publicado: (2026) -
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
por: Liang, Antian, et al.
Publicado: (2025) -
WarpSpeed: A High-Performance Library for Concurrent GPU Hash Tables
por: McCoy, Hunter, et al.
Publicado: (2025) -
Hive Hash Table: A Warp-Cooperative, Dynamically Resizable Hash Table for GPUs
por: Polak, Md Sabbir Hossain, et al.
Publicado: (2025) -
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
por: Huang, En-Ming, et al.
Publicado: (2025)