A Parallel CPU-GPU Framework for Batching Heuristic Operations in Depth-First Heuristic Search
Fuente:
arXiv
Saved in:
| Main Authors: | Futuhi, Ehsan, Sturtevant, Nathan R. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Meta-Heuristic Load Balancer for Cloud Computing Systems
by: Sliwko, Leszek, et al.
Published: (2025)
by: Sliwko, Leszek, et al.
Published: (2025)
Vulcan: Instance-Optimal Systems Heuristics Through LLM-Driven Search
by: Dwivedula, Rohit, et al.
Published: (2025)
by: Dwivedula, Rohit, et al.
Published: (2025)
Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics
by: Zhou, Zijie
Published: (2026)
by: Zhou, Zijie
Published: (2026)
TS-EoH: An Edge Server Task Scheduling Algorithm Based on Evolution of Heuristic
by: Yatong, Wang, et al.
Published: (2024)
by: Yatong, Wang, et al.
Published: (2024)
A Heuristic Algorithm for Shortest Path Search
by: Yu, Huashan, et al.
Published: (2025)
by: Yu, Huashan, et al.
Published: (2025)
ODIN-Based CPU-GPU Architecture with Replay-Driven Simulation and Emulation
by: Dorairaj, Nij, et al.
Published: (2026)
by: Dorairaj, Nij, et al.
Published: (2026)
GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
by: VG, Jithin, et al.
Published: (2025)
by: VG, Jithin, et al.
Published: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
by: Jiang, Xuanlin, et al.
Published: (2024)
by: Jiang, Xuanlin, et al.
Published: (2024)
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
by: Chen, Jiesong, et al.
Published: (2026)
by: Chen, Jiesong, et al.
Published: (2026)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
by: Fan, Jiakun, et al.
Published: (2025)
by: Fan, Jiakun, et al.
Published: (2025)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
by: Lyu, Hongtao, et al.
Published: (2025)
by: Lyu, Hongtao, et al.
Published: (2025)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
by: Chen, Aodong, et al.
Published: (2023)
by: Chen, Aodong, et al.
Published: (2023)
PolyKAN: Efficient Fused GPU Operators for Polynomial Kolmogorov-Arnold Network Variants
by: Yu, Mingkun, et al.
Published: (2025)
by: Yu, Mingkun, et al.
Published: (2025)
Benchmarking of CPU-intensive Stream Data Processing in The Edge Computing Systems
by: Szydlo, Tomasz, et al.
Published: (2025)
by: Szydlo, Tomasz, et al.
Published: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
by: He, Yongchao, et al.
Published: (2025)
by: He, Yongchao, et al.
Published: (2025)
MineDraft: A Framework for Batch Parallel Speculative Decoding
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025)
by: Vellaisamy, Prabhu, et al.
Published: (2025)
Learning Admissible Heuristics for A*: Theory and Practice
by: Futuhi, Ehsan, et al.
Published: (2025)
by: Futuhi, Ehsan, et al.
Published: (2025)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
by: Wu, Qi, et al.
Published: (2026)
by: Wu, Qi, et al.
Published: (2026)
Astra: Efficient and Money-saving Automatic Parallel Strategies Search on Heterogeneous GPUs
by: Wang, Peiran, et al.
Published: (2025)
by: Wang, Peiran, et al.
Published: (2025)
Using Sequential Runtime Distributions for the Parallel Speedup Prediction of SAT Local Search
by: Arbelaez, Alejandro, et al.
Published: (2024)
by: Arbelaez, Alejandro, et al.
Published: (2024)
GPU-Accelerated Primal Heuristics for Mixed Integer Programming
by: Çördük, Akif, et al.
Published: (2025)
by: Çördük, Akif, et al.
Published: (2025)
T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge
by: Wei, Jianyu, et al.
Published: (2024)
by: Wei, Jianyu, et al.
Published: (2024)
A Unified CPU-GPU Protocol for GNN Training
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Harnessing Deep Learning and HPC Kernels via High-Level Loop and Tensor Abstractions on CPU Architectures
by: Georganas, Evangelos, et al.
Published: (2023)
by: Georganas, Evangelos, et al.
Published: (2023)
A Comparative Review of Parallel Exact, Heuristic, Metaheuristic, and Hybrid Optimization Techniques for the Traveling Salesman Problem
by: Alkhalifa, Rabab, et al.
Published: (2025)
by: Alkhalifa, Rabab, et al.
Published: (2025)
Placement Semantics for Distributed Deep Learning: A Systematic Framework for Analyzing Parallelism Strategies
by: Mehta, Deep Pankajbhai
Published: (2026)
by: Mehta, Deep Pankajbhai
Published: (2026)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
by: Ding, Zhimin, et al.
Published: (2024)
by: Ding, Zhimin, et al.
Published: (2024)
Combining GPU and CPU for accelerating evolutionary computing workloads
by: Eynaliyev, Rustam, et al.
Published: (2025)
by: Eynaliyev, Rustam, et al.
Published: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
by: Pan, Yudong, et al.
Published: (2026)
by: Pan, Yudong, et al.
Published: (2026)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024)
by: Kamahori, Keisuke, et al.
Published: (2024)
Towards Safer Heuristics With XPlain
by: Karimi, Pantea, et al.
Published: (2024)
by: Karimi, Pantea, et al.
Published: (2024)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
by: Xu, Lang, et al.
Published: (2024)
by: Xu, Lang, et al.
Published: (2024)
Graph Neural Networks as Ordering Heuristics for Parallel Graph Coloring
by: Langedal, Kenneth, et al.
Published: (2024)
by: Langedal, Kenneth, et al.
Published: (2024)
MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
by: Hwang, Changho, et al.
Published: (2025)
by: Hwang, Changho, et al.
Published: (2025)
UCCL-Zip: Lossless Compression Supercharged GPU Communication
by: Ma, Shuang, et al.
Published: (2026)
by: Ma, Shuang, et al.
Published: (2026)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
by: Lettich, Francesco, et al.
Published: (2024)
by: Lettich, Francesco, et al.
Published: (2024)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
by: Zheng, Wanyi, et al.
Published: (2025)
by: Zheng, Wanyi, et al.
Published: (2025)
Beyond the GPU: The Strategic Role of FPGAs in the Next Wave of AI
by: Jiménez, Arturo Urías
Published: (2025)
by: Jiménez, Arturo Urías
Published: (2025)
Accelerated Digital Twin Learning for Edge AI: A Comparison of FPGA and Mobile GPU
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
Similar Items
-
A Meta-Heuristic Load Balancer for Cloud Computing Systems
by: Sliwko, Leszek, et al.
Published: (2025) -
Vulcan: Instance-Optimal Systems Heuristics Through LLM-Driven Search
by: Dwivedula, Rohit, et al.
Published: (2025) -
Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics
by: Zhou, Zijie
Published: (2026) -
TS-EoH: An Edge Server Task Scheduling Algorithm Based on Evolution of Heuristic
by: Yatong, Wang, et al.
Published: (2024) -
A Heuristic Algorithm for Shortest Path Search
by: Yu, Huashan, et al.
Published: (2025)