ZEUS: An Efficient GPU Optimization Method Integrating PSO, BFGS, and Automatic Differentiation
Fuente:
arXiv
Guardado en:
| Autores principales: | Soos, Dominik, Paterno, Marc, Ranjan, Desh, Zubair, Mohammad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An AD based library for Efficient Hessian and Hessian-Vector Product Computation on GPU
por: Ranjan, Desh, et al.
Publicado: (2024)
por: Ranjan, Desh, et al.
Publicado: (2024)
Distributed Quantum-Enhanced Optimization: A Topographical Preconditioning Approach for High-Dimensional Search
por: Soós, Dominik, et al.
Publicado: (2026)
por: Soós, Dominik, et al.
Publicado: (2026)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
por: Li, Zhonggen, et al.
Publicado: (2025)
por: Li, Zhonggen, et al.
Publicado: (2025)
ACC Saturator: Automatic Kernel Optimization for Directive-Based GPU Code
por: Matsumura, Kazuaki, et al.
Publicado: (2023)
por: Matsumura, Kazuaki, et al.
Publicado: (2023)
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
por: Yang, Zhuoping, et al.
Publicado: (2025)
por: Yang, Zhuoping, et al.
Publicado: (2025)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
por: Qiang, Xinwei, et al.
Publicado: (2026)
por: Qiang, Xinwei, et al.
Publicado: (2026)
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
por: Scheinert, Dominik, et al.
Publicado: (2026)
por: Scheinert, Dominik, et al.
Publicado: (2026)
Heimdall++: Optimizing GPU Utilization and Pipeline Parallelism for Efficient Single-Pulse Detection
por: Xia, Bingzheng, et al.
Publicado: (2025)
por: Xia, Bingzheng, et al.
Publicado: (2025)
Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
por: Qiao, Tong, et al.
Publicado: (2025)
por: Qiao, Tong, et al.
Publicado: (2025)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
por: Boudaoud, Afif, et al.
Publicado: (2026)
por: Boudaoud, Afif, et al.
Publicado: (2026)
Scrutinizing Variables for Checkpoint Using Automatic Differentiation
por: Huang, Xin, et al.
Publicado: (2026)
por: Huang, Xin, et al.
Publicado: (2026)
Optimizing Bloom Filters for Modern GPU Architectures
por: Jünger, Daniel, et al.
Publicado: (2025)
por: Jünger, Daniel, et al.
Publicado: (2025)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
por: Lee, Munkyu, et al.
Publicado: (2024)
por: Lee, Munkyu, et al.
Publicado: (2024)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
por: Islam, Tanzima Z., et al.
Publicado: (2024)
por: Islam, Tanzima Z., et al.
Publicado: (2024)
Efficient Accelerated Graph Edit Distance Computation on GPU
por: Dabah, Adel, et al.
Publicado: (2026)
por: Dabah, Adel, et al.
Publicado: (2026)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
por: Gu, Jianfeng, et al.
Publicado: (2025)
por: Gu, Jianfeng, et al.
Publicado: (2025)
Demeter: Resource-Efficient Distributed Stream Processing under Dynamic Loads with Multi-Configuration Optimization
por: Geldenhuys, Morgan, et al.
Publicado: (2024)
por: Geldenhuys, Morgan, et al.
Publicado: (2024)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
por: Liu, Shifang, et al.
Publicado: (2025)
por: Liu, Shifang, et al.
Publicado: (2025)
A Preliminary Study on Accelerating Simulation Optimization with GPU Implementation
por: He, Jinghai, et al.
Publicado: (2024)
por: He, Jinghai, et al.
Publicado: (2024)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
por: Guo, Runsheng Benson, et al.
Publicado: (2025)
por: Guo, Runsheng Benson, et al.
Publicado: (2025)
Leveraging Mathematical Reasoning of LLMs for Efficient GPU Thread Mapping
por: Maureira, Jose, et al.
Publicado: (2026)
por: Maureira, Jose, et al.
Publicado: (2026)
MERBIT: A GPU-Based SpMV Method for Iterative Workloads
por: Zhang, Qi, et al.
Publicado: (2026)
por: Zhang, Qi, et al.
Publicado: (2026)
GPU-Based Parallel Computing Methods for Medical Photoacoustic Image Reconstruction
por: Yi, Xinyao, et al.
Publicado: (2024)
por: Yi, Xinyao, et al.
Publicado: (2024)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
por: Siavashi, Ahmad, et al.
Publicado: (2025)
por: Siavashi, Ahmad, et al.
Publicado: (2025)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
por: Maurya, Avinash, et al.
Publicado: (2024)
por: Maurya, Avinash, et al.
Publicado: (2024)
MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies
por: Vanecek, Stepan, et al.
Publicado: (2025)
por: Vanecek, Stepan, et al.
Publicado: (2025)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
por: Wu, Hao, et al.
Publicado: (2024)
por: Wu, Hao, et al.
Publicado: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
por: Zhang, WenZheng, et al.
Publicado: (2024)
por: Zhang, WenZheng, et al.
Publicado: (2024)
Optimizing Allreduce Operations for Modern Heterogeneous Architectures with Multiple Processes per GPU
por: Adams, Michael, et al.
Publicado: (2025)
por: Adams, Michael, et al.
Publicado: (2025)
GPU-Accelerated Modified Bessel Function of the Second Kind for Gaussian Processes
por: Geng, Zipei, et al.
Publicado: (2025)
por: Geng, Zipei, et al.
Publicado: (2025)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
por: Doucet, Zachary, et al.
Publicado: (2025)
por: Doucet, Zachary, et al.
Publicado: (2025)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
por: Liu, Yi, et al.
Publicado: (2025)
por: Liu, Yi, et al.
Publicado: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
por: Yu, Minchen, et al.
Publicado: (2023)
por: Yu, Minchen, et al.
Publicado: (2023)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
por: Schieffer, Gabin, et al.
Publicado: (2024)
por: Schieffer, Gabin, et al.
Publicado: (2024)
Proposal of Automatic Offloading Method in Mixed Offloading Destination Environment
por: Yamato, Yoji
Publicado: (2020)
por: Yamato, Yoji
Publicado: (2020)
Optimizing the Variant Calling Pipeline Execution on Human Genomes Using GPU-Enabled Machines
por: Kumar, Ajay, et al.
Publicado: (2025)
por: Kumar, Ajay, et al.
Publicado: (2025)
FastGraph: Optimized GPU-Enabled Algorithms for Fast Graph Building and Message Passing
por: Agarwal, Aarush, et al.
Publicado: (2025)
por: Agarwal, Aarush, et al.
Publicado: (2025)
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
por: Park, Seongyeon, et al.
Publicado: (2025)
por: Park, Seongyeon, et al.
Publicado: (2025)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
por: Zhang, Mingjun, et al.
Publicado: (2025)
por: Zhang, Mingjun, et al.
Publicado: (2025)
Efficient GPU Implementation of Particle Interactions with Cutoff Radius and Few Particles per Cell
por: Algis, David, et al.
Publicado: (2024)
por: Algis, David, et al.
Publicado: (2024)
Ejemplares similares
-
An AD based library for Efficient Hessian and Hessian-Vector Product Computation on GPU
por: Ranjan, Desh, et al.
Publicado: (2024) -
Distributed Quantum-Enhanced Optimization: A Topographical Preconditioning Approach for High-Dimensional Search
por: Soós, Dominik, et al.
Publicado: (2026) -
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
por: Li, Zhonggen, et al.
Publicado: (2025) -
ACC Saturator: Automatic Kernel Optimization for Directive-Based GPU Code
por: Matsumura, Kazuaki, et al.
Publicado: (2023) -
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
por: Yang, Zhuoping, et al.
Publicado: (2025)