Leveraging Mathematical Reasoning of LLMs for Efficient GPU Thread Mapping
Fuente:
arXiv
Saved in:
| Main Authors: | Maureira, Jose, Navarro, Cristóbal A., Ferrada, Hector, Veas-Castillo, Luis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Convex Hull 3D Filtering with GPU Ray Tracing and Tensor Cores
by: Carrasco, Roberto, et al.
Published: (2026)
by: Carrasco, Roberto, et al.
Published: (2026)
CAT: Cellular Automata on Tensor cores
by: Navarro, Cristóbal A., et al.
Published: (2024)
by: Navarro, Cristóbal A., et al.
Published: (2024)
Ray Tracing Cores for General-Purpose Computing: A Literature Review
by: Meneses, Enzo, et al.
Published: (2026)
by: Meneses, Enzo, et al.
Published: (2026)
Thread and Data Mapping in Software Transactional Memory: An Overview
by: Pasqualin, Douglas Pereira, et al.
Published: (2022)
by: Pasqualin, Douglas Pereira, et al.
Published: (2022)
Frustrated with MPI+Threads? Try MPIxThreads!
by: Zhou, Hui, et al.
Published: (2024)
by: Zhou, Hui, et al.
Published: (2024)
AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
by: Park, Seongyeon, et al.
Published: (2024)
by: Park, Seongyeon, et al.
Published: (2024)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
by: Huang, En-Ming, et al.
Published: (2025)
by: Huang, En-Ming, et al.
Published: (2025)
GPU-Resident Gaussian Process Regression Leveraging Asynchronous Tasks with HPX
by: Möllmann, Henrik, et al.
Published: (2026)
by: Möllmann, Henrik, et al.
Published: (2026)
Basic Lock Algorithms in Lightweight Thread Environments
by: Skazhenik, Taras, et al.
Published: (2025)
by: Skazhenik, Taras, et al.
Published: (2025)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
by: Lee, Munkyu, et al.
Published: (2024)
by: Lee, Munkyu, et al.
Published: (2024)
Beyond Thread States: Diagnosing Performance Degradation with eBPF and Thread Dynamics
by: Landau, Diogo, et al.
Published: (2026)
by: Landau, Diogo, et al.
Published: (2026)
Efficient Accelerated Graph Edit Distance Computation on GPU
by: Dabah, Adel, et al.
Published: (2026)
by: Dabah, Adel, et al.
Published: (2026)
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
by: Yang, Zhuoping, et al.
Published: (2025)
by: Yang, Zhuoping, et al.
Published: (2025)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
by: Gu, Jianfeng, et al.
Published: (2025)
by: Gu, Jianfeng, et al.
Published: (2025)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
by: Mo, Zizhao, et al.
Published: (2025)
by: Mo, Zizhao, et al.
Published: (2025)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
by: Zhang, WenZheng, et al.
Published: (2024)
by: Zhang, WenZheng, et al.
Published: (2024)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
by: Li, Zhonggen, et al.
Published: (2025)
by: Li, Zhonggen, et al.
Published: (2025)
ZEUS: An Efficient GPU Optimization Method Integrating PSO, BFGS, and Automatic Differentiation
by: Soos, Dominik, et al.
Published: (2026)
by: Soos, Dominik, et al.
Published: (2026)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
by: Doucet, Zachary, et al.
Published: (2025)
by: Doucet, Zachary, et al.
Published: (2025)
Heimdall++: Optimizing GPU Utilization and Pipeline Parallelism for Efficient Single-Pulse Detection
by: Xia, Bingzheng, et al.
Published: (2025)
by: Xia, Bingzheng, et al.
Published: (2025)
An AD based library for Efficient Hessian and Hessian-Vector Product Computation on GPU
by: Ranjan, Desh, et al.
Published: (2024)
by: Ranjan, Desh, et al.
Published: (2024)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
by: Yu, Minchen, et al.
Published: (2023)
by: Yu, Minchen, et al.
Published: (2023)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
by: Zhao, Yanbo, et al.
Published: (2025)
by: Zhao, Yanbo, et al.
Published: (2025)
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
by: Park, Seongyeon, et al.
Published: (2025)
by: Park, Seongyeon, et al.
Published: (2025)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
by: Zhang, Mingjun, et al.
Published: (2025)
by: Zhang, Mingjun, et al.
Published: (2025)
Efficient GPU Implementation of Particle Interactions with Cutoff Radius and Few Particles per Cell
by: Algis, David, et al.
Published: (2024)
by: Algis, David, et al.
Published: (2024)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
by: Liu, Yunzhao, et al.
Published: (2025)
by: Liu, Yunzhao, et al.
Published: (2025)
The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
by: Amoros, Oscar, et al.
Published: (2025)
by: Amoros, Oscar, et al.
Published: (2025)
A New Family of Thread to Core Allocation Policies for an SMT ARM Processor
by: Navarro, Marta, et al.
Published: (2025)
by: Navarro, Marta, et al.
Published: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
by: He, Yongchao, et al.
Published: (2025)
by: He, Yongchao, et al.
Published: (2025)
Efficient Parallel Execution of Blockchain Transactions Leveraging Conflict Specifications
by: Anjana, Parwat Singh, et al.
Published: (2025)
by: Anjana, Parwat Singh, et al.
Published: (2025)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
by: Qiang, Xinwei, et al.
Published: (2026)
by: Qiang, Xinwei, et al.
Published: (2026)
Accelerating Biclique Counting on GPU
by: Qiu, Linshan, et al.
Published: (2024)
by: Qiu, Linshan, et al.
Published: (2024)
GPU Sharing with Triples Mode
by: Byun, Chansup, et al.
Published: (2024)
by: Byun, Chansup, et al.
Published: (2024)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
by: Kong, Z. Jonny, et al.
Published: (2025)
by: Kong, Z. Jonny, et al.
Published: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
by: Qianli, Liu, et al.
Published: (2025)
by: Qianli, Liu, et al.
Published: (2025)
Similar Items
-
Convex Hull 3D Filtering with GPU Ray Tracing and Tensor Cores
by: Carrasco, Roberto, et al.
Published: (2026) -
CAT: Cellular Automata on Tensor cores
by: Navarro, Cristóbal A., et al.
Published: (2024) -
Ray Tracing Cores for General-Purpose Computing: A Literature Review
by: Meneses, Enzo, et al.
Published: (2026) -
Thread and Data Mapping in Software Transactional Memory: An Overview
by: Pasqualin, Douglas Pereira, et al.
Published: (2022) -
Frustrated with MPI+Threads? Try MPIxThreads!
by: Zhou, Hui, et al.
Published: (2024)