PolyKAN: Efficient Fused GPU Operators for Polynomial Kolmogorov-Arnold Network Variants
Fuente:
arXiv
Guardado en:
| Autores principales: | Yu, Mingkun, Zhong, Heming, Huang, Dan, Lu, Yutong, Jiang, Jiazhi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
por: Wei, Jinhui, et al.
Publicado: (2025)
por: Wei, Jinhui, et al.
Publicado: (2025)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
por: Zhang, Hongbin, et al.
Publicado: (2026)
por: Zhang, Hongbin, et al.
Publicado: (2026)
TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
por: Wu, Shixun, et al.
Publicado: (2025)
por: Wu, Shixun, et al.
Publicado: (2025)
The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
por: Amoros, Oscar, et al.
Publicado: (2025)
por: Amoros, Oscar, et al.
Publicado: (2025)
Fused Breadth-First Probabilistic Traversals on Distributed GPU Systems
por: Neff, Reece, et al.
Publicado: (2023)
por: Neff, Reece, et al.
Publicado: (2023)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
por: Lin, Zejia, et al.
Publicado: (2025)
por: Lin, Zejia, et al.
Publicado: (2025)
Enhancing Physics-Informed Neural Networks with a Hybrid Parallel Kolmogorov-Arnold and MLP Architecture
por: Xu, Zuyu, et al.
Publicado: (2025)
por: Xu, Zuyu, et al.
Publicado: (2025)
Optimizing the Variant Calling Pipeline Execution on Human Genomes Using GPU-Enabled Machines
por: Kumar, Ajay, et al.
Publicado: (2025)
por: Kumar, Ajay, et al.
Publicado: (2025)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
por: Gu, Jianfeng, et al.
Publicado: (2025)
por: Gu, Jianfeng, et al.
Publicado: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
por: Yu, Minchen, et al.
Publicado: (2023)
por: Yu, Minchen, et al.
Publicado: (2023)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
por: Lee, Munkyu, et al.
Publicado: (2024)
por: Lee, Munkyu, et al.
Publicado: (2024)
Optimizing Allreduce Operations for Modern Heterogeneous Architectures with Multiple Processes per GPU
por: Adams, Michael, et al.
Publicado: (2025)
por: Adams, Michael, et al.
Publicado: (2025)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
por: Huang, En-Ming, et al.
Publicado: (2025)
por: Huang, En-Ming, et al.
Publicado: (2025)
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
por: Yang, Zhuoping, et al.
Publicado: (2025)
por: Yang, Zhuoping, et al.
Publicado: (2025)
Efficient Accelerated Graph Edit Distance Computation on GPU
por: Dabah, Adel, et al.
Publicado: (2026)
por: Dabah, Adel, et al.
Publicado: (2026)
Accelerating Biclique Counting on GPU
por: Qiu, Linshan, et al.
Publicado: (2024)
por: Qiu, Linshan, et al.
Publicado: (2024)
PICO: Accelerating All k-Core Paradigms on GPU
por: Zhao, Chen, et al.
Publicado: (2024)
por: Zhao, Chen, et al.
Publicado: (2024)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
por: Huang, Jiajun, et al.
Publicado: (2023)
por: Huang, Jiajun, et al.
Publicado: (2023)
GeoT: Tensor Centric Library for Graph Neural Network via Efficient Segment Reduction on GPU
por: Yu, Zhongming, et al.
Publicado: (2024)
por: Yu, Zhongming, et al.
Publicado: (2024)
Leveraging Mathematical Reasoning of LLMs for Efficient GPU Thread Mapping
por: Maureira, Jose, et al.
Publicado: (2026)
por: Maureira, Jose, et al.
Publicado: (2026)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
por: Zhang, WenZheng, et al.
Publicado: (2024)
por: Zhang, WenZheng, et al.
Publicado: (2024)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
por: Li, Zhonggen, et al.
Publicado: (2025)
por: Li, Zhonggen, et al.
Publicado: (2025)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
por: Guo, Cong, et al.
Publicado: (2024)
por: Guo, Cong, et al.
Publicado: (2024)
GPU-Accelerated Batch-Dynamic Subgraph Matching
por: Qiu, Linshan, et al.
Publicado: (2024)
por: Qiu, Linshan, et al.
Publicado: (2024)
HarMoEny: Efficient Multi-GPU Inference of MoE Models
por: Doucet, Zachary, et al.
Publicado: (2025)
por: Doucet, Zachary, et al.
Publicado: (2025)
Heimdall++: Optimizing GPU Utilization and Pipeline Parallelism for Efficient Single-Pulse Detection
por: Xia, Bingzheng, et al.
Publicado: (2025)
por: Xia, Bingzheng, et al.
Publicado: (2025)
An AD based library for Efficient Hessian and Hessian-Vector Product Computation on GPU
por: Ranjan, Desh, et al.
Publicado: (2024)
por: Ranjan, Desh, et al.
Publicado: (2024)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
por: Liu, Yi, et al.
Publicado: (2025)
por: Liu, Yi, et al.
Publicado: (2025)
ZEUS: An Efficient GPU Optimization Method Integrating PSO, BFGS, and Automatic Differentiation
por: Soos, Dominik, et al.
Publicado: (2026)
por: Soos, Dominik, et al.
Publicado: (2026)
Optimizing Bloom Filters for Modern GPU Architectures
por: Jünger, Daniel, et al.
Publicado: (2025)
por: Jünger, Daniel, et al.
Publicado: (2025)
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
por: Park, Seongyeon, et al.
Publicado: (2025)
por: Park, Seongyeon, et al.
Publicado: (2025)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
por: Zhang, Mingjun, et al.
Publicado: (2025)
por: Zhang, Mingjun, et al.
Publicado: (2025)
Efficient GPU Implementation of Particle Interactions with Cutoff Radius and Few Particles per Cell
por: Algis, David, et al.
Publicado: (2024)
por: Algis, David, et al.
Publicado: (2024)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
por: Liu, Yunzhao, et al.
Publicado: (2025)
por: Liu, Yunzhao, et al.
Publicado: (2025)
MERBIT: A GPU-Based SpMV Method for Iterative Workloads
por: Zhang, Qi, et al.
Publicado: (2026)
por: Zhang, Qi, et al.
Publicado: (2026)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
por: Pilliat, Emmanuel
Publicado: (2026)
por: Pilliat, Emmanuel
Publicado: (2026)
GPZ: GPU-Accelerated Lossy Compressor for Particle Data
por: Li, Ruoyu, et al.
Publicado: (2025)
por: Li, Ruoyu, et al.
Publicado: (2025)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
por: Kumar, Abhishek Vijaya, et al.
Publicado: (2024)
por: Kumar, Abhishek Vijaya, et al.
Publicado: (2024)
SHARe-KAN: Post-Training Vector Quantization for Cache-Resident KAN Inference
por: Smith, Jeff
Publicado: (2025)
por: Smith, Jeff
Publicado: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
por: Wang, Tianyu, et al.
Publicado: (2024)
por: Wang, Tianyu, et al.
Publicado: (2024)
Ejemplares similares
-
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
por: Wei, Jinhui, et al.
Publicado: (2025) -
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
por: Zhang, Hongbin, et al.
Publicado: (2026) -
TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
por: Wu, Shixun, et al.
Publicado: (2025) -
The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
por: Amoros, Oscar, et al.
Publicado: (2025) -
Fused Breadth-First Probabilistic Traversals on Distributed GPU Systems
por: Neff, Reece, et al.
Publicado: (2023)