TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Shixun, Zhai, Yujia, Dai, Huangliang, Zhao, Hairui, Zhu, Yue, Hu, Haiyang, Chen, Zizhong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TurboFFT: A High-Performance Fast Fourier Transform with Fault Tolerance on GPU
by: Wu, Shixun, et al.
Published: (2024)
by: Wu, Shixun, et al.
Published: (2024)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
by: Wu, Shixun, et al.
Published: (2024)
by: Wu, Shixun, et al.
Published: (2024)
FT K-means: A High-Performance K-means on GPU with Fault Tolerance
by: Wu, Shixun, et al.
Published: (2024)
by: Wu, Shixun, et al.
Published: (2024)
FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention
by: Dai, Huangliang, et al.
Published: (2025)
by: Dai, Huangliang, et al.
Published: (2025)
DaggerFFT: A Distributed FFT Framework Using Task Scheduling in Julia
by: Anvari, Sana Taghipour, et al.
Published: (2026)
by: Anvari, Sana Taghipour, et al.
Published: (2026)
A HPX Communication Benchmark: Distributed FFT using Collectives
by: Strack, Alexander, et al.
Published: (2025)
by: Strack, Alexander, et al.
Published: (2025)
Experiences Porting Distributed Applications to Asynchronous Tasks: A Multidimensional FFT Case-study
by: Strack, Alexander, et al.
Published: (2024)
by: Strack, Alexander, et al.
Published: (2024)
DGRO: Diameter-Guided Ring Optimization for Integrated Research Infrastructure Membership
by: Wu, Shixun, et al.
Published: (2024)
by: Wu, Shixun, et al.
Published: (2024)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
by: McFarland, Thomas, et al.
Published: (2025)
by: McFarland, Thomas, et al.
Published: (2025)
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
by: Huang, Jiajun, et al.
Published: (2023)
by: Huang, Jiajun, et al.
Published: (2023)
Accelerating Sparse DNNs Based on Tiled GEMM
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
cuSZ-$i$: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
by: Liu, Jinyang, et al.
Published: (2023)
by: Liu, Jinyang, et al.
Published: (2023)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
by: Bellavita, Julian, et al.
Published: (2026)
by: Bellavita, Julian, et al.
Published: (2026)
Fused Breadth-First Probabilistic Traversals on Distributed GPU Systems
by: Neff, Reece, et al.
Published: (2023)
by: Neff, Reece, et al.
Published: (2023)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
by: Amoros, Oscar, et al.
Published: (2025)
by: Amoros, Oscar, et al.
Published: (2025)
PolyKAN: Efficient Fused GPU Operators for Polynomial Kolmogorov-Arnold Network Variants
by: Yu, Mingkun, et al.
Published: (2025)
by: Yu, Mingkun, et al.
Published: (2025)
Boosting Scientific Error-Bounded Lossy Compression through Optimized Synergistic Lossy-Lossless Orchestration
by: Wu, Shixun, et al.
Published: (2025)
by: Wu, Shixun, et al.
Published: (2025)
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
by: Swann, Ryan, et al.
Published: (2025)
by: Swann, Ryan, et al.
Published: (2025)
LP-GEMM: Integrating Layout Propagation into GEMM Operations
by: Carneiro, César Guedes, et al.
Published: (2026)
by: Carneiro, César Guedes, et al.
Published: (2026)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
by: Pilliat, Emmanuel
Published: (2026)
by: Pilliat, Emmanuel
Published: (2026)
An Optimized Error-controlled MPI Collective Framework Integrated with Lossy Compression
by: Huang, Jiajun, et al.
Published: (2023)
by: Huang, Jiajun, et al.
Published: (2023)
Cultivating Multidisciplinary AI Workforce Development on iTiger GPU Cluster: Practices and Challenges
by: Sharif, Mayira, et al.
Published: (2025)
by: Sharif, Mayira, et al.
Published: (2025)
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
by: Xue, Weicheng, et al.
Published: (2025)
by: Xue, Weicheng, et al.
Published: (2025)
Optimizing Allreduce Operations for Modern Heterogeneous Architectures with Multiple Processes per GPU
by: Adams, Michael, et al.
Published: (2025)
by: Adams, Michael, et al.
Published: (2025)
Stream-K++: Adaptive GPU GEMM Kernel Scheduling and Selection using Bloom Filters
by: Sadasivan, Harisankar, et al.
Published: (2024)
by: Sadasivan, Harisankar, et al.
Published: (2024)
ZCCL: Significantly Improving Collective Communication With Error-Bounded Lossy Compression
by: Huang, Jiajun, et al.
Published: (2025)
by: Huang, Jiajun, et al.
Published: (2025)
Matrix-Free 3D SIMP Topology Optimization with Fused Gather-GEMM-Scatter Kernels
by: Yang, Shaoliang, et al.
Published: (2026)
by: Yang, Shaoliang, et al.
Published: (2026)
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
by: Hu, Huanqi, et al.
Published: (2025)
by: Hu, Huanqi, et al.
Published: (2025)
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
by: Shan, Baodi, et al.
Published: (2024)
by: Shan, Baodi, et al.
Published: (2024)
Host-Side Telemetry for Performance Diagnosis in Cloud and HPC GPU Infrastructure
by: Darzi, Erfan, et al.
Published: (2025)
by: Darzi, Erfan, et al.
Published: (2025)
Dataflow-Oriented Classification and Performance Analysis of GPU-Accelerated Homomorphic Encryption
by: Nozaki, Ai, et al.
Published: (2026)
by: Nozaki, Ai, et al.
Published: (2026)
TX-Digital Twin: Visualizing Supercomputer GPU Performance Data Stream
by: Baskakova, Elena, et al.
Published: (2026)
by: Baskakova, Elena, et al.
Published: (2026)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
GPU Under Pressure: Estimating Application's Stress via Telemetry and Performance Counters
by: Esposito, Giuseppe, et al.
Published: (2025)
by: Esposito, Giuseppe, et al.
Published: (2025)
A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture
by: Yi, Xinyao
Published: (2024)
by: Yi, Xinyao
Published: (2024)
BandPilot: Towards Performance- and Contention-Aware GPU Dispatching in AI Clusters
by: Zhang, Kunming, et al.
Published: (2025)
by: Zhang, Kunming, et al.
Published: (2025)
Slide FFT on a homogeneous mesh in wafer-scale computing
by: van Putten, Maurice H. P. M., et al.
Published: (2024)
by: van Putten, Maurice H. P. M., et al.
Published: (2024)
High-Performance N-Queens Solver on GPU: Iterative DFS with Zero Bank Conflicts
by: Yao, Guangchao, et al.
Published: (2025)
by: Yao, Guangchao, et al.
Published: (2025)
Similar Items
-
TurboFFT: A High-Performance Fast Fourier Transform with Fault Tolerance on GPU
by: Wu, Shixun, et al.
Published: (2024) -
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
by: Wu, Shixun, et al.
Published: (2024) -
FT K-means: A High-Performance K-means on GPU with Fault Tolerance
by: Wu, Shixun, et al.
Published: (2024) -
FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention
by: Dai, Huangliang, et al.
Published: (2025) -
DaggerFFT: A Distributed FFT Framework Using Task Scheduling in Julia
by: Anvari, Sana Taghipour, et al.
Published: (2026)