Dual-Select FMA Butterfly for FFT: Eliminating Twiddle Factor Singularities with Bounded Precomputed Ratios
Fuente:
arXiv
Saved in:
| Main Author: | Bergach, Mohamed Amine |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shortest-Path FFT: Optimal SIMD Instruction Scheduling via Graph Search
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
From 8 Seconds to 370ms: Kernel-Fused SAR Imaging on Apple Silicon via Single-Dispatch FFT Pipelines
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Training Transformers in Cosine Coefficient Space
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Redundant Array Computation Elimination
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Execution time budget assignment for mixed criticality systems
by: Khelassi, Mohamed Amine, et al.
Published: (2023)
by: Khelassi, Mohamed Amine, et al.
Published: (2023)
Can Increasing the Hit Ratio Hurt Cache Throughput? (Long Version)
by: Qiu, Ziyue, et al.
Published: (2024)
by: Qiu, Ziyue, et al.
Published: (2024)
V-Seek: Accelerating LLM Reasoning on Open-hardware Server-class RISC-V Platforms
by: Rodrigo, Javier J. Poveda, et al.
Published: (2025)
by: Rodrigo, Javier J. Poveda, et al.
Published: (2025)
Novel Lower Bounds on M/G/k Scheduling
by: Wang, Ziyuan, et al.
Published: (2025)
by: Wang, Ziyuan, et al.
Published: (2025)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
by: Zhang, Lingqi, et al.
Published: (2025)
by: Zhang, Lingqi, et al.
Published: (2025)
An Upper Bound on the M/M/k Queue With Deterministic Setup Times
by: Williams, Jalani, et al.
Published: (2025)
by: Williams, Jalani, et al.
Published: (2025)
On the Optimization of Singular Spectrum Analyses: A Pragmatic Approach
by: Lopes, Fernando, et al.
Published: (2024)
by: Lopes, Fernando, et al.
Published: (2024)
Universal Workers: A Vision for Eliminating Cold Starts in Serverless Computing
by: Akbari, Saman, et al.
Published: (2025)
by: Akbari, Saman, et al.
Published: (2025)
TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time (Extended Version)
by: Kan, Zeliang, et al.
Published: (2024)
by: Kan, Zeliang, et al.
Published: (2024)
Tail Bounds for Queues with Abandonment: Constant, Moderate, Large Deviations, and Efficient Concentration
by: Wang, Zedong, et al.
Published: (2026)
by: Wang, Zedong, et al.
Published: (2026)
Memshare: Memory Sharing for Multicore Computation in R with an Application to Feature Selection by Mutual Information using PDE
by: Thrun, Michael C., et al.
Published: (2025)
by: Thrun, Michael C., et al.
Published: (2025)
Fine-Grained Clustering-Based Power Identification for Multicores
by: Elshamy, Mohamed R., et al.
Published: (2024)
by: Elshamy, Mohamed R., et al.
Published: (2024)
A Dataset of Performance Measurements and Alerts from Mozilla (Data Artifact)
by: Besbes, Mohamed Bilel, et al.
Published: (2025)
by: Besbes, Mohamed Bilel, et al.
Published: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
by: Liu, Shifang, et al.
Published: (2025)
by: Liu, Shifang, et al.
Published: (2025)
GPU-Accelerated Parallel Selected Inversion for Structured Matrices Using sTiles
by: Fattah, Esmail Abdul, et al.
Published: (2025)
by: Fattah, Esmail Abdul, et al.
Published: (2025)
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation
by: Ke, Chih-Hua
Published: (2026)
by: Ke, Chih-Hua
Published: (2026)
Age of Information with Age-Dependent Server Selection
by: Akar, Nail, et al.
Published: (2025)
by: Akar, Nail, et al.
Published: (2025)
Network Calculus Bounds for Time-Sensitive Networks: A Revisit
by: Jiang, Yuming
Published: (2024)
by: Jiang, Yuming
Published: (2024)
Enhanced Scalability in Assessing Quantum Integer Factorization Performance
by: Lee, Junseo
Published: (2023)
by: Lee, Junseo
Published: (2023)
sTiles: An Accelerated Computational Framework for Sparse Factorizations of Structured Matrices
by: Fattah, Esmail Abdul, et al.
Published: (2025)
by: Fattah, Esmail Abdul, et al.
Published: (2025)
A Novel Hybrid Optical and STAR IRS System for NTN Communications
by: Shang, Shunyuan, et al.
Published: (2025)
by: Shang, Shunyuan, et al.
Published: (2025)
CPINN-ABPI: Physics-Informed Neural Networks for Accurate Power Estimation in MPSoCs
by: Elshamy, Mohamed R., et al.
Published: (2025)
by: Elshamy, Mohamed R., et al.
Published: (2025)
A relação entre a «performance» social e a «performance» económico-financeira
by: Daniel Taborda
Published: (2007)
by: Daniel Taborda
Published: (2007)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2025)
by: Chiang, Hung-Yueh, et al.
Published: (2025)
LOS MÁRGENES TOMAN LA ESCENA. EL USO DE LA PERFORMANCE EN LA LUCHA SUBALTERNA. UNA VISIÓN ANTROPOLÓGICA.
by: Iván Alvarado
Published: (2013)
by: Iván Alvarado
Published: (2013)
Da artificação do sagrado nos museus: entre o teatro e a sacralidade
by: Bruno Brulon
Published: (2013)
by: Bruno Brulon
Published: (2013)
The Configuration Wall: Characterization and Elimination of Accelerator Configuration Overhead
by: Van Delm, Josse, et al.
Published: (2025)
by: Van Delm, Josse, et al.
Published: (2025)
NR Cell Identity-based Handover Decision-making Algorithm for High-speed Scenario within Dual Connectivity
by: Zhu, Zhiyi, et al.
Published: (2025)
by: Zhu, Zhiyi, et al.
Published: (2025)
JSPIM: A Skew-Aware PIM Accelerator for High-Performance Databases Join and Select Operations
by: Tajdari, Sabiha, et al.
Published: (2025)
by: Tajdari, Sabiha, et al.
Published: (2025)
DDSA: Dual-Domain Strategic Attack for Spatial-Temporal Efficiency in Adversarial Robustness Testing
by: Hu, Jinwei, et al.
Published: (2026)
by: Hu, Jinwei, et al.
Published: (2026)
Beyond Technological Usability: Exploratory Factor Analysis of the Comprehensive Assessment of Usability Scale for Learning Technologies (CAUSLT)
by: Lu, Jie, et al.
Published: (2025)
by: Lu, Jie, et al.
Published: (2025)
FIRM SIZE & SUSTAINABLE PERFORMANCE: A LITERATURE REVIEW
by: Asish Kumar Panda
Published: (2025)
by: Asish Kumar Panda
Published: (2025)
O Entrelaçamento dos Estudos Modernos da Performance e as Correntes Atuais em Antropologia
by: Marvin Carlson
Published: (2011)
by: Marvin Carlson
Published: (2011)
Similar Items
-
Shortest-Path FFT: Optimal SIMD Instruction Scheduling via Graph Search
by: Bergach, Mohamed Amine
Published: (2026) -
From 8 Seconds to 370ms: Kernel-Fused SAR Imaging on Apple Silicon via Single-Dispatch FFT Pipelines
by: Bergach, Mohamed Amine
Published: (2026) -
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026) -
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
by: Bergach, Mohamed Amine
Published: (2026) -
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)