FPGA or GPU? Analyzing comparative research for application-specific guidance
Fuente:
arXiv
Salvato in:
| Autori principali: | Purkayastha, Arnab A, Tharwani, Jay, Aggarwal, Shobhit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
di: Tharwani, Jay, et al.
Pubblicazione: (2025)
di: Tharwani, Jay, et al.
Pubblicazione: (2025)
Agentic Operator Generation for ML ASICs
di: Hammond, Alec M., et al.
Pubblicazione: (2025)
di: Hammond, Alec M., et al.
Pubblicazione: (2025)
CPU-less parallel execution of lambda calculus in digital logic
di: Fitchett, Harry, et al.
Pubblicazione: (2026)
di: Fitchett, Harry, et al.
Pubblicazione: (2026)
Exploring Parallelism in FPGA-Based Accelerators for Machine Learning Applications
di: Centeno, Sed, et al.
Pubblicazione: (2025)
di: Centeno, Sed, et al.
Pubblicazione: (2025)
Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using $\mathbb{F}_2$
di: Zhou, Keren, et al.
Pubblicazione: (2025)
di: Zhou, Keren, et al.
Pubblicazione: (2025)
TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
di: Guo, Licheng, et al.
Pubblicazione: (2022)
di: Guo, Licheng, et al.
Pubblicazione: (2022)
Massimult: A Novel Parallel CPU Architecture Based on Combinator Reduction
di: Nicklisch-Franken, Jurgen, et al.
Pubblicazione: (2024)
di: Nicklisch-Franken, Jurgen, et al.
Pubblicazione: (2024)
Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods
di: Fatima, Amel, et al.
Pubblicazione: (2026)
di: Fatima, Amel, et al.
Pubblicazione: (2026)
Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving
di: Zhao, Juntao, et al.
Pubblicazione: (2025)
di: Zhao, Juntao, et al.
Pubblicazione: (2025)
HPU: High-Bandwidth Processing Unit for Scalable, Cost-effective LLM Inference via GPU Co-processing
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
FPGA Innovation Research in the Netherlands: Present Landscape and Future Outlook
di: Alachiotis, Nikolaos, et al.
Pubblicazione: (2025)
di: Alachiotis, Nikolaos, et al.
Pubblicazione: (2025)
EPOCH: Enabling Preemption Operation for Context Saving in Heterogeneous FPGA Systems
di: Malik, Arsalan Ali, et al.
Pubblicazione: (2025)
di: Malik, Arsalan Ali, et al.
Pubblicazione: (2025)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
di: Zhu, Yu, et al.
Pubblicazione: (2025)
di: Zhu, Yu, et al.
Pubblicazione: (2025)
Enabling Time-Aware Priority Traffic Management over Distributed FPGA Nodes
di: Scionti, Alberto, et al.
Pubblicazione: (2025)
di: Scionti, Alberto, et al.
Pubblicazione: (2025)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
di: Liu, Xingyu, et al.
Pubblicazione: (2025)
di: Liu, Xingyu, et al.
Pubblicazione: (2025)
RapidOMS: FPGA-based Open Modification Spectral Library Searching with HD Computing
di: Pinge, Sumukh, et al.
Pubblicazione: (2024)
di: Pinge, Sumukh, et al.
Pubblicazione: (2024)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2024)
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2024)
Context-aware Simopt-Power: Using structural data with simulation metadata to optimise FPGA designs
di: Wadhwa, Eashan, et al.
Pubblicazione: (2026)
di: Wadhwa, Eashan, et al.
Pubblicazione: (2026)
HetGPU: The pursuit of making binary compatibility towards GPUs
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
di: Chung, Euijun, et al.
Pubblicazione: (2026)
di: Chung, Euijun, et al.
Pubblicazione: (2026)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
di: Agrawal, Anirudha, et al.
Pubblicazione: (2024)
di: Agrawal, Anirudha, et al.
Pubblicazione: (2024)
MANOJAVAM: A Scalable, Unified FPGA Accelerator for Matrix Multiplication and Singular Value Decomposition in Principal Component Analysis
di: Ramasubramanian, Srivaths, et al.
Pubblicazione: (2026)
di: Ramasubramanian, Srivaths, et al.
Pubblicazione: (2026)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
di: Elwasif, Wael, et al.
Pubblicazione: (2022)
di: Elwasif, Wael, et al.
Pubblicazione: (2022)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
di: Torres, L. A., et al.
Pubblicazione: (2024)
di: Torres, L. A., et al.
Pubblicazione: (2024)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -- A Critical Review
di: Majeed, Ashiyana Abdul, et al.
Pubblicazione: (2025)
di: Majeed, Ashiyana Abdul, et al.
Pubblicazione: (2025)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
di: Li, Bingyao, et al.
Pubblicazione: (2024)
di: Li, Bingyao, et al.
Pubblicazione: (2024)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
di: McDaniel, Adam, et al.
Pubblicazione: (2026)
di: McDaniel, Adam, et al.
Pubblicazione: (2026)
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
di: Eleftherakis, Panagiotis-Eleftherios, et al.
Pubblicazione: (2026)
di: Eleftherakis, Panagiotis-Eleftherios, et al.
Pubblicazione: (2026)
Cost-Performance Evaluation of General Compute Instances: AWS, Azure, GCP, and OCI
di: Tharwani, Jay, et al.
Pubblicazione: (2024)
di: Tharwani, Jay, et al.
Pubblicazione: (2024)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
di: Jeong, Shinnung, et al.
Pubblicazione: (2025)
di: Jeong, Shinnung, et al.
Pubblicazione: (2025)
EmbBERT: Attention Under 2 MB Memory
di: Bravin, Riccardo, et al.
Pubblicazione: (2025)
di: Bravin, Riccardo, et al.
Pubblicazione: (2025)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
Analyzing a Two-Tier Disaggregated Memory Protection Scheme Based on Memory Replication
di: Volos, Haris, et al.
Pubblicazione: (2025)
di: Volos, Haris, et al.
Pubblicazione: (2025)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
di: Zhou, Zhuoshan, et al.
Pubblicazione: (2026)
di: Zhou, Zhuoshan, et al.
Pubblicazione: (2026)
Simopt -- Simulation pass for Speculative Optimisation of FPGA-CAD flow
di: Wadhwa, Eashan, et al.
Pubblicazione: (2024)
di: Wadhwa, Eashan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
di: Tharwani, Jay, et al.
Pubblicazione: (2025) -
Agentic Operator Generation for ML ASICs
di: Hammond, Alec M., et al.
Pubblicazione: (2025) -
CPU-less parallel execution of lambda calculus in digital logic
di: Fitchett, Harry, et al.
Pubblicazione: (2026) -
Exploring Parallelism in FPGA-Based Accelerators for Machine Learning Applications
di: Centeno, Sed, et al.
Pubblicazione: (2025) -
Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using $\mathbb{F}_2$
di: Zhou, Keren, et al.
Pubblicazione: (2025)