TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
Fuente:
arXiv
Guardado en:
| Autores principales: | Guo, Licheng, Chi, Yuze, Lau, Jason, Song, Linghao, Tian, Xingyu, Khatti, Moazin, Qiao, Weikang, Wang, Jie, Ustun, Ecenur, Fang, Zhenman, Zhang, Zhiru, Cong, Jason |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
por: Prakriya, Neha, et al.
Publicado: (2023)
por: Prakriya, Neha, et al.
Publicado: (2023)
Stream-HLS: Towards Automatic Dataflow Acceleration
por: Basalama, Suhail, et al.
Publicado: (2025)
por: Basalama, Suhail, et al.
Publicado: (2025)
RapidStream IR: Infrastructure for FPGA High-Level Physical Synthesis
por: Lau, Jason, et al.
Publicado: (2024)
por: Lau, Jason, et al.
Publicado: (2024)
A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable Processors
por: Kuper, Reese, et al.
Publicado: (2023)
por: Kuper, Reese, et al.
Publicado: (2023)
ACALSim: A Scalable Parallel Simulation Framework for High-Performance System Design Space Exploration
por: Lin, Wei-Fen, et al.
Publicado: (2026)
por: Lin, Wei-Fen, et al.
Publicado: (2026)
LightningSimV2: Faster and Scalable Simulation for High-Level Synthesis via Graph Compilation and Optimization
por: Sarkar, Rishov, et al.
Publicado: (2024)
por: Sarkar, Rishov, et al.
Publicado: (2024)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
por: Zhang, Niansong, et al.
Publicado: (2025)
por: Zhang, Niansong, et al.
Publicado: (2025)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
por: Yang, Hanchen, et al.
Publicado: (2025)
por: Yang, Hanchen, et al.
Publicado: (2025)
CiFlow: Dataflow Analysis and Optimization of Key Switching for Homomorphic Encryption
por: Neda, Negar, et al.
Publicado: (2023)
por: Neda, Negar, et al.
Publicado: (2023)
Fast Algorithms for Spiking Neural Network Simulation with FPGAs
por: Lindqvist, Björn A., et al.
Publicado: (2024)
por: Lindqvist, Björn A., et al.
Publicado: (2024)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
por: Mao, Shunyu, et al.
Publicado: (2024)
por: Mao, Shunyu, et al.
Publicado: (2024)
DEER: Deep Runahead for Instruction Prefetching on Modern Mobile Workloads
por: Vahdatniya, Parmida, et al.
Publicado: (2025)
por: Vahdatniya, Parmida, et al.
Publicado: (2025)
OmniSim: Simulating Hardware with C Speed and RTL Accuracy for High-Level Synthesis Designs
por: Sarkar, Rishov, et al.
Publicado: (2025)
por: Sarkar, Rishov, et al.
Publicado: (2025)
Selective Parallel Loading of Large-Scale Compressed Graphs with ParaGrapher
por: Esfahani, Mohsen Koohi, et al.
Publicado: (2024)
por: Esfahani, Mohsen Koohi, et al.
Publicado: (2024)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
por: Chen, Qian, et al.
Publicado: (2024)
por: Chen, Qian, et al.
Publicado: (2024)
Iceberg: Enhancing HLS Modeling with Synthetic Data
por: Ding, Zijian, et al.
Publicado: (2025)
por: Ding, Zijian, et al.
Publicado: (2025)
GigaAPI for GPU Parallelization
por: Suvarna, M., et al.
Publicado: (2025)
por: Suvarna, M., et al.
Publicado: (2025)
Parallelizing a modern GPU simulator
por: Huerta, Rodrigo, et al.
Publicado: (2025)
por: Huerta, Rodrigo, et al.
Publicado: (2025)
Heterogeneous Memory Benchmarking Toolkit
por: Ghaemi, Golsana, et al.
Publicado: (2025)
por: Ghaemi, Golsana, et al.
Publicado: (2025)
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
por: Ali, Wajid, et al.
Publicado: (2025)
por: Ali, Wajid, et al.
Publicado: (2025)
Enhancing Instruction Prefetching via Cache and TLB Management
por: Jamet, Alexandre Valentin, et al.
Publicado: (2026)
por: Jamet, Alexandre Valentin, et al.
Publicado: (2026)
ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
por: Zuepke, Alexander, et al.
Publicado: (2026)
por: Zuepke, Alexander, et al.
Publicado: (2026)
Towards CPU Performance Prediction: New Challenge Benchmark Dataset and Novel Approach
por: Liu, Xiaoman
Publicado: (2024)
por: Liu, Xiaoman
Publicado: (2024)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
por: Liu, Songze, et al.
Publicado: (2025)
por: Liu, Songze, et al.
Publicado: (2025)
Recurrent CircuitSAT Sampling for Sequential Circuits
por: Ardakani, Arash, et al.
Publicado: (2025)
por: Ardakani, Arash, et al.
Publicado: (2025)
Introducing the Arm-membench Throughput Benchmark
por: Burth, Cyrill, et al.
Publicado: (2025)
por: Burth, Cyrill, et al.
Publicado: (2025)
Enhancing software-hardware co-design for HEP by low-overhead profiling of single- and multi-threaded programs on diverse architectures with Adaptyst
por: Graczyk, Maksymilian, et al.
Publicado: (2025)
por: Graczyk, Maksymilian, et al.
Publicado: (2025)
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison
por: Li, Ruihao, et al.
Publicado: (2026)
por: Li, Ruihao, et al.
Publicado: (2026)
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation
por: Ke, Chih-Hua
Publicado: (2026)
por: Ke, Chih-Hua
Publicado: (2026)
Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures
por: Anik, Shafayat Mowla, et al.
Publicado: (2026)
por: Anik, Shafayat Mowla, et al.
Publicado: (2026)
Makinote: An FPGA-Based HW/SW Platform for Pre-Silicon Emulation of RISC-V Designs
por: Perdomo, Elias, et al.
Publicado: (2024)
por: Perdomo, Elias, et al.
Publicado: (2024)
AI Load Dynamics--A Power Electronics Perspective
por: Li, Yuzhuo, et al.
Publicado: (2025)
por: Li, Yuzhuo, et al.
Publicado: (2025)
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
por: Wadle, Shayne, et al.
Publicado: (2025)
por: Wadle, Shayne, et al.
Publicado: (2025)
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
por: Ham, Hyungkyu, et al.
Publicado: (2024)
por: Ham, Hyungkyu, et al.
Publicado: (2024)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
por: Bergach, Mohamed Amine
Publicado: (2026)
por: Bergach, Mohamed Amine
Publicado: (2026)
OPTIMA: Design-Space Exploration of Discharge-Based In-SRAM Computing: Quantifying Energy-Accuracy Trade-Offs
por: Seyedfaraji, Saeed, et al.
Publicado: (2024)
por: Seyedfaraji, Saeed, et al.
Publicado: (2024)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
por: Mazzola, Sergio, et al.
Publicado: (2025)
por: Mazzola, Sergio, et al.
Publicado: (2025)
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
por: Qararyah, Fareed, et al.
Publicado: (2025)
por: Qararyah, Fareed, et al.
Publicado: (2025)
Accelerating Transistor-Level Simulation of Integrated Circuits via Equivalence of RC Long-Chain Structures
por: Tang, Ruibai, et al.
Publicado: (2025)
por: Tang, Ruibai, et al.
Publicado: (2025)
The Bicameral Cache: a split cache for vector architectures
por: Rebolledo, Susana, et al.
Publicado: (2024)
por: Rebolledo, Susana, et al.
Publicado: (2024)
Ejemplares similares
-
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
por: Prakriya, Neha, et al.
Publicado: (2023) -
Stream-HLS: Towards Automatic Dataflow Acceleration
por: Basalama, Suhail, et al.
Publicado: (2025) -
RapidStream IR: Infrastructure for FPGA High-Level Physical Synthesis
por: Lau, Jason, et al.
Publicado: (2024) -
A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable Processors
por: Kuper, Reese, et al.
Publicado: (2023) -
ACALSim: A Scalable Parallel Simulation Framework for High-Performance System Design Space Exploration
por: Lin, Wei-Fen, et al.
Publicado: (2026)