FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zeke, Zhang, Jie, Huang, Hongjing, Li, Yingtao, Zhu, Xueying, Sun, Mo, Yang, Zihan, Ma, De, Tang, Huajing, Pan, Gang, Wu, Fei, He, Bingsheng, Alonso, Gustavo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
von: Liu, Lian, et al.
Veröffentlicht: (2026)
von: Liu, Lian, et al.
Veröffentlicht: (2026)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
von: Sun, Mo, et al.
Veröffentlicht: (2024)
von: Sun, Mo, et al.
Veröffentlicht: (2024)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
von: Mei, Linyan, et al.
Veröffentlicht: (2022)
von: Mei, Linyan, et al.
Veröffentlicht: (2022)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
von: Qin, Ruoyu, et al.
Veröffentlicht: (2024)
von: Qin, Ruoyu, et al.
Veröffentlicht: (2024)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
von: li, Fei, et al.
Veröffentlicht: (2026)
von: li, Fei, et al.
Veröffentlicht: (2026)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
PIMDAL: Mitigating the Memory Bottleneck in Data Analytics using a Real Processing-in-Memory System
von: Frouzakis, Manos, et al.
Veröffentlicht: (2025)
von: Frouzakis, Manos, et al.
Veröffentlicht: (2025)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
von: Liu, Fangxin, et al.
Veröffentlicht: (2026)
von: Liu, Fangxin, et al.
Veröffentlicht: (2026)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
von: Zhu, Yu, et al.
Veröffentlicht: (2025)
von: Zhu, Yu, et al.
Veröffentlicht: (2025)
The Feasibility of Implementing Large-Scale Transformers on Multi-FPGA Platforms
von: Gao, Yu, et al.
Veröffentlicht: (2024)
von: Gao, Yu, et al.
Veröffentlicht: (2024)
Security Risks Due to Data Persistence in Cloud FPGA Platforms
von: Zhang, Zhehang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhehang, et al.
Veröffentlicht: (2024)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
von: Kreuzer, Anke, et al.
Veröffentlicht: (2019)
von: Kreuzer, Anke, et al.
Veröffentlicht: (2019)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
Efficient deadlock avoidance for 2D mesh NoCs that use OQ or VOQ routers
von: Papaphilippou, Philippos, et al.
Veröffentlicht: (2023)
von: Papaphilippou, Philippos, et al.
Veröffentlicht: (2023)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity Multicores
von: García-García, Adrián, et al.
Veröffentlicht: (2024)
von: García-García, Adrián, et al.
Veröffentlicht: (2024)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
von: Li, Bohan, et al.
Veröffentlicht: (2026)
von: Li, Bohan, et al.
Veröffentlicht: (2026)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
von: Font, Martí Llopart, et al.
Veröffentlicht: (2026)
von: Font, Martí Llopart, et al.
Veröffentlicht: (2026)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
von: Muntaka, Siddique Abubakr, et al.
Veröffentlicht: (2026)
von: Muntaka, Siddique Abubakr, et al.
Veröffentlicht: (2026)
NetSmith: An Optimization Framework for Machine-Discovered Network Topologies
von: Green, Conor, et al.
Veröffentlicht: (2024)
von: Green, Conor, et al.
Veröffentlicht: (2024)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
Navigating the Landscape of Distributed File Systems: Architectures, Implementations, and Considerations
von: Pan, Xueting, et al.
Veröffentlicht: (2024)
von: Pan, Xueting, et al.
Veröffentlicht: (2024)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
von: Ma, Ke, et al.
Veröffentlicht: (2025)
von: Ma, Ke, et al.
Veröffentlicht: (2025)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
von: Li, Jiamin, et al.
Veröffentlicht: (2025)
von: Li, Jiamin, et al.
Veröffentlicht: (2025)
Evaluating Rapid Makespan Predictions for Heterogeneous Systems with Programmable Logic
von: Wilhelm, Martin, et al.
Veröffentlicht: (2025)
von: Wilhelm, Martin, et al.
Veröffentlicht: (2025)
Context-aware Simopt-Power: Using structural data with simulation metadata to optimise FPGA designs
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2026)
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2026)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2023)
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
von: Liu, Lian, et al.
Veröffentlicht: (2026) -
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
von: Sun, Mo, et al.
Veröffentlicht: (2024) -
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026) -
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026) -
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)