Saved in:
| Main Authors: | Wang, Zeke, Zhang, Jie, Huang, Hongjing, Li, Yingtao, Zhu, Xueying, Sun, Mo, Yang, Zihan, Ma, De, Tang, Huajing, Pan, Gang, Wu, Fei, He, Bingsheng, Alonso, Gustavo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.09318 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
by: Liu, Lian, et al.
Published: (2026)
by: Liu, Lian, et al.
Published: (2026)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
by: Sun, Mo, et al.
Published: (2024)
by: Sun, Mo, et al.
Published: (2024)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
by: Kim, Hyeseong, et al.
Published: (2026)
by: Kim, Hyeseong, et al.
Published: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
by: Zhou, Zhuoshan, et al.
Published: (2026)
by: Zhou, Zhuoshan, et al.
Published: (2026)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
by: Jarmusch, Aaron, et al.
Published: (2026)
by: Jarmusch, Aaron, et al.
Published: (2026)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
by: Qin, Ruoyu, et al.
Published: (2024)
by: Qin, Ruoyu, et al.
Published: (2024)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
by: Mei, Linyan, et al.
Published: (2022)
by: Mei, Linyan, et al.
Published: (2022)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
by: Zhu, Yu, et al.
Published: (2025)
by: Zhu, Yu, et al.
Published: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
by: li, Fei, et al.
Published: (2026)
by: li, Fei, et al.
Published: (2026)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
by: Mo, Zhiwen, et al.
Published: (2026)
by: Mo, Zhiwen, et al.
Published: (2026)
PIMDAL: Mitigating the Memory Bottleneck in Data Analytics using a Real Processing-in-Memory System
by: Frouzakis, Manos, et al.
Published: (2025)
by: Frouzakis, Manos, et al.
Published: (2025)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
by: Liu, Fangxin, et al.
Published: (2026)
by: Liu, Fangxin, et al.
Published: (2026)
The Feasibility of Implementing Large-Scale Transformers on Multi-FPGA Platforms
by: Gao, Yu, et al.
Published: (2024)
by: Gao, Yu, et al.
Published: (2024)
Security Risks Due to Data Persistence in Cloud FPGA Platforms
by: Zhang, Zhehang, et al.
Published: (2024)
by: Zhang, Zhehang, et al.
Published: (2024)
RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
by: Feng, Yinxiao, et al.
Published: (2025)
by: Feng, Yinxiao, et al.
Published: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
by: Kwak, Hyunseok, et al.
Published: (2025)
by: Kwak, Hyunseok, et al.
Published: (2025)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
by: Kreuzer, Anke, et al.
Published: (2019)
by: Kreuzer, Anke, et al.
Published: (2019)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
by: Qiu, Tong Dong, et al.
Published: (2023)
by: Qiu, Tong Dong, et al.
Published: (2023)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
by: Xu, Weihong, et al.
Published: (2025)
by: Xu, Weihong, et al.
Published: (2025)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
by: Negi, Shubham, et al.
Published: (2025)
by: Negi, Shubham, et al.
Published: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
by: Kubo, Tatsuya, et al.
Published: (2025)
by: Kubo, Tatsuya, et al.
Published: (2025)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
by: Chen, Yanru, et al.
Published: (2025)
by: Chen, Yanru, et al.
Published: (2025)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
by: Kurzynski, Marco, et al.
Published: (2025)
by: Kurzynski, Marco, et al.
Published: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Efficient deadlock avoidance for 2D mesh NoCs that use OQ or VOQ routers
by: Papaphilippou, Philippos, et al.
Published: (2023)
by: Papaphilippou, Philippos, et al.
Published: (2023)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
by: Orenes-Vera, Marcelo, et al.
Published: (2023)
by: Orenes-Vera, Marcelo, et al.
Published: (2023)
LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity Multicores
by: García-García, Adrián, et al.
Published: (2024)
by: García-García, Adrián, et al.
Published: (2024)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
by: Li, Bohan, et al.
Published: (2026)
by: Li, Bohan, et al.
Published: (2026)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
by: Font, Martí Llopart, et al.
Published: (2026)
by: Font, Martí Llopart, et al.
Published: (2026)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
by: Muntaka, Siddique Abubakr, et al.
Published: (2026)
by: Muntaka, Siddique Abubakr, et al.
Published: (2026)
NetSmith: An Optimization Framework for Machine-Discovered Network Topologies
by: Green, Conor, et al.
Published: (2024)
by: Green, Conor, et al.
Published: (2024)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
by: Zhang, Zhekai, et al.
Published: (2020)
by: Zhang, Zhekai, et al.
Published: (2020)
Navigating the Landscape of Distributed File Systems: Architectures, Implementations, and Considerations
by: Pan, Xueting, et al.
Published: (2024)
by: Pan, Xueting, et al.
Published: (2024)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
by: Ma, Ke, et al.
Published: (2025)
by: Ma, Ke, et al.
Published: (2025)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
by: Yu, Yanpeng, et al.
Published: (2025)
by: Yu, Yanpeng, et al.
Published: (2025)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
by: Shen, Aofeng, et al.
Published: (2025)
by: Shen, Aofeng, et al.
Published: (2025)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
by: Li, Jiamin, et al.
Published: (2025)
by: Li, Jiamin, et al.
Published: (2025)
Evaluating Rapid Makespan Predictions for Heterogeneous Systems with Programmable Logic
by: Wilhelm, Martin, et al.
Published: (2025)
by: Wilhelm, Martin, et al.
Published: (2025)
Context-aware Simopt-Power: Using structural data with simulation metadata to optimise FPGA designs
by: Wadhwa, Eashan, et al.
Published: (2026)
by: Wadhwa, Eashan, et al.
Published: (2026)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
by: Asquini, Lorenzo, et al.
Published: (2025)
by: Asquini, Lorenzo, et al.
Published: (2025)
Similar Items
-
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
by: Liu, Lian, et al.
Published: (2026) -
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
by: Sun, Mo, et al.
Published: (2024) -
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
by: Kim, Hyeseong, et al.
Published: (2026) -
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
by: Zhou, Zhuoshan, et al.
Published: (2026) -
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
by: Jarmusch, Aaron, et al.
Published: (2026)