FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Bohan, Li, Shengmin, Shi, Xinyu, Yao, Enyi, Catthoor, Francky, Yang, Simei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
von: Wang, Tinglue, et al.
Veröffentlicht: (2025)
von: Wang, Tinglue, et al.
Veröffentlicht: (2025)
Efficient Architecture for RISC-V Vector Memory Access
von: Guan, Hongyi, et al.
Veröffentlicht: (2025)
von: Guan, Hongyi, et al.
Veröffentlicht: (2025)
Accelerating Data Chunking in Deduplication Systems using Vector Instructions
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
Achieving Dependability of AI Execution with Radiation Hardened Processors
von: Taquichiri, Carlos Rafael Tordoya, et al.
Veröffentlicht: (2025)
von: Taquichiri, Carlos Rafael Tordoya, et al.
Veröffentlicht: (2025)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
von: Garg, Raveesh, et al.
Veröffentlicht: (2025)
von: Garg, Raveesh, et al.
Veröffentlicht: (2025)
A New Family of Thread to Core Allocation Policies for an SMT ARM Processor
von: Navarro, Marta, et al.
Veröffentlicht: (2025)
von: Navarro, Marta, et al.
Veröffentlicht: (2025)
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2022)
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2022)
Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300A
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
von: Deng, Yunhao, et al.
Veröffentlicht: (2025)
von: Deng, Yunhao, et al.
Veröffentlicht: (2025)
Vector-Centric Machine Learning Systems: A Cross-Stack Approach
von: Jiang, Wenqi
Veröffentlicht: (2025)
von: Jiang, Wenqi
Veröffentlicht: (2025)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
von: Noh, Si Ung, et al.
Veröffentlicht: (2024)
von: Noh, Si Ung, et al.
Veröffentlicht: (2024)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
AES-SpMM: Balancing Accuracy and Speed by Adaptive Edge Sampling Strategy to Accelerate SpMM in GNNs
von: Song, Yingchen, et al.
Veröffentlicht: (2025)
von: Song, Yingchen, et al.
Veröffentlicht: (2025)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
von: Lin, Bin, et al.
Veröffentlicht: (2024)
von: Lin, Bin, et al.
Veröffentlicht: (2024)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
von: Li, Jiamin, et al.
Veröffentlicht: (2025)
von: Li, Jiamin, et al.
Veröffentlicht: (2025)
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
von: Li, Bingyao, et al.
Veröffentlicht: (2024)
von: Li, Bingyao, et al.
Veröffentlicht: (2024)
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
von: Zou, An, et al.
Veröffentlicht: (2025)
von: Zou, An, et al.
Veröffentlicht: (2025)
SpeedMalloc: Improving Multi-threaded Applications via a Lightweight Core for Memory Allocation
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
von: Feng, Yinxiao, et al.
Veröffentlicht: (2025)
von: Feng, Yinxiao, et al.
Veröffentlicht: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
von: Wang, Zeke, et al.
Veröffentlicht: (2025)
von: Wang, Zeke, et al.
Veröffentlicht: (2025)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
Flex-TPU: A Flexible TPU with Runtime Reconfigurable Dataflow Architecture
von: Elbtity, Mohammed, et al.
Veröffentlicht: (2024)
von: Elbtity, Mohammed, et al.
Veröffentlicht: (2024)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
von: Zheng, Xianzhe, et al.
Veröffentlicht: (2026)
von: Zheng, Xianzhe, et al.
Veröffentlicht: (2026)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
von: Chen, Qian, et al.
Veröffentlicht: (2024)
von: Chen, Qian, et al.
Veröffentlicht: (2024)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
von: Kreuzer, Anke, et al.
Veröffentlicht: (2019)
von: Kreuzer, Anke, et al.
Veröffentlicht: (2019)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
von: Liu, Lian, et al.
Veröffentlicht: (2026)
von: Liu, Lian, et al.
Veröffentlicht: (2026)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Efficient deadlock avoidance for 2D mesh NoCs that use OQ or VOQ routers
von: Papaphilippou, Philippos, et al.
Veröffentlicht: (2023)
von: Papaphilippou, Philippos, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
von: Wang, Tinglue, et al.
Veröffentlicht: (2025) -
Efficient Architecture for RISC-V Vector Memory Access
von: Guan, Hongyi, et al.
Veröffentlicht: (2025) -
Accelerating Data Chunking in Deduplication Systems using Vector Instructions
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025) -
SpArch: Efficient Architecture for Sparse Matrix Multiplication
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020) -
Achieving Dependability of AI Execution with Radiation Hardened Processors
von: Taquichiri, Carlos Rafael Tordoya, et al.
Veröffentlicht: (2025)