FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Tinglue, Li, Yiming, Tang, Wei, Guan, Jiapeng, Guo, Zhenghui, Jiang, Renshuang, Wei, Ran, Li, Jing, Jiang, Zhe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
di: Li, Bohan, et al.
Pubblicazione: (2026)
di: Li, Bohan, et al.
Pubblicazione: (2026)
HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
di: Gao, Yiming, et al.
Pubblicazione: (2025)
di: Gao, Yiming, et al.
Pubblicazione: (2025)
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
di: Wang, Jing, et al.
Pubblicazione: (2025)
di: Wang, Jing, et al.
Pubblicazione: (2025)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
di: Oliveira, Geraldo F., et al.
Pubblicazione: (2025)
di: Oliveira, Geraldo F., et al.
Pubblicazione: (2025)
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
di: Zou, An, et al.
Pubblicazione: (2025)
di: Zou, An, et al.
Pubblicazione: (2025)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
di: Shen, Aofeng, et al.
Pubblicazione: (2025)
di: Shen, Aofeng, et al.
Pubblicazione: (2025)
Simultaneous Many-Row Activation in Off-the-Shelf DRAM Chips: Experimental Characterization and Analysis
di: Yuksel, Ismail Emir, et al.
Pubblicazione: (2024)
di: Yuksel, Ismail Emir, et al.
Pubblicazione: (2024)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
di: Liu, Xingyu, et al.
Pubblicazione: (2025)
di: Liu, Xingyu, et al.
Pubblicazione: (2025)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
di: Yuksel, Ismail Emir, et al.
Pubblicazione: (2023)
di: Yuksel, Ismail Emir, et al.
Pubblicazione: (2023)
RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
di: Feng, Yinxiao, et al.
Pubblicazione: (2025)
di: Feng, Yinxiao, et al.
Pubblicazione: (2025)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
di: Deng, Yunhao, et al.
Pubblicazione: (2025)
di: Deng, Yunhao, et al.
Pubblicazione: (2025)
Kitsune: Enabling Dataflow Execution on GPUs
di: Davies, Michael, et al.
Pubblicazione: (2025)
di: Davies, Michael, et al.
Pubblicazione: (2025)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
di: Noh, Si Ung, et al.
Pubblicazione: (2024)
di: Noh, Si Ung, et al.
Pubblicazione: (2024)
Enabling Mixed criticality applications for the Versal AI-Engines
di: Sprave, Vincent, et al.
Pubblicazione: (2026)
di: Sprave, Vincent, et al.
Pubblicazione: (2026)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
di: Kong, Fanchen, et al.
Pubblicazione: (2025)
di: Kong, Fanchen, et al.
Pubblicazione: (2025)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
di: Asquini, Lorenzo, et al.
Pubblicazione: (2025)
di: Asquini, Lorenzo, et al.
Pubblicazione: (2025)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
EPOCH: Enabling Preemption Operation for Context Saving in Heterogeneous FPGA Systems
di: Malik, Arsalan Ali, et al.
Pubblicazione: (2025)
di: Malik, Arsalan Ali, et al.
Pubblicazione: (2025)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
di: Prakriya, Neha, et al.
Pubblicazione: (2023)
di: Prakriya, Neha, et al.
Pubblicazione: (2023)
UniFormer: Unified and Efficient Transformer for Reasoning Across General and Custom Computing
di: Ran, Zhuoheng, et al.
Pubblicazione: (2025)
di: Ran, Zhuoheng, et al.
Pubblicazione: (2025)
Enabling Time-Aware Priority Traffic Management over Distributed FPGA Nodes
di: Scionti, Alberto, et al.
Pubblicazione: (2025)
di: Scionti, Alberto, et al.
Pubblicazione: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
di: Kubo, Tatsuya, et al.
Pubblicazione: (2025)
di: Kubo, Tatsuya, et al.
Pubblicazione: (2025)
Functionally-Complete Boolean Logic in Real DRAM Chips: Experimental Characterization and Analysis
di: Yuksel, Ismail Emir, et al.
Pubblicazione: (2024)
di: Yuksel, Ismail Emir, et al.
Pubblicazione: (2024)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
Cloud-Native Operation of Roadside Infrastructure Enabling Demand-Driven Collective Perception via V2X
di: Zanger, Lukas, et al.
Pubblicazione: (2026)
di: Zanger, Lukas, et al.
Pubblicazione: (2026)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
di: Mei, Linyan, et al.
Pubblicazione: (2022)
di: Mei, Linyan, et al.
Pubblicazione: (2022)
Efficient Architecture for RISC-V Vector Memory Access
di: Guan, Hongyi, et al.
Pubblicazione: (2025)
di: Guan, Hongyi, et al.
Pubblicazione: (2025)
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
di: Barkhordar, Marzieh, et al.
Pubblicazione: (2026)
di: Barkhordar, Marzieh, et al.
Pubblicazione: (2026)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State Drives
di: Nadig, Rakesh, et al.
Pubblicazione: (2026)
di: Nadig, Rakesh, et al.
Pubblicazione: (2026)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
di: Wang, Xi, et al.
Pubblicazione: (2024)
di: Wang, Xi, et al.
Pubblicazione: (2024)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
di: Zheng, Xianzhe, et al.
Pubblicazione: (2026)
di: Zheng, Xianzhe, et al.
Pubblicazione: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
di: Zhou, Zhuoshan, et al.
Pubblicazione: (2026)
di: Zhou, Zhuoshan, et al.
Pubblicazione: (2026)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
di: Zhang, Qijun, et al.
Pubblicazione: (2026)
di: Zhang, Qijun, et al.
Pubblicazione: (2026)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
di: Li, Jiamin, et al.
Pubblicazione: (2025)
di: Li, Jiamin, et al.
Pubblicazione: (2025)
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
di: Li, Bingyao, et al.
Pubblicazione: (2024)
di: Li, Bingyao, et al.
Pubblicazione: (2024)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
di: Qu, Huanyu, et al.
Pubblicazione: (2025)
di: Qu, Huanyu, et al.
Pubblicazione: (2025)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
di: Chen, Yanru, et al.
Pubblicazione: (2025)
di: Chen, Yanru, et al.
Pubblicazione: (2025)
SpeedMalloc: Improving Multi-threaded Applications via a Lightweight Core for Memory Allocation
di: Li, Ruihao, et al.
Pubblicazione: (2025)
di: Li, Ruihao, et al.
Pubblicazione: (2025)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
di: Liu, Fangxin, et al.
Pubblicazione: (2026)
di: Liu, Fangxin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
di: Li, Bohan, et al.
Pubblicazione: (2026) -
HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
di: Gao, Yiming, et al.
Pubblicazione: (2025) -
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
di: Wang, Jing, et al.
Pubblicazione: (2025) -
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
di: Oliveira, Geraldo F., et al.
Pubblicazione: (2025) -
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
di: Zou, An, et al.
Pubblicazione: (2025)