ACiS: Complex Processing in the Switch Fabric
Fuente:
arXiv
Saved in:
| Main Authors: | Haghi, Pouya, Guo, Anqi, Geng, Tong, Skjellum, Anthony, Herbordt, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
by: Tong, Jianming, et al.
Published: (2024)
by: Tong, Jianming, et al.
Published: (2024)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
by: Dumoulin, Joren, et al.
Published: (2025)
by: Dumoulin, Joren, et al.
Published: (2025)
Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
by: Mamdouh, Ahmed, et al.
Published: (2024)
by: Mamdouh, Ahmed, et al.
Published: (2024)
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
by: Symons, Arne, et al.
Published: (2022)
by: Symons, Arne, et al.
Published: (2022)
Fletch: File-System Metadata Caching in Programmable Switches
by: Liu, Qingxiu, et al.
Published: (2025)
by: Liu, Qingxiu, et al.
Published: (2025)
SPAC: Automating FPGA-based Network Switches with Protocol Adaptive Customization
by: Li, Guoyu, et al.
Published: (2026)
by: Li, Guoyu, et al.
Published: (2026)
Different Perspectives of Memory System Simulation
by: Esmaili-Dokht, Pouya, et al.
Published: (2026)
by: Esmaili-Dokht, Pouya, et al.
Published: (2026)
Ultra Low-Power SDM-based Circuit-Switching for Networks-on-Chip
by: Zaeemi, Meysam, et al.
Published: (2026)
by: Zaeemi, Meysam, et al.
Published: (2026)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
by: Jiang, Aojie, et al.
Published: (2026)
by: Jiang, Aojie, et al.
Published: (2026)
In-place Switch: Reprogramming based SLC Cache Design for Hybrid 3D SSDs
by: Yang, Xufeng, et al.
Published: (2024)
by: Yang, Xufeng, et al.
Published: (2024)
Pack my weights and run! Minimizing overheads for in-memory computing accelerators
by: Houshmand, Pouya, et al.
Published: (2024)
by: Houshmand, Pouya, et al.
Published: (2024)
Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale Integration
by: Feng, Yinxiao, et al.
Published: (2024)
by: Feng, Yinxiao, et al.
Published: (2024)
Performance and Energy Benefits of MRDIMMs
by: Díaz, Pau, et al.
Published: (2026)
by: Díaz, Pau, et al.
Published: (2026)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
Splatonic: Architecture Support for 3D Gaussian Splatting SLAM via Sparse Processing
by: Huang, Xiaotong, et al.
Published: (2025)
by: Huang, Xiaotong, et al.
Published: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
by: Zhong, Linfeng, et al.
Published: (2025)
by: Zhong, Linfeng, et al.
Published: (2025)
FERIVer: An FPGA-assisted Emulated Framework for RTL Verification of RISC-V Processors
by: Qin, Kun, et al.
Published: (2025)
by: Qin, Kun, et al.
Published: (2025)
PIMSYN: Synthesizing Processing-in-memory CNN Accelerators
by: Li, Wanqian, et al.
Published: (2024)
by: Li, Wanqian, et al.
Published: (2024)
Flexible In-NAND Cryptographic Processing for Secure Flash Storage
by: Noh, Seock-Hwan, et al.
Published: (2025)
by: Noh, Seock-Hwan, et al.
Published: (2025)
Containerized In-Storage Processing and Computing-Enabled SSD Disaggregation
by: Kwon, Miryeong, et al.
Published: (2025)
by: Kwon, Miryeong, et al.
Published: (2025)
RTGPU: Real-Time Computing with Graphics Processing Units
by: Gheibi-Fetrat, Atiyeh, et al.
Published: (2025)
by: Gheibi-Fetrat, Atiyeh, et al.
Published: (2025)
FPPS: An FPGA-Based Point Cloud Processing System
by: Zhou, Xiaofeng, et al.
Published: (2026)
by: Zhou, Xiaofeng, et al.
Published: (2026)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
by: Shin, Yongwon, et al.
Published: (2024)
by: Shin, Yongwon, et al.
Published: (2024)
Linear Complexity Fermionic Simulation on Quantum Devices with Hardware Connectivity Constraints
by: Gao, Xiangyu, et al.
Published: (2026)
by: Gao, Xiangyu, et al.
Published: (2026)
DL-PIM: Improving Data Locality in Processing-in-Memory Systems
by: Tian, Parker Hao, et al.
Published: (2025)
by: Tian, Parker Hao, et al.
Published: (2025)
ReGate: Enabling Power Gating in Neural Processing Units
by: Xue, Yuqi, et al.
Published: (2025)
by: Xue, Yuqi, et al.
Published: (2025)
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
by: Wang, Zhao, et al.
Published: (2025)
by: Wang, Zhao, et al.
Published: (2025)
Allspark: Workload Orchestration for Visual Transformers on Processing In-Memory Systems
by: Ge, Mengke, et al.
Published: (2024)
by: Ge, Mengke, et al.
Published: (2024)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
by: Sun, Xiaotian, et al.
Published: (2024)
by: Sun, Xiaotian, et al.
Published: (2024)
Open-source Stand-Alone Versatile Tensor Accelerator
by: Faure-Gignoux, Anthony, et al.
Published: (2025)
by: Faure-Gignoux, Anthony, et al.
Published: (2025)
Compilation and Execution of an Embeddable YOLO-NAS on the VTA
by: Faure-Gignoux, Anthony, et al.
Published: (2026)
by: Faure-Gignoux, Anthony, et al.
Published: (2026)
VerilogCL: A Contrastive Learning Framework for Robust LLM-Based Verilog Generation
by: Tan, Yan, et al.
Published: (2026)
by: Tan, Yan, et al.
Published: (2026)
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
DRIFT: Harnessing Inherent Fault Tolerance for Efficient and Reliable Diffusion Model Inference
by: Wen, Jinqi, et al.
Published: (2026)
by: Wen, Jinqi, et al.
Published: (2026)
Optimizing and Exploring System Performance in Compact Processing-in-Memory-based Chips
by: Chen, Peilin, et al.
Published: (2025)
by: Chen, Peilin, et al.
Published: (2025)
Unlimited Vector Processing for Wireless Baseband Based on RISC-V Extension
by: Jiang, Limin, et al.
Published: (2025)
by: Jiang, Limin, et al.
Published: (2025)
AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
by: Cheng, Feng, et al.
Published: (2025)
by: Cheng, Feng, et al.
Published: (2025)
Instruction-Based Coordination of Heterogeneous Processing Units for Acceleration of DNN Inference
by: Petropoulos, Anastasios, et al.
Published: (2025)
by: Petropoulos, Anastasios, et al.
Published: (2025)
Efficient Sparse Processing-in-Memory Architecture (ESPIM) for Machine Learning Inference
by: He, Mingxuan, et al.
Published: (2024)
by: He, Mingxuan, et al.
Published: (2024)
Similar Items
-
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
by: Tong, Jianming, et al.
Published: (2024) -
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
by: Dumoulin, Joren, et al.
Published: (2025) -
Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
by: Mamdouh, Ahmed, et al.
Published: (2024) -
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
by: Symons, Arne, et al.
Published: (2022) -
Fletch: File-System Metadata Caching in Programmable Switches
by: Liu, Qingxiu, et al.
Published: (2025)