TeraNoC: A Multi-Channel 32-bit Fine-Grained, Hybrid Mesh-Crossbar NoC for Efficient Scale-up of 1000+ Core Shared-L1-Memory Clusters
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Yichao, Fu, Zexin, Fischer, Tim, Li, Yinrong, Bertuletti, Marco, Benini, Luca |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
par: Shen, Diyou, et autres
Publié: (2025)
par: Shen, Diyou, et autres
Publié: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
par: Zhang, Yichao, et autres
Publié: (2026)
par: Zhang, Yichao, et autres
Publié: (2026)
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
par: Zhang, Yichao, et autres
Publié: (2024)
par: Zhang, Yichao, et autres
Publié: (2024)
A 1024 RV-Cores Shared-L1 Cluster with High Bandwidth Memory Link for Low-Latency 6G-SDR
par: Zhang, Yichao, et autres
Publié: (2024)
par: Zhang, Yichao, et autres
Publié: (2024)
A 410GFLOP/s, 64 RISC-V Cores, 204.8GBps Shared-Memory Cluster in 12nm FinFET with Systolic Execution Support for Efficient B5G/6G AI-Enhanced O-RAN
par: Zhang, Yichao, et autres
Publié: (2025)
par: Zhang, Yichao, et autres
Publié: (2025)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
par: Li, Yinrong, et autres
Publié: (2026)
par: Li, Yinrong, et autres
Publié: (2026)
A Dynamic Allocation Scheme for Adaptive Shared-Memory Mapping on Kilo-core RV Clusters for Attention-Based Model Deployment
par: Wang, Bowen, et autres
Publié: (2025)
par: Wang, Bowen, et autres
Publié: (2025)
A 66-Gb/s/5.5-W RISC-V Many-Core Cluster for 5G+ Software-Defined Radio Uplinks
par: Bertuletti, Marco, et autres
Publié: (2025)
par: Bertuletti, Marco, et autres
Publié: (2025)
Optimizing Scalable Multi-Cluster Architectures for Next-Generation Wireless Sensing and Communication
par: Riedel, Samuel, et autres
Publié: (2025)
par: Riedel, Samuel, et autres
Publié: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
par: Colagrande, Luca, et autres
Publié: (2026)
par: Colagrande, Luca, et autres
Publié: (2026)
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
par: Mazzola, Sergio, et autres
Publié: (2025)
par: Mazzola, Sergio, et autres
Publié: (2025)
FlooNoC: A 645 Gbps/link 0.15 pJ/B/hop Open-Source NoC with Wide Physical Links and End-to-End AXI4 Parallel Multi-Stream Support
par: Fischer, Tim, et autres
Publié: (2024)
par: Fischer, Tim, et autres
Publié: (2024)
A Compute&Memory Efficient Model-Driven Neural 5G Receiver for Edge AI-assisted RAN
par: Abdollahpour, Mahdi, et autres
Publié: (2025)
par: Abdollahpour, Mahdi, et autres
Publié: (2025)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
par: Colagrande, Luca, et autres
Publié: (2025)
par: Colagrande, Luca, et autres
Publié: (2025)
TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks
par: Bertuletti, Marco, et autres
Publié: (2026)
par: Bertuletti, Marco, et autres
Publié: (2026)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
par: Mazzola, Sergio, et autres
Publié: (2024)
par: Mazzola, Sergio, et autres
Publié: (2024)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
par: Wang, Yimin, et autres
Publié: (2025)
par: Wang, Yimin, et autres
Publié: (2025)
Demystifying FPGA Hard NoC Performance
par: Liu, Sihao, et autres
Publié: (2025)
par: Liu, Sihao, et autres
Publié: (2025)
Occamy: A 432-Core Dual-Chiplet Dual-HBM2E 768-DP-GFLOP/s RISC-V System for 8-to-64-bit Dense and Sparse Computing in 12nm FinFET
par: Scheffler, Paul, et autres
Publié: (2025)
par: Scheffler, Paul, et autres
Publié: (2025)
Learning Cache Coherence Traffic for NoC Routing Design
par: Xiong, Guochu, et autres
Publié: (2025)
par: Xiong, Guochu, et autres
Publié: (2025)
Fast End-to-End Simulation and Exploration of Many-RISCV-Core Baseband Transceivers for Software-Defined Radio-Access Networks
par: Bertuletti, Marco, et autres
Publié: (2025)
par: Bertuletti, Marco, et autres
Publié: (2025)
fence.t.s: Closing Timing Channels in High-Performance Out-of-Order Cores through ISA-Supported Temporal Partitioning
par: Wistoff, Nils, et autres
Publié: (2024)
par: Wistoff, Nils, et autres
Publié: (2024)
Occamy: A 432-Core 28.1 DP-GFLOP/s/W 83% FPU Utilization Dual-Chiplet, Dual-HBM2E RISC-V-based Accelerator for Stencil and Sparse Linear Algebra Computations with 8-to-64-bit Floating-Point Support in 12nm FinFET
par: Paulin, Gianna, et autres
Publié: (2024)
par: Paulin, Gianna, et autres
Publié: (2024)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
par: Colagrande, Luca, et autres
Publié: (2025)
par: Colagrande, Luca, et autres
Publié: (2025)
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
par: Colagrande, Luca, et autres
Publié: (2026)
par: Colagrande, Luca, et autres
Publié: (2026)
CVA6-VMRT: A Modular Approach Towards Time-Predictable Virtual Memory in a 64-bit Application Class RISC-V Processor
par: Reinwardt, Christopher, et autres
Publié: (2025)
par: Reinwardt, Christopher, et autres
Publié: (2025)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
par: Colagrande, Luca, et autres
Publié: (2025)
par: Colagrande, Luca, et autres
Publié: (2025)
Application of Machine Learning Techniques for Secure Traffic in NoC-based Manycores
par: Lopes, Geaninne, et autres
Publié: (2025)
par: Lopes, Geaninne, et autres
Publié: (2025)
Travel Time Based Task Mapping for NoC-Based DNN Accelerator
par: Chen, Yizhi, et autres
Publié: (2024)
par: Chen, Yizhi, et autres
Publié: (2024)
Systematic Prevention of On-Core Timing Channels by Full Temporal Partitioning
par: Wistoff, Nils, et autres
Publié: (2022)
par: Wistoff, Nils, et autres
Publié: (2022)
A Theoretical Lens for RL-Tuned Language Models via Energy-Based Models
par: Tan, Zhiquan, et autres
Publié: (2025)
par: Tan, Zhiquan, et autres
Publié: (2025)
PAINT: Partial-Solution Adaptive Interpolated Training for Self-Distilled Reasoners
par: Tan, Zhiquan, et autres
Publié: (2026)
par: Tan, Zhiquan, et autres
Publié: (2026)
Self-Supervised On-Policy Distillation for Reasoning Language Models
par: Tan, Zhiquan, et autres
Publié: (2026)
par: Tan, Zhiquan, et autres
Publié: (2026)
2C‐Ternary Content Addressable Memory in Memcapacitor Crossbar Array with NAND Flash Structure
par: Hwiho Hwang, et autres
Publié: (2025)
par: Hwiho Hwang, et autres
Publié: (2025)
MiniFloat-NN and ExSdotp: An ISA Extension and a Modular Open Hardware Unit for Low-Precision Training on RISC-V cores
par: Bertaccini, Luca, et autres
Publié: (2022)
par: Bertaccini, Luca, et autres
Publié: (2022)
Bit Transition Reduction by Data Transmission Ordering in NoC-based DNN Accelerator
par: Chen, Yizhi, et autres
Publié: (2025)
par: Chen, Yizhi, et autres
Publié: (2025)
Ramping Up Open-Source RISC-V Cores: Assessing the Energy Efficiency of Superscalar, Out-of-Order Execution
par: Fu, Zexin, et autres
Publié: (2025)
par: Fu, Zexin, et autres
Publié: (2025)
AraXL: A Physically Scalable, Ultra-Wide RISC-V Vector Processor Design for Fast and Efficient Computation on Long Vectors
par: Purayil, Navaneeth Kunhi, et autres
Publié: (2025)
par: Purayil, Navaneeth Kunhi, et autres
Publié: (2025)
Al0.68Sc0.32N/SiC based metal-ferroelectric-semiconductor capacitors operating up to 1000 °C
par: He, Yunfei, et autres
Publié: (2024)
par: He, Yunfei, et autres
Publié: (2024)
Basilisk: A 34 mm2 End-to-End Open-Source 64-bit Linux-Capable RISC-V SoC in 130nm BiCMOS
par: Sauter, Philippe, et autres
Publié: (2025)
par: Sauter, Philippe, et autres
Publié: (2025)
Documents similaires
-
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
par: Shen, Diyou, et autres
Publié: (2025) -
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
par: Zhang, Yichao, et autres
Publié: (2026) -
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
par: Zhang, Yichao, et autres
Publié: (2024) -
A 1024 RV-Cores Shared-L1 Cluster with High Bandwidth Memory Link for Low-Latency 6G-SDR
par: Zhang, Yichao, et autres
Publié: (2024) -
A 410GFLOP/s, 64 RISC-V Cores, 204.8GBps Shared-Memory Cluster in 12nm FinFET with Systolic Execution Support for Efficient B5G/6G AI-Enhanced O-RAN
par: Zhang, Yichao, et autres
Publié: (2025)