A 1024 RV-Cores Shared-L1 Cluster with High Bandwidth Memory Link for Low-Latency 6G-SDR
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yichao, Bertuletti, Marco, Zhang, Chi, Riedel, Samuel, Vanelli-Coralli, Alessandro, Benini, Luca |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
di: Zhang, Yichao, et al.
Pubblicazione: (2024)
di: Zhang, Yichao, et al.
Pubblicazione: (2024)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
di: Zhang, Yichao, et al.
Pubblicazione: (2026)
di: Zhang, Yichao, et al.
Pubblicazione: (2026)
A 410GFLOP/s, 64 RISC-V Cores, 204.8GBps Shared-Memory Cluster in 12nm FinFET with Systolic Execution Support for Efficient B5G/6G AI-Enhanced O-RAN
di: Zhang, Yichao, et al.
Pubblicazione: (2025)
di: Zhang, Yichao, et al.
Pubblicazione: (2025)
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
di: Shen, Diyou, et al.
Pubblicazione: (2025)
di: Shen, Diyou, et al.
Pubblicazione: (2025)
TeraNoC: A Multi-Channel 32-bit Fine-Grained, Hybrid Mesh-Crossbar NoC for Efficient Scale-up of 1000+ Core Shared-L1-Memory Clusters
di: Zhang, Yichao, et al.
Pubblicazione: (2025)
di: Zhang, Yichao, et al.
Pubblicazione: (2025)
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
di: Mazzola, Sergio, et al.
Pubblicazione: (2025)
di: Mazzola, Sergio, et al.
Pubblicazione: (2025)
Low Latency, High Bandwidth Streaming of Experimental Data with EJFAT
di: Baldin, Ilya, et al.
Pubblicazione: (2025)
di: Baldin, Ilya, et al.
Pubblicazione: (2025)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
di: Li, Yinrong, et al.
Pubblicazione: (2026)
di: Li, Yinrong, et al.
Pubblicazione: (2026)
Shared Memory-Aware Latency-Sensitive Message Aggregation for Fine-Grained Communication
di: Chandrasekar, Kavitha, et al.
Pubblicazione: (2024)
di: Chandrasekar, Kavitha, et al.
Pubblicazione: (2024)
Co-designing a Programmable RISC-V Accelerator for MPC-based Energy and Thermal Management of Many-Core HPC Processors
di: Ottaviano, Alessandro, et al.
Pubblicazione: (2025)
di: Ottaviano, Alessandro, et al.
Pubblicazione: (2025)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
di: Lu, Yao, et al.
Pubblicazione: (2026)
di: Lu, Yao, et al.
Pubblicazione: (2026)
PATSMA: Parameter Auto-tuning for Shared Memory Algorithms
di: Fernandes, Joao B., et al.
Pubblicazione: (2024)
di: Fernandes, Joao B., et al.
Pubblicazione: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
di: Ma, Chenxiang, et al.
Pubblicazione: (2025)
di: Ma, Chenxiang, et al.
Pubblicazione: (2025)
Computation-Bandwidth-Memory Trade-offs: A Unified Paradigm for AI Infrastructure
di: Fan, Yuankai, et al.
Pubblicazione: (2025)
di: Fan, Yuankai, et al.
Pubblicazione: (2025)
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
di: Seo, Eunil, et al.
Pubblicazione: (2025)
di: Seo, Eunil, et al.
Pubblicazione: (2025)
Distributing Context-Aware Shared Memory Data Structures: A Case Study on Singly-Linked Lists
di: Ravishankar, Raaghav, et al.
Pubblicazione: (2024)
di: Ravishankar, Raaghav, et al.
Pubblicazione: (2024)
Optimizing Offload Performance in Heterogeneous MPSoCs
di: Colagrande, Luca, et al.
Pubblicazione: (2024)
di: Colagrande, Luca, et al.
Pubblicazione: (2024)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
Dynamic Client Clustering, Bandwidth Allocation, and Workload Optimization for Semi-synchronous Federated Learning
di: Yu, Liangkun, et al.
Pubblicazione: (2024)
di: Yu, Liangkun, et al.
Pubblicazione: (2024)
On-Package Memory with Universal Chiplet Interconnect Express (UCIe): A Low Power, High Bandwidth, Low Latency and Low Cost Approach
di: Sharma, Debendra Das, et al.
Pubblicazione: (2025)
di: Sharma, Debendra Das, et al.
Pubblicazione: (2025)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
di: Zhao, Yihao, et al.
Pubblicazione: (2025)
di: Zhao, Yihao, et al.
Pubblicazione: (2025)
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
di: Wu, Yongtong, et al.
Pubblicazione: (2026)
di: Wu, Yongtong, et al.
Pubblicazione: (2026)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
di: Shen, Aofeng, et al.
Pubblicazione: (2025)
di: Shen, Aofeng, et al.
Pubblicazione: (2025)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
di: Zhang, Han, et al.
Pubblicazione: (2026)
di: Zhang, Han, et al.
Pubblicazione: (2026)
A 66-Gb/s/5.5-W RISC-V Many-Core Cluster for 5G+ Software-Defined Radio Uplinks
di: Bertuletti, Marco, et al.
Pubblicazione: (2025)
di: Bertuletti, Marco, et al.
Pubblicazione: (2025)
Byzantine-Tolerant Consensus in GPU-Inspired Shared Memory
di: Georgiou, Chryssis, et al.
Pubblicazione: (2025)
di: Georgiou, Chryssis, et al.
Pubblicazione: (2025)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
di: Murat, Antoine, et al.
Pubblicazione: (2024)
di: Murat, Antoine, et al.
Pubblicazione: (2024)
Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
di: Luo, Yizhou, et al.
Pubblicazione: (2024)
di: Luo, Yizhou, et al.
Pubblicazione: (2024)
DRust: Language-Guided Distributed Shared Memory with Fine Granularity, Full Transparency, and Ultra Efficiency
di: Ma, Haoran, et al.
Pubblicazione: (2024)
di: Ma, Haoran, et al.
Pubblicazione: (2024)
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
di: Xu, Yi, et al.
Pubblicazione: (2024)
di: Xu, Yi, et al.
Pubblicazione: (2024)
Enhancing Traffic Safety with AI and 6G: Latency Requirements and Real-Time Threat Detection
di: Horvath, Kurt, et al.
Pubblicazione: (2025)
di: Horvath, Kurt, et al.
Pubblicazione: (2025)
Shared Virtual Memory: Its Design and Performance Implications for Diverse Applications
di: Cooper, Bennett, et al.
Pubblicazione: (2024)
di: Cooper, Bennett, et al.
Pubblicazione: (2024)
ClusterFusion++: Expanding Cluster-Level Fusion to Full Transformer-Block Decoding
di: Jin, ChiHeng, et al.
Pubblicazione: (2026)
di: Jin, ChiHeng, et al.
Pubblicazione: (2026)
Optimizing Memory Allocation in Distributed Clusters with Predictive Modeling
di: Bader, Jonathan, et al.
Pubblicazione: (2026)
di: Bader, Jonathan, et al.
Pubblicazione: (2026)
Solutions for Distributed Memory Access Mechanism on HPC Clusters
di: Meizner, Jan, et al.
Pubblicazione: (2025)
di: Meizner, Jan, et al.
Pubblicazione: (2025)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
di: Zhang, Lingqi, et al.
Pubblicazione: (2025)
di: Zhang, Lingqi, et al.
Pubblicazione: (2025)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
di: Wu, Tian, et al.
Pubblicazione: (2025)
di: Wu, Tian, et al.
Pubblicazione: (2025)
Bandwidth-Aware Network Topology Optimization for Decentralized Learning
di: Shen, Yipeng, et al.
Pubblicazione: (2025)
di: Shen, Yipeng, et al.
Pubblicazione: (2025)
The Carnot Bound: Limits and Possibilities for Bandwidth-Efficient Consensus
di: Lewis-Pye, Andrew, et al.
Pubblicazione: (2026)
di: Lewis-Pye, Andrew, et al.
Pubblicazione: (2026)
An Asynchronous Many-Task Algorithm for Unstructured $S_{N}$ Transport on Shared Memory Systems
di: Elwood, Alex, et al.
Pubblicazione: (2025)
di: Elwood, Alex, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
di: Zhang, Yichao, et al.
Pubblicazione: (2024) -
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
di: Zhang, Yichao, et al.
Pubblicazione: (2026) -
A 410GFLOP/s, 64 RISC-V Cores, 204.8GBps Shared-Memory Cluster in 12nm FinFET with Systolic Execution Support for Efficient B5G/6G AI-Enhanced O-RAN
di: Zhang, Yichao, et al.
Pubblicazione: (2025) -
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
di: Shen, Diyou, et al.
Pubblicazione: (2025) -
TeraNoC: A Multi-Channel 32-bit Fine-Grained, Hybrid Mesh-Crossbar NoC for Efficient Scale-up of 1000+ Core Shared-L1-Memory Clusters
di: Zhang, Yichao, et al.
Pubblicazione: (2025)